0% found this document useful (0 votes)
3 views34 pages

08 Class Basic

Chapter 8 of 'Data Mining: Concepts and Techniques' discusses classification, including supervised and unsupervised learning, model construction, and decision tree induction. It covers various classification methods such as Bayes classification and rule-based classification, along with techniques for improving accuracy like ensemble methods. The chapter also explores model evaluation, attribute selection measures, and the importance of understanding entropy in classification processes.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views34 pages

08 Class Basic

Chapter 8 of 'Data Mining: Concepts and Techniques' discusses classification, including supervised and unsupervised learning, model construction, and decision tree induction. It covers various classification methods such as Bayes classification and rule-based classification, along with techniques for improving accuracy like ensemble methods. The chapter also explores model evaluation, attribute selection measures, and the importance of understanding entropy in classification processes.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Data Mining:

Concepts and Techniques


(3rd ed.)

— Chapter 8 —

Jiawei Han, Micheline Kamber, and Jian Pei


University of Illinois at Urbana-Champaign &
Simon Fraser University
©2011 Han, Kamber & Pei. All rights reserved.
1
Chapter 8. Classification: Basic Concepts

 Classification: Basic Concepts


 Decision Tree Induction
 Bayes Classification Methods
 Rule-Based Classification
 Model Evaluation and Selection
 Techniques to Improve Classification Accuracy:
Ensemble Methods
 Summary
3
Supervised vs. Unsupervised Learning

 Supervised learning (classification)


 Supervision: The training data (observations,
measurements, etc.) are accompanied by labels indicating
the class of the observations
 New data is classified based on the training set
 Unsupervised learning (clustering)
 The class labels of training data is unknown
 Given a set of measurements, observations, etc. with the
aim of establishing the existence of classes or clusters in
the data
4
Prediction Problems: Classification vs.
Numeric Prediction
 Classification
 predicts categorical class labels (discrete or nominal)

 classifies data (constructs a model) based on the training


set and the values (class labels) in a classifying attribute
and uses it in classifying new data
 Numeric Prediction
 models continuous-valued functions, i.e., predicts
unknown or missing values
 Typical applications
 Credit/loan approval:

 Medical diagnosis: if a tumor is cancerous or benign

 Fraud detection: if a transaction is fraudulent

 Web page categorization: which category it is

5
Classification—A Two-Step Process
1. Model construction: describing a set of
predetermined classes
 Each tuple/sample is assumed to belong to a

predefined class, as determined by the class


label attribute
 The set of tuples used for model

construction is training set


 The model is represented as classification

rules, decision trees, or mathematical


formulae
6
Process (1): Model Construction

Classification
Algorithms
Training
Data

NAME RANK YEARS TENURED Classifier


M ike A ssistant P rof 3 no (Model)
M ary A ssistant P rof 7 yes
B ill P rofessor 2 yes
Jim A ssociate P rof 7 yes
IF rank = ‘professor’
D ave A ssistant P rof 6 no
OR years > 6
A nne A ssociate P rof 3 no
THEN tenured = ‘yes’
7
Classification—A Two-Step Process
2. Model usage: for classifying future or unknown objects
 Estimate accuracy of the model

 The known label of test sample is compared with the

classified result from the model


 Accuracy rate is the percentage of test set samples

that are correctly classified by the model..


 correctly classified/number of test samples

 Test set is independent of training set (otherwise

overfitting)
 If the accuracy is acceptable, use the model to classify

new data
 Note: If the test set is used to select models, it is called validation (test) set
8
Process (2): Using the Model in Prediction
IF rank = ‘professor’
OR years > 6
THEN tenured = ‘yes’
Classifier

Testing
Data Unseen Data

(Jeff, Professor, 4)
NAME RANK YEARS TENURED
T om A ssistant P rof 2 no Tenured?
M erlisa A ssociate P rof 7 no
G eorge P rofessor 5 yes
Joseph A ssistant P rof 7 yes
9
Chapter 8. Classification: Basic Concepts

 Classification: Basic Concepts


 Decision Tree Induction
 Bayes Classification Methods
 Rule-Based Classification
 Model Evaluation and Selection
 Techniques to Improve Classification Accuracy:
Ensemble Methods
 Summary
10
Decision Tree Induction: An Example
age income student credit_rating buys_computer
<=30 high no fair no
 Training data set: Buys_computer <=30 high no excellent no
 The data set follows an example of 31…40 high no fair yes
>40 medium no fair yes
Quinlan’s ID3 >40 low yes fair yes
>40 low yes excellent no
 Resulting tree:
31…40 low yes excellent yes
age? <=30 medium no fair no
<=30 low yes fair yes
>40 medium yes fair yes
<=30 medium yes excellent yes
<=30 overcast
31..40 >40 31…40 medium no excellent yes
31…40 high yes fair yes
>40 medium no excellent no

student? yes credit rating?

no yes excellent fair

no yes yes
11
Algorithm for Decision Tree Induction
 Basic algorithm (a greedy algorithm)
 Tree is constructed in a top-down recursive divide-and-conquer manner

 At start, all the training examples are at the root

 Attributes are categorical (if continuous-valued, they are discretized in

advance)
 Examples are partitioned recursively based on selected attributes

 Test attributes are selected on the basis of a heuristic or statistical measure

(e.g., information gain)

12
Algorithm for Decision Tree Induction
 Basic algorithm (a greedy algorithm)
 Conditions for stopping partitioning
 All samples for a given node belong to the same class

 There are no remaining attributes for further partitioning –

majority voting is employed for classifying the leaf


 There are no samples left

13
Brief Review of Entropy

m=2

14
Brief Review of Entropy

 The curve approaches zero at the


extremes. For example: there is no
uncertainty when it is 100% yes or 100%
no. Entropy = 0
 Midpoint:The point where the curve is
highest is approximately:50% yes - 50%
no
 this is the
[Link]:uncertainty is at its
maximum.

m=2

15
Attribute Selection Measures

 The three measures, in general, return good results but


 Information gain:
 Gain ratio:
 Gini index:

16
Attribute Selection Measure:
Information Gain (ID3/C4.5)
 Select the attribute with the highest information gain
 Let pi be the probability that an arbitrary tuple in D belongs to
class Ci, estimated by |Ci, D|/|D|
 Expected information (entropy) needed to classify a tuple in D:
m
Info( D)   pi log 2 ( pi )
i 1
 Information needed (after using A to split D into v partitions) to
classify D: v | D |
InfoA ( D)    Info( D j )
j

j 1 | D |
 Information gained by branching on attribute A

Gain(A)  Info(D)  InfoA(D)


17
Attribute Selection: ID.3 Information Gain
 Class P: buys_computer = “yes”
 Class N: buys_computer = “no”
9 9 5 5
Info( D)  I (9,5)   log 2 ( )  log 2 ( ) 0.940
14 14 14 14

age income student credit_rating buys_computer


<=30 high no fair no
age pi ni I(pi, ni)
<=30 high no excellent no <=30 2 3 0.971
31…40 high no fair yes 31…40 4 0 0
>40 medium no fair yes
>40 low yes fair yes >40 3 2 0.971
>40 low yes excellent no
31…40 low yes excellent yes
<=30 medium no fair no
<=30 low yes fair yes
>40 medium yes fair yes
<=30 medium yes excellent yes
31…40 medium no excellent yes
31…40 high yes fair yes
>40 medium no excellent no
18
Attribute Selection: Information Gain
age pi ni I(pi, ni)
<=30 2 3 0.971 5 4
Infoage ( D)  I (2,3)  I (4,0)
31…40 4 0 0 14 14
>40 3 2 0.971 5
 I (3,2)  0.694
5 14
I (2,3) means “age <=30” has 5 out of 14 samples, with 2 yes’es and 3 no’s. Hence
14

•I(2,3)=0.971
•I(4,0)=0
•I(3,2)=0.971

Gain(age)  Info( D)  Infoage ( D)  0.246


Gain(income)  0.029
Similarly,
Gain( student)  0.151
Gain(credit _ rating)  0.048 19
Gain Ratio for Attribute Selection (C4.5)
 Information gain measure is biased towards attributes with a large number of
values
 C4.5 (a successor of ID3) uses gain ratio to overcome the problem
(normalization to information gain) Info(D)=0.940

=0.286+0.393+0.232 ≈0.911

age income student credit_rating buys_computer


<=30 high no fair no
<=30 high no excellent no
31…40 high no fair yes
Gainincome=Info(D)−Infoincome >40 medium no fair yes
Gainincome=0.940−0.911≈0.029 >40
>40
low
low
yes fair
yes excellent
yes
no
31…40 low yes excellent yes
<=30 medium no fair no
<=30 low yes fair yes
>40 medium yes fair yes
<=30 medium yes excellent yes
31…40 medium no excellent yes
31…40 high yes fair yes
>40 medium no excellent no 21
Gain Ratio for Attribute Selection (C4.5)
 Information gain measure is biased towards attributes with a large number
of values
 C4.5 (a successor of ID3) uses gain ratio to overcome the problem
(normalization to information gain) v |D | | Dj |
SplitInfoA ( D)    log 2 (
j
)
j 1 |D| |D|
 GainRatio(D) = Gain(D)/SplitInfo(D)
 Ex.
age income student credit_rating buys_computer
 gain_ratio(income) = 0.029/1.557 = 0.019 <=30 high no fair no
<=30 high no excellent no
 The attribute with the maximum gain ratio is selected31…40
as the high
splitting attribute
no fair yes
>40 medium no fair yes
>40 low yes fair yes
>40 low yes excellent no
31…40 low yes excellent yes
<=30 medium no fair no
<=30 low yes fair yes
>40 medium yes fair yes
<=30 medium yes excellent yes
31…40 medium no excellent yes
31…40 high yes fair yes
>40 medium no excellent no 22
23
Gini Index (CART)
 If a data set D contains examples from n classes, gini index,
gini(D) is defined as n 2
gini( D)  1  p j
j 1
where pj is the relative frequency of class j in D
 If a data set D is split on A into two subsets D1 and D2, the gini
index gini(D) is defined as
|D1| |D |
gini A ( D)  gini( D1)  2 gini( D 2)
|D| |D|
 Reduction in Impurity:
gini( A)  gini(D)  giniA(D)
 The attribute provides the smallest ginisplit(D) (need to
enumerate all the possible splitting points for each attribute)

24
Computation of Gini Index
 Ex. D has 9 tuples in buys_computer = “yes” and 5 in “no”2 2
   
9 5
gini( D)  1        0.459
 14   14 
 Suppose the attribute income partitions D into 10 in D1: {low, medium} and 4
7 yes 3 no 2 yes 2 no
in D2
 10  4
giniincome{low,medium} ( D)   Gini( D1 )   Gini( D2 )
 14   14 

0.42 0.5

Gini{low,high} is 0.458; Gini{medium,high} is


0.450. Thus, split on the {low,medium}
(and {high}) since it has the lowest Gini
index

25
26
Comparing Attribute Selection Measures

 The three measures, in general, return good results but


 Information gain:
 biased towards multivalued attributes
 Gain ratio:
 tends to prefer unbalanced splits in which one partition is
much smaller than the others
 Gini index:
 biased to multivalued attributes
 has difficulty when # of classes is large
 tends to favor tests that result in equal-sized partitions
and purity in both partitions
27
Overfitting and Tree Pruning
 Overfitting: An induced tree may overfit the training data
 Too many branches, some may reflect anomalies due to

noise or outliers
 Poor accuracy for unseen samples

 Two approaches to avoid overfitting


 Prepruning: Halt tree construction early ̵ do not split a node

if this would result in the goodness measure falling below a


threshold
 Difficult to choose an appropriate threshold

 Postpruning: Remove branches from a “fully grown” tree—

get a sequence of progressively pruned trees


 Use a set of data different from the training data to

decide which is the “best pruned tree”


29
30
Scalability Framework for RainForest

 Separates the scalability aspects from the criteria that


determine the quality of the tree
 Builds an AVC-list: AVC (Attribute, Value, Class_label)
 AVC-set (of an attribute X )
 Projection of training dataset onto the attribute X and
class label where counts of individual class label are
aggregated
 AVC-group (of a node n )
 Set of AVC-sets of all predictor attributes at the node n

31
Rainforest: Training Set and Its AVC Sets

Training Examples AVC-set on Age AVC-set on income


age income studentcredit_rating
buys_computerAge Buy_Computer income Buy_Computer

<=30 high no fair no yes no


<=30 high no excellent no yes no
high 2 2
31…40 high no fair yes <=30 2 3
31..40 4 0 medium 4 2
>40 medium no fair yes
>40 low yes fair yes >40 3 2 low 3 1
>40 low yes excellent no
31…40 low yes excellent yes
AVC-set on
<=30 medium no fair no AVC-set on Student
credit_rating
<=30 low yes fair yes
student Buy_Computer
>40 medium yes fair yes Credit
Buy_Computer

<=30 medium yes excellent yes yes no rating yes no


31…40 medium no excellent yes yes 6 1 fair 6 2
31…40 high yes fair yes no 3 4 excellent 3 3
>40 medium no excellent no
32
BOAT (Bootstrapped Optimistic
Algorithm for Tree Construction)
 Use a statistical technique called bootstrapping to create
several smaller samples (subsets), each fits in memory
 Each subset is used to create a tree, resulting in several
trees
 These trees are examined and used to construct a new
tree T’
 It turns out that T’ is very close to the tree that would
be generated using the whole data set together
 Adv: requires only two scans of DB, an incremental alg.

33
Presentation of Classification Results

May 5, 2026 Data Mining: Concepts and Techniques 34


Visualization of a Decision Tree in SGI/MineSet 3.0

May 5, 2026 Data Mining: Concepts and Techniques 35

You might also like