Data Mining:
Concepts and Techniques
(3rd ed.)
— Chapter 8 —
Jiawei Han, Micheline Kamber, and Jian Pei
University of Illinois at Urbana-Champaign &
Simon Fraser University
©2011 Han, Kamber & Pei. All rights reserved.
1
Chapter 8. Classification: Basic Concepts
Classification: Basic Concepts
Decision Tree Induction
Bayes Classification Methods
Rule-Based Classification
Model Evaluation and Selection
Techniques to Improve Classification Accuracy:
Ensemble Methods
Summary
3
Supervised vs. Unsupervised Learning
Supervised learning (classification)
Supervision: The training data (observations,
measurements, etc.) are accompanied by labels indicating
the class of the observations
New data is classified based on the training set
Unsupervised learning (clustering)
The class labels of training data is unknown
Given a set of measurements, observations, etc. with the
aim of establishing the existence of classes or clusters in
the data
4
Prediction Problems: Classification vs.
Numeric Prediction
Classification
predicts categorical class labels (discrete or nominal)
classifies data (constructs a model) based on the training
set and the values (class labels) in a classifying attribute
and uses it in classifying new data
Numeric Prediction
models continuous-valued functions, i.e., predicts
unknown or missing values
Typical applications
Credit/loan approval:
Medical diagnosis: if a tumor is cancerous or benign
Fraud detection: if a transaction is fraudulent
Web page categorization: which category it is
5
Classification—A Two-Step Process
1. Model construction: describing a set of
predetermined classes
Each tuple/sample is assumed to belong to a
predefined class, as determined by the class
label attribute
The set of tuples used for model
construction is training set
The model is represented as classification
rules, decision trees, or mathematical
formulae
6
Process (1): Model Construction
Classification
Algorithms
Training
Data
NAME RANK YEARS TENURED Classifier
M ike A ssistant P rof 3 no (Model)
M ary A ssistant P rof 7 yes
B ill P rofessor 2 yes
Jim A ssociate P rof 7 yes
IF rank = ‘professor’
D ave A ssistant P rof 6 no
OR years > 6
A nne A ssociate P rof 3 no
THEN tenured = ‘yes’
7
Classification—A Two-Step Process
2. Model usage: for classifying future or unknown objects
Estimate accuracy of the model
The known label of test sample is compared with the
classified result from the model
Accuracy rate is the percentage of test set samples
that are correctly classified by the model..
correctly classified/number of test samples
Test set is independent of training set (otherwise
overfitting)
If the accuracy is acceptable, use the model to classify
new data
Note: If the test set is used to select models, it is called validation (test) set
8
Process (2): Using the Model in Prediction
IF rank = ‘professor’
OR years > 6
THEN tenured = ‘yes’
Classifier
Testing
Data Unseen Data
(Jeff, Professor, 4)
NAME RANK YEARS TENURED
T om A ssistant P rof 2 no Tenured?
M erlisa A ssociate P rof 7 no
G eorge P rofessor 5 yes
Joseph A ssistant P rof 7 yes
9
Chapter 8. Classification: Basic Concepts
Classification: Basic Concepts
Decision Tree Induction
Bayes Classification Methods
Rule-Based Classification
Model Evaluation and Selection
Techniques to Improve Classification Accuracy:
Ensemble Methods
Summary
10
Decision Tree Induction: An Example
age income student credit_rating buys_computer
<=30 high no fair no
Training data set: Buys_computer <=30 high no excellent no
The data set follows an example of 31…40 high no fair yes
>40 medium no fair yes
Quinlan’s ID3 >40 low yes fair yes
>40 low yes excellent no
Resulting tree:
31…40 low yes excellent yes
age? <=30 medium no fair no
<=30 low yes fair yes
>40 medium yes fair yes
<=30 medium yes excellent yes
<=30 overcast
31..40 >40 31…40 medium no excellent yes
31…40 high yes fair yes
>40 medium no excellent no
student? yes credit rating?
no yes excellent fair
no yes yes
11
Algorithm for Decision Tree Induction
Basic algorithm (a greedy algorithm)
Tree is constructed in a top-down recursive divide-and-conquer manner
At start, all the training examples are at the root
Attributes are categorical (if continuous-valued, they are discretized in
advance)
Examples are partitioned recursively based on selected attributes
Test attributes are selected on the basis of a heuristic or statistical measure
(e.g., information gain)
12
Algorithm for Decision Tree Induction
Basic algorithm (a greedy algorithm)
Conditions for stopping partitioning
All samples for a given node belong to the same class
There are no remaining attributes for further partitioning –
majority voting is employed for classifying the leaf
There are no samples left
13
Brief Review of Entropy
m=2
14
Brief Review of Entropy
The curve approaches zero at the
extremes. For example: there is no
uncertainty when it is 100% yes or 100%
no. Entropy = 0
Midpoint:The point where the curve is
highest is approximately:50% yes - 50%
no
this is the
[Link]:uncertainty is at its
maximum.
m=2
15
Attribute Selection Measures
The three measures, in general, return good results but
Information gain:
Gain ratio:
Gini index:
16
Attribute Selection Measure:
Information Gain (ID3/C4.5)
Select the attribute with the highest information gain
Let pi be the probability that an arbitrary tuple in D belongs to
class Ci, estimated by |Ci, D|/|D|
Expected information (entropy) needed to classify a tuple in D:
m
Info( D) pi log 2 ( pi )
i 1
Information needed (after using A to split D into v partitions) to
classify D: v | D |
InfoA ( D) Info( D j )
j
j 1 | D |
Information gained by branching on attribute A
Gain(A) Info(D) InfoA(D)
17
Attribute Selection: ID.3 Information Gain
Class P: buys_computer = “yes”
Class N: buys_computer = “no”
9 9 5 5
Info( D) I (9,5) log 2 ( ) log 2 ( ) 0.940
14 14 14 14
age income student credit_rating buys_computer
<=30 high no fair no
age pi ni I(pi, ni)
<=30 high no excellent no <=30 2 3 0.971
31…40 high no fair yes 31…40 4 0 0
>40 medium no fair yes
>40 low yes fair yes >40 3 2 0.971
>40 low yes excellent no
31…40 low yes excellent yes
<=30 medium no fair no
<=30 low yes fair yes
>40 medium yes fair yes
<=30 medium yes excellent yes
31…40 medium no excellent yes
31…40 high yes fair yes
>40 medium no excellent no
18
Attribute Selection: Information Gain
age pi ni I(pi, ni)
<=30 2 3 0.971 5 4
Infoage ( D) I (2,3) I (4,0)
31…40 4 0 0 14 14
>40 3 2 0.971 5
I (3,2) 0.694
5 14
I (2,3) means “age <=30” has 5 out of 14 samples, with 2 yes’es and 3 no’s. Hence
14
•I(2,3)=0.971
•I(4,0)=0
•I(3,2)=0.971
Gain(age) Info( D) Infoage ( D) 0.246
Gain(income) 0.029
Similarly,
Gain( student) 0.151
Gain(credit _ rating) 0.048 19
Gain Ratio for Attribute Selection (C4.5)
Information gain measure is biased towards attributes with a large number of
values
C4.5 (a successor of ID3) uses gain ratio to overcome the problem
(normalization to information gain) Info(D)=0.940
=0.286+0.393+0.232 ≈0.911
age income student credit_rating buys_computer
<=30 high no fair no
<=30 high no excellent no
31…40 high no fair yes
Gainincome=Info(D)−Infoincome >40 medium no fair yes
Gainincome=0.940−0.911≈0.029 >40
>40
low
low
yes fair
yes excellent
yes
no
31…40 low yes excellent yes
<=30 medium no fair no
<=30 low yes fair yes
>40 medium yes fair yes
<=30 medium yes excellent yes
31…40 medium no excellent yes
31…40 high yes fair yes
>40 medium no excellent no 21
Gain Ratio for Attribute Selection (C4.5)
Information gain measure is biased towards attributes with a large number
of values
C4.5 (a successor of ID3) uses gain ratio to overcome the problem
(normalization to information gain) v |D | | Dj |
SplitInfoA ( D) log 2 (
j
)
j 1 |D| |D|
GainRatio(D) = Gain(D)/SplitInfo(D)
Ex.
age income student credit_rating buys_computer
gain_ratio(income) = 0.029/1.557 = 0.019 <=30 high no fair no
<=30 high no excellent no
The attribute with the maximum gain ratio is selected31…40
as the high
splitting attribute
no fair yes
>40 medium no fair yes
>40 low yes fair yes
>40 low yes excellent no
31…40 low yes excellent yes
<=30 medium no fair no
<=30 low yes fair yes
>40 medium yes fair yes
<=30 medium yes excellent yes
31…40 medium no excellent yes
31…40 high yes fair yes
>40 medium no excellent no 22
23
Gini Index (CART)
If a data set D contains examples from n classes, gini index,
gini(D) is defined as n 2
gini( D) 1 p j
j 1
where pj is the relative frequency of class j in D
If a data set D is split on A into two subsets D1 and D2, the gini
index gini(D) is defined as
|D1| |D |
gini A ( D) gini( D1) 2 gini( D 2)
|D| |D|
Reduction in Impurity:
gini( A) gini(D) giniA(D)
The attribute provides the smallest ginisplit(D) (need to
enumerate all the possible splitting points for each attribute)
24
Computation of Gini Index
Ex. D has 9 tuples in buys_computer = “yes” and 5 in “no”2 2
9 5
gini( D) 1 0.459
14 14
Suppose the attribute income partitions D into 10 in D1: {low, medium} and 4
7 yes 3 no 2 yes 2 no
in D2
10 4
giniincome{low,medium} ( D) Gini( D1 ) Gini( D2 )
14 14
0.42 0.5
Gini{low,high} is 0.458; Gini{medium,high} is
0.450. Thus, split on the {low,medium}
(and {high}) since it has the lowest Gini
index
25
26
Comparing Attribute Selection Measures
The three measures, in general, return good results but
Information gain:
biased towards multivalued attributes
Gain ratio:
tends to prefer unbalanced splits in which one partition is
much smaller than the others
Gini index:
biased to multivalued attributes
has difficulty when # of classes is large
tends to favor tests that result in equal-sized partitions
and purity in both partitions
27
Overfitting and Tree Pruning
Overfitting: An induced tree may overfit the training data
Too many branches, some may reflect anomalies due to
noise or outliers
Poor accuracy for unseen samples
Two approaches to avoid overfitting
Prepruning: Halt tree construction early ̵ do not split a node
if this would result in the goodness measure falling below a
threshold
Difficult to choose an appropriate threshold
Postpruning: Remove branches from a “fully grown” tree—
get a sequence of progressively pruned trees
Use a set of data different from the training data to
decide which is the “best pruned tree”
29
30
Scalability Framework for RainForest
Separates the scalability aspects from the criteria that
determine the quality of the tree
Builds an AVC-list: AVC (Attribute, Value, Class_label)
AVC-set (of an attribute X )
Projection of training dataset onto the attribute X and
class label where counts of individual class label are
aggregated
AVC-group (of a node n )
Set of AVC-sets of all predictor attributes at the node n
31
Rainforest: Training Set and Its AVC Sets
Training Examples AVC-set on Age AVC-set on income
age income studentcredit_rating
buys_computerAge Buy_Computer income Buy_Computer
<=30 high no fair no yes no
<=30 high no excellent no yes no
high 2 2
31…40 high no fair yes <=30 2 3
31..40 4 0 medium 4 2
>40 medium no fair yes
>40 low yes fair yes >40 3 2 low 3 1
>40 low yes excellent no
31…40 low yes excellent yes
AVC-set on
<=30 medium no fair no AVC-set on Student
credit_rating
<=30 low yes fair yes
student Buy_Computer
>40 medium yes fair yes Credit
Buy_Computer
<=30 medium yes excellent yes yes no rating yes no
31…40 medium no excellent yes yes 6 1 fair 6 2
31…40 high yes fair yes no 3 4 excellent 3 3
>40 medium no excellent no
32
BOAT (Bootstrapped Optimistic
Algorithm for Tree Construction)
Use a statistical technique called bootstrapping to create
several smaller samples (subsets), each fits in memory
Each subset is used to create a tree, resulting in several
trees
These trees are examined and used to construct a new
tree T’
It turns out that T’ is very close to the tree that would
be generated using the whole data set together
Adv: requires only two scans of DB, an incremental alg.
33
Presentation of Classification Results
May 5, 2026 Data Mining: Concepts and Techniques 34
Visualization of a Decision Tree in SGI/MineSet 3.0
May 5, 2026 Data Mining: Concepts and Techniques 35