Preparação de Dados em Data Mining
Preparação de Dados em Data Mining
Data Mining
S4
NOVA-IMS 2019/2020
Fernando Lucas Bação
bacao@[Link]
[Link]
Data Preparation
1
18/10/19
Data Preparation
• Objective
• Transform data sets to best expose its information content
• The quality of the models should be better (or at least the same)
after preparation
Data Preparation
2
18/10/19
Data Preparation
Data Preparation
• Naively, one might think that gathering more features never hurts,
since at worst they provide no new information about the class. But
in fact their benefits may be outweighed by the curse of
dimensionality.
3
18/10/19
Data Preparation
• Noisy
• Errors and outliers
• Inconsistent
• E.g. Age=42 Birthday=31/07/1997
• Changes in scales
• Duplicate records with different values
4
18/10/19
Data Preparation
• Missing Data
Inputs
?
?
?
? ?
Records
?
? ?
?
Data Preparation
• Missing Data
Inputs
?
?
?
? ? 10 to 3
Records
?
? ?
?
5
18/10/19
Data Preparation
• Missing Data
Inputs
?
?
?
? ?
Records
?
? ?
?
Data Preparation
6
18/10/19
Data Preparation
Data Preparation
7
18/10/19
Data Preparation
Data Preparation
Edu
Dayswu Income
Age
8
18/10/19
Data Preparation
• Missing data
• Delete variables (you loose information)
• Delete records (potential bias);
• To examine and manually enter a probable value (tedious +
infeasible);
• Automatically fill in with a measure of central tendency (i.e.
mean, median, mode);
• Automatically fill in with a measure of central tendency of a
subset (e.g. men and women);
• To fill in with values from similar individuals (nearest
neighbours);
• Predictive model (linear regression, multiple linear regression);
• Code the missing data explicitly.
Instituto Superior de Estatística e Gestão de Informação
Universidade Nova de Lisboa
Data Preparation
• Missing data
• The most practical approach to the problem is to initially use
the quickest and simplest option;
9
18/10/19
Outlier treatment
Data Preparation
• Outliers
• In statistics, an outlier is an observation point that is distant from
other observations.
10
18/10/19
Data Preparation
Data Preparation
Histograma de Frequência
140
120
100 Data to
80
remove
60
40
20
0
7
4
19
27
36
45
80
89
10
54
63
71
98
1
12
10
11
11
18/10/19
Data Preparation
Frequency
10000
8000 Frequency
6000
800
4000 700
2000 600
500
0 400
7.61
621.405
1235.2
1848.995
2462.79
3076.585
3690.38
4304.175
4917.97
5531.765
6759.355
7373.15
7986.945
9214.535
10442.125
11055.92
11669.715
More
6145.56
8600.74
9828.33
300
200
100
0
7.61
37.5345
67.459
97.3835
127.308
157.2325
187.157
217.0815
276.9305
306.855
336.7795
396.6285
456.4775
516.3265
546.251
576.1755
More
247.006
366.704
426.553
486.402
Instituto Superior de Estatística e Gestão de Informação
Universidade Nova de Lisboa
Data Preparation
fmi:
[Link]
[Link]
[Link]
12
18/10/19
Data Preparation
Data Preparation
1300
1200
1100
1000
900
800
700
600
500
10 20 30 40 50 60 70 80
Series1 Series2
13
18/10/19
Data Preparation
– Self-Organizing Maps
Imbalanced Learning
14
18/10/19
Quick quizz:
Yes No
29
Instituto Superior de Estatística e Gestão de Informação
Universidade Nova de Lisboa
Quick quizz:
Yes No
30
Instituto Superior de Estatística e Gestão de Informação
Universidade Nova de Lisboa
15
18/10/19
Quick quizz:
Yes No
31
Instituto Superior de Estatística e Gestão de Informação
Universidade Nova de Lisboa
• Class Imbalance
16
18/10/19
Imbalanced Learning
• Imbalanced Learning
• An imbalanced learning problem is defined as a classification task
for binary or multi-class datasets where a significant asymmetry
exists between the number of instances for the various classes.
• The Imbalance Ratio (IR), is the ratio between the majority class and
the minority class, (depends on the type of application and for binary
problems values between 100 and 100.000 have been observed)
Imbalanced Learning
• Imbalanced Learning
17
18/10/19
Imbalanced Learning
• Imbalanced Learning
TP True FP False
Predicted Value
Positives Positives
FN False TN True
Negatives Negatives
Imbalanced Learning
• Imbalanced Learning
TN
Negatives
FN False Negatives
True
Negatives
Instituto Superior de Estatística e Gestão de Informação
Universidade Nova de Lisboa
18
18/10/19
Imbalanced Learning
• Imbalanced Learning
Imbalanced Learning
• Imbalanced Learning
19
18/10/19
Imbalanced Learning
• Imbalanced Learning
Actual Value
Positives Negatives
Negatives Positives
TP FP
Predicted Value
True False
Positives Positives
FN TN
False True
Negatives Negatives
Imbalanced Learning
• Imbalanced Learning
Actual Value
Positives Negatives
Negatives Positives
TP FP
Predicted Value
True False
Positives Positives
FN
TN
False True
Negatives Negatives
20
18/10/19
Imbalanced Learning
• Fmeasure
• Gmean
Undersampling
Solutions to Imbalanced
Oversampling
Learning
Hybrid approaches
21
18/10/19
Majority class
Minority class
Majority class
Minority class
22
18/10/19
Majority class
Minority class
Generated sample
Majority class
Minority class
Generated sample
23
18/10/19
47
Instituto Superior de Estatística e Gestão de Informação
Universidade Nova de Lisboa
Imbalanced Learning
• SMOTE
• The idea underlying SMOTE is as simple as it is clever.
• randomly selects another minority class sample x’ from the xknn set.
24
18/10/19
Majority class
Minority class
Informed Oversampling
Majority class
Minority class
25
18/10/19
Informed Oversampling
Majority class
Minority class
j n
x
m
i
Informed Oversampling
Majority class
Minority class
x
x’
j x’
x
m
i
26
18/10/19
Informed Oversampling
Majority class
Minority class
x
x’
xgen
j x’
x xgen
m
i
Informed Oversampling
Majority class
Minority class
x
x’
xgen
27
18/10/19
Informed Oversampling
Majority class
Minority class
28
18/10/19
Discretization
Data Preprocessing
• Discretization
• Divide the range of a continuous variable into intervals
• Some classification algorithms only accept discrete attributes
29
18/10/19
Data Preprocessing
• Discretization
• Equal-width binning
• Divides the range into n intervals of equal size
• If A and B are the minimum and the maximum values of the attribute, the
width of the intervals will be: w=(B-A)/N
• Most simple method
• Outliers may dominate
Data Preprocessing
• Discretization
• Equal-depth binning
• Divides the range into n intervals, each containing approximately the
same number of samples
• Generally preferred avoids clumps
• Gives more intuitive breakpoints
• Shouldn’t break frequent values across bins
30
18/10/19
Data Preprocessing
• Discretization
• Entropy (also called Expected Information) based
discretization
Data Preprocessing
• Discretization
• Entropy (also called Expected Information) based
discretization
• Sort examples in increased order
• Each value forms an interval (m intervals)
• Calculate the entropy measure of each discretization
• Find the binary split boundary that minimizes the entropy function over
all possible partitions. The split is selected as a binary discretization
• Apply the process recursively until some stopping criteria is met
31
18/10/19
Data Preprocessing
• Discretization
• Entropy (also called Expected Information) based
discretization
Very Impure
Data Preprocessing
• Discretization
• Entropy based discretization
32
18/10/19
Data Preprocessing
• Discretization
• Entropy based discretization
• Entropy
• Idea: maximize info
• It measures the purity of a partition:
E = - p log2(p)
0 p 1
Instituto Superior de Estatística e Gestão de Informação
Universidade Nova de Lisboa
Data Preprocessing
• Discretization
• Entropy based discretization
• Partition entropy:
#C
Ent(S) = −∑ pi log 2 ( pi )
i=1
# Sv
Gain(S, A) = Ent(S) − ∑ Ent(Sv )
v∈Valores( A) # S
33
18/10/19
Data Preprocessing
• Discretization
• Entropy based discretization
Data Preprocessing
• Discretization
• Entropy based discretization
34
18/10/19
Data Preprocessing
• Discretization
• Entropy based discretization
# Sv
∑ Ent(Sv ) = 1 (0.469) + 1 (0, 971) = 0.72
2 2
v∈Valores( A) # S
# Sv
Gain(S, A) = Ent(S) − ∑ Ent(Sv )
v∈Valores( A) # S
Data Preprocessing
35
18/10/19
Data Preprocessing
• Reasons:
• Noise Reduction;
• Signal amplification;
• Size Reduction of the Input Space;
• Remove correlated variables
• Remove irrelevant variables
• Constructing ratios and derived variables
36
18/10/19
Data Preprocessing
Data Preprocessing
Not exactly...
37
18/10/19
Data Preprocessing
38
18/10/19
Data Preprocessing
Data Preprocessing
39
18/10/19
Data Preprocessing
Data Preprocessing
40
18/10/19
Data Preprocessing
41
18/10/19
Data Preprocessing
Data Preprocessing
42
18/10/19
Data Preprocessing
Data Preprocessing
1st PC
2nd PC
43
18/10/19
Data Preprocessing
Data Preprocessing
44
18/10/19
Data Preprocessing
Data Preprocessing
45
18/10/19
Data Preprocessing
Data Standardization
46
18/10/19
Data Preprocessing
• Normalization:
Data Preprocessing
47
18/10/19
Data Preprocessing
• Normalization:
Percentage of
people in industry
Data Preprocessing
• Normalization:
⎛ y − min1 ⎞
• Min-Max y' = ⎜ ⎟(max 2 − min 2) + min 2
⎝ max 1 − min1 ⎠
optional
y−µ
• Zscore y' =
std
48
18/10/19
Data Preprocessing
• Normalization:
140000
1
120000
0.8
100000
0.6 80000
60000
0.4
40000
0.2
20000
0 0
Questions?
49