0% acharam este documento útil (0 voto)
12 visualizações49 páginas

Preparação de Dados em Data Mining

Enviado por

r.duzinho10
Direitos autorais
© All Rights Reserved
Levamos muito a sério os direitos de conteúdo. Se você suspeita que este conteúdo é seu, reivindique-o aqui.
Formatos disponíveis
Baixe no formato PDF, TXT ou leia on-line no Scribd
0% acharam este documento útil (0 voto)
12 visualizações49 páginas

Preparação de Dados em Data Mining

Enviado por

r.duzinho10
Direitos autorais
© All Rights Reserved
Levamos muito a sério os direitos de conteúdo. Se você suspeita que este conteúdo é seu, reivindique-o aqui.
Formatos disponíveis
Baixe no formato PDF, TXT ou leia on-line no Scribd

18/10/19

Data Mining
S4

NOVA-IMS 2019/2020
Fernando Lucas Bação
bacao@[Link]
[Link]

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Preparation

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

1
18/10/19

Data Preparation

• Objective
• Transform data sets to best expose its information content

• The quality of the models should be better (or at least the same)
after preparation

• Good data is a prerequisite for good models

• Some techniques are theoretically based, other are just based on


experience

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Preparation

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

2
18/10/19

Data Preparation

• Signal vs. Noise


• In science (physics and telecommunications) noise is defined as
fluctuations and external disturbances in the flow of information
(signal) received;

• An undesired disturbance in relevant information;

• A disturbance that affects a signal and that may distort the


information carried by the signal.

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Preparation

• Signal vs. Noise


• Most databases are large in size (e.g., millions of observations), and
have high dimensionality (e.g., hundreds of variables)

• Naively, one might think that gathering more features never hurts,
since at worst they provide no new information about the class. But
in fact their benefits may be outweighed by the curse of
dimensionality.

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

3
18/10/19

Data Preparation

• Real data suffers from several problems


• Incomplete
• Missing values, lacking attributes of interest, levels of aggregation

• Noisy
• Errors and outliers

• Inconsistent
• E.g. Age=42 Birthday=31/07/1997
• Changes in scales
• Duplicate records with different values

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Treatment of Missing Data

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

4
18/10/19

Data Preparation

• Missing Data

Inputs

?
?

?
? ?
Records
?

? ?
?

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Preparation

• Missing Data

Inputs

?
?

?
? ? 10 to 3
Records
?

? ?
?

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

5
18/10/19

Data Preparation

• Missing Data

Inputs

?
?

?
? ?
Records
?

? ?
?

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Preparation

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

6
18/10/19

Data Preparation

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Preparation

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

7
18/10/19

Data Preparation

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Preparation

Edu

Dayswu Income

Age

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

8
18/10/19

Data Preparation

• Missing data
• Delete variables (you loose information)
• Delete records (potential bias);
• To examine and manually enter a probable value (tedious +
infeasible);
• Automatically fill in with a measure of central tendency (i.e.
mean, median, mode);
• Automatically fill in with a measure of central tendency of a
subset (e.g. men and women);
• To fill in with values from ​similar individuals (nearest
neighbours);
• Predictive model (linear regression, multiple linear regression);
• Code the missing data explicitly.
Instituto Superior de Estatística e Gestão de Informação
Universidade Nova de Lisboa

Data Preparation

• Missing data
• The most practical approach to the problem is to initially use
the quickest and simplest option;

• After achieving some preliminary results we can


comparatively analyze the performance of the model in the full
sample patterns an in those where there was a need to
estimate missing values;

• In the event that the error is significantly higher than in other


data, then we will try to use another method in order to
improve results.

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

9
18/10/19

Outlier treatment

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Preparation

• Outliers
• In statistics, an outlier is an observation point that is distant from
other observations.

• An outlier may be due to variability in the measurement or it may


indicate experimental error; the latter are sometimes excluded
from the data set.

• Extreme cases in one or more variables and with great impact on


the interpretation of results;
• Outliers may come from:
• Unusual but correct situations (the Bill Gates effect),
• Incorrect measurements,
• Errors in data collection;
• Lack of code for missing data.

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

10
18/10/19

Data Preparation

• Outliers (leverage effect)

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Preparation

• Remove Data Outliers


• Automatic limitation (thresholding)
• The imposition of maximum and minimum values ​​
for the variables (age – 0 e 100)

Histograma de Frequência

140

120

100 Data to
80
remove

60

40

20

0
7

4
19

27

36

45

80

89
10

54

63

71

98
1

12
10

11

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

11
18/10/19

Data Preparation

• Remove Data Outliers

Frequency
10000
8000 Frequency
6000
800
4000 700
2000 600
500
0 400
7.61
621.405
1235.2
1848.995
2462.79
3076.585
3690.38
4304.175
4917.97
5531.765

6759.355
7373.15
7986.945

9214.535

10442.125
11055.92
11669.715
More
6145.56

8600.74

9828.33

300
200
100
0

7.61
37.5345
67.459
97.3835
127.308
157.2325
187.157
217.0815

276.9305
306.855
336.7795

396.6285

456.4775

516.3265
546.251
576.1755
More
247.006

366.704

426.553

486.402
Instituto Superior de Estatística e Gestão de Informação
Universidade Nova de Lisboa

Data Preparation

fmi:
[Link]
[Link]
[Link]

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

12
18/10/19

Data Preparation

• Remove Data Outliers


– Normal data distribution
• 3σ +/- Average

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Preparation

• Remove Data Outliers


– In two dimensions…

1300

1200

1100

1000

900

800

700

600

500
10 20 30 40 50 60 70 80

Series1 Series2

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

13
18/10/19

Data Preparation

• Remove Data Outliers

– Cluster Analysis (K-means)

– Self-Organizing Maps

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Imbalanced Learning

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

14
18/10/19

Quick quizz:
Yes No

Is it possible to have a useless classifier with 99%


accuracy?

Is it possible to achieve a 99.9% accuracy with a


trivial classifier?

29
Instituto Superior de Estatística e Gestão de Informação
Universidade Nova de Lisboa

Quick quizz:
Yes No

Is it possible to have a useless classifier with 99% X


accuracy?

Is it possible to achieve a 99.9% accuracy with a X


trivial classifier?

30
Instituto Superior de Estatística e Gestão de Informação
Universidade Nova de Lisboa

15
18/10/19

Quick quizz:
Yes No

Is it possible to have a useless classifier with 99% accuracy?


X
If the minority class is 1%

Is it possible to achieve a 99.9% accuracy with a trivial classifier?


X
If the minority class is 0,1%

31
Instituto Superior de Estatística e Gestão de Informação
Universidade Nova de Lisboa

• Class Imbalance

• Is this frequent in real-world applications?

• Credit Card frauds - ~2% per year.

• HIV prevalence in the USA - ~0.4%.

• Disk drive failures - ~1% per year.

• Factory production defects - ~ 0.1%.

• Business churn - ~3%

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

16
18/10/19

Imbalanced Learning

• Imbalanced Learning
• An imbalanced learning problem is defined as a classification task
for binary or multi-class datasets where a significant asymmetry
exists between the number of instances for the various classes.

• The dominant class is called the majority class (negative cases)


while the rest of the classes are called the minority classes
(positive cases)

• The Imbalance Ratio (IR), is the ratio between the majority class and
the minority class, (depends on the type of application and for binary
problems values between 100 and 100.000 have been observed)

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Imbalanced Learning

• Imbalanced Learning

• Standard learning methods induce a bias in favor of the


majority class during training.

• This happens because the minority classes contribute less to


the maximization of the objective function, which is usually
accuracy.

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

17
18/10/19

Imbalanced Learning

• Imbalanced Learning

• Standard learning methods induce a bias in favor of the


majority class during training.

• This happens because the minority classes contribute less to


the maximization of the objective function, which is usually
accuracy.
Actual Value
Positives Negatives
Negatives Positives

TP True FP False
Predicted Value

Positives Positives

FN False TN True
Negatives Negatives

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Imbalanced Learning

• Imbalanced Learning

• Standard learning methods induce a bias in favor of the


majority class during training.

• This happens because the minority classes contribute less to


the maximization of the objective function, which is usually
accuracy.
Actual Value
Positives Negatives
Positives

TP True Positives FP False Positives


Predicted Value

TN
Negatives

FN False Negatives
True
Negatives
Instituto Superior de Estatística e Gestão de Informação
Universidade Nova de Lisboa

18
18/10/19

Imbalanced Learning

• Imbalanced Learning

• Standard learning methods induce a bias in favor of the


majority class during training.

• This happens because the minority classes contribute less to


the maximization of the objective function, which is usually
accuracy.

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Imbalanced Learning

• Imbalanced Learning

• By optimizing classication accuracy, most algorithms assume a


balanced class distribution

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

19
18/10/19

Imbalanced Learning

• Imbalanced Learning

• Another inherent assumption of many classication algorithms is the


uniformity of misclassication costs

Actual Value
Positives Negatives

Negatives Positives
TP FP
Predicted Value

True False
Positives Positives

FN TN
False True
Negatives Negatives

• Frequently the costs of misclassifying the minority class are much


higher than the costs of misclassification of the majority class

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Imbalanced Learning

• Imbalanced Learning

• Diseases screening tests are a typical situation in in which


false negatives involve a much higher cost than the false
positives.

Actual Value
Positives Negatives
Negatives Positives

TP FP
Predicted Value

True False
Positives Positives
FN
TN
False True
Negatives Negatives

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

20
18/10/19

Imbalanced Learning

• Notes on experimental procedures


• Metrics

• Fmeasure

• Gmean

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Approaches to the Imbalanced Learning Problem

Undersampling

Solutions to Imbalanced
Oversampling
Learning

Hybrid approaches

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

21
18/10/19

Approaches to the Imbalanced Learning Problem

Majority class

Minority class

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Approaches to the Imbalanced Learning Problem


Random Undersampling

Majority class

Minority class

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

22
18/10/19

Approaches to the Imbalanced Learning Problem


Random Oversampling

Majority class

Minority class

Generated sample

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Approaches to the Imbalanced Learning Problem


Hybrid Approach

Majority class

Minority class

Generated sample

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

23
18/10/19

SMOTE: Synthetic Minority Over-sampling


TEchnique

47
Instituto Superior de Estatística e Gestão de Informação
Universidade Nova de Lisboa

Imbalanced Learning

• SMOTE
• The idea underlying SMOTE is as simple as it is clever.

• The basic steps are:

• randomly selecting a minority class instance x;

• then it defines the set of k-nearest neighbors (xknn);

• randomly selects another minority class sample x’ from the xknn set.

• xgen is generated by using a linear interpolation of x and x’, which can be


expressed as:

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

24
18/10/19

General aspects of data collection – acquiring knowledge

Majority class

Minority class

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Informed Oversampling

Majority class

Minority class

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

25
18/10/19

Informed Oversampling

Majority class

Minority class

j n

x
m
i

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Informed Oversampling

Majority class

Minority class

x
x’

j x’

x
m
i

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

26
18/10/19

Informed Oversampling

Majority class

Minority class

x
x’
xgen

j x’

x xgen
m
i

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Informed Oversampling

Majority class

Minority class

x
x’
xgen

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

27
18/10/19

Informed Oversampling

Majority class

Minority class

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

General aspects of data collection

• Use of artificial data:


– It is always preferable to use real data;
• Create data as realistic as possible;
• Make artificial data as representative as possible.

– The quality of the model is constrained by the quality


of the data;
– Creating artificial data translates into the introduction
of some noise.

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

28
18/10/19

Discretization

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Preprocessing

• Discretization
• Divide the range of a continuous variable into intervals
• Some classification algorithms only accept discrete attributes

• Reduce data size


• Prepare for further analysis

• Frequently called binning

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

29
18/10/19

Data Preprocessing

• Discretization
• Equal-width binning
• Divides the range into n intervals of equal size
• If A and B are the minimum and the maximum values of the attribute, the
width of the intervals will be: w=(B-A)/N
• Most simple method
• Outliers may dominate

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Preprocessing

• Discretization
• Equal-depth binning
• Divides the range into n intervals, each containing approximately the
same number of samples
• Generally preferred avoids clumps
• Gives more intuitive breakpoints
• Shouldn’t break frequent values across bins

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

30
18/10/19

Data Preprocessing

• Discretization
• Entropy (also called Expected Information) based
discretization

Very Impure Less Impure Pure

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Preprocessing

• Discretization
• Entropy (also called Expected Information) based
discretization
• Sort examples in increased order
• Each value forms an interval (m intervals)
• Calculate the entropy measure of each discretization
• Find the binary split boundary that minimizes the entropy function over
all possible partitions. The split is selected as a binary discretization
• Apply the process recursively until some stopping criteria is met

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

31
18/10/19

Data Preprocessing

• Discretization
• Entropy (also called Expected Information) based
discretization
Very Impure

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Preprocessing

• Discretization
• Entropy based discretization

Income <= 50K Income > 50K


Age < 25 4 6

Income <= 50K Income > 50K


Age < 25 9 1

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

32
18/10/19

Data Preprocessing

• Discretization
• Entropy based discretization
• Entropy
• Idea: maximize info
• It measures the purity of a partition:

E = - p log2(p)

• Where p is the probability of the examples belong to a specific class


1
Entropy

0 p 1
Instituto Superior de Estatística e Gestão de Informação
Universidade Nova de Lisboa

Data Preprocessing

• Discretization
• Entropy based discretization

• Partition entropy:
#C
Ent(S) = −∑ pi log 2 ( pi )
i=1

• Gain in choosing A attribute:

Gain(Entnew ) = Entinitial − Entnew

# Sv
Gain(S, A) = Ent(S) − ∑ Ent(Sv )
v∈Valores( A) # S

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

33
18/10/19

Data Preprocessing

• Discretization
• Entropy based discretization

Income <= 50K Income > 50K


13 7

Ent(S) = −(13 log 2 (13 )+ 7 log 2 ( 7 )) = 0.403+ 0.530 = 0.934


20 20 20 20

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Preprocessing

• Discretization
• Entropy based discretization

Income <= 50K Income > 50K


Age < 25 4 6

Ent(Age < 25) = −( 4 log 2 ( 4 )+ 6 log 2 (6 )) = 0.529 + 0.442 = 0.971


10 10 10 10

Income <= 50K Income > 50K


Age ≥ 25 9 1

Ent(Age ≥ 25) = −(9 log 2 (9 ) + 1 log 2 ( 1 )) = 0.137 + 0.332 = 0.469


10 10 10 10

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

34
18/10/19

Data Preprocessing

• Discretization
• Entropy based discretization

Income <= 50K Income > 50K


Age < 25 9 1
Age ≥ 25 4 6

# Sv
∑ Ent(Sv ) = 1 (0.469) + 1 (0, 971) = 0.72
2 2
v∈Valores( A) # S

# Sv
Gain(S, A) = Ent(S) − ∑ Ent(Sv )
v∈Valores( A) # S

Ent(S) = −(13 log 2 (13 )+ 7 log 2 ( 7 )) = 0.403+ 0.530 = 0.934


20 20 20 20

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Preprocessing

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

35
18/10/19

Data Preprocessing

• Reasons:
• Noise Reduction;
• Signal amplification;
• Size Reduction of the Input Space;
• Remove correlated variables
• Remove irrelevant variables
• Constructing ratios and derived variables

• Domain-specific knowledge application;


• Normalization;

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Reducing Input Space

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

36
18/10/19

Data Preprocessing

• Additional considerations about data:

• Curse of dimensionality – the input space grows


exponentially with the number of input
variables;

• The larger the input space, the more data and


computing power we need.

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Preprocessing

Three groups, right?

The curse of dimensionality

Not exactly...

When the dimensionality increases, the space becomes more sparse


and it becomes more difficult to find groups

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

37
18/10/19

Data Preprocessing

• Size Reduction of the Input Space (or feature selection):


Two major principles:
Relevance and Redundancy

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Reducing Input Space


Relevancy

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

38
18/10/19

Data Preprocessing

• Size Reduction of the Input Space:

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Preprocessing

• Size Reduction of the Input Space:

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

39
18/10/19

Data Preprocessing

• Size Reduction of the Input Space:


• Heuristic feature selection methods:
• Best single features
• Choose by information gain measures (e.g. entropy)
• A feature is interesting if it reduces uncertainty

No improvement Perfect Split

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Preprocessing

• Size Reduction of the Input Space:


• To create input combinations
• Height2/weight (obesity index)
• Population/area
• Euros spent/nº of purchases
• Euros spent/time as customer
• Debt/income

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

40
18/10/19

Data Preprocessing

• Size Reduction of the Input Space:


• To create input combinations

1. Customer ID (could be 8. Average number of different


anonymous) products purchased per
2. Total revenue for the transaction
customer 9. Relative spend on each
3. Number of transactions per product
customer (frequency) 10. NRS on each product
4. Average time between (and where a product taxonomy
transactions (transaction exists):
interval) 11. Relative spend in each
5. Variance of transaction product subgroup
interval 12. NRS in each product
6. Customer stability index subgroup
(ratio of (5)/(4)) 13. NRS in each product group
7. Days since last visit (recency)
Instituto Superior de Estatística e Gestão de Informação
Universidade Nova de Lisboa

Reducing Input Space


Redundancy

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

41
18/10/19

Data Preprocessing

• Size Reduction of the Input Space:

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Preprocessing

• Size Reduction of the Input Space:


• Principal Component Analysis

• A procedure that uses an orthogonal transformation to


convert a set of observations of possibly correlated
variables into a set of values of linearly uncorrelated
variables called principal components.
• The number of principal components is equal to the
number of original variables.
• This transformation is defined in such a way that the first
principal component has the largest possible variance
(that is, accounts for as much of the variability in the data
as possible), and each succeeding component in turn has
the highest variance.

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

42
18/10/19

Data Preprocessing

• Size Reduction of the Input Space:


• Principal Component Analysis

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Preprocessing

• Size Reduction of the Input Space:


• Principal Component Analysis

1st PC
2nd PC

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

43
18/10/19

Data Preprocessing

• Size Reduction of the Input Space:


• Principal Component Analysis

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Preprocessing

• Size Reduction of the Input Space:


• Principal Component Analysis

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

44
18/10/19

Data Preprocessing

• Size Reduction of the Input Space:


• Principal Component Analysis (careful)

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Preprocessing

• Size Reduction of the Input Space:


• Principal Component Analysis (careful)

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

45
18/10/19

Data Preprocessing

• Size Reduction of the Input Space:


• Principal Component Analysis (careful)

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Standardization

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

46
18/10/19

Data Preprocessing

• Normalization:

• Models assume that the distances in different


directions of the input space have the same
importance.

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Preprocessing

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

47
18/10/19

Data Preprocessing

• Normalization:

• Variables come in many different scales


(percentages, euros, kilos, meters, days…)
Total Population

• In the figure we can see that however different


they are in terms of percentage of population
working in the industry, clusters will always be
defined by the total population

• Normalization: is about adjusting values


measured on different scales to a common scale

Percentage of
people in industry

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Data Preprocessing

• Normalization:

⎛ y − min1 ⎞
• Min-Max y' = ⎜ ⎟(max 2 − min 2) + min 2
⎝ max 1 − min1 ⎠
optional

y−µ
• Zscore y' =
std

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

48
18/10/19

Data Preprocessing

• Normalization:

Min Max Dados Originais


1.2 160000

140000
1

120000

0.8
100000

0.6 80000

60000
0.4

40000

0.2
20000

0 0

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa

Questions?

Instituto Superior de Estatística e Gestão de Informação


Universidade Nova de Lisboa 98

49

Você também pode gostar