0% found this document useful (0 votes)
2 views16 pages

Module 2.docx(ml2)

The document discusses various performance measures for evaluating regression models, including Mean Absolute Error (MAE), Mean Squared Error (MSE), Root Mean Squared Error (RMSE), R-Squared (R²) Score, and Mean Absolute Percentage Error (MAPE). It explains the definitions, formulas, interpretations, advantages, and disadvantages of each metric. Additionally, it covers the Gini Index for decision tree classification and techniques to handle overfitting.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views16 pages

Module 2.docx(ml2)

The document discusses various performance measures for evaluating regression models, including Mean Absolute Error (MAE), Mean Squared Error (MSE), Root Mean Squared Error (RMSE), R-Squared (R²) Score, and Mean Absolute Percentage Error (MAPE). It explains the definitions, formulas, interpretations, advantages, and disadvantages of each metric. Additionally, it covers the Gini Index for decision tree classification and techniques to handle overfitting.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Performance Measures for Regression (5 Marks)

Regression models predict continuous numerical values. To evaluate how well a


regression model performs, different performance metrics are used to compare the
actual values with the predicted values.

Let:

●​ 𝑦𝑖= Actual value

^
●​ 𝑦𝑖= Predicted value

●​ 𝑛= Number of observations

1. Mean Absolute Error (MAE)

Definition:​
MAE is the average of the absolute differences between the actual and predicted
values.

Formula:
𝑛
1 ^
𝑀𝐴𝐸 = 𝑛
∑ ∣𝑦𝑖 − 𝑦𝑖∣
𝑖=1

Interpretation:

●​ Lower MAE indicates better model performance.

●​ It treats all errors equally.

Example:

Absolute
Actual Predicted
Error

10 12 2

20 18 2

30 29 1
2+2+1
𝑀𝐴𝐸 = 3
= 1. 67

Advantages:

●​ Easy to understand.

●​ Less sensitive to outliers.


2. Mean Squared Error (MSE)

Definition:​
MSE is the average of the squared differences between actual and predicted values.

Formula:
𝑛
1 ^ 2
𝑀𝑆𝐸 = 𝑛
∑ (𝑦𝑖 − 𝑦𝑖)
𝑖=1

Example:

Squared
Actual Predicted
Error

10 12 4

20 18 4

30 29 1
4+4+1
𝑀𝑆𝐸 = 3
= 3

Advantages:

●​ Penalizes large errors more heavily.

●​ Widely used in machine learning optimization.

3. Root Mean Squared Error (RMSE)

Definition:​
RMSE is the square root of MSE.

Formula:

𝑛
1 ^ 2
𝑅𝑀𝑆𝐸 = 𝑛
∑ (𝑦𝑖 − 𝑦𝑖)
𝑖=1

Using the previous example:

𝑅𝑀𝑆𝐸 = 3 = 1. 732

Advantages:

●​ Expressed in the same unit as the target variable.

●​ Easy to interpret.
2
4. R-Squared (𝑅 ) Score (Coefficient of Determination)

Definition:​
It measures how well the regression model explains the variability in the data.

Formula:
^ 2
2 ∑(𝑦𝑖−𝑦𝑖)
𝑅 = 1− 2
∑(𝑦𝑖−𝑦ˉ)

where 𝑦ˉis the mean of actual values.

Interpretation:


Meaning
Value

1 Perfect prediction

0 Model performs no better than the mean

Model performs worse than predicting the


<0
mean

Example:​
2
If 𝑅 = 0. 92, the model explains 92% of the variance in the data.

5. Mean Absolute Percentage Error (MAPE)

Definition:​
MAPE measures prediction accuracy as a percentage.

Formula:
𝑛 𝑦𝑖−𝑦𝑖
^
100
𝑀𝐴𝑃𝐸 = 𝑛
∑ ∣ 𝑦𝑖

𝑖=1

Example:

Actual = 100, Predicted = 90


∣100−90∣
𝑀𝐴𝑃𝐸 = 100
×100 = 10%

Advantages:

●​ Easy to understand because it is expressed as a percentage.


Disadvantage:

●​ Cannot be used when actual values are zero.

Comparison of Regression Performance Measures

Best Sensitive to
Metric Formula Unit
Value Outliers

MAE (\frac{1}{n}\sum y-\hat{y} ) 0


1 ^ 2
MSE 𝑛
∑(𝑦 − 𝑦) 0 Yes Squared units

Same as
RMSE 𝑀𝑆𝐸 0 Yes
target

R² 𝑆𝑆𝑟𝑒𝑠
1− 1 No Unitless
Score 𝑆𝑆𝑡𝑜𝑡

Average percentage
MAPE 0% Moderate Percentage
error

Conclusion

Performance measures help evaluate the accuracy of regression models. MAE


measures the average error, MSE and RMSE penalize larger errors more strongly,
R² Score indicates how well the model explains the data, and MAPE expresses
prediction error as a percentage. The choice of metric depends on the application
and the importance of large prediction errors.
Ans.

Given Dataset

Instance Color Wig Num. Ears Emotion

1 G Y 2 S

2 G N 2 S

3 G N 2 S

4 B N 2 S

5 B N 2 H

6 R N 2 H

7 R N 2 H

8 R N 2 H

9 R Y 3 H

Target Variable (Emotion):

●​ S=4

●​ H=5

(i) Find the Root Node using Gini Index

Step 1: Calculate Gini of Parent Node

Total samples = 9

4 2 5 2
𝑃(𝑆) =
4
9
, 𝑃(𝐻) =
5
9
𝐺𝑖𝑛𝑖(𝐷) = 1 − ( ) −( )
9 9
= 1−
16
81

25
81
= 1−
41
81
=
40
81
= 0. 494
Attribute 1: Color

Possible values:

●​ G

●​ B

●​ R

Color = G

Instances = 3

Emotion = S,S,S

Pure node
2 2
𝐺𝑖𝑛𝑖(𝐺) = 1 − (1) − (0) = 0

Color = B

Instances = 2

Emotion = S,H

1 2 1 2
𝑃(𝑆) =
1
2
, 𝑃(𝐻) =
1
2
𝐺𝑖𝑛𝑖(𝐵) = 1 − ( ) −( )
2 2
= 0. 5

Color = R

Instances =4

Emotion = H,H,H,H

Pure node

𝐺𝑖𝑛𝑖(𝑅) = 0

Weighted Gini(Color)
3 2 4 1
= 9
(0) + 9
(0. 5) + 9
(0) = 9
= 0. 111

Attribute 2: Wig
Wig = Y

Instances =2

Emotion = S,H

𝐺𝑖𝑛𝑖(𝑌) = 0. 5

Wig = N

Instances =7

Emotion =

S,S,S,H,H,H,H

3 2 4 2
𝑃(𝑆) =
3
7
, 𝑃(𝐻) =
4
7
𝐺𝑖𝑛𝑖(𝑁) = 1 − ( ) −( )
7 7
= 1−
9
49

16
49
=
24
49
= 0. 490

Weighted Gini(Wig)
2 7
= 9
(0. 5) + 9
(0. 490) = 0. 111 + 0. 381 = 0. 492

Attribute 3: Number of Ears

Possible values:

Ears =2

Instances =8

Emotion

S,S,S,S,H,H,H,H
4 4
𝑃(𝑆) = 8
, 𝑃(𝐻) = 8
𝐺𝑖𝑛𝑖(2) = 0. 5

Ears =3

Instances =1

Emotion =H

Pure node

𝐺𝑖𝑛𝑖(3) = 0
Weighted Gini(Ears)
8 1 4
= 9
(0. 5) + 9
(0) = 9
= 0. 444

Step 2: Compare Gini Values

Weighted
Attribute
Gini

Color 0.111

Wig 0.492

Number of
0.444
Ears

Root Node

The attribute with the lowest Gini Index is selected as the root.

𝑅𝑜𝑜𝑡 𝑁𝑜𝑑𝑒 = 𝐶𝑜𝑙𝑜𝑟

Decision Tree

Color

/ | \

G B R

| | |

S (Split) H

●​ G → S (Pure)

●​ R → H (Pure)

●​ B contains mixed classes (S, H), so it requires further splitting.

(ii) Techniques to Handle Overfitting in Decision Trees

Overfitting occurs when the decision tree learns the training data too closely,
including noise, resulting in poor performance on unseen data.
1. Pre-Pruning (Early Stopping)

●​ Stop tree growth before it becomes too deep.

●​ Set constraints such as:

o​ Maximum tree depth

o​ Minimum samples per node

o​ Minimum information gain

●​ Prevents unnecessary splits.

2. Post-Pruning

●​ Build the full tree first.

●​ Remove branches that contribute little to prediction accuracy.

●​ Improves generalization.

3. Limit Maximum Depth

●​ Restrict the height of the tree.

●​ Smaller trees are less likely to memorize noise.

4. Increase Minimum Samples for Split

●​ Require a minimum number of samples before a node can split.

●​ Avoids creating branches from very few examples.

5. Minimum Samples per Leaf

●​ Ensure each leaf contains a sufficient number of records.

●​ Prevents tiny, highly specific leaf nodes.

6. Cross-Validation

●​ Evaluate the tree using k-fold cross-validation.

●​ Select the tree size that gives the best validation accuracy rather than the
best training accuracy.

7. Feature Selection

●​ Remove irrelevant or redundant features before training.

●​ Reduces unnecessary complexity in the tree.

Final Answer
(i)

●​ Gini(Color) = 0.111

●​ Gini(Wig) = 0.492

●​ Gini(Number of Ears) = 0.444

Since Color has the minimum Gini Index, Color is selected as the root node of
the decision tree.

(ii)

Techniques to avoid overfitting in decision trees include:

1.​ Pre-pruning (early stopping)

2.​ Post-pruning

3.​ Limiting maximum depth

4.​ Increasing minimum samples for split

5.​ Setting minimum samples per leaf

6.​ Cross-validation

7.​ Feature selection

These techniques help improve the model's ability to generalize to unseen data.

Ans. Solution: Construct a Decision Tree using Gini Index

Step 1: Given Dataset

Car No Colour Type Origin Stolen

1 Red Sports Domestic Yes


Car No Colour Type Origin Stolen

2 Red Sports Domestic No

3 Red Sports Domestic Yes

4 Yellow Sports Domestic No

5 Yellow Sports Imported Yes

6 Yellow SUV Imported No

7 Yellow SUV Imported Yes

8 Yellow SUV Domestic No

9 Red SUV Imported No

10 Red Sports Imported Yes

Target Class:

●​ Yes = 5

●​ No = 5

Step 2: Calculate Gini of Parent Node

5 2 5 2
𝐺𝑖𝑛𝑖(𝐷) = 1 − ( ) −( )
10 10
= 1 − 0. 25 − 0. 25 = 0. 5

Step 3: Compute Gini for Each Attribute

Attribute 1: Colour

Red

Instances = 5

(Yes, No, Yes, No, Yes)

Yes = 3

No = 2

3 2 2 2
𝐺𝑖𝑛𝑖(𝑅𝑒𝑑) = 1 − ( ) ( )
5
− 5
= 1−
9
25

4
25
=
12
25
= 0. 48
Yellow

Instances =5

(No, Yes, No, Yes, No)

Yes =2

No =3

𝐺𝑖𝑛𝑖(𝑌𝑒𝑙𝑙𝑜𝑤) = 0. 48

Weighted Gini(Colour)
5 5
= 10
(0. 48) + 10
(0. 48) = 0. 48

Attribute 2: Type

Sports

Instances =6

Yes, No, Yes, No, Yes, Yes

Yes =4

No =2

4 2 2 2
𝐺𝑖𝑛𝑖(𝑆𝑝𝑜𝑟𝑡𝑠) = 1 − ( ) −( )
6 6
= 1−
16
36

4
36
=
16
36
= 0. 444

SUV

Instances =4

No, Yes, No, No

Yes =1

No =3

1 2 3 2
𝐺𝑖𝑛𝑖(𝑆𝑈𝑉) = 1 − ( ) −( )
4 4
= 1−
1
16

9
16
=
6
16
= 0. 375

Weighted Gini(Type)
6 4
= 10
(0. 444) + 10
(0. 375) = 0. 266 + 0. 15 = 0. 416

Attribute 3: Origin

Domestic

Instances =5

Yes, No, Yes, No, No

Yes =2

No =3

𝐺𝑖𝑛𝑖(𝐷𝑜𝑚𝑒𝑠𝑡𝑖𝑐) = 0. 48

Imported

Instances =5

Yes, No, Yes, No, Yes

Yes =3

No =2

𝐺𝑖𝑛𝑖(𝐼𝑚𝑝𝑜𝑟𝑡𝑒𝑑) = 0. 48

Weighted Gini(Origin)
5 5
= 10
(0. 48) + 10
(0. 48) = 0. 48

Step 4: Compare Gini Values

Weighted
Attribute
Gini

Colour 0.480

Type 0.416

Origin 0.480

Root Node
The attribute having the minimum weighted Gini Index is selected.

𝑅𝑜𝑜𝑡 𝑁𝑜𝑑𝑒 = 𝑇𝑦𝑝𝑒

Step 5: Split Further

Branch 1: Type = SUV

Dataset:

Colour Origin Stolen

Yellow Imported No

Yellow Imported Yes

Yellow Domestic No

Red Imported No

Yes =1

No =3

Split on Colour

●​ Yellow → (No, Yes, No)

●​ Red → (No)

This gives the best separation.

Branch 2: Type = Sports

Dataset:

Colour Origin Stolen

Red Domestic Yes

Red Domestic No

Red Domestic Yes

Yellow Domestic No

Yellow Imported Yes

Red Imported Yes


Yes =4

No =2

Split further on Colour, then Origin if required.

Final Decision Tree

Type

/ \

Sports SUV

/ \

Colour Colour

/ \ / \

Red Yellow Red Yellow

| | | |

Mostly Yes Split No Split by Origin

Imported → Mixed

Domestic → No

Final Answer (10 Marks)

(i) Gini Calculations

Weighted
Attribute
Gini

Colour 0.480

Type 0.416

Origin 0.480

Since Type has the lowest Gini Index (0.416), it is selected as the root node.

(ii) Decision Tree

Type
/ \

Sports SUV

/ \

Colour Colour

/ \ / \

Red Yellow Red Yellow

| | | |

Mostly Yes Split No Split

Conclusion: Using the Gini Index criterion, Type is chosen as the root node
because it produces the lowest weighted Gini impurity (0.416), resulting in the best
initial split of the dataset.

You might also like