confusion
matrix
Bishal Patangia
Adjunct Professor
What is a confusion matrix?
It is a tool to evaluate
performance of supervised
machine learning algorithm
when used for classification
problem.
Consider....
You train a ML model for a data set
that is labelled only for YES or NO
for each data row.
Ex. Loan worthiness of a person
against their income, education,
credit score, marital status, number
of children and other features.
But remember....
Only for this problem our classes
are YES and NO. In general
classes can be any two different
things such as Positive or
Negative, Bus or Truck, Spam or
Not Spam, Dog or Cat and so on.
NOTE: Classes do not have to be
necessarily opposite to each other.
Now how do you check the performance
of your model on test data?
YES!
With the help of confusion matrix.
Let us assume your test data contained 100
rows. Then your confusion matrix might
look something like this...
Yes Predicted No
Yes
45 15
Actual
No 10 30
n = 100
Ok, so what
does these
numbers
mean?
Oh before that let us consider ....
YES to be positive label
No to be negative label
Yes Predicted No
Yes
45 15
Actual
No 10 30 TP
n = 100
True Positives
When model predicts YES
and actual value is also YES
Yes Predicted No
Yes
45 15
Actual
No 10 30 FP
n = 100
False Positives or Type 1 error
When model predicts YES
but actual value is NO
Yes Predicted No
Yes
45 15
Actual
No 10 30 TN
n = 100
True Negatives
When model predicts NO
and actual value is also NO
Yes Predicted No
Yes
45 15
Actual
No 10 30 FN
n = 100
False Negatives or Type II error
When model predicts NO
but actual value is YES
In summary...
Yes Predicted No
Yes
TP FN
Actual
No FP TN
n = total no. of predictions
No more confusion...
What was the prediction?
True Positive
True Negative
False Positive
False Negative
Was the prediction right?
Ok!
But how does this help
in evaluating model
performance?
Classification models can
be evaluated using ...
Accuracy Recall
Precision F1 Score
Area Under Curve (AUC)
Accuracy
Total number of correct/true
predictions among all the predictions.
accuracy = TP + TN
n
Recall
Measures classifiers' completeness.
Positive cases correctly predicted
among all actual positive cases.
recall = TP
TP + FN
Precision
Measures classifiers' exactness.
Positive cases correctly predicted
among all predicted positive cases.
precision = TP
TP + FP
F1 Score
Judging the model on accuracy, precision or recall
alone is flawed since most of the ML or data science
problems are context based.
Therefore, F1 Score is usually a safe metric to
evaluate a model. Model with highest F1 Score not
only returns true positive but also increases the
number of true positives.
precision x recall
F1 score = precision + recall
x2
AUC
Before coming to area under curve, it is
important we cover the following.....
Sensitivity Specificity
ROC curve
(Sensitivity vs 1-Specificity)
Sensitivity
It is the True Positive Rate or simply
recall.
Best value = 1
Worst value = 0
Specificity
Number of correct negative predictions
divided by the total number of negatives
= TN/(TN + FP)
It is also called true negative rate (TNR).
False Positive Rate = 1- Specificity
ROC Curve
1---
Sensitivity (TPR)
Random Classifier (Useless)
AUC
Slightly better classifier
More AUC
Better Model
|
0 1 - Specificity (FPR) 1
What if there are more
than two classes?
The Future of Work
using Predictive Models
in HR Analytics
Points to remember....
• Confusion matrix is a classifier evaluation tool.
• Classes need not be opposites.
• Onenegative class is labelled positive, other is
labelled. Just labelled, does not mean anything.
• Accuracy, recall and precision depends on the
context and may not be true judge of the model.
• F1 Score and AUC are more comprehensive
evaluation metrics.