ML Unit2 LectureNotes
ML Unit2 LectureNotes
Contents
1 Introduction to Proximity Measures 3
1.1 What are Proximity Measures? . . . . . . . . . . . . . . . . . . . . . . . 3
1.2 Why Proximity Measures Matter . . . . . . . . . . . . . . . . . . . . . . 3
2 Distance Measures 3
2.1 Properties of a Metric . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
2.2 Minkowski Distance . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4
2.3 Special Cases of Minkowski Distance . . . . . . . . . . . . . . . . . . . . 4
2.3.1 Manhattan Distance (L1 Norm, p = 1) . . . . . . . . . . . . . . . 4
2.3.2 Euclidean Distance (L2 Norm, p = 2) . . . . . . . . . . . . . . . . 5
2.3.3 Chebyshev Distance (L∞ Norm, p → ∞) . . . . . . . . . . . . . . 5
2.4 Comparison of Distance Measures . . . . . . . . . . . . . . . . . . . . . . 5
2.5 Weighted Minkowski Distance . . . . . . . . . . . . . . . . . . . . . . . . 6
2.6 Mahalanobis Distance . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
1
Machine Learning - Unit II 23CSM104
8 KNN Regression 14
8.1 Basic KNN Regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14
8.2 Weighted KNN Regression . . . . . . . . . . . . . . . . . . . . . . . . . . 14
9 Performance of Classifiers 15
9.1 Evaluation Metrics for Classification . . . . . . . . . . . . . . . . . . . . 15
9.1.1 Confusion Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
9.1.2 Key Metrics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
9.1.3 Fβ Score . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
9.2 ROC Curve . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
9.3 Multi-class Evaluation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16
14 Summary 20
2
Machine Learning - Unit II 23CSM104
Definition 1.2 (Similarity). Similarity is a numerical measure of how alike two data
objects are. Higher values indicate greater similarity. Typically ranges from 0 (no simi-
larity) to 1 (identical).
Key Point
Similarity and dissimilarity are complementary concepts:
Proximity Measures
Information Recommendation
Retrieval Systems
2 Distance Measures
Distance measures quantify the dissimilarity between data points in a feature space.
1. Non-negativity: d(x, y) ≥ 0
3
Machine Learning - Unit II 23CSM104
2. Identity: d(x, y) = 0 ⇔ x = y
d(x, z) d(y, z)
x y
d(x, y)
Triangle Inequality: d(x, z) ≤ d(x, y) + d(y, z)
d
!1/p
X
dp (x, y) = |xi − yi |p (1)
i=1
where p ≥ 1 is a parameter.
x2
y(4, 3)
Euclidean
x(1, 1) Manhattan
x1
4
Machine Learning - Unit II 23CSM104
d1 (x, y) = |1 − 4| + |1 − 3|
=3+2=5
This is the most commonly used distance measure, representing the straight-line distance.
5
Machine Learning - Unit II 23CSM104
x2
p=1
p=2
x1
p=∞
Euclidean
Mahalanobis
6
Machine Learning - Unit II 23CSM104
x2
y
x
θ
x1
cos(θ) = similarity
Important
Cosine similarity:
7
Machine Learning - Unit II 23CSM104
Key Point
Pearson correlation is equivalent to cosine similarity of mean-centered data.
|A ∩ B|
J(A, B) = (10)
|A ∪ B|
Jaccard Similarity
A ∩ B = {2, 3}, |A ∩ B| = 2
A ∪ B = {1, 2, 3, 4, 5, 6}, |A ∪ B| = 6
2
J(A, B) = = 0.333
6
8
Machine Learning - Unit II 23CSM104
y
1 0 Total
1 f11 f10 f1+
x
0 f01 f00 f0+
Total f+1 f+0 d
where:
SMC vs Jaccard
For binary vectors x = (1, 0, 0, 0, 1, 0, 0, 1) and y = (1, 1, 0, 0, 1, 0, 1, 0):
Position 1 2 3 4 5 6 7 8
x 1 0 0 0 1 0 0 1
y 1 1 0 0 1 0 1 0
• f11 = 2 (positions 1, 5)
• f00 = 2 (positions 3, 6)
• f10 = 1 (position 8)
Let me recalculate:
9
Machine Learning - Unit II 23CSM104
• f11 = 2 (positions 1, 5)
• f00 = 3 (positions 3, 4, 6)
• f10 = 1 (position 8)
• f01 = 2 (positions 2, 7)
2+3 5
SMC = = = 0.625
8 8
2 2
J= = = 0.4
2+1+2 5
Jaccard f11
f11 +f10 +f01
Russel-Rao f11
d
Key Point
Characteristics of Instance-Based Learning:
10
Machine Learning - Unit II 23CSM104
K-Nearest Weighted
Neighbor Radius-Based KNN
NN
K=3 Class A
Class B
Query
11
Machine Learning - Unit II 23CSM104
6.3 Choice of K
Important
The choice of K significantly affects KNN performance:
• Large K:
K=1 K=5 K = 15
R∗
∗ ∗
R ≤ R1-NN ≤ 2R 1 − (15)
C
12
Machine Learning - Unit II 23CSM104
Key Point
Sparse: 0 neighbors
Dense: 3 neighbors
13
Machine Learning - Unit II 23CSM104
8 KNN Regression
Definition 8.1 (KNN Regression). KNN can be extended to regression by predicting the
average (or weighted average) of the target values of the K nearest neighbors.
Given training data: (1, 3), (2, 5), (3, 7), (5, 8), (6, 10)
Query: xq = 4, K = 3
Distances from xq = 4:
• d(4, 1) = 3, y = 3
• d(4, 2) = 2, y = 5
• d(4, 3) = 1, y = 7 ✓
• d(4, 5) = 1, y = 8 ✓
• d(4, 6) = 2, y = 10 ✓
14
Machine Learning - Unit II 23CSM104
9 Performance of Classifiers
9.1 Evaluation Metrics for Classification
9.1.1 Confusion Matrix
Predicted
Positive Negative
Positive TP FN
Actual
Negative FP TN
TP + TN
Accuracy = (18)
TP + TN + FP + FN
FP + FN
Error Rate = 1 − Accuracy = (19)
TP + TN + FP + FN
TP
Precision = (Positive Predictive Value) (20)
TP + FP
TP
Recall = (Sensitivity, True Positive Rate) (21)
TP + FN
TN
Specificity = (True Negative Rate) (22)
TN + FP
2 × Precision × Recall
F1-Score = (23)
Precision + Recall
9.1.3 Fβ Score
Definition 9.1 (Fβ Score).
Precision × Recall
Fβ = (1 + β 2 ) · (24)
β2 · Precision + Recall
15
Machine Learning - Unit II 23CSM104
TPR
Perfect
Good
Random
FPR
Definition 9.3 (AUC (Area Under ROC Curve)). AUC measures the entire two-dimensional
area underneath the ROC curve.
16
Machine Learning - Unit II 23CSM104
• R2 = 1: Perfect prediction
17
Machine Learning - Unit II 23CSM104
18
Machine Learning - Unit II 23CSM104
Important
As dimensionality increases, the concept of “nearest neighbor” becomes less mean-
ingful because:
• All points become approximately equidistant
2. Choosing K:
19
Machine Learning - Unit II 23CSM104
√
• Start with K = n
• Use cross-validation to tune
• Use odd K for binary classification
14 Summary
Key Takeaways - Unit II
1. Proximity Measures:
2. Distance Measures:
3. Non-Metric Measures:
5. KNN Algorithm:
20
Machine Learning - Unit II 23CSM104
6. Key Considerations:
References
1. Murthy, M. N., & Ananthanarayana, V. S. (2024). Machine Learning Theory and
Practice. Universities Press (India).
3. Tan, P. N., Steinbach, M., & Kumar, V. (2019). Introduction to Data Mining (7th
ed.).
4. Cover, T., & Hart, P. (1967). Nearest neighbor pattern classification. IEEE Trans-
actions on Information Theory, 13(1), 21-27.
21