Machine Learning (Material)
Machine Learning (Material)
Contents
1. Introduction to Machine Learning: .........................................................................................4
Key Parts of Machine Learning ..................................................................................................5
Types of Machine Learning ........................................................................................................5
Applications of Machine Learning .............................................................................................7
2. Machine Learning Classification: ...........................................................................................8
Characteristics of Classification ...............................................................................................8
Types of Classification ..............................................................................................................8
Classification Algorithms ..........................................................................................................9
Evaluation Metrics for Classification ....................................................................................... 10
Steps in Classification Process ............................................................................................... 11
Applications of Classification ................................................................................................. 11
3. Linear Regression ............................................................................................................... 12
Types of Linear Regression...................................................................................................... 12
4. Cost Functions ................................................................................................................... 13
Cost Function for Linear Regression ........................................................................................ 13
Why Mean Squared Error?....................................................................................................... 13
Shape of the Cost Function ..................................................................................................... 13
5. Gradient Descent................................................................................................................ 14
Intuition Behind Gradient Descent .......................................................................................... 14
Gradient Descent Algorithm Steps .......................................................................................... 14
Gradient Descent Update Rule................................................................................................ 14
Learning Rate ......................................................................................................................... 15
Types of Gradient Descent ...................................................................................................... 15
6. Gradient Descent for Classification ..................................................................................... 15
Using Linear Regression for Classification ............................................................................... 15
How It Works .......................................................................................................................... 15
Role of Gradient Descent ........................................................................................................ 15
Limitations ............................................................................................................................. 16
7. Gradient Descent Classification Evaluation: ........................................................................ 16
Why Evaluation is Important ................................................................................................... 16
Evaluation Methods ................................................................................................................ 16
1
MACHINE LEARNING
8. Bias .................................................................................................................................... 20
9. Variance ............................................................................................................................. 21
Characteristics of High Variance ............................................................................................. 21
10. Bias vs Variance (Comparison Table) .............................................................................. 22
11. Bias–Variance Tradeoff ................................................................................................... 22
12. Underfitting vs Overfitting ............................................................................................... 23
13. Bias vs Variance in Linear Regression .............................................................................. 23
14. How to Control Bias and Variance ................................................................................... 23
15. Clustering ...................................................................................................................... 24
Objectives of Clustering ......................................................................................................... 24
Types of Clustering Methods ................................................................................................... 24
Clustering relies on similarity measures such as: .................................................................... 26
16. K-Means Clustering ........................................................................................................ 28
17. K-Median Clustering ....................................................................................................... 29
18. Comparison Table: K-Means vs K-Median ........................................................................ 30
19. Elbow Method ................................................................................................................ 31
Concept ................................................................................................................................. 31
Steps to Apply the Elbow Method ............................................................................................ 31
Advantages ............................................................................................................................ 32
Disadvantages ....................................................................................................................... 32
Applications ........................................................................................................................... 32
20. Artificial Neural Methods ................................................................................................ 32
Key Components of an ANN .................................................................................................... 32
Popular Activation Functions .................................................................................................. 33
Architecture of ANN................................................................................................................ 33
Working Principle ................................................................................................................... 34
Learning Process .................................................................................................................... 34
Training Algorithm .................................................................................................................. 34
Advantages of ANN................................................................................................................. 34
Disadvantages of ANN ............................................................................................................ 35
Applications of ANN ............................................................................................................... 35
21. Feed Forward Neural Networks (FFNN) ........................................................................... 35
Architecture ........................................................................................................................... 35
2
MACHINE LEARNING
3
MACHINE LEARNING
A learning system constructs a model from data and uses it to make predictions or
decisions.
According to Tom Mitchell (1997): “A computer program is said to learn from experience E
with respect to some task T and some performance measure P, if its performance on task
T, as measured by P, improves with experience E.”
Task (T): What the computer is trying to do (e.g., classify emails as spam or not spam).
Experience (E): The data the computer learns from (e.g., past emails).
Performance (P): How we check if it’s learning well (e.g., accuracy of predictions).
4
MACHINE LEARNING
ii. Features: Features are measurable attributes or variables extracted from raw data.
Feature engineering is a critical step that involves selecting, transforming, and
creating features to enhance learning efficiency.
iv. Learning Algorithm: The learning algorithm optimizes the model parameters by
minimizing a loss or error function using techniques such as gradient descent or
optimization heuristics.
i. Supervised Learning:
5
MACHINE LEARNING
Supervised learning involves training a model on labeled data, where both input and
corresponding output are known.
Common tasks:
Algorithms:
Unsupervised learning deals with unlabeled data and aims to discover hidden patterns
or structures.
Common tasks:
• Clustering
• Dimensionality reduction
• Association rule mining
Algorithms:
• K-means Clustering
• Hierarchical Clustering
• DBSCAN
• Principal Component Analysis (PCA)
• Apriori Algorithm
Applications:
• Medical imaging
• Speech recognition
6
MACHINE LEARNING
Algorithms:
• Q-Learning,
• SARSA
• Deep Reinforcement Learning (DQN)
1.2. Advantages
7
MACHINE LEARNING
where xi ∈ R^m is a feature vector with m attributes and yi ∈ {C1, C2..., Ck} is a class
label.
• The task of classification is to learn a function f: R^m → {C1, C2..., Ck} that can
assign a correct class label to new unseen instances.
Characteristics of Classification
i. Supervised Learning:
Classification requires labeled data; each training example must have a known
output.
Types of Classification
i. Binary Classification:
Only two classes exist (e.g., spam vs. non-spam, disease vs. healthy).
8
MACHINE LEARNING
Classification Algorithms
i. Logistic Regression:
o Uses the sigmoid function for binary outcomes and softmax for multiple
classes.
o Probabilistic interpretation
o Finds the optimal hyperplane that maximizes the margin between classes.
9
MACHINE LEARNING
i. Accuracy:
iii. Confusion Matrix: A table summarizing the true vs. predicted labels.
iv. ROC Curve & AUC: Evaluate classifier performance for binary problems at various
thresholds.
10
MACHINE LEARNING
Applications of Classification
1. Medical Diagnosis: Classifying patients as diseased or healthy.
11
MACHINE LEARNING
3. Linear Regression
Linear Regression is a supervised learning algorithm used to predict a numerical value.
• Input (features) → 𝑥
• Output (target) → 𝑦
Example
𝑦 = 𝜃0 + 𝜃1 𝑥
X= Input feature
Y= predicted output
𝜃0 = intercept (bias)
𝜃1 = slop (weight)
𝑦 = 𝜃0 + 𝜃1 𝑥1 + 𝜃2 𝑥2 + ⋯ + 𝜃𝑛 𝑥𝑛
12
MACHINE LEARNING
4. Cost Functions
A cost function measures how well the model fits the data.
Where:
• 𝑦𝑖 → actual value
13
MACHINE LEARNING
5. Gradient Descent
Gradient Descent is an iterative optimization algorithm used to minimize the cost function
by updating model parameters gradually.
b. Compute predictions
c. Calculate cost
d. Compute gradient
e. Update parameters
Where:
• 𝛼→ learning rate
∂𝐽
•
∂𝜃
→ gradient (partial derivative)
14
MACHINE LEARNING
Learning Rate
• Small 𝛼→ slow convergence
How It Works
• Model outputs continuous values
𝟏 if 𝒚
̂ ≥ 𝟎. 𝟓
Class = {
𝟎 if 𝒚
̂ < 𝟎. 𝟓
15
MACHINE LEARNING
Limitations
• Output not bounded between 0 and 1
• Sensitive to outliers
Evaluation Methods
i. Confusion Matrix
16
MACHINE LEARNING
Positive (1) TP FN
Negative (0) FP TN
Where:
ii. Accuracy
Formula
𝑻𝑷 + 𝑻𝑵
Accuracy =
𝑻𝑷 + 𝑻𝑵 + 𝑭𝑷 + 𝑭𝑵
Interpretation
• Easy to understand
Limitation
• Example: predicting all samples as the majority class can still give high accuracy
iii. Precision
Formula
17
MACHINE LEARNING
𝑻𝑷
Precision =
𝑻𝑷 + 𝑭𝑷
Interpretation
• Spam detection
Formula
𝑻𝑷
Recall =
𝑻𝑷 + 𝑭𝑵
Interpretation
• Cancer detection
• Fraud detection
v. F1-Score
Formula
Precision × Recall
F1-score = 𝟐 ×
Precision + Recall
18
MACHINE LEARNING
Interpretation
Formula
𝑻𝑵
Specificity =
𝑻𝑵 + 𝑭𝑷
Interpretation
Formula
19
MACHINE LEARNING
• Precision
• Recall
• Accuracy
ROC Curve
8. Bias
Bias is the error caused by wrong or overly simple assumptions made by the model about
the data.
• Leads to underfitting
20
MACHINE LEARNING
Example
Effect
9. Variance
Variance is the error caused by a model being too sensitive to small changes in the training
data.
• Leads to overfitting
Example
Effect
21
MACHINE LEARNING
o ↓ Bias
o ↑ Variance
o ↑ Bias
o ↓ Variance
Error Decomposition
22
MACHINE LEARNING
o High bias
o Low variance
• Polynomial Regression
o Lower bias
o Higher variance
• Reduce regularization
To Reduce Variance
• Use regularization
15. Clustering
Clustering is an unsupervised learning technique used to group data objects into clusters
such that:
• Objects within the same cluster are more similar to each other
Clustering does not require labeled data, making it useful for discovering hidden patterns
in datasets.
Objectives of Clustering
• Identify natural groupings in data
Examples:
• K-means
• K-median
2. Hierarchical Clustering
24
MACHINE LEARNING
Types:
• Agglomerative (Bottom-Up):
• Divisive (Top-Down):
Output: Dendrogram
3. Density-Based Clustering
Example:
• DBSCAN
4. Grid-Based Clustering
Example:
• STING
5. Model-Based Clustering
25
MACHINE LEARNING
Example:
i. Euclidean Distance
Euclidean distance is the most commonly used distance measure that calculates the
straight-line distance between two data points in multidimensional space. It is widely used
for continuous numerical data.
Formula:
𝒅(𝑿, 𝒀) = √∑(𝒙𝒊 − 𝒚𝒊 )𝟐
Features:
• Sensitive to outliers
Applications:
K-means clustering, KNN, image processing
Manhattan distance, also known as City Block distance, measures distance as the sum of
absolute differences between corresponding attributes.
Formula:
𝒅(𝑿, 𝒀) = ∑ ∣ 𝒙𝒊 − 𝒚𝒊 ∣
26
MACHINE LEARNING
Features:
Applications:
K-median clustering, grid-based path problems
Cosine similarity measures the similarity between two vectors based on the angle between
them rather than physical distance. It is commonly used for high-dimensional and sparse
data.
Formula:
𝑿⋅𝒀
Cosine Similarity =
∥ 𝑿 ∥∥ 𝒀 ∥
Features:
• Ignores magnitude
Applications:
Text mining, document similarity, recommendation systems
Mahalanobis distance measures the distance between two points while considering
correlations among variables. It is scale-invariant and useful for multivariate data analysis.
Formula:
Features:
27
MACHINE LEARNING
Applications:
Anomaly detection, pattern recognition
Objective
Algorithm Steps
3. Assign each data point to the nearest centroid (using Euclidean distance).
Distance Measure
𝒅(𝑿, 𝒀) = √∑(𝒙𝒊 − 𝒚𝒊 )𝟐
Advantages
28
MACHINE LEARNING
Disadvantages
• Requires predefined K
Applications
• Market segmentation
• Document clustering
• Image compression
Objective
• Minimize the sum of absolute distances (Manhattan distance) between points and
cluster center.
𝒅(𝑿, 𝒀) = ∑ ∣ 𝒙𝒊 − 𝒚𝒊 ∣
Algorithm Steps
4. Update medians to minimize total distance (choose point that minimizes total
Manhattan distance).
29
MACHINE LEARNING
Distance Measure
Advantages
Disadvantages
Applications
30
MACHINE LEARNING
It helps avoid arbitrary selection of K and ensures that clusters are meaningful.
Concept
• Clustering algorithms minimize the within-cluster variation (also called Within-
Cluster Sum of Squares, WCSS).
• As K increases:
• The point where the rate of decrease sharply changes direction forms an “elbow”,
which indicates the optimal number of clusters.
𝑾𝑪𝑺𝑺 = ∑ ∑ ∣∣ 𝒙 − 𝝁𝒊 ∣∣𝟐
𝒙∈𝑪𝒊
𝒊=𝟏
Advantages
• Simple and intuitive
Disadvantages
• Elbow may not be clearly visible in some datasets
Applications
• Determining K in K-Means clustering
• Image segmentation
• Customer segmentation
• Market analysis
They consist of interconnected processing units called neurons that can learn patterns
from data. ANNs are widely used in supervised, unsupervised, and reinforcement learning.
32
MACHINE LEARNING
2. Weights
3. Bias
4. Activation Function
Architecture of ANN
1. Input Layer
2. Hidden Layers
33
MACHINE LEARNING
3. Output Layer
Working Principle
1. Each input is multiplied by a weight.
Learning Process
• Supervised Learning: ANN learns from labeled data.
Training Algorithm
• Forward Pass: Input → Hidden layers → Output
Advantages of ANN
• Can model complex, non-linear relationships
34
MACHINE LEARNING
Disadvantages of ANN
• Requires large training datasets
• Computationally expensive
Applications of ANN
• Pattern Recognition: Handwriting, speech, face recognition
• FFNN is mainly used for supervised learning tasks like classification and regression.
Architecture
1. Input Layer
2. Hidden Layer(s)
35
MACHINE LEARNING
3. Output Layer
Working Principle
1. Forward Pass:
2. Error Calculation:
o Compute the difference between predicted and actual outputs using a loss
function (e.g., MSE or cross-entropy).
Key Features
• Unidirectional flow: Data flows from input to output, no loops.
Advantages
• Simple and easy to implement
36
MACHINE LEARNING
• Can approximate any continuous function with enough hidden neurons (Universal
Approximation Theorem)
Disadvantages
• No memory of previous inputs → not suitable for sequential data
Applications
• Handwritten digit recognition (MNIST)
• Image classification
• Unlike Feed Forward Networks, RNNs have memory, meaning the output depends
on current input and previous inputs.
Architecture
1. Input Layer
o Contains neurons whose output is fed back to the same layer or next layer.
37
MACHINE LEARNING
3. Output Layer
Description
• Hidden layer has feedback connections to itself (Elman) or to input layer (Jordan)
Characteristics
Limitations
Description
• Uses memory cells and gates (input, forget, output) to control information flow
Components
Advantages
38
MACHINE LEARNING
Applications
Description
Advantages
Description
Applications
5. Training Methods
39
MACHINE LEARNING
Truncated BPTT
1. Feature Selection
40
MACHINE LEARNING
2. Wrapper Methods
o Computationally intensive.
3. Embedded Methods
Feature Extraction
• Transforms original features into a lower-dimensional space.
41
MACHINE LEARNING
• The tree consists of nodes, branches, and leaves, making it easy to interpret and
visualize.
o Represents the entire dataset and is split into sub-nodes based on a feature.
2. Internal/Decision Nodes
3. Leaf/Terminal Nodes
4. Branches
Working Principle
1. Select the best feature to split the data using a splitting criterion.
42
MACHINE LEARNING
4. Stop when:
Splitting Criteria
1. Information Gain (IG)
o Entropy formula:
2. Gini Index
𝑮𝒊𝒏𝒊 = 𝟏 − ∑𝒑𝟐𝒊
3. Chi-Square Test
2. Regression Tree
43
MACHINE LEARNING
Advantages
• Easy to understand and interpret
Disadvantages
• Prone to overfitting
Applications
• Medical diagnosis: Predict disease based on symptoms
• The closer two points are, the more similar they are considered.
44
MACHINE LEARNING
Key Concept
• Distance-based methods compute a distance metric between points.
3. Cosine Similarity – Measures angle between vectors (high for similar direction)
2. K-Means Clustering
45
MACHINE LEARNING
3. Hierarchical Clustering
5. Outlier Detection
• Points that are far from all other points based on distance are considered outliers
Advantages
• Simple and intuitive
Disadvantages
• Performance depends on distance metric choice
Applications
• Clustering: Customer segmentation, image segmentation
46
MACHINE LEARNING
47