Unit 6 - Machine Learning Algorithm XI AI Notes
Unit 6 - Machine Learning Algorithm XI AI Notes
Notes
1
specialized on training data, can lead to poor performance on new data. Bias in training
data can result in biased predictions, and some models are difficult to interpret, acting
as black boxes. Despite these challenges, ML transforms data into knowledge, enabling
computers to learn, adapt, and make decisions autonomously.
● Artificial intelligence (AI) and machine learning (ML) have significantly impacted
various aspects of our lives. From transportation and finance to healthcare and
entertainment, AI algorithms are pervasive. They power self-driving cars, fraud detection
systems, personalized shopping experiences, and virtual assistants like Siri and Alexa. As
technology continues to evolve, the influence of AI and ML is only expected to grow,
shaping the future of our society and culture.
Activity 1: Autodraw - Experience the power of machine learning with Autodraw! Autodraw
combines machine learning with the creativity of talented artists, allowing you to draw
things quickly and effortlessly. Visit the following link to play the game:
We introduced you to the fascinating world of artificial intelligence (AI) and its various
learning mechanisms. We discussed three main types of machine learning: supervised
learning, unsupervised learning, and reinforcement learning. These terms represent the
algorithms that drive AI systems, serving as the building blocks for programming
intelligent behavior and decision-making processes. Now, let us delve deeper into how these
algorithms shape the landscape of AI applications.
● Supervised learning involves the model learning from labeled data, where the input data is
accompanied by the correct output. The algorithm learns to map input data to output labels
based on example input-output pairs provided during training. The goal is to learn a mapping
function so that the model can make predictions on unseen data. Examples include linear
regression, logistic regression, decision trees, support vector machines, and neural networks.
● Unsupervised learning, on the other hand, deals with unlabelled data, where the
2
algorithm tries to find hidden patterns or structure without explicit guidance. The goal of this
is to explore and discover inherent structures or relationships within the
A. SUPERVISED LEARNING
Supervised learning stands out as one of the foundational types of Machine Learning. It is
a powerful approach that allows machines to learn from labeled data, making
predictions or decisions based on that learning. Within supervised learning, two primary
types of algorithms emerge:
1. Regression – works with continuous data
1. REGRESSION
Understanding Correlation: The Foundation of Regression Analysis
In data analysis, correlation is a fundamental concept that helps us grasp the
relationship between variables, laying the groundwork for predictive modeling and insightful
analysis. Correlation is a measure of the strength of a linear relationship between two
quantitative variables (e.g. price, sales). If the change in one variable appears to be
accompanied by a change in the other variable the two variables are said to be correlated
and this inter dependence is called correlation.
4
Types of Correlation:
Causation
Causation indicates that one event is the result of the occurrence of the other event.
Example: Since there is hot weather, the person will use more sunscreen or eat more ice
cream.
Sometimes, these two events may be correlated also. Example: smoking causes an increase in the
risk of developing lung cancer or it can
correlate with another like-smoking is correlated with alcoholism, but it does not cause
alcoholism. Therefore, we can say causation is not always correlation.
PEARSON’S R
Pearson's correlation coefficient (often denoted as Pearson's r) is one of the crucial
factors to consider when assessing the appropriateness of regression analysis. Pearson's r
measures the strength and direction of the linear relationship between two continuous
variables. In the context of regression analysis, a high degree of correlation between the
independent and dependent variables sug
5
techniques.
The requirements when considering the use of Pearson's correlation coefficient are:
6
1. Scale of measurement should be interval or ratio.
2. Variables should be approximately normally distributed.
3. The association should be linear.
4. There should be no outliers in the data.
Pearson’s r is calculated using the formula:
Solution:
For the Calculation of the Pearson Correlation Coefficient, we will first calculate the
following values:
7
Now the calculation of the Pearson R is as follows:
It is important to note that, regression analysis may not be suitable in certain situations:
1. No Correlation: If there is no correlation between the variables, meaning they change
independently of each other, regression analysis will not provide meaningful insights or
predictions.
2. Non-linear Relationships: While regression can model linear relationships well, it may not
capture more complex, non-linear relationships effectively. In such cases, alternative
techniques like polynomial regression or non-linear regression may be more appropriate.
3. Outliers: Outliers, or extreme data points, can disproportionately influence the
regression model and lead to inaccurate predictions. In the presence of outliers, it is
essential to assess their impact and consider alternative modeling approaches.
4. Violation of Assumptions: Regression analysis relies on certain assumptions, such as the
linearity of relationships and the absence of multicollinearity (high correlation between
predictor variables). If these assumptions are violated, the results of the regression
analysis may be unreliable.
REGRESSION
Regression is a statistical technique used to model the relationship between a
dependent variable and one or more independent variables. Its primary objective is to
understand and predict the value of the dependent variable based on the values of the
independent variables. In simpler terms, regression helps us understand how changes in one
or more variables are associated with changes in another variable.
8
Regression analysis is particularly useful when dealing with continuous data, where
variables can take on any value within a certain range. For example, variables such as height,
temperature, salary, and time are all continuous, meaning they can be measured along a
continuous scale. In regression, these continuous variables are used to predict or explain
the variability in another continuous variable, known as the dependent variable. By
analyzing the relationship between the independent and dependent variables, regression
allows us to make predictions and understand how changes in one variable may impact the
other. This makes regression a powerful tool for forecasting, prediction, and understanding
complex relationships in various fields such as economics, social sciences, and healthcare
When we make a distribution in which there is an involvement of more than one
variable, then such an analysis is called Regression Analysis. It generally focuses on
finding or rather predicting the value of the variable that is dependent on the other.
Let there be two variables x and y. If y depends on x, then the result comes in the form of a simple
regression. Furthermore, we name the variables x and y as:
y – Regression / Dependent / Explained Variable.
It is the variable we want to predict or understand.
9
coefficients allow for making predictions about the dependent variable based on the values of the
independent variable(s) with greater accuracy and reliability. As a result, this is widely used in
regression analysis.
Properties of the Regression line:
● The line minimizes the sum of squared difference between the observed values
(actual y-value) and the predicted value (ŷ value)
● The line passes through the mean of independent and dependent features.
Example 1
In the example of 6 people with different ages and different weight, let us draw the line of best
fit in Excel.
Solution:
Step 1: Select the Age and Weight.
Step 2: Insert a scatter chart and make changes to the following: Trendline Name: Linear,
check Display Equation on Chart X axis minimum: 20
Step 3: Let us verify the values of slope and intercept using slope() and intercept() function in
excel.
Step 4: Click on any cell and type =slope (Now, select the values of Weight, and then type
comma. Now, select the values of Age and press enter.
Step 5: Click on any cell and type =intercept (Now, select the values of Weight, and then type
comma. Now, select the values of Age and press enter.
10
Linear Regression
Linear regression is one of the most basic types of regression in machine learning. The
linear regression model consists of a predictor variable and a dependent variable related
linearly to each other. In case the data involves more than one independent variable, then linear
regression is called multiple linear regression models.
Linear regression is further divided into two types:
a) Simple Linear Regression: The dependent variable's
value is predicted using a single independent variable in
simple linear regression.
b) Multiple Linear Regression: In multiple linear regression,
more than one independent variable is used to predict the
value of the dependent variable.
Applications of Linear Regression:
● Market Analysis: Linear regression helps understand how different factors like pricing,
sales quantity, advertising, and social media engagement relate to each other in the market.
● Sales Forecasting: It predicts future sales by analyzing past sales data along with factors like
marketing spending, seasonal trends, and consumer behavior.
● Predicting Salary Based on Experience: Linear regression estimates a person's salary
based on their years of experience, education, and job role, aiding in recruitment and
compensation planning.
● Sports Analysis: Linear regression analyzes player and team performance by
considering statistics, game conditions, and opponent strength, assisting coaches and team
management in decision-making.
● Medical Research: Linear regression examines relationships between factors like age,
weight, and health outcomes, helping researchers identify risk factors and evaluate
interventions.
Advantages of Linear regression
● Simple technique and easy to implement
● Efficient to train the machine on this model
Disadvantages of Linear regression
1. Sensitivity to outliers, which can significantly impact the analysis.
2. Limited to linear relationships between variables. [Link]
regression-in-machine-learning
11
For Advanced Learners – Python program for Linear regression
Import scipy and draw the line of Linear Regression:
This program:
12
The expected output of the above program would be
REFERENCES
Video links:
● [Link]
● [Link]
● [Link]
● [Link]
2. CLASSIFICATION
For example, let us say, you live in a gated housing society and your society has
separate dustbins for different types of waste: paper waste, plastic waste, food waste and so on.
What you are basically doing over here is classifying the waste into different categories and
then labeling each category. In the picture given below, we are assigning the labels ‘paper’,
‘metal’, ‘plastic’, and so on to different types of waste.
13
Look at the two graphs below and suggest which graph represents the classification
problem.
Graph 1 Graph 2
Types of classification
The four main types of classification are:
1) Binary Classification
2) Multi-Class Classification
3) Multi-Label Classification
14
4) Imbalanced Classification
Classification Binary Multi-Class Multi-Label Imbalanced
Type Classification Classification Classification Classification
Classification Classification tasks
Classification tasks where with unequally
Classification
tasks with more each example distributed class
Description tasks with two
than two class may belong to labels, typically with
class labels.
labels. multiple class a majority and
labels. minority class.
• Email spam
• Face
detection -
classification
spam or not
• Plant species • Photo
• Conversion
classification classification
prediction -
• Optical - objects • Fraud detection
buy or not
character present in the • Outlier detection
Example • Medical test
recognition photo • Medical diagnostic
- Cancer
• Image (bicycle, tests
detected or not
classification apple, person,
• Exam
into thousands etc.)
results -
of classes
pass/fail
15
Steps involved in k-NN
● Select the number K of the neighbors
● Calculate the Euclidean distance of K number of neighbors
● Take the K nearest neighbors as per the calculated Euclidean distance.
● Among these k neighbors, count the number of the data points in each category.
● Assign the new data points to that category for which the number of the neighbor is
maximum.
● Our model is ready.
Applications of KNN:
● Image recognition and classification
● Recommendation systems
● Healthcare diagnostics
● Text mining and sentiment analysis
● Anomaly detection
Advantages of KNN:
● Easy to implement and understand.
● No explicit training phase; the model learns directly from the training data.
● Suitable for both classification and regression tasks.
● Robust to outliers and noisy data.
Limitations of KNN:
● Computationally expensive, especially for large datasets.
● Sensitivity to the choice of distance metric and the number of neighbors (K).
● Requires careful preprocessing and feature scaling.
● Not suitable for high-dimensional data due to the curse of dimensionality. For
advanced learners – Python Program for K Nearest Neighbor Algorithm # importing
libraries
import numpy as nm
import [Link] as mtp
import pandas as pd #importing datasets
data_set= pd.read_csv('user_data.csv')
Website: [Link]
REFERENCES
Video Session:
Classification: [Link]
KNN Algorithm: [Link]
B. UNSUPERVISED LEARNING
3. CLUSTERING
Clustering, or cluster analysis, is a machine learning technique used to group unlabeled
dataset into clusters or groups based on similarity. Clustering aims to organize data points
into groups where points within the same group are more similar to each other than to those in
other groups. It involves finding patterns or structures in the data without the need for
predefined class labels. It does it by finding some similar patterns in the unlabelled dataset
such as shape, size, color, behavior, etc., and divides them as per the presence and absence of
those similar patterns. It is an unsupervised learning method, hence no supervision is
provided to the algorithm, and it deals with the unlabeled dataset.
The clustering technique is commonly used for statistical data analysis.
example: Let us consider the clustering technique using a real-world example. Imagine you are
visiting a shopping center where items are grouped together based on their similarities. For
instance, in the fruits section, you will find apples, bananas, and grapes neatly arranged
together. This organization makes it convenient for shoppers to locate specific items they
are looking for.
Based on colour
Based on size
In a similar way, clustering algorithms group similar data points together based on common
characteristics or features. This approach helps in organizing and making sense of large
datasets in various tasks, such as market segmentation, image recognition, and customer
segmentation.
How Clustering Works:
To cluster data effectively, follow these key steps:
1) Prepare the Data: Select the right features for clustering and make sure the data is ready
by scaling or transforming it as needed.
2) Create Similarity Metrics: Define how similar data points are by comparing their
features. This similarity measure is crucial for clustering.
3) Run the Clustering Algorithm: Apply a clustering algorithm to group the data. Choose one
that works well with your dataset size and characteristics.
4) Interpret the Results: Analyze the clusters to understand what they represent. Since
clustering is unsupervised, interpretation is essential for assessing the quality of the
clusters.
Types of Clustering Methods
Some of the common clustering methods used in Machine learning are:
1) Partitioning Clustering
2) Density-Based Clustering
3) Distribution Model-Based Clustering
4) Hierarchical Clustering
1. Partitioning Clustering
It is a type of clustering that divides the data into
non- hierarchical groups. It is also known as the
centroid- based method. The most common
example of partitioning clustering is the K-
Means Clustering algorithm. In this type, the
dataset is divided into a set of k groups, where k is
used to define the number of pre-defined groups.
The cluster center is created in such a way that
the distance between the data points of one
cluster is minimum as compared to
another cluster centroid.
2. Density-Based Clustering
The density-based clustering method connects the
highly-dense areas into clusters, and the
arbitrarily shaped distributions are formed as long as
the dense region can be connected. This algorithm
does it by identifying different clusters in the
dataset and connects the areas of high
densities into clusters. The dense areas in data
space are divided from each other by sparser
areas. These algorithms can face difficulty in
clustering the data points if the dataset
has varying densities and high dimensions.
3. Distribution Model-Based Clustering
In the distribution model-based clustering
method, the data is divided based on the
probability of how a dataset belongs to a
particular distribution. The grouping is done
by assuming some distributions commonly
Gaussian Distribution.
The example of this type is the Expectation-
Maximization Clustering algorithm that uses
Gaussian Mixture Models (GMM).
4. Hierarchical Clustering
Hierarchical clustering can be used as an
alternative for the partitioned clustering as
there is no requirement of pre-specifying the
number of clusters to be created. In this
technique, the dataset is divided into clusters
to create a tree-like structure, which is also called
a dendrogram. The observations or any number of
clusters can be selected by cutting the tree at the
correct level. The most common example of
this method is the Agglomerative
Hierarchical algorithm.
K- Means clustering
K-Means Clustering is an unsupervised learning algorithm that is used to solve the
clustering problems in machine learning or data science. The k-means algorithm is one of
the most popular clustering algorithms. It classifies the dataset by dividing the samples into
different clusters of equal variances. The number of clusters must be specified in this
algorithm.
Steps involved K-Means Clustering:
The working of the K-Means algorithm is explained in the below steps:
● Select the number K to decide the number of clusters.
● Select random K points or centroids. (It can be other from the input dataset).
● Assign each data point to their closest centroid, which will form the predefined K
clusters.
● Calculate the variance and place a new centroid of each cluster.
● Repeat the third steps, which means reassign each datapoint to the new closest
centroid of each cluster.
● If any reassignment occurs, then go to step-4 else go to FINISH.
● The model is ready.
Activity: Visual AI: This tool allows you to visualize K-means clustering in real-time. Upload
your own data or use provided examples, adjust parameters, and see how clusters change
visually using the link Visualise k-means
Output
3. Which of the given plots is suitable for testing the linear relationship betweena dependent
and independent variable?
a. Bar chart
b. Scatter plot
c. Histograms
d. All of the above
4. Which of the following scatter plots represents a positive correlation?
a. points scattered randomly with no apparent trend
b. points forming a diagonal line and bottom left to top right
c. points forming a diagonal line from top left to bottom right
d. points clustered around a central point
5. Which regression technique is used when there is only one independentvariable?
a. logistic regression
b. multiple linear regression
c. simple linear regression
d. polynomial regression
6. What is one advantage of linear regression analysis?
a. it is robust to outliers
b. it can capture nonlinear relationships between variables
c. it is simple and easy to interpret
d. it is suitable for classification tasks
7. What is supervised learning in Artificial Intelligence?
a. training a computer algorithm on input data that is not labelled.
b. training a computer algorithm on input data that has been labelled fora specific output.
c. training a computer algorithm without any input data
d. training a computer algorithm to perform unsupervised tasks.
8. Which type of classification involves categorizing data into two distinct classes?
a. multi-class classification
b. binary classification
c. unsupervised classification
d. regression classification
9. What is logistic regression commonly used for in binary classification?
a. categorizing observations into multiple classes
b. predicting continuous values for input data
c. categorizing observations into two distinct classes
d. identifying unstructured data patterns
10. What is the primary goal of classification in AI?
a. categorizing data into random groups
b. locating and classifying things or concepts into predefined groups
c. predicting continuous values for input data
d. identifying unstructured data patterns
11. Which algorithm is commonly used for binary classification?
a. Decision trees
b. Support Vector Machine
c. Logistic Regression
d. k-Nearest Neighbors
12. The K-Nearest Neighbors (KNN) algorithm assigns a class to new data point byconsidering:
a. Distance from the data point to a predefined decision boundary
b. Majority vote of its K nearest neighbors in the training data
c. Similarity of the data point to a cluster centroid
d. probability of each class given the data point’s features.
13. What does a classification model in AI ultimately want to achieve?
a. to identify patterns and associations in data
b. to predict continuous numerical values
c. to categorize input data into predefined classes or labels
d. to optimize decision-making processes
14. What are some challenges in applying classification models to real-worldproblems?
a. Data bias and fairness
b. Interpretability and explainability
c. overfitting and underfitting
d. All of the [Link] is clustering?
a. Grouping labeled dataset
b. Dividing data into different clusters
c. Finding linear association between variables
d. Predicting future behaviors of a dependent variable
16. Which type of learning does clustering belong to?
a. Supervised learning
b. Unsupervised learning
c. Semi-supervised learning
d. Reinforcement learning
17. Which method is used to group highly dense areas into clusters?
a. Partitioning clustering
b. Density-based clustering
c. Distribution model-based clustering
d. Hierarchical clustering
18. Which algorithm is an example of partitioning clustering?
a. Mean-shift algorithm
b. DBSCAN algorithm
c. K-Means algorithm
d. Fuzzy clustering algorithm
19. Which clustering method allows data objects to belong to more than one group orcluster?
a. Partitioning clustering
b. Density-based clustering
c. Distribution model-based clustering
d. Fuzzy clustering
20. Which clustering algorithm is sensitive to outliers?
a. K-Means algorithm
b. Mean-shift algorithm
c. DBSCAN algorithm
d. Hierarchical clustering
C. True or False:
1. Clustering is a supervised learning technique.
2. Hierarchical Clustering requires pre-specifying the number of clusters.
3. Fuzzy clustering is a hard clustering method.
4. Classification is an unsupervised learning technique.
5. In k-NN algorithm, k is the number of nearest data points.
6. K-Means algorithm requires specifying the number of clusters.
ANSWERS
A. Multiple Choice Questions
1. a. All of the above
2. c. Regression
3. b. Scatter plot
4. b. points forming a diagonal line and bottom left to top right
5. c. simple linear regression
6. c. it is simple and easy to interpret
7. b. training a computer algorithm on input data that has been labelled for a
specific
output.
8. b. binary classification
9. c. categorizing observations into two distinct classes
10. b. locating and classifying things or concepts into predefined groups
11. c. Logistic Regression
12. b. Majority vote of its K nearest neighbors in the training data
13. c. to categorize input data into predefined classes or label
14. d. All of the above
15. b. Dividing data into different clusters
16. b. Unsupervised learning
17. c. Distribution model-based clustering
18. c. K-Means algorithm
19. d. Fuzzy clustering algorithm
20. a. K-Means algorithm
B. Fill in the blanks
1. Unsupervised Learning
2. Correlation coefficient
3. Linear Regression
4. Outlier
5. K-nearest neighbors (KNN) algorithm
6. unlabelled dataset
7. centroid-based method
8. low point density
9. Specified
10. Data Analysis
C. True or False:
1. False 2. False 3. False 4. False 5. True 6. True