Skymachine Learning
Skymachine Learning
Authored by
Srikanth S
Lalitha Y Divya S R
[Link], MCA [Link]., [Link]
Assistant Professor Assistant Professor
Department of Computer Science DeparU uent of Computer Science
Vijaya College, Jayanagar V ijaya College,R V Road
Bengaluru Bengaluru
Skyward Publishers
# 157, 7th Cross, 3rd Main Road, Chamarajpet,
Bengaluru-18. Phone ; 080-43706620 / 080-26603535
Mob: 9611185999
E-mail: [Link]@[Link]
Website: [Link]
A Text Book of “Machine Learning” - by Srikanth S, Lalitha Y, Divya S R, Mrs. Roopa H R, &
Dr. Kadli Nanjundeshwara as per the New NEP Syllabus for VI Semester BCA, Bengaluru City University &
Bangalore University.
© Authors
Copy Right;No part o f this book may be reproduced, stored in a retrieval system, or transmitted, in any
form or by any means, without the previous permission o f the copyright holders. Every effort has been
made to avoid errors or omissions in this publication. In spite of this, some errors might have crept in.
Any mistake, error or discrepancy noted may be brought to our notice which shall be taken care in the
next edition. The publisher shall not verify the originality, authenticity, ownership, non-infringement
o f the data, content, and information. The Authors are the sole owners o f the copyrights o f the Work It
shall be Authors sole responsibility to ensure the lawfulness o f the content and publisher is not responsi
ble for any copyright issues. It is notified that publisher will not be responsiblefo r any damage or loss o f
action anyone, o f any kind, in any manner, there from all disputes are subject to Bengaluru jurisdiction
only.
Disclaimer: S l^ a r d Publishers has exercised due care and caution in collecting all the data before
publishing the book In spite o f this, if any omission, inaccuracy or printing error occurs with regards
to the data contained in this book, Styward Publishers will not be held responsible or liable. S i^ ard
Publishers will be grateful fo r your suggestions which will be o f great help for other readers.
ISBN: 978-93-95085-78-6
P r i c e 275/ -
Published by:
Skyward Publishers
#157,7th Cross, 3rd Main Road, Chamarajpet
Bangalore-18. Phone: 080-26603535/43706620,
Mob: 9611185999
E-mail: [Link]@[Link]
Website: [Link]
DTP By
Nirmala & Mary, Skyward Team
PREFACE
Welcome to the world of Machine Learning! This book has been meticulously crafted to serve
as a comprehensive guide for students pursuing the 6 * Semester BCA course at Bengaluru City
University and Bangalore University, in alignm ent with the latest National Education Policy (NEP)
syllabus. Machine Learning has emerged a s a transformative field th at empowers computers
to learn from data and make intelligent decisions without being explicitly programmed.
Understanding the fundamentals of M achine Learning is essential in today's data-driven world,
and this book aims to equip you with the knowledge and skills necessary to excel in this domain.
- Authors
IV
SYLLABUS
vi
1.17.1 Jupyter Notebook 1-45
2.4.2 E xam p le: Implementing ML in Real E state for House Price Prediction 2.12
2.6.1 Why Visualizing the Data is Needed During Data Preparation? 2.22
VII
2.7.5 Data Spitting 2.46
VIII
3 .6 .4 Applications of Naive Bayes Classifiers 3.47
4.1 Introduction
4.3 Clustering
IX
4.5.3 K-Means Clustering for Image Segm entation 4.20
FUNDAMENTALS OF
':fSj'^ ?VV
: ‘It
r
Contents
■ ■ ■ in tr o d u c tio n t o M a c h i n e L e a rn in g
The rise of "big data" has changed how we use technology. With more personal com puters and
w ireless devices around, we now create and use lots of data. Every time we do som ething online, like
shopping or surfing the web, we create important data that can be used to customize products and
services based on what we like.
Think about a big superm arket th a t tracks a lot of sales data every day By looking at this data, the store
tries to guess what custom ers like to buy. Customers also w ant products that suit them . Customers
shopping habits aren't com pletely random; they change over time and place, but understanding
patterns can predict what custom ers might want.
Algorithms play a pivotal role in addressing computational challenges. While algorithm s are readily
available for tasks like sorting and seraching complexities arise when no algorithm exists for tasks
such as predicting custom er behavior or detecting spam emails. That's where m achine learning
com es in by teaching com puters o r machines to find patterns in data and make predictions without
specific instructions.
Machine Learning is like teaching a computer to learn and make decisions on its own by showing it
examples and patterns, ju st like how we learn from our experiences Machine learning likes to look at
big data to find useful inform ation or make predictions. Industries such as banking, manufacturing,
healthcare, and telecom use m achine learning to catch fraud, improve processes, diagnose illnesses,
and manage networks.
Machine Learning is primarily a concept and a field of study
within the broader domain of artificial in te llig e n c e . Machine
learning is about extracting knowledge from data. It is a research
field at the intersection of stadstics, artificial intelligence, and
computer science and is also known as p re d ictiv e a n a ljtic s
or statistical le arn in g . Machine Learning involves developing
algorithms and models that enable computers to learn from
data and make predictions or decisions without bein g explicitly
programmed for each task.
Machine learning methods are now widely used in various aspects of daily life. They are utilized in
suggesting movies, recom m ending food or products, identifying individuals in photos for security or
tagging purposes, optimizing search engine results for relevance, personalizing social media feeds,
predicting user behavior for targeted advertising, enhancing virtual assistants like Siri or Alexa,
improving healthcare diagnostics through image analysis, and automating fraud detection in financial
transactions on modern w ebsites and [Link] such as Facebook, Amazon, and Netfiix uses
multiple machine learning m odels across various sections o f their websites.
m m y What is Learning?
Learning is the process of acquiring knowledge or skills through study, experience, o r teaching. It
involves the ability to understand, retain, and apply new information or behaviors. It is a fundamental
human activity, essential for personal growth and development, adaptation to changes, and problem
solving in daily life.
Fundamentals o f INachine Learning ^ 1.3
The term "Intelligence" is defined b y the words knowledge, skill, and memory. Human begin learning
by memorizing. After some tim e, h e realizes that the sim ple ability to remember som ething is not
intelligence. Then he practices converting the data stored in his memory into knowledge and applies
it to develop problem-solving abilities in the real world.
Examples Learning
Academic Learning: A student learns about the laws of physics in a classroom setting through lectures,
textbooks, and experiments.
Skill Acquisition: An individual learns to play the guitar through practice, observing others, and
perhaps taking music lessons.
Social Learning: A child learns social norms and behaviors by observing and imitating their parents
or peers.
Professional Development: An employee learns new software or a new system at work through
training sessions and by using the software in practical tasks.
Adaptive Learning: An individual learns to adapt to a new culture and language by living in a foreign
country, interacting with locals, and experiencing the culture firsthand. ______________________
U n d e r s ta n d in g H u m a n L e a r n in g
Human learning is something w e do naturally as we go through life. Human le a r n in g m eans
g ain in g new knowledge, s k ills , an d behaviors by e x p e rie n cin g things, w atching o th e rs , and
b e in g taught. Unlike machines, humans can think, reason, and adjust what they know to different
situations. Our brains are am azing organs that help us take in information through our senses, make
links between things, and rem em ber them. Human learning isn't just about one area bu t includes
many mental skills like solving problems, thinking creatively, and understanding em otions.
U n d e r s ta n d in g M a c h in e L e a r n in g
Machine learning is the branch o f Artificial Intelligence th at focuses on developing models and
algorithms that let com puters learn from data and improve from previous experience without
being explicitly programmed fo r every task. In simple words, ML teaches the system s to think and
understand like humans by learning fi'om the data.
Unlike traditional programming, w here instructions are explicitly given to the computer, machine
learning algorithms have the ability to learn and improve from pattern s in data. These algorithm s are
trained using labelled datasets, w here the desired output is know n, and then applied to new, unseen
data to make predictions or decisions.
One o f the key strengths of m achine learning lies in its ability to process and analyse vast amounts
of data quickly and accurately This has led to ground breaking advancements in various fields, such
as image and speech recognition, natural language processing, and autonomous vehicles. Machine
learning algorithms can detect com plex patterns that may n o t be obvious to humans, leading to
precise predictions and improved decision-making processes.
I l l j l l l l G o a ls o f M a c h in e L e a r n i i ^
■ K a H is to ry o f M a c h in e L e a r n i n g
Before som e years (about 4 0 -5 0 years), machine learning was scien ce fiction, but today it is the part of
our daily life. Machine learning is making our day to day life easy from self-driving c a rs to Amazon
v irtu a l a ssista n t "Alexa". However, the idea behind machine learning is so old and has a long history.
Machine learning has a rich history that dates back to the m id -20th century Let us discuss the key
m ilestones and developments in the history of machine learning:
1. 1 9 5 0 s - 19 6 0 s: The foundation of machine learning was laid during this period with the work
o f pioneers like Alan Turing, who proposed the concept o f a "learning machine" in his paper
"Computing Machinery and Intelligence" in 1950.
• Arthur Samuel, a pioneer in the field of machine learning, created a program in 1952
that enabled an IBM com puter to play checkers at a high level of performance.
• A rthur Sam uel coined the term "Machine L e a rn in g " in 1959.
2 . 1 9 7 0 s - 1 9 8 0 s: This era saw the emergence o f sym bolic A1 and expert system s, where
knowledge was represented explicitly in the form o f rules. However, limitations in handling
uncertainty and complexity led to a shift towards m ore data-driven approaches.
f undamenfals of Machine Learning
3 . 1 9 9 0 s : The 1990s marked the resurgence of interest in neural networks with the development
o f backpropagation algorithms for training deep neural networks. Support vector m achines
(SVMs) also gained popularity as powerful tools for classification tasks.
4 . 2 0 0 0 s : Grouping methods like random forests and gradient boosting became popular for
com bining multiple models to improve predictions. M achine learning started being used in
various real-world applications.T his period also saw the increasing use of data mining and
m achine learning in various applications, including recommendation systems and natural
language processing.
5. 2 0 1 0 s - P resent: The past decade has been characterized by the widespread adoption of deep
learning due to advancement o f computational power and the availability of large datasets.
Deep learning models, particularly convolutional neural networks (CNNs) and recurrent
neural networks (RNNs), have achieved remarkable success in tasks such as image recognition,
sp eech recognition, and language translation.
6. F u tu re Trends: The future o f machine learning is likely to be shaped by developments in areas
such as reinforcement learning, generative adversarial netw orks (GANs), and explainable AI.
T h ere will also be a focus on ethics, fairness, and making Al systems easier to understand and
tru st.
W h a t is M ach in e L e a r n i n g ?
Machine learning is a subfield of artificial intelligence, which is broadly defined as the capability o f a
m achine to imitate intelligent human behaviour. Artificial intelligence systems are used to perform
com plex tasks in a way that is sim ilar to how humans solve problem s.
In the real world, we are surrounded by humans who can leam everything from their experiences
with th e ir learning capability, and we have computers or m achines which work on our instructions.
But can a m achine also learn fi-om experiences or past data like a human does? So here com es the
role o f M ach in e Learning.
lcanlearneverytWngV
Human
C automatkafly from j
eqperiences.
Canuleam?
j Machine
M achine learning is programming computers to optimize a perform ance criterion using exam ple data
or past experience. We have a model defined up to some param eters, and learning is the execution
of a computer program to optim ize the parameters of the m odel using the training d ata or past
experience. The model may be predictive to make predictions in the future, or descriptive to gain
knowledge from data, or both.
Machine learning (ML) is a branch of artificial intelligence (AI) that enables machines to
automatically learn from data and previous experiences in order to identify patterns and make
predictions with minimal human intervention.
The goal of machine learning is to develop systems that can automatically learn and adapt to new
information and tasks, ultimately improving their performance over time._____________
Machine learning algorithms m ake a model using past data to help predict or decide things without
direct instructions. By using historical data, these algorithms com bine statistics and com puter science
to create predictive models. The more data we give them, the better they perform. If a m achine gets
more data, it can learn and improve its predictions.
4. Self-Driving Cars: Self-driving cars use machine learning algorithms to analyze data from sensors and
cameras to navigate roads, detect obstacles, and make driving decisions. The system learns from real-
world driving experiences to improve safety and efficiency.
5. Fraud Detection: Banks and credit card companies use machine learning to detect fraudulent
transactions by analyzing patterns in spending behavior. The system learns to identify unusual
activities and flag them for further investigation.
6. Personalized Product Recommendations: E-commerce platforms like Amazon, flipkart and eBay
utilize machine learning to suggest products based on browsing history and purchase behavior. These
systems learn from user interactions to recommend items that align with user preferences,
7. Dsmamic Pricing Strategies: Online marketplaces like eBay, amazon and Airbnb use machine learning
to adjust prices in real-time based on factors such as demand, competition, and customer behavior. By
analyzing vast amounts of data, these systems can set optimal prices to maximize revenue and attract
customers, leading to a more competitive pricing strategy in the e-commerce industry.
In the competitive landscape of e-commerce, personalized recommendations play a crucial role in enhancing
user experience and driving customer engagement. Companies like Amazon leverage machine learning
algorithms to analyze user behavior, understand preferences, and deliver tailored product suggestions. Let's
explore how Amazon utilizes machine learning in the context of a user browsing Dell laptops on their platform.
Step 1: U ser Browsing Dell Laptop on Amazon
A user visits the Amazon website and explores a variety of Dell laptops, comparing models,
specifications, and prices.
Step 2: Tracking User Behavior
Amazon's system tracks the user's browsing activity, capturing details of the Dell laptop models
viewed, time spent on each product page, and interactions like adding items to the cart or wish list.
Step 3: Data Collection and Analysis
Machine learning algorithms analyze the collected data, processing the user's interactions with Dell
laptops to identify preferences and patterns.
Step 4: Building User Profile Based on the analyzed data
Amazon creates a user profile that includes preferences for Dell laptops, budget constraints, desired
features (e.g., screen size, processor speed), and relevant past purchase history.
Step 5: Recommendation Generation
Utilizing the user profile and machine learning models, Amazon's recommendation system generates
personalized suggestions, recommending other Dell laptops or related accessories that align with the
user's preferences.
Step 6: Dlsplajing Recommendations
Upon the user's return to the Amazon platform, personalized recommendations are showcased on the
______ homepage or product pages, displaying relevant Dell laptops based on the user's previous interactions.
Machine Learning
G X Machii
The online streaming platforms like Netflix or Amazon Prime Vidoe, the utilization of machine learning
algorithms has revolutionized the way users discover and engage with content By analyzing user behavior and
preferences, platforms can offer personalized recommendations, enhancing the overall viewing experience.
Let's delve into a case study that illustrates how Netflix leverages machine learning to understand user
preferences and optimize content recommendations.
Step 1: User Watching Movies on Netflix
A user accesses their Netflbc account and begins watching movies across various genres, including
action, comedy, and drama.
Step 2: Tracking User Viewing Behavior
Netflix's system tracks the user's viewing patterns, capturing details such as the genres of movies
watched, viewing session durations, and interactions like adding movies to the watchlist or providing
ratings.
Step 3: Data Collection and Analysis
Machine learning algorithms analyze the collected data on user viewing habits to discern preferences
and viewing trends.
Step 4: Building User Profile Based on the analyzed data
Netflix constructs a user profile that encompasses preferred genres, favorite actors or directors,
viewing habits (e.g., binge-watching), and movie ratings provided by the user.
Step 5: Recommendation Generation
Leveraging the user profile and machine learning models, Netflix's recommendation system generates
personalized suggestions, recommending movies or TV shows in similar genres or featuring preferred
actors to align with the user's tastes.
Step 6: Displaying Recommendations
When the user explores the Netflix library, personalized recommendations are showcased on the
homepage or category pages, presenting relevant content based on the user's viewing history and
preferences.________ _________ _____ ___________ ______________ _______ _______________ _____
;T5i ' , ' , , , --'‘.y i^ ^ fF u n d a m e n ta ls o f Maehin^Uarning
F e a t u r e s o f M a c h in e L e a r n i n g
Machine learning (ML) is a subset of artificial intelligence that focuses on building systems that
learti from data, identify patterns, and make decisions with minimal hum an intervention. Some key
features of m achine learning along with exam ples are listed below:
1. A d a p ta b ility : Machine learning m odels can adapt to new data and changing environments,
making th e m versatile for various applications.
E x a m p le : In autonomous vehicles, m achine learning algorithms continuously learn from real
time sen so r d ata to adapt to different driving conditions and improve decision-making.
2. A u to m a tio n : Machine learning enables automation of tasks by allow ing systems to learn from
data and m ak e decisions without explicit programming instructions.
Exam ple : In email spam detection, machine learning algorithm s can automatically classify
incoming em ails as spam or non-spam based on patterns in the content.
3. S c a la b ility : Machine learning algorithm s are designed to handle large volumes of data and
scale w ith it. As more data becom es available, machine learning models can update their
predictions and decisions, often improving in accuracy.
Exam ple : In social media platforms, machine learning algorithm s analyze vast amounts of
user-gen erated content to personalize feeds and recommendations for millions of users.
4. P e rso n a liz a tio n : Machine learning enables personalized experiences by tailoring
recom m endations and content based on individual preferences.
Exam ple: In streaming services like Spotify, machine learning algorithm s analyze user listening
habits to c re a te personalized playlists and recommendations.
5. P re d ictiv e A n a l)^ c s : Machine learning excels in making predictions and forecasts based on
historical d ata.
Exam p le : Financial institutions use m achine learning to predict stock market trends, assess
loan risks, an d detect fraudulent transactions based on historical data. In weather forecasting,
machine learn in g models analyze past weather patterns and cu rren t atmospheric conditions
to pred ict fu tu re weather outcomes with improved accuracy.
6. C o n tin u o u s Im p ro v em en t: As m achine learning algorithms are exposed to new data, they
are able to independently improve th eir performance.
......
/■;*^’■V ' ^ ' ? - V % ' ‘ ‘ . . : ;I'3k-^V
B ia m p le ; The voice assistan t technologies like Siri and Alexa, which becom e m o re accurate in
understanding and processing user requests over time.
7. Decision M a k in g : ML can assist in making decisions with minimal human intervention.
Example : In h ealthcare, machine learning models help in diagnosing diseases and
recommending tre atm e n t plans based on patient data and trends found in h isto rical data.
8. Pattern R e c o g n itio n : ML excels at recognizing patterns and regularities in d ata.
Example : Facial recognition technology uses machine learning to id en tify and venfy
individuals based on digital images of their faces.
9. Efficiency: Machine learning can enhance efficiency by automating rep etitiv e tasks and
optimizing processes.
Example : In manufacturing, machine learning algorithms analyze production data to predict
equipment failures and schedule maintenance proactively, reducing dow ntim e and improving
productivity. _______________________
m y T ra d itio n a l P r o g r a m m i n g A p p ro a c h V s M a c h in e L e a r n in g A p p r o a c h
In software development, th ere are two main approaches to solving tasks re la te d to pattern
recognition or decision-m aking: the traditional programming approach and the m achine learning
(ML) approach.
The traditional program m ing is suitable for tasks w ith clear rules and stru ctu res and machine
learning (ML) approach o ffers a more dynamic and adaptive approach for tasks mvolving pattern
recognition or decision-m aking in complex and evolving environments.
1. Traditional P ro g ra m m in g Approach:
. M ethod : In th e traditional programming approach, developers m anually write explicit
instructions (algorithm s) to solve a specific task or problem.
. P rocess : D evelopers analyze the problem, define rules and conditions, and create a set
of instructions (cod e) that the computer follows to produce the d esired output.
. P attern R e c o g n itio n : In tasks involving pattern recognition o r decision-making,
developers n eed to anticipate all possible scenarios and explicitly program rules to
handle each case.
. L im itatio ns : T h is approach is effective for tasks with clear and w ell-defined rules but
can be challenging for complex problems w ith ambiguous patterns o r evolving data.
2. M achine L ea rn in g A pproach:
• M ethod : In contrast, the machine learning approach involves train in g algorithms on
data to learn patterns and make decisions without explicit programm ing.
. P ro cess : Instead of manually coding rules, machine learning algorithm s analyze large
datasets, identify patterns, and adjust th eir models based on feed b ack to improve
accuracy.
. P attern R e c o g n itio n : Machine learning algorithms excel at recognizing complex
patterns and adapting to new information without the need for pred efin ed rules.
F u n d a m e n ia is^ M o ch in e L e a rn in g A | .n
Advantages : This approach is highly effective for tasks where patterns are difficult to
define explicitiy or v^^hen the data is constantly changing or evolving.
Flexibility : Traditional programming requires developers to anticipate all scen ario s and
write explicit rules, while m achine learning adapts to new patterns and data automatically.
Scalability : Machine learning can handle large and complex datasets more efficiently than
traditional programming, making it suitable for a wide range of applications.
A daptability : Machine learning models can evolve and improve over time as they learn
from more data, whereas traditional programs may require manual updates to accomm odate
changes.
12. Adaptation to Fluctuadng DaU. Environments : ML's Beribility makes it ideal for [Link]
requiring real-time adaptability. ^
13. Discovery and Uaming : ML algorithms reveal new correlations and trends, contnbufng
human learning and knowledge expansion^
programmed.
Note
A computer program which learns from experience is called a machine learning program or simply a learning
program. Sudi a program is sometimes also referred to as a learner._________ ______________ ___________
Machine learning problem s typically involve:
1. Task o r O b jectiv e ( T ) : Defining what needs to be achieved, such as classification, regression,
clustering, or p attern recognition.
2. Perform ance M e tric (P ): Determining how the performance of the m achine learning model
will be evaluated su ch as accuracy, precision, recall, or F I score.
3. Training E x p e rie n c e fE): Providing the model with a dataset of input features and
corresponding lab els to leam from during the training process.
By formulating a m achine learning problem with a clear task, performance m etric, and training
experience, developers and data scientists can design and implement effective machine learning
solutions to address a w ide range of real-world challenges and applications.
1. P red iction Models: ML uses statlstica: techniques to build models that can predict future
outcom es based on historical d ata.
2 . L earn in g from Data: ML system s learn ftom pr«rtous d atasets The more data they have,
b e tte r they can make predictions by recognizing patterns m the data.
actual results. ,
4 Feedback and Iteration: M achine learning is a continuous process. The
are compared to real outcom es, and any differences are used to improve the model. T h .s may
involve retraining the model w ith new data or adjusting its settings.
• f^n^[Link] diapram :
— -- ------------- r '
r
Building w Output
Training Machine Learning ---- ^ Logical Models
W
Input Past Data- Algorithm
^ J
!
New Data
Learn from Data
1 in p u t Past DaU: Historical d ata is fed into the ML algorithm for training purposes.
E xam p le : The system is fed w ith historical data on custom er purchases, mcludmg items
bought, purchase frequency, an d spending patterns.
2 Training: The algorithm u n d erg oes a tiaining phase w here it learns fro m the data prodded.
E xam p le • During the train in g phase, the algorithm learns from the past purchase data to
identify trends such as popu lar products, seasonal buying patterns, and customer p reference .
3 M achine Learning A lgorith m : This refers to the core s e t of rules and statistical processes
“ let^ealgorithm to learn from d,[Link] tbatproces
4 Bnildlns Logical M odels: A fter th e training phase, the system constructs logical models based
r !h ™ d patterns and relationships in the data. These models represent the insights
gained from the training data. , , . .
E xam p le : Based on the train in g data, the system creates logical models that ^how how
different factors like prod uct preferences and buying frequency influence custom er behavior.
5 O utput When new data is introduced to the system, the logical models created durm g training
° e t e d r p r e d i c t outcom es. The algorithm applies the learned patterns to the new data to
H o w M a c h in e L e a rn in g is b e i n g u se d ?
Machine learning is a modern innovation th at has enhanced many industrial and professional
processes as well as our daily lives. It’s a su b set o f artificial intelligence (Al), which focuses on using
statistical techniques to build intelligent com puter systems to learn from available databases.
With m achine learning, computer system s can take all the custom er data and utilise it. It operates
on what’s been programmed while also adjusting to new conditions or changes. Algorithms adapt to
data, developing behaviours that were not programmed in advance.
Examples
1. Image Recognition
Image recognition is a well-known and widespread example of machine learning in the real world. It
can identify an object as a digital image, based on the intensity of the pixels in black and white images
or colour images.
4. Statistical Arbitrage
5. Predictive Analytics
Machinelearningcan classify [Link],whicharethe„defin^by
i/Vhentheclassificati(
Real-world examples o f Predictive Analytics:
. Predicting whether a transaction is fraudulent or legitimate
Improve prediction systems to calculate the possibility of f a ^
Predictive analytics is one of the most promising examples of machine leammg. It s applicable for
everything: from product development to real estate pricing.
K i l l T y p e s o f M a c h i n e L ie a r n in g
!• several types o f m achine learning, each with special characteristics and applications. Som
o f the main types of m achine learning algonthms are as follows.
Supervised learning is a tj^ e of machine learning where the algorithm learns to map input data to
the correct output by being trained on labeled examples. In supervised learning, the training data
consists of input-output pairs, where the input data is accompanied by the corresponding correct
output or label. The goal of supervised learning is for the algorithm to learn a mapping function
from the input to the output so that it can make accurate predictions on new, unseen data^_______
1. Labeled Data: In supervised learning, the training data is labeled, meaning th a t each input
data point is associated w ith the correct output or targ et label. For example, in a spam email
detection system, each em ail is labeled as either spam or not spam.
2. Training Phase: During th e training phase, the algorithm uses the labeled data to learn the
relationship between the input features and the corresponding output labels. The algorithm
adjusts its internal param eters based on the training data to minimize the erro r betw een its
predictions and the true labels.
3 . Testing Phase: Once the model is trained, it is tested on a separate test d ataset th a t contains
new data that it has not seen before.
4 . Prediction: Once the m odel is trained, it can be used to make predictions on new, unseen
data. The model takes the input data and uses the learned mapping function to predict the
corresponding output label.
Machine JLcwrning ■■'d'h
5 . Evaluation: The perform ance of a supervised learning model is evaluated by comparing its
predictions on a separate test dataset with the true labels. Common evaluation m etrics include
accuracy, precision, recall, and FI score, depending on the nature of the problem.
In supervised learning, models are trained using labelled dataset, where the model learns about each type
of data. Once the training process is completed, the model is tested on the basis of test data (a subset of the
training set], and then it predicts the output.
The working of Supervised learning can be easily understood by the below example and diagram:
Labeled Data
Prediction
Square
1
Model Training
A Triangle
Laboles
O " □
Hexagon y V r Square
Test Data
Triangle
Suppose we have a dataset of different types of shapes which includes square, rectangle, triangle, and Polygon.
Now the first step is that we need to train the model for each shape.
• Labeled Data:
If the given shape has four sides, and all the sides are equal, then it will be labelled as a Square.
- If the given shape has three sides, then it will be labelled as a triangle.
- If the given shape has six equal sides dien it will be labelled as hexagon.
• Training Phase : The machine learning model is trained on this labeled dataset to learnthe patterns
and relationships between the input features (number of sides) and the corresponding class labels.
• Testing Phase: Once the model is trained, it is tested on a separate test dataset that contains new
instances of shapes that it has not seen before. The model takes the features of each shape in the test
set [e.g., number of sides) and predicts the class label based on the learned mapping function from the
training phase.
• Prediction: For each new shape in the test set, the model predicts the class label based on the learned
criteria (e.g., number of sides). If the shape has four equal sides, the model would predict it as a square.
If it has three sides, it would predict it as a triangle, and so on based on the defined rules.
•_Evaluation: The performance of the model is evaluated based on how accurately it predicts the class
______labels for the shapes in the test set
upervised machine learning algorithms can be categorized into two main types: Regression and
lassification.
^ —----— ————— ——— — — — — -Fundamentals o f M achine'
.................................
1. C lassification A lgorithm s: Classification algorithm s are used when the output variable is
categorical, m eaning it falls into distinct cla sse s or categories. The goal is to predict the class
label of new data points based on the p atte rn s learned from the training data. For example,
classifying em ails as spam or not spam, o r predicting whether a patient has a high risk of
heart disease o r not. Classification algorithm s learn to map the input features to one of the
predefined classes.
A list of co m m o n classification a lg o rith m s used in supervised learn in g:
o Logistic Regression ° Decision Trees
o Random Forest ° Support Vector Machines (SVM)
o Naive Bayes ° K-Nearest Neighbors (KNN)
1. Spam Filtering: In email filtering, a classification algorithm can be used to classify incoming
emails as either spam or non-spam based on the content and features of the email.
Algorithm: A common algorithm used for this task is Naive Bayes, which calculates the
probability of an email being spam or non-spam based on the presence of certain keywords or
features in the email content
2. Customer Chum Prediction: Predicting whether a customer is likely to churn (cancel their
subscription or service) based on historical data and customer behavior.
Algorithm: Random Forest is a popular algorithm for this task, as it can handle complex
relationships in the data and provide insights into the factors influencing customer chum.
3. Sentim ent Analysis: Analyzing text data to determine the sentiment (positive, negative,
neutral) expressed in reviews, social media posts, or customer feedback.
Algorithm: Support Vector Machines (SVM) are commonly used for sentiment analysis tasks
due to their ability to handle high-dimensional data and find optimal decision boundaries
between classes.
4. Image Classification: Classifying images into predefined categories such as animals, objects,
or scenes based on their visual features.
Algorithm: Convolutional Neural Networks (CNNs) are widely used for image classification
tasks due to their ability to learn hierarchical features from images and achieve state-of-die-art
performance.
5. Medical Diagnosis: Predicting the presence or absence of a disease based on patient symptoms,
medical history, and test results.
Algorithm: Decision Trees can be used for medical diagnosis tasks to create interpretable rules
for identifying patterns in patient data and making diagnostic decisions.
6. Fraud Detection: Identifying fraudulent transactions or activities in financial systems to
prevent financial losses.
Algorithm: Logistic Regression is commonly used for fraud detection due to its ability to model
binary outcomes and provide probabilities of fraudulent behavior based on transaction data.
2. R egression A lg o rith m s: Regression algorithm s are used when the relationship between the
input variables and the continuous output variable needs to be predicted. The goal is to estimate
a continuous value based on input featu res For example, predicting the price of a house based
^ ch in e Learning
on its size, location, and amenities, or forecasting the sales of a prod u ct Regression algorithms
leam to map the input features to a continuous numerical value.
A lis t o f com m o n reg ressio n alg o rith m s used in supervised le a rn in g :
o Linear Regression o Polynomial Regression
o Ridge Regression o Lasso Regression
o Decision tree Regression o Random Forest Regression
o Support Vector Regression (SVR) o Gradient Boosting Regression
1. House Price Prediction: Predicting the selling price of a house based on features like area,
number of bedrooms, location, etc.
Algorithm: Linear Regression is commonly used for house price prediction as it models the
relationship between the input features and the house price with a linear equation.
2. Stock Price Forecasting: Forecasting the future price of a stock based on historical stock data,
market trends, and other relevant factors.
Algorithm: Time Series Forecasting techniques like ARIMA (AutoRegressive Integrated Moving
Average) or Prophet can be used for stock price prediction to analyze and predict stock price
movements over time.
3. Demand Forecasting: Predicting the demand for a product or service in the future based on
historical sales data, market trends, and external factors.
Algorithm: Random Forest Regression can be employed for demand forecasting tasks to
capture complex relationships in the data and predict future demand levels accurately.
4. Temperature Prediction: Forecasting temperature values for weather forecasting applications
based on historical weather data and meteorological factors.
Algorithm: Support Vector Regression (SVR) is a regression algorithm that can be used
for temperature prediction tasks by finding the optimal hjfperplane to predict continuous
temperature values.
5. Sales Revenue Prediction: Estimating future sales revenue for a company based on historical
sales data, marketing campaigns, and economic indicators.
Algorithm: Gradient Boosting Regression algorithms like XGBoost or LightGBM are effective
for sales revenue prediction tasks as they can handle large datasets and capture non-linear
relationships in the data.
• P re d ictiv e Analytics: Utilized for forecasting outcomes such as sales figures, custom er
retention rates, and stock m arket trends.
• M ed ical Diagnosis: Helps in identifying diseases and medical conditions from patient data.
• Fraud D etection: Identifies and flags potentially fi^udulent transactions.
• A utonom ou s Vehicles: Enables vehicles to detectand respond to objects in dieir surroundings.
• E m ail S p am Detection; Classifies incom ing emails as eitiier spam or legitimate.
• Q u ality Control in M anu factu rin g: Used to inspect products for defects and maintain quality
standards.
• C red it Scoring: Assesses the cred it risk associated with borrow ers to predict loan default
probabilities.
• G am ing: Analyzes player behavior, character recognition, and NPC creation in gaming
environments.
. C u sto m er Su p p o rt Automates task s in customer service for improved efficiency.
. W e a th e r Forecasting: Predicts m eteorological parameters like temperature and precipitation
for w eather forecasts.
. S p o rts Analytics; Analyzes player performance, predicts gam e outcomes, and optim izes
strategies for sports teams.
O Predictive Accuracy: Supervised models can make accurate predictions by learning patterns from
labeled training data.
O Interpretable Results: Models provide insights into the relationship between input features and
target variables.
O Evaluation Metrics: Clear evaluation metrics such as accuracy, precision, recall, and FI score enable
easy model performance assessment.
O Feature Importance: Identify the most relevant features that contribute to the prediction outcome.
O Generalization: Supervised models can generalize well to unseen data, making them suitable for real-
world applications. ____________ ^
____________ ___________________ ______________________
C Data Dependency: Supervised learning requires labeled training data, which can be time-consuming
and expensive to acquire, especially for large datasets.
C Overfitting: Models trained on labeled data may overfit, capturing noise or irrelevant patterns in the
training set and leading to poor generalization on unseen data.
C Limited Flexibility: Supervised models are constrained by the features and labels provided in the
training data, limiting their ability to adapt to new or unseen patterns.
A
U n s u p e r v is e d L e a r n i n g
The word Unsupervised implies " n o t being done o r a c tin g u n d e r supervision". Unsupervised
learning is a learning method in which a machine learns w ithou t any supervision. The training is
provided to the machine with the s e t of data that has not been labelled, classified, or categorized, and
the algorithm needs to act on that data without any supervision.
The prim ary goal of Unsupervised learning is often to d iscov er hidden patterns, sim ilarities, or
clu sters within the data, which can then be used for various purposes, such as data exploration,
visualization, dimensionality reduction, and more.
Unsupervised learning is a type of machine learning where the algorithm learns to identify patterns
and relationships in data without being explicitly trained on labeled examples. Unlike supervised
learning, unsupervised learning algorithms work on unlabeled data, where the algorithm tries to
find hidden structures or patterns within the data. The goal of unsupervised learning is to explore
the data and extract meaningful insights without the need for predefinedjabels^^^^^^^^^^^^
1. U nlabeled Data: In unsupervised learning, the algorithm works with unlabeled data, meaning
that the input data points do not have corresponding output labels. The algorithm s task is to
discover the underlying structure or patterns in the d ata on its own.
2. Clustering: Clustering is a common technique in unsupervised learning where the algorithm
groups similar data points together based on th eir features or characteristics. Clustering
algorithms aim to partition the data into clusters such th a t data points within the sam e cluster
are more similar to each o th er than to those in other clusters.
3. D im ensionality R ed u ctio n : Dimensionality reduction techniques are used in unsupervised
learning to reduce the num ber of features in the data w hile preserving important information.
This helps in visualizing high-dimensional data and rem oving noise or redundant information.
4 . Anomaly D etection: Unsupervised learning algorithm s can also be used for anom aly detection,
where the algorithm identifies data points that deviate significantly from the norm o r expected
behavior. Anomalies are data points that are rare or unusual compared to the m ajority of the
data.
5. A ssociation Rule L earn in g : Association rule learning is another technique in unsupervised
learning that discovers interesting relationships or associations between variables in large
datasets. It is commonly used in market basket analysis to identify patterns in consumer
behavior.
Fundamentals of Machine Learning ^ 1.23
The below diagram illustrates a conceptual representation of an unsupervised learning model in machine
learning.
M odel
1. Input Data: This is the dataset given to the model. In unsupervised learning, the data has no labels. The
shapes (circles, triangles) represent different data points with their features.
2. Model: This is the core of the machine learning system. The model finds patterns or insights in the
input data without labeled guidance. In unsupervised learning, common tasks include clustering,
where the model groups similar data points together, and dimensionality reduction, where the model
simplifies inputs by removing redundant features to make data more understandable.
3. Output Data: This is what the model produces after processing the input In unsupervised learning, the
output is not predicted labels but insights from the data. The output data, shown as shapes, indicates
that the model may have categorized or clustered the data based on their characteristics.
The diagram gives a general idea of how unsupervised learning processes data to uncover its structure.
Unsupervised machine learning algorithms can be categorized majorly into three main types:
C lustering, D im ensionality R ed uction and A sso ciation .
1. Clustering : Clustering, algorithms are used to group similar data points together based on
their inherent characteristics or features. The goal is to discover natural groupings or clusters
within the data w ithout any predefined class lab els. Clustering algorithms learn to identify
patterns in the data and assign data points to clu sters. Common tasks include grouping
customers based on purchasing behavior or clu stering documents based on content similarity.
List of com mon c lu s te rin g algorithm s used in u n su p erv ised learning:
• K-Means Clustering
• Hierarchical Clustering
• DBSCAN (Density-Based Spatial Clustering o f Applications with Noise)
• Gaussian Mixture Models
2. Image Segmentation: Sep aratin g an image into different regions for image analysis and object
recognition.
Algorithm: Hierarchical clustering can be utilized in image segmentation for grouping pixels
or features that exhibit similar characteristics.
3. A n o m a ly Detection: Identifying miusual patterns or outiiers that do not conform to expected
behavior, such as fraud or network intrusions.
Algorithm: DBSCAN is effective for anomaly detection as it can find outiiers in datasets witii
noise and varied densities.__________________
2. Dimensionality Reduction Algorithms: Dimensionality reduction algorithm s are used to
reduce the number of input features in a dataset w hile presening im portant information. The
goal is to simplify the data by eliminating redundant or irrelevant features, m aking it easier to
visualize and analyze. Dimensionality reduction techniques are beneficial for high-dimensiona
data visualization and feature selection.
List of common dimensionaUty reducdon algorithms used in unsupervised learning:
• Principal Component Analysis (PCA)
• t-Distributed Stochastic Neighbor Embedding (t-SNE)
• Singular Value Decomposition (SVD)
• Independent Component Analysis (ICA)
. M achine Learning
Examples
1. Feature Selection: Reducing the number of input variables to simplify models and eliminate
redundancy.
Algorithm: PCAisfrequentlyused to tt-ansformhigh-dimensionaldataintoalower-dimensional
• FP-growth Algorithm
of Machine Learning
E )
1. Market Basket Analysis: Identifies products often purchased together to optimize marketing
strategies.
Algorithm: The Apriori algorithm is widely used for this application. It operates by identifying
the frequent individual items in the database and extending tiiem to larger item sets as long as
those item sets appear sufficiently often in the database.
2. Fraud Detection: Discovers patterns in data that may indicate fraudulent behavior.
Algorithm: Sequential pattern discovery using equivalence classes (SPADE] algorithm can be
used for detecting suspicious sequences of transactions.
3. Recommendation Systems: Recommends items based on user's past behavior.
Algorithm: Collaborative filtering approaches, which may include matrix factorization or
neural network-based recommendation algorithms, utilize association patterns to predict user
preferences.
4. Cross-Marketing: Finds associations between product categories to drive cross-promotional
strategies.
Algorithm: Association Rule Learning (ARL) algorithms, such as Apriori or FP-Growth, help
retailers bundle products in promotions effectively
5. Catalog Design: Arranges items in a catalog to maximize the discoveiy of associated items.
Algorithm: The Eclat algorithm uses transaction id intersections to improve computational
speed, ______________________________ __________________________
C u sto m er Behaviour A nalysis: Uncover patterns and in sig h ts for better marketing and
product recommendations.
^ C o n ten t Recom m endation: Classify and tag content to m ake it easier to recommend sim ilar
item s to users.
^ E x p lo ra to ry Data A nalysis (ED A ): Explore data and gain insights before defining specific
tasks.
O Discover Hidden Patterns: Uncover hidden patterns and structures in data without the need for
labeled outcomes.
0 Data Exploration: Facilitate data exploration and visualization by reducing high-dimensional data to
lower dimensions.
O Anomaly Detection: Identify outliers or anomalies in datasets that deviate from normal patterns.
O Feature Extraction: Extract essential features from data to improve model performance and reduce
dimensionality.
O Scalability: Easily handle large volumes of data and adapt to new data without the need for manual
labeling._________ ____________ ____________________________ _— ----------------- —----------------- —
C Interpretability: Models may be harder to interpret compared to supervised learning models due to
the lack of explicit labels.
C Evaluation Metrics: Lack of clear evaluation metrics for unsupervised learning tasks can make model
performance assessment challenging.
C Domain Knowledge: Requires domain expertise to interpret and validate the discovered patterns
effectively.
C Computational Complexity: Some unsupervised algorithms can be computationally intensive and
time-consuming, especially for large datasets.
S e m i-S u p e rv is e d M a c h i n e L e a r n in g
Semi-Supervised learning is a type o f Machine Learning algorithm th a t lies between Supervised and
Unsupervised machine learning. It represents the intermediate ground between Supervised (W ith
Labelled training data] and Unsupervised learning (with no labelled training data) algorithm s and
uses the combination of labelled and unlabelled datasets during th e training period.
Although Semi-supervised learning is the middle ground b etw een supervised and unsupervised
learning and operates on the data th at consists of a few labels, it m ostly consists of unlabelled data. It
is com pletely different from supervised and unsupervised learning as they are based on the presence
& absence o f labels. In semi-supervised learning, the algorithm is train ed on a dataset that contains a
small am ount of labeled data and a larger amount of unlabeled data.
m
Fundamentals of Machine learning ^ ^ 1.27
Semi-supervised learning is a machine learning paradigm where the algorithm learns from a
combination of labeled and unlabeled data. Unlike supervised learning that relies solely on labeled
examples and unsupervised learning that operates on unlabeled data, semi-supervised learning
strikes a balance between the two. The goal is to leverage the labeled data to guide the learning
process and utilize the unlabeled data to discover underl}dng patterns and relationships within the
data.
To overcom e the drawbacks of supervised learning and unsupervised learning algorithms, the
concept o f Semi-supervised learning is introduced. The main aim o f semi-supervised learning is to
effectively u se all the available data, rath er than only labelled data like in supervised learning. Initially,
similar d ata is clustered along with an unsupervised learning algorithm , and further, it helps to label
the unlabelled data into labelled data. It is because labelled data is a comparatively more expensive
acquisition th an unlabelled data.
We can im agine these algorithms with an example. Supervised learning is where a student is under
the su p erv isio n of an instructor at hom e and college. Further, if th a t student is self-analysing the
same co n ce p t without any help from the instructor, it comes u n der unsupervised learning. Under
sem i-supervised learning, the student has to revise himself after analysing the same concept under
the guidance o f an instructor at college.
In sem i-supervised learning, we w ork w ith a mix of labeled and unlabeled data to train our algorithm.
The key com ponents are:
1. L a b e le d and Unlabeled D ata: We have some data with lab els (like pictures of cats labeled as
"c a t") and a lot of data without labels. The algorithm learns from both types of data to improve
its understanding.
2. L ab el Propagation: This technique spreads the known labels to similar unlabeled data points.
If a labeled picture of a cat looks similar to an unlabeled picture, we can assume the unlabeled
on e is also a c a t This helps the algorithm learn more from th e unlabeled data.
3. Pseudo-Labeling: The algorithm predicts labels for the unlabeled data based on its current
understanding. These predicted labels are then used to train th e model further. It's like making
ed u cated guesses to teach the algorithm.
4 . Self-Training: The algorithm trains on the labeled data, th en uses its knowledge to predict
la b els for unlabeled data. If it's confident about these predictions, it adds them to the labeled
d a ta se t for future training. This process helps the model learn more from the unlabeled data.
5. Transfer Learning: In sem i-supervised learning, we can use knowledge from pre-trained
m od els on labeled data to assist in tasks with limited labeled examples. This transfer of
know ledge helps the algorithm learn more efficiently and improve its performance in the
sem i-supervised task.
Example |How SemrSupervised Learning Works?__________ _-------------------------------------------------
Let's consider an example of seml-supervlsed leamlni using the scenario of classilying entails as either spam
or not spam.
1 Labeled and Unlabeled Data: We have a small set of labeled emails where some are marked as spam
■ and others as non-spam. The majority of emails in our dataset are not labeled as spam or non-spam.
2. LabelPropagation:ThealgorithmlooksatthelabeledemailsandtheircharacteristicsOikekeywords,
senderinformation)[Link],hrun abe edem a s
For instance, if a labeled email with the word "discount" is marked as spam, similar unlabeled emails
with "discount" may also be considered spam.
3. Pseudo-Ubeling: The algorithm predicts labeU for the unlabeled emails based on its InlM
understanding. If It predicts that an email is likely spam based on its content, it assigns a pseudo la
of "spam" to that email.
4. Self-Training: The algorithm trains on the labeled emails and uses this knowledge to predict labels for
the unlabeled emails. If it is confident in its predictions (e.g., high probability of an email being spam),
it adds these emails with pseudo-labels to the training set for further training iterations.
5. Transfer Learning: If there are pre-ti^ined models for email classification taste with labeled data,
we can leverage their knowledge to improve our semi-supervised learning model s performance. The
insights gained from these pre-trained models can help our algorithm better understand and classify
emails in the semi-supervised setting.
By combining these techniques in a semi-superrised learning approach, our algorithm can
a larger set of emaiU as spam or non-spam, ewn with limited labeled data, leadmgto improved email flltenng
T here are several types or categories ofsemi-supetvised m achine learning algorithms. Here are some
com m on ones:
1 Self-Training A lgorithm s: These algorithms iteratively train on the labeled data and then use
' the model to predict labels for the unlabeled data. T h e high-confidence predictions are added
to the labeled dataset for further training.
2 Co-Training A lgorithm s: In co-training, the algorithm trains multiple models on different
subsets of features or data. Each model then provides predictions for the unlabeled data, and
the agreement between th e models helps in labeUng th e unlabeled instances.
3. Semi-Supervised Support Vector MaciUnes (S3VM): S3VM extends traditional Support
Vector Machines [SVM) to incorporate unlabeled d ata in the learaing process. It aim s to And a
decision boundary that n o t only separates the labeled d ata but also considers the distribution
o f the unlabeled data.
4 Graph-Based A lgorith m s: These algorithms re p rese n t the data as a graph w here nodes are
' data points and edges represent relationships b etw een them. By propagating labels^ through
the graph,thesealgorithmscanleverage the structure ofthedataforsem i-supervisedlearm ng.
underlying distribution o f the data. They can generate new data points and help in improving
the model's understanding of the data distribution.
6 . Low-Density S e p a ra tio n Algorithms: These algorithm s aim to find a decision boundary
that separates high-density regions (labeled data) fi-om low-density regions (unlabeled data].
By considering the density of the data points, th ese algorithms can effectively classify both
labeled and unlabeled instances.
Each type of semi-supervised learning algorithm has its strengths and is suitable for different types
o f datasets and learning tasks. Researchers and practitioners choose the most appropriate algorithm
b ased on the characteristics o f the data and the specific learning objectives.
Semi-supervised learning has various applications across different domains due to its ability to
leverage both labeled and unlabeled data efficiently. Here are some common applications of semi-
supervised learning:
1. Text Classification: In natural language processing tasks like sentiment analysis, document
categorization, or spam detection, sem i-supervised learning can be used to improve
classification accuracy by utilizing a combination o f labeled and unlabeled text data.
2 . Image Recognition: Semi-supervised learning is beneficial in image recognition tasks where
there is a large am ount o f unlabeled image data. By training on a small set of labeled images
and propagating labels to similar unlabeled images, the model can learn to recognize patterns
and objects more effectively.
3 . Speech R ecognition: Semi-supervised learning can Enhance speech recognition systems
by utilizing both labeled and unlabeled speech data. By leveraging the sim ilarities between
labeled and unlabeled speech samples, the model can improve its accuracy in transcribing
speech.
4 . Anomaly D etection: In cybersecurity and fraud detection, semi-supervised learning can help
identify anomalies in data by learning the norm al patterns from labeled data and detecting
deviations in the unlabeled data.
5 . Drug Discovery: In the pharmaceutical industry, semi-supervised learning can be applied to
predict the properties o f new drug compounds by training on a small set of labeled compounds
and leveraging the vast amount of unlabeled chem ical data available.
6 . Recom m endation System s: Semi-supervised learning can enhance recommendation systems
by utilizing both explicit user ratings (labeled d ata) and implicit user behavior (unlabeled
data) to provide m ore personalized and accurate recommendations.
7 . Medical Image A nalysis: In medical imaging tasks such as tumor detection or disease
diagnosis, sem i-supervised learning can assist in analyzing large volumes of medical images
by combining labeled images with similar unlabeled images to improve diagnostic accuracy
8 . Social Network A nalysis: Semi-supervised learning can be used in social network analysis
to predict connections o r identify communities w ithin a network by leveraging both labeled
connections and the netw ork structure of unlabeled data.
AdvantageB and Disadvanatges [Link] M a t o l ^ n g
Advantages o f ijemi-oupei
kdvantages ot Semi-Superviscd
viwu Machine Learning ...................................................... ...........
................. .
"Efficient Use Data: Setni-^pervised learning allows for the utilization of large amounts
ofunlabeled data,Which isoftenmoreabundantandeasiertoobtainthanlabeled data. This canlea
improved model performance without the need for extensive labehng efforts.
3 Cost-Effective: By reducing the reliance on labeled data, semi-supervised learning can be more
cost-effective compared to supervised learning, especially m scenarios where a e mg a i
consuming or expensive. .
=r :r r r r
unlabeled instances. . . r
O Performance: Seml-supeivised learning can lead to improved model
especially in cases where labeled daU is scarce. By incorporaSng unlabeled data, the model can
more robust representations and make better predictions.
S Scalability:Semi-supervisedlearningtechnlquescanscalewelltolargedatasets,astheycaneffectively
I------ r - nf iinlabeled data to enhance the model’s learning process^ -----------
^ ^f-'...
_ _ ------ T-;—-T— -------- 7 ^ ------ , - • ... TT'T^Tsmr’ i - - ^
r a l D isady^teges o f Semi-Supervised Machine Leari^^ng y^v-. ____
e OuaitoofUnlabeledData:Theeffectivenessofseml-supervisedlearningheavilyreliesontheqmlity
and reTevance of the unlabeled data. If the unlabeled data is noisy or contains irrelevant mformaOon,
aiiu leicvain^t; \ji --------------------
can negatively impact the model's performance.
C Model Complexity: Implementing semi-supervised learning algorithms can be more complex than
[Link],[Link],on
or pseudo-labeling.
e Risk of Overntting: In some cases, semi-supervised learning models may be P™ "' “
^ ^ d llly when the unlabeled data Introduces noise or biases that are not effecttveiy handled dunng
1111111 R e i n f o r c e m e n t L e a rn in g
T-»or*fr^rmanrP.
improves its perform ance.
- J ' Fa^am entals of Machine Learning Mi .31
Trial, error, an d d elay are the most relevant characteristics of rein fo rcem en t learning. In this
technique, the m odel keeps on increasing its performance using Rew ard Feedback to learn the
behaviour or pattern.
E n v i r o n m erit:
R e w a r d , S t:a te
A c tio n
A g e n t:
1. Agent: The entity that learns and m akes decisions based on the environm ent's feedback. The
agent takes actions in the environment to achieve a specific goal.
2. Environm ent: The external system with which the agent interacts. T h e environment provides
feedback to th e agent in the form of rew ards or penalties based on th e agent's actions.
3. Actions: T h e s e t of possible choices th at the agent can take in a given state. The agent selects
actions based on its policy, which defines how it chooses actions in different states.
4. State: The cu rren t situation or configuration of the environment at a particular time step. The
agent's actio n s influence the transition from one state to another.
5. Rewards: Numeric feedback provided by the environment to the ag en t after each action. The
agent's objectiv e is to maximize the cumulative reward it receives o v er time.
6. Policy: The strategy or set of rules that the agent uses to select actio n s in different states. The
policy can b e deterministic or stochastic.
Reinforcement learning algorithms aim to learn an optimal policy th a t maximizes the expected
cumulative rew ard over time.
Jig Example ' How Reinforcement Learning Works?
> Gam e P la 5ang: RL can teach agents to play games, even com plex ones.
> R o b o tics: RL can teach robots to perform tasks autonomously.
> A u to n om ou s Vehicles: RL can help self-driving cars navigate and make decisions.
> R eco m m en d atib n Systems: RL can enhance recommendation algorithms by learning user
preferences.
> H e a lth ca re : RL can be used to optim ize treatment plans and d rug discoveiy.
> N atural Language P rocessing (N LP): RL can be used in dialogue systems and chatbots.
> F in a n ce a n d Trading: RL can be used for algoritiimic trading.
> Supply C hain and Inventory M a n a g e m e n t RL can be u sed to optimize supply chain
operations.
bIs of Machine Learning
O It has autonomous decision-making that is well-suited for tasks and that can learn to make a sequence
of decisions, like robotics and game-playing.
O This technique is preferred to achieve long-term results that are very difficult to achieve.
O It is used to solve a complex problems that cannot be solved by conventional techniques.---------------
Supervised learning is not close to true Artificial Unsupervised learning is more close to the true
intelligence as in this, we first train the model for each Artificial Intelligence as it learns similarly as a child
data, and then only it can predict the correct output learns daily routine things by his experiences.
Machine learning is a buzzword for today's technology, and it is growing very rapidly day by day. We
are using machine learning in our daily life even w ithout knowing it such as Google Maps, Google
assistant, Alexa, etc.
Below are some most trending real-world applications o f Machine
1. Image R ecognition:
Image recognition is one of the most common applications of machine learning. It is used to
identify objects, persons, places, digital images, etc. The popular use case of im age recognition
and face detection is, A utom atic fiiend tagging suggestion.
Facebook provides us a feature of auto friend tagging suggestion. Whenever we upload a photo
with our Facebook friends, then we automatically get a tagging suggestion w ith name, and the
technology behind this is machine learning’s face d etectio n and reco g n ition algo rith m .
It is based on the Facebook project named "Deep Face,” which is responsible for face recognition
and person identification in the picture.
2. Speech R ecognition
While using Google, we get an option of "S e arch by voice," it comes under sp eech recognition,
and it's a popular application of machine learning.
Speech recognition is a process of converting voice instructions into text, and it is also known
as "Speech to text", o r "Com puter speech recognition ."
At present, machine learning algorithms are widely used by various applications of speech
recognition. Google assistan t, Siri, C ortana, and Alexa are using sp eech recognition
technology to follow the voice instructions.
'" •''•^’■‘iA ■ ■ Fundamentals of Machine Learning
3. Traffic Prediction:
If we want to visit a n ew place, we take help o f Google Maps, which shows us the correct path
with the shortest ro u te and predicts the traffic conditions.
It predicts the traffic conditions such as w hether traffic is cleared, slow-moving, or heavily
congested with the h elp o f two ways:
o Real Time lo c a tio n o f the vehicle form Google Map app and sensors
o Average tim e h a s taken on past days at the sam e time.
Everyone who is u sing Google Map is helping this app to make it better. It takes information
from the user and sen d s back to its database to improve the performance.
4 . Product R ecom m en d atio n s:
Machine learning is w idely used by various e-com m erce and entertainment com panies such
as Amazon, Netflix, etc., for product recom m endation to the user. W henever w e search for
some product on A m azon , then we started getting an advertisement for the sam e product
while internet surfing on the same browser and this is because of machine learning.
Google understands th e user interest using various machine learning algorithm s and suggests
the product as per cu stom er interest.
As similar, when w e use Netflix, we find som e recommendations for entertainm en t series,
movies, etc., and this is also done with the help o f machine learning.
5. Self-Driving C ars:
One of the most exciting applications of machine learning is self-driving cars. M achine learning
plays a significant ro le in self-driving cars.
Tesla, the most p opu lar car manufacturing com pany is working on self-driving car.
It is using unsupervised learning method to train the car models to detect people and objects
while driving.
6. Email Spam and M alw are Filtering:
Whenever we receive a new email, it is filtered automatically as important, norm al, and spam.
We always receive an important mail in our inbox with the important symbol and spam emails
in our spam box, and th e technology behind this is Machine learning.
Below are some spam filters used by Gmail:
o Content Filter
o Header filter
o General black lists filter
o Rules-based filters
o Permission filters
Some machine learn in g algorithms such as M ulti-Layer Perceptron, D ecision tre e , and Naive
Bayes classifler a re used for email spam filtering and malware detection.
7. Virtual P erson al A ssistan t:
We have various virtu al personal assistants such as Google assistant, Alexa, C ortana, Sin. As
the name suggests, th ey help us in finding the inform ation using our voice instruction. These
1 .3 6
assistants can help us in various ways ju st by our voice instructions such as Play music, call
someone. Open an email, Scheduling an appointment, etc.
These virtual assistants use machine learning algorithms .These assistant record our voice
Instructions, send it over the server on a cloud, and decode it using Machine U am m g
algorithms and act accordingly.
8. Online F rau d Detection:
Machine learning is making our online transaction safe and secure by detecting fraud
transaction.
Whenever we perform some online transaction, there may be various ways that a fra';‘dulent
transaction can take place such as fake accou n ts, fake ids, and s te a l m o n ey in A e middle of
a transaction. So to detect this. Feed Forw ard Neural netw ork helps us by checking whether
it is a genuine transaction or a fraud transaction.
Foreach genuine transaction, the output is converted into some hash values, and these values
become the input for the next round.
For each genuine transaction, there is a specific pattern which gets change for the fraud
transaction hence, it detects it and makes our online transactions m ore secure.
9. Sto ck M a rk et Trading:
Machine learning is widely used in stock market trading. In the stock market, there is always
a risk of up and downs in shares, so for this machine learning’s lo n g sh o rt term m em ory
n eu ral n e tw o rk is used for the prediction of stock market trends.
Google's GNMT (Google Neural Machine Translation) provide this feature, which is a Neural
Machine Learning that translates the text into our familiar language, and it called as automatic
translation.
The technology behind the automatic translation is a sequence to sequence learning algonthm,
which is used with image recognition and translates the text fi-om one language to another
language.
R f l T M a c h i n e l e a r n i n g L ife C y c le
HScfflne learning has given the computer systems the abilities to automatically
explicitly programmed. But how does a machine learning system work? So, it can be described us g
the life cycle o f m achine learning.
Machine learning life cycle is a cyclic process to build an efficient machine learning project The main
purpose of the life cycle is to find a solution to the problem or project.
I m
Fundamentals of Machine Lea
Machine learning life cycle involves seven m ajor steps, which are given below :
• Gathering D ata • Data preparation
• Data W rangling • Analyse Data
• Train the m od el • Test the model
• Deployment
In the complete life cycle process, to solve a problem, we create a m achine learning system called
“model" and this m odel is created by providing "training". But to train a m odel, we need data, hence,
life cycle starts b y collecting data.
1. Gathering Data:
Data G athering is the first step of the machine learning life cycle. T he goal of this step is to
identify and obtain all data-related problems.
In this step , w e need to identify the different data sources, as data can be collected from
various so u rce s such as files, d atab ase , in tern et, or m obile d ev ice s. It is one of the most
important step s of the life cycle. The quantity and quality of the collected data will determine
the efficiency o f the output. The more will be the data, the more accu rate will be the prediction.
This step includes the below tasks:
o Identify various data sources
o Collect data
o Integrate the data obtained from different sources
By perform ing the above task, we get a coherent set of data, also called as a dataset, It will be
used in fu rth e r steps.
2. Data Preparation
After collectin g the data, we need to prepare it for further steps. Data preparation is a step
where we p u t our data into a suitable place and prepare it to use in our machine learning
training.
In this step , first, we put all data together, and then randomize the ordering of data
This step ca n be further divided into two processes:
o DataExploration: It is used to understand the nature o f data th at we have to work
with. W e need to understand the characteristics, format, and quality of data. A better
Machine Learning - ^ ■
3. D ata Wrangling
Data wrangling is the process o f cleaning and converting raw data into a useable
important steps of the com plete process. Cleaning of data is required to address qu ty
issu 0s
It is not necessary that data w e have collected is always o f our use as some of the d ata may not
be usefiil. In real-world applications, collected data may have various issues, including.
o Missing Values
o Duplicate data
o Invalid data
o Noise
So, we use various filtering techniques to dean the data.
It is mandatory to detect and remove the above issues because it can negatively affect the
N ow ftrc ta n e d and prepared data is passed on to the analysis step This step involves:
N o w "th e °tL step is to train the model, in this step we train our model to improve its
performance for better outcom e of the problem.
We use datasets to train th e model using various m achine learning algorithm s. Training a
m L e U s required so that it can understand the various patterns, rules, and, features.
project or problem.
i Fundanieiitalj«#lto 1.39
7 . D eploym ent
The last step of machine learn in g life cycle is deplo 5m ient, where we deploy the m odel in the
real-world system.
If the above-prepared m odel is producing an accurate result as per our req u irem en t with
acceptable speed, then we deploy the model in the real system. But before deploying the
project, we will check w h eth er it is improving its perform ance using available data o r not. The
deployment phase is sim ilar to making the final report for a project.
Machine learning faces several challenges that can im pact the performance and reliab ility of
models. Each of these challenges plays a critical role in the su ccess of machine learning applications
and requires careful consideration and mitigation strategies to ensure accurate pred ictions and
generalization to new data.
1. In su fficien t Quantity o f T r a in in g D a ta :
Machine learning model generally require large am ounts of data to perform w ell. With
insufficient training data, m od els have a limited ability to learn, leading to poor perform ance
on unseen data. Insufficient training data can hinder the ability of machine learning m odels to
learn complex patterns and m ake accurate predictions.
Exam ple: Consider a facial recognition system designed to identify individuals at an event. If
the training data only includes a few images per person, the model might not learn to generalize
well, leading to inaccurate identification. *
2 . N onrepresentative T ra in in g D a t a :
Nonrepresentative training d ata means that the data used to teach a machine learning model
does not show a complete pictu re of what the model will face in the real world. This can lead
to the model making m istakes o r having biases because it hasn't learned from a w ide enough
range of examples. To avoid this, it's important to train the model on a diverse set o f d ata that
covers all possible scenarios it might encounter. This way, the model can learn m ore effectively
and make better predictions w hen faced with new, unseen data.
Exam ple: Consider a scen ario where a machine learning model is being trained to classify
different types of animals b ased on their features. If the training dataset only includes images
o f dogs and cats but lacks im ages o f birds and fish, the model may struggle to accurately classify
these missing animal types w hen presented with them in real-world scenarios. This lack of
representation in the training data can lead to the model making errors or show ing biases
towards the animals it was train ed on, resulting in unreliable predictions. To ensure th e model
can effectively classify a w ide range of animals, it is crucial to train it on a diverse d ataset that
includes various animal sp ecies with different characteristics and features.
3. Poor-Quality D a ta :
Poor-quality data Including m issing values, errors, o r inconsistencies can introduce noise and
biases into machine learning models. It affects the model performance and reliability. Data
cleaning and preprocessing a re crucial to address issues related to poor-quality data.
M a e h in e ^ rn in g : ■ .r., ‘
Why Python?
P )^ o n is used in numerous d ata science and machine learning applications due to its versatility and
user-fnendly nature. It seam lessly combines the robust capabilities of general-purpose programming
languages with the sim plicity o f domain-specific scripting languages like MATLAB or R. W ith a rich
ecosystem of libraries caterin g to data loading, visualization, statistics, natural language processing,
image processing, and m ore. Python equips data scientists with a diverse toolkit encom passing both
general and specialized functionalities.
The flexibility of Python extend s to its interactive nature, allowing users to engage directly with the
code through interfaces like term inals or popular tools such as Jupyter Notebook. This interactive
capability is particularly advantageous in the iterative nature of machine learning and d ata analysis,
where insights are derived from the data itself.
Python has emerged as the preferred language for m achine learning due to its exceptional qualities.
P}^hon is renowned for its simplicity, readability, and extensive library ecosystem tailored for data
science and machine learning tasks. Libraries like NumPy, Pandas, SciWt-learn, TensorFlow, and
PyTorch provide powerful to o ls for data manipulation, visualization, and model building.
Python is the preferred language for machine learning due to several key reasons.
• Rich Ecosystem o f Libraries: Python includes a vast array of libraries and frameworlcs specifically
created for machine learning and data science tasks. Libraries like NumPy, Pandas, Scikit-learn,
TensorFlow, and PyTorch provide powerful tools for data manipulation, visualization, and building
machine learning m odels.
• Ease of Learning and Use: Python is popular for its simplicity and readability. Its clean S3mtax and
extensive documentation enable developers to write code efficiently, accelerating the development
process. ^ ___________________________________
f ,42
Community Support: Python has a large and active community ot developers and data sa en u s^
who contribute to open-sounre projects and offer support through forums, tutorials, and onhne
resources. This collaborative environment fosters knowledge sharing and mnovation.
Versatility: Python is a ver^tile language that extends beyond machine learning to various
domains like web development, automation, scientific computing, and more. Its flexibility makes it
a valuable skill for professionals working across diverse fields.
integration Capabilittes: Python seamlessly Integrates with other languages and tools, faciltoting
easy integration with existing systems and technologies. It can be combmed with languages like C/
C++, Java, and R for building machine learning applications.
Scalability: Python offers scalability for machine learning projects by handling of large datasets
and complex algorithms. It supports deployment on different platforms, including cloud services,
to facilitate scalable machine learning model training and deployment
. Platform Independence: Python is platform-independent, meaning that code written in Python
can run on different operating systems without modification. It provides the flexibility and ease of
deployment across various platforms.
IR ra S c ik it-le a r n
S d k lt-le a m is a widely used open-source m ach in e learning library in Python that provides a simple
and efficient tool for data analysis and modeling- It is built on NumPy, SciPy, and Matplotlib which are
popular libraries for scientific computing and d ata visualization in Python. Scikit-learn is designe o
be user-friendly, accessible to both beginners and experts, and offers a wide range of machine learning
algorithms and tools for various tasks such as classification, regression, clustering, dimensionality
reduction, and more.
What is Scikit-learn ?
Scikit-learn is a popular open-source machine learning libraiy in Python that offers a comprehensive
set of tools and algorithms for data analysis, modeling and machine learning tasks. It is built on
foundational libraries like NumPy, SciPy. and Matplodib. Scikit-learn provides a user-fiiendly and
efficient framework for both beginners and experts in thejeld o fd ata^ aen ce^ ^ ^
3. Model E valu ation and Selection: S cikit-learn provides robust tools for model evaluation,
parameter tuning, and selection. Techniques such as cross-validation, grid search, and
performance m etrics like accuracy, p recision, recall, and F I score help users assess and
optimize the performance of their m achine learning models.
4. P rep rocessin g and Feature Engineering: The library includes utilities for data preprocessing,
feature scaling, feature selection, and transformation. These capabilities enable users to
prepare and clean their data effectively b e fo re training machine learning models to improve
model perform ance and generalization.
5. Integration w ith NumPy and Pandas: Scikit-learn seamlessly integrates with NumPy arrays
and Pandas DataFrames to facilitate d ata manipulation and com patibility with other data
science tools in the Python ecosystem. T h is integration stream lines the workflow for data
preprocessing and model building.
6. Extensive Documentation and C om m u n ity Support: Scikit-learn offers detailed
documentation, tutorials, and examples to assist users in understanding and utilizing machine
learning algorithm s effectively. The lib ra ry benefits from a strong community of users and
contributors who provide support, sh a re knowledge, and contribute to its development,
enhancing its usability and reliability.
7. Scalability a n d Perform ance: While prim arily designed for small to medium-sized datasets,
Scikit-learn offers scalability through integration with parallel processing libraries like Dask
and joblib. T his scalability feature allow s users to handle larger datasets efficiently and
leverage distributed computing resources w hen necessary
J J J j I n s t a i r i n g scikit-learn
To install scikit-leam on windows follow the ste p s given below:
Prerequisites
Python: Ensure P)^hon is installed on th e system. Scikit-learn is com patible with Python 3.6
or higher.
pip: Ensure pip is installed, which is the package installer for Python.
import sklearn
print(sklearn. ^version— )
> » ' a k le a m
p r i i t '( 3 k l e a t n * ^ v e r 3 i o n ^ ) .
» >
> »
Fu ndam entals of Learning ^ 1.4 5
Scikit-leam has dependencies on several other Python Ubraries, which are usually installed
automatically when installing Scikit-learn using pip. Some key dependencies include:
• N um Py: Fundamental package for scientific computing with Python.
• S ciP y : Library for mathematics, science, and engineering.
• jo b lib : Library for ligjitweight pipelining in Python.
• th re a d p o o lctl: Library for controlling the number of threads used by native libraries.
E s s e n t i a l L ib r a rie s a n d T o o ls
Understanding scikit-leam is im portant for machine learning applications. However, the additional
libraries such as NumPy, SciPy, pandas, and matplotlib enhance the overall experience. The Jupyter
Notebook, an interactive programming environment, is introduced for improved workflow. Proficiency
in these tools is essential for maxim izing the benefits of scikit-learn in real-world scenarios.
Understanding and utilizing these tools can significantly enhance the workflow and productivity of
data scientists and machine learning practitioners.
What are the Essential Libraries and Tools required for Machine Learning Projects ?
? l
Essential libraries and tools for effective implementation of machine learning projects is very important.
Understanding and utilizing these tools can significandy enhance the workflow and productivity of data
scientists and machine learning practitioners. ^
1. scikit-leam : A widely-used machine learning libraiy in Python that provides a simple and efficient
tool for data analysis and modeling.
2. NumPy: Fundamental package for scientific computing in P3rthon, providing support for large,
multi-dimensional arrays and matrices.
3. SciPy: SciPy is a library for mathematics, science, and engineering. It offers modules for optimization,
integration, interpoladon, and more. It is built on NumPy.
4 . pandas: Data manipulation and analysis library that offers data structures and functions to
efficiently work with structured data.
5. matplotlib: It is a libraiy for creating static, animated, and interactive visualizations In Python,
essential for data visualization tasks.
6. Jupyter Notebook: It is interactive web-based tool for creating and sharing documents that contain
live code, equations, visualizations, and narrative text. The Jupyter Notebook is an interactive
environment for running code in the browser. _________________ __
J u p y t e r ffateb o o k ~
Jupyter Notebook is an open-source w eb application that allows you to create and share documents
that contain live code, equations, visualizations, and narrative text. It is widely used in data science,
machine learning, scientific computing, and educational purposes.
01 .j^chiije^rning . •
Key Features
Installation
Jupyter Notebook can be installed using Python's package manager pip:
pip i n s t a l l notebook
• After installation, it can be started with the com m and;
jupyter notebook
This command launches a local web server and opens a notebook interface in the default web
browser
I Fundanienials o f M a ch in e Lc
NiimPy
NumPy is a core library used in various fields such as machine learning, data science, and scientific
research due to its powerful array manipulation capabilities.
In machine learning, NumPy arrays are used for storin g and manipulating data, serving as inputs to
machine learning algorithm s for tasks like classification, regression, and clustering.
Key Feature
3. Im age P ro ce ssin g and Com puter G rap h ics: NumPy arrays facilitate the storage and
manipulation o f pixel values for images. It can be used in various image processing tasks, such
as filtering, transformation, and visualization.
4. S cien tific Sim ulation s: Its efficient com putation capabilities make it ideal for simulations in
physics, chemistry, and engineering, w h ere large arrays of data and numerous computations
are common.
In stall Num Py : To install NumPy, type 'p ip install numpy' in the Command Prompt and
press 'Enter'. This command instructs pip to download and install the numpy package from
the P 5^ o n Package Index (PyPI).
5. Confirm th e In sta lla tio n ; After the installation process is complete, the successful installation
of NumPy can be verified. Type 'pjrthon' in the Command Prompt to open Python's interactive
mode, then type 'im port numpy a s n p ' and press 'Enter'. If no erro r message is displayed, it
confirms th at NumPy has been successfully installed. To exit the interactive mode, type 'exitQ'.
Example 1
import numpy as np Matrix Addition:
[[ 6 8]
[10 12 ]]
a = [Link]([[l, 2 ], [3 , 4]])
Matrix Subtraction:
b = [Link]([[5, 6], [7, 8]])
[[-4 -4]
[-4 -4]]
print("Matrix Addition:\n", a + b) Matrix Multiplication (element
printC'Matrix Subtraction:\n", a - b) wise) :
print("Matrix Multiplication (ele«ent-wise) :\n ", a* b) [[ 5 12]
printC'Matrix Multiplication:\n", [Link](a, b)) [2132]]
print("Tnanspose of a:\n", [Link](a)) Matrix Multiplication:
[[19 22]
[43 50]]
Transpose of a:
[[1 3]
[2 4]]
. Fundamentals of Machine I
Explanation
In the above example, the various matrix operations are performed using the NumPy library.
Specifically, the matrices 'a' and 'b' are added and subtracted element-wise, yielding the sum and
difference respectively.
The operation a * b performs element-wise multiplication.
The [Link](a, b) function carries out matrix multiplication by following the rules of linear algebra.
The [Link](a) function flips the matrix 'a' over its diagonal, effectively swapping its row and
column indices, to create a new matrix.
Explanation
In the above example, multiple statistical measures are calculated on the array 'a' using NumPy’s built-
in functions.
The standard deviation of the array, represented by [Link](a), measures the amount of variation or
dispersion of the set of values.
The variance of the array, calculated by [Link](a) is another measure of dispersion, which is essentially
the square of the standard deviation.
The median found using [Link](a) is the middle value in the sorted list of numbers that separates
the higher half from the lower half of the data set
The percentile is computed with [Link](a, 50), which represents the value below which a given
percentage of the data falls. In this case, the 50th percentile is calculated, which is also known as the
median.
S ciP y
SciPy is an open-source Python library th at is used for scientific and technical computing. It builds on
top of NumPy and provides a wide range of functions for numerical integration, optimization, signal
processing, linear algebra, statistics, and much more. SciPy is a powerful tool for scientific computing
and is widely used in various fields, including machine learning, physics, engineering, and biology
1 Inteeration and Optimization: SciPy Includes functions for numerical integration
' in ter^ lation , and opamizatlon. These capabilities are essential for sol«ng optimization
problems in machine learning, such as parameter tuning in algorithms like support
m achines CSVM) o r n e u r a l n e t w o r k s .
2 Signal Processing: SciPy offers tools for signal processing tasks like filtering, spectral ana^sis
fn^w a^form generation. These hinctions are valuable forprocessingandanalyzingsignalsm
These operations are crucial for many machine learning algorithms that involve matr
computations. .
4 Statistics- SciPy includes statistical functions for probability distributions, hypothesis testing
a^d d e s c r i p ~ ^ These functions are useMl for d ata analysis, model evaluation, and
understanding the significance of results in machine le a rn in g experiments.
5 S p arse Matrices: SciPy supports sparse matrix representations and provides efficien
■ afgorithms for working with large, sparse datasets. /"^trices are c o m j n l y
mwhine learning for tasks like collaborative filtering text minmg and graph analysis.
6 im ag e Processing: SciPy includes modules for image processing tasks such as filtering edge
A c t i o n , and morphology. These ftmcaons are beneficial for preprocessmg image data m
machine learning applications like computer vision and object recognlOon.
7 interoperability with N um Py: SciPy seamlessly integrates with NumPy, making it “
■ combinethearraymanipulationcapabiilttesofNumPywiththeadvancedsciennficcomput^^^^
functions of SciPy This integration enhances the efficiency and productivity
learning workflows.
1. Scientific Analysis: For task s that require precise calculations and date
physics and chemistry, SciPy provides robust algorithm s th a t are dependable and efficien t
2. Engineering Applications: Many engineering disciplines use SciPy for simulating real-world
processes, optimizing system s, and analyzing data.
3. Academic Research: Researchers in fields like econom ics, sociology, and psychology utilize
SciPy’s statistical tools to analyze experimental data.
4 . Im age Processing: SciPy's sub-package ndimage supports tasks in multi-dimensional image
processing widely used in fields such as medical image analysis and computer vision.
P r e r e q u is ite s : SciPy requires Python and NumPy. Before installing SciPy, ensure that Python and
Install SciPy: We can install SciPy using pip. Python’s package manager. To install SciPy, open
command prompt or term inal and type:
pip install scipy
Verify Installation: To confirm that SciPy has been successfully installed, launch Python in
interactive mode and try importing SciPy:
import scipy
If no error appears, the installation is successful.
x+2y=8
3x+4y-18
This system can be written in matrix form as:
[1 2 x 8
3 4 y 18
terms (B).
The matrix A represents the coefficients of the variables in the system of linear equations, and vector
B is the right-hand side of these equations.
The solve function computes the exact solution to the linear equations, offering a direct method
to handle such problems without manually implementing Gaussian elimination or other solving
techniques.
The script will output: Solution of the system: [2. 3.]
This means x=2 and y=3 are the solutions to the system of equations.______________________________
’I*
■ I
Pandas
s r ; r. ; ™ r r : i r i C i p . :
Pandas i
transform ation, exploration, and analysis.
structures.
d a ts t^ i— otyforanalysisandexporting .
‘ = = ~ = S = = S S =
• ^ ” 5 H = .S ~ = =
• i-3 S = 3 = r 2 H = r ~
•= = = = = — = = = =
■ s r ~ 3 - = s s s =
- i ‘r:FunaamentaIs of Machine LMrhing'/^
m m k
1. Data Cleaning: Data scientists often spend a large amount of their tim e cleaning data, and
Pandas provides powerful tools to perform this task efficiently.
2. Data E xp loration a n d Analysis: Pandas provides a high-level, flexible, and fast tool for data
analysis.
3. Data V isualization: It seamlessly integrates with the data visualization libraries such as
Matplotlib to plot data directly from data fram es.
4. Building M achine Learning Models: Before building models, data needs to be preprocessed
and transformed effectively. Pandas is often used for these tasks, ensuring that data is in the
correct form and ready for training models.
5. Tim e S e rie s A nalysis: Pandas has built-in support for handling time series data. Whether it's
resampling of tim e series data to convert frequencies, generating date ranges, or shifting and
lagging values. Pandas has tools to handle all th ese tasks effectively.
It's essential to have Python and pip (Python's package installer) pre-installed to install Pandas on a
system running Windows. The installation steps are:
1. Open Com mand P ro m p t: The Command Prom pt can be opened by searching for 'cmd' in the
Start menu and clicking on the Command Prom pt app.
2. Install P and as : To install Pandas, type 'p ip in sta ll pandas' in the Command Prompt and
press 'Enter'.
3. Confirm th e In s ta lla tio n : After the installation process is complete, the successful installation
of Pandas can be verified. Type 'python' in the Command Prompt to open Python's interactive
mode, then type 'im p o rt pandas as pd' and press 'Enter'. If no error m essage is displayed, it
:r a T r .“ e ; "
the respective data.
6. C u sto m izatio n and Styling: Matplotlib provides extensive customization options for styling
plots, such as setting plot colors, line styles, markers, fonts, and plot sizes. Users can create
visually appealing plots by adjusting the appearance of elem en ts to suit their preferences.
7. S u b p lo ts and Figures: Matplotlib supports the creation o f m ultiple subplots within a single
figure, allowing users to display multiple plots in a grid layout. This feature is useful for
com paring different datasets or visualizing related inform ation in a single plot window.
8. In te ra c tiv e Plotting: Matplotlib can be used in interactive m ode to create dynamic plots that
resp ond to user interactions, such as zooming, panning, and selecting data points. Interactive
p lottin g is beneficial for exploring data and gaining insights through visual exploration.
1. D ata Visualization; It is widely used for exploring and understanding data through
visualizations, especially where the data is time-series data o r ordered categories.
2. S c ie n tific Plotting: In academic and scientific publications, Matplotlib is used extensively to
cre ate high-quality plots, charts, and figures.
3. A lg o rith m Visualization: For data scientists and developers, visualizing the algorithm's
beh av ior can be crucial for diagnosing problems, and M atplotlib provides the tools necessary
to c re a te these visualizations.
4. In te ra c tiv e Applications: M atplotlib can be used to create desktop graphical user interfaces
or dynam ic dashboards for visualizing data in real-time.
5. E d u ca tio n : Its ease of use and wide range of plotting capabilities make it an excellent tool for
teach in g concepts in data science, statistics, and com putational mathematics.
Installation
The step -by-step process of installing Matplotlib using pip in the com m and prompt:
1. Open th e command prompt
2. Type in the following command to install Matplotiib:
pip install matplotlib
3. P ress the Enter key to execute the command. This will s ta rt the installation process for
M atplotlib.
4. W ait fo r the installation process to complete. We should see a message indicating that
M atplotlib has been successfully installed.
5. To verify the installation, we can type in the following com m and:
pip show matplotlib
6. This should display information about the installed version o f Matplotlib, including its location
and version number.
7. We can now start using Matplotlib in a Python projects. To use Matplotlib in a code, we need to
im p o rt the library using the following line of code:
import [Link] as pit
MacKine Uarniiig'.'
R e v ie w Q u estio n s
21. What is NumPy ? Why it is needed for ML? Explain its features.
22. What is Pandas? Why it is needed for ML? Explain its features.
23. What is Pandas? Why it is needed for ML? Explain its features.
24. What is Jupyter Notebook? Why it is needed for ML? Explain its features.
25. What is SciPy? Why it is needed for ML? Explain its features.
26. What is Matplotlib? Why it is needed for ML? Explain its features.
' !■v:»;&
Introduction
Data Preparation in Machine Learning
Working w ith Real Data
Look at the Big Picture
Get the Data
Discover and Visualize the Data to Gain Insights
Prepare the Data for Machine Learning Algorithms
M I B Introduction i onino
Data preparation Is a fundamental s t e p T h e quality and
transforming, and organizing raw data im oact the performance and accuracy of ML
r r r r j j ™ -
; r s : . : r x r r r : : ' = ' . r - . - « . _ -
machine learning model produces more reliable and accurate [Link] .
i M H n u ^ n g o f P g it^ l^ M ^ in e L e a rn i^ . ^ ^,
— i r.5 ^ tA > ro r ie s ^ n 5 a ta in M a c h in e L e a r n i n g
In machine learning, datasets are typically divided into three main subsets for model development
and evaluation: training data, testing data, and validation data.
DATA in Machine^
Learning ,
1. Training Data:
• Definition: Training data is the initial dataset used to train a Machine Learning model,
com prising input features and corresponding target labels or outcomes.
• Example: In a housing price prediction project, a dataset of 1,000 houses with features
like square footage, number of bedroom s, and location, along with their actual selling
prices, serves as the training data.
• Explanation: The model learns from the training data by analyzing the relationships
betw een the input features (square footage, bedrooms, location) and the target labels
(selling prices). Through processes like gradient descent, the model adjusts its internal
param eters to minimize prediction errors and enhance accuracy.
2. Validation Data:
• Definition: VaUdation data is a separate dataset utilized to fine-tune the model's
hyperparam eters and evaluate its performance during training.
Machine l^irning . V?- ’ "
• Exam ple: For the housing price prediction task, a su b set of 200 houses with sim ilar
features but different p rices is designated as the validation set.
• Explanation: During training, the model's perform ance with various hyperparam eter
configurations (e.g., learning rate adjustments) is assessed using the validation set.
By comparing the m odel's performance on the validation data for each configuration,
optimal hyperparameters are selected to improve the m odel's predictive accuracy.
3. T estin g Data:
• D efinition: Testing data is a dataset employed to assess the model's perform ance and
generalization capabilities on unseen data.
• Exam ple: Following training and validation, a fresh d ataset of 300 houses, com pletely
new to the model with undisclosed selling prices, is reserved for testing. f
• Explanation: The model is evaluated on the testing s e t to gauge its ability to predict
selling prices of unseen houses. Testing data provides an unbiased evaluation o f the
model's performance in real-world scenarios, indicating its capacity to generalize
accurately to new instances.
• The training data teaches the model to recognize patterns, the validation data helps fine-tune the
model's settings, and the testing data evaluates the model's performance on unseen instances. Each
type of data serves a specific purpose in the Machine Learning workflow, ensuring that the model is
trained effectively, optimized for performance, and capable of making accurate predictions on new
data.
W h at is Data Preparation ? ^
• Data preparation is defined as a gathering, combining, cleaning, and transforming raw data to
make accurate predictions in Machine learning projects. It is the later stage of the machine learning
lifecycle, which comes after data collection. ________
i ■‘Ottta Preparation 2.5
I m p o r ta n c e a n d B e n e f i t s o f D a ta P r e p a r a t i o n
P r e p a r a t io n I s s u e s in M a c h in e L e a r n i n g
Various issues have been reported during the data preparation step in machine learning as follows:
> M issin g data: Missing data o r incomplete records is a prevalent issue found in m ost datasets.
Instead of appropriate data, sometimes records contain empty cells, values [e.g., NULL or
N/A), or a specific character, such as a question mark, etc.
> O u tliers or Anom alies: ML algorithms are sensitive to the range and distribution o f values
w hen data comes from unknown sources. These values can spoil the entire machine learning
training system and the perform ance of the model. Hence, it is essential to detect these outliers
o r anomalies through techniques such as visualization technique.
> U nstru ctu red Data F o rm at: Data comes from various soiirces and needs to be extracted into
a different format Hence, before deploying an ML project, always consult with domain experts
o r import data from known sources.
> L im ited Features: W henever data comes from a single source, it contains limited features, so
it is necessary to import data from various sources for feature enrichment or build multiple
features in datasets.
> U nd erstanding Featu re E n gin eerin g : Features engineering helps develop additional content
in the ML models, increasing model performance and accuracy in predictions.
S t e p s in D a ta P r e p a r a t i o n P r o c e s s
Data Preparation is a crucial step in the Machine Learning process, and it involves various key steps to
ensure the data is suitable for training models effectively. The key steps involved in data preparation
are listed below:
1. D ata Collection: The initial step involves gathering raw data from a variety of sources such
as databases, file systems, sensors, or external APIs. This process lays the foundation for
subsequent data processing and analysis.
E xam p le: Gathering custom er information from a GRM system (such as customer IDs, purchase
history), transaction data from a database (including order amounts, timestamps), and social
media interactions from an API (like customer engagem ent metrics, comments) to analyze
custom er behavior for a targeted marketing campaign.
Ml
2 . Data Cleaning: Data cleaning is a fundamental step in ± e data preparation process that
involves identifying and rectifying errors, inconsistencies, and missing values in the datasetto
ensure its quality and reliability for subsequent analysis and modeling.
> Handling M issin g Values: Dealing with m issing data by either filling them with
appropriate values (e.g., mean, median, mode) or removing them to prevent inaccuraaes
in the model.
Example ; If som e customer records have m issing age and income columns. Replace
missing values in the "Age" column with the m ean age o f the avaiUble data and missmg
values in the "Incom e" column with the mean income.
> Filtering O u tliers: Identifying and removing anomalies that significantly deviate from
the rest of the data using statistical methods like Z-scores or IQR.
Example: Identifying and removing outliers in the customer age field, such as entries
with ages over 1 0 0 years, to ensure the data's accuracy.
3. Data Transformation: Data transformation involves converting and standardizing the dataset
to make it more suitable for Machine Learning algorithm s by ensuring consistency, reducing
redundancy, and improving interpretability.
> Normalization and Standardization: Scaling data to a standard range (e.g., 0 to
1 for norm alization) or transforming data to have zero mean and unit variance
(standardization) to ensure consistency, especially for models sensitive to feature scales.
. Think of normalization as adjusting values to fit within a specific range, like
resizing a photo to fit a frame. It ensures all data points are on a sim ilar scale,
preventing any one type of data from overshadowing others.
. Standardization is like making data follow a standard pattern, such as ensuring all
ingredients in a recipe are measured in the same units. Ithelps m odels understand
and com pare different types of data m ore easily
Example : Scaling customer purchase amounts to a standard range (e.g., 0 to 1) to ensure
consistency in the analysis, especially when comparing with other features like customer
engagement levels.
> Encoding C ateg orical Variables: Converting categorical variables into a form at suitable
for ML algorithms, such as one-hot encoding or label encoding.
. O ne-H ot Encoding: Imagine creating a list of checkboxes for different categories,
where each checkbox is either ticked (1 ) or unticked (0). This method helps the
model understand and use categorical data effectively.
. L abel Encod ing: Think of assigning a unique number to each category, like giving
each type o f fh iit a specific code. This simplifies the data for the model to process,
making it easier to work with different categories.
Example: Converting categorical data like customer segmentation (e.g., premium,
standard, basic) into numerical values using one-hot encoding for model compatibility.
4 Data Reduction: Data reduction techniques aim to simplify and con d en ^ the dataset while
retaining its essential information, making it m ore manageable and efficient for Machine
Il»
Data Preparation
Encoding. This technique creates binary columns for each category within the "Region" .
feature, assigning a value o f 1 if the customer belongs to th at region and 0 otherwise. By
encoding categorical variables in this manner, we ensure th at the model can effectively
utilize this information for making predictions w ithout introducing any ordinal
relationship between the regions.
> F e a tu re Creation: Feature creation involves generating new features by combining
existing ones or extracting additional information from the data. This process aims to
provide the model with more relevant and informative input variables.
Creating new features can capture complex relationships in the data that the original
features may not fully represent. It can lead to improved model performance and better
generalization to unseen data.
Feature creation techniques include polynomial features, interaction terms, domain-
specific feature engineering, and text or image feature extraction. These methods
help enrich the dataset with new information that can enhance the model's predictive
capabilities.
Exam p le : Consider com bining customer attributes like purchase history, browsing
behavior, and demographic information to create a new feature representing overall
custom er engagement This feature creation process helps in capturing complex
relationships within the data and providing a more com prehensive view of customer
interactions, enabling better insights for marketing strategies and customer segmentation.
6. D ata Sp ittin g : Data splitting in m achine learning is the process o f dividing the data into
separate subsets to be used at different stages of model building and evaluation. The primary
goal o f data splitting is to ensure that the model trained on one set o f data can generalize well
to new, unseen data. This helps avoid problems like overfitting, w here a model performs well
on the training data but poorly on new data.
Com m on Types of Data Splits
> T rain in g Data: This is the largest portion of the dataset and is used to train the model.
The model learns to identify patterns and make decisions based on this data. Typically,
about 70-80% of the entire d ataset is allocated to the training set.
> V alidation Data: This subset is used to tune the model's hyperparameters and make
decisions about which models or configurations to use. It acts as a check to avoid
overfitting on the training d ataset. Typically, about 1 0 -1 5 % o f the dataset is reserved for
validation.
> T e st Data: This is used to evaluate the final model's perform ance after it has been trained
and validated. The test set should be a completely independent dataset that the model
has not seen during training o r validation. This helps provide an unbiased evaluation o f
how well the model is expected to perform in the real world. Typically, about 10-15% of
the dataset is used as the te st set.
Data Preparatio
What is Data Collection ? Explain the Key Steps Involved in Data Collection.
%
Data collection is the process of gathering relevant information o r data from various sources
to be used for analysis, decision-making, or research purposes. It is a crucial step in any data-
driven p ro ject as the quality and accuracy of the collected data directly impact the outcomes of
subsequent analyses and modeling.
The Key Steps involved in data collection :
1. Define Objectives: Clearly define the objectives and goals o f the data collection process.
Understand what specific inform ation is needed and how it will be used to achieve the
desired outcomes.
2. Identify Data Sources: D etermine the sources from which the data will be collected.
Sources can include databases, surveys, sensors, web scraping, social media, public
records, etc.
3. Design Data Collection M ethods: Choose appropriate methods for collecting data
based on the objectives and sources. Common methods include surveys, interviews,
observations, experiments, and automated data collection tools.
4 . Develop Data Collection Tools: Create tools such as questionnaires, forms, sensors, or
softw are applications to collect data efficiently and accurately. Ensure that the tools are
designed to capture the required information effectively.
5. D ata Collection: Implement the data collection process according to the defined
m ethods and tools. Collect data from the identified sources while ensuring data quality,
consistency, and relevance to the objectives.
6. D ata Storage and M anagem ent: Organize and store the collected data in a secure and
accessible manner. Establish data management protocols to maintain data integrity and
confidentiality
7. D ata Documentation: Document the data collection process, including details of
sources, methods, tools used, and any modifications made during the process. This
documentation is essential for transparency and r e p r o d u c i b i l i t y . __________
Numerous open datasets are available online across various domains to provide valuable resources
for experim entation and study. These resources offer a wide range of datasets for machine learning
practitioners to explore and utilize in th eir projects by covering diverse topics and domains.
1. UC Irv in e M achine L earning R e p o sito ry : A well-established repository hosting datasets
specifically for machine learning experiments.
W eb site: [Link]
2. K aggle D a ta s e ts : Known for hosting competitions, Kaggle also offers a diverse collection of
datasets contributed by users and organizations, spanning various domains from econom ics
to image data.
W e b s ite : [Link]
3. A m azon's AWS D atasets : Amazon Web Services (AWS) provides a vast array o f public
datasets that can be seamlessly integrated with cloud-based applications.
W e b s ite : [Link]
4. A dd itional R e so u rce s:
• W ikipedia's List o f M ach in e Learning D atasets : A resource listing various d atasets
suitable for machine learning projects.
W e b site :h ttp s ://e n .w [Link]/w iki/List_of_datasets_for_m achine-learning_
research
• Q [Link] : A platform w here users can find discussions and recommendations on
datasets for machine learning.
W e b s ite : [Link]
• Datasets S u b red d it: A subreddit dedicated to sharing and discussing datasets across
different domains.
W ebsite: [Link] w w .reddiLcom /r/datasets/ ____________________________
By working with this real-world dataset, machine learning practitioners can explore predictive modeling tasks
related to healthcare and medical diagnostics, gaining insights into the factors influencing liver disease in
the Indian population. Analyzing this data can lead to the development of predictive models that assist in
early detection and management of liver-related conditions, showcasing the practical application of machine
learning in healthcare.
This dataset can be utilized for agricultural applications and precision farming in India to help farmers make
informed decisions about crop selection and optimizing agricultural productivity. By analyzing this real-world
data, machine learning models can provide valuable insights and recommendations for crop cultivation,
contributing to sustainable farming practices and enhancing agricultural outcomes in diverse regions of India.
8. Deplojmient:
• Deploy the trained model into a production environment where it can receive new data
inputs and make predictions on house prices.
9. M onitoring and Maintenance:
• Monitor the model's performance regularly in the live environment.
. Retrain the model periodically with new data to keep it updated and accurate.
10. Feedback Integration:
. Gather feedback from real estate agents, buyers, and sellers on the model's predictions.
. Use feedback to improve the model's accuracy and refine the implementation process.
(jretthe Data
Working on a m achine learning project begins by obtaining the necessary data, which serves as tlie
foundation for all subsequent modeling and analysis. Let us understand the structured approach to
effectively gather data for ML projects:
1. S ettin g Up Your Environment:
. S ystem Preparation: Ensure Python is installed on the system. If not, it can be
downloaded and installed from Python's official website.([Link]
. W o r k s p a c e C r e a t i o n r - A - d e d i c a t e d - d i r e c t o i i ^ o r m a c h i n e l e a r n i n g p r o je c t s s h o u ld be
created for organizational clarity. This can be set up using Command Prompt:
DATA_URL="[Link]
housing/housing.C SV "
DATA_PATH = [Link]("C:\\ML_Projects", "datasets", "housing")
if __name__ == "__^main__
fetch_data()
The CSV is downloaded to the C:\ML_Projects as shown below.
Explanation
The above Python script automates the downloading of a dataset from a specified URL and stores it
locally.
import os: This module provides a way of using operating system dependent functionality like
reading or writing to the file system.
. import urllib-request: This module is used for opening and reading URLs, specifically, it is
used here to download data from the internet.
. [Link]: A string that holds the URL of the dataset. This URL points to a CSV file hosted on
GitHub.
. DATA PATH: A string that specifies the local directory path where the dataset will be saved. It
uses o'[Link] to construct the path. This function is platform-independent and ensures the
path is correctly formatted for the operating system.
• fetch_data: The function is defined to handle the downloading of data.
. [Link](data_path, exist_ok=True): Ensures that the directory specified by [Link]
exists. If it doesn’t, this function will create all necessary directories in the path. The parameter
exist_ok=True allows the command to succeed even if the directory already exists.
. csv_path: Constructs the full local file path where the CSV file will be saved after downloading.
. [Link](data_url, csv_path): Downloads the file from data_url and saves it
to the location specified by csv_path.
. printfD ata downloaded to:", csv_path): Outputs a message to the console indicating where
the data has been downloaded.
. Main Block - if _ n a m e _ == "_m ain_": This block ensures that the fetch_dataQ function is
called only when the script is run directly__________ _________ ______________
L o a d th e D a ta a n d E x p l o r e th e D a ta
When working with machine learning projects, the first m ajor step after acquiring the data is to load
it into a usable format for analysis and preprocessing. The m ost common format for data storage is
the CSV (comma-separated values) file, which can easily be loaded into a Pandas DataFrame. After
loading the data, exploring it helps in understanding the structure, content, and initial insights that
guide further data manipulation and analysis.
1. Load the Data : The process begins with importing necessary libraries and loading the data
into a DataFrame. This is tj^ ically done using Pandas due to its robustness and ease o f use for
handling structured data.
2. Exploring the D ata : Once the data is loaded into a DataFrame, it's crucial to explore it to
understand its characteristics, such as the number o f features, rows, possible m issing values,
and the type of data (num erical or categorical).
Example Load the Data
import pandas as pd
import os
# Main block to ensure the script runs only when directly executed
if __name__ == "__main__" :
data = load_data(DATA_PATH) # Load the data
explore_data(data) # Explore the data
Explanation <
• Import Libraries: The script starts by importing necessary libraries, pandas is used for data
manipulation and analysis, and os helps in handling file and directory paths.
• Define Data Path: DATA_PATH holds the directory path where the CSV file is stored. This makes the
script more flexible and easier to modify if the data location changes.
• Load Data Function: load_data function takes the path to the CSV file, constructs the full path to the
file, reads it using pd.read_csvQ, and returns the DataFrame. This DataFrame contains all the data from
the CSV file, ready for analysis.
• Explore Data Function: [Link] function takes a DataFrame as input and performs three key
operations: __________________________________ _________
Head of tiie DataFrame: Shows the first five rows using [Link], providing a quick view
dataset structure and initial rows.
Stattsacal summary: Uses daK,.describeO to output summary statlsdcs tijat describe the uumen,:al
flelds<,[Link]>it,mean,stdCstandarddeviatloii),[Link].
Separating a portion o f the dataset for testing is essential as it allows for assessing how well a model
will perform on new data that it has not seen during training. This process sim ulates real-world
scenarios and helps in determining the model's reliability and generalization capabilities.
1. Allocation: W hen creating a test set, it is im portant to allocate a specific portion of the dataset,
typically around 20% , for testing purposes. This allocation ensures that there is a separate
subset of data reserved exclusively for evaluating the model's performance.
2. Purpose: The te st set serves as a critical com ponent in the model development process by
providing a m eans to assess how well the model generalizes to new, unseen data. By withholding
a portion of the data for testing, we can evalute the model's effectiveness in making predictions
on data it has n ot been trained on.
3. Independence: The test set should be independent of the training data to avoid any biases
that may arise from using the same data for both training and evaluation. This independence
ensures that the model's performance is assessed on truly unseen instances, enhancmg its
reliability in real-world applications.
4. Evaluation: Testing the model on a separate te st set allows for a com prehensive evaluation
of its predictive capabilities. By com paring the model's performance on the test set to ite
performance on the training data, we can gain insights into its ability to generahze and make
accurate predictions on new data.
2 .2 0 ^ - Woehine Ufflrning
5. V alidation: The test set acts as a validation mechanism for the model, providmg a benchm ark
for assessing its performance and identifying any potential issues such as overfitting or
underfitting. This validation step is crucial in ensuring the m odel's robustness and reliability
in practical scenarios.
6. Consistency: To maintain consistency in model evaluation, it is recommended to set a random
seed when splitting the data into training and test sets. This practice ensures that the test set
remains consistent across different ru ns of the model, allowing for reliable comparison and
assessm ent of model performance.
1. Random Sam pling: This method is suitable for large d atasets w here randomly selecting
data points for the test set ensures th a t it is representative of the overall dataset. It helps in
evaluating the model's performance on a diverse set of instances.
2. Stratified Sam pling: When certain characteristics or categories in the data are crucial for the
model's predictions, using this technique ensures that these attributes are well-represented in
the test set. It is particularly useful for sm aller datasets w here maintaining the distribution of
important features is essential for accu rate evaluation.______________
strat_train_set = [Link][train_index]
strat_test_set = [Link][test_index]
# Remove the income_cat attribute so the data is back to its original state
for set_ in (strat_train_set, strat_test_set):
set_.drop("income_cat", axis=l, inplace=True)
4.1250 0.002907
3.1250 0.002907
15.0001 0.002665
3.8750 0.002422
2.1250 0.002422
3.7831 0.000242
2.8056 0.000242
2.1270 0.000242
2.1518 0.000242
4.1111 0.000242
Name: medianjncome, Length: 3446, dtype; float64
Explanation
Load Data : The load_housing_data function is crucial as it allows the seamless reading of structured
data (like CSV files) into a Python environment where it can be easily manipulated and analyzed.
This function uses the pandas library, which is a powerful tool for data analysis and manipulation.
Specifically, pandas.read_csvQ is used to load the data from a CSV file into a DataFrame. A DataFrame
is a 2-dimensional labeled data structure with columns of potentially different types of data.
By loading data into a DataFrame, users can take advantage of the various data manipulation
capabilities of pandas to clean, transform, and preprocess the data effectively before any analysis or
model training.
Income Categoiy : Creating an incom [Link] column is essential for performing stratified sampling
based on the median income of the households in the dataset Stratified sampling is a method of
sampling that involves dividing a population into smaller groups, known as strata, that share a similar
attribute.
The function uses [Link] to categorize the medianjncome into specified bins. This categorization
helps in ensuring that the sampling of data for model training and testing reflects the overall distribution
of income categories in the entire dataset_______________________________ ___ ________________
Stratification ensures that each category of income is properly represented in both training and test
sets, which helps in building a model that performs well across different income groups, thereby
reducing sampling bias.
• Stratified Sampling: It is used to divide the data into a training set and a test set while maintaining a
consistent percentage of samples for each categoiy of income across both sets.
StratifiedShuffleSplit from skleam.model_selection provides a way of ensuring that the data is
randomly split in such a way that the income categoiy proportions are preserved in both training and
test datasets as compared to the full dataset
This technique helps in maintaining the statistical properties of the original data, which can be critical
for the predictive performance of the machine learning model, especially on unseen data.
• Clean Up : It is used to clean up the data by removing the income_cat column after the stratified
sampling is done. The dropQ method from pandas is used on both training and test sets to remove the
income_cat column, reverting the dataset to its original state before stratification.
This step is crucial for keeping the data tidy and ensuring that only the original attributes are used for
further analysis and model training.
• Printing Proportions; It is to verify that the stratification was performed correctly. The script prints
the proportion of each income category writhin the test set using value_count$0 and normalizing the
results with lenQ to get the percentage.
This confirmation step is essential to ensure that the stratified sampling process has been executed as
expected, providing confidence in the robustness of the subsequent training and validation processes.
4 . E rro r Id e n tificatio n : Early in the process, visualizing data can reveal errors in data collection
and processing such as biases and inconsistencies that need to be addressed before further
analysis.
5. Inform ing P rep ro cessin g D ecision s : Insights gained from visualizing data can direct the
preprocessing steps like normalization, handling m issing values, or feature engineering.
For example, if data visualization shows that some variables have a non-linear relationship,
poljmomial features might be created to model these effects.
6. Facilitatin g C om m unication : Visualizations make it easier to communicate findings and
data characteristics to stakeholders, who may not be familiar with the technical aspects of
data science but need to understand the basis of decisions or models built from the data.
2. Box Plots:
• P u rp ose: Box plots summarize the distribution of data by displaying quartiles, outliers,
and the overall spread of values in a com p act visual format.
• Usage; Box plots are valuable in data preparation for comparing the distribution of
features across different categories o r groups and identifying potential outliers or
variations.
• Exam ple: In a retail dataset, box plots can b e used to compare sales figures across different
product categories. Detecting outliers in sales data for specific product categories can
prompt further investigation into data quality issues or anomalies before analysis.
4. H eatmaps:
• P u rp ose : Heatmaps use color gradients to represent data values in a matrix format,
making it easier to identify patterns and relationships in large datasets.
II Data Preparation
U sage: Heatmaps are beneficial in data preparation for visualizing correlations between
features, identifying clusters, or detecting anomalies.
E xam p le: Creating a heatmap o f feature correlations in custom er survey data can
highlight strong relationships betw een satisfaction scores and purchase behavior. This
can guide feature selection strategies during data preprocessing to enhance model
perform ance.
5. Scatter P lo ts:
• P u rp o se: Scatter plots display the relationship between two variables by plotting data
points on a Cartesian plane, helping to identify patterns, trends, and correlations.
• U sage: Scatter plots are useful in data preparation for exploring associations between
variables and detecting oudiers or data inconsistencies.
• E xam p le: Plotting a scatter plot o f custom er age against purchase frequency in a sales
dataset can reveal any linear or non-linear relationships betw een age and buying
behavior. This insight can guide feature engineering decisions during data preprocessing
to capture relevant patterns for predictive modeling.
Scatterplot o f %Fat vs BM f
45
AO
^ 3S
tE
as 30
25
20
15
15 20 25 35
w m m ^
6. Line C h arts:
• P u rp o se : Line charts depict data trends over time by connecting data points with hnes,
making them ideal for tracking changes and patterns in sequential data.
. U sage: Line charts are valuable in data preparation for visualizing temporal trends,
m onitoring data quality metrics, o r tracking preprocessing steps over time.
• E xam p le: Tracking the evolution o f data cleaning efforts using a line chart of missing
value percentages over successive data preparation stages can help in assessing the
effectiveness of data cleaning techniques and ensuring data quality before model training.
Product A, Product B and Total Product Sold
Aditya, 20,BCA,4,88,172,70
Bhavika,19,BBA,2, missing,158,50
Chirag, 21,BCom,5,9 2 ,180,80
Deepa,22,BCA,6,94,160,missing
Esha,20,BBA,3,missing,70,54
Farhan,19,BCA,l,85,165,65
Gita,21, BCom,6,78,170,75
Hitesh, 20,BCA,4,101,165,170
Ila,18,BBA,l,87,190,48
Dai,22,BCom,5,85,180,82
Kavya, 19,BCA,2,missing,159,110
Lalit,20,BBA,3,95,175,80
Mira, 2 1 ,BCom,6,82,missing,70
Nikhil,22,BCA,5,93,174,76
dm,19,BBA,2,75,168,60
Priya,18,BCom,1,95,154,55
Raj, 21, BCA, 5,90,182,83
Sunita, 20,BBA,3, missing,160,120
Tarun,19,BCom,4,88,80,40
Usha,21,BCA,6,92,164,59___________ ___________ ________ __________________ _
Example Data Visualization
import pandas as pd
import [Link] as pit
import numpy as np
[Link](rlU(len([Link])), ™ ’=ation-«)
[Link](range(len([Link])), [Link])
[Link]('Correlation Heatmap’)
plt.tight_layout()
[Link]
def main():
data = load_data()
prepared_data = prepare_data(data)
plot_data(prepared_data)
if __name__ == — •
Weight Distnbution Cburae Bmilltnent
Explanation:
The chart provides a comprehensive look at the distribution and relationship between various features in a
student dataset
1. Attendance Percentage Distribution: This histogram likely shows the frequency of students across
different ranges of attendance percentages. Peaks in the histogram can indicate common attendance
rates, while gaps might suggest less common attendance behaviors.
2. Weight Distribution : The box plot for weight distribution probably displays the spread of student
weights. Outliers may appear as individual points, indicating students with weights significantly higher
or lower than the rest.
3. Height Distribution : Similar to the weight distribution, the box plot for height helps identify the
median, quartiles, and potential outliers in student heights. This can help pinpoint if any students are
unusually tall or short, which might be outliers needing verification.
4. Height vs Weight Scatter P lo t: This scatter plot likely illustrates the relationship between students'
heights and weights. A clear trend, like an upward slope, would suggest a positive correlation, where
taller students tend to be heavier.
5. Course Enrollment: A bar chart showing the number of students enrolled in different courses would
indicate the popularity or selection frequency of each course, like BCA, BBA, or BCom.
relationship between variables, like age and semester correlatmg with progression
program.
Each c h a rt helps in identifying different data discrepancies: t+pm.= thatreauire
.Thehistogr^mcan reveal Ifattendancedata is [Link]
further investigation.
. The box plots for weight and height can highlight outliers that might be due to data entry er ^
. The scatter plot can help find any unusual relationship between height and weight which might
conform to medical standards.
. The bar chart can suggest if them’s an imbalance in course enrollments that may affect class soes
resource allocation.
building^_____________ __________
Prepare
irrepaic? the D ata for Machine Learning Algorithms
----------------
m m B m m
m achine learning model d u r in g training and evaluation. „„rW i„w this
Missing values in data refer to the absence of inform ation or data points for certain observations or
attributes in a dataset. Handling missing values is crucial in data preprocessing to ensure the qualitv
and reliability of the m achine learning model.
Example: The Titanic Passengers dataset has m issing values in the Age and Cabin columns. The
passenger information has been extracted from various historical sources. In this case the missinc
values couldn't be found in the sources.
Passengerd Survived Pclass Gender Age SibSp Parch Ticket Fare Cabin Embarked
1 0 3 Male 22 1 0 A/5 21171 7.25 ( ) s
2 1 1 Female 38 1 0 PC 17599 71.2833 C85 c
3 1 3 Female 26 0 0 STON/02.311282 7.925 ( > s
4 1 1 Female 35 1 0 113803 53.1 C123 s
5 0 Male 35 0
3
0 373450 8.05 ,P" N
\ s
6 0 3 Male ( 0 0 330877 8.4583 '
Q
Missii^values
Common M ethods to H and le Missing V a lu e s:
• Deletion: Involves removing entire rows with missing values. While simple, it can lead to loss
of valuable data.
• M ean/M edian/M ode Im putation: Replace missing values with the mean, median, or mode
of the respective feature. This method is sim ple but may distort the original distribution.
. Forward Fill/ B ack w ard Fill: Fill missing values with the most recent non-missing value
(forward fill] or the next non-missing value (backward fill) along the column.
. K-Nearest N eig h b o rs (KNN) Im putation: Predict missing values based on the values of the
nearest neighbors in the feature space.
• Prediction M odels: Use machine learning algorithms to predict m issing values based on
other features in the d ataset This approach can be effective but requires m ore computational
resources.
Let's consider an example dataset with missing values in the "Age" and -Income" columns:
ID Gender Age Income Region
1 Male 35 50000 East
2 Female NaN 6 0000 West
3 Male 45 NaN North
4 Female 30 7 0000 South
■ -':^L<*'
In this example, we have missing values represented as "NaN" in the ’Age' and "Income- columns. To handle
these missing values:
1 Identify Missing Values: Look for cells in the dataset that contain "NaN" or any other placeholder
indicating missing data. In our example, we have missing values in the "Age" and Income columns.
2. Choose Imputation Method: Decide on the imputation method to fill in the missmg values. For
simplicity, let's use mean imputation in this example.
3. Calculate Mean Values: Calculate the mean age and mean income from the available data in the
respective columns.
4. Replace Missing Values: Replace the missing values in the "Age" column with the mean age and the
missing values in the "Income" column with the mean income.
5. Updated Dataset:
import pandas as pd
from [Link] import Simplelmputer
df = [Link](data)
2. H andling Outliers
Outliers are data points that significantly differ from other observations in a dataset. These data points
can skew statistical analyses and m achine learning models, leading to inaccurate results. Outliers can
occur due to various reasons such as measurement errors, data entry mistakes, or genuine extrem e
values in th e data.
E xam ple : In a dataset containing information about individuals, such as their age, it is com m on to
encounter outliers, such as ages above 100 years. While som e individuals may indeed be over 100
years old, extrem e ages can impact statistical analyses and m achine learning models if not handled
appropriately.
Id en tification of Outliers:
• Visual methods like box plots, scatter plots, and histograms can help identify outliers. Statistical
m ethods such as z-scores, IQR (Interquartile Range), and Tukey's method can be used to detect
outliers.
neededtoensurethattherem ovalofoutliersdoesnotbiastheanalysis.
r = S r r = = = 2 E S r
import pandas as pd
import numpy as np
Data Transformation
Data transformation is a fundamental process in data preprocessing that involves modifying the
original data to make it m ore suitable for analysis or modeling. This transformation can help improve
the quality of the data, address issues like skewness or outliers, and enhance the performance of
m achine learning algorithms.
....
, Machine Learning ' '' -’.iip '■ ''■^3^'^':^ ' * I't ■
■‘ ; ‘ -■-^V-?-'’^.'V.^
Data transformation is a crucial step in the machine learning (ML) pipeline because it ensures that the
input data is in a suitable format for modeling, which helps improve the performance and accuracy of the
models. Some key reasons why data transformation is necessary in ML are given below.
• Feature Scaling: Feature scaling is a method used in data preprocessing for machine learning that
involves adjusting the range of the features in the data. This technique is essential because many
machine learning algorithms perform better or converge faster when features are on a relatively
similar scale and close to normally distributed.
Different features in the dataset might have different units and scales. For example, age might
range from 0 to 100, while salary might range from thousands to crores. Algorithms that rely on
the distance between data points, like k-nearest neighbors (KNN) and support vector machines
(SVM), can be biased towards features with larger scales. Transformations like Min-Max scaling
or Standardization help normalize the data, ensuring that each feature contributes equally to the
model's predictions.
• Handling Skewed Data: Many machine learning algorithms assume data is normally distributed.
If the data is skewed, transformations like logarithmic, square root, or Box-Cox can help reduce
skewness, making the patterns in the data more interpretable and accessible to the model
• Encoding Categorical Variables: Machine learning models generally work with numerical data.
Categorical data, such as gender (Male/Female) or state names, need to be converted to numerical
formats using techniques like one-hot encoding or label encoding. This conversion allows
algorithms to process the data effectively.
• Feature Extraction: Transforming data can help in extracting more meaningful features which
might not be captured directly from raw data.
• Improving Model Performance: Proper data transformation can lead to better model performance.
1. N orm alization :
• Normalization is a type of data transform ation thatscales the values of numerical features
to a standard range, typically betw een 0 and 1.
• It helps in bringing all features to a similar scale, preventing certain features from
dominating the model due to their larger magnitude.
• Normalization involves scaling num erical features to a standard range, like transforming
house prices from 100,000 to 2 0 0 ,0 0 0 to a range between 0 and 1. By normalizing data,
features with different scales, such as square footage and price, are brought to a common
scale for fair comparison.
2. S tan d ard izatio n :
• Standardization is another data transform ation technique th at centers the data around a
mean of 0 and scales it to have a standard deviation of 1. It is like converting heights and
weights to z-scores.
• It m akes the data follow a standard normal distribution, which can be beneficial for
algorithm s that assume normally distributed data.
3. Log Tran sform ation :
• Log transformation is applied to skewed data, like converting income values to their
logarithm ic form to handle extrem e values.
• It helps in making the data more symmetrical and reducing the im pact of extreme values,
especially in positively skewed distributions.
4. Encoding Categorical Variables :
• Converting categorical variables into numerical representations through techniques like
one-hot encoding or label encoding is a form of data transform ation.
- One-hot Encoding: Creates a new binary column for each category level.
- Label Encoding: Assigns a unique integer based on the alphabetical ordering of
the categories.
• One-hot encoding transforms categorical variables into binary values, such as converting
"color" categories like red, blue, and green into Os and Is.
• It allows categorical data to be used in machine learning models that require numerical
input.
• E xam p le : Converting categorical columns to numerical columns
Gender Male Female
Male 1 0
Female 0 1
Male 1 0
Male 1 0
Female f 0 1
Male 1 0
Female 0 1
Female 0 1
Let's consider a dataset containing the following information about students' exam scores in two subjects:
Math and English. The Math Score ranges from 0 to 100, while the English Score ranges from 0 to 50.
• Dataset: Student ID, Math Score, English Score
Student ID: 1,2,3,4, 5
Math Score: 85 ,7 0 ,9 0 ,6 5 ,8 0
English Score: 40,30,45,25, 35
m
Machine Learning^
- o b L cav e: [Link] fte exam scores « . a scale between 0 and 1 for bofl. Matt, and Engllsl, scores.
Data T r a n s f o r m a t i o n Steps: n i
. Feature Scaling: Nonnaltee the Math Score and English Score values to a range bet»aen 1.
To normalize the scores, we can use a simple formula;
N orm alized Value = (Value - Min Value] / (Max Value - Min Value)
df = [Link](data)
print("DataFrame before transformation: )
print(df)
])
■r
vPahM^repflwition
cat_pipeline = Pipeline([
('encoder'j OneHotEncoderO)
])
fulljaipeline = ColumnTransformer([
('num', num_pipeline, ['Age', 'Income']),
('cat', catjjipeline, ['Gender', 'Region'])
])
SSIiSIIISSl8
DataFrame before transformation:
Age Gender Income Region
0 35 Male 50000 East
1 28 Female 60000 West
2 45 Male 55000 North
Explanation
In the above code, a sample DataFrame is created with columns for 'Age', 'Gender', 'Income', and 'Region'. The
data is then transformed using ColumnTransformer with separate pipelines for numerical and categorical
columns.
. Numerical Pipeline (StandardScaler): The num_pipeline uses StandardScaler to standardize the
numerical features 'Age' and 'Income'. Standardization involves centering the data around the mean
and scaling to unit variance.
. Categorical Pipeline (O neHotEncoder): The cat_pipeline uses OneHotEncoder to encode the
categorical features 'Gender' and 'Region' into binary vectors. This process converts categorical
variables into a format suitable for machine learning algorithms.
. ColumnTransformer (Full Pipeline): The fulLpipeline combines the numerical and categorical
_______ pipelines to apply the transformations to the respective columns in the DataFrame.----------------------
m m
Transformed Data Display: The transformed data is stored in transformed_data and then converted
into a new DataFrame [Link] with columns for standardized 'Age' and 'Income' as well as one-
hot encoded 'Gender' and 'Region'.
Output Interpretation:
• After appl}ang the transformation process using ColumnTransformer with StandardScaler for numerical
features and OneHotEncoder for categorical features, the data is transformed as follows:
- [Link]: Represents the standardized (scaled) values of the 'Age' column.
- Scaledjncome: Represents tiie standardized (scaled) values of the 'Income' column.
- [Link] and [Link]: One-hot encoded representation of the 'Gender' column
where 'Female' and 'Male' are encoded as binary values.
- [Link], [Link], and [Link]: One-hot encoded representation of the 'Region'
column where 'East', 'North', and 'West' are encoded as binary values.
• Individual 1:
- Scaled_Age: -0.143346
- Scaledjncome: -1.224745
- Gender. Male (encoded as 0.0 for Female and 1.0 for Male)
- Region; East (encoded as 1.0 for East, 0.0 for North, and 0.0 for West)
• Individual 2:
- Scaled_Age: -1.146764
- Scaledjncome: 1.224745
- Gender: Female (encoded as 1.0 for Female and 0.0 for Male)
- Region: West (encoded as 0.0 for East, 0.0 for North, and 1.0 for West)
• Individuals:
- Scaled_Age: 1.290110
- Scaledjncome: 0.000000
- Gender: Male (encoded as 0.0 for Female and 1.0 for Male)
- Region: North (encoded as 0.0 for East, 1.0 for North, and 0.0 for West)
This transformation process standardizes numerical features and converts categorical features into a format
suitable for machine learning algorithms, ensuring that the data is appropriately prepared for model training
and evaluation
m m p b a t a Reduction
Data reduction is a critical step in preparing data for efficient analysis, especially in contexts involving
large datasets or complex models. The process of data reduction involves diminishing the amount of
data that needs to be processed and analyzed without significantly sacrificing valuable information.
’ 3^^;''^'. • Data Prepar
The main reasons why data reduction is important in data processing and machine learning are listed
below;
• Improves Efficiency: Reducing the size of the data set can significantiy decrease the computational
resources required for processing. This leads to faster training times for machine learning models
and quicker execution of data analysis tasks.
• Reduces Storage Requirements: By minimizing the data volume, data reduction techniques
help in lowering storage space requirements. This is particularly important for businesses or
applications where data storage costs are a concern.
• Enhances Model Performance: In machine learning, reducing the number of input features
(dimensionality reduction) helps in removing irrelevant or redundant features, which can improve
the model's accuracy and performance. Techniques such as Principal Component Analysis (PCA)
and feature selection are commonly used to achieve this.
• Mitigates Overfitting: Overfitting occurs when a model learns not only the valid patterns but also
the noise in the training data. By reducing the number of features or the complexity of the data, the
risk of overfitting is reduced, making the model more generalizable to new, unseen data.
• Simplifies Data Visualization: Reducing the number of dimensions or the volume of data can
simplify data visualization, making it easier to identify patterns and trends. Visualizing fewer
variables or data points can help in drawing more meaningful conclusions without the distraction
of noise.
• Improves Data Quality: Data reduction can help in improving the quality of data by focusing
on the most relevant attributes. This can be particularly fmportant in scenarios where the data
includes irrelevant or extraneous information that could lead to poor decision-making.
• Cost-effective Data Management Managing large volumes of data can be costly, not just in terms
of storage, but also in terms of the computational cost required for data processing and analysis.
Data reduction helps in managing these costs more effectively._________________
A dataset contains information about students, including name, age, date of birth, exam scores, study hours,
extracurricular activities, and academic performance. The goal is to predict student grades based on these
features.
Select exam scores, study hours, and extracurricular activities as the most influential features for predicting
student grades. By selecting key features, such as exam scores and study hours, the dataset is reduced to
essential predictors. The simplified dataset improves model performance, interpretability, and efficiency in
predicting student grades. The streamlined dataset accelerates model training and enhances decision-making
for academic performance analysis.______________ _________________________________ ______________ ______
})
# Binary target based on some condition
data['Pass'] = (data['TestScores'] + data['Assignments'] + data['ProjectScore'] / 3
> I50).astype(int)
# Feature scaling
scaler = StandardScaler()
fe a tu re s_ sc a le d = scaler.fit_transform([Link]', axis=l))
# Print results
printC'Original Data Shape:", [Link]( 'Pass', axis=l).shape)
print("Data Shape After Feature Selection:", selected_features.shape)
print("Selected Features:", selected_feature_names.tolist())
print("Data Shape After PCA:", features_pca.shape)
analysis (PCA).
1. imports an d Data Creation : The script starts by importing necessary libraries: numpy, pandas,
and several modules from scikit-learn. It then generates a synthetic dataset with 100 observations of
student data, including various performance metrics like study hours, attendance, and scores. A binary
target (Pass) is computed based on a custom formula to simulate pass/fail outcomes based on test
scores, assignments, and project scores.
2. Feature Scaling : All features are scaled using StandardScaler, which normalizes the data to have a
mean of zero and a standard deviation of one. This is an important preprocessing step, especially for
PCA and many machine learning algorithms that are sensitive to the scale of the input data.
3. Feature Selection : SelectKBest with ANOVA F-test ([Link] is used to select the top 3 features that
have the highest statistical significance in relation to the target (Pass). This method evaluates each
feature's influence on the target variable and picks the most influential ones.
4. PCA for Dimensionality Reduction : PCA is applied to the scaled data to reduce its dimensionality
’ to 2 principal components. This step transforms the data into a new coordinate system, reducing the
number of features while attempting to keep the most significant variance in the data.
Output I n te r p F e te itio n :
. Original D ata Shape: (1 0 0 ,6 ): The original dataset consists of 100 samples, each witii 6 features. These
features include StudyHours, Attendance, Participation, ProjectScore, TestScores, and Assignments
This is the full dataset before any transformations or reductions are applied..-----------------------------
Machine learning . -V ‘
. Data Shape After Feature Selection: ( 1 0 0 ,3 ) : After applying feature selection using the SelectKBest
method with an ANOVA F-test, the number of features in the dataset has been reduced to 3. The dataset
still contains 100 samples, indicating that no data points were removed—only features were reduced
This reduction focuses on retaining only the most statistically significant features in relation to the
target variable (Pass). This helps in simplifying the model, potentially improving model performance
by reducing overfitting, and decreasing the computational load for further processing.
• Selected Features: ['ProjectScore', 'TestScores', ’A ssignm ents']: The three features selected as most
relevant are ProjectScore, TestScores, and Assignments. These were determined to have the strongest
statistical relationship with the student's ability to pass (as defined by the target variable).
This output provides insight into which factors are most influential for student success in this sjmthetic
dataset. For instance, how well a student performs in projects, tests, and assignments are key indicators
of their likelihood to pass, according to the model.
. Data Shape After PCA: (1 0 0 , 2 ) : After applying PCA, the dimensionality of the dataset is further
reduced to just 2 principal components from the originally scaled 6 features. This transformation
results in a new dataset that still has 100 samples, but now each sample is represented by only 2
derived features.
PCA helps in reducing the dimensionality while attempting to preserve as much of the data's variability
as possible. These 2 principal components capture the essence of the dataset's information, reducing
the complexity and enhancing computational efficiency. This is particularly useful for visualization,
further analysis, or as input into machine learning algorithms that may perform better with lower
dimensional data.
F e a tu r e E n g in e e rin g
Feature engineering is a fundamental process in the field of m achine learning where raw data is
transform ed into formatted datasets th at machine learning algorithm s can work with more effectively.
This process involves creating new features from existing data, transforming data into m ore useful
formats, or enhancing the quality of data to improve the accuracy and efficiency of predictive models.
1. F eatu re Creation: This involves creating new variables from existing data to provide
additional insight to the m odels. This might involve com bining features, deriving new metrics
from existing data, or aggregating data over time or space.
Exam ple : For a dataset containing student attendance and grades, creating a feature that
represents the average grade over the past three tests m ight predict future perform ance better
than individual test scores.
2. F eatu re T ran sfo rm atio n : Transforming features to enhance their predictive power or making
them more suitable for models. Common transform ations include normalization, scaling,
appljang mathematical functions like logarithms or exponentials, and more.
Exam ple: In a student dataset, transforming the 'StudyHours' feature from raw hours to
categories such as 'Low', 'Medium', and 'High' based on defined thresholds (e.g., 0-3, 4-6, 7+
hours) can sometimes provide clearer signals for predicting student performance.
Data Preparafi
{ m
The main reasons why feature engineering is important in data processing and machine learning are listed
below:
• Improves Model Perform an ce: Well-engineered features provide a better representation of
patterns in the data, improving model accuracy and performance.
• Reduces Model Complexity: By effectively capturing the underl3dng signals in the data, simpler
models can be used, or models can converge faster on the optimal solution.
• Enhances Data Interpretability: Good features can help to understand the influence and relation
of variables to the prediction outcomes, providing insights into the process.
• Adaptability Across Various Models: Effective feature engineering can make a dataset more
adaptable across different t)T3es of machine learning models, potentially leading to better
performance without changing the underlying algorithms.______________________________________
})
# Adjust GPA based on study hours; assuming more study hours slightly improves GPA
data['AdjustedGPA'] = data['GPA'] + (data['NormalizedStudyHours'] .astype(int) - 2) * 0.1
Data Spitting
Data splitting in machine learning is the process of dividing the data into separate subsets to be used
at different stages of model building and evaluation. The prim ary goal of data splitting is to ensure
that the model trained on one set o f data can generalize well to new, unseen data. This helps avoid
problems like overfitting, w here a model performs well on the training data but poorly on new data.
1. Training Data: This is the largest portion of the d ataset and is used to train the model. The
model learns to identify patterns and make decisions based on this data. Typically, about 70-
80% of the entire dataset is allocated to the training set.
2. Validation D ata: This subset is used to tune the model's hyperparameters and make decisions
’ about which models o r configurations to use. It acts as a check to avoid overfitting on the
training dataset Typically, about 10-15% of the dataset is reserved for validation.
3. T est Data: This is used to evaluate the final model’s performance after it has been trained and
validated. T h e te st set should be a completely independent dataset that the model has not seen
during training or validation. This helps provide an unbiased evaluation of how w ell the model
is expected to perform in the real world. Typically, about 10-15% of the d ataset is used as the
test set.
M e t h o d s o f D ata Sp littin g
1. Random Splitting: This is the most common method where data points are randomly
assigned to the training, validation, and test sets. This method assumes that all data points are
independent and identically distributed.
2. Stratified Splitting: In scenarios where the dataset includes categories that are unevenly
distributed, such as in classification problems with imbalanced classes, stratified splitting
ensures that each class is proportionately represented in the training, validation, and te st sets.
This helps in building a model that is fair and has learned adequately from all classes.
3 . T im e-based Splitting: For tim e-series data, where tem poral patterns and dependencies are
important, data is split based on time. For example, the model may be trained on data from the
past year, validated on the following month, and tested on the month after th a t
The main reasons why data spitting is important in data processing and machine learning are listed below:
• Avoiding Oveifitting: When a model is trained extensively on a particular set of data, there is a
risk that it learns the noise and specific details of the training data to an extent that it negatively
impacts the performance on new data. By using separate training and testing datasets, it is possible
to minimize the risk of overfitting.
• Model Validation: Data splitting allows for a validation set that can be used to fine-tune model
parameters (hyperparameters). This process Is essential f jr identifying the best model settings
because it prevents tweaking the model based on the test set, which could lead to biased assessments
of its effectiveness.
• Assessing Model Performance: A test set, separate from the training data, provides an unbiased
evaluation of a final model’s performance. This is critical for understanding how a model is likely to
perform in practical scenarios, ensuring that the evaluations reflect true predictive performance on
unseen data.
• Improving Model Robustness: By training a model on a diverse training set and validating it on
different subsets of the data, the robustness and reliability of the model are enhanced. Data splitting
ensures that the model can handle variations in the data and accurately predict outcomes across a
range of scenarios.
• Effective Tuning and Comparison: Data splitting facilitates rigorous comparisons between
different models and configurations under consistent conditions. Each model is given the same
opportunity to learn from specific data and validated and tested on identical sets, making
comparisons feir and decisions more informed.
# split the dataset Into training (60%), validation (20%), and test (20%)
, First, split into training and test
[Link], [Link], [Link], y_test ■ train_test_split(x, y. [Link]
. ....................
state=42) # 0.25 x 0.8 = 0.2
Explanation
1. Data P re p a ra tio n : The array X con tain s h)q)othetical pairs o f features Qike study hours and
grades), w hile the array y contains b inary labels indicating pass ( 1 ] or fail ( 0 ).
2. Data S p iittin g :
• The data is initially divided into training (80% of the original data] and test sets (2 0 %
o f the original data) using train_test_split.
• The training data is further sp lit to carve out a validation set, taking 25% of the training
data (which constitutes 2 0 % o f the original dataset). The proportions are managed
to ensure the final split percentages are maintained as intended (60% training, 2 0 %
validation, 20 % test).
3. Output:
• The code prints the sizes of each dataset to confirm the splitting proportions.
• It also prints the actual training, validation, and test sets with their respective features
and labels.
model accuracy.
estimate.
Step 6 : Final Model Training
After Identifying the b e st model and hyperpararaeters, the final model is trained on the en
training dataset to leverage all available data for optimal learning.
E x a m p le : The optimized random forest model, with fine-tuned
validation, is trained on the com plete student dataset to maximize its predictive capabilities.
Example 1 A Pji;hoii Code to Select and Train the Model -;Linear Regression-Model
Objective : The objective of this program is to demonstrate the use of a Linear Regression model to
predict students' performance based on the number of study hours. By training the model on dummy data
representing study hours and corresponding scores, the program aims to provide users with a tool to input
study hours and receive a predicted score within the valid range of 0 to 100. This program serves as a simple
educational example to showcase the prediction capabilities of a machine learning model in the context of
student performance prediction.
# Im port n ecessary l i b r a r i e s
im port numpy as np
from sk le a rn .lin ear_m o d el im port L inearR egression
2 Select and Train Model: A Unear Regression model Is selected and trained on the dummy data [Link]
' the 'fitO' method. The model learns the relationship between study hours and scores.
3 Accept U s e r I n p u t a n d Predict O u t c o m e :
. The code enters a loop «here It prompts the user to Input the number of study hours.
. lftheuserlnputlsnegative,am [Link],ve.
. The model then predlctsthe student's score based on the input study hours using the [Link]
. TheTedlcted score Is constrained to be between 0 and 100 using the 'm a.0' and 'n-inO'
functions.
Finally, the program displays the predicted score for the input study houj;s.
rc ra ^ tr— :^^^
Bedrooms,Bathrooms,SquareFootage,Location,Pnce
3,2,1500,1,3000000
4,3,2000,2,4500000
2,1,800,3,2250000
3,2,1600,1,3200000
4,3,2500,2,4750000
5,4,3000,3,5000000
3.2,1800,1,3500000 _______________ _
2,1,1000,3,2100000
3,2,1700,2,3300000
4,3,2400,1,4100000
2,1,850,3,2000000
5,4,2900,2,4800000
3,2,1400,1,3100000
4,2,2200,3,4300000
2,1,900,1,2150000
5,4,2800,2,4600000
3,2,1900,3,3400000
4,3,2100,1,4000000
5,3,2750,2,4950000
2,1,1100,3,2300000
3,2,1600,1,3150000
4,3,2300,2,4200000
5,3,2650,3,4850000
3,1,1300,1,2750000
4,2,2200,2,4150000
2,1,950,3,2050000
5,4,3100,1,5000000
3,2,1450,2,3250000
4,3,2350,3,4450000
2,1,1200,1,2400000
Step 2: Write a Python code to read the CSV file, select the model and train the model to predict the house
price. Save this file as "[Link]"
import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from [Link] import RandomForestRegressor
from [Link] import mean_squared_error
# Create a DataFrame for the input features with appropriate column names
user_features = [Link]({
'Bedrooms': [bedrooms],
'Bathrooms': [bathrooms],
'SquareFootage': [square_footage],
'Location': [location]
})
C:\ML_Projects>python [Link]
Model trained and evaluated. RMSE on test set: 252114.48
T^^^Da^Praparafjon
1 0
C:\ML_Projects>python [Link]
Model trained and evaluated. RMSE on test set: 252114.48
P M « P T T iP in a g
Introduction
Types o f Supervised Machine Learning
Some Sample Datasets
K-Nearest Neighbors (K-NN) Algorithm
Linear Models
Naive Bayes Classifiers
Decision Trees
Review Questions
M B Introduction
Supervised machine learning is .earn from S e s X t "s
learning. Supervised maclime leammg computer with labeled examples
r : . g ?.- r . = :r r ,r .
z ;r r d :2 T h f m “a c " S ^ ^ ^ ^
o t e Z l t applies the same concept as a .tudent learns in the superv,s,on of the teacher,
r m i t c h a p " we will describe supervised learning in more detail and explam several popular
supervised learning algorithms. __ _________
B E T Types r>f Supervised Machine Learning
“ e^sed machine learning aasslflcation and Regression are two h-ndamental types of tasks
_ ^.1 __ K oin cr n rp d irte d :
" Classification
\f Supervised | 1 1
1 Learning J■
Regression
I __
rlAc<; labels while regression tocuses on picuicuiiiB
Classification is about predicting dis •« -,tHnn and regression depends on the nature
X T —
H H || Classitlcation
i“re^rrfrors"Tnrd“ ^^^
In classification problems, the goal is to p red ict the categorical class labels of new or unseen input
data based on past observations. The outpu t variable is a category, such as "spam" or "not spam" for
email classification, "male" or "female" for gen der classification or "cat," "dog," or "horse" for image
classification.
1. Classifying emails as either Spam o r Not Spam: Predicting whether an email is spam (class 1) or not
spam (class 0). The output variable is discrete, taking on two distinct values (0 or 1) representing the
two classes.
2. Classifying Images of Fruits : Classifying images of fruits into categories such as apple (class 0),
banana (class 1), or orange (class 2). The output variable is discrete, with multiple distinct values (0
or 1 or 2 ) representing the three classes (apple, banana, orange).
3. Sentiment Analysis: Analyzing text data to determine the sentiment of a review (positive, negative,
neutral). The sentiment labels (positive, negative, neutral) are discrete categories assigned to the input
text
4. Medical Diagnosis: Predicting the presence of a disease based on patient symptoms and test results.
The diagnosis categories (e.g., disease present, no disease) are discrete labels assigned to the patient
data.
5. Customer Retention: Predicting whether a customer will renew a subscription or not (churn
prediction).
1. D iscrete Output Variables: The term "discrete" in the context o f machine learning refers to
a type of variable that has specific and separate values, as opposed to continuous variables,
which can take any value within a range. Discrete variables are countable and have distinct
categories or values, which cannot be subdivided meaningfully.
Exam ples o f Discrete Outputs :
• G ender Classification: Male, Fem ale
• Loan Approval: Approved, Rejected
• Movie Genre Classification; Action, Romance, Thriller, Comedy, Drama
2. Supervised Learning: Classification is a supervised learning approach, meaning it relies on
labeled training data to learn the relationship between input features and the target classes.
3-4
^ e r r a r b e ^ r T l :^ ” —
e. r :f :d sp ee. - o ^ - =
7. Healthcare and B iom ed ical R esearch: Classification algorithms are used in healthcare for
disease diagnosis, patient risk stratification, and medical image anafysis. By classifying medical
data, healthcare professionals can make accurate diagnoses and treatment decisions.
8 . Customer Segm en tation: Businesses use classification to segment customers into different
groups based on demographics, behavior, or p references. This segmentation helps in tailoring
marketing campaigns, improving custom er satisfaction, and increasing retention rates.
9. Quality Control an d Anomaly D etectio n : Classification is employed in quality control
processes to classify products as defective o r non-defective. It is also used for anomaly
detection to identify unusual patterns or outliers in data, signaling potential issues.
10. Automated Decision-M aking: With the advancem ent of artificial intelligence and machine
learning, classification algorithms are integrated into automated decision-making systems.
These systems can classify data in real-tim e and make autonomous decisions based on
predefined rules and models.
These classification algorithms offer a diverse s e t o f tools for solving a wide range of classification
tasks in machine learning. Each algorithm has its strengths and is suitable for different types of data
and problem domains. Experimenting with these algorithm s and understanding their characteristics
can help in selecting the m ost appropriate approach for a given classification problem.
1. Logistic Regression:
• Description: Logistic regression is a lin ear classification algorithm used for binary
classification tasks. It estimates the probability that a given input belongs to a particular
class.
• Advantages: Simple, interpretable, w orks w ell for linearly separable data.
• Application: Spam detection, custom er churn prediction.
2. Support Vector Machines (SVM):
• Description: SVM is a versatile classification algorithm that finds the optimal hyperplane
to separate classes in the feature space. It can handle linear and non-linear classification
tasks.
• Advantages: Effective in high-dimensional spaces, works well with clear margin of
separation.
• Application: Text categorization, image recognition.
3. Decision Trees:
• Description: Decision trees are non-linear classifiers that recursively split the data based
on feature values to make predictions. They create a tree-like structure of decisions.
Advantages: Easy to interpret, can handle b oth numerical and categorical data.
• Application: Customer segmentation, m edical diagnosis.
Machine Lrarhing'
/ A ^ l ^ a b l e o f l e a r n i n g i n t r i c a t e patterns,suitahleforlargedatasets.
Measuringperf— indassificationin;ivesevalua^ng^^^^^^^
positives CFP), tru e negatives (TN), and false nega ives performance of a
Confusion M atrix : A confusion matrix is a . ^e of an algorithm by displaying the
1. A ccu racy : Accuracy is the most com monly used metric for evaluating classification models.
It m easu res the proportion of correct predictions made by the m od el out of the total number
of predictions. It is calculated as the ratio of the number o f c o r re c t predictions to the total
num ber o f predictions.
Accuracy = (TP + TN) / (TP + FP + TN + FN)
A high accuracy score indicates that the model is making co rre ct predictions most of the time.
However, accuracy can be misleading when the class distribution is imbalanced.
2. P re c isio n : Precision measures the proportion of true positives am ong the instances that the
model predicted as positive. It is calculated as the ratio o f the n u m b er of true positives to the
total nu m b er of instances predicted as positive.
P recision = TP / (TP + FP)
A high precision score indicates that the model is making few er false positive predictions. It is
useful w hen the cost of false positives is high.
3. R e c a ll: Recall measures the proportion of true positives am ong th e instances that are actually
positive. It is calculated as the ratio o f the number of true positives to the total number of
actual positive instances.
Recall = TP / (TP + FN)
A high recall score indicates that the model is capturing a m ajority of the actual positive
instances. It is useful when the cost o f false negatives is high.
4. F I s c o r e : F I score is the harmonic mean o f precision and recall. It provides a balance between
the two m etrics and is particularly useful when the class (jistribu tio n is imbalanced.
F I score = 2 * (p recisio n * recall) / (p re c is io n + recall)
A high F I score indicates that the model has both good precision and recall. It is useful when
both false positives and false negatives are equally important
Example
Suppose we have a binary classification problem where we are predicting whether an email is spam (positive
class) or not spam (negative class). After appljnng a classification algorithm to a set of emails, we can construct
a confusion matrix to evaluate its performance.
Actual/Predicted Predicted Not Spam Predicted Spam
Actual Not Spam True Negative (TN) False Positive (FP)
Actual Spam False Negative (FN) True Positive (TP)
In this example:
• True Negative (TN) : Emails correctly predicted as not spam.
• False Positive (FP): Emails incorrectly predicted as spam (actually not spam).
• False Negative (FN) : Emails incorrectly predicted as not spam (actually spam).
• True Positive (TP): Emails correctly predicted as spam.
Suppose we have 100 emails, out of which 30 are spam and 70 are not spam. After applying a classification
algorithm, we obtain the following confusion matrix:
T
In this example:
• True Negative (TN) = 65 (65 emails were correctly predicted as not spam).
• False Positive (FP) = 5 ( 5 emails were incorrectly predicted as spam).
• False Negative (FN) = 10 (10 emails were incorrectfy predicted as not spam).
• True Positive (TP) = 20 (20 emails were correcdy predicted as spam).
By anal3^ing these values, we can calculate various performance metrics such as accuracy, precision, recall,
and FI score to assess the effectiveness of the classification algorithm in distinguishing between spam and not
spam emails.
1. Accuracy:
Accuracy = (TP + TN) / (TP + TN + FP + FN)
Accuracy = (20 + 65) / (20 + 65 + 5 + 10) = 85 /100 = 0.85 or 85%
The accuracy of 85% indicates that the model is making correct predictions for 85% of the total
instances. This means that the model is performing well in terms of overall classification accuracy
2. Precision :
Precision = TP / (TP + FP)
Precision = 20 / (20 + 5) = 20 / 25 = 0.8 or 80%
The precision of 80% indicates that out of all the instances predicted as positive, 80% of them are
actually positive. This means that the model is making fewer false positive predictions.
3. Recall (Sensitivity):
Recall = TP/(TP+ FN)
Recall = 2 0 / (20 + 10) = 20 / 30 = 0.67 or 67%
The recall of 67% indicates that out of all the actual positive instances, the model is able to capture
67% of them. This means that the model is missing some of the actual positive instances.
4. F I S co re:
F I Score = 2 * (Precision * Recall) / (Precision + Recall)
F I Score = 2 * (0.8 * 0.67) / (0.8 + 0.67)
F I Score = 2 * (0.536) / (1.47) = 1.072 / 1.47 = 0.73 or 73%
The F I score of 73% indicates that the model has a good balance between precision and recall. This
means that the model is making accurate positive predictions while capturing a majority of the actual
positive instances. However, it is important to note that the FI score is lower than the accuracy score,
which suggests that the class distribution may be imbalanced_______________________________
Regression
Regression is a type of supervised learning technique in machine learning where the goal is to predict
continuous or quantitative outputs based on input features. Unlike classification, where the output is
categorical, regression models predict a continuous value.
. ; -'-V -
Regression is a process of finding the correlations between dependent and independent variables.
It helps in predicting the continuous variables such as prediction o f M arket Trends, prediction of
House prices, etc. The task of the Regression algorithm is to find the mapping function to map the
input variable(x] to the continuous output variable(y).
Regression algorithms are used if there is a relationship betw een the input variable and the output
variable. It is used for the prediction o f continuous variables. The goal of regression tasks is to predict
a continuous numbe or a real number. If there is continuity betw een possible outcomes, then the
problem is a regression problem.
~1 ConHnuous O u tp u t V ariables; The term "continuous" refers to output v ariab les that can e
■ any value v,ithin a range. These are quantiflable and can he subdivided into fin e r increments,
which are not restricted to separate categories.
Ejcamples of Continuous Outputs:
. House P ric e P rediction: Predicting property prices based on location, size, and other
features The predicted house price is a continuous variable that re p rese n ts the monetary
value of a house. The output is a real number that can vaiyacross a wide ran ge, depending
on the input features.
. sto c k P ric e Forecasting; Estimating future stock prices based on h isto rical data and
market indicators. Predicted stock prices are continuous values that can fluctuate minute
by minute. Each predicted value is a specific numeric figure that re p rese n ts the stock
price at a future time.
2 Supervised L ea rn in g ; Regression also relies on labeled training data w here e a ch input feature
set i s " “ ed with a continuous output value. This data is used to train the m odel to understand
and predict the relationship between input variables and the continuous ou tcom e.
3. L in e a r vs. N o n -lin ear; The relationship can be linear (simple linear reg re ssio n l
(polynomial regression, logistic regression for binao- outcomes m a
L \ e l and choosing the correct type of regression model depends on th e nature of the
relationship betw een variables.
4 D ecision B o u n d a rie s; in regression, the decision boundary can be con sid ered as the line
or c u ™ that b e st fits the data points in the feature space. This concept is nacre nuanced m
regression as the - f i f direcUy predicts a value rather than categonzing.
5 P erfo rm an ce E valuation M etrics; Performance in regression (asks is a s s e s s ^
focusing on how close the predicted values are to the actual values. Com m on M etrics nclude
Mean Squared E rro r [MSE), Root Mean Squared Error (RMSE). Mean A bsolute Er [ ),
R-squared (Coefficient of Determination).
l’ s im p le L iiiia r R egressio n ; Simple linear regression predicts a resp o nse variable usmg a
l g ! e f e l r e . It assumes a linear relationship between the independent variable and the
target variable. .. j u
Exam ple; Predicting a student's exam score based on the number of hours th e y studied Here,
the number of hours studied is the single feature used to predict the exam sco re.
2 M ultiple L in e a r Regression; Multiple linear regression involves using multiple features
To X t a re s p o n s f variable. It extends simple linear regression to incorporate several
independent variables.
Exam ple; Predicting house prices based on features like square footage, n u m b er of bedroom ,
l o c a J n , and age o f the house. In this case, multiple features are con sid ered to estimate the
selling price o f a property.
' •- ' Supervised teaming^ 3.11
Regression algorithm s a re used to predict continuous values based on input features. Some common
regression algorithm s w ith their descriptions, advantages, and applications:
1. Linear R e g re s sio n :
• D e scrip tio n : Linear regression models the relationship b etw een the independent
variables an d the continuous target variable by fitting a linear equation to the data.
• Advantages: Simple, interpretable, computationally efficient.
• A p p licatio n : Predicting house prices, estim ating sales revenue.
2. Polynomial R eg re ssio n :
• D escrip tio n : Extends linear regression by adding polynomial term s to the model,
allowing it to capture non-linear relationships.
• Advantages: Can model a broader range o f data shapes than lin ear regression.
• A p p licatio n : Situations where the relationship between variables is curved, such as
growth ra te s, trajectories, and other natural phenomena.
3. Ridge R e g re ssio n :
• D escrip tio n : Ridge regression is a regularized form of linear regression that adds a
penalty te r m to the cost function to prevent overfitting by shrinking th e coefficients.
• Advantages: Handles multicollinearity, reduces model complexity.
• A p p licatio n : Stock price prediction, risk analysis.
4. Lasso R e g re ssio n :
• D e scrip tio n : Lasso regression is another regularized linear regression technique that
uses the L I norm penalty for feature selection by shrinking som e coefficients to zero.
• Advantages: Feature selection, interpretable models.
• A p p licatio n : Marketing spend optimization, medical cost prediction.
3.11 l^ ^ to ^ iiy ;^ '«
Regression analysis plays a crucial role in machine learning and data science fo r predicting continuous
outcomes based on input variables. Some key importance and benefits o f regression in machine
learning are:
1 P re d ictiv e Modeling: Regression models are essential for m aking predictions and forecasting
' fixture trend s based on historical data. They help in understanding th e relationship between
input features and tiie target variable.
2. In te rp re ta b ility : Regression models are often easy to interpret, especially linear regression,
as they provide coefficients that indicate the impact of each feature o n the target vanable. This
interpretability is valuable for decision-making and understanding th e driving factors behind
predictions.
3 F e a tu re Selectio n : Regression analysis can help in identifying th e m ost important features
’ tiiat influence the target variable. Techniques like Lasso reg ression can perform featiire
selection by shrinking coefficients to zero, leading to a more con cise and relevant model.
4 . M odel Evaluation: Regression models provide mettles such as RMSE (Root Mean Square
Error) and R-squared to evaluate the performance o f the model. T h e se metrics help in assessing
the accuracy and reliability of predictions.
5. Handling Non-linear Relationships: Techniques like polynom ial regression, decision tree
reg ression , and support vector regression can capture non-linear relationships between
variables, allowing for more flexible modeling of complex data patterns.
6 . R isk A sse ssm e n t: Regression models are widely used in risk assessment and financial
analysis to predict outcomes such as credit risk, stock prices, and insurance claims. They help
in quantifying and managing risks effectively.
7. Optimization and Decision Making: Regression models can b e used for optimization tasks,
such as determining the optimal pricing strategy, resource allocation, or process improvement
based on predictive insights.
8. Scalability and Efficiency: Regression algorithms can handle large datasets efficiently making
them su itab le for real-world applications with high-dimensional data and a large number of
observations.
9. Generalization: Well-constructed regression models can generalize well to unseen data,
m aking th e m reliable for making predictions on new instances o r in production environments.
10. Versatilily: Regression techniques are versatile and can be applied to various domains such
as h ealth care, marketing finance, and engineering, making them a fundamental tool in data
analysis and decision support systems.
When evaluating the performance o f regression models, various m etrics are used to assess how well
tiie model p re d icts continuous outcomes. Here are some common perform ance evaluation metrics in
regression:
1. Mean Squared Error (MSE):
• D e scrip tio n : MSE calculates the average of the squared differences between predicted
v alu es and actual values.
• Formula: MSE = E(yi - yO^ / n
• Advantages: Penalizes large errors, provides a m easure o f model accuracy.
• Disadvantages: Sensitive to outliers.
2. Root Mean Squared Error (RMSE):
• D e scrip tio n : RMSE is the square root of the MSE, providing a measure of the standard
d eviation of the residuals.
• Formula: RMSE = V(I(yi - yO^ / n)
• Advantages: Interpretable in the same units as the targ et variable.
• Disadvantages: Same as MSE, sensitive to outliers.
3. M ean A b solu te Error (MAE):
• D e scrip tio n : MAE calculates the average of the absolute differences between predicted
valu es and actual values.
• Formula: MAE = EJyi - ^i| / n
Sis:
>:MachrneiMming -‘r/.. - f e ? ■-
In Regression, we try to find the b est fit line, which can In Classification, we tiy to find the decision boundary,
predict the output more accurately. which can divide the dataset into different classes.
Regression algorithms can be used to solve the Classification Algorithms can be used to solve
regression problems such as Weather Prediction, classification problems such as Identification of spam
House price prediction, etc. emails. Speech Recognition, Identification of cancer
cells, etc.
Regression algorithms can be further categorized into Classification algorithms can be divided into binary
linear regression (example, Ordinary Least Squares) classifiers (example. Logistic Regression, Support
and non-linear regression (example. Decision Trees, Vector Machines) and multi-class classifiers (example,
Random Forest). Random Forest, K-Nearest Neighbors).
Common evaluation metrics for regression include Evaluation metrics for classification include Accuracy,
Mean Squared Error (MSE), Root Mean Squared Error Precision, Recall, FI Score, and Area Under the ROC
(RMSE), and R-squared. Curve (AUC-ROC).
Regression models are often more interpretable as Classification models may focus more on decision
they provide coefficients indicating the impact of input boundaries and class predictions rather than feature
features on the output importance.
Some regression algorithms may be computationally Certain classification algorithms can handle large
intensive for large datasets due to the complexity of datasets efficlentiy, making them suitable for scalable
fitting curves. applications.
participation,culturaLevents,overall_grade
1,85,78,92,95,Yes,No^
2,72,65,80,88 ,No,Yes,B
3,90,85,88,92,Yes,Yes,A
4,78,70,75.85,No,No,C
5,95,88,94,98,Yes,Yes,A
6,68,72,70,80,No,No,D
7,82,75,85,90,Yes,Yes,B
8,88,82,90,94,Yes,No,A
9,75,68,72,82,No,Yes,C
10,93,90,87,96,Yes,Yes,A
11,70,62,78,86,No,No,C
12,84,80,86,9l,Yes,Yes,B
13,77,72,74,84,No,No,C
14,91,86,89,93,Yes,YesA
15,73,68,70,81,No,Yes,C
2 Stock M arket T ren d s Dataset; The dataset focuses on the stock m arket trends of Indian
companies. This dummy dataset includes information on various com panies such as
stock prices, trading volumes, market indices, P /E ratios, dividend yields, and sector-wise
performance. It can be used for analyzing stock m arket trends, exploring co rre la o m s widi
L n o m i c indicators, and practicing predictive modeling for future trend predictions. TVpe the
below data and Save the file as stockm [Link]._________________ ___________ ________
company_name,stock_price,trading_volume,market_index,pe_ratio,dividend_yield,sector.
performance
Company A,1200,50000,15000,25,2.5,0utperforming
Company £,700,30000,11000,15,1.5,Underperforming
Company F,850,40000,13000,19,1.9,Neutral
Company G,1050,45000.14000,21,2.1,0utperforming
Company 1,1000.43000,12500,20,2.0,Neutral
Company 1,800.38000.11500,17,1.7,Underperforming
Company K,950,44000,13000,19,1.9,Neutral
Company L,1150.47000,14000,23,2.3,0utperforming
Company M,7 2 0 3 1 0 0 0 ,10500 ,16 ,1 .6,Undeq3erfomiing
Company N,880.39 000,12000,18,1.8,Neutral
Company 0,1020.46000,13s00,21,2.1,0utperforming
3. W eadier P a tte r n s D ataset: This dummy dataset includes information on w eather conditions
such as tem p eratu re, humidity, wind speed, and the weather condition for each day. It can be
used for anal3^ in g w eather patterns, studying the impact of weather on various activities, and
building pred ictive models for weather [Link] the below data and Save the file as
w ead ier_d [Link].
The K-Nearest N eighbors (K-NN) algorithm is a fundamental machine learning technique that operates
on the principle o f sim ilarity. It is a versatile and intuitive method used for b oth classification and
regression tasks. In K-NN, the classification of a new data point is determined by th e majority class of
its nearest neighbors in th e training dataset Similarly, for regression tasks, the algorithm predicts the
value of a new data p o in t based on the average o f the target values of its clo sest neighbors. K-NN is
known for its sim plicity and effectiveness in scenarios where the underlying d ata distribution is not
well-defined or w hen lin e a r separation is not feasible.
Vjlhachine learning
Suppose there are two categories, i.e., C ategory A and Category B, and we have a new data point
x l, so this data point will lie in which o f these categories. To solve this type of problem, we
need a K-NN algorithm. With the help o f K-NN, we can easily identify the category or class o f a
particular dataset. Consider the below d iagram ;
\
9 O
Category B Category B
Example
In the scenario where we have an image of a creature that exhibits similarities to both cats and dogs, but we
need to determine whether it belongs to the cat or dog category, the K-Nearest Neighbors (K-NN] algorithm can
be employed. By leveraging its similarity-based approach, the K-NN model will analyze the features of the new
image and compare them to existing images of cats and dogs in the dataset. Based on the closest resemblance
or similarity to features of known cat and dog images, the algorithm will classify the new image into either the
cat or dog category. This process of identifying the category of the creature in the image showcases how K-NN
utilizes the concept of similarity to make accurate classifications in machine learning tasks.
KNN Classifier
The class with the highest frequ en cy among the K neighbors is selected as the predicted class
for the new data point.
5. For Regression:
In regression tasks, the algorithm calculates the average (o r weighted average) of the targ et
values o f the K nearest neighbors.
This average value serves as th e predicted value for the new data point, providing a continuous
output rather than discrete cla sse s as in classification.
K -N N A lg o r ith m
The K-NN working can be explained on th e basis of the below algorithm :
• As a general guideline, start with small values of K (e.g., K=3 or K=5] and gradually increase the
value while monitoring the model's performance. This iterative approach can help in finding
an optimal K value that balances bias and variance in the model.
Let's consider an example with a dataset for classifying fruits based on two features: sweetness and acidity. We
will use the K-Nearest Neighbors (KNN) algorithm to classify a new fruit based on its sweetness and acidity
values.
• Example Training Data:
Fruit 1: Sweetness 8, Acidity 3 - Type: Apple
Fruit 2: Sweetness 6, Acidity 2 - Type: Apple
Fruit 3: Sweetness 3, Acidity 7 - 13^)6: Lemon
Fruit 4: Sweetness 2, Acidity 8 - Type: Lemon
• New Data Point: Sweetness 5, Acidity 4
Find the type of fmit using KNN algorithm.
Solution: . c j-
. For the new data point with Sweetness 5 and Acidity 4. the KNN algorithm would classify it by finding
its nearest neighbors based on Euclidean distance in the 2D feature space of Sweetness and Aadity, and
then determining the majority class among those neighbors to assign the type of fhnt.
. Let's choose K = 3 for this example.
. Calculate the Euclidean distance between the new data point and all data points in the training set
Distance from (5,4) to Fruit 1 (8,3) : sqrtCC8-5)> * (3-4)>) = sqrt[9 1 1 ) = sqrt(lO) » 3.16
DisUnce from (5,4) to Fnilt 2 (6,2) : sqrtC(6-5)= * (2-4)’) = sqrt(l t 4) = sqrt(5) = 2.24
Distance Irom (5,4) to [Link] 3 (3,7) : sqrtCC3-S)' + = *<lrt(4 + 9) = sqrt(13) » 3.61
Distance from (5,4) to Fruit 4 (2,8) : sqrt((2-5)» t (8-4)>) = sqrt(9 * 16 ) = sqrt(25) = S
13,12,15
. Select Majority Class:
Among these nearest neighbors; 12,13, and 15 are PASS.
Since all 3 nearest neighbors are classified as PASS, the majority class is PASS.
Therefore, based on the majority class rule, the new student with Hath 6 , CS 8 , and English 6 will be con-ectly
classified as PASS using the KNN algorithm with K = 3
# User input for sweetness and acidity values of the new fruit
sweetness = float(input ("Enter sweetness value (1-10): "))
acidity = float(input ("Enter acidity value (1-10): "))
X_new = [Link]([[sweetness, acidity]])
print(f"The predicted price for the fruit with sweetness {sweetness} and acidity
{acidity} is: Rs.{predicted_price[0]:.2f}")
Output
Enter sweetness value (1-10): 5
Enter acidity value (1-10): 4
The predicted price for the fruit with sweetness 5.0 and acidity 4.0 is: Rs.76.67
Explanation
T raining Data: The example training data consists o f arrays representing sw eetness and
acidity values of fruits (X_train) and their corresponding prices (y_train). This data is used to
train the KNeighborsRegressor model.
U ser Input: The code prom pts the user to input sw eetness and acidity values for a new fruit.
These values are stored in an array X_new for predicting the price of the new fruit based on
the trained model.
M odel Initialization: The code initializes a KNeighborsRegressor object with n_neighbors=3,
setting up the KNN algorithm to consider the 3 nearest neighbors when making predictions
for the new student.
M odel Training: The model is then fitted with the training data (X_train, y_train) to learn the
relationships between fruit characteristics and prices.
P rediction and Output: The model predicts the price o f the new fruit (X_new) using the
predictQ method. The predicted price is displayed._______________________________________
3 . Retail. KNN is utilized in retail for custom er segm entation, personalized recommendations,
and market basket analysis. By identifying sim ilar customer profiles or recommending
products based on past purchases, KNN enhances th e shopping experience and boosts sales.
4 . Social Media: Social media platforms leverage KNN for friend recommendations, content
filtering, and sentim ent analysis. By analyzing u ser interactions and preferences, KNN suggests
connections, filters news feeds, and categorizes u se r sentiments to enhance user engagement
5. Environmental S cie n c e : In environmental science, KNN is employed for tasks such as species
classification, pollution monitoring, and clim ate modeling. By analyzing environmental data
and patterns, KNN helps researchers predict sp ecies distribution, detect pollution hotspots,
and model climate changes.
o n Linear Models
Linear models are a fundamental class of algorithm s in supervised machine learning that make
predictions by computing a linear combination of the input features. In a linear model, the relationship
betw een the input features and the target variable is represented as a linear function.
T h e general form of a linear model can be expressed as: Y = C„ + C X + _______ + C X
O i l * n n
In the formula, Y and X^,X^..... represent the variables in th e dataset C^, Cj.™C„ are the regression
coefficients that we estimate from the dataset
1. Y (D ep end entV ariable):
• Y is the dependent variable, also known as the response variable or targ et variable. It's
what you are trying to predict or explain.
• In practical term s, Y could be som ething like the price of a house, the weight of an
individual, the mileage of a car, or any o th er variable that depends on other factors.
2 . X (Independent V ariab les);
• Xj, Xj.....X^ are the independent variables, also known as predictors o r explanatory
variables. These are the variables we use to predict Y.
• Each X, represents a different feature or characteristic. For example, in a model predicting
house prices, X^ might represent the size o f the house, X^ might represent the number of
bedrooms, X3 might represent the age o f the house, and so on.
3. C (Coefficients):
• Cg, Cj are the coefficients or param eters o f the model. They quantify the relationship
between each independent variable and the dependent variable.
• C„ is a special coefficient known as the in te rcep t It represents the expected value of Y
when all the X variables are equal to zero.
• CjC^....Cj^are the slopes for the respective X variables. They represent how much Y is
ej^ected to change with a one-unit change in the corresponding X variable, holding all
< other variables constant.
Linear models are characterized by their simplicity and interpretability, making them widely used in
various machine learning tasks. These models are efficient, easy to implement, and provide insights
into the importance of different features in making predictions.
• Example 1 : Consider a housing price prediction task where the goal is to predict the price of a
house based on features like area, number of bedrooms, and location. Alinear regression model
can be trained to estimate the house price by learning the coefficients for each feature and an
intercept term.
• Example 2 : Predicting student scores based on study hours. Gwen the number of hours a
student studies, a linear regression model can predict the exam score.___________ ^
__________
Algorithm: Linear Regression is a common algorithm for regression tasks th at fits a linear
relationship between features and the target variable.
»1 * ■ ■
■
■
"
''■
’'•9
4
PI'’’'■
•'''
.................../- ' v' - ' - Su p ew s^ ^
2. C lassification w itli L inear M o d e ls: In classification tasks, linear models separate classes by
defining a lin ear decision boundary in the feature space to classify data points into different
categories.
Linear Regression
Linear Regression is a supervised machine learning algorithm used fo r predicting a continuous
numerical output based on one or more input features. The algorithm aim s to find the best-fitting
linear relationship between the input features and the target variable.
1. M odel Representation: In linear regression, the relationship b etw een the input features (X)
and the target variable (Y) is represented by a linear equation o f th e form:
Y = C„ + C,X, + ............ + C„X
W here
• Y is the predicted output,
• Cp ,Cj ,C 2 are the coefficients (weights) to be learned,
• Xj ,X^ r- ,\ are the input features.
2. O b jectiv e: The goal of linear regression is to find the values o f coefficients ,C^ ,C^ ,Cj that
minimize the difference between the predicted values and the actu al target values.
3. T ra in in g : The algorithm learns the optimal values of coefficients by minimizing a cost function,
t 3Apically the Mean Squared Error (MSE), which measures th e average squared difference
betw een predicted and actual values.
4. P re d ic tio n : Once trained, the model can make predictions on n ew data by plugging in the
input features into the learned equation.
Consider a simple linear regression model predicting the price of a car based on its mileage and age:
• Y: Car price (in dollars)
• XjZ Car mileage (in thousands of miles)
• X^: Age of the car (in years)
• C^, Cj, C^: Coefficients to be estimated from data
The model might look something like this:
• Car Price=Cj + x Mileage + x Age
Here, Cj tells us how much the car price decreases for every additional thousand miles driven, and C^ tells us
how much the car price decreases for every additional year of age, with C^ indicating the base price of the car.
The objective is to create a linear regression model that predicts the price of a car, in lakhs of INR, based on
two main factors: its mileage and age. ____________________
Supervised Learning,|T 3.31
Consider a simple linear regression model predicting the price of a house based on its size (in square feet)
and age:
• Y: House price (in lakhs)
• X^: House size (in square feet)
• X^: Age of the house (in years)
• Cg, Cj, C^: Coefficients to be estimated from data
The model might look something like this:
• House Price = C^ + C^* Size + C^ * Age
Given D ata Sample:
Let's calculate the coefficients CO, Cl, and C2 for the linear regression model predicting the house price based
on size and age using the given data sample step by ^ p :
Step 1: Calculate the Mean Values:
Mean Size = (1500 + 1200 + 1800 + 1000 +1400) / 5 = 1380 sq. ft
MeanAge = (5 + 2 + 4 + l + 3 ) / 5 = 3years
Mean Price = (200 + 180 + 220 + 160 + 190) / 5 = 190 lakhs
Step 2: Calculate the Covariance and Variance:
Covariance(Size. Price) = Z((Size - Mean Size) * (Price - Mean Price)) / (n-1)
Covariance(Size, Price) = [(1500-1380)(200-190) + (1200-1380)(180-190) + (1800-1380)(220-190)
+ (1000-1380)(160-190) + (1400-1380)(190-190)]/4 = 6750
Covariance(Age, Price) = I((Age - Mean Age) * (Price - Mean Price)) / (n-1)
Covariance(Age, Price) = [(5-3)(200-190) + (2-3)(180-190) + (4-3)(220-190) + (1-3)(160-190) +
(3-3) (190-190)] /4 = 30
Variance(Size) = Z((Size - Mean Size)^) / (n-1)
Variance(Size) = [(1500-1380)^ + (1200-1380)^ + (1800-1380)^ + (1000-1380)^ + (1400-1380)^] / 4
= 92000
Variance(Age) = Z((Age - Mean Age)^) / (n-1)
Variance(Age) = [(5-3)^ + (2-3)^ + (4-3)^ + (1-3)^ + (3-3)^] / 4 = 2.5
Step 3: Calculate the Coefficients:
Cl = Covariance(Size, Price) / Variance(Size) = 6750 / 92000 = 0.073
C2 = Covariance(Age, Price) / Variance(Age) = 30 / 2.5 = 12
CO = Mean Price - Cl * Mean Size - C2 * Mean Age
= 190 - 0.073 * 1380 - 1 2 * 3 = 190 - 69 - 30 = 53.26 lakhs
Model Training and Coefficients:
. The linear regression model has been trained using the given data of house size, age, and corresponding
prices.
. Intercept (CJ = 53.26 lakhs: CO represents the intercept of the linear regression model. In this case,
it is 91 lakhs. When both the size and age of the house are zero, the predicted house price is 53.26
lakhs. However, in real-world scenarios, this interpretation may not be meaningful as houses cannot
have zero size or age.
• Cj = 0.073: Cl is the coefficient associated with the size (sq. ft) feature in the linear regression model.
For every one square feet increase in the size of the house, the predicted house price is expected to
increase by 0.073 lakhs (Rs.7300), assuming the age remains constant
. =12: C2 is the coefficient associated with the age (years) feature in the linear regression model. For
every one year increase in the age of the house, the predicted house price is expected to increase by 12
lakhs, assuming the size remains constant
Predicted House Price Calculation: Using the derived coefficients from the linear regression model, the
predicted price of a house with 1300 sq. ft size and 2 years old is calculated as follows:________________
• House Price = C, + * Size + * Age
House Price = 53.26 + 0.073 * 1300 +12*2=53.26 + 94.9 +24=172.16 lakhs
This calculation suggests that, under the model derived from the given data, a house with these specifications
(1300 sq. ft size and 2 years old) is predicted to have a price of approximately 172.16 lakhs. This predicted
price reflects the combined effect of the house’s size and age on its overall value in the real estate market
Note: The differences in predicted prices between the manual calculation and the scikit-learn LinearRegression
model can be attributed to various factors such as the model complexity, feature scaling, model assumptions,
data variability, and the handling of the intercept term.
To align the manual calculation with the model predictions, one would need to adjust the manual calculation
methodology to match the assumptions and processes used by the scildt-leam LinearRegression model.
3,34 X ‘’-‘i •tachine Learning
.Explanation
1. Data Preparation : The given data sample consists of house features (size and age) and the target
variable (price). The features (size and age) are stored in the variable X, while the target prices are
stored in the variable y after separating them from the data sample.
2. Model T rain in g: A Linear Regression model is created using scikit-learn's LinearRegression class and
trained on the features (size and age) and target prices from the given data sample. The model learns
the relationship between the features and the target variable during the training process.
3. User Input: The program prompts the user to enter the size and age of a new house for which they want
to predict the price. The user inputs are stored in the variables new_size and new_age after converting
them to floating-point numbers.
4. Prediction : The trained Linear Regression model is used to predict the price for the new house based
on the user-provided size and age. The model's predictQ method is called with the new feature values
to obtain the predicted price for the new data point
5. Output: Finally, the program prints the predicted price for the new house with the given size and age
in a formatted string, displaying the input values and the predicted price in Lakhs. This allows users
to quickly get an estimate of the house price based on the provided features using the trained linear
regression model.
Logistic Regression
One common method for using regression for classification is logistic reg ression . A logistic regression
is actually a classification algorithm that predicts the probability of an ob servation belonging to a
certain class. The logistic regression model uses a logistic function to map th e output to a probability
value between 0 and 1, making it suitable for binary classification tasks.
For example, in a b inary classification problem w here the goal is to predict w h e th e r an email is spam
or not spam, logistic regression can be used to model the probability o f an em ail being spam based on
features such as the presence of certain keywords, email length, or sen d er information. The output
of the logistic regression model can then be interpreted as the probability o f th e email belonging to
the spam class.
In multi-class classification tasks, multinomial logistic regression can be used to predict the probability
of an observation belonging to each class within the dataset This allow s fo r the classification of
observations into multiple categories based on the highest predicted probability.
It's important to note that while regression for classification can be a u seful technique, there are
also dedicated classification algorithms, such as decision ft-ees, su p p ort vector machines, and
neural networks, th at are specifically designed for handling classification ta sk s and may outperform
regression-based approaches in certain scenarios.
Jig Example Binary Classification using Logistic Regression:
Consider a binary classification problem where the task is to predict whether a student will pass (class 1) or
fail (class 0) an exam based on the number of hours studied. The dataset contains the number of hours studied
by each student and whether they passed or failed.
Data:
• Independent Variable (X): Number of hours studied
Dependent Variable (Y): Pass (1) or Fail (0) ____________________________ _
Supervised le a m in g ^ ^ 3 35
3. 3 6 V M ach in e Leqrning' H f ^ ^ W ^ ^ f e »
# Output
print ("Predicted probabilities for each class:")
for i, prob in enuraerate(multi_prediction[0]):
print(fProbability of class {i}: {prob}")
Output
Enter the size of the fruit (0-1): 1
Enter the color of the fruit (0-1): 1
Predicted probabilities for each class:
Probability of class 0: 0.3660619419625619
Probability of class 1: 0.2542368791495831
Probability of class 2: 0.37970117888785493
• Market Basket Analysis: Predicting which products are likely to be purchased together
in retail settings.
3. Finance:
. Risk Assessment: Predicting credit risk, loan default probabilities, insuran ce claim
likelihood, etc.
• Stock Market Analysis: Forecasting stock prices o r identifying trading opportunities.
4. Healthcare:
• Disease Prediction: Using logistic regression to predict the likelihood of a p atien t having
a particular disease based on symptoms and medical history
• Drug Response Prediction: Predicting how patients will respond to different treatm ents
based on their characteristics.
5. Recommendation Systems:
• Collaborative Filtering: Linear models can be used in recommendation system s to
predict user preferences based on historical data.
6. Natural Language Processing (NLP):
• Text Classification: Logistic Regression is commonly used for sentim ent analysis, spam
detection, and text categorization tasks.
Naive Bayes classification is a probabilistic method for categorizing data points based on Bayes'
theorem, which establishes a connection between the present data and existing assumptions or
beliefs. It involves updating p rior beliefs (prior probabilities of class labels) based on the observed
evidence (features of the input data) to make predictions.
It is commonly used for binary and multi-class classification problems. Naive Bayes classificationis
particularly popular in natural language processing tasks like spam filtering and document
categorization.
Naive Bayes classification depends on the principle of Bayes’ Theorem. Before moving to the Naive
Bayes, it is important to know about Bayes’ theorem.
Bayes'Theorem
Bayes' theorem describes th e probability of an event, based on prior knowledge o f conditions that
might be related to the event. In the context of classification, Bayes' theorem is used to calculate
the probability of a class label given the observed features. In the context of classification, it can be
expressed as:
Bayes' Theorem Formula: P(A|B) = (P(B|A) x P(A))/ P(B)
where:
N a iv e B a y e s C la s s ifie r
Naive Bayes is a sim ple and powerful classification algorithm based on Bayes' Theorem with an
assumption o f independence between features. The assumption of independence between features
means that the presence of a particular feature in a class is independent of the presence of any other
feature. This assum ption simplifies the calculation o f probabilities by assum ing th at the effect of one
feature on the class is independent of the presence o f other features.
'5;'jr: V -. - '
Supervised Learn il
Example: Consider a text classification task w here we want to classify em ails as spam or not spam
based on the p resen ce of two features: the words "discount" and "offer". The independence assumption
implies that the occurrence of the word "discount" in an email does not affect the occurrence of the
word "offer" in th e sam e email when determining if the email is spam or not.
The Naive Bayes Classifier is a probabilistic machine learning model that's used for classification tasks. It
is based on Bayes' Theorem with the assumption that means that the presence of a particular feature in a
class is independent of the presence of any other feature.
The Naive Bayes Classifier algorithm is comprised of two words Naive and Bayes:
• Naive: It is called Naive because it assumes that the occurrence of a certain feature is independent
of the occurrence of other features. Such as if the fruit is identified on the bases of color, shape, and
taste, then red, spherical, and sweet fruit is recognized as an apple. Hence each feature individually
contributes to identify that it is an apple without depending on each other.
• Bayes: It is called Bayes because it depends on the principle of Bayes’ Theorem.
It is widely used in text classification, spam filtering, and recommendation systems due to its efficiency
and effectiveness in handling high-dimensional data.
The Naive Bayes Classifier works by calculating the posterior probability of each class label given the
input features using Bayes' Theorem. It assumes feature independence, simplifying the calculation
by considering each feature's contribution to the class probability Independently. By multlpl3 dng the
likelihood of each feature given the class label with the prior probability o f the class, the classifier
orninq>-*v..'--,!'.Sfr^ ■ -. ~, . * --. o'--vs>?i ■■</*)»■.?v
determines the m ost probable class for th e input data. This approach enables efficient and effective
classification, making Naive Bayes a popu lar choice for text classification, spam filtering, and other
machine learning tasks.
Working Principle:
1. Bayes' Theorem: Bayes' Theorem calculates the probability o f a hypothesis (class label) given
the data (features]. Mathematically, it is represented as:
P(A|B) = (P(B|A )xP(A ))/P(B)
where;
o P (A|B) : The probability o f class A given the data B. (Posterior Probability)
o P (BIA) : The Probability o f d ata B given the class A. (Likelihood Probability)
o P (A) : Prior probability o f cla ss A.
o P (B): Prior Probability o f class B.
2. Naive Bayes Assumption:
• Naive Bayes simplifies the computation of P(BIA) by assuming that all features in B (such
as words in an email) are independent of each other given the class A. This assum ption
allows the model to tre a t each feature separately, which simplifies the calculations
drastically.
3. Classification Process:
. Given a set of features X = {x^, x , , ..., x j and a set of class labels C = {c,, c^.....c J , the Naive
Bayes classifier predicts th e m ost probable class label for the input features.
• The classifier calculates th e posterior probability for each class label and selects the class
with the highest probability.
4. Model Training - Calculating Probabilities
• Calculate Prior Probabilities: This is the probability o f each class in the training dataset
like P(A), P(B) etc..
• Calculate Likelihoods P(B,|A): This involves calculating the probability of each feature
B. given each class A.
5. Calculating Likelihood Product
. Given a set of features X = {x^ , x^ ,...x^}, the likelihood of the features given a class is
calculated by multiplying th e probabilities of each independent feature:
The likelihood product P(X|C J is calculated by assuming feature independence:
P(X1CJ = P (x JC J X P (x,| C J X...............X P(xJC,)
• Each term P (xJC J is the probability of feature x, given class C^.
6. Calculating Probability of the features P(X)
• The generic formula for th e total probability of a feature set X in the context o f Naive
Bayes classification is given by:
P(X) = P (X I Cj ) x P ( c J + PC X I c, )xP(c,) + ............ + P( X I c, )x P (c J
.'S u p e r v ise d Learnin g
Where,
o P(X|Cj) is the probability of observing the feature set X given class c ..
o P(Cj) is the p rio r probability of class c..
o k is the total n u m b e r of classes.
7. Calculating Posterior Probabilities for Classification
• For each class C^, calcu late the posterior probability that a given set of features X belongs
to class using the form ula derived from Bayes' theorem :
P(CJX) = (P (X | C JX P (C J3/ P (X )
where:
o P(C^|X) is the p o sterio r probability of class given the features X.
o P(X|CJ is the likelihood of the features given class C^.
o P(C J is the p rio r probability of class
o P(X) is the probability of the features.
8. D ecision Rule:
• The Naive Bayes cla ssifier selects the class label C^ that maximizes the posterior
probability P(CJX).
• Select the class label w ith the highest posterior probability as the predicted class.
Example A Python Code fpy Spam Eniail Qet^gtion using thfiN^ve Bayes classification algorithm
# Import necessary libraries ,,
from sklearn.naive_bayes import MultinomialNB
from sklearn.feature_extraction.text import CountVectorizer
import numpy as np
# Given Data
X_train = [Link](["free money", "click here for free", "lottery", "buy now", "amazing
offer"])
y_train = [Link]([1, 1, 1, 0, 0]) # i for spam, 0 for non-spam
Enter the text of the new email: Buy iphone at amazing offer
The email is classified as: Not Spam
Explanation
1. Data Preparation: The code initializes training data X_train containing email text and y_train
containing corresponding labels (1 for spam, 0 for non-spam). It uses CountVectorizer to transform the
text data into numerical features. This step converts the text data into a matrix of token counts.
2. Model Training: A Multinomial Naive Bayes classifier (MultinomialNB) is instantiated and trained
on the transformed training data (X_train_counts] and labels (y_train). Naive Bayes classifiers are
commonly used for text classification tasks like spam detection due to their simplicity and effectiveness
with text data.
3. User Input and Prediction: The code prompts the user to enter the text of a new email for classification.
The new email text is transformed using the same CountVectorizer instance to convert it into numerical
features (new_email_counts).
4. Classification: The trained classifier predicts the class of the new email by calling predict on the
transformed new email data (new_emaiI_counts). If the predicted class is 1, the email is classified as
"Spam"; otherwise, it is classified as "Not Spam".
D e c is io n T r e e s
Decision tree-based algorithms use a tree-like model to make decisions based on input data. The
tree-like model consists of a series o f nodes that represent decisions or tests on the input data, and
branches that represent the possible outcomes of those decisions or tests. The leaves o f th e tree
represent the final decision or prediction.
The process of building a decision tree-based algorithm involves selecting the best attrib u te to split
the data at each node, based on a m easure of information gain or impurity reduction. T he goal is
to create a tree that is as small as possible while still accurately classifying or predicting th e target
variable.
There are several popular decision tree-based algorithms, including IDS, C4.5, and CART. Each
algorithm has its own strengths and weaknesses, and the choice of algorithm depends on th e specific
problem and data set.
Decision tree-based algorithms are widely used in a variety of applications, including classification,
regression, and feature selection. T hey are particularly useful for problems with a large num ber
of features or complex decision boundaries, as they can capture non-linear relationships and
interactions between features.
One o f the main advantages o f d ecisio n tree-based algorithms is their interpretability. The resulting
tree can be easily visualized and understood, making it useful for explaining the reasoning behind
the model's predictions. However, decision tree-based algorithm s can also be prone to overfitting,
especially when the tree is too large or the data set is noisy. Regularization techniques, such as
pruning or ensemble methods, ca n help to mitigate this issue.
1. Information Gain
In physics and m athem atics, entropy is referred to as the randomness or the im purity in a
system. In information theory, it refers to the impurity in a group of examples.
^ c ii in e le a r n i n g '■ ' ^
An entropy is a m easure of the impurity or random ness of a dataset and it is used in decision
tree algorithms to evaluate the effectiveness o f attributes in partitioning the data into more
homogeneous subsets with respect to the ta rg e t variable. A lower entropy indicates a more
homogeneous subset, while a higher entropy in d icates a more heterogeneous subset.
we follow these steps:
To calculate the information gain for attribute A,
1. Calculate the entropy of the original dataset D.
Entropy (D) = - Zi Pi logj p
where D^, D^,..., Di^are the subsets of D th a t correspond to each value o f attribute A, and
|Dj|, ID jI,..., |D^| are the sizes of those su b sets.
3. Compute the information gain using the formula:
Gain(A) = E ntropy(D ) - Entropy^(D)
The attribute A with the highest inform ation gain, Gain(A), is chosen as the splitting
attribute at a particular node in the d ecisio n tree. This means that the attribute that
provides the m ost reduction in entropy o r th e most effective partitioning of the data is
selected for splitting at that node.
Problem 1: Suppose we have a dataset of 10 examples, each with a binary class label ("Yes" or "No"].
There are 6 examples with a class label of "Yes" and 4 examples with a class label of "No." The entropy
of the dataset is:
Entropy(D) =-6/10 log2(6/10) - 4/10 log2(4/10) = 0.971
This indicates that the dataset is relatively impure or random, with a high degree of uncertainty about
the class labels.
Problem 2: Suppose we have a dataset of 10 examples, each with a binary class label ("Yes" or "No").
There are 10 examples with a class label of "Yes" and 0 examples with a class label of "No." The entropy
of the dataset is:
Entropy(D) = -10/10 log2(10/10) - 0/10 log2(0/10) = 0
This indicates that the dataset is completely pure o r homogeneous, with no uncertainty about the dass
labels.
Problem 3: Suppose we have a dataset of 10 examples, each with a binary class label ("Yes" or "No").
There are 5 examples with a class label of "Yes" and 5 examples with a class label of "No." The entropy
of the dataset is:
Entropy(D) = -5/10 log2(5/10) - 5/10 log2(5/10) = 1
This indicates that the dataset is relatively impure or random, with a high degree of uncertainty about
the class labels.
Let's calculate the information gain step by step for this problem.
1. Calculate the entropy of the entire dataset based on the class label "PlayTennis :
-f There are 6 examples with TlayTennis=Yes" and 4 examples with "PlayTennis=No".
-f The entropy of the dataset is:
♦ Entropy(D) = -6/10 log2(6/10) - 4/10 log2(4/10) = 0.9709
2. Calculate the information gain for the attribute "Outlook":
There are 4 examples with "Outlook=Sunny", 3 examples with "Outlook=Overcast", and 3 examples
with "Outlook=Rainy".
-f Calculate the entropy of each subset based on the class label PlayTennis :
♦ Subset with "Outlook=Sunny":
^ There are 2 examples with "PlayTennis=No" and 2 examples with "PlayTennis=Yes".
The entropy of this subset is;
A Entropy(Dl) = -2/4 logZ(2/4) - 2/4 log2(2/4) = 1
♦ Subset with "Outlook=Overcast":
^ There are 0 examples with "PlayTennis=No" and 3 examples with "PlayTennis=Yes".
•/ The entropy of this subset is:
Entropy(D2) = -0/3 lcg2C0/3) -3/3 log2(3/3) = 0
♦ Subset with "Outlook=Rainy":
•y There are 2 examples with "Pla5 ^ennis=No" and 1 example with PlayTennis=Yes .
The entropy of this subset is:
^ Entropy(D3) = -2/3 log2(2/3) -1/3 log2(l/3) = 0.9183
Calculate the weighted average entropy of the subsets:
Weighted Average Entropy=(4/10 * Entropy(Dl)) + (3/10 * Entropy(D2)) + (3/10 * Entropy(D3))
= (4/10 * 1) + (3/10 * 0) + (3/10 * 0.9183)
= 0.5509
-f Calculate the information gain of the attribute "Outlook":
Information Gain(Outiook) = Entropy(D) - Weighted Average Entropy
= 0.9709 - 0.5509
= 0.4200
3. Calculate the information gain for the attribute "Temperature":
There are 4 examples with "Temperature=Hot", 2 examples with "Temperature=Mild", and 4 examples
with "Temperature=Coor'.
-f Calculate the entropy of each su b set based on the class label "PlayTennis :
2. Gain Ratio
The Gain Ratio m easures the effectiveness o f a particular attribute in classifying the data. It takes into
account both the information gain and the sp lit information, providing a more balanced assessm ent
of the attributes usefulness in the decision-m aking processof building a^iecision-tFeerThe^Gain Ratio
is a metric used in decision tree algorithms, particularly in the C4.5 algorithm.
By using the Gain Ratio, decision tree algorithm s can make more informed decisions about which
attributes to use for splitting, leading to m ore accurate and generalizable models.
3 .5 4 ^ '.
Gain(A)
The Gain Ratio is defined as: J " Splitlnfo^ (D)
W here:
• GainCA) is the information gain of attribute A on d ataset D.
• Splitlnfo^(D) is the spUt inform ation of attribute A on d a ta se t D.
V |D. I
fl DJl
The Splitlnfo^(D) is calculated as: SplitInfoA(D) = - ^ - j ^ x l o g 2
1D|
W here:
. |Djl/lDl acts as the weight o f the jth partition.
• V is the number of discrete values in attribute A.
Example“
Understanding Gain Ratio
_______ ' -------------
Suppose we have a dataset of students with two attributes: "Study Hours" and "Pass/Fail". We want to build a
decision tree to predict whether a student will pass or fail based on the number of study hours.
1. Information Gain:
^ Information Gain measures how much the "Study Hours" attribute helps in predicting the "Pass/
Fail" outcome.
Higher Information Gain indicates that "Study Hours" is more useful for making decisions in the
decision tree.
2. Split Information:
Split Information measures the uncertainty caused by different splits on the "Study Hours
attribute.
^ It considers the number of study hour ranges and how many students fall into each range.
3. GainRatio:
The GainRatio is the ratio of Information Gain to Split Information.
^ It balances the usehilness of "Study Hours" with the potential uncertainty introduced by its
different splits.
For example, if splitting the data based on "Study Hours" leads to a high Information Gain but also mtroduces
a lot of uncertainty due to many different study hour ranges, the GainRatio will help in evaluating whether the
split is worth it
In this way the GainRatio guides the decision tree algorithm in choosing the most effective attributes for
making decisions, leading to more accurate predictions.______________ ________ _____________________
3
• Ppaii - g (3 students failed out of 5)
Calculate Entropy(D):
'2^ '2 ' '3^ '3 '
Entropy(D) = - •log2 •10g2
.5 , .5 , ,5,
Using base-2 logarithms:
'2^
Entropy(D)=- (-2 .3 2 2 )- (-1.585)
15,
Entropy(D3»0.971
Step 3: Calculate Information Gain (Gain(Study Hours))
Information Gain measures how much the "Study Hours" attribute helps in reducing uncertainty
(Entropy) in predicting "Pass/Fail." The formula for Information Gain is as follows;
|DJ
Gain(Study Hours) = Entropy(D) - ^ . Entropy(D,)
\m )
Where:
• k is the number of possible values of the attribute.
• D|is the subset of data where the attribute has the i'*' value.
In our case, "Study Hours" has four possible values: 1, 2, 3, and 4.
For "Study Hours = 1" (D l):
• Entropy(Dl) = 0 (Since there's only one student, the entropy is 0.)
For "Study Hours = 2" (D2):
» Entropy(D2) = 1______
For "Study Hours = 3" p 3 ) :
• Entropy[D3) = 0 (Since there's only one student, the entropy is 0.)
For "Study Hours = 4" (D4):
• Entropy(D4)=0 (Since Aere's only one student, the entropy is 0.)
Now, calculate Gain(Study Hours):
Gain(StudyHours)=0.971 - i . 0 + - . l + - . 0 + i .0
5 5 5 5
Gain(StudyHours)= 0 .9 7 1 --
D
Gain(StudyHours) = 0.971-0.4
Gain(StudyHours]=0.571
Step 4: Calculate Split Information (SplitInfo(Study H ours))
Split Information measures the uncertainty introduced by different splits on the "Study Hours
attribute. The formula for Split Information is as follows:
SpIitInfo(StudyHours) = - X t , ^
Splitlnfo(StudyHours) ~ 1.921
Step 5: Calculate GainRatio (GainRatio(Study H ours))
GainRatio is the ratio of Information Gain to Split Information;
, „ ^ Gain(StudyHours)
GainRatloCStudyHours)=j^|.^n,„([Link])
0.571
GainRatio(StudyHours) =
1.921
GainRatio(StudyHours)w0.297
So, the GainRatio for the "Study Hours" attribute is approximately 0.297. This value helps us evaluate the
usefulness of "Study Hours" in making decisions in the decision tree while considering all possible values of
"Study Hours," including "4." Higher GainRatio values indicate that the attribute is more valuable for splitting
the data while considering the potential uncertainty introduced by different splits.______________ .
The Gini index, used in the CART algorithm, is a m easu re o f impurity or uncertainty in a dataset It
is commonly used to evaluate the quality of a particu lar sp lit in a decision tree. The Gini index for a
dataset D is calculated based on the probabilities o f e a ch class in the dataset
Another decision tree algorithm CART (Classification an d Regression Tree) uses the Gini method to
create split points.
Glni(D) = l - X : . P | '
Where p. is the probability th at a tuple in D belongs to d a s s C,.
The Gini Index considers a binary split for each attribute. We can compute a weighted sum of the
impurity of each partition. If a binary split on attribu te A partitions data D into D1 and D2, the Gini
index of D is;
In interacting with a discrete-valued attribute, the splitting attribute is chosen from the subset that
5 aeldsthe lowest gini index for the given value. W hen dealing with continuous-valued characteristics,
the approach is to consider every pair of neighboring values as a potential split point; the splitting
point is determined by selecting the point with the low est gini index.
AGini(A) = Gini(D) - GiniJD)
The splitting attribute is determined by taking the attribute with the lowest Gini index.
GiniIndex
Let's understand a simple example of calculating the Gini Index and selecting the splitting attribute using a
decision tree. We'll start with a basic dataset and explain each step in detail.
Step 1: The Dataset
Imagine we have a dataset of fhiits with two attributes: "Color" and "Class" (whether the fruit is
"Apple" or "Banana"). We want to build a decision tree to classify these fhiits based on their color.
Sample dataset:
C olor ■■■' Class :
Fruiti Red Apple
Fruit2 Yellow Banana
Fruits Red Apple
Fruit4 Yellow Banana
Step 2: Calculate Gini Index (Gini(D))
The Gini Index measures the impurity or uncertainty in the dataset D. The formula for Gini Index
is as follows:
Calculate-Gini(D):
Gini(D) = l -
Simplify:
^ Jsa^ l^ .fflach in e learninjir
1 1^
Gini(D) = l - ---1---
4 4
Gini(D) = l - -
Gini(D] = i
Calculate Gini(Dl):
(r
Gini(Dl) = l -
Simplify: Gini(Dl) = 1 - (1 + 0)
Gini(Dl) = 0
For "Color = Yellow" (D2):
2
• Peanana = bananas out o f 2 yellow fruits).
2
Calculate Gini(D2):
(2 \
Glni(D2) = l - +
Simplify:
Gini(D2) = 1 - (0 + 1)
Gini(D2) = 0
Step 4: Calculate AGini(Color)
Now, calculate the reduction in impurity (AGini) for the "Color" attribute:
f l Di O
AGini(Color)= Gini(D) - Gini(D)
i D| j
Where:
|D||is the size of subset Dj created by the split (example, "Red" and "Yellow" subsets).
|D1 is the size of the original dataset D.
Gini(D^ is the Gini Index for subset D, calculated in Step 3.
AGini(Color) = ^ - - .0 +
4
Simplify:
AGlni(Color) = " (0 + 0)
AGini(Color) = ^
So, in our decision tree, we would split the data based on the "Color" attribute, specifically into "Red" and
“Yellow" subsets, because it results in the greatest reduction in impurity (Gini Index). This process continues
recursively to build the decision tree.
IDS Algorithm
The IDS algorithm is a classic decision tree algorithm that is used to build a decision tree from a
dataset. T he goal of the algorithm is to create a tree that can predict th e class label of instances based
on the attrib u te values. The algorithm selects the best attribute a t each node of the tree based on
inform ation gain, which measures the effectiveness of an attribute in classifying the training data.
The IDS algorithm works by recursively partitioning the dataset into subsets based on the values
of the attrib u tes. At each node of the tree, the algorithm selects the attribute that provides the m ost
inform ation gain, which is a measure o f how much the attribute red uces the uncertainty about the
class labels. T he attribute with the highest information gain is chosen as the splitting attribute for the
node.
The inform ation gain is calculated using the entropy of the d ataset b efo re and after the split. Entropy
is a m easure o f the impurity of a set of examples, where a set is consid ered pure if all examples belong
to the sam e class.
m
1. Simplicity: ID3 is relatively simple to understand and implement, making it accessible for beginners
and useful for educational purposes.
2. Handles Categorical Data: IDS is well-suited for handling categorical attributes and class labels,
making it effective for classification tasks involving non-numeric data.
3. Interpretability: The resulting decision tree is easy to interpret and visualize, allowing users to
understand the decision-making process and the rules used for classification.
4. Feature Selection: 1D3 inherently performs feature selection by choosing the most informative
attributes for splitting, which can help in identifying the most relevant features for classification.
C4.5 Algorithm
The C4.5 algorithm is a popular decision tree algorithm developed by Ross Quinlan as an extension
of the earlier IDS algorithm. It addresses some of the limitations of IDS and introduces several
improvements.
C4.5 is a decision tree based algorithm used for constructing decision trees from a dataset Its
primary purpose is to perform classification tasks by creating a decision tree that can be used to
make predictions about the class label of new instances based on their attribute values.
The C4.5 algorithm uses a top-down, greedy approach to construct a decision tree. It employs the
concept of information gain and entropy to determine the best attribute for splitting at each node of
the tree. The algorithm aims to create a tree that maximizes the information gain at each split, leading
to more accurate and efficient classification.
'M
The C4.5 decision tree algorithm, an improvement over IDS, introduces several key enhancements
and strategies to build more accurate decision trees. The key improvements are listed below:
1. Handling Missing Data:
C4.5 handles missing data by simply ignoring it during the calculation of gain ratio. When
building the decision tree, the gain ratio is calculated based only on the records that have
a value for the attribute in question. To classify a record with a missing attribute value, the
algorithm can predict the value for that item based on the known attribute values of other
records.
JSg Example Missing Data
Let's consider a simple example where we have a dataset for predicting whether a person will buy a
product based on their age and income. However, some entries have missing values for the income
attribute.
Suppose we have the following dataset:
Age Income Will Buy
25 30,000 Yes
35 50,000 No
45 Yes
30 40,000 No
In this example, the income value for the third entry is missing. When building the decision tree using
C4.S, the gain ratio for splitting based on income would be calculated based only on the available
records (1", 2"“* and 4 “'’ entries). If a new record with a missing income value needs to be classified, C4.5
can predict the income value for that item based on the known attribute values of other records in the
dataset __________________ _
2. Continuous Data:
C4.5 addresses the handling of continuous data by dividing the data into ranges based
on the attribute values found in the training sample. This allows the algorithm to
effectively work with continuous attributes, unlike IDS, which primarily handles discrete
attributes.
Example Continuous Data
Suppose we have a dataset of housing prices with two attributes: square footage and price. The square
footage attribute is continuous, and we want to build a decision tree to predict the price of a house
based on its square footage.
u=n Pruning
Suppose we have a dataset for predicting whether a customer will purchase a product based on their
age and income level. We use C4.5 to build a decision tree, and the resulting tree is as follows:
If age < 30 and income = high, then purchase
If age < 30 and income = low, then no purchase
If age >= 30 and income = high, then purchase
If age >= 30 and income = low, then no purchaseNow, let's say we have a validation dataset
that we use to evaluate the performance of the decision tree. We find that the decision tree has a high
accuracy on the training dataset, but a lower accuracy on the validation dataset. This suggests that the
decision tree is overfitting to the training dataset
To address this issue, we can prune the decision tree by removing subtrees that do not improve its
performance on the validation dataset. For example, we might consider removing the subtree for the
condition "age < 30," since it only applies to a small subset of the data and may be overfitting to noise
in the training dataset
After pruning the decision tree, we might end up with the following simplified tree:
I f income = high, then purchase
If income = low, then no purchase _______________________________________ _________
• •
This pruned decision tree is simpler and more generalizable than the original decision tree, since it is
less likely to overfit to noise in the training dataset By pruning the decision tree, we have improved its
performance on the validation dataset and made it more suitable for predicting whether a customer
will purchase a product based on their age and income level.
4. Rules:
C4.5 allows for classification via decision trees or rules generated from them. It also offers
techniques to simplify complex rules. For example, it can replace the left-hand side of a rule
with a simpler version if all records in the training set are treated identically. Additionally, an
"otherwise" type of rule can be used to indicate what should be done if no other rules apply.
□
HOg EYainpli' Rules
Suppose we have a dataset for predicting whether a customer will purchase a product based on their
age and income level. We use C4.5 to build a decision tree, and the resulting tree is as follows:
If age < 30 and income = high, then purchase
v' If age < 30 and income = low, then no purchase
^ If age >=30 and income = high, then purchase
^ If age >= 30 and income = low, then no purchase
We can use C4.5 to generate rules from this decision tree. For example, the first rule would be:
^ If age < 30 and income = high, then purchase
Similarly, we can generate rules for the other branches of the decision tree. For example, the second
rule would be:
If age < 30 and income = low, then no purchase
We can also use C4.5 to simplify complex rules. For example, suppose we have the following rule:
If age < 30 and income = high and gender = female and education = college, then purchase
This rule is complex and may be overfitting to noise in the training dataset. To simplify this rule, C4.5
can replace the left-hand side of the rule with a simpler version if all records in the training set are
treated identically In this case, we might simplify the rule to:
^ If age < 30 and income = high, then purchase
Finally, we can use an "otherwise" type of rule to indicate what should be done if no other rules apply
For example, we might have the following rule:
'/ If no other rules apply, then no purchase
These rules generated from the decision tree can be used to classify new customers based on their age
and income level. By generating rules from the decision tree, we can simplify the classification process
and make it more interpretable.___________ _______ __________________________________ _____
5. Splitting:
the is s u e of-overfittingjjy taking into account the cardinality of each division
C 4 .5 - a d d r e s s e s
when selecting the best attribute for splitting. It uses the GainRatio instead of Gain for splitting
purposes. The GainRatio compensates for the skewness of the GainRatio value toward sphts
where the size of one subset is close to that of the starting one, ensuring a larger than average
information gain.
Example .".\ Splitting
Suppose we have a dataset of students and we want to build a decision tree to predict whether a student
will pass o r fail an exam based on two attributes: study hours per week and attendance percentage. The
dataset has the following distribution:
• Pass: 60 students
• Fail: 40 students
We want to decide which attribute to use for the first split in the decision tree. We calculate the
GainRatio for both attributes (study hours and attendance percentage) to determine the best attribute
for splitting.
For the study hours attribute, the dataset can be split into two subsets:
• Subset 1: Study hours <5
• Subset 2: Study hours >=5
The information gain for this split is calculated using entropy measures. Similarly, we calculate the
information gain for the attendance percentage attribute by splitting the dataset based on different
attendance percentage thresholds.
After calculating the information gain for both attributes, we compute the GainRatio for each attribute.
The GainRatio takes into account the cardinality of each division and compensates for the skewness of
the GainRatio value toward certain splits.
Suppose we find that the GainRatio for the study hours attribute is 0.6, and the GainRatio for the
attendance percentage attribute is 0.8. Based on these values, C4.5 would select the attribute with the
highest GainRatio (in this case, the attendance percentage attribute) for the first split in the decision
tree.
By using GainRatio, C4.S ensures that the attribute selected for splitting takes into account the size of
the subsets and compensates for any skevraess in the information gain values, thereby addressing the
issue of overfitting and improving the accuracy of the decision tree.
This example demonstrates how C4.5 uses GainRatio to make informed decisions about attribute
selection for splitting, ultimately leading to the construction of more effective and generalized decision
trees. __________________________________________________
7. Continue the process o f calculating information gain ratio, selecting splitting attributes, and creating
child nodes until the tree is fully constructed.
8. Pruning (optional):
•_______ After the tree is constructed, pruning techniques can be applied to reduce overfitting and improve
________ generalization.____________
m
1. Versatility: The C4.5 algorithm can handle both categorical and continuous attributes, making it
suitable for a vvide range o f datasets.
2. HandlingMissing Data: The C4.5 algorithm can handle missing data by ignoring missing values during
attribute selection and prediction.
3. Reduced Overfitting: The C4.5 algorithm uses pruning techniques to reduce overfitting and improve
generalization.
4. Easy to Interpret: The decision tree generated by the C4.5 algorithm is easy to interpret and can
provide Insights into the underlying data._____________________________________________
1. Computationally Expensive: The C4.5 algorithm can be computationally expensive, especially for
large datasets virith many attributes.
2. Sensitive to Noisy Data: The C4.5 algorithm is sensitive to noisy data, which can lead to overfitting and
inaccurate predictions.
3. Biased towards Attributes with ManyValues: The C4.5 algorithm tends to favor attributes with many
values, which can lead to overfitting and inaccurate predictions.
4. Limited to Binary Classification: The C4.5 algorithm is limited to binary classification problems and
cannot handle multi-class classification problems without modifications. _______________________
The CART (Classification and Regression Trees) algorithm is a popular decision tree algorithm used
for both classification and regression tasks. CART is a recursive partitioning algorithm that recursively
splits the dataset into subsets based on the values of input variables. It constructs binary trees where
each non-leaf node represents a decision based on a feature, and each leaf node represents the output
(class label or numerical value).
The purpose of the CART algorithm is to create a predictive model that can be used for both
classification and regression tasks. It aims to partition the input space into regions that are as
homogeneous as possible vdth respect to the target variable.
The algorithm uses a top-down greedy approach to recursively split the dataset based on the feature
that provides the best split, as determined by a criterion such as the Gini impurity for classification
tasks or the reduction in variance for regression tasks. The process continues until a stopping criterion
is met, such as reaching a maximum tree depth or having nodes vdth a minimum number of samples.
How it works?
1. Select the B est Split: The algorithm evaluates all possible splits for each feature and selects
the one that maximizes the homogeneity of the resulting subsets.
^ For Classification: Calculate the Gini impurity for each feature and select the split that
maximizes the homogeneity of the resulting subsets.
Example CART
Suppose we have a dataset of weather conditions and corresponding activities, and we want to build a decision
tree to
TemperatiuB * Humidity Activity <.
-■ “ Outlook
high false no
sunny hot
hot high true no
sunny
high false yes
overcast hot
high false yes
rainy mild
normal false yes
rainy cool
normal true no
rainy cool
normal true yes
overcast cool
high false no
sunny mild
normal false yes
sunny cool
normal 1 false yes
rainy mild
S u p erv ised iMrnii
The root node is "Outlook," and it has three branches for "Sunny," "Overcast," and "Rainy."
Under "Sunny," there's a decision based on "Humidity."
Under "Overcast," the decision is straightforward, resulting in a "Yes" prediction.
Under "Rainy," there's a decision based on "Windy."___________________________
Limitations<of CART
1. Tendency to Overfit: CART decision trees can grow very large and complex, leading to overfitting on
the training data. Pruning techniques are often required to address this issue.
2. Sensitive to Small Variations in D ata: Small changes in the input data can lead to significantly different
tree structures, making the model less robust
3. Binary Splits Only: CART creates binary trees, which may not capture more complex relationships
present in the data that require multiway splits.
4. Not Suitable for Unbalanced D ata: CART may produce biased trees when dealing with unbalanced
datasets, where one class is much more prevalent than the others.
5. Greedy Algorithm: fcART uses a greedy algorithm to select the best split at each node, which may not
always lead to the globally optimal tree structure.________________________ _______________ ____
import pandas as pd
from [Link] import DecisionTreeClassifier
from [Link] import OneHotEncoder
3. Healthcare Diagnostics:
Decision Trees are utilized in healthcare for diagnostics and disease prediction. These
algorithms can analyze patient data, sjmiptoms, and test results to assist in diagnosing medical
conditions, recommending treatments, and predicting patient outcomes.
4. Predictive Maintenance:
Decision-based algorithms are used in predictive maintenance applications to anticipate
equipment failures and schedule maintenance activities proactively. By analyzing sensor data
and historical maintenance records, these algorithms can predict when machinery is likely to
malfunction.
5. Marketing Campaign Optimization:
Decision Trees are applied in marketing to optimize campaign strategies and target specific
customer segments effectively By analyzing customer demographics, behavior, and responses
to past campaigns, decision-based algorithms can help businesses tailor marketing efforts for
better engagement and conversion rates.
6. E-commerce Product Recommendations:
Decision-based algorithms power recommendation systems in e-commerce platforms to
suggest products to users based on their browsing history, purchase behavior, and preferences.
These algorithms enhance the user experience and increase sales by offering personalized
product recommendations.
7. Chum Prediction:
Decision Trees are used in churn prediction models to forecast customer attrition or churn.
By analyzing customer interactions, usage patterns, and feedback, decision-based algorithms
can identify customers at risk of leaving a service or subscription, allowing businesses to take
proactive retention measures.
8. Energy Consumption Forecasting:
Decision-based algorithms are employed in energy consumption forecasting to predict
electricity demand and optimize energy distribution. These algorithms anal3rze historical
consumption data, weather patterns, and other factors to forecast energy usage and improve
resource planning.
O Decision Trees are easy to interpret and understand, making them valuable for explaining the reasoning
behind predictions to stakeholders and domain experts.
O Decision-based algorithms can provide insights into feature importance, helping in feature selection
and understanding theimpacL of variables^on the model^predictions.
O Decision Trees can capture non-linear relationships between features and the target variable, making
them suitable for complex datasets with non-linear patterns.______
3.74 V :^;Madiine L earning
0 Decision Trees can handle missing values in the data without the need for imputation techniques,
simplifying the preprocessing steps.
3 Decision-based algorithms are robust to outliers in the data and can handle noisy data witiiout
significantly impacting model performance._______________________________________
UNSUPERVISED LEARNING
^ sn sn n D
•■ Introduction
•- Types of Unsupervised Learning
•- Clustering
•- K-Means Clustering Algorithm
•■ Using Clustering for Image Segmentation
•■ Using Clustering for Preprocessing
Using Clustering for Semi-Supervised Learning
- DBSCAN
Other Clustering Algorithms.
Review questions.
V .
In tr o d u c tio n
In many aspects, unsupervised learning differs greatly from supervised machine learning. This sort of
machine learning does not require supervision. It means that in unsupervised machine learning, we
train the machine with a dataset that hasn’t been labelled or trained, and the machine then predicts
the results without any supervision of the underlying data.
Unsupervised learning is the training of a machine using information that is neither classified nor
labelled and allowing the algorithm to act on that information without guidance. Here the task of the
machine is to group unsorted information according to similarities, patterns, and differences without
any prior training of data
The unsupervised learning algorithm's main goal is to classify or group the unsorted dataset
according to the pattern, similarities, and differences that it can identify in the data. The machines
are given instructions to find hidden pattern in the input dataset, and the findings are then analysed.
Unsupervised learning includes all types of machine learning scenarios where there is no predefined
output or instructor to guide the learning algorithm. In unsupervised learning, the learning algorithm
is just shown the input data and asked to extract knowledge from this data.
B k I clustering
Clustering is a type of unsupervised learning method. Clustering is also known as grouping is
the process of categorizing a set of objects in a way that objects within the same group or cluster are
more similar to each other than to those in other groups. This is similar to dividing data objects into
subclasses based on their similarities.
Clustering is a unsupervised learning method that involves grouping similar data points together
based on their characteristics. Unlike classification, clustering does not involve predefined groups or
labels, but rather relies on finding similarities between data points to group them into clusters.
There are different definitions for clusters, but they generally involve a set of like elements that are
distinct from elements in other clusters. One common definition is that the distance between points
in a cluster is less than the distance between a point in the cluster and any point outside it
A term closely aligned with clustering is "database segmentation," where similar tuples or records
within a database are grouped togeflier. This segmentation aims to partition or segment the database
into distinct components, providing users with a more generalized perspective of the data. In this
context, there is no explicit differentiation between segmentation and clustering.
Determining how to cluster data is not always straightforward, and there are different methods and
algorithms for clustering.
Definitions of Clustering
L
Clustering is the task of dividing the population or data points into a number of groups such
that data p o i n t s in4he^ame-groups_are more similar to other data points in the same group and
dissimilar to the data points in other groups.
Clustering is basically a collection of objects on the basis of similarity and dissimilarity between
them.
Clustering or Cluster analysis is the method of grouping the entitiesjased^onjimila^
Unsu
W h a t is C lu s te r in g ?
Clustering is a technique used to group similar data objects together based on their characteristics
or attributes. The goal of clustering is to identify patterns or structures in the data that can help in
understanding the underl)dng relationships and associations between the data objects. Clustering
is an unsupervised learning technique, meaning that it does not require labeled data and can be
used to explore and discover patterns in the data without prior knowledge of the data’s structure.
Important Note
Clustering is somewhere similar to the classification algorithm, but the difference is the type of dataset that
we are using. In classification, we work with the labeled data set, whereas in clustering, we work with the
unlabelled dataset.
<)r- --
M achin e Learning ; ' . . ,. *^x 'Ot'^j"''-
Background : A retail company wants to improve its marketing strategies by better understanding
its diverse customer base. The company has a large database of customer transaction data including
purchase history, demographic information, and online behavior.
Application of Clustering: The company applies clustering techniques to segment its customer base
into distinct groups based on various attributes such as purchasing behavior, age, location, and product
preferences. By using clustering algorithms, the company can identify natural groupings within the
customer data without predefined categories.
Benefits:
Targeted Marketing : With the identified customer segments, the company can tailor its
marketing campaigns to address the specific needs and preferences of each group. For
example, different segments may respond better to different types of promotions or product
recommendations.
Product Customization : Understanding the distinct preferences of each segment allows the
company to customize its product offerings to better meet the demands of different customer
groups. This can lead to increased customer satisfaction and loyalty,
v' Resource Allocation: By knowing the characteristics of each segment, the company can allocate
resources more effectively by focusing on the segments with the highest potential for sales and
customer engagement.
Customer Retention: The insights gained fi'om customer segmentation can help in developing
targeted retention strategies such as personalized loyalty programs or communication strategies
tailored to the needs of each segment
Conclusion: Through the application of clustering for customer segmentation, the retail company can
gain valuable insights into its customer base, leading to more effective marketing strategies, improved
customer satisfaction, and ultimately, increased sales and profitability.
Applications of Clustering
Clustering, as a key technique in unsupervised learning and it is used in applications across various
domains and industries. Some common applications of clustering include:
1. Customer Segmentation:
Businesses use clustering to segment customers based on their purchasing behavior,
demographics, or preferences. This segmentation helps in targeted marketing, personalized
recommendations, and improving customer satisfaction.
2. Image Segmentation:
In image processing, clustering is used for segmenting images into regions with similar
characteristics such as color, texture, or intensity. This is valuable in medical imaging, object
recognition, and computer vision applications.
3. Anomaly Detection:
Clustering algorithms can identify outliers or anomalies in datasets, which is crucial for fraud
detection, network security, and quality control in manufacturing processes.
M achim l ^ m i i g .
I
4. Document Clustering:
Text documents can be clustered based on their content to group similar documents together.
This is useful in information retrieval, document organization, and topic modeling.
5. Genomics and Bioinformatics:
Clustering is applied in genomics to group genes with similar expression patterns or in protein
clustering for structural analysis. It helps in understanding genetic relationships and biological
functions.
6. Recommendation Systems:
Clustering techniques are used in recommendation systems to group users or items with
similar preferences. This enables personalized recommendations in e-commerce, streaming
services, and social media platforms.
7. Spatial Data Analysis:
Clustering is utilized in geographical data analysis to identify spatial patterns, cluster locations,
and regional trends. It is valuable in urban planning, resource allocation, and environmental
studies.
8. Market Research:
Clustering assists in market segmentation, where customers are grouped based on their
buying behavior, demographics, or psychographics. This information helps businesses in
product positioning, pricing strategies, and targeted advertising.
9. Healthcare Analytics:
Clustering is used in healthcare for patient segmentation, disease clustering, and medical
image analysis. It aids in personalized medicine, treatment planning, and healthcare resource
optimization.
Clustering Attributes
Clustering attributes refer to the characteristics or features of the data that are used to group similar
1. Partitioning Clustering
Partitional clustering is a type of clustering algorithm that divides a dataset into non-overiapping
clusters, where each data point belongs to only one cluster. The goal of partitional clustering is to
group similar data points together and separate dissimilar data points into different clusters.
jMachine Learmng ■ r?“.
In other words, partitional clustering aims to partiUon the data into a s e t of dusters, where each
cluster contains data points that are similar to each other and dissim ilar to data points in other
clusters. The number of clusters is usually determined by the user, and th e algonthm tries to find the
best partition of the data into the specified number of clusters.
A company may use partitional clustering to group customers into segments based on their purchase history,
such as frequency of purchases, amount spent, or types of products purchased. The algorithm would assign
each customer to the closest cluster based on their purchasing behavior, and each cluster would represent a
segment of customers with similar purchasing behavior.
Once the customers are segmented, the company can tailor their marketing strategies to each segment For
example, they may offer discounts or promotions to customers in a particular segment to encourage them
to make more purchases. They may also use different marketing channels o r messaging for each segment to
better target their marketing efforts.
Density-Based Clustering identify clusters based on the density of d a ta points in a dataset High
density region s contain data points that are densely packed with m an y neighboring points within
a specified radius, indicating cohesive clusters. In contrast, low d ensity regions have fewer nearby
points, potentially representing noise or outliers. Unlike partitioning clustering where the number
of clusters is predefined, density-based methods group dosely packed points and distinguish hlgh-
density clu sters from low-density areas, offering flexibility to capture clusters of varjing shapes and
sizes while effectively handling complex data distributions and outliers. ’ ■
Example
Example: In a retail scenario, a company may use DBSCAN to duster custom ers based on their shopping
behavior. The algorithm can identify clusters of customers who frequently shop together in certain areas of
a store (high-density regions) and separate them from customers *vho shop less frequently or in different
areas (low -density regions). By segmenting customers based on shopping patterns rather tfian predefined
categories, the company can tailor marketing strategies and promotions to each cluster effectively.
In the distribution model-based clustering method, the data is divided based on the probability of
how a d ataset belongs to a particular distribution. - i ^
JS g E x a m p le
Consider a library with a diverse collection of books covering various subjects such as science histoiy,
literature, and a rt Initially, each book is treated as a separate entity, representing a distinct cluster. However,
as the clustering process progresses, books with similar subject m atter or content a r e grouped together to
form dusters representing specific topics or genres. This process continues, with clustei^ of books being
further grouped into broader categories, such as "science and technology," "humanities," or fiction.
At each level of the hierarchy, the clustering process identifies similarities betiA^een books and organizes tiiem
into meaningful groups based on their content This hierarchical structure allows library visitors to navigate
t h e c o l l e c t i o n b a s e d o n t i i e i r i n t e r e s t s ,s t a r t i n g f r o m s p e c i f i c b o o k s a n d g r a d u a l l y m o v i n g t o b r o a d e r c a t e g o n e s
and topics. It also enables librarians to manage and organize the collection more effectivej^
“ ----------------------- -------------------
T ypes o f H ierarchical Clustering Algorithm s:
-h Agglomerative: This bottdm-up approach starts with each data point as an individual cluster
and then merges the d o sestp air of clusters until all points are merged into a single cluster.
^ Divisive: This top-down approach begins with all data points in a single cluster, which is then
split recursively into smaller clusters.
' - ' Unsoperviiwl learnir
In fuzzy clustering, data objects are allowed to belong to multiple clusters simultaneously, with each
object having membership coefficients indicating the degree to which it belongs to each cluster.
Unlike traditional hard clustering methods where data points are assigned to a single cluster, fuzzy
clustering assigns membership values that represent the likelihood of a data point belonging to
different clusters. The Fuzzy C-means algorithm, also known as the Fuzzy k-means algorithm, is a
popular example of fuzzy clustering.
It is an iterative algorithm that divides the unlabeled dataset into k different dusters in such a way that each
dataset belongs only one group that has similar properties.
The algorithm takes the unlabeled dataset as input, divides the dataset into k-number of clusters, and repeats
the process until it does not find the best clusters. The value of k should be predetermined in this algorithm.
The algorithm starts by randomly selecting K points from the dataset as the initial centroids of the
clusters. In each iteration, each data point is assigned to the nearest centroid, and the centroid of each
cluster is updated as the m ean of all the data points assigned to that cluster. This process is repeated
until convergence, which is achieved when the centroids no longer change or the change is below
a certain threshold. The output of the algorithm is K clusters, with each cluster containing the data
points that are closest to its centroid.
The K-means algorithm is widely used for clustering datasets with multiple attributes and can be
applied to a variety of domains, such as image segmentation, customer segmentation, and anomaly
detection.
# Create a dataset with Indian student performance data (student name, study hours,
exam score)
students = [Link]([['Srikanth', 100], ['Snigdha', 75], ['Mary', 35], ['Nirmala',
55], ['Raju', 85], ['Rama',30], ['Sita', 45], ['Lava', 65], ['Kusha', 25], ['Hanuman',
50]])
Cluster 2:
Student Name: Nirmala^Exam Score: 55
Student Name: Sita,Exam Score: 45
Student Name: Lava,Exam Score: 65
Student Name: HanumanjExam Score: 50
Cluster 3:
Student Name: Mary,Exam Score: 35
Student Name: Rama,Exam Score: 30
Student Name: Kusha,Exam Score: 25
Explanation
Data Preparation: The code initializes a dataset with student names and exam scores. Each student's
data is represented as an array with ± e student's name and exam score.
Feature Extraction: It extracts only the numeric feature (exam score) from the dataset for clustering.
By selecting only the exam scores as the feature for clustering, the code prepares the data for the
K-Means clustering algorithm.
K-Means Clustering: The code applies the K-Means clustering algorithm with n_clusters=3 to cluster
the students based on their exam scores into three clusters. K-Means is an unsupervised machine
learning algorithm that partitions the data into K clusters based on similarity.
Cluster Assignment: After clustering, the code assigns each student to a cluster based on the clustering
results. It creates a dictionary clusters where each key represents a cluster label, and the corresponding
value is a list of students belonging to that cluster.
Output Display: Finally, the code displays the clusters along with the student names and exam scores
in each cluster. It iterates through the clusters dictionary and prints the student names and exam scores
for each cluster, providing insights into how the students are grouped based on their exam performance.
M o d iin e L e arn ii^
K-Means is computationally efficient and can handle large datasets with low computational co st
The algorithm is easy to implement and interpret, making it accessible to users of different skill levels.
K-Means can utilize various distance metrics to measure data point similarity.
It is well-suited for scenarios where clusters are spherical and exhibit similar variance.
Results from K-Means are easily interpretable, aiding in understanding data grouping.
The algorithm can be robust to outliers, minimizing their impact on clustering results.
K-Means can be parallelized to speed up the clustering process by distributing the computation across
multiple processors or nodes^______________________________________________________________
Limits of K-M eans Clustering Algorithm
Understanding the limitations is crucial for selecting the appropriate clusteriflg algorithm based on the
characteristics of the data and the desired clustering outcomes.
• Multiple Runs for Optimal Solutions: K-Means may converge to suboptimal solutions due to its
sensitivity to the initial random centroids. To mitigate this, the algorithm needs to be run multiple times
with different initializations to improve the chances of finding the best clustering solution.
• Manual Selection of Number of Clusters: One of the challenges with K-Means is the need to specify
the number of clusters (K) beforehand. Determining the optimal number of clusters can be subjective
and may require domain knowledge or trial-and-error, making it a cumbersome task.
• Limited Cluster Shape Flexibility: K-Means assumes that clusters are spherical and of sim ilar size and
density. When clusters have varying sizes, different densities, or nonspherical shapes, K-Means may
struggle to accurately cluster the data, leading to suboptimal results.
• Inability to Handle Complex Cluster Shapes: In scenarios where the clusters exhibit complex shapes
such as ellipsoids, K-Means may fail to capture the true underlying structure of the data. This limitation
is evident when the clusters have different dimensions, orientations, and densities, as K-Means tends to
produce suboptimal or incorrect cluster assignments.
• Alternative Clustering Algorithms: Depending on the nature of the data and the shapes o f the clusters,
different clustering algorithms like Gaussian Mixture Models (GMM) may outperform K-Means. GMM is
more flexible in capturing clusters with varying shapes and densities, making it a suitable alternative
______ for datasets with non-spherical clusters.
K -M e a n s C l u s t e r i n g for Im ag e S e g m e n ta t i o n
The K-means algorithm is a popular clustering technique used in image segm enUUon to partition an
image into K clusters based on p ix e l similarities
K -M e q n «
■VHrevm* Clustering for...
Image
w ^Segmentation
^ -- ---------- ---------------
1. Initialization: The algorithm starts by randomly initializing K duster centers in the feature space,
where K is the predefined number of dusters.
2 . - r - " - " ' step: Each pixel in the image is assigned to the cluster whose centroid is closest to It in
'te rm s of feature similarity. The distance m etric,olten Euclidean distance,is used to measures,milanty.
3. Update Step: After assigning ail pixels to clusters, the centroids of the clusters are recalculated based
on the m ean o f th e p ixel values within each cluster. ----------------------- -------------------------------- ---
4. Iteration: Steps 2 and 3 are repeated iteratively until convergence criteria are met, such as a maximum
number of iterations or minimal change in cluster assignments.
5. Segmentation: Once the algorithm converges, each pixel in the image is associated with the cluster it
belongs to. It is effectively segmented the image into distinct regions based on pixel similarity.______
E xam ple A P^hon Code to Demonstrate Image Segmentation Using K-Means Clustering Method.
import [Link] as pit
from [Link] import KMeans
from [Link] import imread
from [Link] import resize
import nuoqpy as np
plt.tight_layout()
[Link] ______ _________ ___________________' .a— -------- — ----------------
4 .2 2 ^
HI
Output
Explanation
1. Image Loading and Reshaping:
• The code starts by loading an image from the given path [Link] using the imread function
from the [Link] module.
• The image is then reshaped into a two-dimensional array where each row represents a pixel and
each column represents the color channels (Red, Green, and Blue).
2. K-Means Clustering:
• KMeans clustering from scikit-leam is applied to the array of pixels with n_clusters=5, meaning
the image will be segmented into 5 different regions based on the color of the pixels.
• The fit method of the KMeans object is used to compute the clusters.
3. Centroid Assignment:
• Each pixel in the image is assigned to the nearest cluster centroid after the K-means algorithm
converges. The RGB values of each pixel are replaced with the RGB values of the centroid of the
cluster it belongs to, resulting in a segmented image where each segment has a uniform color.
4. Image Display:
• The code sets up a figure with two subplots using matplotlib to display both the original and the
segmented images side by side.
• The original image is displayed on the left, and the segmented image is displayed on the right
with the title 'Segmented Image when K=5'.
5. Legend Creation:
• A legend is created to help identify which colors correspond to which clusters.
• For each cluster, a colored dot is plotted with a label "Cluster {idx}", where {idx} is the index of the
cluster. This dot's color represents the color of the cluster centroid in RGB space.
Note: To run this code, the scikit-image library needs to be installed using pip install scildt-image
# Assign clusters and calculate the distance from each point to its assigned cluster
# Detect anomalies
outliers = X[min_distances > threshold]
' ■ m - •• •••
• • •
’ • '-m •;
- 3 - 2 - 1 0 1 2 3
Feature 1
Explanation'
1. Data Generation: The code begins by generating synthetic data using make_blobs from Scikit-learn,
simulating a dataset with 300 samples grouped into 4 clusters.
2. Lustering with K-Means: K-Means clustering is applied to the dataset specifying 4 clusters. Each data
point is then assigned to the nearest cluster.
3. Distance Calculation: The distance of each data point to its nearest cluster center is calculated. This
distance metric helps in identifying how far a point is from the cluster's core.
4. Setting a Threshold: A threshold is set at the 95th percentile of these distances. Data points whose
distance to the nearest duster center exceeds this threshold are considered anomalies.
5. Anomaly Detection: Points are labeled as anomalies if their distance to the nearest cluster center is
greater than the defined threshold. These are visualized in red, whereas normal points are in blue.
6 . Visualization: The data, along with identified anomalies and cluster centers, are visualized using
Matplotlib. This visual representation helps to clearly see the normal data, the anomalies, and the
cluster centers.
Clustering plays a crucial role in enhancing semi-supervised learning by leveraging the inform ation
present in both labeled and unlabeled data. In many practical applications, acquiring a large am ount
of labeled data can be expensive or im practical or time consuming. Clustering techniques offer a way
to extract valuable insights from the unlabeled data to improve the m odel's performance.
The process of using clustering in sem i-supervised learning involves several key steps to effectively
leverage the information from both labeled and unlabeled data. A structured approach to incorporating
clustering in semi-supervised learning:
1. D ata P rep aratio n :
• S p lit th e Data: Divide the d ataset into labeled and unlabeled portions. The labeled data
contains input features along with corresponding output labels, while the unlabeled data
lacks explicit labels.
2. C lu stering U nlabeled D ata:
. Apply Clustering A lg o rith m s: Utilize unsupervised clustering algorithms [e.g., KMeans,
DBSCAN) on the unlabeled data to identify clusters based on similarities in feature
space.
3. P seudo-Labeling:
. A ssign Pseudo-Labels: For each cluster generated by the clustering algorithm, assign
pseudo-labels to the data points within the cluster. This can be done by considering the
m ajority label of the labeled data points in the same cluster.
. L abel P ropagation: Propagate th e pseudo-labels to the unlabeled data points w ithin the
clusters for effectively creatin g a partially labeled dataset.
m
4. Model Training:
• Combine Labeled and Pseudo-Labeled Data: Merge the labeled data with the newly
pseudo-labeled data to create an augmented training set.
• T rain the Model: Use a semi-supervised learning algorithm (e.g., self-training, co
training) to train the m odel on the combined dataset by leveraging both labeled and
pseudo-labeled data for learning.
5. Model Evaluation:
• Validate the Model: Evaluate the trained model on a separate validation set to assess its
performance and generalization to unseen data.
• Fine-Tuning: Iterate on the model training process by adjusting h3 q)erparameters and
clustering parameters as needed to improve performance.
6. P red iction and Inference:
• Make Predictions: Use th e trained model to make predictions on new, unseen data,
leveraging the knowledge gained from both labeled and unlabeled data.
• In corp orate Feedback: Continuously update and refine the model based on feedback
and new labeled data to enhance its predictive capabilities.
Example A Python Code to Use Clustering in Semi-Sup'ervised Learning
import numpy as np
import pandas as pd
from [Link] import KMeans
from [Link] import StandardScaler
from [Link] import DecisionTreeClassifier
from [Link] import accuracy_score
classifier = DecisionTreeClassifier(random_state=42)
[Link](X_train, y_train)
Initial Data:
student ID Exam Score Attendance Label
0 1 85 90 Pass
1 2 35 55 Fail
2 3 90 95 Pass
2 4 75 85 None
4 5 25 40 None
5 6 78 ________ 88 None
-^Unsupervisedj
1 2 35 55 0 .0 -•fitr/.'
2 3 90 95 1 .0 ' '> . .aSc ::
3 4 75 85 1 .0
4 5 25 40 0 .0
5 6 78 88 1 .0
Explanation
1. Data Preparation and Initial Display;
• The code begins by creating a pandas DataFrame from a dictionary that includes students’ IDs
their exam scores, attendance rates, and some labeled data ('Pass', 'Fail', and None for uhlabeled).
• The initial data is printed to give an overview of what is being processed. This hel^s In
understanding the dataset structure before any operations are applied.
2. Data Preprocessing t
• Label Conversion: The categorical labels 'Pass' and 'Fail' are converted into numerical fonnat
(1 for 'Pass', 0 for 'Fail'). This conversion is necessary for mathematical operations and model
training that will follow.
• Feature Scaling: The 'Exam Score' and 'Attendance' features are scaled using StahdaidScaler
from scikit-learn. Scaling is crucial as K-Means clustering is sensitive to the scale'df data, and
scaling ensures that each feature contributes equally to the distance calculations in the clustering
process. ’ , . . . I ,,,
DBSCAN (Density-Based Spatial Clustering of Applications with Noise) is a clustering algorithm in machine
learning that groups together points that are closely packed based on a density criterion. It is particularly
useful for identifying clusters of varying shapes and sizes in a dataset, while also being robust to noise
and outliers.
m l Importance o f DBSCAN
1. Density-Based: DBSCAN works on the idea of density connectivity and density reachability.
It groups together points that are closely packed together,. It mark the points that lie alone in
low-density regions as outliers.
2. Robust to Noise: DBSCAN is robust to noise and can identify outliers as noise points, making
it suitable for datasets with noise or outliers.
3. Handles Clusters o f Varying Shapes and D en sities: DBSCAN can identify clusters of arbitrary
shapes and sizes, unlike K-means, which assum es spherical clusters.
4. No Need to Specify Number of Clusters: Unlike K-means, DBSCAN does not require specifying
the number of clusters beforehand, making it m ore flexible.
5. Efficient: DBSCAN is computationally efficient and can scale well to large datasets.
H ow D B S C A N W o rk s ?
In the DBSCAN algorithm, a circle with a radius epsilon is drawn around each data point and the data point
is classified into Core Point, Border Point, or Noise Point. The data point is classified as a core point if it has
min_samples of data points with epsilon radius. If it has points less than [Link] it is known as Border
Point and if there are no points inside epsilon radius it is considered a Noise Point.
Let us understand working through an example.
In the above figure, we can see that point A has no
points inside epsilon(e) radius. Hence it is a Noise
Point. Point B has min_samples(=4) number of
points with epsilon(e) radius, thus it is a Core Point
While the point C has only 1 ( less than minPoints)
point, hence it is a Border Point..
JS q Example Working of DBSCAN Algorithm
Suppose we have a dataset of points representing customers in a shopping mall based on their spending score
and annual income. We want to group these customers into clusters using DBSCAN.
1. Core Points, Border Points, and Noise Points:
• Core Points: A core point could be a customer who has at least 5 other customers within a
distance of 10 units. These core points act as central hubs in a cluster.
• Border Points: Border points are customers who are reachable from core points but do not have
enough neighbors to be core points themselves.
• Noise Points: Noise points are customers who do not belong to any cluster, perhaps because
they are outliers in terms of spending score and income.
2. Parameters:
• Epsilon (eps): Let's set epsilon to 10 units, meaning points within a distance of 10 units are
considered neighbors.
• Minimum Samples (min_samples): We require at least 5 points within the epsilon radius for a
point to be considered a core point.
3. Algorithm Steps:
• Initialization: Start by randomly selecting a customer who has not been visited as the Initial
point for forming a cluster.
• Expand:
o For each core point or border point, expand the cluster by adding neighboring customers
recursively based on the epsilon and min_samples criteria.
o Check If the neighboring customers meet the criteria to be core points or border points
and add them to the cluster.
• Termination: The algorithm stops when all customers have been visited and clustered.
4. Output:
• Clusters: Customers grouped together based on their spending score and Income density.
• Noise: Outliers or customers who do not fit well into any cluster.
In this example, DBSCAN would help Identify clusters of customers with similar spending behaviors and
income levels, while also highlighting outliers who do not conform to any specific cluster pattern.
# DBSCAN algorithm
dbscan = DBSCAN(eps=0,5, inin_samples=2 )
clusters = dbscan.fit_predict(data_scaled)
class_member_mask = (clusters == k)
# Plot outliers
xy = data[~class_member_mask]
pit•plot(xy[:, 0 ], xy[:, 1 ], 'o', markerfacecolor=tuple(col),
markeredgecolor='k ', markersize=6 )
Feature 1
Explanation,
1. Data Preparation
. The daaset consists of manually specified points in a 2D space repi^sendng '
The cooKiinates include typical clusters and distinct outliers (e.g„ points at [25, 80] [6 ,
70]).
2. Feature Scaling , .
. Before applying DBSCAN, the data is standardized using [Link] scaling adjusts each
f e a ^ t v T z e r o mean and unit variance, which is crucial because DBSCAN is sensitive to fte
" between points. Standardizing the data helps prevent features with larger scales from
dominating how clusters are formed.
3 . DBSCAN Clustering
. DBSCAN is initiated with an eps value of 0.5 and [Link] of 2. These parameters dictate th
clustering behavior; .
. eps (epsilon) is the maximum distance between two points for one to be considered as m t e
neighborhood of the other.
. min samples is the minimum number of points required to form a dense region (i.e., a duster).
. The'algorithm is expected to identify core points, border points, and ouUiers based on these
parameters.
4. Identification of Clusters and Outliers
• The DBSCAN algorithm categorizes the points into clusters and noise (outliers).
. In this dataset; Two main clusters are identified in denser regions of the dataset, where points are
close together. . ^ .
. Two points located at [25. 80) and [60, 70] are marked as noise because they do not meet th<
__________ renulred density fdefined by eps and min_samples) to be mcluded m a cluster-----------------------
II ‘SUnsu'peryised Learning
E)
5. Visualization
• The results are plotted using matplotlib, where different clusters are marked with different
colors, and outliers are colored in black.
• This visual representation helps illustrate DBSCAN's effectiveness at separating closely-knit
groups from sparse points, enhancing the understanding of cluster formation and outlier
________ detection in spatial data.___________________________________________________________
Applications of DBSCAN
s L m e t e r Robustness: DBSCAN Is less sensitive to its parameten>, such as the neighboi^ood radius
■ Ceps) and minlmun. number of points ([Link]), compared to other clustenng algonthms.
6 Handles Uneven Cluster Densities: DBSCAN can handle clusters with vanring densities, making i
suitable for datasets where clusters have different densities._______________________________________
| O j ^ D i s a d ^ t a ^ s:^ th eD B S C A N ^ gciirith m ^
1 Sensitive to Parameters: While DBSCAN is less sensitive to parameters compared to some ottier
■ clustering algorithms, choosing the right values for epsilon (eps) and minimum points ([Link])
can still be challenging and may impact the clustering results.
2 Difflcuity with Varying Density: DBSCAN may struggle with datasets where clusBrs have varying
" a l l t r e l i e l a s i n g l e e p s i , o n valueto define thenelghborhood radius forall points.
provided by Scikit-Learn, which cater to a variety of needs and data charact*nsO cs.
1. AggjomerativeClustering: *S
• D escription: Agglomerative clustering starts with individual instances as separate
clu sters and Iteratively merges the closest pair of d usters u n til a ll instances belong to a
single cluster. This process creates a hierarchy of clusters, re p rese n te d as a dendrogram.
• Advantages: It can capture clusters of various shapes and sizes, does not require
sp ecifyin g the number of clusters beforehand, and is su itab le for datasets with a large
n u m ber of instances.
• Scalability: Agglomerative clustering can scale well to large d atasets if a connectivity
m atrix is provided, indicating which instances are neighbors. W ithout a connectivity
m atrix, th e algorithm may not scale efficiently for large d atasets.
2. BIRCH (Balanced Iterative Reducing and Clustering using Hierarchies):
• D escription: BIRCH is designed for handling large datasets efficien tly by building a tree
stru ctu re during training. It uses a compact representation to assign new instances to
clu sters without storing all instances in memory.
• Advantages: BIRCH can handle large datasets with lim ited m em ory, making it suitable
for sce n a rio s where memory usage is a concern. It provides re s u lts comparable to batch
K-M eans clustering.
• Lim itation: BIRCH works best for datasets with a moderate n u m b e r of features, typically
less th a n 2 0 , due to the tree structure's complexity.
3. Mean-Shift:
• D escription: Mean-Shift clustering is a non-parametric clu sterin g algorithm that
iterativ ely shifts data points towards the mode of the d ata distribution. It identifies
clu sters b y finding density peaks in the data.
• Advantages: Mean-Shift can discover clusters of arbitrary sh a p e s and sizes without
req u irin g the number of clusters as an input parameter. It is effectiv e in handling datasets
with irregu lar cluster shapes.
• Lim itation: The computational complexity of Mean-Shift is q u ad ratic (0(m ''2)], making
it less suitable for large datasets due to its high com putational cost.
4. Affinity Propagation:
• D escription: Affinity Propagation identifies exemplars in the data that represent clusters
and assig n s data points to these exemplars based on sim ilarity m easures. It uses message
passing to determine the exemplars and cluster assignments.
• Advantages: Affinity Propagation can automatically d eterm in e th e number of clusters
and is effective in identifying clusters of varying sizes and shapes.
• Lim itation: The algorithm's computational complexity is q u ad ratic (0(m ''2)), making it
less e fficien t for large datasets due to its high computational dem ands.
.............. . .........................
• r „ r r e ;~ “ L ;m ,^
.
in idendiyiDg clusters In graph data, such as sodal networks.
l i f l l R e v ie w Q u estion s
2. W hat is clustering?
3. Give tw o examples of using clustering to solve real life problem s.
vS ni.
B a :
- Install and set up Python and essential libraries like NumPy and pandas.
•■ Introduce scikit-learn as a machine learning library.
-■ Install and set up scikit-learn and other necessary tools.
^ Write a program to Load and explore the dataset of .CVS and excel files using pandas.
^ Wnte a program to Visualize the dataset to gain insights using Matplotlib o r Seaborn by
plotting scatter plots, b a r charts. ,
^ ^Iling^ to Handle missing data, encode categorical variables, and perform feature
- Write a program to im plem ent a k-Nearest Neighbours (k-NN) classifier using scikitlearn and
Tram the classifier on th e dataset and evaluate its performance.
- Write a program to im plem ent a linear regression model for regression tasks and
Train the model on a d ataset with continuous target variables.
- Write a program to implement a decision tree classifier using scikit-learn and visualize the
decision tree and understand its splits.
Program 1 InstaU and set up Python and essential libraries like NumPy and pandas.
Setting up Python and essential libraries on a Windows system for machine learning involves a series of
straightforward steps thatpreparetheenvironmentfordataanalysisandalgorithmdevelopmentBymstalhng
Python along with NumPy and Pandas, users can handle a wide array of data manipulation tasks efficien y
Follow the below steps to set up Python and essential libraries such as NumPy and Pandas for machme
learning on Windows.
Step 1: Install Python
Download Python: Go to the official Python website at [Link], navigate to the "Downloads"
section, and download the latest version for Windows. Choose the executable installer.
install Python: Execute the downloaded file. It is crucial to check the box labeled "Add Python to
PATH" at the start of the installation wizard. Select "Customize installation" and ensure all options,
including "pip", are selected. In the "Advanced Options," choose "Install for all users" and set the
installation path to C:\Python. Proceed by clicking "Install".
Program 4 Wnte a program t o Load [Link]^,tiie_dataset o f CSV and excel files using pandas.
Step 1: Creating CSV and Excel Files with Dummy Data
• Create CSV File: Open a text editor like Notepad or any other code editor. Enter the following data
Name, Age,Score
Srikanth,28,85
Snigdha,22,78
Mary,31,92
Save this file as sample_data.csv in the C:\ML_Projects directory.
• Create Excel File: We can use Microsoft Excel or Google Sheets to create this file. Enter the below data:
’'"Course. Sem
Rajesh BCA 1
Ramesh BCA 2
Swati BCOM 1
Fiorina BCOM 3
Pooja BBA 2
Raghu BBA 4
Save this file as sample_data.xlsx in the C:\ML_Projects directory.
Step 2: Python Code to Load and Explore the Data
import pandas as pd
CSV.
student I D ,S tu dy H o u r s , E x a m Score
1 ,5 ,8 2
2 ,2 ,4 8
3 ,8 ,9 0
4 ,1 ,3 5
5 ,3 ,5 0
6 ,4 ,6 6
7 ,9 ,9 5
8 ,6 ,7 5
9 ,7 ,8 8
1 0 .0 .5 .3 0
1 1 ,1 0 ,9 6
.
12 0 .2 0
1 3 ,1 2 ,9 8
Step 2: Python Code:
import pandas as pd
import [Link] as pit
alpha=0.7)
[Link]('Study Hours vs. Exam Scores')
[Link]('Study Hours')
[Link] Scores')
[Link](True)
I.
Explanation
1. Data Loading: The script uses pandas to load a CSV file containing students' study hours and exam
scores from a specified path.
2. Scatter Plot: Matplotlib is employed to create a scatter plot, plotting 'Study Hours' against 'Exam
Scores' to visually explore the relationship between these variables.
3. Bar Chart Setup: The data is categorized into bins based on study hours using [Link]. It then
calculates the average exam score for each category and uses this data to generate a bar chart showing
average scores by study hour range.
4. Visualization Configuration: Both plots are configured in a single figure, with clear titles and labels
for axes, using pltsubpIotQ to arrange them side by side for easy comparison.
5. Display: The script concludes with pltshowQ to display the configured plots, providing insights into
how stu(fy time correlates with academic performance through both detailed and aggregated views.
# Create du m y data
data = { ^___________________________
■Age': [25, 30, None, 28, 35],
■Gender': ['Female', 'Male', 'Male', 'Female', 'Male'],
'Income': [50000, 60000, 45000, None, 70000]
}
df = [Link](data)
# Feature scaling
scaler = StandardScaler()
scaled_data = scaler.fit_transform(df [['Age', 'Income']])
if predicted_outcome[0] = = 1 :
printC'Based on the exam scores provided, the student is predicted to pass. )
else: .
printC'Based on the exam scores provided, the student is predicted to fail. )______^_______
A ; Lab Programs ^ a .II
O utput
Output 1:
Accuracy on the test set: 1.00
Enter Exam Score 1: 45
Enter Exam Score 2: 50
Based on the exam scores provided, the student is predicted to fail.
Output 2:
Accuracy on the test set: 1.00
Enter Exam Score 1: 75
Enter Exam Score 2: 89
Based on the exam scores provided, the student is predicted to pass.
i'Ezplanation
1. Data Preparation:
The code initializes a numpy array X with exam scores as features and y with binary pass/fail labels.
It represents a simple dataset where each row corresponds to a student's exam scores and pass/fail
outcome.
2. Model Training and Evaluation:
It splits the data into training and testing sets using train_test_split with a test size of 20 % and a
random seed for reproducibility.
The code initializes a K-Nearest Neighbors (KNN) classifier with n_neighbors=3 and trains it on the
training data (X_train, y_train).
The model's performance is evaluated by predicting on the test set [X_test) and calculating the accuracy
using accuracy_score.
3. User Interaction:
The code prompts the user to input exam scores for a new student using inputQ function.
It prepares the user input as a numpy array userjnput to make a prediction using the trained KNN
classifier.
4. Prediction and Output:
The code predicts the outcome (pass/fail) for the new student based on the input exam scores using
the trained KNN classifier.
It then prints a message indicating whether the student is predicted to pass or fail based on the model's
prediction.
# Dummy house price prediction data: features (house size, number of bedrooms) and
target variable (house price)
X = [Link]([[1000, 2], [1500, 3], [1200, 2], [1800, 4], [900, 2], [2000, 3]])
y = [Link]([300000, 400000, 350000, 500000, 280000, 450000])
A .1 2 k ;
Write a program to implement a decision, tree classifier using scikit-learn and visualize
the decision tree and understand its splits. ’ . . -. -
import numpy as np
from [Link] import DecisionTreeClassifier, plot_tree
from [Link] import export_text
import [Link] as pit
1. Data Preparation:
The code defines a custom dummy dataset for fruit classification with features representing Weight
and 'Texture' of fruits and the target variable 'Fruit Type' (e.g., Apple, Orange, Melon).
The features and target labels are stored in NumPy arrays X and y, respectively.
2. Decision Tree Classifier Initialization:
A Decision Tree Classifier is initialized with random_state=42 to ensure reproducibility of results.
The classifier is then trained on the custom dummy dataset using the fitQ method.
3. Visualization of Decision Tre« Splits:
The export_text function is used to generate text-based rules of the Decision Tree Classifier based on
the features provided.
These rules provide insights into how the Decision Tree makes splits based on the Weight and
'Texture' features to classify different types of fruits.
4. Plotting the Decision Tree;
The [Link] function is utilized to visualize the Decision Tree structure graphically.
The Decision Tree is displayed with filled nodes, and the feature names ('Weight', 'Texture') and class
names (unique fruit types) are specified for better interpretation.
5. Displaying the Decision Tree Visualization:
A Matplotlib figure is created witfi a specific size to accommodate the Decision Tree plot
The Decision Tree visualization is shown using plt-showQ, allowing us to observe the tree structure
and decision-making process visually._____________ _______________________________________
import numpy as np
import [Link] as pit
from [Link] import KMeans
Explanation
1. Dummy Data Generation:
The program generates dummy customer data with features representing 'Age' and 'Income' of
customers.
K-Means Clustering:
K-Means clustering is applied to the customer data with n_clusters=3 to create 3 clusters.
The algorithm assigns each data point to one of the clusters based on the similarity of features.
2. Visualization;
The clusters are visualized using a scatter plot where each point represents a customer.
Different clusters are distinguished by colors, and cluster centers (centroids) are marked in red.
3. Plot Interpretation:
The plot helps visualize how customers are grouped into clusters based on their 'Age' and 'Income'.
Centroids represent the center of each cluster, showing the average 'Age' and 'Income' values for
customers in that cluster.
.a : - ; * , ; , ; . , . ■;
l^ O ^ E M V E S T X ^ p A P E R S
S ection -*/!
I. Answer any Four questions. Each question carries Two morks ( 4 X 2 = 8)
1. W h at is Machine Learning? Give an example.
2. W h at is Scikit-Ieam ?
3. W h at is Labeled Data and Unlabeled Data? Give an exam ple.
4. W h at is Classification? Give an example.
5. W h at is Clustering?
6. WhatisDBSCAN?
s e c tio n s
II. Answer any Four question. Each question carries Five marks 4 X 5 = 20)
7. W hy Use Machine Learning ?
8. W rite the Applications of Machine Learning.
9. W h at is Feature Engineering? Explain the Key Components of Feature Engineering.
10. H ow Naive Bayes Classifier works?
11. How K-Means Clustering Works? Write an algorithm.
12. W rite a Python Code to Demonstrate K-Means Clustering.
I Section~C
III. Answer any Four questions. Each question carries Eight marks 4 X 8 = 32)
13. Explain the Types of Machine Learning.
14. Explain the Essential Libraries and Tools required for M achine Learning Projects.
15. a) Discuss the Sources of Real-World Data.
b) Explain the Process of Selecting and Training a M achine Learning Model
16. Explain How to Discover and Visulaize the Data to Gain Insights in Data Preparation.
17. a) W rite a Python Code to Demonstrate Classification Tasks using CART,
b) W rite the applications of K-Means Clustering.
18. a ) How Clustering is Used In Semi-Supervised Learning?
b) W rite a note on a] Mean-Shift b) Affinity Propagation
. -
Section-;^
I. Answer any Four questions. Each question carries Two marks (4X2 = 8 )
II. Answer any Four question. Each question carries Five marks 4 X 5 = 20)
SecUon-C
III. Answer any Four questions. Each question carries Eight marks { 4 X 8 = 32)
I
Q u estio n Pap ers b.3
Section-1^
I. Answer any Four questions. Each question carries Two marks (4X2 = 8)
Sectiofi-'B
II. Answer any Four question. Each question carries Five marks { 4 X 5 = 20)
E m i
III. Answer any Four questions. Each question carries Eight marks ( 4 X 8 = 32)
(4X2 = 8 )
I. Answer any Four questions. Each question carries Two marks
4 X 5 = 20)
II. Answer any Four question. Each question carries Five marks
SecUofi-C
( 4 X 8 = 32)
III. Answer any Four questions. Each question carries Eight marks
13. Explain the types of Supervised and Unsupervised Machine Learning Algorithms.
14. W hat are NumPy and Pandas ? Why it is needed for ML? Explain its features.
15. a) Why Visualizing the Data is Needed During Data Preparation?
b) How to Load the Data and Explore the Data in ML?
16. How K-Nearest Neighbors (K-NN) Works ? Explain with an example for both classification and
regression tasks
17. a] W rite the Advantages and Disadvantages of Decision Tree Based Algorithms
b] How Clustering is used in Preprocessing?
18. Explain K-Means Clustering for Image Segmentation. W rite an algorithm.
1