Artificial Intelligence
By
Dr. Manoj Kumar
University School of Automation and Robotics
GGSIP University, East Campus, Delhi, India
Lecture-27 Machine Learning 1
What is machine learning?
“Learning is any process by which a system improves performance from experience.”
- Herbert Simon
➢ A branch of artificial intelligence,
concerned with the design and
development of algorithms that allow
computers to evolve behaviors
based on empirical data.
➢ As intelligence requires knowledge, it
is necessary for the computers to
acquire knowledge.
➢ Machine learning refers to a system
capable of the autonomous
acquisition and integration of
knowledge
What is machine learning?
➢ What if human can train the machines to learn from the past data and do what humans can do and much
faster?
What is machine learning?
What is machine learning?
What is machine learning?
More the data > Better Model > Higher accuracy
What is machine learning?
➢ The name machine learning was coined in 1959 by Arthur Samuel Tom M. Mitchell provided a widely quoted,
more formal definition of the algorithms studied in the machine learning:
An agent/computer program is said to learn from experience with
respect to some class of tasks (T), and a performance measure
(P), if its [the learner’s: agent/computer program ] performance at
tasks in the class, as measured by P, improves with experience
(E). “A well-defined learning task is given by <P,T,E>.
It could be diagnosing the illness of patient: the number of patients didn’t have adverse reaction of medicine you gave
Writing exams : The number of marks you got.
When you learning to improve your performance based on experience is
known as: Inductive Learning
When Do We Use Machine Learning?
ML is used when: No human experts
industrial/manufacturing control.
• Human expertise does not exist (navigating on Mars) mass spectrometer analysis, drug design, astronomic
• Humans can’t explain their expertise (speech recognition) discovery.
• Models must be customized (personalized medicine) Black-box human expertise
face/handwriting/speech recognition.
• Models are based on huge amounts of data (genomics) driving a car, flying a plane.
Rapidly changing phenomena credit
scoring, financial modeling.
diagnosis, fraud detection.
Need for customization/personalization
personalized news reader.
movie/book recommendation.
Defining the Learning Task
Improve on task T, with respect to performance metric P, based on experience E
T: Playing checkers
P: Percentage of games won against an arbitrary opponent
E: Playing practice games against itself
T: Recognizing hand-written words
P: Percentage of words correctly classified
E: Database of human-labeled images of handwritten words
T: Driving on four-lane highways using vision sensors
P: Average distance traveled before a human-judged error
E: A sequence of images and steering commands recorded while observing a human driver.
T: Categorize email messages as spam or legitimate.
P: Percentage of email messages correctly classified.
E: Database of emails, some with human-given labels
History of machine Learning
➢ Machine learning is a field of computer science that deals with the development of algorithms that can
learn from data.
History of machine Learning
➢ Machine learning is a field of computer science that deals with the development of algorithms that can
learn from data.
1943: The First Neutral Network with Electric Circuit
The first neutral network with electric circuit was developed by Warren McCulloch and Walter Pitts in 1943.
The goal of the network was to solve a problem that had been posed by John von Neumann and others: how
could computers be made to communicate with each other?
This early model showed that it was possible for two computers to communicate without any human
interaction. This event is important because it paved the way for machine learning development.
History of machine Learning
1950: Turing Test
The Turing Test is a test of artificial intelligence proposed by mathematician Alan Turing. It involves
determining whether a machine can act like a human, or if humans can’t tell the difference between human
and machine given answers.
The goal of the test is to determine whether machines can think intelligently and demonstrate some form of
emotional capability. It does not matter whether the answer is true or false but whether it is considered
human or not by the questioner. There have been several attempts to create an AI that passes the Turing Test,
but no machine has yet successfully done so.
The Turing Test has been criticized because it measures how much a machine can imitate a human rather than
proving their true intelligence.
History of machine Learning
1952: Computer Checkers
Arthur Samuel was a pioneer in machine learning and is credited with creating the first computer program to
play championship-level checkers. His program, which he developed in 1952, used a technique called alpha-
beta pruning to measure the chances of winning a game. This method is still widely used in games today. In
addition, Samuel also developed the minimax algorithm, which is a technique for minimizing losses in games.
History of machine Learning
1957: Frank Rosenblatt – The Perceptron
Frank Rosenblatt was a psychologist who is most famous for his work on machine learning. In 1957, he
developed the perceptron, which is a machine learning algorithm. The Perceptron was one of the first
algorithms to use artificial neural networks, widely used in machine learning.
It was designed to improve the accuracy of computer predictions. The goal of the Perceptron was to learn from
data by adjusting its parameters until it reached an optimal solution. Perceptron’s purpose was to make it easier
for computers to learn from data and to improve upon previous methods that had limited success.
History of machine Learning
1967: The Nearest Neighbor Algorithm
The Nearest Neighbor Algorithm was developed as a way to automatically identify patterns within large
datasets. The goal of this algorithm is to find similarities between two items and determine which one is closer
to the pattern found in the other item. This can be used for things like finding relationships between different
pieces of data or predicting future events based on past events.
In 1967, Cover and Hart published an article on “Nearest neighbor pattern classification.” It is a method of
inductive logic used in machine learning to classify an input object into one of two categories. The pattern
classifies the same items that are classified in the same categories as its nearest neighbors. This method is
used to classify objects with a number of attributes, many of which are categorical or numerical and may have
overlapping values.
History of machine Learning
1974: The Backpropagation
Backpropagation was initially designed to help neural networks learn how to recognize patterns. However,
it has also been used in other areas of machine learning, such as boosting performance and generalizing
from data sets to new instances. The goal of backpropagation is to improve the accuracy of a model by
adjusting its weights so that it can more accurately predict future outputs.
Paul Werbos laid the foundation for this approach to machine learning in his dissertation in 1974, which is
included in the book “The Roots of Backpropagation“.
History of machine Learning
1979: The Stanford Cart
The Stanford Cart is a remote-controlled robot that can move independently in space. It was first developed in
the 1960s and reached an important milestone in its development in 1979. The purpose of the Stanford
Cart is to avoid obstacles and reach a specific destination: In 1979, “The Cart” succeeded for the first time in
traversing a room filled with chairs in 5 hours without human intervention.
History of machine Learning
The AI Winter in the History of Machine Learning
AI has seen a number of highs and lows over the years.
The low point for AI was known as the AI winter, which
happened in the late 70s to the 90s. During this time,
research funding dried up and many projects were shut
down due to their lack of success. It has been
described as a series of hype cycles that have led to
disappointment and disillusionment among
developers, researchers, users, and media.
History of machine Learning
The Rise of Machine Learning in History
The rise of machine learning in the 21th century is a result of Moore’s Law and its exponential growth. When
computing power was becoming more affordable, it became possible to train AI algorithms using more data,
which resulted in an increase of the accuracy and efficiency of these algorithms.
1997: A Machine Defeats a Man in Chess
In 1997, the IBM supercomputer Deep Blue defeated chess grandmaster Garry Kasparov in a match. It was
the first time a machine had beaten an expert player at chess and it caused great concern for humans in the
chess community. This was a landmark event as it showed that AI systems could surpass human
understanding in complex tasks.
This marked a magical turning point in machine learning because the world now knew that mankind had
created its own opponent- an artificial intelligence that could learn and evolve on its own.
History of machine Learning
2002: Software Library Torch
Torch is a software library for machine learning and data science. Torch was created by Geoffrey Hinton,
Pedro Domingos, and Andrew Ng to develop the first large-scale free machine learning platform. In 2002, the
founders of Torch created it as an alternative to other libraries because they believed that their specific needs
were not met by other libraries. As of 2018, it has over 1 million downloads on Github and is one of the most
popular machine learning libraries available today.
Keep in mind: No longer in active development, however, PyTorch can be used, which is based on the Torch
Library.
History of machine Learning
2006: Geoffrey Hinton, the father of Deep Learning
In 2006, Geoffrey Hinton published his “A Fast Learning Algorithm for Deep Belief Nets.” This paper was
the birth of deep learning. He showed that by using a deep belief network, a computer could be trained
to recognize patterns in images.
Hinton’s paper described the first deep learning algorithm that can achieve human-level performance on
difficult and complex pattern recognition tasks.
History of machine Learning
2011: Google Brain
Google Brain is a research group of Google devoted to artificial intelligence and machine learning. The group
was founded in 2011 by Google X and is located in Mountain View, California. The team works closely with other
AI research groups within Google such as the DeepMind group that has developed AlphaGo, an AI that defeated
the world champion at Go. Their goal is to build machines that can learn from data, understand language,
answer questions in natural language, and have common sense reasoning.
The group is, as of 2021, led by Geoffrey Hinton, Jeff Dean and Zoubin Ghahramani and focuses on deep
learning, a model of artificial neural networks that is capable to learn complex patterns from data automatically
without being explicitly programmed.
History of machine Learning
2014: DeepFace
DeepFace is a deep learning algorithm which was originally developed in 2014 and is part of the company “Meta”.
The project received significant media attention after it outperformed human performance on the well-known
“Faces in the Wild” test.
DeepFace is based on a deep neural network, which consists of many layers of artificial neurons and weights that
connect each layer to its neighboring ones. The algorithm takes as input a training data set of photographs, with
each photo annotated with the identity and age of its subject. The team has been very successful in recent years
and published many papers on their research results. They have also trained several deep neural networks that
have achieved significant success in pattern recognition and machine learning tasks.
History of machine Learning
2017: ImageNet Challenge – Milestone in the History of Machine Learning
The ImageNet Challenge is a competition in computer vision that has been running since 2010. The
challenge focuses on the abilities of programs to process patterns in images and recognize objects with
varying degrees of detail.
In 2017, a milestone was reached. 29 out of 38 teams achieved 95% accuracy with their computer vision
models. The improvement in image recognition is immense.
State of the art Applications of machine Learning
State of the art Applications of machine Learning
State of the art Applications of machine Learning
State of the art Applications of machine Learning
State of the art Applications of machine Learning
State of the art Applications of machine Learning
State of the art Applications of machine Learning
History of Machine Learning Extension
History of Machine Learning Extension
History of Machine Learning Extension
Data in Machine Learning
➢ Machine learning depends largely on test data.
➢ A large amount of data is required for ML.
Types of Data in ML
Categorical Numerical
Types of Data in ML
(Attributes of the data set)
Types of Data in ML
Types of Data in ML
Types of Data in ML
Types of Data in ML
Types of Data in ML
Types of Data in ML
4. Colour: Red, Blue etc.
Types of Data in ML
Types of Data in ML
The data is arranged in an
order so we can find its median
value.
Types of Data in ML
Types of Data in ML
Types of Data in ML
Types of Data in ML
Types of Data in ML
Time-Series Data
Time-series data in machine
learning falls under the category
of quantitative data. This is
because time-series data is
numerical and can be measured
or counted. It typically consists of
real numbers that represent a
certain variable over different
points in time. This type of data
is often used in forecasting
models, trend analysis, and other
statistical methods within
machine learning.
Types of Data in ML
Data Attributes in ML
Data Attributes in ML
Machine learning Activities
Machine learning Activities
Data Preprocessing in ML
Data Preprocessing in ML
Data Preprocessing in ML
Data Preprocessing in ML
What is data ?
Data are raw facts that have not been processed to
explain their meaning.
There are three types of Data
What is structured data?
Structured data is data whose elements are addressable for effective analysis. It has been organized into a
formatted repository that is typically a database. It concerns all data which can be stored in
database SQL in a table with rows and columns. They have relational keys and can easily be mapped into
pre-designed fields. Today, those data are most processed in the development and simplest way to
manage information. Example: Relational data.
What is structured data?
Data
What is Unstructured Data?
Unstructured data is a data which is not organized in a predefined manner or does not have a predefined data
model, thus it is not a good fit for a mainstream relational database. So for Unstructured data, there are
alternative platforms for storing and managing, it is increasingly prevalent in IT systems and is used by
organizations in a variety of business intelligence and analytics applications. Example: Word, PDF, Text, Media
logs.
What is Unstructured Data?
What is Semi-structured data
Semi-structured data is a type of data that is not purely structured, but also not completely
unstructured. It contains some level of organization or structure, but does not conform to a rigid schema
or data model, and may contain elements that are not easily categorized or classified.
Advantages of Semi-structured Data:
•The data is not constrained by a fixed schema
•Flexible i.e Schema can be easily changed.
•Data is portable
•It is possible to view structured data as semi-structured data
•Its supports users who can not express their need in SQL
•It can deal easily with the heterogeneity of sources.
•Flexibility: Semi-structured data provides more flexibility in terms of data storage
and management, as it can accommodate data that does not fit into a strict,
predefined schema. This makes it easier to incorporate new types of data into an
existing database or data processing pipeline.
Comparison
Properties Structured data Semi-structured data Unstructured data
It is based on Relational database It is based on XML/RDF(Resource It is based on character and binary
Technology
table Description Framework). data
Matured transaction and various Transaction is adapted from DBMS No transaction management and no
Transaction management
concurrency techniques not matured concurrency
Versioning over tuples or graph is
Version management Versioning over tuples,row,tables Versioned as a whole
possible
It is more flexible than structured
It is schema dependent and less It is more flexible and there is
Flexibility data but less flexible than
flexible absence of schema
unstructured data
It is very difficult to scale DB It’s scaling is simpler than structured
Scalability It is more scalable.
schema data
Robustness Very robust New technology, not very spread —
Structured query allow complex Queries over anonymous nodes are
Query performance Only textual queries are possible
joining possible
AI Canonical Architecture
Unstructured and Structured Data
BIG DATA
Data which are very large in size is called Big Data. Normally we work on data of size MB(WordDoc ,Excel) or
maximum GB(Movies, Codes) but data in Peta bytes i.e. 10^15 byte size is called Big Data. It is stated that almost
90% of today's data has been generated in the past 3 years.
Sources of Big Data
•Social networking sites: Facebook, Google, LinkedIn all these sites generates huge amount of
data on a day to day basis as they have billions of users worldwide.
•E-commerce site: Sites like Amazon, Flipkart, Alibaba generates huge amount of logs from
which users buying trends can be traced.
•Weather Station: All the weather station and satellite gives very huge data which are stored
and manipulated to forecast weather.
•Telecom company: Telecom giants like Airtel, Vodafone study the user trends and accordingly
publish their plans and for this they store the data of its million users.
•Share Market: Stock exchange across the world generates huge amount of data through its
daily transaction.
BIG DATA
BIG DATA
BIG DATA Characteristics
•Volume: the size and amounts of big data that companies manage
and analyze
•Value: the most important “V” from the perspective of the business,
the value of big data usually comes from insight discovery and pattern
recognition that lead to more effective operations, stronger customer
relationships and other clear and quantifiable business benefits
•Variety: the diversity and range of different data types, including
unstructured data, semi-structured data and raw data
•Velocity: the speed at which companies receive, store and manage
data – e.g., the specific number of social media posts or search
queries received within a day, hour or other unit of time
•Veracity: the “truth” or accuracy of data and information assets,
which often determines executive-level confidence
The additional characteristic of variability can also be considered:
•Variability: the changing nature of the data companies seek to
capture, manage and analyze – e.g., in sentiment or text analytics,
changes in the meaning of key words or phrases
BIG DATA
➢ Big data and machine learning are like a dynamic duo, each having its strengths.
➢ They're not rivals; instead, they work well together.
➢ When you use them together, amazing things can happen.
➢ Think of big data's 5Vs (volume, velocity, variety, veracity, and value) - machine learning helps
handle them and make accurate predictions.
➢ On the flip side, when creating machine learning models, big data plays a role by giving top-notch
data and enhancing learning methods through analytics.
➢ It's like they bring out the best in each other!
➢ There is no secret that almost all organizations, such as Google, Amazon, IBM, Netflix, etc., have
already discovered the power of big data analytics enhanced by machine learning.
➢ Machine Learning is a very crucial technology, and with big data, it has become more powerful for
data collection, data analysis, and data integration.
➢ All big organizations use machine learning algorithms for running their business properly.
BIG DATA
We can apply machine learning algorithms to every element of Big data operation, including:
•Data Labeling and Segmentation
•Data Analytics
•Scenario Simulation
In machine learning algorithms, we need multiple varieties of data for training a machine and predicting accurate results.
However, sometimes it becomes difficult to manage these bulkified data.
So, it becomes a challenge to manage and analyze Big Data.
Further, this unstructured data is useless until it is well interpreted.
Thus, to use information, there is a need for talent, algorithms, and computing infrastructure.
Machine Learning enables machines or systems to learn from past experience and use data received from big data, and
predict accurate results.
Hence, this leads to generating improved quality business operations and building better customer relationship
management.
Big Data helps machine learning by providing a variety of data so machines can learn more or multiple samples or training
data.
In such ways, businesses can accomplish their dreams and get the benefit of big data using ML algorithms. However, for
using the combination of ML and big data, companies need skilled data scientists.
How Does Big Data Analytics Work?
Companies need to work around analytics applications, partner with data scientists and engage
with other data analysts to extract relevant and valid insights from big data. In addition, they must
have an enhanced understanding of all available data. Finally, the analytics team also needs to
clarify what they want to extract from the data.
The team needs to take care of :
•Cleansing,
•Profiling,
•Transformation,
•Validation of data sets.
These are some of the most important initial steps taken in data analysis.
Once all the big data has been prepared and gathered for interpretation, a combination of
advanced data science and analytics disciplines is applied through different machine learning
tools.
This will help to generate results that lead to businesses growth and development.
Leveraging Machine Learning
Leveraging Machine Learning:
•It involves understanding how to effectively use machine learning algorithms and techniques to solve real-world problems. It
covers the application of machine learning in various domains, such as healthcare, finance, marketing, and more. How to
choose appropriate machine learning models, preprocess data, train models, and evaluate their performance? Additionally, it
may include discussions on the ethical considerations and challenges associated with leveraging machine learning in different
industries.
Algorithms and Techniques:
[Link] Learning Algorithms:
1. Regression: Predicting a continuous outcome.
2. Classification: Assigning labels to data points (e.g., spam or not spam).
[Link] Learning Algorithms:
1. Clustering: Grouping similar data points together (e.g., customer segmentation).
2. Dimensionality Reduction: Reducing the number of features while preserving essential information.
[Link]-Supervised and Self-Supervised Learning:
1. Learning from a combination of labeled and unlabeled data.
2. Learning from data without explicit labels, using the inherent structure of the data.
[Link] Learning:
1. Teaching models to make decisions by trial and error, receiving feedback in the form of rewards or penalties.
[Link] Networks and Deep Learning:
1. Applying deep neural networks for complex tasks like image recognition, natural language processing, and speech
recognition.
BIG DATA Real-World Problems
Real-World Problems:
[Link]:
1. Predicting patient outcomes based on medical records.
2. Personalized treatment recommendations.
[Link]:
1. Credit scoring and risk assessment.
2. Fraud detection in financial transactions.
[Link]:
1. Customer segmentation and targeted advertising.
2. Predicting customer churn and recommending retention strategies.
[Link]:
1. Predictive maintenance to reduce equipment downtime.
2. Quality control in production processes.
5.E-commerce:
1. Recommender systems for personalized product recommendations.
2. Fraud detection in online transactions.
BIG DATA
[Link] Vehicles:
1. Object detection and recognition for safe navigation.
2. Decision-making algorithms for route planning.
[Link] Language Processing (NLP):
1. Sentiment analysis for customer reviews.
2. Language translation and chatbot applications.
[Link] and Video Analysis:
1. Facial recognition for security.
2. Object detection in video surveillance.
[Link] Prediction:
1. Predicting weather patterns and climate trends.
2. Analyzing environmental data for sustainable practices.
[Link]:
1. Personalized learning plans based on student performance.
2. Early detection of learning difficulties.
What is analytics?
• "Analytics in general, involves the use of mathematical or scientific methods to generate insight from
data"
Data Analytics
Data analytics is the process of examining, cleaning, transforming, and modeling data with the goal of discovering
useful information, drawing conclusions, and supporting decision-making. It involves the use of various techniques
and tools to analyze patterns, trends, and relationships within datasets.
Here are key aspects of data analytics:
[Link] Collection:
•Gathering relevant data from various sources, which can include databases, spreadsheets, sensors, social
media, and more.
[Link] Cleaning:
•Ensuring that the collected data is accurate, complete, and free from errors or inconsistencies.
[Link] Transformation:
•Converting raw data into a format suitable for analysis. This may involve normalization, aggregation, or other
preprocessing steps.
[Link] Analysis:
•Applying statistical and mathematical methods, as well as using tools and algorithms, to uncover patterns,
correlations, and trends in the data.
[Link] Visualization:
•Representing the results of data analysis through charts, graphs, and other visualizations to make complex
information more understandable.
[Link] Analytics:
•Summarizing and interpreting historical data to understand what has happened. It involves the examination of
past events and their characteristics.
BIG DATA
[Link] Analytics:
•Using statistical algorithms and machine learning techniques to make predictions about future
events based on historical data.
[Link] Analytics:
•Recommending actions to optimize outcomes based on the insights gained from data analysis.
[Link] Intelligence:
•Leveraging data to support business decision-making by providing actionable insights and
relevant information.
[Link] Data Analytics:
•Analyzing large and complex datasets, often characterized by the three Vs (volume, velocity,
and variety), using specialized tools and technologies.
Data analytics plays a crucial role in various fields, including business, healthcare, finance,
marketing, and science. It helps organizations make informed decisions, identify opportunities, and
address challenges by extracting meaningful insights from data. The growing importance of data
analytics has led to the development of specialized roles and tools to meet the increasing demand
for skilled professionals in this field.
BIG DATA
BIG DATA
BIG DATA
BIG DATA
Descriptive vs Predictive Analytics:
•Descriptive analytics involves analyzing historical data to understand what has
happened in the past. It focuses on summarizing and interpreting data to gain
insights into trends and patterns. This type of analytics is useful for reporting and
data visualization.
•Predictive analytics, on the other hand, aims to forecast future outcomes based
on historical data and statistical algorithms. It involves the use of machine
learning models to make predictions and identify potential trends. Predictive
analytics is valuable for making informed decisions and planning for the future.
BIG DATA
BIG DATA
BIG DATA
BIG DATA
BIG DATA