Module 5:
Planning, Learning
Pratik Oak
Department of CSE AIML
Learning check-points
➔ Planning
➔ Learning categories- Supervised, unsupervised,
semi-supervised and reinforcement learning
12-2 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
AI Objectives
■ Make machines smarter (primary goal)
■ Understand what intelligence is
■ Make machines more intelligent and useful
■ Signs of intelligence…
■ Learn or understand from experience
■ Make sense out of ambiguous situations
■ Respond quickly to new situations
■ Use reasoning to solve problems
■ Apply knowledge to manipulate the environment
12-3 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Test for Intelligence
Turing Test for Intelligence
■ A computer can be
considered to be smart
only when a human
interviewer, “conversing”
with both an unseen
human being and an
unseen computer, can
not determine which is
which.
- Alan Turing
12-4 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Symbolic Processing
■ AI …
■ represents knowledge as a set of symbols, and
■ uses these symbols to represent problems, and
■ apply various strategies and rules to manipulate
symbols to solve problems
■ A symbol is a string of characters that stands for
some real-world concept (e.g., Product, consumer,…)
■ Examples:
■ (DEFECTIVE product)
■ (LEASED-BY product customer) - LISP
■ Tastes_Good (chocolate)
12-5 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
AI Concepts
■ Reasoning
■ Inferencing from facts and rules using heuristics or other
search approaches
■ Pattern Matching
■ Attempt to describe and match objects, events, or processes
in terms of their qualitative features and logical and
computational relationships
■ Knowledge Base
12-6 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Evolution of artificial intelligence
12-7 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Artificial vs. Natural Intelligence
■ Advantages of AI
■ More permanent
■ Ease of duplication and dissemination
■ Less expensive
■ Consistent and thorough
■ Can be documented
■ Can execute certain tasks much faster
■ Can perform certain tasks better than many people
■ Advantages of Biological Natural Intelligence
■ Is truly creative
■ Can use sensory input directly and creatively
■ Can apply experience in different situations
12-8 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
The AI Field…
■ AI provides the
scientific
foundation for
many commercial
technologies
12-9 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
AI is often transparent in many
commercial products
■ Anti-lock Braking Systems (ABS)
■ Automatic Transmissions
■ Video Camcorders
■ Appliances
■ Washers, Toasters, Stoves
■ Help Desk Software
■ Subway Control…
12-10 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Planning
• We studied how to take actions in the
world (search)
• We studied how to represent objects,
relations, etc. (logic)
• Now we will combine the two!
12-11 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Planning
Planning is an integral part of automation
Recommended clip from Charlie Chaplin’s Modern
Times to see what can go wrong:
[Link]
12-12 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Planning
Always focus on understanding-
Action
➔ Preconditioning
➔ Effects
12-13 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Planning to write a paper
12-14 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
General Purpose Vs Special Purpose Planning
Planning: find a sequence of actions to achieve a goal
General purpose: symbolic descriptions of the problems and
the domain. The plan generation algorithm the same
Advantage: - opportunity to have clear semantics
Disadvantage: - symbolic description requirement
Domain Specific: The plan generation algorithm depends on
the particular domain
Advantage: - can be very efficient
Disadvantage: - lack of clear semantics
- knowledge-engineering for plan generation
12-15 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Types of planning
Forward State Space Planning Backward State Space Planning
(FSSP) (BSSP)
1 Initial State to new state Target State to sub-goal state
2 Large Branching factor Small Branching factor
3 Algorithm- sound Not sound algorithm (inconsistent)
12-16 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Partial order planning (POP)
★ Order is partial
★ Doesn’t specify which action will come first
out of two actions
★ Problem decomposition to work in
non-operative environment
12-17 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
POP of wearing shoe
12-18 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
POP as search problem
Set of actions:
These are the steps of plan.
For e.g.: Set of Actions = {Start, Rightsock, Rightshoe ,Leftsock, Leftshoe, Finish}
Set of ordering constraints/preconditions:
i. Preconditions are considered as ordering constraints.(i.e. without performing action “x” we
cannot perform action “y”)
ii. For e.g.: Set of ordering = {Right-sock <right-shoe; left-sock<left-shoe}="" that="" is="" in=""
order="" to="" wear="" shoe,="" first="" we="" should="" wear="" a="" sock.<="" p="">
Set of causal links:
Action A achieves effect “E” for action B
12-19 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Total Order Planning (TOP)
➢ Explore linear sequences of actions from start to
goal state
➢ No problem decomposition
12-20 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Comparison of POP & TOP for wearing shoes
12-21 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
POP TOP
Sequence of actions not fully Sequence of action fully
specified specified
Determines order of actions Determines order of action
dynamically before plan is executed
More flexible to planning More deterministic to planning
Used in high degree of Used for well-defined
uncertainty problems
12-22 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Learning categories
➔ Supervised
➔ Unsupervised
➔ Semi-supervised
➔ Reinforcement learning
12-23 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Supervised Learning
Known labels
12-24 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Supervised Learning Steps
12-25 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Supervised Learning Algorithm Categories
Algorithms->
12-26 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Supervised Learning Pros & Cons
Pros
Cons
12-27 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Un-supervised Learning
Unknown labels
12-28 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Unsupervised Learning Algorithm Categories
Algorithms->
12-29 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Unsupervised Learning Pros & Cons
Pros
Cons
12-30 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Semi-supervised Learning
12-31 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Semi-supervised Learning: When to use?
12-32 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Semi-supervised Learning Algorithm Categories
Algorithms
12-33 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Semi-supervised Learning Real Life Applications
Fraud detection
Semi-supervised learning can be used to train systems to identify fraud and legitimate
cases by starting with a small set of labeled transactions.
Customer segmentation
Semi-supervised learning can be used to define initial segments based on a small
labeled dataset, and then refine and expand those categories with a larger pool of
unlabeled data.
Image classification
Semi-supervised learning can be used to train a model on a small subset of labeled
images, and then use that model to predict labels for the rest of the images.
12-34 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Semi-supervised Learning Real Life Applications
Text document classification
When there's a large amount of data, like emails or product reviews, semi-supervised
learning can be used to label a small portion of the data and then use the rest to refine
the model.
Medical image analysis
Semi-supervised learning can be used to supplement medical experts' analysis of images
like MRIs and X-rays with unlabeled images.
Speech recognition
Semi-supervised learning can be used to combine labeled speech data with unlabeled
audio to improve the model's ability to understand what's being said.
12-35 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Reinforcement Learning
12-36 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Reinforcement Learning contd…
12-37 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Elements of Reinforcement Learning
12-38 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Reinforcement Learning Algorithms
● Deep Q-Learning Network (DQN)
● Deep Deterministic Policy Gradient (DDPG)
● Twin Delayed Deep Deterministic (TD3)
● Asynchronous Advantage Actor-Critic (A3C)
● Advantage Actor-Critic (A2C)
● Markov decision process (MDP)
● Bellman equation
● Dynamic programming
● Value iteration
● Policy iteration
12-39 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Reinforcement Learning Real Life Applications
Self-driving cars
In a simulated environment, a self-driving car learns which actions to take in different situations,
such as navigating city traffic. Once deployed in the real world, the car can refine its learned
policy with new data.
Chatbots
Self-improving chatbots use reinforcement learning to select responses based on user
conversations. For example, MILABOT uses reinforcement learning to converse with humans
about small talk topics.
Personalized treatment plans
In healthcare, reinforcement learning can be used to create personalized treatment plans for
patients with long-term illnesses.
12-40 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
Reinforcement Learning Real Life Applications
Optimizing inventory levels
Walmart uses reinforcement learning to optimize price reductions for unsold inventory.
Manufacturing efficiency
Reinforcement learning can be used to optimize manufacturing processes, leading to
increased production efficiency and reduced waste.
Energy consumption optimization
Financial trading strategies
Game AI
Smart grid management
12-41 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall
References
1. 2011 Pearson Education, Inc. Publishing as Prentice Hall
2. N. P. Padhy, ―Artificial Intelligence and Intelligent Systems, Oxford
University Press.
12-42 Copyright © 2011 Pearson Education, Inc. Publishing as Prentice Hall