Artificial Intelligence and Embedded
System
Dr. Arijit Ukil
Principal Scientist, TCS Research
Arijit Ukil
Supervised Learning Pipeline
[Link]
Arijit2 Ukil
Supervised Learning Pipeline
[Link]
Arijit3 Ukil
Unsupervised Learning Pipeline
[Link]
Arijit4 Ukil
Geometric View of ML Problems
[Link]
Arijit5 Ukil
Broader Classes of Learning (Detour)
Machine learning models can be classified
into two categories:
▪ Generative
▪ Discriminative
[Link]
Arijit6 Ukil
Broader Classes of Learning (Detour)
Discriminative models:
• learn decision boundaries that separate different
classes.
• maximize the conditional probability: P(Y|X) —
Given an input X, maximize the probability of
label Y.
• are meant explicitly for classification tasks.
Examples include: Logistic regression, Random
Forest, Neural Networks, Decision Trees, etc.
[Link]
Arijit7 Ukil
Broader Classes of Learning (Detour)
[Link]
Arijit8 Ukil
Broader Classes of Learning (Detour)
Generative models:
• maximize the joint probability: P(X, Y)
• learn the class-conditional distribution P(X|Y)
• are typically not meant for classification tasks.
Examples include: Naive Bayes, Gaussian Mixture
Models, Generative Adversarial Networks
(GAN),etc.
[Link]
Arijit9 Ukil
Broader Classes of Learning (Detour)
[Link]
Arijit10 Ukil
Broader Classes of Learning (Detour)
[Link]
Arijit11 Ukil
Data Types (Supervised Learning)
Source- [Link]
Arijit12 Ukil
Types of Probability
Joint Probability
Probability of two events occurring together:
𝑃 𝐴∩𝐵 .
Marginal Probability
Probability of a single event, regardless of others:
𝑃 𝐴 = 𝑃 𝐴 ∩ 𝐵
𝐵
Conditional Probability
Probability of one event given another:
𝑃 𝐴∩𝐵
𝑃 𝐴∣𝐵 =
𝑃 𝐵
Arijit13 Ukil
Types of Probability
Arijit14 Ukil
Notation
Source- [Link]
Arijit15 Ukil
Notation
Source- [Link]
Arijit16 Ukil
Data Generating Process
Source- [Link]
Arijit17 Ukil
Data Generating Process- i.i.d
Source- [Link]
Arijit18 Ukil
Regression vs Classification
Source- [Link]
Arijit19 Ukil
What is a Model?
Source- [Link]
Arijit20 Ukil
What is a Hypothesis?
Source- [Link]
Arijit21 Ukil
What is a Hypothesis?
Source- [Link]
Arijit22 Ukil
Parametrization
Source- [Link]
Arijit23 Ukil
Parametrization
Source- [Link]
Arijit24 Ukil
Example: Univariate Linear Functions
Source- [Link]
Arijit25 Ukil
Example: Bivariate Quadratic Functions
Source- [Link]
Arijit26 Ukil
Example: Bivariate Quadratic Functions
Source- [Link]
Arijit27 Ukil
What is a Learner?
Source- [Link]
Arijit28 Ukil
What is a Learning?
Source- [Link]
Arijit29 Ukil
ML Loss
Source- [Link]
Arijit30 Ukil
ML Loss
Source- [Link]
Arijit31 Ukil
Supervised Learning Revisited
[Link]
Arijit32 Ukil
Un/Supervised Learning Revisited
[Link]
Arijit33 Ukil
Un/Supervised Learning Revisited
[Link]
Arijit34 Ukil
Probabilistic Modeling
[Link]
Arijit35 Ukil
Probabilistic Modeling
[Link]
Arijit36 Ukil
Parameter Estimation in Probabilistic Modeling
[Link]
Arijit37 Ukil
Parameter Estimation in Probabilistic Modeling
[Link]
Arijit38 Ukil
Parameter Estimation in Probabilistic Modeling
[Link]
Arijit39 Ukil
Parameter Estimation in Probabilistic Modeling
[Link]
Arijit40 Ukil
Parameter Estimation in Probabilistic Modeling
[Link]
Arijit41 Ukil
Log Likelihood
[Link]
Arijit42 Ukil
Maximum Log Likelihood
[Link]
Arijit43 Ukil
Maximum Log Likelihood
[Link]
Arijit44 Ukil
Maximum Log Likelihood
[Link]
Arijit45 Ukil
Maximum Log Likelihood
[Link]
Arijit46 Ukil
Maximum-a-Posteriori (MAP) Estimation
[Link]
Arijit47 Ukil
Maximum-a-Posteriori (MAP) Estimation
[Link]
Arijit48 Ukil
Maximum-a-Posteriori (MAP) Estimation
[Link]
Arijit49 Ukil
Maximum-a-Posteriori (MAP) Estimation
[Link]
Arijit50 Ukil
Maximum-a-Posteriori (MAP) Estimation
[Link]
Arijit51 Ukil
Maximum-a-Posteriori (MAP) Estimation
[Link]
Arijit52 Ukil
Maximum-a-Posteriori (MAP) Estimation
[Link]
Arijit53 Ukil
General Problem Setup
• In supervised learning setting, we find a model or function
ℎθ . , parameterized by 𝜃 that describes the random vector
𝒙 associated with label or target 𝑦 with joint distribution
𝑝𝑑𝑎𝑡𝑎 𝒙, 𝑦 .
• 𝕏 𝑇𝑟𝑎𝑖𝑛 = 𝒙1 , 𝒙2 , … , 𝒙𝑁 associated with labels 𝑌𝑇𝑟𝑎𝑖𝑛
= 𝑦1 , 𝑦 2 , … , 𝑦 𝑁 .
• 𝑥 1 , 𝑦1 , 𝑥 2 , 𝑦 2 , … , 𝑥 𝑁 , 𝑦 𝑁 ~𝑝𝑑𝑎𝑡𝑎 under i.i.d.
assumption.
• Model learning is imposing independent and identically
distribution (i.i.d.) condition.
Arijit54 Ukil
General Problem Setup
• In machine learning, the principal aim is to minimize an
objective function that penalizes the model ℎθ when it makes
mistake, which is denoted by loss function ℒ ℎθ 𝒙 , 𝑦 .
• We minimize the expected risk:
ℛ ℎ =𝔼 𝒙,𝑦 ~𝑝𝑑𝑎𝑡𝑎 ℒ ℎθ 𝒙 , 𝑦 .
• We do not have complete idea of 𝑝𝑑𝑎𝑡𝑎 𝒙, 𝑦 .
• We simply know the training dataset 𝕏 𝑇𝑟𝑎𝑖𝑛 , 𝑌𝑇𝑟𝑎𝑖𝑛 .
• Hence, we focus on empirical risk minimization (ERM):
1 𝑁
ℛ𝑒𝑚𝑝 ℎ = σ𝑛=1 ℒ ℎθ 𝒙𝑛 , 𝑦 𝑛 .
𝑁
Arijit55 Ukil
General Problem Setup
• Considering negative log-likelihood as the loss function
under maximum likelihood estimation (MLE) principle,
which is a special case of ERM, the MLE cost function is:
𝒥 𝜃 = 𝔼 𝒙,𝑦 ~𝑝ො𝑑𝑎𝑡𝑎 − log 𝑝𝜃 𝑦 𝒙 and the optimization
problem: 𝜃 ∗ =argmin 𝒥 𝜃 .
𝜃
• when 𝑁 is small, it is not practical to assume the closeness of
𝑝Ƹ 𝑑𝑎𝑡𝑎 and 𝑝𝑑𝑎𝑡𝑎 .
• Consequently, the learned model ℎθ is not properly trained as
𝒥 𝜃 is poorly estimated.
• Practical ML problems often suffer from training data
scarcity issue due to the expenses associated with the expert-
intervened annotation efforts.
Arijit56 Ukil
Thank You
Arijit57 Ukil
Q&A
Arijit58 Ukil