Bayes Classifier
Reyyan Yeniterzi
Bayesian Reasoning
In machine learning, goal is to determine the best hypothesis from
the hypothesis space given the observe data
What is the best hypothesis?
The most probable hypothesis
Bayes theorem provides a direct way to calculate this probability
2 Machine Learning - Reyyan Yeniterzi
Bayes Optimal Classifier
Learn the function 𝑓 which maps 𝑋 to 𝑌
𝑋: features
𝑌: target classes (labels)
Suppose you know 𝑃(𝑌|𝑋) exactly, how should you classify?
3 Machine Learning - Reyyan Yeniterzi
Bayes Optimal Classifier
Learn the function 𝑓 which maps 𝑋 to 𝑌
𝑋: features
𝑌: target classes (labels)
Suppose you know 𝑃(𝑌|𝑋) exactly, how should you classify?
Bayes optimal classifier:
𝑦ො = 𝑎𝑟𝑔𝑚𝑎𝑥𝑦 𝑃(𝑦|𝑥)
4 Machine Learning - Reyyan Yeniterzi
Remember the Bayes Theorem
likelihood prior
posterior
𝑃(𝑥|𝑦)𝑃(𝑦)
𝑃(𝑦|𝑥) = 𝑎𝑟𝑔𝑚𝑎𝑥𝑦𝜖𝑌
𝑃(𝑥)
normalization
6 Machine Learning - Reyyan Yeniterzi
Remember the Bayes Theorem
likelihood prior
posterior
𝑃(𝐷|ℎ)𝑃(ℎ)
𝑃(ℎ|𝐷) = 𝑎𝑟𝑔𝑚𝑎𝑥ℎ𝜖𝐻
𝑃(𝐷)
𝐷 is the data and ℎ is an hypothesis normalization
with the hypothesis space 𝐻
7 Machine Learning - Reyyan Yeniterzi
Bayes Theorem
Prior probability:
the initial probability that hypothesis ℎ holds before observing the
training data 𝐷
reflects any background knowledge that we have on ℎ
If there is no such prior background knowledge?
Assign the same prior probability to each candidate hypothesis
Posterior probability: (depends on the data)
reflects our confidence that ℎ holds after we have seen the training data
𝐷
8 Machine Learning - Reyyan Yeniterzi
An example
9 Machine Learning - Reyyan Yeniterzi
An example
10 Machine Learning - Reyyan Yeniterzi
An example
Given a new patient with positive lab results, should we diagnose
the patient as having cancer or not?
11 Machine Learning - Reyyan Yeniterzi
An example
Given a new patient with positive lab results, should we diagnose
the patient as having cancer or not?
12 Machine Learning - Reyyan Yeniterzi
An example
Given a new patient with positive lab results, should we diagnose
the patient as having cancer or not?
Answer is obvious, but the exact posterior probability can be
determined after normalization
13 Machine Learning - Reyyan Yeniterzi
An example
Prior
Posterior
Posterior probability of cancer is significantly higher than its prior
probability
The most probable hypothesis is still that the patient does not have
cancer
14 Machine Learning - Reyyan Yeniterzi
Remember the Bayes Theorem
likelihood prior
posterior
𝑃(𝑥|𝑦)𝑃(𝑦)
𝑃(𝑦|𝑥) = 𝑎𝑟𝑔𝑚𝑎𝑥𝑦𝜖𝑌
𝑃(𝑥)
normalization
15 Machine Learning - Reyyan Yeniterzi
A Bayesian Classifier
𝑃(𝑥|𝑦)𝑃(𝑦)
𝑦ො = 𝑎𝑟𝑔𝑚𝑎𝑥𝑦𝜖𝑌
𝑃(𝑥)
Does 𝑃(𝑥) have any affect on the choice of 𝑦?
ො
16 Machine Learning - Reyyan Yeniterzi
A Bayesian Classifier
𝑃(𝑥|𝑦)𝑃(𝑦)
𝑦ො = 𝑎𝑟𝑔𝑚𝑎𝑥𝑦𝜖𝑌
𝑃(𝑥)
Does 𝑃(𝑥) have any affect on the choice of 𝑦?
ො
Therefore drop it
𝑦ො = 𝑎𝑟𝑔𝑚𝑎𝑥𝑦𝜖𝑌 𝑃(𝑥|𝑦)𝑃(𝑦)
17 Machine Learning - Reyyan Yeniterzi
A Bayesian Classifier
prior
𝑦ො = 𝑎𝑟𝑔𝑚𝑎𝑥𝑦𝜖𝑌 𝑃(𝑥|𝑦)𝑃(𝑦)
likelihood
18 Machine Learning - Reyyan Yeniterzi
A Bayesian Classifier
𝑦ො = 𝑎𝑟𝑔𝑚𝑎𝑥𝑦𝜖𝑌 𝑃(𝑥|𝑦)𝑃(𝑦)
Each data instance is represented with set of features (attributes)
𝑃 𝑥 𝑦 = 𝑃(𝑥1 , 𝑥2 , … 𝑥𝑁 |𝑦)
How many parameters do we need to estimate?
19 Machine Learning - Reyyan Yeniterzi
Parameters of 𝑃 𝑥 𝑦
𝑃 𝑥 𝑦 = 𝑃(𝑥1 , 𝑥2 , … 𝑥𝑁 |𝑦)
How many parameters do we need to estimate?
Lets simplify
Assume all attributes and labels are Boolean variables
So at this case how many parameters do we need to estimate?
20 Machine Learning - Reyyan Yeniterzi
Parameters of 𝑃 𝑥 𝑦
How many rows are there in this table
21 Machine Learning - Reyyan Yeniterzi
Parameters of 𝑃 𝑥 𝑦
How many rows are there in this table
22 Machine Learning - Reyyan Yeniterzi
Parameters of 𝑃 𝑥 𝑦
How many rows are there in this table
23 Machine Learning - Reyyan Yeniterzi
Parameters of 𝑃 𝑥 𝑦
We also need another table for 𝑦 = 1
24 Machine Learning - Reyyan Yeniterzi
Parameters of 𝑃 𝑥 𝑦
We also need another table for 𝑦 = 1
25 Machine Learning - Reyyan Yeniterzi
Parameters of 𝑃(𝑦)
What about 𝑃(𝑦)?
26 Machine Learning - Reyyan Yeniterzi
Parameters of 𝑃(𝑦)
What about 𝑃(𝑦)?
1
27 Machine Learning - Reyyan Yeniterzi
Parameters of 𝑃 𝑥 𝑦 and 𝑃(𝑦)
𝑃 𝑥 𝑦 = 𝑃 𝑥1 , 𝑥2 , … 𝑥𝑁 𝑦
# 𝑝𝑎𝑟𝑎𝑚𝑒𝑡𝑒𝑟𝑠 = ( 𝐹 𝑁 −1)𝐶
𝑃 𝑦
# 𝑝𝑎𝑟𝑎𝑚𝑒𝑡𝑒𝑟𝑠 = 𝐶 − 1
28 Machine Learning - Reyyan Yeniterzi
Parameters of 𝑃 𝑥 𝑦
Can we reduce the parameters of 𝑃 𝑥 𝑦 ?
29 Machine Learning - Reyyan Yeniterzi
Parameters of 𝑃 𝑥 𝑦
Can we reduce the parameters of 𝑃 𝑥 𝑦 ?
Yes, by using a simplifying assumption
Conditional Independence (Naïve Bayes) Assumption
Assume that the feature probabilities 𝑃(𝑥𝑖 |𝑦) are independent given the
class 𝑦
30 Machine Learning - Reyyan Yeniterzi
Parameters of 𝑃 𝑥 𝑦
Can we reduce the parameters of 𝑃 𝑥 𝑦 ?
Yes, by using a simplifying assumption
Conditional Independence (Naïve Bayes) Assumption
Assume that the feature probabilities 𝑃(𝑥𝑖 |𝑦) are independent given the
class 𝑦
P 𝑥1 , 𝑥2 … 𝑥𝑁 𝑦 = 𝑃 𝑥1 𝑦 𝑃 𝑥2 𝑦 … 𝑃 𝑥𝑁 𝑦
31 Machine Learning - Reyyan Yeniterzi
Naïve Bayes Assumption
Conditional Independence (Naïve Bayes) Assumption
Assume that the feature probabilities 𝑃(𝑥𝑖 |𝑦) are independent given the
class 𝑦
P 𝑥1 , 𝑥2 … 𝑥𝑁 𝑦 = 𝑃 𝑥1 𝑦 𝑃 𝑥2 𝑦 … 𝑃 𝑥𝑁 𝑦
P 𝑥1 , 𝑥2 … 𝑥𝑁 𝑦 = 𝑎𝑟𝑔𝑚𝑎𝑥𝑦𝜖𝑌 ෑ 𝑃 𝑥 𝑦
𝑥𝜖𝐹
32 Machine Learning - Reyyan Yeniterzi
Naïve Bayes Assumption
Conditional Independence (Naïve Bayes) Assumption
What is the parameter size now?
again for Boolean features for 𝑥1 , 𝑥2 , … 𝑥𝑁
33 Machine Learning - Reyyan Yeniterzi
Naïve Bayes Assumption
Conditional Independence (Naïve Bayes) Assumption
What is the parameter size now?
again for Boolean features for 𝑥1 , 𝑥2 , … 𝑥𝑁
2𝑁
Compared to previous case (2(2𝑁 − 1)), it is much more efficient
34 Machine Learning - Reyyan Yeniterzi
Naïve Bayes Assumption
Conditional Independence (Naïve Bayes) Assumption
This assumption is not always correct but makes the problem much simpler
to solve
In practice with this assumption we can solve problems with high
accuracies
35 Machine Learning - Reyyan Yeniterzi
Naïve Bayes In a Nutshell
36 Machine Learning - Reyyan Yeniterzi
Naïve Bayes In a Nutshell
Training step:
Prediction step:
37 Machine Learning - Reyyan Yeniterzi
Naïve Bayes Example (From Mitchell)
38 Machine Learning - Reyyan Yeniterzi
Naïve Bayes Example (From Mitchell)
39 Machine Learning - Reyyan Yeniterzi
Naïve Bayes Example (From Mitchell)
40 Machine Learning - Reyyan Yeniterzi
Naïve Bayes Example: Training
41
Naïve Bayes Example: Testing
Here is a new instance:
X = (sunny, cool, high, strong)
42 Machine Learning - Reyyan Yeniterzi
Naïve Bayes Example: Testing
Outlook Play=Yes Play=No Temperature Play=Yes Play=No
Sunny 2/9 3/5 Hot 2/9 2/5
Overcast 4/9 0/5 Mild 4/9 2/5
Rain 3/9 2/5 Cool 3/9 1/5
Humidity Play=Yes Play=No Wind Play=Yes Play=No
High 3/9 4/5 Strong 3/9 3/5
Normal 6/9 1/5 Weak 6/9 2/5
P(Play=Yes) = 9/14 P(Play=No) = 5/14
43
Naïve Bayes Example: Testing
Use these probabilities to estimate the likelihood and posterior
Likelihood
44 Machine Learning - Reyyan Yeniterzi
Naïve Bayes Example: Testing
Use these probabilities to estimate the likelihood and posterior
Likelihood
Posterior
45 Machine Learning - Reyyan Yeniterzi
Naïve Bayes Example: Testing
Use these probabilities to estimate the likelihood and posterior
Likelihood
Posterior
46 Machine Learning - Reyyan Yeniterzi
Some practical details
Multiplying many probabilities. Danger of …
47 Machine Learning - Reyyan Yeniterzi
Some practical details
Multiplying many probabilities. Danger of underflow!
Solution:
Take the log
p1 * p2 = e log(p1)+log(p2)
Perform all computations by summing logs of probabilities rather than
multiplying probabilities
48 Machine Learning - Reyyan Yeniterzi
Some practical details
Multiplying many probabilities. Danger of underflow!
Solution:
Take the log
p1 * p2 = e log(p1)+log(p2)
Perform all computations by summing logs of probabilities rather than
multiplying probabilities
49 Machine Learning - Reyyan Yeniterzi
Bayesian Reasoning
Provides a probabilistic approach to inference
Combines prior knowledge (probabilities) with observed data
Therefore requires prior probabilities
Bayesian learning is effective in some particular tasks
Naïve Bayes for text classification
50 Machine Learning - Reyyan Yeniterzi
Bayes Classifier with Continuous Features
Gaussian Naïve Bayes Classifier
51
Classification: Diabetes Example
Classify whether a patient has diabetes (𝑌 = 1) or not (𝑌 = 0)
based on the test results
Given the test results (𝑋), compute the posterior probability using
Bayes rule
52 Machine Learning - Reyyan Yeniterzi
Classification: Diabetes Example
Lets start with a simple case, where we have only one test result
White blood cell count
𝑃(𝑋|𝑌)
Number of patients
White blood cell count
53 Machine Learning - Reyyan Yeniterzi
Remember the Normal Distribution
54 Machine Learning - Reyyan Yeniterzi
Gaussian Bayes Classifier
Assume the white blood cell count is normally distributed.
Notice we assume different Gaussians for each classes
These are class conditional Gaussians
55 Machine Learning - Reyyan Yeniterzi
Classification: Diabetes Example
How can I fit Gaussian distribution to my data?
56 Machine Learning - Reyyan Yeniterzi
Gaussian Bayes Classifier
Recall MLE for the Gaussians
57 Machine Learning - Reyyan Yeniterzi
Classification: Diabetes Example
Doctor has prior P 𝑌 = 0 = 0.8
A new patients comes in, the test returns result 𝑋 = 48
Does the patient have diabetes?
58 Machine Learning - Reyyan Yeniterzi
Classification: Diabetes Example
Compute 𝑃(𝑋 = 48|𝑌 = 0) and 𝑃(𝑋 = 48|𝑌 = 1) via our
estimated Gaussian distributions
We have the prior given P 𝑌 = 0 = 0.8
Compute posterior 𝑃(𝑌 = 0|𝑋 = 48) via Bayes rule using the prior
P(X|Y=0) (no diabetes)
P(X|Y=1) (diabetes)
59 Machine Learning - Reyyan Yeniterzi
Classification: Diabetes Example
Add a second test result: Plasma Gluocose value
Now our input is two dimensional
Plasma glucose value
White blood cell count
60 Machine Learning - Reyyan Yeniterzi
Multivariate Gaussian Distribution
Multiple measurements (sensors)
𝑑 inputs/features/attributes
𝑁 instances/observations/examples
61 Machine Learning - Reyyan Yeniterzi
Multivariate Gaussian Distribution
Parameters
Mean
Covariance
62 Machine Learning - Reyyan Yeniterzi
Multivariate Gaussian Distribution
63 Machine Learning - Reyyan Yeniterzi
Bivariate Normal
64 Machine Learning - Reyyan Yeniterzi
Bivariate Normal
As Σ becomes larger, the Gaussian becomes more spread-out
As Σ becomes smaller, the distribution becomes more compressed
65 Machine Learning - Reyyan Yeniterzi
Bivariate Normal
66 Machine Learning - Reyyan Yeniterzi
Bivariate Normal
The leftmost figure shows the familiar standard normal distribution
As we increase the off-diagonal entry in Σ, the density becomes more
“compressed” towards the 45◦ line
67 Machine Learning - Reyyan Yeniterzi
Bivariate Normal
68 Machine Learning - Reyyan Yeniterzi
Bivariate Normal
69 Machine Learning - Reyyan Yeniterzi
70 [Link]
Gaussian Naïve Bayes Classification
Assume conditional independence given the class labels
72 Machine Learning - Reyyan Yeniterzi
Example: Reading the Mind
Is the person reading a sentence or viewing a picture?
Reading the word “Hammer” or “Apartment”?
Viewing a vertical or horizontal line?
73 Machine Learning - Reyyan Yeniterzi
Example: Reading the Mind
Input is an image showing neural activity each of 20.000 locations
in the brain
Each value is a real value
74 Machine Learning - Reyyan Yeniterzi
Example: Reading the Mind
Input: FMRI image
Gaussian Naïve Bayes House
or
Bottle
75 Machine Learning - Reyyan Yeniterzi
Example: Reading the Mind
X is continuous
Remember the Naïve Bayes model
76 Machine Learning - Reyyan Yeniterzi
Example: Reading the Mind
X is continuous
Remember the Naïve Bayes model
Common approach: Assume 𝑃(𝑋𝑖 |𝑌 = 𝑦𝑘 ) follows a Gaussian
(Normal) distribution
77
Gaussian Naïve Bayes In a Nutshell
Training step:
Prediction step:
78
Gaussian Naïve Bayes In a Nutshell
Training step:
Prediction step:
79
Estimating Parameters
80
Example: Reading the Mind (Training)
81 Machine Learning - Reyyan Yeniterzi
Example: Reading the Mind
Y is the mental state
X is the neural activity seen from fMRI
The mean of the Gaussians defining P(Xi|Y=“bottle”)
82 Machine Learning - Reyyan Yeniterzi
Example: Reading the Mind
Results
83 Machine Learning - Reyyan Yeniterzi
Naïve Bayes
Other Issues
84
Naïve Bayes Missing Values
Suppose you don’t have a value for some feature for an example in
the training data
Applicant’scredit history unknown
Medical test has not been performed on a patient
How to calculate the likelihood? How would you handle missing
feature values?
85 Machine Learning - Reyyan Yeniterzi
Naïve Bayes Missing Values
Suppose you don’t have a value for some feature for an example in
the training data 𝑋𝑗
How to calculate the likelihood? How would you handle missing
feature values?
Easy with Naïve Bayes. Ignore the attribute
86 Machine Learning - Reyyan Yeniterzi
Naïve Bayes Missing Values
87 Machine Learning - Reyyan Yeniterzi
The Conditional Independence Assumption
Usually features are not conditionally independent
In practice it often works well.
It does not produce accurate probability estimates when its independence
assumptions are violated, it may still (and often) pick the correct maximum-
probability class in many cases [Domingos&Pazzani, 1996].
Typically handles noise well since it does not even focus on completely fitting
the training data.
88 Machine Learning - Reyyan Yeniterzi
Naïve Bayes
Incremental updates
Training is fast (linear in the number of examples, features and classes)
If the model is going to be updated very often as new data come, you
may implement it such that it allows cheap incremental updates.
For example: Store raw counts instead of probabilities
New example of class k:
For each feature update the counts based on the example
Update the class counts, update the number of training data
When need to classify compute the probabilities
89 Machine Learning - Reyyan Yeniterzi
Naïve Bayes
Advantages:
Very fast, low storage requirements
Robust to irrelevant features
Optimal if the conditional independence assumptions hold
A good dependable baseline for text classification
90 Machine Learning - Reyyan Yeniterzi
Acknowledgements
Tom Mitchell
Victor Lavrenko
Richard Zemel
Raquel Urtasun
Sanja Fidler
91 Machine Learning - Reyyan Yeniterzi
Any Questions?