KASHIF JAVED
EED, UET, Lahore
1
Lecture 6
Decision Theory; Generative
and Discriminative Models
Readings: KASHIF JAVED
▪ [Link] EED, UET, Lahore
▪ Duda [Link], “Pattern Classification”, 2 Edition, Wiley, 2001
nd
2
Decision theory aka Risk Minimization
• Decision theory is based on quantifying the tradeoffs between various
classification decisions using probability and the costs that accompany
such decisions
• It assumes that the decision problem is posed in probabilistic terms, and
that all of the relevant probability values are known
KASHIF JAVED
EED, UET, Lahore
3
Decision theory aka Risk Minimization
• One aspect of probabilistic data is that sometimes a point in feature space
doesn’t have just one class
• Suppose your data is adult men and women with just one feature: height
• You want to train a classifier that takes in an adult’s height and returns a
classification, man or woman
KASHIF JAVED
EED, UET, Lahore
4
Decision theory aka Risk Minimization
• Suppose you are asked to predict the gender of a 5’5” adult (test set)
• Well, your training set includes some 5’5” women and some 5’5” men
• What should you do?
KASHIF JAVED
EED, UET, Lahore
5
Decision theory aka Risk Minimization
• In your feature space, you have
two training points at the same
location with different classes
• More generally, the height
distributions of men and women
overlap
KASHIF JAVED
EED, UET, Lahore
6
Decision theory aka Risk Minimization
• Obviously, in that case, you can’t
draw a decision boundary that
classifies all points with 100%
accuracy
KASHIF JAVED
EED, UET, Lahore
7
Decision theory aka Risk Minimization
• Multiple sample points with
different classes could lie at same
point: we want a probabilistic
classifier
KASHIF JAVED
EED, UET, Lahore
8
Decision theory aka Risk Minimization
• Problem: Predict whether a person being examined is suffering from
cancer
KASHIF JAVED
EED, UET, Lahore
9
Decision theory aka Risk Minimization
• In decision-theoretic terminology, we would say that as each person is
examined, nature is in one or the other of the two possible states: Either
the person has cancer or the person has no cancer
• We let 𝑌 denote the state of nature, with 𝑌 = 1 for cancer and 𝑌 = −1 for
no cancer
• Because the state of nature is so unpredictable, we consider 𝑌 to be a
KASHIF
variable that must be described JAVED
probabilistically
EED, UET, Lahore
10
Decision theory aka Risk Minimization
• We assume that there is some a priori probability (or simply prior) 𝑃(𝑌 = 1)
that the next person is suffering from cancer and some probability 𝑃(𝑌 =
− 1) that he is not
• These prior probabilities reflect our prior knowledge of how likely we are to
observe a person with cancer or no cancer before the person is actually
examined
KASHIF JAVED
EED, UET, Lahore
11
Decision theory aka Risk Minimization
• Suppose we were forced to make a decision about the cancer disease of a
person that will appear next without being allowed to see him
• For the moment, we assume that any incorrect classification entails the
same cost, and that the only information we are allowed to use is the
value of the prior probabilities
• With this little information, it is logical to use the following decision rule:
Decide 𝑐𝑎𝑛𝑐𝑒𝑟 𝑖𝑓 𝑃(𝑌 = 1) >KASHIF JAVED
𝑃(𝑌 = −1); otherwise decide 𝑛𝑜 𝑐𝑎𝑛𝑐𝑒𝑟
EED, UET, Lahore
12
Decision theory aka Risk Minimization
• This rule makes sense if we are to judge just one person, but if we are to
judge many persons, using this rule repeatedly may seem a bit strange
• We would always make the same decision even though we know that two
types of persons will appear
• How well it works depends upon the values of the prior probabilities
KASHIF JAVED
EED, UET, Lahore
13
Decision theory aka Risk Minimization
• If 𝑃(𝑌 = 1) is very much greater than 𝑃(𝑌 = −1), our decision in favor of
cancer will be right most of the time
• If 𝑃 𝑌 = 1 = 𝑃(𝑌 = −), only a fifty-fifty chance of being right
• In general, the probability of error is the smaller of 𝑃(𝑌 = 1) and 𝑃(𝑌 =
− 1)
KASHIF JAVED
EED, UET, Lahore
14
Decision theory aka Risk Minimization
• Mostly, we are not asked to make decisions with so little information
• We might for instance use a calorie measurement 𝑋 to improve our
classifier
• Different people have different calorie intakes, and we express this
variability in probabilistic terms
KASHIF JAVED
EED, UET, Lahore
15
Decision theory aka Risk Minimization
• Consider 𝑋 to be a discrete random variable whose distribution depends
on the state of nature (𝑌) and is expressed as 𝑃(𝑋|𝑌)
• This is the class-conditional probability density function—the probability
density function for 𝑋 given that the state of nature is 𝑌
KASHIF JAVED
EED, UET, Lahore
16
Decision theory aka Risk Minimization
• Suppose 10% of population has cancer, 90% doesn’t.
• Probability distributions for calorie intake, 𝑃(𝑋|𝑌) (class-conditional
probability density function):
Calories (𝑋) < 1200 1200-1600 > 1600
Cancer (𝑌 = 1) 20% 50% 30%
No cancer(𝑌 = −1) 1% 10% 89% P(X|Y)
KASHIF JAVED
EED, UET, Lahore
• These numbers are made up. Please don’t take them as medical advice
17
Decision theory aka Risk Minimization
• Evidence: All the people who are eating 𝑋 calories/day.
• 𝑃(𝑋) = 𝑃(𝑋|𝑌 = 1) 𝑃(𝑌 = 1) + 𝑃(𝑋|𝑌 = −1) 𝑃(𝑌 = −1)
• 𝑃(1200 ≤ 𝑋 ≤ 1600) = 0.5 × 0.1 + 0.1 × 0.9 = 0.14
• You meet guy eating 𝑋 = 1400 calories/day. Guess whether he has
cancer?
KASHIF JAVED
EED, UET, Lahore
18
Decision theory aka Risk Minimization
• If you’re in a hurry, you might see that 50% of people with cancer eat 1,400
calories, but only 10% of people with no cancer do, and conclude that
someone who eats 1,400 calories probably has cancer
• But that would be wrong, because that reasoning fails to take the
prevalence of cancer—the prior probabilities into account
KASHIF JAVED
EED, UET, Lahore
19
Bayes’ Theorem
• The (joint) probability density of finding a pattern that is in category 𝑌 = 1
and has feature value 𝑋 can be written:
𝑃 𝑋, 𝑌 = 1 = 𝑃 𝑌 = 1 𝑋 𝑃 𝑋 = 𝑃 𝑋 𝑌 = 1 𝑃(𝑌 = 1)
• Rearranging these leads us to Bayes formula:
𝑃 𝑋 𝑌 = 1 𝑃(𝑌 = 1)
𝑃 𝑌=1𝑋 =
𝑃(𝑋)
KASHIF JAVED
EED, UET, Lahore
20
Bayes’ Theorem
• 𝑃(𝑌 = 1) is called the prior 𝑃 𝑋 𝑌 = 1 𝑃(𝑌 = 1)
probability that the next person has 𝑃 𝑌=1𝑋 =
𝑃(𝑋)
cancer
𝑃 𝑌 = 1 + 𝑃 𝑌 = −1 = 1 𝑙𝑖𝑘𝑒𝑙𝑖ℎ𝑜𝑜𝑑 × 𝑝𝑟𝑖𝑜𝑟
𝑃𝑜𝑠𝑡𝑒𝑟𝑖𝑜𝑟 =
𝑒𝑣𝑖𝑑𝑒𝑛𝑐𝑒
KASHIF JAVED
EED, UET, Lahore
21
Bayes’ Theorem
• Bayes formula shows that by 𝑃 𝑋 𝑌 = 1 𝑃(𝑌 = 1)
observing the value of 𝑋 we can 𝑃 𝑌=1𝑋 =
𝑃(𝑋)
convert the prior probability to the
posteriori probability—the
probability of the state of nature
𝑙𝑖𝑘𝑒𝑙𝑖ℎ𝑜𝑜𝑑 × 𝑝𝑟𝑖𝑜𝑟
being 𝑌 = 1 given that feature 𝑃𝑜𝑠𝑡𝑒𝑟𝑖𝑜𝑟 =
value 𝑋 has been measured 𝑒𝑣𝑖𝑑𝑒𝑛𝑐𝑒
KASHIF JAVED
EED, UET, Lahore
22
Bayes’ Theorem
• The likelihood of 𝑌 = 1, w.r.t 𝑋 is 𝑃 𝑋 𝑌 = 1 𝑃(𝑌 = 1)
a term chosen to indicate that, 𝑃 𝑌=1𝑋 =
𝑃(𝑋)
other things being equal, the
category 𝑌 = 1 for which 𝑃(𝑋|𝑌 =
1) is large is more "likely" to be the
𝑙𝑖𝑘𝑒𝑙𝑖ℎ𝑜𝑜𝑑 × 𝑝𝑟𝑖𝑜𝑟
true category 𝑃𝑜𝑠𝑡𝑒𝑟𝑖𝑜𝑟 =
𝑒𝑣𝑖𝑑𝑒𝑛𝑐𝑒
KASHIF JAVED
EED, UET, Lahore
23
Bayes’ Theorem
• The formula is the product of the 𝑃 𝑋 𝑌 = 1 𝑃(𝑌 = 1)
likelihood and the prior probability 𝑃 𝑌=1𝑋 =
𝑃(𝑋)
that is most important in
determining the posterior
probability
𝑙𝑖𝑘𝑒𝑙𝑖ℎ𝑜𝑜𝑑 × 𝑝𝑟𝑖𝑜𝑟
𝑃𝑜𝑠𝑡𝑒𝑟𝑖𝑜𝑟 =
𝑒𝑣𝑖𝑑𝑒𝑛𝑐𝑒
KASHIF JAVED
EED, UET, Lahore
24
Bayes’ Theorem
• The evidence factor, 𝑃(𝑋), can be 𝑃 𝑋 𝑌 = 1 𝑃(𝑌 = 1)
viewed as merely a scale factor 𝑃 𝑌=1𝑋 =
𝑃(𝑋)
that guarantees that the posterior
probabilities sum to one as all
good probabilities must
𝑙𝑖𝑘𝑒𝑙𝑖ℎ𝑜𝑜𝑑 × 𝑝𝑟𝑖𝑜𝑟
𝑃𝑜𝑠𝑡𝑒𝑟𝑖𝑜𝑟 =
𝑒𝑣𝑖𝑑𝑒𝑛𝑐𝑒
• It is unimportant as far as making a
decision is concerned
KASHIF JAVED
EED, UET, Lahore
25
Bayes’ Theorem
• If we have an observation 𝑋 for which 𝑃 𝑌 = 1 𝑋 > 𝑃(𝑌 = −1|𝑋), we
would naturally be inclined to decide that the true state of nature is 𝑌 = 1
• We can calculate the probability of error whenever we make a decision
𝑃 𝑌=1𝑋 𝑖𝑓 𝑤𝑒 𝑑𝑒𝑐𝑖𝑑𝑒 𝑌 = −1
𝑃 𝑒𝑟𝑟𝑜𝑟 𝑋 = ቊ
𝑃 𝑌 = −1 𝑋 𝑖𝑓 𝑤𝑒 𝑑𝑒𝑐𝑖𝑑𝑒 𝑌 = 1
KASHIF JAVED
EED, UET, Lahore
26
Bayes’ Theorem
• We can minimize the probability of error by
Decide 𝑌 = 1 if 𝑃(𝑌 = 1|𝑋) > 𝑃(𝑌 = −1|𝑋); otherwise 𝑌 = −1
• This is called Bayes decision rule for minimizing the probability of error
• Under this rule: 𝑃(𝑒𝑟𝑟𝑜𝑟|𝑋) =KASHIF
min[𝑃(𝑌JAVED
= 1|𝑋), 𝑃(𝑌 = −1|𝑋)]
EED, UET, Lahore
27
Bayes’ Theorem
• By eliminating the scale factor 𝑃(𝑋), we obtain the following completely
equivalent decision rule:
Decide 𝑌 = 1 if 𝑃 𝑋 𝑌 = 1 𝑃(𝑌 = 1) > 𝑃 𝑋 𝑌 = −1 𝑃(𝑌 = −1);
otherwise 𝑌 = −1
KASHIF JAVED
EED, UET, Lahore
28
Bayes’ Theorem
• Consider a few special cases:
• If for some 𝑋 we have 𝑝(𝑋|𝑌 = 1) = 𝑝(𝑋|𝑌 = −1), then that particular
observation gives us no information about the state of nature
• In this case, the decision hinges entirely on the prior probabilities
• If 𝑃(𝑌 = 1) = 𝑃(𝑌 = −1), then the states of nature are equally probable
• In this case the decision is based entirely on the likelihoods 𝑝(𝑋|𝑌)
KASHIF JAVED
EED, UET, Lahore
29
Bayes’ Theorem
• In general, both of these factors are important in making a decision,
and the Bayes decision rule combines them to achieve the minimum
probability of error
KASHIF JAVED
EED, UET, Lahore
30
Bayes’ Theorem
𝑃 𝑋 𝑌 = 1 𝑃(𝑌 = 1) 0.05
𝑃 𝑌=1𝑋 = =
𝑃(𝑋) 0.14
𝑃 𝑋 𝑌 = −1 𝑃(𝑌 = −1) 0.09
𝑃 𝑌 = −1 𝑋 = =
𝑃(𝑋) 0.14
5
𝑃 𝐶𝑎𝑛𝑐𝑒𝑟 1200 ≤ 𝑋 ≤ 1600 𝑐𝑎𝑙𝑠 = ≈ 36%
14
• So, we probably shouldn’t diagnose
KASHIFcancer
JAVED
EED, UET, Lahore
31
Decision theory aka Risk Minimization
• BUT . . . we’re assuming that we want to maximize the chance of a correct
prediction — not always the right assumption
• For many applications, our objective will be more complex than simply
minimizing the number of misclassifications
KASHIF JAVED
EED, UET, Lahore
32
Decision theory aka Risk Minimization
• If you’re developing a cheap screening test for cancer, it would be better to
make fewer false negatives, even if this was at the expense of making
more false positives
• A false negative might mean somebody misses an early diagnosis and
dies of a cancer that could have been treated if caught early
• A false positive just means that you spend more money on more accurate
tests KASHIF JAVED
EED, UET, Lahore
33
Decision theory aka Risk Minimization
• Loss or cost function:
▪ a single, overall measure of loss incurred in taking any of the available
decisions or actions
▪ allows us treat situations in which some kinds of classification mistakes are
more costly than others
• Our goal is to minimize the total loss incurred
KASHIF JAVED
EED, UET, Lahore
34
Decision theory aka Risk Minimization
• A loss function 𝐿(𝑧, 𝑦) specifies badness if classifier predicts 𝑧, true class is
𝑦.
1 if 𝑧 = 1, 𝑦 = −1 false positive is bad
e.g., 𝐿 𝑧, 𝑦 = ൞5 if 𝑧 = −1, 𝑦 = 1 false negative is BAAAAAD
0 if 𝑧 = 𝑦
• Using this loss function, if we guess that there is no cancer, and if we are
wrong, we will be wrong by 𝑃 𝑌 = 1 𝑋 = 36% and incur a loss of 5
KASHIF JAVED
EED, UET, Lahore
35
Decision theory aka Risk Minimization
• A loss function 𝐿(𝑧, 𝑦) specifies badness if classifier predicts 𝑧, true class is
𝑦.
1 if 𝑧 = 1, 𝑦 = −1 false positive is bad
e.g., 𝐿 𝑧, 𝑦 = ൞5 if 𝑧 = −1, 𝑦 = 1 false negative is BAAAAAD
0 if 𝑧 = 𝑦
• A 36% probability of loss 5 is worse than a 64% prob. of loss 1, so, we
recommend further cancer screening.
KASHIF JAVED
EED, UET, Lahore
36
Decision theory aka Risk Minimization
• Defs: The loss function above is asymmetrical.
• The 0-1 loss function is 1 for incorrect predictions
0 for correct
• 0-1 penalizes all wrong answers the same
• Another example where you want a very asymmetrical loss function is for
spam detection — putting a good email in the spam folder is much worse
than putting spam in your inbox
KASHIF JAVED
EED, UET, Lahore
37
Decision theory aka Risk Minimization
• Let 𝑟 ∶ ℝ𝑑 → ±1 be a decision rule, aka classifier:
• a function that takes input 𝑥 and outputs 1 (“in class”) or -1 (“not in class”)
• The risk for r is the expected loss over all values of 𝑥, 𝑦:
𝑅 𝑟 = 𝐸[ 𝐿(𝑟 𝑋 , 𝑌)]
The lossJAVED
KASHIF function is the cost you
pay UET,
EED, make decision 𝑟(𝑋),
if you Lahore
but the true state is 𝑌
38
Decision theory aka Risk Minimization
𝑅 𝑟 = 𝐸[ 𝐿(𝑟 𝑋 , 𝑌)]
= σ𝑥 𝐿 𝑟 𝑥 , 1 𝑃 𝑌 = 1 𝑋 = 𝑥 + 𝐿 𝑟 𝑥 , −1 𝑃 𝑌 = −1 𝑋 = 𝑥 𝑃(𝑋 = 𝑥)
= 𝑃 𝑌 = 1 𝐿 𝑟 𝑥 ,1 𝑃 𝑋 = 𝑥 𝑌 = 1
𝑥
+ 𝑃 𝑌 = −1 𝐿 𝑟 𝑥 , −1 𝑃 𝑋 = 𝑥 𝑌 = −1
𝑥
• Our problem is to find a decision rule against 𝑃(𝑌) that minimizes the overall
risk or expected loss
KASHIF JAVED
EED, UET, Lahore
39
Decision theory aka Risk Minimization
• The Bayes decision rule aka Bayes classifier is the fn 𝑟 ∗ that minimizes
functional 𝑅(𝑟). Assuming 𝐿(𝑧, 𝑦) = 0 for 𝑧 = 𝑦:
1 if 𝐿 −1, 1 𝑃 𝑌 = 1 𝑋 = 𝑥 > 𝐿 1, −1 𝑃 𝑌 = −1 𝑋 = 𝑥
• 𝑟∗ 𝑥 = ቊ
−1 otherwise
• The 𝑟 ∗ is called the Bayes risk.
• When 𝐿 is symmetric, the big, KASHIF
key principle
JAVED you should memorize is pick
EED, UET,
the class with the biggest posterior Lahore
probability
40
Decision theory aka Risk Minimization
• The Bayes decision rule aka Bayes classifier is the fn 𝑟 ∗ that minimizes
functional 𝑅(𝑟). Assuming 𝐿(𝑧, 𝑦) = 0 for 𝑧 = 𝑦:
1 if 𝐿 −1, 1 𝑃 𝑌 = 1 𝑋 = 𝑥 > 𝐿 1, −1 𝑃 𝑌 = −1 𝑋 = 𝑥
• 𝑟∗ 𝑥 = ቊ
−1 otherwise
• If the loss function is asymmetric, then you must weight the posteriors with
the losses
KASHIF JAVED
EED, UET, Lahore
41
Decision theory aka Risk Minimization
• In cancer example:
• 𝑟∗ 𝑥 = 1 for 𝑥 ≤ 1600
• 𝑟 ∗ 𝑥 = −1 for 𝑥 > 1600
• The Bayes risk, aka optimal risk, is the risk of the Bayes classifier
• In our cancer example, the last expression for risk 𝑅 gives:
𝑅(𝑟 ∗ ) = 0.1(5 × 0.3) + 0.9(1 × 0.01 + 1 × 0.1) = 0.249
• No decision rule gives a lowerKASHIF
risk JAVED
EED, UET, Lahore
42
Decision theory aka Risk Minimization
• It is interesting that, if we really know all these probabilities, we really can
construct an ideal probabilistic classifier
• But in real applications, we rarely know these probabilities; the best we
can do is use statistical methods to estimate them
• Deriving/using 𝑟 ∗ is called risk minimization
KASHIF JAVED
EED, UET, Lahore
43
Continuous Distributions
• Suppose 𝑋 has a continuous probability
density function (PDF).
• prob. that random variable:
𝑥2
𝑋 ∈ 𝑥1 , 𝑥2 = න 𝑓 𝑥 𝑑𝑥
𝑥1
• shaded area
KASHIF JAVED
EED, UET, Lahore
44
Continuous Distributions
• Area under whole curve
∞
= 1 = න 𝑓 𝑥 𝑑𝑥
−∞
• Expected value of 𝑔(𝑋): 𝐸[𝑔(𝑋)]
∞
= න 𝑔 𝑥 𝑓 𝑥 𝑑𝑥
−∞
KASHIF JAVED
EED, UET, Lahore
45
Continuous Distributions
∞
• Mean 𝜇 = 𝐸 𝑋 = −∞ 𝑥 𝑓 𝑥 𝑑𝑥
• Variance 𝜎 2 = 𝐸 𝑋 − 𝜇 2
= E 𝑋 2 − 𝜇2
KASHIF JAVED
EED, UET, Lahore
46
Decision theory aka Risk Minimization
• Perhaps our cancer statistics look like this.
KASHIF JAVED
EED, UET, Lahore
47
Decision theory aka Risk Minimization
• Let’s go back to the 0-1 loss function for a moment
• In other words, suppose you want a classifier that maximizes the chance
of a correct prediction
• The wrong answer would be to look where these two curves cross and
make that be the decision boundary
• As before, it’s wrong because it doesn’t consider the prior probabilities.
KASHIF JAVED
EED, UET, Lahore
48
Decision theory aka Risk Minimization
• Suppose 𝑃(𝑌 = 1) = 1/3, 𝑃(𝑌 = −1) = 2/3, 0-1 loss:
• To maximize the chance, you’ll predict correctly whether somebody has
cancer, the Bayes decision rule looks up 𝑥 on this chart and picks the
curve with the highest probability
• In this example, that means you pick cancer when x is left of the optimal
decision boundary, and no cancer when x is to the right.
KASHIF JAVED
EED, UET, Lahore
49
Decision theory aka Risk Minimization
• Suppose 𝑃(𝑌 = 1) = 1/3, 𝑃(𝑌 = −1) = 2/3, 0-1 loss:
• To maximize the chance, you’ll predict correctly whether somebody has
cancer, the Bayes decision rule looks up 𝑥 on this chart and picks the
curve with the highest probability.
• In this example, that means you pick cancer when x is left of the optimal
decision boundary, and no cancer when x is to the right.
Effect of different priors is KASHIF JAVED
that the threshold of EED, UET, Lahore
decision moves toward the
mean of the less likely class 50
Decision theory aka Risk Minimization
• Define risk as before, replacing summations with integrals.
• 𝑅 𝑟 = 𝐸[ 𝐿(𝑟 𝑋 , 𝑌)]
= 𝑃 𝑌 = 1 න 𝐿 𝑟 𝑥 , 1 𝑓 𝑋 = 𝑥 𝑌 = 1 𝑑𝑥
+ 𝑃 𝑌 = −1 න 𝐿 𝑟 𝑥 , −1 𝑓 𝑋 = 𝑥 𝑌 = −1 𝑑𝑥
KASHIF JAVED
EED, UET, Lahore
51
Decision theory aka Risk Minimization
• For Bayes decision rule, Bayes risk is the area under minimum of functions
• Assuming 𝐿(𝑧, 𝑦) = 0 𝑓𝑜𝑟 𝑧 = 𝑦:
• 𝑅 𝑟 ∗ = =𝑦𝑛𝑖𝑚 ±1 𝐿 −𝑦, 𝑦 𝑓 𝑋 = 𝑥 𝑌 = 𝑦 𝑃 𝑌 = 𝑦 dx
KASHIF JAVED
EED, UET, Lahore
52
Wrong decisions, so loss non-zero
Decision theory aka Risk Minimization
• Two different views of the same 2D Gaussians. Note the Bayes optimal
decision boundary, which is white at right
𝑓(𝑋|𝑌)
KASHIF JAVED
EED, UET, Lahore
𝑥2
𝑥1 53
3 ways to Build Classifiers
(1) Generative models (e.g., LDA)
▪ Assume sample points come from probability distributions, different for each
class
▪ Guess form of distributions
▪ For each class 𝐶, fit distribution parameters to class 𝐶 points, giving 𝑓(𝑋|𝑌 =
𝐶)
▪ For each 𝐶, estimate 𝑃(𝑌 = 𝐶)
▪ Bayes’ Theorem gives 𝑃(𝑌|𝑋)
▪ If 0-1 loss, pick class 𝐶 that maximizes 𝑃(𝑌 = 𝐶|𝑋 = 𝑥) equivalently,
KASHIF JAVED
maximizes 𝑓(𝑋 = 𝑥|𝑌 = 𝐶) 𝑃(𝑌 = 𝐶)
EED, UET, Lahore
54
3 ways to Build Classifiers
(2) Discriminative models (e.g., logistic regression)
▪ Model 𝑃(𝑌|𝑋) directly
(3) Find decision boundary (e.g., SVM)
▪ Model 𝑟(𝑥) directly (no posterior)
KASHIF JAVED
EED, UET, Lahore
55
3 ways to Build Classifiers
• Advantage of (1 & 2): 𝑃(𝑌|𝑋) tells you probability your guess is wrong
[This is something SVMs don’t do.]
• Advantage of (1): you can diagnose outliers: 𝑓(𝑋) is very small
• Disadvantages of (1): often hard to estimate distributions accurately;
real distributions rarely match standard ones.
KASHIF JAVED
EED, UET, Lahore
56
Decision theory aka Risk Minimization
• A generative model is a full probabilistic model of all variables, whereas a
discriminative model provides a model only for the target variables that we
want to predict
• It’s important to remember that we rarely know precisely the value of any
of these probabilities.
• There is usually error in all of these probabilities.
KASHIF JAVED
EED, UET, Lahore
57
Decision theory aka Risk Minimization
• In practice, generative models are most popular when you have
phenomena that are well approximated by the normal distribution, and you
have enough sample points that you can approximate the shape of the
distribution well
KASHIF JAVED
EED, UET, Lahore
58