Conditionalization
JAN. 22
Conditionalization
Conditionalization: For any time ti and later time tj, if proposition E in £
represents everything the agent learns between ti and tj, and cr(E) > 0, then for
any H in £,
crj(H) = cri(H|E)
cr𝑖 (𝐻 & 𝐸)
= (RATIO FORMULA)
cr𝑖 (𝐸)
cr𝑖 (𝐸|𝐻)× cri(𝐻)
= (BAYES’S THEOREM)
cr𝑖 (𝐸)
Conditionalization
1. ZEROING: Assign credence 0 to all
P E cr1 cr2
state-descriptions inconsistent with
T T 1/6 1/3
evidence learned.
T F 1/3 0
2. RESCALING: Multiply each remaining F T 1/3 2/3
non-zero credence by the same F F 1/6 0
constant so that they all sum to 1.
Conditionalization
•What if the agent receives no new evidence between t1 and t2?
•Bayesians represent an empty evidence set as a tautology.
crj(H) = cri(H|⊤) = cri(H)
Base Rate Fallacy
One in 1000 people have a particular disease. You have a test for the disease
which is 90% accurate. If you apply the test to a subject who has the disease, it
will yield a positive result 90% of the time, and if you apply the test to a subject
who lacks the disease, it will yield a negative result 90% of the time?
You randomly select a person and apply the test. The test yields a positive result.
How confident should you be that this subject actually has the disease?
Base Rate Fallacy
𝑐𝑟1(𝑃|𝐷)×𝑐𝑟1(𝐷)
•cr2(D) = cr1(D|P) =
𝑐𝑟1 𝑃𝐷 ×𝑐𝑟1 𝐷 +𝑐𝑟1 𝑃 ~𝐷 ×𝑐𝑟1(~𝐷)
0.90(0.001)
= = 0.0009/0.1008 ≈ 0.009 (0.9%)
0.90 0.001 +0.10(0.999)
Base Rate Fallacy
Bayes Factor: ratio of the likelihood of the hypothesis to the likelihood of the
catchall.
cr𝑗 (𝐻) cr𝑖 (𝐻|𝐸) cr𝑖 (𝐻) cr𝑖 (𝐸|H)
= = ×
cr𝑗 (~𝐻) cr𝑖 (~𝐻|𝐸) cr𝑖 (~𝐻) cr𝑖 (𝐸|~𝐻)
Odds for H after update = Odds for H before update multiply by the Bayes factor.
Base Rate Fallacy
Initial odds for D: cr(D) : cr(~D) = 1 : 999
cr𝑖 (𝑃|𝐷) 9/10
Bayes factor: = =9
cr𝑖 (𝑃|~𝐷) 1/10
Odds for D after update: initial odds × = 9 : 999
Bayes factor
Consequences of Conditionalization
•If an agent’s credence distribution satisfies the probability axioms and ratio
formula and updates by Conditionalization, then her resulting credence
distribution will satisfy the probability axioms as well.
•Conditionalization is cumulative:
If cr2(H) = cr1(H|E) and cr3(H) = cr2(H|E’), then cr3(H) = cr1(H|E & E’)
•Conditionalization is commutative: Conditionalizing first on E and then E’ has
the same effect as conditionalizing on E’ and then E.
Consequences of Conditionalization
•Conditionalization creates certainties:
crj(E) = cri(E|E) = 1
For any proposition P entailed by E, crj(P) = cri(P|E) = 1
•Conditionalization retains certainties:
If cri(H) = 1, then crj(H) = 1
Consequences of Conditionalization
Problem 4.2. Prove that conditionalizing retains certainties. In other words,
prove that if cri(H) = 1 and crj is generated from cri by Conditionalization, then
crj(H) = 1 as well.
Worries About Certainty
Implication of CONDITIONALIZATION: If an agent updates by conditionalizing
throughout her life, any piece of evidence which she learns at any point will
remain certain for her ever after.
Can we ever be certain in contingent propositions?
•For any contingent proposition P we may claim to be certain about, we can
imagine ways in which we turn out to be wrong about P or at least unable to
conclusively rule it out.
•If we are certain in a proposition, we should bet our life against a penny that P.
Yet there seems to be no contingent proposition on which we would take up
such a bet.
Worries About Certainty
•REGULARITY PRINCIPLE: In a rational credence distribution, no logically contingent
proposition receives unconditional credence 0.
•REGULARITY fixes an agent’s doxastic possibility set as the full set of logical
possibilities.
PRINCIPLE OF TOTAL EVIDENCE
•Probabilistic relations can be non-monotonic: H might be highly probable given
E, but improbable given E & E’.
•PRINCIPLE OF TOTAL EVIDENCE: A rational agent’s credence distribution takes into
account all the evidence she possesses.
Observation Selection Effect
•An agent may fail to take into all evidence by failing to take into account the
manner in which she comes to gain the evidence.
Observation Selection Effect
•LIKELIHOOD PRINCIPLE: If pr(E|Ha) > pr(E|Hb) then E favours Ha over Hb.
Suppose we catch 50 fish from a lake,
E = All the fish are more than 6 inches long.
H1 = All the fish in the lake are more than 6 inches long.
H2 = Only half the fish in the lake are more than 6 inches long.
pr(E|H1) > pr(E|H2)
O: The net we use has 6-inch holes.
pr(E|H1 & O) = pr(E|H2 & O) = 1
Survivorship Bias
Monty Hall Problem
Imagine you’re on a game show. There are three doors, one with a prize behind
it. You are allowed to pick one at random. Suppose you picked door A. Before
revealing what’s behind the door you picked, the host opens one of the other
doors. In your case, the host opens door B. Crucially, since the host doesn’t want
to give away the game, they will always open an empty door. They then ask if
you would like to switch to door C.
Would you choose to switch?
Monty Hall Problem
cr1 cr2
Prize behind A
1/6 1/3
cr1 cr2 & host opens B
Prize behind A 1/3 1/2 Prize behind A
1/6 0
& host opens C
Prize behind B 1/3 0
Prize behind B
Prize behind C 1/3 1/2 1/3 0
& host opens C
Prize behind C
1/3 2/3
& host opens B
Initial Priors
•What happens if we view the process of updating credences backwards, working
from an agent’s present credences back through her past credences?
•Since each conditionalization strictly added evidence, the earlier distributions
contain successively less and less contingent information as we travel back.
•An agent’s initial prior distribution (“ur-prior”) is the credence distribution she
has when she has no evidence.
Initial Priors
•Initial prior should satisfy the Kolmogorov axioms and Ratio Formula.
•Initial prior should be regular.
•A rational agent’s credence distribution at any given time is her initial prior
distribution conditional on her total evidence at that time. (CONDITIONALIZATION is
cumulative)