0% found this document useful (0 votes)
38 views88 pages

Understanding Causal Relationships

Uploaded by

mahnaf837
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
38 views88 pages

Understanding Causal Relationships

Uploaded by

mahnaf837
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Chapter 7

Key terms
Causal argument: the attempt to establish a causal connection between two
factors (i.e., anything that can stand in a causal relation such as events,
situations, or features of objects).

Causal mechanism: the specific way in which one event causes another. For
example, the causal mechanism by which smoking causes cancer involves the
formation of DNA adducts by the carcinogens from cigarette smoke that are
taken into the body.

Clustering illusion: a form of pattern-seeking in which people tend to think


that random distributions over an area are clustering too much to be random.

Common cause: Two events, A and B, are correlated due to common


cause when some third event C is responsible for both of them, and that's why
they occur together at a higher rate than alone.

Correlation: If A occurs at a higher rate when B occurs than it does otherwise,


we say A and B have a positive binary correlation. If A occurs to a greater
degree when B occurs to a greater degree, we have a positive scalar
correlation. Negative correlations are statistical relationships in the opposite
direction: A occurs at a lower rate when B occurs than it does otherwise, or A
occurs to a lesser degree as B occurs to a greater degree. If the term
“correlated” is used without specifying positive or negative, we assume that
the term refers to a positive correlation.

Double-blind study: an experiment is double-blind when neither the subject


nor the experimenter is aware of which subjects belong to the control arm
and which belong to the experimental arm of the trial. This experimental
design helps to rule out experimenter effects as a possible explanation for
observed differences in outcomes between the control group and the group
receiving the intervention.

Immediate vs. distal causes: A distal cause of x is one that


is effective through intermediate causes. A proximate cause is of x is one that
is immediately responsible for the event.

Mere chance (as an explanation for a correlation): when there is a genuine


correlation between two factors but there is no causal connection between
them. We often identify correlations due to mere chance by assessing the
plausibility of the causal mechanism required for a causal connection
between the relevant factors.

Pattern-seeking: the tendency to be over-sensitive to patterns even in scarce


data that could be entirely random.

Placebo-controlled: when an experiment is placebo-controlled, the control


group (the participants in the experiment who are not given the intervention
being tested) receives a placebo treatment. This experimental design helps to
rule out the placebo effect as a possible explanation for observed differences
in outcomes between the group receiving the intervention and the control
group.

Placebo effect: a positive effect arising from the expectation that an


intervention (usually medical or dietary) will be effective. This effect works
entirely through a subject's psychology. For example, when subjects take pills
that they believe are effective for pain or depression, some report that the
pills are effective even when they are not in fact biologically active.

Post hoc ergo propter hoc: Latin for “after this therefore because of this”. It's
the name given to the fallacy of assuming that because event B happens after
event A, it must have been caused by A.
Randomized controlled trial: in this kind of experiment, subjects are
randomly divided into two groups, and some intervention (e.g., a drug) is
applied to members of one group only. This procedure helps to rule out other
factors (aside from the intervention being tested) that might explain
differences in the observed outcomes between the two groups.

Regression to the mean: the tendency, when selecting a data point that lies
outside the mean, for adjacent data points to lie closer to the mean. This
tendency can result in misleading correlations.

Reverse causation: when we propose that A causes B in order to explain a


correlation between them, but in fact the correlation is explained by the fact
that B causes A.

Robust evidence: evidence that stems from a wide range of experiments (i.e.,
from different sources and from different kinds of experiments). This helps to
ensure that we are drawing on lots of data, and that the result does not stem
from some flaw in the study's design, or error on the part of the
experimenters.

Side effect (as an explanation for a correlation): when two factors, A and B,
are correlated due to some additional consequence of the presence of one of
the factors, which is not the alleged causal mechanism. For example, a drug
may be correlated with a reported reduction in pain, even if a fake pill with no
active ingredients would be just as effective. This doesn't mean that the pill
has no effect, but that its effect is not due to the drug it contains, but due
instead to a side effect of taking the drug: namely, the expectation that the pill
will be effective.

Statistical significance: a correlation in our sample is said to have statistical


significance if it is sufficiently unlikely to have occurred by chance given that
there's no correlation in the larger population. The threshold for statistical
significance is often set by convention in each field with respect to a p-value.
For example, the social sciences tend to use a threshold of a p-value of .05;
this tells us that there is only a 5% chance that we would see a correlation of
this big in our population without there being a correlation in the larger
population.

7.1 Causal thinking


Consider the last time you saw a fight in a movie. One actor's fist launches
forward; the other actor staggers back. You didn't make a conscious inference
that one event caused the other—that judgment feels almost as automatic as
perception itself. You just seem to see the causal connection.

In reality, of course, you know that punches in movies are actually just
choreographed with little or no actual contact. But knowing this doesn't stop
System 1 from making automatic causal inferences. In other words, it's a kind
of cognitive illusion: System 1 will continue to infer causation even when
System 2 knows it isn't really there.

Now imagine a red dot on your computer screen. A green dot moves towards
it, and as soon as they touch, the red dot moves away at the same speed and
in the same direction. If the timing is right, you can't avoid feeling like
you saw a causal connection, a transfer of force. But if, instead, they don't
touch and the red dot only moves after pausing for a moment, it feels like the
red dot moved itself. Consciously, we know the dots are just pixels on a screen
that don't transfer force at all. But we can't shake the sense that we're seeing
a direct causal connection in one case and not the other.

An instinct for causal stories


Our perceptual judgments are just one example of a more general fact about
our minds: we can't help thinking about the world in terms of causal stories.
This tendency is critical to one of our superpowers in the animal kingdom: our
ability to understand our environment and shape it to our liking.
However, the cognitive processes we use to arrive at these causal stories were
honed during much simpler times. Our ancient ancestors noticed countless
simple patterns in their environment: when they ate certain plants, they got
sick; when they struck some flint, there were sparks; when the sun went
down, it got darker and colder. And these things still happen in our world. But
our natural bias towards simple causal stories can lead us to oversimplify
highly complex things like diseases, economies, and political systems.

In addition, our minds are so prone to perceiving patterns that we often seem
to find them even in random, patternless settings—for example, we see faces
in the clouds and animal shapes in the stars. In the environments of our
ancestors, there was a great advantage to finding genuine patterns, and little
downside to over-detecting them. (Is that vague shape in the shadows a
predator, or nothing at all? Better to err on the side of caution!)

The result is a mind that's perhaps a bit too eager to find patterns. For
example, if you consider the sequence "2...4...," the next number probably
just pops into your head. But hang on—what is the next number? Some
people think of 6, while others think of 8 or even 16. (Do we add 2 to the
previous number, double it, or square it?) As soon as System 2 kicks in, we
realize the answer could be any of these. For a moment, though, the answer
may seem obvious. System 1 completes the pattern with the simple
earnestness of a retriever bringing back a stick.

Even when our observations have no pattern at all, we can't help but suspect
some other factor at work. For example, suppose we map recent crimes
across a city, yielding something like the picture on the left. Most of us would
find it suspicious that there are so many clusters: we'd want to know what's
causing the incidents to collect in some places and not others. But in fact this
picture shows a completely random distribution. Real randomness generates
clusters, even though we expect each data point to have its own personal
space. This is called the clustering illusion, and it can send us seeking causal
explanations where none exist.

One thing after another


The most rudimentary error we make about causation is to infer that B was
caused by A just because B happened after A. This is known as the fallacy
of post hoc ergo propter hoc—“after this, therefore because of this." To help
you avoid this fallacy, just remember that everyone who commits the fallacy
of post hoc ergo propter hoc eventually ends up dead!

At some level, we know it's absurd to assume that one event caused another
just because it happened first; but it's a remarkably easy mistake to make on
the fly. Often the problem is one of communication. A report that two events
occurred in sequence is often taken to convey that a causal relationship links
them. For example, suppose I say "A fish jumped, and the water rippled." It's
fairly clear that I'm suggesting the fish caused the ripples. Now suppose I say,
"After meeting my new boyfriend, my grandma had a heart attack." The
speaker might just be reporting a sequence of events, but it's still natural for
an audience to seek a potential causal connection.

As always, this lack of clarity in our language can be exploited. By reporting a


sequence of events, you can insinuate a causal relationship without explicitly
stating it. This is why politicians like to mention positive things that have
happened since they took office, and negative things that have happened
since their opponents took office. The message gets conveyed, even if they
don't state the causal connections explicitly. They can leave that part up to
the automatic associations of their listeners. And if challenged for evidence,
they need only defend the claim they were making explicitly: "All I was saying
is that one thing happened before the other! Draw your own conclusions!"

Complex causes
The causal stories that come naturally to us are often very simple. For
example, "The cause of the fire was a match." But in our complex world, many
of the things we want to understand don't arise from a single cause. There
may be no answer to the question, "What was the cause?"—not because there
was no cause, but because there were too many interconnected causes—each
of which played a part.

In fact, it's rare for anything to have a single cause. When we talk about the
cause of an event, we usually mean the one factor that is somehow most out
of the ordinary. For example, if a fire breaks out, in most ordinary
contexts, the cause is a source of ignition like a match. But there are other
contexts where the presence of fuel or even oxygen might be the most out-of-
the-ordinary factor. For example, imagine an experiment in a vacuum under
extremely high temperatures. If the researchers were depending on a lack of
oxygen to keep things from burning up, then the unexpected presence of
oxygen would count as the cause.

We can also distinguish between the immediate causes of events and


the distal causes that explain the immediate causes. For example, a certain
drought might have a clear immediate cause, such as a long-term lack of rain.
But it would be useful to know what caused that lack of rain: perhaps a
combination of changes in regional temperatures and wind patterns. And
these factors in turn may be part of a global change in climate that is largely
due to rising levels of greenhouse gases. Each causal factor is a node in a
network that itself has causes, and combines with other factors to bring about
its effects. The real causal story is rarely ever a simple one.

7.2 Causes and correlations


Most errors in causal reasoning are subtler than simply assuming that two
events that happen in sequence must be causally related. Instead, we start by
sensing a pattern over time: maybe we've seen several events of one kind that
follow events of another kind. Or we notice that a feature which comes in
degrees tends to increase when another feature is present. In other words, we
notice a correlation between two kinds of events, or two kinds of features.
From this repeated pattern, we then infer a causal relationship.

In fact, this inference is so natural that the language we use to report a


correlation often gets straightforwardly interpreted as reporting causation.
When people hear that two things are "associated" or "linked" or "related,"
they often misinterpret that as a claim about a causal connection. But in the
sciences, these expressions are typically used to indicate a correlation that
may or may not be causal.

So how can we tell if two apparently correlated factors are causally related?
(Factors include anything that can stand in causal relationships—events,
situations, or features of objects.) Correlations can certainly
provide evidence of causation, but we need to be very careful when evaluating
that evidence. Because there are many ways that we can go wrong in making
this kind of inference, it is important to isolate three inferential steps that
must be made:

[Link] observed a correlation between A and B;


[Link] is a general correlation between A and B (inferred from 1); and
3.A causes B (inferred from 2).

Because we could go wrong at any step, we should become less confident


with each interim conclusion. This means that a good argument of this sort
requires very strong evidence at each step. (We'll look at the relevant rule for
probability in Chapter 8, but just to get a sense of how inferential weakness
can compound, suppose you're 80% confident that the first statement is true,
and 80% confident that the second is true given that the first is true. In that
case, you should only be 64% confident that they are both true. And if you're
80% confident that the third is true given that the first two are true, you
should only be about 51% confident that all three statements are true [1].)

Correlation doesn't imply causation, but it does waggle its


eyebrows suggestively and gesture furtively while mouthing
'look over there.'
—Randall Munroe

So a good causal argument from correlation requires that we establish three


things with a high degree of confidence: (1) that a correlation exists in the
cases we've observed; (2) that this means there is a general correlation that
holds beyond the cases we've observed; and (3) that this general correlation is
not misleading: it really results from A causing B.

We'll go through these steps one at a time, but first it's worth getting clear on
exactly what correlations are.

The nature of correlation


So what is a correlation, exactly? Here's the definition. There is a positive
binary correlation between factors A and B when, on average:

 A occurs at a higher rate when B occurs than it does otherwise.

And if we replace "higher" in this definition with "lower," then there is


a negative (or inverse) binary correlation between the two factors. (If we don't
specify and just say that two things are "correlated," we mean they are
positively correlated.) Three crucial things are worth clarifying.

1. The definition has to do with rates or proportions. So, for example, take two
very common traits in the US, such as owning a phone and loving pizza.
Assume that most phone owners love pizza, and vice versa. If a correlation
between A and B just meant that A occurs more often with B than it does
otherwise, we could conclude that these two traits are correlated: there are
more phone owners who love pizza than phone-owners who don't love pizza.
But that's only because so many people love pizza!

But—and this is the most important bit—knowing the statistical facts above,
you should be pretty confident that a random American loves pizza. But if you
learned that they own a phone, that wouldn't give you any more information
about whether they love pizza. (Or at least, not unless you know something I
don't!) That's because, as far as we know, the proportion of phone owners
who love pizza is no different from the proportion of people in general who
love pizza. So learning that they own a phone isn't any evidence that they love
pizza: it could be that they are just as likely to own a phone whether or not
they love pizza.

To know that there was a correlation, we'd need to know that the rate of
phone-ownership is higher among pizza lovers than it is in general; or that the
rate of pizza-loving is higher among phone owners than it is in general. So, a
useful rule of thumb that can help identify correlations is to ask yourself
whether, if you knew the statistical facts and nothing else, learning that factor
A is present gives you evidence that factor B is also present (or vice versa). In
this case, the two bullet points above don't give us any reason to think
someone is more likely to own a cell phone after learning that they are male.

2. Correlation is symmetrical—if it holds in one direction, it also holds in the


other. In other words, if A occurs at a higher rate when B occurs than it does
otherwise, then B occurs at a higher rate when A occurs than it does
otherwise. This may not seem obvious at first, but it's true.

Suppose 10% of people love anchovies and 40% of people love wine. And
suppose that there is also a correlation between loving anchovies and loving
wine. This means that there's a higher proportion of anchovy-lovers among
wine-lovers than there is overall. So, more than 10% of wine-lovers
are also anchovy-lovers, which means that more than 4% of people overall
love both things.

In that case, what percentage of anchovy-lovers are also wine-lovers? Well, we


know that more than 4% of people overall love both, and that's already more
than 40% of the anchovy lovers, who are only 10% of all the people. So there's
a higher proportion of wine-lovers among anchovy lovers than there is overall,
which means there is also a correlation between loving wine and loving
anchovies.

In the image above, only 4.5 percent of people love both, but that's enough
for a correlation. Try other percentages to see for yourself that it's impossible
to make wine-loving correlated with anchovy-loving but not vice-versa.

3. Finally, in the definition above, we're assuming the correlation we're talking
about has to do with factors we're treating as all-or-nothing rather than as
coming in degrees. So the rate of a factor simply has to do with how often it's
present and absent. For example, in the example above, we are treating wine-
loving and anchovy-loving as simple yes/no factors, and not specifying exactly
how much people love those things. This means we need a kind of arbitrary
cut-off point for how much love is required to count as loving wine. A
correlation between two all-or-nothing factors is called a binary correlation.

But it might be even more interesting to ask whether there is


a scalar correlation rather than just a binary correlation—that is, is loving wine
more correlated with loving anchovies more, all the way up to really intense
love? We say there is a positive scalar correlation between factors A and B
when:

 on average, A occurs to a greater degree when B occurs to a greater


degree.
(If we replace only one instance of the word "greater" with "lesser," then one
factor increases as the other decreases, giving us a negative (or inverse) scalar
correlation.)

It often happens that two factors can be understood either as scalar or as


binary, as with wine-loving, so we can ask about both kinds of correlation. For
example, the height and diameter of trees both come in degrees. And at least
on average, the greater a tree's diameter, the greater its height (and vice
versa). In this sense, the two features are correlated. But it would make no
sense to say that height occurs at a higher rate with diameter, because all
trees have both height and diameter. Unlike a binary correlation, which
relates all-or-nothing factors, this is correlation is scalar.

Unfortunately, the distinction between binary and scalar correlations is a bit


trickier than it seems at first, because the same underlying correlation can
often be measured in either a scalar or binary way. For example, although
height and diameter are scalar, we could make our calculations simpler by
selecting an arbitrary cutoff for what it takes to count as a tall tree and what it
takes to count as a wide tree. That would give us two binary features, and we
can then report a binary correlation: tallness occurs at a higher rate among
wide trees than among non-wide trees. This isn't the best way to measure the
correlation (it would be more precise to treat it as scalar), but even scientists
take shortcuts.

Illusory correlations
We turn now to the first of three ways in which we can wrongly conclude from
an apparent correlation that a causal relationship exists between two factors.
Recall the three inferential steps from above:

[Link] observed a correlation between A and B;


[Link] is a general correlation between A and B; and
3.A causes B.
The first kind of error is that we're wrong about (1): it only seems to us like A
and B correlate in our observed sample.

But why might we get the false impression that A and B correlate in our
sample? Taking just the case of binary correlation, we may be overestimating
the rate at which A occurs along with B in our sample, or underestimating the
rate at which it occurs without B in our sample—or both. This might happen
for various reasons, such as motivated reasoning, selective recall, and
selective noticing.

For example, recall the idea that people behave strangely more often when
the moon is full. That's a claim about the correlation between two factors. I
might think that I have good evidence for that claim because I think a
correlation exists between full moons and strange behavior in my experience,
and then I can generalize from my experience.

The problem is that if I'm subject to selective noticing, I might be wrong about
the correlations in my own observations. Maybe the only time I think about
the moon hypothesis is when I happen to notice that someone is behaving
strangely and there's a full moon. I simply don't notice the times when there's
a full moon and no one is behaving strangely, or the times when there's
strange behavior and no full moon. As a result, it can feel like the rate of
strange behavior is higher in my experience during a full moon, even though
it's not. If I were to carefully tally up the occurrences of strange behavior that I
observe, along with phases of the moon, I'd see that no correlation actually
exists.

Another kind of mistake is simply that we fail to think proportionally. For


example, suppose we've only observed Bob when it's cold and we notice that
he has worn a hat 70% of the time. Can we conclude that there's a correlation
in our observations between his wearing a hat and cold temperatures? Of
course not! What if he wears a hat 70% of the time regardless of the
temperature? In that case, there's no special correlation between his hat
wearing and the cold: he just loves wearing hats.

If we are told, "Most of the time when it's cold, Bob wears a hat," it's easy to
forget that this is not enough to establish a correlation. To infer a correlation,
we have to assume that Bob doesn't also wear a hat most of the time even
when it's not cold. Maybe this is a safe assumption to make, but maybe not.
The point is that if we just ignore it, we are neglecting the base rate, a mistake
we encountered in the previous chapter.

Again, it can help to think in terms of evidence. To establish a correlation


between A and B, we must not only check how common B is given that A
occurs, but also how likely not-B is given that A occurs. This should remind
you of the strength test for evidence. If Bob's wearing a hat is correlated with
the cold, then his wearing a hat should be more likely given that it's cold than
given that it's not cold. This in turn means that seeing Bob wear a hat
is evidence that it's cold out: you should become more confident that it's cold
out when you see Bob wearing a hat. (And, because correlation is
symmetrical, knowing that it's cold out is evidence that Bob is wearing a hat.)

Consider a final example—this time, one of selective recall. As we saw in a


previous chapter, when asked whether Italians tend to be friendly, we search
our memory harder for examples of friendly Italians than for examples of
unfriendly Italians. So even if the question is just asking about
whether most Italians are friendly, we do a pretty bad job: we'd need to
consider both kinds of examples. But things are worse than that, because the
loose generalization "Italians tend to be friendly" is usually a claim
about correlation. And establishing a correlation requires considering four
kinds of cases: friendly Italians, unfriendly Italians, friendly non-Italians, and
unfriendly non-Italians. All four of those combinations are relevant if we are
comparing the rate of friendliness among Italians with the rate of friendliness
among non-Italians. So if the question is about a correlation, our selective
search for cases in which Italians are friendly is absurdly inadequate.

Generalizing correlations
Suppose we've avoided these errors and correctly identified a correlation in
our experience. The next point at which our causal inference can flounder is
when we generalize from our sample to conclude that a correlation exists in
the general population. (After all, our observations usually only constitute a
small sample of the relevant cases.) In the previous chapter, we saw several
reasons why our sample might fail to match the wider set of cases. All the
same lessons apply when we're generalizing about correlations from a sample
—for example, we need to be aware of sampling biases, participation biases,
response biases, and so on.

However, there is one important difference when we're dealing with


correlations. When estimating the proportion of individuals with some
feature, we said that a "sufficiently large" sample is one that gives us a
sufficiently narrow confidence interval. But when we are interested in a
correlation between two features, we want a sample large to make our
correlation statistically significant.

For example, if we're estimating the rate of respiratory problems in a country,


we need a sample large enough to give us a narrow confidence interval; but if
we want to know whether a correlation exists between respiratory problems
and air pollution, we need a sample large enough that we have a good chance
of finding a correlation that is statistically significant.

So what does this mean, exactly? A correlation found in a sample


is statistically significant when we'd be sufficiently unlikely to find a
correlation at least this large in a sample of this size without there
being some correlation in the larger population. We can work this out by
supposing that there is no correlation in the larger population and then
simulate taking many random samples of this size, and working out what
proportion of those samples would show a correlation of at least the size that
we observe, merely by chance.

As with confidence intervals, the threshold for a statistically significant


correlation is somewhat arbitrary. By convention, sufficiently unlikely in the
social sciences means there's less than a 5% chance of seeing a correlation of
this size or larger in our sample without there being some correlation in the
larger population. This corresponds to a p-value of .05. (In areas like physics,
however, the threshold is often more stringent.)

So how strong is the evidence from a study that finds a statistically significant
correlation? Note that if H = there is a correlation in the population and the
evidence = there is a correlation of at least this size in the sample, then
statistical significance ensures a low value for the probability of the evidence
given not-H—namely .05. Usually we can also assume that we'd be much
more likely to see this correlation if there really is a correlation in the
population as a whole, meaning that the probability of the evidence given H is
fairly high in comparison. In that case, statistical significance translates into a
fairly high strength factor for the evidence provided by our sample. But note
that the strength factor is not exactly overwhelming. A sample correlation that
is just barely statistically significant will have at best a strength factor of 20 in
favor of the generalization.

To make this point more vivid, imagine we find a barely statistically significant
correlation in our sample. As we've seen, this means roughly a 5% chance of
seeing a correlation like this in our sample even if there's no correlation in the
population as a whole. So if twenty studies like ours were conducted, we
should expect one to find a statistically significant correlation even if there's
absolutely no correlation in the population!
If we also take into account the file drawer effect and bias for surprising
findings in scientific journals, we should be even more careful. When we see a
published study with a surprising result and a p-value just under .05, we
should keep in mind that may have been conducted that found no exciting or
significant results and went unpublished. This means that the evidence
provided for a surprising correlation by a single study with that level of
significance may be far from conclusive.

This selection effect only compounds for science reporting in the popular
media. Studies with surprising or frightening results are far more likely to
make their way into the popular media than those with boring results. In
addition, such studies are often reported in highly misleading ways—for
example, by interpreting correlations as though they established causation.
For these reasons, if we're not experts in the relevant field, we should be very
careful when forming opinions from studies we encounter in the popular
media. It can help to track down the original study, which is likely to contain a
much more careful interpretation of the data, often noting weaknesses in the
study itself, and rarely jumping to causal conclusions. But even this will not
erase the selection effect inherent in the fact that we're only looking
at this study—rather than other less exciting ones—because we heard about it
in a media report.

This is one of many reasons why there is really no substitute for consulting the
opinions of scientific experts, at least if there is anything close to a consensus
in the field. The experts have already synthesized the evidence from a wide
variety of studies, so they're in a much better position to assess
the real significance of new studies. Relatedly, we can look for a meta-
analysis on the question—a type of study that tries to integrate the evidence
from all the available studies on the topic.

7.3 Misleading correlations


We've now reached the final and most difficult step in our three-step
argument. Having established a general correlation between two factors, we
still have to infer a causal relationship between them. Unfortunately, even
real correlations are often misleading in various ways.

There is no general recipe for ruling out misleading correlations, partly


because theorists disagree about the precise nature of causation. But we can
all agree on some common ways in which a correlation can be misleading. In
this section, we'll focus on five in particular:

 reverse causation
 common cause
 a side effect (e.g., placebo)
 regression to the mean
 mere chance

Let's consider each of these in turn.

Reverse causation
Sometimes, when A and B are genuinely correlated, we get the direction of
causation wrong: we think that A causes B when actually B causes A.

For example, suppose we study people's overall level of happiness and also
whether they get and stay married. We find a correlation between being
happy and being married. Given this, it's tempting to conclude that marriage
increases happiness. But what if the causal relationship goes the other way
around—happy people are just more likely to get and stay married?

We can't rule this possibility out just by observing the correlation. But there
are other things we could do. For example, we could study whether getting
married is related to change in happiness over time rather than overall
happiness. If we find a greater increase in happiness over time for people who
got married, that would provide more evidence that marriage is the cause. On
the other hand, we should be extremely wary about a simple causal story
like Marriage makes people happier. For example, it could be that the people
who got married used to be less happy because they wanted to be married;
but the people who didn't get married aren't the sort of people who would
have been happier if they got married anyway. Only a randomized controlled
study could really rule this possibility out, and that would be impossible to
run! (We'd have to take a group of people and randomly pick some to get
married and some to stay unmarried.)

"I wish they didn't turn on the seatbelt sign so much! Every time
they do, it gets bumpy."
—Billy, from Family Circus (Bil Keane)

As you'll recall, correlation is symmetrical. But it's interesting to note how we


state correlations when we want to suggest a causal relationship: we always
mention the alleged cause first. So, for example, there is obviously a
correlation between how serious a fire is and how many firefighters go to the
fire. But stating the correlation the other way round suggests a causal
relationship in the wrong direction:

 The greater the number of firefighters who go to a fire, the worse the
fire is!

Likewise, for many other real correlations:

 The more a student is tutored, the worse their grade!


 The more diets you go on, the less likely you are to be normal weight!

These statements are actually all true if we take them as merely reporting
correlations. But they also suggest causal relationships—in the wrong
direction. The fact that there is a standard way to communicate causal
relationships simply by stating a correlation is a telling sign about our
tendency to look for causal stories and run the two things together.
Sometimes a correlation could be taken to provide evidence for a causal claim
in either direction. In that case, it's possible to influence which causal
conclusion people jump to, simply by choosing which of the correlated things
you mention first. And of course, when media outlets report scientific results,
they have an incentive to spin the correlation in the most exciting way
possible.

So, for example, what's the most sensational way to report a study finding
that couples in their 40s who look younger than they actually are have sex
more often than couples in their 40s who look older than they actually are?
One could report that result in many ways, for example by saying that looking
younger is correlated with having more sex, or having more sex is correlated
with looking younger. Not only did media outlets choose the second way of
reporting, they often just leapt to the causal conclusion that "sex is the secret
to looking younger," a conclusion with no good evidence for it at all. The idea
that having more sex could cause us to look younger might be appealing, but
the correlation found in the study is just as consistent with the conclusion
that looking younger leads to more frequent sex.

Then again, maybe what's going on is slightly more complicated: perhaps


people who have the time and motivation to maintain their youthful
appearance are also more likely to have the time and motivation to maintain
an active sex life. In that case, the correlation is due to a common cause, which
is the topic of our next section.

Common cause
Probably the most important type of misleading correlation is when two
factors are correlated due to a common cause—that is, a third factor that
influences both of the others. The mistake is to think that A and B are
correlated because A causes B, when actually the correlation is due to the fact
that C causes both A and B.
Suppose a child who lives in a temperate zone believes that snow comes in
the winter because the clouds want to cover up the trees that have lost their
leaves. The child has noticed a real correlation: every year after the leaves fall
from the trees, the snow falls from the clouds. And there is a real causal
relationship here, just not directly between those two factors. Instead, both
factors share a common cause: falling temperatures.

Common causes are a major problem for non-randomized studies that find
correlations between two factors in a large population. For example, suppose
we find that swimming is correlated with better health outcomes than
running or playing most team sports. This by itself should not make us very
confident that swimming actually improves health more than those other
sports, especially since we know that swimming requires access to a
swimming pool, which in turn may require pool fees, etc. In other words, it's
plausible that swimming is at least somewhat influenced by income, which
we know also influences health outcomes. So the correlation we've
discovered might just be due to a common cause: income affects people's
recreational activities, and separately affects their health outcomes. We can
try to control for this effect by comparing health outcomes only across people
of the same income. But this may not solve the more general problem of
possible common causes in our study: there could be other factors related to
socioeconomic status that affect recreational activities and also health, and
that aren't quite captured by income.

Take another example. Suppose we want to know whether broccoli is good


for people's health, so we take a large group of people, look at how much
broccoli they eat, and then track their health outcomes. A major problem with
this approach is that broccoli is widely considered to be a healthy food. So even
if it's not actually affecting people's health, we should expect the kind of
person who eats more broccoli to also be the kind of person who does other
things that are considered healthy—like exercising, refraining from tobacco,
etc. We could try to control for all of those variables by only comparing people
who are the same in every other respect but differ in their broccoli intake. But
being the sort of person who cares about health or safety might lead to other
differences in lifestyle that we haven't thought of. This means it's very difficult
to rule out the possibility that being that sort of person is a common cause of
both broccoli consumption and increased health. The only way to be sure that
broccoli itself causes increased health would be to use a randomized trial.
(More on that below.)

The same holds for many other behaviors that are widely considered
beneficial, such as buying a car that seems safe. Suppose we find out that
Volvos are less likely to be involved in fatal accidents than Fords. Is this good
evidence that Volvos are safer than Fords? Ironically, the reason it's not is
precisely that Volvos have a reputation for safety, which means Volvo owners
are more likely to be safety-conscious to begin with, and safety-conscious
drivers are less likely to be in fatal accidents. (Luckily, we can also compare
cars using crash-test ratings that take the driver out of the equation.)

"I used to think correlation implied causation. Then I took a


statistics class. Now I don't."
"Sounds like the class helped."
"Well, maybe."
—Randall Munroe, [Link]

Some factors are so pervasive that they end up being common causes for a
great many correlations. In addition to socioeconomic status, consider global
trends like population growth and economic progress. Together, these have
lead to a steady increase in a great many measurable factors, so that we
would find at least a rough correlation over the last twenty years between
such apparently unrelated things as the number of avocados consumed in
Michigan, the average quality of wifi in Spain, and the number of haircuts per
capita given in India every day.
Here's a final and well-known example. Many people have used anecdotal
evidence to argue that a causal link exists between vaccines and autism. Since
certain signs of autism tend to occur around the same time as the
recommended age for the MMR vaccine, there are many parents who can
report that soon after the vaccination, they noticed signs of autism in their
child. However, because age is a likely common cause, a proper test would
compare children of the same age who are vaccinated with those who are not.
The most comprehensive studies are quite definitive that vaccines do not
cause autism. (However, as this SMBC comic ironically points out, there could
be a causal relationship in the other direction!)

Side effects
Sometimes there is a genuine causal relationship between A and B, but it's
not the relationship we expected. In particular, B may be caused by a side
effect of A.

For example, a study may find that a drug is correlated with reported
reduction in pain, but also that a fake pill with no active ingredients is equally
effective. This is not the same as saying the drug has no effect. Instead, its
effect is not due its chemical composition—it's due to the expectation that it
will be effective. When a treatment is effective due to this kind of expectation,
that's known as the placebo effect.

Note that placebo treatments needn't be pills. They can be anything that we
expect to be effective. Suppose I think that a walk in the park will help my
headache. If I take a walk and then felt better, it could be the fresh air and
exercise that helped, or it could be the expectation that I'd feel better. An even
more likely explanation, perhaps, is that my walk made me feel better
through a combination of these factors.
Recently another form of non-pill placebo has become popular: silicone wrist
bands with holograms whose alleged "holographic power" that will increase
your strength and balance. (These have been sold under various brands and
cost significantly more than a band of silicone should!) You can find a video of
a very thorough placebo-controlled test of one such band here. (In the study,
subjects don't know if they're holding the hologram band or a placebo
band.) Unsurprisingly, subjects did equally well with both, providing strong
evidence that any effect from the hologram band was a placebo effect.

Don't get me wrong—I'm not saying the hologram bands didn't work. In fact,
my guess is that people really do tend to perform better with such bands, at
least while they are thinking about them. The mechanism of expectation is
extremely powerful, especially for things like athletic performance, where the
individual's psychological state is enormously important. That's why many
people—no doubt including many of the people marketing the bands—
believe the bands work. They really do work! They just work by way of the
placebo effect, not by holographic power. (In contrast, placebo treatments
tend not to work for things like reducing the size of a tumor.)

This is one reason why using placebo treatments is an important part of


testing medical interventions: it's the only way of being sure that the
treatment has efficacy beyond that of the placebo effect. (And as you might
imagine, there is some controversy about prescribing an intervention just
because it's likely to have a positive placebo effect.)

Regression to the mean


Another important type of misleading correlation is due to regression to the
mean. This is the tendency, when selectively picking data points that lie
outside the mean, for adjacent data points to lie closer to the mean. (Not to
be confused with my brief aside about Spiders Georg in the previous chapter
—that was a digression to the meme.)
This image shows us a value that generally increases over time but also
bounces around a lot above and below its overall trend. Now suppose we
were to selectively choose points on the blue line that are significantly higher
than the red line. Even though the overall trend is upwards, we would expect
the value to fall a bit from those local high points. Likewise, if we were to start
from only points that are much lower than the trendline, we would expect
subsequent data points to be higher. This is what it means for data to regress
towards the mean.

Regression to the mean is an extremely common thing, but can be very


misleading. For example, when I have an extremely bad day, I tend to have a
better day the next day, due to regression to the mean. Now suppose I apply
some intervention only on the bad days. (It could be anything: taking a pill,
meditating, or calling my mom.) If I do this every time I have a terrible day, it
will create a pattern between applying the intervention and then feeling
better a bit later. This is a real correlation, but it's only due to a kind of
selection effect. I didn't randomly select days to apply the intervention, and
then test whether I felt better. I only selected days that were especially bad—
i.e., days when I was likely to regress to the mean the next day regardless.

Or suppose you're the mayor of a city and you want to give people the false
impression that you've improved the traffic situation. You can find the
intersections that had particularly high numbers of accidents this year, as
compared to previous years. Then install an intervention, like a traffic camera,
a new light pattern, or extra signs. Chances are the following year the rate of
accidents will drop at those intersections, even if the interventions had no
effect! The magic is in picking the right intersections—regression to the mean
will do the rest.

The same illusion should be expected for any interventions that are applied
only in cases when things are especially bad. It doesn't matter whether the
intervention is alternative medicine, conventional medicine, supernatural
healing, or a new investment strategy. These things may or may not be
effective: but we get no evidence that they're effective when we see the
improvements that we should expect anyway due to regression to the mean.

Here's a final example. Are punishments more effective than rewards in


influencing children's behavior? If we think in terms of regression to the
mean, we can see that there is likely to be a two-fold illusion:

 Particularly terrible behavior is likely to be followed by better behavior


regardless of our intervention, so punishment will tend to
look more effective than it is.
 Particularly wonderful behavior is likely to be followed by worse
behavior regardless of our intervention, so reward will tend to
look less effective than it is.

Taken together, these are likely to reinforce the idea that punishment is more
effective than reward, even if it's not.

So how does regression to the mean relate to the placebo effect? Even though
they are quite different things, they often operate together. This is because
many misleading correlations involve some kind of intervention (e.g., a
medicine) applied to particularly bad cases. But then:

 People often expect the intervention to work, which in turn can give
rise to a placebo effect, even if the intervention is not otherwise
effective.
 If we start from particularly bad cases, we should expect things to
regress to the mean even if the intervention is not effective.

Luckily, both types of misleading correlation can be ruled out using the same
kind of study, as we'll see below.

Mere chance
The issue here is not that our sample showed a correlation just by chance—
let's suppose we've ruled out that possibility by using a sufficiently large and
unbiased sample. The issue is that there might even be a correlation in the
population as a whole merely by chance. Luckily, such correlations tend to be
very narrow and strange. For example, consider the following graph, which
tracks two factors in the US over a decade:

This chart shows a highly significant scalar correlation for this particular
period of time: for the most part, the greater the number of letters in the
winning word in a given year, the greater the number of people killed in the
US by venomous spiders. It's extremely unlikely that this correlation over a
decade would happen by mere chance.

Given this, can we simply reject the hypothesis that the correlation is due to
chance? Not at all, because however unlikely it is that this correlation would
happen by chance, a causal relationship between the two factors is even less
likely. It would be ludicrous to think we could reduce the number of people
being killed by venomous spiders by reducing the length of the winning word
in the spelling bee—or vice versa. There just is no causal mechanism that
could explain the connection. (A causal mechanism is the specific way in
which one event causes another. For example, the causal mechanism by
which smoking causes cancer involves the formation of DNA adducts by the
carcinogens from cigarette smoke that are taken into the body.) In Chapter 9,
we will look more carefully at a number of factors that should be taken into
account when assessing the plausibility of hypotheses.

If we had to predict whether the correlation would continue into the future,
we should predict that it will not. So how can we explain the correlation?
Simple: even though it was very unlikely for it to happen by chance, it just
did. After all, sometimes very unlikely things do happen! And as a matter of
fact, I was not very surprised when I encountered this correlation, because I
found it with a powerful selection effect: I went looking on the internet for an
example of bizarre correlations. My search brought me to a website
called Spurious Correlations, whose owner had sifted through vast troves of
data to come up with the weirdest correlations he could find. Given the
complex data sets he was using, there were bound to be some very unlikely
chance correlations.

One lesson here is that we should be wary of charts that overlay multiple
different axes with different scales. Part of what makes these correlations look
so impressive is that the y-axes have been truncated and scaled to emphasize
the fact that the factors move up and down together.

The more important lesson, however, is that we should not be surprised to


see very unlikely coincidences when they are brought to our attention
through a process that sifts through a great many events specifically to search
for coincidences. Every day in United States, there are more than 300 people
who experience events so incredibly unlikely that they would only happen to
one in a million people on a given day. If we could search through the country
to find those events, we'd have no shortage of material to amaze us [2].

Of course, when people experience extremely strange events, they are likely
to talk about them, and others will often amplify their voice so that we do end
up hearing about them. Given our inter-connectedness, we should therefore
expect to hear about some genuine events that are so unlikely to happen by
chance that they seem to be better explained by ghosts, aliens, or other X-
factors unknown to science. Should we conclude that such events are
evidence for those X-factors? "After all," some say, "the probability of this
event occurring just by chance through natural causes is incredibly low. So
this event is strong evidence for X."

That might look like a pretty good argument until we remember the selection
effect involved. We have an enormous pool of possible events, coupled with a
process that sifts through them and focuses on the most improbable ones. We
should expect such a method to give rise to data that is very misleading.
(Note, also, that the argument fails to carefully assess the probability of the
evidence given the X-factor hypothesis: for example, does the hypothesis that
ghosts exist really make it likely that this event would happen to this person,
given the zillions of possible ways for ghosts to manifest? [3])

Evidence & experiments


Recall the general form of an argument from correlation to causation:

[Link] observed a correlation between A and B;


[Link] is a general correlation between A and B; and
3.A causes B.

Suppose our observations do show a correlation. Can they take us all the way
from (1) to (3)? We need a large enough study that avoids sampling bias to
establish (2), and we need our evidence to somehow distinguish between a
genuine causal relationship and all the misleading correlations we've
discussed in §7.3.

The hypothesis we are testing is that A causes B. What we want is evidence


with a high strength factor, either for or against that hypothesis. This
means we want evidence that meets one of these conditions:

 it's far more likely if A causes B than if it doesn't; or


 it's far more likely if A does not cause B than if it does.

Ideally, we can devise an experiment that will give us an unmistakable result


no matter what: evidence that could basically only have been observed if A
causes B, or evidence that could basically only have been observed if A does
not cause B. In practice, this is very hard to obtain because there are so many
ways for a correlation to link A to B even if A does not cause B. What we want
is a study that can simultaneously rule all of these possibilities out.
This is why the gold standard for establishing causation from correlation is a
double-blind randomized controlled trial. In a randomized controlled trial,
subjects are randomly selected and placed into two groups. The treatment
under investigation is applied only to one group, but both groups are followed
and assessed for relevant changes. In a double-blind study , neither the
subjects nor the experimenters that interact with or assess them know which
group the subjects are in. (Usually subjects in the treatment arm are aware of
the treatment—such as a medication; so blinding the trial requires the control
group to receive a placebo treatment.)

Note how many possible sources of error this eliminates, by ensuring that no
factors other than the treatment might explain differences in the observed
outcomes between the two groups. To illustrate, suppose we ran a study like
this and found a significant difference between the groups, with effect B
occurring in the treatment group and not the control group.

 Bias in the sample. Since subjects are randomized, there's no sampling


bias between the two groups. Since the experimenters are blinded,
they can't accidentally bias the sample after the randomization by
treating the groups differently.
 Reverse causation. The correlation can't be due to reverse causation,
because the cause of treatment is the randomizing procedure of the
experiment itself.
 Common cause. There could be a common cause for (i) being selected
for the study and also for (ii) exhibiting B. But this would affect both
groups equally, and we found a difference between the groups.
 Placebo. Because the subjects don't know which group they're in, any
placebo effect would show up in both groups.
 Regression to the mean. The subjects could be regressing to the mean—
if, for example, they were selected due to having some condition. But
again, regression to the mean would affect both groups equally.
In short, the design of the study ensures that we'd be unlikely to get the result
we did if A does not cause B. So a study like this provides strong evidence that
it A does cause B. On the flip side, if a study like this shows no difference
between the groups, that's a result we'd only expect if A does not cause B, so
we get strong evidence that it does not.

Of course, our evidence would be even stronger if it stemmed from several


different sources, all pointing in the same direction. This would help to ensure
that the result doesn't stem from any particular flaw in the study's design, any
error on the part of the experimenters, or sheer bad luck. Evidence that stems
from such a wide variety of experiments can be described as robust.

Chapter 8
Key terms
Ad hoc modification: the process of altering our theories to account for new
evidence, without properly acknowledging that this makes them more
complex, and thus less plausible, than the original formulation of the theory.

Base rate: the statistical rate of occurrence of a feature or event in general.

Coherence: a criterion of theory choice. The more a theory fits with what we
know about the world, the more coherent it is with our background
knowledge. This gives a label to the point that we want all of our beliefs to fit
well together. If a theory clashes with other things that we have reason to be
confident about, then absent other evidence, we should find it unlikely. On
the other hand, if we get evidence that it's true, we'll also become less
confident in the beliefs that it clashes with.

Heads I win, tails we're even: This error involves failing to treat a new fact as
evidence against our favored position, even though its opposite would have
readily been welcomed as evidence for our position. There are various ways
we might make this error, including by ignoring the evidence or being
inconsistent with what we assign to the strength factor values. But the key to
this particular pitfall is the inconsistency between how we respond to the
evidence at hand, and how we would have responded to the opposite
evidence.

Neglect of priors: when we get new evidence for a claim, we need to pay
attention both to the strength of that evidence and the probability of that
claim before we got the evidence. (In our updating rule, that initial (or "prior")
probability gets converted into prior odds.) Often our initial probability comes
from a statistical fact called a base rate. When we take evidence into account
but ignore the base rate, that's base rate neglect. Someone who commits
base rate neglect may still be paying attention to all the evidence and using
the full strength factor. One way to diagnose base rate neglect in someone's
reasoning is to ask whether accounting for the base rate would have changed
their conclusion.

Neglect of total evidence: when someone updates selectively on some new


evidence but not other bits of new evidence that are relevant. Such a person
may still take the base rate into account and properly update using the full
strength factor with the evidence they do pay attention to.

Opposite evidence rule: to help us avoid ignoring evidence against a view


that we find plausible, it can be useful to ask ourselves how we would have
reacted to the opposite observation. If we would have treated the opposite
observation as evidence for our view, then we should treat the evidence we
have as evidence against our view. (Though the amount of evidence can be
different.) In other words, if E is evidence for H, then not-E is evidence for not-
H. Forgetting this leads to the error we call heads I win, tails we're even.

Planning Fallacy: the tendency to systematically underestimate how long it


will take to complete a project involving many steps. This tendency might be
explained by neglect of the conjunction rule.
Simplicity: this criterion of theory choice tells us that we should find theories
more likely the less complexity they add to our overall view of the world
(a.k.a. "Occam's Razor"). For example, if a theory were to involve accepting a
new type of object with an accompanying system of facts about that type of
object, that theory would be more complex than a theory without that type of
object, all else equal.

Updating: adjusting one's degree of confidence in a claim after getting


evidence for or against that claim. Updating should be performed according
to the updating rule.

Updating rule: the rule for properly updating one's degree of confidence in a
claim after getting evidence. The rule is: prior odds × strength factor = new
odds. This is equivalent to a combination of two rules this text does not
specifically covered in this text: Bayes' theorem and conditionalization.

8.1 How to update


The elements
Updating means adjusting our confidence in proportion to the strength of the
evidence. This means that there are two distinct elements—something that
gets adjusted, and something that does the adjusting. What gets adjusted is
the degree of confidence you had in the hypothesis before getting the
evidence: we call this your prior degree of confidence. And what does the
adjusting is the strength factor of the new evidence. The process of updating
involves these two distinct things coming together to yield a new degree of
confidence.

Recall from the Mindset chapter that it's critical to decouple these two things:
to set aside one's prior confidence when assessing the strength of the new
piece of evidence. That helps us avoid evaluating the evidence in a biased
way. And if we focus on the two questions that make up the strength test, we
can see that they are entirely independent from how confident we are in the
hypothesis, because they are suppositional:

 suppose H is true: how likely is this evidence?


 suppose H is false: how likely is this evidence?

The first question asks you to treat H as certain, and the second one asks you
to treat not-H as certain; neither one has anything to do with how confident
you actually are in H. So actually thinking in these terms should help us to
decouple the strength of the evidence from our prior confidence.

The rule
Ok, so once we have these two independent values, how exactly do they
combine to give us a new degree of confidence? It's easy to see that it can't be
as simple as multiplying or adding. For example, if my prior confidence is 99%
and I get evidence with a strength factor of 10, it makes no sense to say that
my new confidence should be 990% or 109%.

Luckily, the rule for updating is simple if we express our prior confidence as
a ratio rather than a fraction. This is common in gambling, where probabilities
are expressed as ratios, which are called odds. For example, if there's a race
between two horses and the bets reflect an equal probability of winning,
people will say the odds are one-to-one or 1 : 1.

Let's just review what you learned in 6th grade about converting fractions to
ratios. Suppose you have two apples and three bananas. Then the fraction of
all the fruits that are apples is 2/5. But the ratio of apples to bananas is 2 : 3.
The ratio directly compares the number of apples to the number of non-
apples, while the fraction compares the number of apples to the number of all
of the fruits (including the apples). So, in a fraction, the denominator is the
number you would get from adding both sides of the ratio.
So how does this apply to probability? Let's say you randomly select a fruit
from the 5 fruits. The probability that you pick an apple, expressed as a
fraction, is 2/5. And the same probability expressed in terms of odds is 2 : 3.
The way we say this out loud is "the odds of picking an apple are 2 to 3".

You can visualize the probability of H as involving lots of possible situations,


where H is true in some of them and false in others. Then you could express
the probability in two ways. As a fraction, it's the proportion of situations
where H is true out of all of the situations. As odds, it's the ratio of situations
where H is true vs. situations where it's false.

Once we express our prior confidence as odds, the updating rule is very
simple:

Updating rule: prior odds × strength factor = new odds

So, for example, let's say your prior confidence in H is 3/4. That means the
prior odds of H are three to one (or 3 : 1). We can visualize this as three
situations where H is true for every one situation where H is false. In other
words, H is three times more likely to be true than not true (which is why we
say the odds are three to one).

Now suppose we get evidence supporting H with a strength factor of 5. The


way we combine these two numbers is that we take the side of the odds
representing H—in this case, the left side of the odds—and multiply that
number by 5. You can see how the updating rule happens in the image to the
right.

So we end up with odds of 15 to 1, or expressed as a fraction, our new degree


of confidence is 15/16.

Detectives and fruit


Let's apply this rule to some cases.

 You're a detective in search of a fugitive. At this time of day, there is an


80% chance that he's at a neighborhood dive called the Cadillac
Lounge. The place has two rooms, and if he's there at all, there is an
equal chance of him being in either room. As you enter, you find that
he's not in the first room. How confident should you now be that he's
in the second?

People's intuitive answers to this question are usually wrong. In fact, most
people don't even know whether the detective should now think the fugitive
is probably in the lounge, or probably not in the lounge. Some people even
think the detective's confidence shouldn't change after seeing the first room.
But that can't be right. Just imagine you're checking through the lounge and
eventually have just the closets left: it would be crazy at that point to still be
80% sure that he's there! Surely you've been getting evidence all along that
he's not, so your confidence should slowly drop from 80% as you check more
fully.

Luckily, we can just use the updating rule. First, let's work out the prior odds.
Initially, there's an 80% probability he's in the lounge, and a 20% probability
that he's not. This means he's four times more likely to be there than not: the
odds of his being there are 4 : 1. (Equivalently, the odds of his not being there
are 1 : 4.)

Next, we need the strength factor of the evidence, which is that he's not in the
first room. Supposing he's in the lounge, the probability you would get this
evidence is 1/2. Supposing he's not in the lounge, you were certain to get that
evidence. So you are twice as likely to get the evidence if he's not in the
lounge: it has a strength factor of 2.
Since we got evidence supporting his not being in the lounge, we simply
multiply the odds of his not being there by the strength of the evidence.
Remember, we multiply the side of the odds for which we get evidence.

The odds of his being in the lounge are now 4 : 2, which is the same ratio as 2 :
1. So we started off thinking he was four times more likely to be in the lounge
than not, and after getting evidence with a strength factor of two, we now
think he's only twice as likely to be there than not. Expressed as a fraction, the
probability he's in the lounge is now 2/3. You should still think he's probably
in the lounge, but you shouldn't be nearly as confident as you were before.

Next, let's take the case from Chapter 1 that was used to illustrate how we are
naturally poor at thinking about probability:

 Two bags are placed on a table. One contains two apples, and the other
contains one apple and one orange. You have no idea which bag is
which. You reach into one, and grab something at random. It's an
apple. Now what's the probability that the bag you reached into has
an orange in it?

The answer that jumps out is 50%, but a simple application of the evidence
test shows that can't be right. Grabbing an apple was more likely if this was
the bag with two apples, so your observation is evidence that this is the bag
with two apples.

Applying the strength test, we can see that the observation was certain to
happen if this was the bag with two apples, but only 50% likely if this was the
bag with an orange. This means you have evidence with a strength factor of 2
that this is the bag with two apples. Starting with prior odds of 1 to 1, we get
new odds of 1 to 2 that this the bag with an orange, or a probability of 1/3.
Test your physician
Let's go back to the medical case that stumped so many Harvard physicians:

 Among people with no symptoms, one person has Disease X for every
thousand people who don't. A random person with no symptoms is
tested and gets a positive result. The test always gives a positive result
when someone has the disease, and only gives a positive result 5% of
the time when they don't. How likely is it that this person has the
disease?

It's pretty easy to see what the right answer is if we start by imagining a large
number of patients. We know that, among asymptomatic people, there is one
person who has the disease for every 1000 people who don't. So let's
represent those people visually: the purple square is the person who has it.

Now suppose we test all of these people. How many will test positive? The
one person who has the disease will definitely test positive. But in addition,
5% of the people who don't have the disease will test positive. So, 50 people
will test positive out of the 1000 people who don't have it!

I'll use blue to cover all 51 people who test positive:

Now it's easy to see how likely it is that a random person who tests positive
actually has the disease. We can ignore all the people who don't test positive,
since we're only interested in the people with a positive test like our subject.
Out of those people, how many actually have the disease? In other words,
what's the probability that a random person from the blue-covered area has
the disease? Clearly, the answer is 1/51, or just under 2%. So all those doctors
who thought the patient almost definitely had the disease were way off!
So how does this fit with our updating rule? The method we just used, in
which we imagined 1001 people and thought about which of them would test
positive, is actually very similar to the updating rule. We started off with a
thousand people who did not have the disease, and then narrowed them
down to the 5% that had a positive test. And we started with one person
who did have the disease, and that person definitely have a positive test. This
left us with a ratio, among people with a positive test, of 1 person with the
disease to every 50 without it.

What happens if we use the updating rule instead? We start with prior odds of
having the disease of 1 : 1000, reflecting the initial ratio of people with the
disease. Next, we work out the strength factor of testing positive. A positive
test is guaranteed if someone has the disease, so the probability of the
evidence given the hypothesis is 1. And the probability of the evidence
supposing the person does not have the disease is 5%. So the evidence is 20
times more likely to occur if someone has the disease than if they don't. In
other words, it has a strength factor of 20 in favor of having the disease.

We now multiply the prior odds of 1 : 1000 by the strength factor of 20, giving
us odds of 20 : 1000, which reduces to 1 : 50. Clearly, the two methods of
working through this problem take the same initial ratio and transform it in
equivalent ways, giving us the same result.

8.2 Combination rules


We're going to take one more step into the rules of probability so that we can
handle more cases. But first I'd like to introduce some simple notation that
will help us express some of these ideas more exactly. It's just two letters and
two symbols:

 "P" means "the probability of"


 The upright bar ( | ) means supposing/given
 The tilde ( ~ ) means not

Remember that the strength test for evidence is "How much more likely is this
if H is true than if it's not?" We can now restate that test using our notation,
where "E" is some new fact we've learned:

The strength test: How many times greater is P(E | H) than P(E | ~H)?

In other words, the strength factor of E in favor of H is equal to P(E | H) divided


by P(E | ~H).

Conjunctions
We can also use this notation to state some rules about how probabilities
combine. First, if we have two hypotheses A and B, what's the probability
that both are true? In other words, what's the probability of the conjunction of
the two hypotheses?

We can represent probability visually using area: the proportion of a total area
taken up by a hypothesis corresponds to the probability of that hypothesis.
Now, out of all the scenarios where A and B are true, the conjunction is true
where they overlap.

In this case, we can see that the probability of A is 1/4, represented by having
A cover 1/4 of the total area. And we can also see that B covers about a third of
the total area, but it covers half of the area that A covers. If we look at just the
area that A covers, we are looking at how things are given A. So the probability
of B given A is 1/2. But then it follows that the part where A and B overlap is
1/8 of the total area. So the probability of the conjunction (A & B) is 1/8.

This gives us a general rule to find the probability of a conjunction of any two
claims: just multiply the probability of one of them by the probability of the
other one given the first one. (That is, either multiply the probability of A by
the probability of B given A; or multiply the probability of B by the probability
of A given B.)

Conjunction rule: P(A and B) = P(A) × P(B | A)

In some cases, supposing that one hypothesis is true makes no difference to


the probability of of the other. In such cases, we say that they are
probabilistically independent. In that case, since P(B | A) is the same as P(B),
the probability of the conjunction is just P(A) × P(B).

For example, suppose the probability that it's raining in New York today is 1/4
and the probability that it's raining in Manila is 1/3. What is the probability
that it's raining in both places? It's pretty safe to suppose that rain in one
place is totally unrelated to rain in the other. In that case, we can simply
multiply P(New York) by P(Manila) to get the probability of the conjunction,
which is 1/12.

This image helps to illustrate why. The rain-in-NYC area covers 1/4 of the total
area, and the rain-in-Manila area covers 1/3 of the total area, to represent the
prior probabilities. Because the hypotheses are independent, if we suppose
that one is true, it doesn't change the probability of the other. When we use
area to represent probability, supposing H is the same as zooming in on H. So
independence means that if rain-in-NYC covers 1/4 of the total area, it also
covers 1/4 of the rain-in-Manila area. And it follows that the overlap area must
be 1/12 of the total area. (Note that it if the overlap area is 1/12 of the total
area, rain-in-Manila must also cover 1/3 of the rain-in-NYC, illustrating that
independence always goes both ways: if supposing A doesn't affect the
probability of B, the reverse must also be true.)

One important lesson to take-away is that a conjunction of two claims must


always be less probable than either claim on its own, unless one of them
entails the other. (If A entails B, then B covers all of the A area, so the overlap
area will be the same size as the A area.) And in general, a claim gets less
probable the more conjunctive it is. So even if each element of the claim is
itself quite probable, the conjunction may still be improbable. Unfortunately,
we are pretty bad at keeping track of how unlikely our hypotheses get as they
become more conjunctive. For example, you may recall from Chapter 6 that
after hearing a description of a woman named Linda, a large majority of
people ranked Linda is a bank teller and is active in the feminist movement as
more probable than Linda is a bank teller. Somehow our brains fail to notice
that it's impossible for the more conjunctive claim to be more probable than
the simpler one.

One way this error shows up in our practical lives is that we tend to
underestimate how long it will take to complete a project involving many
steps—a pitfall called the planning fallacy. In studies, when people are asked
to predict how long it would take them to finish a long project, fewer than a
third actually finish by the time they predicted. And only about half finish by
the time they predicted they would finish "if everything went as poorly as it
possibly could" [3]!

Suppose that your homework project has five parts that you must complete
without getting stuck on any of them. For the first part, there's an 80% chance
that you won't get stuck. Given that you don't get stuck on the first, there's an
80% chance you won't get stuck on the second, and so on. Then it may seem
like there's a good chance that you won't get stuck at all. But in fact, the
probability that you won't get stuck is .8 to the power of 5, or less than 1/3. So
there's actually a 2/3 chance that you will get stuck! The planning fallacy is
just another case of failing to keep track of how unlikely a hypothesis gets as
more conjunctions are added to it.

Disjunctions
Next, if we have two hypotheses A and B, what's the probability that at least
one of them is true? In other words, what's the probability of the disjunction of
the two hypotheses? In the simplest case, we simply add up the probabilities
of A and B. For example, what is the probability of rolling either a 5 or a 6 on a
6-sided die? Picture the situation using relative area on a "chocolate bar" to
represent probability:

Each number has a probability of 1/6 of being rolled and so takes up 1/6th of
the probability space. Obviously, the probability of rolling either a 5 or a 6 is
2/6. But simply adding the two probabilities only works when A and B
are mutually exclusive, meaning they can't both be true. You can't roll both a
5 and a 6 on the same roll: that's why there's no overlap between the two
areas on the chocolate bar.

What if the two claims are not mutually exclusive, though? In that case, their
probabilities overlap. If we were to just add them up, we'd be counting the
overlap area twice. For example, consider the NYC-Manila example again:

What's the probability that it's raining in at least one of the two places (i.e. the
total colored area)? If we just add up the probability of rain in New York and
the probability of rain in Manila, we'd be double-counting the days when it's
raining in both places. We only want to count the dark purple area once.

So, to get the probability of one or the other happening, we just add the
probabilities of the two events and then subtract the probability of the
conjunction (i.e., the overlap). In this case, this means we add 1/3 and 1/4,
giving us 7/12, and then subtract 1/12. So the probability that it's either
raining in New York or in Manila is 6/12 or 1/2. (And indeed, you can see that
the colored area comprises half of the chocolate bar as a whole.)

Using our formal notation, the general addition rule for disjunctions is this:
Disjunction rule: P(A or B) = P(A) + P(B) − P(A and B)

Note that in cases where the claims are mutually exclusive, there's no overlap
area, so we would just be subtracting zero at the end. In that case, we can
ignore the last part of this rule, and just add the probabilities of A and B.

8.3 Updating pitfalls


Recall the two probability pitfalls we discussed in Chapter 5: one-sided
strength testing and heads I win, tails we're even. These are worth reminding
ourselves about before considering three more pitfalls.

In one-sided strength testing, we simply notice that an observation fits quite


well with our first or favored view, and so treat it as actually supporting that
view. This, of course, forgets that the observation is only evidence if
it's more likely given H than given ~H: to assess that comparison we need to
consider how we'd expect things to look if the opposite view were true. (And
that, as we saw, is a clear benefit of the de-biasing technique that
psychologists call considering the opposite view.)

The mistake we call heads I win, tails we're even involves failing to recognize
that if an observation would support H, then the opposite observation would
support not-H. This means, if we are considering some test of H, we can't treat
one outcome of that test as evidence for our favored view unless we would
have treated the opposite outcome as evidence against our favored view.
Likewise, we can't treat an observation as neutral if we would have
considered its opposite as evidence for our favored view. (And thinking
carefully about how we'd have responded to the opposite result is a benefit of
the de-biasing technique that psychologists call considering the opposite
evidence.)

Now that we have a more advanced understanding of how to adjust our


degrees of belief in proportion to the evidence, we can identify three
additional errors that people commonly make when they consider new
evidence:

 neglect of priors
 neglect of total evidence
 ad hoc modification

Neglect of priors
In the example of Disease X, even though our test provided fairly strong
evidence that the patient had the disease, the evidence was not enough to
make it likely that the patient had the disease. The very low base rate of the
disease—that is, its general prevalence in the relevant pool of individuals—far
outweighed the new evidence. But even though the physicians who were
asked about the case were told the base rate of the disease, most of them
failed to give it enough weight when updating. And since the base rate is what
should set their prior odds in this case, this is an instance of neglect of priors.

How common would Disease X need to be in the general population for us to


feel 95% confident that a patient who tests positive actually has it—as half the
doctors said we should be? The answer is that fully half the asymptomatic
population would need to have Disease X! That would give us a prior odds of
1:1, and updating with evidence that has a strength factor of 20 would give us
new odds of 20:1, or a probability of just over 95%. In the image below, you
can see 501 people with the disease who have a positive test, and only 25
people without the disease who have a positive test (that's 5% of the
remaining 500). If you randomly select someone with a positive test, there's a
95% chance they have the disease.
The guess that these physicians made fits with a more general pattern: when
people neglect to actually think about the prior odds of a hypothesis, they
often reason as though the prior odds were 1:1.

Another kind of case in which we frequently forget about base rates is when
we encounter relative comparisons of unlikely events. For example, suppose
you discover that you have a 50% increased risk of a certain disease. How
much should this concern you? It depends on the base rate: for example, if
only one in a million people get this disease, your chances have gone from 1 in
a million to 1.5 in a million! Or suppose you hear that shark attacks are up by
300% this year. Looking at that "300" makes it hard to recall that tripling a tiny
number can still result in a tiny number! When something is rare, even a very
significant relative increase may fail to achieve significance as
an absolute increase. (This is effectively what is happening when we have
pretty strong evidence but the base rate is so low that the absolute increase in
our confidence is pretty minor.)

Lotteries are a good illustration of the importance of priors. Suppose that


Fred, an ordinary guy, wins a lottery with a million tickets. The probability that
he'd win if he's an ordinary player is 1 in a million; but he'd definitely win if he
successfully rigged the lottery in his favor. So Fred's winning is very strong
evidence that he rigged the lottery: it has a strength factor of a million! And
yet when someone like Fred wins the lottery, it doesn't seem like any reason
at all to think the lottery was rigged. What's going on?

The priors are key. We know lotteries are safe and happen all the time with no
one being suspected of cheating. So the prior probability that someone would
rig this lottery is fairly low to begin with, let's say 1/1000. And let's suppose we
consider each of the ticket holders equally likely to rig the lottery, so that
1/1000 gets divided by a million to arrive at the prior that Fred would rig the
lottery: 1 in a billion. This means that when he wins, that strength factor of a
million yields a probability of 1/1000 that he rigged it. But at the same time,
the probability that someone successfully rigged the lottery hasn't changed at
all: it remains 1/1000. What we now know is that if someone successfully
rigged the lottery, it was Fred. [4]

For a final example, imagine we're investigating a murder and we have a DNA
sample from the murderer. We scan a large national database of DNA samples
from random people… and we find a match! Only 1 in a million people has
this particular DNA sequence, so the chance of being a match if you're
innocent is one in a million. But if you're guilty, you're definitely a match. So
this evidence has a strength factor of a million!

Does this mean the person with the matching DNA is probably guilty? Actually,
no. Not if they are just a random citizen and don't have any connection at all
to the crime. After all, in a country of 325 million people, there should be
about 325 matches! And they can't all be guilty. To put it in terms of the
updating rule, that person's prior odds were at most 1 to 325 million: even
multiplying the odds by a million only gets us to odds of 1 to 325. This
highlights the difference between the chance of being a match if you're
innocent and the chance of being innocent if you're a match! The former can be
as low as one in a million while the latter is still pretty high.

Neglect of total evidence


In real life, a DNA match is likely to be identified from a small number of
suspects, not from a large national database. Suppose one man is not only a
DNA match but also knows the victim and left a footprint at the scene. There
may be 325 other people in the country with matching DNA, but it's extremely
unlikely that any of the others were at the scene of the crime.

Now suppose the suspect picks on one of these facts and says: "So my
footprints were found near the crime scene—but you found 10 pairs of
footprints there! So there's a 90% chance I'm innocent!" That would
be neglect of total evidence: i.e., ignoring the total evidence available and
only updating on some of it. If all we had were the suspect's footprints at the
crime scene, we should think he's probably innocent. But we know a lot more
than that and we need to update on all of the facts. The combined evidence
has a strength factor that establishes guilt beyond a reasonable doubt.

The famous murder defense of OJ Simpson provides an apt example of this


pitfall. When the prosecution presented evidence that Simpson had been
violent toward his wife, the defense pointed out that, statistically, only one in
2500 abused women is actually murdered by her husband [5]. The point of
using this statistic was to suggest that this evidence only brought the
probability of Simpson being the murderer to 1 in 2500. But this ignores a key
piece of evidence—the fact that Nicole Simpson was murdered! The real
question is: of women who were abused by their husbands and later
murdered, how many were murdered by their husbands? And the answer to
that question is: almost 90%. The point is that we can't pick and choose which
evidence to update on. It's the total evidence that counts.

Ad hoc modification
As in the example from Chapter 5, Fred believes that he has ESP, claiming that
he can use his powers to guess which type of card a person has randomly
pulled from a deck. To test his claim, we set up some controlled conditions to
see if his guesses are better than chance.

This time, instead of just claiming that he got unlucky, Fred responds to the
failed test by modifying what he believes about his ESP. Rather than
becoming less confident about having ESP, Fred just takes himself to have
learned that he must have a form of ESP that can't be tested. He then points
out that the test isn't any evidence at all against his new claim:

I have a special form of ESP that doesn't work under experimental


conditions.
After all, supposing this were true, we'd expect the test to fail! This is called ad
hoc modification: responding to new evidence by modifying one's hypothesis
to make it fit the evidence better, without recognizing that reduces the
plausibility of the hypothesis. It's very frustrating when others do it, even
though it's barely noticeable when we do it ourselves.

But what's wrong with ad hoc modification, exactly? It's true that our test isn't
evidence against Fred's new claim. The problem is that he's treating the new
claim as though it were just as plausible to begin with as his original one. But,
in fact, it's a far weirder and more specific claim and therefore far less
plausible than the original. So, before the test, whatever probability we
assigned to Fred having ESP, we should have assigned a lower (arguably,
much lower) probability to his having ESP that doesn't work under
experimental conditions.

Here's a more mundane example. You're 50% confident that your friend is
home. You knock on her door but she doesn't answer. At first you feel like this
is strong evidence that she's not home; but then it occurs to you that not
answering the door is also evidence that she's in the shower. And if she's in
the shower, she's home. So why not just switch to that hypothesis instead,
and continue to be just as confident that she's home?

Let's apply the updating rule. Before getting the evidence, let's say your prior
was 1/10 that she's in the shower given that she's home (and 1/20 overall).
You know that she'd definitely answer if she's home, unless she's in the
shower. So the evidence that she didn't answer the door is 10 times more
likely if she's not home than if she's home, meaning it has a strength factor of
10 in favor of her not being home. This means your odds should go up to 10 : 1
that she's not home. But her failing to answer is also evidence that she's in the
shower! You've narrowed the possible situations down to two: 10 : 1 odds that
she's not home, and 1 : 10 odds that she's in the shower. The weird-sounding
part is that the very same observation can be evidence against a general
hypothesis (she's home) and also for a specific version of that hypothesis
(she's home but in the shower). But in this case, even though you get some
evidence that she's in the shower, you get much stronger evidence that she's
not home at all—and indeed you should become pretty confident that she's
not home.

In a defensive mindset, it's very common for people to refuse to acknowledge


evidence against a theory, and instead fall back on a highly specific version of
their theory that manages to fit the evidence. But it's absolutely critical to
recognize that unless our evidence deductively entails that a hypothesis is
false, it is always possible to retreat to a more complex and bizarre version of
the hypothesis that fits the evidence perfectly.

For example, even if you directly see your friend somewhere other than home,
you could always retreat to the hypothesis that she's actually home but you're
just dreaming. So the fact that there is some more specific version of our
claim that fits the evidence means literally nothing except that the evidence
doesn't deductively entail that our claim is false—and remember, we almost
never get absolute proof of anything in real life.

If there's a single most important error made by conspiracy theorists, it's


probably this. As new facts crop up that should cast doubt on the conspiracy
theory, it morphs and adapts to explain the available evidence, and somehow
in that process the conspiracy theorists fail to update on the evidence against
the theory in general. It's as though all of the prior probability they had
assigned to the original theory gets to carry over to the ever more bizarre
modifications of it. But that way of responding to evidence will allow us to be
completely unresponsive to evidence against any theory.

8.4 Assessing priors


As we've seen, how confident we should be after getting new evidence for a
claim is a matter of two things: (i) our prior degree of confidence in the claim,
and (ii) the strength of the evidence. The more implausible a claim is to begin
with, the more evidence we need before becoming confident that it's true. But
so far we haven't said much about where we get our prior confidence.

For example, suppose a random person comes up to you and declares that he
can see the future. You happen to have 20-sided die in your pocket, so you
challenge her to predict the outcome of your next roll. She says the die will
come up 14. You roll the die and it comes up 14. The stranger smiles and walks
away. You just got evidence with a strength factor of 20 (and a p-value of .05)
that she can actually see the future. Clearly, it should take far more evidence
to convince you that she can see the future, so even after updating you should
be very confident that her prediction was a coincidence. This means your
prior degree of confidence that she is clairvoyant should have been extremely
low to begin with. But why, exactly? You presumably aren't working from a
statistical fact about the proportion of people who can see the future.

So far, we've worked mostly with cases where there is some statistical base
rate to go on: 80% of the time on an evening like this, the fugitive is in the
lounge; 1% of people in the general population have Disease X; etc. But for
many claims, we need to be able to assign reasonable priors without relying
on any clear statistical base rates:

 Neanderthals figured out how to create fire.


 There are real ghosts that sometimes haunt places.
 The universe is filled with a hard-to-detect form of matter ("dark
matter").
 When Caesar crossed the Rubicon, he intended to start a civil war.

We can get evidence for and against each of these claims, and updating on
evidence requires us to assign a prior probability to the claim we're
evaluating. But in these cases, it's unclear how there could be something like
a statistical base rate to help us assign prior probabilities to them. So we have
to use other criteria to evaluate them prior to getting new evidence. We'll
focus on two such criteria: coherence and simplicity.

Coherence
Suppose we just discovered the very first evidence that Neanderthals could
make fire. Our discovery is what looks like a fire pit in a Neanderthal cave, and
doesn't seem likely to just be sticks from a forest fire or something else. In
other words, it's far more likely that we'd find this pit if Neanderthals could
make fire than if they couldn't—let's say 10 times more likely. Now how
confident should we be that Neanderthals could make fire?

To answer this, we need to assess how likely the claim was independently of
that evidence. And there doesn't seem to be any statistical fact that we can
use as a base rate: for all we know, mastering fire is something that has only
ever happened once. But this doesn't mean we simply had no idea how
confident to be that Neanderthals could make fire. We may not be able to
assign a very specific prior probability, but we know it's much higher than one
in a million and much lower than one. We can also compare it to our prior
confidence in other theories, such as the theory that Tyrannosaurus Rex could
make fire. It would take much stronger evidence to convince us of that theory,
meaning that it has a lower prior probability!

Why is this? Well, one of these theories simply fits much better with other
things that we already know about the world—in other words, it coheres with
our background knowledge. From fossil records, we have a rough sense of how
much brain-power and manual dexterity the two species likely had. If we
discovered that Tyrannosaurs had mastered fire, we'd have to revise much of
what we thought we knew about them. The more a hypothesis clashes with
other things we are confident about, the less plausible we should consider it
prior to getting evidence for it.
To take another example, suppose we know that one of these two claims is
true:

H1: All ravens are black


H2: Almost all ravens are black.

To test them, we take a random sample of a thousand ravens and find that
they're all black. This provides us with some evidence for H1: if it's true, all the
ravens in the sample would have to be black, while if H2 is true, some ravens
in the sample might not have been black. But this doesn't mean we can just
conclude that H1 is more likely than H2: we haven't considered their prior
probabilities! Statistical generalization, done properly, is just another form of
updating: it requires assessing both the strength factor of the
evidence and the prior probability of the hypothesis.

So we have to first ask how well the two theories fit with our background
knowledge about biology. Well, one thing we know is that rare forms of color
variation (e.g., albinism) occur in most species of animals, including birds.
And, the biological basis for color variation in ravens is probably not that
different from the basis for color variation in other animals. So H2 coheres
better with our background knowledge than H1 does.

Before our sample, then, we should have considered it likely that some ravens
are albino. And although our sample did provide some evidence for H1, it
didn't provide very strong evidence. (Typically about one in five thousand
members of a species will be albino, so if there are some rare albino ravens,
they could easily have failed to show up in our sample.) Given our priors, we
would need much stronger evidence before deciding that H1 is more likely
than H2.

Simplicity
For our second criterion, let's take a very different example. Consider the fact
that some people feel certain they have seen or heard ghosts in certain
locations like old, abandoned houses. One way to explain this fact is that such
people actually have seen or heard ghosts. Another way to explain this fact is
by a combination of known factors, such as:

 noises and movements due to drafts, old houses resettling, rats, owls,
bats, etc.,
 hypnagogic (half-dreaming) states
 selective noticing of spooky things
 confirmation bias more generally
 the effect of fear on false perception of other minds [6]
 altered states from mild carbon monoxide poisoning (in some old
buildings)
 ... and so on.

So we have two hypotheses to explain the evidence. Whichever one of our two
hypotheses is true, we would expect people to feel like they had ghostly
encounters. So how can we distinguish these hypotheses?

Not Occam's Razor


Well, it's often considered a fundamental principle of reasoning
that simpler theories are better and, indeed, more likely to be true. The
criterion of simplicity tells us that we should find claims less likely the more
complexity they add to our overall view of the world. (This sometimes goes by
the name of "Occam's Razor" because it cuts away complexity.) This isn't a
matter of how long it takes to state the theory, but of how much more
complicated the world needs to be in order for the theory to be accurate. For
example, it might seem like the theory that uses the known factors above is
more complicated than the ghost theory, because it appeals to many things
instead of just one. But we already know these things exist, and we already
understand how they work. So using them in our explanation doesn't add any
complexity at all to our theory of the world as a whole.
On the other hand, it's worth thinking about just how complex the typical
ghost hypothesis is. (It's also worth bearing in mind that half of Americans
believe it.) First, it requires us to accept a whole new category of things in the
world that don't interact normally with other things. But, in addition, we also
have to believe a lot of very specific claims about what they're like:

 they like to haunt dark and abandoned places


 they like to moan, knock, creak, and hoot but otherwise don't interact
with things much
 they can't be detected with any ordinary scientific instruments (except
maybe certain devices that generate static, which is convenient
because if you listen to random static hard enough and with your
expectations primed, your brain will pick out words.)

These are all additional bits of complexity in the theory, because there's no
reason why the spirits of the dead, if they existed, couldn't have behaved
completely differently. They might have just introduced themselves and
politely told us what they want. They might have sounded like ordinary
people, or chickadees, or kazoos—or nothing at all. And if their goal were to
frighten people, wouldn't it be even more frightening to manifest themselves
clearly on camera or in the middle of a crowded street, and leave perfectly
legible but terrifying notes? Also why do they like old run-down houses?
Nobody else likes those places.

Of course, it's not impossible that the spirits of the dead would have a thing
for imitating owls and loose window shutters, but the point is that all of these
facts add to the complexity of the theory. It requires not only that we accept
strange new entities, but that we accept an elaborate system of weirdly
specific facts about them. However implausible it is that the dead hang
around in spirit form, we should find this specific ghost theory to be less
plausible still.
As you'll recall from the previous section, this last point actually follows from
the logic of probability: a conjunction of two claims must always be less
probable than either claim on its own, unless one of them entails the other.

In general, we consider it a good principle to explain the


phenomena by the simplest hypothesis possible.
—Ptolemy's Almagest (2nd century AD)

In the case of the ghost hypothesis, we need to set aside our familiarity with a
particular way of thinking about ghosts and actually look at each part of the
theory independently. Imagine you really hadn't heard anything about ghosts
before, and then were asked to predict how the spirits of the dead would
behave, if there were any. You would presumably have been very unsure and
wouldn't have put a very high probability on any of the specific traits that the
standard theory attributes to ghosts. Given this, the laws of probability say
that the whole conjunction should have a very low prior probability. But
somehow, when that big conjunction is put forward as an explanation for
people's ghostly experiences, we mainly notice that it would explain the
evidence if it were true, and not how wildly implausible it is.

If we fail to notice that theories get less plausible the more conjunctive they
are, we can easily fall into ad hoc modification. For example, suppose we try
to devise a test to provide us with evidence that ghosts don't exist. Whatever
results we get, the ghost theory can just be made more complex to
accommodate that evidence. (Indeed, it turns out that ghosts conveniently
hate any kind of controlled experiment or even the presence of skeptics
like James Randi. So, naturally, they can't be detected in controlled
conditions.)

Decoupling, redux
At some point, you may have seen a flowchart of the scientific method that
presents a chronological order for the different steps that scientists
undertake: raising questions, forming hypotheses, making observations,
analyzing results, etc. But different flowcharts have these steps in different
orders. When they put "form a hypothesis" before "make observations," they
have in mind the kind of case where we can design a study that will test our
hypotheses. As we've seen, the best way to ensure that this happens is often
to run a prospective, randomized controlled trial.

But not all science works this way. Sometimes we can't help but form our
theories after getting some of our main observations. These are the situations
that people have in mind when their flowcharts put "make observations"
before "form a hypothesis". For example, the Big Bang theory, the theory of
evolution, the heliocentric theory of the solar system—these are all theories
that were formed in order to explain observations that had already been
made. (In our sense, "theory" just means any claim or group of claims whose
truth we don't directly observe, but would, if true, help explain some things
we do observe. A theory can certainly have overwhelming evidence in its
favor, like the theory that the Earth goes around the sun, and the theory that
smoking causes lung cancer.)

So how can we talk about our prior confidence in a hypothesis if we hadn't


even considered it before getting the evidence? The answer is that we have to
carefully assess the plausibility of the hypothesis independently of this
evidence, using the same criteria we would have used if we hadn't yet gotten
the evidence. For example, consider a detective assessing various theories
after being called to a crime scene. She gets the evidence first and only then
begins to craft theories to explain it.

But shouldn't it worry us that these theories were designed to explain the
evidence? Well, this is only a problem if we fail to decouple the probability of
the hypothesis from the strength of the evidence, as described in the Mindset
chapter. But failing to decouple can happen in either direction. As we have
seen, starting with a hypothesis that we find plausible can skew our
assessment of how strong our subsequent evidence is. On the other hand, if
we start by getting strong evidence and only then assess the hypothesis, we
might not notice its complexity or fail to think of all the alternative ways
things could have been (i.e. possibility freeze).

So although the theory of natural selection was created to explain Darwin's


observations of species variation, he could still update properly as long as he
decoupled these two elements:

 How probable is it that inherited trait variation and selective survival


over long enough periods would eventually lead to species variation?
 Supposing this does happen, how likely is it that we'd observe these
particular facts? And supposing this doesn't happen, how likely is it
that we'd observe these facts?

The key is to remember that a hypothesis can be implausible even if it would


do a good job of explaining our observations if it were true, as with the
complex ghost theory. But in fact, part of what made Darwin's theory so
powerful is that it explains so much with a very simple hypothesis: it doesn't
appeal to a new unexplained "force" of evolution, only to things we know
about anyway, such as mutations, survival pressure, selection effects, and
very long spans of time.

So, if decoupling is really what matters, and not the order of assessment, why
do some scientific fields insist that one's hypothesis must be formulated
before evaluating the evidence? One reason is that carefully specifying a
hypothesis allows us to devise the right kind of experiment to confirm or
disconfirm the hypothesis. Another reason has to do with a selection effect
that happens when you are effectively testing multiple hypotheses at the
same time. Recall from the chapter on Causes that some scientific fields have
a conventional threshold for how strong the evidence from an experiment
must be to be taken seriously. For example, one standard threshold is that
there must be less than a 5% chance of a correlation of this size in our sample
if there's no correlation in the larger population. And statistically, this means
we should expect that out of 20 results just barely reaching this threshold, one
of them is the result of bad luck.

Now suppose we test a drug by giving it to a treatment group and a control


group and just seeing what happens, rather than specifying a hypothesis in
advance. We check our sample and see if taking the drug correlates in the
sample with improvements in cholesterol, depression, weight, anxiety,
muscle aches, etc., for a total of 20 conditions. Suppose one of them looks
promising and we just check whether that correlation was sufficiently unlikely
to happen by chance. The problem is that we just effectively ran 20
experiments all together, so we should expect one of them to give us a false
positive result even if the drug does nothing! If we pretend this was our
hypothesis all along and don't reveal the fact that there was a huge selection
effect involved, we are misrepresenting the strength of the evidence. One way
to avoid this is to ensure that researchers specify their hypothesis in advance.

Chapter 9
Diminishing marginal utility: when something has diminishing marginal
utility, each additional unit provides less and less utility. For example, the
difference between having $0 and having $1000 is very stark, giving one
access to food and possibly shelter. So, that first $1,000 has a lot of marginal
utility. However, the difference between having $1,000,000 and $1,001,000 is
not so stark; one’s life probably wouldn’t change appreciably with that extra
$1,000. This demonstrates that money has diminishing marginal utility: the
value of additional units of money declines the more we've already acquired.
Endowment effect: when we think of something as belonging to us—a
possession—we value it more than if it's only potentially ours. For example,
whereas I might have only been willing to pay $3 for a mug before acquiring it,
upon acquiring it and thinking of it as mine, I value it more highly.

Expected monetary value: A measure of how much money we stand to gain


with different possible outcomes of a choice, weighted by how likely those
outcome are. To calculate the expected monetary value: for each outcome,
multiply its monetary value in dollars by its probability. Then, add up the
results.

Expected utility: A measure of how much of what we care about is achieved


by the different possible outcomes of a choice, weighted by how likely those
outcomes are. To calculate expected utility, we multiply the utility of each
possible outcome by the probability of each outcome if that choice is made.
Then, to get the expected utility of that choice, we add up the value of the
various possible outcome. (See util and utility.)

Honoring sunk costs: taking unrecoverable costs into account when


estimating an option's expected value. Sunk costs include any loss or effort
we've endured in the past, not just monetary costs. For example, it would be a
mistake to take into account how much you paid for a now-ruined shirt when
considering whether or not to discard it. Better to ask whether you would
acquire the ruined shirt for free if it wasn't previously yours. Sometimes when
honoring sunk costs we are trying to avoid feeling that a past decision wasn't
a good one. But whether a decision in the past was good is now
unchangeable: and anyway, a good decision is one that makes the best choice
given what you know and value at the time of the decision. You can't reach
back in time and "save" a decision if it was bad, nor can you do anything to
make it bad if it was good.

New vs. Old Risks: we tend to reliably deviate from rational decision-making
by assigning more value to the avoidance of new risks as compared to old
risks. This can be explained by how old risks seem to us to be more
manageable while new risks seem scarier. For example, people tend to worry
much more about physical harm from strangers (e.g., terrorists), even though
most murder victims are killed by someone they knew personally.

Outcome framing: the same outcome can be described as a loss or a gain,


and these different frames can reliably influence human decision-making. For
example, physicians will tend to avoid a procedure with a 10% mortality rate
but choose a procedure with a 90% survival rate, even though these two rates
are equivalent. The first description has a loss frame, which makes it seem
less desirable than the second description, which has a gain frame.

Possibility and certainty effects: we tend to reliably deviate from rational


decision-making by overvaluing the mere possibility of good outcomes and
the certainty of avoiding bad outcomes. For example, we tend to be willing to
pay more to go from a 0% to a 1% chance of winning a free vacation than we'd
pay to go from a 4% to 5% chance of winning a free vacation.

Sunk costs: see honoring sunk costs.

Temporal discounting: the tendency to value a given outcome less the


further into the future it is from the present moment; that is, to weigh the
costs and benefits of near outcomes more heavily than those of distant
outcomes. We often have to make decisions that involve a trade-off between
benefits now and costs later, or vice versa. A case of simple temporal
discounting would be that I value getting a donut in five minutes more than
getting a donut in a week. Sometimes this leads us to make different
assessments of what to do, depending on how close to the choice we are, or
whether it's in the past or future. In that case, our temporal discounting
is time-inconsistent.

Time-inconsistent discounting: when temporal discounting leads to


inconsistency over time about whether one choice is better than another. This
means that we predictably tend to disagree with our past and/or future selves
about what we should do or should have done, even with no new information.
For example, right now I would like my future self to refrain from eating a
donut when offered it. But I can predict that when my future self in a month is
faced with a donut, he will value eating the donut at that time more than the
later health benefits of not eating it. So, as the donut-time approaches, which
decision has a higher expected utility flips, and I change my mind about which
choice is better. And then after eating the donut, it flips again.

Utility & utils: a measure of how much of what we care about is achieved by
an outcome. There is no absolute scale for how to measure utility, but it is
important that we maintain the same proportions in our evaluations. So, if we
value an outcome, A, twice as much as an outcome, B, then we should assign
twice as much utility to B as to A. We can do this with utils, a "dummy" unit
allowing us to calculate the relative value of decisions. The absolute value of
utils used in a calculation is not important, but it's crucial that we preserve
the ratio of different values. For example, if we value going to the theater
twice as much as going to the park, then the value of utils we assign to going
to the theater should be twice as high as the value of utils we assign to going
to the park.

9.1 The logic of decisions


When we make a rational decision, we have to envision the various possible
outcomes of our choices. We also have to make two kinds of judgments about
those outcomes: we have to assess how good they are, and also
how probable they are. If a choice has a possible outcome that is both very
good and very unlikely, its goodness doesn't count very heavily in favor of the
choice. Together, our judgments about the goodness and probability of
outcomes provide us with at least some of our reasons for making one choice
over the other.
Decision theory is about the relationship between these reasons and the
choices they support. The structure of this relationship is called the logic of
decisions.

As you’ll recall, our primary question when discussing the logic of arguments
was: "Holding fixed the truth of these premises, how strongly do they support
the conclusion?" To answer that question, we set aside whether the premises
were actually true.

Similarly, when discussing the logic of a decision, we will ask: "Holding fixed
these judgments about the probability and goodness of outcomes, how
strongly do they support this choice?" This doesn’t mean we shouldn't think
carefully about the probability of possible outcomes, and how good or bad
they are. In this chapter, though, we'll focus on the question of what we
should do given a set of probabilities and value judgments about the possible
outcomes [1].

Possible outcomes
Every decision you make involves two or more choices—and each choice can
have multiple possible outcomes. We can represent this with a decision tree:

The black square is the way things are to begin with. Each circle represents a
choice we could make—maybe they are actions we can take, or things we can
say, or even the choice to do nothing at all. And the squares directly linked to
each circle represent possible outcomes of the choice. In this example, each
choice could lead to two possible outcomes, so we have four possible
outcomes in total. Note that we are treating all the outcomes as mutually
exclusive: even if one of the orange outcomes looks a lot like one of the blue
ones—say, in both we gain $1—they're at least differentiated by the fact that
we made the blue or orange choice, respectively.
Notice that for each choice, we can ask about the probability that an outcome
will occur if a particular choice is made. For example, here's one set of
probabilities we might assign to the various outcomes if the corresponding
choices are made.

Since we are assuming that the possible outcomes are mutually exclusive, the
probabilities of the outcomes for each choice should add up to one. Note that
we don’t assign probabilities to the choices themselves—that is, the circles—
because it’s up to us which one to pick. We're not trying to guess the
probability of our own actions; we're trying to decide which action is best.

Expected monetary value


Once we've considered all possible outcomes of a choice, and assessed their
goodness and probability, we need a way to combine those judgments.

Let's start with an easy kind of decision—a purely financial decision in which
the only concern is how much money you end up with. Each outcome has
a monetary value—the amount of money (in dollars, say) that we gain or lose
in that outcome. And we can weight this number by the outcome's
probability: simply multiply the two numbers together. This gives us each
outcome's weighted monetary value.

Of course, each choice in our example has more than one possible outcome.
But since these are already weighted by probability, we can simply add them
all up, giving us the expected monetary value of a choice—that is, the sum of
the weighted monetary values of all of its possible outcomes. Again, to find
the expected monetary value of a choice, follow two simple steps:

 For each possible outcome, multiply its monetary value in dollars by its
probability.
 Add up the results!
Let's try this on our simple decision between the blue choice and the orange
choice.

So, in this example, the top blue outcome has a weighted monetary value of
80 cents, while the bottom blue outcome has a weighted monetary value of 60
cents. Adding these together gives us the expected monetary value of the
choice represented by the blue circle—namely $1.40.

Meanwhile, the top orange outcome has a weighted monetary value of $3,
while the bottom orange outcome has a weighted monetary value of negative
$1. So these add up to $2, meaning that the orange choice has a significantly
higher expected monetary value than the blue choice.

So does this mean the orange choice is better? Obviously not, if the outcomes
involve more than just money. Maybe the blue outcomes strengthen a
friendship, whereas the orange outcomes hurt someone. But let's suppose
this isn't true, and the only difference in the outcomes is a monetary gain or
loss. (And, obviously, whatever effects that gain or loss has on our lives.) In
that case, it might seem that we should follow this rule:

 Pick the choice that has the highest expected monetary value.

But actually, this is a bad rule, even if our choice is just about money. The
reason for this is that money has a feature called diminishing marginal
utility, which we'll define in the next section.

Mo money, less marginal utility


Imagine you're a middle-aged person with about $100,000 in savings. Now ask
yourself if you'd be willing to play this game: a coin is tossed, and the results
are as follows:
if heads, you get an additional $100,000, plus $100;
if tails, you lose your original $100,000

The expected monetary value of playing the game is $50,050 minus $50,000—
i.e., $50. Meanwhile, refusing to play has an expected value of $0. So, if we
follow the rule that tells us we should pick the choice with the highest
expected monetary value, we should accept the terms of this game.

Would you play this game? Obviously not! If you lost all your money, you'd be
hungry and destitute—that's a huge change in the downward direction, as far
as your quality of life goes. In contrast, the upward movement from doubling
your money would not be nearly as large. In other words, the badness
of tails is far greater than the goodness of heads. That is, the benefit to
doubling your money would not be as great as the loss of losing it all.

If this is so, we can't measure the outcomes of the coin toss simply by using
weighted monetary value, even if the toss only has monetary outcomes. We
must consider the utility that money has to us— that is, how much of what we
care about a unit of money provides. But this depends largely on how many
dollars we already have! For someone with no assets at all, $100,000 makes an
enormous difference. For someone who already has $100,000, another
$100,000 makes much less difference. And for a multi-millionaire, an
additional $100,000 makes very little difference indeed.

I can understand about having millions of dollars... but once


you get much beyond that, I have to tell you, it's the same
hamburger.
—Bill Gates

What Gates means by "it's the same hamburger" is that your ability to enjoy
life does not just keep increasing at the same rate the more money you have.
The difference between a $30 hamburger and a $25 hamburger is much
smaller than the difference between a $5 hamburger and no hamburger at all.
After a certain point, apparently, there's not much remaining upside in the
hamburger department.

Likewise, the difference in benefit between being homeless and having a


$100,000 home is much larger than that between having a $2 million home
and having a $2.1 million home. As homes get more expensive, each
additional thousand dollars in cost makes less of a difference to the lives of
the people living in them.

In short, each additional dollar provides less of what we really care about than
the previous dollar did. This is what's meant by diminishing marginal
utility. You may be gaining more total utility with each dollar, but
the additional utility of the next (that is, "marginal") dollar goes down the
more dollars you have: it diminishes.

In fact, research on well-being and life satisfaction indicates that each


additional doubling of income per year provides about the same benefit, at
least for incomes under $100,000 per adult. (This is true both globally and in
the US.) This means that an additional dollar going to someone who is poor
provides much more well-being than an additional dollar going to someone
who is rich. The utility of income, measured in terms of well-being or life
satisfaction, looks like the graph below. (The graph is scaled for the US:
worldwide, most people live on much less than $10,000 per year [2].)

As you can see, an additional $10,000 adds far more to life satisfaction on the
left side of the graph than it does on the right. What about the marginal utility
of money for incomes higher than $100,000? The best evidence suggests that
little or no additional gain in well-being can be had from more income after
that point: the upward trend actually plateaus there, or just below [3].
For some types of goods, the utility of an additional quantity of
something even turns negative. Consider, for example, what happens as you
keep taking bites of cake:

At first, each bite of cake has positive value. But after a while, each additional
bite starts to reduce the overall goodness of the situation. This means that,
when we look at the marginal utility of each bite—i.e., the additional value of
the next bite—it drops below zero at some point.

The dotted line represents that moment when each additional bite of cake
just makes things worse.

The value of everything else


The utility of an outcome is simply shorthand for: how much of what we care
about is gained or lost in the outcome.

To fully grasp the idea of utility, consider how many different kinds of things
we care about. So far, we've considered the quality of life and happiness that
money can provide for us, and those are certainly things we care about. But
we care about many other things too. For instance, many of us care about
things like acquiring knowledge, telling the truth, and keeping promises—
even beyond the happiness that these things can provide. And we also care
about other people's well-being. This may be most obvious when thinking of
family and friends, but most of us also care at least a bit about the well-being
of people we don't know personally.

(A brief aside about motives. It's often assumed that, at a fundamental level,
we care only about ourselves, and we help others only as a way to show off or
make ourselves feel good. But the best and most recent science does not
support this idea. Instead, like at least some other animals, it seems we often
have genuinely altruistic motives. The fact that helping others can make us
feel good or increase our status—or even confer some evolutionary advantage
—doesn't mean that's our goal when we help others. On the contrary: in part,
at least, our aim is often simply to keep others from suffering [4].)

So: many things matter to us—they make outcomes better or worse by our
lights—and we often have to make decisions involving trade-offs among
them. The idea is that to compare our choices, we should start by comparing
the value of the outcomes, taking into account everything we care about. We
have to make judgments about how good we consider various outcomes to be
in comparison to each other— even when the outcomes involve very different
things that matter to us. For example, how much better or worse would it be
to spend this evening quietly sitting in a park than to spend it at a physics
lecture? We don't normally try to nail down comparisons like this, but the idea
behind the concept of utility is that we should be able to do so. (Otherwise, if
we can't decide which of two options is better, and by how much, the logic of
decisions may not have much to say that will help our decision—just as
deductive logic won't help us to assess a conclusion if we have no idea
whether the premises are true.)

When considering the utility of outcomes, we don't really need a unit of value
—all we need is the ability to compare them in relative terms. (For example,
we should be able to say how much better or worse one outcome is than
another—for example, "twice as bad," "equally good," "twice as good.") But in
making our comparisons of expected value, it can be helpful to use a kind of
"dummy" unit for utility. We call that unit the util.

So, for example, suppose that Bob's first dollar is 1000 times more valuable to
him than his millionth dollar. That is, he gets 1000 times the value in his life
from his first dollar than from his millionth. To represent these values in our
calculations, we can use any size units for our utils as long as they preserve
that 1000-to-1 ratio. So, we could say that his first dollar is worth 1000 utils,
and his millionth dollar is worth 1 util. Or we could say that his first dollar is
worth 30,000 utils, and his millionth dollar is worth 30 utils. Either way, our
comparisons of outcomes will remain the same, and so will our conclusions
about how Bob should act. What we're really representing is the comparative
utility between his first dollar and his millionth.

Before moving on, let me address some potential concerns about this
framework for thinking about decisions.

First, expressing what we care about in numerical units like this can sound
very cold and robotic, but remember that these are not real units, just
placeholders for a comparison of how good outcomes are. As long as we think
it makes sense to say that outcome A is twice as good as outcome B, for
example, then we can use utils to represent that kind of relative judgment. As
long as I can compare a holiday, a warm cup of soup, paying a certain price,
and so on, in terms of how much more valuable one is than another, then I
can think of the comparison in terms of utils.

Second, using utils might sound unrealistically specific. But, as with


probability, the util is just a tool to represent our fuzzy attitudes about things.
Normal people don't think of their probability judgments in terms of decimals
like "0.86", but they do say things like "I'm pretty sure that X," or "I suspect
that X," or "I doubt that X." Similarly, with utils—unless our outcomes involve
something easy to quantify, we might not really think that outcome A is
exactly 5/9ths as good as outcome B. But saying "B is a little more than twice
as good as A" doesn't seem like an unrealistic as a value judgment.

Third, this framework does not assume that the ends justify the means. After
all, the framework doesn't restrict how we place value on an outcome. So if
we care about not using immoral means, we will place less value on outcomes
that are achieved through immoral means. For example, if we value not
stealing, then any outcome in which we stole something has some additional
negative value added to it. If we would never steal under any circumstances
because we consider stealing that bad, we can represent this value judgment
by treating any outcome in which we steal as having infinite negative utility
for us.

Expected utility
Now instead of working with expected monetary value, we can now work
with expected utility. To compare choices, we simply: (i) assign utils and
probabilities to outcomes, (ii) multiply them together to obtain the weighted
utility of each outcome, and (iii) for each choice, add up the weighted utility of
possible outcomes. This gives us the expected utility of that choice.

More formally, suppose choice A has just two possible outcomes: O1 and O2. If
we use "U(O1)" to represent the utility of O1, then the expected utility of A is:

P(O1 | A) U( O1 ) + P(O2 | A ) U(O2)

On the other hand, if A has more than two possible outcomes we just add
them all to get the expected utility of A:

P(O1 | A) U(O1) + P(O2 | A) U(O2) + … + P(On | A) U(On)

We can now state one of the central ideas of decision theory:

The best choice is the one with the highest expected utility.

This rule tells us to make the choice with the best mix of outcomes that
achieve what we care about, and high probabilities for those outcomes. As
we'll see, thinking about decisions in terms of expected utility can be
extremely useful in revealing why some choices are better than others, and in
helping us to avoid decision-making errors. (Some philosophers and
economists have challenged the idea that this rule applies to literally every
decision, and even the idea that everything we care about in outcomes can be
compared [5]. However, even these thinkers will acknowledge that this
framework is useful for many ordinary decisions and for understanding the
errors in decision-making that we'll be discussing in the next section.)

As a simple example, thinking about our choices in terms of expected utility


sheds light on why we shouldn't play the coin-toss game with our $100,000,
even though it has positive expected monetary value. The reason not to play
is that, for any reasonable way of representing the diminishing marginal utility
of money, the game has extremely negative expected utility.

Or consider the fact that buying insurance is almost always a choice with
negative expected monetary value. Suppose I own a home outright, and I'm
considering flood insurance. After taking into account my home's
replacement value and location, the insurance company suggests I pay them
a monthly premium in exchange for which they'll replace my home if it's
destroyed in a flood.

Of course, the company's goal is to make money, so they would only offer me
this insurance if doing so had a positive expected monetary value for them. In
other words, the money they make from premiums if there's no flood,
multiplied by the probability that there's no flood, should be more than
they'd lose in payouts if there is a flood, multiplied by the probability that
there is a flood. But then isn't this a losing proposition for me? It means that
there's a negative expected monetary value for me in buying the insurance.
Am I simply betting on different odds than the insurance company is—betting
that a flood is more likely than they think? That would be a bad idea:
insurance companies are very good at estimating risks and potential damage.

In fact, even if I agree with the insurance company about the probabilities of
all of the outcomes, it is often both in the company's best interest to sell me
insurance, and also in my best interest to buy it. This becomes clear when we
start thinking in terms of expected utility rather than expected monetary
value. Because of the diminishing marginal utility of money, I am spending
money that is less valuable to me to protect against the possibility of losing
money that is more valuable to me.

If we imagine the value of my home as a stack of bills, I should be willing to


pay with the not-very-valuable dollars on the top of the stack in order to
protect the very-valuable dollars at the bottom. Meanwhile, the company has
such an enormous stack that even if they had to pay out the value of my home,
the value to the company of each remaining dollar in their stack wouldn't
change much. The company effectively has no diminishing marginal utility
and is just looking to maximize the monetary value of the sale. So the
transaction can have positive expected utility for us both—a win-win!

Likewise, for any form of insurance—usually the expected loss from paying
the premium is greater than the expected gain from a payout. But buying
insurance often has positive expected utility nonetheless.

9.2 Decision Pitfalls


Let's use our framework for thinking about decisions to help diagnose some
common types of irrational decision-making:

 outcome framing
 new vs. old risks
 endowment effect
 possibility & certainty effects
 honoring sunk costs
 time-inconsistent utilities

We'll consider these one at a time.

Outcome framing
A single decision can be presented in various ways, and you won't be
surprised to know that people's decisions can be reliably influenced by how
the outcomes are framed. For example, suppose you're a gameshow guest
participating in this game:

Gain game: You are ‘given’ $1,000. Then you must choose between:

A: $500 more, guaranteed


B: A 50% chance of gaining $1000 more

The expected monetary value of each option is $1500. But most people
choose A in this scenario: it's the "safe choice" and doesn't turn on a risky coin
toss. This choices focuses on what feels like potential gain.

But what about a slightly different game?

Loss game: You are 'given' $2000. Then you must choose between:

A: A guaranteed loss of $500


B: A 50% chance of losing $1000

Again, the expected monetary value of each option is $1500. But this time
most people don't choose A with its safe $1500—they strongly tend towards B,
the riskier choice. This game focuses on what feels like potential loss.

In both games, option B gives you a 50% chance of walking away with $1000
and a 50% chance of walking away with $2000. The only difference is
the outcome framing, which can heavily influence people's willingness to take
risk. System 1 represents the Gain Game as a choice between two possible
gains—good either way. But System 1 represents the Loss Game as a choice
between a guaranteed loss and a 50% chance at a loss. Taking a chance to
avoid any loss sounds good. System 1 doesn't fully adjust for the fact that the
potential "loss" in B is twice as large, or that it's actually all a gain in the
context of the game.
Here's another example. Suppose you're a doctor helping a patient choose a
treatment for lung cancer: surgery vs. radiation. Surgery has better long-term
outcomes if the patient survives the first month, but there's a chance of fatal
complications in the first month. Now consider these descriptions of that
chance:

Frame A: The one-month survival rate is 90%


Frame B: The mortality rate in the first month is 10%

When physicians were asked to decide between surgery vs. radiation using
Frame A, the vast majority recommended surgery. When they were presented
with Frame B, only about half recommended surgery. Frame A emphasizes the
survivors, so that physicians will naturally focus on that 90% and how they
have better long-term outcomes. But Frame B emphasizes the 10% who die in
the first month, and that seems like an unacceptable loss to many people.
Same choice, different frame.

There is some evidence that, in choices like this, a conflict arises between
Systems 1 and 2. Neuroscientists have performed functional imaging (fMRI) on
the brains of people making such choices. When people are strongly
influenced by the frame, the amygdala (a part of the brain associated with
emotion) is more likely to be active. When people resist being influenced by
the frame—like the physicians who recommended surgery even in Frame B—a
different area is likely to activate. That area, the anterior cingulate cortex, is
associated with self-control [6].

New vs. old risks


In our discussion of selection effects in Chapter 5, we noted that people are
far more concerned about certain kinds of risk than others. We focused in
particular on the fact that people worry about risks that press our instinctive
fear buttons—plane crashes, bombs, shootings. Ironically, even the fact that a
bad outcome is rare can cause it to be reported more widely on the news:
road accidents involving one or two vehicles are so common that no one
reports them, even though they kill tens of thousands of people every year.
But if a plane crashes, there's universal coverage of it on the news. This
increases many people's fear of flying, even though being a passenger on a
commercial airline is much safer than driving.

The logic of decision-making instructs us to cut through all of that and focus
with laser precision on two things: how good or bad the outcomes
are and how probable they are. This means resisting our emotional reactions
to certain risks, and actively taking into account the availability effect
involved with media coverage, since both of these factors can skew our
judgments about probability. But our irrationality extends beyond making
mistakes about probability. Even when explicitly presented with the
probability of a risk, people often make decisions irrationally.

For example, in several studies, subjects have been asked how much they'd
pay to eliminate an already present risk of N%, and how much they'd pay
to avoid taking on a risk of N%. Since the choice was being made about
whether to eliminate or avoid the risk going forward, decision theory tells us
that these options should have the same utility. But people are willing to pay
far more to avoid taking on a new risk than to eliminate an old risk, even when
they understand that the probability of something bad occurring is the same
in either case.

What is likely happening here is that System 1 is more comfortable with risks
that we've already survived—at least so far. These feel familiar and
manageable. But a new risk feels scarier, even if the probability and badness
of the outcome are the same. (This effect holds even if the risks in these
studies are fictional and the subjects are just pretending to be familiar with
the "existing risk".)
Because unfamiliar risks push our emotional buttons, they tend to attract
greater media coverage than familiar risks, creating a feedback loop of fear.
One need only turn on the news to see this effect at work. It's a new gang or
group of terrorists. It's a new disease, or a new chemical with an unfamiliar
name that we're all ingesting (even if the harm caused is estimated by experts
to be negligible or zero). Far less scary-seeming are familiar mass killers that
grind out hundreds of thousands of deaths each year: smoking, poor diet, lack
of exercise, particulate pollution [7]. It took a massive information campaign
to help Americans recognize the danger of smoking in spite of its normalcy
and familiarity. But putting another log on the familiar fire, how many of us
ever stop to wonder how harmful woodsmoke is [8]?

Our instinct is to treat the many deaths due to familiar causes as though they
were less terrible than the few deaths caused by scary new risks. It's possible
to assign utilities to outcomes in such a way that dying from cancer is less bad
for you than dying from a terrorist attack at the same age. But doing so strikes
me as deeply confused on reflection. After all—if you're killed by something
familiar, you still pretty much die. (Also, I'd take an instant death from a
terrorist attack or shooting in place of an agonizingly slow death from cancer
any day.)

The most dramatic example of a familiar risk that we have learned to accept is
the cause that kills more people than anything else: senescence, a.k.a. aging,
the primary risk factor for the leading causes of death worldwide. Senescence
is not the continuation of the process of development from childhood to
adulthood; it's the accumulation at a cellular level of damage that arises as a
side-effect of normal biological functions. That damage in turn causes the
collection of chronic diseases and syndromes known as "old age". Of course, a
risk is only worth taking into account if it could possibly be avoided, and we
can't yet do much to prevent or repair the damage of senescence. But
researchers have already dramatically extended the "healthspan" of several
animal species—i.e., the period of their lives before the onset of age-related
decline. And in many cases, we know that evolutionary processes have altered
the rate at which species age.

For these reasons, a growing number of scientists believe that radical


healthspan extension will eventually be possible for humans. Still, this idea
has generated very little public enthusiasm, especially in comparison to the
harm inflicted by the diseases of aging. Perhaps there are many reasons to be
skeptical about the project of extending human healthspan, or even to
oppose it morally. But very few people have given it any serious
consideration. Could this be because we're so familiar with all the suffering
and death caused by aging that we hardly notice it? Or perhaps because our
seemingly inevitable decline and demise is so terrifying that we try to give it a
positive spin or else avoid thinking about it entirely?

The endowment effect


When we think of something as belonging to us—a possession—we value it
more than if it is only potentially ours. This straightforward little bias is known
as the endowment effect.

A famous example involved a study in which one group of subjects was given
a free mug, and offered various amounts of money for it [9]. Another group
was allowed to examine the mugs but had to choose between the mug and
the money. The result was that the first group didn't want to part with their
mugs for less than an average of $7.12, whereas the second group chose the
money rather than the mug at an average offer of $3.12. In other words, the
value of the mug more than doubled in peoples' minds simply due to the idea
that they possessed it: it was their mug. (Their own... their precious...)

Maybe you have a shirt that you know, realistically, you will never wear again.
But even if it has no sentimental value, the mere fact that something is yours
can make letting it go more difficult.
A good heuristic here is to pretend it isn't already yours. Ask yourself: would I
take this if it were free at a garage sale? (Or, if you can fetch a price for it, ask
whether you'd pay that price now at a garage sale.)

Maybe you'd think—I don't have room for this old junk! In that case, you're
just subject to an endowment effect when you keep it.

The possibility and certainty effects


Another way in which we reliably deviate from rational decisions is by
overvaluing two things:

 the mere possibility of good outcomes; and


 the certainty of avoiding bad outcomes.

Suppose you face a very small chance of a very bad outcome. How much
would you pay to:

A: Reduce risk from .006 to .003?


B: Reduce risk from .001 to 0?

Decision theory will tell you that, however much negative utility that bad
outcome has, you should be willing to pay three times more for (A) than for
(B). But in studies, people are actually willing to pay more for (B) than for (A).
System 1 would very much like to eliminate the risk—to be certain that this
won't happen. Certainty is something System 1 understands—but it's not very
good with numbers, so those very low probabilities all pretty much feel the
same. This is the certainty effect.

The flip side of the coin is that we also over-emphasize the mere possibility of
a positive outcome. For instance, how much would you pay to increase the
chance of winning a free vacation...
A: ... from 4% to 5%?
B: ... from 0% to 1%?

Decision theory will tell you that these are equally valuable—at least as far as
the value of actually enjoying the vacation is concerned. (There is something
to be said for the idea that when people buy lottery tickets part of what
they're paying for is the fun of picturing themselves as a winner.) But, of
course, people are willing to pay more for (B). Again, System 1 understands
what it means for an outcome to go from being impossible to being possible.
The small shifts between probabilities, however, are much harder to grasp
intuitively. When this fact about our psychology influences our decision-
making, it's called the possibility effect.

The possibility and certainty effects work together to help explain why legal
settlements for frivolous claims are often surprisingly high. Such claims have
a low probability of success, but when they're successful they can sometimes
bring in very high amounts. The plaintiff wants to keep this high reward a
possibility, and thus may be reluctant to settle even for a reasonable sum.
Meanwhile, the defendant wants the certainty of avoiding that great loss and
will be eager to settle even for a higher-than-reasonable sum.

Honoring sunk costs


Imagine you paid $100 for a ticket to see a band in concert tonight. You can't
sell it at this point. Then you hear that the band has been putting on a pretty
bad show, and there's going to be a huge blizzard tonight. You really hate
driving in snow, and there's a risk of getting stuck or having an accident.

The negative outcomes are fairly bad and fairly likely, and outweigh the mild
enjoyment of the concert you're likely to have. So let's assume that sitting at
home actually has a higher expected utility in this case.
Still, many of us in a situation like this feel like we should go. Why? We have
the sense that we are otherwise losing money—throwing away $100. Notice,
though, that when you sketch a decision tree for the decision between going
to the concert or staying home, the $100 doesn't show up at all! That's
because you already spent it—it's a sunk cost and should make no difference
to your decision going forward. Nothing you do will change the fact that you
spent that $100. The only rational question remaining is whether the concert
is worth attending, given that you happen to have a ticket.

That feeling of wanting to make good on the spent $100 is called honoring
sunk costs—taking unrecoverable costs into account when estimating the
expected value of an option. (By "costs" I mean any kind of loss or effort we've
endured in the past, not just monetary costs.) It's not completely clear why we
feel the urge to honor sunk costs, but it likely has something to do
with wanting to have been right all along. If you stay home, there's a sense in
which you made the wrong decision when you bought the ticket—at least
given what you know now. If you go, then there's at least a chance that you'll
enjoy yourself and the blizzard will be fine. So there is a chance that you'll
look like you made all good decisions in retrospect.

Of course, this is silly. What makes a decision good at a given time depends on
what your reasonable degrees of belief are at that time. It has nothing to do
with whether the results of that decision happen to work out well by chance.
Whether or not it was the right decision to buy the ticket doesn't depend at all
on what you do now, since you had no way of knowing whether there was
going to be a blizzard. The way to avoid honoring sunk costs is to stop
thinking about our past decisions as though we can render them good or bad
in retrospect. Your past decisions were good if you picked the best
choice given what you knew and valued at the time of the decision. You can't
reach back in time and "save" a decision if it was bad, nor can you do anything
to make it bad if it was good.
A useful heuristic for overcoming this tendency to honor sunk costs is to ask
ourselves: What would I do if there had been no earlier costs? In the case of the
concert ticket, that would mean asking, "What would I do if I had just found
this ticket, or received it for free?" And the answer is : you'd stay home and
avoid the blizzard.

No matter how far you've gone down the wrong road, turn
back.
—Turkish proverb

In investing, the tendency to honor sunk costs is sometimes called "throwing


good money after bad." Suppose you made an investment in a stock and it fell
massively in price. You might be tempted to think: if I buy more now at this
cheaper price, then even if the stock only goes back to where it was before, I'll
have made a return on my investment! The appeal of this idea comes from the
appeal of somehow keeping the initial investment from having been a bad
idea. But again—that's silly. If it was a bad idea, it was a bad idea regardless of
what you do now. Apply the heuristic: pretend someone gave me this crappy
stock. Now I can decide whether to keep it, sell it, or buy more without feeling
like I need to honor my decision to buy it in the first place.

We can also find arguments in politics that assume it's rational to honor sunk
costs. For example, it would be embarrassing to abandon a project that has
already cost taxpayers greatly, even if it's found that the additional cost
would not be worth the benefit. The argument that it's worth continuing
because so much has been spent already makes sense only if you place great
value on the politicians' avoidance of embarrassment!

Sometimes when a company is on the wrong track, the previous CEO gets
fired. Even if that CEO knows what to do going forward, it can often be too
tempting for them to double-down on their old decisions in order to "save"
them. Best to have someone who can start from scratch and pull the plug on
any projects that are no longer worthwhile. But it's also possible to harness
this mindset without actually replacing yourself as the decision-maker.
Pretend at every point that you're starting anew—you're the new CEO of your
life, not trying to make good on old decisions, just trying to start where you
are now.

Time-inconsistent utilities
Our final decision pitfall explains many of the situations where we struggle
with ourselves about what we should do—working hard, eating right,
exercising, sleeping enough, etc. It also helps explain that bane of
contemporary life: procrastination. The pitfall I have in mind is simply that we
tend to value things differently at different times, in a way that leads us into
conflict with ourselves—our past and future selves.

Take a simple example. Right now, I feel like eating a donut. I'd certainly have
a hard time resisting if I were offered one. But if you asked me right now
whether I want a donut later today, I'd have a much easier time resisting. And
if you asked me right now whether I'd like a donut in a week, I'd say "no
thanks." In fact, I might ask you to please keep that donut away from my
future self. I'd rather skip the calories!

So what's going on? If we were to draw a graph for how much I now value
getting a donut at different times, it would look like this:

The graph above shows the value that I currently place on various possible
future events—my having a donut at various future times. The further into the
future these possible donuts get, the less I value them. This is called temporal
discounting—placing lower value on things the further into the future they
are.
We could also graph how I value the same event—say my having a donut at
noon on Friday—at various times as I approach it. Then the graph looks like
this:

The previous graph showed me fixed at a single time, thinking about donuts
that I might get at different points in the future. This graph shows me, at
different times, thinking about the same donut I might eat at noon on Friday.
As I get closer to potentially eating that maple-glazed deliciousness, I come to
value it more and more.

So what's the problem? Well, right now it's Monday and I can tell that I
disagree with my Friday self about the value of that donut. My present self and
my future self both value the more distant future benefits of avoiding junk
calories. But I can see that my Friday self will value the pleasure of eating a
donut on Friday more than the more distant future benefits of avoiding it. In
contrast, because my present self places very little value on Friday's donut,
that value is outweighed by the future benefits of not eating it. So if I can, I will
even take steps to avoid my future self having easy access to one.

To take another example, suppose I'm comparing the value of two things in
the future: lazing around on Friday, or having gone for a run on Friday. (Not
only do I feel good immediately after, but there will also be small but real
health benefits in the more distant future.) A week beforehand, I want my
future self to go for a run because I value the benefits of having gone for a run
more than I value the benefit of lazing around on Friday.

Monday: utility of Friday run > utility of Friday laziness

This situation is represented by the fact that the green line is higher than the
blue line on the left-hand side of this graph:
Unfortunately, as time passes and I move rightward on this graph, the benefit
of lazing around on Friday morning starts to feel more significant because
Friday morning gets closer.

Friday morning: utility of Friday laziness > utility of Friday run

I have now entered the fail zone, the area of the graph where one value curve
pops up and temporarily overrides the other. For a short time—basically just
while lazing around—I treat lazing around as more valuable than the benefits
of running, then kick myself afterwards because I once again value the
benefits of having gone for a run more than the benefits of having lazed
around. In short, the only time I value lazing around on Friday more than
running on Friday is the very time at which I actually get to decide what to do
on Friday.

When we place a great deal of weight on benefits at the present moment, we


end up discounting benefits in the future in such a way that we are
inconsistent across time. This means that we end up disagreeing with our past
and future selves [10]. This is called time-inconsistent discounting.

Another way to make this vivid is to test it with money. Suppose we ask
people:

Do you want $100 now or $120 in a month?

Most would take the $100 now, even though 20% per month is an amazing
rate of return. People have a discount curve for getting the money that drops
very quickly during the first month. (A 'discount curve' is a chart like the first
one about donuts above, showing the current utility of donuts at various
future times.) But suppose we ask them:

Do you want $100 in 12 months, or $120 in 13 months?


That 1-month difference doesn't feel so large any more, since it's far out in the
future along the horizontal slope of the utility curve. So they go for the $120 in
13 months.

This pair of decisions is problematic, because if I choose $120 in 13 months,


then in 12 months I will disagree with my previous decision—at that point I
suddenly decide that having the $100 after 12 months would have been a
better choice. After all, at that point, I'll prefer $100 immediately to $120 in a
month. The shape of our temporal discounting, revealed by our answers to
these two questions, is time-inconsistent.

Our many selves


It's sometimes useful to think of myself as comprised of various temporary
selves over time—time-slices of myself. Almost all my time slices want me to
go for a run and avoid donuts on Friday, except for the one that really matters,
which is the time-slice in charge of what to do on Friday. That, in a nutshell, is
the greatest impediment many of us face to more productive and successful
lives. We have met the enemy, and he is us: or rather, whatever time-slice of
us is in charge of what to do right now.

In a sense, all my time-slices form a community in which each time-slice gets


to be a temporary dictator for a short time. Generally, all the time-slices are
against eating donuts and in favor of exercising. But each time-slice also has a
soft spot for having donuts and lazing around at one particular time—namely,
the very time where he happens to be in charge. The result is that every other
time-slice looks on in horror as the dictatorship is continuously passed to the
one time-slice with the worst possible values to be in charge at that very
moment.

How could such a community ever successfully complete a long-term project?


The guy in charge never cares as much about the long-term as everyone else:
he's always the most short-sighted individual in the community. One strategy
for solving this kind of problem is for past time-slices to reach into the future
and influence the decisions of future time-slices. Here are four practical ways
to do that:

 restrictions
 costs
 resolutions
 rewards

Let's consider these in order.

First, if I know I might be tempted by a donut tomorrow, I can take steps now
to restrict my access to donuts tomorrow. My future self is likely to be
annoyed by this, but he doesn't know what's best for us! A famous example of
restricting one's future self comes from the Odyssey: when his ship is about to
pass the Sirens, Odysseus makes sure he won't jump overboard in pursuit of
their singing by having his men tie him to the mast, so he could hear their
song and still survive.

If I can't restrict my future self's access to donuts, I might still be able to add a
cost to eating them, which might be enough to dissuade him. For example, if I
know that there's one donut left and that I'll be tempted by it, I might promise
someone that I will save it for them. Then when my future self considers
eating it, he'll have to consider not only the health benefits of resisting the
donut (which, in the moment, are not enough to outweigh the temptation)
but also the additional cost of breaking that promise.

A similar strategy sometimes works using self-promises or resolutions. If I


make a very serious commitment to myself that I will get my homework done
a day early, my future self might think twice about violating that commitment
—at least if I convince myself that it matters to be the sort of person who
keeps resolutions! Breaking the resolution becomes a cost that might be large
enough to keep my future self from skipping the run. The strategy of racking
up "streaks" or "chains" of actions works in a similar manner. If I've
accumulated an unbroken streak of workouts, my future self will likely
consider it a cost to break that streak. Failing to run on this occasion becomes
a much bigger deal, potentially emblematic of my inability to get fit by
committing to a habit. Thinking of the chain in this way may be enough to
push my Friday morning self out the door. (In addition, he will know that
our even later selves will be angry with him for breaking the chain!)

Finally, if the long-term benefits of working out won't be enough to motivate


my future self, I may be able to add additional short-term rewards that will
sweeten the choice. For example, suppose I tend to be motivated by hanging
out with a friend of mine. Then I may be able to motivate my future self to get
out of bed by making a date to work out that morning with my friend. (This
might work as a potential cost as well as a reward, if I expect my friend to fault
me for not showing up!)

You might also like