100% found this document useful (7 votes)
3K views225 pages

Naked Statistics

In 'Naked Statistics', Charles Wheelan aims to make statistics accessible and engaging by focusing on intuition rather than complex mathematics. The book covers various statistical concepts, including descriptive statistics, correlation, probability, and regression analysis, using real-world examples to illustrate their relevance. Wheelan emphasizes the importance of understanding statistics to navigate data-driven decisions and avoid misleading conclusions.

Translated by

ScribdTranslations
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
100% found this document useful (7 votes)
3K views225 pages

Naked Statistics

In 'Naked Statistics', Charles Wheelan aims to make statistics accessible and engaging by focusing on intuition rather than complex mathematics. The book covers various statistical concepts, including descriptive statistics, correlation, probability, and regression analysis, using real-world examples to illustrate their relevance. Wheelan emphasizes the importance of understanding statistics to navigate data-driven decisions and avoid misleading conclusions.

Translated by

ScribdTranslations
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Machine Translated by

Google

naked statistics
STRIPPING THE DREAD FROM THE DATA

Charles Wheelan
BEST SEILING AUTHOR OF NAKED ECONOMICS
Machine Translated by
Google

naked statistics
Taking the fear out of data

CARLOS WHEELAN

W. W. Norton & CompanyNew


York I London
Machine Translated by
Google

Dedication
Machine Translated by
Google

For Katrina
Machine Translated by
Google

Content

Cover
Title page
Dedication

Introduction: Why I hated calculus but I love statistics

1 What's the point?

2 Descriptive Statistics: Who was the greatest baseball player of all time?
Appendix to Chapter 2

3 Misleading description: “He has a great personality!” and other true but wildly misleading
statements

4 Correlation: How does Netflix know which movies I like?


Appendix to Chapter 4

5 Basic probability: Don't buy the extended warranty for your $99 printer

5½ The Monty Hall Problem

Six problems with probability: How overconfident math geeks almost destroyed the global financial
system

7 The importance of data: “Garbage in, garbage out”

8 The central limit theorem: the Lebron James of statistics


Machine Translated by
Google

9 Inference: Why my statistics teacher thought I might have cheated


Appendix to Chapter 9

10 Polls: How we know 64 percent of Americans support


the death penalty (with a sampling error of ± 3 percent)
Appendix to Chapter 10

11 Regression Analysis: The Miracle Elixir


Appendix to Chapter 11

12 Common regression mistakes: the mandatory warning label

13 Program Evaluation: Will Going to Harvard Change Your Life?

Conclusion: Five questions that statistics can help answer


Appendix: Statistical software
Grades
Expressions of gratitude
Index

Copyright

Also by Charles Wheelan


Machine Translated by
Google

Introduction
Why I hated calculus but I love statistics

I have always had an uncomfortable relationship with mathematics. I don't like numbers for numbers'
sake. I'm not impressed by fancy formulas that have no real-world application. In particular, I disliked
high school calculus for the simple reason that no one bothered to tell me why I needed to learn it.
What is the area under a parabola? Who cares?
In fact, one of the greatest moments of my life occurred during my senior year of high school, at the end of the first
semester of Advanced Placement Calculus. I was working on the final exam, certainly less prepared for the exam than I
should have been. (I had been accepted into my first choice university a few weeks earlier, which had sapped what little
motivation I had for the course.) As I looked at the final exam questions, they seemed completely unfamiliar to me. I don't
mean to say that I had trouble answering the questions. I mean, I didn't even recognize what they were asking me. He
was no stranger to being unprepared for exams, but, to paraphrase Donald Rumsfeld, he usually knew what he didn't
know. This exam seemed even more Greek than usual. I flipped through the exam pages for a while and then more or
less gave up. I walked to the front of the classroom, where my calculus teacher, whom we will call Carol
Smith was supervising the exam. "Lady. Smith,” I said, “I don’t recognize a lot of the things on the test.”

Suffice it to say that Mrs. Smith didn't like me much more than she liked me. Yes, I can now admit that I sometimes
used my limited powers as student body president to schedule school-wide assemblies only to have Mrs. Smith's calculus
class cancelled. Yes, my friends and I received flowers for Mrs. Smith during class from a “secret admirer” just so we
could laugh in the back of the room while she looked around in embarrassment. And yes, I stopped doing homework
once I entered college.

So when I approached Ms. Smith in the middle of the exam and told her that the material seemed unfamiliar to me,
she was, well, unsympathetic. "Carlos," he said.

out loud, apparently to me, but facing the rows of desks to make sure the entire class could hear, "if you had
studied, the material would be much more familiar to you." This was a compelling point.
So I crept back to my desk. After a few minutes, Brian Arbetter, a much better calculus student than I,
walked to the front of the room and whispered a few things to Mrs. Smith. She answered him in a whisper and
then something truly extraordinary happened. "Class, I need your attention," Mrs. Smith announced. "It seems
Machine Translated by
Google

I gave you the second semester exam by mistake." We were far enough along in the testing period that the
entire exam had to be aborted and rescheduled.
I can't fully describe my euphoria. I would go on in life and marry a wonderful woman. We have three
healthy children. I have published books and visited places like the Taj Mahal and Angkor Wat. Still, the day
my calculus teacher got what she deserved is one of the five most important moments in my life. (The fact that
I almost failed the final make-up exam did not significantly diminish this wonderful life experience.)

The calculus exam incident tells you a lot of what you need to know about my relationship with
mathematics, but not everything. Oddly enough, I loved physics in high school, even though physics is largely
based on the same calculus I refused to do in Mrs. Smith's class. Because? Because physics has a clear
purpose. I distinctly remember my high school physics teacher showing us during the World Series how we
could use the basic acceleration formula to estimate how far a home run had been hit. That's great, and the
same formula has many more socially significant applications.
Once I got to college, I really enjoyed probability, again because it gave me insight into interesting real-life
situations. In retrospect, I now recognize that it wasn't the math that bothered me in calculus class; it was that
no one saw fit to explain the meaning of it. If you are not fascinated by the elegance of the formulas alone
(which I am emphatically not), then it is simply a bunch of tedious, mechanistic formulas, at least as I was
taught them.
This brings me to statistics (which, for the purposes of this book, includes probability). I love statistics.
Statistics can be used to explain everything from DNA testing to the idiocy of playing the lottery. Statistics can
help us identify factors associated with diseases such as cancer and heart disease; it can help us detect
cheating on standardized tests. Statistics can even help you win at game shows. There was a famous show
during my childhood called Let's Make a Deal, with its equally famous host, Monty Hall. At the end of each
day's show, a successful player would stand with Monty in front of three large doors: Door No. 1, door no. 2,
and Door No. 3. Monty Hall explained to the player that there was a very desirable prize behind one of the
doors (something like a new car) and a
goat behind the other two. The idea was simple: the player chose one of the doors and placed the contents behind that
door.
As each player stood in front of the doors with Monty Hall, they had a 1 in 3 chance of choosing which door would
open to reveal the valuable prize.
But Let's Make a Deal had a twist that has delighted statisticians ever since (and perplexed everyone else). After the
player chose a door, Monty Hall would open one of the two remaining doors, always revealing a goat. As an example,
let's suppose that the player has chosen door no. 1. Monty would then open door no. 3; the live goat would be standing
there on the stage. There would still be two closed doors, numbers. 1 and 2. If the valuable prize were behind the no. 1,
the contestant would win; if he was behind no. 2, would lose. But then things got more interesting: Monty would turn to
the player and ask if he would like to change his mind and switch doors (from number 1 to number 2 in this case).
Remember, both doors were still closed and the only new information the contestant received was that a goat appeared
behind one of the doors he did not choose.

Should I change?
Machine Translated by
Google

The answer is yes. Because? That's in Chapter 5½.

The paradox of statistics is that it is everywhere (from batting averages to presidential polls), but the discipline itself has a
reputation for being uninteresting and inaccessible. Many statistics books and classes are too loaded with math and
jargon. Believe me, the technical details are crucial (and interesting), but it's Greek if you don't understand the intuition.
And you may not even care about intuition if you're not convinced that there's some reason to learn it. Each chapter of
this book promises to answer the basic question I asked (unsuccessfully) my high school calculus teacher: What's the
point of this?

This book is about intuition. It lacks math, equations, and graphs; when they are used, I promise they will serve a
clear and enlightening purpose. Meanwhile, the book contains plenty of examples to convince you that there are great
reasons to learn these things. Statistics can be really interesting and most of it isn't that difficult.

The idea for this book was born not long after my unfortunate experience in Mrs. Smith's AP Calculus class. I went to
graduate school to study economics and public policy. Even before the program began, I was assigned (unsurprisingly) to
a “math camp” along with most of my classmates to prepare us for the quantitative rigors that would follow. For three
weeks, we learned math all day in a windowless basement classroom (really).
One of those days, I had something very close to a professional epiphany. Our

The instructor was trying to teach us the circumstances under which the sum of an infinite series converges to
a finite number. Stay with me here for a minute because this concept will become clear. (Right now you're
probably feeling like I did in that windowless classroom.) An infinite series is a pattern of numbers that
continues forever, such as 1 + ½ + ¼ + ⅛ The three dots mean the .pa. tron continues to infinity. .
This is the part we were having trouble understanding. Our instructor was trying to convince us, using
some proof I have long forgotten, that a series of numbers can go on forever and still add up to
(approximately) a finite number. One of my classmates, Will Warshauer, wanted nothing to do with it, despite
the impressive mathematical demonstration. (To be honest, I was a little skeptical myself.) How can something
that is infinite add up to something that is finite?
Then I had an inspiration, or more accurately, an intuition of what the instructor was trying to explain. I
turned to Will and told him what I had just worked out in my head. Imagine that you have placed yourself
exactly 2 feet from a wall.

Now move half the distance to that wall (1 foot), so that you are standing 1 foot away.
From 1 foot away, move again half the distance to the wall (6 inches or ½ foot). And from 6 inches away, do it again
(move 3 inches or ¼ foot). Then do it again (move 1½ inches or ⅛ of a foot). Etc.
Little by little you will get quite close to the wall. (For example, when you are 1/1024 of an inch from the
wall, you will move half the distance, or another 1/2048 of an inch.) But you will never hit the wall, because by
definition each move will take you only half the remaining distance. In other words, you will get infinitely closer
to the wall but never hit it. If we measure your movements in feet, the series can be described as 1 + ½ + ¼ +
Machine Translated by
Google

⅛ ...
Therein lies the idea: although you will continue moving forever (each move will take you half the remaining
distance to the wall), the total distance you travel can never be more than 2 feet, which is your initial distance
from the wall. For mathematical purposes, the total distance traveled can be approximated as 2 feet, which is
very useful for calculation purposes. A mathematician would say that the sum of this infinite series is 1 foot +
½ foot + ¼ foot + ⅛ foot. . . converges at 2 feet, which is what our instructor was trying to teach us that day.

The thing is, I convinced Will. I convinced myself. I don't remember the math that shows that the sum of an
infinite series can converge to a finite number, but I can always look it up online. And when it does, it will
probably make sense. In my experience, intuition makes math and other technical details easier.

more understandable, but not necessarily the other way around.


The goal of this book is to make the most important statistical concepts more intuitive and accessible, not only for
those of us forced to study them in windowless classrooms, but for anyone interested in the extraordinary power of
numbers and data.

Now, having argued that the basic tools of statistics are less intuitive and accessible than they should be, I am going to
point out a seemingly contradictory point: statistics can be too accessible in the sense that anyone with data and a
computer can perform sophisticated statistical procedures with a few keystrokes. The problem is that if the data are poor
or if statistical techniques are used incorrectly, the conclusions can be extremely misleading and even potentially
dangerous. Consider the following hypothetical news story from the Internet: People who take short breaks at work are
much more likely to die from cancer. Imagine that headline popping up as you browse the web.
According to a seemingly impressive study of 36,000 office workers (a huge data set!), those workers who reported
leaving their offices to take regular ten-minute breaks during the workday were 41 percent more likely to develop cancer
over the next five years than workers who did not report leaving their offices for regular ten-minute breaks during the
workday. who do not leave their offices during the working day. Clearly we need to act on these kinds of findings,
perhaps some kind of national awareness campaign to avoid short-term work interruptions.
Or maybe we just need to think more clearly about what many workers do during that ten-minute break. My
professional experience suggests that many of those workers who report leaving their offices for short breaks are
huddled outside the building entrance smoking cigarettes (creating a haze of smoke through which the rest of us have to
walk to enter or exit). . Furthermore, I would infer that it is probably the cigarettes, not the short breaks at work, that are
causing the cancer. I made up this example just to make it particularly absurd, but I can assure you that many real-life
statistical abominations are almost as absurd once deconstructed.

Statistics are like a high-caliber weapon: useful when used correctly and potentially disastrous in the wrong hands.
This book won't make you an expert in statistics; it will teach you enough care and respect for the field not to do the
statistical equivalent of blowing someone's mind.
This is not a textbook, which is liberating in terms of the topics that need to be covered and the ways in which they
Machine Translated by
Google

can be explained. The book has been designed to present the most relevant statistical concepts for everyday life. How do
scientists come to the conclusion that something causes cancer? How do surveys work (and what can go wrong)? Who
“lies with statistics” and how

do they do it? How does your credit card company use data about what you're buying to predict whether
you're likely to default on a payment? (Seriously, they can do that.)

If you want to understand the numbers behind the news and appreciate the extraordinary (and growing)
power of data, here's what you need to know. In the end, I hope to persuade you of the observation first made
by the Swedish mathematician and writer Andrejs Dunkels: it is easy to lie with statistics, but it is difficult to tell
the truth without them.

But I have even bolder aspirations than that. I think you might enjoy the statistics. The underlying ideas are
fabulously interesting and relevant. The key is to separate the important ideas from the arcane technical
details that can get in the way. Those are the bare statistics.
Machine Translated by
Google

CHAPTER 1

What's the point?

I have noticed a curious phenomenon. Students will complain that statistics are confusing and irrelevant. Then the same
students will leave the classroom and talk happily at lunch about batting averages (during the summer) or wind chill factor
(during the winter) or grade point averages (always). They will acknowledge that the National Football League's "passer
rating" (a statistic that condenses a quarterback's performance into a single number) is a somewhat flawed and arbitrary
measure of a quarterback's game-day performance. The same data (completion rate, average yards per pass attempt,
percentage of touchdown passes per pass attempt, and interception rate) could be combined in a different way, such as
giving greater or lesser weight to any of those inputs, to generate different information. but an equally credible measure of
performance. However, anyone who has watched football recognizes that it is useful to have a single number that can be
used to summarize a quarterback's performance.
Is the quarterback's rating perfect? No. Statistics rarely offer a single "right" way to do something. Does it provide
meaningful information in an easily accessible way? Absolutely. It's a good tool to make a quick comparison between the
performance of two quarterbacks on a given day. I am a fan of the Chicago Bears. During the 2011 playoffs, the Bears
played the Packers; the Packers won. There are many ways I could describe that game, including pages and pages of
analysis and raw data. But here is a more succinct analysis. Chicago Bears quarterback Jay Cutler had a passer rating of
31.8. In contrast, Green Bay quarterback Aaron Rodgers had a passer rating of 55.4. Similarly, we can compare Jay
Cutler's performance to an earlier season game against Green Bay, when he had a passer rating of 85.6. That tells you a
lot of what you need to know to understand why the Bears beat the Packers earlier in the season but lost to them in the
playoffs.
That's a very useful synopsis of what happened on the field. Does it make things simpler? Yes, that is both the
strength and the weakness of any descriptive statistic. One number tells you Jay Cutler was outgunned by Aaron Rodgers
in the Bears' playoff loss. On the other hand, that number won't tell you if a quarterback had a bad play, such as throwing
a perfect pass that was batted away by the receiver and then intercepted, or if he "advanced" in a certain key. plays
(since every completion carries equal weight, whether it's a crucial third down or a meaningless play at the end
of the game), or if the defense was terrible. Etc.

The funny thing is that the same people who are perfectly comfortable talking about statistics in the context
of sports, the weather, or grades will freeze with anxiety when a researcher starts explaining something like the
Gini index, which is a standard tool in economics for measuring income inequality. I'll explain what the Gini
index is in a moment, but for now the most important thing to recognize is that the Gini index is like the passer
index. It is a useful tool for reducing complex information into a single number. As such, it has the strengths of
most descriptive statistics, namely that it provides an easy way to compare income distribution in two countries,
or in a single country at different times.
Machine Translated by
Google

The Gini index measures how equitably wealth (or income) is shared within a country on a scale from zero
to one. The statistic can be calculated for wealth or for annual income, and can be calculated at the individual
level or at the household level. (All of these statistics will be highly correlated but not identical.) The Gini index,
like the passer rating, has no intrinsic meaning; it is a comparison tool. A country in which all households had
the same wealth would have a Gini index of zero. In contrast, a country in which a single household owned all
the country's wealth would have a Gini index of one. As you can probably guess, the closer a country is to one,
the more unequal its distribution of wealth. The United States has a Gini index of 0.45, according to the Central
Intelligence Agency (a great collector of statistics, by the way).
1
1
So what?

Once that number is put into context, it can tell us a lot. For example, Sweden has a Gini index of 0.23.
Canada's is .32. China's is .42. Brazil's is .54. He *
South Africa's is .65. By looking at those numbers, we get a sense of where the United States falls relative to the
rest of the world when it comes to income inequality. Also
we can compare different moments in time. The Gini index for the United States was 0.41 in 1997 and rose to
0.45 over the next decade. (The most recent CIA data is from 2007.) This tells us objectively that while the
United States became richer during that time period, the distribution of wealth became more unequal.
Again, we can compare changes in the Gini index across countries over roughly the same time period.
Inequality in Canada remained broadly unchanged over the same period. Sweden has had significant
economic growth over the past two decades, but the Gini index in Sweden actually fell from 0.25 in 1992 to
0.23 in 2005, meaning that Sweden became richer and more equal over that
Machine Translated by
Google

period.
Is the Gini index the perfect measure of inequality? Absolutely not, just as passer rating is not a perfect measure of
quarterback performance. But it certainly gives us valuable information about a socially significant phenomenon in a convenient
format.

We have also slowly backtracked on our way to answering the question posed in the chapter title: What's the point? The
point is that statistics helps us process data, which is really just a fancy name for information. Sometimes data is trivial in the
grand scheme of things, as is the case with sports statistics. Sometimes they offer insights into the nature of human existence,
as is the case with the Gini index.
But, as any good infomercial would point out, that's not all! Hal Varian, Google's chief economist, told the New York Times
that being a statistician will be 2. I'll be the first to admit that we economists have a distorted definition of "sexy." will be “the sexy job” for the
next decade. Sometimes Still, consider the following disparate questions: How can we detect schools that cheat on their tests?
standardized?
How does Netflix know what kind of movies you like?
How can we know which substances or behaviors cause cancer, given that
Can't we perform experiments that cause cancer in humans?
Does praying for surgical patients improve their outcomes?
Is there really an economic benefit to earning a degree from a highly selective college or university?
What is causing the increasing incidence of autism?
Statistics can help answer these questions (or, we hope, can soon).
The world is producing more and more data, faster and faster. However, as noted by the
New York Times, “data is simply the raw material of knowledge.” the most powerful tool 3* The statistics are
we have to use data for some meaningful purpose, whether it's identifying underrated
baseball players or paying teachers more fairly. Here's a quick tour of how statistics can bring meaning to raw data.

Description and comparison A score


bowling is a descriptive statistic. So is batting average. Most American sports fans over the age of five are already familiar with
the field of descriptive statistics. We use numbers, in sports and everywhere else in life, to summarize information. How good a
baseball player was Mickey Mantle? He was a .298 career hitter. For a baseball fan, that's a significant statement, which is
remarkable.
4
when you think about it, because it sums up an eighteen-season career.
(I guess there's something slightly depressing about having the job of a lifetime.)

(It collapsed into a single number.) Of course, baseball fans have also come to recognize that descriptive
statistics other than batting average can better summarize a player's value on the field.
We evaluate the academic performance of high school and college students using a grade point average or
GPA. A letter grade is assigned a point value; typically an A is worth 4 points, a B is worth 3, a C is worth 2,
and so on. Upon graduation, when high school students apply for college and college students look for jobs,
the grade point average is a useful tool to assess their academic potential. Someone who has a GPA of 3.7 is
clearly a better student than someone from the same school with a GPA of 2.5. That makes it a good
descriptive statistic. It is easy to calculate, easy to understand and easy to compare between students.
But it's not perfect. GPA does not reflect the difficulty of courses that different students have taken. How do
Machine Translated by
Google

we compare a student with a 3.4 GPA in classes that seem relatively easy and a student with a 2.9 GPA who
has taken calculus, physics, and other difficult subjects? I went to a high school that tried to solve this problem
by placing greater emphasis on difficult classes, so that an A in an “honors” class was worth five points instead
of the usual four. This caused its own problems. My mother quickly recognized the distortion caused by this
GPA “solution.” For a student who takes a lot of honors classes (me), any A in a non-honors course, like gym
or health education, would actually lower my GPA, even though it is impossible to get better than an A in those
classes. As a result, my parents forbade me from taking driver's education in high school, lest even a perfect
performance diminish my chances of getting into a competitive college and writing popular books. Instead, they
paid to send me to a private driving school, at night during the summer.

Was that crazy? Yeah. But a theme of this book will be that overreliance on any one descriptive statistic can
lead to misleading conclusions or result in undesirable behavior. My original draft of that sentence used the
phrase "oversimplified descriptive statistics," but I deleted the word "oversimplified" because it is redundant.
Descriptive statistics exist to simplify, which always entails some loss of nuance or detail. Anyone who works
with numbers must recognize this.

Inference
How many homeless people live on the streets of Chicago? How often do married people have sex? These
may seem like wildly different types of questions; in fact, both can be answered (not perfectly) by using basic statistical
tools. A key function of statistics is to use the data we have to make informed guesses about broader issues for which we
do not have complete information. In short, we can use data from the "known world" to make informed inferences about
the "unknown world."
Let's start with the issue of homelessness. It is expensive and logistically difficult to count the homeless population in a
large metropolitan area. However, it is important to have a numerical estimate of this population for the purposes of
providing social services, obtaining eligibility for state and federal income, and obtaining representation in Congress. An
important statistical practice is sampling, which is the process of collecting data for a small area—say, a handful of census
tracts—and then using that data to make an informed judgment, or inference, about the homeless population of the city as
a whole. whole. Sampling requires far fewer resources than trying to count an entire population; if done correctly, it can be
just as accurate.

A political survey is a form of sampling. A research organization will attempt to contact a sample of households that
are broadly representative of the general population and ask them their opinions on a particular issue or candidate.
Obviously, this is much cheaper and faster than trying to contact every household in an entire state or country. Polling and
research firm Gallup estimates that a methodologically sound survey of 1,000 households will produce roughly the same
results as a survey that attempted to contact every household in the United States.
Here's how we find out how often Americans are having sex, with whom, and what kind. In the mid-1990s, the National
Opinion Research Center at the University of Chicago conducted a remarkably ambitious study of American sexual
behavior. The results were based on detailed in-person surveys conducted with a large, representative sample of
American adults. If you read on, Chapter 10 will tell you what they learned. How many other statistics books can promise
Machine Translated by
Google

you that?

Risk assessment and other probability-related events


Casinos make money in the long run, always. That doesn't mean they are making money at any given time. When the
bells and whistles ring, some high roller has just won thousands of dollars. The entire gambling industry is based on
games of chance, meaning that the outcome of any particular roll of the dice or card is uncertain. At the same time, the
underlying probabilities of the relevant events (hitting 21 in blackjack or spinning red in roulette) are known. When the
underlying odds favor the casinos (as they always do), we can be increasingly confident that the “house” will come
out ahead as the number of bets placed grows larger and larger, even as those bells and whistles continue to
ring.
This turns out to be a powerful phenomenon in areas of life far beyond casinos. Many companies must
assess the risks associated with a variety of adverse outcomes. They can't make those risks go away
completely, just as a casino can't guarantee that you won't win every hand of blackjack you play. However, any
company facing uncertainty can manage these risks by designing processes so that the probability of an
adverse outcome, from an environmental catastrophe to a defective product, is acceptably low. Wall Street
firms often assess the risks posed by their portfolios under different scenarios, weighting each of those
scenarios based on its likelihood. The 2008 financial crisis was precipitated in part by a series of market events
that had been considered extremely improbable, such as all the players in a casino playing blackjack all night.
Later in the book I will argue that these Wall Street models were flawed and that the data they used to assess
underlying risks was too limited, but the important point here is that any model for addressing risk must be
based on probability.
When individuals and businesses cannot make unacceptable risks go away, they seek protection in other
ways. The entire insurance industry is based on charging customers to protect them against some adverse
outcome, such as a car accident or a house fire. The insurance industry does not make money by eliminating
these events; cars crash and houses burn down every day. Sometimes cars even crash into houses, causing
them to burn down. Instead, the insurance industry makes money by charging premiums that are more than
enough to pay expected payouts for car accidents and home fires. (The insurance company may also try to
reduce your expected payouts by encouraging safe driving, fencing around pools, installing smoke detectors in
every bedroom, etc.)
Probability can even be used to detect cheating in some situations. Caveon Test Security specializes in
what it describes as "data forensics" for 5
a former test developer finds For example, the company (which was founded by for the SAT) will score tests at
a school or testing site in which the number of identical wrong answers is highly unlikely, typically a pattern that
would occur by chance less than once in a million. Mathematical logic arises from the fact that we cannot learn
much when a large group of students answers a question correctly. That's what they're supposed to do; they
could be cheating or they could be clever. But when those same test takers get an answer wrong, not everyone
should always get the same answer wrong. If they do, it suggests that they are copying each other (or
share responses via text). The company also looks for tests in which the test taker scores significantly better on
difficult questions than on easy questions (suggesting that the test taker knew the answers in advance) and
Machine Translated by
Google

tests in which the number of “bad to right” cross-outs is significantly greater than the number of “good to wrong”
cross-outs (suggesting that a teacher or administrator changed the answer sheets after the test).
Of course, you can see the limitations of using probability. A large group of test takers could have the same
incorrect answers by coincidence; in fact, the more schools we test, the more likely we are to observe such
patterns simply by chance. A statistical anomaly does not prove that a crime has been committed. Delma
Kinney, a fifty-year-old Atlanta man, won $1 million in a 6-chance lottery game in 2008 and then another $1
million in an instant-win game in 2011. The chance of that happening to the same person is in the range of 1 in
25 trillion. We can't arrest Mr. Kinney for fraud based on that calculation alone (although we might ask if he has
any relatives who work for the state lottery). Probability is a weapon in an arsenal that requires good judgment.

Identifying important relationships


(statistical detective work)
Does smoking cigarettes cause cancer? We have an answer to that question, but the process to answer it was
not as simple as you might think. The scientific method dictates that if we are testing a scientific hypothesis, we
should conduct a controlled experiment in which the variable of interest (e.g., smoking) is the only thing that
differs between the experimental group and the control group. If we observe a marked difference in some
outcome between the two groups (for example, lung cancer), we can safely infer that the variable of interest is
what caused that outcome. We can't do that kind of experiment on humans. If our working hypothesis is that
smoking causes cancer, it would be unethical to assign recent college graduates to two groups, smokers and
nonsmokers, and then see who gets cancer at the twentieth reunion. (We may conduct controlled experiments
on humans when our hypothesis is that a new drug or treatment may improve their health; we may not
deliberately expose human subjects when we expect an adverse outcome.)
*
Now, I might point out that it is not necessary to conduct an ethically dubious experiment to observe the
effects of smoking. Couldn't we just skip all this fancy methodology and compare cancer rates at the 20th
reunion between those who have smoked since graduating and those who haven't?
No. Smokers and nonsmokers are likely to be different in ways other than their smoking behavior. For example,
smokers are more likely to engage in other habits, such as excessive drinking or poor eating, that lead to adverse health
outcomes. If smokers are particularly sick at the 20th meeting, we would not know whether to attribute this result to
smoking or to other harmful things that many smokers do. We would also have a serious problem with the data on which
we base our analysis. Smokers who have become seriously ill with cancer are less likely to attend the 20th meeting.
(Dead smokers definitely won't show up.) As a result, any analysis of the health of the 20th reunion attendees (smoking-
related or anything else) will be seriously flawed by the fact that the healthiest members of the class are the ones most
likely to show up. The further the class is from graduation, say a fortieth or fiftieth reunion, the more severe this bias
becomes.
We cannot treat humans like laboratory rats. As a result, statistics read a lot like good detective work. Data provides
clues and patterns that can ultimately lead to meaningful conclusions. You’ve probably seen one of those jaw-dropping
Machine Translated by
Google

police procedural shows like CSI: NY, where highly attractive detectives and forensic experts pore over clues (DNA from a
cigarette butt, teeth marks on an apple, a single fiber from a car’s floor mat). and then use the evidence to catch a violent
criminal. The appeal of the show is that these experts don't have the conventional evidence used to find the bad guy, like
an eyewitness or a surveillance video tape. So instead they turn to scientific inference. Statistics basically do the same
thing. The data presents disorganized clues: the crime scene. Statistical analysis is the detective work that transforms raw
data into a meaningful conclusion.
After Chapter 11, you'll appreciate the TV show I hope to host: CSI: Regression Analysis, which would be just a slight
departure from those other action-packed police procedurals. Regression analysis is the tool that allows researchers to
isolate a relationship between two variables, such as smoking and cancer, while holding constant (or “controlling”) the
effects of other important variables, such as diet, exercise, weight, etc. When you read in the newspaper that eating a bran
muffin every day will reduce your chances of getting colon cancer, you need not fear that some unfortunate group of
human experimental subjects have been force-fed bran muffins in the basement of a federal laboratory somewhere while
the control group in the next building is given bacon and eggs. Instead, the researchers will collect detailed information on
thousands of people, including how often they eat bran muffins, and then use regression analysis to do two crucial things:
(1) quantify the observed association between eating bran muffins and getting colon cancer (e.g., a
(hypothetical finding that people who eat bran muffins have a 9 percent lower incidence of colon cancer,
controlling for other factors that may affect disease incidence); and (2) to quantify the likelihood that the
association between bran muffins and a lower rate of colon cancer observed in this study is simply a
coincidence (a quirk in the data from this sample of people) rather than a meaningful insight into the
relationship between diet and health.
Of course, CSI: Regression Analysis will star actors and actresses who are far more attractive than the
academics who normally pore over that data. These beauties (all of whom would have PhDs, despite being
only twenty-three years old) would study large data sets and use the latest statistical tools to answer important
societal questions: What are the most effective tools for combating violent crime? Who are most likely to
become terrorists? Later in the book we will discuss the concept of a “statistically significant” finding, meaning
that the analysis has uncovered an association between two variables that is unlikely to be due to chance
alone. For academic researchers, this type of statistical finding is “smoking gun.” In CSI: Regression Analysis, I
imagine a researcher working late into the night in the computer lab because of her daytime commitment as a
member of the United States Olympic beach volleyball team. When you get the printout of your statistical
analysis, you see exactly what you were looking for: a large, statistically significant relationship in your data set
between some variable that you had hypothesized might be important and the onset of autism. You must share
this trailer immediately!
The researcher takes the printout and runs down the hall, slowed down a little by the fact that she is
wearing high heels and a relatively small, tight black skirt.
She finds her male companion, who is inexplicably fit and tanned for a guy who works fourteen hours a day in a
basement computer lab, and shows him the results. He runs his fingers along his neatly trimmed goatee, grabs
his 9mm Glock pistol from his desk drawer and slides it into the holster beneath his $5,000 Hugo Boss suit
(also inexplicable given his starting academic salary of $38,000 a year). Together, the regression analysis
experts walk briskly to see their boss, a grizzled veteran who has overcome failed relationships and a drinking
Machine Translated by
Google

problem. ..
Well, you don't have to watch the TV drama to appreciate the importance of this kind of statistical research.
Almost every societal challenge we care about has been based on the systematic analysis of large data sets.
(In many cases, the collection of relevant data, which is costly and time-consuming, plays a crucial role in this
process, as will be explained in Chapter 7.)
I may have embellished my characters in CSI: Regression Analysis, but not the kinds of important questions
they might examine. There is an academic literature.
on terrorists and suicide bombers, a topic that would be difficult to study using human subjects (or lab rats, for that matter).
One of those books, What Makes a Terrorist, was written by one of my statistics professors in graduate school. The book
draws its conclusions from data collected on terrorist attacks around the world. An example of conclusion: terrorists are
not desperately poor or poorly educated. The author, Princeton economist Alan Krueger, concludes: “Terrorists tend to
come from well-educated, middle-class or high-income families.” 7 Why? Well, that exposes one of the limitations.
from regression analysis. We can isolate a strong association between two variables by using statistical analysis, but
we cannot necessarily explain why that relationship exists, and in some cases we cannot know for certain whether the
relationship is causal, meaning that a change in one variable is actually causing a change in the other. In the case of
terrorism, Professor Krueger hypothesizes that since terrorists are motivated by political goals, those with the most
education and the most resources have the greatest incentive to change society. These people may also feel particularly
irritated by the suppression of freedom, another factor associated with terrorism. In Krueger's study, countries with high
levels of political repression have more terrorist activity (holding other factors constant).

This discussion brings me back to the question posed by the chapter title: What's the point? The point is not to do
mathematics or to dazzle friends and colleagues with advanced statistical techniques. The point is to learn things that
inform our lives.

Lies, damned lies and statistics


Even under the best of circumstances, statistical analysis rarely reveals "the truth." We are usually building a
circumstantial case based on imperfect data. As a result, there are numerous reasons why intellectually honest people
might disagree with statistical results or their implications. At the most basic level, we may disagree with the question
being answered. Sports enthusiasts will forever argue about “the greatest baseball player of all time” because there is no
objective definition of “best.” Sophisticated descriptive statistics can inform this question, but they will never answer it
definitively. As will be noted in the next chapter, more socially significant issues fall prey to the same basic challenge.
What's happening to the economic health of the American middle class? That answer depends on how you define both
“middle class” and “economic health.”
There are limits to the data we can collect and the types of experiments we can run. Alan Krueger's study of terrorists
did not follow thousands of young people over several decades to see which of them evolved into terrorists.

It's simply not possible. Nor can we create two identical nations (except one is highly repressive and the other is not) and
then compare the number of suicide bombers that emerge in each. Even when we can conduct large controlled
experiments with humans, they are neither easy nor cheap. Researchers conducted a large-scale study on whether or not
Machine Translated by
Google

prayer reduces post-surgical complications, which was one of the questions posed earlier in this chapter. That study cost
$2.4 million. (For the results, you'll have to wait until Chapter 13.)
Secretary of Defense Donald Rumsfeld famously said, “You go to war with the army you have, not the army you might
want or wish to have later.” Whatever one thinks of Rumsfeld (and the Iraq war he was explaining), that aphorism also
applies to research. We perform statistical analysis using the best available data, methodologies and resources. The
approach is not like addition or long division, where the correct technique produces the “correct” answer and a computer is
always more accurate and less fallible than a human. Statistical analysis is more like good detective work (hence the
commercial potential of CSI: Regression Analysis). Intelligent and honest people will often disagree about what the data is
trying to tell us.

But who says that everyone who uses statistics is intelligent or honest? As mentioned, this book began as an homage
to How to Lie with Statistics, which was first published in 1954 and has sold over a million copies. The reality is that you
can lie with statistics. Or you may make unintentional mistakes. In any case, the mathematical precision associated with
statistical analysis can disguise serious nonsense. This book will discuss many of the most common statistical errors and
misrepresentations (so you can recognize them, not use them).

So, going back to the title chapter, what is the point of learning statistics?
To summarize huge amounts of data.
To make better decisions.
Respond to important social issues.
Recognize patterns that can refine the way we do everything from selling diapers to catching criminals.
Catch the cheaters and prosecute the offenders.
Evaluate the effectiveness of policies, programs, medications, medical services. procedures and other innovations.
And to detect the scoundrels who use these same powerful tools for nefarious purposes.
If you can do all that while looking great in a Hugo Boss suit or a black skirt-shorts, then you too could be the next star
of CSI: Regression Analysis.

* The Gini index is sometimes multiplied by 100 to convert it to a whole number. In that case, the United States would have a Gini index of 45.
* Historically, the word “data” has been considered plural (e.g. e.g., “The data is very encouraging”). The singular is “data,” which would refer to a single
data point, such as a person’s response to a single question in a survey. Using the word "data" as a plural noun is a quick way to tell anyone doing
serious research that you are familiar with statistics. That said, many grammar authorities and many publications, such as the New York Times, now
accept that “data” can be singular or plural, as demonstrated by the passage I quoted from the Times.
* This is a gross simplification of the fascinating and complex field of medical ethics.
Machine Translated by
Google

CHAPTER 2

Descriptive Statistics
Who was the best baseball player of all time?

Let us reflect for a moment on two seemingly unrelated questions: (1) What
What is happening to the economic health of the American middle class? and (2) Who was the best baseball player of all
time?
The first question is deeply important. It tends to be at the center of presidential campaigns and other social
movements. The middle class is the heart of America, so the economic well-being of that group is a crucial indicator of the
nation's overall economic health. The second question is trivial (in the literal sense of the word), but baseball enthusiasts
can argue about it endlessly.
What the two questions have in common is that they can be used to illustrate the strengths and limitations of descriptive
statistics, which are the numbers and calculations we use to summarize raw data.
If I want to prove that Derek Jeter is a great baseball player, I can sit down and describe every at-bat in every major
league game he ever played. That would be raw data, and it would take some time to digest, given that Jeter has played
seventeen seasons with the New York Yankees and has had 9,868 at-bats.

Or I can just tell you that at the end of the 2011 season, Derek Jeter had a career batting average of .313. This is a
descriptive statistic or a “summary statistic.”

Batting average is a gross simplification of Jeter's seventeen seasons. It is easy to understand, elegant in its simplicity,
and limited in what it can tell us.
Baseball experts have a wealth of descriptive statistics that they consider more valuable than batting average. I called
Steve Moyer, president of Baseball Info Solutions (a company that provides a lot of raw data on Moneyball guys), to ask:
(1) What are the most important statistics for evaluating baseball talent? and (2) Who was the best player of all time? I'll
share his response once we have more context.
Meanwhile, let's return to the less trivial issue: the economic health of the middle class. The ideal would be to find the
economic equivalent of a batting average, or something even better. We would like to have a simple but accurate measure
of what the economic well-being of the typical American worker has looked like.

changing in recent years. Are the people we define as middle class getting richer, poorer, or just staying put? A
reasonable answer (though by no means the “correct” answer) would be to calculate the change in per capita
income in the United States over a generation, which is roughly equivalent to thirty years. Per capita income is
a simple average: total income divided by population size. By that measure, median income in the United
States rose from $7,787 in 1980 to $26,487 in 2010 (the latest year for which the government has data).
1
Voila! Congratulations to us.
Machine Translated by
Google

There's just one problem. My quick calculation is technically correct, and yet completely wrong in terms of
the question I set out to answer. To begin with, the above figures are not adjusted for inflation. (A per capita
income of $7,787 in 1980 is equivalent to about $19,600 when converted to 2010 dollars.) This is a relatively
quick solution. The biggest problem is that the average income in the United States is not equal to the income
of the average American. Let's break down that clever little phrase.

Per capita income simply takes all the income earned in the country and divides it by the number of people,
which tells us absolutely nothing about who earned how much of that income, in 1980 or in 2010. As Occupy
Wall Street members would point out, explosive income growth for the top 1 percent can significantly increase
per capita income without putting more money into the pockets of the remaining 99 percent. In other words,
average income can rise without helping the average American.
As with the consultation on baseball statistics, I have sought outside experts on how we should measure
the health of the American middle class. I asked two prominent labor economists, including President Obama's
chief economic adviser, what descriptive statistics they would use to assess the economic well-being of a
typical American. Yes, you will get that answer too once we take a quick tour of the descriptive statistics to give
it more meaning.
From baseball to revenue, the most basic task when working with data is summarizing a large amount of
information. There are about 330 million residents in the United States. A spreadsheet containing the name
and income history of every American would contain all the information we could want about the country's
economic health, but it would also be so unwieldy that it would tell us nothing at all. The irony is that more data
can often present less clarity. So we simplify. We perform calculations that reduce a complex set of data into a
handful of numbers that describe that data, much as we might encapsulate a complex, multifaceted Olympic
gymnastics performance into one number: 9.8.
The good news is that these descriptive statistics give us a manageable and meaningful summary of the
underlying phenomenon. That's what this chapter is about.
about. The bad news is that any simplification invites abuse. Descriptive statistics can be like online dating
profiles: technically accurate, yet quite misleading.

Suppose you're at work, idly surfing the Web when you stumble upon a fascinating diary account of Kim
Kardashian's failed seventy-two-day marriage to professional basketball player Kris Humphries. You've just
finished reading about the seventh day of marriage when your boss shows up with two huge data files. One file
contains warranty claims information for each of the 57,334 laser printers your company sold last year. (For
each printer sold, the file documents the number of quality problems that were reported during the warranty
period.) The other file has the same information for each of the 994,773 laser printers its main competitor sold
during the same period.
Your boss wants to know how your company's printing presses compare in terms of quality to those of the
competition.
Luckily, the computer you've been using to read up on the Kardashian marriage has a basic statistics
package, but where do you start? Your instincts are probably right: the first descriptive task is usually to find
some measure of the “center” of a data set, or what statisticians might describe as its “central tendency.” What
Machine Translated by
Google

is the typical quality experience of your printers compared to competitors? The most basic measure of the
“middle” of a distribution is the mean or average. In this case, we want to know the average number of quality
issues per printer sold for your company and for your competitor. You would simply add up the total number of
quality issues reported for all printers during the warranty period and then divide it by the total number of
printers sold. (Remember, the same printer can have multiple problems while under warranty.) I would do that
for each company, creating one important descriptive statistic: the average number of quality issues per printer
sold.

Let's say it turns out that your competitors' printers have an average of 2.8 quality-related issues per printer
during the warranty period, compared to the average of 9.1 defects reported by your company. That was easy.
You've just taken data on a million printers sold by two different companies and boiled it down to the essence
of the problem: your printers break down frequently. Clearly it's time to send a quick email to your boss
quantifying this quality gap and then get back to day eight of Kim Kardashian's marriage.
Or maybe not. I was deliberately vague earlier when I referred to “half” of a distribution. The mean, or
average, turns out to have some problems in that regard, namely that it is prone to being distorted by "outliers,"
which are observations

which are further away from the center. To understand this concept, imagine ten guys sitting on bar stools at a middle-
class drinking establishment in Seattle; each of these guys makes $35,000 a year, making the median annual income of
the group $35,000. Bill Gates walks into the bar with a talking parrot perched on his shoulder. (The parrot has nothing to
do with the example, but it somehow adds flavor to things.) Let's assume, for the sake of example, that Bill Gates has an
annual income of $1 billion. When Bill sits on the eleventh stool at the bar, the average annual income of the bar's
customers increases to approximately $91 million.
Obviously, none of the original ten drinkers are any richer (although it might be reasonable to expect Bill Gates to buy a
round or two). If I were to describe the patrons of this bar with an average annual income of $91 million, the statement
would be both statistically correct and wildly misleading. This isn't a bar where billionaires hang out; it's a bar where a
bunch of guys with relatively low incomes sit next to Bill Gates and his talking parrot. The sensitivity of the mean to outliers
is why we should not measure the economic health of the American middle class by looking at per capita income.

Because there has been explosive growth in incomes at the top end of the distribution (CEOs, hedge fund managers, and
athletes like Derek Jeter), average income in the United States could be heavily skewed toward the mega-rich, making it
look a lot like the income of the mega-rich. bar stools with Bill Gates at the end.
For this reason, we have another statistic that also indicates the “mean” of a distribution, although in a different way:
the median. The median is the point that divides a distribution in half, meaning that half of the observations lie above the
median and half lie below. (If there are an even number of observations, the median is the midpoint between the two
middle observations.) If we go back to the bar stool example, the median annual income of the ten guys who originally sat
at the bar is $35,000. When Bill Gates walks in with his parrot and sits on a stool, the average annual income of the
eleven is still $35,000. If you literally imagine lining up the bar patrons on stools in ascending order of their income, the
income of the man sitting on the sixth stool represents the median income of the group. If Warren Buffett comes in and sits
on the twelfth stool next to Bill Gates, the median remains unchanged.
Machine Translated by
Google

*
For distributions without major outliers, the median and mean will be similar. I've included a hypothetical summary of
quality data from competing printers. In particular, I have presented the data in what is known as a frequency distribution.
The number of quality issues per printer is shown at the bottom; the height of each bar represents the percentages of
printers sold with that number of quality issues. For example, 36 percent of competing printers had two quality defects
during the warranty period.

Because the distribution includes all possible quality outcomes, including zero defects, the proportions must
sum to 1 (or 100 percent).

Frequency distribution of quality complaints from competitors


Printers

Zero One Two Three Four Five Six Seven Eight Nine Ten or more
Quality problems per printer

Since the distribution is almost symmetrical, the mean and median are relatively close to each other. The
distribution is slightly skewed to the right due to the small number of printers with many reported quality
defects. These outliers shift the mean slightly to the right but have no impact on the median. Suppose that just
before you rush off to send your quality report to your boss, you decide to calculate the average number of
quality problems experienced by your company's printers and those of your competitors. With a few
keystrokes, you will get the result. The median number of quality complaints for your company's printers is 2;
The median number of quality complaints for your company's printers is 1.
Hey? The average number of quality complaints per printer at your company is actually lower than that of
your competition. Since the Kardashian marriage is getting monotonous and you are intrigued by this finding,
you print out a frequency distribution for your own quality problems.

Frequency distribution of quality complaints in your company


Machine Translated by
Google

Quality problems per printer


What is clear is that your company does not have a uniform quality problem; you have a “lemon” problem; a
small number of printers have a large number of quality complaints. These outliers inflate the mean but not the
median. The most important thing from a production point of view is that you don't need to restructure the entire
manufacturing process; you just need to find out where the low-quality printers are coming from and... * fix it.
Neither the median nor the mean are difficult to calculate; the key is determining which measure of the
“mean” is most accurate in a particular situation (a phenomenon that is easily exploited). Meanwhile, the
median has some useful relatives. As we have already discussed, the median divides a distribution in half.
The distribution can be divided into quarters or quartiles. The first quartile consists of the bottom 25 percent of
observations; the second quartile consists of the next 25 percent of observations; and so on. Or the distribution
can be divided into deciles, each containing 10 percent of the observations. (If your income is in the top decile
of the U.S. income distribution, you would be earning more than 90 percent of your coworkers.) We can go
even further and divide the distribution into hundredths or percentiles. Each percentile represents 1 percent of
the distribution, so the first percentile represents the bottom 1 percent of the distribution and the 99th percentile
represents the top 1 percent of the distribution.

The benefit of these types of descriptive statistics is that they describe where a particular observation
stands in comparison to all the others. If I tell you that your child scored in the 3rd percentile on a reading
comprehension test, you immediately know that the family should spend more time at the library.
You don't need to know anything about the test itself or how many questions your child answered correctly.
The percentile score provides a ranking of your child's score relative to that of all other test takers. If the test
was easy, then most test takers will have a large number of correct answers, but your child will have fewer
correct answers than most others. If the test was extremely difficult, all test takers will have a low number of
correct answers, but your child's score will be even lower.
This is a good point to introduce some useful terminology. An “absolute” score, number, or figure has some
intrinsic meaning. If I shoot 83 on eighteen holes of golf, that's an absolute figure. I can do it on a day when the
temperature is 58 degrees, which is also an absolute number. Absolute numbers can usually be interpreted
without context or additional information. When I tell you I shot 83, you don't need to know what other golfers
shot that day to evaluate my performance. (The exception might be if the conditions are particularly terrible, or
if the course is especially difficult or easy.) If I finish ninth on the golf course

tournament, that's a relative statistic. A “relative” value or number only has meaning in comparison to
something else, or in some broader context, such as in comparison to the eight golfers who shot better than
me. Most standardized tests produce results that are meaningful only as relative statistics. If I tell you that a
third-grade student in an Illinois elementary school scored 43 out of 60 on the math portion of the Illinois State
Achievement Test, that absolute score doesn't mean much. But when I convert it to a percentile (that is, put
that raw score into a distribution with the math scores of all the other third graders in Illinois), it becomes very
meaningful. If 43 correct answers fall within the 83rd percentile, then this student is doing better than most of
his or her peers across the state. If you are in the 8th percentile, then you are struggling. In this case, the
percentile (the relative score) is more significant than the number of correct answers (the absolute score).
Machine Translated by
Google

Another statistic that can help us describe what might otherwise be a mess of numbers is the standard
deviation, which is a measure of how spread out the data is relative to its mean. In other words, how spread
out are the observations? Suppose I collected data on the weights of 250 people on a plane bound for Boston,
and I also collected the weights of a sample of 250 Boston Marathon qualifiers. Now let's assume the average
weight of both groups is about the same, say 155 pounds. Anyone who has ever been squeezed into a row on
a crowded flight, fighting for the armrest, knows that many people on a typical commercial flight weigh more
than 155 pounds. But you may remember from those same unpleasant, crowded flights that there were lots of
crying babies and misbehaving children, all of whom have enormous lung capacity but not much mass. When it
comes to calculating the average weight on the flight, the weight of the 320-pound football players on either
side of the middle seat is probably offset by the screaming little baby on the other side of the row and the six-
year-old kicking the back of the seat. your seat from the back row.
Based on the descriptive tools presented so far, the weights of airline passengers and marathon runners
are almost identical. But they are not. Yes, the weights of the two groups have roughly the same “mean,” but
airline passengers have much more dispersion around that midpoint, meaning their weights are farther from
the midpoint. My eight-year-old son might point out that marathon runners all seem to weigh the same, while
airline passengers have some tiny people and some oddly large ones.
Airline passengers' weights are "more evenly distributed," which is an important attribute when it comes to
describing the weights of these two groups.
The standard deviation is the descriptive statistic that allows us to assign a single number to this dispersion
around the mean. The formulas for calculating the
The standard deviation and variance (another common measure of dispersion from which the standard
deviation is derived) are included in an appendix at the end of the chapter. For now, let's think about why
measuring dispersion matters.
Suppose you walk into the doctor's office. He has been feeling fatigued since his promotion to head of print
quality in North America. Your doctor draws your blood, and a few days later your assistant leaves you a
message on your answering machine to tell you that your HCb2 (a fictitious blood chemical) count is 134. You
rush online and discover that the average HCb2 count for a person your age is 122 (and the median is about
the same). Holy cow! If you're like me, you'd eventually write a will. You would write tearful letters to your
parents, spouse, children, and close friends. You could take up skydiving or try to write a novel really quickly.
You would send your boss a hastily worded email comparing him to a certain part of human anatomy, ALL
CAPS.
None of these things may be necessary (and your email to your boss could go very wrong). When you call
the doctor's office back to schedule your hospice care, the physician's assistant tells you that your count is
within the normal range. But how could that be? “My count is 12 points higher than average!” you repeatedly
shout to the receiver.
“The standard deviation of the HCb2 count is 18,” the technician tells him dryly.

What the hell does that mean?


There is natural variation in the HCb2 count, as with most biological phenomena (e.g., (e.g. height). While
the average count for the fake chemical may be 122, many healthy people have higher or lower counts. The
Machine Translated by
Google

danger arises only when the HCb2 count increases or decreases excessively. So how do we figure out what
“excessively” means in this context? As we have already noted, the standard deviation is a measure of
dispersion, which means that it reflects how closely the observations are clustered around the mean. For many
typical data distributions, a high proportion of the observations lie within one standard deviation of the mean
(meaning they are in the range of one standard deviation below the mean to one standard deviation above the
mean). To illustrate with a simple example, the average height of American adult male is 5 feet 10 inches. The
standard deviation is about 3 inches. A high proportion of adult men are between 5 feet 7 inches and 6 feet 1
inch tall.
Or, to put it another way, any man in this height range would not be considered abnormally short or tall.
Which brings us back to your worrying HCb2 results. Yes, your count is 12 above average, but that's less than
one standard deviation, which is the blood chemical equivalent of about 6 feet.
Tall, not particularly unusual. Of course, far fewer observations are within two standard deviations of the mean, let alone
three or four standard deviations. (In the case of height, an American male who is three standard deviations above
average in height would be 6 feet 7 inches or taller.)
Some distributions are more dispersed than others. Therefore, the standard deviation of the weights of the 250 airline
passengers will be greater than the standard deviation of the weights of the 250 marathon runners. A frequency
distribution of airline passenger weights would literally be wider (more spread out) than a frequency distribution of
marathon runner weights. Once we know the mean and standard deviation of any collection of data, we have tremendous
intellectual traction. For example, suppose I tell you that the mean score on the SAT math test is 500 with a standard
deviation of 100.
As with height, most students taking the test will be within one standard deviation of the mean, or between 400 and 600.
How many students do you think will score 720 or higher? Probably not many, since they are more than two standard
deviations above the mean.
In fact, we can do even better than “not many.” This is a good time to introduce one of the most important, useful and
common distributions in statistics: the normal distribution. Normally distributed data are symmetrical around their mean in
the familiar bell shape.
The normal distribution describes many common phenomena. Imagine a frequency distribution describing popcorn
popping on the stove. Some kernels start popping early, maybe one or two per second; within ten or fifteen seconds, the
kernels are exploding frantically. Then, gradually, the number of grains exploding per second fades away at about the
same rate at which the burst began. American men's heights are more or less normally distributed, meaning they are
roughly symmetrical around the mean of 5 feet 10 inches.
Each SAT test is specifically designed to produce a normal distribution of scores with a mean of 500 and a standard
deviation of 100. According to the Wall Street Journal, Americans even tend to park in a normal distribution at malls, with
most cars parked directly in front of the mall entrance (the “peak” of the normal curve), with “queues” of cars heading to
the right and left of the entrance.
The beauty of the normal distribution (its power, finesse, and Michael Jordan-like elegance) comes from the fact that
we know by definition exactly what proportion of the observations in a normal distribution lie within one standard deviation
of the mean (68.2 percent), within two standard deviations of the mean (95.4 percent), within three standard deviations
(99.7 percent), and so on. This may seem like a triviality. In fact, it is the basis on which much of the statistics are built. We
Machine Translated by
Google

will return to this point in much greater depth later in the book.

The normal distribution

The mean is the middle line often represented by the Greek letter µ.
The standard deviation is often represented by the Greek letter σ. Each band represents one standard
deviation.

Descriptive statistics are often used to compare two numbers or quantities. I'm an inch taller than my brother;
today's temperature is nine degrees above the historical average for this date; etc. These comparisons make
sense because most of us recognize the scale of the units involved. An inch isn't much when it comes to a
person's height, so you can infer that my brother and I are roughly the same height. In contrast, nine degrees is
a significant temperature deviation in almost any climate and at any time of year, so nine degrees above
average makes for a much hotter day than usual. But suppose I told you that granola cereal A contains 31
milligrams more sodium than granola cereal B. Unless you know a lot about sodium (and granola cereal
serving sizes), that statement won't be particularly informative. Or what if I told you that my cousin Al made
$53,000 less this year than last year? Should we be worried about Al? Or is he a hedge fund manager for
whom $53,000 is a rounding error in his annual compensation?
In both the sodium and income examples, we are missing context. The easiest way to give meaning to
these relative comparisons is by using percentages. Would it mean something if I told you that granola bar A
has 50 percent more sodium than granola bar B, or that Uncle Al's income fell 47 percent last year. Measuring
change as a percentage gives us some sense of scale.

You probably learned how to calculate percentages in fourth grade and will be tempted to skip the next few
paragraphs. That seems fine to me. But first make a simple one.

exercise for me. Suppose a department store sells a dress for $100. The assistant manager marks down all merchandise
by 25 percent. But then that assistant director gets fired for being . , . .... .... .
at a bar with Bill Gates, and the new assistant manager raises all prices by 25 percent. What is the final price of the
dress? If you said (or thought) $100, then you better not skip any paragraphs.
Machine Translated by
Google

The final price of the dress is actually $93.75. This is not just a fun parlor trick that will earn you applause and
adulation at cocktail parties. Percentages are useful, but also potentially confusing or even misleading. The formula for
calculating a percentage difference (or change) is: (new figure – original figure)/original figure. The numerator (the top of
the fraction) gives us the size of the change in absolute terms; the denominator (the bottom of the fraction) is what puts
this change into context by comparing it to our starting point. At first, this seems simple, like when the assistant store
manager reduces the price of a $100 dress by 25 percent. Twenty-five percent of the original price of $100 is $25; that's
the discount, which reduces the price to $75. You can plug the numbers into the formula above and do some simple
manipulations to get to the same place: ($100 – $75)/$100 = 0.25, or 25 percent.

The dress is selling for $75 when the new assistant manager demands that the price be increased by 25 percent.
That's where many of the people reading this paragraph probably made a mistake. The 25 percent markup is calculated
as a percentage of the reduced price of the new dress, which is $75. The increase will be 0.25 ($75), or $18.75, which is
how the final price ends up at $93.75 (and not $100). The point is that a percentage change always gives the value of a
number in relation to something else. So we better understand what that something more is.

I once invested some money in a company my college roommate started. As this was a private company, there were
no requirements as to what information had to be provided to shareholders. Several years passed without any information
about the fate of my investment; my former roommate remained rather secretive on the subject. Finally, I received a letter
in the mail informing me that the company's profits were 46 percent higher than the previous year. There was no
information on the size of those gains in absolute terms, meaning I still had no idea how my investment was performing.
Suppose the company earned 27 cents last year, which is practically nothing. This year the company made 39 cents,
which is practically nothing. However, the company's earnings grew from 27 cents to 39 cents, which technically
represents a 46 percent increase. Obviously, the letter to shareholders would have been more depressing if it had pointed
out that the company's cumulative earnings over two years were less than the cost of a cup of Starbucks.
coffee.
To be fair to my roommate, he eventually sold the company for hundreds of millions of dollars, allowing me
to make a 100 percent return on my investment. (Since you have no idea how much I invested, you also have
no idea how much money I made, which reinforces my point here nicely!)
Let me make a further distinction. Percentage change should not be confused with a change in percentage
points. Rates are usually expressed in percentages. The sales tax rate in Illinois is 6.75 percent. I pay my
agent 15 percent of the royalties from my book. These rates apply to a certain amount, such as income in the
case of the income tax rate. Obviously rates can go up or down; less intuitively, rate changes can be described
in very different ways. The best example of this was a recent change to Illinois' personal income tax, which was
raised from 3 to 5 percent. There are two ways to express this tax change, and both are technically accurate.
The Democrats, who designed this tax increase, pointed out (correctly) that the state income tax rate was
increased by 2 percentage points (from 3 percent to 5 percent). Republicans pointed out (also correctly) that
the state income tax had increased by 67 percent. [This is a useful proof of the formula from a few paragraphs
back: (5 – 3)/3 = 2/3, which rounds to 67 percent.
Machine Translated by
Google

Democrats focused on the absolute change in the tax rate; Republicans focused on the percentage change
in the tax burden. As noted, both descriptions are technically correct, though I would argue that the Republican
description more accurately conveys the impact of the tax change, since what I will have to pay to the
government (the amount I care about, rather than how it is calculated) has actually increased by 67 percent.

Many phenomena defy perfect description with a single statistic. Suppose quarterback Aaron Rodgers throws
for 365 yards but no touchdowns. Meanwhile, Peyton Manning throws for just 127 yards but three touchdowns.
Manning generated more points, but Rodgers presumably set up touchdowns by marching his team down the
field and keeping the other team's offense off the field. Who played better? In Chapter 1, I discussed the NFL's
passer rating, which is the league's reasonable attempt to address this statistical challenge. Passer rating is an
example of an index, which is a descriptive statistic made up of other descriptive statistics. Once these different
performance measures are consolidated into a single number, that statistic can be used to make comparisons,
such as ranking quarterbacks on a particular day, or even over the course of an entire career. If baseball had a
similar rating, then the question of the greatest player of all time would be settled. Or yes?
The advantage of any index is that it consolidates a lot of complex information into a single number. Then we can rank things
that would otherwise defy simple comparison: anything from quarterbacks to colleges to beauty pageant contestants. In the Miss
America pageant, the overall winner is a combination of five different competitions: personal interview, swimsuit, evening gown,
talent, and question on stage. (The participants themselves vote separately for Miss Congeniality.)

Unfortunately, the disadvantage of any index is that it consolidates a lot of complex information into a single number. There
are countless ways to do this; each has the potential to produce a different result. Malcolm Gladwell makes this point
2 brilliantly in a
New Yorker article criticizing our compelling need to classify things.
(It falls particularly badly on college rankings.) Gladwell offers the example of Car and Driver's rating of three sports cars: the
Porsche Cayman, the Chevrolet Corvette, and the Lotus Evora.
Using a formula that included twenty-one different variables, Car and Driver ranked the Porsche number one.
But Gladwell points out that “exterior styling” accounts for just 4 percent of the overall Car and Driver score, which seems
ridiculously low for a sports car. If style is given more weight in the overall ranking (25 percent), then the Lotus comes out on top.

But wait. Gladwell also notes that a car's sticker price has relatively little weight in Car and Driver's formula. If value is
weighted more heavily (so the ranking is based equally on price, exterior styling, and vehicle features), the Chevy Corvette takes
the number one spot.
Any index is very sensitive to the descriptive statistics that are improvised to construct it and to the weight given to each of
these components. As a result, indexes range from useful but imperfect tools to complete charades. An example of the former is
the United Nations Human Development Index or HDI. The HDI was created as a measure of economic well-being that is
broader than income alone. The HDI uses income as one of its components, but also includes measures of life expectancy and
educational attainment. The United States ranks eleventh in the world in terms of economic output per capita (behind several oil-
rich nations such as Qatar, Brunei and Kuwait), but fourth in the world in human development.

3
It is true that the HDI rankings would change slightly if the index components were
reconfigured, but no reasonable change will move Zimbabwe up the rankings beyond Norway. The HDI provides a practical and
Machine Translated by
Google

reasonably accurate snapshot of living standards around the world.

Descriptive statistics give us insight into the phenomena that interest us. In that

In this sense, we can return to the questions posed at the beginning of the chapter. Who is the best baseball
player of all time? More importantly for the purposes of this chapter, what descriptive statistics would be most
useful in answering that question? According to Steve Moyer, president of Baseball Info Solutions, the three
most valuable statistics (other than age) for evaluating any non-pitcher would be:

1. On-base percentage (OBP), sometimes called on-base average (OBA): Measures the proportion of
the time a player reaches base successfully, including walks (which are not counted in batting average).

2. Slugging percentage (SLG): Measures batting power by calculating the total number of bases
reached per at-bat. A single counts as 1, a double is 2, a triple is 3, and a home run is 4. Thus, a batter
who hit a single and a triple in five at-bats would have a slugging percentage of (1 + 3)/5 or .800.
3. At bat (AB): Puts the above into context. Any downer can have impressive stats for one or two games.
A superstar compiles impressive “numbers” in thousands of plate appearances.

In Moyer's opinion (without hesitation, I might add), the greatest baseball player of all time was Babe Ruth
because of his unique ability to hit and pitch. Babe Ruth still holds the major league career slugging record
4
at .690.
What is happening to the economic health of the American middle class? Once again, I gave the floor to the
experts. I emailed Jeff Grogger (a colleague of mine at the University of Chicago) and Alan Krueger (the same
Princeton economist who studied terrorists and now chairs President Obama's Council of Economic Advisers).
Both gave variations of the same basic answer. To assess the economic health of the American “middle class,”
we must examine changes in median wages (adjusted for inflation) over the past several decades. They also
recommended examining changes in wages at the 25th and 75th percentiles (which can reasonably be
interpreted as the upper and lower bounds for the middle class).

One more distinction is necessary. When assessing economic health, we can look at income or wages.
They are not the same thing. A wage is what we are paid for a fixed amount of work, such as an hourly or
weekly wage. Revenue is the sum of all payments from different sources. If workers take on second jobs or
work longer hours, their income can increase without changing their wages. (In fact, income can rise even if
wages are falling, as long as a worker works enough hours at work.) However, if individuals have to work
harder to earn more, it is difficult to assess the overall effect on their incomes. welfare. the salary

It is a less ambiguous measure of how Americans are compensated for the work they do; the higher the wage, the more
workers earn per hour of work.

All that said, here's a chart of American wages over the past three decades. I also added the 90th percentile to
Machine Translated by
Google

illustrate the changes in wages of middle-class workers over this time period compared to workers at the top of the
distribution.

Source: “Changes in the Distribution of Workers’ Hourly Wages from 1979 to 2009,” Congressional Budget Office,
February 16, 2011. The data in the chart can be found at
[Link]

Several conclusions can be drawn from these data. They do not present a single “correct” answer regarding the
economic fate of the middle class.
We are told that the typical worker, an American worker earning the median wage, has been “working in place” for nearly
thirty years. Workers in the 90th percentile have fared much, much better. Descriptive statistics help frame the issue. What
we do about it, in any case, is an ideological and political question.

APPENDIX TO CHAPTER 2

Data for printer defect charts.


Machine Translated by
Google

Ten or
Zero One Two Three Four Five Sir Seven Fight Nine more

Froqucncy of
competitor's 12 14 36 or 8 6 5 J 0 2 1
defects

Ten or
Zero One Two Three Four Five Six Seven Eight Nine more

Frequency
of your 25 JI 9 4 J 0 0 1 1 0 26
defects

Formula for variance and standard deviation.


Variance and standard deviation are the most common statistical mechanisms for measuring
and describe the dispersion of a distribution. The variance, which is often represented by the
2
symbol σ, is calculated by determining how far the observations are within a distribution from the mean. The
problem, however, is that the difference between each observation and the mean is squared; the sum of those
squared terms is then divided by the number of observations.
Specifically:

For any set of n observations x1, x2, x3 .. . xn with mean u, Variance = o2 = I(x1 - u)2 + (x2 - u)2 + (x3 - u)2
+ ... (xn - u)2]/n

Because the difference between each term and the mean is squared, the formula for calculating the
variance gives special weight to observations that are far from the mean or outliers, as illustrated by the
following table of student heights.

Height H eight (p Distance from the


Group Distance from the mean = Absolute
1
(=70 (En-w)2 Group 2 = 70 (kn-w2
inches) mean = Absolute inches) value of (Xn -
value of <%-»• P>*

Nick 74 4 16 Sahar 65 5 25
Elana 66 4 16 Maggie 68 2 4

Dinah 68 2 4 Faisal 69 1 1

Rebecca 69 1 1 Ted 70 0 0

Ben 73 J 9 Jeff 71 1 1
Charu 70 0 0 Daffodil 75 5 25

Total = 14 Total = 46 =14 Total = 56


Variance = 56/6
Variance =
46/6 = 7.7 = 9.3

Standard ratio Standard

=,7.7=28 deviation = 9.3=3

* The absolute value is the distance between two figures, regardless of direction, so it is always positive. In this case, it represents the
number of inches between the individual's height and the average.
* With twelve customers in the bar, the median would be the midpoint between the income of the man on the sixth stool and the income of the
man on the seventh stool. Since they both earn $35,000, the median is $35,000.
If one earned $35,000 and the other $36,000, the median for the entire group would be $35,500.
Machine Translated by
Google

Both groups of students have an average height of 70 inches. The heights of students in both groups also
differ from the mean by the same number of total inches: 14. According to this measure of dispersion, the two
distributions are identical.
However, the variance for Group 2 is greater due to the weight given in the variance formula to values that are
particularly far from the mean (Sahar and Narciso in this case).
Variance is rarely used as a descriptive statistic on its own. Instead, variance is more useful as a step
toward calculating the standard deviation of a distribution, which is a more intuitive tool as a descriptive
statistic.
The standard deviation of a set of observations is the square root of the variance:

For any set of n observations x1, . . . xn with mean µ, x2, x3 standard deviation = σ = square root of
this total amount =

V[(†1 - p)2 + (*2 - P)2 * (x3 - p)2 + ■ • • (In “ M)2Vn

† Manufacturing update: It turns out that nearly all of the faulty printers were manufactured at a plant in Kentucky where workers had removed
parts from the assembly line to build a bourbon distillery. Both perpetually drunk employees and randomly missing parts on the assembly line
appear to have compromised the quality of the printers produced there.
* Amazingly, this person was one of ten people with annual incomes of $35,000 who were sitting on bar stools when Bill Gates walked in with
his parrot. Imagine!
Machine Translated by
Google

CHAPTER 3

Misleading description “He has a great


personality!” and other true but wildly misleading
statements

For
anyone who's ever thought about dating, the phrase "he's got a great personality" often sets off alarm bells — not
because the description is necessarily inaccurate, but because of what it might not reveal, like the fact that the guy has a
criminal record. or that their divorce “is not entirely final.” We don't doubt that this guy has a great personality; we fear that
a true statement, great personality, is being used to mask or obscure other information in a way that is seriously
misleading (assuming most of us would rather not date ex-offenders who are still married). The statement is not a lie per
se, meaning it would not convict him of perjury, but it could still be so inaccurate as to be false.

And the same goes for statistics. Although the field of statistics has its roots in mathematics and mathematics is exact,
using statistics to describe complex phenomena is not exact. That leaves a lot of room to obscure the truth. Mark Twain
pointed out that there are three kinds of lies: lies, damned lies, and statistics.
* .......................................... .... ................................................... . .
As explained in the last chapter, most of the phenomena that interest us can be described in multiple ways.
Once there are multiple ways to describe the same thing (e.g., "has a great personality" or "was convicted of securities
fraud"), the descriptive statistics we choose to use (or not use) will have a profound impact. in the impression we are
leaving. Someone with nefarious motives can use perfectly valid facts and figures to support completely questionable or
illegitimate conclusions.

We should start with the crucial distinction between "precision" and "accuracy." These words are not interchangeable.
Precision reflects the accuracy with which we can express something. In a description of the length of your trip, “41.6
miles” is more accurate than “about 40 miles,” which is more accurate than “a long, fucking way.” If you ask me how far
the nearest gas station is and I tell you it's 1,265 miles east, that's an accurate answer. Here's the problem: that answer
can be completely inaccurate if the gas station is in the other direction. On the other hand, if I tell you,

“Drive about ten minutes until you see a hot dog stand. The gas station will be a few hundred meters later on
the right. If you pass Hooters, you’ve gone too far,” my answer is less precise than “1,265 miles east,” but
significantly better because it sends you in the direction of the gas station.
Accuracy is a measure of whether a figure is, broadly speaking, consistent with the truth; hence the danger of
confusing precision with exactness. If an answer is precise, greater precision is usually better. But no amount
of accuracy can compensate for inaccuracy.

In fact, precision can mask inaccuracy by giving us a false sense of certainty, either inadvertently or
Machine Translated by
Google

deliberately. Joseph McCarthy, the red-baiting senator from Wisconsin, reached the height of his reckless
accusations in 1950 when he alleged not only that the U.S. State Department was infiltrated with communists
but that it had a list of their names. During a speech in Wheeling, West Virginia, McCarthy waved a piece of
paper in the air and declared: “I have here in my hand a list of 205, a list of names that were made known to
the Secretary of State as members of the Communist Party and yet they are still working and shaping policy in
1
the State Department.”
It turns out the paper didn't have any names, but the specificity of the accusation gave it credibility, even though it was a
blatant lie.
I learned the important distinction between precision and accuracy in a less malicious context. One year for
Christmas my wife bought me a golf rangefinder to calculate distances on the course from my golf ball to the
hole. The device works by using a sort of laser; I stand next to my ball on the fairway (or rough) and point the
rangefinder at the flag on the green, at which point the device calculates the exact distance I'm supposed to hit
the ball. This is an improvement over standard yardage markers, which give distances only to the center of the
green (and are therefore accurate but less precise). With my Christmas gift rangefinder I was able to tell that I
was 147.2 yards from the hole. I hoped the precision of this nifty technology would improve my golf game.

Instead, it worsened considerably.


There were two problems. First, I used the stupid device for three months before I realized it was set to
meters instead of yards; all the seemingly accurate calculations (147.2) were wrong. Secondly, I would
sometimes inadvertently aim the laser beam at the trees behind the green, instead of at the flag marking the
hole, so that my “perfect” shot would land exactly the distance it was supposed to: right over the green toward
the hole. forest. The lesson for me, which applies to all statistical analysis, is that even the most precise
measurements or calculations must be checked against common sense.
To take an example with more serious implications, many of Wall Street firms' risk management models prior
to the 2008 financial crisis were quite accurate. The concept of “value at risk” allowed companies to accurately quantify the
amount of company capital that could be lost under different scenarios. The problem was that the super-fancy models
were equivalent to setting my rangefinder in meters instead of yards. The mathematics was complex and arcane. The
answers he produced were reassuringly accurate. But the assumptions about what might happen to global markets that
were built into the models were simply wrong, making the conclusions wildly inaccurate in ways that destabilized not just
Wall Street but the entire global economy.
Even the most precise and accurate descriptive statistics can suffer from a more fundamental problem: a lack of clarity
about what exactly we are trying to define, describe, or explain. Statistical arguments have much in common with bad
marriages; the litigants often talk over each other. Let's consider an important economic question: How healthy is
American manufacturing? It is often said that huge numbers of American manufacturing jobs are being lost to China, India
and other low-wage countries. One also hears that high-tech manufacturing still thrives in the United States and that the
United States remains one of the world's leading exporters of manufactured goods. Which is it? This would seem to be a
case where a solid analysis of good data could reconcile these competing narratives. Is American manufacturing profitable
and globally competitive, or is it shrinking in the face of intense foreign competition?
Both. The British magazine The Economist reconciled both seemingly conflicting views on American manufacturing
with the following chart.
Machine Translated by
Google

“The Recovery of the Rust Belt,” March 10, 2011

The apparent contradiction lies in how the “health” of the American manufacturing industry is defined. In terms of
output (the total value of goods produced and sold), the U.S. manufacturing sector grew steadily in the 2000s, took a big
hit during the Great Recession, and has since rebounded strongly. This is consistent with data from the CIA World Factbook
showing that the United States is the world's third-largest manufacturing exporter, behind China and Germany.
The United States remains a manufacturing powerhouse.
But The Economist chart has a second line, which is manufacturing employment. The number of manufacturing jobs in the
United States has been steadily declining; approximately six million manufacturing jobs have been lost over the past decade.
Together, these two stories (rising manufacturing output and falling employment) tell the whole story. Manufacturing in the United
States has become increasingly productive, meaning factories are producing more with fewer workers. This is good from a global
competitiveness standpoint, because it makes American products more competitive with manufactured goods from low-wage
countries. (One way to compete with a company that can pay its workers $2 an hour is to create a manufacturing process so
efficient that a worker earning $40 can make twenty times as much.) But there are far fewer manufacturing jobs, which is terrible
news for the displaced workers who depended on those wages.

Since this is a book about statistics, not manufacturing, let’s get back to the main point, which is that the “health” of American
manufacturing—something seemingly easy to quantify—depends on how one decides to define health: output or employment?
In this case (and in many others), the fuller story comes from including both figures, as The Economist wisely decided to do in its
chart.

Even when we agree on a single measure of success—say, student test scores—there is plenty of statistical wiggle room.
See if you can reconcile the following hypothetical statements, all of which could be true:
Politician A (the challenger): “Our schools are getting worse! Sixty percent of our schools had lower test scores this year than
last year.”
Politician B (the headline): “Our schools are getting better! “Eighty percent of our students scored higher on tests this year
than last year.”
Here's a hint: not all schools necessarily have the same number of students. If you take another look at the seemingly
contradictory statements, what you'll see is that one politician is using schools as his unit of analysis (“Sixty percent of our
schools…”), and the other is using students as the unit. of analysis (“Eighty percent of our students…”). The unit of analysis is
the entity that the statistics compare or describe: the academic performance of one of them and the performance of the students
of the other. It is entirely possible for most students to improve and most schools to get worse, if the students showing
Machine Translated by
Google

improvement are from very large schools. To make this example more intuitive, let's do the same exercise using the American
states:

Politician A (populist): “Our economy is in shit! “Thirty states had revenue declines last year.”
Politician B (more elitist): “Our economy is showing appreciable gains: seventy percent of Americans had increasing
incomes last year.”
What I would infer from those statements is that the largest states have the healthiest economies: New
York, California, Texas, Illinois, etc. The thirty states with declining median incomes are likely to be much
smaller: Vermont, North Dakota, Rhode Island, etc. Given the disparity in state size, it is quite possible that
most states will do worse while most Americans will do better. The key lesson is to pay attention to the unit of
analysis. Who or what is described? Is that different from the “who” or “what” described by someone else?

Although the above examples are hypothetical, here is a crucial statistical question that is not: Is
globalization improving or worsening income inequality around the world? According to one interpretation,
globalization has only exacerbated existing income inequalities; in 1980, richer countries (measured by GDP
per capita) tended to grow faster between 1980 and 2000 than poorer countries. Rich countries simply got
richer, suggesting that trade, outsourcing, foreign investment and the other components of “globalization” are
merely tools for the developed world to extend its economic hegemony. Down with globalization! Down with
globalization!
But wait a moment. The same data can (and should) be interpreted completely differently if the unit of
analysis is changed. We don't care about poor countries; we care about the poor. And it turns out that a high
proportion of the world's poor live in China and India. Both countries are huge (with a population of over a
billion); both were relatively poor in 1980. China and India have not only grown rapidly over the past few
decades, but they have done so in large part because of their greater economic integration with the rest of the
world. They are “rapid globalizers,” as The Economist has described them.
Since our goal is to alleviate human misery, it makes no sense to give China (population 1.3 billion) the same
weight as Mauritius (population 1.3 billion) when examining the effects of globalization on the poor.
The unit of analysis should be people, not countries. What really happened between 1980 and 2000 is very
similar to my fake school example above. Most of the world's poor lived in two giant countries that grew
extremely rapidly as they became more integrated into the global economy. A proper analysis draws a
completely different conclusion about the benefits of globalization for the world's poor. As The Economist
notes, “if you look at people, not countries, global inequality is falling rapidly.”
Telecommunications companies AT&T and Verizon have recently engaged in an advertising battle that
exploits this kind of ambiguity about what is being described. Both companies provide cell phone service. One
of the main concerns of most mobile phone users is the quality of service in the locations where they are likely
to make or receive phone calls. Therefore, a logical point of comparison between the two companies is the size
and quality of their networks.
While consumers just want decent cell phone service in lots of places, both AT&T and Verizon have come up
with different metrics to measure the somewhat amorphous demand for “decent cell phone service in lots of
Machine Translated by
Google

places.”
Verizon launched an aggressive advertising campaign touting the geographic coverage of its network; you may
remember maps of the United States that showed the large percentage of the country covered by Verizon's
network compared to the relatively insignificant geographic coverage of AT&T's network.
Verizon's chosen unit of analysis is geographic area covered, because the company has more.
AT&T responded by launching a campaign that changed the unit of analysis.
Their billboards announced that “AT&T covers 97 percent of Americans.” Note the use of the word "Americans"
instead of "United States." AT&T focused on the fact that most people don't live in rural Montana or the Arizona
desert. Since the population is not evenly distributed across the physical geography of the United States, the
key to good cellular service (the campaign implicitly argued) is to have a network where callers actually live and
work, not necessarily where they go to camp. However, as someone who spends a fair amount of time in rural
New Hampshire, I sympathize with Verizon on this one.
Our old friends, the mean and the median, can also be used for nefarious purposes.
As you will recall from the previous chapter, both the median and the mean are measures of the “center” of a
distribution, or its “central tendency.” The mean is a simple average: the sum of the observations divided by the
number of observations. (The average of 3, 4, 5, 6 and 102 is 24). The median is the midpoint of the
distribution; half of the observations lie above the median and half lie below. (The median of 3, 4, 5, 6, and 102
is 5.) Now, the intelligent reader will see that there is a considerable difference between 24 and 5. If, for some
reason, I wanted to describe this group of numbers in a way that makes them look big, I would focus on the
mean. If I want it to appear smaller, I will quote the median.

Now let's see how this plays out in real life. Consider George W. Bush's tax cuts. Bush, which were touted
by the Bush administration as being good for most American families. While pushing the plan, the
administration noted that 92 million Americans would receive an average tax cut of more than
$1,000 ($1,083 to be precise). But was that summary of the tax cut accurate? According to the New York Times, "The data
doesn't lie, but some of it is fake."

Would 92 million Americans get a tax cut? Yeah.


Would most of those people get a tax cut of around $1,000? No. The median tax cut was less than $100.
A relatively small number of extremely wealthy people were eligible for large tax cuts; these large numbers distort the
average, making the average tax cut appear larger than what most Americans would likely receive. The median is not sensitive
to outliers and, in this case, is probably a more accurate description of how the tax cuts affected the typical household.
Of course, the median can also do its part to disguise things because it is not sensitive to outliers. Suppose you have a life-
threatening illness. The good news is that a new drug has been developed that may prove effective. The downside is that it is
extremely expensive and has many unpleasant side effects.
“But does it work?” you ask. The doctor informs him that the new drug increases the average life expectancy of patients with his
disease by two weeks. This is not encouraging news; the medication may not be worth the cost and inconvenience. His
insurance company refuses to pay for the treatment; he has a pretty good case based on average life expectancy figures.
Machine Translated by
Google

However, in this case the median can be a terribly misleading statistic. Suppose that many patients do not respond to the new treatment but that a large

number of patients, say 30 or 40 percent, are completely cured. This success would not be reflected in the median (although the average life expectancy of those

taking the drug would seem very impressive). In this case, outliers (those who take the drug and live a long time) would be very relevant to your decision. And this

is not just a hypothetical case. Evolutionary biologist Stephen Jay Gould was diagnosed with a form of cancer that had a median survival of eight months; he died

of a different, unrelated type of cancer twenty years later. 3 Gould later wrote a famous article entitled “The Median Is Not the Message,” in which he argued that

his scientific knowledge of statistics saved him from the erroneous conclusion that he would necessarily be dead within eight months. The definition of the median

tells us that half of patients will live at least eight months, and possibly much, much longer than that. The mortality distribution is “skewed to the right,” which is

more than a technicality if one has the disease.

In this example, the defining characteristic of the median: that it does not weight observations based on how far they are
from the midpoint, only based on

Whether they are up or down, it turns out to be their weakness. On the contrary, the mean is affected by
dispersion. From an accuracy standpoint, the question of median versus mean revolves around whether
outliers in a distribution distort what is being described or, on the contrary, are an important part of the
message.
(Once again, judgment trumps math.) Of course, nothing says you have to choose the median or the mean. Any thorough
statistical analysis would likely show both. When only the median or mean appears, it may be for brevity, or it may be
because someone is trying to “persuade” you with statistics.

Those of a certain age may remember the following exchange (as I recall it) between the characters played by
Chevy Chase and Ted Knight in the movie Caddyshack. The two men meet in the locker room after both have
just left the golf course:
TED KNIGHT: What did you shoot?
CHEVY CHASE: Oh, I don't keep score.
TED KNIGHT: So how do you compare to other golfers?
CHEVY CHASE: For height.
I'm not going to try to explain why this is funny. I will say that a lot of statistical shenanigans arise from
“apples and oranges” comparisons. Let's say you're trying to compare the price of a hotel room in London with
the price of a hotel room in Paris. You send your six-year-old to the computer to do some Internet research, as
he or she is much faster and better than you. His son informs him that hotel rooms in Paris are more
expensive, around 180 per night; a comparable room in London costs 150 per night.
You would probably explain to your child the difference between pounds and euros and then send him or
Machine Translated by
Google

her back to the computer to find the exchange rate between the two currencies so you could make a
meaningful comparison. (This example has a loose basis in truth: After I paid 100 rupees for a cup of tea in
India, my daughter wanted to know why everything in India was so expensive.)
Obviously, the numbers on currencies of different countries mean nothing until we convert them into
comparable units. What is the exchange rate between the pound and the euro or, in the case of India, between
the dollar and the rupee?
This seems like a painfully obvious lesson, but one that is routinely ignored, especially by politicians and
Hollywood studios. These people clearly recognise the difference between euros and pounds; however, they
overlook a more subtle example of apples and oranges: inflation. A dollar today is not the same as a dollar
sixty years ago; it buys much less. Due to inflation, something that
Machine Translated by
Google

cost $1 in 1950, would cost $9.37 in 2011. As a result, any monetary comparison between 1950 and 2011 without adjusting for
changes in the value of the dollar would be less accurate than comparing euro and pound figures, since the euro and pound are
closer to each other in value than a 1950 dollar is to a 2011 dollar.

This is such an important phenomenon that economists have terms to indicate whether figures have been adjusted for
inflation or not. Nominal figures are not adjusted for inflation. A comparison of the nominal cost of a government program in 1970
with the nominal cost of the same program in 2011 simply compares the size of the checks the Treasury wrote in those two
years, without any recognition that a dollar in 1970 bought more things than a dollar in 2011. If we spent $10 million on a
program in 1970 to provide housing assistance to veterans and $40 million on the same program in 2011, the federal
commitment to that program has actually decreased. Yes, spending has increased in nominal terms, but that does not reflect the
changing value of the dollars being spent. One dollar in 1970 is worth $5.83 in 2011; the government would need to spend $58.3
million on veterans housing benefits in 2011 to provide support comparable to the $10 million it spent in 1970.

The actual figures, on the other hand, are adjusted for inflation. The most commonly accepted methodology is to convert all
figures to a single unit, such as 2011 dollars, to make an “apples-to-apples” comparison. Many websites, including the U.S.
Bureau of Labor Statistics,
In the U.S., they have simple inflation calculators that will compare the value of a dollar at different points in time.
For a real-life example (yes, pun intended) of how statistics can look different when adjusted for *_
* for
inflation, check out the following chart of the US federal minimum wage. U.S., which represents both the
nominal value of the minimum wage and its real purchasing power in 2010 dollars. .
Machine Translated by
Google

Source: [Link]

The federal minimum wage (the number posted on the bulletin board in some remote corner of your office)
is set by Congress. This salary, currently $7.25, is a nominal figure. Your boss doesn't have to make sure that
$7.25 can buy you as much as it did two years ago; he or she just has to make sure that you get paid at least
$7.25 for every hour of work you do. What's important is the number on the check, not what that number can
buy.
However, inflation over time erodes the purchasing power of the minimum wage (and all other nominal
wages, which is why unions often negotiate “cost-of-living adjustments”). If prices rise faster than Congress
raises the minimum wage, the real value of that minimum hourly pay will fall. Supporters of a minimum wage
should be concerned about the real value of that wage, since the goal of the law is to guarantee low-wage
workers a minimum level of consumption for an hour of work, not to give them a check with a large amount.
who buys less than before. (If that were the case, then we could pay low-wage workers in rupees.)
Hollywood studios may be the most egregiously oblivious to the distortions caused by inflation when comparing figures at different points in time, and

deliberately so. What were the five highest-grossing films (domestic) of the five periods in 2011?

[Link] (2009)
2. Titanic (1997)
3. The Dark Knight (2008)
4. Star Wars Episode IV (1977)
[Link] 2 (2004)
Now you might feel like that list looks a little suspicious. They were successful movies, but Shrek 2? Was it
really a bigger commercial success than Gone with the Wind? The Godfather? Jaws? No, no and no.
Hollywood likes to make each blockbuster seem bigger and more successful than the last. One way to do this
would be to quote box office receipts in Indian rupees, which would inspire headlines like: “Harry Potter breaks
box office record with weekend takings of 1.3 billion!” But even the most foolish movie buffs would be
suspicious of figures that are large just because they are quoted in a currency with relatively little purchasing
power. Instead, Hollywood studios (and the journalists who report on them) simply use nominal figures, making
recent films appear successful largely because ticket prices are higher now than they were ten, twenty, or fifty
Machine Translated by
Google

years ago. (When Gone with the Wind was released in 1939, a ticket cost about $0.50.) The most accurate
way to compare commercial success over time would be to adjust ticket receipts for inflation. Earning $100
million in 1939 is much more impressive than earning $500 million in 2011. So what are the highest-grossing
films in the US of all time, adjusted for inflation?
6

1. Gone with the Wind (1939)


2. Star Wars Episode IV (1977)
3. The Sound of Music (1965)
4. ET (1982)
5. The Ten Commandments (1956)
In real terms, Avatar falls to 14th place; Shrek 2 falls to 31st place.
Even comparing apples to apples leaves plenty of room for mischief. As discussed in the last chapter, an
important function of statistics is to describe changes in quantities over time. Are taxes going up? How many
cheeseburgers are we selling compared to last year? How much have we reduced arsenic in our drinking
water? We often use percentages to express these changes because they give us a sense of scale and
context. We understand what it means to reduce the amount of arsenic in drinking water by 22 percent, while
few of us would know whether reducing arsenic by one microgram (the absolute reduction) would be a
significant change or not. Percentages don't lie, but they can exaggerate. One way to make growth appear
explosive is to use percentage change to describe some change relative to a very low starting point. I live in
Cook County, Illinois. One day I was shocked to learn that the portion of my tax dollars funding the suburban
Cook County Tuberculosis Sanitarium District was scheduled to increase by 527 percent. However, I cancelled
my massive anti-tax demonstration (which was actually still in the planning phase) when I heard about it.

It would cost me less than a good turkey sandwich. The Sanatorium District for Tuberculosis treats approximately one
hundred cases a year; it is not a large or expensive organization. The Chicago Sun-Times noted that for the typical
homeowner, the tax bill would go from $1.15 to $6. qualify a growth figure by noting that 7 Sometimes researchers say it is
coming “from a low base,” meaning that any increase will seem large by comparison.

Obviously the other side is true. A small percentage of a huge sum can be a big amount. Suppose the Secretary of
Defense reports that defense spending will grow by only 4 percent this year. Great news! Not really, given that the
Department of Defense's budget is nearly $700 billion. Four percent of $700 billion is $28 billion, which is enough to buy a
lot of turkey sandwiches. In fact, that seemingly insignificant 4 percent increase in the defense budget is more than
NASA's entire budget and roughly the same as the budgets of the Departments of Labor and Treasury combined.
Similarly, your kind-hearted boss might point out that, to be fair, all employees will receive the same raise this year: 10
percent.
What a magnanimous gesture, except if your boss makes $1 million and you make $50,000, your raise will be $100,000
and his will be $5,000. The statement “everyone will get the same 10 percent raise this year” sounds much better than “my
raise will be twenty times yours.” Both things are true in this case.
Machine Translated by
Google

Any comparison of a quantity that changes over time must have a start point and an end point. Sometimes these
points can be manipulated in ways that affect the message. I once had a professor who liked to talk about his “Republican
slides” and his “Democratic slides.” He was referring to data on defense spending, and what he meant was that he could
organize the same data in different ways to please Democratic or Republican audiences. For his Republican audiences,
he would offer the following slide with data on increases in defense spending under Ronald Reagan. Clearly, Reagan
helped restore our commitment to defense and security, which in turn helped win the Cold War. No one can look at these
figures and not appreciate Ronald Reagan's ironclad determination to take on the Soviets.

Defense spending in billions, 1981-1988


Machine Translated by
Google

350

Yo
300
250

200

150

100
50

1981 1982 1983 1984 1985 1986 1987 1988

For Democrats, my former professor simply used the same (nominal) data, but over a longer time frame.
For this group, he said Jimmy Carter deserves credit for initiating defense preparation. As the next
“Democratic” slide shows, defense spending increases from 1977 to 1980 show the same basic trend as
increases during the Reagan presidency. Thank goodness that Jimmy Carter, an Annapolis graduate and
former naval officer, began the process of making America strong again!

Defense spending in billions, 1977-1988


Fountain:

[Link]?
span=usgs302&year=1988&view=1&expand=30&expandC=&units=b&fy=fy12&local=s&state=US&pie=#usgs302.

While the primary purpose of statistics is to present a meaningful picture of the things we care about, in
many cases we also expect to act on these numbers. NFL teams want a simple measure of quarterback
quality so they can find and recruit talented players out of college. Companies measure the performance of
their employees so they can promote those who are valuable and fire those who are not. There is a common
business aphorism: "You can't manage what you can't measure." TRUE. But you'd better be absolutely sure
that what you're measuring is actually what you're trying to manage.
Machine Translated by
Google

Consider the quality of the school. It is crucial to measure this, as we would like to reward and emulate “good” schools
while sanctioning or fixing “bad” schools. (And within each school, we have the similar challenge of measuring teacher
quality, for the same basic reason.) The most common measure of quality for both schools and teachers is test scores. If
students get impressive scores on a well-designed standardized test, then presumably the teacher and the school are
doing a good job. On the contrary, poor test scores are a clear sign that many people should be fired, sooner rather than
later. These statistics can help us a lot in fixing our public education system, right?
Mistaken. Any evaluation of teachers or schools that relies solely on test scores will present a dangerously inaccurate
picture. The students who walk through the front doors of different schools have very different backgrounds and abilities.
We know, for example, that the education and income of a student's parents have a significant impact on performance,
regardless of which school they attend. The statistic we're missing in this case turns out to be the only one that matters
for our purposes: How much of a student's performance, good or bad, can be attributed to what happens inside the
school (or inside a particular classroom)?

Students who live in wealthy, highly educated communities will do well from the moment their parents drop them off at
school on their first day of kindergarten. The other side is also true. There are schools with extremely disadvantaged
populations where teachers may be doing a remarkable job, but student test scores will still be low, though not as low as
they would have been if the teachers had not done a good job. What we need is some measure of “value added” at the
school level, or even at the classroom level. We don't want to know the absolute level of student achievement; we want to
know to what extent student achievement has been affected by the educational factors we are trying to assess.

At first glance, this seems like an easy task, as we can simply give students a pretest and a posttest. If we know
students' test scores when they enter a particular school or classroom, then we can measure their performance at the
end and attribute the difference to what happened in that school or classroom.
Unfortunately, I am wrong again. Students with different abilities or backgrounds may also learn at different rates.
Some students will grasp the material faster than others for reasons that have nothing to do with the quality of teaching.
So if students at Prosperous School A and Poor School B begin studying algebra at the same time and at the same level,
the explanation for the fact that students at Prosperous School A perform better in algebra a year later may be that the
teachers are better, or it may be that the students were able to learn faster, or both. Researchers are working

Develop statistical techniques that measure the quality of instruction in ways that adequately account for students' different
backgrounds and abilities. Meanwhile, our attempts to identify the “best” schools can be ridiculously misleading.

Each fall, several Chicago newspapers and magazines publish a ranking of the region’s “best” high schools, usually based
on state test score data. Here's the laugh-out-loud part from a statistical standpoint: Several of the high schools that consistently
occupy the top spots in the rankings are selective enrollment schools, meaning that students must apply to get in, and only a
small proportion of those students are accepted. One of the most important admission criteria is standardized test scores. So
let's summarize: (1) these schools are being recognized as “excellent” for having students with high test scores; (2) to get into
such a school, one must have high test scores. This is the logical equivalent of giving an award to the basketball team for doing
Machine Translated by
Google

such an excellent job of producing tall students.

Even if you have a solid indicator of what you are trying to measure and manage, the challenges are not over. The good news is
that “statistics-based management” can improve the underlying behavior of the person or institution being managed. If you can
measure the proportion of defective products coming off an assembly line, and if those defects are a function of things
happening on the plant floor, then some kind of bonus for workers tied to a reduction in defective products would presumably
change behavior in the right kinds of ways. Each of us responds to incentives (even if it's just praise or a better parking spot).
Statistics measure the results that matter; incentives give us a reason to improve those results.

Or, in some cases, just to make the statistics look better. That's the bad news.
If school administrators are evaluated (and perhaps even compensated) on the basis of the high school graduation rate of
students in a particular school district, they will focus their efforts on increasing the number of students who graduate. Of
course, they can also devote some effort to improving the graduation rate, which is not necessarily the same thing. For
example, students who drop out of school before graduating may be classified as “dropouts” rather than dropouts. This isn't just
a hypothetical example; it's a charge that was brought against former Education Secretary Rod Paige during his tenure as
Houston's schools superintendent. Paige was hired by President George W.

Bush will become U.S. education secretary because of his remarkable success in Houston in reducing the dropout rate and
improving test scores.
If you're aware of the little business aphorisms I keep throwing out

By the way, here's another one: "It's never a good day when 60 Minutes shows up at your door." Dan Rather and the 60
Minutes II team took a trip to Houston and discovered that the manipulation of statistics was far more impressive than the
educational improvement.
8
High schools routinely classified students who dropped out of high school as transferring
to another school, returning to their home country, or leaving to earn a General Equivalency Diploma (GED), none of which
count as dropouts in official statistics. Houston reported a citywide dropout rate of 1.5 percent in the year examined; 60 Minutes
estimated the true dropout rate to be between 25 and 50 percent.
The statistical shenanigans with test scores were equally impressive. One way to improve test scores (in Houston or
anywhere else) is to improve the quality of education so that students learn more and do better. This is a good thing. Another
(less virtuous) way to improve test scores is to prevent the worst students from taking them. If the scores of the lowest-
performing students are removed, the average test score for the school or district will increase, even if the rest of the students
show no improvement. In Texas, the state achievement test is given in tenth grade. There was evidence that Houston schools
were trying to prevent weaker students from reaching the tenth grade. In one particularly egregious example, a student spent
three years in ninth grade and was then promoted directly to eleventh grade, a deviously clever way to prevent a weak student
from taking a tenth-grade benchmark exam without forcing him to drop out (which would have shown up in a different statistic).

It is not clear that Rod Paige was complicit in this statistical cheating during his tenure as Houston superintendent; however,
Machine Translated by
Google

he did implement a rigorous accountability program that awarded cash bonuses to principals who met his dropout and test
score goals and fired or demoted principals who failed to meet their targets. The directors definitely responded to the incentives;
that is the most important lesson. But you'd better be absolutely sure that the people being evaluated can't possibly be made
better (statistically) in ways that are not consistent with the goal at hand.

New York State learned this the hard way. The state introduced “scorecards” that assess the mortality rates of patients of
cardiologists who perform coronary angioplasty, a common treatment for heart disease. 9 This seems a perfectly reasonable
and useful use of descriptive statistics. It is important to know the proportion of a cardiologist's patients who die in surgery, and
it makes sense for the government to collect and publicize such data, since individual consumers would otherwise not have
access to it. So is this a good policy? Yeah, apart from the fact that he probably ended up killing people.

Obviously, cardiologists care about their “scorecard.” However, the easiest way for a surgeon to improve his mortality
rate is not by killing fewer people; presumably most doctors are already trying very hard to keep their patients alive. The
easiest way for a doctor to improve his mortality rate is to refuse to operate on the sickest patients. According to a survey
by the University of Rochester School of Medicine and Dentistry, the scorecard, which ostensibly serves patients, may
also harm them: 83 percent of cardiologists surveyed said that because of public mortality statistics, some patients who
might benefit from angioplasty might not receive the procedure; 79 percent of physicians said some of their personal
medical decisions had been influenced by the knowledge that mortality data is collected and made public. The sad
paradox of this seemingly useful descriptive statistic is that cardiologists responded rationally by denying care to the
patients who needed it most.

A statistical index has all the potential dangers of any descriptive statistic, plus the distortions introduced by combining
multiple indicators into a single number. By definition, any index will be sensitive to how it is constructed; it will be affected
both by the measures that are included in the index and by how each of those measures is weighted. For example, why
doesn't the NFL passer rating include any measure of third-down completions? And in the case of the Human
Development Index, how should a country's literacy rate be weighted relative to per capita income? In the end, the
important question is whether the simplicity and ease of use introduced by lumping many indicators into a single number
outweighs the inherent inaccuracy of the process. Sometimes that answer can be no, which brings us back (as promised)
to U.S. News & World Report (USNWR) college rankings.

The USNWR rankings use sixteen indicators to rate and rank America's colleges, universities, and professional
schools. In 2010, for example, the National University and Liberal Arts College Rankings used “student selectivity” as 15
percent of the index; student selectivity, in turn, is calculated based on a school’s acceptance rate, the proportion of
incoming students who were in the top 10 percent of their high school class, and the average SAT and ACT scores of
incoming students. The benefit of the USNWR rankings is that they provide a lot of information about thousands of
schools in a simple and accessible way. Even critics admit that much of the information collected about American
colleges and universities is valuable. Prospective students should know an institution's graduation rate and average class
size.
Of course, providing meaningful information is a completely different task than grouping all that information into a
single classification that aims to
Machine Translated by
Google

be authoritarian. Critics say the rankings are poorly constructed, misleading and detrimental to students' long-term interests.
"One of the concerns is simply that this is a list that purports to rank institutions in numerical order, which is a level of precision
that those data simply don't support," says Michael McPherson,
. ......................................... ... . 10 .. . .
the former president of Macalester College in Minnesota. Why contributions
of alumni should count for 5 percent of a school's score? And if it is important, why doesn't it count as ten percent?
According to US News & World Report, “Each indicator is assigned a weight (expressed as a percentage) based on our
judgments about which quality measures are most important.”
11
Judgment is one thing; arbitrariness is another. The most weighted variable in the
ranking of national universities and faculties is “academic reputation”. This reputation is determined based on
a “peer evaluation survey” completed by administrators from other colleges and universities and a survey of
high school counselors. In his general critique of rankings, Malcolm Gladwell offers a scathing (if humorous)
critique of peer review methodology. He cites a questionnaire sent by a former Michigan chief justice to about
100 lawyers asking them to rank ten law schools in order of quality. Penn State was one of the law schools on
the list; lawyers ranked it near the middle. At the time, Penn State did not have a law school.
12
Despite all the data USNWR collects, it's not obvious that the rankings measure what prospective students should care
about: how much learning is taking place at a given institution? Football fans may quibble with the makeup of passer rating, but
no one can deny that its components (completions, yards, touchdowns and interceptions) are an important part of a
quarterback's overall performance. That's not necessarily the case with the USNWR criteria, most of which focus on inputs (e.g.,
what kinds of students are admitted, how much faculty are paid, the percentage of faculty who work full time) rather than
educational outcomes. Two notable exceptions are the freshman retention rate and the graduation rate, but even those
indicators do not measure learning. As Michael McPherson notes, "We really don't learn anything from US News about whether
the education they received during those four years actually enhanced their talents or enriched their knowledge."

All of this would still be a harmless exercise, were it not for the fact that it seems to encourage behaviors that are not
necessarily good for students or higher education. For example, one statistic used to calculate rankings is financial resources
per student; the problem is that there is no corresponding measure of how well that money is being spent. An institution that
spends less money to obtain better results

(and therefore can charge a lower tuition) is penalized in the classification process. Colleges and universities
also have an incentive to encourage large numbers of students to apply, including those with no realistic
hopes of getting in, because that makes the school appear more selective. This is a waste of resources for
schools that solicit fake applications and for students who end up applying without any meaningful chance of
being accepted.
Since we're about to move on to the probability chapter, I'm betting that the US News & World Report
rankings aren't going away anytime soon. As Leon Botstein, president of Bard College, has noted, “People
love easy answers.
13 What is the best place? Number 1."
Machine Translated by
Google

The overall lesson of this chapter is that statistical misconduct has very little to do with bad mathematics. In
any case, impressive calculations may hide nefarious motives. The fact that you calculated the mean correctly
does not alter the fact that the median is a more accurate indicator. Judgment and integrity turn out to be
surprisingly important. Detailed knowledge of statistics does not deter crime, just as detailed knowledge of the
law does not prevent criminal behavior. With both statistics and crime, the bad guys often know exactly what
they're doing!

* Twain attributed this line to British Prime Minister Benjamin Disraeli, but there is no record that Disraeli ever said or wrote it.
* Available at [Link]
Machine Translated by
Google

CHAPTER 4

Correlation

How does Netflix know what movies I like?

Netflix insists I will like the film Bhutto, a documentary that offers an “in-depth and sometimes incendiary look
at the life and tragic death of former Pakistani Prime Minister Benazir Bhutto.” I will probably like the Bhutto
movie. (I added it to my queue). The Netflix recommendations I've seen in the past have been fantastic. And
when someone recommends me a movie I've already seen, it's usually one I've really enjoyed.
How does Netflix do that? Is there some massive team of interns at corporate headquarters who used a
combination of Google and interviews with my family and friends to determine whether I would like a
documentary about a former Pakistani prime minister? Of course not. Netflix has simply mastered some very
sophisticated statistics. Netflix doesn't know me. But it does know what movies I liked in the past (because I
rated them). Using that information, along with other customers' ratings and a powerful computer, Netflix can
make surprisingly accurate predictions about my tastes.
I'll get back to Netflix's specific algorithm for making these selections; for now, the important thing is that it's
all based on correlation. Netflix recommends movies similar to other movies I've liked; It also recommends
movies that have been highly rated by other customers whose ratings are similar to mine.
Bhutto was recommended because of my five-star ratings for two other documentaries, Enron: The Smartest
Guys in the Room and Fog of War.
Correlation measures the degree to which two phenomena are related to each other. For example, there is
a correlation between summer temperatures and ice cream sales. When one goes up, so does the other. Two
variables are positively correlated if a change in one is associated with a change in the other in the same
direction, such as the ratio between height and weight. Taller people weigh more (on average); shorter people
weigh less. A correlation is negative if a positive change in one variable is associated with a negative change
in the other, such as the relationship between exercise and weight.
The tricky thing about these types of associations is that not all observations fit the pattern. Sometimes
short people weigh more than tall people. Sometimes people who don't exercise are thinner than people who
exercise all the time.
Still, there is a significant relationship between height and weight, and between exercise and weight.

If we were to make a scatter plot of the heights and weights of a random sample of American adults, we would expect
to see something like this:
Machine Translated by
Google

Scatter plot for height and weight

Height (inches)

If we were to create a scatter plot of the association between exercise (measured in minutes of intensive exercise per
week) and weight, we would expect a negative correlation, with those who exercise more tending to weigh less. But a
pattern consisting of dots scattered across the page is a somewhat unwieldy tool. (If Netflix tried to make movie
recommendations to me by plotting ratings of thousands of movies from millions of customers, the results would bury
headquarters in scatterplots.) Instead, the power of correlation as a statistical tool is that we can summarize an
association between two variables in a single descriptive statistic: the correlation coefficient.

The correlation coefficient has two fabulously attractive features. First, for mathematical reasons that have been
relegated to the appendix, it is a single number ranging from –1 to 1. A correlation of 1, often described as perfect
correlation, means that every change in one variable is associated with an equivalent change in the other variable in the
same direction.

A correlation of –1, or perfect negative correlation, means that every change in one variable is associated with an
equivalent change in the other variable in the opposite direction.

The closer the correlation is to 1 or –1, the stronger the association. A correlation of 0 (or close) means that the
variables have no significant association with each other, such as the relationship between shoe size and SAT.
scores.

The second attractive feature of the correlation coefficient is that it has no associated units. We can
calculate the correlation between height and weight, even though height is measured in inches and weight in
pounds.
We can even calculate the correlation between the number of televisions high school students have in their
homes and their SAT scores, which I assure you will be positive. (More on that relationship in a moment.) The
correlation coefficient does something seemingly miraculous: it collapses a complex mess of data measured in
different units (like our scatterplots of height and weight) into a single, elegant descriptive statistic.
As?
Machine Translated by
Google

As always, I have included the most common formula for calculating the correlation coefficient in the
appendix at the end of the chapter. This is not a statistic you are going to calculate by hand. (After you have
entered the data, a basic software package such as Microsoft Excel will calculate the correlation between two
variables.) Still, intuition is not that difficult. The formula for calculating the correlation coefficient does the
following:

1. Calculate the mean and standard deviation of both variables. If we stick to the height and weight
example, we would know the average height of the people in the sample, the average weight of the
people in the sample, and the standard deviation for both height and weight.
2. Convert all data so that each observation is represented by its distance (in standard deviations) from
the mean. Stay with me; it's not that complicated. Suppose the mean height in the sample is 66 inches
(with a standard deviation of 5 inches) and the mean weight is 177 pounds (with a standard deviation of
10 pounds). Now suppose you are 72 inches tall and weigh 168 pounds. We can also say that his height
is 1.2 standard deviations above the mean for height [(72 – 66)/5)] and 0.9 standard deviations below
the mean for weight, or –0.9 for the purposes of the formula [(168 – 177)/10]. Yes, it is unusual for
someone to be above average in height and below average in weight, but since you paid good money
for this book, I thought I should at least make you tall and thin.

Notice that your height and weight, which were previously expressed in inches and pounds, have been
reduced to 1.2 and –0.9. This is what makes units disappear.
3. Here I will wave my hands and let the computer do the work. The formula then calculates the ratio of
height to weight for all individuals in the sample, measured in standard units. When individuals in the
sample are tall, say, 1.5 or 2 standard deviations above the mean, what does their weight tend to be,
measured in standard deviations from the mean?

the average for weight? And when individuals are close to the average in terms of height, what are their weights
measured in standard units?

If the distance from the mean of one variable tends to be broadly consistent with the distance from the mean of the
other variable (for example, people who are far from the mean for height in any direction also tend to be far from the
mean in the same direction for weight), then we would expect a strong positive correlation.
If the distance from the mean for one variable tends to correspond to a similar distance from the mean for the second
variable in the other direction (for example, people who are well above the mean in terms of exercise tend to be well
below the mean in terms of weight), then we would expect a strong negative correlation.
If two variables do not tend to deviate from the mean in any meaningful pattern (e.g., shoe size and
exercise), then we would expect little or no correlation.

You suffered a lot in that section; We'll be back to movie rentals soon.
Before we head back to Netflix, though, let's reflect on another aspect of life where correlation matters: the SAT. Yes, that
SAT. The SAT Reasoning Test, formerly known as the Scholastic Aptitude Test, is a standardized test consisting of three
Machine Translated by
Google

sections: mathematics, reading, and writing. You probably took the SAT, or will soon.
You probably haven't thought deeply about why you had to take the SAT. The purpose of the test is to measure academic
ability and predict college performance. Of course, one might reasonably ask (especially those who dislike standardized
tests): Isn't that what high school is for? Why is a four-hour test so important when college admissions officers have
access to four years of high school grades?
The answer to these questions is hidden in chapters 1 and 2. High school grades are an imperfect descriptive statistic.
A student who earns mediocre grades while taking a difficult schedule of math and science classes may have more
academic ability and potential than a student at the same school with better grades in less challenging classes.
Obviously, there are even greater potential discrepancies between schools. According to the College Board, which
produces and administers the SAT, the test was created to “democratize college access for all students.” That seems fine
to me. The SAT provides a standardized measure of ability that can be easily compared across all students applying to
college. But is it a good measure of skill? If we want a metric that can be easily compared across students, we could also
have all high school seniors run the 100-yard dash, which is cheaper and easier than administering the SAT. The
problem, of course, is that 100-yard dash performance is not correlated with college performance. It's easy to get the
data; it just won't tell us anything meaningful.

So how well is the SAT doing in this regard? Unfortunately for future generations of high school students, the SAT
does a reasonably good job of predicting freshman-year college grades. The College Board publishes the relevant
correlations. On a scale of 0 (no correlation) to 1 (perfect correlation), the correlation between high school GPA and
freshman college GPA is .56. (To put that in perspective, the correlation between height and weight for adult men in the
United States is about 0.4.) The correlation between the composite SAT _................................................ . .. X
................................... —1
The score (critical reading, math, and writing) and freshman college GPA is also .56. That would seem to be an argument
in favor of abandoning the SAT, since the test does not appear to predict college performance any better than high school
grades. In fact, the best predictor of all is a combination of SAT scores and high school GPA, which has a .64 correlation
with freshman-year college grades. I'm sorry.

A crucial point in this general discussion is that correlation does not imply causation; a positive or negative association
between two variables does not necessarily mean that a change in one of the variables is causing the change in the
other. For example, I earlier alluded to a likely positive correlation between a student's SAT scores and the number of
televisions his or her family owns. This does not mean that overly anxious parents can improve their children's test scores
by buying five additional televisions for the house. It also probably doesn't mean that watching a lot of TV is good for
academic performance.

The most logical explanation for such a correlation would be that highly educated parents can afford many televisions
and tend to have children with better than average outcomes. Both televisions and test scores are probably caused by a
third variable, which is parental education. I cannot prove the correlation between televisions in the home and SAT
scores. (The College Board does not provide such data.) However, I can show that students from wealthy families have
higher average SAT scores than students from less wealthy families. According to the College Board, students with family
incomes over $200,000 have an average SAT math score of 586, compared to an average SAT math score of 460 for
Machine Translated by
Google

students with family incomes of $20,000 or less. Meanwhile, it is also likely that
Families with incomes over $200,000 are more likely to have more televisions in their (multiple) households than families
with incomes of $20,000 or less.

I started writing this chapter many days ago. Since then, I have had the opportunity to watch the Bhutto documentary.
Wow! This is an extraordinary film about an extraordinary family. The original images, which span from the partition of
India and Pakistan in 1947 to the assassination of Benazir Bhutto in 2007, are extraordinary. Bhutto's voice is effectively
woven throughout the film in the form of speeches and interviews. Anyway, I gave the movie five stars, which is pretty
much what Netflix predicted.
At the most basic level, Netflix is exploiting the concept of correlation. First, I rate a set of movies. Netflix compares my
ratings to those of other customers to identify those whose ratings are highly correlated with mine. Those customers
usually like the movies I like. With this set up, Netflix can recommend movies that like-minded customers have rated
highly but that I haven't seen yet.
That's the "big picture." The actual methodology is much more complex. In fact, Netflix ran a contest in 2006 in which
the public was invited to design a mechanism that would improve Netflix's existing recommendations by at least 10
percent (meaning the system was 10 percent more accurate at predicting how a customer would rate a product). movie
after watching it).
The winner would take home $1,000,000.
Each individual or team that registered for the contest received “training data” consisting of more than 100 million
ratings of 18,000 movies from 480,000 Netflix customers. A separate set of 2.8 million ratings was “withheld,” meaning
Netflix knew how customers rated these movies, but contest entrants did not. Competitors were judged based on how
well their algorithms predicted actual customer opinions of these withheld films. Over three years, thousands of teams
from more than 180 countries submitted proposals. There were two requirements to enter. First, the winner had to license
the algorithm to Netflix. And second, the winner had to “describe to the world how you did it and why it works.” 3 In 2009,
Netflix announced a winner: a seven-person team made up of statisticians and computer scientists from the United
States,
Austria, Canada and Israel. Unfortunately, I cannot describe the winning system, not even in an appendix. The article
explaining the system is ninety-two pages long.

I am impressed by the quality of Netflix


recommendations. Still, the system is just a super-fancy variation on what people have been doing since the dawn of
cinema: finding someone with similar tastes and asking them for a recommendation. You tend to like what I like and
dislike what I don't like, so what did you think of George Clooney's new movie?

That is the essence of correlation.

APPENDIX TO CHAPTER 4 To calculate


the correlation coefficient between two sets of numbers, would perform the following steps, each of which is illustrated by
using data on heights and weights of 15 hypothetical students in the following table.

1. Convert each student's height to standard units: (height –


Machine Translated by
Google

mean)/standard deviation.
2. Convert each student's weight to standard units: (weight – mean)/standard deviation.
3. Calculate each student's product of (weight in standard units) × (height in standard units). You
should see that this number will be larger in absolute value when a student's height and weight are
relatively far from the mean.

4. The correlation coefficient is the sum of the products calculated above divided by the number of
observations (15 in this case). The correlation between height and weight for this group of students
is .83. Since the correlation coefficient can range from –1 to 1, this is a relatively high degree of positive
correlation, as would be expected with height and weight.

TO B C D AND F
Height in Weight in
Student Height Weight (Weight in standard units)
^randani units srantiard units

Nick 74 193 1.21 0.99 1.19

Flat 66 133 -0.63 -0.67 0.42

Dinah 68 155 -0.17 -0.06 0.01

Rebecca 69 147 0.06 -029 -04)2

Ben 73 175 0.98 0.49 0.48

Charu 70 128 0.29 -0.81 -0-24

Sahar 60 100 -2.00 -159 3.18

Maggie 63 128 -1.32 -0^1 1.07


Faisal 67 170 -0.40 0.35 -0.14

Ted 70 182 0.29 0.68 0.20

Nariso 70 178 0.29 0.57 0.17

Katrina 70 118 0.29 -1.09 -0.32

CJ 75 227 1.44 1.93 2.77


Sophia 62 115 -1.54 -1.17 1.81

W1U 74 211 UI 1.49 1.80

Mean 68.73 157.33 Totals 12.39


Standard
Deviation 436 3612 Correlation coefficient = Total/ n = 12 39/15 = 0.83

The formula for calculating the correlation coefficient requires a slight deviation from the notation. The
number ∑, known as the plus sign, is a useful character in statistics. It represents the sum of the quantity that
follows it. For example, if there is a set of observations x1 , x2 , x3 , and x4 , then ∑ (xi ) tells us to sum the
four observations: x1 + x2 + x3 + x4 . Therefore, ∑ (xi ) = x1 + x2 + x3 + x4 .
Our formula for the mean of a set of i observations

could be represented as follows: mean = ∑ (xi )/n.


Z (x:) —
We can make the formula even more adaptable by writing ii , which adds the quantity x1 + x2 + x3 + . . .
xn , or, in other words, all the terms beginning with x1 (because i = 1) through xn (because i = n). Our formula
for the mean of a set of n observations could be represented as follows:

mean =L “In
Machine Translated by
Google

Given this general notation, the formula for calculating the correlation coefficient, r, for two variables x and
y is as follows:

r _1 (xK)(yi-) i-1 G Gy
Machine Translated by
Google

where

n = the number of observations; is the mean of the variable


And x; is the mean of the variable
and; σx is the standard deviation of the variable
x; σy is the standard deviation of the variable y.

Any statistical software program with statistical tools can also calculate the correlation coefficient
between two variables. In the student height and weight example, using Microsoft Excel produces the
same correlation between height and weight for all fifteen students as the manual calculation in the
table above: 0.83.

* You can read it at [Link]


Machine Translated by
Google

CHAPTER 5

Basic probability
Don't buy the extended warranty on your $99 printer

In 1981, Joseph Schlitz Brewing Company spent $1.7 million on what seemed like a surprisingly bold and risky marketing
campaign for its ailing Schlitz brand. At halftime of the Super Bowl, in front of 100 million people around the world, the
company broadcast a live taste test in which Schlitz beer was pitted against 1 million people.
a key competitor, Michelob. Even bolder, the company did not choose drinkers
of beer at random to evaluate the two beers; he selected 100 Michelob drinkers.
This was the culmination of a campaign that spanned the NFL playoffs. The
2
A total of five live televised taste tests were conducted, each involving 100 consumers of a competing brand

(Budweiser, Miller or Michelob) performing a blind taste test between their supposed favourite beer and Schlitz. Each of
the beer tastings was aggressively promoted, as was the playoff game during which it would take place (e.g., “Watch
Schlitz v. “Bud, live during the AFC playoffs”).
The marketing message was clear: Even beer drinkers who think they like another brand will prefer Schlitz in a blind
taste test. For the Super Bowl venue, Schlitz even hired a former NFL referee to oversee the test. Given the risk involved
in conducting blind taste tests in front of large live TV audiences, you might assume Schlitz produced a spectacularly
delicious beer, right?
Not necessarily. Schlitz only needed one mediocre beer and a solid grasp of statistics to know that this ploy (a term I
don't use lightly, even when it comes to beer advertising) would almost certainly work in his favor. Most beers in the
Schlitz category taste pretty much the same; ironically, that is exactly the fact that this advertising campaign exploited.
Suppose the typical beer drinker on the street can't tell Schlitz from Budweiser from Michelob from Miller. In that case, a
blind taste test between any two beers is essentially a coin flip. On average, half of the tasters will choose Schlitz and the
other half will choose the beer they find “challenging.” This fact alone would probably not make an advertising campaign
particularly effective. (“You can’t tell the difference, so you might as well drink Schlitz.”) And Schlitz would by no means
want to conduct this test among its own loyal customers; about half of these Schlitz drinkers would choose the
competitor's beer. It looks bad when beer drinkers supposedly most committed to your brand choose a competitor in a
blind taste test, which is
Machine Translated by
Google

exactly what Schlitz was trying to do to its competitors.


Schlitz did something smarter. The genius of the campaign was to conduct the taste test exclusively among beer
drinkers, who stated that they preferred a competing beer. If the blind taste test is really just a coin flip, then about half of
Budweiser, Miller, or Michelob drinkers will end up choosing Schlitz. That makes Schlitz look really good. Half of Bud
drinkers like Schlitz better!
And it looks particularly good at halftime of the Super Bowl with a former NFL referee (in uniform) performing the test.
Still, it's live TV. Even if Schlitz's statisticians had determined through a bunch of private pre-testing that the typical
Michelob drinker would choose Schlitz 50 percent of the time, what if the 100 Michelob drinkers they took the test at the
Super Bowl halftime show up as quirky drinkers? Yes, blind taste testing is the equivalent of flipping a coin, but what if the
majority of tasters chose Michelob simply by chance? After all, if we lined up the same 100 guys and asked them to flip a
coin, they'd likely get 85 or 90 tails. That kind of bad luck in the taste test would be a disaster for the Schlitz brand (not to
mention a waste of the $1.7 million from live television coverage).

Statistics to the rescue! If there were some kind of statistics superhero, he would have broken into Schlitz corporate
headquarters and revealed the details of what statisticians call a binomial experiment (also called a Bernoulli trial). The
key features of a binomial experiment are that we have a fixed number of trials (e.g. e.g., 100 tasters), each with two
possible outcomes (Schlitz or Michelob), and the probability of "success" is the same in each
test. (I assume the probability of choosing one beer or the other is 50 percent, and define *.
this is
“success” as a taster choosing Schlitz.) We also assume that all “tests” are independent,
meaning that a blind taster’s decision has no impact on any other taster’s decision.

With just this information, a statistical superhero can calculate the probability of all the different outcomes of the 100
trials, such as 52 Schlitz and 48 Michelob or 31 Schlitz and 69 Michelob. Those of us who are not statistical superheroes
can use a computer to
do the same. The odds of the 100 Michelob tasters chose Schlitz, an impressive number given that all of the men who took
the live blind taste test had professed to be Michelob drinkers. It was very likely that a result at least as good would occur. If the
taste test is really like flipping a coin, then basic probability tells us that there was a 98 percent chance that at least 40 of the
tasters would choose Schlitz, and an 86 percent chance that at least 45 of the tasters would do so. † In theory, this wasn't a

taste testers
harvest were
very risky tactic at all.
in 1,267,650,600,228,229,401,496,703,205,376. There was probably a higher chance that all the testers
would be killed at halftime by an asteroid. More importantly, the same basic calculations can give us the
cumulative probability for a variety of outcomes, such as the chances that 40 or fewer raters will choose
Schlitz.
These figures would clearly have allayed the fears of Schlitz's marketing people.

Suppose Schlitz would have been happy if at least 40 of the 100


So what happened to Schlitz? At halftime of the 1981 Super Bowl, exactly 50 10 percent of Michelob drinkers chose Schlitz
in the blind taste test.
There are two important lessons here: probability is a remarkably powerful tool, and many of the top beers of the 1980s
were indistinguishable from one another. This chapter will focus primarily on the first lesson.
Machine Translated by
Google

Probability is the study of events and outcomes that involve an element of uncertainty. Investing in the stock market involves
uncertainty. The same thing happens when you flip a coin, which can come up heads or tails. Flipping a coin four times in a row
involves additional layers of uncertainty, because each of the four flips can result in heads or tails. If you flip a coin four times in
a row, I can't know for sure the outcome in advance (neither can you). However, I can determine in advance that some
outcomes (two heads, two tails) are more likely than others (four heads). As the Schlitz folks found, such probability-based
insights can be extremely useful. In fact, if you can understand why the probability of getting four heads in a row on a fair coin is
1 in 16, you can (with a little work) understand everything from how the insurance industry works to whether a professional
football team should kick the ball. extra point after a touchdown or seeking a two-point conversion.

Let's start with the easy part: many events have known probabilities. The probability of getting heads on a fair coin is ½. The
probability of rolling a one on a single die is Other events have probabilities that can be inferred on the basis of past data. The
probability of successfully kicking an extra point after a touchdown in professional football is 0.94, meaning that kickers make,
on average, 94 of every 100 extra point attempts. (Obviously, this figure may vary slightly for different kickers, under different
weather circumstances, etc., but it's not going to change radically.) Simply having and appreciating this type of information can
often clarify decision-making and make risks explicit.

For example, the Australian Transport Safety Board published a report quantifying the death risks for different modes of
transport. Despite widespread fear of flying, the risks associated with commercial air travel are minimal. Australia has not had a
fatality on commercial aircraft since the 1960s, so the death rate per 100 million kilometres travelled is essentially zero. The rate
for drivers is 0.5 fatalities.

for every 100 million kilometers traveled. The really impressive number is that of motorcycles, if you aspire to be an organ
donor. The mortality rate is thirty-five times higher among
.... ......................... 3
motorcycles than between cars.
In September 2011, a 6.5-ton NASA satellite plummeted to Earth and was expected to break up once it hit the Earth's
atmosphere. What were the chances of being hit by debris? Should I have kept the kids at home and not going to school?
NASA space scientists estimated that the probability of a particular person being hit by a piece of the falling satellite was
1 in 21 trillion.
However, the chances that any person anywhere on Earth could be
*
reached were 1 in 3.20A01. In the end, the satellite broke up on re-entry, but scientists are not entirely sure where all the
pieces ended up. be hurt. The 4 no-reported probabilities don't tell us for certain what will happen; they tell us what is
likely to happen and what is less likely to happen. Sensible people can use these types of numbers in business and in life.
For example, when you hear on the radio that a satellite is falling to Earth, you should not rush home on a motorcycle to
warn your family.

When it comes to risk, our fears don't always match what the numbers tell us we should fear. One of the
most surprising findings of Freakonomics, by Steve Levitt and Stephen Dubner, was that backyard swimming
pools are far more dangerous than guns in the closet. 5 Levitt and Dubner estimate that a child under the age
Machine Translated by
Google

of ten is one hundred times more likely to die in a swimming pool than in a firearm accident. † An intriguing
paper by three Cornell researchers, Garrick Blalock, Vrinda Kadiyali, and Daniel Simon, found that thousands
of Americans may have died since the 9/11 attacks because they were afraid of flying. Know that driving is
dangerous. As more Americans chose to drive rather than fly after 9/11, there were an estimated 344
additional traffic deaths per month in October, November, and December 2001 (taking into account the
average number of deaths and other factors that normally contribute to death). traffic accidents, such as
weather). This effect dissipated over time, presumably as fear of terrorism diminished, but the study's authors
estimate that the 9/11 attacks may have caused more than 2,000 driving deaths.

Sometimes probability can also tell us after the fact what probably happened and what probably did not happen, as in the
case of DNA analysis. When CSI: Miami technicians find a trail of saliva in an apple core near a murder victim, that saliva
doesn't have the killer's name on it, even when a very attractive technician views it under a powerful microscope. Instead,
saliva (or hair, skin, or a bone fragment) will contain a segment of DNA. Each DNA

The segment, in turn, has regions, or loci, that can vary from one individual to another (except in the case of
identical twins, who share the same DNA). When the medical examiner reports that a DNA sample is a
“match,” that is only part of what the prosecution has to prove. Yes, the loci analyzed in the DNA sample from
the crime scene must match the loci in the DNA sample taken from the suspect. However, prosecutors must
also prove that the match between the two DNA samples is not a mere coincidence.
Humans share similarities in their DNA, just as we share other similarities: shoe size, height, eye color.
(More than 99 percent of all DNA is identical among all humans.) If researchers have access to only a small
sample of DNA in which only a few loci can be analyzed, it is possible that thousands or even millions of
individuals share that genetic fragment. . Therefore, the more loci that can be tested and the more natural
genetic variation there is at each of those loci, the more certain the match will be. Or, to put it another way, the
DNA sample is less likely to match more than one person. 7

To understand this, imagine that your “DNA number” consists of your phone number attached to your
Social Security number. This nineteen-digit sequence uniquely identifies you. Think of each digit as a “spot”
with ten possibilities: 0, 1, 2, 3, etc. Now suppose crime scene investigators find the remnant of a “DNA
number” at the crime scene: 4 5 9 4 0 This exactly matches 9co8n1s7u “number . - -
of DNA”. Are you guilty?
You should see three things. First, anything other than a complete genome-wide match leaves some room
for uncertainty. Second, the more “loci” that can be tested, the less uncertainty remains. And third, context
matters. This coincidence would be extremely convincing if you were also caught speeding away from the
crime scene with the victim's credit cards in your pocket.
When researchers have unlimited time and resources, the typical process involves testing thirteen different
loci. The chances of two people sharing the same DNA profile at all thirteen loci are extremely low. When DNA
was used to identify remains found at the World Trade Center after 9/11, samples found at the scene were
compared to samples provided by relatives of the victims. The probability required to establish a positive
identification was one in a billion, meaning that the probability that the discovered remains belonged to
someone other than the identified victim had to be judged to be one in a billion or less. Later in the search, this
rule was relaxed as there were fewer unidentified victims for the remains to be confused with.
Machine Translated by
Google

When resources are limited, or the available DNA sample is too small or too contaminated to analyze
thirteen loci, things become more interesting and controversial. The Los Angeles Times published a series in
2008 examining the use of DNA as criminal evidence. 8 In particular, the Times questioned whether the
probabilities typically used by law enforcement underestimate the likelihood of coincidental matches. (Since no
one knows the DNA profile of the entire population, the probabilities presented in court by the FBI and other
law enforcement entities are estimates.) The intellectual backlash was instigated when a crime lab analyst in
Arizona running tests against the state's DNA database discovered two unrelated criminals whose DNA
matched at nine loci; according to the FBI, the chances of a nine-loci match between two unrelated people are
1 in 113 billion. Subsequent searches in other DNA databases yielded more than a thousand human pairs with
genetic matches at nine or more loci. I will leave this issue for defense attorneys and law enforcement to
resolve. For now, the lesson is that the dazzling science of DNA analysis is only as good as the probabilities
used to back it up.

It is often extremely valuable to know the probability of multiple events occurring. What is the probability that
the power goes out and the generator does not work? The probability of two independent events occurring is
the product of their respective probabilities. In other words, the probability of Event A and Event B occurring is
the probability of Event A multiplied by the probability of Event B. An example makes it much more intuitive. If
the probability of getting heads on a fair coin is ½, then the probability of getting heads twice in a row is ½ × ½,
or ¼. The probability of getting three heads in a row is ⅛, the probability of four heads in a row is 1/16, and so
on. (You should see that the probability of getting four crosses in a row is also 1/16.) This explains why your
school or office system administrator is constantly checking in on you to improve the “quality” of your
password. If you have a six-digit password using only numeric digits, we can calculate the number of possible
passwords: 10 × 10 × 10 × 10 × 10 × 10, which is equal to 10 or 1,000,000. These seem like a lot of
possibilities, but a computer could still do all 1,000,000 possible combinations in a fraction of a second.

So, let's say your system administrator badgers you enough to include letters in your password.

At this point, each of the 6 digits now has 36 combinations: 26 letters and 10 digits. The number of possible passwords exceeds two billion. If your size grows to

36 × 36 × 36 × 36 × 36 × 36, or 36, the administrator requires eight digits and urges you to 6

,
University of Chicago, the number of potential passwords is 8. go up use symbols like #, @, % and !, as does the a 46 or just over 20 trillion.

There is a crucial distinction here. This formula is applicable only if the events are independent,
meaning that the outcome of one has no effect on the outcome of another. For example, the probability that
you get heads on the first toss does not change the probability that you get heads on the second toss. On the
other hand, the probability of rain today is not independent of whether it rained yesterday, since storm fronts
can last for days. Similarly, the probability of crashing your car today and crashing your car next year are not
independent. Whatever caused your failure this year could also cause your failure next year; you may be
prone to drunk driving, drag racing, texting while driving, or just plain bad driving. (This is why your car
insurance rates go up after an accident – it's not simply that the company wants to recover the money it paid
on the claim; rather, it now has new information about your likelihood of being in an accident in the future. ,
who, after having driven through the garage door, has climbed in).
Suppose you are interested in the probability of one event or the other occurring: outcome A or outcome B
Machine Translated by
Google

(again assuming they are independent). In this case, the probability of getting A or B consists of the sum of
their individual probabilities: the probability of A plus the probability of B.
For example, the probability of rolling a 1, 2, or 3 on a single die is the sum of its individual probabilities: + + =
= ½. This should make intuitive sense. There are six possible outcomes when rolling a die. The numbers 1, 2,
and 3 together represent half of those possible outcomes. So you have a 50 percent chance of rolling a 1, 2,
or 3. If you are playing craps in Las Vegas, the probability of rolling a 7 or 11 on a single roll is the number of
combinations that add up to 7 or 11 divided. by the total number of combinations that can be rolled with two
dice, or * %e .
Probability also allows us to calculate what could be the most useful tool in all managerial decision making,
particularly in finance: expected value. The expected value takes basic probability one step further. The
expected value or payoff of some event, for example the purchase of a lottery ticket, is the sum of all the
different outcomes, each weighted by its probability and payoff. As always, an example makes this clear.
Suppose you are invited to play a game in which you roll a single die. The payout for this game is $1 if you get
a 1; $2 if you get a 2; $3 if you get a 3; and so on. What is the expected value for a single roll of the die? Each
possible outcome has a probability, so the expected value is: ($1) +
% ($2) + ($3) + ($4)1+ ($5) + ($6) = or $3.50. % 2%/6
,
At first glance, the expected value of $3.50 might seem like a relatively meaningless figure. After all, you
can't actually win $3.50 on a single roll of the dice (since your payout has to be a whole number). In fact, the
expected value turns out to be extremely powerful because it can tell you whether a particular event is
“fair”, given its price and the expected result. Suppose you have the opportunity to play the above game for $3
per spin. Does it make sense to play? Yes, because the expected value of the outcome ($3.50) is greater than
the cost of playing ($3.00). This doesn't guarantee that you'll make money playing once, but it does help clarify
which risks are worth taking and which aren't.
We can take this hypothetical example and apply it to professional football. As noted above, after a
touchdown, teams have the option of kicking an extra point or attempting a two-point conversion. The former
involves kicking the ball through the uprights from the three-yard line; the latter involves running or passing
into the end zone from the three-yard line, which is significantly more difficult. Teams can choose the easy
option and get one point, or they can choose the harder option and get two points. To do?
Statisticians can't play football or date cheerleaders, but they can provide statistical guidance to football
coaches. 9 As noted above, the probability of making the kick after a touchdown is 0.94. This means that the
expected value of a point after the attempt is also 0.94, since it equals the reward (1 point) multiplied by the
probability of success (0.94). No team ever scores 0.94 points, but this number is useful in quantifying the
value of attempting this option after a touchdown relative to the alternative, which is the two-point conversion.
The expected value of “going for two” is much lower: .74. Yes, the reward is higher (2 points), but the success
rate is dramatically lower (0.37). Obviously, if there is one second left in the game and a team is two points
behind after scoring a touchdown, it has no choice but to go for a two-point conversion. But if a team's goal is
to maximize points scored over time, then kicking the extra point is the strategy that will accomplish that.
The same basic analysis can illustrate why you should never buy a lottery ticket. In Illinois, the odds
associated with the game's various possible payouts are printed on the back of each ticket. I bought a $1
Machine Translated by
Google

instant ticket. (Personal note: Is this tax deductible?) On the back, in very, very small print, are the odds of
winning various cash prizes or a free new ticket: 1 in 10 (free ticket); 1 in 15 ($2); 1 in 42.86 ($4); 1 in 75 ($5);
and so on up to the 1 in 40,000 chance of winning $1,000. I calculated the expected payout for my instant
ticket by adding up each possible cash prize weighted by its*pRroebsaulbtailiqduaedm. A $1 lottery ticket has an
expected payout of about $0.56, making it an absolutely miserable way to spend $1. As luck would have it, I won $2.
Despite my $2 prize, buying the ticket was stupid. This is one of the crucial lessons of probability. Good
decisions (as measured by underlying probabilities) can turn out to be bad. And bad decisions, like spending
A dollar in the Illinois Lottery can still pay off, at least in the short term. But in the end probability triumphs. An
important theorem known as the law of large numbers tells us that as the number of trials increases, the
average of the results will get closer and closer to its expected value. Yes, I won $2 playing the lottery today.
And tomorrow I could make $2 again. But if I buy thousands of $1 lottery tickets, each with an expected payout
of $0.56, then I have a near-mathematical certainty that I will lose money. By the time I've spent a million
dollars on tickets, I'll end up with something surprisingly close to $560,000.
The law of large numbers explains why casinos always make money in the long run. The odds associated
with all casino games favor the house (assuming the casino can successfully prevent blackjack players from
counting cards). If enough bets are placed over a long enough time, the casino will surely win more than it
loses. The law of large numbers also demonstrates why Schlitz found it much better to conduct 100 blind taste
tests at halftime of the Super Bowl rather than just 10. See "Probability Density Functions" for a Schlitz-type
test with 10, 100, and 1000. essays.
(Though it sounds fancy, a probability density function simply plots the varying outcomes along the x-axis and
the expected probability of each outcome on the y-axis; the weighted probabilities (each outcome multiplied by
its expected frequency) will add up to 1.) Again, I'm assuming that the taste test is like a coin toss, and that
each taster has a 0.5 chance of choosing Schlitz. As you can see below, the expected result converges
around 50 percent of tasters choosing Schlitz as the number of tasters increases. At the same time, the
probability of obtaining a result that deviates markedly from 50 percent decreases dramatically as the number
of trials increases.

10 tests

100 tests
Machine Translated by
Google

1,000 essays

I previously stipulated that Schlitz executives would be happy if 40 percent or more of Michelob drinkers
chose Schlitz in the blind taste test. The following figures reflect the probability of obtaining that result as the
number of tasters increases:

10 blind tasters: 0.83


100 blind tasters: 0.98
1,000 blind tasters: .9999999999
1,000,000 blind tasters: 1

By now, the intuition behind the chapter subtitle is obvious: “Don’t buy the extended warranty on your $99
printer.” Okay, maybe that's not so obvious. Let me back up. The entire insurance industry is based on
probability. (A warranty is just a form of insurance.) When you insure something, you agree to receive specific
compensation in the event of a clearly defined contingency. For example, your car insurance will replace your
car if it is stolen or crushed by a tree. In exchange for this guarantee, you agree to pay a fixed amount of
money during the period in which you are insured. The key idea is that in exchange for a regular, predictable
payment, you have transferred the risk of having your car stolen, crushed or even totaled to the insurance
company.

because of your own poor driving.


Why are these companies willing to take such risks? Because they will make huge profits in the long run if they price
Machine Translated by
Google

their premiums correctly.


Obviously, some cars insured by Allstate will be stolen. Others will be destroyed when their owners drive over a fire
hydrant, as happened to my high school girlfriend. (He also had to replace the fire hydrant, which is a lot more expensive
than you might think.) But most cars insured by Allstate or any other company will be fine. To make money, the insurance
company only needs to collect more premiums than it pays out in claims. And to do so, the company must have a solid
understanding of what is known in industry jargon as the “expected loss” of each policy. This is exactly the same concept
as expected value, just with a touch of insurance. If your car is insured for $40,000 and the chances of it being stolen in a
given year are 1 in 1,000, then the expected annual loss on your car is $40. The annual premium for the theft coverage
portion must be more than $40.
At that point, the insurance company becomes like the casino or the Illinois lottery. Yes, there will be payments, but in the
long run what comes in will be more than what goes out.

As a consumer, you must recognize that insurance will not save you money in the long run. What it will do is prevent
unacceptably high losses, such as replacing a $40,000 car that was stolen or a $350,000 house that burned down.
Buying insurance is a statistically “bad bet” because, on average, you will pay the insurance company more than you will
receive in return. However, it can still be a sensible tool to protect yourself against outcomes that would otherwise ruin
your life. Ironically, someone as rich as Warren Buffett can save money by not buying car insurance, home insurance, or
even health insurance because he can afford anything bad that might happen to him.
Which finally brings us back to your $99 printer! We will assume that you have
TO...................... ■ ., r. - .- *
I just picked out the perfect new laser printer at Best Buy or some other retailer.
When you get to the checkout, the sales assistant will offer you a number of extended warranty options. For another $25
or $50, Best Buy will repair or replace the printer if it breaks in a year or two. Based on your understanding of probability,
insurance, and basic economics, you should be able to immediately assume all of the following: (1) Best Buy is a for-profit
company that seeks to maximize profits. (2) The sales assistant is eager for you to purchase the extended warranty. (3)
From numbers 1 and 2, we can infer that the cost of the warranty to you is higher than the expected cost of fixing or
repairing the printer to Best Buy. If this were not the case, Best Buy would not be so aggressive in trying to sell it to you.
(4) If your $99 printer breaks down and you have to pay out of pocket to fix or replace it, this will not significantly change
your
life.
On average, you will pay more for an extended warranty than you would pay to have the printer repaired. The broader
lesson (and one of the central lessons of personal finance) is that you should always insure yourself against any adverse
contingency that you cannot comfortably bear. You should avoid buying insurance for everything else.

Expected value can also help us untangle complex decisions involving many contingencies at different times. Suppose a
friend of yours asks you to invest $1 million in research examining a new cure for male pattern baldness. You would
probably wonder what the probability of success would be; you will get a complicated answer. This is a research project,
so there's only a 30 percent chance the team will discover a cure that works. If the team does not find a cure, you will get
back $250,000 of your investment, as those funds will have been set aside to bring the drug to market (testing, marketing,
Machine Translated by
Google

etc.). Even if researchers succeed, there is only a 60 percent chance that the U.S.
The Food and Drug Administration will approve the new miracle cure for baldness as safe for use in humans. Even then,
if the drug is safe and effective, there is a 10 percent chance that a competitor will come to market with a better drug at
about the same time, wiping out any potential benefit. If all goes well (the drug is safe, effective, and uncompetitive), then
the best estimate of the return on its investment is $25 million.
Should you make the investment?
This seems like a mess of information. The potential profit is huge (25 times your initial investment), but there are
many potential dangers. A decision tree can help organize this type of information and, if the probabilities associated with
each outcome are correct, give you a probabilistic assessment of what you should do. The decision tree maps out each
source of uncertainty and the probabilities associated with all possible outcomes. The end of the tree gives us all the
possible payoffs and the probability of each. If we weight each payoff according to its probability and add up all the
possibilities, we get the expected value of this investment opportunity. As always, the best way to understand this is to
take a look.

(.3)(.6)(.9)($25
million)
= (.162)($25 million)
= $4,050,000

(,3)(.6)(.l)(S0)
= 0.018 (SO)
= $0
(.3)(.4)($0)
= 0.12 ($0)
= S0

= (.7)
($250,000)
= $175,000
Expected payoff = $4,050,000 + $0 + Í0 +
$175,000
= $4,225,000
The investment decision
This particular opportunity has an attractive expected value. The weighted payment is $4.225 billion. Still, this investment
may not be the smartest thing to do with the college tuition money you've saved for your children.
The decision tree lets you know that your expected profit is much greater than what you are asked to invest. On the other hand,
the most likely outcome, meaning the one that will happen most often, is that the company will not discover a cure for baldness
and you will only recover $250,000. Your appetite for this investment may depend on your risk profile. The law of large numbers
suggests that an investment firm, or a wealthy individual like Warren Buffet, should look for hundreds of opportunities like this
one with uncertain outcomes but attractive expected returns. Some will work; many will not. On average, these investors will
make a lot of money, just like an insurance company or a casino.
If the expected reward is in your favor, it is always better to run more tests.
Machine Translated by
Google

The same basic process can be used to explain a seemingly counterintuitive phenomenon. Sometimes it doesn't make
sense to screen the entire population for a rare but serious disease, such as HIV/AIDS. Suppose we can test for some rare
disease with a high degree of accuracy. As an example, suppose the disease affects 1 in 100,000 adults and the test is 99.9999
percent accurate. The test never produces a false negative (meaning it never misses someone who has the disease); however,
about 1 in 10,000 tests performed on a healthy person will produce a false positive, meaning the person tests positive but does
not actually have the disease. The surprising result here is that, despite the impressive accuracy of the test, most people who
test positive will not have the disease. This will create enormous anxiety among those who test false positive; it may also waste
finite healthcare resources on follow-up testing and treatment.

If we look at the entire US adult population, or roughly 175 million people,


Machine Translated by
Google

the decision tree looks like this:

Widespread detection of a rare disease

Have disease and test


positive

Have disease and test


negative
Do not have disease
and test positive

Do not have disease and test negative

1,750
People with illness 1,750 = .09 = 9%
19,250
Those told them 1,750 + 17,500 have the disease

Only 1,750 adults suffer from the disease. They all test positive. More than 174 million adults do not suffer
from the disease. Of this healthy group that takes the test, 99.9999 get the correct result that they do not have
the disease. Only 0.0001 get a false positive. But 0.0001 out of 174 million is still a big number. In fact, an
average of 17,500 people will get false positives.
Let's see what that means. A total of 19,250 people are reported to be suffering from the disease; only 9
percent of them are actually sick! And that's with a test that has a very low false positive rate. Without getting
too off topic, this should give you an idea of why cost containment in health care sometimes means less
screening for disease in healthy people, not more. In the case of a disease like HIV/AIDS, public health
officials often recommend that available resources be used to screen higher-risk populations, such as gay men
or intravenous drug users.

Sometimes probability helps us by pointing out suspicious patterns. Chapter 1 introduced the problem of
institutionalized cheating in standardized testing and one of the companies eradicating it, Caveon Test
Security. The Securities and Exchange Commission (SEC), the government agency responsible for enforcing
federal laws related to securities trading, uses a similar methodology to catch inside traders. (Insider trading
involves the illegal use of private information, such as a law firm's knowledge of an impending corporate
takeover, to trade stock or other securities of the targeted companies.) The SEC uses powerful computers to
examine hundreds of millions of stock transactions and look for suspicions. 10 The SEC will also investigate
Machine Translated by
Google

investment managers' disappointing earnings. with unusually high returns over long periods of time. (Both
economic theory and historical data suggest that it is extremely difficult for a single investor to average
outperformance year after year.) Of course, smart investors are always trying to anticipate good and bad news
and devise perfectly legal strategies that will consequently beat the market. Being a good investor does not
necessarily make you a criminal. How does a computer tell the difference? I called the SEC's enforcement
division several times to ask what particular patterns are most likely to indicate criminal activity. They still
haven't called me back.

In the 2002 film Minority Report, Tom Cruise plays a "pre-crime" detective who is part of an office that uses
technology to predict crimes before they are committed.
Well friends, that's not science fiction anymore. In 2011, the New York Times
11 published the
following headline: “Send the police before there is a crime.” The story described how detectives were sent to
a parking lot in downtown Santa Cruz by a computer program that predicted there was a high probability of car
break-ins at that location that day. Police later arrested two women who were looking out of the car windows.
One had outstanding arrest warrants; the other was carrying illegal drugs.
The Santa Cruz system was designed by two mathematicians, an anthropologist and a criminologist. The
Chicago Police Department has created an entire predictive analytics unit, in part because gang activity, the
source of much of the city's violence, follows certain patterns. The book Data Mining and Predictive Analysis:
Intelligence Gathering and Crime Analysis, a statistics guide for law enforcement, begins enthusiastically: “It is
now possible to predict the future when it comes to crime, such as identifying crime trends, anticipating hot
spots in the community, refining resource deployment decisions, and ensuring the greatest protection for
citizens in the most efficient manner.”
(Look, I read this kind of stuff so you don't have to.)
“Predictive policing” is part of a broader movement called predictive analytics. Crime will always involve an element of
uncertainty, as will determining who will crash your car or default on your mortgage. Probability helps us avoid those risks. And
information refines our understanding of the relevant probabilities. Companies facing uncertainty have always sought to quantify
their risks. Lenders ask for things like income verification and credit score. However, these blunt lending instruments are
beginning to look like the predictive equivalent of a caveman's stone tools. The confluence of massive amounts of digital data
and cheap computing power has generated fascinating insights into human behavior. Insurance officials correctly describe their
business as “risk transfer,” so they should better understand the risks being transferred to them.

Companies like Allstate are dedicated to learning about things that might otherwise seem like random trivia:

• Drivers between the ages of twenty and twenty-four are the most likely to be involved in a traffic accident.
fatal accident. • The most commonly stolen car in Illinois is the Honda Civic (as opposed to , . . -,.*
Chevrolet full-size pickup trucks in
Alabama). • Texting while driving causes accidents, but state laws prohibiting the practice don't seem to stop drivers from
doing it. In fact, such laws could even make matters worse by encouraging drivers to hide their phones and thus take
their eyes off the road while texting.
Machine Translated by
Google

Credit card companies are at the forefront of this type of analysis, both because they know a lot about our spending habits
and because their business model relies heavily on finding customers who are hardly a good credit risk. (Customers who pose
the highest credit risks tend to be losers because they pay their bills in full each month; customers who carry large balances at
high interest rates are the biggest winners, as long as they don't default on their payments.) .) One of the most intriguing studies
of who is likely to pay a bill and who is likely to walk away was conducted by JP Martin, “a math-loving executive” at Canadian
Tire, a large retailer that sells a wide range of 13 When Martin analyzed the data: automotive products and other retail goods.
every transaction made with a previous one: found that what customers purchased was a remarkably accurate predictor of their
subsequent payment behaviour when used in conjunction with traditional tools like income and credit history.

A New York Times article titled “What Does Your Credit Card Company Know About You?” described some of Martin’s most
intriguing

Findings: “People who bought cheap, generic car oil were much more likely to skip a credit card payment than someone
who bought expensive, brand-name oil. People who bought carbon monoxide monitors for their homes or those little felt
pads that keep chair legs from scratching floors almost never defaulted. Anyone who bought a car accessory with a
chrome skull or a 'Mega Thruster exhaust system' was very likely to eventually default on their bill.”

Probability gives us tools to face the uncertainties of life. You shouldn't play the lottery. You should invest in the stock
market if you have a long-term investment horizon (because stocks typically have the best long-term returns). You should
get insurance for some things, but not for others. Probability can even help you maximize your winnings on game shows
(as will be shown in the next chapter).
That said (or written), probability is not deterministic. No, you shouldn't buy a lottery ticket, but you could still win
money if you do. And yes, probability can help us catch cheaters and criminals, but when used inappropriately it can also
send innocent people to jail. That's why we have Chapter 6.

1 I have in mind “The Six Sigma Man.” The lowercase Greek letter sigma, σ, represents the standard deviation.
Six Sigma Man is six standard deviations above the norm in terms of statistical ability, strength, and intelligence.
2 For everyone
For these calculations, I used a handy online binomial calculator, at [Link]

1 NASA also noted that even falling space debris is government property. Apparently it's illegal to keep a satellite souvenir, even if it lands in your
backyard.
2 Levitt and Dubner's calculations are as follows. Each year, approximately 550 children under the age of ten drown and 175 children under the age of
ten die in firearm accidents. The rates they compare are 1 drowning per 11,000 residential swimming pools compared to 1 firearm death per “more than
one million” guns. For teenagers, I suspect the numbers may change dramatically, because they are better swimmers and more likely to cause tragedy
if they come across a loaded gun. However, I have not checked the data on this point. * There are 6 ways to roll a 7 with two dice: (1,6); (2,5); (3,4);
(6,1); (5,2); and (4,3). There are only 2 ways to roll an 11: (5,6) and (6,5).

Meanwhile, there are 36 possible rolls in total with two dice: (1,1); (1,2); (1,3); (1,4); (1,5); (1,6). And (2,1); (2,2); (2,3); (2,4); (2,5); (2,6). And (3,1);
(3,2); (3,3); (3,4); (3,5); (3,6). And (4,1); (4,2); (4,3); (4,4); (4,5); (4,6). And (5,1); (5,2); (5,3); (5,4); (5,5); (5,6). And finally, (6,1); (6,2); (6,3); (6,4); (6,5);
and (6,6).
So the probability of rolling a 7 or an 11 is the number of possible ways to roll either of those two numbers divided by the total number of possible
rolls with two dice, which is 8/36. By the way, much of the previous research on probability was done by gamblers to determine exactly this sort of thing.
1 The total expected value for the $1 Illinois Dugout Doubler ticket (rounded to the nearest cent) is as follows: 1/15 ($2) + 1/42.86 ($4) + 1/75 ($5) +
Machine Translated by
Google

1/200 ($10) + 1/300 ($25) + 1/1,589.40 ($50) + 1/8,000 ($100) + 1/16,000 ($200) + 1/48,000 ($500) + 1/40,000 ($1,000) = $0.13 + $0.09 + $0.07 +
$0.05 + $0.08 + $0.03 + $0.01 + $0.01 + $0.01 + $0.03 = $0.51. However, there is also a 1/10 chance of getting a free ticket, which has an expected
payout of $0.51, so the overall expected payout is $0.51 + 0.1 ($0.51) = $0.51 + $. 05 = $.56.
2 Earlier in the book I used an example involving drunk employees producing faulty laser printers. You'll have to forget about that example here and
assume that the company has fixed its quality issues.
3 Since I've warned you to be rigorous with descriptive statistics, I feel compelled to point out that the most commonly stolen car is not necessarily the
type of car that is most likely to be stolen. A high number of Honda Civics are reported stolen because there are so many of them on the road; the
chances of any given Honda Civic being stolen (which is what car insurance companies care about) can be quite low. Conversely, even if 99
percent of all Ferraris were stolen, Ferraris would not make the “most commonly stolen” list, because there just aren’t that many of them to
steal.
Machine Translated by
Google

CHAPTER 5½

The Monty Hall Problem

The “Monty Hall problem” is a famous probability-related puzzle faced by contestants on the
game show Let's Make a Deal, which premiered in the United States in 1963 and is still
broadcast in some markets around the world. (I remember watching the show every time I was
home sick since elementary school.) The program's gift to statisticians was described in the
introduction. At the end of each day's show, one contestant was invited to stand alongside
presenter Monty Hall in front of three large doors: Door No. 1, door no. 2, and Door No. 3. Monty
explained to the contestant that there was a very desirable prize behind one of the doors and a
goat behind the other two doors. The player chose one of the three doors and received as a prize
whatever was behind it. (I don't know if the participants actually kept the goat; for our purposes,
let's assume that most players preferred the new car.)

The initial probability of winning was simple. There were two goats and a car. As the
participant stood in front of the doors with Monty, he or she had a 1 in 3 chance of choosing
which door would open to reveal the car. But as noted above, Let's Make a Deal had a twist,
which is why the show and its host have been immortalized in probability literature. After the
contestant chose a door, Monty would open one of the two doors that the contestant had not
chosen, always revealing a goat. At that point, Monty would ask the contestant if he or she would
like to change his or her choice: move from the locked door he or she had originally chosen to the
other remaining locked door.

As an example, suppose the contestant originally chose door no. 1. Monty would then open
door no. 3; a live goat would be standing there on a stage. There would still be two closed doors,
numbers. 1 and 2. If the valuable prize were behind the no. 1, the contestant would win; if he was
behind no. 2, would lose. It was then that Monty would turn to the player and ask if he would like
to change his mind and switch gates from no. 1 to no. 2 in this case.
Remember, both doors are still closed. The only new information the contestant has received is
that a goat appeared behind one of the doors that he did not open.

Should I change?
Yeah. The contestant has a 1/3 chance of winning if he keeps his initial. choice and 2/3
chance of winning if it changes. If you don't believe me read on.
I admit that this answer seems completely counterintuitive at first. It would seem that the
contestant has a one-third chance of winning no matter what he does.
Machine Translated by
Google

There are three closed doors. At the beginning, each door has a one in three chance of winning
the valuable prize. What does it matter if I move from one closed door to another?

The answer lies in the fact that Monty Hall knows what is behind every door. If the contestant
chooses Door No. 1 and there is a car behind, then Monty can open the no. 2 or not. 3 to show a
goat.
If the contestant chooses Door No. 1 and the car is behind no. 2, then Monty opens no. 3.
If the contestant chooses Door No. 1 and the car is behind no. 3, then Monty opens no. 2.
By switching after opening a door, the contestant gains the benefit of choosing two doors
instead of one. I will try to persuade you in three different ways that this analysis is correct.
The first is empirical. In 2008, New York Times columnist John Tierney wrote about the Monty Hall
phenomenon. feature that The Times later built into an interactive site. which allows you to play the game

yourself, including deciding whether or not to change. (There are even goats and cars coming out
from behind the doors.) The game tracks your success when you change doors after making your
initial decision compared to when you don't. * Try it yourself.
I paid one of my sons to play it 100 times, changing each time. I paid his brother to play 100 times without
changing. The one who changed won 72 times; the one who did not change won 33 times. They both
received two dollars for their efforts.
Data from Let's Make a Deal episodes suggests the same.
According to Leonard Mlodinow, author of The Drunkard's Walk, contestants who changed their
choice earned about twice as much as those who did.
2 no.
My second explanation comes down to intuition. Suppose the rules were changed slightly.
Suppose the contestant starts by choosing one of three doors: no. 1, no. 2, or not. 3, as it is
normally played. But then, before any door opens to reveal a goat, Monty says, "Would you like
to give up your choice in exchange for the other two doors you didn't choose?"
So if you chose Gate No. 1, you could get rid of that door in exchange for what's behind the no. 2
and no. 3. If you chose no. 3, you could change to no. 1 and no. 2.
Etc.
That wouldn't be a particularly difficult decision. Obviously you will have to give up one door in
exchange for two, as it increases your chances of winning from 1/3 to 2/3. Here's the intriguing
part: that's exactly what Monty Hall lets you do in the actual game after you reveal the goat. The
basic idea is that if you had to choose two doors, one of them would always have a goat behind
it. When he opens a door to reveal a goat before asking you if you want to change, he's doing
you a huge favor! He's saying (in effect): "There's a two-thirds chance that the car is behind one
of the doors you didn't choose, and look, it's not that one!"

Think of it this way. Suppose you chose door no. 1. Monty then offers you the option to take Gates 2
and 3. You accept the offer, give up one door and get two, meaning you can reasonably expect to win the
Machine Translated by
Google

car 2/3 of the time. At that time, what would happen if Monty opened door no. 3—one of your doors—to
reveal a goat? Should you feel less confident about your decision? Of course not. If the car was behind the
no. 3, would have opened the no. 2! He hasn't shown you anything.

When the game is played normally, Monty actually gives you a choice between the door you originally
chose and the other two doors, only one of which could have a car behind it. When he opens a door to
reveal a goat, he's simply doing you the courtesy of showing you which of the other two doors doesn't have
a car in it. You have the same chance of winning in the following two scenarios: 1. Choosing Door No. 1,
then agree to switch to Door no. 2 and Door No. 3 before opening any door.

2. Choosing door no. 1, then agree to switch to Door no. 2 after Monty reveals a goat behind door
no. 3 (or choose number 3 after revealing a goat behind number 2).

In both cases, switching gives you the advantage of having two doors instead of one, and therefore you
can double your chances of winning, from 1/3 to 2/3.

My third explanation is a more extreme version of the same basic intuition.


Suppose Monty Hall offers you a choice of 100 doors instead of just three. After choosing the door, say no.
47, opens another 98 doors with goats behind them. Now there are only two doors that remain closed, no.
47 (your original choice) and another, let's say, no. 61. Should you change?
Of course you should. There is a 99 percent chance that the car was behind one of the doors you did
not originally choose. Monty did you the favor of opening 98 of those doors that you did not choose, all of
which he knew were not

have the car behind them. There is only a 1 in 100 chance that your original choice was correct (#47).
There is a 99 in 100 chance that your original choice was not correct. And if your original choice was not
correct, then the car is behind the other door, no. 61. If you want to win 99 out of 100 times, you should
switch to no. 61.

In short, if you're ever a contestant on Let's Make a Deal, you should definitely switch doors when Monty
Hall (or his replacement) gives you the option. The most applicable lesson is that your instinct about
probability can sometimes lead you astray.

* Can you play at [Link]


_r=2&oref=slogin&oref=slogin.
Machine Translated by
Google

CHAPTER 6

Problems with probability


How overconfident math geeks almost destroyed the global financial
system

Statistics can't be smarter than the people who use them. and in some In some cases, they can make smart
people do stupid things. One of the most irresponsible uses of statistics in recent memory involved the
mechanism for measuring risk on Wall Street before the 2008 financial crisis. At that time, companies across
the financial industry used a common barometer of risk, the Value at Risk or VaR model. In theory, VaR
combined the elegance of an indicator (packing a lot of information into a single number) with the power of
probability (associating an expected gain or loss with each of the company's assets or trading positions). The
model assumed that there are a variety of possible outcomes for each of the firm's investments. For example, if
the company owns shares of General Electric, the value of those shares may go up or down. When VaR is
calculated for a short period of time, say, a week, the most likely result is that the stock will have approximately
the same value at the end of that period as it had at the beginning. There is less chance of stocks rising or
falling by 10 percent. And an even smaller chance of them going up or down 25 percent, and so on.

Based on past data on market movements, the company's quantitative experts (often called "quants" in the
industry and "rich nerds" elsewhere) might assign a dollar figure, say $13 million, that represented the most the
company could lose. in that position during the time period under examination, with 99 percent probability. In
other words, 99 times out of 100 the company would not lose more than $13 million on a particular trading
position; 1 time out of 100, it would.
Remember that last part, because it will be important soon.
Before the 2008 financial crisis, companies relied on the VaR model to quantify their overall risk. If a single
trader had 923 different open positions (investments that could rise or fall in value), each of those investments
could be evaluated as described above for General Electric stock; from there, the total risk of the trader's
portfolio could be calculated. The formula even took into account correlations between different positions. For example,
if two investments had expected returns that were negatively correlated, a loss in one would likely have been offset by a gain in
the other, making the two investments together less risky than either one separately. Typically, the head of the trading desk
would know that bond trader Bob Smith has a 24-hour VaR (the value at risk over the next 24 hours) of $19 million, again with
99 percent probability.
The most Bob Smith could lose in the next 24 hours would be $19 million, 99 times out of 100.
So, better yet, the aggregate risk to the company could be calculated at any time by taking the same basic process one step
further. The underlying mathematical mechanics are obviously fabulously complicated, as companies had a dizzying array of
investments in different currencies, with different amounts of leverage (the amount of money borrowed to make the investment),
trading in markets with varying degrees of liquidity, and so on. Despite all that, company executives apparently had an accurate
Machine Translated by
Google

measure of the magnitude of the risk the company had taken at any given time. As former New York Times business writer Joe
Nocera explained, “The great appeal of VaR, and its big selling point for people who are not quants, is that it expresses risk as a
single number, a dollar figure, nothing less.”

1. .......................................................
At JP Morgan, where the VaR model was developed
and refined, the daily VaR calculation was known as the “4:15 report” because it would be on the desks of senior executives
every afternoon at 4:15, just after the US financial markets. had closed for the day.

Presumably this was a good thing, as generally more information is better, especially when it comes to risks. After all,
probability is a powerful tool. Isn't this the same kind of calculation Schlitz executives made before spending big bucks on blind
taste tests at halftime of the Super Bowl?
Not necessarily. VaR has been called “potentially catastrophic,” “a fraud,” and many other things that don’t fit into a familiar
statistics book like this one. In particular, the model has been blamed for the onset and severity of the financial crisis. The main
criticism of VaR is that the underlying risks associated with financial markets are not as predictable as a coin toss or even a blind
taste test between two beers. The false precision built into the models created a false sense of security. VaR was like a faulty
speedometer, which is arguably worse than no speedometer at all. If you rely too much on a faulty speedometer, you will miss
other signs that your speed is unsafe. On the other hand, if there is no speedometer, you have no choice but to look around for
clues as to how fast you are actually going.

Circa 2005, with VaR hitting desks at 4:15 every weekday, Wall Street was driving pretty fast. Unfortunately, there
were two major problems with the risk profiles encapsulated by VaR models. First, the underlying probabilities
on which the models were built were based on past market movements; however, in financial markets (unlike
beer tasting), the future does not necessarily resemble the past. There was no intellectual justification for
assuming that market movements between 1980 and 2005 were the best predictor of market movements after
2005. In some ways, this lack of imagination resembles the military's periodic mistaken assumption that the
next war will look like the last. In the 1990s and early 2000s, commercial banks used home mortgage lending
models that assigned zero probability to large declines in home prices. Never before have home prices fallen
as much and as fast as they did beginning in 2007. But that's what happened. Former Federal Reserve
Chairman Alan Greenspan explained to a congressional committee after the fact: “But the entire intellectual
edifice collapsed in the summer [of 2007] because the data fed into risk management models typically covered
only the past two decades, a period of euphoria. If, instead, the models had been more appropriately adapted
to historical periods of stress, capital requirements would have been much higher and the financial world would
be in much better shape, in my view.”
3

Second, even if the underlying data could accurately predict future risk, the 99 percent assurance offered
by the VaR model was dangerously useless, because it is the 1 percent that is really going to screw you up.
Hedge fund manager David Einhorn explained: "This is like an airbag that works all the time, except when you
have a car accident." If a company has a value at risk of $500 million, that can be interpreted to mean that it
has a 99 percent chance of losing no more than $500 million over the specified time period. Well, hello, that
also means the company has a 1 percent chance of losing more than $500 million (much, much more in some
circumstances). In fact, the models had nothing to say about how bad that 1 percent scenario might turn out.
Machine Translated by
Google

Very little attention was paid to “tail risk,” the small risk (named after the tail of the distribution) of some
catastrophic outcome. (If you're driving home from a bar with a blood alcohol level of .15, there's probably less
than a 1 percent chance of having an accident and dying; that doesn't mean it's a sensible thing to do.) Many
companies have compounded this mistake. making unrealistic assumptions about their preparedness for rare
events. Former Treasury Secretary Hank Paulson has explained that many companies assumed they would be
able to get cash in the event of a 4-asset bankruptcy. need selling. But during a crisis, all other companies need cash
too, so they all try to sell the same types of assets. It is the management equivalent of
“I don’t need to stock up on water because if there’s a natural disaster, I’ll just go to the supermarket and buy some.” Of course,
after an asteroid hits your town, fifty thousand other people are also trying to buy water; when you get to the supermarket, the
windows are broken and the shelves are empty.
The fact that you never considered that your city could be crushed by a massive asteroid was exactly the problem with VaR.
Here is New York Times columnist Joe Nocera again, summarizing the thoughts of Nicholas Taleb, author of The Black Swan:
The Impact of the Highly Improbable and a scathing critic of VaR: “The greatest risks are never the ones you can see and
measure, but the ones you can’t see and therefore can never measure.” Those that seem so outside the bounds of normal
probability that you can't imagine they could happen in your lifetime... although, of course, they do happen, more often than you
might think.

In some ways, the VaR debacle is the opposite of the Schlitz example in Chapter 5. Schlitz operated with a known
probability distribution.
Whatever data the company had on the likelihood that blind tasters would choose Schlitz was a good estimate of how similar live
tasters would perform at halftime. Schlitz even got around his problem by running the entire test on men who said they liked the
other beers better. Even if no more than twenty-five Michelob drinkers chose Schlitz (an almost unbelievably low result), Schlitz
could still claim that one in four beer drinkers should consider switching. Perhaps most importantly, this was all just beer, not the
global financial system. Wall Street quants made three fundamental mistakes. First, they confused precision with accuracy. VaR
models were like my golf rangefinder when it was set to meters instead of yards: accurate and incorrect. False precision led Wall
Street executives to believe they had risk under control when they didn't.

Second, the estimates of the underlying probabilities were wrong. As Alan Greenspan pointed out in testimony cited earlier in
this chapter, the relatively calm and prosperous decades before 2005 should not have been used to create probability
distributions of what might happen in markets in the decades that followed. This is the equivalent of walking into a casino and
thinking you will win at roulette 62 percent of the time because that's what happened the last time you played. It would be a long
and expensive evening. Third, companies neglected their “tail risk.” VaR models predicted what would happen 99 times out of
100. This is how probability works (as will be repeatedly emphasized in the second half of the book). Unlikely things happen. In
fact, over a long enough period of time, they're not even that unlikely. People get struck by lightning all the time. My mother has
had three holes in one.

Statistical arrogance at commercial banks and ultimately on Wall Street contributed to the most severe global financial
contraction since the Great Depression. The crisis that began in 2008 destroyed trillions of dollars in U.S. wealth, pushed
Machine Translated by
Google

unemployment to more than 10 percent, created waves of foreclosures and corporate bankruptcies, and saddled
governments around the world with massive debts as they struggled to contain the economic damage. This is a sadly
ironic result, given that sophisticated tools such as VaR were designed to mitigate risk.

Probability offers a powerful and useful set of tools, many of which can be used correctly to understand the world or
incorrectly to wreak havoc on it. Continuing with the “statistics as a powerful weapon” metaphor I’ve used throughout the
book, I’ll paraphrase the gun rights lobby: probability doesn’t make mistakes; people who use probability make mistakes.
The remainder of this chapter will catalogue some of the most common errors, misunderstandings, and ethical dilemmas
related to probability.

Assuming that events are independent when they are not. The probability of getting heads on a fair coin is ½. The
probability of getting two heads in a row is (½) or ¼, since 2
The probability of two independent events occurring is the product of their individual probabilities. Now that you are armed
with this powerful knowledge, let's say you have been promoted to head of risk management at a major airline. Your
assistant informs you that the probability of an airplane engine failing for any reason during a transatlantic flight is 1 in
100,000. Given the number of transatlantic flights, this is not an acceptable risk. Fortunately, every plane that makes such
a trip has at least two engines.

His assistant has calculated that the risk of both engines shutting down over the Atlantic must be 2
be (1/100,000) or 1 in 10 billion, which is a reasonable security risk. This would be a good time to tell your assistant to use
up his vacation days before he gets fired. The two engine failures are not independent events. If an airplane flies through
a flock of geese during takeoff, both engines are likely to be similarly compromised. The same would be true of many
other factors that affect a jet engine's performance, from weather to improper maintenance. If one engine fails, the
probability of the second engine failing will be significantly greater than 1 in 100,000.
Does this seem obvious? It was not obvious during the 1990s, when British prosecutors committed a serious
miscarriage of justice due to an inappropriate use of probability. As in the hypothetical jet engine example, the statistical
error was assuming that several events were independent (like flipping a coin) rather than dependent (when a certain
outcome makes a similar outcome more likely in the future). However, this mistake was real and innocent people

As a result, they were sent to prison.


The error arose in the context of sudden infant death syndrome (SIDS), a phenomenon in which a perfectly healthy
baby dies in its crib. (The British refer to SIDS as “sudden infant death.”) SIDS was a medical mystery that attracted more
attention as infant deaths from other causes became less common. These infant deaths were * Because they were so
mysterious and little understood that they aroused suspicion. Sometimes that suspicion was justified. SIDS was
sometimes used to cover up parental neglect or abuse; a post-mortem examination cannot necessarily distinguish natural
deaths from those involving foul play. British prosecutors and courts became convinced that one way to separate crime
from natural deaths would be to focus on families where multiple sudden deaths occurred. Sir Roy Meadow, a prominent
British paediatrician, was a frequent expert witness on this point. As the British magazine The Economist explains: “What
became known as Meadow’s Law – the idea that one child death is a tragedy, two are suspicious and three are murder –
is based on the notion that if an event is rare, then two or more instances of it in the same family are so improbable that
they are unlikely to be the result of chance.” 5 Sir Meadow told the jury that the chance of two babies in a family dying
suddenly from natural causes was extraordinary: 1 in 73 million. He explained the calculation: since the incidence of a
Machine Translated by
Google

sudden death is rare, 1 in 8,500, the probability of having two sudden deaths in the same family would be (1/8,500), that
is, approximately 1 in 73 million. This stinks of foul play. That's what juries decided, sending many parents to prison based
on this testimony about crib death statistics (often without any medical evidence to support abuse or neglect). In some
cases, babies were separated from their parents at birth due to the unexplained death of a sibling.

The Economist explained how a misunderstanding about statistical independence became a flaw in Meadow's
testimony:

There is an obvious flaw in this reasoning, as the Royal Statistical Society, protector of its ridiculed subject, has
pointed out. The probability calculation works well, as long as it is certain that crib deaths are completely random
and not linked by some unknown factor. But with something as mysterious as sudden deaths, it's entirely possible
that there is a link: something genetic, for example, that would make a family that had suffered one sudden death
more likely, not less, to suffer another. And ever since those women were convicted, scientists have been
suggesting that there may be such a link.

In 2004, the British government announced that it would review 258 trials in which parents had been convicted of
murdering their young children.
Not understanding when events ARE independent. A different type of error occurs when events that are
independent are not treated as such. If you find yourself in a casino (a place, statistically speaking, you
shouldn't go), you'll see people staring wistfully at their dice or cards and declaring that they're “beaten.” If the
roulette ball has landed on black five times in a row, it is clear that it must now land on red. No no no! The
probability of the ball landing on a red number remains unchanged: 16/38. Conversely, this belief is sometimes
called the "gambler's fallacy." In fact, if you flip a normal coin 1,000,000 times and get 1,000,000 heads in a
row, the probability of getting tails on the next flip is still ½. The very definition of statistical independence
between two events is that the outcome of one has no effect on the outcome of the other. Even if you don't find
the statistics convincing, you might wonder about physics: How can it be that flipping a series of tails in a row
makes it more likely that the coin will come up heads on the next toss?

Even in sports, the notion of streaks can be illusory. One of the most famous and interesting academic
papers related to probability refutes the common notion that basketball players periodically develop a streak of
good shots during a game, or “a hot hand.” Certainly, most sports fans would say that a player who makes a
shot is more likely to make the next shot than a player who just missed. No, according to research by Thomas
Gilovich, Robert Vallone and Amos Tversky, who tested the hot hand in three different ways. 6
First, they
analyzed shot data from Philadelphia 76ers home games during the 1980-81 season. (At the time, similar data
was not available for other NBA teams.) They found "no evidence of a positive correlation between the
outcomes of successive shots." Second, they did the same with the Boston Celtics' free throw data, which
produced the same result. And finally, they ran a controlled experiment with members of the Cornell men's and
women's basketball teams.
Machine Translated by
Google

Players made an average of 48 percent of their field goals after taking their last shot and 47 percent after
missing. For fourteen of twenty-six players, the correlation between making one shot and then making the next
shot was negative.
Only one player showed a significant positive correlation between one shot and the next.

That's not what most basketball fans will tell you. For example, 91 percent of basketball fans surveyed at
Stanford and Cornell by the paper's authors agreed with the statement that a player is more likely to make his
next shot after making his last two or three shots than after missing his last one. two or three shots. The
importance of the “hot hand” role lies in the difference between perception and empirical reality. The authors note:
"People's intuitive conceptions of randomness systematically deviate from the laws of chance." We see patterns where none
may actually exist.
Like cancer clusters.

Groups happen. You've probably read the story in the newspaper, or maybe seen the news: A statistically unlikely number of
people in a particular area have contracted a rare form of cancer. It must be the water, or the local power plant, or the cell phone
tower. Of course, any of those things could actually be causing adverse health outcomes. (How statistics can identify such
causal relationships will be explored in later chapters.) But this group of cases may also be the product of pure chance, even
when the number of cases seems highly improbable. Yes, the chance of five people in the same school, church or workplace
getting the same rare form of leukemia may be one in a million, but there are millions of schools, churches and workplaces. It is
not very unlikely that five people would contract the same rare form of leukemia in one of those places. We just don't think about
all the schools, churches and workplaces where this hasn't happened. To use a different variation of the same basic example,
the probability of winning the lottery may be 1 in 20 million, but none of us is surprised that someone wins, because millions of
tickets have been sold. (Despite my general aversion to lotteries, I admire Illinois' slogan: "Someone's going to win the lottery, it
might as well be you.")

Here is an exercise I do with my students to make the same basic point.


The larger the class, the better it works. I ask everyone in the class to take out a coin and stand up. We all flip the coin; whoever
turns their head has to sit down.
Assuming we start with 100 students, about 50 will be seated after the first round. Then we do it again, after which there are
about 25 left standing. Etc. Most of the time, there will be a student at the end who has flipped five or six tails in a row. At that
point, I ask the student questions like, “How did you do it?” and “What are the best training exercises to flip so many tails in a
row?” or “Is there a special diet that helped you achieve this impressive feat?” These questions provoke laughter because the
class has just watched the entire process unfold; they know that the student who flipped six tails in a row has no special talent
for flipping coins. He or she turned out to be the one who ended up with a lot of tails. However, when we see an anomalous
event like that out of context, we assume that something other than randomness must be responsible.
Machine Translated by
Google

The prosecutor's fallacy. Suppose you hear testimony in court to the effect that: (1) a DNA sample found at a crime scene
matches a sample taken from the defendant; and (2) there is only a one in a million chance that the sample recovered at
the crime scene matches someone other than the defendant.
(For the purposes of this example, you can assume that the prosecution's odds are correct.) Based on that evidence,
would you vote to convict?

I hope not.
The prosecutor's fallacy occurs when the context surrounding statistical evidence is neglected. Here are two
scenarios, each of which could explain the DNA evidence being used to prosecute the defendant.
Defendant 1: This defendant, a scorned lover of the victim, was arrested three blocks from the crime scene carrying
the murder weapon. After his arrest, the court forced him to provide a DNA sample, which matched a sample taken from a
hair found at the crime scene.
Defendant 2: This defendant was convicted of a similar crime in a different state several years ago. As a result of that
conviction, his DNA was included in a national DNA database of more than one million violent offenders. The DNA sample
taken from the hair found at the crime scene was checked against that database and compared to this individual, who has
no known association with the victim.
As noted above, in both cases the prosecutor can rightly say that the DNA sample taken from the crime scene
matches that of the accused and that there is only a one in a million chance that it matches that of anyone else. But in the
case of defendant 2, there's a good chance he could be that random person, the one in a million guy whose DNA just
happens to be similar to the real killer's. Because the chances of finding a match in a million are relatively high if you run
the sample through a database with samples of one million people.

Reversion to the mean (or regression to the mean). You may have heard of the Sports Illustrated curse, whereby
individual athletes or teams featured on the cover of Sports Illustrated subsequently see their performance decline. One
explanation is that appearing on the cover of the magazine has some adverse effect on subsequent performance. The
most statistically sound explanation is that teams and athletes appear on their cover after some abnormally good period
(such as a twenty-game winning streak) and that their subsequent performance simply returns to normal or the mean.
This is the phenomenon known as reversion to the mean. Probability tells us that any outlier (an observation that is
particularly far from the mean in one direction or another) is likely to be followed by results that are more consistent with
the long-run average.

Reversion to the mean may explain why the Chicago Cubs always seem to pay huge salaries to free agents who
subsequently disappoint fans like me. Players can negotiate huge salaries with the Cubs after one or two exceptional
seasons. Putting on a Cubs uniform doesn't necessarily make these players worse off (though I wouldn't necessarily rule it
out); rather, the Cubs pay a lot of money for these superstars at the end of an exceptional period (an atypical year or two)
after which their performance for the Cubs returns to something closer to normal.

The same phenomenon may explain why students who perform much better than usual on some type of test will, on
average, perform slightly worse on a retest, and students who perform worse than usual will tend to perform slightly better
Machine Translated by
Google

when they take the test again. One way to think of this reversion to the mean is that performance (both mental and
physical) consists of some underlying effort related to talent plus an element of luck, good or bad. (Statisticians would call
this random error.) In any case, those individuals who perform well above average over some period have probably had
luck on their side; those who perform well below average have probably had bad luck. (In the case of a test, think of
students guessing right or wrong; in the case of a baseball player, think of a hit that might go wrong or landing on one foot
just right for a triple.) or ends up very unlucky (as inevitably will happen), the resulting performance will be closer to the
average.
Imagine I'm trying to put together a team of superstar coin flippers (under the mistaken impression that talent matters
when it comes to coin flipping). After watching a student throw six tails in a row, I offer him a ten-year, $50 million contract.
Needless to say, I will be disappointed when this student only throws 50 percent tails in those ten years.
At first glance, mean reversion may seem contrary to the “gambler’s fallacy.” After the student throws six tails in a row,
should he throw heads or no heads? The probability of getting heads on the next toss is the same as always: ½. Just
because you've flipped many tails in a row doesn't make it any more likely that you'll get heads on the next flip. Each
release is an independent event.
However, we can expect the results of subsequent tosses to be consistent with what probability predicts, which is half
heads and half tails, rather than what it has been in the past, which is all tails. It is almost certain that someone who has
thrown all tails will start throwing more heads over the next 10, 20, or 100 throws. And the more changes, the more the
result will resemble the 50-50 average result predicted by the law of large numbers. (Or, alternatively, we should start
looking for evidence of fraud.)

As a curious note, researchers have also documented a Businessweek report phenomenon. When CEOs receive
high-profile awards, including being named one of Businessweek’s “Best Managers,” their companies subsequently
underperform over the next three years, as measured by both accounting earnings and stock price. However, unlike the
Sports Illustrated effect, this effect appears to be more than a reversion to the mean. According to Ulrike Malmendier and
Geoffrey Tate, economists at the University of California at Berkeley and UCLA, respectively, when CEOs achieve
“superstar” status, they become distracted by their newfound prominence.
7
They write their memoirs. They are invited to sit in the outside meetings. They start
looking for trophy wives. (The authors propose only the first two explanations, but the last one also seems plausible to
me.)
Malmendier and Tate write: "Our results suggest that media-induced superstar culture leads to behavioral distortions
beyond mere reversion to bad behavior." In other words, when a CEO appears on the cover of Businessweek, he sells his
shares.

Statistical discrimination. When is it okay to act on what probability tells us is likely to happen, and when is it not okay? In
2003, Anna Diamantopoulou, the European Commissioner for Employment and Social Affairs, proposed a directive stating
that insurance companies cannot charge different rates to men and women, because it violates the European Union's
principle of equal treatment.
8
For insurers, however, gender-based premiums do not constitute discrimination; they
are just statistics. Men tend to pay more for car insurance because they have more accidents. Women pay more for
annuities (a financial product that pays a fixed sum monthly or annually until death) because they live longer.
Obviously, many women have more accidents than many men, and many men live longer than many women. But, as
explained in the last chapter, insurance companies don't care. They only care about what happens on average, because if
Machine Translated by
Google

they do well, the company will make money. What is interesting about the European Commission's policy banning gender-
based insurance premiums, which is being implemented in 2012, is that the authorities are not pretending that gender is
not related to the risks insured; they are simply declaring that disparate rates based on sex are unacceptable.
*

At first, that seems like an annoying nod to political correctness. On second thought, I'm not so sure. Remember all
that awesome stuff about preventing crimes before they happen? Probability can take us to some intriguing but disturbing
places in this regard. How should we react when our probability-based models tell us that meth smugglers from Mexico
are most likely to be Hispanic men between the ages of eighteen and thirty who drive red pickup trucks between 9:00 p.m.
and midnight, when we also know that the vast majority of Hispanic men who fit that profile are not smugglers?

methamphetamine? Yes, I used the word profiling, because that is the least glamorous description of the
predictive analytics I described so brilliantly in the last chapter, or at least one potential aspect of it.
Probability tells us what is more likely and what is less likely. Yes, this is just basic statistics: the tools
described in the last few chapters. But they are also statistics with social implications. If we want to catch
violent criminals, terrorists, drug dealers and others with the potential to cause enormous harm, then we must
use every tool at our disposal. Probability can be one such tool. It would be naive to think that gender, age,
race, ethnicity, religion, and country of origin taken together tell us nothing about anything related to law
enforcement.
But what we can or should do with such information (assuming it has any predictive value) is a
philosophical and legal question, not a statistical one.
Every day we receive more and more information about more and more things. Is it okay to discriminate if the
data tells us that we will be right much more often than we will be wrong? (This is the origin of the term
“statistical discrimination” or “rational discrimination.”) The same type of analysis that can be used to determine
that people who buy birdseed are less likely to default on their credit cards (yes, that's really true) can be
applied everywhere else in life.
How much of that is acceptable? If we can build a model that correctly identifies drug dealers 80 times out of
100, what happens to the poor in the 20 percent? Because our model will harass them again and again.

The broader point here is that our ability to analyze data has become much more sophisticated than our
thinking about what we should do with the results. You may agree or disagree with the European
Commission's decision to ban gender-based insurance premiums, but I promise you that it will not be the last
difficult decision of this kind. We like to think of numbers as "cold, hard facts." If we do the math right, then we
should have the correct answer. The most interesting and dangerous reality is that sometimes we can get the
math right and end up getting it wrong in a dangerous direction. We can blow up the financial system or harass
a twenty-two-year-old white male standing on a particular corner at a particular time of day, because,
according to our statistical model, he is almost certainly there to buy drugs. For all the elegance and precision
of probability, there is no substitute for thinking about what calculations we are making and why we are making
them.
Machine Translated by
Google

* SIDS remains a medical mystery, although many of the risk factors have been identified. For example, infant deaths can be dramatically reduced by
putting children to sleep on their backs.

* The policy change was ultimately precipitated by a 2011 ruling by the European Court of Justice that different bonuses for men and women
constitute sex discrimination.
Machine Translated by
Google

CHAPTER 7

The importance of data


"Garbage in, garbage out"

In the spring of 2012, researchers published a surprising finding in the prestigious journal Science. According to this cutting-edge
research, when male fruit flies are repeatedly rejected by females, they drown their sorrows in alcohol.
The New York Times described the study in a front-page article: “They were fledgling young males, and they attacked not once,
not twice, but a dozen times with a group of attractive females hovering nearby. So they did what so many men do after being
repeatedly rejected: they got drunk and used alcohol as a balm for unfulfilled desire.”
1

This research advances our understanding of the brain's reward system, which in turn may help us find new strategies to
deal with drug and alcohol dependence. One substance abuse expert described reading the study as "looking back in time, to
the very origins of the reward circuitry that drives fundamental behaviors like sex, eating and sleeping."
Since I am not an expert in this field, I had two slightly different reactions when reading about the despised fruit flies. First, it
made me feel nostalgic for college.
Secondly, my inner researcher started to wonder how fruit flies get drunk. Is there a miniature fruit fly bar, with a variety of fruit-
based liqueurs and a fruit fly-sympathetic bartender? Is country western music playing in the background? Do fruit flies like
country western music?
It turns out the design of the experiment was devilishly simple. A group of male fruit flies were allowed to mate freely with
virgin females. Another group of males was released among female fruit flies that had already mated and were therefore
indifferent to amorous advances from the males. Both groups of male fruit flies were then offered feeding straws that offered a
choice between standard fruit fly food, yeast and sugar, and the "hard stuff": yeast, sugar and 15 percent alcohol.
Males who had spent days trying to mate with indifferent females were significantly more likely to drink alcohol.
Despite the minor nature of these results, they have important implications for humans. They suggest a connection between
stress, chemical responses in the brain and the appetite for alcohol. However, the results are not a triumph of statistics. They are
a triumph of data, which made relatively basic statistics

possible analysis. The genius of this study was finding a way to create a group of sexually satiated male fruit
flies and a group of sexually frustrated male fruit flies, and then finding a way to compare their drinking habits.
Once the researchers did that, crunching the numbers was no more complicated than a typical high school
science fair project.
Data is to statistics what a good offensive line is to a star quarterback. In front of every star quarterback is a
good group of blockers. They generally don't get much credit. But without them, you'll never see a star
quarterback. Most statistics books assume that you use good data, just as a cookbook assumes that you don't
buy rancid meat or rotten vegetables. But even the best recipe won't save a meal that starts with spoiled
ingredients. The same is true of statistics; no amount of sophisticated analysis can compensate for
Machine Translated by
Google

fundamentally flawed data. Hence the expression “garbage in, garbage out.” Data deserves respect, and so do
offensive linemen.

We typically ask our data to do one of three things. First, we can demand a sample of data that is
representative of some larger group or population. If we are trying to measure voter attitudes toward a
particular political candidate, we will need to interview a sample of likely voters who are representative of all
voters in the relevant political jurisdiction. (And remember, we don't want a sample that is representative of
everyone who lives in that jurisdiction; we want a sample of those who are likely to vote.) One of the most
powerful findings in statistics, which will be explained in greater depth. What we will see in the next two
chapters is that inferences made from reasonably large and correctly drawn samples can be just as accurate
as trying to obtain the same information from the entire population.
The simplest way to gather a representative sample of a larger population is to randomly select some
subset of that population. (Surprisingly, this is known as a simple random sample.) The key to this
methodology is that each observation in the relevant population must have an equal chance of being included
in the sample. If you plan to survey a random sample of 100 adults in a neighborhood with 4,328 adult
residents, your methodology must ensure that each of those 4,328 residents has an equal chance of ending up
as one of the 100 adults surveyed. Statistics books almost always illustrate this point by drawing colored
marbles from an urn. (In fact, it is the only place where you see the word “urn” used with any regularity.) If there
are 60,000 blue marbles and 40,000 red marbles in a giant urn, then the most likely composition of a sample of
100 marbles drawn at random from the urn would be 60 blue marbles and 40 red marbles. If we did this more
than once, there would obviously be deviations from sample to sample: some might have 62 blue marbles and
38 red marbles,
or 58 blue and 42 red. But the chances of drawing a random sample that deviates wildly from the composition of the
marbles in the urn are very, very low.
Now, it is true that there are some practical challenges here. Most of the populations we care about tend to be more
complicated than an urn full of marbles. How exactly would a random sample of the U.S. adult population be selected for
inclusion in a telephone survey? Even a seemingly elegant solution like a random phone dialer has potential flaws. Some
people (particularly those with low incomes) may not have a phone. Others (particularly high-income individuals) may be
more likely to screen calls and choose not to answer. Chapter 10 will describe some of the strategies that polling firms use
to overcome these types of sampling challenges (most of which became even more complicated with the advent of cell
phones). The key idea is that a correctly drawn sample will resemble the population from which it is drawn. In terms of
intuition, one can imagine tasting a pot of soup with a single spoonful. If you've stirred the soup properly, a single spoonful
can tell you the flavor of the entire pot.
A statistics text will include many more details about sampling methods. Survey and market research firms spend their
days figuring out how to obtain good, representative data from diverse populations in the most cost-effective way. By now,
you should appreciate several important things: (1) A representative sample is a fabulously important thing, as it opens the
door to some of the most powerful tools that statistics has to offer. (2) Getting a good sample is harder than it seems. (3)
Many of the most egregious statistical claims are caused by good statistical methods applied to bad samples, not the
other way around. (4) Size matters and the bigger the better. The details will be explained in the next chapters, but it
should be intuitive that a larger sample will help smooth out any abnormal variations. (A bowl of soup will be an even
Machine Translated by
Google

better test than a tablespoon.) One crucial caveat is that a larger sample will not compensate for errors or “biases” in your
composition. A bad sample is a bad sample. No supercomputer or fancy formula is going to rescue the validity of your
national presidential poll if the respondents come solely from a telephone survey of Washington, DC residents.
Washington, DC residents don't vote like the rest of America; calling 100,000 DC residents instead of 1,000 won't fix that
fundamental problem with your poll. In fact, it could be argued that a large, biased sample is worse than a small, biased
sample because it will give a false sense of confidence regarding the results.

The second thing we often ask of data is that it provide some source of comparison. Is a new drug more effective than
current treatment? Are ex-offenders who receive job training less likely to return to prison than ex-convicts?
Convicts who do not receive such training? Do students who attend charter schools perform better than similar
students who attend regular public schools?
In these cases, the goal is to find two groups of subjects that are broadly similar except for the application of
any “treatment” that interests us. In the context of social sciences, the word "treatment" is broad enough to
encompass anything from being a sexually frustrated fruit fly to receiving an income tax refund. As with any
other application of the scientific method, we attempt to isolate the impact of a specific intervention or attribute.
This was the genius of the fruit fly experiment. The researchers discovered a way to create a control group (the
males that were mated) and a "treatment" group (the males that were knocked down); the subsequent
difference in their drinking behaviors can be attributed to whether they were sexually slighted or not.
In the physical and biological sciences, creating treatment and control groups is relatively straightforward.
Chemists can make small variations from one test tube to another and then study the difference in results.
Biologists can do the same with their petri dishes. Even most animal tests are simpler than trying to get fruit
flies to drink alcohol. We can have one group of rats exercise regularly on a treadmill and then compare their
mental acuity in a maze with the performance of another group of rats that did not exercise. But when humans
get involved, things get more complicated. A sound statistical analysis often requires a treatment and control
group, but we can't force people to do the things we make lab rats do. (And many people don't like it when
even lab rats do these things.) Do repeated concussions cause serious neurological problems later in life? This
is a really important question. The future of football (and perhaps other sports) depends on the answer.
However, this is a question that cannot be answered by experiments on humans. So unless we can teach fruit
flies to wear helmets and execute the spreading offensive, we need to find other ways to study the long-term
impact of head trauma.
A recurring challenge in research with human subjects is creating treatment and control groups that differ
only in that one group receives the treatment and the other does not. For this reason, the “gold standard” of
research is randomization, a process by which human subjects (or schools, or hospitals, or whatever we are
studying) are randomly assigned to either the treatment or control group. We do not assume that all
experimental subjects are identical.
Instead, chance becomes our friend (once again), and we assume that randomization will evenly divide all
relevant characteristics between the two groups: both characteristics we can observe, such as race or income,
and also confounding characteristics we can't measure or haven't experienced. considered, such as
perseverance or faith.
Machine Translated by
Google

The third reason we collect data is, to quote my teenage daughter, "just because." Sometimes we don't have a specific
idea of what we will do with the information, but we suspect that at some point it will be useful. This is similar to a crime
scene detective demanding that all possible evidence be captured so that it can later be sorted through for clues. Some of
these tests will be useful, some will not. If we knew exactly what would be useful, we probably wouldn't need to do the
research in the first place.

You probably know that smoking and obesity are risk factors for heart disease. You probably don't know that a long-
running study among residents of Framingham, Massachusetts, helped clarify those relationships. Framingham is a
suburban city of about 67,000 people about twenty miles west of Boston.
To non-researchers, it is best known as a Boston suburb with reasonably priced housing and convenient access to the
impressive and upscale Natick Mall. To researchers, Framingham is best known as the home of the Framingham Heart
Study, one of the most successful and influential longitudinal studies in the history of modern science.

A longitudinal study collects information on a large group of subjects at many different times, for example, once every
two years. The same participants can be interviewed periodically for ten, twenty, or even fifty years after their entry into
the study, creating a remarkably rich trove of information. In the case of the Framingham study, researchers collected
information on 5,209 adult Framingham residents in 1948: height, weight, blood pressure, education, family structure, diet,
smoking, drug use, and so on. Importantly, the researchers have collected follow-up data on the same participants ever
since (and also data on their offspring, to examine genetic factors linked to heart disease). Framingham data have been
used to produce more than two thousand scholarly articles since 1950, including nearly a thousand between 2000 and
2009.

These studies have produced crucial findings for our understanding of cardiovascular disease, many of which we now
take for granted: cigarette smoking increases the risk of heart disease (1960); physical activity reduces the risk of heart
disease and obesity increases it (1967); high blood pressure increases the risk of stroke (1970); high levels of HDL
cholesterol (hereafter known as “good cholesterol”) reduce the risk of death (1988); people with parents and siblings who
have cardiovascular disease have a significantly increased risk of developing the disease (2004 and 2005).

Longitudinal data sets are the research equivalent of a Ferrari. Data are particularly valuable when it comes to
exploring causal relationships that may take years or decades to develop. For example, the Perry Preschool Study began
in the late 1960s with a group of 123 three- and four-year-old African-American children from poor families. Participating
children were randomly assigned to a group that received an intensive preschool program and a comparison group that
did not. The researchers then measured various outcomes for both groups over the next forty years. The results make a
compelling case for the benefits of early childhood education. Students who received the intensive preschool experience
had higher IQs at age five. They were more likely to graduate from high school. They had higher incomes at age forty. In
contrast, participants who did not receive the preschool program were significantly more likely to have been arrested five
or more times before age forty.
It's not surprising that we can't always have the Ferrari. The research equivalent of a Toyota is a cross-sectional data
set, which is a collection of data collected at a single point in time. For example, if epidemiologists are looking for the
cause of a new disease (or an outbreak of an old one), they might collect data on everyone affected in the hope of finding
Machine Translated by
Google

a pattern that leads to the source. What have they eaten? Where have they traveled? What else do they have in
common?
Researchers can also collect data from people who do not have the disease to highlight contrasts between the two
groups.
In fact, all this interesting talk about cross-sectional data reminds me of the week before my wedding, when I became
part of a data set. I was working in Kathmandu, Nepal, when I tested positive for a little-known stomach illness called
“blue-green algae,” which had been found in only two places in the world. Researchers had isolated the pathogen that
caused the disease, but were still unsure what type of organism it was, since it had never been identified before. When I
called home to tell my fiancée about my diagnosis, I recognized that there was bad news. The disease had no known
means of transmission and no known cure and could cause extreme fatigue and other unpleasant side effects lasting from
a few days to many months.
* WITH ONLY
a week to the wedding, yes, this could be a problem. Would I have complete control of my digestive system as I walked
down the aisle? Maybe.
But then I really tried to focus on the good news. At first, it was thought that “blue-green algae” were not deadly. And
secondly, tropical disease experts from as far away as Bangkok had taken a personal interest in my case. How cool is
that? (Also, I did an excellent job of repeatedly steering the discussion back to wedding planning: "Enough about my
incurable disease. Tell me more about the flowers.)
I spent my last hours in Kathmandu completing a thirty-page survey describing every aspect of my life: Where did I
eat? What did I eat? How did I cook? Did I go swimming? Where and how often? All the others who had been
diagnosed with the disease was doing the same thing. The pathogen was eventually identified as a waterborne form of
cyanobacteria. (These bacteria are blue and are the only type of bacteria that get their energy from photosynthesis, hence
the original description of the disease as “blue-green algae.”) The disease was found to respond to treatment with
traditional antibiotics, but, interestingly, not to some of the newer ones. All these discoveries came too late to help me, but
I was lucky enough to recover quickly anyway. I had almost perfect control of my digestive system on the wedding day.

Behind every important study there is good data that made the analysis possible. And behind every bad study. . . Well,
keep reading. People often talk about “lying with statistics.” I would say that some of the most egregious statistical errors
involve lying with the data; the statistical analysis is fine, but the data on which the calculations are made are false or
inappropriate. Below are some common examples of “garbage in, garbage out.”

Selection bias. Pauline Kael, a longtime film critic for The New Yorker, reportedly said after Richard Nixon was elected
president: “Nixon couldn’t have won. “I don't know anyone who voted for him.” The quote is most likely apocryphal, but it's
a beautiful example of how a lousy sample (the liberal friend group) can offer a misleading snapshot of a larger population
(voters across America). And it introduces the question that one should always ask: How did we choose the sample or
samples that we are evaluating? If each member of the relevant population does not have an equal chance of ending up in
the sample, we will have a problem with the results that arise from that sample. A ritual of presidential politics is the Iowa
poll, in which Republican candidates descend on Ames, Iowa, in August of the year before a presidential election to court
participants, each of whom pays $30 to cast a vote in the poll. The Iowa poll doesn't tell us much about the future of the
Republican candidates. (The poll has predicted only three of the last five Republican candidates.) Because? Because
Machine Translated by
Google

Iowans who pay $30 to vote in polls are different from other Iowa Republicans; and Iowa Republicans are different from
Republican voters across the country.

Selection bias can be introduced in many other ways. A consumer survey at an airport will be biased by the fact that
people who fly are likely to be wealthier than the general public; a survey at a rest stop on Interstate 90 may have the
opposite problem. Both surveys are likely biased by the fact that people who are willing to take a survey in a public place
are different from people who would prefer not to be bothered. If you ask 100 people in a public place to complete a short
survey and 60 are willing to answer your questions

Those 60 are likely different in significant ways from the 40 who spent time without making eye contact.
One of the most famous statistical errors of all time, the infamous Literary Digest survey of 1936, was
caused by a biased sample. That year, Kansas Governor Alf Landon, a Republican, ran for president against
incumbent President Franklin Roosevelt, a Democrat. Literary Digest, an influential weekly news magazine at
the time, mailed a survey to its subscribers and to car and telephone owners whose addresses could be
extracted from public records. In total, the Literary Digest poll included 10 million likely voters, an
astronomically large sample size. As polls with good samples get bigger, they get better, as the margin of error
gets smaller. As the polls with bad samples continue to grow, the pile of garbage becomes bigger and smellier.
Literary Digest predicted that Landon would defeat Roosevelt with 57 percent of the popular vote. In fact,
Roosevelt won in a landslide, with 60 percent of the popular vote and forty-six of the forty-eight states in the
electoral college. Literary Digest's sample was "junk": The magazine's subscribers were wealthier than the
average American and therefore more likely to vote Republican, as were households with telephones and
automobiles in 1936.
2

We can end up with the same basic problem when comparing outcomes between a treatment group and a
control group if the mechanism for classifying individuals into one group or the other is not random. Let's
consider a recent finding in the medical literature about the side effects of prostate cancer treatment. There are
three common treatments for prostate cancer: surgical removal of the prostate; radiation therapy; or
brachytherapy (which involves implanting radioactive “seeds” near the cancer). Impotence is a common side
effect of prostate cancer “seeds.” Researchers have documented the sexual function of men receiving each of
the three treatments. A study of 1,000 men found that two years after treatment, 35 percent of men in the
surgery group were able to have sexual intercourse, compared with 37 percent in the radiation group and 43
percent in the brachytherapy group.
Can one look at these data and assume that brachytherapy is less likely to harm a man's sexual function? No no no. The study authors
explicitly caution that we cannot conclude that brachytherapy is better at preserving sexual function, since men who receive this treatment are
generally younger and fitter than men who receive the other treatment. The purpose of the study was simply to document the degree of sexual
side effects across all treatment types.

A related source of bias, known as self-selection bias, will arise whenever individuals volunteer to be in a
treatment group. For example, inmates who volunteer for a drug treatment group are differentiated from other
Machine Translated by
Google

inmates because they have volunteered to be in a drug treatment program. If participants in this program are
more likely to stay out of prison after their release than other prisoners, that's great, but it tells us absolutely
nothing about the value of the drug treatment program. These former inmates may have turned their lives
around because the program helped them quit drugs. Or they may have changed their lives because of other
factors that also made them more likely to volunteer for a drug treatment program (like having a really strong
desire not to go back to prison). We cannot separate the causal impact of one (the drug treatment program)
from the other (being the kind of person who volunteers for a drug treatment program).
Publication bias. Positive findings are more likely to be published than negative ones, which can skew the
results we see. Suppose you've just conducted a rigorous longitudinal study in which you conclusively
conclude that playing video games does not prevent colon cancer. He followed a representative sample of
100,000 Americans for twenty years; participants who spent hours playing video games had roughly the same
incidence of colon cancer as participants who did not play video games at all. We will assume that your
methodology is impeccable. Which prestigious medical journal is going to publish your results?

None, for two reasons. First of all, there is no solid scientific reason to believe that video games have any
impact on colon cancer, so it is not obvious why this study was being conducted. Second, and more relevant
here, the fact that something does not prevent cancer is not a particularly interesting finding. After all, most
things don't prevent cancer. Negative findings are not particularly attractive, either in medicine or elsewhere.
The net effect is to distort the research that we see or don't see. Suppose one of your graduate students
has conducted a different longitudinal study. She finds that people who spend a lot of time playing video games
have a lower incidence of colon cancer. Now that's interesting! That's exactly the kind of finding that would
catch the attention of a medical journal, the popular press, bloggers, and video game manufacturers (who
would put labels on their products extolling the health benefits of their products). It wouldn't be long before
Tiger Moms across the country were “protecting” their children from cancer by snatching books from their
hands and forcing them to play video games.

Of course, an important recurring idea in statistics is that unusual things happen from time to time, simply by
chance. If you do 100 studies, chances are one of them will come up with results that are complete nonsense, such as a statistical
association between playing video games and a lower incidence of colon cancer. Here's the problem: the 99 studies that find no link
between video games and colon cancer won't be published because they're not very interesting. The only study that finds a statistical
link will be published and receive a lot of follow-up attention. The source of bias does not arise from the studies themselves but from the
biased information that actually reaches the public. Someone reading the scientific literature on video games and cancer will find only
one study, and that one study will suggest that playing video games can prevent cancer. In fact, 99 out of 100 studies found no such
link.

Yes, my example is absurd, but the problem is real and serious. Here is the first sentence of a New York Times article about the
publication bias surrounding drugs to treat depression: “The makers of antidepressants like Prozac and Paxil never published the
results of about a third of the drug trials they conducted to win government approval. deceive doctors and consumers about the true
effectiveness of drugs.” 4 It turns out that 94 percent of the studies with positive results on the effectiveness of these drugs were
published, while only 14 percent of the studies with non-positive results were published. For patients suffering from depression, this is a
big problem. When all studies are included, antidepressants are better than placebo only by "a modest margin."
Machine Translated by
Google

To combat this problem, medical journals now typically require that any study be registered at the beginning of the project to be
eligible for publication later. This gives editors some evidence about the proportion of positive and non-positive findings. If 100 studies
are submitted that propose to examine the effect of skateboarding on heart disease, and ultimately only one is submitted for publication
with positive results, editors can infer that the other studies had non-positive results (or at least they can investigate this possibility). .

Recall bias. Memory is a fascinating thing, although it is not always a great source of good data. We have a natural human impulse to
understand the present as a logical consequence of things that happened in the past: cause and effect. The problem is that our
memories turn out to be "systematically fragile" when we try to explain some particularly good or bad outcome in the present. Consider
a study that looks at the relationship between diet and cancer. In 1993, a Harvard researcher compiled a data set that included a group
of women with breast cancer and a group of women of the same age who had not been diagnosed with cancer.

Women in both groups were asked about their eating habits earlier in life. The study yielded clear results: women with breast cancer
were significantly more likely to have had high-fat diets when they were younger.
Oh, but this wasn't actually a study on how diet affects your chances of getting cancer. This was a study of how cancer
affects a woman's memory of her diet earlier in her life. All of the women in the study had completed a dietary survey years
earlier, before any of them were diagnosed with cancer. The surprising finding was that women with breast cancer recalled a
much higher-fat diet than they actually ate; women without cancer did not. The New York Times Magazine described the
insidious nature of this withdrawal bias:

A breast cancer diagnosis hadn't just changed a woman's present and future; it had altered her past. The women with
breast cancer had (unconsciously) decided that a high-fat diet was a likely predisposition to their disease and
(unconsciously) recalled a high-fat diet. It was a pattern poignantly familiar to anyone who knows the history of this
stigmatized disease: These women, like thousands of women before them, had searched their own memories for a cause
and then invoked that cause in

5 memory.

Recall bias is one of the reasons why longitudinal studies are often preferred over cross-sectional studies. In a longitudinal
study, data is collected at the same time. At age five, a participant can be asked about his or her attitudes toward school. Then,
thirteen years later, we can revisit that same participant and determine whether he or she dropped out of high school. In a cross-
sectional study, where all data are collected at one point in time, we have to ask an eighteen-year-old who dropped out of high
school how he or she felt about school at age five, which is inherently less reliable.

Survivorship bias. Suppose the principal of a high school reports that the test scores of a particular group of students have
increased steadily for four years. The sophomores' scores for this class were better than those of the freshmen. The third year
scores were even better and the senior year scores were the best of all. We will stipulate that there is no cheating, including any
creative use of descriptive statistics. Each year, this cohort of students has done better than the previous year, by every possible
measure: mean, median, percentage of students on grade level, etc.

Would you (a) nominate this school leader as “principal of the year” or (b) demand more data?
I say "b". I smell survivorship bias, which occurs when some or many of the observations are falling out of the sample,
Machine Translated by
Google

changing the composition of the observations that remain and therefore affecting the results of any analysis. Suppose our
director is really horrible. The students at his school do not learn anything; every year half of them drop out. Well, that
could do very good things for school test scores, without any individual student doing better. If we make the reasonable
assumption that the worst performing students (with the lowest test scores) are the most likely to drop out of school, then
the average grades of the students who are left behind will steadily increase as more and more students drop out. (If you
have a room with people of different heights, forcing short people out will increase the average height in the room, but it
won't make anyone taller.)

The mutual fund industry has aggressively (and insidiously) exploited survivorship bias to make its returns appear
better to investors than they really are. Mutual funds typically compare their performance to a key benchmark for stocks,

the Standard & Poor's 500, which is an index of 500 leading U.S. mutual companies. It is said * ~ . behind
the index if the fund outperforms the index if its performance is better than it, or its performance is worse. An easy and
cheap option for investors who don't want to pay a mutual fund manager is to buy an S&P 500 index fund, which is a
mutual fund that simply buys shares of the 500 stocks in the index. Mutual fund managers like to believe that they are
smart investors, able to use their knowledge to pick stocks that will outperform a simple index fund. In fact, it is relatively
difficult to outperform the S&P 500 over a consistent period of time. (The S&P 500 is essentially an average of all the large
stocks that are traded, so, simply as a matter of mathematics, we would expect about half of actively managed mutual
funds to outperform the S&P 500 in a given year and the other half to underperform.) It doesn't seem very good to lose to
a meaningless index that simply buys 500 stocks and holds them. Without analysis. Without sophisticated macroeconomic
forecasts. And, to the delight of investors, there are no high management fees.

What should a traditional mutual fund company do? Fake data to the rescue!
This is how you can “beat the market” without beating the market. A large mutual company will open many new actively
managed funds (meaning experts pick the stocks, often with a particular focus or strategy). As an example, suppose a
mutual fund company opens twenty new funds, each of which has roughly a 50 percent chance of outperforming the S&P
500 in a given year. (This assumption is consistent with long-term data.) Now, basic probability suggests that only ten of
the firm's new funds will beat the S&P 500 in the first year; five funds will beat it two years in a row; and two or three will
beat it three years in a row.

Here comes the clever part. At that point, new mutual funds with unimpressive returns relative to the S&P
500 quietly close. (Its assets are integrated into other existing funds.) The firm can then heavily advertise the
two or three new funds that have “consistently outperformed the S&P 500,” even if that performance is the
stock-picking equivalent of pulling three heads in a row. . Subsequent performance of these funds is likely to
revert to the mean, albeit after investors have accumulated. The number of mutual funds or investment gurus
that have consistently outperformed the S&P 500 over a long period. . . - * is surprisingly small.

Healthy user bias. People who take vitamins regularly are likely to be healthy, because they are the kind of
people who take vitamins regularly! Whether vitamins have any impact is a separate issue. Consider the
following thought experiment. Suppose public health officials promulgate the theory that all new parents should
put their children to bed in only purple pajamas, because that helps stimulate brain development. Twenty years
later, longitudinal research confirms that having worn purple pajamas as a child has an overwhelmingly large
positive association with success in life. We found, for example, that 98 percent of Harvard freshmen wore
purple pajamas as children (and many still do), compared with just 3 percent of inmates in the Massachusetts
Machine Translated by
Google

state prison system.


Of course, purple pajamas don't matter, but having the kind of parents who put their children in purple
pajamas does matter. Even when we try to control for factors like parental education, we're still left with
unobservable differences between parents who obsess over dressing their kids in purple pajamas and those
who don't. As New York Times health editor Gary Taubes explains: “At its simplest, the problem is that people
who faithfully engage in activities that are good for them — taking a medicine as prescribed, for example, or
eating what they believe to be a healthy diet — are fundamentally different from those who don’t.” 6 This effect
can potentially confound any study attempting to assess the actual effect of perceived healthy activities, such
as exercising regularly or eating kale. We think we are comparing the health effects of two diets: kale versus no
kale. In fact, if the treatment and control groups are not randomly assigned, we are comparing two diets
consumed by two different types of people. We have a treatment group that is different from the control group
in two ways, instead of just one.

If statistics is detective work, then data are the clues. My wife spent a year teaching high school students in
rural New Hampshire. One of his students was arrested for breaking into a hardware store and stealing some
tools. The police were able to solve the box because (1) it had just snowed and there were footprints in the snow leading
from the hardware store to the student's house; and (2) the stolen tools were found inside. Good leads help.
As good data. But first you have to get good data, and that's a lot harder than it sounds.

* At that time, the illness had a mean duration of forty-three days with a standard deviation of twenty-four days.
* The S&P 500 is a good example of what an index can and should do. The index consists of the stock prices of the 500 leading US companies. US,
each weighted by market value (so that larger companies have more weight in the index than smaller ones). The index is a simple and accurate indicator
of what is happening with the stock prices of the largest U.S. companies at any given time.
* For a very interesting discussion of why you should probably buy index funds instead of trying to beat the market, read A Random Walk Down Wall
Street by my former professor Burton Malkiel.
Machine Translated by
Google

CHAPTER 8

The central limit theorem


The Lebron James of statistics

Sometimes statistics seem almost magical. We can draw broad and powerful conclusions from relatively little data. In
some ways we can gain meaningful insight into a presidential election by polling just a thousand American voters. We can
test a hundred chicken breasts for salmonella in a poultry processing plant and conclude, from that sample alone, whether
the entire plant is safe or not. Where does this extraordinary power to generalize come from?

Much of it comes from the central limit theorem, which is the Lebron James of statistics, if Lebron were also a
supermodel, a Harvard professor, and a Nobel Peace Prize winner. The central limit theorem is the “powerhouse” of many
statistical activities that involve using a sample to make inferences about a large population (such as a survey or a
salmonella test). These types of inferences may seem mystical; in fact, they are just a combination of two tools we have
already explored: probability and appropriate sampling. Before we dive into the mechanics of the central limit theorem
(which is not that complicated), here is an example to give you a general intuition.

Suppose you live in a city that hosts a marathon. Runners from all over the world will be competing, which means
many of them do not speak English. Race logistics require runners to register on race morning, after which they are
randomly assigned to buses that will take them to the starting line. Unfortunately one of the buses gets lost on the way to
the race.
(Okay, you'll have to assume that no one has a cell phone and that the driver doesn't have a GPS navigation device;
unless you want to do a lot of nasty math right now, just do it.) As a civic leader of this city, you join the search team.

As luck would have it, near your home you stumble upon a broken-down bus with a large group of disgruntled
international passengers, none of whom speak English. This must be the missing bus! You're going to be a hero! Except
you have one nagging doubt: the passengers on this bus are, well, very big.
At a quick glance, it is estimated that the average weight of this group of passengers must be over 220 pounds. There is
no way for a random group

of marathon runners could be this heavy. He sends his message by radio to search headquarters: “I think it's the wrong
bus. Keep watching."
A closer look confirms your initial impression. When a translator arrives, you discover that this broken-down bus was
headed to the International Sausage Festival, which is also taking place in your city that same weekend. (For the sake of
verisimilitude, it's entirely possible that sausage festival participants also wear sweatpants.)

Congratulations. If you can understand how someone who takes a quick look at the weights of the passengers on a
Machine Translated by
Google

bus can infer that they are probably not on their way to the starting line of a marathon, then you understand the basic idea
of the central limit theorem. The rest is just fleshing out the details. And if you understand the central limit theorem, most
forms of statistical inference will seem relatively intuitive to you.

The central principle underlying the central limit theorem is that a large, correctly drawn sample will resemble the
population from which it is drawn. Obviously there will be variation from sample to sample (for example, each bus heading
to the start of the marathon will have a slightly different mix of passengers), but the likelihood of any sample deviating
wildly from the underlying population is very low. This logic is what allowed him to make a quick judgment when he
boarded the broken-down bus and saw the average circumference of the passengers on board. Many important people
run marathons; there are probably hundreds of people weighing more than 200 pounds in any given race. But most
marathon runners are relatively thin. Therefore, the probability that so many of the larger riders were randomly assigned to
the same bus is very, very low. One could conclude with a reasonable degree of confidence that this was not the missing
marathon bus. Yes, you could have been wrong, but probability tells us that most of the time you would have been right.

That is the basic intuition behind the central limit theorem. When we add some statistical details, we can quantify the probability that
he is right or wrong. For example, we might calculate that in a marathon field of 10,000 runners with an average weight of 155 pounds,
there is less than a 1 in 100 chance that a random sample of 60 of those runners (our missed bus) would have an average weight of
155 pounds. 220 pounds or more. For now, let's stick with intuition; there will be plenty of time to do the math later. The central limit
theorem allows us to make the following inferences, all of which will be explored in greater depth in the next chapter.

1. If we have detailed information about some population, then we can make powerful inferences about any
suitably drawn sample from that population.

population. For example, suppose the principal of a school has detailed information about the
standardized test scores of all the students in the school.
your school (mean, standard deviation, etc.). That is the relevant population. Now suppose a school
district bureaucrat arrives next week to administer a similar standardized test to 100 randomly selected
students. The performance of those 100 students, the sample, will be used to evaluate the performance
of the school as a whole.
How confident can the principal be that the performance of those 100 randomly selected students will
accurately reflect how the entire student body has performed on similar standardized tests? Quite.
According to the central limit theorem, the average test score for a random sample of 100 students will
normally not deviate markedly from the average test score for the entire school.
2. If we have detailed information about a correctly drawn sample (mean and standard deviation), we
can make surprisingly accurate inferences about the population from which that sample was drawn.
Basically, this works in the opposite direction of the previous example, putting us in the shoes of the
school district bureaucrat who is evaluating several schools in the district. Unlike the school principal, this
Machine Translated by
Google

bureaucrat does not have (or does not trust) the standardized test score data that the principal has for all
students at a particular school, which is the relevant population. Instead, it will administer a similar test to
a random sample of 100 students at each school.

Can this administrator be reasonably confident that the overall performance of any given school can
be fairly assessed based on the test scores of a sample of only 100 students from that school? Yeah.
The central limit theorem tells us that a large sample will typically not deviate markedly from its
underlying population, meaning that the sample results (scores of the 100 randomly chosen students)
are a good proxy for the results of the general population (the student body). in a particular school). Of
course, that's how polls work. A methodologically sound survey of 1,200 Americans can tell us a lot
about how the entire country thinks.
Think about it: if not. 1 above is true, no. 2 must also be true and vice versa. If a sample normally
resembles the population from which it is drawn, then it must also be true that a population will normally
resemble a sample drawn from that population. (If children normally resemble their parents, then parents
should also resemble their children.)
3. If we have data describing a particular sample and data about a particular population, we can infer
whether or not that sample is consistent with a sample likely to be drawn from that population. This is the
example of the missing bus described at the beginning of the chapter. We know the average weight
(more or less) of the participants in the marathon. And we know the average weight (more or less) of the
passengers on the broken down bus. The central limit theorem allows us to calculate the probability that
a particular sample (the chubby people on the bus) was drawn from a given population (the marathon
field). If that probability is low, then we can conclude with a high degree of confidence that the sample
was not taken from the population in question (for example, the people on this bus don't really look like a
group of marathon runners heading for the start). line).
4. Finally, if we know the underlying characteristics of two samples, we can infer whether both samples
were likely drawn from the same population. Let's go back to our (increasingly absurd) bus example. We
now know that a marathon is being held in the city, in addition to the International Sausage Festival.
Suppose both groups have thousands of participants, and both groups operate buses, all loaded with
random samples of marathon runners or sausage enthusiasts. Suppose further that two buses collide.
(I've already admitted that the example is absurd, so read on.) As a civic leader, you arrive at the scene
and are tasked with determining whether or not both buses were headed to the same event (sausage
festival or marathon). ). Miraculously, no one on any of the buses speaks English, but the paramedics
give you detailed information about the weight of all the passengers on each bus.

From that alone, one can infer whether the two buses were likely headed to the same event or to
different events. Again, think about this intuitively. Suppose the average weight of passengers on a bus
is 157 pounds, with a standard deviation of 11 pounds (meaning that a high proportion of passengers
weigh between 146 and 168 pounds).
Now suppose that the passengers on the second bus have a mean weight of 211 pounds with a
standard deviation of 21 pounds (meaning that a high proportion of the passengers weigh between 190
Machine Translated by
Google

pounds and 232 pounds).


Forget about statistical formulas for a moment and just use logic: does it seem likely that the passengers
on those two buses randomly came from the same population?
No. It seems much more likely that one bus will be full of marathon runners and the other will be full of
sausage enthusiasts. In addition to the difference in average weight between the two buses, you can
also see that the variation in weights between the two buses is very large compared to the variation in
weights within each bus. The people who weigh one standard deviation above the mean on the “thin”
bus are 168 pounds, which is less than the people who are one standard deviation below the mean on
the “other” bus (190 pounds). This is a telltale sign (both statistically and logically) that the two samples
likely came from different populations.

If all of this makes intuitive sense, then you are 93.2 percent of the way to ..,
understanding
the central limit theorem. We need to go a step further to put some technical weight behind intuition. Obviously,
when you poked your head inside the broken-down bus and saw a group of big people in sweatpants, you had
a “gut feeling” that they weren’t marathon runners. The central limit theorem allows us to go beyond that hunch
and assign a degree of confidence to our conclusion.

For example, some basic calculations will allow me to conclude that 99 times out of 100 the average weight
of any randomly selected bus of marathoners will be within nine pounds of the average weight of the entire
marathon field. That's what gives statistical weight to my hunch when I stumble upon the broken-down bus.
These riders have an average weight twenty-one pounds heavier than the average marathon weight,
something that should only occur by chance less than once in 100. As a result, I can reject the hypothesis that
this is a missing weight. marathon bus with 99 percent confidence, meaning I should expect my inference to be
correct 99 times out of 100.
And yes, probability suggests that, on average, I will be wrong 1 time in 100.

All of this analysis stems from the central limit theorem, which, from a statistical standpoint, has LeBron James-
like power and elegance. According to the central limit theorem, sample means of any population will be
distributed approximately as a normal distribution around the population mean. Wait a moment while we
analyze that statement.

1. Suppose we have a population, like our marathon field, and we are interested in the weights of its
members. Any sample of riders, such as each bus of sixty riders, will have a mean.
2. If we take repeated samples, such as selecting random groups of sixty runners from the field over and
over again, then each of those samples will have its own mean weight. These are the sample media.
3. Most sample means will be very close to the population mean. Some will be a little taller. Some will be
a little lower. By chance alone, very few will be significantly taller than the population average, and very
few will be significantly shorter.

Cue the music, because this is where it all comes together in a powerful crescendo. . .
4. The central limit theorem tells us that sample means will be distributed approximately as a normal distribution around
the population mean.
Machine Translated by
Google

The normal distribution, as you will recall from Chapter 2, is the bell-shaped distribution (for example, the height of adult males) in which
68 percent of the observations lie within one standard deviation of the mean, 95 percent lie within two standard deviations, and 95
percent lie within two standard deviations. soon.

5. All of this will be true no matter what the underlying population distribution is. The population from which the samples
are drawn does not have to have a normal distribution for the sample means to be normally distributed.

Let's think about some real data, say, the distribution of household income in the United States. Household income is not
normally distributed in the United States; instead, it tends to be skewed to the right. No household can earn less than $0 in any
given year, so that must be the lower limit of the distribution. Meanwhile, a small group of households can earn astonishingly
high annual incomes – hundreds of millions or even billions of dollars in some cases. As a result, we would expect the
distribution of household income to have a long right tail, something like this:

Annual household income

The median household income in the United States is approximately $51,900; the median family income is $70,900. 1
(People like Bill Gates shift income
familiar half to the right, just as he did when he entered the bar in Chapter 2). Now suppose we take a random sample of 1,000
American households and collect information on annual household income. Based on the above information and the central limit
theorem, what can we infer about this sample?

It turns out, quite a lot. First, our best guess about what the meaning of any sample is will be the mean of the population from
which it is drawn. The goal of a representative sample is to resemble the underlying population. A properly
drawn sample will, on average, resemble the United States. There will be hedge fund managers, homeless
people, police officers, and everyone else, all roughly in proportion to their frequency in the population. We
would therefore expect the median household income for a representative sample of 1,000 American
households to be about $70,900. Is that exactly it? No. But it shouldn't be tremendously different either.
If we took several samples of 1,000 households, we would expect the different sample means to cluster
around the population mean, $70,900. We would expect some means to be higher and some to be lower.
Could we get a sample of 1,000 households with a median household income of $427,000? Sure, that's
possible, but highly unlikely. (Remember, our sampling methodology is sound; we're not conducting a survey in
the parking lot of the Greenwich Country Club.) It's also highly unlikely that an adequate sample of 1,000
American households would have a median income of $8,000.
Machine Translated by
Google

All of that is just basic logic. The central limit theorem allows us to go one step further by describing the
expected distribution of those different sample means as they cluster around the population mean. Specifically,
the sample means will form a normal distribution around the population mean, which in this case is $70,900.
Remember, the shape of the underlying population doesn't matter. The distribution of household income in the
United States is quite skewed, but the distribution of sample means will not be. If we took 100 different
samples, each with 1,000 households, and graphed the frequency of our results, we would expect those
sample means to form the familiar "bell-shaped" distribution around $70,900.

The greater the number of samples, the closer the distribution will be to the normal distribution. And the
larger the size of each sample, the narrower that distribution will be. To test this result, let's do a fun
experiment with real data on the weight of real Americans. The University of Michigan conducts a longitudinal
study called Americans' Changing Lives, which involves detailed observations of several thousand American
adults, including their weight. The weight distribution is slightly skewed to the right, because it is biologically
easier to be 100 pounds overweight than 100 pounds underweight. The average weight of all adults in the
study is 162 pounds.

Using basic statistical software, we can instruct the computer to take a random sample of 100 people from
the Changing Lives data. In fact, we can do this over and over again to see how the results match what the
central limit theorem would predict. Here is a graph of the distribution of 100 sample means (rounded to the
nearest pound) randomly generated from the Changing Lives data.

100 sample means, n = 100

The larger the sample size and the more samples that are taken, the closer the distribution of sample means will
approximate the normal curve. (As a general rule, the sample size must be at least 30 for the central limit theorem to be
valid.) This makes sense. A larger sample is less likely to be affected by random variation. A sample of 2 can be highly
biased by one particularly large or small person. In contrast, a sample of 500 will not be unduly affected by a few
particularly large or small individuals.
We are now very close to making all our statistical dreams come true! The sample means are distributed
approximately like a normal curve, as described above. The power of a normal distribution comes from the fact that we
know approximately what proportion of observations will lie within one standard deviation above or below the mean (68
percent); what proportion of observations will lie within two standard deviations above or below the mean (95 percent);
Machine Translated by
Google

and so on. This is powerful stuff.


Earlier in this chapter I pointed out that we could intuitively infer that a busload of passengers with an average weight
twenty-five pounds greater than the average weight of the entire marathon was probably not the missing busload of
runners. To quantify that intuition (to be able to say that this inference will be correct 95 percent of the time, or 99 percent,
or 99.9 percent), we need just one more technical concept: the standard error.

The standard error measures the dispersion of the sample means. How accurately do we expect the sample means to
cluster around the population mean? There is some potential confusion here, since we have now introduced two different
measures of dispersion: the standard deviation and the standard error. Here's what you need to remember to keep them
in order:

1. The standard deviation measures the dispersion in the underlying population.


In this case, you could measure the dispersion of weights for all participants in the Framingham Heart Study,
or the dispersion around the mean for the entire marathon field.
2. The standard error measures the dispersion of the sample means. If we draw repeated samples from
100 participants in the Framingham Heart Study, what will be the spread of those sample means?
3. This is what unites the two concepts: The standard error is the standard deviation of the sample
means! Isn't it great?
A large standard error means that the sample means are widely distributed around the population mean; a
small standard error means that they are relatively closely clustered. Here are three real-life examples from
Changing Lives data.

100 sample means, n = 20


Machine Translated by
Google

100 sample means, n = 100

Female population only/100 Sample means, n = 100

The second distribution, which has a larger sample size, is more closely clustered around the mean than
the first distribution. The larger sample size makes it less likely that the sample mean will deviate markedly
from the population mean. The final set of sample means is drawn from only a subset of the population, the
women in the study. Since the weights for women in the data set are less diffuse than the weights for all
individuals in the population, it stands to reason that the weights for samples drawn from only women are less
diffuse than for samples drawn from the entire Changing Lives population. (These samples are also clustered
around a slightly different population mean, as the mean weight of all women in the Changing Lives study is
different from the mean weight of the entire study population.)
The pattern you saw above is generally valid. Sample means will cluster more closely around the population
mean as the size of each sample increases (for example, our sample means were more clustered when we
sampled 100 rather than 30). And sample means will cluster less closely around the population mean when the
underlying population is more dispersed (for example, our sample means for the entire Changing Lives
population were more dispersed than the sample means for just the women in the study).
If you have followed the logic so far, then the formula for the standard error follows naturally:
yn = s{n,
where s is the standard deviation of the population from which the sample is drawn SE is the sample size.
Keep your head on straight!
Don't let the appearance of the letters ruin basic intuition. The standard error will be large when the standard
deviation of the underlying distribution is large.
Machine Translated by
Google

A large sample drawn from a widely dispersed population is likely to be very scattered; a large sample from a
population closely clustered around the mean is also likely to be closely clustered around the mean. If we still
look at weight, we would expect the standard error of a sample drawn from the entire Changing Lives
population to be larger than the standard error of a sample drawn only from men in their twenties. This is why
the standard deviations are in the numerator.
Similarly, we would expect the standard error to decrease as sample size increases, since large samples are less likely to be
distorted by extreme outliers. That is why the sample size (n) is in the denominator. (The reason we take the square root of n will
be left for a more advanced text; the basic relationship is what is important here.)

In the case of the Changing Lives data, we actually know the population standard deviation; often that's not the case. For
large samples, we can assume that the sample standard deviation is reasonably close to the population standard deviation.
*
Finally, we have come to the reward of all this. Because sample means are normally distributed (thanks to the central limit
theorem), we can harness the power of the normal curve. We expect that approximately 68 percent of all sample means will lie
within one standard error of the population mean; 95 percent of the sample means will lie within two standard errors of the
population mean; and 99.7 percent of the sample means will lie within three standard errors of the population mean.

Frequency distribution of sample means

So let's go back to a variation of our missed bus example, only now we can replace intuition with numbers. (The example
itself will still be absurd; the next chapter will have many less absurd real-world examples.) Suppose the Changing Lives study
has invited all the individuals in the study to meet in

Boston for a weekend of data collection and revelry. Participants are randomly loaded onto buses and transported
Machine Translated by
Google

between testing facility buildings, where they are weighed, measured, pricked, spiked, etc. Surprisingly, one bus goes
missing, an event that is broadcast on the local news. Around this time, you're returning from the Sausage Festival when
you see a crashed bus on the side of the road. The bus apparently swerved to avoid a wild fox crossing the road, and all
passengers are unconscious but not seriously injured. (I need them not to communicate for the example to work, but I
don't want their injuries to be too worrying.) Paramedics on scene inform you that the average weight of the 62
passengers on the bus is 194 pounds. Additionally, the fox that the bus swerved to avoid was cut slightly and appears to
have a broken hind leg.
Fortunately, you know the mean weight and standard deviation of the entire Changing Lives population, have a
working knowledge of the central limit theorem, and know how to administer first aid to a wild fox. The mean weight of
Changing Lives participants is 162; the standard deviation is 36.
From that information, we can calculate the standard error for a sample of 62 people (the number of unconscious
passengers on the bus): s/62 = 36/7.9, or 4.6.
The difference between the sample mean (194 pounds) and the population mean (162 pounds) is 32 pounds, or much
more than three standard errors. We know from the central limit theorem that 99.7 percent of all sample means will lie
within three standard errors of the population mean. That makes it extremely unlikely that this bus represents a random
group of Changing Lives participants.
In his duty as a civic leader, he calls the study officials to tell them that this probably isn't the missing bus, only now he can
offer statistical evidence, rather than just "a hunch." He tells the Changing Lives people that he can reject the possibility
that this is the missing bus with a confidence level of 99.7 percent. And because you're talking to researchers, they really
understand what you're talking about.

Their analysis is further confirmed when paramedics perform blood tests on the bus passengers and discover that the
mean cholesterol level for all bus passengers is five standard errors above the mean cholesterol level of the Changing
Lives study participants. This suggests, correctly as will later be seen, that the unconscious passengers are involved in the
Sausage Festival.
[There is a happy ending. When the bus passengers regained consciousness, Changing Lives study officials offered
them counseling about the dangers of a diet high in saturated fat, leading many of them to adopt more heart-healthy
eating habits. Meanwhile, the fox was nursed back to health at a wildlife reserve. . . and ... . . . . . * , local wildlife
and was eventually released back into the wild.]
I've tried to stick to the basics in this chapter. You should note that for the central limit theorem to apply, sample sizes
must be relatively large (more than 30 as a general rule). We also need a relatively large sample if we are to assume that
the standard deviation of the sample is approximately the same as the standard deviation of the population from which it is
drawn. There are many statistical corrections that can be applied when these conditions are not met, but that's all the icing
on the cake (and maybe even a pinch of the icing on the cake).
The “big picture” here is simple and tremendously powerful: 1.

If large random samples are drawn from any population, the means of those samples will be normally distributed
around the population mean (regardless of what the underlying population distribution looks like). ).
2. Most sample means will be reasonably close to the population mean; the standard error is what defines
"reasonably close."
3. The central limit theorem tells us the probability that a sample mean lies within a certain distance from the
Machine Translated by
Google

population mean. It is relatively unlikely for a sample mean to be more than two standard errors from the population
mean, and extremely unlikely for it to be three or more standard errors from the population mean.
4. The less likely it is that a result was observed by chance, the more confident we can be in assuming that some
other factor is at play.

That's pretty much what statistical inference is all about. The central limit theorem is what makes most of this possible.
And until LeBron James wins as many NBA championships as Michael Jordan (six), the central limit theorem will be much
more impressive than he is.

* Note the clever use of false precision here.


* When the population standard deviation is calculated from a smaller sample, the formula is modified slightly: this helps to account for the fact that the
spread in a small sample may underestimate the spread of the entire population. This is not very relevant to the main points of this chapter.

* My colleague at the University of Chicago, Jim Sallee, makes a very important critique of the missing bus examples. He points out that very few buses
get lost. So if we are looking for a missing bus, any bus that appears missing or crashed will likely be that bus, regardless of the weight of passengers on
the bus. He is right. (Think about it: If you lost your child in a supermarket and the store manager told you there was a missing child near register six, you
would immediately conclude that it was probably your child.) So, let's add one more element of absurdity to these examples and pretend that buses get
lost all the time.

CHAPTER 9

Inference
Why my statistics teacher thought I
might have cheated

In the spring of my senior year of college, I took a statistics class. I wasn't particularly enamored of statistics or
most math-based disciplines at the time, but I had promised my dad that I would take the course if I could leave
school for ten days to go on a family trip to the Soviet Union. So, I basically took statistics in exchange for a trip
to the USSR. This turned out to be a great deal, both because I liked statistics more than I thought and
because I was able to visit the USSR in the spring of 1988. Who would have known that the country would not
exist in its communist form for long? more extensive?
This story is actually relevant to the chapter; the thing is, I wasn't as dedicated to my statistics course during
the semester as I could have been. Among other responsibilities, I was also writing an honors thesis that was
due approximately halfway through the semester. We had regular quizzes in the statistics course, many of
which I ignored or failed. I studied a little for the midterm exam and did pretty well, literally. But a few weeks
before the end of the semester, two things happened. First, I finished my thesis, which gave me a lot of free
time. And second, I realized that statistics wasn't as difficult as I had thought. I started studying the statistics
book and doing the work earlier in the course. I got an A on the final exam.
That's when my statistics professor, whose name I've long forgotten, called me into his office. I don't
remember exactly what he said, but it was something like, "You actually did a lot better in the final than you did
in the midterm." This was not a congratulatory visit during which I was recognized for finally doing some
Machine Translated by
Google

serious work in class. There was an implicit (though not explicit) accusation in his citation; the expectation was
that I would explain why I performed better on the final exam than on the midterm. In short, this guy suspected
that I might have cheated. Now that I have been teaching for many years, I am more sympathetic to your line of
thinking. In almost every course I have taught, there is a surprising degree of correlation between a student's
performance on the midterm exam and the final exam. It is very unusual for a student
score below average on the midterm exam and then near the top of the class on the final.
I explained that I had finished my thesis and had taken the class seriously (doing things like reading the assigned
textbook chapters and doing homework). He seemed pleased with this explanation and I left, still somewhat uneasy about
the implied accusation.

Believe it or not, this anecdote embodies much of what you need to know about statistical inference, including both its
strengths and potential weaknesses. Statistics cannot prove anything with certainty. Instead, the power of statistical
inference comes from observing some pattern or outcome and then using probability to determine the most likely
explanation for that outcome.
Suppose a strange gambler comes to town and offers you a bet: you win $1,000 if he rolls a six on a single die; you win
$500 if he rolls anything else—a pretty good bet from your point of view. He then proceeds to roll ten sixes in a row, taking
$10,000 from you.

One possible explanation is that he got lucky. An alternative explanation is that he cheated in some way. The
probability of rolling ten sixes in a row on a fair die is approximately 1 in 60 million. You can't prove he cheated, but you
should at least inspect the die.
Of course, the most likely explanation is not always the correct one.
Extremely strange things happen. Linda Cooper is a South Carolina woman who was struck by lightning four times.
(The Federal Water Management Administration)
ER estimates that the chance of being struck by lightning just once is 1 in 600,000.) Linda Cooper's insurance company
cannot deny her coverage simply because her injuries are statistically improbable. Going back to my undergraduate
statistics exam, the professor had reasonable grounds to be suspicious. He saw a pattern that was highly unlikely; this is
exactly how investigators detect cheating on standardized tests and how the SEC detects insider trading. But an unlikely
pattern is just an unlikely pattern unless it is corroborated by additional evidence. Later in this chapter we will discuss the
errors that can arise when probability leads us astray.
For now, we should appreciate that statistical inference uses data to address important questions. Is this a new drug
effective in treating heart disease? Do mobile phones cause brain cancer? Note that I am not claiming that statistics can
answer these types of questions unambiguously; instead, inference tells us what is likely and what is unlikely.
Researchers cannot prove that a new drug is effective in treating heart disease, even when they have data from a
carefully controlled clinical trial. After all, it is quite possible that there is random variation in patient outcomes in the
treatment and control groups that is unrelated to the new drug. If 53 out of 100 patients taking the new heart
disease drug showed marked improvement compared with 49 out of 100 patients given a placebo, we would
not immediately conclude that the new drug is effective. This is a result that can easily be explained by random
variation between the two groups rather than by the new drug.
Machine Translated by
Google

But suppose instead that 91 out of 100 patients receiving the new drug show marked improvement,
compared with 49 out of 100 patients in the control group. It is still possible that this impressive result is not
related to the new drug; patients in the treatment group may be particularly lucky or resistant. But now that is a
much less likely explanation. In the formal language of statistical inference, researchers would probably
conclude the following: (1) If the experimental drug had no effect, we would rarely see this amount of variation
in outcomes between those receiving the drug and those taking the placebo. . (2) Therefore, it is highly unlikely
that the drug will have no positive effect.
(3) The alternative (and most likely) explanation for the observed data pattern is that the experimental drug has
a positive effect.
Statistical inference is the process by which data speaks to us and allows us to draw meaningful
conclusions. This is the reward! The goal of statistics is not to do countless rigorous mathematical calculations;
the goal is to gain insight into significant social phenomena. Statistical inference is really just the union of two
concepts we've already discussed: data and probability (with a little help from the central limit theorem). In this
chapter I have taken an important methodological shortcut. All examples will assume that we are working with
large, properly extracted samples. This assumption means that the central limit theorem applies and the mean
and standard deviation of any sample will be approximately the same as the mean and standard deviation of
the population from which it is drawn. Both of these things make our calculations easier.
Statistical inference does not depend on this simplifying assumption, but various methodological
workarounds for dealing with small samples or imperfect data often hinder understanding the bigger picture.
The purpose here is to introduce the power of statistical inference and explain how it works. Once you get that,
it's pretty easy to add complexity.

One of the most common tools in statistical inference is hypothesis testing. Actually, I have already introduced
this concept, just without the fancy terminology. As noted above, statistics alone cannot prove anything;
instead, we use statistical inference to accept or reject explanations on the basis of their relative likelihood. To
be more precise, any statistical inference begins with an implicit or explicit null hypothesis. This is our initial
assumption, which will be rejected or not based on further statistical analysis. If we reject the null hypothesis, we usually
accept some alternative hypothesis that is more consistent with the observed data. For example, in a court of law the
initial assumption, or null hypothesis, is that the defendant is innocent. The prosecution's job is to persuade the judge or
jury to reject that assumption and accept the alternative hypothesis, which is that the defendant is guilty. As a matter of
logic, the alternative hypothesis is a conclusion that must be true if we can reject the null hypothesis. Let's consider some
examples.

Null hypothesis: This new experimental drug is no more effective at preventing malaria than a placebo.
Alternative hypothesis: This new experimental drug may help prevent malaria.

The data: One group is randomly chosen to receive the new experimental drug and a control group receives a placebo.
At the end of a period of time, the group receiving the experimental drug has far fewer cases of malaria than the control
group. This would be an extremely unlikely outcome if the experimental drug had no medical impact. As a result, we reject
the null hypothesis that the new drug has no impact (beyond that of a placebo) and accept the logical alternative, which is
our alternative hypothesis: this new experimental drug may help prevent malaria.
Machine Translated by
Google

This methodological approach is so strange that we should give one more example. Again, note that the null
hypothesis and the alternative hypothesis are logical complements. If one is true, the other is not. Or, if we reject one
statement, we must accept the other.

Null hypothesis: Substance abuse treatment for prisoners does not reduce their
Rate of new arrests after release from prison.

Alternative hypothesis: Substance abuse treatment for prisoners will make it less likely that they will be arrested again
after their release.
The (hypothetical) data: Prisoners were randomly assigned to two groups; the “treatment” group received substance
abuse treatment and the control group did not. (This is one of those interesting times when the treatment group actually
gets treated!) After five years, both groups have similar readmission rates.
, , ....................................................................... * . , . , , , .
In this case, we cannot reject the null hypothesis. The data has not given us any
reason to discard our initial assumption that substance abuse treatment is not an effective tool to prevent ex-offenders
from returning to prison.
It may seem counterintuitive, but researchers often create a null hypothesis in the hopes of being able to reject it. In
both examples above, a research “success” (finding a new malaria drug or reducing recurrence) involved rejecting the null
hypothesis. The data made this possible only in one of the cases (the

anti-malaria drug).

In a court of law, the threshold for rejecting the presumption of innocence is the qualitative assessment that the accused is
“guilty beyond a reasonable doubt.”
The judge or jury must define exactly what that means. Statistics leverage the same basic idea, but instead “guilty beyond
a reasonable doubt” is defined quantitatively. Researchers often ask: If the null hypothesis is true, what is the probability
that we would observe this data pattern by chance? To use a familiar example, medical researchers might ask: If this
experimental drug has no effect on heart disease (our null hypothesis), what is the likelihood that 91 out of 100 patients
who receive the drug will show improvement compared with only 49 out of 100? Do patients receive a placebo? If the data
suggest that the null hypothesis is extremely unlikely (as in this medical example), then we should reject it and accept the
alternative hypothesis (that the drug is effective in treating heart disease).
In that regard, let's revisit the Atlanta standardized cheating scandal that is alluded to at several points in the book.
Atlanta's test results were first flagged due to a large number of "bad to right" erasures. Obviously, students taking
standardized tests erase answers all the time. And some groups of students may be particularly lucky in their changes,
without necessarily having to cheat. For that reason, the null hypothesis is that the standardized test scores for any
particular school district are legitimate and that any irregular pattern of deletions is simply a product of chance. We
certainly do not want to punish students or administrators because an unusually high proportion of students made sensible
changes to their answer sheets in the final minutes of an important state exam.
But “unusually high” doesn’t quite describe what was happening in Atlanta. Some classrooms had answer sheets
where the number of erasures from incorrect to correct was twenty to fifty standard deviations above the state norm. (To
put this in perspective, remember that most observations in a distribution typically fall within two standard deviations of the
mean.) So what were the chances that the Atlanta students would just erase a large number of incorrect answers and
Machine Translated by
Google

replace them with correct answers as if nothing had happened? a matter of chance? The official who analyzed the data
described the probability of the Atlanta pattern occurring without cheating as roughly equal to the chance that 70,000
tall.
people would show up to a football game at the Georgia Dome and all be over seven feet Could it happen? Yeah. Is it
likely? Not so much.
Georgia officials still couldn't convict anyone of misconduct, just as my teacher couldn't (and shouldn't) have had me
expelled from school because
My final statistics exam grade was out of sync with my mid-semester grade. Atlanta officials could not prove any cheating
was taking place. However, they could reject the null hypothesis that the results were legitimate. And they could do so with
a "high degree of confidence," meaning that the observed pattern was almost impossible among normal test takers.
Therefore, they explicitly accepted the alternative hypothesis, which is that something fishy was going on. (I suspect they
used language that seemed more official.) In fact, subsequent investigation uncovered the “smoking drafts.” There were
reports of teachers changing answers, giving away answers, allowing low-scoring children to copy from high-scoring
children, and even pointing out answers while standing in front of students' desks. The most egregious cheating involved a
group of teachers who held a weekend pizza party during which they reviewed test papers and changed students'
answers.

In the Atlanta example, we could reject the null hypothesis of “no cheating” because the pattern of test results was
wildly unlikely in the absence of foul play. But how implausible does the null hypothesis have to be before we can reject it
and invite some alternative explanation?
One of the most common thresholds researchers use to reject a null hypothesis is 5 percent, which is often written in
decimal form: 0.05. This probability is known as the significance level and represents the upper limit of the probability of
observing some data pattern if the null hypothesis were true. Stay with me for a moment, because it's really not that
complicated.
Let's consider a significance level of .05. We can reject a null hypothesis at the 0.05 level if there is less than a 5
percent chance of obtaining a result at least as extreme as what we would have observed if the null hypothesis were true.
A simple example can make this much clearer. I hate to do this to you, but assume once again that you've been assigned
missed bus duties (partly due to your valiant efforts in the last chapter). Only now he's working full-time for the Changing
Lives study researchers, and they've provided him with excellent data to help inform his work. Each bus operated by the
study organisers has approximately 60 passengers, so we can treat the passengers on any given bus as a random
sample drawn from the entire Changing Lives population. One morning you wake up to the news that a pro-obesity
terrorist group has hijacked a bus in the Boston area.
Their job is to drop from a helicopter onto the roof of the moving
bus, sneak inside through the emergency exit, and then stealthily determine whether passengers are Changing Lives
participants based solely on their weight. (Seriously, this is no more far-fetched than most action-adventure plots, and it's
far more educational.)
When the helicopter takes off from the command base, you are given a machine.
pistol, several grenades, a watch that also functions as a high-resolution video camera, and the data we calculated in the
last chapter on the mean weight and standard error of the samples taken from the Changing Lives participants. Any
random sample of 60 participants will have an expected mean weight of 162 pounds and a standard deviation of 36
Machine Translated by
Google

pounds, since that is the mean and standard deviation of all the participants in the study (the population). With that data,
we can calculate the standard error of the sample mean: At Mission Control, the following distribution is scanned onto the
inside of your right retina, so you can consult it after you sneak onto the moving bus and secretly weigh all the
passengers. inside.

Distribution of sample means

Mean weight for the sample

As the distribution above shows, we would expect about 95 percent of all 60-person samples drawn from Changing
Lives participants to have a mean weight within two standard errors of the population mean, or roughly between . ... *
.................................................................................................. . . 153 and 171 pounds. In contrast, only 5 out of 100
times a sample of
60 people randomly selected from the Changing Lives participants would have an average weight of either over 171
pounds or under 153 pounds.
(You are performing what is known as a “two-tailed” hypothesis test; the difference between this and a “one-tailed” test will
be discussed in an appendix at the end of the chapter.) Your supervisors on the counter-terrorism task force have decided
that .05 is the level of importance for your mission. If the average weight of the 60 passengers on the hijacked bus is
greater than 171 or less than 153, then you will reject the null hypothesis that the bus contains Changing Lives
participants, accept the alternative hypothesis that the bus contains 60 people headed somewhere else, and wait.
additional orders.

You successfully board the moving bus and secretly weigh all the passengers. The mean weight of this
sample of 60 people is 136 pounds, which is more than two standard errors below the mean. (Another
important clue is that all the passengers are children dressed in "Glendale Hockey Camp."
t-shirts.)

According to your mission instructions, you can reject the null hypothesis that this bus contains a random
sample of 60 Changing Lives study participants at a significance level of 0.05. This means (1) the mean weight
on the bus falls within a range that we would expect to observe only 5 times out of 100 if the null hypothesis
were true and this really was a bus full of Changing Lives passengers; (2) we can reject the null hypothesis at
Machine Translated by
Google

the .05 significance level; and (3) on average, 95 times out of 100 you will have correctly rejected the null
hypothesis, and 5 times out of 100 you will have been wrong, meaning you will have concluded that this is not
a bus full of Changing Lives participants, when in fact it is. It turns out that this sample of Changing Lives
people has a particularly high or low average weight relative to the average of the study participants as a
whole.

The mission is not over yet. Your supervisor at mission control (played by Angelina Jolie in the film version
of this example) asks you to calculate a p-value for your result. The p-value is the specific probability of
obtaining a result at least as extreme as the observed one if the null hypothesis is true. The average weight of
the passengers on this bus is 136, which is 5.7 standard errors below the average of the Changing Lives study
participants. The probability of obtaining a result at least as extreme if it were really a sample of Changing
Lives participants is less than 0.0001. (In a research paper, this would be reported as p<.0001.) Once the
mission is complete, you jump off the moving bus and land safely in the passenger seat of a convertible driving
in an adjacent lane.

[This story also has a happy ending. Once the pro-obesity terrorists learn more about their city's
International Sausage Festival, they will agree to abandon violence and work peacefully to promote obesity by
expanding and promoting sausage festivals around the world.]

If the 0.05 significance level seems somewhat arbitrary, that is because it is. There is no single standardized
statistical threshold for rejecting a null hypothesis. Both 0.01 and 0.1 are also reasonably common thresholds
for performing the type of analysis described above.

Obviously, rejecting the null hypothesis at the 0.01 level (meaning that there is less than a 1 in 100 chance
of observing an outcome in this range if the null hypothesis were true) carries more statistical weight than
rejecting the null hypothesis.

hypothesis at level .1 (meaning there is less than a 1 in 10 chance of observing this outcome if the null
hypothesis were true). The pros and cons of different significance levels will be discussed later in this chapter.
For now, the important point is that when we can reject a null hypothesis with some reasonable level of
significance, the results are said to be "statistically significant."
Here's what that means in real life. When you read in the paper that people who eat twenty bran muffins a
day have lower rates of colon cancer than people who don't eat prodigious amounts of bran, the underlying
academic research probably looked something like this: (1) In some big data pool, researchers determined that
people who ate at least twenty bran muffins a day had a lower incidence of colon cancer than people who
didn't eat much bran. (2) The researchers' null hypothesis was that eating bran muffins has no impact on colon
cancer. (3) The disparity in colon cancer outcomes between those who ate a lot of bran muffins and those who
did not could not easily be explained by chance alone. More specifically, if eating bran muffins has no true
association with colon cancer, the probability that such a wide difference in cancer incidence between bran
eaters and non-bran eaters would occur by chance is less than some threshold, such as 0.05. (Researchers
should set this threshold before performing their statistical analysis to avoid choosing a threshold after the fact
Machine Translated by
Google

that is convenient for the results to appear significant.) (4) The academic article probably contains a conclusion
that says something like this: “We found a statistically significant association between daily consumption of
twenty or more bran muffins and a reduced incidence of colon cancer. These results are significant at the .05”
level.
When I later read about that study in the Chicago Sun-Times while eating bacon and eggs for breakfast, the
headline was probably more direct and interesting: “20 Bran Muffins a Day Help Keep Colon Cancer Away.”
However, that newspaper headline, while much more interesting to read than the academic article, may also be
introducing a serious inaccuracy. The study does not actually claim that eating bran muffins reduces a person's
risk of getting colon cancer; it simply shows a negative correlation between bran muffin consumption and colon
cancer incidence in a large data set. This statistical association is not sufficient to demonstrate that bran
muffins improve health outcomes. After all, the kind of people who eat bran muffins (particularly twenty a day!)
can do many other things that reduce their cancer risk, such as eating less red meat, exercising regularly,
getting cancer screenings, etc.
(This is the “healthy user bias” from Chapter 7.) Is it the bran muffins at work here, or are there other behaviors
or personal attributes that people who eat a lot of bran muffins share? This distinction between correlation and
Causality is crucial for the proper interpretation of statistical results. We will revisit the idea that “correlation does not equal
causation” later in the book.
I should also point out that statistical significance says nothing about the size of the association. People who eat a lot
of bran muffins may have a lower incidence of colon cancer, but how much lower? The difference in colon cancer rates
between bran muffin eaters and non-bran muffin eaters may be trivial; the finding of statistical significance means only that
the observed effect, however small, is unlikely to be a coincidence. Suppose you come across a well-designed study that
found a statistically significant positive relationship between eating a banana before the SAT and scoring higher on the
math portion of the test. One of the first questions you'll want to ask is: How big is this effect? It could easily be 0.9 points;
on a test with an average score of 500, that's not a life-changing number. In Chapter 11 we will return to this crucial
distinction between size and significance when it comes to interpreting statistical results.
Meanwhile, the finding that there is "no statistically significant association" between two variables means that any
relationship between the two variables can reasonably be explained by chance alone. The New York Times recently
published an exposé about technology companies selling software they claim improves student performance, when the
data suggests otherwise. 3 According to the article, Carnegie Mellon University sells a software program called Cognitive
Tutor with this bold claim: “Revolutionary math curriculum. “Revolutionary results.”
However, an evaluation of Cognitive Tutor by the U.S. Department of Education found that US concluded that the product
had “no discernible effects” on high school students’ test scores. (The Times suggested that the appropriate marketing
campaign should be “Undistinguished maths curricula. “Unproven results”). In fact, a study of ten software products
designed to teach skills like math or reading found that nine of them “had no statistically significant results.” effects on test
scores.” In other words, federal investigators cannot rule out mere chance as a cause of any variation in the performance
of students who use these software products and those who do not.

Let me pause here to remind you why all of this is important. A Wall Street Journal article from May 2011 was headlined
“Link Between Autism and Brain Size.” This is an important step forward, as the causes of autism spectrum disorder
remain elusive. The first sentence of the Wall Street Journal article, which summarizes an article published in the Archives
Machine Translated by
Google

of General Psychiatry, reports: “Children with autism have larger brains than children without the disorder, and growth
appears to occur earlier than age 2, a study finds. New study published on Monday.”
4
Based on brain imaging of 59 children with autism and 38 children without autism, researchers at the
University of North Carolina reported that children with autism have brains that are up to 10 percent
larger than those of children of the same age without autism.
Here's the relevant medical question: Is there a physiological difference in the brains of young children with
autism spectrum disorder? If so, this information could lead to a better understanding of the causes of the
disorder and how it can be treated or prevented.
And here's the relevant statistical question: Can researchers make broad inferences about autism spectrum
disorder in general based on a study of a seemingly small group of children with autism (59) and an even
smaller control group (38)—just 97? topics in total? The answer is yes. The researchers concluded that the
probability of observing the differences in total brain size that they found in their two samples would be only 2
in 1,000 (p = 0.002) if there were in fact no real difference in brain size between children with and without ASD
in the general population.
I then located the original study in the Archives of General Psychiatry. The methods5 used by these
researchers are no more sophisticated than the concepts we have covered so far. I'll give you a quick rundown
of the rationale behind this socially and statistically significant result. First, it must be acknowledged that each
group of children, the 59 with autism and the 38 without autism, constitutes a reasonably large sample drawn
from their respective populations: all children with and without autism spectrum disorder. The samples are
large enough for the central limit to apply. If you've already tried to forget about the last chapter, I'll remind you
what the central limit theorem tells us: (1) sample means for any population will be distributed approximately as
a normal distribution around the true population mean; (2) we would expect the sample mean and sample
standard deviation to be approximately equal to the mean and standard deviation of the population from which
they are drawn; and (3) about 68 percent of the sample means will lie within one standard error of the
population mean, about 95 percent will lie within two standard errors of the population mean, and so on.
In less technical language, all this means that any sample should closely resemble the population from
which it is drawn; while every sample will be different, it would be relatively rare for the mean of a correctly
drawn sample to deviate greatly from the mean of the relevant underlying population. Similarly, we would also
expect two samples drawn from the same population to look very similar to each other. Or, to think of the
situation somewhat differently, if we have two samples that have extremely different means, the most likely
explanation is that they come from different populations.
Here is a quick and intuitive example. Suppose your null hypothesis is that male professional basketball
players have the same mean height as the rest of the adult male population. You randomly select a sample of
50 professional basketball players and a sample of 50 men who do not play professional basketball. Suppose
the mean height of your basketball sample is 6 feet 7 inches and the mean height of the non-basketball players
is 5 feet 10 inches (a difference of 9 inches). What is the probability of observing such a large difference in
mean height between the two samples if there is actually no difference in average height between professional
basketball players and all other men in the general population? The non-technical answer: very, . ■ * very, very
low.
The autism research paper has the same basic methodology. The article compares several measures of
Machine Translated by
Google

brain size among samples of children. (Brain measurements were made with MRI scans at age two, and again
between ages four and five.) I will focus on a single measurement, total brain volume. The researchers' null
hypothesis was presumably that there are no anatomical differences in the brains of children with and without
autism. The alternative hypothesis is that the brains of children with autism spectrum disorder are
fundamentally different. Such a discovery would still leave many questions, but it would point to a direction for
future research.

In this study, children with autism spectrum disorder had a mean brain volume of 1,310.4 cubic centimeters;
children in the control group had a mean brain volume of 1,238.8 cubic centimeters. Thus, the difference in
mean brain volume between the two groups is 71.6 cubic centimeters. How likely would this result be if there
were actually no differences in average brain size in the general population between children who have an
autism spectrum disorder and those who do not?

You may remember from the last chapter that we can create a standard error for each of our samples:
where ' s , is ' the sample standard deviation and n is the number of observations. The research work gives us
these figures. The standard error for total brain volume for the 59 children in the autism spectrum disorder
sample is 13 cubic centimeters; the standard error for total brain volume for the 38 children in the control group
is 18 cubic centimeters. You will recall that the central limit theorem tells us that for 95 samples out of 100, the
sample mean will be within two standard errors of the true population mean, in one direction or the other.

As a result, we can infer from our sample that 95 times out of 100 the interval of 1310.4 cubic centimeters ±
26 (which is two standard errors) will contain the average brain volume of all children with autism spectrum
disorder. This expression is called a confidence interval. We can say 95 percent.
Machine Translated by
Google

Confidence that the range of 1284.4 to 1336.4 cubic centimeters contains the average total brain volume of
children in the general population with autism spectrum disorder.

Using the same methodology, we can say with 95 percent confidence that the interval 1238.8 ± 36, or
between 1202.8 and 1274.8 cubic centimeters, will include the average brain volume of children in the general
population who do not have autism spectrum disorder.
Yes, there are a lot of numbers here. Maybe you just threw the book across the room. *. ■. r
....
If not, or if you later went and retrieved the book, what you should notice is that our
confidence intervals do not overlap. The lower limit of our 95 percent confidence interval for the average brain
size of children with autism in the general population (1284.4 cubic centimeters) is even larger than the upper
limit of the 95 percent confidence interval for the average brain size of young children in the general population.
population without autism (1274.8 cubic centimeters), as illustrated in the following diagram.

95% confidence intervalfor 95% confidence interval for children with autism spectrum disorder / TO
1202.8 the non-autismgeneral
1238.8 1274.8 1284-4 1310.4 1336.4
population
Yo--------------------1---------------------1
naked statistics-------------------------------------2
Taking the fear out of data---------------------------------------------------------------------------------------------2
CARLOS WHEELAN-------------------------------------------------------------------------------------------------2
Content----------------------------------------------------------------------------------------------------------------------5
Introduction----------------------------------------------------------------------------------------------------------------------7
What's the point?------------------------------------------------------------------------------------------------------------12
Inference-------------------------------------------------------------------------------------------------------------15
The normal distribution-----------------------------------------------------------------------------------------------30
Formula for variance and standard deviation.-------------------------------------------------------------------36
Misleading description “He has a great personality!” and other true but wildly misleading statements- -38
Correlation----------------------------------------------------------------------------------------------------------------55
How does Netflix know what movies I like?-------------------------------------------------------------------------55
Basic probability------------------------------------------------------------------------------------------------------------63
Don't buy the extended warranty on your $99 printer---------------------------------------------------------------63
The investment decision----------------------------------------------------------------------------------------------73
Widespread detection of a rare disease--------------------------------------------------------------------------75
The Monty Hall Problem---------------------------------------------------------------------------------------------------79
Problems with probability-----------------------------------------------------------------------------------------------83
The importance of data "Garbage in, garbage out"------------------------------------------------------------------93
The central limit theorem----------------------------------------------------------------------------------------------------104
The Lebron James of statistics-------------------------------------------------------------------------------------104
100 sample means, n = 100--------------------------------------------------------------------------------------110
Machine Translated by
Google

Inference-----------------------------------------------------------------------------------------------------------------------115
Why my statistics teacher thought I might have cheated----------------------------------------------------115
Distribution of sample means------------------------------------------------------------------------------------120
Difference in sample means--------------------------------------------------------------------------------------133
Difference in sample means (Measured in standard errors)-------------------------------------------------134
Difference in sample means (Measured in standard errors)-------------------------------------------------135
The judgment should inform whether a one- or two-tailed hypothesis is more appropriate for the
analysis being performed.----------------------------------------------------------------------------------------135
Vote-----------------------------------------------------------------------------------------------------------------------136
How do we know that 64 percent of Americans support the death penalty (with a sampling error of ± 3
percent)?------------------------------------------------------------------------------------------------------------136
Regression analysis----------------------------------------------------------------------------------------------------------147
The miraculous elixir------------------------------------------------------------------------------------------------147
Regression equation for weight----------------------------------------------------------------------------------165
Common regression errors--------------------------------------------------------------------------------------------------167
The mandatory warning label------------------------------------------------------------------------------------167
Program Evaluation Will going to Harvard change your life?------------------------------------------------------176
Conclusion--------------------------------------------------------------------------------------------------------------------188
Appendix----------------------------------------------------------------------------------------------------------------------200
statistical software-------------------------------------------------------------------------------------------------200
*----------------------------------------------------------------------------------------------------------------------200
Was------------------------------------------------------------------------------------------------------------------200
IBMSPSS--------------------------------------------------------------------------------------------------------------202
Grades--------------------------------------------------------------------------------------------------------------------------203
Chapter 1: What's the point?----------------------------------------------------------------------------------------203
Chapter 2: Descriptive Statistics--------------------------------------------------------------------------------------204
Chapter 3: Misleading Description--------------------------------------------------------------------------------------205
Chapter 4: Correlation----------------------------------------------------------------------------------------------------206
Chapter 5: Basic Probability-------------------------------------------------------------------------------------------207
Chapter 6: Problems with probability--------------------------------------------------------------------------------208
Chapter 7: The importance of data-----------------------------------------------------------------------------------209
Chapter 8: The Central Limit Theorem---------------------------------------------------------------------------------210
Chapter 9: Inference-------------------------------------------------------------------------------------------------------211
Chapter 10: Survey-----------------------------------------------------------------------------------------------------212
Chapter 11: Regression Analysis-------------------------------------------------------------------------------------213
Machine Translated by
Google

Chapter 12: Common Regression Errors-------------------------------------------------------------------------------214


Chapter 13: Program Evaluation--------------------------------------------------------------------------------------215
Conclusion------------------------------------------------------------------------------------------------------------215
Expressions of gratitude-----------------------------------------------------------------------------------------------216
Index----------------------------------------------------------------------------------------------------------------------219

This is the first clue that there may be an underlying anatomical difference in the brains of young children
with autism spectrum disorder. Still, it's just a hint. All of these inferences are based on data from fewer than
100 children. Maybe we just have extravagant samples.
One last statistical procedure can make all this happen. If statistics were an Olympic event like figure
skating, this would be the final program, after which euphoric fans throw bouquets of flowers onto the ice. We
can calculate the exact probability of observing a mean difference at least this large (1310.4 cubic centimeters
versus 1238.8 cubic centimeters) if there really is no difference in brain size between children on the autism
spectrum and all other children in the general population. We can find a p-value for the observed difference in
means.
So that you don't have to throw the book across the room again, I've included the formula in an appendix.
The intuition is quite simple. If we draw two large samples from the same population, we would expect them to
have very similar means. In fact, our best guess is that they will have identical means. For example, if I were to
select 100 NBA players and they had an average height of 6 feet 7 inches, then I would expect another random
sample of 100 NBA players.

The NBA will have an average height close to 6 feet 7 inches. Well, maybe the two samples are an inch or two apart. But
it is less likely that the means of the two samples will be 4 inches apart, and even less likely that there will be a difference
of 6 or 8 inches. It turns out that we can calculate a standard error for the difference between two sample means; this
standard error gives us a measure of the spread we can expect, on average, when we subtract one sample mean from the
other. (As noted above, the formula is in the chapter appendix.) The important thing is that we can use this standard error
to calculate the probability that two samples come from the same population. Here's how it works:

1. If two samples are drawn from the same population, our best estimate of the difference between their means is
zero.
2. The central limit theorem tells us that in repeated samples, the difference between the two means will be
distributed approximately as a normal distribution.
(So, have you fallen in love with the central limit theorem yet or not?)
3. If the two samples really come from the same population, then in about 68 cases out of 100, the difference
between the two sample means will be within one standard error of zero. And in about 95 cases out of 100, the
difference between the two sample means will be within two standard errors of zero. And in 99.7 cases out of 100,
the difference will be within three standard errors of zero, which happens to be what motivates the conclusion of the
autism research paper we started with.
Machine Translated by
Google

As noted above, the difference in mean brain size between the sample of children with autism spectrum disorder and
the control group is 71.6 cubic centimeters. The standard error of that difference is 22.7, meaning that the difference in
means between the two samples is more than three standard errors from zero; we would expect such an extreme result
(or more) only 2 times out of 1,000 if these samples were drawn from an identical population.

In the article published in Archives of General Psychiatry, the authors report a p-value of 0.002, as I mentioned above.
Now you know where it came from!

For all the wonders of statistical inference, there are some major drawbacks. They are derived from the example that
introduced the chapter: my suspicious statistics teacher. The powerful process of statistical inference is based on
probability, not some kind of cosmic certainty. We don't want to send people to jail just for doing the equivalent of drawing
two royal flushes in a row; it can happen, even if someone isn't cheating. As a result, we are faced with a fundamental
dilemma when it comes to any kind of hypothesis testing.

This statistical reality came to a head in 2011, when the Journal of Personality and Social Psychology
prepared to publish a scholarly article that, on the surface, looked like thousands of other scholarly articles.
6
A Cornell professor
explicitly proposed a null hypothesis, performed an experiment to test his null hypothesis, and then rejected the
null hypothesis with a significance of 0.05 on the basis of the experimental results. The result was a huge stir,
both in scientific circles and in mainstream media such as the New York Times.
Suffice it to say that articles in the Journal of Personality and Social Psychology do not typically attract big
headlines. What exactly made this study so controversial? The researcher in question was testing humans'
ability to exercise extrasensory perception, or ESP. The null hypothesis was that ESP does not exist; the
alternative hypothesis was that humans have extrasensory powers. To study this question, the researcher
recruited a large sample of participants to examine two “curtains” placed on a computer screen. A software
program randomly places an erotic photo behind one curtain or another. In repeated trials, study participants
were able to choose the curtain with the erotic photo behind it 53 percent of the time, whereas chance says
they would get it right only 50 percent of the time. Because of the large sample size, the researcher was able to
reject the null hypothesis that ESP does not exist and instead accept the alternative hypothesis that ESP may
allow individuals to sense future events. The decision to publish the paper was widely criticised on the grounds
that a single statistically significant event could easily be a product of chance, especially when there is no other
evidence to corroborate or even explain the finding. The New York Times summed up the criticism: “Claims
that defy almost all the laws of science are by definition extraordinary and therefore require extraordinary
evidence.

Failure to take this into account (as conventional social science analyses do) makes many findings appear far
more significant than they really are.”
One answer to this kind of nonsense would seem to be a more rigorous threshold for defining *
Machine Translated by
Google

statistical significance, such as 0.001. But that creates its


own problems. Choosing an appropriate significance level involves an inherent trade-off.

If our burden of proof for rejecting the null hypothesis is too low (e.g., .1), we will periodically find ourselves
rejecting the null hypothesis when it is actually true (as I suspect was the case with the ESP study). In
statistical language, this is known as a type I error. Consider the example of an American court, where the null
hypothesis is that a defendant is not guilty and the threshold for rejecting that null hypothesis is “guilty beyond
a reasonable doubt.” Suppose we

If we relax that threshold to something like "a strong feeling that the guy did it." This will ensure that more
criminals go to jail, and more innocent people too. In a statistical context, this is equivalent to having a
relatively low level of significance, such as 0.1.

Well, 1 in 10 isn't exactly wildly unlikely. Consider this challenge in the context of the approval of a new
cancer drug. For every ten drugs we approve with this relatively low burden of statistical evidence, one of them
doesn't actually work and showed promising results in the trial simply by chance. (Or, in the court example, for
every ten defendants we find guilty, one of them was actually innocent.) A type I error involves wrongly
rejecting a null hypothesis. Although the terminology is somewhat contradictory, this is also known as a "false
positive." Here's a way to reconcile the jargon. When you go to the doctor to get tested for a disease, the null
hypothesis is that you don't have that disease. If the laboratory results can be used to reject the null
hypothesis, then the test is said to be positive. And if the result is positive but you are not actually sick, then it
is a false positive.
In any case, the lower our statistical load to reject the null hypothesis, the more likely this is to happen.
Obviously we would prefer not to approve ineffective cancer drugs or send innocent defendants to prison.
But there is a tension here. The higher the threshold for rejecting the null hypothesis, the more likely it is
that we will fail to reject a null hypothesis that should be rejected. If we need five eyewitnesses to convict every
criminal defendant, then many guilty defendants will be unjustly released. (Of course, fewer innocent people
will go to prison.) If we adopt a significance level of 0.001 in clinical trials of all new cancer drugs, then we will
minimize the approval of ineffective drugs. (There is only a 1 in 1,000 chance of wrongly rejecting the null
hypothesis that the drug is no more effective than a placebo.) However, we now introduce the risk of not
approving many effective drugs because we have set the bar for approval so high. This is known as a type II
error or false negative.

What kind of mistake is worse? That depends on the circumstances. The most important point is that you
recognize the compensation. There is no statistical “free lunch.” Let us consider these non-statistical situations,
all of which involve a trade-off between type I and type II errors.

1. Spam filters. The null hypothesis is that any particular email message is not spam. Your spam filter
looks for clues that can be used to reject that null hypothesis for any particular email, such as huge
distribution lists or phrases like "penis enlargement." A Type I error would involve a detection sending an
Machine Translated by
Google

email message that is not actually spam (a false positive). A Type II error would involve letting spam pass through
your inbox filter (a false negative). Given the costs of missing an important email compared to the costs of receiving
an occasional message about herbal vitamins, most people would probably err on the side of allowing Type II
errors. An optimally designed spam filter should require a relatively high degree of certainty before rejecting the null
hypothesis that an incoming email is legitimate and blocking it.

2. Cancer detection. We have numerous tests available for early detection of cancer, such as mammograms
(breast cancer), PSA testing (prostate cancer), and even full-body MRIs for anything else that might seem
suspicious. The null hypothesis for anyone undergoing this type of screening is that there is no cancer. Screening is
used to reject this null hypothesis if the results are suspicious. It has always been assumed that a type I error (a
false positive that turns out to be nothing) is far preferable to a type II error (a false negative that misses a cancer
diagnosis). Historically, cancer detection has been the opposite of the spam example. Physicians and patients are
willing to tolerate a fair number of type I errors (false positives) to avoid the possibility of a type II error (missing a
cancer diagnosis). More recently, health policy experts have begun to question this view because of the high costs
and serious side effects associated with false positives.

3. Capture terrorists. Neither a Type I error nor a Type II error is acceptable in this situation, which is why society
continues to debate the appropriate balance between fighting terrorism and protecting civil liberties. The null
hypothesis is that an individual is not a terrorist. As in a normal criminal context, we do not want to make a Type I
error and send innocent people to Guantanamo Bay. However, in a world with weapons of mass destruction, letting
even a single terrorist go free (a Type II error) can be literally catastrophic. This is why (whether they approve or
not) the United States holds suspected terrorists in Guantanamo Bay based on less evidence than would be
necessary to convict them in an ordinary criminal court.

Statistical inference is neither magical nor infallible, but it is an extraordinary tool for making sense of the world. We
can gain great insight into many phenomena in life simply by determining the most likely explanation. Most of us do this all
the time (e.g., “I think that college student passed out on the floor surrounded by beer cans has had too much to drink”
instead of “I think that college student passed out on the floor surrounded by beer cans has been poisoned by
terrorists”).

Statistical inference simply formalizes the process.

APPENDIX TO CHAPTER 9

Calculate the standard error for a difference in means


Machine Translated by
Google

Formula for comparing two averages

x - y —> numerator yields the size of the difference in means f si + si —> denominator yields the standard error for a difference V —
nnr in mean between two samples

where = sample mean x =


) sample mean and
sx = sample standard deviation x
sy = sample standard deviation and
nx = number of observations in the sample x
ny = number of observations in the sample and

Our null hypothesis is that the two sample means are equal. The above formula calculates the observed
difference in means relative to the size of the standard error for the difference in means. Again we rely heavily
on the normal distribution. If the underlying population means are really the same, then we would expect the
difference in the sample means to be less than one standard error about 68 percent of the time; less than two
standard errors about 95 percent of the time; and so on.
In the autism example from the chapter, the difference in the mean between the two samples was 71.6
cubic centimeters with a standard error of 22.7. The ratio of that observed difference is 3.15, meaning that the
two samples have means that are separated by more than 3 standard errors. As noted in the chapter, the
probability of obtaining samples with such different means if the underlying populations have the same mean is
very, very low. Specifically, the probability of observing a difference in means of 3.15 standard errors or greater
is 0.002.

Difference in sample means

One- and two-tailed hypothesis testing In this


chapter presented the idea of using samples to test whether male professional basketball players are the same
height as the general population. I perfected a detail. Our null hypothesis is that male basketball players have
the same average height as men in the general population. What I overlooked is that we have two possible
alternative hypotheses.
Machine Translated by
Google

An alternative hypothesis is that male professional basketball players have a different average height than
the general male population; they could be taller than other men in the population or shorter. This was the
approach he took when he boarded the hijacked bus and weighed the passengers to determine if they would
participate in the Changing Lives study. The null hypothesis that bus participants were part of the study could
be rejected if the average weight of the passengers was significantly higher than the overall average of
Changing Lives participants or if it was significantly lower (as turned out to be the case). Our second alternative
hypothesis is that male professional basketball players are, on average, taller than other men in the population.
In this case, the prior knowledge we bring to this question tells us that basketball players cannot be shorter
than the general population. The distinction between these two alternative hypotheses will determine whether
we do a one-tailed hypothesis test or a two-tailed hypothesis test.

In both cases, let's assume that we are going to do a significance test at the 0.05 level. We will reject our
null hypothesis if we observe a difference in heights between the two samples that would occur 5 times out of
100 or less if all of these guys really were the same height. So far, so good.
This is where things get a little more nuanced. When our alternative hypothesis is that basketball players
are taller than other men, we will do a one-tailed hypothesis test. We will measure the difference in average
height between our sample of male basketball players and our sample of normal men. We know that if our null
hypothesis is true, then we will observe a difference of 1.64 standard errors or more only 5 times out of 100.
We reject our null hypothesis if our result falls within this range, as shown in the following diagram.

Difference in sample means (Measured in standard errors)

Now let's review the other alternative hypothesis: that male basketball players might be taller or shorter than
the general population. Our general approach is the same. Again, we will reject our null hypothesis that
basketball players are the same height as the general population if we get a result that would occur 5 times out
of 100 or less if there really was no difference in heights. The difference, however, is that we now have to
consider the possibility that basketball players are shorter than the general population. Therefore, we will reject
our null hypothesis if our sample of male basketball players has a mean height that is significantly greater or
less than the mean height of our sample of normal men. This requires a two-tailed hypothesis test. The cutoff
Machine Translated by
Google

points for rejecting our null hypothesis will be different because we now have to account for the possibility of a
large difference in sample means in either direction: positive or negative. More specifically, the range in which
we will reject our null hypothesis has been split between the two tails. We will still reject our null hypothesis if
we get a result that would occur 5 percent of the time or less if basketball players were the same height as the
general population; only now we have two different ways to end up rejecting the null hypothesis.
We will reject our null hypothesis if the mean height of the sample of male basketball players is so much
greater than the mean of normal men that we would observe such a result only 2.5 times out of 100 if the
basketball players really were all the same height. others.
And we will reject our null hypothesis if the mean height of the sample of male basketball players is so
much shorter than the mean of normal men that we would observe such a result only 2.5 times out of 100 if the
basketball players were really the same height as everyone else.
Together, these two contingencies add up to 5 percent, as the following diagram illustrates.
Difference in sample means
(Measured in standard errors)

XY
Difference in sample means

The judgment should inform whether a one- or two-tailed hypothesis is more


appropriate for the analysis being performed.

* As a matter of semantics, we have not shown that the null hypothesis is true (that substance abuse treatment has no effect). It may be extremely effective for
another group of prisoners. Or perhaps many more prisoners in this treatment group would have been arrested again if they had not received the treatment. In any
case, based on the data collected, we simply have not been able to reject our null hypothesis. There is a similar distinction between “failing to reject” a null
hypothesis and accepting it. Just because one study failed to disprove that substance abuse treatment has no effect (yes, a double negative) does not mean one
should accept that substance abuse treatment is useless. There is a significant statistical distinction here. That said, research is often designed to inform policy,
and correctional officials, who have to decide where to allocate resources, might reasonably accept the position that substance abuse treatment is ineffective until
they are convinced otherwise. Here, as in so many other areas of statistics, judgment matters.

* This example is inspired by real events. Obviously many details have been changed for reasons of national security. I can neither confirm nor deny my own
involvement.
* To be precise, 95 percent of all sample means will lie within 1.96 standard errors above or below the population mean.
* There are two possible alternative hypotheses. One is that male professional basketball players are taller than the general male population. The other is simply
that male professional basketball players have a different average height than the general male population (leaving open the possibility that male basketball
Machine Translated by
Google

players are actually shorter than other men). This distinction has little impact when performing significance tests and calculating p-values. It is explained in more
advanced texts and is not important for our general discussion here.

* I admit that I once tore a statistics book in half out of frustration.


* Another answer is to try to replicate the results in additional studies.

CHAPTER 10

Vote
How do we know that 64 percent of Americans
support the death penalty (with a sampling error
of ± 3 percent)?

In late 2011, the New York Times published a front-page article reporting that
1
“a deep sense of anxiety and doubt about the future hangs over the nation.” The story delved deep into the American
psyche and offered insights into public opinion on issues ranging from the performance of the Obama administration to the
distribution of wealth. Here's a snapshot of what Americans said in the fall of 2011:

• A shocking 89 percent of Americans said they distrust the government to do the right thing, the highest level of
distrust ever
registered. • Two-thirds of the public said wealth should be distributed more evenly in the country. •
Forty-three percent of Americans said they generally agreed with the views of the Occupy Wall Street movement,
an amorphous protest movement that began near Wall Street in New York and was spreading to other cities in the *
country. A slightly higher percentage, 46 percent,
He said the views of people involved in the Occupy Wall Street movement “generally reflect the views of the
majority of Americans.”
• Forty-six percent of Americans approved of Barack Obama's handling of his job as president, and an identical
46 percent disapproved of his job performance. • Only 9 percent
percent of the public approved of the way Congress was handling its work. • Although the primaries
With the presidential election set to begin in just two months, roughly 80 percent of Republican primary voters said
it was “still too early to say who they will support.”

These are fascinating numbers that provided significant insight into American opinions a year before the presidential
race. Still, one might reasonably ask: How do we know all this? How can we draw such radical conclusions?

About the attitudes of hundreds of millions of adults? And how do we know if these broad conclusions are accurate?
The answer, of course, is that we conduct surveys. Or in the example above, the New York Times and CBS News
might run a poll. (The fact that two competing news organizations are collaborating on a project like this is the first clue
that conducting a methodologically sound national survey is not cheap.) I have no doubt that you are familiar with the
Machine Translated by
Google

results of the surveys. It may be less obvious that survey methodology is just another form of statistical inference. A
survey (or poll) is an inference about the opinions of a population based on the opinions expressed by some sample
drawn from that population.
The power of probing arises from the same source as our previous sampling examples: the central limit theorem. If we
take a large, representative sample of American voters (or any other group), we can reasonably assume that our sample
will closely resemble the population from which it is drawn. If exactly half of American adults disapprove of gay marriage,
then our best guess about the attitudes of a representative sample of 1,000 Americans is that about half of them will
disapprove of gay marriage.
By contrast, and more importantly from a polling standpoint, if we have a representative sample of 1,000 Americans
who feel a certain way—such as the 46 percent who disapprove of President Obama's job performance—then we can
infer from that sample that the general population is likely to feel the same way. In fact, we can calculate the probability
that our sample results deviate greatly from the true attitudes of the population. When we read that a survey has a “margin
of error” of ±3 percent, this is actually the same kind of 95 percent confidence interval that we calculated in the previous
chapter. Our “95 percent confidence” means that if we conducted 100 different polls on samples drawn from the same
population, we would expect the responses we got from our sample in 95 of those polls to be within 3 percentage points in
one direction or the other of the true sentiment of the population. In the context of the job approval question in the New
York Times/CBS poll, we can be 95 percent confident that the true proportion of all Americans who disapprove of
President Obama's job rating is in the range of 46 percent ± 3 percent, or between 43 percent and 49 percent. If you read
the fine print of the New York Times/CBS poll (as I urge you to do), here’s roughly what it says: “In theory, in 19 out of 20
cases, the overall results based on such samples will differ by no more than 3 percentage points in either direction from
what would have been obtained by attempting to interview all American adults.”

A fundamental difference between a survey and other forms of sampling is that the sample statistic we are interested in
will not be a mean (e.g., a mean). e.g., 187 pounds) but rather a percentage or proportion (e.g., 187 pounds) e.g., 47
percent of voters, or 0.47). In other respects, the process is identical. When we have a large, representative sample (the
survey), we would expect the proportion of respondents in the sample who feel a certain way (for example, the 9 percent
who think Congress is doing a good job) to be roughly equal to the proportion of all Americans who feel that way. This is
no different from assuming that the mean weight of a sample of 1,000 American men should be approximately equal to the
mean weight of all American men. Still, we expect some variation in the percentage approving of Congress from sample to
sample, just as we would expect some variation in the mean weight when taking different random samples of 1,000 men.
If the New York Times and CBS had conducted a second poll (asking the same questions to a new sample of 1,000
American adults) it is highly unlikely that the results of the second poll would have been identical to the results of the first.
On the other hand, we should not expect the responses of our second sample to differ much from the responses given by
the first. (To return to the metaphor used above, if you taste a spoonful of soup, stir the pot, and then taste again, both
spoonfuls will taste similar.)
The standard error is what tells us how much dispersion we can expect in our results from sample to sample, which in this
case means survey to survey.
The formula for calculating a standard error for a percentage or proportion is slightly different from the formula
presented above; the intuition is exactly the same. For any properly drawn random sample, the standard error is equal to
where p is the P(1-P)n, proportion of respondents expressing a particular opinion, (1 – p) is the proportion of respondents
expressing a different opinion, and n is the total number of respondents in the sample. sample. You should see that the
standard error will decrease as the sample size increases, since n is in the denominator. The standard error also tends to
be smaller when p and (1 – p) are far apart. For example, the standard error will be smaller in a survey in which 95
Machine Translated by
Google

percent of respondents express a certain opinion than in a survey in which opinions tend to split 50-50. This is just math,
since (0.05)(0.95) = 0.047, while (0.5)(0.5) = 0.25; a smaller number in the numerator of the formula leads to a smaller
standard error.
As an example, suppose a simple “exit poll” of 500 representative voters on Election Day reveals that 53 percent voted
for the Republican candidate; 45 percent of voters voted for the Democrat; and 2 percent supported a third-party
candidate. If we use the Republican candidate as our proportion of interest, the standard survey would be this exit would
do
/(53)(1-.53)500 = /(.53)(47)/500 = ¿25/500 = ¿0005 = .02236.
For simplicity, we will round the standard error of this exit poll to 0.02. So far, that's just a number. Let's analyze why
that number is important. take on the
The polls have just closed and you work for a television network that wants to declare a winner in the race before the full
results are available. You are now the network's official data analyst (having read two-thirds of this book) and your
producer wants to know if it is possible to "call the race" on the basis of this exit poll.
You explain that the answer depends on the confidence that people on the network would like to have in the
advertisement or, more specifically, the risk they are willing to take of being wrong. Remember, the standard error gives
us an idea of how often we can expect our sample proportion (the exit poll) to be reasonably close to the true population
proportion (the election result). We know that about 68 percent of the time we can expect the sample proportion (the 53
percent of voters who said they voted Republican in this case) to be within one standard error of the true final count. As a
result, you tell your producer “with 68 percent confidence” that your sample, which shows the Republican got 53 percent of
the vote ± 2 percent, or between 51 and 55 percent, has captured the true count for the Republican candidate. Meanwhile,
the same exit poll shows that the Democratic candidate has received 45 percent of the vote. If we assume that the vote
count for the Democratic candidate has the same standard error (a simplification I'll explain in a minute), we can say with
68 percent confidence that the exit poll sample, which shows the Democrat with 45 percent of the vote ± 2 percent, or
between 43 and 47 percent, contains the Democrat's true count. According to this calculation, the Republican is the
winner.

The graphics department rushes to create a sleek three-dimensional image that it can display on the screen to its
viewers: Republican

53% Democrat 45% Independent 2%


(Margin of error 2%)

At first, his producer is impressed and excited, largely because the graphic above is three-dimensional, multi-colored,
and can rotate on the screen.
Yet when he explains that roughly 68 times out of 100 his exit poll results will be within one standard error of the true
election result, his producer, who has twice been sent to anger management programs by the courts, points out the
obvious math: 32 times out of 100 his exit poll will not be within one standard error of the true election result. And?
You explain that there are two possibilities: (1) the Republican candidate could have received even more votes than
your poll predicted, in which case you
Machine Translated by
Google

He will still have called the elections correctly. Or (2) there is a reasonably high probability that the Democratic
candidate received many more votes than your poll reported, in which case your fancy 3-D multi-colored
spinning chart will have reported the wrong winner.
His producer throws a cup of coffee across the room and uses several phrases that violate his probation.
She screams: "How can we be [redacted] sure that we have the [redacted] correct result?"
Ever the statistics guru, he points out that he cannot be sure of any result until all the votes are counted.
However, it can offer a 95 percent confidence interval. In this case, your rotating, three-dimensional, multi-color
chart will be wrong, on average, only 5 times out of 100.
His producer lights a cigarette and seems to relax. He decides not to mention the ban on smoking in the
workplace, as it turned out to be disastrous last time.
However, he does share some bad news. The only way the broadcaster can have more confidence in its poll
results is by widening the “margin of error.” And when that happens, there is no longer a clear winner in the
election. You show your boss the fancy new chart:
Republican 53% Democrat 45%
Independent 2%
(Margin of error 4%)
We know from the central limit theorem that about 95 percent of the sample proportions will lie within two
standard errors of the true population proportion (which is 4 percent in this case). So if we want to have more
confidence in our survey results, we need to be less ambitious in what we predict. As the graph above
illustrates (without the 3-D and color), with a 95 percent confidence level, the television station can announce
that the Republican candidate has won 53 percent of the vote ± 4 percent, or between 49 and 57 percent of the
vote. votes cast. Meanwhile, the Democratic candidate received 45 percent ± 4 percent, or between 41 and 49
percent of the votes cast.
And yes, now you have a new problem. With a confidence level of 95 percent, the possibility that the two
candidates are tied with 49 percent of the vote each cannot be rejected. This is an unavoidable trade-off; the
only way to be more confident that your poll results will be consistent with the election outcome without new
data is to become more timid in your predictions. Think in a non-statistical context. Suppose you tell a friend
that you are “pretty sure” that Thomas Jefferson was the third or fourth president. How can you be more
confident in your historical knowledge? Being less specific. Are
It is “absolutely positive” that Thomas Jefferson was one of the first five presidents.

Your producer tells you to order a pizza and get ready to stay at work all night. At that moment, statistical good fortune shines
upon you. The results of a second exit poll reach his desk with a sample of 2,000 voters. These results show the following:
Republican (52 percent); Democrat (45 percent); Independent (3 percent). Your producer is now completely exasperated as this
poll suggests the gap between the candidates has narrowed, making it even more difficult for you to call the race on time. But
wait! You point out (heroically) that the sample size (2000) is four times larger than the sample of the first survey. As a result, the
standard error will be significantly reduced. The new standard error for the Republican candidate is 0.01.
/.52(48)/2,000,
If your producer is still comfortable with a 95 percent confidence level, he or she can declare the Republican candidate the
Machine Translated by
Google

winner. With their new standard error of .01, the 95 percent confidence intervals for the candidates are as follows: Republican:
52 ± 2, or between 50 and 54 percent of the votes cast; Democrat: 45 ± 2, or between 43 and 47 percent of the votes cast.
There is no longer overlap between the two confidence intervals. You can predict on the air that the Republican candidate will be
the winner; more than 95 times out of 100 you will be

. * correct.
But this case is even better than that. The central limit theorem tells us that 99.7 percent of the time a sample proportion will
lie within three standard errors of the true population proportion. In this election example, our 99.7 percent confidence intervals
for the two candidates are: Republican, 52 ± 3 percent, or between 49 and 55 percent; Democrat, 45 ± 3 percent, or between 42
and 48 percent. If you report that the Republican candidate has won, there is a small chance that you and your producer will be
fired, thanks to your new sample of 2,000 voters.

You should see that a larger sample size results in a smaller and smaller standard error, which is how large national surveys
can end up with surprisingly accurate results. On the other hand, smaller samples obviously generate larger standard errors and
therefore a larger confidence interval (or “margin of sampling error,” to use survey jargon). The fine print of the New York
Times/CBS poll notes that the margin of error for questions about the Republican primary is 5 percentage points, compared with
3 percentage points for other questions in the poll. These questions were only asked of self-identified Republican primary and
caucus voters, so the sample size for this subset of questions was reduced to 455 (compared with 1,650 adults for the rest of the
survey).

As always, I've simplified a lot of things in this chapter. You may have

I recognized that in my previous election example, the Republican and Democratic candidates should each
have their own standard error. Think about the formula again: the sample size, n, is the same for both
candidates, but p and (1 – p) will be slightly different. In the second exit poll (with the sample of 2,000 voters),
the standard error for the Republican is that for the Democrat. Of course, for (-52(.48//2,000 = .01117; all
intents and purposes, those two numbers are the same. For that reason, I have adopted a common
convention, which is to take the higher of the two standard errors and use it for all candidates. If anything, this
introduces a bit more caution into our confidence intervals.

Many national surveys that ask multiple questions will go a step further. In the case of the New York
Times/CBS poll, the standard error should technically be different for each question, depending on the answer.
For example, the standard error for finding that 9 percent of the public approves of the way Congress is
handling its job should be smaller than the standard error for the question finding that 46 percent of the public
approves of the way President Obama has handled his job. work, since .09 × (.91) is less than .46 × (.54) —
0819 versus .2484. (The intuition behind this formula is explained in a chapter appendix.)

Since it would be confusing and inconvenient to have a different standard error for each question, surveys
of this nature will typically assume that the sampling proportion for each question is 0.5 (or 50 percent),
generating the largest possible standard error for any given question. sample size and then adopt that standard
error for . . . . . . ... . * calculate the margin of sampling error for the entire survey.
Machine Translated by
Google

When done correctly, surveys are amazing tools. According to Frank Newport, editor in chief of the Gallup
Organization, a survey of 1,000 people can provide meaningful and accurate information about attitudes across
the country.
Statistically speaking, he is right. But to get those meaningful and accurate results, we have to conduct a
proper survey and then interpret the results correctly, which is much easier said than done. Poor poll results
are usually not due to bad math in calculating standard errors. Poor survey results are often due to a biased
sample, bad questions, or both. The mantra “garbage in, garbage out” applies in spades when it comes to
sampling public opinion. Below are key methodological questions one should ask when conducting a survey or
reviewing the work of others.

Is this an accurate sample of the population whose opinions we are trying to measure? Chapter 7 discussed
many common data-related challenges.
However, I will point out once again the danger of selection bias, particularly self-selection. Any survey that
relies on individuals selected for sampling, such as a call-in radio program or a voluntary Internet survey, will
capture only the opinions of those who make the effort to express their opinions. These are likely to be people
who feel particularly strongly about an issue, or those who have a lot of time on their hands. None of these
groups are likely to be representative of the general public. I once appeared as a guest on a call-in radio show.
One caller to the show emphatically stated on air that my views were “so wrong” that he pulled his car off the
road and found a pay phone to call the show and register his dissent. I'd like to think that listeners who didn't
pull their cars off the road to watch the show felt differently.
Any method of collecting opinions that systematically excludes some segment of the population is also
prone to bias. For example, mobile phones have introduced a number of new methodological complexities.
Professional polling organizations make every effort to survey a representative sample of the relevant
population. The New York Times/CBS poll was based on telephone interviews conducted over six days with
1,650 adults, 1,475 of whom said they were registered to vote.
I can only guess at the rest of the methodology, but most professional surveys use some variation of the
following techniques. To ensure that the adults who answer the phone are representative of the population, the
process begins with probability, a variation on drawing marbles from an urn. A computer randomly selects a set
of fixed telephone exchanges. (A central office is an area code plus the first three digits of a telephone
number.) By choosing randomly from the country's 69,000 residential exchanges, each in proportion to its
share of all telephone numbers, the survey is likely to yield a generally representative result. Geographical
distribution of the population. As the fine print explains, "The exchanges were chosen so as to ensure that each
region of the country was represented in proportion to its share of all telephone numbers." For each selected
exchange, the computer added four random digits. As a result, both listed and unlisted numbers will end up on
the final list of households to be called. The survey also included “random dialing of cell phone numbers.”
For each number dialed, an adult is designated to respond using a “random procedure,” such as asking for
the youngest adult currently at home. This process has been refined to produce a sample of respondents that
resembles the adult population in terms of age and gender. Most importantly, the interviewer will attempt to
make several calls at different times of the day and night to reach each selected phone number. These
repeated attempts (up to ten or twelve calls to the same number) are an important part of obtaining an
Machine Translated by
Google

unbiased sample. Obviously, it would be cheaper and easier to make random calls to different numbers until a
sufficiently large sample of adults had picked up the phone and answered the relevant questions. However,
such a sample would be biased toward people who are most likely to be at home and answer the phone: the
unemployed, the elderly, and so on. That's fine as long as you're willing to qualify your poll results like this:
President Obama's approval rating is 46 percent among the unemployed, the elderly, and others who are just
eager to answer random phone calls.
One indicator of the validity of a survey is the response rate: What proportion of respondents who were
chosen to be contacted ultimately completed the survey? A low response rate may be a warning sign of
possible sampling bias. The more people who choose not to respond to the survey, or who simply cannot be
contacted, the greater the chance that this large group is different in some material way from those who did
respond to the questions. Pollsters can check for “nonresponse bias” by analyzing available data on
respondents whom they were unable to contact. Do you live in a particular area? Do they refuse to answer for
any particular reason? Are they more likely to be from a particular racial, ethnic, or income group? This type of
analysis can determine whether or not a low response rate will affect the survey results.

Have the questions been asked in a way that provides accurate information on the topic of interest? Soliciting
public opinion requires more nuance than measuring test scores or putting respondents on a scale to
determine their weight. Survey results can be extremely sensitive to how a question is asked. Let's take a
seemingly simple example: What proportion of Americans support capital punishment? As the chapter title
suggests, a solid and consistent majority of Americans approve of the death penalty. According to Gallup,
every year since 2002, more than 60 percent of Americans have said they favor the death penalty for a person
convicted of murder. The percentage of Americans who support capital punishment has fluctuated in a
relatively narrow range from a high of 70 percent in 2003 to a low of 64 percent at several different times. The
polling data is clear: Americans support the death penalty by a wide margin.
Or not. American support for the death penalty plummets when life imprisonment without parole is offered
as an alternative. A 2006 Gallup poll found that only 47 percent of Americans considered the death penalty to
second
be the appropriate punishment for murder, compared with 48 percent who preferred the death penalty to a
life sentence. That's not just a statistic to amuse the prison crowd. party; it means that there is no longer
majority support for capital punishment when life imprisonment without parole is a credible alternative. When
we request
Public opinion, the formulation of the question and the choice of language can be of enormous importance.
Politicians often exploit this phenomenon by using surveys and focus groups to test “words that work.” For example,
voters are more likely to support “tax relief” than “tax cuts,” even though the two phrases describe the same thing.
Similarly, voters are less concerned about “climate change” than “global warming,” even though global warming is a form
of climate change. Obviously politicians are trying to manipulate voters' responses by choosing non-neutral words. If
pollsters are to be considered honest intermediaries who produce legitimate results, they must avoid language that could
affect the accuracy of the information collected. Similarly, if responses are to be compared over time (for example, how
consumers feel about the economy today compared to how they felt a year ago), then the questions that elicit that
information over time should be the same or very similar.
Polling organizations like Gallup often conduct “split-sample tests,” in which variations of a question are tested on
Machine Translated by
Google

different samples to assess how small changes in wording affect respondents’ responses. For experts like Frank Newport
of Gallup, the answers to each question present significant data, but may seem inconsistent. Attitudes toward capital
punishment—even when those responses change dramatically when life without parole is offered as an option—tell us
something important. The key point, Newport says, is to view any survey results in context. No single question or survey
can capture the full depth of public opinion on a complex issue.

Are respondents telling the truth? Surveys are like online dating: there is a little leeway in the veracity of the information
provided. We know that people hide the truth, especially when the questions they ask are embarrassing or sensitive.
Respondents may overstate their income or inflate the number of times they have sex in a typical month. They may not
admit that they don't vote. They may hesitate to express opinions that are unpopular or socially unacceptable. For all
these reasons, even the most carefully designed survey depends on the integrity of respondents' responses.
Election polls depend crucially on separating those who will vote on election day from those who will not. (If we're
trying to gauge the likely winner of an election, we don't care about the opinions of anyone who isn't going to vote.)
People often say they are going to vote because they think that is what pollsters want to hear. Studies that have compared
self-reported voting behavior with voter records consistently find that between a quarter and a third of respondents say
they voted when in fact they did notUlonahifcoiremroand. and minimize this potential
The bias is asking whether a respondent voted in the last election or in the last election. Respondents who
have voted consistently in the past are more likely to vote in the future. Similarly, if there is a concern that
respondents may be hesitant to express a socially unacceptable response, such as a negative view of a racial
or ethnic group, the question can be phrased in a more subtle way, such as asking “whether people you know
hold such a view.”
One of the most sensitive surveys of all time was a study conducted by the National Opinion Research
Center (NORC) at the University of Chicago called "The Social Organization of Sexuality: Sexual Practices in
the United States." It quickly became known as the “Sex Study.” The 5 The formal description of the study
included phrases such as "the organization of behaviors that constitute sexual transactions" and "sexual
partnerships and behavior across the life span." (I'm not even sure what a “life course” is; spell check says it's
not a word.) I am oversimplifying when I write that the survey sought to document who does what and to whom
and how often. The purpose of the study, which was published in 1995, was not simply to enlighten us all about
the sexual behavior of our neighbors (although that was part of it), but also to assess how sexual behavior in
the United States would likely affect the spread of HIV/AIDS.
If Americans are hesitant to admit that they don't vote, you can imagine how keen they are to describe their
sexual behavior, especially when it might involve illicit activities, infidelity, or just plain weird stuff. The
methodology of the Sex Study was impressive. The research was based on ninety-minute interviews with
3,342 adults chosen as representative of the adult population of the United States. Nearly 80 percent of
selected respondents completed the survey, leading the authors to conclude that the findings are an accurate
report of American sexual behavior (or at least what we were doing in 1995).
Since you've endured a chapter on survey methodology and now almost an entire book on statistics, you're
entitled to take a look at what they found (none of which is particularly shocking). As one critic noted: "There's
a lot less sexual behavior than we might think." 6
Machine Translated by
Google

• People generally have sex with other people like them. Ninety percent of the couples shared the same
race, religion, social class, and general age group. • The typical respondent engaged in sexual activities
“a few times a month,” although there was wide variation. The number of sexual partners since age
eighteen ranged from zero to over 1,000. • Approximately 5 percent of men and 4 percent of women
reported some activity
sexual with a same-sex partner. • Eighty percent of respondents had a sexual partner in the previous
year or none at all. • Respondents with a sexual partner were happier than those who had none or had7
multiple couples. • A quarter of married men and 10 percent of married women reported having engaged
in sexual activity
extramarital sexual activity. • Most people do it the old-fashioned way: Vaginal intercourse was the most
attractive sexual activity for both men and women.
A review of the Sex Study offered a simple but powerful criticism: the conclusion that the survey's accuracy
represents the sexual practices of adults in the United States "assumes that NORC respondents reflected the
population from which they were drawn and provided their data." exactly 8 That phrase could also be the
conclusion of all answers”.
this chapter. At first glance, the most suspicious thing about polls is that the opinions of
so few can tell us about the opinions of so many. But that's the easy part. One of the most basic statistical
principles is that an adequate sample will resemble the population from which it is drawn. The real challenge of
polling is twofold: finding and reaching the right sample; and getting information from that representative group
in a way that accurately reflects what its members believe.

APPENDIX TO CHAPTER 10

Why is the standard error larger when p (and 1 – p) are close to 50 percent?
Here's the intuition for why the standard error is larger when the proportion responding in a particular way (p) is
close to 50 percent (which, as a matter of mathematics, means that 1 – p will also be close to 50 percent). Let's
say you're conducting two surveys in North Dakota. The first poll is designed to measure the mix of
Republicans and Democrats in the state. Suppose the true political mix of North Dakota's population is evenly
split 50-50, but your poll finds 60 percent Republican and 40 percent Democrat. Their results are off by 10
percentage points, which is a large margin. However, you have made this huge mistake without making an
unimaginably large mistake in data collection. He has overcounted Republicans relative to their true incidence
in the population by 20 percent [(60 – 50)/50]. And in doing so, he has also underestimated the Democrats by
20 percent [(40 – 50)/50]. That could happen, even with decent survey methodology.
Their second survey is designed to measure the fraction of Native Americans in North Dakota's population.
Let us assume that the true proportion of natives
Americans in North Dakota make up 10 percent, while non-Native Americans make up 90 percent of the state's population. Now
let's consider how bad the data collection would have to be to produce a survey with a sampling error of 10 percentage points.
This could happen in two ways. First, you might find that 0 percent of the population is Native American and 100 percent is not
Native American. Or you might find that 20 percent of the population is Native American and 80 percent is non-Native American.
Machine Translated by
Google

In one case, all Native Americans were overlooked; in the other, twice their actual incidence in the population was found. These
are really serious sampling errors. In both cases, your estimate is 100 percent wrong: either [(0 – 10)/10] or [(20 – 10)/10].

And if only 20 percent of Native Americans were missed (the same degree of error as in the Republican-Democratic poll), the
results would be 8 percent Native Americans and 92 percent non-Native Americans, representing just 2 percentage points more
Native Americans. the real division of the population.
When py 1 – p are close to 50 percent, relatively small sampling errors are magnified into large absolute errors in the survey
result.
When p 1 – p are closer to zero, the opposite occurs. Even relatively
Large sampling errors produce small absolute errors in the survey results.
The same 20 percent sampling error skewed the poll result between Democrats and Republicans by 10 percentage points,
while it skewed the poll of Native Americans by just 2 percentage points. Since the standard error in a survey is measured in
absolute terms (for example, ±5 percent), the formula recognizes that this error is likely to be larger when p and 1 – p are close
to 50 percent.

* According to their website, “Occupy Wall Street is a people-powered movement that began on September 17, 2011, in Liberty Square
in Manhattan’s financial district and has spread to more than 100 cities in the United States and actions in more than 1,500 cities
around the world. . Occupy Wall Street is fighting against the corrosive power of big banks and multinational corporations over the
democratic process, and Wall Street's role in creating an economic collapse that has caused the biggest recession in generations. The
movement is inspired by the popular uprisings in Egypt and Tunisia, and aims to expose how the richest 1% of people are writing the
rules of an unfair global economy that is closing off our future.”
* We expect the Republican candidate's true vote count to be outside the poll's confidence interval about 5 percent of the time. In those
cases, their true vote count would be less than 50 percent or more than 54 percent. However, if he gets more than 54 percent of the
vote, his station has not made a mistake in declaring him the winner. (He has only underestimated the margin of his victory.) As a result,
the probability that his poll will lead him to wrongly declare the Republican candidate the winner is only 2.5 percent. * The formula for
calculating the standard error of a survey that I have presented here assumes that the survey is conducted on a random sample of the
population. Sophisticated polling organizations may deviate from this sampling method, in which case the formula for calculating the
standard error will also change slightly. However, the basic methodology remains the same.

CHAPTER 11

Regression analysis
The miraculous elixir

Can stress at work kill you? Yeah. There is compelling evidence that the rigors of work can lead to premature death,
especially from heart disease. But it's not the kind of stress you're probably imagining. CEOs, who routinely have to make
high-stakes decisions that determine the fate of their companies, are at significantly less risk than their secretaries, who
dutifully answer phones and perform other tasks as instructed. How can that make sense? It turns out that the most
dangerous type of job stress comes from having “low control” over one’s responsibilities. Several studies of thousands of
British civil servants (the Whitehall studies) have found that workers who have little control over their jobs (meaning they
have minimal say in what tasks are performed or how they are performed) have a significantly higher mortality rate than
other civil service workers with greater decision-making authority. According to this research, it's not the stress associated
Machine Translated by
Google

with important responsibilities that will kill you; it's the stress associated with being told what to do and having little say in
how or when to do it.

This is not a chapter about work stress, heart disease or British civil servants. The relevant question regarding the
Whitehall studies (and others like them) is how researchers can reach such a conclusion. Clearly this cannot be a random
experiment. We cannot arbitrarily assign human beings to different jobs, force them to work at those jobs for many years,
and then measure who dies at greater rates. (Ethical concerns aside, we would presumably wreak havoc on the British
civil service if we were to distribute posts at random.) Instead, researchers have collected detailed longitudinal data on
thousands of people in the British civil service; these data can be analysed to identify significant associations, such as the
connection between "low control" jobs and coronary heart disease.

A simple association is not enough to conclude that certain types of jobs are bad for your health. If we simply look at
low-ranking workers in the British civil service hierarchy as having higher rates of heart disease, our results would be
confounded by other factors. For example, we would expect low-level workers to have less education than senior officials
in the bureaucracy. They may be more likely to smoke (perhaps due to job frustration). They may have had a less
healthy childhood, which diminished their job prospects. Or your lower wage may limit your access to health
care. Etc. The point is that any study that simply compares the health outcomes of a large group of British
workers (or any other large group) won't really tell us much. Other sources of variation in the data are likely to
obscure the relationship we care about. Is “low job control” really causing heart disease? Or is it some
combination of other factors shared by people with low job control, in which case we may be missing the real
threat to public health altogether?
Regression analysis is the statistical tool that helps us face this challenge. Specifically, regression analysis
allows us to quantify the relationship between a particular variable and an outcome we are interested in while
controlling for other factors. In other words, we can isolate the effect of one variable, such as having a certain
type of job, while holding constant the effects of other variables. The Whitehall studies used regression
analysis to measure the health impacts of low job control among people who are similar in other ways, such as
smoking behaviour. (In fact, lower-level workers smoke more than their superiors; this explains a relatively
small amount of the variation in heart disease in the Whitehall hierarchy.)
Most of the studies you read about in the paper are based on regression analysis. When researchers
conclude that children who spend a lot of time in daycare are more likely to have behavioral problems in
elementary school than children who spend that time at home, the study did not randomly assign thousands of
infants to daycare or home care with a parent. The study also did not limit itself to comparing the elementary
school behavior of children who had different early childhood experiences without acknowledging that these
populations are likely to be different in other key ways. Different families make different decisions about child
care because they are different. In some homes there are two parents present; some do not. Some have two
working parents; some don't.
Some households are wealthier or more educated than others. All of these things affect decisions about child
care and affect how well children in those families perform in elementary school. When done correctly,
regression analysis can help us estimate the effects of daycare apart from other things that affect young
Machine Translated by
Google

children: family income, family structure, parents' education, etc.


Now, there are two key phrases in that last sentence. The first is "when done correctly." With adequate data
and access to a personal computer, a six-year-old child could use a basic statistics program to generate
regression results. Personal computing has made the mechanics of regression analysis almost simple. The
problem is that the mechanics of regression analysis are not the hard part;
The difficult part is determining which variables should be considered in the analysis and how to best do it. Regression
analysis is like one of those sophisticated powerful tools. It is relatively easy to use, but difficult to use well and potentially
dangerous if used incorrectly.

The second important sentence above is "help us estimate." Our study of child care does not give us a “correct”
answer about the relationship between day care and later school performance. Instead, it quantifies the observed
relationship for a particular group of children over a particular time period. Can we draw conclusions that could be applied
to the general population? Yes, but we will have the same limitations and caveats as with any other type of inference.
First, our sample has to be representative of the population we care about. A study of 2,000 young children in Sweden
won't tell us much about the best policies for early childhood education in rural Mexico. And secondly, there will be
variations from sample to sample. If we conduct multiple studies on children and child care, each study will produce
slightly different findings, even if the methodologies are all sound and similar.

Regression analysis is similar to surveys. The good news is that if we have a large representative sample and a sound
methodology, the relationship we observe for our sample data is not likely to deviate much from the true relationship for
the entire population. If 10,000 people who exercise three or more times per week have much lower rates of
cardiovascular disease than 10,000 people who do not exercise (but are similar in all other important ways), then there is
a good chance that we would see a similar association. between exercise and cardiovascular health for the general
population. That's why we do these studies.
(The point is not to tell those who do not exercise and are sick at the end of the study that they should have exercised.)
The bad news is that we are not definitively proving that exercise prevents heart disease. Instead, we are rejecting the
null hypothesis that exercise has no association with heart disease, based on some statistical threshold that was chosen
before the study was conducted. Specifically, the study authors would report that if exercise is not related to
cardiovascular health, the probability of observing such a marked difference in heart disease between athletes and
nonathletes in this large sample would be less than 5 in 100, or below some other threshold of statistical significance.

Let's pause for a moment and wave our first giant yellow flag. Suppose this particular study compared a large group of
individuals who play squash regularly with those in an equally sized group who do not exercise at all. Playing squash
provides a good cardiovascular workout. However, we also know that squash players tend to be wealthy enough to belong
to squash clubs.
Machine Translated by
Google

courts. Wealthy people may have greater access to healthcare, which can also improve cardiovascular health.
If our analysis is sloppy, we may attribute health benefits to playing squash when in fact the real benefit comes
from being wealthy enough to play squash (in which case playing polo would also be associated with better
heart health, even though the horse is getting more exercise). from work).
Or maybe causality goes in the other direction. Could having a healthy heart “cause” exercise? Yeah. Sick
people, especially those with early forms of heart disease, will find it much harder to exercise. They will
probably be less likely to play squash regularly. Again, if the analysis is sloppy or oversimplified, the claim that
exercise is good for health may simply reflect the fact that people who start out in poor health find it difficult to
exercise. In this case, playing squash doesn't make anyone healthier; it simply separates the healthy from the
unhealthy.
There are so many potential pitfalls in regression that I have devoted the next chapter to the most egregious
mistakes. For now, we'll focus on what can go right.
Regression analysis has the uncanny ability to isolate a statistical relationship that interests us, such as that
between job control and heart disease, while also taking into account other factors that might confound the
relationship.
How exactly does this work? If we know that low-level British civil servants smoke more than their superiors,
how can we discern how much of their poor cardiovascular health is due to their low-level jobs and how much
is due to smoking? These two factors seem inextricably intertwined.
Regression analysis (done correctly!) can untangle them. To explain the intuition, I need to start with the
basic idea underlying all forms of regression analysis, from the simplest statistical relationships to the complex
models cobbled together by Nobel Prize winners. In essence, regression analysis seeks to find the “best fit” for
a linear relationship between two variables. A simple example is the relationship between height and weight.
Taller people tend to weigh more, although this is obviously not always the case. If we were to plot the heights
and weights of a group of graduate students, you might remember what they looked like in Chapter 4:

Scatter plot for height and weight


Machine Translated by
Google

Height (inches)

If you were asked to describe the pattern, you might say something like, “Weight seems to increase with height.” This
is not a terribly revealing or specific statement. Regression analysis allows us to go one step further and “fit a line” that
better describes a linear relationship between the two variables.
Many possible lines generally match the height and weight data.
But how do we know which is the best line for this data? In fact, how exactly would we define “better”? Regression
analysis typically uses a methodology called ordinary least squares, or OLS. The technical details, including why OLS
produces the best fit, will have to be left for a more advanced book. The key point lies in the “least squares” part of the
name; OLS fits the line that minimizes the sum of the squared residuals. That's not as complicated as it sounds.
Each observation in our height and weight data set has a residual, which is its vertical distance from the regression line,
except for those observations that lie directly on the line, for which the residual is equal to zero. (In the diagram below, the
residual is marked for a hypothetical person A.) It should be intuitive that the larger the sum of the residuals overall, the
worse the fit of the line. The only non-intuitive twist to OLS is that the formula takes the square of each residual before
adding them all up (which increases the weight given to observations that lie particularly far from the regression line, or
“outliers”).

Ordinary least squares "fits" the line that minimizes the sum of squared residuals, as illustrated below.

Best fit line for height and weight


Machine Translated by
Google

If the technical details have given you a headache, you might be forgiven for simply understanding the
conclusion: ordinary least squares gives us the best description of a linear relationship between two variables.
The result is not just a line but, as you may recall from high school geometry, an equation that describes that
line. This is known as a regression equation and takes the following form: y = a + bx, where y is the weight in
pounds; a is the y-intercept of the line (the y value when x = 0); b is the slope of the line; and x is the height in
inches. The slope of the line we fitted, b, describes the "best" linear relationship between height and weight for
this sample, as defined by ordinary least squares.
The regression line certainly does not perfectly describe all the observations in the data set. But it is the
best description we can find of what is clearly a significant relationship between height and weight. It also
means that each observation can be explained as WEIGHT = a + b(HEIGHT) + e, where e is a “residual” that
captures the variation in each individual's weight that is not explained by height. Ultimately, it means that our
best estimate for the weight of any person in the dataset would be a + b(HEIGHT). Even though most
observations do not lie exactly on the regression line, the residual still has an expected value of zero, since any
person in our sample is just as likely to weigh more than the regression equation predicts as to weigh less.
Enough of the theoretical jargon! Let's look at some actual height and weight data from the Changing Lives
study, but first I need to clarify some basic terminology. The variable being explained (the weight in this case)
is known as the dependent variable (because it depends on other factors). The variables we use to explain our
dependent variable are known as explanatory variables since they explain the outcome we care about. (To
make matters more difficult, explanatory variables are sometimes also called independent variables or control
variables.) Let's start by using height to explain
weight among Changing Lives participants; We will add other potentials later. explanatory factors. There are 3,537 adults
participating in the Changing Lives study. This is our number of observations, or n. (Sometimes a research paper may
state that n = 3537.) When we run a simple regression on the Changing Lives data with weight as the dependent variable
and height as the only explanatory variable, we get the following results:

WEIGHT = –135 + (4.5) × HEIGHT IN INCHES

a = –135. This is the y-intercept, which by itself has no particular meaning. (If you take it literally, a person who is zero
Machine Translated by
Google

inches tall would weigh at least 135 pounds; obviously this is nonsense on several levels.) This figure is also known as a
constant, because it is the starting point for calculating the weight of all observations. in the studio.

b = 4.5. Our estimate of b, 4.5, is known as the regression coefficient, or in statistical jargon, the “height coefficient,”
because it gives us the best estimate of the relationship between height and weight among Changing Lives participants.
The regression coefficient has a convenient interpretation: a one-unit increase in the independent variable (height) is
associated with a 4.5-unit increase in the dependent variable (weight). For our data sample, this means that a 1-inch
increase in height is associated with a 4.5-pound increase in weight.

Therefore, if we had no other information, our best estimate for the weight of a person who is 5 feet 10 inches (70 inches)
tall in the Changing Lives study would be – 135 + 4.5 (70) = 180 pounds.
This is our reward, as we have now quantified the best fit for the linear relationship between height and weight for
Changing Lives participants. The same basic tools can be used to explore more complex relationships and more socially
significant issues. For any regression coefficient, you will generally be interested in three things: sign, size, and
significance.
Sign. The sign (positive or negative) of the coefficient of an independent variable tells us the direction of its association
with the dependent variable (the result we are trying to explain). In the simple case above, the height coefficient is
positive. Taller people tend to weigh more. Some relationships will work in the other direction. I would expect the
association between exercise and weight to be negative. If the Changing Lives study included data on something like
“miles traveled per month,” I’m pretty sure the coefficient on “miles traveled” would be negative. Running more is
associated with weighing less.

Size. How large is the observed effect between the independent variable and the dependent variable? Is it of a
magnitude that matters? In this case, each inch of height is associated with 4.5 pounds, which is a considerable
percentage of a

typical body weight of a person. To explain why some people weigh more than others, height is clearly an
important factor. In other studies, we may find an explanatory variable that has a statistically significant impact
on our outcome of interest (meaning the observed effect is probably not a product of chance), but that effect
may be so small as to be trivially or socially insignificant. insignificant.
For example, suppose we are examining the determinants of income. Why do some people make more money
than others? Explanatory variables are likely to be things like education, years of work experience, etc. In a
large data set, researchers might also find that people with whiter teeth earn $86 more per year than other
workers, ceteris paribus. (“Ceteris paribus” comes from Latin and means “under equal conditions”). The
positive and statistically significant coefficient of the variable “white teeth” means that the individuals compared
are similar in other aspects: same education, same work experience, etc. (I'll explain in a moment how we
accomplished this tantalizing feat.) Our statistical analysis has shown that whiter teeth are associated with $86
in additional annual income per year and that this finding is unlikely to be a mere coincidence. This means (1)
that we have rejected the null hypothesis that really white teeth have no association with income with a high
degree of confidence; and (2) if we look at other data samples, we are likely to find a similar relationship
Machine Translated by
Google

between nice teeth and higher income.


So what? We found a statistically significant result, but not one that is particularly meaningful. For starters,
$86 a year is not a life-changing sum of money. From a public policy standpoint, $86 is probably also less than
it would cost to whiten one person's teeth each year, so we can't even recommend that young workers make
such an investment. And, although I still have one more chapter to go, I would also be concerned about some
serious methodological problems. For example, having perfect teeth may be associated with other personality
traits that explain the wage advantage; the income effect may be caused by the type of people who care about
their teeth, not the teeth themselves. For now, the point is that we need to take note of the size of the
association we observe between the explanatory variable and our outcome of interest.

Meaning. Is the observed result an aberration based on a peculiar sample of data, or does it reflect a
significant association that is likely to be observed in the population as a whole? This is the same basic
question we have been asking ourselves in the last few chapters. In the context of height and weight, do we
think we would observe a similar positive association in other representative samples of the population? To
answer this question, we use the basic inference tools that have already been introduced. Our regression
coefficient
It is based on an observed relationship between height and weight for a particular sample of data. If we were to
test another large sample of data, we would almost certainly get a slightly different association between height
and weight, and therefore a different coefficient. The relationship between height and weight observed in the
Whitehall data is likely to be different from the relationship between height and weight observed for participants
in the Changing Lives study. However, we know from the central limit theorem that the mean of a large,
correctly drawn sample will usually not deviate much from the mean of the population as a whole. Similarly, we
can assume that the observed relationship between variables such as height and weight will not typically vary
wildly from sample to sample, assuming these samples are large and appropriately drawn from the same
population.

Think intuitively: It's highly unlikely (though still possible) that we'd find that each centimeter of height was
associated with an additional 4.5 pounds among Changing Lives participants, but that there was no association
between height and weight in a different representative sample of 3,000 American adults.
This should give you a first idea of how we will test whether the results of our regression are statistically
significant or not. As with surveys and other forms of inference, we can calculate a standard error for the
regression coefficient. The standard error is a measure of the likely dispersion that we would observe in the
coefficient if we performed the regression analysis on repeated samples drawn from the same population. If we
were to measure and weigh a different sample of 3,000 Americans, we might find in further analysis that each
inch of height is associated with 4.3 pounds. If we did it again with another sample of 3,000 Americans, we
might find that each inch is associated with 5.2 pounds. Once again, the normal distribution is our friend. For
large data samples, such as our Changing Lives dataset, we can assume that our various coefficients will be
normally distributed around the "true" association between height and weight in the US adult population.
Starting from that assumption, we can calculate a standard error for the regression coefficient that gives us an
idea of how much dispersion we should expect in the coefficients from one sample to another. I won't go into
Machine Translated by
Google

the formula for calculating the standard error here, because it will take us in a direction that involves a lot of
math and because all basic statistical packages will calculate it for you.

However, I must warn that when we are working with a small sample of data (such as a group of 20 adults
instead of the more than 3,000 people in the Changing Lives study) the normal distribution is no longer willing
to be our friend.
Specifically, if we repeatedly perform regression analyses on different small samples, we can no longer
assume that our various coefficients will be normally distributed around the "true" association between height
and weight in the US adult population. Instead, our coefficients will continue to be distributed around the “true”
association between height and weight for the US adult population in what is known as the t distribution.
(Basically, the t-distribution is more spread out than the normal distribution and therefore has “fatter tails.”)
Nothing else changes; any basic statistical software package will easily handle the additional complexity
associated with using t distributions. For this reason, the t-distribution will be explained in more detail in the
chapter's appendix.

Sticking with large samples for now (and the normal distribution), the most important thing to understand is
why the standard error matters. As with surveys and other forms of inference, we expect more than half of our
observed regression coefficients to lie within one standard error of the standard deviation. [Link].
cpiaórná[Link]. Roughly 95 percent will be within two errors. With that, we're almost home, because now
we can do a little hypothesis testing. (Seriously, did you think you were done with hypothesis testing?) Once
we have a coefficient and a standard error, we can test the null hypothesis that there is, in fact, no relationship
between the explanatory variable and the dependent variable (meaning that the true association between the
two variables in the population is zero).

In our simple height and weight example, we can test how likely it is that we will find in our Changing Lives
sample that each centimeter of height is associated with 4.5 pounds if there really is no association between
height and weight in the general population. I ran the regression using a basic statistics program; the standard
error in the height coefficient is 0.13. This means that if we were to do this analysis repeatedly (say with 100
different samples), then we would expect our observed regression coefficient to be within two standard errors
of the true population parameter about 95 times out of 100.
Therefore, we can express our results in two different but related ways. First, we can construct a 95 percent
confidence interval. We can say that 95 times out of 100 we expect our confidence interval, which is 4.5 ± 0.26,
to contain the true population parameter. This is the range between 4.24 and 4.76. A basic statistics package
will also calculate this range. Second, we can see that our 95 percent confidence interval for the true
association between height and weight does not include zero. Therefore, we can reject the null hypothesis that
there is no association between height and weight for the general population with a 95 percent confidence
level. This result can also be expressed as statistically significant at the 0.05 level; there is only a 5 percent
chance that we are wrongly rejecting the null hypothesis.
Machine Translated by
Google

In fact, our results are even more extreme than that. The standard error (0.13) is extremely low relative to
the size of the coefficient (4.5). A rough rule of thumb is that the coefficient is likely to be statistically significant
when it is at least twice the size of the standard error. A package of

Statistics also calculates a p-value, which is 0.000 in this case, meaning that there is essentially zero chance of
getting a result as extreme as the one we've observed (or more) if there is no true association between height
and weight. and weight in the general population. Remember, we have not shown that taller people weigh
more in the general population; we have simply shown that our results for the Changing Lives sample would be
highly anomalous if that were not the case.

Our basic regression analysis produces another notable statistic: the R is a measure of the total amount of
variation explained by the regression equation. We know that we have a wide 2 which
variation in the weight of our Changing Lives sample. Many of the people in the ,

sample weigh more than the average for the group as a whole; many weigh less. The R tells us how much of
that variation around the mean is associated solely with differences in height. The answer in our case is 0.25,
or 25 percent. The most significant point may be that 75 percent of the weight variation in our sample remains
unexplained. Clearly, there are factors other than height that could help us understand the weight of Changing
Lives participants. This is where things get more interesting.
I admit that I began this chapter by selling regression analysis as the miracle elixir of social science
research. So far, all I've done is use a statistics package and an impressive data set to show that tall people
tend to weigh more than short people. A short visit to a shopping mall would probably have convinced him of
the same. Now that you understand the basics, we can unleash the true power of regression analysis. It's time
to take off the training wheels!

As promised, regression analysis allows us to unravel complex relationships in which multiple factors affect
some outcome we care about, such as income, test scores, or heart disease. When we include multiple
variables in the regression equation, the analysis gives us an estimate of the linear association between each
explanatory variable and the dependent variable while holding other dependent variables constant, or
“controlling” for these other factors. Let's stick with the weight for a while. We have found an association
between height and weight; we know that there are other factors that can help explain weight (age, sex, diet,
exercise, etc.). Regression analysis (often called multiple regression analysis when more than one explanatory
variable is involved, or multivariate regression analysis) will give us a coefficient for each explanatory variable
included in the regression equation. In other words, among people of the same sex and height, what
is the relationship between age and weight? Once we have more than one explanatory variable, we
can no longer represent these data in two dimensions. (Try to imagine a graph that represents the
weight, sex, height, and age of each participant in the Changing Lives study.)
However, the basic methodology is the same as in our simple height and weight example. As we add
explanatory variables, a statistical package will calculate the regression coefficients that minimize the
total sum of squared residuals in the regression equation.
Let's work with the Changing Lives data for now, and then I'll come back and give an intuitive
explanation of how this statistical parting of the Red Sea might work. We can start by adding one
more variable to the equation that explains the weights of Changing Lives participants: age. When we
Machine Translated by
Google

run the regression including height and age as explanatory variables for weight, this is what we get.
WEIGHT = –145 + 4.6 × (HEIGHT IN INCHES) + 0.1 × (AGE IN YEARS)
The age coefficient is 0.1. This can be interpreted to mean that each additional year of age is
associated with an additional 0.1 pound of weight, holding height constant. For any group of people of
the same height, on average those who are ten years older will weigh one pound more. This is not a
huge effect, but it is consistent with what we tend to see in life. The coefficient is significant at the
0.05 level.
You may have noticed that the height coefficient has increased slightly. Once age is in our
regression, we have a more refined understanding of the relationship between height and weight.
Among people who are the same age in our sample, or who “keep age constant,” each additional inch
of height is associated with 4.6 pounds of weight gain.
Let's add one more variable: sex. This will be slightly different because sex can only accept two
possibilities, male or female. How do you put M or F in a regression? The answer is that we use what
is called a binary variable or dummy variable. In our dataset, we enter a 1 for participants who are
female and a 0 for those who are male. (This is not intended to be a value judgment.) The sex
coefficient can then be interpreted as the effect on weight of being female, ceteris paribus. The
coefficient is –4.8, which is not surprising. We can interpret this to mean that for people of the same
height and age, women typically weigh 4.8 pounds less than men. Now we can begin to see some of
the power of multiple regression analysis. We know that women tend to be shorter than men, but our
coefficient takes this into account since we have already controlled for height. What we have isolated here
is the effect of being a woman. The new regression becomes:
WEIGHT = –118 + 4.3 × (HEIGHT IN INCHES) + 0.12 (AGE IN YEARS) – 4.8 (IF SEX IS FEMALE)

Our best estimate of the weight of a fifty-three-year-old woman who is 5 feet 5 inches tall is: –118 + 4.3 (65)
+ 0.12 (53) – 4.8 = 163 pounds.
And our best guess for a thirty-five year old male who is 6 feet 3 inches tall is: 118 + 4.3 (75) + 0.12 (35) =
209 pounds. We skip the last term in our regression result (–4.8) since this person is not female.

Now we can start trying things that are more interesting and less predictable. What about education? How
could that affect weight? I would hypothesize that better educated people are more health conscious and will
therefore weigh less, ceteris paribus. We also didn't test any exercise measures; I assume that, holding other
factors constant, people in the sample who exercised more would weigh less.
What about poverty? Does being low-income in the United States have effects on weight? The Changing
Lives study asks whether participants receive food stamps, which is a good measure of poverty in the United
States. Finally, I am interested in race. We know that people of color have different life experiences in the
United States because of their race. There are cultural and residential factors associated with race in the
United States that have implications for weight. Many cities are still characterized by a high degree of racial
segregation; African Americans may be more likely than other residents to live in “food deserts,” which are
areas with limited access to grocery stores that sell fresh fruits, vegetables, and other produce.
Machine Translated by
Google

We can use regression analysis to separate the independent effect of each of the possible explanatory
factors described above. For example, we can isolate the association between race and weight, holding
constant other socioeconomic factors such as educational attainment and poverty. Among people who
graduated from high school and are eligible for food stamps, what is the statistical association between weight
and being black?
At this point, our regression equation is so long that it would be cumbersome to print the results here in their
entirety. Academic papers often include large tables summarizing the results of several regression equations. I
have included a table with the complete results of this regression equation in the appendix of this chapter. In
the meantime, here are the highlights of what happens when we add up education, exercise, poverty
(measured by food stamp receipt) and
race towards our equation.
All of our original variables (height, age, and sex) remain significant. The coefficients change little as we add
explanatory variables. All our new variables are statistically significant at the 0.05 level. The R in the 2 regression ranges
from zero to 0.29. (Remember, an R equation does not predict any individual's weight better than the mean, which means
our regression equation rose by 0.25 from the mean of any individual's sample weight.) Much of the sample is perfectly
fine; R is the weight of each person in the variation in weight between individuals remains unexplained.

Education turns out to be negatively associated with weight, as hypothesized. Among participants in the Changing
Lives study, each year of education is associated with -1.3 pounds.
Not surprisingly, exercise is also negatively associated with weight. The Changing Lives study includes an index that
rates each study participant according to their level of physical activity. Individuals in the bottom quintile for physical
activity weigh, on average, 4.5 pounds more than other adults in the sample, ceteris paribus. Those in the bottom quintile
of physical activity weigh, on average, almost 9 pounds more than adults in the top quintile of physical activity.

People receiving food stamps (the poverty indicator in this regression) are heavier than other adults. Food stamp
recipients weigh an average of 5.6 pounds more than other Changing Lives participants, ceteris paribus.
The racial variable is particularly interesting. Even after controlling for all the other variables described up to this point,
race is still very important when it comes to explaining weight. Non-Hispanic Black adults in the Changing Lives sample
weigh, on average, about 10 pounds more than other adults in the sample. Ten pounds is a lot of weight, both in absolute
terms and compared to the effects of the other explanatory variables in the regression equation. This is not a peculiarity of
the data. The p-value for the dummy variable for non-Hispanic blacks is 0.000 and the 95 percent confidence interval
extends from 7.7 pounds to 16.1 pounds.

What's going on? The honest answer is that I have no idea. Let me reiterate a point that was buried earlier in a
footnote: I'm just playing with data here to illustrate how regression analysis works. The analyses presented here are to
true academic research what street hockey is to the NHL. If this were a real research project, there would be weeks or
months of follow-up analysis to test this finding. What I can say is that I have demonstrated why multiple regression
analysis is the best tool we have for finding meaningful patterns in large numbers. complex data sets. We start with a
ridiculously banal exercise: quantifying the relationship between height and weight. Before long we were knee-
Machine Translated by
Google

deep in issues of real social importance.


In that regard, I can offer you an actual study that used regression analysis to investigate a socially
significant issue: gender discrimination in the workplace. The funny thing about discrimination is that it's hard to
observe directly. No employer explicitly states that someone is paid less because of their race or gender or that
someone has not been hired for discriminatory reasons (which would presumably leave the person in a
different job with a lower salary).
Instead, what we see are racial and gender wage gaps that may be the result of discrimination: whites earn
more than blacks; men earn more than women; and so on. The methodological challenge is that these
observed gaps may also be the result of underlying differences between workers that have nothing to do with
workplace discrimination, such as the fact that women tend to choose more part-time jobs. How much of the
wage gap is due to factors associated with productivity at work, and how much of the gap, if any, is due to
workforce discrimination? No one can claim that this is a trivial matter.
Regression analysis can help us answer this question. However, our methodology will be a little more
indirect than it was with our analysis explaining weight. Since we cannot measure discrimination directly, we
will look at other factors that traditionally explain wages, such as education, experience, occupational field, etc.
The arguments in favor of discrimination are circumstantial: if a significant wage gap persists after controlling
for other factors that typically explain wages, then discrimination is a likely culprit. The larger the unexplained
portion of any pay gap, the more suspicious we should be. As an example, consider a paper by three
economists examining the wage trajectories of a sample of about 2,500 men and women who graduated with
an MBA from the University of Chicago Booth School of Business. Upon graduation, male and female
graduates have
very similar average starting salaries: $130,000 for men and $115,000 for women. However, after ten years in
the workforce, a huge gap has opened up; on average, women earn a staggering 45 percent less than their
male classmates: $243,000 versus $442,000. In a broader sample of more than 18,000 MBA graduates who
entered the workforce between 1990 and 2006, being a woman was associated with 29 percent lower
earnings. What happens to women once they enter the workforce?
According to the study's authors (Marianne Bertrand of the Booth School of Business and Claudia Goldin
and Lawrence Katz of Harvard), discrimination is not a likely explanation for most of the gap. The gender pay
gap fades as the authors add more explanatory variables to the analysis. For example, men
take more finance classes in the MBA program and graduate with higher grade point averages. When these data are
included as control variables in the regression equation, the unexplained share of the gap between men's and women's
earnings falls to 19 percent. When variables are added to the equation to account for post-MBA work experience,
particularly outside the workforce, the unexplained portion of the gender pay gap drops to 9 percent. And when
explanatory variables for other job characteristics, such as employer type and hours worked, are added, the unexplained
portion of the gender pay gap falls to less than 4 percent.

For workers who have been in the workforce more than ten years, the authors can ultimately explain all but 1 percent
of the gender wage gap with factors unrelated to workplace discrimination. They conclude: “We identify three proximate
reasons for the large and growing gender earnings gap: differences in pre-MBA education; differences in career
interruptions; and differences in weekly hours. These three determinants can explain most of the gender differences over
Machine Translated by
Google

the years following completion of the MBA.”

I hope I have convinced you of the value of multiple regression analysis, particularly the research insights that arise from
being able to isolate the effect of an explanatory variable while controlling for other confounding factors. I have not yet
provided an intuitive explanation of how this statistical “miracle elixir” works. When we use regression analysis to assess
the relationship between education and weight, ceteris paribus, how does a statistical package control for factors such as
height, sex, age and income when we know that our Changing Lives participants are not identical in these other respects?

To understand how we can isolate the effect on weight of a single variable, say education, imagine the following
situation. Suppose all the Changing Lives participants are gathered in one place, say, Framingham, Massachusetts. Now
suppose that men and women are separated. And then let's assume that both men and women are divided by height.
There will be a room of six-foot-tall men. Next door will be a room for 6-foot-1-inch men, and so on for both sexes. If we
have enough participants in our study, we can further subdivide each of those rooms by income. Over time we will have
many rooms, each containing individuals who are identical in every respect except education and weight, which are the
two variables we care about. There would be a room of forty-five-year-old men, 5 feet 5 inches tall, making $30,000 to
$40,000 a year. Next to them would be all the 45-year-old women who are 5 feet 5 inches tall and make between $30,000
and $40,000 a year. And so on (and so on).
There will still be some weight variation in each room; people who are the

People of the same sex and height and with the same income will weigh different amounts, although there will presumably
be much less variation in weight in each room than in the overall sample. Our goal now is to see how much of the
remaining variation in weight in each room can be explained by education. In other words, what is the best linear
relationship between education and weight in each room?

The final challenge, however, is that we don't want different coefficients in each “room.” The objective of this exercise
is to calculate a single coefficient that best expresses the relationship between education and weight for the entire sample,
keeping other factors constant. What we would like to calculate is the unique education coefficient that we can use in each
room to minimize the sum of the squared residuals of all rooms combined. What education coefficient minimizes the
square of the unexplained weight of each individual in all rooms? This becomes our regression coefficient because it is the
best explanation of the linear relationship between education and weight for this sample when we hold sex, height, and
income constant.

Plus, you can see why big data sets are so useful. They allow us to control many factors and at the same time have
many observations in each “room”. Obviously, a computer can do all this in a fraction of a second without having to gather
thousands of people in different rooms.

Let's end the chapter where we started, with the connection between work stress and coronary heart disease. The
Whitehall studies of British civil servants attempted to measure the association between employment status and death
from coronary heart disease in subsequent years. One of the first studies followed 17,530 civil servants for seven and a
half years.
2 The authors concluded: “Men in
Machine Translated by
Google

the lower occupational grades were shorter, heavier for their height, had higher blood pressure, higher plasma glucose,
smoked more, and reported less leisure-time physical activity than men in the higher grades. However, when the influence
on mortality of all these factors plus plasma cholesterol was taken into account, the inverse association between
employment status and [coronary heart disease] mortality remained strong.” The “assignment” they refer to for these other
known risk factors is done using regression analysis.
The study shows that, holding other health factors constant (including height, which is a decent indicator of
early childhood health and nutrition), working in a “menial” job can literally kill you.
Skepticism is always a good first response. At the beginning of the chapter I wrote that “low control” jobs are bad for
your health. That may or may not be synonymous with being low on the management totem pole. A follow-up study

Using a second sample of 10,308 British civil servants, an attempt was made to further explore this distinction. 3 Once
again, workers were divided into management grades (high, intermediate, and low), only this time participants were also
given a fifteen-item questionnaire assessing their level of “latitude of decision or control.” These included questions such
as "Do you have a choice about how you do your work?" and categorical responses (ranging from "never" to "often") to
statements such as "I can decide when to take a break." The researchers found that "low control" workers had a
significantly higher risk of developing coronary heart disease over the course of the study than "high control" workers.
However, the researchers also found that workers with rigorous job demands were not at increased risk for developing
heart disease, nor were workers who reported low levels of social support at work. Lack of control seems to be the cause
of death, literally.

Whitehall studies have two characteristics typically associated with solid research. First, the results have been
replicated elsewhere. In the public health literature, the idea of “low control” has evolved into a term known as “job strain,”
which characterizes jobs with “high psychological workload demands” and “low decision-making latitude.” Thirty-six
studies on the topic were published between 1981 and 1993; most found a significant positive association between job
strain and heart disease. 4

Second, the researchers searched for and found biological evidence to support the mechanism by which this particular
type of job stress causes poor health. Working conditions that involve rigorous demands but little control can trigger
physiological responses (such as the release of stress-related hormones) that increase the risk of heart disease over the
long term. Even animal research plays a role; low-status monkeys and baboons (which bear some resemblance to civil
servants at the bottom of the authority chain) have physiological differences from their high-status peers that put them at
greater cardiovascular risk. 5 All things being equal, it's better not to be a low-status baboon, which is a point I try to get
across to my children as often as possible, particularly my son. The broader message here is that regression analysis is
arguably the most important tool researchers have for finding meaningful patterns in large data sets. We generally can't
conduct controlled experiments to learn about job discrimination or the factors that cause heart disease. Our insights into
these socially significant issues and many others come from the statistical tools covered in this chapter. In fact, it would
not be an exaggeration to say that a high proportion of all important research done in the social sciences over the last half
century (particularly since the advent of cheap computing power) is based on regression analysis.
Regression analysis supercharges the scientific method; as a result, we are healthier, safer, and better informed.
So what could possibly go wrong with this powerful and impressive tool? Keep reading.
Machine Translated by
Google

APPENDIX TO CHAPTER 11
The t distribution

Life gets a little more complicated when we do our regression analysis (or other forms of statistical inference) on a small
sample of data. Suppose we were analyzing the relationship between weight and height based on a sample of just 25
adults, rather than using a huge data set like the Changing Lives study.
Logic suggests that we should have less confidence in generalizing our results to the entire adult population from a
sample of 25 than from a sample of 3,000.
One of the themes throughout the book has been that smaller samples tend to generate greater dispersion in results. Our
sample of 25 will still give us meaningful information, as will a sample of 5 or 10, but how meaningful is it?
The t-distribution answers that question. If we analyze the association between height and weight for repeated
samples of 25 adults, we can no longer assume that the various coefficients we obtain for height will be normally
distributed around the "true" coefficient for height in the adult population. They will still be distributed around the true
coefficient for the entire population, but the shape of that distribution will not be our familiar bell-shaped normal curve.
Instead, we have to assume that repeated samples of only 25 will produce a greater spread around the true population
coefficient, and hence a distribution with “fatter tails.” And repeated samples of 10 will produce even greater spread than
that, and therefore even fatter tails. The t-distribution is actually a series or “family” of probability density functions that
vary depending on the size of our sample. Specifically, the more data we have in our sample, the more “degrees of
freedom” we have in determining the appropriate distribution with which to evaluate our results. In a more advanced class,
you will learn exactly how to calculate degrees of freedom; for our purposes, they are approximately equal to the number
of observations in our sample. For example, a basic regression analysis with a sample of 10 and a single explanatory
variable has 9 degrees of freedom. The more degrees of freedom we have, the more confident we can be that our sample
represents the real population and the “stricter” our distribution will be, as illustrated in the following diagram.

As the number of degrees of freedom increases, the t distribution converges to the normal distribution. That
is why when we work with large data sets, we can use the normal distribution for our various calculations.
Machine Translated by
Google

The t-distribution simply adds nuance to the same process of statistical inference we have been using
throughout the book. We are still formulating a null hypothesis and then testing it against some observed data.
If the data we observe would be highly improbable if the null hypothesis were true, then we reject the null
hypothesis. The only thing that changes with the t-distribution are the underlying probabilities for evaluating the
observed outcomes. The “fatter” the tail on a particular probability distribution (e.g., the t-distribution for eight
degrees of freedom), the more dispersion we would expect in our observed data simply as a matter of chance,
and therefore the less certain we can be. by rejecting our null hypothesis.

For example, suppose we are running a regression equation and the null hypothesis is that the coefficient of
a particular variable is zero. Once we have the regression results, we would calculate a t-statistic, which is the
ratio of the observed coefficient to the standard error of that coefficient. then evaluated against * This t-statistic
is any t-distribution that is appropriate for the sample size of the data (since this is largely what determines the
number of degrees of freedom). When the t-statistic is sufficiently large, meaning that our observed coefficient
is far from what the null hypothesis would predict, we can reject the null hypothesis at some level of statistical
significance. Again, this is the same basic process of statistical inference that we have been employing
throughout the book.

The fewer the degrees of freedom (and therefore the “fatter” the tails of the relevant t distribution), the
larger the t statistic will have to be for us to reject the null hypothesis at any given significance level.
In the hypothetical regression example described above, if we had four degrees of freedom, we would
need a t-statistic of at least 2.13 to reject the null hypothesis at the 0.05 level (in a one-tailed test).
However, if we have 20,000 degrees of freedom (which essentially allows us to use the normal
distribution), we would need only a t-statistic of 1.65 to reject the null hypothesis at the 0.05 level in
the same one-tailed test.

Regression equation for weight


Standard p-valuc (two- 95% Confidence
Mistake tailed test) Interval
Variable Coefficient t-statistic
Height 4.4 .2 21.4 .000 4.0 to 4.8

Age 08 .03 2.2 .026 DI to .2


Sex -5.7 1.7 -3.4 .001 -9.0 to -2.4

Years of Education
Attainment -.7 .2 -3.S .000 -1.1 to -3

Bottom Quintile of 3.7 1.4 2.6 .009 .9 to 6.5


Physics Activity
Dummies fair
Receiving Food 5.6 2.1 2.7 .007 1.5 0 9.7

Stamps

Non-Hispanic Black 9.7 IJ 7.2 .000 7Dw 12.3

Intercept -117

* You should consider this exercise as “fun with data” rather than an authoritative exploration of any of the relationships described in the regression
equations below. The purpose here is to provide an intuitive example of how regression analysis works, not to conduct meaningful research on the
weight of Americans.
* “Parameter” is a fancy term for any statistic that describes a characteristic of some population; the average weight of all adult males is a parameter of
Machine Translated by
Google

that population. So is the standard deviation. In the example here, the true association between height and weight for the population is a parameter of
that population.
* When the null hypothesis is that a regression coefficient is zero (as is often the case), the ratio of the observed regression coefficient to the standard
error is known as the t statistic. This will also be explained in the appendix of the chapter.

* Broader discriminatory forces in society may affect the careers women choose or the fact that they are more likely than men to interrupt their careers
to care for their children. However, these important questions are separate from the more specific question of whether women are paid less than men for
doing the same jobs.
* These studies differ slightly from the regression equations presented earlier in this chapter. The outcome of interest, or dependent variable, is binary
in these studies. A participant either has some type of heart-related health problem during the study period or does not have any. As a result,
researchers use a tool called multivariate logistic regression. The basic idea is the same as that of the ordinary least squares models described in this
chapter. Each coefficient expresses the effect of a particular explanatory variable on the dependent variable while holding constant the effects of other
variables in the model. The key difference is that all the variables in the equation affect the likelihood of some event, such as having a heart attack,
occurring during the study period. In this study, for example, workers in the low control group were 1.99 times more likely to suffer “any coronary event”
during the study period than workers in the high control group after controlling for other coronary risk factors.

* The most general formula for calculating a t statistic is as follows:

_b-b
5 SE

where the observed is efficient , boh is then the null hypothesis for that coefficient, and sebis the standard error for the observed coefficient b.
Machine Translated by
Google

CHAPTER 12

Common regression errors


The mandatory warning label

This is one of the most important things to remember when doing research involving regression analysis: try not to kill anyone.
You can even put a little sticky note on your computer monitor: “Don’t kill people with your research.”
Because some very smart people have inadvertently broken that rule.
Beginning in the 1990s, the medical establishment coalesced around the idea that older women should take estrogen
supplements to protect against heart disease, osteoporosis, and other conditions associated with menopause. In
In 2001, about 15 million women were prescribed estrogen in the belief that it would make them healthier. Because? Because
research at the time (using the basic methodology outlined in the last chapter) suggested that this was a sensible medical
strategy. In particular, a longitudinal study of 122,000 women (the Nurses' Health Study) found a negative association between
estrogen supplements and heart attacks. Women taking estrogen suffered one-third as many heart attacks as women not taking
estrogen. This wasn't a couple of teenagers using dad's computer to watch porn and run regression equations. The Nurses'
Health Study is led by Harvard Medical School and the Harvard School of Public Health.

Meanwhile, scientists and doctors offered a medical theory as to why hormone supplements might be beneficial to women's health. A woman's ovaries

produce less estrogen as she ages; if estrogen is important to the body, making up for this deficit in old age could protect a woman's long-term health. Hence the

name of the treatment: hormone replacement therapy. Some researchers even began to suggest that older men should be given an estrogen booster 2 .

And then, as millions of women were prescribed hormone replacement therapy, estrogen was subjected to the most rigorous
form of scientific scrutiny: clinical trials. Rather than searching a large data set, like the Nurses' Health Study, for statistical
associations that may or may not be causal, a clinical trial is a controlled experiment. One sample is given a treatment, such as
hormone replacement; another sample is given a placebo.

Clinical trials showed that women taking estrogen had a higher incidence of heart disease, stroke, blood clots, breast cancer,
and other adverse health outcomes. Estrogen supplements had some benefits, but those benefits were far outweighed by other
risks. Beginning in 2002, physicians were advised not to prescribe estrogen to their elderly patients. The New York Times
Magazine posed a sensitive but socially significant question: How many women died prematurely or suffered strokes or breast
cancer because they were taking a pill their doctors had prescribed to keep them healthy?
The answer: “A reasonable estimate would be tens of thousands.” 3

Regression analysis is the hydrogen bomb of the statistical arsenal. Anyone with a personal computer and a large data set can
be a researcher in their own home or cubicle. What could go wrong? All kinds of things.
Regression analysis provides precise answers to complicated questions. These answers may or may not be accurate. In the
wrong hands, regression analysis will produce misleading or simply wrong results. And, as the estrogen example illustrates,
even in the right hands this powerful statistical tool can lead us accelerating dangerously in the wrong direction. The rest of this
chapter will explain the most common regression “mistakes.” I put “errors” in quotes because, as with all other types of statistical
Machine Translated by
Google

analysis, intelligent people can knowingly exploit these methodological points for nefarious purposes.

Here is a list of the "top seven" of the most common abuses of an otherwise extraordinary tool.

.. . .. .. ....................................................................... *. .....................................................................................
Use regression to analyze a nonlinear relationship. Have you ever read the label?
warning on a hair dryer, the part that warns: Do not use in the bathtub? And you think, "What kind of idiot uses a hair dryer in the
bathtub?" It is an electrical appliance; Do not use electrical appliances near water. They are not designed for that. If regression
analysis had a similar warning label, it would say: Do not use when there is no linear association between the variables you are
analyzing. Remember, a regression coefficient describes the slope of the “line of best fit” to the data; a line that is not straight will
have different slopes in different places. As an example, consider the following hypothetical relationship between the number of
golf lessons I take during a month (an explanatory variable) and my average score on an eighteen-hole round during that month
(the dependent variable). As can be seen from the scatter plot, there is no consistent linear relationship.

Effect of golf lessons on score

Or $100 $200 $300 $400


Money spent on lessons per month

There is a pattern, but it cannot be easily described with a single straight line.
The first few golf lessons seem to make my score drop quickly. There is a negative association between the
lessons and my scores in this section; the slope is negative. More lessons produce lower scores (which is good
in golf).
But then when I get to the point where I'm spending $200-$300 a month on lessons, the lessons don't seem
to have much of an effect. There is no clear association in this segment between additional instruction and my
golf scores; the slope is zero.

And eventually, the lessons seem to backfire. Once I spend $300 per month on instruction, incremental
lessons are associated with higher scores; the slope is positive in this range. (Later in this chapter I will discuss
the distinct possibility that bad golf may be causing the lessons, rather than the other way around.)
Machine Translated by
Google

The most important point here is that we cannot accurately summarize the relationship between lessons
and scores with a single coefficient. The best interpretation of the pattern described above is that golf lessons
have several different linear relationships with my scores. You can see that; a stats package won't do it. If you
plug this data into a regression equation, the computer will give you a single coefficient. That coefficient will not
accurately reflect the true relationship between the variables of interest. The results you get will be the
statistical equivalent of using a hair dryer in the bathtub.

Regression analysis is intended to be used when the relationship between variables is linear. A textbook or
advanced statistics course will guide you through the other central assumptions underlying regression analysis.
As with any other tool, the further one deviates from its intended use, the less effective or even potentially
dangerous it will be.
Correlation does not equal causation. Regression analysis can only demonstrate an association between two
variables. As I mentioned before, we cannot prove with statistics alone that a change in one variable is causing
a change in the other. In fact, a sloppy regression equation can produce a large, statistically significant
association between two variables that have nothing to do with each other. Suppose we were looking for
potential causes for the rising rate of autism in the United States over the past two decades. Our dependent
variable (the outcome we are trying to explain) would be some measure of the incidence of autism per year,
such as the number of cases diagnosed per 1,000 children of a given age. If we included annual per capita
income in China as an explanatory variable, we would almost certainly find a positive and statistically
significant association between rising incomes in China and rising autism rates in the United States over the
past twenty years.
Because? Because both have increased considerably over the same period. However, I highly doubt that a
severe recession in China would reduce the autism rate in the United States. (To be fair, if you saw a strong
relationship between rapid economic growth in China and autism rates in China alone, you might start looking
for some environmental factor related to economic growth, such as industrial pollution, that might explain the
association.)
The type of false association between two variables that I just illustrated is just one example of a more
general phenomenon known as spurious causality.
There are several other ways in which an association between A and B can be misinterpreted.

Reverse causality. A statistical association between A and B does not prove that A causes B. In fact, it is
entirely plausible that B is causing A. I mentioned this possibility earlier in the golf lesson example. Suppose
that when I build a complex model to explain my golf scores, the golf lessons variable is consistently
associated with worse scores. The more lessons I take, the worse I shoot! One explanation is that I have a
really bad golf instructor. A more plausible explanation is that I tend to take more lessons when I play poorly;
the bad golf is prompting more lessons, not the other way around. (There are some simple methodological
solutions to a problem of this nature. For example, you might include golf lessons in one month as an
explanatory variable for golf scores in the next month.)

As noted earlier in this chapter, causality can go in both directions. Suppose you conduct research showing
that states that spend more money on K-12 education have higher economic growth rates than states that
spend less on K-12 education. A positive and significant association between these two variables does not
Machine Translated by
Google

provide any insight into which direction the relationship lies. starts running. Investments in K-12 education
could generate economic growth. On the other hand, states that have strong economies can afford to spend more on
K-12 education, so the strong economy could be causing the spending on education. Or, spending on education could
boost economic growth, making additional spending on education possible; the causality could go both ways.

The point is that we should not use explanatory variables that might be affected by the outcome we are
trying to explain, or else the results will be hopelessly confounded. For example, it would be inappropriate to
use the unemployment rate in a regression equation explaining GDP growth, since unemployment is clearly
affected by the GDP growth rate. Or, to put it another way, a regression analysis finding that reducing
unemployment will boost GDP growth is a silly and meaningless finding, since boosting GDP growth is usually
necessary to reduce unemployment.
We should have reasons to believe that our explanatory variables affect the dependent variable and not the
other way around.
Omitted variable bias. You should be skeptical the next time you see a big headline proclaiming, "Golfers More
Prone to Heart Disease, Cancer and Arthritis!" I wouldn't be surprised if golfers had a higher incidence of all of
those diseases than non-golfers; I also suspect that golf is probably good for your health because it provides
socialization and moderate exercise. How can I reconcile those two statements? Very easily. Any study
attempting to measure the health effects of golf practice must adequately control for age. In general, people
play more golf as they get older, especially when they retire. Any analysis that ignores age as an explanatory
variable will overlook the fact that golfers, on average, will be older than non-golfers. Golf doesn't kill people;
old age is killing people, and they happen to enjoy playing golf while doing it. I suspect that when age is
inserted into the regression analysis as a control variable, we will get a different result. Among people of the
same age, golf may be slightly preventive of serious diseases. That's a pretty big difference.
In this example, age is an important “omitted variable.” When we leave age out of a regression equation
that explains heart disease or some other adverse health outcome, the variable “playing golf” takes on two
explanatory roles instead of just one. It tells us the effect of playing golf on heart disease and it tells us the
effect of advancing age on heart disease (since golfers tend to be older than the rest of the population). In
statistical jargon, we would say that the golf variable is “capturing” the effect of age. The problem is that these
two effects are mixed. At best, our results are a confusing mess. At worst, you mistakenly assume that golf is bad
for your health, when in fact the opposite is likely true.
Regression results will be misleading and inaccurate if the regression equation omits an important explanatory
variable, particularly if other variables in the equation “capture” that effect. Suppose we are trying to explain school quality.
This is an important result to understand: what makes schools good? Our dependent variable, the quantifiable measure of
quality, would probably be test scores. We would almost certainly examine school spending as an explanatory variable in
hopes of quantifying the relationship between spending and test scores. Do schools that spend more get better results? If
school spending were the only explanatory variable, I have no doubt that we would find a large and statistically significant
relationship between spending and test scores. But that finding, and the implication that we can spend money to get to
better schools, is deeply flawed.
There are many potentially significant omitted variables here, but the crucial one is parental education. Well-educated
Machine Translated by
Google

families tend to live in prosperous areas that spend a lot of money on their schools; these families also tend to have
children who do well on tests (and poor families are more likely to have struggling students). If we do not have some
measure of student socioeconomic status as a control variable, our regression results will likely show a large positive
association between school spending and test scores, when in fact those results may be a function of the type of students
who walk through the school doors, not the money spent on the building.

I remember a college professor pointing out that SAT scores are highly correlated with the number of cars a family
owns. He suggested that the SAT was therefore an unfair and inappropriate tool for college admissions. The SAT has its
flaws but the correlation between scores and family cars is not what worries me the most. I'm not too worried about rich
families being able to put their kids through college by buying three more cars. The number of cars in a family's garage is
an indicator of their income, education, and other measures of socioeconomic status. The fact that rich kids score better
on the SAT than poor kids is nothing new. (As noted above, the average SAT Critical Reading score for students from
families with household incomes over $200,000 is 134 points higher than the average score for students from households
with incomes under $20,000.)
4
The biggest concern should be whether the SAT is “trainable” or not. How much can students improve their
scores by taking private SAT prep classes? Wealthy families are clearly better able to send their children to test
preparation classes. Any causal improvement between these classes and SAT scores would favor students from wealthy
families relative to disadvantaged students of equal ability (who presumably could also have improved their scores with a
prep class but never had that opportunity).

Highly correlated explanatory variables (multicollinearity). If a regression equation includes two or more explanatory
variables that are highly correlated with each other, the analysis will not necessarily be able to discern the true relationship
between each of those variables and the outcome we are trying to explain. An example will make this clear. Suppose we
are trying to measure the effect of illegal drug use on SAT scores. Specifically, we have data on whether the participants
in our study have ever used cocaine and also on whether they have ever used heroin. (Presumably we would have many
other control variables as well.) What is the impact of cocaine use on SAT scores, holding other factors constant, including
heroin use? And what is the impact of heroin use on SAT scores, controlling for cocaine use and other factors?
The coefficients on heroin and cocaine use may not be able to tell us that. The methodological challenge is that people
who have used heroin have probably also used cocaine. If we put both variables into the equation, we will have very few
individuals who have used one drug but not the other, leaving us with very little variation in the data with which to calculate
their independent effects.
Think for a moment about the mental images used to explain regression analysis in the last chapter. We split our data
sample into different “rooms” in which each observation is identical except for one variable, which then allows us to isolate
the effect of that variable while controlling for other potential confounders. We may have 692 individuals in our sample
who have used both cocaine and heroin. However, we may have only 3 people who have used cocaine but not heroin and
2 people who have used heroin and not cocaine. Any inference about the independent effect of one drug or another will be
based on these small samples.

We are unlikely to obtain significant coefficients on either the cocaine or heroin variable; we may also obscure the
broader and more important relationship between SAT scores and use of either of these drugs. When two explanatory
Machine Translated by
Google

variables are highly correlated, researchers typically use one or the other in the regression equation, or they may create
some type of composite variable, such as "used cocaine or heroin." For example, when researchers want to control for a
student's general socioeconomic background, they may include variables for both “mother's education” and “father's
education,” as this inclusion provides important information about the home educational environment.

However, if the goal of the regression analysis is to isolate the effect of either the mother's or the father's education, then
putting both variables into the equation is more likely to confuse the issue than to clarify it. The correlation between the
educational attainment of a husband and his wife is so high that we cannot rely on regression analysis to obtain
coefficients that meaningfully isolate the effect of either parent's education (just as it is difficult to separate the impact of
cocaine use from the impact of heroin use).

Extrapolating beyond the data. Regression analysis, like all forms of statistical inference, is designed to give us insights
into the world around us. We look for patterns that are valid for the general population. However, our results are valid only
for a population similar to the sample in which the analysis was performed. In the last chapter, I created a regression
equation to predict 2 of my model's final weight based on a set of 0.29 variables, meaning it did a decent job of explaining
the variation in independents. The R was weighted for a large sample of individuals, all of whom were adults.

So what happens if we use our regression equation to predict the likely weight of a newborn? Let's try it. My daughter
was 21 inches tall when she was born. We will say that his age at birth was zero; he had no education and did no
exercise. She was white and feminine. The regression equation based on the Changing Lives data predicts that her birth
weight should have been negative 19.6 pounds.
(He weighed 8½ pounds.)
The authors of one of the Whitehall studies mentioned in the last chapter were surprisingly explicit in reaching their
narrow conclusion: “Low control in the work environment is associated with an increased risk of future coronary heart
disease among men and women employed in government offices.” (italics added).

Data mining (too many variables). If omitting important variables is a potential problem, then presumably the solution
should be to add as many explanatory variables as possible to a regression equation. No.
Your results may be compromised if you include too many variables, particularly extraneous explanatory variables
without theoretical justification.
For example, one should not design a research strategy based on the following premise: since we do not know what
causes autism, we should include as many potential explanatory variables as possible in the regression equation just to
see what might turn out to be statistically significant. ; Then maybe we'll get some answers. If you put enough junk
variables into a regression equation, it is likely that one of them will reach the threshold of statistical significance simply by
chance. The additional danger is that junk variables are not always easily recognized as such.
Clever researchers can always construct an after-the-fact theory as to why some curious variable that is actually
nonsense turns out to be statistically significant.

To clarify this point, I often do the same coin toss exercise I explained during the discussion on probabilities. In a class of
about forty students, I will have each student flip a coin. Any student who flips tails is eliminated; the rest flip again. In the second
round, those who flip tails are again eliminated. I continue the rounds of flipping until a student has flipped five or six heads in a
Machine Translated by
Google

row.
You might remember some of the silly follow-up questions: “What’s your secret?” Is it on the wrist? Can you teach us to turn
heads all the time? Maybe it's that Harvard sweatshirt you're wearing.
Obviously the series of heads is just luck; all the students have seen what happened. However, this is not necessarily how
the result could or would be interpreted in a scientific context. The probability of getting five heads in a row is 1/32, or 0.03. This
is comfortably below the 0.05 threshold we typically use to reject a null hypothesis. Our null hypothesis in this case is that the
student has no special talent for flipping heads; the lucky series of heads (which is sure to happen to at least one student when I
start with a large group) allows us to reject the null hypothesis and adopt the alternative hypothesis: this student has a special
ability to flip heads. Once he has accomplished this impressive feat, we can study him for clues about his success in the toss: his
throwing form, his athletic training, his extraordinary concentration while the coin is in the air, etc.

And it's all nonsense.


This phenomenon can even affect legitimate research. The accepted convention is to reject a null hypothesis when we
observe something that would happen by chance only 1 in 20 times or less if the null hypothesis were true. Of course, if we run
20 studies, or if we include 20 junk variables in a single regression equation, then on average we will get 1 statistically significant
spurious finding. The New York Times Magazine captured this tension beautifully in a quote from Richard Peto, a medical
statistician and epidemiologist: “Epidemiology is so beautiful and provides such important insight into human life and death, but
there is an incredible amount of junk published.”
6

Even the results of clinical trials, which are typically randomized experiments and therefore the gold standard of medical
research, should be viewed with some skepticism. In 2011, the Wall Street Journal ran a front-page article on what it described
as one of the “dirty little secrets” of medical research: “Most results, including those appearing in top-tier peer-reviewed journals,
can’t be reproduced.”
7
(A peer-reviewed journal is a publication in which other experts in the same field review studies and articles
for methodological soundness before approving them for publication; such publications are considered the gatekeepers of
academic research.) One reason for this “dirty little secret” is the positive publication bias described in Chapter 7. If
researchers and medical journals pay attention to the positive findings and ignore the negative ones, then they can publish
the one study that finds a drug is effective and ignore the nineteen in which it has no effect. . Some clinical trials may also
have small sample sizes (such as those for rare diseases), increasing the chances that random variation in the data will
receive more attention than it deserves. On top of that, researchers may have some conscious or unconscious bias, either
due to a strongly held prior belief or because a positive finding would be better for their career. (No one gets rich or
famous by proving that it doesn't cure cancer.)
For all these reasons, a surprising amount of expert research turns out to be wrong. John Ioannidis, a Greek physician
and epidemiologist, reviewed forty-nine (8) leading medical journals. cited in the medical literature at least three thousand
times. However, about a third of the research was later refuted by subsequent work. (For example, some of the studies he
reviewed promoted estrogen replacement therapy.) Dr. Ioannidis estimates that about half of published scientific papers
will turn out to be incorrect.
9
His research was published in the Journal of the American Medical Association, one of the
journals in which the articles he studied had appeared. This creates a certain mind-blowing irony: if Dr. Ioannidis's
Machine Translated by
Google

research is correct, then there is a good chance that his research is wrong.

Regression analysis remains an incredible statistical tool. (Okay, maybe my description as a “miracle elixir” in the last
chapter was a bit hyperbolic.)
Regression analysis allows us to find key patterns in large data sets, and those patterns are often the key to important
research in medicine and social sciences. Statistics provide us with objective standards for evaluating these patterns.
When used correctly, regression analysis is an important part of the scientific method. Consider this chapter as the
mandatory warning label.
All of the various specific warnings on that label can be summed up in two key lessons. First, designing a good
regression equation (figuring out which variables should be examined and where the data should come from) is more
important than the underlying statistical calculations. This process is known as equation estimation or specifying a good
regression equation. The best researchers are those who can think logically about what variables should be included in a
regression equation, what might be missing, and how the final results can and should be interpreted.

Second, like most other statistical inferences, regression analysis constructs only a circumstantial case. An association
between two variables is like a fingerprint at a crime scene. It points us in the right direction, but it is rarely enough to
condemn. (And sometimes a fingerprint at a crime scene doesn't belong to the perpetrator.) Any regression analysis
needs a theoretical foundation: why are the explanatory variables in the equation? What phenomena from other
disciplines can explain the observed results? For example, why do we think that wearing purple shoes would improve
performance on the math portion of the SAT or that eating popcorn can help prevent prostate cancer? The results should
be replicated, or at least be consistent with other findings.

Even a miracle elixir won't work if it's not taken as directed.

* There are more sophisticated methods that can be used to adapt regression analysis for use with nonlinear data. However, before using those
tools, it is necessary to understand why using the standard ordinary least squares method with non-linear data will give you a meaningless
result.
Machine Translated by
Google

CHAPTER 13

Program Evaluation Will going


to Harvard change your life?

Brilliant social scientists are not brilliant because they can do complex calculations in their heads or because
they make more money on Jeopardy than less brilliant researchers (although both of these feats may be true).
Brilliant researchers (those who appreciably change our knowledge of the world) are often individuals or teams
who find creative ways to conduct "controlled" experiments. To measure the effect of any treatment or
intervention, we need something to compare it to. How would going to Harvard affect your life?
Well, to answer that question, we need to know what happens to you after you go to Harvard and what
happens to you after you don't go to Harvard. Obviously we can't have data on both. However, clever
researchers find ways to compare some treatments (e.g., going to Harvard) with the counterfactual, which is
what would have happened in the absence of that treatment.
To illustrate this point, let's consider a seemingly simple question: Does putting more police officers on the
streets deter crime? This is a socially important issue, as crime imposes enormous costs on society. If
increased police presence reduces crime, either through deterrence or by capturing and imprisoning bad
actors, then investments in additional police officers could yield large benefits. On the other hand, police
officers are relatively expensive; if they have little or no impact on reducing crime, then society could make
better use of its resources elsewhere (perhaps with investments in crime-fighting technology, such as
surveillance cameras).
The challenge is that our seemingly simple question – what is the causal effect of more police officers on
crime? – is very difficult to answer. By this point in the book, I should acknowledge that we cannot answer this
question simply by examining whether jurisdictions with high numbers of police officers per capita have lower
crime rates. Zurich is not Los Angeles. Even a comparison of large American cities will be deeply flawed; Los
Angeles, New York, Houston, Miami, Detroit and Chicago are all different places with different demographic
and crime challenges.
Our usual approach would be to try to specify a regression equation that controls for these differences.
Unfortunately, even multiple regression analysis is not

we will save ourselves here. If we try to explain crime rates (our dependent variable) using police officers per
capita as the explanatory variable (along with other controls), we will have a serious reverse causality problem.
We have strong theoretical reason to believe that putting more police officers on the streets will reduce crime,
but it is also possible that crime may "cause" police officers, in the sense that cities that experience crime
waves will hire more police officers. We could easily find a positive but misleading association between crime
and police: places with the most police officers have the worst crime problems. Of course, places with a lot of
doctors also tend to have the highest concentration of sick people. These doctors do not make people sick;
Machine Translated by
Google

they are placed where they are most needed (and at the same time the sick are moved to places where they
can receive adequate medical care). I suspect there are a disproportionate number of oncologists and
cardiologists in Florida; banishing them from the state will not make the retirement population any healthier.
Welcome to program evaluation, which is the process by which we seek to measure the causal effect of
some intervention, from a new cancer drug to a job placement program for high school dropouts. Or put more
police on the streets. The intervention we are interested in is often called a “treatment,” although that word is
used more broadly in a statistical context than in ordinary language. A treatment can be literal treatment, such
as some type of medical intervention, or it can be something like attending college or receiving job training
upon release from prison. The point is that we are seeking to isolate the effect of that single factor; ideally, we
would want to know how the group receiving that treatment fares compared to some other group whose
members are identical in every other respect except the treatment.

Program evaluation offers a set of tools to isolate treatment effect when cause and effect are elusive. That's
how Jonathan Klick and Alexander Tabarrok, researchers at the University of Pennsylvania and George Mason
University, respectively, studied how putting more police officers on the streets affects the crime rate. His
research strategy made use of the terrorist alert system. Specifically, Washington, DC responds to days of
“high alert” for terrorism by placing more officers in certain areas of the city, as the capital is a natural target for
terrorism. We can assume that there is no relationship between street crime and the terrorist threat, so this
increase in police presence in DC is unrelated to the conventional crime rate, or is “exogenous.” The
researchers' most valuable insight was to acknowledge the natural experiment here: What happens to common
crime on days of "high alert" for terrorism?
The answer: The number of crimes committed when the terrorist threat was Orange (high alert and more
police) was approximately 7 percent lower than when the
The terrorist threat level was Yellow (high alert but no additional police precautions). The authors also found
that the decline in crime was most pronounced in the police district that receives the most police attention on
high-alert days (because it includes the White House, the Capitol and the National Mall). The important thing is
that we can answer difficult but socially significant questions; we just have to be smart about it. These are
some of the most common approaches to isolating the effect of a treatment.

Randomized controlled experiments. The simplest way to create a treatment and control group is to (hopefully)
create a treatment and control group. There are two major challenges to this approach. First, there are many
types of experiments that we cannot perform on people. This limitation (I hope) will not go away anytime soon.
As a result, we can only conduct controlled experiments on human subjects when there is reason to believe
that the treatment effect has a potentially positive outcome. Often this is not the case (for example,
“treatments” like experimenting with drugs or dropping out of high school), which is why we need the strategies
presented at the end of the chapter.
Secondly, there is much more variation among people than among laboratory rats. The effect of the
treatment we are testing could easily be confounded by other variations in the treatment and control groups –
there will surely be tall people, short people, sick people, healthy people, men, women, criminals, alcoholics,
investment bankers, and so on. How can we ensure that differences between these other characteristics do not
Machine Translated by
Google

spoil the results? I have good news: this is one of those rare cases in life where the best approach involves the
least work!
The optimal way to create any treatment and control group is to randomly distribute study participants between
the two groups. The nice thing about randomization is that it will usually distribute non-treatment variables more
or less evenly between the two groups: both characteristics that are obvious, such as sex, race, age, and
education, and unobservable characteristics that might otherwise show up. ruin the results.
Think about it: If we have 1000 women in our prospective sample, when we randomly split the sample into
two groups, the most likely outcome is that 500 women will end up in each. Obviously we can't expect that
exact split, but again probability is our friend. The probability of a group getting a disproportionate number of
women (or a disproportionate number of individuals with any other characteristic) is low. For example, if we
have a sample of 1,000 people, half of whom are women, there is less than a 1 percent chance of having fewer
than 450 women in one group or the other. Obviously, the larger the samples, the more effective randomization
will be in creating two
similar groups.
Medical trials typically aim to conduct randomized, controlled experiments. Ideally, these clinical trials are
double-blind, meaning that neither the patient nor the doctor knows who is receiving the treatment and who is
receiving a placebo.
Obviously, this is impossible with treatments such as surgical procedures (hopefully the cardiac surgeon knows
which patients will undergo bypass surgery). However, even with surgical procedures, it is still possible to
prevent patients from knowing whether they are in the treatment or control group. One of my favorite studies
involved evaluating a certain type of knee surgery for pain relief. The treatment group received surgery. The
control group underwent "sham" surgery in which the surgeon made three small incisions in the knee and
"pretended to operate."
*
It turned out that real surgery was no more effective than sham
surgery in relieving knee pain. 1

Randomized trials can be used to test some interesting phenomena. For example, do prayers offered by
strangers improve post-surgical outcomes? Reasonable people have widely varying opinions on religion, but a
study published in the American Heart Journal conducted a controlled study that examined whether patients
recovering from heart bypass surgery would have fewer postoperative complications if a large group of
strangers prayed for their safe and speedy recovery.
2
The study involved 1,800 patients and members of three religious congregations
across the country. The patients, all of whom received coronary bypass surgery, were divided into three
groups: one group was not prayed for; one group was prayed for and told so; the third group was prayed for,
but participants in that group were told they might or might not receive prayers (thus controlling for the placebo
effect of prayer). Meanwhile, members of religious congregations were asked to offer prayers for specific
patients by first name and the first initial of their last name (e.g., Charlie W.).
Parishioners were given freedom to pray, as long as the prayer included the phrase “for a successful surgery
with a speedy, healthy and uncomplicated recovery.”
Machine Translated by
Google

AND? Could prayer be the cost-effective solution to America's health care challenges? Probably not. The
researchers found no difference in the rate of complications within thirty days after surgery between those who
were offered prayers compared to those who were not offered prayers. Critics of the study pointed to a
possible omitted variable: sentences from other sources. As the New York Times summarized: “Experts said
the study could not overcome perhaps the biggest obstacle to studying prayer: the unknown amount of prayer
each person received daily from friends, relatives and congregations around the world who offered prayer for the sick and dying.”

Experimenting on humans can get you arrested or possibly brought before an international criminal court. You
should be aware of this.
However, there is still room in the social sciences for randomized, controlled experiments with “human
subjects.” A famous and influential experiment is the Tennessee Project STAR experiment, which tested the
effect of smaller classes on student learning. The relationship between class size and learning is enormously
important. Nations around the world are struggling to improve educational outcomes. If smaller classes
promote more effective learning, ceteris paribus, then society should invest in hiring more teachers to reduce
class sizes. At the same time, hiring teachers is expensive; if students in smaller classes do better for reasons
unrelated to class size, then we could end up wasting a huge amount of money.
The relationship between class size and student achievement is surprisingly difficult to study. Schools with
small class sizes generally have greater resources, which means that both students and teachers are likely to
be different from students and teachers at schools with larger class sizes. And within schools, smaller classes
tend to be smaller for a reason. A principal may assign difficult students to a small class, in which case we
might find a spurious negative association between smaller classes and student achievement. Or veteran
teachers may choose to teach small classes, in which case the benefit of small classes may come from the
teachers who choose to teach them and not from the lower student-teacher ratio.
Beginning in 1985, Tennessee's Project STAR conducted a controlled experiment to test the effects of smaller
classes. 3 (Lamar Alexander was governor of Tennessee at the time; he later went on to become secretary of
education under President George H.W. Bush.) In kindergarten, students from seventy-nine different schools
were randomly assigned to a small class (ages 13 to 17). students), a regular class (22 to 25 students), or a
regular class with a regular teacher and a teacher's assistant. Teachers were also randomly assigned to
different classrooms. Students remained in the class type to which they were randomly assigned until third
grade. Various realities of life eroded randomization. Some students logged into the system in the middle of the
experiment; others left. Some students were moved from one class to another for disciplinary reasons; some
parents successfully lobbied to have students moved to smaller classes. Etc.
Still, Project STAR remains the only randomized test of the effects of smaller classes. The results turned
out to be statistically and socially significant. Overall, students in small classes scored 0.15 standard deviations
better on standardized tests than students in regular-sized classes; black students in small classes made gains that
were twice as large. Now the bad news. The Project STAR experiment cost approximately $12 million. The study on the
effect of prayer on post-surgical complications cost $2.4 million. The best studies are like anything else: they cost a lot of
money.

Natural experiment. Not everyone has millions of dollars to set up a large randomized trial. A cheaper alternative is to
exploit a natural experiment, which occurs when random circumstances somehow create something resembling a
Machine Translated by
Google

randomized, controlled experiment. This was the case with our Washington, DC, police example at the beginning of the
chapter.
Sometimes life creates a treatment and control group by accident; when that happens, researchers are eager to take
advantage of the results. Consider the surprising but complicated link between education and longevity. People who
receive more education tend to live longer, even after controlling for things like income and access to health care. As the
New York Times noted: “The one social factor that researchers agree is consistently linked to longer lives in every country
where it has been studied is education. It is more important than race; it erases any income effects.”
4
But so far, that's just a correlation. Does more education, ceteris
paribus, produce better health? If you think of education itself as the “treatment,” will getting more education make you live
longer?

This would seem to be an almost impossible question to study, since people who choose to receive more education
are different from those who do not. The difference between high school graduates and college graduates is not just four
years of schooling. There could easily be some unobservable characteristics shared by people who pursue an education
that also explain their longer life expectancy. If that is the case, offering more education to those who would have chosen
less education will not actually improve their health. Health improvement would not be a function of incremental education;
it would be a function of the type of people who pursue that incremental education.

We cannot conduct a randomized experiment to solve this puzzle, because that would imply that some participants
would leave school earlier than they would like. (Try explaining to someone that they can't go to college, ever, because
they're in the control group.) The only possible test of a causal effect of education on longevity would be some kind of
experiment that forced a large segment of the population to stay in school longer than its members would choose. This is
at least morally acceptable, since we expect a positive effect from the treatment. Still, we cannot force children to stay in
school; that is not the

American style.
Ah, but it is. Every state has some type of minimum schooling law, and at different times in history those
laws have changed. That kind of exogenous change in educational attainment (meaning it is not caused by the
individuals being studied) is exactly the kind of thing that makes researchers swoon with excitement. Adriana
Lleras-Muney, a graduate student at Columbia, saw the potential for research in the fact that different states
have changed their minimum schooling laws at different times. He went back in history and studied the
relationship between when states changed their minimum schooling laws and subsequent changes in life
expectancy in those states (reviewing lots of census data). There was still a methodological challenge; if
residents of a state live longer after the state raises its minimum schooling law, we cannot attribute the
longevity to the additional schooling. Life expectancy generally increases with time. People lived longer in 1900
than in 1850, no matter what the states did.

However, Lleras-Muney had a natural control: states that did not change their minimum schooling laws. Their work
amounts to a giant laboratory experiment in which Illinois residents are forced to stay in school for seven years, while their
Machine Translated by
Google

Indiana neighbors can drop out after six years. The difference is that this controlled experiment was made possible by a
historical accident, hence the term "natural experiment."

What happened? The life expectancy of adults who reached age thirty-five was extended by an additional year and a
half simply by attending one more year of school. 5 Lleras-Muney's results have been replicated in other countries where
variations in compulsory schooling laws have created similar natural experiments. Some skepticism is necessary. We do
not yet understand the mechanism by which additional schooling leads to longer lives.

Non-equivalent control. Sometimes the best option available for studying the effect of a treatment is to create non-
randomized treatment and control groups. Our hope/expectation is that the two groups will be broadly similar even though
circumstances have not allowed us the statistical luxury of randomization. The good news is that we have a treatment
group and a control group. The bad news is that any non-random assignment creates at least the possibility of bias. There
may be unobserved differences between treatment and control groups related to how participants are assigned to one
group or the other. Hence the name “non-equivalent control”.

A nonequivalent control group can still be a very useful tool. Let us reflect on the question posed in the title of this
chapter: Is there a significant advantage in life for

Attending a highly selective college or university? Obviously, Harvard, Princeton and Dartmouth graduates around the
world are doing very well. On average, they earn more money and have broader life opportunities than students who
attend less selective institutions. (A 2008 study by [Link] found that the median salary for Dartmouth graduates
with ten to twenty years of work experience was $134,000, the highest of any undergraduate institution; Princeton ranked
6
second with a median of $131,000.) As I hope you realize by now
heights, these impressive numbers tell us absolutely nothing about the value of a Dartmouth or Princeton education.
Students who attend Dartmouth and Princeton are talented when they apply; that's why they get accepted. They would
probably do well in life no matter what college they went to.
What we don't know is the therapeutic effect of attending a place like Harvard or Yale. Do graduates of these elite
institutions do well in life because they were so talented when they entered campus? Or do these colleges and universities
add value by accepting talented individuals and making them even more productive? Or both?

We cannot perform a randomized experiment to answer this question. Few high school students would agree to be
randomly assigned to a college; Harvard and Dartmouth would not be particularly interested in accepting students who
were randomly assigned to them, either. We seem to be left without any mechanism to test the value of the treatment
effect. Cunning to the rescue! economists *

Stacy Dale and Alan Krueger found a way to answer this question by taking advantage of the fact that many students
apply to multiple colleges. Some of those students are accepted to a highly selective school and choose to attend that
school; others are accepted to a highly selective school but choose to attend a less selective college or university. Bingo!
We now have a treatment group (those students who attended highly selective colleges and universities) and a
Machine Translated by
Google

nonequivalent control group (those students who were talented enough to be accepted into such a school but chose to
attend a less selective institution).

Dale and Krueger studied longitudinal data on the income of both groups. This isn't a perfect apples-to-apples
comparison, and income is clearly not the only life outcome that matters, but its findings should ease the anxieties of
overburdened high school students and their parents. Students who attended more selective colleges scored about the
same as students of seemingly similar ability who attended less selective schools. The only exception was students from
low-income families, who earned more if they attended a selective college or university. The Dale and Krueger approach
is an elegant way to resolve the problems of treatment effect (spending four years at an elite institution) and
selection effect (the most talented students are admitted to such institutions). In a summary of the research for
the New York Times, Alan Krueger indirectly answered the question posed in the title of this chapter:
“Recognize that your own drive, ambition and talents will determine your success more than the name of the
college on your diploma. " 8

Difference in differences. One of the best ways to observe cause and effect is to do something and then see
what happens. After all, this is how babies and toddlers (and sometimes adults) learn about the world. My
children learned very quickly that if they threw pieces of food around the kitchen (cause), the dog would
eagerly run after them (effect). Presumably, the same power of observation can help inform the rest of life. If
we cut taxes and the economy improves, then the tax cuts must have been responsible.
Maybe. The huge potential danger of this approach is that life tends to be more complex than throwing
chicken nuggets around the kitchen. Yes, we may have cut taxes at a specific time, but there were other
“interventions” unfolding during roughly the same period: more women were going to college, the Internet and
other technological innovations were raising the productivity of American workers, the Chinese currency was
undervalued, the Chicago Cubs fired their general manager, and so on. What happened after the tax cut
cannot be attributed solely to the tax cut. The challenge with any kind of “before and after” analysis is that just
because one thing follows another doesn’t mean there is a causal relationship between the two.
A “difference in differences” approach can help us identify the effects of some intervention by doing two
things. First, we look at “before” and “after” data for any group or jurisdiction that received the treatment, such
as unemployment figures for a county that implemented a job training program. Second, we compared that
data to unemployment figures during the same period for a similar county that did not implement any such
program.
The important assumption is that the two groups used for analysis are largely comparable except for
treatment; as a result, any significant differences in outcomes between the two groups can be attributed to the
program or policy being evaluated. For example, suppose a county in Illinois implements a job training program
to combat high unemployment. Over the next two years, the unemployment rate continues to rise. Does that
make the program a failure? Who knows?

Effect of job training on unemployment in County A


Machine Translated by
Google

Time

Other broader economic forces may be at play, including the possibility of a prolonged economic crisis. A
difference-in-differences approach would compare the change in the unemployment rate over time in the
county we are evaluating to the unemployment rate in a neighboring county with no job training program; the
two counties should be similar in all other important ways: industry mix, demographics, etc. How does the
unemployment rate change over time in the county with the new job training program relative to the county that
did not implement such a program? We can reasonably infer the treatment effect of the program by comparing
changes in the two counties over the study period: the "difference in differences." The other county in this study
effectively acts as a control group, allowing us to leverage data collected before and after the intervention. If
the control group is good, it will be exposed to the same broader forces as our treatment group. The difference-
in-differences approach can be particularly illuminating when the treatment initially appears ineffective
(unemployment is higher after the program is implemented than before), however, the control group shows us
that the trend would have been even worse in the absence of the intervention. .

Effect of job training on unemployment in County A, with County B as a comparison


Machine Translated by
Google

County B

Job training begins


in County A
'Time

Discontinuity analysis. One way to create a treatment and control group is to compare the results of a group
that just barely qualified for an intervention or treatment with the results of a group that just missed the eligibility
cutoff and did not receive the treatment. Those individuals who fall just above and below some arbitrary cutoff,
such as a test score or a minimum family income, will be nearly identical in many important respects; the fact
that one group received the treatment and the other did not is essentially arbitrary. As a result, we can
compare their results in ways that provide meaningful insights into the effectiveness of the relevant
intervention.

Suppose a school district requires summer school for struggling students. The district would like to know if
the summer program has any long-term academic value. As always, a simple comparison between students
who attend summer school and those who do not would be worse than useless. Students who attend summer
school are there because they are struggling. Even if the summer school program is highly effective,
participating students will likely fare worse in the long run than students who were not required to attend
summer school. What we want to know is how struggling students perform after attending summer school
compared to how they would have done if they had not attended summer school. Yes, we could do some sort
of controlled experiment where struggling students are randomly selected to attend summer school or not, but
that would involve denying the control group access to a program that we think would be helpful.

Instead, treatment and control groups are created by comparing those students who just missed the
threshold for summer school with those who just missed it. Think about it: students who fail a midterm are
appreciably different from students who do not fail the midterm. But students who earn a 59 percent (a failing
grade) are not appreciably different from those students who earn a 60 percent (a passing grade). If those who fail the
midterm are enrolled in some treatment, such as mandatory tutoring for the final exam, then we would have a reasonable
treatment and control group if we compare the final exam scores of those who barely failed the midterm (and received
tutoring). with the scores of those who barely passed the midterm exam (and did not receive tutoring).
This approach was used to determine the effectiveness of incarceration of juvenile offenders as a deterrent to future
crime. Obviously, this type of analysis cannot simply compare the recidivism rates of juvenile offenders who are
incarcerated with the recidivism rates of juvenile offenders who received lighter sentences. Juvenile offenders who are
Machine Translated by
Google

sent to prison tend to commit more serious crimes than juvenile offenders who receive lighter sentences; that's why they
go to prison. Nor can we create a treatment and control group by randomly distributing prison sentences (unless you want
to risk twenty-five years in the big house the next time you make an illegal right turn on red). Randi Hjalmarsson, now a
researcher at the University of London, took advantage of the rigid sentencing guidelines for juvenile offenders in
Washington state to better understand the causal effect of a prison sentence on future criminal behavior.
Specifically, he compared the recidivism rate of those juvenile offenders who were “barely” sentenced to prison with the
recidivism rate of those youth who “barely” got a pass (which usually involved a fine or probation).
9

Washington's criminal justice system creates a grid for each convicted offender that is used to administer a sentence.
The x-axis measures the crimes previously assigned to the offender. For example, each prior felony counts as one point;
each prior misdemeanor counts as a quarter point. The point total is rounded down to a whole number (which will be
important in a moment). Meanwhile, the y-axis measures the severity of the current crime on a scale ranging from E (least
serious) to A+ (most serious). A convicted juvenile's sentence is calculated literally by finding the appropriate box on the
grid: an offender with two prior offense points who commits a Class B felony will receive fifteen to thirty-six months in a
juvenile jail. A convicted felon with only one point for prior offenses who commits the same crime will not be sent to prison.
This discontinuity is what motivated the research strategy. Hjalmarsson compared the outcomes of convicted offenders
who were just above and below the threshold for a prison sentence. As he explains in the article, “if there are two
individuals with a current offense class of C+ and [prior] adjudication scores of 2¾ and 3, then only the latter individual will
be sentenced to state prison.”
For investigative purposes, those two individuals are essentially the same, until one of them goes to jail.
And at that point, their behavior seems to diverge markedly. Juvenile offenders who go to prison are much less
likely to be convicted of another crime (after they leave prison).
We care about what works. This is true in medicine, in economics, in business, in criminal justice... in
everything. Causality is a tough nut to crack, however, even in cases where cause and effect seem surprisingly
obvious. To understand the true impact of a treatment, we need to know the “counterfactual,” which is what
would have happened in the absence of that treatment or intervention. The counterfactual is often difficult or
impossible to observe. Consider a non-statistical example: Did the US invasion of Iraq make the United States
safer?
There is only one intellectually honest answer: we will never know. The reason we will never know is that
we do not know – and cannot know – what would have happened if the United States
The United States would not have invaded Iraq. It is true that the United States did not find any weapons of
mass destruction. But it's possible that the day after the United States failed to invade Iraq, Saddam Hussein
might have gotten into the shower and said to himself: "I could really use a hydrogen bomb." I wonder if the
North Koreans will sell me one? After that, who knows?
Of course, it's also possible that Saddam Hussein had stepped into that very shower the day after the US
failed to invade Iraq and said to himself, "I could really use some..." at which point he slipped on a bar of soap.
He hit his head on a marble ornament and died. In that case, the world would have been rid of Saddam
Hussein without the enormous costs associated with an American invasion. Who knows what would have
happened?
Machine Translated by
Google

The purpose of any program evaluation is to provide some kind of counterfactual against which a treatment
or intervention can be measured. In the case of a randomized controlled experiment, the control group is the
counterfactual. In cases where a controlled experiment is impractical or immoral, we need to find some other
way to approximate the counterfactual. Our understanding of the world depends on finding clever ways to do
so.

* Participants knew they were participating in a clinical trial and could receive the sham surgery.

* Researchers love to use the word "explode." It has a specific meaning in terms of taking advantage of some data-related opportunity. For example,
when researchers encounter some natural experiment that creates a treatment and control group, they will describe how they plan to "exploit variation in
the data." † Potential for bias exists here. Both groups of students are talented enough to get into a highly selective school. However, one group of
students chose to go to the school and the other group did not. The group of students who chose to attend a less selective school may be less
motivated, less hard-working, or different in other ways that we cannot observe. If Dale and Krueger had found that students who attended a highly
selective school had higher lifetime earnings than students who were accepted to that school but

Instead, we went to a less selective college, and we still couldn't be sure if the difference was due to the selective school or the type of student
who chose to attend said school when given the option. However, this potential bias is of little importance in Dale and Krueger's study due to its
direction.
Dale and Krueger find that students who attended highly selective schools did not earn significantly more in life than students who were
accepted but went elsewhere even though students who declined to attend a highly selective school may have had attributes that led them to
earn money. less in life apart from his education. If anything, the bias here causes the findings to exaggerate the pecuniary benefits of attending
a highly selective college, which turn out to be insubstantial anyway.
Machine Translated by
Google

Conclusion
Five questions that statistics
can help answer

Not long ago, it was much more difficult to gather information and much more
expensive to analyze. Imagine studying the data from a million credit card transactions in an era
(just a few decades ago) when there were only paper receipts and no personal computers to
analyze the accumulated data.
During the Great Depression, there were no official statistics with which to measure the depth of the
economic problems. The government collected no official data on either gross domestic product
(GDP) or unemployment, meaning politicians were trying to do the economic equivalent of
navigating through a forest without a compass. Herbert Hoover declared the Great Depression over
in 1930, based on inaccurate and outdated data available. In his State of the Union address, he told
the country that two and a half million Americans were out of work. In fact, five million Americans
were unemployed and unemployment was increasing by 100,000 people every week. As James
Surowiecki recently observed in The New Yorker, “Washington was making policy in the dark.”
1

We are now inundated with data. For the most part, that's a good thing. The statistical tools
presented in this book can be used to address some of our most important societal challenges. In
that sense, I thought it would be appropriate to end the book with questions, not answers. As we try
to digest and analyze staggering amounts of information, here are five important (and admittedly
random) questions whose socially meaningful answers will involve many of the tools presented in
this book.

WHAT IS THE FUTURE OF FOOTBALL?

In 2009, Malcolm Gladwell posed a question in a New Yorker article that at first struck me as
unnecessarily sensationalist and provocative: How different are dogfighting and football? 2 The
connection between the two activities arose from the fact that quarterback Michael Vick, who had
served time in prison for his

involvement in a dogfighting ring, had been reinstated in the National Football League just as information was
beginning to emerge that football-related head trauma may be associated with depression, memory loss,
Machine Translated by
Google

dementia and other neurological problems later in life. Gladwell's central premise was that both professional
football and dogfighting are inherently devastating to the participants. By the end of the article, I was
convinced that I had raised an intriguing point.

This is what we know. There is growing evidence that concussions and other brain injuries associated with
football can cause serious and permanent neurological damage. (Similar phenomena have been observed in
boxers and hockey players.) Many former NFL standouts have publicly shared their post-football battles with
depression, memory loss and dementia. Perhaps the most poignant was Dave Duerson, the former Chicago
Bears safety and Super Bowl winner, who committed suicide by shooting himself in the chest; he left explicit
instructions to his family to study his brain after his death.

In a telephone survey of 1,000 randomly selected former NFL players who had played at least three years
in the league, 6.1 percent of former players over the age of fifty reported that they had been diagnosed with
“dementia, Alzheimer’s disease, or other memory-related illness.” disease." That's five times the national
average for that age group. For younger players, the diagnosis rate was nineteen times higher than the
national average. Hundreds of former NFL players have sued both the league and football helmet
manufacturers for allegedly hiding information about the dangers of head trauma.
One of the researchers studying the impacts of brain trauma is Ann McKee, who heads the
neuropathology laboratory at the Veterans Affairs Hospital in Bedford, Massachusetts. (Incidentally, McKee
also does the neuropathology work for the Framingham Heart Study.) Dr. McKee has documented the buildup
of abnormal proteins called tau in the brains of athletes who have suffered brain trauma, such as boxers and
football players. This leads to a condition known as chronic traumatic encephalopathy or CTE, which is a
progressive neurological disorder that has many of the same symptoms as Alzheimer's.

Meanwhile, other researchers have been documenting the connection between football and brain trauma.
Kevin Guskiewicz, who heads the Sports Concussion Research Program at the University of North Carolina,
has installed sensors inside the helmets of North Carolina football players to record the force and nature of
blows to the head. According to their data, players routinely receive blows to the head with a force equivalent
to hitting the windshield in a car accident at twenty-five miles per hour.

This is what we don't know. Has any evidence of brain injuries been discovered so far?

Representative of the long-term neurological risks faced by all professional football players? Or could it simply
be a “cluster” of adverse outcomes that constitutes a statistical aberration? Even if it turns out that football
players face significantly higher risks of neurological disorders in the future, we would still need to investigate
causality. Could the type of men who play football (and boxing and hockey) be prone to these kinds of
problems? Is it possible that other factors, such as steroid use, could contribute to neurological problems later
in life?

If the accumulating evidence suggests a clear causal link between playing football and long-term brain
Machine Translated by
Google

injury, players (and parents of younger players), coaches, attorneys, NFL officials, and perhaps even parents
of younger players will have to address a paramount question. Government regulators: Is there any way to
play soccer that reduces most or all of the risk of head trauma? If not, then what? This is the point behind
Malcolm Gladwell's comparison between football and dogfighting. He explains that dog fighting is abhorrent to
the public because the dog owner willingly subjects his dog to a competition that ends in suffering and
destruction. "And why?" he asks. “For the entertainment of an audience and the possibility of earning a
payday. In the 19th century, dog fighting was widely accepted by the American public. But we no longer
consider that type of transaction to be morally acceptable in a sport.”
Almost all of the types of statistical analysis described in this book are currently used to determine whether
professional football as we know it now has a future.

WHAT (IF ANYTHING) IS CAUSING THE DRAMATIC INCREASE IN AUTISM INCIDENCE?

In 2012, the Centers for Disease Control reported that 1 in 88 American children had been diagnosed with an
autism spectrum disorder (based on data from 4
2008). The diagnosis rate had increased from 1 in 110 in 2006 to 1 in 150 in 2002, or almost double in less
than a decade. Autism spectrum disorders (ASD) are a group of developmental disabilities characterized by
atypical development in socialization, communication, and behavior. The "spectrum" indicates that autism
encompasses a wide range of behaviorally defined conditions. 5 Boys are five times more likely to be
diagnosed with an ASD than girls (meaning the incidence in boys is even higher than 1 in 88).
The first intriguing statistical question is whether we are experiencing an autism epidemic, a “diagnosis epidemic,” or some combination of the

two. 6 In previous decades, children with an autism spectrum disorder had two. symptoms that may not have been diagnosed, or their

developmental challenges

could have been described more generally as a “learning disability.” Doctors, parents and teachers are now much
more aware of the symptoms of ASD, which naturally leads to more diagnoses regardless of whether the incidence
of autism is increasing or not.

In any case, the surprisingly high incidence of ASD represents a serious challenge for families, schools and the
rest of society. The average lifetime cost of treating an autism spectrum disorder for a single individual is $3.5
million.
7
Despite what is clearly an epidemic, we know surprisingly little about the causes of this
condition. Thomas Insel, director of the National Institute of Mental Health, said: “Is it cell phones? Ultrasound? Diet
sodas? Every parent has a theory. At this point, we just don't know.” 8 What is different or unique about the
lives and backgrounds of children with ASD? What are the most significant physiological differences between
children with and without ASD? Is the incidence of ASD different between countries?

If so, why? Traditional statistical detective work is finding clues.


A recent study by researchers at the University of California, Davis, identified ten places in California with autism
rates double the rates of surrounding areas; each of
Machine Translated by
Google

9
Autism clusters are a neighborhood with a concentration of white, highly educated parents. Is that a clue or a
coincidence?
Or could it reflect that relatively privileged families are more likely to be diagnosed with autism spectrum disorder?
The same researchers are also conducting a study in which they will collect dust samples from the homes of 1,300
families with an autistic child to test for chemicals or other environmental contaminants that may play a causal role.
Meanwhile, other researchers have identified what appears to be a genetic inheritance.
10 component of autism by studying ASD among identical and fraternal twins.
The likelihood of two children from the same family having an ASD is higher among identical twins (who share the
same genetic makeup) than among fraternal twins (whose genetic similarity is the same as that of normal siblings).
This finding does not rule out important environmental factors, or perhaps the interaction between environmental
and genetic factors. After all, heart disease has a significant genetic component, but clearly smoking, diet, exercise,
and many other environmental and behavioral factors matter, too.

One of the most important contributions of statistical analysis so far has been to debunk false causes, many of
which have arisen because of a confusion between correlation and causation. An autism spectrum disorder usually
appears suddenly between a child's first and second birthday. This has led to a widespread belief that childhood
vaccines, particularly the MMR vaccine,

Measles, mumps, and rubella (MMR) are causing the increasing incidence of autism.
Dan Burton, a member of Congress from Indiana, told the New York Times: “My grandson received nine injections in one day, seven of which

contained thimerosal, which, as you know, is 50 percent mercury, and shortly thereafter he became autistic. .”

Scientists have strongly refuted the false association between thimerosal and ASD. Autism rates did not
decline when thimerosal was removed from the MMR vaccine, nor are autism rates lower in countries that
never used the vaccine. However, the false connection persists, causing some parents to refuse to vaccinate
their children. Ironically, this offers no protection against autism and at the same time puts children at risk for
other serious diseases (and contributes to the spread of those diseases throughout the population).

Autism poses one of the greatest medical and social challenges of our time. We understand very little
about the disorder in relation to its enormous (and possibly growing) impact on our collective well-being.
Researchers are using all the tools in this book (and many more) to change that.

HOW CAN WE IDENTIFY AND REWARD GOOD TEACHERS AND SCHOOLS?

We need good schools. And we need good teachers to have good schools. It therefore logically follows that
we should reward good teachers and good schools, while firing bad teachers and closing bad schools.
How exactly do we do that?
Test results give us an objective measure of student performance. However, we know that some students
will do much better on standardized tests than others for reasons that have nothing to do with what happens
inside a classroom or a school. The seemingly simple solution is to evaluate schools and teachers based on
the progress their students make over a period of time. What did students know when they started in a
Machine Translated by
Google

particular classroom with a particular teacher?


What did they know a year later? The difference is the “added value” in that classroom.

We can even use statistics to get a more refined sense of this added value by taking into account
demographic characteristics of students in a given classroom, such as race, income, and performance on
other tests (which can be a measure of aptitude). . If a teacher makes significant progress with students who
have typically struggled in the past, then he or she can be considered highly effective.

Voila! We can now assess teacher quality with statistical precision. And good schools, of course, are only
those that are full of effective teachers.

How do these practical statistical evaluations work in practice? In 2012, New York City took the step of
publishing ratings for all 18,000 public school teachers based on a “value-added assessment” that measured
progress in their performance.
12
students' test scores taking into account various student characteristics. The Los Angeles Times published a similar
set of rankings for Los Angeles teachers in 2010.

In both New York and Los Angeles, the reaction has been loud and mixed. Arne Duncan, the secretary of
U.S. education has generally been supportive of these types of value-added assessments. They provide information
where none existed before. After the Los Angeles data was released, Secretary Duncan told the New York Times: “Silence
is not an option.” The Obama administration has provided financial incentives for states to develop value-added indicators
for paying and promoting teachers. Proponents of these evaluation measures rightly point out that they represent a huge
potential improvement over systems in which all teachers are paid according to a uniform salary plan that gives no weight
to any measure of classroom performance.

On the other hand, many experts have warned that this type of teacher evaluation data has large margins of
error and can produce misleading results.
The union representing New York City teachers spent more than $100,000 on a newspaper advertising campaign
based on the headline “This is no way to grade a teacher.” 13 Opponents argue that value-added assessments create
false accuracy that will be abused by parents and public officials who do not understand the limitations of this type
of assessment.

This seems to be a case where everyone is right, to some extent. Doug Staiger, an economist at Dartmouth
College who works extensively with value-added data for teachers, cautions that such data are inherently “noisy.” A
given teacher's results are often based on a single test taken on a single day by a single group of students. All sorts
of factors can lead to random fluctuations: anything from a particularly difficult group of students to a broken air
conditioning unit making noise in the classroom on test day. The correlation in year-to-year performance for a single
Machine Translated by
Google

teacher using these indicators is only about 0.35. (Interestingly, the correlation in year-over-year performance of
Major League Baseball players is also around 0.35, as measured by the batting average of hitters and the
cumulative performance average of pitchers.)
14

Data on teacher effectiveness is useful, Staiger says, but it is only one tool in the teacher performance
evaluation process. Data becomes “less noisy” when authorities have more years of data for a particular teacher
with different

classrooms of students (just as we can learn more about an athlete when we have data from more games and
more seasons). In the case of New York City teacher ratings, system leaders had been trained on the
appropriate use of value-added data and its inherent limitations. The public did not receive that information. As
a result, teacher evaluations are too often viewed as a definitive guide to distinguishing between “good” and
“bad” teachers. We like rankings (just think of US News & World Report's college rankings), even when the
data doesn't support such accuracy.
Staiger offers a final caveat of a different sort: We'd better be sure that the outcomes we're measuring,
such as the results of a particular standardized test, really correspond to what we care about in the long run.
Some unique data from the Air Force Academy suggests, unsurprisingly, that test scores that shine now may
not be golden in the future. The Air Force Academy, like other military academies, randomly assigns its cadets
to different sections of standardized core courses, such as introduction to calculus. This randomization
eliminates any potential selection effects when comparing the effectiveness of professors; over time, we can
assume that all professors get students with similar aptitudes (unlike most universities, where students with
different abilities may select into or out of different courses). The Air Force Academy also uses the same
curriculum and exams in each section of a particular course. Scott Carrell and James West, professors at the
University of California, Davis, and the Air Force Academy, used this elegant arrangement to answer one of
the most important questions in higher education: which teachers are most effective? 15 The answer: less
experienced and less knowledgeable teachers
degrees from fancy universities. These teachers have students who generally score better on standardized
tests in introductory courses. They also get better student evaluations for their courses. It is clear that these
young, motivated professors are more committed to their teaching than the older, grumpy professors with
PhDs from places like Harvard. The old folks must be using the same yellowed teaching notes they used in
1978; they probably think PowerPoint is an energy drink, except they don't know what an energy drink is
either. Obviously the data tells us that we should fire these old people, or at least let them retire with dignity.

But wait. Let's not fire anyone yet. The Air Force Academy study yielded another significant finding:
student performance over a broader horizon.
Carrell and West found that in math and science, students who had more experienced (and more
credentialed) instructors in introductory courses performed better in their required follow-up courses than
students who had less experienced instructors in introductory courses. A logical interpretation is that less
experienced instructors are more likely to “teach the
Machine Translated by
Google

“test” in the introductory course. This results in impressive test scores and happy students when it comes to
completing the instructor evaluation.
Meanwhile, the grumpy old professors (whom we almost fired just a paragraph ago) are focusing less on the test
and more on the important concepts, which are what matters most in later courses and life after the Air Force
Academy.

It is clear that we need to evaluate teachers and professors. We just have to make sure we do it right. The long-
term policy challenge, rooted in statistics, is to develop a system that rewards a teacher's real added value in the
classroom.

WHAT ARE THE BEST TOOLS TO COMBAT GLOBAL POVERTY?

We know surprisingly little about how to make poor countries less poor. It is true that we understand the things that
distinguish rich countries from poor ones, such as their educational levels and the quality of their governments. And
it is also true that we have seen countries like India and China transform economically in recent decades. But even
with this knowledge, it is not obvious what steps we can take to make places like Mali or Burkina Faso less poor.
Where should we start?

French economist Esther Duflo is transforming our understanding of global poverty by adapting an old tool to
new purposes: the randomized controlled experiment. Duflo, who teaches at MIT, literally conducts experiments on
different interventions to improve the lives of the poor in developing countries.
For example, one of the long-standing problems of schools in India is teacher absenteeism, particularly in small,
one-teacher rural schools. Duflo and co-author Rema Hanna tested a smart technology-based solution in a random
sample of 60 single-teacher schools in the Indian state. 16 teachers from these 60 experimental schools were
offered a bonus with cameras from Rajasthan. for good attendance. Here's the creative part: Teachers were given
tamper-proof date and time stamps. They proved that they had appeared every day by taking a photo with their
students.
17

Absenteeism was halved among teachers in the experimental schools compared with teachers in a randomly
selected control group of 60 schools.
Student test scores increased and more students graduated to the next level of education. (I bet the photos are
adorable too!)
One of Duflo's experiments in Kenya involved giving a randomly selected group of farmers a small subsidy to
buy fertilizer immediately after harvest. Previous evidence suggested that fertilizers significantly increase crop
yields. Farmers were aware of this benefit, but when it came time to plant a new crop, they often did not have
enough money left from the last harvest to buy fertilizer. This perpetuates what is known as the
“poverty trap,” as subsistence farmers are too poor to be less poor. Duflo and his co-authors found
that a small subsidy (free delivery of fertilizer) offered to farmers when they still had cash after the
harvest increased fertilizer use by 10 to 20 percentage points compared with use in a control group.
Machine Translated by
Google

Esther Duflo has even waded into the gender war. Who is more responsible when it comes to
managing family finances, men or women? In rich countries, these are the kinds of things couples
can argue about in marriage counseling. In poor countries, this can literally determine whether
children get enough to eat. Anecdotal evidence dating back to the dawn of civilisation suggests that
women place a high priority on the health and well-being of their children, while men are more likely
to drink their way to their wages at the local pub (or whatever the caveman equivalent was). At
worst, this anecdotal evidence simply reinforces old stereotypes. At best, it's difficult to prove,
because a family's finances are intertwined to some extent. How can we separate how husbands
and wives choose to spend community resources?
Duflo did not shy away from this delicate issue. 19 On the contrary, he found
a fascinating natural experiment. In Côte d'Ivoire, women and men in a family often share
responsibility for some crops. For long-standing cultural reasons, men and women also grow their
own cash crops. (Men grow cocoa, coffee, and a few other things; women grow bananas, coconuts,
and a few other crops.) The beauty of this arrangement from a research standpoint is that men's and
women's crops respond to rainfall patterns in different ways. . In years when cocoa and coffee are
doing well, men have more disposable income to spend. In years when bananas and coconuts are
doing well, women have more extra money.
Now we just need to address a delicate question: Are the children in these families better off in
years when the men's crops do well or in years when the women get a particularly bountiful harvest?
The answer: When women are doing well, they spend some of their extra money on more food
for the family. Men don't. Sorry guys.
In 2010, Duflo was awarded the John Bates Clark Medal. This award is given by the American
Economic Association to the best economist under the age of forty.
Among expert economists, this prize is considered more prestigious than the
Nobel Prize in Economics because it was historically awarded only every two years. (Following
Duflo's award in 2010, the medal is now awarded annually.) In any case, the Clark Medal is the
MVP.
Award for people with thick glasses (metaphorically speaking).
Duflo is conducting an evaluation of the program. His work, and that of others now using his
methods, is literally changing the lives of the poor. From a statistical perspective, Duflo's work has
encouraged us to think more broadly about how randomized, controlled experiments (long
considered the province of laboratory sciences) can be used more broadly to uncover causal
relationships in many other areas of life.
WHO KNOWS WHAT ABOUT YOU?

Last summer we hired a new nanny. When he arrived home, I began to explain our family background: “I am a
professor, my wife is a teacher. . .”
“Oh, I know,” the nanny said with a wave of her hand. "I googled you."
I was simultaneously relieved not to have to finish my rant and slightly alarmed by how much of
my life could be reconstructed from a brief Internet search. Our ability to collect and analyze vast
Machine Translated by
Google

amounts of data—combining digital information with cheap computing power and the Internet—is
unique in human history. We're going to need some new rules for this new era.
Let's put the power of data into perspective with just one example from retailer Target. Like most
companies, Target strives to increase profits by understanding its customers. To do this, the
company hires statisticians to perform the kind of “predictive analytics” described earlier in the book;
they use sales data combined with other information about consumers to determine who buys what
and why. None of this is inherently bad, because it means that the target near you is likely to have
exactly what you want.
But let's consider for a moment just one example of the kind of things that statisticians working in
the windowless basement of corporate headquarters can figure out. Target has learned that
pregnancy is a particularly important time in terms of developing purchasing patterns. Pregnant
women develop “retail relationships” that can last for decades. As a result, Target wants to identify
pregnant women, particularly those in their second trimester, and bring them into its stores more
frequently. A New York Times writer followed Target's predictive analytics team as it attempted to
find and attract 20 pregnant shoppers.
The first part is easy. Target has a baby shower registry where pregnant women register to
receive baby gifts before the birth of their children. These women are already Target shoppers and
have effectively told the store that they are pregnant. But here's the statistical twist: Target found
that other women who demonstrate the same purchasing patterns are likely also pregnant.
For example, pregnant women often switch to unscented lotions. They start buying vitamin
supplements. They start buying extra-large bags of cotton balls. Target's predictive analytics gurus
identified twenty-five products that together made up a "pregnancy prediction score." The goal of
this analysis was to send pregnancy-related coupons to pregnant women in hopes of attracting them
as long-term Target shoppers.
How good was the model? The New York Times Magazine published a story about a
Minneapolis man who walked into a Target store and demanded to see a manager. The man was
furious that his high school-aged daughter was being bombarded with pregnancy-related Target
coupons. “She’s still in high school and you’re sending her coupons for baby clothes and cribs? Are
you trying to encourage her to get pregnant? the man asked.
The store manager apologized profusely. He even called the father several days later to
apologize again. Only this time the man was less angry; it was his turn to apologize. “It turns out
there have been some activities going on in my house that I wasn’t fully aware of,” the father said.
"She will be born in August."
Target's statisticians had discovered that his daughter was pregnant before he did.

that's their business. . . and neither is your business. It may seem more than a little intrusive. For
that reason, some companies now hide how much they know about you. For example, if you are a
pregnant woman in your second trimester, you may receive some coupons in the mail for cribs and
diapers, along with a discount on a lawn mower and a coupon for free bowling socks with the
purchase of any pair of bowling shoes. It seems fortuitous to you that pregnancy-related coupons
Machine Translated by
Google

arrived in the mail along with the rest of the trash. In fact, the company knows that you don't bowl or
mow your own lawn; you're just covering your tracks so that what you know doesn't seem so creepy.
Facebook, a company with virtually no physical assets, has become one of the most valuable
companies in the world. For investors (unlike users), Facebook has one huge asset: data. Investors
don't love Facebook because it lets them reconnect with their prom dates. They love Facebook
because every mouse click generates data about where users live, where they shop, what they buy,
who they know, and how they spend their time. For users hoping to reconnect with their prom dates,
corporate data collection can cross the boundaries of privacy.
Chris Cox, Facebook's vice president of product, told the New York Times:
21
"The challenge of the information age is what to do with it."
Yeah.
And in the public sphere, the marriage of data and technology becomes even more complicated.
Cities around the world have installed thousands of security cameras in public places, some of which will
soon feature facial recognition technology. Law enforcement authorities can follow any car wherever it goes (and keep
extensive records of where it has been) by attaching a global positioning device to the vehicle and then tracking it via
satellite. Is this a cheap and efficient way to monitor potential criminal activities? Or is the government using technology to
trample on our personal freedom? In 2012, the U.S. Supreme Court unanimously decided that it was the latter, ruling that
law enforcement officials can no longer place tracking devices on private vehicles without a *
court order.
Meanwhile, governments around the world maintain huge DNA databases that are a powerful tool for solving crimes.
Whose DNA should be in the database?
That of all convicted criminals? That of each person arrested (whether or not ultimately convicted)? Or a sample of each of
us?
We are only just beginning to wrestle with the problems that lie at the intersection of technology and personal data,
neither of which were terribly relevant when government information was stored in dusty basement filing cabinets rather
than in digital databases that anyone can search from anywhere.
Statistics are more important than ever because we have more meaningful opportunities to make use of data. However,
formulas will not tell us which uses of data are appropriate and which are not. Mathematics cannot replace judgment.

In that sense, let's end the book with some word associations: fire, knives, cars, depilatory cream. Each of these things
has an important purpose. Each one improves our lives. And each of them can cause serious problems when abused.

You can now add statistics to that list. Go ahead and use data wisely and appropriately!

* I was not eligible for the 2010 award because I was over forty years old. Besides, he hadn't done anything to deserve it.
* United States v. Jones.
Machine Translated by
Google

Appendix
statistical software

I suspect you won't be doing your statistical analysis with pencil, paper and calculator. Here's a
quick tour of the most commonly used software packages for the types of tasks described in this
book.

Microsoft Excel
Microsoft Excel is probably the most widely used program for calculating simple statistics such as
mean and standard deviation. Excel can also perform basic regression analysis. Most computers
come equipped with Microsoft Office, so Excel is probably on your desktop right now. Excel is easy
to use compared to more sophisticated statistical software packages. Basic statistical calculations
can be performed using the formula bar.
Excel cannot perform some of the advanced tasks that more specialized programs can perform.
However, there are Excel extensions that you can purchase (and some that you can download for
free) that will expand the statistical capabilities of the program. A great advantage of Excel is that it
offers easy ways to display two-dimensional data with visually appealing charts. These charts can
be easily placed into Microsoft PowerPoint and Microsoft Word.
*
Was
Stata is a statistical package used worldwide by research professionals; its interface has a serious,
academic feel. Stata has a wide range of capabilities for performing basic tasks, such as creating
data tables and calculating descriptive statistics. Of course, that's not why university professors and
other serious researchers choose Stata. The software is designed to handle sophisticated statistical
testing and data modeling that goes far beyond the types of things described in this book.
Stata is ideal for those who have a solid understanding of statistics (a basic knowledge of
programming also helps) and those who don't need fancy formatting, just the answers to their
statistical queries. Stata is not the best choice if your goal is to produce quick graphs from your data.
Expert users

Say that Stata can produce good graphs but that Excel is easier to use for that purpose.

Stata offers several different stand-alone software packages. You can license the product for one year
Machine Translated by
Google

(after one year, the software no longer runs on your computer) or license it forever. One of the cheapest
options is Stata/IC, which is designed for "students and researchers with moderately sized data sets." There is
a discount for users who are in the educational sector. Even then, a single-user annual license for Stata/IC
costs $295 and a perpetual license costs $595. If you plan to launch a satellite to Mars and need to do some
really serious numerical calculations, you can look into more advanced Stata packages, which can cost
thousands of dollars.

SAS †
SAS has broad appeal not only to professional researchers but also to business analysts and
engineers due to its wide range of analytical capabilities. SAS sells two different statistical packages.
The first is called SAS Analytics Pro, which can read data in virtually any format and perform
advanced data analysis. The software also features good data visualization tools, such as advanced
mapping capabilities. It's not cheap. Even for those in the education and government sectors, a single
commercial or individual license for this package costs $8,500, plus an annual licensing fee.
The second SAS statistical package is SAS Visual Data Discovery. It has an easy-to-use interface that
requires no coding or programming knowledge while providing advanced data analysis capabilities. As the
name suggests, this package is intended to allow the user to easily explore data with interactive visualization.
You can also export data animations to presentations, web pages, and other documents. This one isn't cheap
either. A single commercial or individual license for this package costs $9,810, plus an annual license fee.
SAS sells some specialized management tools, such as a product that uses statistics to detect fraud and
financial crime.

This may seem like a character from a James Bond movie. In fact, R is a popular, free or “open source”
statistical package. It can be easily downloaded and installed on your computer within minutes. There is also
an active "R community" that shares code and can offer help and guidance when needed.
R is not only the cheapest option, but it is also one of the most malleable packages of all those described
here. Depending on your perspective, this

Flexibility is either frustrating or one of R's greatest assets. If you are new to statistical software, the
program offers almost no structure. The interface won't help you much. On the other hand,
programmers (and even people who have a basic familiarity with coding principles) may find the lack
of structure liberating. Users are free to tell the program to do exactly what they want it to do, even
to work with external programs.

*
IBMSPSS
IBM SPSS has something for everyone, from statistical experts to less statistically savvy business
analysts. IBM SPSS is good for beginners because it offers a menu-driven interface. IBM SPSS also
Machine Translated by
Google

offers a range of tools or “modules” designed to perform specific functions, such as IBM SPSS
Forecasting, IBM SPSS Advanced Statistics, IBM SPSS Visualization Designer, and IBM SPSS
Regression. Modules can be purchased individually or combined into packages.

The most basic package offered is IBM SPSS Statistics Standard Edition, which allows you to
calculate simple statistics and perform basic data analysis, such as identifying trends and building
predictive models. A single fixed-term commercial license costs $2,250. The premium package,
which includes most modules, costs $6,750. Discounts are available for those working in education.
sector.

* See [Link] †
See [Link]
* See [Link]
Machine Translated by
Google

Grades

Chapter 1: What's the point?


1 Central Intelligence Agency, The World Factbook,
[Link]
2 Steve Lohr, “For Today’s Graduate, Just One Word: Statistics,” New York Times,
August 6, 2009.
3 Ibid.
4 [Link],
[Link]
5 Trip Gabriel, “Gimmicks Find an Adversary in Technology,” New York Times,
December 28, 2010.
6 Eyder Peralta, “Atlanta Man Wins Lottery For Second Time In Three Years,” NPR News
(blog), November 29, 2011.
7 Alan B. Krueger, What Makes a Terrorist: The Economics and Roots of Terrorism
(Princeton: Princeton University Press, 2008).
Machine Translated by
Google

Chapter 2: Descriptive Statistics

1 US Census Bureau US, Current Population Survey, Social and Economic Supplements
annuals, [Link]
2 Malcolm Gladwell, “The Order of Things,” The New Yorker, February 14, 2011.
3 CIA, World Factbook and United Nations Development Programme, Human
Development Report 2011, [Link]
4 [Link].
Machine Translated by
Google

Chapter 3: Misleading Description


1 Robert Griffith, The Politics of Fear: Joseph R. McCarthy and the Senate, 2nd ed. (Amherst: University of
Massachusetts Press, 1987), p. 49.
2 “Catching Up,” Economist, August 23, 2003.
3 Carl Bialik, “When the Median Doesn’t Mean What It Seems,” Wall Street Journal, May 21–22, 2011.
4 Stephen Jay Gould, “The Median Isn't the Message,” with a prefatory note and a postscript
by Steve Dunn, [Link]
median_not_msg.html.
5 See [Link]
6 Box Office Mojo ([Link]), June 29, 2011.
7 Steve Patterson, “527% tax hike may surprise some, but it only costs about $5,” Chicago Sun-Times, December 5,
2005.
8 Rebecca Leung, “'The Texas Miracle': 60 Minutes II investigates claims that Houston schools falsified dropout rates,”
[Link], August 25, 2004.

9 Marc Santora, "Cardiologists Say Rankings Influence Surgical Decisions," New York Times, January 11, 2005.
10 Interview with National Public Radio, August 20, 2006, [Link]
storyId=5678463.
11 See [Link] college-rankings#4.
12 Gladwell, "The Order of Things."
13 Interview with National Public Radio, February 22, 2007, http://
[Link]/templates/story/[Link]?storyId=7383744.
Machine Translated by
Google

Chapter 4: Correlation
1 College Board, Questions
FAQs, [Link] predictors-with-first-
[Link].
2 College Board, Report on the overall profile of the college-bound senior population
university 2011, [Link]
3 See [Link]
Machine Translated by
Google

Chapter 5: Basic Probability

1 David A. Aaker, Brand Equity Management: Capitalizing on the Value of a Brand (New York: Free Press,
1991).
2 Victor J. Tremblay and Carol Horton Tremblay, The US Brewing Industry. US: Economic and Data Analysis
(Cambridge: MIT Press, 2005).
3 Australian Transport Safety Bureau Discussion Paper, “Cross Modal Safety Comparisons”, 1 January 2005.
4 Marcia Dunn, “1 in 21 trillion chances of being hit by a satellite,” Chicago SunTimes, September 21, 2011.
5 Steven D. Levitt and Stephen J. Dubner, Freakonomics: A Rogue Economist Explores the Hidden Side of
Everything (New York: William Morrow Paperbacks, 2009).

6 Garrick Blalock, Vrinda Kadiyali, and Daniel Simon, “Driving Fatalities after 9/11: A Hidden Cost of Terrorism”
(unpublished manuscript, December 5, 2005).

7 General information on genetic testing comes from the Human Genome Project Forensics,
[Link] uman_Genome/elsi/[Link].

8 Jason Felch and Maura Dolan, “FBI Resists Scrutiny of ‘Parties’,” The

Angeles Times, July 20, 2008.


9 David Leonhardt, “In Football, 6 + 2 Usually Equals 6,” New York Times, January 16, 2000.

10 Roger Lowenstein, “The War on Insider Trading: Beware the Market Winners,” New York Times Magazine,
September 22, 2011.
11 Erica Goode, “Send in the Police Before There’s a Crime,” New York Times, August 15, 2011.
12 Insurance risk data comes from all of the following: “Teen Drivers,” Insurance Information Institute, March
2012; “Texting Laws and Collision Claim Rates,” Insurance Institute for Highway Safety, September 2010; “Hot
Wheels,” National Insurance Crime Bureau, August 2, 2011.
13 Charles Duhigg, “What Does Your Credit Card Company Know About You?” New York Times Magazine, May
12, 2009.

Chapter 5½: Monty Hall Problem 1 John Tierney, “And


“Behind Door No. 1, a Fatal Flaw,” New York Times, April 8, 2008.

2 Leonard Mlodinow, Drunkard's Walk: How Chance Rules Our Lives (New York: Vintage
Books, 2009).
Machine Translated by
Google

Chapter 6: Problems with probability


1 Joe Nocera, “Risk Mismanagement,” New York Times Magazine, January 2, 2009.

2 Robert E. Hall, “The Long Slump,” American Economic Review 101, no.
3 (April 2011): 431–69.
4 Alan Greenspan, Testimony before the House Committee on Oversight and Government Reform, October 23,
2008.
5 Hank Paulson, Address at Dartmouth College, Hanover, NH, August 11, 2011.

6 “The Probability of Injustice,” Economist, January 22, 2004.


7 Thomas Gilovich, Robert Vallone, and Amos Tversky, “The Hot Hand in Basketball: On the Misperception of
Random Sequences,” Cognitive Psychology 17, no. 3 (1985): 295– 314.
8 Ulrike Malmendier and Geoffrey Tate, “Superstar CEOs,” Quarterly Journal of Economics 124, no. 4
(November 2009): 1593–638.
9 “The Price of Equality,” Economist, November 15, 2003.
Machine Translated by
Google

Chapter 7: The importance of data


1 Benedict Carey, “Learning from the Scorned and Drunken Fruit Fly,” New York Times, March 15, 2012.
2 Cynthia Crossen, “1936 Poll Fiasco Brought ‘Science’ to Election Polling,” Wall Street Journal, October 2,
2006.
3 Tara Parker-Pope, “Chances for Sexual Recovery Vary Widely After Prostate Cancer,” New York Times,
September 21, 2011.
4 Benedict Carey, “Researchers Find Biases in Reports of Drug Trials,” New York Times, January 17, 2008.
5 Siddhartha Mukherjee, "Do mobile phones cause brain cancer?" New York Times, April 17, 2011.
6 Gary Taubes, “Do We Really Know What Makes Us Healthy?” New York Times, September 16, 2007.
Machine Translated by
Google

Chapter 8: The Central Limit Theorem


1 US Census Bureau US
Machine Translated by
Google

Chapter 9: Inference
1 John Friedman, Out of the Blue: A History of Lightning: Science, Superstition, and Amazing Stories
of Survival (New York: Delacorte Press, 2008).

2 “Low Marks All Round”, Economist, July 14, 2011.


3 Trip Gabriel and Matt Richtel, “Inflating the Software Report Card,” New York Times, October 9,
2011.
4 Jennifer Corbett Dooren, “Link in Autism, Brain Size,” Wall Street Journal, May 3, 2011.
5 Heather Cody Hazlett et al., “Early Brain Overgrowth in Autism Associated with Increased Cortical
Surface Area Before 2 Years of Age,” Archives of General Psychiatry 68, no. 5 (May 2011): 467–76.
6 Benedict Carey, “Top Journal Plans to Publish Article on ESP, and Psychologists Are Outraged,”
New York Times, January 6, 2011.
Machine Translated by
Google

Chapter 10: Survey


1 Jeff Zeleny and Megan Thee-Brenan, “New Poll Finds a Deep Distrust of Government,” New York Times,
October 26, 2011.
2 Lydia Saad, “Americans Remain Firm in Support of Death Penalty,” [Link], November 17, 2008.
3 Telephone interview with Frank Newport, November 30, 2011.
4 Stanley Presser, “Sex, Samples, and Response Errors,” Contemporary Sociology 24, no. 4 (July 1995): 296–
98.
5 The results were published in two different formats, one more academic than the other. Edward O. Lauman,
The Social Organization of Sexuality: Sexual Practices in the United States (Chicago: University of Chicago
Press, 1994); Robert T. Michael, John H. Gagnon, Edward O. Laumann and Gina Kolata, Sex in America: A
Definitive Survey (New York: Grand Central Publishing, 1995).

6 Kaye Wellings, book review in British Medical Journal 310, no. 6978 (February 25, 1995): 540.
7 John DeLamater, “The NORC Sex Survey,” Science 270, no. 5235 (20 October 1995): 501.
8 Presser, “Sex, Samples, and Response Errors.”
Machine Translated by
Google

Chapter 11: Regression Analysis


1 Marianne Bertrand, Claudia Goldin and Lawrence F. Katz, “Dynamics of the Gender Gap for Young
Professionals in the Corporate and Financial Sectors,” NBER Working Paper 14681, January 2009.
2 MG Marmot, Geoffrey Rose, M. Shipley and P.J.S. Hamilton, “Employment level and coronary heart disease in
British civil servants.”
Journal of Epidemiology and Community Health 32, no. 4 (1978): 244–49.
3 Hans Bosma, Michael G. Marmot, Harry Hemingway, Amanda C.
Nicholson, Eric Brunner and Stephen A. Stansfeld, “Low occupational control and risk of coronary heart disease in
Whitehall II (prospective cohort)
Study”, British Medical Journal 314, no. 7080 (22 February 1997): 558–65.

4 Peter L. Schnall, Paul A. Landesbergis and Dean Baker, “Job Strain and Cardiovascular Disease,” Annual
Review of Public Health 15 (1994): 381–411.

5 MG Marmot, H. Bosma, H. Hemingway, E. Brunner and S. Stansfeld, “Contribution of job control and other risk
factors to social variations in disease incidence”
“coronary arteries”, Lancet 350 (July 26, 1997): 235–39.
Machine Translated by
Google

Chapter 12: Common Regression Errors


1 Gary Taubes, “Do We Really Know What Makes Us Healthy?” New York Times Magazine, September 16, 2007.
2 “Vive la Diferencia”, Economist, October 20, 2001.
3 Taubes, “Do we really know?”
4 College Board, 2011 Total Group Profile of College-Bound Seniors Report,
[Link]
5 Hans Bosma et al., “Low occupational control and risk of coronary heart disease in the Whitehall II study
(prospective cohort),” British Medical Journal 314, no. 7080 (22 February 1997): 564.

6 Taubes, “Do we really know?”


7 Gautam Naik, “Scientists’ Elusive Goal: Replicating Study Results,” Wall Street Journal, December 2, 2011.
8 John PA Ioannidis, “Conflicting and Initially Stronger Effects in Highly Cited Clinical Research,” Journal of the
American Medical Association 294, no. 2 (July 13, 2005): 218–28.
9 “Scientific Accuracy and Statistics”, Economist, September 1, 2005.
Machine Translated by
Google

Chapter 13: Program Evaluation


I Gina Kolata, “Arthritis Surgery for Bad Knees Is Cited as a Sham,” New York Times,
II July 2002.
2 Benedict Carey, “A Long-Awaited Medical Study Questions the Power of Prayer,” New York Times, March 31, 2006.
3 Diane Whitmore Schanzenbach, “What Have Project STAR Researchers Learned?” Harris School Working Paper,
August 2006.
4 Gina Kolata, “A Surprising Secret to a Long Life: Staying in School,” New York Times, January 3, 2007.
5 Adriana Lleras-Muney, “The Relationship between Education and Adult Mortality in the United States,” Review of
Economic Studies 72, no. 1 (2005): 189–221.

6 Kurt Badenhausen, “The Best Colleges to Get Rich,” [Link], July 30, 2008.
7 Stacy Berg Dale and Alan Krueger, “Estimating the Benefit of Attending a More Selective College: An Application of
Selection on Observables and Unobservables,” Quarterly Journal of Economics 117, no. 4 (November 2002): 1491–
527.

8 Alan B. Krueger, “Kids Smart Enough to Get into Elite Schools May Not Need to Bother,” New York Times, April 27,
2000.
9 Randi Hjalmarsson, “Juvenile prisons: a path to the straight and narrow or to more hardened criminality?” Journal of
Law and Economics 52, no. 4 (November 2009): 779–809.

Conclusion
1 James Surowiecki, “A Billion Prices Now,” The New Yorker, May 30, 2011.

2 Malcolm Gladwell, “Offensive Play,” The New Yorker, October 19, 2009.
3 Ken Belson, “NFL Roundup; Concussion Suits Joined,” New York Times, February 1, 2012.
4 Shirley S. Wang, “Autism diagnoses rise sharply in the US. “U.S.”, Wall Street Journal, March 30, 2012.
5 Catherine Rice, “Prevalence of Autism Spectrum Disorders,” Autism and Developmental Disabilities Monitoring
Network, Centers for Disease Control and Prevention, 2006, [Link]
mmwrhtml/[Link].

6 Alan Zarembo, “Autism Boom: An Epidemic of Disease or Discovery?” [Link], December 11, 2011.
7 Michael Ganz, “The Lifetime Distribution of the Incremental Societal Costs of Autism,” Archives of Pediatrics &
Adolescent Medicine 161, no. 4 (April 2007): 343–49.

8 Gardiner Harris and Anahad O'Connor, “On the Cause of Autism, It's Parents Versus Research,” New York
Times, June 25, 2005.
9 Julie Steenhuysen, "Study reveals 10 autism clusters in California," Yahoo! News, January 5, 2012.
10 Joachim Hallmayer et al., “Genetic Heritability and Shared Environmental Factors among Twin Pairs with
Machine Translated by
Google

Autism,” Archives of General Psychiatry 68, no. 11 (November 2011): 1095–102.


11 Gardiner Harris and Anahad O'Connor, “On the Cause of Autism, It’s Parents Versus Research,” New York
Times, June 25, 2005.
12 Fernanda Santos and Robert Gebeloff, “Teacher Quality Widely Diffused, Ratings Indicate,” New York Times,
February 24, 2012.
13 Winnie Hu, “As Teachers’ Grades About to Be Released, Union Launches Campaign to Discredit Them,” New
York Times, February 23, 2012.
14 T. Schall and G. Smith, "Are Baseball Players Regressing to the Mean?" American Statistician 54 (2000): 231–
35.
15 Scott E. Carrell and James E. West, “Does Teacher Quality Matter?”
“Evidence from Random Assignment of Students to Professors,” National Bureau of Economic Research Working
Paper 14081, June 2008.
16 Esther Duflo and Rema Hanna, “Monitoring Works: Getting Teachers to Come to School,” National Bureau of
Economic Research Working Paper 11880, December 2005.
17 Christopher Udry, “Esther Duflo: 2010 John Bates Clark Medalist,” Journal of Economic Perspectives 25, no. 3
(Summer 2011): 197–216.
18 Esther Duflo, Michael Kremer and Jonathan Robinson, “Nudging Farmers to Use Fertilizers: Theory and
Experimental Evidence from Kenya,”
National Bureau of Economic Research Working Paper 15131, July 2009.
19 Esther Duflo and Christopher Udry, “Intrahousehold Resource Allocation in CÔte d'Ivoire: Social Norms,
Separated Accounts and Consumption Choices,” working paper, December 21, 2004.
20 Charles Duhigg, “How Companies Learn Your Secrets,” New York Times Magazine, February 16, 2012.
21 Somini Sengupta and Evelyn M. Rusley, “The value of personal data?
“Facebook Ready to Figure It Out,” New York Times, February 1, 2012.

Expressions of gratitude

This book was conceived as an homage to an earlier WW Norton classic, How to Lie with
Statistics by Darrell Huff, written in the 1950s and which has sold over a million copies. That
book, like this one, was written to demystify statistics and persuade ordinary readers that what
they don't understand about the numbers behind the headlines can hurt them. I hope I have
done justice to Mr. Huff's classic. In any case, I would love to have sold a million copies in fifty
years!
I am continually grateful to WW Norton, and Drake McFeely in particular, for allowing me to write
books that address important issues in a way that is understandable to non-specialist readers. Drake
has been a close friend and supporter for over a decade.
Jeff Shreve is the WW Norton guy who made this book a reality.
Machine Translated by
Google

Knowing Jeff, one might think he would be too kind to impose the multiple deadlines involved in
producing a book like this. It's not true. Yes, he really is that nice, but somehow his gentle
nudge seems to do the job. (For example, these thanks are due tomorrow morning.) I
appreciate having a friendly foreman who keeps things moving.
My greatest debt of gratitude is to the many men and women who performed the important
research and analysis described in this book. I am not a statistician or a researcher. I am simply
a translator of other people's interesting and meaningful work. I hope I have conveyed
throughout this book how important good research and sound analysis are in making us
healthier, wealthier, safer and better informed.
In particular, I would like to acknowledge the extensive work of Princeton economist Alan
Krueger, who has made intelligent and significant research contributions on topics ranging from
the roots of terrorism to the economic benefits of higher education. (His findings on both topics
are pleasantly contradictory.) Most importantly (for me) Alan was one of my statistics professors
in graduate school; I have always been impressed by his ability to successfully balance
research, teaching, and public service.

Jim Sallee, Jeff Grogger, Patty Anderson, and Arthur Minetz read earlier drafts of the manuscript and
made numerous helpful suggestions. Thank you for saving me from myself! Frank Newport of Gallup
and Mike Kagay of the New York Times were kind enough to take the time to explain to me the
methodological nuances of polling. Despite all your efforts, the mistakes that remain are mine.

Katie Wade was a tireless research assistant. (I've always wanted to use the word “indefatigable,”
and finally, this is the perfect context.) Katie is the source of many of the anecdotes and examples that
illuminate concepts throughout the book. No Katie, there are no funny examples.
I have wanted to write books since I was in elementary school. The person who allows me to do that
and make a living doing it is my agent, Tina Bennett.
Tina embodies the best of the publishing business. She loves making meaningful work happen while
tirelessly promoting the interests of her clients.

And finally, my family deserves credit for tolerating me while I was writing this book.
(Chapter deadlines were posted on the refrigerator.) There's evidence that I become 31 percent more
irritable and 23 percent more exhausted when I approach (or miss) deadlines for important books. My
wife, Leah, is the first, best, and most important editor of everything I write. Thank you for that, and for
being such a smart, supportive, and fun partner in all our other endeavors.
The book is dedicated to my eldest daughter, Katrina. It's hard to believe that the child who was in a
crib when I wrote Naked Economics can now read chapters and provide meaningful feedback. Katrina,
you are a parent's dream, as is Sophie and CJ, who will soon be reading chapters and manuscripts as
well.
Machine Translated by
Google

Index

Page numbers in italics refer to figures.

“Absolute” score, 22, 23, 48 in


percentages, 28
academic reputation, 56
precision, 37–38, 84, 99
ACT, 55
African Americans, 201 years, 198,
199, 204
Air Force Academy, 248 a
49 airline passengers, 23 to 24, 25
Alabama, 87
alcohol, 110–11, 114
Alexander, Lamar, 230
algorithms, 64–65
Allstate, 81, 87
Alzheimer's disease, 242, 243
American Economic Association, 251
American Heart Journal, 229
The Changing Lives of Americans, 135, 135, 136, 137, 138–41, 150–52, 166, 192, 193,
195, 196, 199, 200, 201, 202, 204, 208, 221 annuities,
107 antidepressants,
121

Arbetter, Brian, x
Archives of General Psychiatry, 155, 156–60 arsenic, 48 lines of
assembly, 53
AT&T, 42
at bats, 32
Australian Transport Safety Board, 71–72
Austria, 65
autism, 4, 221, 244–46
brain size and, 155–60, 165
Avatar (film), 47, 48
average, see income
Machine Translated by
Google

average average, 16–17, 18–19, 27, 55 yards


average per pass attempt, 1

baboons, 207
banks, 170n
Bardo College, 57
Baseball Info Solutions, 16, 31–32 players
Baseball, 5, 248 best
of all time, 13, 15, 30,
31–32 basketball, streaks, 103
batting averages, xii, 1, 4, 5, 15–16 curve
bell-shaped, 20, 25–26, 26, 133, 134, 136, 208, 209
Bernoulli trial, 70
Bertrand, Marianne, 203
Bhutto, Benazir, 58, 64
Bhutto (film), 58–60, 64
biases,
113 binary variables, 200
binomial experiment, 70
Black Swan, The: The Impact of the Highly Improbable (Taleb), 98–99
Blalock, Garrick, 72–73
Blind taste tests, 68–71, 79, 79, 80, 97, 99
blood pressure, 115, 116 algae
green-blues, 116–17
Boston Celtics, 103
Boston Marathon, 23–24, 25
Botstein, Leon, 57
bowling scores,
4 boxers, 242,
243 cancer of
brain, 145 brain size, autism and,
155–60, 165 bran muffins, 11, 153–54
Brazil, 3
breast cancer, 122, 163
Brunei, 31
Budweiser, 68, 69
Buffett, Warren, 19, 81–82, 84 Office
from Labor Statistics, USA. US, 46 Burton,
Dan, 246 Bush, George
HW, 230 Bush, George W., 43,
53 Businessweek, 107

Caddyshack (film), 44–45


Machine Translated by
Google

calculus, ix–xi, xii


California, 41 years old
California, University of, 245
Canada, 3, 65
Canadian Tire, 88
cancer, 44, 162, 226
brain, 145
bran muffins and, 11, 153–54
mom, 122, 163
causes of, xv, 4
cell phones and, 145
colon, 11, 153–54 diet
and, 122 prostate,
163, 224 exams of
detection, 163
smoking and, xiv, 9–10, 11 groups
of cancer, 103–4 accidents
automotive, 8, 76
cardiology, 54–55
cardiovascular diseases, consult insurance
Car for heart disease, 81, 107
Carnegie Mellon University, 155
Carrell, Scott, 248–49 cars,
72
Carter, Jimmy, 50
casinos, 7, 79, 84, 99, 102
causality, causality, 225–40, 243 as
not involved in correlation, 63, 154, 215–17, 245–46 reverse, 216–17
Cave Test Security, 8
CBS News, 170, 172, 177, 178 service of
Cellular telephony, 42
Centers for Disease Control, 244
central limit theorem, 127–42, 146, 195 under study
on autism, 156, 158, 159 in surveys,
130, 170–71, 174 sampling and, 127–30,
139 central tendency, 17–18 , 19, 20–
21, 20, 21, 34 see also mean; median

Executive Directors, 19

Changing Lives, Check Out Charter Schools Changing Lives


Of the Americans, 113
Chase, Chevy, 44–45 do
trap, 89 author
Machine Translated by
Google

accused of, 143–44, 149 in evidence


standardized, 4, 8–9, 86, 145, 148–49
Chevrolet, 87
Chevrolet Corvette, 30-31
Chicago, University of, 75, 182, 203
Chicago Bears, 1-2
Chicago Cubs, 105–6, 235–36
Chicago Police Department, 87
Chicago Sun-Times, 154 decisions
on child care, 187
China, 39
coin of, 235
cholesterol, 116
chronic traumatic encephalopathy (CTE), 243 public officials,
185–87, 195, 205–7, 221 clarity, 38–39 change
climate, 180
groupings, groupings,
103–4 of the sample means, 138
cocaine, 219–20 cocoa, 251
coefficients, 196, 197,
199, 208, 220
in height, 193 regression, 193, 195, 196
size of, 197–
98 see also correlation coefficient
(r) coffee, 251 Tutor
cognitive, 155 coin toss, 71, 75–
76, 100, 104, 221–22 gambler's fallacy and, 106–
7 Cold War, 49 College
Board, 62, 63, 64 college degrees, 4 college rankings, 30 colon cancer, 11, 153–54 commercial banks, 97
Party
Communist, 37 completion rate, 1
confidence, 149 confidence interval, 158, 171–75, 176–77 Congress, U.S. USA, 171
constant, 193 control, 11, 198 non-equivalent, 233–40 see also
regression analysis control group, 114, 126, 227–28, 238–39 as counterfactual, 240 controlled experiments,
ethics and 9 variables
control, see explanatory variables Cook County, Ill., 48–49 Cooper, Linda, 144–45 University of
Cornell, 103 angioplasty
coronary, 54–55 coronary bypass surgery, 229–30, 58–67 exercise and weight correlation,
60 height and weight, 59, 59, 60, 61, 63, 65–67, 189–204, 190, 191,
208 negative, 60, 62 not implying causality, 63,
154, 215–17, 245–46 perfect, 60 in sports streaks, 103 correlation coefficient (r), 60–61
Machine Translated by
Google

cost of living adjustments, 47 deaths


sudden, 101–2
Ivory Coast, 251
Council of Economic Advisers, 32
Cox, Chris, 254
credit cards, xv, 88,
241 risks
credit, 88
credit score, 87 crimes,
criminals, 11, 14, 89 police officers and deterrence
of, 225–26, 227
predictions, 86–87 data
cross-sectional, 115-16 recall bias and, 123
Cruise, Tom, 86
CSI: Miami, 73
cumulative probability, 165
coins, 96
Cutler, Jay, 2
cyanobacteria, 117

Dale, Stacy, 234–35


The Dark Knight (film), 47
Dartmouth College, 233, 234, 247 data,
xv, 3, 115–17, 241–42, 252, 255 for
comparison, 113–15
transversal, 115–16
disagreements about, 13–14
normal distribution of, 25– 26 poor,
xiv as
representative, 111–13
sampling of, 6–7, 111–13
statistics and, 4, 111
summary of, 5, 14, 15–16, 17 data,
problems with, 117–26 bias of
healthy user, 125–26, 154 bias
publication bias, 120–22, 223
selection bias, 118–19, 178
survival, 123–25
data mining, 221–23
Data Mining and Predictive Analytics: Intelligence Gathering and Crime
Analysis, 87
data tables, 258
quotes, 36
misleading descriptions, 36–57 for
Machine Translated by
Google

coronary angioplasty, 54–55 media as,


42–43 medium as,
42–44 tests
standardized as, 51–52, 53–54 deciles, 22
Defense
Department, USA US, 49–50, 50 dementia, 242
Democratic Party,
USA US: Defense spending and, 49
tax increases and, 29

dependent variables, 192, 193–94, 197, 198, 199, 206n, 216, 217, 226 depression, 121, 242 descriptive statistics, 15–
35 central tendency discovered in,
18–22 dispersion measured in, 23–25 topics
framed by, 33 economic health of
the middle class measured by,
16–17 in Stata, 258 as summary, 15–16, 30

Diamantopoulou, Anna, 107 diapers,


14 diet, 115,
198 cancer and 122
difference in
differences, 235–37 discontinuity analysis,
238–40 dispersion, 23–24, 196, 210 average
affected by, 44 sample means,
136

Disraeli, Benjamin, 36
distrust, 169
DNA databases, 254
DNA testing, xi, 10
criminal evidence of, 74–75, 105 loci in, 73–75

prosecutor's fallacy and 105 fights


dogs, 242, 243–44 driving,
71–72, 73 rate of
abandonment, 53, 54 drugs,
9, 43–44, 115
drug smugglers, 108,
Machine Translated by
Google

109 drug treatment, 120


The Drunkard's Walk, El (Mlodinow), 92
Dubner, Stephen, 72 years old
Duerson, Dave, 242
Duflo, Esther, 250–52
dummy variables, 200
Duncan, Arne, 247
Dunkel, Andrés, xv

Economist, 39–40, 41, 42, 101–2


education, 115, 194, 200, 201, 204–5, 216–17, 218, 220, 249 income and 233–
35 longevity and 231–33
educational level, 31
Department of Education, U.S.
U.S., 155 Egypt, 170n Einhorn, David, 98
elections, USA
USA, 1936, 118–19
employment, 39–40 Enron: The Smartest
Guys in the Room
(film), 59 epidemiology, 222–23 estimation, 187–88
estrogen supplements,
211–12 ET (film),
48 European Commission, 107–8, 109
euros, 45 ex
convicts, 113, 147–48, 227 exercise, 125–
26, 198, 201

heart disease and, 188–89


weight and, 60, 62 surveys
at the polling station, 172–73
expected loss, 81
expected value, 77
of football plays, 77–78
of lottery tickets, 78–79
of investment in drugs for pattern baldness
male, 82–84 explanatory variables, 192, 193–94, 197, 198, 199,
203, 217 highly
correlated, 219–20 guarantees
expanded, 80–81, 82 sexual activity
extramarital, 183 points
Machine Translated by
Google

extra, 71, 77–78


extrapolation, 220–21 extrasensory perception (ESP), 161

Facebook, 254
False negatives (type II errors), 84, 162–64
False positives (type I errors), 84–85, 162–64
family structure, 115
fat tails, 208, 209–
10 FBI, 74–
75 Federal Emergency Management Administration, 144 films, highest grossing, 47–48 crisis
financial 2008, 7–8, 38, 95–100, 109 industry
financial, 7–8, 38, 95–100, 109 fires, 8
Fog
war (film), 59
Food and Drug Administration,
USA USA, 83 deserts
food, 201 coupons
of food,
200, 201 football, 51 extra point versus two point conversion
points in, 71, 77–78 trauma
cranial in, 114, 242–44
pin rating in, 1–
2, 3, 56 investment
foreign, 41 “4:15 report”, 96, 97 Framingham
Heart Study,
115–16, 136, 243 fraud, 86, 107 Freakonomics
(Levitt and Dubner), 72 distribution of
frequency, 20, 20, 25 rate of
freshman retention, 56–57 fruit flies, 110–11, 113, 114
Gallup Organization, 7, 177, 180, 181 fallacy
of the player, 102, 106–7 game, 7, 79,
99, 102
Gates, Bill, 18-19, 27, 134
gay marriage, 171
GDP, 217
Gender as an explanatory variable, 198, 199–200, 204, 205
gender discrimination, 107–8, 202–4
General Electric, 95, 96
General Equivalency Diploma (GED), 53 factors
genetics, 115 genetics,
245
Machine Translated by
Google

Germany, 39
Gilovich, Thomas, 103
Gini index, 2-3
Gladwell, Malcolm, 30–31, 56, 242, 243–44 globalization,
41 warm-up
global, 180
The Godfather, The (film), 47
Goldin, Claudia, 203 golf,
217
golf lessons, 214–15, 214
golf rangefinder, 38, 99
Gone with the Wind (film), 47, 48
google, 4
Gould, Stephen Jay, 44
public debt, average of
ratings from 99 to 100, 1, 5 to 6, 63
Graduation rate, 56 to 57
Great Britain, Probabilistic fallacy in the criminal justice system in, 100-102
Great Depression, 99, 241
Great Recession, 39
Green Bay Packers, 1-2
Greenspan, Alan, 97, 99
Grogger, Jeff, 32
gross domestic product (GDP), 241
Guantanamo Bay, 164
guilty beyond a reasonable doubt, 148
weapons, 72
Guskiewicz, Kevin, 243

Hall, Monty, xi–xii, 90–94


Hanna, Rema, 250
Harvard University, 211, 225, 233, 234, 249
“HCb2” count, 24–
25 HDL cholesterol, 116 head trauma, 114, 242–44
Health care, 189
cost containment
in, 85 health insurance, 82 bias of
healthy user, 125–26, 154
heart disease, 145, 148,
198, 217–18 supplements
estrogen and, 211 exercise and, 188–89
Framingham Study on,
115–16 , 136 component
Machine Translated by
Google

genetic of, 245 stress and,


185–87, 205–7
hedge fund managers, 19
height, 115, 204, 205 average, 25, 26, 35, 156–57, 159, 166–68 weight correlated with,
59, 59, 60, 61,
63, 65–67, 189–204, 190, 191, 208 heroin, 219–20
highly explanatory variables
correlated, 219–20
school dropout
secondary school, 226–27
HIV/AIDS, 84–85 , 182
Hjalmarsson,
Randi, 239 players
hockey, 242 executions
mortgages, 99
homeless people,
6 home mortgages,
97 insurance of
owners, 82
homosexuality, 182
Honda Civic, 87 Hoover, Herbert, 241 hormones, 207 hot hands, 102–3 Houston, Tex., 53–54 How I
Hussein, Saddam, 240 tests
of hypothesis, 146–48, 197

IBM SPSS, 260


Illinois, 29, 41, 87, 236 lottery,
78–79, 81 imprisonment,
239–40 incentives, 53 income,
32, 114, 194, 198,
204, 205 education and 235 per capita, 16–17, 18 –
19, 27, 55, 216 inclination
on the right, 133–34, 133 income inequality,
2–3, 41–42 income tax,
29, 114 income verification, 87
independent events:
gambler's fallacy and, 106–7
misunderstanding of, 102–3
probability of both happening, 75–
76, 100 probability that either
of the two happens, 76–77 independent variables, see explanatory variables
Machine Translated by
Google

India, 39, 45, 64


indicator, 95
infinite series, xii–xiii inflation,
16, 45–46, 47 information, 96–
97
Island, Thomas, 245
insurance, 8, 71, 81–82, 84, 89, 144–45 equality of
gender and, 107–8 interception rate,
1

internet, 235
Internet Surveys, 178
intervention, 225, 226–27 intuition,
xii–xiv
Ionnidis, John, 223
Iowa Informal Poll, 118

Iraq War, 240

Jaws (film), 47 jet engines,


100 Jeter, Derek, 15,
19 job placement program,
226 job training, 113, 227, 236–37, 236,
237 John Bates Clark Medal, 251
Journal of Personality and Social Psychology, 160–61 Journal of
the American Medical Association, 223 J.P. Morgan, 96
sentences, 57
criminals
juveniles, 239–40 Kadiyali,
Vrinda, 72–73 Kael, Pauline,
118 Kathmandu, Nepal,
116–17 Katz, Lawrence, 203
Kenya, 250 Kinney,
Delma, 9
Click, Jonathan, 227
Knight, Ted, 44–45
Krueger, Alan, 12–13, 32, 234–35 Kuwait, 31

Landon, Alf, 118–19


laser printers, central tendency explained by, 17–18, 19, 20–21, 20, 21, 34 law of
large numbers, 78–79, 84, 107 minimum
Machine Translated by
Google

squares, 190–94
legal system, 148, 162
“lemon” problem, 21
Let's Make a Deal, xi–xii, 90–94
leukemia, 104
leverage, 96
Levitt, Steve, 72
life expectancy, 31, 43
liquidity, 96
literacy rate, 55
Literary compendium, 118-19
Lleras-Muney, Adriana, 232–33
longevity, education and, 231–33
longitudinal studies:
Changing lives, 135, 135, 136, 137, 138–41, 150–52, 166, 192, 193, 195,
196, 199, 200, 201, 202, 204, 208, 221 on education, 116 on
education and income, 231–33 bias of
healthy user and 125 on heart disease,
115–16 recall bias
and, 122–23 Los Angeles,
California, 247 Los Angeles
Times,
74, 247 lottery: double
winner of, 9 irrationality of the game, xi, 78–79, 81, 89 Lotus Evora, 30–
31 luck, 106

malaria, 146–47, 148


male pattern baldness, 82–84, 83
Malkiel, Burton, 125n
Malmendier, Ulrike, 107
mammograms, 163
Manning, Peyton, 30
Mantle, Mickey, 5
manufacturing, 39–40, 39
marathon runners, 23 –24, 25
marathons, 127–29
marbles in urn, probability and, 112, 178–79 margin
error, see research companies
of
confidence interval market,
Machine Translated by
Google

113 Martin, JP, 88


Mauricio, 41
McCarthy, Joseph, 37
McKee, Ann, 243
McPherson, Michael, 56–57
Meadow, Roy, 101–2 Law
Meadow, 101–2 average,
18–19, 146, 196n affected
due to dispersion, 44 under study
of autism, 156 theorem
from the central limit and, 129, 131, 132

in correlation coefficient,
67 formula for,
66 tall Americans,
25, 26 of
income, 134 median
vs., 18, 19, 44 in Microsoft
Excel, 257 possible deception of, 42–
43 standard error for the difference of,
164–65 measles, mumps, and rubella (MMR),
245–46
median, 21 average
vs., 18, 19, 44
Outliers and 43 possible deception
from, 42–44 “The median is not the message” (Gould) ,
44 loss of
memory, 242 men, management of
money for, 250–51
Michelob, 68–71, 80 Court
Michigan Supreme Court, 56
Microsoft Excel, 61, 67, 257 class
media, 13, 15, 16–17, 32
according to the
measured by median, 19
Miller, 68, 69 minimum wage,
46–47 Minority Report
(film), 86th competition
of Miss America, 30 Mlodinow,
Leonard, 92
models, financial, 7–8, 38, 95–100
monkeys, 207
Machine Translated by
Google

Monty Hall Problem, xi–xii, 90–94


motorcycles, 72 Moyer,
Steve, 16, 31–32
magnetic resonance imaging scans, 163
multicollinearity, 219–20 multinationals,
170n multiple regression analysis, 199–204, 226 multivariate logistic regression, 206n mutual funds,

NASA, 72
National Football League, 1–2, 3, 51, 56, 242–44 Institute
National Mental Health Center, 245 National Center
National Opinion Research Council (NORC), 7, 181–83 Natives
Americans, 184
Natural experiments, 231–33
NBA, 103
negative correlation, 60, 62
Netflix, 4, 58–59, 60, 62, 64–65
Newport, Frank, 177
New York, 41, 54–55
New York, NY, 247
New Yorker, 30–31, 118, 241, 242 New
York Times, 4, 43, 86–87, 91–92, 96, 98–99, 110, 125, 154–55, 161, 169,
170, 172, 177, 178, 229–30, 231, 235, 246, 247, 254
New York Times Magazine, 88, 122, 222, 253
Nixon, Ricardo, 118
Nobel Prize in Economics, 251
Nocera, Joe, 96, 98–99
Nominal figures, 45–46, 47
non-equivalent control, 233–40
nonlinear relationships, 214–15
normal distribution, see bell curve
North Carolina, University of, 155–60
North Dakota, 41, 183–84
null hypothesis, 146–48, 149, 156–57, 166, 188 threshold of
rejection of, 149–50, 152, 153, 161–64, 197, 222
Nurses' Health Study, 211

Obama, Barack, 17, 32


job qualification of, 169, 170, 171, 177,
179 obesity, heart disease
and, 115–16 observations, 67, 192–93
Occupy Wall Street, 16, 169–70 bias
of omitted variable, 217–19
base percentage, 31
“one-tailed” hypothesis testing, 151, 166–68
Machine Translated by
Google

Ordinary least squares (OLS), 190–


94 osteoporosis, 211
outliers, 18–20, 21
mean reversion of, 105
median insensitivity to, 43
sample mean and, 138
in variance, 34
production, 39–
40 outsourcing, 41

Paige, Rod, 53, 54


Pakistan, 64
parameters, 196, 197
passer rating, 1–2,
3, 56 patterns,
14 Paulson, Hank, 98
Paxil, 121
[Link], 233
peer review survey,
56 Penn State, 56
economic output per capita, 31
per capita income, 16–17, 18–19, 27, 55, 216 percentage
of touchdown passes per pass attempt, 1 percentage, 27–
28, 29 exaggeration by, 48
formula for, 28
percentiles, 22 , 23
perfect correlation,
60 negative correlation
perfect, 60 Preschool Study of
Perry, 116 Peto, Richard, 222
Philadelphia 76ers, 103
physics, x placebo, 148, 228
effect
placebo, 229 agents
of police and deterrence
of crime, 225–26, 227 polling companies, 113 surveys,
xv, 169–84 precision
of the answers in,
181–83 central limit theorem and, 130,
170–71, 174 confidence interval in, 171–75 output, 172–73
Machine Translated by
Google

margin of error in, 171


methodology of, 178–82
poorly done, 178
presidential, xii
proportion used in, 171–
72 regression analysis vs.,
188 response rate of, 179–
80 sample size in,
172, 175 sampling in, 6 –7,
111–13 bias of
selection in, 178 on the activity
sexual, 6, 7, 181–83 standard error
in, 172–76, 195
phone, 112 Porsche
Cayman,
30–31 pounds, 45 poverty,
200–201, 249 –52
poverty trap, 250
prayer, 4, 13, 229–30, 231
precision, 37–38, 99, 247
predictive analytics, 252–54 surveillance
predictive policing, 86–
87, 108 pregnancy, 252–54
Princeton University, 233
printers, warranties on, 80–81,
82 prisoners, treatment
of drugs for, 120
information
private, 86
probability, 68–89
cumulative,
165 in gambling, 7 lack of
determinism in,
89 limits of, 9 and marbles in urn, 112, 178–79 utility of, xi probability, problems with, 95–109 in British justice
system, 100–
102 grouping,
104–5 and 2008 financial crisis, 7–8, 38, 95–100, 109
prosecutor's fallacy, 104–5
reversion to the mean, 105–7
Machine Translated by
Google

statistical discrimination, 107–9


probability density function,
79 productivity,
235 elaboration
of profiles, 108 prosecutor's fallacy,
104–5 prostate cancer, 163, 224
Prozac, 121
PSA test,
163 publication bias, 120–22,
223 p-value, 152, 157n, 159, 160, 197–98

Qatar, 31
how many, 95, 99
quarterbacks, 1–2
quartiles, 22

R (computer program),
259 r (correlation coefficient), 60–
61 calculation of, 65–67 race,
114, 200–201 program
of radio calls, 178
Rajasthan, India, 250
random error, 106
randomization, 114–15
randomized controlled trials, 227–29 group of
control as counterfactual in, 240 on curing
poverty, 250–52 ethics and, 227–
28, 240 on prayer and
surgery, 229 –30 on the size of
the school, 230–31
Random Walk Down Wall Street, A (Malkiel), ranking 125n, 30–
31, 56, 248
Rather, Dan, 53
rational discrimination, 108
Reagan, Ronald, 49, 50
real figures, 46
recall bias,
122–23 regression analysis, 10–12, 185–
211 difficulty of, 187
on gender discrimination,
202–4 height and weight, 189–204, 190, 191
in Microsoft Excel, 257
multiple, 199–204, 226
polls vs.,
188 standard error in, 195–
Machine Translated by
Google

97 Whitehall studies, 185–87, 195,


205–7 regression analysis, errors in, 187, 189, 211–
24 correlation confused with causality, 215–
16 data extraction,
221–23 extrapolation,
220–21 highly explanatory variables
correlated, 219–20 with relationships
non-linear, 214 –15 bias of
omitted variable, 217–19 coefficient of
regression, 193, 195, 196 regression equation,
191–92, 198, 201, 222
relative statistics,
22–23, 27 percent
as, 28th Match
Republican,
USA US: Spending on
defense and, 49
surveys of, 170
tax increases and,
29 waste, 190–94 response rate, 179–
80 causality
reverse, 216–17 reversal
(regression) to the mean,
105–7 Rhode Island,
41 “skewed to the right”, 44,
133 evaluation of
risks, 7–8 risk management,
98 Rochester, University
of, 55 Rodgers,
Aaron, 2, 29 Royal Statistical Society, 102 Rumsfeld, Donald, 14 rupees, 45, 47 Ruth, Babe, 32

Sallee, Jim, 141n


sample means, 132–33, 139, 139, 142, 150, 151, 167, 168
in the study of
autism, 156
grouping of, 138
dispersion of, 136
outliers and, 138 sampling,
6–7, 111–
13, 134 bad, 113 limit theorem
central and, 127–
30 homeless people, 6 size
of, 113, 172, 175, 196, 220
Santa Cruz,
California, 86–87 SAS, 258–59
Machine Translated by
Google

SAT scores, 55, 60, 62–63, 220


home televisions and 63–64
income and 63–64, 218–
19 in math test, 25, 224 average and
deviation
standard in, 25, 26 satellites, 72 Beer
Schlitz, 68–71, 79, 79,
80, 97, 99 schools,
187, 246–49
quality of, 51–52
size of, 230 –31
Science, 110–11 dashboards, 54–55 Commission
of Stock Market and Security, 86, 145
Selection bias, 118–19, 178
behavior of
self-reported vote, 181, 182 self-selection, 178 11 of
September 2001, terrorist attacks
of, 72–73, 74 “Sex Study,” 181–83
sexual behavior: of the
fruit flies, 110–
11, 113, 114 self-report of, 6, 7,
181–83
Shrek 2 (film), 47, 48
sigma, see
standard deviation sign, 193 significance, 193, 195–96
size vs., 154 significance level, 149–50, 152, 153, 157n, 166, 199 Simon, Daniel, 73 simple random sample, 112
Six
Sigma Man, 70n 60
Minutes II, 53 size, 193–94
significance vs., 154
slugging percentage, 31–32 Smith, Carol, ix–x, xii smoking, 116 cancer
caused by, xiv, 9 –10, 11
heart disease and, 115–16, 186–87, 189–90
smoking behavior, 115
socially insignificant effects, 194
“The Social Organization of Sexuality: Sexual Practices in the United States,” 181–83
sodium, 27
Sounds of Music, The (film), 48
South Africa, 3
Soviet Union, 49, 143
spam filters, 163
sports, streaks, 102-3
Sports Concussion Research Program, 243
Sports Illustrated, 107
Machine Translated by
Google

pumpkin, 188–89
Staiger, Doug, 247-48
Standard & Poor's 500, 123–24
standard deviation, 23–25, 146, 196n in the
Atlanta cheating scandal, 149 in
the autism study, 156 in the curve
bell, 133 central limit theorem and
129, 131 in the coefficient of
correlation, 67
formula for, 35 in Microsoft
Excel, 257 in distribution
normal, 26 standard error, 136–42, 152,
184 in autism study,
159–60 for mean difference, 164–
65 formula for, 138, 150, 172, 176–77 in
surveys, 172–76
in regression analysis, 195–97
standardized tests, 246
central limit theorem and 129–30 do
trap, 4, 8–9, 86, 145, 148–49 as an indicator
misleading, 51–52, 53–54 relative statistics
produced by , 22–23 see also SAT scores units
standard, 65
Stanford University, 103
Star Wars Episode IV (film), 47, 48
Stata, 258
statistical discrimination, 107–9
Statistical examples:
Author accused of cheating, 143–44, 149
author's investment, 28–29 income
average, 16–17, 18–19, 27 height of the
basketball players, 156–57, 159, 166–68 central limit theorem and, 127–28, 130, 131, 132 central tendency, 17–18,
19, 20–21, 20, 21 credit risks, 88 crimes and, 86–87
on the effectiveness
of the teachers,
248–49
Framingham Study, 115–16, 136
golf rangefinder, 38, 99
male pattern baldness medicine,
Machine Translated by
Google

82–84, 83 urn marbles, 112, 178–79


Netflix algorithm, 4, 58–59, 60, 62, 64–65
Perry Preschool Study, 116
predictive policing, 86–87, 108
of rare diseases, 84–85, 85
Schlitz beer, 68–71, 79, 79, 80
sodium, 27
standard deviation, 23–25 value
at risk, 38, 95–97, 98–100 see also
longitudinal studies
Statistically significant findings, 11, 153, 154–55, 194, 221
statistical software, 257–60
statistics:
data vs., 111

as detective work, 10–11,


14 errors in,
14 lack of certainty in, 144–46
lie with, 14
misleading conclusions from, xiv, 6 relative
versus absolute, 22–23
reputation of, xii, 1
as a summary, 5 , 14, 15–16, 17
ubiquity of, xii
undesirable behavior caused
by, 6 utility of, xi, xii, xv, 3, 14
see also misleading description; descriptive statistics steroids,
243 stock market,
71, 89 streaks,
102–3 stress, heart disease and, 185–
87, 205–7
accidents
cerebrovascular, 116
Student selectivity, 55 Substance abuse, 110–11 Cook County Suburban Tuberculosis Sanitarium District, 48 –
49 death syndrome
sudden infant seizure
(SMSL), 101–2 plus sign,
66 summer school, 238
Super Bowl, 68–71, 79, 97 Supreme Court,
USA USA, 254 surgery,
Machine Translated by
Google

prayer and, 4, 13, 229–30, 231


Surowiecki,
James, 241 survivorship bias, 123–25 Sweden, 3 pools, 72

Tabarrok, Alexander, 227


tail risk,
98, 99 tails, fat, 208, 209–10
Taleb, Nicholas, 98–99
Objective, 52–54
Tate, Geoffrey, 107
tau, 243
Taubes, Gary, 125
taxes, income, 29,
114 tax cuts, 43,
114, 180, 235 t-distribution, 196,
208–11 teaching quality,
51 teachers, 246–49
absenteeism between, 250 salary
of, 4
Technology companies, 154–55
telecommunications, 42
Telephone surveys, 112
televisions, 63–64
Ten Commandments, The (film), 48
Tennessee Project STAR Experiment, 230–31
terrorism, terrorists, 163–64
alert system for, 227
causes of, 11, 12–13 of
September 11, 72–
73, 74 risks
of, 73 scores of
tests, 53, 198 averages
reversal and, 106 see also SAT scores; standardized tests
Texas, 41, 54
texting while
leads, 88 thimerosal, 246
Tierney, John, 91–92
Titanic (film), 47
touchdowns, 1
exchange,
41 treatment, 9, 113–14, 225, 226–27 group
of treatment, 114, 126, 227–29, 238–39 coefficient of
Machine Translated by
Google

actual population, 208 population parameter


real, 196, 197 t-statistic, 197n

Tunisia, 170n
Tversky, Amos, 103
Twain, Mark, 36
twin studies,
245 two-point conversions, 71, 77–78
“two-tailed” hypothesis testing, 151, 166–68
Type I errors (false positives), 84–85, 162–64
Type II errors (false negatives), 84, 162–64

uncertainty, 71, 74
unemployment, 217, 236–37, 236, 241 unions,
47, 247 Index of
United Nations Human Development, 31, 55 States
United, 65 Index of
Gini of, 3
manufacturing in, 39–40, 39 media
height in, 25, 26 class
average in, 13, 15, 16–17, 32 production
economic per capita of, 31 unit of
analysis, 40–42 US News &
World Report, 55–57, 248

vaccines, 245–46
Vallone, Robert, 103
value-added evaluations, 247–48
value at risk, 38, 95–97, 98–100
variables, 224
dependent, 192, 193–94, 197, 198, 199, 206n, 216, 217, 226 explanatory (independent), 192, 193–94, 197, 198,
199, 203, 217 highly correlated, 219–20
Varian, Hal, 4
variance, 24
formula for, 34–35
outliers in, 34
Verizon, 42 years old
Vermont, 41
Veterans housing, 45–46
Vick, Michael, 242
vitamins, 125
Machine Translated by
Google

voting, self-reporting behavior, 181, 182

Copyright

Copyright © 2013 by Charles Wheelan

All rights reserved. Printed in the United States of America.


First edition

For information regarding permission to reproduce selections from this book, write to Permissions, WW Norton & Company, Inc., 500
Fifth Avenue

For information on special discounts for


bulk purchases, please contact WW Special Sales
Norton at specialsales@[Link] or 800-233-4830

Manufactured by Courier Westford


Production manager: Anna Oler

ISBN 978-0-393-07195-5 (hardcover) eISBN


978-0-393-08982-0

WW Norton & Company, Inc.


500 Fifth Avenue, New York, NY 10110
[Link]

WW Norton & Company Ltd.


Castle House, 75/76 Wells Street, London W1T 3QT
Machine Translated by
Google

Also by Charles Wheelan

10½ Things No Commencement Speaker Has Ever SaidNaked


Economics: Stripping Back the Dismal Science

Common questions

Powered by AI

Regression analysis is a powerful tool in social science research, capable of finding patterns in complex data sets and providing insights into societal challenges . However, its reliability is contingent on proper usage and the quality of the data and design . It helps identify associations between variables but cannot definitively establish causation or the reasons behind them . Regression is based on linear relationships, so it may not be suitable for nonlinear data without proper adjustments . Additionally, the presence of multicollinearity and small sample sizes can further complicate the interpretability and accuracy of the results . Experts must carefully choose variables and account for potential biases for reliable outcomes . Despite these challenges, regression analysis remains essential for uncovering insights that are not easily observable through direct observation ."}

The Central Limit Theorem (CLT) is foundational in statistics as it states that the distribution of sample means will approximate a normal distribution as the sample size becomes large, regardless of the shape of the population distribution. This is because a large sample size reduces random variation, causing the sample means to cluster closely around the population mean. The CLT is used to predict how sample means distribute, allowing statisticians to infer population characteristics from samples .

Statistical manipulation in school performance reporting occurs when the impact of external factors on student outcomes, such as socioeconomic background, is ignored in favor of simplistic metrics like test scores. Schools with affluent students may appear to perform better merely due to students' backgrounds rather than actual instructional quality . This kind of misrepresentation can lead to inaccurate evaluations, where schools appear more or less effective than they are because the statistics fail to account for differing student needs and backgrounds . Additionally, selective reporting or using inappropriate benchmarks, reminiscent of survivorship bias in the mutual fund industry, can also distort the perceived performance of schools . These manipulations skew public perception and can lead to misinformed educational policies and practices. Moreover, statistics like GPA can oversimplify a student's performance and fail to reflect course difficulty or the nuances of educational experiences . Overall, the misuse of statistics in this context emphasizes the necessity for careful, critical application and interpretation of education data to capture the true educational impact , as well as the risk of drawing incorrect conclusions from oversimplified statistics .

The standard error measures the dispersion of sample means around the population mean and reflects how much these sample means are expected to vary from one sample to another. It decreases with larger sample sizes and is influenced by the distribution of the underlying population . In contrast, the standard deviation measures the dispersion of individual data points within a single dataset and describes how much the data points differ from the overall mean of the population . The standard deviation is used to understand the spread in the data itself, while the standard error helps to understand the accuracy of a sample mean as an estimate of the population mean ."}

A primary ethical concern with controlled experiments involving humans is the balance between experimental integrity and participant welfare. An example is the large-scale prayer study assessing post-surgical outcomes; despite its importance, the high cost raises questions about resource allocation compared to direct patient care investments. Ethical guidelines require transparency, informed consent, and minimization of harm, ensuring that experiments do not exploit or unduly risk participant welfare for scientific gains .

The challenge of defining core concepts significantly affects the interpretation of statistics related to socially significant issues. Variable definitions lead to differing statistical outcomes, impacting interpretation. For instance, statistics like descriptive statistics simplify data but also risk losing nuance, affecting interpretations of complex social issues . Oversimplified metrics, such as the Human Development Index or quarterback ratings, provide practical insights but remain imperfect measures, susceptible to interpretation bias depending on how components are weighted . Furthermore, lack of clear definitions for concepts like "middle class" or "economic health" makes it difficult to make definitive conclusions, as statistical tools can be manipulated to support various narratives . Thus, the way core concepts are defined shapes the statistical tools used and influences the conclusions drawn from data.

Rankings of schools based on test scores can be misleading because they often fail to account for the selection mechanisms and inherent advantages of certain schools. For instance, selective enrollment schools, which admit students based on high standardized test scores, are often ranked as 'excellent' due to the high scores of their students. However, this recognition can be misleading as it does not necessarily reflect the quality of education provided, but rather the selectivity of the student body entering the school .

In multiple regression analysis, dummy variable encoding is used to include categorical variables, like sex, by creating binary variables. For example, sex can be encoded with 1 representing female and 0 representing male. The regression coefficient for the dummy variable then reflects the difference in the dependent variable, such as weight, associated with the categorical distinction, holding all other variables constant .

The role of incentives in statistics-based management is to drive behaviors that align with desired outcomes, often using data to inform decisions and shape strategies. Incentives can motivate specific actions when statistical analysis shows potential benefits or improvements, such as efficiency in business processes or targeted interventions in social policies . However, these incentives can lead to undesirable outcomes like inappropriate data use or manipulation, resulting in "statistical discrimination" or profiling, where certain groups are unfairly targeted based on generalized traits rather than individual behaviors. This can perpetuate biases and result in adverse impacts on particular demographics . Additionally, placing too much emphasis on statistics without considering their limitations can lead to oversimplifications that hide crucial details, fostering incorrect or misleading conclusions .

The 'hot hand' in sports is a classic example of the divergence between perception and statistical reality because human intuition often misinterprets random events as having patterns or streaks. This phenomenon is where players are perceived to have a streak of successful outcomes, like scoring multiple times in a game, leading observers to believe the player is "hot." However, statistical analyses often show that such streaks are consistent with chance alone, rather than a true increase in performance capability. Despite this, the "hot hand" belief persists because people tend to see patterns in randomness, failing to differentiate between random variation and actual changes in skill level. Precision in measuring such phenomena is often overshadowed by intuitive narratives, leading to misinterpretations of what the data actually describe .

You might also like