Advanced Statistics
Lecture 5
Quentin Lippmann
1 / 54
Introduction to Causal Inference
The Endogeneity Problem
• We are interested in y = β0 + β1 x1 + β2 x2 + ... + u
• We need hypotheses to interpret βj as the causal effect of xj
• Crucial hypothesis: E (u/x ) = 0
• The error term should not include variables correlated with x that
have an impact on y
• In practice: very hard to satisfy
• If violated: βj is biased and corresponds to a correlation
2 / 54
The Endogeneity Problem
Formally
• We want to estimate:
Y = β0 + β1 X + u
3 / 54
The Endogeneity Problem
Formally
• We want to estimate:
Y = β0 + β1 X + u
• When E [u/X ] 6= 0 → X is endogenous
3 / 54
The Endogeneity Problem
Formally
• We want to estimate:
Y = β0 + β1 X + u
• When E [u/X ] 6= 0 → X is endogenous
• When E [u/X ] = 0 → X is exogenous
3 / 54
The Endogeneity Problem
Formally
• We want to estimate:
Y = β0 + β1 X + u
• When E [u/X ] 6= 0 → X is endogenous
• When E [u/X ] = 0 → X is exogenous
3 / 54
The Three Sources of Endogeneity
1 Omitted Variables
4 / 54
The Three Sources of Endogeneity
1 Omitted Variables
4 / 54
The Three Sources of Endogeneity
1 Omitted Variables
• Some determinants of Y are unobserved and remain in the error term
• These determinants are correlated with x
• Ex: y = wage; x = education; error term = ability
4 / 54
The Three Sources of Endogeneity
1 Omitted Variables
• Some determinants of Y are unobserved and remain in the error term
• These determinants are correlated with x
• Ex: y = wage; x = education; error term = ability
2 Reverse Causality
4 / 54
The Three Sources of Endogeneity
1 Omitted Variables
• Some determinants of Y are unobserved and remain in the error term
• These determinants are correlated with x
• Ex: y = wage; x = education; error term = ability
2 Reverse Causality
• Y also partly/entirely causes X
• Ex: y = gdp; x = institutions
4 / 54
The Three Sources of Endogeneity
1 Omitted Variables
• Some determinants of Y are unobserved and remain in the error term
• These determinants are correlated with x
• Ex: y = wage; x = education; error term = ability
2 Reverse Causality
• Y also partly/entirely causes X
• Ex: y = gdp; x = institutions
3 Measurement error
4 / 54
The Three Sources of Endogeneity
1 Omitted Variables
• Some determinants of Y are unobserved and remain in the error term
• These determinants are correlated with x
• Ex: y = wage; x = education; error term = ability
2 Reverse Causality
• Y also partly/entirely causes X
• Ex: y = gdp; x = institutions
3 Measurement error
• True variable is X ∗ but we only observe X = X ∗ + µ
• Error term contains measurement issues
4 / 54
Illustrating omitted variable bias
• An example where we assume we know parameters
True model: log(wage) = β0 + 0.08educ + 0.30ability + ε
Estimated model: log(wage) = β0 + 0.08educ + u .
• Suppose more educated people have higher ability ⇒ ability ≈ 0.5educ + v
5 / 54
Illustrating omitted variable bias
• An example where we assume we know parameters
True model: log(wage) = β0 + 0.08educ + 0.30ability + ε
Estimated model: log(wage) = β0 + 0.08educ + u .
• Suppose more educated people have higher ability ⇒ ability ≈ 0.5educ + v
• Substitute ability into the true model
5 / 54
Illustrating omitted variable bias
• An example where we assume we know parameters
True model: log(wage) = β0 + 0.08educ + 0.30ability + ε
Estimated model: log(wage) = β0 + 0.08educ + u .
• Suppose more educated people have higher ability ⇒ ability ≈ 0.5educ + v
• Substitute ability into the true model
log(wage) = β0 + 0.08educ + 0.30(0.5educ + v ) + ε
log(wage) = β0 + (0.08 + 0.15) educ + (0.30v + ε)
| {z } | {z }
0.23 u
• If you regress only on educ: the slope you pick up is about 0.23
5 / 54
Illustrating omitted variable bias
• An example where we assume we know parameters
True model: log(wage) = β0 + 0.08educ + 0.30ability + ε
Estimated model: log(wage) = β0 + 0.08educ + u .
• Suppose more educated people have higher ability ⇒ ability ≈ 0.5educ + v
• Substitute ability into the true model
log(wage) = β0 + 0.08educ + 0.30(0.5educ + v ) + ε
log(wage) = β0 + (0.08 + 0.15) educ + (0.30v + ε)
| {z } | {z }
0.23 u
• If you regress only on educ: the slope you pick up is about 0.23
• Why bigger than 0.08?
• The part of ability that moves with education gets bundled into the
education coefficient
• If the omitted variable raises Y and increases with X ⇒ upward bias
5 / 54
Introduction to Causal Inference
Solving the Endogeneity Problem
• Solving the endogeneity problem ⇒ causal inference
6 / 54
Introduction to Causal Inference
Solving the Endogeneity Problem
• Solving the endogeneity problem ⇒ causal inference
• How do researchers solve the problem?
6 / 54
Introduction to Causal Inference
Solving the Endogeneity Problem
• Solving the endogeneity problem ⇒ causal inference
• How do researchers solve the problem?
• Effect of increasing the minimum wage on employment
6 / 54
Introduction to Causal Inference
Solving the Endogeneity Problem
• Solving the endogeneity problem ⇒ causal inference
• How do researchers solve the problem?
• Effect of increasing the minimum wage on employment
• Effect of having children on wages
6 / 54
Introduction to Causal Inference
Solving the Endogeneity Problem
• Solving the endogeneity problem ⇒ causal inference
• How do researchers solve the problem?
• Effect of increasing the minimum wage on employment
• Effect of having children on wages
• Effect of migration on employment
6 / 54
Introduction to Causal Inference
Solving the Endogeneity Problem
• Solving the endogeneity problem ⇒ causal inference
• How do researchers solve the problem?
• Effect of increasing the minimum wage on employment
• Effect of having children on wages
• Effect of migration on employment
6 / 54
Introduction to Causal Inference
Solving the Endogeneity Problem
• Solving the endogeneity problem ⇒ causal inference
• How do researchers solve the problem?
• Effect of increasing the minimum wage on employment
• Effect of having children on wages
• Effect of migration on employment
• Using experimental or quasi-experimental techniques
• Randomized controlled trials
• Difference in differences
• Regression discontinuity
• Instrumental variables
• Matching
• The rest of this course is devoted to these techniques
6 / 54
1 The Endogeneity Problem
2 Motivation
3 Terminology
4 RCTs and Covariates
5 Different RCT Designs
6 Power Calculations
7 Internal and External Validity
7 / 54
The 2019 Nobel Prize in Economics
• In 2019: A. Banerjee, E. Duflot and M. Kremer were awarded the
Nobel prize in economics
8 / 54
The 2019 Nobel Prize in Economics
• In 2019: A. Banerjee, E. Duflot and M. Kremer were awarded the
Nobel prize in economics
• Why?
9 / 54
The 2019 Nobel Prize in Economics
• In 2019: A. Banerjee, E. Duflot and M. Kremer were awarded the
Nobel prize in economics
• Why? “For their experimental approach to alleviating global
poverty”
• “The laureates have used a new approach to obtaining reliable
answers about the best ways to fight global poverty
9 / 54
The 2019 Nobel Prize in Economics
• In 2019: A. Banerjee, E. Duflot and M. Kremer were awarded the
Nobel prize in economics
• Why? “For their experimental approach to alleviating global
poverty”
• “The laureates have used a new approach to obtaining reliable
answers about the best ways to fight global poverty
• “It involves dividing this issue into smaller, more manageable,
questions
9 / 54
The 2019 Nobel Prize in Economics
• In 2019: A. Banerjee, E. Duflot and M. Kremer were awarded the
Nobel prize in economics
• Why? “For their experimental approach to alleviating global
poverty”
• “The laureates have used a new approach to obtaining reliable
answers about the best ways to fight global poverty
• “It involves dividing this issue into smaller, more manageable,
questions
• “They have shown that these smaller, more precise, questions are
often best answered via carefully designed experiments among the
people who are most affected
9 / 54
The 2019 Nobel Prize in Economics
• In 2019: A. Banerjee, E. Duflot and M. Kremer were awarded the
Nobel prize in economics
• Why? “For their experimental approach to alleviating global
poverty”
• “The laureates have used a new approach to obtaining reliable
answers about the best ways to fight global poverty
• “It involves dividing this issue into smaller, more manageable,
questions
• “They have shown that these smaller, more precise, questions are
often best answered via carefully designed experiments among the
people who are most affected
• "As a direct result of one of their studies, more than five million
Indian children have benefitted from effective programmes of
remedial tutoring in schools"
9 / 54
The 2019 Nobel Prize in Economics
• In 2019: A. Banerjee, E. Duflot and M. Kremer were awarded the
Nobel prize in economics
• Why? “For their experimental approach to alleviating global
poverty”
• “The laureates have used a new approach to obtaining reliable
answers about the best ways to fight global poverty
• “It involves dividing this issue into smaller, more manageable,
questions
• “They have shown that these smaller, more precise, questions are
often best answered via carefully designed experiments among the
people who are most affected
• "As a direct result of one of their studies, more than five million
Indian children have benefitted from effective programmes of
remedial tutoring in schools"
• Today: What is this experimental approach?
9 / 54
An Example: the Balsakhi Program (Banerjee et al.
2007)
• Question: how do we improve children’s learning outcomes?
10 / 54
An Example: the Balsakhi Program (Banerjee et al.
2007)
• Question: how do we improve children’s learning outcomes?
• Motivation: we know how to get children into school but we don’t
know how to improve school quality
• More spending on textbooks:
10 / 54
An Example: the Balsakhi Program (Banerjee et al.
2007)
• Question: how do we improve children’s learning outcomes?
• Motivation: we know how to get children into school but we don’t
know how to improve school quality
• More spending on textbooks: no impact on children’s test score
• Additional teachers:
10 / 54
An Example: the Balsakhi Program (Banerjee et al.
2007)
• Question: how do we improve children’s learning outcomes?
• Motivation: we know how to get children into school but we don’t
know how to improve school quality
• More spending on textbooks: no impact on children’s test score
• Additional teachers: no impact on children’s test scores
10 / 54
An Example: the Balsakhi Program (Banerjee et al.
2007)
• Question: how do we improve children’s learning outcomes?
• Motivation: we know how to get children into school but we don’t
know how to improve school quality
• More spending on textbooks: no impact on children’s test score
• Additional teachers: no impact on children’s test scores
• Policy: hire a young woman ("balsakhi") from the community to
help children who are lagging behind for 2 hours per day
10 / 54
An Example: the Balsakhi Program (Banerjee et al.
2007)
• Question: how do we improve children’s learning outcomes?
• Motivation: we know how to get children into school but we don’t
know how to improve school quality
• More spending on textbooks: no impact on children’s test score
• Additional teachers: no impact on children’s test scores
• Policy: hire a young woman ("balsakhi") from the community to
help children who are lagging behind for 2 hours per day
• Findings: substantial positive effect on children’s academic
achievement driven entirely by children who went to the balsakhi
• Contrasts with other policies that had no effect
10 / 54
An Example: the Balsakhi Program (Banerjee et al.
2007)
• Question: how can they say that the program had a positive
effect?
• It could be that only specific schools participated in the program
• Answer:
11 / 54
An Example: the Balsakhi Program (Banerjee et al.
2007)
• Question: how can they say that the program had a positive
effect?
• It could be that only specific schools participated in the program
• Answer: Randomized Controlled Trial
• Participation in the program was randomized at the level of a school
• Schools were randomly assigned into the treatment and the control
group
• Large sample of over 15,000 students
• Two different cities
• The Balsakhi program has since been adapted, re-evaluated, and
scaled up across India to affect more than 30m children
11 / 54
1 The Endogeneity Problem
2 Motivation
3 Terminology
4 RCTs and Covariates
5 Different RCT Designs
6 Power Calculations
7 Internal and External Validity
12 / 54
Terminology
• We are interested in Y = β0 + β1 T + u
• Where T is called a treatment
• Vaccine
• Job training
• Support classes
• Two groups
• Treatment group: Ti = 1
• Control group: Ti = 0
• We can also control for additional variables X
13 / 54
1. Intent-to-Treat Effect
• Two quantities of interest
• First quantity: Intent-to-Treat Effect (ITT)
• Effect of offering a product/service
• We usually cannot force people to take the treatment
• The issue of compliance is central to RCTs
• People can refuse to take up the treatment
• Needs to be taken into account
• How do we estimate the ITT?
• Simple: compare outcome means for Treatment and Control groups
T C
• ITT: βb = Y − Y
• The ITT is estimated by regressing Yi = α + βTi + ²i
14 / 54
2. Treatment-on-the-Treated Effect
• What is the effect of a program on those who took it?
• Second quantity: Treatment-on-the-Treated Effect (ToT)
• Corresponds to the Average Treatment on the Treated (ATT)
• ITT: measures the effect of offering job training
• ToT: measures the effect of taking up the job training
• If compliance is 100%: ATT = ITT
15 / 54
2. Treatment-on-the-Treated Effect
• How do we measure the ToT?
T =1 T =0
Y −Y ITT
ToT = T =1 T =0
=
T −T takeuprate
• Formally: corresponds to an instrumental variable approach
• The allocation in the treatment or the control group corresponds to
the instrument
• We’ll come back to that later during this course
16 / 54
Time for a Quizz
17 / 54
1 The Endogeneity Problem
2 Motivation
3 Terminology
4 RCTs and Covariates
5 Different RCT Designs
6 Power Calculations
7 Internal and External Validity
18 / 54
What are Control Variables Used for in RCTs?
• In RCT:
• ITT (ToT) can be measured with a simple difference in group
means
• Why do we need other control variables?
19 / 54
What are Control Variables Used for in RCTs?
• In RCT:
• ITT (ToT) can be measured with a simple difference in group
means
• Why do we need other control variables?
1 Randomization (Balance) checks
2 Uncertainty reduction
3 Heterogeneous treatment effects
19 / 54
1. Balance checks
• What are balance checks?
20 / 54
1. Balance checks
• What are balance checks?
• Balance checks are used to measure whether the treatment’s
assignment is really random
20 / 54
1. Balance checks
• What are balance checks?
• Balance checks are used to measure whether the treatment’s
assignment is really random
• Why are they used?
20 / 54
1. Balance checks
• What are balance checks?
• Balance checks are used to measure whether the treatment’s
assignment is really random
• Why are they used?
• RCTs are supposed to solve E (u/x ) = 0
• Balance checks provide a statistical measure
• To what extent my treatment and control groups differ?
20 / 54
1. Balance checks
• What are balance checks?
• Balance checks are used to measure whether the treatment’s
assignment is really random
• Why are they used?
• RCTs are supposed to solve E (u/x ) = 0
• Balance checks provide a statistical measure
• To what extent my treatment and control groups differ?
• How do you compute them?
20 / 54
1. Balance checks
• What are balance checks?
• Balance checks are used to measure whether the treatment’s
assignment is really random
• Why are they used?
• RCTs are supposed to solve E (u/x ) = 0
• Balance checks provide a statistical measure
• To what extent my treatment and control groups differ?
• How do you compute them?
• Take pre-treatment variables such as age, gender, income, etc.
• Compute the difference in group means between treatment and
group means for each variable
• You should not find any statistical difference
20 / 54
1. Balance checks
• What are balance checks?
• Balance checks are used to measure whether the treatment’s
assignment is really random
• Why are they used?
• RCTs are supposed to solve E (u/x ) = 0
• Balance checks provide a statistical measure
• To what extent my treatment and control groups differ?
• How do you compute them?
• Take pre-treatment variables such as age, gender, income, etc.
• Compute the difference in group means between treatment and
group means for each variable
• You should not find any statistical difference
• What are their main limits?
20 / 54
1. Balance checks
• What are balance checks?
• Balance checks are used to measure whether the treatment’s
assignment is really random
• Why are they used?
• RCTs are supposed to solve E (u/x ) = 0
• Balance checks provide a statistical measure
• To what extent my treatment and control groups differ?
• How do you compute them?
• Take pre-treatment variables such as age, gender, income, etc.
• Compute the difference in group means between treatment and
group means for each variable
• You should not find any statistical difference
• What are their main limits?
• You can only run balance checks on observable characteristics
• Unobservable characteristics may still differ
20 / 54
2. Heterogeneous Treatment Effects
• In many settings: policymakers will care about effects on
particular subgroups
• by age, gender, education level, etc.
• Compute the effect of the interaction between the treatment and
one of these variables
• Limit: multiple hypothesis testing
• Probability of finding a statistically significant result increases with
the number of hypotheses tested
21 / 54
1 The Endogeneity Problem
2 Motivation
3 Terminology
4 RCTs and Covariates
5 Different RCT Designs
6 Power Calculations
7 Internal and External Validity
22 / 54
Different RCT Designs
• There are different ways to implement a RCT
• It depends on many factors such as
• The political context
• Project capacities
• The treatment
• Randomization can also occur at different levels
• Individual, school, city, etc.
• Most common designs
1 Lottery
2 Phase In
3 Encouragement Design
23 / 54
1. Lotteries
An example
• Motivation: perinatal depression ranges from 10 to 20% of the
population
• Often undiagnosed
• Important consequences on mental health
• Important consequences on economic outcomes
• Question:
24 / 54
1. Lotteries
An example
• Motivation: perinatal depression ranges from 10 to 20% of the
population
• Often undiagnosed
• Important consequences on mental health
• Important consequences on economic outcomes
• Question: can psychotherapy reduce it?
• For how long?
• Long term impacts?
• Source: Baranov, Victoria, Sonia Bhalotra, Pietro Biroli, and Joanna Maselko, "Maternal
Depression, Women’s Empowerment, and Parental Investment: Evidence from a
Randomized Controlled Trial." American Economic Review, 2020
24 / 54
1. Lotteries
An example: design
• How can we assess the impact of psychotherapy?
• Compare people with and without psychotherapy
• What is the specification? Does E (u/x ) = 0?
25 / 54
1. Lotteries
An example: design
• How can we assess the impact of psychotherapy?
• Compare people with and without psychotherapy
• What is the specification? Does E (u/x ) = 0?
• Randomized Controlled Trial:
• Provide a treatment randomly
• Cognitive behavioral therapy to perinatally depressed women
• Design:
• Rural Punjab in Pakistan from April 2005 to March 2006
• 20 villages received the treatment
• 20 villages did not
• All 3,518 pregnant women assessed for prenatal depression
• 463 in treatment vs 440 in control
25 / 54
1. Lotteries
26 / 54
1. Lotteries
Balance Checks
27 / 54
1. Lotteries
28 / 54
1. Lotteries
29 / 54
2. Phase In Design
• Problem: often impossible to implement a program everywhere at
the same time
• Because of financial and administrative reasons
• Solution: phase-in the program over several years
• Solution: phase in design
• Decompose the whole area in a large number of similar units
• Randomly draw the order of introduction of the program
30 / 54
2. Phase In Design
• Advantages
• Ethical: treatment has to be delivered to everybody (ex: medication)
• Cooperation: local authorities have an incentive to cooperate and
maintain contact with researchers
• Disadvantages
• Anticipation of future benefits can change control group behaviour
• ex: knowing that I will have access to cheap microcredit, I delay
investing now
• Impossible to estimate long-run effects
• Phases need to be long enough to allow for lagged treatment effects
31 / 54
2. Phase In Design: An Example
• Motivation: Worms infect 1 in 4 people worldwide, especially
school-age children in Africa
• Reason: poor sanitation / hygiene
• Important negative consequences on schooling
• Program: Annual mass-deworming program in Kenyan Schools
• Design
• 75 schools participated (30,000 pupils)
• Phased in over three years
• Group 1 in 1998, group 2 in 1999 and group 3 in 2000
• Identification
• 1998: Group 1 as T, groups 2+3 as C
• 1999: Groups 1+2 as T, group 3 as C
• Source: Miguel, Kremer (2004), "Worms: Identifying Impacts on Education and Health in
the Presence of Treatment Externalities", Econometrica 72(1)
32 / 54
2. Phase in Design: Balance Checks
33 / 54
2. Phase In Design: An Example
• Compliance:
• 78% of eligible pupils received medical treatment
• Those who did not take it benefited from positive externalities
• Main Results
• School absence decreases from 25% in C-schools to 16% in T-schools
• More than 1/3 reduction in absences
• Cost-effectiveness
• Cost of deworming: $0.49 per pupil per year
• $3.50 per extra school year
34 / 54
2. Phase In Design: An Example
• Strong externalities at play
• Children infect each other
• This is why ITT and ATT distinction does not matter here
• Long-term effects
• Evidence that wages are higher and long-term health is better for
those who benefited from deworming as a child
35 / 54
3. Encouragement Design
• Problem: Randomizing access to a program may not be feasible
• We want to know if using fertilizer increases farmer income
• We cannot force them to adopt the technology
• We want to know whether doorstep discussions influence voters
• We cannot force them to have a discussion
• Encouragement Design
• Randomize who receives incentives to get the treatment
• Marketing, information, reminders, financial incentives
• We need different take-up rates of treatment between the two
groups
36 / 54
3. Encouragement Design
• Question: does door-to-door canvassing influence voters?
• Problem:
• Individuals who accept to have a discussion are different from those
who don’t
• We cannot force individuals to have a discussion
• Solution: Randomized experiment
• Randomize the door-to-door strategy by precinct
• Treatment: precincts where canvassers are sent
• Control: precincts where no canvassers are sent
• Compare the vote share at the precinct level
• Context: 2012 French presidential campaign
• Source: V. Pons (2018), "Will a Five-Minute Discussion Change Your Mind? A Countrywide
Experiment on Voter Choice in France", American Economic Review, 2018
37 / 54
3. Encouragement Design
38 / 54
3. Encouragement Design: Balance Checks
39 / 54
3. Encouragement Design
Impact on turnout
40 / 54
3. Encouragement Design
Impact on left-wing candidate’s vote share
41 / 54
3. Encouragement Design
• Results
• Turnout: no impact
• Vote share: François Hollande’s vote share increased by 3.2 p.p. in
precincts targeted by canvassers
• Why is it an encouragement design?
• Not possible to force individuals to have a discussion
• Only a portion of voters will accept it → similar to taking up a
treatment
• Randomize incentives to get the treatment
42 / 54
1 The Endogeneity Problem
2 Motivation
3 Terminology
4 RCTs and Covariates
5 Different RCT Designs
6 Power Calculations
7 Internal and External Validity
43 / 54
What Are Power Calculations?
• We want to test H0 : β = 0 against H1 : β 6= 0
• Power of an experiment design ≡ Probability that we will be
able to reject H0 if H1 is correct
• Why it matters?
• When planning an experiment, one needs to plan the sample size
• The power gives us an indication on the sample size needed to
statistically identify a given treatment effect
• Without adequate power ⇒ likely to find non significant effect
• Failure to find a statistically significant effect can be misinterpreted
as the failure of the program, rather than the failure of the evaluation
44 / 54
Key Parameters
• Minimum detectable effect size
• Effect size below which we cannot precisely distinguish the effect
from zero
• Need to think about the expected effect
• MDES increases ⇒ Necessary sample size decreases
• Standard deviation of population outcome
• Measures the variability of the data
• SD increases ⇒ Necessary sample size increases
• Statistical confidence/Precision
• Type 1 errors: null hypothesis is true but is rejected (false positive)
• Type 2 errors: null hypothesis is false but fails to be rejected (false
negative)
• The greater the precision ⇒ Necessary simple size increases
• Sample size
45 / 54
Time for a Quizz
46 / 54
1 The Endogeneity Problem
2 Motivation
3 Terminology
4 RCTs and Covariates
5 Different RCT Designs
6 Power Calculations
7 Internal and External Validity
47 / 54
Criteria to Assess the Validity of a RCT
• Internal validity:
• Is the effect causal?
• Is the selection bias really zero?
• Do we measure what we intended to measure?
48 / 54
Criteria to Assess the Validity of a RCT
• Internal validity:
• Is the effect causal?
• Is the selection bias really zero?
• Do we measure what we intended to measure?
• External validity
• Can the results be extrapolated to a larger population?
• Are they just about a very specific context?
48 / 54
Threats to Internal Validity
• Failure of randomization
• Randomize based on non-random variables e.g. letters of the
alphabet
49 / 54
Threats to Internal Validity
• Failure of randomization
• Randomize based on non-random variables e.g. letters of the
alphabet
• Non-compliance with experimental protocol
• Individuals want to be part of the treatment or the control depending
on the policy
• They may pressure experimenters
49 / 54
Threats to Internal Validity
• Failure of randomization
• Randomize based on non-random variables e.g. letters of the
alphabet
• Non-compliance with experimental protocol
• Individuals want to be part of the treatment or the control depending
on the policy
• They may pressure experimenters
• Attrition
• If an experiment lasts several months/years ⇒ individuals may move
or drop out of the experiment
• One of the worst problems in practice
49 / 54
Threats to Internal Validity
• Small Samples
• Usually costly to run a RCT ⇒ researchers work with small samples
• Does not cause bias but leads to imprecision
50 / 54
Threats to Internal Validity
• Small Samples
• Usually costly to run a RCT ⇒ researchers work with small samples
• Does not cause bias but leads to imprecision
• Hawthorne effect
• Subjects behaving differently because they know they are being
studied
• ex: Workers’ productivity increases because they know they are being
watched and not because of the treatment
• This is why medical trials include placebos
50 / 54
Threats to Internal Validity
• Small Samples
• Usually costly to run a RCT ⇒ researchers work with small samples
• Does not cause bias but leads to imprecision
• Hawthorne effect
• Subjects behaving differently because they know they are being
studied
• ex: Workers’ productivity increases because they know they are being
watched and not because of the treatment
• This is why medical trials include placebos
• Observer Effect
• Researchers unconsciously influence the subjects
• This is why researchers do not know who got the treatment in
medical trials
50 / 54
Threats to External Validity
• General Equilibrium Effects
• Would the results hold if the policy was scaled up to all the people of
a country
• ex:. impact of job training? ⇒ What happens if everyone
participates in the job training? Positive impacts on
employment/wage outcome likely to be much smaller
51 / 54
Threats to External Validity
• General Equilibrium Effects
• Would the results hold if the policy was scaled up to all the people of
a country
• ex:. impact of job training? ⇒ What happens if everyone
participates in the job training? Positive impacts on
employment/wage outcome likely to be much smaller
• RCT results are context-specific
• Non-representative sample
• ex: most RCTs in developing countries come from a relatively small
number of countries that are stable and willing to
• Non-representative treatment
• ex: when scaled up, treatments are usually different
51 / 54
Which matters the most?
• RCTs are designed to ensure internal validity
• They should be able to deliver causal results
• External validity remains debatable
• It is often the main limit of RCTs
• It can be addressed by comparing many different RCTs but this is
costly
52 / 54
1 The Endogeneity Problem
2 Motivation
3 Terminology
4 RCTs and Covariates
5 Different RCT Designs
6 Power Calculations
7 Internal and External Validity
53 / 54
Summary
• What are RCTs?
• How do they solve the selection problem?
• Intent to treat and Treatment on the treated
• The roles of covariates
• The different RCT Designs
• Power Calculations
• Internal and External validity
54 / 54