RM - Module 1
RM - Module 1
Module 1
Did Swachh Bharat reduce diarrhoea?
Attribution
• Development programs: typically designed to change outcomes
• Measurement:
• Usually inputs and outputs are tracked
• whether intended outcomes actually improve?
• changes in well-being that can be attributed to a particular project, program, or policy
• Evidence-based policymaking
• results guide target-setting, accountability, budget allocation, and program design
• Counterfactual:
• what would have happened to participants had they not participated?
[valid comparison group]
• Purpose: impact/process/cost/simulation/monitoring
‘How’ of evaluations
• Prospective evaluations
• evaluation designed ex ante; baseline before exposure; pre-specified outcomes
with counterfactual strategy
• Retrospective evaluations
• evaluation designed ex post; counterfactual reconstructed from observational
variation / secondary data
• Prospective:
• Government pilots SBM in selected districts first; collect pre and post
rollout OD and health data
• Retrospective
• Researchers use NFHS 2015–16 and 2019–21 to compare districts with
differing SBM exposure
‘How’ of evaluations
• Efficacy studies
• small-scale, tightly controlled tests proving a concept can work under ideal
conditions (think “proof of concept”)
• strong internal validity
• Effectiveness studies
• generate evidence under real-world implementation conditions
• greater external validity
The distinction matters because a program that works in a controlled trial may
fail when implemented by overburdened government bureaucracies.
‘How’ of evaluation
• Five approaches that complement impact evaluation
• monitoring
• provides administrative data to verify implementation fidelity and dosage
• ex ante simulations
• project likely effects before launch
• process/implementation evaluations
• examine how a program is implemented
• Often mixed methods
• Knowledge
• capable of generating significant learning; addressing a genuine knowledge gap
• Implementation
• supported by operational integrity and tied to decision-making commitments
• Feasible
• time and budget constraints
Criterion SBM
• Result chain: cement floors → reduce parasites and dust → improved child health →
reduced maternal stress →increased maternal happiness
• Embedded Assumptions
• Dirt floors are a significant source of parasites/dust.
• Child illness meaningfully affects maternal stress.
• Maternal happiness is responsive to reductions in stress.
• No major offsetting effects (e.g., maintenance burden, social stigma).
Mechanism Link Specific Measurable Attributable Realistic Targeted
Prevalence of
Lab-confirmed Compare
Cement floors → soil-transmitted Detectable Direct biological
stool sample treatment vs.
Reduced helminth within 6–12 exposure
(binary infection control
parasites infection months pathway
status) households
(children <5)
Random Direct
Improved child Maternal Validated scale Low-cost survey-
assignment psychological
health → Reduced perceived stress (e.g., Perceived based
identifies causal caregiving
maternal stress score Stress Scale) instrument
effect mechanism
Reduced
Self-reported life Standard Randomized
maternal stress → Standard survey Direct well-being
satisfaction (0– subjective well- exposure to
Increased instrument outcome
10 scale) being question intervention
happiness
Quick recap exercise
Rural Girls’ Secondary School Scholarship (monthly)
• Identify:
• the key counterfactual question the evaluation must answer
• at least one link in the chain that you think is most uncertain and would benefit from a
mechanism experiment
• SMART indicators for the final outcome
• Policy context: evaluates the “Graduation” anti-
poverty program (the transfer of a productive asset with
consumption support, training, and coaching plus
savings encouragement and health education and/or
services ) across six countries (Ethiopia, Ghana,
• Policy context: evaluates a school-based Honduras, India, Pakistan, and Peru)
deworming program across 75 primary schools
(~30,000 pupils) in rural Kenya
• Design: six parallel RCTs covering ~10,495 households
• Design: cluster-randomised phase-in. • theory of change: the poverty-trap model
• Schools were randomly assigned to three • measures outcomes at every link in the causal
groups receiving treatment in staggered chain identified from ToC
sequence • consumption, food security, productive and
• at any point in time, not-yet-treated schools household assets, financial inclusion, time
serve as the counterfactual. use, income and revenues, physical health,
mental health, political involvement, and
• Results: 25% reduction in absenteeism in women’s empowerment
treatment schools
• significant cross-school externalities: • Results: statistically significant impacts on 10 key
disease-transmission spillovers outcomes or indices
Four types of comparisons (Tilly, 1997; Macrosociology)
• Individualising comparison
• examines a small number of cases to highlight what makes each distinctive
• Universalising comparison
• aims to show that a phenomenon has essentially identical features across all instances
examined
• Variation-finding comparison
• identifies systematic differences across cases to establish a principle of variation in character or
intensity
• Encompassing comparison
• locates different instances at varying positions within the same overarching system (e.g., the
capitalist world-economy) and explains their characteristics as functions of their relationships
to that system
Type Core Aim What It Does What It Is Not
Compares a small number of
Individualising cases to highlight how each Not aimed at deriving general
Clarify distinctiveness
Comparison follows a distinctive historical laws.
trajectory.
Demonstrates that different cases
Universalising Establish common Not claiming total empirical
share the same underlying
Comparison properties uniformity.
process, structure, or mechanism.
Identifies systematic variation
Variation-Finding Explain patterned across cases to derive a principle
Not merely descriptive contrast.
Comparison differences explaining differences in outcome
or intensity.
Situates cases within a larger
interconnected system and
Encompassing Explain relational Not treating cases as
explains their features as
Comparison position independent units.
consequences of their structural
position within that system.
Classify the Comparisons
A historian examines the
A political scientist argues
unique institutional features
that all democratic
of Indian federalism to
transitions share three
explain why India’s
common features: elite
democratic trajectory differs
splits, mass mobilisation,
from other post-colonial
and negotiated pacts.
states.
A world-systems scholar
A researcher compares land explains why East Asian
reform outcomes across five states industrialised while
Latin American countries to Sub-Saharan African states
identify which factors explain did not by locating both
variation in redistribution. within the global division of
labour.
A Tale of Two Cultures (Mahoney & Goertz, 2006)
• Quantitative and qualitative research not points on a single continuum
• distinct methodological cultures with different values, beliefs, and norms.
• each tradition’s practices are internally coherent
• Key differences
• Effects-of-causes vs causes-of-effects
• Additive regression vs Boolean combinations
• Population inference vs case explanation
• Probabilistic vs necessary/sufficient causation
Evidence-based Policymaking (EBPM)
• vague, aspirational term rather than a description of how policy actually works
• “Policy” encompasses the sum total of government action from signals of intent to final outcomes.
• “Evidence” is an argument or assertion backed by information, with scientific evidence occupying a privileged
but contested position atop a hierarchy of methods.
• “Policymakers” are not a monolithic group but a dispersed population operating within policy communities.
• Issue of evidence:
• Problem evidence vs. solution evidence
• Source of confusion: conflating evidence about the size of a problem (e.g., the epidemiological link between smoking and
cancer) with evidence about the effectiveness of a solution (e.g., whether higher tobacco taxes reduce consumption).
• Scientists who identify problems are not necessarily best positioned to design solutions.
• Issue of rationality
• Comprehensive rationality (an idealised model where policymakers have clear preferences, full information,
and choose optimally) and bounded rationality (the reality where policymakers face unclear aims, limited
information, and rely on heuristics).
• Evidence now taken for granted—such as the link between tobacco and disease—has taken decades to be
accepted within government.
Assumptions of comprehensive rationality
• societal values are not faithfully reflected in policymakers’ values
• shared power: small number of central actors do not control the process
• 1970s: Ground water (GW) overexploitation acknowledged in the agenda [Model Groundwater
Bill (1970, revised 1972)]
• Low adoption rates (exception: TN)
• Top barriers
• poor access to timely research
• mutual mistrust between researchers and policymakers
• policymakers’ lack of research skills
• Top facilitators
• researcher-policymaker collaboration
• clear, accessible outputs
“….we find that as documents become more publicly accessible, they increasingly
communicate doubt. This discrepancy is most pronounced between advertorials and
all other documents. For example, accounting for expressions of reasonable doubt,
83% of peer-reviewed papers and 80% of internal documents acknowledge that
climate change is real and human-caused, yet only 12% of advertorials do so, with
81% instead expressing doubt. We conclude that ExxonMobil contributed to
advancing climate science—by way of its scientists’ academic publications—but
promoted doubt about it in advertorials.”
Evidence != influence on policy (despite internal scientific consensus!)
• Why:
• Research is a critical input to effective policies, yet…
• Research != policy making
Qual Research 101
Patton (2015)
• Core types of qualitative data
1. Design strategies
Criterion-i To identify and select all cases that meet some predetermined criterion of
importance
Criterion-e To identify and select all cases that exceed or fall outside a specified criterion
Maximum variation Important shared patterns that cut across cases and derived their significance
from having emerged out of heterogeneity
Strategy Objective
However, we don’t observe the same unit in two states of the world!
Quant Research 101
Causality
For unit i:
On average, how much would the outcome change if the whole population received the treatment
instead of none receiving it?
Average Treatment Effects: ATE
• average causal effect of the treatment for units that actually receive the
treatment.
• It measures the expected difference between:
• the outcome for treated units with treatment 𝑌 1 ,and
• the outcome those same units would have had without treatment 𝑌 0 .
• Thus, ATT focuses only on the subpopulation that participates in the program.
For those who actually received the treatment, how much did the treatment
change their outcomes on average?
Average Treatment Effects: ATT
• average causal effect that the treatment would have had for units that did not
receive the treatment.
• It measures the expected difference between:
• the outcome those units would have had if treated 𝑌 1 ,and
• the outcome they actually experienced without treatment 𝑌 0 .
• Thus, ATC focuses on the subpopulation that did not participate in the program.
If the untreated group had received the treatment, what would their average
treatment effect have been?
Average Treatment Effects: ATC
For households that did not participate in MGNREGA, how much
would their income have changed if they had participated?
• Selection bias
• treatment status is correlated with potential outcomes i.e. treated and
untreated groups may differ systematically even in the absence of
treatment
• Households participating in MGNREGA (D=1)
• Households not participating (D=0)
• Observed difference:
• 𝐸 Income ∣ 𝐷 = 1 − 𝐸 Income ∣ 𝐷 = 0
• Design
• subject to uncertainty
Minimum Wages and Employment: A Case Study of the Fast-Food
Industry in New Jersey and Pennsylvania (Card, David & Alan Krueger,1994),
American Economic Review, 84(4): 772–793)
Banerjee, Abhijit; Esther Duflo; Rachel Glennerster; and Dhruva Kothari (2010).
“Improving Immunisation Coverage in Rural India: Clustered Randomised
Controlled Evaluation.” BMJ, 340:c2220.
Quasi-Experimental Designs
Was the Card & Krueger study a true experiment? Why or why not?
For example…
▪ Does a job-training program reduce unemployment?
• “Air Quality, Infant Mortality, and the Clean Air Act of 1970.”
Quarterly Journal of Economics, 118(3), 1121–1167.
Sampling in the Logic of Scientific Inference
• ability to generalize findings depends on how that subset is selected
• Evaluating a national job-training program with 10 million workers.
• Coverage error: occurs when the sampling frame does not fully represent
the target population
INFEREN
ANALYSI CE
MEASUR S
OBTAINE EMENT /
D CODING
SELECTE
D RESPON
SAMPLIN SES /
G SAMPLE
SAMPLIN OBSERVA
G FRAME DESIGN TIONS
TARGET
POPULAT
ION
Sampling Frames
• operational list of units from which a sample is drawn
• urban residents – municipal address registry
• students – school enrollment lists
• voters – electoral rolls
• must satisfy:
1. completeness (all units included)
2. accuracy (correct classification)
3. no duplication (each unit appears once)
Probability Sampling
• every population unit to have a known, non-zero probability of
selection
• unbiased estimation and robust statistical inference
• Say,
Population = 10,000 households
Sample = 500 households
Probability of selection = 0.05
• Standard Error
• estimated standard deviation of the sampling distribution of a statistic
• Think: how much a sample estimate would vary if we repeatedly drew samples
from the same population (sampling variability)?
• CI = 𝜃መ ±zα/2×SE(𝜃መ )
• 90% - 1.64
• 95% - 1.96
• 99% - 2.58
1. Simple Random Sampling
• Each possible sample of size n has an equal probability of being
selected.
• How?
1. List population units
2. Assign numbers
3. Use random selection
• Properties:
- Unbiased estimator of population mean
- Sampling error decreases with larger sample sizes (n increases)
2. Systematic Sampling
• Units are selected at regular intervals from a list.
• How?
1. Calculate interval k = N/n
2. Choose a random starting point
3. Select every kth unit
• Example:
Population = 10,000
Sample = 200
Interval k = 50
3. Stratified Sampling
• Population divided into subgroups (strata) such as gender, region,
or income group.
• Types:
• Proportionate stratification – sample mirrors population proportions.
• Disproportionate stratification – some groups oversampled for analytical
purposes.
4. Cluster Sampling
• Clusters are natural groups like villages, schools, or city blocks.
Clusters are sampled first, and units within clusters are then
surveyed.
• Advantages:
- Reduced cost
- Easier fieldwork
• Disadvantages:
- Lower statistical precision due to intra-cluster correlation.
5. Multistage Sampling
• large surveys sample in multiple stages
• Example:
Stage 1 – districts
Stage 2 – villages
Stage 3 – households
• Common modes:
Face-to-face surveys
Telephone surveys
Mail surveys
Online surveys
External validity
RECAP:
Reliability = consistency of measurement
Validity = accuracy of measurement
Indian datasets
Design:
1) RCT in chosen villages;
2) Observational at national/state level (NSSO panel data);
3) Qualitative data analysis in chosen village (farmer FGDs).