0% found this document useful (0 votes)
9 views105 pages

RM - Module 1

The document discusses research methods for public policy, focusing on the evaluation of programs like Swachh Bharat Mission (SBM) in India and Progresa in Mexico. It emphasizes the importance of counterfactuals in assessing program impacts, the distinction between prospective and retrospective evaluations, and the need for evidence-based policymaking. Additionally, it highlights challenges in the policy process, including the complexity of policymaking environments and barriers to using evidence effectively.

Uploaded by

divya.reddy
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
9 views105 pages

RM - Module 1

The document discusses research methods for public policy, focusing on the evaluation of programs like Swachh Bharat Mission (SBM) in India and Progresa in Mexico. It emphasizes the importance of counterfactuals in assessing program impacts, the distinction between prospective and retrospective evaluations, and the need for evidence-based policymaking. Additionally, it highlights challenges in the policy process, including the complexity of policymaking environments and barriers to using evidence effectively.

Uploaded by

divya.reddy
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Research Methods for Public Policy

Module 1
Did Swachh Bharat reduce diarrhoea?
Attribution
• Development programs: typically designed to change outcomes

• Measurement:
• Usually inputs and outputs are tracked
• whether intended outcomes actually improve?
• changes in well-being that can be attributed to a particular project, program, or policy

• Evidence-based policymaking
• results guide target-setting, accountability, budget allocation, and program design

• Counterfactual:
• what would have happened to participants had they not participated?
[valid comparison group]

• What is the counterfactual outcome we need for assessing SBM?


• same households, same time, but no toilets
SBM

• Program goal: Eliminate open defecation in rural India by providing


household toilets and promoting sanitation behaviours

• Key Government-Reported Outcomes


• Toilet Construction and Coverage (Government Reports)
• all villages self-reported/declared themselves ODF by 2 Oct 2019
• over 11 crore toilets and 2.23 lakh community sanitary complexes built across
states/UTs since 2014
• ODF
• More than 6 lakh villages were declared ODF by 2019 based on SBM dashboards and
official statements.
• Usage (National Annual Rural Sanitation Survey (2019–20))
• ~95.2% of rural households with access to a toilet reported using it
• ~99.6% of these households had water available for toilet use
Why counterfactual matters

• Measuring inputs/outputs (toilets built) does not alone show


health or behavioural impact.

• Counterfactual: What health, sanitation behaviours, and disease


outcomes would the same households have had without SBM-
induced toilets and sanitation promotion?
• same households, same time, but no toilets/ODF intervention
Evaluation Design Map
• Timing: prospective and retrospective

• Setting: efficacy and effectiveness

• Purpose: impact/process/cost/simulation/monitoring
‘How’ of evaluations
• Prospective evaluations
• evaluation designed ex ante; baseline before exposure; pre-specified outcomes
with counterfactual strategy

• Retrospective evaluations
• evaluation designed ex post; counterfactual reconstructed from observational
variation / secondary data

Which is superior? prospective?


• baseline data to verify group comparability
• defining success measures early forces programmatic clarity
• treatment/comparison groups can be identified before assignment
occurs
SBM

• Prospective:
• Government pilots SBM in selected districts first; collect pre and post
rollout OD and health data

• Retrospective
• Researchers use NFHS 2015–16 and 2019–21 to compare districts with
differing SBM exposure
‘How’ of evaluations
• Efficacy studies
• small-scale, tightly controlled tests proving a concept can work under ideal
conditions (think “proof of concept”)
• strong internal validity

• Effectiveness studies
• generate evidence under real-world implementation conditions
• greater external validity

The distinction matters because a program that works in a controlled trial may
fail when implemented by overburdened government bureaucracies.
‘How’ of evaluation
• Five approaches that complement impact evaluation

• monitoring
• provides administrative data to verify implementation fidelity and dosage

• ex ante simulations
• project likely effects before launch

• process/implementation evaluations
• examine how a program is implemented
• Often mixed methods

• cost-benefit and cost-effectiveness analyses


• compare costs against returns
‘When’ of evaluation
• Relevance
• the program must be strategically relevant

• Knowledge
• capable of generating significant learning; addressing a genuine knowledge gap

• Implementation
• supported by operational integrity and tied to decision-making commitments

• Feasible
• time and budget constraints
Criterion SBM

SBM is a national flagship mission linked to sanitation, public health,


Relevance
gender dignity, and SDG commitments.

Uncertainty existed about sustained toilet usage, behavior change, and


Knowledge
health impacts (e.g., child morbidity, mortality).

Variation across districts in toilet construction quality, behavior change


Implementation Integrity campaigns, and ODF verification affects whether outcomes can be
meaningfully attributed.

Large-scale datasets (NFHS, Census, SBM dashboard) enable


Feasibility retrospective analysis; however, no nationwide prospective baseline
was embedded in original design.
Progresa, Mexico
• Conditional cash transfer program (1997)
• Oportunidades (2002)
• Prospera (2012-2018)

• Eligible poor and vulnerable households who


• send their children to primary and secondary schools
• whose mothers and children receive regular preventive care at local
health clinics
• grants to improve food consumption and nutritional supplements for
young children and pregnant and lactating mothers

• Reached 20 percent of the total population of Mexico

• Rigorous evaluation built into the rollout


• survived a change in the ruling political party
• contributed to global adoption of CCTs: adopted in 52 countries
‘How’ of evaluation
• Theory of change
• articulates the causal logic through which an intervention is expected to
generate desired outcomes results (aim and how)
• Sequence of change, causal mechanism linking each stage, assumptions for
each link to hold, contextual conditions under which the intervention is
expected to work
• Results chain: Inputs → Activities → Outputs → Outcomes → Final Outcomes
• makes explicit the assumptions and mechanisms linking each stage

• Mechanism experiments: evaluations designed to test specific


causal links rather than overall program effects
• SMART indicators (Specific, Measurable, Attributable, Realistic, Targeted)
• anticipated effect sizes must be carefully considered for sufficient statistical
power
Mexico’s Piso Firme program: replacing dirt floors with cement

• Result chain: cement floors → reduce parasites and dust → improved child health →
reduced maternal stress →increased maternal happiness

• Mechanisms (multiple domains)


• Environmental health
Cement floors reduce soil-transmitted helminths and indoor dust exposure.
• Health-to-psychological
Improved child health reduces caregiving burden and health-related anxiety.
• Psychological well-being
Reduced stress increases reported maternal happiness.

• Embedded Assumptions
• Dirt floors are a significant source of parasites/dust.
• Child illness meaningfully affects maternal stress.
• Maternal happiness is responsive to reductions in stress.
• No major offsetting effects (e.g., maintenance burden, social stigma).
Mechanism Link Specific Measurable Attributable Realistic Targeted

Prevalence of
Lab-confirmed Compare
Cement floors → soil-transmitted Detectable Direct biological
stool sample treatment vs.
Reduced helminth within 6–12 exposure
(binary infection control
parasites infection months pathway
status) households
(children <5)

Random Direct
Improved child Maternal Validated scale Low-cost survey-
assignment psychological
health → Reduced perceived stress (e.g., Perceived based
identifies causal caregiving
maternal stress score Stress Scale) instrument
effect mechanism

Reduced
Self-reported life Standard Randomized
maternal stress → Standard survey Direct well-being
satisfaction (0– subjective well- exposure to
Increased instrument outcome
10 scale) being question intervention
happiness
Quick recap exercise
Rural Girls’ Secondary School Scholarship (monthly)

Policy Goal: Increase secondary school completion among rural girls

• Draw a results chain (Inputs → Activities → Outputs → Outcomes → Final Outcomes).

• Identify:
• the key counterfactual question the evaluation must answer
• at least one link in the chain that you think is most uncertain and would benefit from a
mechanism experiment
• SMART indicators for the final outcome
• Policy context: evaluates the “Graduation” anti-
poverty program (the transfer of a productive asset with
consumption support, training, and coaching plus
savings encouragement and health education and/or
services ) across six countries (Ethiopia, Ghana,
• Policy context: evaluates a school-based Honduras, India, Pakistan, and Peru)
deworming program across 75 primary schools
(~30,000 pupils) in rural Kenya
• Design: six parallel RCTs covering ~10,495 households
• Design: cluster-randomised phase-in. • theory of change: the poverty-trap model
• Schools were randomly assigned to three • measures outcomes at every link in the causal
groups receiving treatment in staggered chain identified from ToC
sequence • consumption, food security, productive and
• at any point in time, not-yet-treated schools household assets, financial inclusion, time
serve as the counterfactual. use, income and revenues, physical health,
mental health, political involvement, and
• Results: 25% reduction in absenteeism in women’s empowerment
treatment schools
• significant cross-school externalities: • Results: statistically significant impacts on 10 key
disease-transmission spillovers outcomes or indices
Four types of comparisons (Tilly, 1997; Macrosociology)
• Individualising comparison
• examines a small number of cases to highlight what makes each distinctive

• Universalising comparison
• aims to show that a phenomenon has essentially identical features across all instances
examined

• Variation-finding comparison
• identifies systematic differences across cases to establish a principle of variation in character or
intensity

• Encompassing comparison
• locates different instances at varying positions within the same overarching system (e.g., the
capitalist world-economy) and explains their characteristics as functions of their relationships
to that system
Type Core Aim What It Does What It Is Not
Compares a small number of
Individualising cases to highlight how each Not aimed at deriving general
Clarify distinctiveness
Comparison follows a distinctive historical laws.
trajectory.
Demonstrates that different cases
Universalising Establish common Not claiming total empirical
share the same underlying
Comparison properties uniformity.
process, structure, or mechanism.
Identifies systematic variation
Variation-Finding Explain patterned across cases to derive a principle
Not merely descriptive contrast.
Comparison differences explaining differences in outcome
or intensity.
Situates cases within a larger
interconnected system and
Encompassing Explain relational Not treating cases as
explains their features as
Comparison position independent units.
consequences of their structural
position within that system.
Classify the Comparisons
A historian examines the
A political scientist argues
unique institutional features
that all democratic
of Indian federalism to
transitions share three
explain why India’s
common features: elite
democratic trajectory differs
splits, mass mobilisation,
from other post-colonial
and negotiated pacts.
states.

A world-systems scholar
A researcher compares land explains why East Asian
reform outcomes across five states industrialised while
Latin American countries to Sub-Saharan African states
identify which factors explain did not by locating both
variation in redistribution. within the global division of
labour.
A Tale of Two Cultures (Mahoney & Goertz, 2006)
• Quantitative and qualitative research not points on a single continuum
• distinct methodological cultures with different values, beliefs, and norms.
• each tradition’s practices are internally coherent

• Goal: enhance cross-tradition communication

• Key differences
• Effects-of-causes vs causes-of-effects
• Additive regression vs Boolean combinations
• Population inference vs case explanation
• Probabilistic vs necessary/sufficient causation
Evidence-based Policymaking (EBPM)
• vague, aspirational term rather than a description of how policy actually works
• “Policy” encompasses the sum total of government action from signals of intent to final outcomes.
• “Evidence” is an argument or assertion backed by information, with scientific evidence occupying a privileged
but contested position atop a hierarchy of methods.
• “Policymakers” are not a monolithic group but a dispersed population operating within policy communities.

• Issue of evidence:
• Problem evidence vs. solution evidence
• Source of confusion: conflating evidence about the size of a problem (e.g., the epidemiological link between smoking and
cancer) with evidence about the effectiveness of a solution (e.g., whether higher tobacco taxes reduce consumption).
• Scientists who identify problems are not necessarily best positioned to design solutions.
• Issue of rationality
• Comprehensive rationality (an idealised model where policymakers have clear preferences, full information,
and choose optimally) and bounded rationality (the reality where policymakers face unclear aims, limited
information, and rely on heuristics).
• Evidence now taken for granted—such as the link between tobacco and disease—has taken decades to be
accepted within government.
Assumptions of comprehensive rationality
• societal values are not faithfully reflected in policymakers’ values

• shared power: small number of central actors do not control the process

• facts and values cannot be cleanly separated

• organisations do not comprehensively search for information

• policymaking does not follow a linear sequence


• often solutions exist before problems arise (the “garbage can model”)
• problems, solutions, participants, and choice opportunities happen to meet
Policy cycle as heuristic
agenda setting → formulation → legitimation → implementation → evaluation

• Misleadingly simple description of how policy is made

• Real-world: interacting cycles

• Non-linear, multi-level and multi-decade, recursive policy making

• 1970s: Ground water (GW) overexploitation acknowledged in the agenda [Model Groundwater
Bill (1970, revised 1972)]
• Low adoption rates (exception: TN)

• Subsequent feedbacks: Model Bills in 1992, 1996, and 2005


• uneven adoption amid protests (e.g., Karnataka)
• Atal Bhujal Yojana (2019) and Jal Jeevan Mission integrations for recharge/monitoring under ongoing federal
tensions.
Key tenets for researchers seeking to
influence policies
• Policymakers: boundedly rational, relying on cognitive shortcuts

• Policymaking environment: complex and unpredictable

• Distinction between uncertainty (incomplete information, partially


addressable by better evidence) and ambiguity (fundamentally
different ways of understanding a problem, not solvable by more
data)
Oliver, K., Innvar, S., Lorenc, T., Woodman, J. & Thomas, J. (2014). “A
systematic review of barriers to and facilitators of the use of evidence by
policymakers.” BMC Health Services Research, 14, 2.
• Systematic review: 145 studies across health, criminal justice, transport, and drug
policy

• Top barriers
• poor access to timely research
• mutual mistrust between researchers and policymakers
• policymakers’ lack of research skills

• Top facilitators
• researcher-policymaker collaboration
• clear, accessible outputs

Policymakers cannot process all available evidence: bounded rationality;


Researchers and policymakers define “evidence” differently: uncertainty/ambiguity distinction
Supran, G. & Oreskes, N. (2017). “Assessing ExxonMobil’s
Climate Change Communications (1977–2014).”
Environmental Research Letters, 12(8)
• Systematic content analysis of 187 ExxonMobil communications spanning nearly
four decades

“….we find that as documents become more publicly accessible, they increasingly
communicate doubt. This discrepancy is most pronounced between advertorials and
all other documents. For example, accounting for expressions of reasonable doubt,
83% of peer-reviewed papers and 80% of internal documents acknowledge that
climate change is real and human-caused, yet only 12% of advertorials do so, with
81% instead expressing doubt. We conclude that ExxonMobil contributed to
advancing climate science—by way of its scientists’ academic publications—but
promoted doubt about it in advertorials.”
Evidence != influence on policy (despite internal scientific consensus!)

Powerful actors can manufacture ambiguity competing interpretations of a problem to exploit


bounded rationality in the policy process.
Logic of inquiry
• What does it mean to draw valid inferences about policy effects?

• What are the different traditions of explanation in the social


sciences?

• And what happens when rigorous evidence meets the messy


realities of policymaking?
Aim of the course
• What: Research thinking

• How: Tools of research


• Basics of qualitative methods
• Basics of quantitative methods
• Intro to mixed methods

• Why:
• Research is a critical input to effective policies, yet…
• Research != policy making
Qual Research 101
Patton (2015)
• Core types of qualitative data

• detailed observations (field notes describing activities, behaviours, and


settings)

• in-depth open-ended interviews (yielding direct quotations)

• documents/artifacts (written materials and records)

themes, patterns, concepts, and insights


Strategic Themes in Qualitative Inquiry
12 strategic themes organised into three groups

1. Design strategies

• naturalistic inquiry (studying real-world situations without


manipulation)

• emergent design flexibility (adapting as understanding deepens)

• purposeful sampling (selecting information-rich cases for insight


rather than statistical generalisation)
Strategic Themes in Qualitative Inquiry
2. Data collection strategies

• qualitative data (thick description, direct quotations, document review)

• personal experience and engagement (the researcher’s direct


immersion)

• empathic neutrality and mindfulness (seeking understanding without


judgment)

• dynamic systems thinking (attention to process, change, and system


dynamics)
Strategic Themes in Qualitative Inquiry
3. Analysis and reporting strategies

• unique case orientation (respecting individual case details before cross-case


analysis)

• holistic perspective (understanding phenomena as complex systems)

• inductive analysis & creative synthesis (immersion in specifics to discover patterns)

• context sensitivity (placing findings in social and historical context)

• voice, perspective, and reflexivity (awareness of the analyst’s own perspective)


Additional recommended reading for those interested in qualitative
research: Patton (2015) Chapters 4 & 5
Purposeful Sampling (Palinkas et al., 2015)
• identification and selection of cases that are “information rich” for
the most effective use of limited resources

• knowledgeable about or experienced with a phenomenon of interest

• availability and willingness to participate

• ability to communicate experiences and opinions in an articulate,


expressive, and reflective manner
Strategy Objective

Criterion-i To identify and select all cases that meet some predetermined criterion of
importance
Criterion-e To identify and select all cases that exceed or fall outside a specified criterion

Typical case To illustrate or highlight what is typical, normal or average

Homogeneity To describe a particular subgroup in depth, to reduce variation, simplify analysis


and facilitate group interviewing
Snowball To identify cases of interest from sampling people who know people that generally
have similar characteristics who, in turn know people, also with similar characteristics
Extreme or deviant To illuminate both the unusual and the typical
case
Intensity Same objective as extreme case sampling but with less emphasis on extremes

Maximum variation Important shared patterns that cut across cases and derived their significance
from having emerged out of heterogeneity
Strategy Objective

Critical case To permit logical generalization and maximum application of information


because if it is true in this one case, it’s likely to be true of all other cases
Theory-based To find manifestations of a theoretical construct to elaborate and examine the
construct and its variations
Confirming & To confirm the importance and meaning of possible patterns and checking out
disconfirming case the viability of emergent findings with new data and additional cases
Stratified To capture major variations rather than to identify a common core, although
purposeful the latter may emerge in the analysis
Purposeful To increase the credibility of results
random
Opportunistic or To take advantage of circumstances, events and opportunities for additional
emergent data collection as they arise
Convenience To collect information from participants who are easily accessible to the
researcher
…versus sampling in quant

• probabilistic or random sampling

• ensure the generalizability of findings (minimizing the potential for bias


in selection and to control for the potential influence of known and
unknown confounders)

• established formulae for avoiding Type I and Type II errors

Think generalizable and breadth


Most Significant Change
• Participatory evaluation method
• Broad domains of change & reporting period • Reviewers
• Different organisational levels
• Identify stakeholders
• Program participants, community members,
program staff • Advantages:
• Capture complexities
• Stories of significant change
• Adaptations
• Phrasing questions and structured participation
• ToC, SNA
• Potential mixed methods
• Feedback and verification design

Dialogical, story-based monitoring and evaluation


• Identify which changes were most meaningful
MSC

• What would be the


recommended sampling
approach?

• What sample size do you think is


adequate for this method?

• How would you analyse the


stories collected?
Bennett's Hierarchy (Bennett, 1976)
MGNREGA reduces rural poverty.

• Is this descriptive, predictive, or causal?

• What must be true for this to be causal?


• “what-if highway” (Huntington-Klein)

However, we don’t observe the same unit in two states of the world!
Quant Research 101
Causality
For unit i:

𝑌𝑖 (1)= outcome if treated

𝑌𝑖 (0)= outcome if not treated

Causal effect: τ_i = Y_i(1) − Y_i(0)

• We typically only observe one of these in the real world.


• Causal inference = solving a missing data problem
• Empirical work estimates average causal effects across groups.
Average Treatment Effects

Average Treatment Effect (ATE)

ATE = E[Y(1) − Y(0)]

• average causal effect of a treatment across the entire population of interest


• It measures the expected difference between:
• the outcome if everyone in the population were treated 𝑌 1
• the outcome if everyone in the population were untreated 𝑌 0 .
• Equivalently, it is the expected value of the individual treatment effect across all units in the
population.

On average, how much would the outcome change if the whole population received the treatment
instead of none receiving it?
Average Treatment Effects: ATE

If every rural household participated in MGNREGA instead of none


participating, what would the average income change be?

ATE = E[Income if household participates] − E[Income if household does not participate]

• Why This Is Difficult to Estimate


• For each household we observe only one outcome (counterfactual is missing):
• Income with participation, or
• Income without participation.
• Estimating ATE therefore requires credible comparisons (e.g., random assignment or valid research
designs).
Average Treatment Effects

Average Treatment Effect on the Treated (ATT):

ATT = E[Y(1) − Y(0) | D=1]

• average causal effect of the treatment for units that actually receive the
treatment.
• It measures the expected difference between:
• the outcome for treated units with treatment 𝑌 1 ,and
• the outcome those same units would have had without treatment 𝑌 0 .
• Thus, ATT focuses only on the subpopulation that participates in the program.

For those who actually received the treatment, how much did the treatment
change their outcomes on average?
Average Treatment Effects: ATT

For households that actually participated in MGNREGA, how much


higher (or lower) was their income compared to what their income
would have been if they had not participated?

For participating households, we observe income with


participation, but their income without participation is
counterfactual.
Average Treatment Effects
Average Treatment Effect on the Control (ATC)
𝐴𝑇𝐶 = 𝐸 𝑌 1 − 𝑌 0 ∣ 𝐷 = 0

• average causal effect that the treatment would have had for units that did not
receive the treatment.
• It measures the expected difference between:
• the outcome those units would have had if treated 𝑌 1 ,and
• the outcome they actually experienced without treatment 𝑌 0 .
• Thus, ATC focuses on the subpopulation that did not participate in the program.

If the untreated group had received the treatment, what would their average
treatment effect have been?
Average Treatment Effects: ATC
For households that did not participate in MGNREGA, how much
would their income have changed if they had participated?

For these households we observe income without participation, but


income with participation is counterfactual.
Identification issue
The difficulty in estimating these quantities arises because:
• For treated units we observe: 𝑌𝑖 = 𝑌𝑖 1
• For untreated units we observe: 𝑌𝑖 = 𝑌𝑖 0

But the estimands require expectations involving both potential


outcomes.

missing counterfactual problem


Selection Bias
Observed differences between treated and untreated groups combine
causal effects with selection bias.

Observed difference: E[Y | D=1] − E[Y | D=0]

This equals: ATT + Selection Bias

Selection bias arises when treated and untreated groups differ


systematically even in the absence of treatment.
Selection Bias and ATT
E[Y∣D=1] − E[Y∣D=0] = E[Y(1) − Y(0)∣D=1] + E[Y(0)∣D=1] − E[Y(0)∣D=0]

E[Y∣D=1] − E[Y∣D=0] = E[Y(1) − Y(0)∣D=1] + E[Y(0)∣D=1] − E[Y(0)∣D=0]

• Selection bias
• treatment status is correlated with potential outcomes i.e. treated and
untreated groups may differ systematically even in the absence of
treatment
• Households participating in MGNREGA (D=1)
• Households not participating (D=0)

• Observed difference:
• 𝐸 Income ∣ 𝐷 = 1 − 𝐸 Income ∣ 𝐷 = 0

• This difference reflects:


• the causal effect of MGNREGA on income, plus
• differences between households that choose to participate and those that do
not (e.g., poorer households may be more likely to participate)

Thus the comparison cannot be interpreted as the causal effect without


addressing selection bias.
Random Assignment
Randomized experiments eliminate selection bias because treatment
assignment is independent of potential outcomes.
• 𝑌 1 ,𝑌 0 ⊥ 𝐷

• If treatment is randomly assigned:


𝐸 𝑌 0 ∣𝐷=1 =𝐸 𝑌 0 ∣𝐷=0
and
𝐸 𝑌 1 ∣𝐷=1 =𝐸 𝑌 1 ∣𝐷=0
Thus treated and control groups are statistically comparable.
• Suppose villages are randomly assigned to receive early MNREGA
rollout.
• Then comparing:
• 𝐸 Income ∣ MNREGA village − 𝐸 Income ∣ Control village
• provides an unbiased estimate of the average causal effect of
MNREGA on household income.
• Because assignment is random, participating and non-participating
villages are similar on average before the program.
Causal Inference requires…
• Clearly defined treatment
• Clearly defined outcome
• A counterfactual
• An identification strategy
• Defensible assumptions
Recap
• “Midday meals improve learning outcomes.”
• “Air pollution increases infant mortality.”
• “Women’s political reservation improves governance quality.”

For each, answer the following:


• Is it causal?
• What is the counterfactual?
• Identify the likely source of bias.
• Propose one identification strategy.
Muralidharan & Sundararaman, “Teacher Performance Pay:
Experimental Evidence from India” (Journal of Political
Economy, 2011).

• Design

• randomized evaluation of a teacher performance pay


program in government-run rural primary schools in
Andhra Pradesh.

• random assignment ⇒ treated and control schools


comparable “on average.”
Cole, “Do voters demand responsive governments? Evidence
from Indian disaster relief” (Journal of Development
Economics, 2012)
• Design:

• Exogenous shocks: rainfall

• outside political control

• isolate the relationship between:


• disaster shocks
• government relief
• voter responses.
Research Design
• connects a research question to the evidence needed to answer it

• valid descriptive and causal inferences about social phenomena

Research Question → Research Design → Data → Inference

• determines what evidence will be collected, how data will be


measured, which cases will be examined, and how inference will
be drawn from the evidence
Research Design
• For policy
• How can we produce credible evidence about policy problems?

• Does a job-training program reduce unemployment?


• Does air pollution regulation improve public health?
• Do cash transfers reduce poverty?

• Poor research design can produce misleading conclusions.


Logic of Scientific Inference
• observe a subset of the world and use it to infer broader patterns

• Descriptive inference – identifying facts and patterns in the


world
• limited observations?

• Causal inference – identifying cause–effect relationships

• subject to uncertainty
Minimum Wages and Employment: A Case Study of the Fast-Food
Industry in New Jersey and Pennsylvania (Card, David & Alan Krueger,1994),
American Economic Review, 84(4): 772–793)

Why might comparing New Jersey and Pennsylvania approximate


the counterfactual?
Key principles
• Number of observations: increase where possible
• More observations provide more information and improve inference.
• What if only two restaurants were studied?

• Case selection: not based on outcomes


• For example, studying only successful policies creates biased conclusions.

• Explanatory variables: increase variation


• Variation helps identify causal relationships.
• What if all states had identical policies?

• Research design: transparent (ethics!)


• clearly explain data sources, case selection, and analytical methods.
• must be replicable
• Why do KKV argue that increasing the number of observations
improves inference? Are there situations where more observations
might not improve a study?

• What happens if cases are selected based on outcomes?

• How does variation in explanatory variables help identify causal effects?


Experiments
• Strong internal validity for causal effects:
• Random assignment (selection bias)
• Manipulation of treatment
• Comparison between treatment and control groups

Banerjee, Abhijit; Esther Duflo; Rachel Glennerster; and Dhruva Kothari (2010).
“Improving Immunisation Coverage in Rural India: Clustered Randomised
Controlled Evaluation.” BMJ, 340:c2220.
Quasi-Experimental Designs

• In many policy contexts, randomization is impossible or unethical.


• Natural experiments
• Regression discontinuity designs
• Difference-in-differences analysis
• Interrupted time-series designs
These designs attempt to approximate the counterfactual using
observational data.

Was the Card & Krueger study a true experiment? Why or why not?
For example…
▪ Does a job-training program reduce unemployment?

▪ What can be possible research designs?


▪ Experiment
▪ Quantitative observational study
▪ Qualitative study

Comment on validity of the results of each of these.


Recap: Social Science Research (Bryman)
• Research:
• structured but iterative process
• rarely a linear sequence
• revise questions and methods during the process

• Research design links research questions to evidence.


• Descriptive and causal inference
• Experiments provide strong causal evidence but are often infeasible.
• Observational studies require careful design to address bias.
• Both qualitative and quantitative methods contribute to rigorous research.

Reviewing Developing Designing


Identifying a Collecting Analyzing Writing up
the theoretical research
research problem data data findings
literature concepts methods
Review exercise
• “Public understandings of air pollution: the ‘localisation’ of
environmental risk.”
Global Environmental Change, 11(2), 133–145.

• “Lung Cancer, Cardiopulmonary Mortality, and Long-Term Exposure


to Fine Particulate Air Pollution.”
Journal of the American Medical Association (JAMA), 287(9), 1132–
1141.

• “Air Quality, Infant Mortality, and the Clean Air Act of 1970.”
Quarterly Journal of Economics, 118(3), 1121–1167.
Sampling in the Logic of Scientific Inference
• ability to generalize findings depends on how that subset is selected
• Evaluating a national job-training program with 10 million workers.

Can we collect data from everyone? If not, what do we do?

• Population → Sampling Frame → Sampling Design → Sample → Inference


• Bias?
• Does a government cash transfer reduce poverty?
• Population: All households eligible for the program.
• Observed Data: Survey of 1,500 households.
Population
• Population: the full set of units about which the researcher wants
to make claims

• Target Population: the group the researcher intends to generalize


to

• Study Population: the population actually represented in the


sampling frame
Sampling
• Inference:
• assumption that the sample represents the broader population

• Population: all MSMEs in India

• Coverage error: occurs when the sampling frame does not fully represent
the target population
INFEREN
ANALYSI CE
MEASUR S
OBTAINE EMENT /
D CODING
SELECTE
D RESPON
SAMPLIN SES /
G SAMPLE
SAMPLIN OBSERVA
G FRAME DESIGN TIONS
TARGET
POPULAT
ION
Sampling Frames
• operational list of units from which a sample is drawn
• urban residents – municipal address registry
• students – school enrollment lists
• voters – electoral rolls

• must satisfy:
1. completeness (all units included)
2. accuracy (correct classification)
3. no duplication (each unit appears once)
Probability Sampling
• every population unit to have a known, non-zero probability of
selection
• unbiased estimation and robust statistical inference
• Say,
Population = 10,000 households
Sample = 500 households
Probability of selection = 0.05

• enables estimation of sampling variability (standard errors,


confidence intervals).
Recap
• Sampling error= 𝜃መ −θ
• difference between the sample estimate and the true population value

• Standard Error
• estimated standard deviation of the sampling distribution of a statistic
• Think: how much a sample estimate would vary if we repeatedly drew samples
from the same population (sampling variability)?

• For a simple random sample,


ˉ 𝑠
𝑆𝐸 𝑥 =
𝑛
where
𝑠= sample standard deviation
𝑛= sample size
Recap
• Confidence interval:
• Width determined by SE

• CI = 𝜃መ ±zα/2×SE(𝜃መ )
• 90% - 1.64
• 95% - 1.96
• 99% - 2.58
1. Simple Random Sampling
• Each possible sample of size n has an equal probability of being
selected.

• How?
1. List population units
2. Assign numbers
3. Use random selection

• Properties:
- Unbiased estimator of population mean
- Sampling error decreases with larger sample sizes (n increases)
2. Systematic Sampling
• Units are selected at regular intervals from a list.

• How?
1. Calculate interval k = N/n
2. Choose a random starting point
3. Select every kth unit

• Example:
Population = 10,000
Sample = 200
Interval k = 50
3. Stratified Sampling
• Population divided into subgroups (strata) such as gender, region,
or income group.

• Types:
• Proportionate stratification – sample mirrors population proportions.
• Disproportionate stratification – some groups oversampled for analytical
purposes.
4. Cluster Sampling
• Clusters are natural groups like villages, schools, or city blocks.
Clusters are sampled first, and units within clusters are then
surveyed.

• Advantages:
- Reduced cost
- Easier fieldwork

• Disadvantages:
- Lower statistical precision due to intra-cluster correlation.
5. Multistage Sampling
• large surveys sample in multiple stages

• Example:
Stage 1 – districts
Stage 2 – villages
Stage 3 – households

• balances representativeness and logistical feasibility.


Sampling Weights
• When units have different probabilities of selection, weights
correct for unequal representation.

• Weight = (1 / probability of selection)


Sampling weights: examples
Village type Population households (HHs)
Large villages 800
Small villages 200
Total 1000

• Calculate population shares.


• Sampling design: 40 HHs from each village.
• Calculate probabilities of selection.
• Calculate sampling weights.
• Electricity access in sample: 90% in large village; 50% in small village.
• Calculate the weighted estimates.
• Calculate weighted mean.
Sampling weights: examples
Group Population
General caste 700
SC 300
Total 1000

• Calculate population shares.


• Sampling design : 50% in gen; 50% SCs
• Calculate probabilities of selection.
• Calculate sampling weights.
• Toilet access in sample: 80% in gen; 40% SCs
• Calculate the weighted estimates.
• Calculate weighted mean.
Non-Sampling Errors
Coverage incomplete sampling frame
Error
Non-respo selected units decline participation
nse Error
Measurem poorly designed questions
ent Error
Processing data entry or coding mistakes
Error
If respondents misreport income, is that a sampling error?
Survey Mode Effects
• Different data collection modes affect responses.

• Common modes:
Face-to-face surveys
Telephone surveys
Mail surveys
Online surveys

• Would responses differ across survey modes? Why?


• Mode influences response rates and social desirability bias.
Validity
Internal validity

• Whether the causal inference is correct.

External validity

Which validity issue • Whether findings can be generalized to other contexts.


might arise in the
minimum wage study? Construct validity
Why?
• Whether the measurement accurately captures the
theoretical concept.

Statistical conclusion validity

• Whether statistical analysis supports the causal claim.


Reliability
• consistency of measurement.
• Types:
• Test-retest reliability
• Inter-coder reliability
• Internal consistency

• A measure can be reliable but still invalid if it consistently produces incorrect


results.

RECAP:
Reliability = consistency of measurement
Validity = accuracy of measurement
Indian datasets

Dataset Sampling design


stratified two-stage sample
National Family Health Survey (NFHS)
design
Periodic Labour Force Survey (PLFS); National Sample Survey (NSS;
stratified multi-stage design
HCES); Annual Health Survey (AHS)
Longitudinal Aging Survey in India (LASI); Time Use Survey (India);
multistage stratified sampling
District Level Household and Facility Survey (DLHS)
Socio-Economic and Caste Census (SECC); Agriculture Census complete enumeration
Indian datasets
• What to bear in mind:
• Survey weights
• Standard errors (clustered)
• Strata identifiers

• Why multi stage sampling frames are preferred:


• Representativeness
• Census frame
• Cost constraints
• Econometrics friendly!
Back to Qual Research…
Conceptual Frameworks
• map relationships between key variables and guide data
collection and analysis

• typically iterative and may evolve as new evidence emerges during


research.
Sampling issue Consequence

Incomplete sampling frame weakens representativeness; affects external validity (generalizability)

Hard to make design-based statistical inference (cannot accurately


Unknown probability of selection
calculate sampling error and weights); weakens representativeness

Can introduce nonresponse bias if respondents differ systematically


Low response rate / selective nonresponse
from nonrespondents.

Reduces precision; increases standard errors; and widens confidence


Small sample size
intervals.

Reduces precision relative to simple random sampling (SRS) due to


Clustered sample with correlated units intra-cluster correlation; if ignored in analysis, standard errors may be
underestimated.

Can narrow analytic insight and miss important variation, negative


Poorly chosen qualitative sample
cases, or causal processes.

Severely limits representativeness; inference beyond the sample is


Convenience/Snowball sampling
usually unjustified
Summary (Module 1)
Dimension KKV Brady & Collier
Logic of inference Shared across methods
More observations,
What strengthens Process evidence,
explicit design, bias
inference context, case knowledge
reduction
Causal-process
Key evidence Dataset observations
observations
More observations where Purposive selection for
Case selection
possible explanatory leverage
General patterns and Mechanisms and
Comparative strength
bias control contextual explanation
Concluding exercise
Do PM-KISAN direct benefit transfers reduce rural poverty?

Design:
1) RCT in chosen villages;
2) Observational at national/state level (NSSO panel data);
3) Qualitative data analysis in chosen village (farmer FGDs).

You might also like