0% found this document useful (0 votes)
3 views26 pages

Sampling Methods in Public Policy Research

Uploaded by

sendadds2
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views26 pages

Sampling Methods in Public Policy Research

Uploaded by

sendadds2
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

PLCY 750: Research Methods for Public

Policy

Data Collection Methods: Sampling I

E. Johnson
09-09-25
What happened?

2
Sampling matters

3
Probability and non-probability
sampling
Types of non-probability sampling
• Voluntary sampling
• “Do you have a minute to talk about the environment?”
• Convenience sampling
• “UNC Psychology study seeks student participants”
• Snowball sampling
• Chain sampling or respondent driven sampling
• Gatekeepers
• Organized crime, drug users, homeless
• Quota sampling
• Need a certain number from each group (e.g. freshman, sophomores, etc.)

4
Probability and non-probability
sampling
• Probability sampling
• Uses chance to select units in sample
• Like drawing random numbers from a bingo
wheel
• Basis for the methods we use to assess
precision of quantitative research results

5
Probability sampling procedures
• Simple random sampling (SRS)
• Systematic sampling
• Stratified sampling
• Use if you need to ensure even distribution across sub groups
• Use if there is similarity within subgroups and differences
between subgroups on key variables of interest?
• Enables disproportionate sampling
• Cluster sampling
• Save time/money/resources by randomly selecting clusters
from a more manageable frame then sample from within
clusters only

6
Random Selection vs. Random
Assignment
External validity is the degree to which the results of a study can be generalized beyond the study sample to other populations,
settings, or [Link] answers: “If this works here, will it also work elsewhere (or for the broader group we care about)?

Theoretical
Population

External related to
Random Selection
Validity generalizability

Evaluation
Sample

Random Random Internal


Assignment Assignment Validity
not dependent on sample
Treatment Control it has to do with causal inference
Group Group

Internal validity refers to the extent to which a study can demonstrate a clear, trustworthy causal relationship between an
independent variable (cause) and a dependent variable (effect), without interference from confounding factors
Generalizability
• Definition?
External validity is the degree to which the results of a study can be generalized beyond the study sample to other populations,
settings, or [Link] answers: “If this works here, will it also work elsewhere (or for the broader group we care about)?

• How confident are you that your results represent the population
of interest?
• (do you even have a clear idea of who your target population are?)

• Why are we talking about generalizability in a class on sampling?


We haven’t even gotten to methods yet!

Sampling frame:

A sampling frame is the list or source from which a sample is drawn to represent a population

8
Basic concepts in sampling

9
Basic concepts in sampling--
Population: parameter :: Sample:
statistic
• Population Parameters:
• μ - population mean
• σ² - population variance
• π - population proportion

• Sample Statistics: Or called estimate

• X̄ - sample mean (estimator/estimate)


• s² - sample variance
• p̂ - sample proportion

• Population: parameter :: Sample: statistic

10
Basic concepts in sampling--
Population: parameter :: Sample:
statistic

• Standard Deviation: σ - population standard deviation


s - sample standard deviation
• Standard Error: SE - standard error SE(X̄) - standard
error of the sample mean SE(p̂) - standard error of the
sample proportion
11
Population of Interest
• How to define?
• Not always a human population
• Can be a universe of entities or collective entirety of some measurable unit,
e.g. an ocean, forest, airshed
• May not be large – e.g. population of UNC freshman vs. population of all
U.S. undergraduates vs. all U.S. adults

• Sample is drawn from population


• Inference is made based on statistics that estimate population parameters
(numerical characteristic of a population which we do not measure)

12
Identify target population and a
possible descriptive research
question
Sampling Frames:

1. Registrar’s list of all UNC students


2. All UNC students entering the Dean Dome for a game
3. List of all registered voters
4. List of all U.S. electric utility companies

13
Bias – systematic differences
between the sample and population
• Sampling Bias
1. Coverage Bias Occurs
frame.
when some members of the population are systematically excluded from the sampling

2. Non-response bias Happens when selected participants do not respond, and their non-participation is
related to the study topic.

3. Response bias – respondent answers Occurs when participants respond


inaccurately, dishonestly, or misleadingly.
incorrectly/dishonestly
• Poorly worded questions
• Social desirability bias
• Bradley and Hawthorne effects
• Explain how each may have contributed to the polls’
general failure to predict a Trump victory
• When do coverage and nonresponse problems cause bias?
14
Response rate

• Response rate = (contact rate)*(cooperation rate)


• Contact rate = share of frame/sample reached with successful
opportunity to participate
• Cooperation rate = share of those contacted who participated

• Example
• 50% of addresses from my sampling frame agree to be
surveyed
• Only 60% of those who agree actually complete the survey
• Response rate = 30%

Johnson PLCY 581 - 2018


15
Non-response bias
A causal model here tries to explain the likelihood (propensity) that an individual will respond to a survey based on different determinants.
Broadly, it looks like:𝑃(Response)=𝑓(Respondent characteristics,Survey design,Contextual factors)P(Response)=f(Respondent
characteristics,Survey design,Contextual factors)

16
X could also refer to propensity to
respond (or coverage propensity)
• Z here is something that jointly explains both the
outcome and the propensity to respond
P Y

Corr(x,z)>0 corr(x,z)<0
Corr(y,z)>0 Positive (upward) bias Negative (downward)
bias
Corr(y,z)<0 Negative (downward) Positive (upward) bias
bias

17
Coverage Bias
• Definition
• Sampling frame is systematically different from the population
• Ex) List of registered voters may be wealthier and older than those planning
to register on election day
• When are we worried about it?
• When the systematic differences are related to the outcome, e.g. electoral
preferences
• What can we do about it?
• See if you can figure out what it is that’s driving the propensity to be
covered and the outcome
• TRY to correct the sample frame or weight responses according to other
data on the population (not usually possible– Need measure of Z!!!)

18
Non-response Bias
• Definition
• Propensity to respond is related to what the survey is trying to measure
• Ex) people who recycle are more likely to respond to survey about their
recycling habits
• What can we do about it?
• See if you can figure out what it is that’s driving the propensity to respond
• Be clear about direction of bias and what might be causing it
• weight responses according to other data on the population

19
Probability and non-probability
sampling
Types of non-probability sampling
• Voluntary sampling
• “Do you have a minute to talk about the environment?”
• Convenience sampling
• “UNC Psychology study seeks student participants”
• Snowball sampling
• Chain sampling or respondent driven sampling
• Gatekeepers
• Organized crime, drug users, homeless
• Quota sampling
• Need a certain number from each group (e.g. freshman, sophomores, etc.)

20
Cluster vs. Stratified: Differences
• In cluster sampling, the cluster is treated as the sampling unit
so analysis is done on a population of clusters (at least in the
first stage).

• In stratified sampling, the analysis is done on elements within


strata.

• In stratified sampling, a random sample is drawn from each of


the strata, whereas in cluster sampling only the selected
clusters are studied.

• The main objective of cluster sampling is to reduce costs by


increasing sampling efficiency. This contrasts with stratified
sampling where the main objective is to ensure coverage of
key strata and to increase precision by reducing sampling error.
21
Cluster vs. Stratified

22
Clarifying Example: need random
sample of Durham households
Stratified Cluster

Steps 1) Divide city into districts 1) Divide city into districts


(strata). (clusters).
2) Draw random sample of 2) Draw random sample of
households from each districts.
district. 3) Draw random sample of
households from each
district.

Reason To ensure desired number of To make it easier (and less


for Use households in each district costly) to do door-to-door
(and increase precision). surveys.

23
Clarifying Example #2: need random
sample of UNC undergrads
Stratified Cluster

Steps 1) Define strata for UNC 1) Define clusters as UNC


students by year/level classes.
(e.g., 1st, 2nd, 3rd, 4th, , and 2) Draw a random sample of
“Super seniors”). classes from the course
2) Draw a random sample of schedule.
students from within each 3) Administer survey to all
year. students in the randomly
selected classes.

Reason for To ensure desired number of To make it easier (and less


Use students in each year (and costly) to distribute surveys.
increase precision).

24
UNC 2016 Class Profile
All reported races/ethnicities:

• Asian/Asian American 14%


• Black/African American 11%
• Caucasian/White 71%
• Hispanic/Latino/Latina 7%
• Native American 2%
• Pacific Islander 0.1%

25
Post-stratification adjustment -
Weighting
• Why?
• Probability of being in the sample is not the same for all strata (e.g. in
oversampling)
• But you need your statistics to represent the whole pop.
• How?
• Compare select characteristics of the sample with the known values of
those characteristics in the population.
• Correct estimates with weights. (lowers effective sample size = likely
loss of precision because it will change your standard errors)

• Example: Females often respond to surveys at a higher rate


(0.60) than men. But, females only represent 0.50 of the
underlying population. So, down-weight their responses by
0.50/0.60 = 0.833 (when computing whole-sample statistics,
e.g., means)
26

You might also like