Lecture 1: Sampling Methods
WEEK 2
From Sample to
Population
Recall the basic process of research:
1. Collect a sample from your population of interest
2. Calculate statistics from your sample that attempt to answer your
research question(s)
3. Use the statistics calculated on your sample (e.g. mean, proportion,
corrleation) as estimates of the true population values
• e.g. the mean height in your sample is used to estimate the mean height
in the entire population of interest
• Generalize your study results to the entire population of interest
Key assumption: your sample is representative of the population
“Random Sampling Error” and “Bias”
2 reasons why the value of a statistic calculated from a sample differs from the
true population value:
1. “random sampling error”: by chance, the sample you collect might have a
higher/lower mean than that of the entire population
• Suppose the population has 1 million members, and you recruit a sample of 1000
subjects. Your friend recruits a different set of 1000 subjects. By chance, the
mean in your sample will likely differ from the mean in your friends sample, and
both of your means will likely differ from the true population mean.
• As your sample size gets larger, the statistics calculated from your sample will get
closer to the true population values (i.e. “random sampling error” will diminish)
2. Bias in the study design (there are many types of bias!)
When is your sample NOT representative?
Selection Bias: a systematic tendency to favor including subjects with particular
characteristics while excluding other types of subjects
i.e. some members from the population of interest were more likely to be
included in your sample compared to others
Suppose your population of interest is all U.S. citizens, and you use telephone
calls to administer a survey about an upcoming election poll
• What about Americans who do not own a telephone?
• What if the type of people who respond are very different from the type
of people who do not respond? “Non-response bias”
How to ensure the sample is representative?
Randomly sampling (“selecting”) subjects from the population of interest
tends to produce samples that are more representative of the population
4 main types of “Random sampling” methods:
• Simple random sampling
• Systematic random sampling
• Stratified random sampling
• Cluster random sampling
“Simple” Random Sampling
Identify and create a list of all possible subjects in the population of interest.
Then use a CPU to randomly select “N” number of subjects from the list
• Requires that all members of the population are identifiable and able to be
contacted (not always possible!)
Each member of the population has an equal probability of being selected
Since all members of the population are equally likely to be chosen, the sample is
representative of the entire population
“Systematic” Random Sampling
Select every “k”th subject from the population of interest
• recruit every 10th patient that visits a hospital clinic
• Give a survey to every 20th customer that enters a store
Unlike simple RS, systematic RS does not require that every member of the
population be identified/listed
• Although if you did have a list, you could simple recruit every kth name on the list
Potential bias if there are ordered or cyclical patterns:
• E.g. time of day or days of the week: if you only recruit patients on Mon. and
Wed., or only in the mornings, you might underrepresent certain types of patients
who tend to go at other times
“Stratified” Random Sampling
Divide the population into groups (“strata”) with similar characteristics. Then
take a simple random sample from each group.
Common stratification variables: sex, age group, race, or hospital clinic
Used when the researcher wants to ensure a representative sample from
each subgroup (e.g. to improve accuracy when comparing groups or making
inferences about particular minority groups)
Requires information about the stratification variable(s) for all potential
subjects in the population (may not be easily available)
“Cluster” Random Sampling
Divide population into “clusters” and take a random sample from each cluster
Clusters are usually defined geographically: e.g. you might divide a city into
neighborhoods, and neighborhoods into blocks, then blocks into houses
• Randomly choose a set of neighborhoods, then randomly choose sets
of blocks within each neighborhood, then randomly choose sets of
houses within each block, then send mail surveys to these houses
Similar to systematic RS, cluster RS does not require that all members of the
population be identified/listed
Non-Random Sampling
“convenience sample”: participants who are “readily available” are enrolled until
the desired sample size is achieved
No other criteria to the sampling method except that people be available and
willing to participate
E.g. trying to recruit every patient that walks into a hospital clinic
Putting study participation flyers in elevators or other common areas
Problem: not all members of the population have the same probability of being
included in the study the sample may not be representative of the population
Subjects who are “readily available” may differ from subjects in the population
who are not readily available
Summary
The value of a statistic calculated from your sample can differ from the true
population value due to 1) random sampling error, or 2) bias in study design
A major cause of bias occurs when your sample is not well representative of
the population of interest: “selection bias”
The 4 “random sampling” methods help ensure that the sample is
representative of the population of interest
Non-random “convenience sampling”: greater risk for selection bias
Most journal articles will have a “Table 1” that describes characteristics of
their sample. This will give you an idea of the actual type of “population” that
their study results can generalize to.
Paper 1: From your homework 1
Activity 15 (Polleverywhere):
Which of the sampling schemes
was used to obtain the sample
in this paper?