WHAT IS SAMPLING?
Sampling is the process of selecting a group of individuals from a population in order to
study them and characterize the population as a whole.
It’s a pretty simple idea. Let’s say we want to know something about a population—the
percentage of people in Mexico who smoke, for example. One way to go about this would
be to call up everyone in Mexico (122 million people) and ask them if they smoke. The other
way would be to get a subgroup of individuals together (1,000 people, for example) and ask
them if they smoke, and then use this information as an approximation of the information
we really want. This group of 1,000 people who make it possible for us to understand the
behavior of Mexicans in general is called a sample, and the way we select them is
called sampling.
In the above definition, we used two new terms that will be essential throughout this series
of posts:
Population (sometimes called a universe): All of the individuals that we want to study or
characterize. In the example above, our sample is the population of Mexico, but we could
talk about any sort of population, however broad or specific. For example, if we want to
know how much Mexican smokers smoke on average, our population would be “smokers in
Mexico.”
Sample: This is the group of individuals in the population that we select to study—through a
survey, for example.
WHY DOES SAMPLING WORK?
Sampling is useful because we can pair it with an inverse process known as generalization.
To understand a population, the steps we follow are: (1) select a sample from the
population, (2) measure certain data or an opinion for all individuals in the sample and (3)
project the result we observe in the sample onto the population. This projection or
extrapolation is called generalization of results
Generalization of results necessarily adds a certain degree of error to those results. Imagine
that we asked a random sample of 1,000 Mexicans if they smoke, and 25% say that they do.
If 25% of 1,000 randomly selected Mexicans smoke, simple logic would suggest that we
would get about the same result if we asked all 122 million Mexicans. Since we chose our
sample at random, however, it’s possible that we chose a higher proportion of smokers than
that of the population. It's also possible that smokers are underrepresented in our sample.
Random sampling might give us a different percentage than the actual percentage—maybe
25.2% of the population smokes, for example. Generalization of results from a sample, then,
means accepting a certain amount of error, as illustrated below.
Fortunately, we can use statistics to quantify the error inherent in generalizing from a
sample to a population. We do this using two parameters: the margin of error, which is the
maximum difference we’d expect between the data we observe in the sample and the
actual data in the population, and the confidence level, which is our degree of certainty that
the real data for the population falls within our margin of error.
If, for our hypothetical study, we ask a sample of 471 Mexicans if they smoke, our results
will have a margin of error of ±5% with a confidence level of 97%. This is the proper way of
expressing results when sampling.
SAMPLE SIZE
What sample size do I need to study a given population? That depends on the size of the
population and on the amount of error that you’re ready to accept, as we explain in this
post. Greater precision will require a larger sample. If you want absolute certainty in your
results, down to the last decimal, your sample has to be as big as your population.
But sample sizes have one key characteristic that explains why sampling is popular in so
many fields: as we study larger populations, the necessary sample size represents a
diminishing percentage of that population.
There’s a didactic explanation of this phenomenon at [Link], an interesting blog
on mathematics (in Spanish). Imagine we’re conducting a survey to learn a certain
percentage of something (let’s stick with the percentage of smokers) with a predetermined
degree of error—a margin of error of 5% with a confidence level of 95%, for example. If the
population to be studied consisted of just 100 people, we would need a 79.5-person sample
(that is, 79.5% of the population, a pretty sizable chunk of the whole population). If the
population consisted of 1,000 people, we would need a 277.7-person sample (27.7% of the
population). If the population consisted of 100,000 people, we would need a 382.7-person
sample (3.83% of the population).
As you can see, the necessary sample size does increase with the size of the population, but
not proportionally. The rate of increase tapers off as the sample represents a smaller and
smaller percentage of the total population. In fact, once the population reaches a certain
size (around 100,000 individuals), the sample size needn’t grow any more. Check out the
table below for a few examples:
Sample size necessary for a 5% margin of error with a 95% confidence level.
Population Necessary sample
%
size
10 10 100%
100 80 80%
1.000 278 27,8%
10.000 370 3,7%
100.000 383 0,38%
1.000.000 384 0,038%
10.000.000 385 0,004%
100.000.000 385 0,0004%
Based on the information above, we know that however large the population, with 385
individuals we can estimate any percentage with the same level of error (a 5% margin of
error with a confidence level of 95%). This is why sampling is such a powerful tool: it enables
us to make extremely precise assertions about a large number of individuals by studying just
a small number of them.
On the other hand, as we saw with the 100-person population above, sampling doesn’t
work well with small populations. If there are ten students in my class, I need to learn each
of their opinions to learn the class’s global opinion; I can't skip a single one. If we have a ten-
person population and we don’t want to exceed our proposed level of error, we need to
survey every individual.
ADVANTAGES AND DISADVANTAGES OF SAMPLING
Below, we’ve summarized the main advantages and disadvantages of using sampling to
study a population.
Advantages:
We don't have to study as many individuals, and we don’t need as many resources (time and
money).
Data manipulation is much simpler. If a 1,000-person sample is enough, why go analyzing a
file with millions of records?
Disadvantages:
We introduce (controlled) error to our results, since the nature of sampling requires
generalization of results.
We run the risk of biased results due to a poorly selected sample. For example, if individuals
are selected for our sample in a nonrandom way, it could have a serious effect on our
results.
THE SIMPLE RANDOM SAMPLE: DEFINITION AND ALTERNATIVES
The theory behind sampling is based on the concept of the simple random sample. In a
simple random sample, individuals are selected from the population in a completely random
fashion. This implies that all individuals have identical (nonzero) probability of being
selected for our sample.
But theory is one thing and practice is another. It is only possible to select truly random
samples in extremely controlled contexts. On the other hand, when we have populations
made up of groups that are homogeneous (among themselves), we can take advantage of
these groupings to improve the quality of our sample (or to reduce the necessary sample
size).
In the next few posts, we’ll talk about different forms of sampling, starting with the two
large umbrella categories: random and nonrandom sampling. See you then!
--------------------------------------------------------------------------------------------------------------------------
----------------------------------------------------
Random and non-random sampling
In a recent post, we learned about sampling and the advantages it offers when we want to
study a population. Today, we're going to take a look at the two main sampling methods.
Let’s start by defining the concept of a sampling frame.
SAMPLING FRAME
A sampling frame is a list of elements that make up the population that we want to study.
The sample is drawn from this list. The elements to be studied could be individuals, but they
could also be households, institutions or anything else that might be investigated. The
elements within the sampling frame are known as sampling units.
Let’s look at an example. Suppose we want to gauge customers’ satisfaction with a
particular business. To create our sampling frame, we could access the business’s computer
system and pull up a list of everyone who has purchased a product in the past year. Every
individual on that list would be considered a sampling unit. We could then select our sample
by choosing a group of these customers.
The proportion of the sampling frame included in the sample is known as the sampling
fraction. We saw in an earlier post, that this fraction, along with the sample size, determines
the precision of the results that we will obtain by surveying our sample.
RANDOM SAMPLING
We’re dealing with random sampling whenever the following conditions are met:
(1) Every element in our population has a nonzero probability of being selected as part of
the sample.
(2) We have accurate knowledge of this probability, known as the inclusion probability, for
each element in the sampling frame.
If both of these criteria are met, it is possible to obtain unbiased results about the
population from studying the sample. To obtain unbiased results, it may sometimes be
necessary to use weighting methods; such weighting is possible precisely because we know
each individual's probability of being included in the sample. Samples obtained under these
conditions are also known as random samples.
The above definition leads us to conclude that we can only create a random sample if we
have a sampling frame. A national census, a database of mailing addresses within a city and
a list of a business’s customers are all examples of sampling frames that make random
sampling possible. In each of the above cases, the population to be studied is different: the
residents of a country, the households in a city and a business’s customers, respectively.
Once we have our sampling frame, the random sampling method defines the exact method
we will use to select our sample; for example, simple random sampling, systematic
sampling, stratified sampling, disproportional stratified sampling, cluster sampling, and so
on.
NON-RANDOM SAMPLING
All that said, it’s not easy to meet the criteria imposed by random sampling:
(1) It is relatively unusual to have a sampling frame available to you when you’re conducting
market studies.
(2) Ensuring that every individual in a population has a nonzero probability of being selected
is just as difficult to accomplish; knowing every sampling unit’s exact inclusion probability is
even more difficult. The individuals that cannot be selected as part of a sample are generally
referred to as excluded units.
For these reasons—and to minimize costs—researchers often turn to other sampling
methods, known as nonrandom sampling. When using these alternative methods,
researchers generally select elements for the sample based on hypotheses about the
population of interest, known as selection criteria. For example, if we’re selecting our
sample by stopping people on the street, attempting to stop an equal number of men and
women (to coincide with the presumed gender distribution in the population) would be a
criterion of nonrandom sampling.
In these cases, since the selection of units for the sample isn’t random, we shouldn’t talk
about error estimates. In other words, a nonrandom sample tells us about a population, but
we don’t know how precisely: we can’t determine a margin of error or a confidence level.
These types of sampling methods include availability sampling, sequential sampling, quota
sampling, discretionary sampling and snowball sampling.
SAMPLING ERRORS
As we said above, it is impossible to know the margin of error we’ll have in a study (results
from a survey, for example) when we use nonrandom sampling. This includes surveys
conducted by selecting passersby on the street and interviewing them face-to-face, by
making telephone calls at random or obtained through online panels. None of these cases
fulfills the criteria for random sampling: a sampling frame with units for which we can
calculate the probability of being selected for our sample. When we conduct live surveys on
the street, we don’t have access to a list of the individuals who make up the population.
When we conduct telephone interviews, although we have a list of telephone numbers, not
everyone has a landline or a listed number. When we obtain responses from an online
panel, individuals who do not have Internet access cannot be selected, and so their inclusion
probability is zero.
Nevertheless, we regularly come across studies conducted using these methods that state a
margin of error and a confidence level. Formally speaking, this is an incorrect practice, but
researchers tend to use it in order to give some indication of the influence that the sample
size has on the precision of the results. It would be more accurate to say, “if this were a
random sample, it would have margin of error equal to X.”
There is a wide range of opinions over the usefulness of stating a margin of error under
these circumstances, as expressed in a debate described in the next post.
In the next few posts, we’ll take a look at each of the sampling methods in turn: how they
work, what they are used for and what kind of results they provide.
--------------------------------------------------------------------------------------------------------------------------
Random sampling: simple random sampling
Continuing with our series of posts on sampling, today we'll review the first random
sampling method: simple random sampling. This is one of the most popular sampling
methods, and it serves as a reference for many others, even though, as we’ve said before, in
practice it can be difficult to implement.
DEFINITION
Simple random sampling (SRS) is a sampling method in which all of the elements in the
population—and, consequently, all of the units in the sampling frame—have the same
probability of being selected for the sample. It would be along the lines of having a fair raffle
among every individual in the population: we give everyone raffle tickets with unique
sequential numbers, put them all in a basket and draw numbers from the basket at random.
The individuals whose numbers are selected become our sample. Obviously, in practice,
these methods can be automated using computers.
Depending on whether or not the individuals in the population can be selected for the
sample more than once, we distinguish between SRS with replacement and SRS without
replacement. If we sample with replacement, the fact that an individual was randomly
selected for our sample does not prevent that same individual from being chosen again in
the next selection. This would be the equivalent of putting the raffle ticket back in the
basket after every draw. If, on the other hand, we choose to sample without replacement,
an individual selected for the sample is not eligible for the next drawing in the raffle.
To replace, or not to replace? That is the question. It’s a simple math problem. In his
book Statistical Sampling (2005), César Pérez López presents a clear comparison between
the two methods. In terms of both estimation precision and minimum sample size required
to obtain a given level of precision, we can firmly conclude that simple random sampling
without replacement is more efficient.
To understand this result, let’s start with the following expression for sample size in an SRS
without replacement. The formula below shows the relationship between the sample sizes
required to achieve a given level of precision in a finite versus an infinite population:
where n0 is the sample size required with an infinite population and N is the size of the finite
population. It can be shown that the sample size required when sampling with replacement
(nr) is equal to the sample size required when sampling without replacement from an
infinite population. (nr=n0). In that case, we can state that
Therefore, the sample size required when we sample without replacement is smaller than
the sample size required when we sample with replacement. This makes sense intuitively: if
we sample with replacement and the same individual is selected more than once, the effect
is similar to reducing the sample size, since the presence of the repeatedly selected
individual(s) makes our sample less diverse. Finally, the two sampling methods coincide if
the population is infinite, since in that case the odds of selecting an individual more than
once in the same sample would be infinitely small.
BENEFITS OF SIMPLE RANDOM SAMPLING
The development of computer science has enabled us to quickly design simple random
samples that are extremely reliable. Random numbers generated by software—or, strictly
speaking, pseudo-random numbers—are increasingly reliable.
So, by using SRS, we can be sure that we are drawing representative samples, so the only
error that could affect our results is chance. Most importantly, we can accurately calculate
the likelihood of this error (or at least delimit it). Check out our next post for more
information.
DRAWBACK TO SIMPLE RANDOM SAMPLING
The only drawback to SRS is the difficulty of putting it into practice in real-life research.
Remember: since it is a probability method, we need a sampling frame that includes all
individuals, and all of them need to be selectable for our sample. This is a tough
requirement for most real-life market and opinion studies to meet, so researchers for such
studies are often forced to use other methods.
In our next post, we’ll take a look at another very popular random sampling
method: stratified sampling. See you then!
--------------------------------------------------------------------------------------------------------------------------
------------------------------------------------------------------------------------------------------
Random sampling: stratified sampling
In an earlier post, we saw the definition, advantages and drawback of simple random
sampling. Today, we’re going to take a look at stratified sampling.
This method, which is a form of random sampling, consists of dividing the entire population
being studied into different subgroups or discrete strata (the plural form of the word), so
that an individual can belong to only one stratum (the singular). Once the strata have been
defined, in order to create a sample, we select individuals by applying a sampling method to
each of the strata separately. If, for example, we use simple random sampling for every
stratum, we’re using what’s called stratified random sampling (StratRS). We can also use
other sampling methods for every stratum, such as systematic sampling and random
sampling with replacement.
Strata tend to be homogeneous groups of individuals, while groups are heterogeneous
among themselves. For example, if we’re expecting very different behavior between men
and women in a study, it would be convenient to define two strata, one for each gender. If
we have selected these strata correctly, (1) men should behave similarly to each other, (2)
women should behave similarly to each other and (3) men and women should exhibit
dissimilar behavior.
If this condition (strata are internally homogeneous and heterogeneous among themselves)
is met, then using StratRS reduces sampling errors and improves the precision of our results
when we base a study on our sample.
It is fairly common to define strata based on certain variables that are characteristic of the
population, such as age, sex, class or geographic region. These variables make it easy to
divide the sample into mutually exclusive groups and enable us to discern different
behaviors within the population.
TYPES OF STRATIFIED SAMPLING
Depending on the size that we assign to the strata, we can talk about different sorts of
stratified sampling. It’s also common to talk about different ways of “allocating” the sample
into strata.
(1) Proportional stratified sampling
When we select a characteristic of the individuals in order to define the strata, the resulting
subpopulations are often different sizes. Let’s say that we want to study the percent of the
Mexican population that smokes, and we decide that age would be a good criterion to
stratify (that is, we think that smoking habits might vary significantly by age). We define
three strata: under twenty, twenty to forty-four and forty-five and up. When we divide the
population of Mexico into these three strata, we don’t expect all three groups to be the
same size. And, in fact, the real data confirm this:
* Stratum 1 - Mexican population younger than twenty: 42.4 million (41.0%)
* Stratum 2 - Mexican population between twenty and forty-four: 37.6 million (36.3%)
* Stratum 3 - Mexican population older than forty-four: 23.5 million (22.7%)
If we use proportional stratified sampling, the sample should consist of strata that maintain
the same proportions as the population. If, for this example, we want to create a sample of
1,000 individuals, the strata must have the following sizes:
Stratum Population Proportion Sample
1 42,4M 41,0% 410
2 37,6M 36,3% 363
3 23,5M 22,7% 227
2) Uniform stratified sampling
We talk about uniform allocation when we assign the same sample size to all of our defined
strata, regardless of those strata’s weight within the population. A uniform stratified
sampling of the above example would yield the following sample for each stratum:
Stratum Population Proportion Sample
1 42,4M 41,0% 334
2 37,6M 36,3% 333
3 23,5M 22,7% 333
This method favors strata that have less weight in the population by affording them the
same level of importance as the more relevant strata. This reduces the global effectiveness
of our sample (the results will be less precise), but it enables us to study individual
characteristics of each stratum with greater precision. In our example, if we want to make
some specific statement about the population of Stratum 3 (those older than forty-four), we
could reduce sampling errors by using a 333-unit sample, rather than a 227-unit sample
(which we would use in proportional stratified sampling).
(3) Optimal stratified sampling (with respect to standard deviation)
In this case, the size of the strata in the sample is not proportional to the population. Rather,
the size of the strata is proportional to the standard deviation of the variables being studied.
In other words, the strata with the greatest internal variability are the largest, so that the
whole sample better represents the groups that are most difficult to study.
EFFECTIVENESS OF DIFFERENT STRATIFIED SAMPLINGS
The inevitable questions are: When is it best to use stratification? Which kind of
stratification is best?
Proportional stratified sampling always produces the same number of sampling errors as
simple random sampling, or fewer. This means that it is more precise. We achieve equality
when the averages or proportions that we are studying are equal in all strata. Therefore,
stratification is more beneficial when the strata are more diverse among themselves.
Optimal stratified sampling is always as precise or more precise than proportional
stratified sampling. These methods are equally precise and totally equivalent when the
typical variations within each stratum are equal. Therefore, optimal stratification is more
beneficial when there is greater variation within each group, a situation in which we could
reduce the sample size of the most homogeneous group and favor those that are more
heterogeneous. Then again, this is a more complex method that requires a lot of
information about the sample before we even get started—information that we often don't
possess.
SAMPLE SIZES REQUIRED FOR EACH METHOD
Now we've seen how stratification can be beneficial. Not only can these methods be used to
more precisely estimate averages (e.g. the average number of cigarettes consumed by
smokers in Mexico) and proportions (e.g. the percent of the Mexican population that
smokes); they also enable us to reduce the sample size necessary to make an estimate with
a predetermined level of error.
The table below summarizes the sample size required for each method, based on the
maximum level of error that we are willing to accept and the characteristics of the
population itself, which we will assume is infinite in size (if it is finite, a correcting factor
must be applied).
To understand the above table, you have to know that:
Z = The deviation from the mean value that we will accept in order to achieve our desired
confidence level. Depending on the confidence level we want, we'll use a cutoff value based
on the quantiles of the Gaussian distribution. The most common values are:
90% confidence level -> Z=1.645
95% confidence level -> Z=1.96
99% confidence level -> Z=2.575
L is the number of strata into which we divide the sample and h represents a particular
stratum. Therefore, h can vary between 1 and L strata.
p is the proportion of the total population that we are trying to determine (e.g. the percent
of the Mexican population that smokes). Therefore, (1-p) is the complementary proportion
of the population; that is, the proportion of the population to which the sought criterion
does not apply (non-smokers). Likewise, ph represents that proportion within each stratum.
σ2 is the variance of the data we’re seeking (in the case of estimating averages) within the
total population. σh2 is the variance within every stratum.
e is the accepted margin of error.
Wh is the stratum’s weight within the sample (the size of the stratum with respect to the
whole sample). If we’re talking about proportional stratification, every Wh is equal to the
proportion represented by that stratum in the population. If we’re talking about optimal
stratification, every Wh is calculated based on the dispersion within each stratum.
Using the formulas above, it is possible to demonstrate that these different stratification
methods only reduce the sample size if the values p and σ vary between strata. Otherwise,
all of the expressions are equivalent. Let’s take a look at an example: if we take the
expression of sample size required to estimate an average by means of optimal stratified
sampling (ignoring the parameter Z in this case)
and we consider all of the strata’s variances to be equal (σh=σ) and further consider the
strata to be identical in size (Wh=1/L), we get
We hope that this post has helped clarify how useful stratified sampling can be. In our next
few posts, we’ll tackle systematic sampling.