SAMPLING METHODS
Introduction
Sampling consists of selecting some part of a population to observe so that one may estimate something
about the whole population.
Example 1: Suppose that a researcher wishes to estimate the average diastolic blood pressure (DBP)
of citizens of city A. The researcher may not have enough time, money, or any other resources to
enumerate the entire population. In such cases the researcher can measure the DBP of sample of
subjects. The part of the population included in the study must be representative.
Example 2: To estimate the amount of recoverable oil in a region, a few (highly expensive) sample
holes are drilled. The researchers cannot drill the holes everywhere. So based on the data from a sample
of holes, the amount of recoverable oil.
Technical Terms
Population: The population or universe is an aggregate of elements about which the inference is to be
made.
Unit: The unit is a member of population from which the information is sought.
Sampling Units: The sampling units are non-overlapping collections of elements of the population.
Sampling frame: A list of all units in the population to be sampled is termed as sampling frame.
Sample: A subset of population selected from a sampling frame to draw inferences about a population
characteristic is called sample.
Parameter and Statistic: A parameter is a number describing a whole population (e.g., population
mean), while a statistic is a number describing a sample (e.g., sample mean).
A population is the entire group that one wants to draw conclusions about .A sample is the specific
group that you will collect data from. The size of the sample is always less than the total size of the
population.
In research, a population doesn’t always refer to people. It can mean a group containing elements of
anything you want to study, such as objects, events, organizations, countries, species, organisms, etc.
1
Need for Sampling
Collection of information on every unit in the population for the characteristics of interest is known as
complete enumeration or census. The money, manpower, and time required for carrying out a census
will generally be large, and there are many situations where with limited means complete enumeration
is not possible. There are also instances where it is not feasible to enumerate all units due to their
perishable nature. In all such cases, the investigator has no alternative except resorting to a sample
survey. The number of units (not necessarily distinct) included in the sample is known as the sample
size and is usually denoted by n, whereas the number of units in the population is called population
size and is denoted by N.
The advantages of a sample survey over complete enumeration are given below:
1) Greater Speed The time taken for collecting and analysing the data for a sample is much less
than that for a complete enumeration. Often, we come across situations where the information
is to be collected within a specified period. In such cases, where time available is short or the
population is large, sampling is the only alternative.
2) Greater Accuracy A census usually involves a huge and unwieldy organization and, therefore,
many types of errors may creep in. Sometimes, it may not be possible to control these errors
adequately. In sample surveys, the volume of work is considerably reduced. On account of this,
the services of better trained and efficient staff can be obtained without much difficulty. This
will help in producing more accurate results than those for complete enumeration.
3) More Detailed Information As the number of units in a sample are much less than those in
census, it is, therefore, possible to observe/interview each and every sample unit intensively.
Also, the information can be obtained on a greater number of variables. However, in complete
enumeration such an effort becomes comparatively difficult.
4) Reduced Cost Because of lesser number of units in the sample in comparison to the population,
considerable time, money, and energy are saved in observing the sample units in relation to the
situation where all units in the population are to be covered.
From the above discussion, it is seen that the sample survey is more economical, provides more
accurate information, and has greater scope in subject coverage as compared to a complete
enumeration.
Sampling and Non-Sampling Errors
Sampling Error: Sampling error is the error that arises in a data collection process as a result of taking
a sample from a population rather than using the whole population. Sampling error is one of two
2
reasons for the difference between an estimate of a population parameter and the true, but unknown,
value of the population parameter. The sampling error for a given sample is unknown but when the
sampling is random, for some estimates (for example, sample mean, sample proportion) theoretical
methods may be used to measure the extent of the variation caused by sampling error.” Sampling error
is mainly cause due to the reason that sample not whole population.
Non-Sampling Error: Non-sampling error is the error that arises in a data collection process as a result
of factors other than taking a sample. Non-sampling errors have the potential to cause bias in polls,
surveys, or samples. There are many different types of non-sampling errors and the names used to
describe them are not consistent. This may be due to poor sampling method, measurement errors, and
behavioural effect.
Sampling error is one which occurs due to unrepresentativeness of the sample selected for observation.
Conversely, non-sampling error is an error arise from human error, such as error in problem
identification, method or procedure used, etc.
BASIS FOR
SAMPLING ERROR NON-SAMPLING ERROR
COMPARISON
Meaning Sampling error is a type of error, An error occurs due to sources other
occurs due to the sample selected than sampling, while conducting
does not perfectly represents the survey activities is known as non-
population of interest. sampling error.
Cause Deviation between sample mean and Deficiency and analysis of data
population mean
Type Random Random or Non-random
Occurs Only when sample is selected. Both in sample and census.
Sample size Possibility of error reduced with the It has nothing to do with the sample
increase in sample size. size.
The non-sampling errors are unavoidable in census and surveys. The data collected by complete
enumeration in census is free from sampling error but would not remain free from non-sampling errors.
The data collected through sample surveys can have both – sampling errors as well as non-sampling
errors. The non-sampling errors arise because of the factors other than the inductive process of
inferring about the population from a sample.
3
Non sampling errors can occur at every stage of planning and execution of survey or census. It occurs
at the planning stage, fieldwork stage as well as at tabulation and computation stage. The main sources
of the non-sampling errors are
• lack of proper specification of the domain of study and scope of the investigation,
• incomplete coverage of the population or sample,
• faulty definition,
• defective methods of data collection and
• tabulation errors.
Non-sampling errors may be broadly classified into three categories.
(a) Specification errors: These errors occur at planning stage due to various reasons, e.g.,
inadequate and inconsistent specification of data with respect to the objectives of surveys/census,
omission or duplication of units due to imprecise definitions, faulty method of
enumeration/interview/ambiguous
schedules etc.
(b) Ascertainment errors: These errors occur at field stage due to various reasons e.g., lack of
trained and experienced investigations, recall errors and other type of errors in data collection, lack
of adequate inspection and lack of supervision of primary staff etc.
(c) Tabulation errors: These errors occur at tabulation stage due to various reasons, e.g.,
inadequate scrutiny of data, errors in processing the data, errors in publishing the tabulated results,
graphs etc.
Ascertainment errors may be further sub-divided into
(i) Coverage errors owing to over-enumeration or under-enumeration of the population or the
sample, resulting from duplication or omission of units and from the non-response.
(ii) Content errors relating to the wrong entries due to the errors on the part of investigators and
respondents.
Sampling Procedures
The method which is used to select the sample from a population is known as sampling procedure.
These procedures can be put into two categories – probability/random sampling and
nonprobability/non-random sampling.
Non-Probability Sampling Techniques
Convenience sampling:
It is primarily determined by convenience to the researcher.
4
This can include factors like:
• Ease of access
• Geographical proximity
• Existing contact within the population of interest
Convenience samples are sometimes called “accidental samples,” because participants can be
selected for the sample simply because they happen to be nearby when the researcher is
conducting the data collection.
Example: Convenience sampling
Suppose that a researcher is investigating the association between daily weather and daily
shopping patterns. To collect insight into people’s shopping patterns, the researcher decides to
stand outside a major shopping mall in your area for a week, stopping people as they exit and
asking them if they are willing to answer a few questions about their purchases.
Quota Sampling:
In quota sampling, the researcher selects a predetermined number or proportion of units, called a quota.
The quota should comprise subgroups with specific characteristics (e.g., individuals, cases, or
organizations) and should be selected in a non-random manner. The subgroups, called strata, should
be mutually exclusive.
Example: Suppose that a researcher is seeking opinions about the design choices on a website, but do
not know how many people use it. The researcher may decide to draw a sample of 100 people,
including a quota of 50 people under 40 and a quota of 50 people over 40. This way, the researcher
gets perspective of both age groups.
Snow ball sampling:
It is used when the population is hard to reach, or there is no existing database or other sampling frame
to find them. Research about socially marginalized groups such as drug addicts, homeless people, or
sex workers often uses snowball sampling. To conduct a snowball sample, the researcher starts by
finding one person who is willing to participate in the research and then by asking them to introduce
to others. Alternatively, the research may involve finding people who use a certain product or have
experience in the area of interest. In these cases, one can also use networks of people to gain access to
your population of interest.
Example: A researcher is studying homeless people living in a city. He/she starts by attending a housing
advocacy meeting, striking up a conversation with a homeless woman. Then the researcher explains
the purpose of research and she agrees to participate. She invites the researcher to a parking lot serving
as temporary housing and offers to introduce the researcher around.
5
In this way, the process of snowball sampling begins. The researcher started by attending the meeting,
where he/she met someone who could then put him/her in touch with others in the group.
Judgement/Purposive Sampling:
Purposive (judgmental) sampling: Purposive sampling is a blanket term for several sampling
techniques that choose participants deliberately due to qualities they possess. It is also called
judgmental sampling, because it relies on the judgment of the researcher to select the units (e.g.,
people, cases, or organizations studied). Purposive sampling is common in qualitative and mixed
methods research designs, especially when considering specific issues with unique cases.
Advantages of non-probability sampling
Depending on research design, there are advantages to choosing non-probability sampling.
• Non-probability sampling does not require a sampling frame, so subjects are often readily
available. This can make non-probability sampling quicker and easier to carry out.
• Non-probability sampling allows the researcher to target groups within population. In certain
types of research, it is vital that certain units be included in sample. For example, many kinds
of medical research rely on people with a specific health issue.
• Although it is not possible to make statistical inferences from the sample to the population,
non-probability sampling methods can provide researchers with the data to make other types
of generalizations from the sample being studied.
Disadvantages of non-probability sampling
Non-probability sampling has some downsides as well. These include the following:
• Non-probability samples are extremely unlikely to be representative of the population studied.
This undermines the generalizability and validity of your results.
• Non-probability samples are at risk of several kinds of research bias:
• As some units in the population have no chance of being included in the sample, under coverage
bias is likely.
• Furthermore, since the selection of units included in the sample is often based on ease of access,
sampling bias is common as well.
• While the subjective judgment of the researcher in choosing who makes up the sample can be
an advantage, it also increases the risk of observer bias.
Probability Sampling Techniques
1) Simple Random Sampling
2) Systematic Random Sampling
6
3) Stratified Random Sampling
4) Cluster Random Sampling
Simple Random Sampling:
A simple random sample is a randomly selected subset of a population. In this sampling method, each
member of the population has an exactly equal chance of being selected.
This method is the most straightforward of all the probability sampling methods, since it only involves
a single random selection and requires little advance knowledge about the population. Because it uses
randomization, any research performed on this sample should have high internal and external validity,
and be at a lower risk for research biases like sampling bias and selection bias.
When to use simple random sampling?
Simple random sampling is used to make statistical inferences about a population. It helps ensure high
internal validity: randomization is the best method to reduce the impact of potential confounding
variables. In addition, with a large enough sample size, a simple random sample has high external
validity: it represents the characteristics of the larger population. However, simple random sampling
can be challenging to implement in practice. To use this method, there are some prerequisites:
• A complete list of every member of the population.
• Access to each member of the population if they are selected.
• The time and resources to collect data from the necessary sample size.
Simple random sampling works best if you have a lot of time and resources to conduct your study, or
if you are studying a limited population that can easily be sampled.
The steps in performing a simple random sampling procedure:
Step 1: Define the population
Start by deciding on the population. It is important to ensure that there is an access to every individual
member of the population, so that one can collect data from all those who are selected for the sample.
Step 2: Decide on the sample size
Next, The sample size must be decided. Although larger samples provide more statistical certainty,
they also cost more and require far more work. There are several potential ways to decide upon the
size of sample, but one of the simplest involves using a formula with your desired confidence interval
7
and confidence level, estimated size of the population, and the standard deviation of variable of
interest.
Step 3: Randomly select the sample
This can be done in one of two ways: the lottery or random number method. In the lottery method, one
chooses the sample at random by “drawing from a hat” or by using a computer program that will
simulate the same action. In the random number method, every individual is assigned a number. By
using a random number generator or random number tables, the subjects are randomly selected. The
random number function (RAND) in Microsoft Excel can also be used to generate random numbers.
Step 4: Collect data from sample
The final step is to collect the data. To ensure the validity of findings, the researcher needs to make
sure every individual selected participates in the study. If some drop out or do not participate, this
could bias the findings.
Systematic Random Sampling:
Is a probability sampling method in which researchers select members of the population at a regular
interval (or k) determined in advance.
If the population order is random or random-like (e.g., alphabetical), then this method gives a
representative sample that can be used to draw conclusions about your population of interest.
When to use systematic sampling?
Systematic sampling is a method that imitates many of the randomization benefits of simple random
sampling, but is slightly easier to conduct.
The advantage of this method is that it can be used even when the list of all subjects in the population
is not available.
Example: You run a department store and are interested in how you can improve the store experience
for your customers. To investigate this question, you ask an employee to stand by the store entrance
and survey every 20th visitor who leaves, every day for a week.
Although you do not necessarily have a list of all your customers ahead of time, this method should
still provide you with a representative sample of your customers since their order of exit is essentially
random.
Order of the population: When using systematic sampling with a population list, it’s essential to
consider the order in which population is listed to ensure that the sample is valid. If the population is
in ascending or descending order, using systematic sampling should still gives a representative sample,
as it will include participants from both the bottom and top ends of the population.
8
The researcher should not use systematic sampling if your population is ordered cyclically or
periodically, as your resulting sample cannot be guaranteed to be representative.
Example: Alternating list
Your population list alternates between men (on the even numbers) and women (on the odd numbers).
You choose to sample every tenth individual, which will therefore result in only men being included
in your sample. This would obviously be unrepresentative of the population.
Example: Cyclically ordered list
You are sampling from a population list of approximately 1000 hospital patients. The list is divided
into 50 departments of around 20 patients each. Within each department, the list is ordered by age,
from youngest to oldest. This results in a list of 20 repeated age cycles.
If you sample every 20th individual, because each department is ordered by age, your population will
consist of the oldest person in each one. This will most likely not provide a representative sample of
the entire hospital population (high generalizability).
Systematic sampling also begins with the complete sampling frame and assignment of unique
identification numbers. However, in systematic sampling, subjects are selected at fixed intervals, e.g.,
every third or every fifth person is selected. The spacing or interval between selections is determined
by the ratio of the population size to the sample size (N/n). For example, if the population size is
N=1,000 and a sample size of n=100 is desired, then the sampling interval is 1,000/100 = 10, so every
tenth person is selected into the sample. The selection process begins by selecting the first person at
random from the first ten subjects in the sampling frame using a random number table; then 10th
subject is selected.
If the desired sample size is n=175, then the sampling fraction is 1,000/175 = 5.7, so we round this
down to five and take every fifth person. Once the first person is selected at random, every fifth person
is selected from that point on through the end of the list.
With systematic sampling like this, it is possible to obtain non-representative samples if there is a
systematic arrangement of individuals in the population. For example, suppose that the population of
interest consisted of married couples and that the sampling frame was set up to list each husband and
then his wife. Selecting every tenth person (or any even-numbered multiple) would result in selecting
all males or females depending on the starting point. This is an extreme example, but one should
consider all potential sources of systematic bias in the sampling process.
9
Stratified Random Sampling:
In a stratified sample, researchers divide a population into homogeneous subpopulations called strata
(the plural of stratum) based on specific characteristics (e.g., race, gender identity, location, etc.). Every
member of the population studied should be in exactly one stratum.
Each stratum is then sampled using another probability sampling method, such as cluster sampling or
simple random sampling, allowing researchers to estimate statistical measures for each sub-population.
Researchers rely on stratified sampling when a population’s characteristics are diverse and they want
to ensure that every characteristic is properly represented in the sample. This helps with the
generalizability and validity of the study, as well as avoiding research biases like under coverage bias.
To use stratified sampling, the population must be divided into mutually exclusive and exhaustive
subgroups. That means every member of the population can be clearly classified into exactly one
subgroup.
Stratified sampling is the best choice among the probability sampling methods when subgroups will
have different mean values for the variable(s) under study.
A stratified sample includes subjects from every subgroup, ensuring that it reflects the diversity of the
population. This is the advantage of stratified sampling when compare to other sampling techniques.
In stratified sampling, we split the population into non-overlapping groups or strata (e.g., men and
women, people under 30 years of age and people 30 years of age and older), and then sample within
each strata. The purpose is to ensure adequate representation of subjects in each stratum.
Sampling within each stratum can be by simple random sampling or systematic sampling. For example,
if a population contains 70% men and 30% women, and we want to ensure the same representation in
the sample, we can stratify and sample the numbers of men and women to ensure the same
representation. For example, if the desired sample size is n=200, then n=140 men and n=60 women
could be sampled either by simple random sampling or by systematic sampling.
10
Cluster Random Sampling:
Cluster random sampling is a probability sampling method where researchers divide a large population
into smaller groups known as clusters, and then select randomly among the clusters to form a sample.
Cluster sampling is typically used when the population and the desired sample size are particularly
large.
Various types of Cluster Sampling Techniques:
Single Stage Cluster Sampling:
A single-stage cluster is a type of cluster sampling where each unit of the chosen clusters is sampled.
Researchers will first divide the total sample into a predetermined number of clusters based on how
large they want each cluster to be. Then, they randomly select and sample from the clusters and collect
data from each individual unit in the selected clusters.
Example:
Single Stage Cluster Sampling:
In two-stage cluster sampling, researchers will only collect data from a random subsample of
individual units within each of the selected clusters to use as the sample. This technique is less precise
than single-stage sampling and should only be used when it is too challenging or expensive to test the
entire cluster.
Example:
Multi Stage Cluster Sampling:
This type of cluster sampling involves the same process as double-stage sampling, except with a few
extra steps. In multi-stage sampling, researchers will continue to randomly sample elements from
within the clusters until they reach a manageable sample size.
Example:
Steps in Cluster Sampling:
• First, choose the target population that you wish to study and determine your desired sample
size.
• Then, divide your sample into clusters. When forming the clusters, make sure each cluster’s
population is diverse, has a similar distribution of characteristics to the distribution of the
population as a whole, and has the same number of members. The goal is to form clusters that
are representative of the total population as a whole.
• Next, select clusters by a random selection process. It is important to randomly select from
the clusters to preserve your results’ validity. The number of clusters selected is based on how
large the sample size is.
11
• In single-stage sampling, collect data from each individual unit of the clusters you selected in
Step 3.
• In the case of double-stage or multi-stage sampling, you randomly select individual units from
within the selected clusters to use as your sample. You will then collect your data from each of
these individual units. Double-stage and multi-stage clustering tend to be easier than single-
stage because you will work with a much smaller sample.
Advantages
• Time and cost-efficient: Cluster sampling is cheaper and quicker than other sampling
methods. For example, it reduces travel expenses for wide geographical populations.
• High external validity: If your population is clustered properly to represent every possible
characteristic of the entire population, your clusters will accurately reflect the entire
population.
• Practicality and ease: This type of sampling process enables researchers to study large
populations that would otherwise be too challenging or complicated to analyze otherwise.
Limitations
• High sampling error: When the clusters do not mirror the population’s characteristics or serve
as a mini-representation of the population as a whole, there will be less statistical certainty
and accuracy. This error is even greater when you use more stages of clustering.
• Complexity: Planning study designs for cluster sampling usually requires more attention
because researchers need to determine how to divide up a larger population efficiently and
properly.
12