Sampling
Sampling is a cornerstone of quantitative research [Link]'s the systematic
process of selecting a subset of individuals or items from a larger group to make
inferences about the entire group. The rationale behind sampling lies in the
impracticality or impossibility of studying every single element within a population due to
constraints of time, resources, and accessibility. By carefully selecting a representative
sample, researchers can gain valuable insights into the characteristics of the larger
population without needing to collect data from every member. This efficiency makes
sampling an indispensable tool across various fields, including social sciences, public
health, market research, and quality control.
Basic Terminology
A clear understanding of fundamental terms is essential for comprehending the
intricacies of sampling:
• Population (N): The population, often referred to as the study population, is the
complete set of all individuals, objects, events, or measurements that possess a
common observable characteristic of interest to the researcher. It is the entire
group about which a researcher wishes to draw conclusions. For instance, if a
researcher is studying the academic performance of university students in
Bangladesh, the population would be all university students in Bangladesh. It's
crucial to define the population precisely, as an ill-defined population can lead to
flawed research outcomes.
• Sample: A sample is a carefully chosen, smaller, and representative subset of
the population. It serves as a miniature version of the larger group, designed to
reflect its characteristics accurately. The goal is for the sample to be a
microcosm of the population, allowing generalizations to be made. For example,
from the population of all university students in Bangladesh, a researcher might
select a sample of 1000 students.
• Parameter: A parameter is a numerical characteristic that describes an entire
population. Since it describes the entire population, parameters are typically fixed
values, though often unknown. Examples include the population mean (μ) and
population standard deviation (σ). For instance, the average GPA of all university
students in Bangladesh would be a population parameter. Researchers usually
try to estimate these unknown parameters using sample statistics.
• Statistic: In contrast to a parameter, a statistic is a numerical characteristic that
describes a sample. Statistics are calculated from sample data and are used to
estimate population parameters. The sample mean (xˉ) and sample standard
deviation are examples of statistics. If a researcher calculates the average GPA
of the 1000 sampled university students, that would be a sample statistic. The
relationship between a statistic and a parameter is crucial for inference: statistics
are used to make educated guesses about parameters.
• Sample Size (n): This term refers to the exact number of individuals, cases, or
units included in the sample from the study population. It is denoted by 'n'.
Determining the appropriate sample size is a critical step in research design, as it
directly impacts the reliability and generalizability of the findings. A sample size
that is too small might not be representative, while an excessively large sample
size can be resource-intensive without significantly increasing accuracy.
• Sampling Design/Strategy: This refers to the specific method or procedure
employed to select elements from the population to form the sample. The choice
of sampling design is critical because it dictates how representative the sample is
and, consequently, the validity of the inferences drawn. Different research
questions, available resources, and population characteristics necessitate
different sampling strategies. A well-chosen sampling design enhances the
credibility of the research.
• Sampling Unit/Element: A sampling unit, also known as a sampling element, is
an individual member or entity of the population that is considered for selection
into the sample. This could be a person, a household, a school, an organization,
or any other defined entity depending on the research scope. Each sampling unit
is distinct and serves as the basic element from which the sample is constructed.
• Sampling Frame: The sampling frame is an exhaustive list or directory of all the
sampling units within the study population from which the sample will be
selected. It acts as a practical representation of the theoretical population. An
ideal sampling frame is complete, accurate, and up-to-date, ensuring that every
element of the population has a chance of being included. If the sampling frame
is incomplete or contains inaccuracies (e.g., includes elements not in the
population or excludes elements that should be), it can introduce bias into the
sample, leading to inaccurate results. For instance, a university's student roster
could serve as a sampling frame for a study on its students.
Purpose of Sampling
The primary purpose of sampling in quantitative research is to facilitate the drawing of
inferences or generalizations about a larger population based on data collected from a
smaller, manageable subset. This fundamental principle underpins much of empirical
research. Without sampling, studying large populations would be prohibitively
expensive, time-consuming, and often logistically impossible.
Specifically, sampling serves several key objectives:
1. Efficiency and Resource Optimization: Studying an entire population (a
census) is often impractical due to vast numbers, geographical spread, and
resource limitations (financial, human, time). Sampling provides a cost-effective
and time-efficient alternative, allowing researchers to complete studies within
realistic constraints. For example, surveying every eligible voter in a national
election is unfeasible; thus, polls rely on samples.
2. Feasibility of Data Collection: For some populations, a complete enumeration
might be physically impossible. Consider studies involving destructive testing
(e.g., testing the lifespan of light bulbs) where every item cannot be tested.
Sampling allows for testing a subset without destroying the entire population.
3. Enhanced Data Quality: By focusing resources on a smaller sample,
researchers can often collect more detailed, accurate, and higher-quality data per
unit than would be possible in a large-scale census. This can involve more in-
depth interviews, specialized measurements, or rigorous follow-ups that are not
scalable to an entire population.
4. Drawing Inferences and Generalizations: The ultimate goal of sampling is to
use the characteristics observed in the sample to make statistically valid
conclusions about the unknown characteristics (parameters) of the entire
population. This process, known as statistical inference, allows researchers to
predict prevalence, identify trends, test hypotheses, and estimate outcomes for
the larger group with a measurable degree of certainty.
5. Reduced Non-sampling Error: While sampling introduces sampling error (due
to studying a part rather than the whole), it can help reduce non-sampling errors
that arise from issues like interviewer fatigue, data entry mistakes, or non-
response, which are often exacerbated in large-scale data collection efforts. A
smaller, more controlled sample allows for better training of data collectors and
more rigorous quality checks.
In essence, sampling is the art and science of selecting a few (a sample) from a bigger
group (the sampling population) to serve as a reliable basis for estimating or predicting
an unknown piece of information, situation, or outcome regarding the larger group.
Principles of Sampling
Factors Affecting Inferences from a Sample
The reliability and accuracy of inferences drawn from a sample are significantly
influenced by two primary factors:
1. The Size of the Sample: This is perhaps the most intuitive factor. Generally,
larger samples provide more certainty and precision in the findings compared to
smaller ones. A larger sample is more likely to capture the variability present in
the population, thereby reducing the margin of error and increasing the
confidence in the generalizations made. As a rule, the larger the sample size,
the more accurate the findings are likely to be. This is because as 'n' (sample
size) increases, the sampling distribution of the sample mean (for example)
becomes narrower, meaning sample means are more tightly clustered around
the true population mean. However, there's a point of diminishing returns where
increasing the sample size further yields only marginal improvements in accuracy
but significantly increases costs and effort.
2. The Extent of Variation in the Sampling Population: The homogeneity or
heterogeneity of the study population with respect to the characteristics under
investigation profoundly impacts the required sample size and the certainty of
inferences.
1. Homogeneous Population: If a population is highly homogeneous
(uniform or similar) regarding the characteristics being studied, even a
relatively small sample can provide a reasonably good and accurate
estimate. In an extreme case, if all elements in a population are identical,
selecting just one element would provide an absolutely accurate estimate.
For example, if all products coming off an assembly line are identical in
quality, testing one would tell you about all of them.
2. Heterogeneous Population: Conversely, if the population is
heterogeneous (dissimilar or diversified) with respect to the characteristics
under study, a larger sample is necessary to obtain the same level of
accuracy. This is because a more varied population requires a larger
number of observations to adequately capture the full range of differences
present. The greater the variation (or standard deviation) in the study
population, the higher the standard error for a given sample size in your
estimates, leading to greater uncertainty. Therefore, to achieve similar
precision in a heterogeneous population, researchers must draw a larger
sample. For instance, studying opinions on a complex social issue across
a diverse national population would require a much larger sample than
studying the average height of male basketball players in a specific
league.
Understanding these two factors helps researchers make informed decisions about
sample size, balancing the need for precision with practical constraints.
Steps in Sample Design
Designing a sample involves several systematic steps:
1. Type of Universe: First, clearly define the entire group or set of objects you want
to study. This "universe" can be:
o Finite: A countable population, like the residents of a city or the number of
workers in a factory.
o Infinite: An uncountable population, such as the number of stars in the
sky or listeners of a specific radio program.
2. Sampling Unit: Determine the basic unit that will be sampled. This could be a:
o Geographical unit: State, district, village.
o Construction unit: House, flat.
o Social unit: Family, club, school.
o Individual: A single person.
3. Source List/Sampling Frame: Create or identify the comprehensive list from
which the sample will be drawn. This list should ideally include all the sampling
units in your defined universe.
Criteria for Selecting a Sampling Procedure
When choosing a sampling procedure, researchers must be mindful of two main
sources of incorrect inferences: systematic bias and sampling error.
A. Systematic Bias
This type of bias arises from flaws in the sampling procedures themselves. Here are
common causes:
1. Inappropriate Sampling Frame: The list or source from which the sample is
drawn (the sampling frame) may not accurately represent the population. For
example, using an outdated phone book for a current population study.
2. Defective Measuring Device: The tools used to collect data can introduce bias.
In surveys, a poorly worded questionnaire or a biased interviewer can lead to
systematic errors.
3. Non-respondents: If a significant portion of the selected sample does not
participate, and their characteristics differ from those who do respond, it can lead
to biased results.
4. Indeterminacy Principle: Individuals may behave differently when they know
they are being observed compared to their natural behavior. This "observer
effect" can introduce bias.
5. Natural Bias in Reporting Data: People sometimes intentionally misrepresent
information. For instance, they might understate income for tax purposes but
overstate it to appear more affluent.
B. Sampling Errors
These are the random variations that occur when estimating population parameters
from a sample. They naturally decrease as the sample size increases and when the
population being studied is more homogeneous (less varied).
Key Takeaway: When selecting a sampling procedure, prioritize methods that lead to a
relatively small sampling error and effectively help to control systematic bias.
Additional Considerations for Sampling Procedure Selection:
1. Sample Size: The sample size should be optimum, meaning it's large enough to
ensure efficiency, representativeness, reliability, and flexibility in the study's
outcomes.
2. Parameters of Interest: Clearly identify the specific characteristics or measures
of the population that are the focus of your research. This guides the choice of
sampling method.
3. Budgetary Constraint: Practical cost considerations heavily influence the
feasibility of different sampling procedures.
4. Sampling Procedure: The specific method chosen (e.g., random sampling,
stratified sampling) should align with the research objectives and constraints.
Characteristics of a Good Sample Design
A good sample design is crucial for obtaining accurate and reliable research findings.
Here are its key characteristics:
• Truly Representative Sample: The sample must accurately reflect the
characteristics of the entire population being studied. This ensures that the
findings from the sample can be generalized to the larger group.
• Small Sampling Error: The design should minimize the random variations
between the sample estimates and the true population values. A smaller
sampling error means more precise results.
• Viable within Budget: The chosen sample design needs to be practical and
achievable given the financial resources available for the research study.
• Controlled Systematic Bias: A good design effectively controls or minimizes
systematic errors that can skew results, ensuring the data collected is accurate
and unbiased.
• Generalizability with Confidence: The sample should be designed in such a
way that the results obtained from studying it can be confidently applied to the
entire population (the "universe") with a reasonable level of certainty.
Types of Sampling
Sampling methods are broadly categorized into two main types, distinguished by
whether every element in the population has a known, non-zero chance of being
selected:
1. Probability Sampling (Random Sampling): This is a sampling technique
where each element in the population has an equal and independent chance of
selection into the sample of research interest. The key characteristic is that the
selection process is random, meaning it is free from human bias and allows for
the calculation of sampling error. This makes probability sampling ideal for
quantitative research where the goal is to generalize findings from the sample to
the population with a known level of confidence. Since every unit has a
calculable probability of inclusion, statistical inferences can be rigorously made
about the population parameters.
2. Non-Probability Sampling (Non-Random Sampling): In contrast, non-
probability sampling is a technique where the researcher selects samples based
on subjective judgment, convenience, or specific criteria rather than random
selection. This method is less stringent and does not guarantee that every
element of the population has an equal chance of being selected. Consequently,
the results obtained from non-probability samples cannot be generalized to the
entire population with the same statistical confidence as probability samples.
This method often depends heavily on the expertise and discretion of the
researchers. It is widely used in qualitative research where the aim is to explore
in-depth understanding, generate hypotheses, or identify specific cases rather
than make statistical generalizations.
Types of Probability Sampling Techniques
Probability sampling techniques are essential for studies aiming for statistical
generalizability. They include:
• 1. Simple Random Sampling (SRS):
o Description: Simple Random Sampling is considered the most
fundamental and commonly used method of probability sampling. In SRS,
every single element or unit within the defined population has an exactly
equal and independent chance of being selected for the sample. This
independence means that the selection of one unit does not influence the
selection of any other unit. The absence of bias is its defining
characteristic, making it a powerful method for achieving
representativeness.
o Methods of Selection: There are several practical ways to draw a simple
random sample:
▪ The Fishbowl Draw: This classic method involves writing each
population unit's name or number on separate slips of paper,
placing them in a container (like a fishbowl), thoroughly mixing
them, and then drawing the desired number of slips blindfolded.
This technique is often used in lotteries.
▪ Computer Programs/Random Number Generators: Modern
research largely relies on specialized software or statistical
packages (e.g., R, Python, SPSS, SAS) that can generate random
numbers efficiently. Researchers assign a unique number to each
element in the sampling frame and then use the software to
randomly select the required number of unique identifiers. This
method is highly efficient and eliminates human error.
▪ Table of Randomly Generated Numbers: Before the widespread
availability of computing power, researchers used pre-generated
tables of random numbers. They would assign numbers to each
population unit and then use the random number table to pick units
corresponding to the numbers in the table.
o Example: Consider a class of 80 students from which a sample of 20 is
to be selected using SRS. The first step is to assign a unique number from
1 to 80 to each student. Then, using a method like the fishbowl draw or a
random number generator, 20 unique numbers are selected. The
students corresponding to these 20 numbers form the sample. This
technique ensures that every student has an equal chance of being
included, making the sample highly representative of the class.
o Advantages: High external validity (generalizability), unbiased selection,
easy to understand.
o Disadvantages: Requires a complete and accurate sampling frame, can
be impractical for very large populations or geographically dispersed ones,
may result in an unrepresentative sample if the population has distinct
subgroups (though this is a matter of chance, not bias).
• 2. Systematic Sampling:
o Description: Systematic sampling involves selecting elements from a
sampling frame at regular intervals after a random starting point. The
sampling frame is first organized or listed in some order. Then, a sampling
interval (k) is determined by dividing the population size (N) by the desired
sample size (n) (i.e., k = N/n). A random starting point is chosen within the
first interval, and subsequent elements are selected at every k-th position
from that point onward.
o Process:
1. Determine the population size (N) and desired sample size (n).
2. Calculate the sampling interval (k = N/n).
3. Randomly select a starting number between 1 and k (inclusive).
This is the first element of your sample.
4. Select every k-th element thereafter until the desired sample size
is reached.
o Example: Suppose you have 50 students in a class (N=50) and want to
select a sample of 10 (n=10) using systematic sampling.
1. Calculate the interval width: k = 50/10 = 5. This means you will
select one element from every five.
2. Using SRS, randomly select a number between 1 and 5. Let's say
you selected 3.
3. The first student in your sample is student number 3.
4. Subsequent students will be 3+5=8, 8+5=13, 13+5=18, and so on,
until you have 10 students (3, 8, 13, 18, 23, 28, 33, 38, 43, 48).
o Advantages: Simpler and more convenient than SRS, especially for large
populations, and it ensures good coverage of the population if the list is
randomly ordered. It avoids the clustering that can sometimes occur by
chance in SRS.
o Disadvantages: If there's a hidden pattern or periodicity in the sampling
frame that aligns with the sampling interval, it can lead to a biased sample.
For example, if every 10th house on a list is a corner house, and your
interval is 10, you might only sample corner houses. Due to its initial
random selection followed by a fixed pattern, systematic sampling is
sometimes classified as a 'mixed' sampling procedure.
• 3. Stratified Sampling:
o Description: Stratified sampling is employed when the study population
is heterogeneous or exhibits significant variability with respect to key
characteristics relevant to the research (e.g., gender, age, income,
education level, geographic region). The core idea is to divide the
population into distinct, non-overlapping subgroups, called 'strata,' where
elements within each stratum are homogeneous (similar) concerning the
characteristic of interest. After forming strata, a simple random sample is
drawn independently from each stratum.
o Purpose: This method ensures that all important subgroups within the
population are represented in the sample, thereby improving the
representativeness of the sample and reducing sampling error, especially
when there are significant differences between subgroups.
o Types of Stratified Sampling:
▪ Proportionate Stratified Sampling: In this method, the number of
elements selected from each stratum is directly proportional to its
size in the total population. For example, if a university population is
60% female and 40% male, and a sample of 100 students is
needed, the sample would include 60 females and 40 males,
reflecting the population proportions. This ensures the sample is a
true miniature of the population.
▪ Disproportionate Stratified Sampling: Here, the number of
elements selected from each stratum does not necessarily reflect
its proportion in the total population. This approach is used when a
particular stratum is small but of particular research interest, or
when there is greater variability within certain strata. For instance, a
researcher might oversample a smaller ethnic minority group to
ensure sufficient data for meaningful analysis, even if that group is
a small proportion of the total population. Weights are often used in
the analysis phase to adjust for the oversampling and ensure
generalizability to the overall population.
o Advantages: Ensures representation of key subgroups, increases
precision by reducing sampling error (especially for heterogeneous
populations), allows for separate analyses within each stratum.
o Disadvantages: Requires knowledge of population characteristics to form
strata, requires a complete sampling frame that includes information for
stratification, can be complex to implement, especially with many strata.
• 4. Cluster Sampling:
o Description: Cluster sampling is a probability sampling method where
the population is divided into naturally occurring groups or clusters (e.g.,
geographical areas like districts, cities, wards, or institutions like schools,
hospitals). Instead of sampling individual elements, the researcher
randomly selects a certain number of these clusters. Once a cluster is
selected, all elements within that selected cluster are then included in the
sample and investigated. The crucial assumption is that each cluster
should ideally be a mini-representation of the population as a whole,
meaning the variability within clusters should be similar to the variability in
the entire population.
o Process:
1. Divide the population into distinct clusters.
2. Randomly select a sample of clusters.
3. Include all elements from the selected clusters in the final sample.
o Example: To survey student opinions on a new curriculum in a large
school district, it would be impractical to list every student. Instead, the
researcher might identify all schools in the district (clusters). A random
sample of schools is selected, and then all students within those selected
schools are surveyed.
o Advantages: Highly efficient and cost-effective, especially for
geographically dispersed populations, as it reduces travel and listing
costs. It doesn't require a complete sampling frame of individual elements,
only a list of clusters.
o Disadvantages: Can be less precise than SRS or stratified sampling if
clusters are not truly heterogeneous within themselves (i.e., if clusters are
internally homogeneous but different from each other). This leads to a
higher sampling error compared to other probability methods for the same
sample size. Requires careful design to ensure clusters are as
representative as possible.
6. Multi-Stage Sampling:
o Description: Multi-stage sampling is a more complex form of cluster
sampling, involving successive stages of sampling, where larger units are
sampled first, and then smaller units within the selected larger units are
sampled. This technique is particularly useful for very large and
geographically widespread populations where a single-stage sampling
method would be impractical or impossible.
o Process: It involves multiple levels of clustering and random selection.
For instance, in the first stage, primary sampling units (PSUs) are
randomly selected. In the second stage, secondary sampling units (SSUs)
are randomly selected from within the chosen PSUs, and so on, until the
ultimate sampling units are reached.
o Example: To investigate the working efficiency of nationalized banks
across Bangladesh:
▪ Stage 1: Randomly select a few large primary sampling units, such
as divisions in the country.
▪ Stage 2 (Two-stage design): From the selected divisions,
randomly select certain districts. Then, interview all banks within
these chosen districts. This would form a two-stage sampling
design where the ultimate sampling units are clusters of districts.
▪ Stage 3 (Three-stage design): If, instead of interviewing all banks
within the selected districts, the researcher randomly selects
specific towns within those districts and then interviews all banks in
the chosen towns, this constitutes a three-stage sampling design.
o Advantages: Highly practical and efficient for large-scale national
surveys, reduces costs and logistical challenges, avoids the need for a
complete list of all elements in the population.
o Disadvantages: Can be more complex to design and analyze than single-
stage methods. Each stage of sampling introduces potential for sampling
error, meaning the total sampling error can be higher compared to single-
stage probability methods of comparable size. Requires careful
consideration of sampling at each stage to ensure overall
representativeness.
Types of Non-Probability Sampling Techniques
Non-probability sampling methods do not rely on random selection and are often used
in qualitative research, pilot studies, or when probability sampling is not feasible. The
inability to calculate sampling error limits the generalizability of findings to the
population. These methods include:
• 1. Convenience Sampling (Haphazard Sampling or Accidental Sampling):
o Description: Convenience sampling involves selecting participants who
are readily available, easily accessible, geographically proximate, or
willing to participate in the study. The selection is based purely on the
researcher's ease of access, without any systematic or random procedure.
o Example: Conducting a survey by standing outside a supermarket and
interviewing shoppers as they exit. Or, collecting data from students in a
specific classroom because they are easily available to the researcher.
o Advantages: Extremely affordable, quick, and easy to implement. It is
useful for preliminary research, pilot studies, or generating initial ideas and
hypotheses.
o Disadvantages: Highly prone to selection bias, as the sample is unlikely
to be representative of the population. Findings cannot be reliably
generalized to the larger population, making it unsuitable for studies
requiring statistical inference. The researcher has little control over the
representativeness of the sample.
• 2. Purposive or Judgmental Sampling:
o Description: In purposive sampling, also known as judgmental sampling,
the researcher deliberately selects participants based on their specific
characteristics, expertise, or knowledge that are directly relevant to the
study's purpose. The researcher uses their judgment to choose
"information-rich cases" that are expected to provide unique and valuable
insights. This method is typically used in qualitative research where the
goal is to gain in-depth understanding from specific individuals or groups
rather than broad generalization.
o Example: If studying the experiences of expert software developers, a
researcher would purposefully select individuals who have extensive
experience and a deep understanding of software development, rather
than randomly selecting people from the general population. Or,
interviewing key informants who are well-informed about a particular
phenomenon of interest.
o Advantages: Allows researchers to target specific, knowledgeable
individuals or groups, providing rich and relevant qualitative data. Cost-
effective for reaching a specific niche.
o Disadvantages: High potential for researcher bias, as the selection
depends entirely on the researcher's judgment. The generalizability of
findings is severely limited, as the sample is not statistically
representative. Reliability depends heavily on the researcher's expertise in
selecting appropriate participants.
• 3. Quota Sampling:
o Description: Quota sampling involves setting specific quotas for different
subgroups within the population and then conveniently selecting
participants until each quota is met. The researcher decides on the
proportions of various subgroups (e.g., age groups, genders, income
levels) that need to be represented in the sample, similar to stratified
sampling. However, unlike stratified sampling, the selection within each
subgroup is non-random (usually convenience-based).
o Example: For a study on employee career goals in an organization of
500 employees, if the researcher wants equal representation of males and
females, they might set a quota of 250 males and 250 females. They
would then interview the first 250 available males and 250 available
females until the quotas are filled.
o Advantages: Relatively quick and inexpensive to implement, especially
when a sampling frame is unavailable. Ensures representation of specific
characteristics in the sample, which can be useful for comparing
subgroups.
o Disadvantages: Introduces selection bias because elements are not
randomly selected within the quotas. The chosen participants might not be
representative of their respective subgroups (e.g., only easily accessible
males/females might be selected). Generalizability to the population is
limited.
• 4. Snowball Sampling (Networking Sampling):
o Description: Snowball sampling is a non-probability technique primarily
used when the target population is rare, difficult to locate, or hidden (e.g.,
individuals with specific rare diseases, members of underground
organizations, or professionals in niche fields). It starts with identifying one
or a few initial participants (the "seeds") who meet the study criteria.
These initial participants are then asked to refer the researcher to other
individuals they know who also meet the criteria. The sample grows like a
snowball rolling downhill, gathering more participants through referrals.
o Example: If studying the experiences of undocumented immigrants, a
researcher might start by contacting one or two individuals who fit the
profile. These individuals then refer the researcher to others within their
network, who in turn refer more, and so on.
o Advantages: Highly effective for reaching hidden or hard-to-access
populations where a traditional sampling frame does not exist or is
impractical. It can uncover participants who might otherwise remain
unknown to the researcher.
o Disadvantages: High potential for bias, as the sample is likely to be
homogenous, comprising individuals connected within the same social
networks. This limits the diversity of the sample and thus its
generalizability. It's impossible to determine the sampling error or how
representative the sample is of the larger hidden population. The sample
size is often small and dependent on the willingness of initial participants
to refer others.