Sampling
Sampling is indispensable technique of social science research; the research work cannot be
undertaken without use of sampling. Before research work, the investigator has to decide
whether the entire population is to be made subject for data collection or a particular group is
to be selected as representative of the entire population. The study of the total population is
not possible and it is also impracticable. The concept of sampling has been introduced with a
view to making the research findings economical and accurate.
The former method when the entire population is taken into account is called ‘Census Method’.
On the other hand, when a small group is taken into account as representative of the whole is
called ‘Sampling Method’. In social science research, it is not possible to study all the units in
the universe or population and it is not required even. Hence, a researcher selects samples. A
sample is a portion of the people drawn from a larger population. A sample simply means a
smaller representation of the larger whole. Sampling is the selection part of an aggregate or
totality known as population, on the basis of which a decision concerning the population is
made. Thus, we can say that a finite subset of statistical individuals in a population is called a
sample and the number of individuals in a sample is called sample size.
Definition
In a simple sense, sampling refers to the method used to select a given number of people (or
things) from a population.
According to Mildred Parton, “Sampling method is the process or the method of drawing a
definite number of the individuals, cases or the observations from a particular universe,
selecting part of a total group for investigation.”
P.V. Young - “A statistical method is a miniature picture or cross-selection of the entire group
or aggregate from which sample is taken.
Bogardus – “Sampling is the selection of certain percentage of a group item according to a pre-
determined plan.”
According to Manheim, “a sample is part of the population which is studied in order to make
inferences about the whole population”.
William [Link] and [Link] – “A sample, as the name implies is smaller representative
of a larger whole”.
FEATURES OF SAMPLING
Sampling Method selects for its study only a percentage of the entire population. One of the
important features of samples is that it should have the attribute of representativeness.
1
[Link]: Representativeness is one of the important features of samples is that it
should have the attribute or representativeness .A representative sample is one in which the
odds are considered good enough that the selected sample is sufficiently representative to
justify generalizing from the sample of the population .If the sample is similar to the universe
in all respects representative may be absolute .When selecting a sample the researcher must
have complete details of the population .Haphazard method will lead to waste of time and
resources.
[Link]: Another important feature of sample is Reliability. A sample is said to be reliable
when it is free from bias. Reliability of the sample very much depends on homogeneity of the
sample. Homogeneity of the sample simply means that the sample should have all the traits
that the universe exhibits. Reliability of the sample will be doubled if it’s not a representative
sample
[Link]: Accuracy is defined as the degree to which bias is absent from the sample. An
accurate (Unbiased) sample is one which exactly represents the population. It is free from any
influence that causes any difference between sample value and population value.
[Link]: The sample must yield precise estimate. Precision is measured by the standard
error or standard deviation of the sample estimate. The smaller the standard error or estimate,
the higher is the precision of the sample
[Link]: A good sample must be adequate in size in order to be reliable. The sample should be
of such that the inferences drawn from the sample are accurate to a given level of confidence
[Link]: A sample is said to be adequate when it is of sufficient size to allow confidence
in the stability of its characteristics. According to Goode and Hatt, “a sample is adequate when
it is of sufficient size to allow confidence in the stability of its characteristics.” in other words
it should not be too large than is necessary.
[Link] Bias: According to Moser and Kalton bias in selection of sample can arise: If the
sample is done by a non –random method which generally means that the selection is
consciously or unconsciously influenced by human choice.
If the sampling frame (list, index or other population record) which serves as the basis for
selection does not cover the population adequately, completely or accurately.
If some section of the population is impossible to find or refuse to cooperate.
Sampling advantages
Economy of time: In this method, we study representative units and get the results, which we
would get after the study of the universe. Thus, we save time.
2
Economy of resources: Since in sampling method, we make a selective study of a
representative unit and the advantage of saving resources. Although we have resources, yet we
get the same results as are achieved after the study of the entire universe.
Detailed study: Because in the sampling method, the area of the study is small and so it is not
possible for us to make a detailed and intensive study.
Accuracy of study: We have the advantage of detailed and intensive study and so the results
are generally more accurate and reliable. Accuracy is achieved more because of the area of the
study is small and so we are able to control all possible situations that are required for study.
Administrative convenience: Social research generally deals with human beings and social
groups who have their own ways. If the group that we are studying is large, it is not possible
for us to put them together and make a proper study. In sampling method, since the area of the
study is small, it is possible for carryout the work in an efficient manner.
Difficulties of the census method are not faced: Census Method” implies the study of the
entire population. The field of the social research, deals with the human beings and social
groups, and so the ‘census method’ of study is not only difficult but sometimes impossible. The
sampling method or sampling meets the difficulties of the ‘census method’ and makes the
impossibilities of the method as possible. In census method, it is not possible to cover the entire
universe, but through ‘sampling method’ the whole universe can be covered in an
administratively convenient manner and hence we save time, money, and energy.
DISADVANTAGE OF SAMPLING
Possibilities of bias and prejudices: If the method of sampling is faulty or the nature of the
phenomenon has its faults, our study of problem may be biased and prejudiced. The result can
be defective if prejudices and bias creeps into this method.
lack of representation: It is not easy to select a representative sample, particularly when the
phenomenon, which we are studying, is a social phenomenon, which is of a very complex
nature. Because of this difficulty, sampling method is said to be disadvantageous.
Need for specific and specialized knowledge: In order to employ successfully and
scientifically the sampling method, it is necessary for the investigator and this lack of specific
and specialized knowledge poses difficulties in proper use of sampling method.
IMPORTANCE OF SAMPLING
Sampling is a critical aspect of social research because it allows researchers to study a subset
of the population in order to make inferences about the larger population. Here are some of the
key reasons why sampling is significant in social research:
3
Cost-effectiveness: Conducting research on an entire population can be prohibitively
expensive, time-consuming, and sometimes impossible. Sampling allows researchers to study
a smaller subset of the population that is still representative of the larger population, which can
be more cost-effective and practical.
Generalizability: The goal of social research is often to make generalizations about a
population based on the findings from a sample. Sampling techniques can help ensure that the
sample is representative of the population, which can increase the generalizability of the
findings
Precision: Sampling techniques can help ensure that the sample is not biased, which can
increase the precision of the findings. For example, if a sample is selected using a random
sampling technique, it is less likely to be biased towards any particular group or characteristic
in the population
Ethical considerations: In some cases, it may be unethical or impractical to study an entire
population. Sampling can help researchers avoid putting the entire population at risk or causing
undue harm
Feasibility: In some cases, it may simply be impossible to study the entire population, such as
in cases where the population is very large, geographically dispersed, or otherwise difficult to
access. Sampling can make it feasible to study the population by focusing on a smaller, more
manageable subset.
Reduce errors: Sampling can help reduce errors and increase the accuracy of research
findings. For example, if a researcher studies the entire population, they may encounter errors
such as measurement error, response bias, or sampling bias. By using a sample, researchers can
minimize these errors and increase the precision of their findings.
Variety of methods: There are various sampling techniques available to researchers, and
different methods may be appropriate depending on the research question, population size, and
other factors. By carefully selecting the appropriate sampling method, researchers can obtain a
sample that is representative of the population and minimize biases.
Accessibility: Sampling allows researchers to study populations that may be difficult to access
or rare. For example, researchers may be interested in studying a particular group of people
with a rare disease or a group that is geographically dispersed. By using a sampling technique,
researchers can access these populations and study them more easily.
Time-efficiency: Sampling can save time in social research by allowing researchers to obtain
a smaller sample size that can be studied more easily and quickly than the entire population.
4
Replicability: Sampling makes it possible for other researchers to replicate studies and test the
validity of findings. If the sampling method is described in detail, other researchers can use the
same method to obtain a similar sample and test the validity of the original findings. This
increases the reliability and credibility of social research.
Thus, sampling is significant in social research because it allows researchers to obtain accurate,
representative, and generalizable findings while minimizing costs, time, and ethical concerns.
Basic concepts
Parameters: A parameter is a numerical value that describes a characteristic of the whole
population. It is to be noted that parameters are usually unknown and that they are used to
summarize the whole population. It is fixed , unknown value that researchers are interested in
estimating or testing hypothesis about
Parameters can include measures such as the population mean, population proportion,
population standard deviation etc. For example, if studying the average height of all adults in
a country, the parameter of interest would be the population mean height.
Since it is usually impractical or impossible to measure the entire population, researchers often
collect data from a sample and use statistical methods to estimate the population parameter.
It is important to distinguish from parameters from statistics, as parameters describe the
population, while statistics are values calculated from sample data that estimate or describe the
corresponding population parameter.
Parameters are conventionally denoted by Greek alphabets. Greek letter mu (μ) and the
population standard deviation by the Greek letter sigma (σ). Parameters are fixed constants,
that is, they do not vary like variables. However, their values are usually unknown because it
is infeasible to measure an entire population.
It is important to note that the value of a parameter is computed from all the population
observations. Thus, the parameter 'mean income' is calculated from all the income figures of
different individuals that constitute the population. Similarly, for the calculation of the
parameter 'correlation coefficient of heights and weights’, we require the values of all the pairs
of heights and weights in a population.
Thus, we can define a parameter as a function of the population values. If θ is a parameter
that we want to obtain from the population values X1, X2, …. Xn, then
Statistic
5
When studying census and sample surveys, we learned that it is not always possible to collect
information from every unit of a population. Due to limitations such as time, cost, and
accessibility, calculating the exact population parameter (like population mean or population
standard deviation) may not be feasible.
In such cases, we draw a sample from the population and use the information from the sample
to estimate the population characteristics. The numerical value that we calculate from the
sample data is called a statistic.
In such situations, we try to get some idea about the parameter from the information obtained
from a sample drawn from the population. This sample information is summarized in the form
of a statistic. For example, sample mean or sample median or sample mode is called a statistic.
Thus, a statistic is calculated from the values of the units that are included in the sample.
A statistic is any numerical measure that is computed from the values of the units included in
a sample. It is a function of the sample observations.
Estimator and Estimate
The main purpose of calculating a statistic is to estimate a population parameter (such as
population mean, population proportion, or population variance).
Estimator: The basic purpose of a statistic is to estimate some population parameter. The
procedure followed or the formula used to compute a statistic is called an estimator and the
value of a statistic so computed is known as an estimate. An estimator is the rule, method, or
formula that we use to compute a statistic from the sample data.
It is a mathematical function of sample observations.
Estimate: An estimate is the numerical value obtained when we apply the estimator
(formula) to a particular sample.
ˉ 1
Consider the formula: 𝑋 = 𝑛 ∑𝑛𝑖=1 𝑥𝑖
This formula for calculating the sample mean is an estimator. This formula is an estimator of
the population mean.
Population: The entire group of individuals or items that you want to study. This can include
people, objects, or events. In social scientific research, a population is the cluster of people
you are most interested in; it is often the “who” that you want to be able to say something about
at the end of your study. For example, the population could be all students in a school or all
manufactured products in a factory. The total number of sampling units in the population is the
population size, generally denoted by N. The population size can be finite or infinite (N is
large).
6
Census: The complete count of the population is called a census. The observations on all the
sampling units in the population are collected in the census. For example, in India, the census
is conducted every tenth year, and observations of all the persons staying in India are collected.
Sample: A sample is the group of people you successfully recruit from your sampling frame
to participate in your study. A sample consists only of a portion of the population units. Such
a collection of units is called the sample. When all the salient features of the population are
present in the sample, then it is called a representative sample. If you are a participant in a
research project—answering survey questions, participating in interviews, etc.—you are part
of the sample of that research project.
Sampling Unit: A sampling unit is the single element that we select during sampling. A
sampling unit refers to the specific entity or group selected from a large population for the
purpose of data collection in research. It can be an individual, a household, an organization or
even a geographical are, depending on the research design and objectives. Understanding the
sampling unit is crucial as it directly impact the validly and relatability of the research finding.
A well-defined sampling unit helps to ensure that the sample accurately reflects the broader
population, which is essential for drawing valid conclusions
Sampling Error: A sampling error is the difference between the characteristics of a sample
(like its mean or proportion) and those of the entire population it represents. It occurs because
a sample is only a subset of the population, so it may not perfectly reflect the population’s true
values. Sampling errors occur because the sample is not representative of the population or is
biased in some way. It arises naturally whenever researchers use a sample instead of studying
the entire population.
Sampling Error: When we use statistics to make predictions, we often face differences
between what we expect and what happens. These differences are called errors. Even with
statistically sound sampling methods, the specific sample drawn may not perfectly mirror all
characteristics of the entire population. This inherent discrepancy is the sampling error.
Formula
The general formula for sampling error (for the mean) is: where σ = population standard
deviation, n = sample size. This formula shows that increasing sample size reduces sampling
error.
Types of Sampling Errors
There are different categories of sampling errors.
7
• Population-specific Error: A population-specific error occurs when a researcher
doesn’t understand who to survey.
• Selection Error: Selection error occurs when the survey is self-selected, or when only
those participants who are interested in the survey respond to the questions. Researchers
can attempt to overcome selection errors by finding ways to encourage participation.
• Sample Frame Error: A sample frame error occurs when a sample is selected from
the wrong population data.
• Nonresponse Error: A nonresponse error occurs when a useful response is not
obtained from the surveys because researchers were unable to contact potential
respondents (or potential respondents refused to respond).
• Random Error: It is due to the natural variation that occurs when a random sample is
selected from a population. It results from chance factors and is an inherent part of the
sampling process. The magnitude of unexpected errors can be reduced by increasing
the sample size.
• Systematic Error: It is due to factors that systematically bias the sample in a particular
direction. It is not due to chance factors; it can occur when the sampling method or
sampling frame is flawed. For example, if a researcher only collects samples from one
geographic region, this can lead to a biased sample if the population of interest is spread
across multiple areas.
Non-Sampling Error
Non-sampling error, on the other hand, is an error that occurs due to factors other than the
sampling process, such as errors in data collection, processing, or analysis. Non-sampling error
can arise from any stage of the research process, from study design to data analysis. Non-
sampling error, can be accidental or systematic and is often more challenging to quantify and
address.
Sampling frame
An intermediate point between the overall population and the sample that is drawn for the
research is called a sampling frame. A sampling frame is a list of all the units of the population
from which researchers draw a sample. All the sampling units in the sampling frame have
identification particulars. For example, all the students in a particular university listed and their
roll numbers constitute the sampling frame. Similarly, the list of households with the name of
the head of the family or house address constitutes the sampling frame. In another example,
the residents of a city area may be listed in more than one frame - as per automobile registration
and the listing in the telephone directory.
8
Developing Sampling Frames: Developing a sampling frame means creating a complete,
accurate list of all the elements (individuals, households, organizations, etc.) in the target
population from which a sample will be drawn. It is the foundation of reliable survey research
because the quality of the sampling frame directly affects the validity of the results. A well-
built frame minimizes sampling error and ensures that survey findings truly reflect the
population.
A sampling frame is the list or database that contains all members of the population you want
to study. It acts as the bridge between the target population and the sample you select. Example:
If your population is “all university students,” the sampling frame might be the official student
enrollment list.
Steps in Developing a Sampling Frame
1. Identify the Target Population: Define clearly who you want to study (e.g., residents
of Bengaluru North, small businesses in India, etc.).
2. Define Sampling Units: Decide the unit of analysis (individuals, households,
schools, companies).
3. Choose Data Sources: Use reliable sources such as census records, membership
lists, registries, or databases. For surveys, sources might include electoral rolls, phone
directories, or organizational records.
4. Construct the Frame: Compile the list ensuring completeness (all members
included) and accuracy (no duplicates, no outdated entries).
5. Check for Coverage Errors: Ensure no one is wrongly excluded or included.
Example: A phone directory misses people without landlines → coverage error.
6. Pilot Test the Frame: Run a small test sample to check representativeness and
identify gaps.
7. Evaluate and Update: Sampling frames must be regularly updated to remain
valid. Outdated frames lead to bias and unreliable results.
Common Challenges
• Incomplete lists: Some members of the population may not be recorded.
• Duplicates: Same unit appearing multiple times.
• Outdated information: People move, businesses close, numbers change.
• Accessibility issues: Legal or privacy restrictions may limit access to data.
Types of Sampling Frames
1. List -based sampling frames: Uses a comprehensive list of population members (e.g., voter
rolls, school registers, hospital patient lists).
9
2. Area-Based Sampling Frame: A frame based on geographic divisions (regions, districts,
blocks). Examples: Census tracts used for household surveys and City maps divided
into neighborhoods for urban studies.
3. Dual-Frame Sampling Frame: Combines two different frames to improve coverage
and reduce bias. Often used when one frame alone misses certain groups. Example:
Telephone surveys may use both landline and mobile phone directories to reach a broader
population.
4. Institutional Sampling Frame: Uses organizations or institutions as the basis. Examples:
Hospital patient records for health studies or School registers for educational research.
5. Constructed Sampling Frame: Built from scratch when no ready-made list exists.
Examples: Door-to-door enumeration in a village or social media scraping to identify
online communities.
Conclusion
Sampling frames can be list-based, area-based, dual-frame, institutional, or constructed. The
choice depends on your population, resources, and research goals. A well-developed frame
ensures representativeness and minimizes sampling error.
10