SAMPLING
CENSUS AND SAMPLE SURVEY
All items in any field of inquiry constitute a ‘Universe’ or ‘Population.’ A complete
enumeration of all items in the ‘population’ is known as a census inquiry. It can be presumed
that in such an inquiry, when all items are covered, no element of chance is left and highest
accuracy is obtained. But in practice this may not be true. Even the slightest element of bias in
such an inquiry will get larger and larger as the number of observation increases. Moreover,
there is no way of checking the element of bias or its extent except through a resurvey or use
of sample checks. Besides, this type of inquiry involves a great deal of time, money and energy.
Therefore, when the field of inquiry is large, this method becomes difficult to adopt because of
the resources involved. At times, this method is practically beyond the reach of ordinary
researchers. Perhaps, government is the only institution which can get the complete
enumeration carried out. Even the government adopts this in very rare cases such as population
census conducted once in a decade. Further, many a time it is not possible to examine every
item in the population, and sometimes it is possible to obtain sufficiently accurate results by
studying only a part of total population. In such cases there is no utility of census surveys.
However, it needs to be emphasised that when the universe is a small one, it is no use resorting
to a sample survey. When field studies are undertaken in practical life, considerations of time
and cost almost invariably lead to a selection of respondents i.e., selection of only a few items.
The respondents selected should be as representative of the total population as possible in order
to produce a miniature cross-section. The selected respondents constitute what is technically
called a ‘sample’ and the selection process is called ‘sampling technique.’ The survey so
conducted is known as ‘sample survey’. Algebraically, let the population size be N and if a part
of size n (which is < N) of this population is selected according to some rule for studying some
characteristic of the population, the group consisting of these n units is known as ‘sample’.
Researcher must prepare a sample design for his study i.e., he must plan how a sample should
be selected and of what size such a sample would be.
IMPORTANT TERMS RELATED TO SAMPLING
Population
In a statistical investigation the interest usually lies in the assessment of the general magnitude
and the study of variation with respect to one or more characteristics relating to individuals
belonging to a group. This group of individuals under study is called as population or universe.
So, population is defined as, “an aggregate of objects, animate or inanimate under study.” It
may be finite or infinite.
Sample
A finite subset of statistical individuals in a population is called sample. It is quite often used
in our day to day life.
For example: Assessing the quality of foodgrain by taking a handful of it from the bag.
WHAT IS SAMPLING?
A Sampling is a part of the total population. It can be an individual element or a group of
elements selected from the population. Although it is a subset, it is representative of the
population and suitable for research in terms of cost, convenience and time. The sample group
can be selected based on a probability or a non-probability approach. A sample usually consists
of various units of the population. The size of the sample is represented by ‘n’.
Sampling is the act, process, or technique of selecting a representative part of a population for
the purpose of determining the characteristics of the whole population. In other words, the
process of selecting a sample from a population using special sampling techniques called
sampling. It should be ensured in the sampling process itself that the sample selected is
representative of the population. Though the sampling is not new but sampling theory has been
developed recently. People know or not but they have been using the sampling technique in
their day to day life.
For example:
A housewife tests a small quantity of rice or wheat to see whether it is of a good quality and
gives the generalised result about the whole rice kept in the bag or vessel. The result arrived at
most of the times is 100% correct.
Sample size- the number of individuals in the sample is called sample size
However the researcher has to key the following points in his mind while deciding the size of
the sample.
• Nature of the Universe: When the items of the universe are homogenous, a small
sample can serve the purpose, suppose they are heterogeneous, a large sample would
be required. Number of groups: When a researcher forms class – groups a large sample
is necessary as a small sample might not be able to give a reasonable number of items
in each class-group.
• Nature of study: When the researcher examines the items very intensively and
continuously then the sample should be small. He may prefer general survey when the
size of the sample is large but a small sample is considered appropriate in technical
surveys.
• Sample Technique: The researcher has to decide the sampling tools while determining
the size of the sample A small random sample is better than a larger but badly selected
sample.
• Accuracy and confidence level: A researcher requires a large size sample when the
accuracy or the level of precision is to be kept high. To get more accuracy for a fixed
significance level the samples size has to be increased fourfold.
• Resources available: What amount of time and financial resources are available to the
researcher will determine the size of sample, with sufficient time and large volume of
funds available the sample size could be large otherwise it should be small.
• Miscellaneous factors: In addition to the above considerations the following points to
be considered by a researcher. Nature of units sizes of the population size of
questionnaire availability or trained investigators the conditions under which the
sample is being conducted the time available for completion of the study.
The need of a sample:
There are large economic benefits of selecting a sample rather than conducting a census. The
cost of taking a census survey may go up to lakhs of rupees interviewing all 5000 employees
of an organization. It can be identified what is to be known by choosing a sample of few
hundred. The quality of a study conducted with a sample is usually more than with a population.
Research findings provide competent evidence of this opinion.
In one study, more than 90 percent of the total survey error was from non-sampling sources,
and only 10 percent was error from random sampling. The results of a study from are quicker
than from a study of census. The speed of execution reduces the time between recognizing of
a requirement for data and the data within reach.
When the population of the study is small and the variability high, as well as the components
completely different from each other, the census study is more appropriate. If the universe is
small and the variability is high, the selected sample may not be representative of the universe.
The results drawn from the sample are not accurate as estimates of the population values.
When the sample is taken appropriately, however, some sample elements underestimate the
parameters and some other overestimates them. Variations in these values act in opposition to
each other this counteraction arises in sample values that is usually near to the population value.
For these offsetting effects to happen, however, there must be adequate members in the sample,
and they must be drawn in a way that neither underestimation nor overestimation occurs.
CRITERIA OF SELECTING A SAMPLING PROCEDURE
In this context one must remember that two costs are involved in a sampling analysis viz., the
cost of collecting the data and the cost of an incorrect inference resulting from the data.
Researcher must keep in view the two causes of incorrect inferences viz., systematic bias and
sampling error.
Systematic Bias: A systematic bias results from errors in the sampling procedures, and it cannot
be reduced or eliminated by increasing the sample size. At best the causes responsible for these
errors can be detected and corrected.
Usually a systematic bias is the result of one or more of the following factors:
• Inappropriate sampling frame: If the sampling frame is inappropriate i.e., a biased
representation of the universe, it will result in a systematic bias.
• Defective measuring device: If the measuring device is constantly in error, it will result
in systematic bias. In survey work, systematic bias can result if the questionnaire or the
interviewer is biased. Similarly, if the physical measuring device is defective there will
be systematic bias in the data collected through such a measuring device.
• Non-respondents: If we are unable to sample all the individuals initially included in the
sample, there may arise a systematic bias. The reason is that in such a situation the
likelihood of establishing contact or receiving a response from an individual is often
correlated with the measure of what is to be estimated.
• Indeterminancy principle: Sometimes we find that individuals act differently when kept
under observation than what they do when kept in non-observed situations. For
instance, if workers are aware that somebody is observing them in course of a work
study on the basis of which the average length of time to complete a task will be
determined and accordingly the quota will be set for piece work, they generally tend to
work slowly in comparison to the speed with which they work if kept unobserved. Thus,
the indeterminancy principle may also be a cause of a systematic bias.
• Natural bias in the reporting of data: Natural bias of respondents in the reporting of data
is often the cause of a systematic bias in many inquiries. There is usually a downward
bias in the income data collected by government taxation department, whereas we find
an upward bias in the income data collected by some social organisation. People in
general understate their incomes if asked about it for tax purposes, but they overstate
the same if asked for social status or their affluence. Generally in psychological surveys,
people tend to give what they think is the ‘correct’ answer rather than revealing their
true feelings.
Sampling Error
Sampling errors are the random variations in the sample estimates around the true population
parameters. Since they occur randomly and are equally likely to be in either direction, their
nature happens to be of compensatory type and the expected value of such errors happens to
be equal to zero. Sampling error decreases with the increase in the size of the sample, and it
happens to be of a smaller magnitude in case of homogeneous population.
Sampling error can be measured for a given sample design and size. The measurement of
sampling error is usually called the ‘precision of the sampling plan’. If we increase the sample
size, the precision can be improved. But increasing the size of the sample has its own limitations
viz., a large sized sample increases the cost of collecting data and also enhances the systematic
bias. Thus the effective way to increase precision is usually to select a better sampling design
which has a smaller sampling error for a given sample size at a given cost. In practice, however,
people prefer a less precise design because it is easier to adopt the same and also because of
the fact that systematic bias can be controlled in a better way in such a design.
These errors are broadly classified as sampling errors and non-sampling errors.
1. Biased errors
2. Unbiased errors
Biased Errors:
Biased errors are understood as the inference of the investigators likes and dislikes in the
process of sampling. For sample if an investigator has to collect data from a specific group
also. This may because of investigator’s urge to complete the work early or failure to
understand the purpose of the survey. Such a mistake may result in collection of wrong data
which eventually will result only in wrong conclusions or inferences about the population. The
following are the reasons for biased errors.
• Faulty process of selection: This refers to a situation when the investigator does not
apply the randomness in his choice or selection of the sample elements from the
population.
• Faulty collection of information; Adoption of faulty method of collecting information
may cause errors. This will happen if the scope is not clear.
• Faulty method of analysis: This will happen when the researcher is not having
knowledge about the usage of tools.
Un Biased Errors:
Non-sampling errors are those errors, which are not due to any sampling process. It is due to
several other causes. Such errors are most due to the following reasons:
• Investigators may collect data without using complete schedules or proper
measurement. As a result data collected may not be relevant at all.
• Faulty method of interview or observations may also contribute to non-sampling errors.
Using of untrained and un skilled investigators.
In brief, while selecting a sampling procedure, researcher must ensure that the procedure causes
a relatively small sampling error and helps to control the systematic bias in a better way.
CHARACTERISTICS OF A GOOD SAMPLE DESIGN
From what has been stated above, we can list down the characteristics of a good sample design
as under:
1. Sample design must result in a truly representative sample.
2. Sample design must be such which results in a small sampling error.
3. Sample design must be viable in the context of funds available for the research study.
4. Sample design must be such so that systematic bias can be controlled in a better way.
5. Sample should be such that the results of the sample study can be applied, in general,
for the universe with a reasonable level of confidence.
SAMPLING DESIGN- A sample design is a definite plan for obtaining a sample from a given
population. It refers to the technique or the procedure the researcher would adopt in selecting
items for the sample. Sample design may as well lay down the number of items to be included
in the sample i.e., the size of the sample. Sample design is determined before data are collected.
There are many sample designs from which a researcher can choose. Some designs are
relatively more precise and easier to apply than others. Researcher must select/prepare a sample
design which should be reliable and appropriate for his research study.
STEPS IN SAMPLE DESIGN:
1. Objective: the first step of the sampling design is to define the objectives of survey in
clear and concrete terms. The sponsors or the researchers of the survey should confirm
that the objectives are commensurate with the money, manpower, and time limit
available for the survey.
2. Population: in order to meet the objectives of the survey, what should be the
population? This question should be answered in the second step. The population
should be clearly defined.
3. Sampling units and frame: a decision has to be taken concerning a sampling unit
before selecting sample. Sampling unit may be a geographical one such as state, district,
village, etc., or a construction unit such as house, flat, etc., or it may be a social unit
such as family, club, school, etc., or it may be an individual. The researcher will have
to decide one or more such units that he has to select for his study. The list of sampling
units is called as “frame” or sampling frame. Sampling frame contains the names of all
items of a universe (in case of finite universe only). Such a list should be
comprehensive, correct, reliable and appropriate. It is extremely important for the
source list to be as representative of the population as possible.
4. Size of sample: this refers to the number of items to be selected from the universe to
constitute a sample. This is a major problem before a researcher. The size of the sample
should neither be excessively large, nor too small. It should be optimum, an optimum
sample is one which fulfils the requirements of efficiency, representativeness, reliability
and flexibility. While deciding the size of sample, researcher must determine the desired
population as also an acceptable confidence level for the estimate. The size of the
population variance needs to be considered as in case of larger variance usually a bigger
sample is needed. The size of population must be kept in view for this also limits the
sample size. The parameters of interest in a research study must be kept in view, while
deciding the size of the sample. Cost too dictate the size of the sample that we can draw.
As such, budgetary constraint must invariably be taken into consideration when we
decide the sample size.
5. Parameters of interest: statistical constants of the population are called as parameters,
eg., population mean, population proportion, etc., when we do census survey we get the
actual value of parameters. On the other hand, when we do sample survey we get the
estimates of the unknown population parameters in place of their actual values.
In determining the sample design, one must consider the question of the specific
population parameters which is of interest. For instance, we may be interested in
knowing some average or the other measure concerning the population. There may also
be important sub- groups in the population about whom we would like to make
estimates. All this has a strong impact upon the sample design we would accept.
6. Data collection: no irrelevant information should be collected and no essential
information should be discarded. The objectives of the survey should be very much
clear in the mind of surveyor.
7. Non- respondents: Because of practical difficulties, data may not be collected for all
the sampled units. This non- response tends to change the results. The reasons for non-
response should be recorded by the investigator. Such cases should be handled with
caution.
8. Selection of proper sampling design: the researcher must decide the type of sample
he will use ie., he must decide about the technique to be used in selecting the items for
sample. In fact, this technique or procedure stands for the sample design itself. There
are several sample designs out of which the researcher must choose one for his study.
Obviously, he must select that design which, for a given sample size and for a given
cost, has a smaller sampling error.
9. Organising field work: the success of a survey depends on the reliable field work.
There should be efficient supervisory staff and trained personnels for the field work.
10. Pilot survey: it is always helpful to try out the research design on a small scale before
going to the field. This is called “Pilot Survey” or “Pretest”. It might give the better
idea of practical problems and troubles.
11. Budgetary constraint: cost considerations, from practical point of view, have a major
impact upon decisions relating to not only the size of the sample but also to the type of
sample. This fact can even lead to the use of non-probability sample.
SAMPLING METHODS/ TECHNIQUES/ TYPES:
In Statistics, the sampling method or sampling technique is the process of studying the
population by gathering information and analyzing that data. It is the basis of the data where
the sample space is enormous.
There are several different sampling techniques available, and they can be subdivided into two
groups. All these methods of sampling may involve specifically targeting hard or approach to
reach groups.
1. Probability Sampling
• Simple random sampling
• Systematic sampling
• Stratified sampling
• Multistage cluster sampling
2. Non-probability Sampling
• Convenience sampling
• Quota sampling
• Judgment sampling
• Snowball sampling
Probability Sampling
A sampling in which every member of the population has a calculable and non-zero probability
of being included in the sample is known as probability sampling. Methods of random selection
consistent with both the probabilities of inclusion are used in forming estimates from the
sample. The probability of selection need not be equal for members of the population. If the
purpose of the research is to arrive at conclusions or make predictions affecting the population
as a whole, then the choice of a probabilistic sampling approach is desirable. This method is
more time consuming and expensive than the non-probability sampling method. The benefit of
using probability sampling is that it guarantees the sample that should be the representative of
the population.
Probability Sampling Types
Probability Sampling methods are further classified into different types, such as simple random
sampling, systematic sampling, stratified sampling, and clustered sampling.
Simple Random Sampling
A sampling process where each element in the target population has an equal chance or
probability of inclusion in the sample is known as simple random sampling. In small
populations random sampling is done without replacement to avoid the instance of a unit being
sampled more than once. Since the item selection entirely depends on the chance, this method
is known as “Method of chance Selection”. As the sample size is large, and the item is chosen
randomly, it is known as “Representative Sampling”.
For example, if a sample of 15,000 names is to be drawn from the telephone directory, then
there is equal chance for each number in the directory to be selected. These numbers (serial no.
of names) could be randomly generated by the computer or picked out of a box. These numbers
could be later matched with the corresponding names thus fulfilling the list.
Suppose we want to select a simple random sample of 200 students from a school. Here, we
can assign a number to every student in the school database from 1 to 500 and use a random
number generator to select a sample of 200 numbers.
Systematic Sampling
In some instances, the most practical way of sampling is to select every ith item on a list. This
type is known as systematic sampling. An element of randomness is introduced into this kind
of sampling by using random numbers to pick up the unit with which to start. For instance, if
a 4% sample is desired, the first item would be selected randomly from the 1st 25 and thereafter
after 25th item would automatically be included in the sample. Thus, in systematic sampling
only the first unit is selected randomly and the remaining units of the sample are selected at
fixed intervals. Although a systematic sample is not a random sample in the strict sense of the
term, but it is often considered reasonable to treat systematic sample as if it were a random
sample.
Stratified Sampling
If a population firm which a sample is to be drawn does not constitute a homogeneous group,
stratified sampling technically is generally applied in order to obtain a representative sample.
Under this, the population is divided into several sub- populations that are individually more
homogeneous than the total population (the different sub- populations are called “strata”) and
then we select items from each stratum to constitute a sample. Since each stratum is more
homogeneous than the total population, we are able to get more precise estimates for each
stratum and by estimating more accurately each of the component parts, we get a better estimate
of the whole. In brief, stratified sampling results in more reliable and detailed information.
The following 3 questions are highly relevant in the context of stratified sampling
a) How to frame strata? The strata be formed on the basis of common characteristics of
the items to be put in each stratum. This means that the various strata be formed in such
a way as to ensure elements being most homogeneous within each stratum and most
heterogeneous between the different strata. Thus, strata are purposively formed and are
usually based on past experience and personal judgement of the researcher. One should
always remember that careful consideration of the relationship between the
characteristics of the population and the characteristics to be estimated are normally
used to define the strata. At times, pilot study may be conducted for determining a more
appropriate and efficient stratification plan. We can do so by taking small samples of
equal size from each of the proposed strata and then examining the variances within
and among the possible stratifications.
b) How should items be selected from each stratum? We can say that the usual method,
for selection of items for the sample from each stratum, resorted to is that of simple
random sampling. Systematic sampling can be used if it is considered more appropriate
in certain situations.
c) How many items be selected from each stratum or how to allocate the sample size of
each stratum? We usually follow the method of proportional allocation under which the
sizes of the samples from the different strata are kept proportional to the sizes of the
strata.
Cluster Sampling
Clustering involves grouping the population into various clusters and selecting few
clusters for study. Cluster sampling is suitable for conducting research studies that cover
large geographic area. Once the cluster is formed the researcher can either go for one
stage, two stages, or multistage cluster sampling. In single stage, all the elements from
each selected are studied, whereas in two stages, the researchers use random to select few
elements from clusters. Multistage sampling involves selecting a sample in two or more
successive stages. Here the cluster selected in the first stage can be divided into cluster
units.
For example consider the case where a company decides to interview 400 households
about the likeability of its new detergent in a metropolitan city. To minimize the resources
and time researchers divide the city into separate blocks say 40, each block consist of
heterogeneous units. The researcher may opt for the two stage cluster sampling if he finds
that individual clusters have little heterogeneity to other clusters. Similarly a multistage
cluster sampling involves three or more sampling steps, it differs from stratified sampling
that is done in cluster in contrast to elements within strata as is the case in the stratified
sampling. Elements are randomly selected from each stratum in case of stratified sampling
whereas only selected clusters are studied in cluster sampling.
Non- Probability Sampling
The non-probability sampling method is a technique in which the researcher selects the
sample based on subjective judgment rather than the random selection. In this method, not
all the members of the population have a chance to participate in the study. It involves the
selection of units based on factors other than random chance. It is also known as deliberate
sampling and purposive sampling. For example, a scheme whereby units are selected
purposefully would yield a non-random sample. In a general sense, it is an umbrella term,
which includes any sample that does not conform to the requirements of a probability
sampling. Convenience sampling, quota sampling, judgment sampling and snowball
sampling are few examples of non- probability sampling.
Convenience Sampling
The selection of units from the population based on their easy availability and accessibility
to the researcher is known as convenience sampling. For example, imagine a Co., that
surveys a sample of its employees to know the acceptance for a new flavor of potato chips
that it plans to introduce in the market. This type of sampling is a typical example of
convenience sampling as the criterion for selecting a sample is convenience and
availability. Although this type of research is easy and cost effective, the findings of the
sample survey cannot be generalized to the entire population, as the sample is not
representative. As there is no set criterion for selecting the sample, there is a scope for
research being influenced by the bias of the researcher. As in the above ex, the researcher
may conduct a sample survey involving its own employees to find whether the market,
would accept the product.
In researching customer support services in a particular region, we ask your few customers
to complete a survey on the products after the purchase. This is a convenient way to collect
data. Still, as we only surveyed customers taking the same product. At the same time, the
sample is not representative of all the customers in that area.
Quota Sampling
In the quota sampling method, the researcher forms a sample that involves the individuals to
represent the population based on specific traits or qualities. The researcher chooses the sample
subsets that bring the useful collection of data that generalizes the entire population.
In quota sampling, the entire population is segmented into mutually exclusive groups. The
number of respondents (quota) that are to be drawn from each of several categories is
specified in advance and the final selection of respondents is left to the interviewer who
proceeds until the quota for each category is filled. Quota sampling finds extensive use in
commercial research where the main objective is to ensure that the sample represents in
relative proportion, the people in the various categories in the population, such as gender,
age group, social class, ethnicity and region of residence. For example, if a researcher
wants to segment the entire population based on gender, then he would have two
categories of respondents, that is, males and females. If he plans to collect a sample of 30,
he may allot a quota of 15 for male and 15 for female respondents. Therefore, the
researcher will stop administering the questionnaire to females after he interviews the 15th
female respondent, that is, when the quota of 15 females is filled.
Judgmental Sampling
The selection of a unit, from the population based on the judgment of an experienced
researcher, is known as judgment or purposive sampling. Here, the sample units are
selected based on population’s parameters. It is often noticed that companies frequently
select certain preferred cities during test marketing their products. This is because they
consider the population of that particular city to be representative of the total population
of the country. The same is the case with the selection of specific shopping malls that
according to the researcher’s judgment attract a reasonable number of customers from
different sections of the society. Polling results predicted on television is also a result of
judgment sampling. Researchers select those districts that have voting patterns close to
the overall state or country in the previous year. The judgment of the researcher is based
on the assumption that the past voting trends of selected sample districts are still
representative of the political behavior of the state’s population. For example, certain
companies test market their new product launches in cities like Mumbai and Bangalore,
because the profile of these cities is representative of the total Indian population.
Snowball Sampling
Sampling procedures that involve the selection of additional respondents are known as
snowball sampling. This sampling technique is used against low incidence or rare
populations. Sampling is a big problem in this case, as the defined population from which
the sample can be drawn is not available. Therefore, the process sampling depends on the
chain system of referrals. Suppose, SG sports Ltd., a manufacturer of sports equipment
plans to survey 100 senior players through its new website for getting their feedback on
the quality of its products.
However, keeping track of such senior squash players can be very difficult, as their
presence may be very rare or low. Therefore, it collects the details of the first 200 visitors
to its website, to list if any of them is a squash player or knows a squash player. If the
visitor is a squash player, then he is requested to refer the names of at least 3 other players
known to him. The referred names of the squash players are then called upon for further
referrals and this goes on until the sample size of 100 adult players is reached. Although
small sample sizes and low costs are the clear advantages of snowball sampling, bias is
one of its disadvantages. The referral names obtained from those sampled in the initial
stages may be similar to those initially sampled. Therefore, the sample may not represent
a cross-section of the total population. It may also happen that visitors to the site or
interviewers may refuse to disclose the names of those whom they know.