Sampling theory
• Selecting a small representative part from the larger whole under
study.
• Larger whole is – population
• Small part selected from the larger whole is – sample
• Ex: in opinion poll, a relatively small number of people are
interviewed , to get idea about the opinion of whole community
Parameter and statistics
• While studying a population, we are usually interested in one or more
unknown characteristics of the population: parameters
• The corresponding sample quantities (such as sample mean and variance),
often called sample statistics or statistics.
• Sometimes it is not possible from practical point of view to measure and
observe entire population, so we draw a sample from it
• Purpose of studying sampling theory is: to evolve sampling procedures so
that maximum efficiency can be achieved at a minimum cost.
• Procedure to observe the entire population is called complete
enumeration or census.
• A study of the inferences made concerning a population by using
samples drawn from it, together with indications of the accuracy of
such inferences by using probability theory, is called statistical
inference.
Population
• Finite and infinite
• Books in library; stars, yield of paddy field
• Population may be existent or hypothetical
• Electrical fuses; rolling of die
• Sample is a subset of population- scaled down version of population
• Sampling- without replacement or with replacement
• Sampling where each member of the population may be chosen more than
once is called sampling with replacement, while if each member cannot be
chosen more than once it is called sampling without replacement.
Sampling Distribution
• Consider all possible samples of size N that can be drawn from a given population (either
with or without replacement). For each sample, we can compute a statistic (such as the
mean and the standard deviation) that will vary from sample to sample. A statistic, being
based on sample observation, is a random variable. In this manner we obtain a
distribution of the statistic that is called its sampling distribution.
• Sampling Unit
• The ultimate unit sampled from the population and observed is called a sampling unit.
• Sampling frame:
• It is the complete list of population units from which the sample is drawn.
• For household survey, the list of all street addresses, for an agricultural survey, a list of all
farmers or a map of areas containing farms can serve as the sampling frame.
• Some popular sampling frames are census report, voters' list etc. Failing to include all the
units of the target population in the sampling frame is called under coverage.
• Sampling fluctuations:
• Sampling fluctuations are the differences in the values of a statistic
observed for different samples.
• Standard error (s.e.) :
• Standard deviation of a statistic in its sampling distribution is known
as its standard error. It is so called because it refers to the precision of
the statistic as an estimator.
Data Collection Methods
• The design of an enquiry the collection of data deserve serious attention.
• Careful and detailed planning in the initial stages can lead to saving in time and
money and improved accuracy.
• As complete a plan as possible should be drawn up before the actual collection of
information begins, specifying what data are to be obtained, from whom and by
what methods.
• There should also be full and unambiguous definitions of terms, clear
instructions to investigators and respondents and, maybe, some indications of
the mode of analysis of the results.
Census vs Sampling
• Although the plan should be complete, it should not be totally rigid, for some
adjustments to the plan will be inevitable as the collection and analysis of data
proceed.
• A fundamental question to be considered at the outset is whether the collection
of data should be done by complete enumeration or by sampling.
• In the former case, each and every individual of the group to which the data are
to relate is covered, and information gathered for each individual separately.
• In the latter, only some individuals forming a representative part of the group are
covered, either because the group is too large or because the items on which
information is sought are too numerous.
• Complete enumeration may lead to greater accuracy and greater refinement in
analysis, but it may be a very expensive and time-consuming operation. A sample
designed and taken with care can produce results that may be sufficiently
accurate for the purpose of the enquiry, and it can save much time and money
• In some cases, a combination of the census and sample methods may
by advisable. Thus frequent sample surveys may be used, as in
demography, to fill the gaps between censuses taken at regular
intervals.
• Or, some simple questions may be asked of every one, while more
complicated questions may be put only to a proportion (say 5 per
cent or 10 per cent) of all respondents in an enquiry.
• Note, again, that the information sought may be gathered, from the
individuals of the whole group (called the population) or from those
of the sample, by one of three methods—
• the questionnaire method, the interviewer method and the method
of direct observation.
The questionnaire method
• In economic and social enquiries, information is almost always collected by having someone to fill
up a form or questionnaire.
• But a matter to be decided is whether the forms should be completed by an enumerator or
investigator who collects data by asking questions and noting down answers, or whether these
should be left with the respondent to be filled up on his own.
• In the questionnaire method, each informant (or respondent) is provided with a questionnaire,
usually sent by mail with return postage prepaid, and is asked to supply the information in the
form of answers to the questions.
• Obviously, this method can be effective only when the informants have attained a certain level of
education.
• It can work, for instance, when a daily newspaper decides to conduct an opinion poll among its
readers on some topical issue.
• The drawback of the method is that the informants may not evince sufficient interest in the
enquiry even if they are sufficiently enlightened. Consequently, the data may involve a high
percentage of non-response and thus fail to reflect the true state of the field of enquiry
The interviewer method
• In the interviewer method, enumerators go from one informant to another and elicit the
required information. This method is used in population censuses. Also, it is the method that has
to be employed in case the informants are not all literate or, even if literate, have not attained the
requisite educational level.
• For instance, if one is interested in family income and expenditure on different items, one may
arrange to interview the head of each family and collect the information sought from him.
• The data collected by this method are likely to be more accurate, since a tactful investigator may
persuade the informant to supply the required information and the meaning of each question
may be properly explained to him so that the answers may be correct and to the point.
• Whichever of the two methods may be used, the questions and the accompanying instructions to
enumerators and respondents have to be very carefully designed.
• It is necessary that each question be clearly phrased and capable of an
unambiguous answer.
• The instructions must take into account all possibilities, even remote ones.
• The way a question is put may well influence the answer, as those who have
conducted public opinion polls will bear out.
• A device often to advantage is to insert a question meant primarily to produce
answers to other questions.
• For instance, the relationship of the members of the household to one another
may be used to check the stated age figures.
• It is for this reason that forms often include apparently unnecessary or irrelevant
questions.
• The report of a statistical enquiry should include the layout of the form used.
Method of direct observation
• In the method of direct observation, the enquirer or his assistants get the data
directly from the field of enquiry without having to depend on the co-operation of
informants.
• When data are needed on the height and weight of, say, 200 college students, they
will be approached individually and the height (say in cm.) of each measured with a
tape and the weight (say in kg.) measured with weighing balance.
• If data are needed on the sentence-length of a novel by, say, Bankimchandra, the
enquirer himself will go through the book and note for each sentence the length,
i.e. the number of words contained therein.
• On the other hand, if data are required on the incidence of blindness among a
group of people, one will just observe each member of the group and note
whether he or she is or is not blind.
• The direct method of data collection may, therefore, involve either measurement
or counting or bare observation
Basic Principles of Sample Surveys
• The basic principles of sample survey are (i) validity and (ii) optimization.
• By validity of a sampling procedure we mean that the sample should be so
selected that valid estimators are available and their properties (like
unbiasedness, minimum variance etc.) can be studied objectively. This
principle will be satisfied by a probability sample.
• The principle of optimization takes into account the efficiency of the
sampling procedure and cost of the survey.
• Efficiency of an estimator is measured by the inverse of its sampling
variance.
• The cost is generally measured in terms of expenditure incurred in the
survey. The principle of optimization ensures that a fixed level of efficiency
can be achieved for a minimum cost, or a maximum level of efficiency can
be achieved for a fixed cost.
Bias in Sample Survey
• Bias is defined as the systematic prejudice in one direction.
• It is the name given to the course which influence the choice of sampling units and / or
the observations, and make them something other than what they actually should be.
• Bias may occur due to imperfect instruments, personal angularities of the investigator,
defective sampling techniques, wrong choice of the estimators or some other causes.
• Wherever there exists any scope of applying personal choice or judgement on the part of
the investigator, bias is almost certain to creep in.
• This can be seen by conducting a simple experiment. You can ask your friend to tell you
'at random' one hundred digits including zero and then count the number of even and
odd numbers. If the numbers are really random, then the number of even numbers and
odd numbers should be almost same. But in most cases the numbers are significantly
different which implies the presence of some bias.
• In a survey we can broadly see two types of bias: (i) sampling bias and (ii) non-sampling
bias.
sampling bias and non-sampling bias
• Bias originated from sampling is called sampling bias.
• Procedural bias is the non-sampling bias.
• Procedural bias can be in the form of response bias, where people do not tend to
respond properly; observational bias, where sample chosen are not representative of the
population. Very often either of the two occurs. There can be other types of procedural
biases also, like, non-response bias, where people do not respond at all; and interviewer
bias, where the interviewer collects the information with a biased frame of mind.
• In complete enumeration, only non-sampling bias can occur.
• Non-response bias (non-sampling bias) is the distortion or change in the observation that
can arise, because a large number of units selected for the sample by some random
mechanism do not respond or refuse to respond.
• Response bias or Interviewer bias (non-sampling bias) is the bias that can arise due to
wrong response or over-or under-statement of the fact. The way the question is asked
may also affect the response.
Sampling bias
• Selection bias (a sampling bias) is the systematic tendency to exclude
or include certain type of unit.
• Certain sampling techniques help reduce bias because of how the
sample was selected.
• A sampling technique is biased if it produces results that
systematically vary from the truth about the population.
• There are some other sources of sampling bias like bias Due to faulty
demarcation of sampling units, bias due to wrong choice of the
statistic, bias due to substitution etc.
• One threat to the validity of the conclusion in sample survey is 'bias'.
Advantage of Sampling over Census
• 1. Reduction of cost: Sampling can provide reliable information at less cost than
census. Although the cost of collecting an observation may be higher in sample
survey, the total cost is expected to be smaller since a sample consists of a part of
the total population. If a population is very large, census can be very expensive.
• 2. Greater speed: Using sample, the data can be collected more quickly, hence
the estimates can be published in time. In census, the collection of data may take
so long that the information gathered may be no longer needed or useful.
• 3. Greater scope: A complete enumeration or census requires a large o
administrative set-up and involves many persons in data collection. Some
enquiries may even require highly trained personnel or specialized equipment for
collecting data which, sometimes, may make census unmanageable, whereas in
sample survey we may have better coverage and may not face the problems
stated above due to the smaller scale of the investigation.
• 4. Greater accuracy: Estimates based on sample are often more
accurate than those based on census, because investigators can be
more careful while collecting data for fewer units. In sample, more
attention can be given on data quality through the training of
personnel and monitoring. A complete enumeration or census
requires a large administrative set-up and involves many persons in
data collection. With this huge administrative complexity and the
pressure to produce timely estimates, many types of errors can easily
creep into the census data.
• 5. Measure of precision of the estimates: In complete enumeration
there is no way of gauging the magnitude of errors incorporated in
the estimates. But a sample, drawn on the basis of a properly
designed survey method, permits quantitative assessment of errors
involved in the estimates.
• 6. Greater applicability: In some studies the observations are obtained by
destroying the units. For example, to obtain the average life-hour of
electric bulbs, we have to measure the total hours the bulbs survive until
they burnt out. A cookie must be pulverized in order to determine the fat
content etc. In such cases drawing of sample is essential, since census
destroys the entire population. However, there may be some situations
where a census study, in which observations are coming through a
destructive test, is a must.
• For examining if any egg contains salmonella bacteria, it is necessary to
destroy it to get confirmed. Thus, if we want to know whether all eggs from
a particular poultry farm are free from salmonella, once few eggs are found
to carry the bacteria, we cannot rely on sampling, since by sample study it
is impossible to answer with absolute confidence, which is necessary to
avoid health hazard, we have to consider a census study and destroy all
eggs so that they cannot produce any further chicken carrying salmonella
bacteria.
• While studying an infinite or hypothetical population sampling is inevitable.
• However, when it is essential to gather information from all the units
of a population irrespective of cost or time, a census is required, for
example, population census, where relevant information are collected
from all the units of the population.
• There are many sampling methods available. We mention a few
commonly used simple sampling schemes. The choice between these
sampling methods depends on (1) the nature of the problem or
investigation, (2) the availability of good sampling frames (a list of all
of the population members), (3) the budget or available financial
resources, (4) the desired level of accuracy, and (5) the method by
which data will be collected, such as questionnaires or interviews.
Errors in Sample Data
• Irrespective of which sampling scheme is used, the sample
observations are prone to various sources of error that may seriously
affect the inferences about the population.
• Some sources of error can be controlled. However, others may be
unavoidable because they are inherent in the nature of the sampling
process
• The errors can be classified as sampling errors and non sampling
errors.
Sampling errors and Non sampling errors
• Sampling errors occur because the sample is not an exact
representative of the population.
• Sampling error is due to the differences between the characteristics
of the population and those of a sample from the population.
• For example, we are interested in the average test score in a large
statistics class of size, say, 80. A sample of size 10 grades from this
resulted in an average test score of 75.
• If the average test for the entire 80 students (the population) is 72,
then the sampling error is 75 − 72 = 3.
Non sampling errors
• Non sampling errors occur in the collection, recording and processing
of sample data.
• For example, such errors could occur as a result of bias in selection of
elements of the sample, poorly designed survey questions,
measurement and recording errors, incorrect responses, or no
responses from individuals selected from the population.
Types of Sampling
• Random Sampling
• Non- Random Sampling
Non-probability or Non- Random Sampling
• Non-probability sampling procedure is a sampling procedure which does
not attach any definite probabilistic basis for selecting items in the sample
from the population.
• These samples produce biased results that systematically differ from the
truth about the population.
• For example, suppose, a truck-load of potatoes has arrived in, and because
of convenience, a sample of few baskets of potatoes are taken from the top
of the truckload to examine the quality of the potatoes. Clearly, the
potatoes taken from the top may not be representative of the entire
truckload, since we do not have any potatoes from the bottom or middle of
the truckload in the sample which may have been damaged in the
shipment.
Purposive sampling
• This is based on the intention or the purpose of study. This method of
sampling is also known as subjective or judgment sampling method.
• Accordingly, investigator himself purposively chooses certain items which
to his judgment are best representatives of the universe. Here the
selection is deliberate and based on own idea of the investigator about the
sample units.
• As such under this method, the chance of inclusion of some items in the
sample very high while that of others is very low.
• However, for better selection of items under this method certain criteria of
selection is first laid down and then the investigator is allowed to make the
selection of the items of his own accord within the orbit of those criteria.
Purposive sampling
• Here the sample units are selected with definite purpose in view.
• Participants are selected based on specific characteristics, such as knowledge,
experiences, or other criteria that are relevant to the study. This method is often
used in qualitative studies when researchers want to gain in-depth insights from
specific groups or individuals
• For example, if we want to get opinion on an issue of IR, then best we ask career
diplomats and academics specializing in international relations
• if we want to give the picture that the standard of living has increased in the city
of New Delhi, we may take individuals in the sample from rich and posh localities
like Defense Colony, South Extension, Golf Links, Jor Bagh, Chanakyapuri, Greater
Kailash etc. and ignore the localities where low income group and the middle
class families live.
• The second ex. sampling suffers from the drawback of favoritism and nepotism
and does not give a representative sample of the population.
• Results obtained from the analysis of purposively selected sample
may be very useful, but data so obtained are not amenable to any
formal statistical techniques like unbiased estimation of population
parameters, standard error of estimator etc. are not possible to
obtain.
• Hence one of the disadvantages in such sampling is that the
inferences generated using a non-probability sample cannot be
generalized for larger population. Such samples certainly do not
possess the characteristics of random sample and criticized of having
selection bias.
Convenience sampling
• Convenience sampling is a non-probability sampling scheme where
selection of the sample units is made in accordance with the convenience
of the researcher.
• For example, in some study related to the trend of bank loans taken by the
farmers, neighborhood banks are selected, because they are nearer to the
location where the researcher is located.
• This method is often used when researchers need to quickly gather data,
and when practicality is more important than specific characteristics
• Purposive sampling aims to identify the best people to answer the research
question, while convenience sampling can help generate a hypothesis or
research question.
Quota sampling
• Quota sampling is also a non-probability sampling procedure, where the
researchers are simply given quotas to be filled from different strata for
information.
• a researcher selects a predetermined number or proportion of participants from
a population based on mutually exclusive criteria. The process involves dividing
the population into subgroups, or strata, and then recruiting participants until the
quota is reached
• Here a specified number of different types of units are selected and information
are collected. It is to be noted that if the sampling frames for different strata are
not available, then quota sampling is used by fixing a sample quota for each
stratum.
• For example, a researcher may want to interview ten workers and decides that
five must be male and five must be female. In quota sampling, an advantage of
stratification is achieved without having any element of probability sampling.
• Suppose you want to gauge consumer interest
in a new meal kit delivery service in north west
Delhi.
• Depending on your research goals, you can
divide your population into several strata, such
as:
• Dietary preferences
• Age group
• Area code
• Let’s say you want to focus on dietary
preferences. You divide the population into
meat eaters, vegetarians, and vegans, drawing
a sample of 600 people.
• Since the company wants to cater to all
consumers, you set a quota of 200 people for
each dietary group. In this way, all dietary
preferences are equally represented in your
research, and you can easily compare these
groups.
Snowball sampling
• Snowball sampling is another type of non-probability sampling, which is used
when the researchers do not know from whom the information are to be
gathered.
• Snowball sampling is used in interviews, where the researchers would ask the
interviewees about the 'interest of study' and the names of other people who can
also provide sufficient information on the same. The process of interview will
continue in this manner until no new names are suggested by the interviewees.
• It is also known as Chain sampling, referral sampling, cold calling etc.
• It's a type of non-probability sampling where researcher himself does not selects
all elements to be included in a sample. Here the sample group is said to grow
like a rolling snowball. It is basically useful where either there is no information
about the subjects or elements of the study
Random sampling
• One way in which a representative sample may be obtained is by a
process called random sampling, according to which each member of
a population has an equal chance of being included in the sample.
• One technique for obtaining a random sample is to assign numbers to
each member of the population, write these numbers on small pieces
of paper, place them in an urn, and then draw numbers from the urn,
being careful to mix thoroughly before each drawing.
• An alternative method is to use a (Tippet’s) table of random numbers
specially constructed for such purposes.
Systematic sampling
• A systematic sample is a sample in which every Kth element in the
sampling frame is selected after a suitable random start for the first
element. We list the population elements in some order (say
alphabetical) and choose the desired sampling fraction
• STEPS FOR SELECTING A SYSTEMATIC SAMPLE
• 1. Number the elements of the population from 1 to N.
• 2. Decide on the sample size, say n, that we need.
• 3. Choose K = N/n.
• 4. Randomly select an integer between 1 to K.
• 5. Then take every Kth element.
• If the population has 1000 elements arranged in some order and we
decide to sample 10% (i.e., N=1000 and n=100), then
K=1000/100=10. Pick a number at random between 1 and K=10
inclusive, say 3. Then select elements numbered 3, 13, 23, ... , 993.
• Systematic sampling is widely used because it is easy to implement. If
the list of population elements is in random order to begin with, then
the method is similar to simple random sampling.
• If, however, there is a correlation or association between successive
elements, or if there is some periodic structure, then this sampling
method may introduce biases. Systematic sampling is often used to
select a specified number of records from a computer file.
• In simple random sampling, each data point has an equal probability
of being chosen. Meanwhile, systematic sampling chooses a data
point per each predetermined interval.
Stratified sampling
• A stratified sample is a modification of simple random sampling and
systematic sampling and is designed to obtain a more representative
sample, but at the cost of a more complicated procedure.
• Compared to random sampling, stratified sampling reduces sampling error.
A sample obtained by stratifying (dividing into non overlapping groups) the
sampling frame based on some factor or factors and then selecting some
elements from each of the strata is called a stratified sample.
• Here, a population with N elements is divided into s subpopulations. A
sample is drawn from each subpopulation independently. The size of each
subpopulation and sample sizes in each subpopulation may vary
• This method is useful when the population can be divided into subgroups
that are expected to have different mean values for the variable being
studied. Stratified sampling can improve the accuracy of results by ensuring
that all subgroups are represented proportionally.
STEPS FOR SELECTING A STRATIFIED SAMPLE
• 1. Decide on the relevant stratification factors (sex, age, income, etc.).
• 2. Divide the entire population into strata (subpopulations) based on the
stratification criteria. Sizes of strata may vary.
• 3. Select the requisite number of units using simple random sampling or
systematic sampling from each subpopulation. The requisite number may
depend on the subpopulation sizes.
• Examples of strata might be males and females, undergraduate students
and graduate students, managers and nonmanagers, or populations of
clients in different racial groups such as African Americans, Asians, whites,
and Hispanics. Stratified sampling is often used when one or more of the
strata in the population have a low incidence relative to the other strata.
In a population of 1000 children from an area school, there are
600 boys and 400 girls. We divide them into strata based on
their parents’ income as shown in Table
Classification of School Children Boys Girls
Poor 120 240
Middle class 150 100
Rich 330 60
USES OF STRATIFIED SAMPLING
In addition to providing information about the whole population, this
sampling scheme provides information about the subpopulations, the
study of which may be of interest.
For example, in a U.S. presidential election, opinion polls by state may
be more important in deciding on the electoral college advantage than
a national opinion poll.
Stratified sampling can be considerably more precise than a simple
random sample, because the population is fairly homogeneous within
each stratum but there is a sizable variation between the strata.
Cluster sampling
• In cluster sampling, the sampling unit contains groups of elements called
clusters instead of individual elements of the population.
• A cluster is an intact group naturally available in the field. Unlike the
stratified sample where the strata are created by the researcher based on
stratification variables, the clusters naturally exist and are not formed by
the researcher for data collection.
• Cluster sampling is also called area sampling.
• To obtain a cluster sample, first take a simple random sample of groups and
then sample all elements within the selected clusters (groups).
• Cluster sampling is convenient to implement. However, because it is likely
that units in a cluster will be relatively homogeneous, this method may be
less precise than simple random sampling.
• Suppose we wish to select a sample of about 10% from all fifth-grade
children of a county. We randomly select 10% of the elementary schools
assumed to have approximately the same number of fifth-grade students
and select all fifth-grade children from these schools.
• This is an example of cluster sampling, each cluster being an elementary
school that was selected.
• The clusters should ideally each be mini-representations of the population
as a whole.
• Cluster sampling is best for large, spread-out populations and all units of
each selected group are included in the sample
• Cluster sampling can reduce travel and logistical expenses and improve
cost-effectiveness and operational efficiency.
• Cluster sampling can be prone to biased samples and higher sampling
error.
Sampling With And Without Replacement
• If we draw a number from an urn, we have the choice of replacing or not replacing the number
into the urn before a second drawing.
• In the first case the number can come up again and again, whereas in the second it can only
come up once.
• Sampling where each member of the population may be chosen more than once is called
sampling with replacement, while if each member cannot be chosen more than once it is called
sampling without replacement.
• Populations are either finite or infinite.
• If, for example, we draw 10 balls successively without replacement from an urn containing 100
balls, we are sampling from a finite population; while if we toss a coin 50 times and count the
number of heads, we are sampling from an infinite population.
• A finite population in which sampling is with replacement can theoretically be considered
infinite, since any number of samples can be drawn without exhausting the population.
• For many practical purposes, sampling from a finite population that is very large can be
considered to be sampling from an infinite population.