Research Design:
Sampling Design:
Sampling Strategies:
Sample Size Determination:
Confidence Interval
𝑚𝑒𝑎𝑛 ± 𝑡𝛼/2 ∗ 𝑠𝑒(𝑚𝑒𝑎𝑛)
𝑚𝑒𝑎𝑛 ± 𝑡𝛼/2 ∗ √𝑣𝑎𝑟𝑖𝑎𝑛𝑐𝑒 𝑜𝑓 𝑋/𝑛
𝑚𝑒𝑎𝑛 ± 𝑡𝛼/2 ∗ 𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣 𝑜𝑓 𝑋/√𝑛
𝑀𝑎𝑟𝑔𝑖𝑛 𝑜𝑓 𝐸𝑟𝑟𝑜𝑟 = 𝑡𝛼/2 ∗ 𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣 𝑜𝑓 𝑋/√𝑛
INTERVAL ESTIMATE OF A POPULATION MEAN: σ UNKNOWN
𝑥̅ ± 𝑡𝛼/2 ∗ 𝑆/√𝑛
Margin of Error:
𝐸 = 𝑡𝛼/2 ∗ 𝑆/√𝑛
Derive the equation for sample size
2
√𝑛 = (𝑡𝛼 ) ∗ 𝑆/𝐸
2
2
(𝑡𝛼 ) ∗ 𝑆 2
2
𝑛=
𝐸2
Proportion:
𝑝̅ (1 − 𝑝̅)
𝑝̅ ± 𝑡𝛼/2 ∗ √
𝑛
Margin of Error
𝑝̅ (1 − 𝑝̅ )
𝐸 = 𝑡𝛼/2 ∗ √
𝑛
2
(𝑡𝛼 ) 𝑝̅(1 − 𝑝̅)
2
𝑛=
𝐸2
Sampling Efficiency:
stratification,
cluster sampling and
sampling in stages: Multi-stage Sampling
Stratified Clustered Multi-stage sampling strategy:
∑(𝑌𝑖 −𝑌̅)2
Variance of 𝑌 = 𝑛−1
𝑣𝑎𝑟𝑖𝑎𝑛𝑐𝑒 𝑜𝑓 𝑌 235.1909
Variance of 𝑌̅= = = .0866584
𝑛 2714
∑(𝑌𝑖 −𝑌̅)2
= (𝑛−1)
Standard Error of 𝑌̅ =square root (Variance of 𝑌̅)= 𝑠𝑞𝑟𝑡(. 0866584) = .294378
Sampling
1. Sampling is a technique by which a part of the population is selected and results from
this fraction are generalised on the whole population from which the part or sample has been
selected.
A.1 Sample design
2.
In general, sampling theory is concerned with how, for a given population, the estimates from the
survey and the sampling errors associated with them are related to the sample size and structure.
In practice sample design involves
• the determination of sample size,
• structure and
• takes into account costs of the survey.
Procedures of selection, implementation and estimation
− Each element in the population should be represented in the frame from which the sample is
to be selected.
− The selection of the sample should be based on a random process which gives each unit a
specified probability of selection.
− All and only selected units must be enumerated.
− In estimating population parameters from the sample, the data from each unit/element must be
weighted in accordance with its probability of selection.
Significance of probability sampling to large-scale household surveys
− It permits coverage of the whole target population in sample selection.
− It reduces sampling bias.
− It permits generalization of sample results to the population from which the sample is
selected.
− It has been argued that it allows the surveyor to present results without having to apologise
for using non-scientific methods (Kish, 1965).
− It allows the calculation of sampling errors, which are reliability measures.
Basic requirements for designing a probability sample
− The target population must be clearly defined.
− There must be a sampling frames or frames in case of multi-stage samples.
− The objectives of the survey must be unambiguously specified in terms of:
a. Survey content
b. Analytical variables
c. Level of dissagregation (e.g. do you need estimates or data at national, rural, urban,
provincial, district, etc. levels?).
− Budget and field constraints should be taken into account.
− Precision requirements must be spelled out in order to determine the sample size.
Random selection of units reduces the chance of getting a non-representative sample.
Randomisation is a safe way to overcome the effects of unforeseen biasing factors.
A.2 Basics of probability sampling strategies
A.2.1 Simple random sampling
17.
Simple random sampling (SRS) is a probability sample selection method where each element of
the population has an equal chance/probability of selection.
Selection of the sample can be with or without replacement.
This method is rarely used in large-scale household surveys because it is costly in terms of listing
and travel. It can be regarded as the basic form of the population structure. SRS is attractive for
being simple in terms of selection and estimation procedures (e.g. sampling errors).
18. While SRS is not very much used in practice, it is basic to sampling theory mainly
because of its simple mathematical properties. Most statistical theories and techniques, therefore,
assume simple random selection of elements. Indeed all other probability sample selections may
be seen as restrictions on SRS, which suppress some combinations of population elements. SRS
serves two functions:
− Sets a baseline for comparing the relative efficiency of other sampling techniques.
− It can be used as the final method for selecting the elementary units, in the context of the
more complex designs such as clustering and stratified sampling.
a. Simple random sampling with replacement (SRSWR).
b. Simple random sampling without replacement (SRSWOR).
Stratification
47. Stratified sampling is a method in which the sampling units in the population are divided
into groups called strata.
Stratification is usually done in such a way that the population is subdivided into heterogeneous
groups which are internally homogeneous. In general, when sampling units are homogeneous
with respect to the auxiliary variable termed stratification variable, the variability of strata
estimators is usually reduced. Further there is considerable flexibility in stratification in the sense
that the sampling and estimation procedures can be rightly different from stratum to stratum.
48. In stratified sampling, therefore, we group together units/elements which are more or less
similar, so that the variance δ h 2 within each stratum is small, at the same time it is essential
that the means ( ) h x of the different strata are as different as possible. An appropriate estimate
for the population as a whole is obtained by suitably combining stratum-wise estimators of the
characteristic under consideration.
A.2.3.1. Advantages of stratified sampling
49. The main advantage of stratified sampling is the possible increase in the precision of
estimates and the possibility of using different sampling procedures in different strata.
− In case of skewed populations since larger sampling fractions may be necessary for selecting
from the few large units, resulting in giving greater weight to few extremely large units for
reducing the sampling variability.
− When a survey organization has several field offices in various regions into which the
country has been divided for administrative purposes it may be useful to treat the regions as
strata, so as to facilitate the organization of fieldwork.
− When estimates are required within specific margins of error, not only for the whole
population, but also for certain sub-groups such as provinces, rural or urban, gender, etc.
Through stratification such estimates can conveniently be provided.
51. Summary of steps followed in stratified sampling:
− The entire population of sampling units is divided into internally homogeneous but externally
heterogeneous sub-populations.
− Within each stratum, a separate sample is selected from all sampling units in the stratum.
− From the sample obtained in each stratum, a separate stratum mean (or any other statistic) is
computed. The strum means are then properly weighted to form a combined estimate for the
population.
− Usually proportionate sampling within strata is used when overall, e.g. national estimates are
the objective of the survey and the survey is multipurpose.
− Disproportionate sampling is used when sub-group domains have priority, in cases where
estimates for sub-national areas are wanted with equal reliabilities.
A.2.4 Cluster sampling
59. The discussions in the previous sections have so far been about sampling methods in
which elementary sampling units were considered as arranged in a list from a frame in such a
way that individual units could be selected directly from a frame.
60.
In Cluster Sampling, the higher units e.g. enumeration areas (see chapter 2) of selection contain
more than one elementary unit. In this case, the sampling unit is the cluster.
61. For example, to select a random sample of households in a city a simple method is to have a
list of all households. This may not be possible as in practice there may be no complete frame of
all households in the city. In order to go round this problem, clusters in the form of blocks Ward)
could be formed. Then a sample of blocks (e.g. 5 Ward) could be selected, subsequently a list of
households in the selected blocks made. If need be, in each block a sample of households say
10% could be drawn.
A.2.4.1. Some reasons for using cluster sampling
a. Clustering reduces travel and other costs of data collection.
b. It can improve supervision, control, follow-up coverage and other aspects that have an impact
on the quality of data being collected.
c. The construction of the frame is made cheaper as it is done in stages.
For instance, in multi-stage sampling discussed in chapter 2 a frame covering the entire
population is required only for selecting PSUs i.e. clusters at the first stage. At any lower stage, a
frame is required only within the units selected at the preceding stage. In addition, frames of
larger and higher stage units tend to be more durable and therefore usable over longer period of
time. Lists of small units such as households and particularly of people tend to become obsolete
within a short period of time.
d. There is administrative convenience in the implementation of the survey.