Household Survey Management Guide
Household Survey Management Guide
LECTURE 1: INTRODUCTION
1.1. Introduction
Household surveys are an important source of socio-economic data for households
and individuals in developing and transition countries. Important indicators to inform
and monitor development policies are often derived from such surveys. In developing
countries, they have become a dominant form of data collection, supplementing, or
sometimes even replacing other data collection programmes and civil registration
systems.
In addition to national surveys funded out of regular national budgets, there are a large
number of household surveys being conducted in developing and transition countries
that are sponsored by international agencies, for the purposes of constructing and
monitoring national estimates of characteristics or indicators of interest to the
agencies, and also for making international comparisons of these indicators. Most
such surveys are conducted on an ad hoc basis, but there is renewed interest in the
establishment of ongoing multi-subject, multi-round integrated programmes of
surveys, with technical assistance from international organizations, such as the United
Nations and the World Bank, in all stages of survey design, implementation, analysis
and dissemination. Prominent examples of household surveys conducted by
international agencies in developing countries are the Demographic and Health
Surveys (DHS), carried out by ORC Macro for the United States Agency for
International Development (USAID); the Living Standards Measurement Study
(LSMS) surveys, conducted with technical assistance from the World Bank, and the
Multiple Indicator Cluster Surveys (MICS) conducted by the United Nations Children’s
Fund (UNICEF).
Household surveys are among three major sources of social and demographic
statistics in many countries. It is recognized that population and housing censuses are
also a key source of social statistics, but they are conducted, usually, at long intervals
of about ten years. The third source is administrative record systems. Household
surveys provide a cheaper alternative to censuses for timely data and a more relevant
and convenient alternative to administrative record systems. They are used for
collection of detailed and varied socio-demographic data pertaining to conditions
under which people live, their well-being, activities in which they engage, demographic
characteristics and cultural factors which influence behaviour, as well as social and
economic change.
1
Despite the listed advantages of the household surveys, data generated through
household surveys are used in complementary manner with data from other sources
such as censuses and administrative records. Many countries have in place household
survey programmes that include both periodic and adhoc surveys. It is advisable that
the household survey programme be part of an integrated statistical data collection
system of a country.
2
STAT 444: SURVEY ORGANISATION AND MANAGEMENT
LECTURE 2: PLANNING AND EXECUTION OF SURVEYS
3
Source: United Nations (2005). Designing Household Survey Samples: Practical
Guidelines. New York (Chapter 2)
4
survey subject matter should be clearly understood by those involved in designing the
survey operations. Further, interviewers must thoroughly master the practical
procedures that may lead to the successful collection of accurate data.
5
STAT 444: SURVEY ORGANISATION AND MANAGEMENT
LECTURE 3: SAMPLING STRATEGIES
3.1. INTRODUCTION
Almost all sample designs for household surveys, both in developing and developed
countries, are complex because of their multi-stage, stratified and clustered features.
In addition, national-level household sample surveys are often general-purpose in
scope, covering multiple topics of interest to the government and this also adds to their
complexity. Most of the surveys are based on multistage stratified area probability
sample designs. These designs are used primarily for frame development and for
clustering interviews to reduce cost.
The units selected at the first stage (i.e., Primary Sampling Units (PSUs)), are
frequently constructed from enumeration areas identified and used in a preceding
national population and housing census. The units selected within each selected PSU
are referred to as second-stage units, units selected at the third stage are referred to
as the third-stage units, and so on. For households in developing and transition
countries, second-stage units are typically dwelling units or households, and units
selected at the third stage are usually persons. In general, the units selected at the
last stage in a multistage design are referred to as the ultimate sampling units.
Despite the many similarities discussed above, sample designs for surveys in
developing and transition countries are not identical across countries, and may vary
with respect to, for example, the target populations, content and objectives, the
number of design strata, sampling rates within strata, sample sizes within PSUs, and
the number of PSUs selected within strata. In addition, the underlying populations may
vary with respect to their prevalence rates for specified population characteristics, the
degree of heterogeneity within and across strata, and the distribution of specific
subpopulations within and across strata.
6
Multiple Indicator Cluster Survey 2017/18 (Ghana MICS 2017/18), the urban and rural
areas within each region were identified as the main sampling strata and a two-stage
sample design was used for the selection of households. The primary sampling units
(PSUs) selected at the first stage were the enumeration areas (EAs) defined for the
2010 PHC enumeration. A listing of households was conducted in each sample EA,
and a sample of households was selected at the second stage. (Read on the
methodology for Ghana MICS 2017/18).
Stratification partitions the units in the population into mutually exclusive and
collectively exhaustive subgroups or strata. Separate samples are then selected from
each stratum. A primary purpose of stratification is to improve the precision of the
survey estimates. In this case, the formation of the strata should be such that units in
the same stratum are as homogeneous as possible and units in different strata are as
heterogeneous as possible with respect to the characteristics of interest to the survey.
Other benefits of stratification include (i) administrative convenience and flexibility and
(ii) guaranteed representation of important domains and special subpopulations.
Implicit stratification variables sometimes used for PSU selection include residential
area (low- income, moderate-income, high-income), expenditure category (usually in
quintiles), ethnic group and area of residence in urban areas; and area under
cultivation, number of poultry or cattle owned, proportion of non-agricultural workers,
etc., in rural areas. For socio-economic surveys, implicit stratification variables include
the proportion of households classified as poor, the proportion of adults with secondary
or higher education, and distance from the centre of a large city. Variables used for
implicit stratification are usually obtained from census data.
7
3.2.3. Sample selection of PSUs
For household surveys in developing and transition countries, PSUs are often small
geographical area units within the strata. If census information is available, PSUs may
be the enumeration areas identified and used in the census. Since the PSUs affect the
quality of all subsequent phases of the survey process, it is important to ensure that
the units designated as PSUs are of good quality and that they are selected for the
survey in a reasonably efficient manner. For PSUs to be considered of good quality,
they must, in general:
i. Have clearly identifiable boundaries that are stable over time;
ii. Cover the target population completely;
iii. Have a measure of size for sampling purposes;
iv. Have data for stratification purposes;
v. Be large in number.
(Student to do further reading Sample selection of PSUs in the reading text)
PPS sampling yields unequal probabilities of selection for PSUs. Essentially, the
measure of size of the PSU determines its probability of selection. However, when
combined with an appropriate subsampling fraction for selecting households within
selected PSUs, it can lead to an overall self-weighting sample of households in which
all households have the same probability of selection regardless of the PSUs in which
they are located. Its principal attraction is that it can lead to approximately equal
sample sizes per PSU. For household surveys, a good example of a PPS size variable
for the selection of PSUs is the number of households. For farm surveys, a PPS size
measure that is frequently used is the size of the farm. For business surveys, typical
PPS measures of size include the number of employees, number of establishments
and annual volume of sales. Probability proportional to size (PPS) measures of size
are likely to change over time, and this fact must be taken into consideration in the
sample design process.
8
Consider a sample of households, obtained from a two-stage design, with 𝒂 PSUs
selected at the first stage and a sample of households at the second stage. Let the
measure of size (for example, the number of households at the time of the last census)
of the 𝒊𝒕𝒉 PSU be 𝑴𝒊 . If the PSUs are selected with PPS, then the probability 𝑃𝑖 of
selecting the 𝒊𝒕𝒉 PSU is given by:
𝑀𝐼
𝑃𝑖 = 𝑎 ×
∑𝑖 𝑀𝑖
Now, let 𝑷𝒋|𝒊 denote the conditional probability of selecting the 𝑗𝑡ℎ household in the 𝑖𝑡ℎ
PSU, given that the 𝑖𝑡ℎ PSU was selected at the first stage. Then, the selection
equation for the unconditional probability 𝑃𝑖𝑗 of selecting the 𝑗𝑡ℎ household in the 𝑖𝑡ℎ
PSU under this design is:
𝑃𝑖𝑗 = 𝑃𝑖 × 𝑃𝑗|𝑖
𝑓
𝑃𝑗|𝑖 =
𝑃𝑖
If the measures of size of the PSUs are the true sizes, and there is no change in the
measure of size between sample selection and data collection, and if b households
are selected in each sampled PSU, then we obtain a self-weighting sample of
households with a probability of selection given by
𝑀𝐼 𝑏 𝑎𝑏
𝑃𝑖𝑗 = 𝑎 × ∑ × =∑ =𝑓
𝑖 𝑀𝑖 𝑀𝑖 𝑖 𝑀𝑖
where f is a constant.
The problem with this procedure is that the true measures of size are rarely known in
practice. However, it is often possible to obtain good estimates, such as population
and household counts from a recent census, or some other reliable source. This allows
us to apply the procedure known as probability-proportional-to-estimated-size (PPES)
sampling. There are two choices for PPES sampling in a two-stage design with
9
households selected at the second stage: either (a) select households at a fixed rate
in each sampled PSU; or (b) select a fixed number of households per sampled PSU.
PPES sampling of households at a fixed rate is implemented as follows. Let the true
values of the measure of size be denoted by Ni, and assume that the values Mi are
good estimates of Ni. We then apply the sampling rate b/Mi to the ith PSU to obtain a
sample size of
𝑏
𝑏𝑖 = 𝑀 × 𝑁𝑖
𝑖
Note that subsampling within PSUs at a fixed rate (inversely proportional to the
measures of size of the PSUs) involves the determination of a rate for each sampled
PSU so that, together with the PSU selection probability, we obtain an equal-
probability sample of households, regardless of the actual size of the PSUs. However,
this procedure does not provide control over the subsample sizes, and hence the
overall sample size. More households will be sampled from PSUs with larger-than-
expected numbers of households, and fewer households will be sampled from PSUs
with smaller-than-expected numbers of households. This has implications for the
fieldwork organization. In addition, if the measures of size are so out of date that the
variation in the realized samples is extreme, there may be a need for a change in the
sampling rate so as to obtain sample sizes that are a bit more homogeneous across
PSUs, which would entail some degree of departure from a self-weighting design.
The second procedure, selecting a fixed number of households per PSU, avoids the
disadvantage of variable sample sizes per PSU but does not produce a self-weighting
sample. However, if the measures of size are updated immediately prior to sample
selection of PSUs, they may provide good enough approximations that will lead to an
approximately self-weighting sample of households.
10
Sometimes the listings are of dwelling units and then all households in selected
dwelling units are included if a dwelling unit is sampled. The objective of this listing
step is to create an up-to-date sampling frame from which households can be selected.
Prior to sample selection in each sampled PSU, the listed households may be sorted
with respect to geography and other variables deemed strongly correlated with the
survey variables of interest. Then, households are sampled from the ordered list by an
equal probability systematic sampling procedure. The households may be selected
within sampled PSUs at sampling rates that generate equal overall probabilities of
selection for all households or at rates that generate a fixed number of sampled
households in each PSU. Frequently, the ultimate sampling units are households and
information is collected on the selected households and all members of those
households. For special modules covering incomes and expenditures, for which
households are the units of analysis, a knowledgeable respondent is often selected to
be the household informant. For subjects considered sensitive for persons within
households (for example, domestic abuse), a random sample of persons (frequently
of one person) is selected within each sampled household.
Clustering reduces the cost of data collection considerably, but correlations among
units in the same cluster inflate the variance (lower the precision) of survey estimates,
compared with a design in which households are not clustered. Thus, the challenge
for the survey designer is to achieve the right balance between the cost savings and
the corresponding loss in precision associated with clustering. The inflation in variance
of survey estimates attributable to clustering contributes to the so-called design effect.
The design effect represents the factor by which the variance of an estimate based on
a simple random sample of the same size must be multiplied to take account of the
complexities of the actual sample design due to stratification, clustering and weighting.
It is defined as the ratio of the variance of an estimate based on the complex design
relative to that based on a simple random sample of the same size.
An expression for the design effect (due to clustering) for an estimate [for example, an
estimated mean (𝑦̅)] is given approximately by:
𝐷2 (𝑦̅) = 1 + (𝑏 − 1)𝜌
where 𝐷2 (𝑦̅) denotes the design effect for the estimated mean (𝑦̅), 𝜌 is the intra-class
correlation, and 𝑏 is the average number of households to be selected from each
11
cluster, that is to say, the average cluster sample size. The intra-class correlation is a
measure of the degree of homogeneity (with respect to the variable of interest) of the
units within a cluster. Since units in the same cluster tend to be similar to one another,
the intra-class correlation is almost always positive. For human populations, a positive
intra-class correlation may be due to the fact that households in the same cluster
belong to the same income class; may share the same attitudes towards the issues of
the day; and are often exposed to the same environmental conditions (climate,
infectious diseases, natural disaster, etc.).
Failure to take account of the design effect in the estimates of standard errors can lead
to invalid interpretation of the survey results. It should be noted that the magnitude of
𝐷2 (𝑦̅) is directly related to the value of b, the cluster sample size, and the intra-class
correlation (𝜌). For a fixed value of 𝜌, the design effect increases linearly with 𝑏. Thus,
to achieve low design effects, it is desirable to use as small a cluster sample size as
possible.
In general, the optimum number of households to be selected in each PSU will depend
on the data-collection cost structure and the degree of homogeneity or clustering with
respect to the survey variables within the PSU. Assume a two-stage design with PSUs
selected at the first stage and households selected at the second stage. Also, assume
a linear cost model for the overall cost related to the sampling of PSUs and households
given by:
𝐶 = 𝑎𝐶1 + 𝑎𝑏𝐶2
where 𝐶1 and 𝐶2 are, respectively, the cost of an additional PSU and the cost of an
additional household; and 𝑎 and 𝑏 denote, respectively, the number of selected PSUs
and the number of households selected per PSU (Cochran, 1977, p. 280). Under this
cost model, the optimum choice for 𝑏 that minimizes the variance of the sample mean
(see Kish, 1965, sect. 8.3.b) is approximately given by:
𝐶1 (1−𝜌)
𝑏𝑜𝑝𝑡 = √
𝐶2 𝜌
Table I gives the optimal subsample size (b) for various cost ratios C1/C2 and intraclass
correlation. Note that all other things being equal, the optimal sample size decreases
(that is to say, the sample is more broadly spread across clusters) as the intra-class
correlation increases and as the cost of an additional household increases relative to
that of a PSU.
12
Table I: Optimal subsample sizes for selected combinations of cost ratio and
intra-class
Correlation
The cost model used in the derivation of the optimal cluster size is an oversimplified
one but is probably adequate for general guidance. Since most surveys are multi-
purpose in nature, involving different variables and correspondingly different values of
𝜌, the choice of 𝑏 often involves a degree of compromise among several different
optima.
In the absence of precise cost information, table I can be used to determine the optimal
number households to be selected in a cluster for various choices of cost ratio and
intra-class correlation. For instance, if it is known a priori that the cost of including a
PSU is four times as great as that of including a household, and that the inter-class
correlation for a variable of interest is 0.05, then it is advisable to select about nine
households in the cluster. Note that the optimum number of households to be selected
in a cluster does not depend on the overall budget available for the survey. The total
budget determines only the number of PSUs to be selected.
In general, the factors that need to be considered in determining the sample allocation
across PSUs and households within PSUs include the precision of the survey
estimates (through the design effect), the cost of data collection and the fieldwork
organization. If travel costs are high, as is the case in rural areas, it is preferable to
select a few PSUs and many households in each PSU. On the other hand, if, as in
urban areas, travel costs are lower, then it is more efficient to select many PSUs and,
then, fewer households within each PSU. On the other hand, in rural areas, it may be
more efficient to select more households per PSU. These choices must be made in
such a way as to produce an efficient distribution of workload among the interviewers
and supervisors.
13
The target population for most household surveys comprises the civilian non-
institutionalized population. To obtain the desired data from this target population,
interviews are often conducted at the household level. In general, only persons
considered permanent residents of the household are eligible for inclusion in the
surveys. Permanent residents of a household who are away temporarily, such as
persons on vacation, or temporarily in a hospital, and students living away from home
during the school year, are generally included if their household is selected. Students
living away from home during the school year are not included in the survey if sampled
at their school-time residence because data for such students would be obtained from
their permanent place of residence. Groups that are generally excluded from
household surveys in developing and transition countries include members of the
armed forces living in barracks or in private homes; persons in prisons, hospitals,
nursing homes or other institutions; homeless people; and nomads. Most of these
groups are generally excluded because of the practical difficulties usually encountered
in collecting data from them. It should be noted that the decision on whether or not to
exclude a group needs to be made taking into consideration the survey objectives.
(a) Non-Coverage
The term non-coverage. refers to the failure of the sampling frame to cover all the
target population, as a result of which some sampling units have no probability of
inclusion in the sample. Non-coverage is a major concern for household surveys
conducted in developing and transition countries. Evidence of the impact of non-
coverage can be seen from the fact that sample estimates of population counts based
on most surveys in developing and transition countries fall well short of population
estimates from other sources.
There are three levels of non-coverage: the PSU level, the household level and the
person level. For developing and transition countries, non-coverage of PSUs is a less
serious problem than non-coverage of households and of eligible persons within
sampled households. Non-coverage of PSUs occurs, for example, when some regions
of a country are excluded from a survey on purpose, because they are inaccessible,
owing to war, natural disaster or other causes.
Also, remote areas with very few households or persons are sometimes removed from
the sampling frames for household surveys because they represent a small proportion
of the population and so have very little effect on the population figures. Non-coverage
14
is a more serious problem at the household and person levels. Households or persons
may be erroneously excluded from the survey as the result of the complex definitional
and conceptual issues regarding household structure and composition. There is
potential for inconsistent interpretation of these issues by different interviewers or
those responsible for creating lists of households and household members. Therefore,
strict operational instructions are needed to guide interviewers on who is to be
considered a household member and on what is to be considered a household or a
dwelling unit. As a means of addressing this problem, the quality of the listing of
households and eligible persons within households should be made a key area for
methodological work and training in developing and transition countries.
(c) Blanks
The problem of blanks arises when some listings on the sampling frame contain no
elements of the target population. For a list frame of dwelling units, a blank would
correspond to an empty dwelling. This problem also arises in instances where one is
sampling particular subgroups of the population, for instance, women who had given
birth last year. Some households that were listed and sampled will not contain any
women who gave birth last year. If possible, blanks can be removed from the frame
before sample selection. However, this is not cost-effective in many practical
applications. A more practical solution is to identify and eliminate blanks after sample
selection. However, eliminating blanks means that the realized sample will be smaller
and of variable size.
15
uniquely identified with their first visit to a watering hole after a given date, with later
visits being treated as blanks. Otherwise, the weights of the sampled units need to be
adjusted to account for the duplicates.
It should also include data that may be useful for stratification, such as ethnic and
racial composition, median expenditure or expenditure quintiles, etc. If properly
maintained, the master sampling frame can be used to service an integrated system
of surveys including repeated surveys. See chapter V for details about the construction
and maintenance of master sampling frames.
A domain may also be generally defined as any subset of the population for which
separate estimates are planned in the survey design. A domain could be a stratum, a
combination of strata, an administrative region, or urban, rural or other subdivisions
within these regions. Domains can also be demographic subpopulations defined by
such characteristics as age, race and sex.
It is important that the number of domains of interest for a particular survey be kept at
a moderate level since the sample size required to provide reliable estimates for each
of a large number of domains would necessarily be very large, and indirectly
increasing the overall sample size required, and the cost of the survey.
16
3.4.2. Sample allocation under Domain Estimation
In order to achieve precise survey estimates for domains of interest, there is the need
for samples of adequate sizes to be allocated to the various domains. The problem
arises when equal precision is desired for domains with widely varying population
sizes. If estimates are desired at the same level of precision for all domains, then an
equal allocation (that is to say, the equal sample size per domain) is the most
efficient strategy. The disadvantage of such an allocation is that it can lead to a serious
loss of efficiency for national estimates. To overcome this problem, proportionate
allocation, which uses equal sampling fractions in each domain, is frequently the
most suitable allocation for national estimates. When domains differ markedly in size
and when both national and domain estimates are required, some compromise
between equal allocation and equal sampling fractions is required. Kish (1988)
proposed a compromise between proportional and equal allocation based on an
allocation proportional to
𝑛√𝑊ℎ2 + 𝐻 −2
where 𝑛 is the overall sample, size, 𝑊ℎ is the proportion of the population in stratum ℎ
and 𝐻 is the number of strata.
17
in that level over time between two points in time (for instance, the change in
the poverty rate between two points in time). At this point, the discussion will focus
on the precision of survey estimates in the context of estimation of the level of a
characteristic at a point in time. In particular, the characteristic of interest will be the
percentage of households in poverty (i.e., the poverty rate).
𝑛 𝑝(100 − 𝜌)
𝑠𝑒(𝑝) = √𝑑 2 (𝑝) × (1 − ) ×
𝑁 𝑛
where 𝑛 denotes the overall number of households for the domain of interest, 𝑁
denotes the total number of households in the domain and 𝑑 2 (𝑝) denotes the
estimated design effect associated with the complex design of the survey. 𝑛/𝑁, is the
sampling fraction (proportion of the population that is in the sample) and the factor
[1 − (𝑛 / 𝑁)] is the finite population correction factor (fpc). (the proportion of the
population not included in the sample). The fpc represents the adjustment made to the
standard error of the estimate to account for the fact that the sample is selected without
replacement from a finite population.
𝑠𝑒(𝑝) 𝑛 (100 − 𝜌)
𝑐𝑣(𝑝) = √𝑑 2 (𝑝) × (1 − ) ×
𝑝 𝑁 𝑛𝑝
18
and editing can be done in a fashion that is efficient in terms of both time and money.
In addition, large sample size will imply a large number of staff working on the study.
In such a case, if the interviewers or field personnel are not well trained and
experienced, there may be problems in the data collection and subsequent editing of
the data. Consequently, the data available for analysis will be of low quality, and policy
makers will have low confidence in the decisions being made on the basis of these
data. Another concern that affects the quality of data is non-response. larger sample
sizes make it more difficult and expensive to minimize survey non-response. It is
important to keep survey non-response as low as possible, in order to reduce the
possibility of large biases in the survey estimates. With a smaller sample, it will be
much easier and more cost-effective to revisit households that initially chose not to
participate, in an attempt to persuade them to do so.
19
STAT 444: SURVEY ORGANISATION AND MANAGEMENT
LECTURE 4: QUESTIONNAIRE DESIGN FOR HOUSEHOLD SURVEYS
4.1. Introduction
Household surveys provide a wealth of information on many aspects of life.
Developing a questionnaire is critical part of any household survey design since the
usefulness of household survey data depends heavily on the quality of the survey,
both in terms of questionnaire design and actual implementation in the field. Any
mistake or inaccuracies in the wording or flow of questions may lead to undesired
biases which affects the quality of the survey. Sufficient time allocated to designing
questionnaire and testing the questionnaire helps to mitigate the errors and biases.
The lecture seeks to discuss the most important issues to be considered when
designing survey questionnaires in household surveys.
20
ii. Type 2: Questions that connects household characteristics with
government policies and programmes in order to examine the coverage
of those programmes. An example of this type of question is:
What proportion of households participate in a particular programme, and how
do the characteristics of these households compare with those of households
that do not participate in the programme?
Once a set of questions to be answered has been agreed upon, the questions can be
expressed as objectives of the survey. The next step is to rank these objectives in
order of importance. If the number of objectives is large, it is quite possible that the
survey will not be able to collect all the information needed to achieve all of them
because of low budgets, capacity limitation and other constraints. When this happens,
objectives that have low priority (relative to the effort required to collect the information
needed to attain them) should be dropped. In this process of deciding what objectives
the survey will meet, one must check whether other data that already exist can be
used to answer the question associated with the objective. Any objective that can be
met using existing data from other sources should be dropped from the list of
objectives for the new survey.
Another point to be noted is that some survey designers prefer to express the set of
questions or objectives in terms of a set of tables to be completed using the survey
data. This approach, which is often referred to as the tabulation plan. Tabulation plans
works best with the first three types of questions. More generally, the way in which the
data collected in a household survey will be used to answer the questions (attain the
objectives) can be referred to as the "data analysis plan".
21
i. The financial resources available to undertake the survey limits both how
many households can be surveyed and how much time interviewers can spend
with any given household, which in turn limits how many questions can be
asked of a given household. In general, there are different combinations of
sample size (number of households surveyed) and the amount of information
that one can obtain from each household. In particular, for a given quantity of
financial resources, one can increase the sample size only by decreasing the
amount of information collected from each household, and vice versa. Clearly,
this has implications for the number of objectives of the survey and the precision
of those objectives (that is to say, the accuracy of the answers to the underlying
questions): a small sample size can allow one to collect more data per
household and thus answer more questions of interest, but the precision of
those answers will be lower owing to the lower sample size. A related point is
that the quality of the data, in the sense of the accuracy of the information, will
also be affected by the resources available. For example, if funds are available
to allow each interviewer more time to complete a questionnaire of a given size,
the additional time could be used to return to the household to correct errors or
inconsistencies in the data that are detected after an interview has been
completed.
ii. The capacity of the organization that will implement the survey: Large
sample sizes or highly detailed household questionnaires may exceed the
capacity of the implementing organization to undertake the survey at the
desired level of quality. The larger the sample size, the greater the number of
interviewers and data entry staff that will be necessary to hire and train
(assuming that the amount of time required to complete the survey cannot be
extended), which means that the organization may have to reduce the minimum
acceptable qualifications for interviewers and data entry staff in order to hire the
requisite number. Similarly, more extensive household questionnaires will
require more training and more competent staff, and well-trained, highly
competent interviewers and data entry staff are often in short supply in
developing countries.
22
4.3.1. Module Approach
A household survey questionnaire is usually composed of several parts, often called
modules. A module consists of one or more pages of questions that collect information
on a particular subject, such as housing, employment or health. More generally, in
almost any household survey questionnaire that has several questions on a given
topic, such as the education of each household member, it is convenient to put those
questions together on one or more pages of the questionnaire and to refer to that page
or those pages as the module for that topic; for example, the questions on education
mentioned above would become the "education module". In this way, the entire
questionnaire can be viewed as a collection of modules, perhaps as few as 3 or as
many as 15 or 20, depending on the number of topics covered by the questionnaire.
Each module contains several questions, sometimes only 5 or 6, but other times as
many as 50 or even more than 100. Very large modules, such as those with more than
50 questions, should be further divided into sub-modules that focus on particular
topics. For example, a large module on employment could be divided into the following
sub-modules: primary job, secondary job, and employment history. In any event, the
overall number of questions on a questionnaire should be kept to the minimum
required to elicit the desired information. The module approach is convenient because
it allows the design of the questionnaire to be broken down into two steps. The choice
of modules and the details of each module will vary greatly, depending on the
objectives of, and the constraints faced by, the survey.
Almost all household surveys collect information on the number of people belonging
to the household, and some very basic information on them, such as their age, sex
and relationship to the head of the household. These questions can be put into a short
one page "household roster" module. This module should be one of the first modules
-- and in most cases, the first module -- in the questionnaire. Many household survey
questionnaires will later ask questions of individual household members on topics such
as education, employment, health and migration. Any such topics for which about five
or more questions are asked, should probably be put into a special module on that
topic. If only one, two or three questions are asked, it may be more convenient to
include them in the household roster, or perhaps in another module that asks
questions of individual household members.
Types of Modules
Almost all of the modules in a household survey can be divided into two main types:
i. those that ask questions of individual members: For modules that ask
questions of individual household members, it should be noted that the
questions that are asked of individual household members need not be the
same for each member; many household surveys have questions that apply
only to some types of household members, such as children younger than
five years of age or women of childbearing age.
ii. those that ask general questions about the household: Examples of this
type of questions are questions on the characteristics of the dwelling in
which the household lives and questions on the expenditures of the
household as a whole on food and non-food items.
23
Fig. 4.1: Modules from the Ghana Multiple Indicator Cluster Survey 2017/18
i. Modules that consist of questions that are relatively easy to answer and
that pertain to topics that are not sensitive are to come on top of the
questionnaire. Usually the household roster comes as the first module,
since basic information on household members is usually not a sensitive
topic. Starting the interview with simple questions on non-sensitive
topics will help the interviewer put the household members at ease and
develop a rapport with them. This will give the interviewer as much time
as possible to gain the confidence of the household members, which will
increase the probability that they will answer the sensitive questions fully
and truthfully. In addition, if sensitive questions cause the household
members to stop the interview, at least all of the non-sensitive
information will already have been obtained.
ii. modules that are likely to be answered by the same household member
is grouped together. For example, questions on food and non-food
expenditure should be together because it is likely that one person in the
household is best able to answer both types of questions. This allows
that person to answer all the questions of these modules that he or she
can, and then end his or her participation, leaving other household
members to answer the remaining modules. The general point here is to
use the household members’ time efficiently, which will be appreciated
and thus will increase their co-operation. It is also likely to save the
interviewer’s time because each respondent need be called only once to
make his or her contribution to the interview.
iii. In most all cases, the questions should be written out on the
questionnaire so that the interviewer can conduct the interview by
reading each question from the questionnaire. This ensures that the
same questions are asked of all households. The alternative is for a
24
survey questionnaire to be designed as a form with minimal wording,
which requires each interviewer to pose questions using his or her own
words. This approach it leads to many errors. For example, suppose that
a module on employment has a "question" that simply reads "main
occupation". This is unclear. Does it refer to the occupation on the day
or week of the interview, or the main occupation during that past 12
months? For persons with two occupations, is the main occupation the
one that has the highest income or the one for which the hours or days
worked is the highest? This confusion can be avoided if the question is
written out in detail, as in the following example: "During the past seven
days, what kind of work did you do? If you had more than one kind of
work, tell me the one for which you worked the most hours during the
past seven days."
iv. the questionnaire should include precise definitions of all key concepts
used in the survey questionnaire, primarily to allow the interviewer to
refer to the definition during the interview when unusual cases are
encountered. In addition, the questionnaire should contain some
instructional comments for the interviewer; More elaborate instructions
and explanations of terms should be provided in an interviewer manual.
v. keep questions as short and simple as possible, using common,
everyday terms. If the question is complicated, break it down into two or
more separate questions. Thus avoid double barreled questions: For
example, the following question:
During the past seven days, were you employed for wages or other
remuneration, or were you self-employed in a household enterprise,
were you engaged in both types of activities simultaneously, or were
you engaged in neither activity?
can be replaced with the following two separate questions using less technical
terms:
1. During the past seven days, did you work for pay for someone
who is not a member of this household?
2. During the past seven days, did you work on your own account,
for example, as a farmer or a seller of goods or services?
vi. In addition, all questions should be checked carefully to ensure that they
are not leading. or otherwise likely to induce the respondent to give
biased responses. Eg: In the upcoming elections, will you vote for
the party in power to continue its good works? Will you vote for the
government in power despite its bad performance?
vii. the questionnaire should be designed so that the answers to almost all
questions are pre-coded. Such questions are often called “closed
ended questions” by survey designers. For example, the responses to
questions for which the answer is either yes or no can be recorded in the
questionnaire as "1" for yes and "2" for no. This is easier for the
interviewer, who needs to write only a single digit instead of an entire
word or phrase.9 More importantly, it bypasses the “Coding” step in
which questionnaires with the interviewers. (often illegible) handwritten
25
responses consisting of one or more words are given to an office “coder”
who then writes out numerical codes for those responses. In some cases
the questions should be asked in ways that allow the respondent to
answer in his or her own words. This is known as open-ended
questions.
i. The coding scheme for answers should be consistent across questions.
For example, in almost all household surveys there are many questions
for which the answer is either yes or no. The numerical codes for all such
questions in the questionnaire should always be the same, for example,
“1” for yes and “2” for no. Once this (or some other) coding rule is
established, it should be used for all yes or no responses to questions
on the questionnaire. Thus, the interviewer will learn that he or she
should always code 1 for yes and 2 for no for all yes or no questions
in the questionnaire.
ii. the survey questionnaire should include “skip codes” which indicate
which questions are not to be asked of the household, based on the
answers to previous questions.
4.4.1. Pre-testing
This is the first stage of field-testing the draft questionnaire and it involves trying out
selected sections (modules) of the questionnaire on a small number of households
(for example, 10-15), to obtain an approximate idea of how well the draft questionnaire
pages’ work. This can be done more than once, starting in the early stages of the
questionnaire design process.
26
Another important aspect of the pilot test is that it does not only the draft questionnaire
but also the entire fieldwork plan, including supervision methods, data entry, and
written materials such as interviewer manuals. Only by testing the entire process can
the team be assured that the survey is ready for implementation. A useful last step is
to undertake a “quick analysis” of the data collected in the pilot test to check for
problems that may otherwise be overlooked.
Reading assignment:
The importance of conducting pre-testing
The importance of conducting a pilot test
27
STAT 444: SURVEY ORGANISATION AND MANAGEMENT
LECTURE 5: NON-SAMPLING ERRORS IN HOUSEHOLD SURVEYS
5.1. Introduction
The notion of error as applied to a statistic or estimate of some unknown target quantity
(or parameter) refers to the difference between the estimate (say, , 𝑌̂) and the
theoretical “true parameter value” (say, 𝑌) that would be obtained or reported if all
sources of error were eliminated. Another term for error is deviation but the term error
is so entrenched or common. To illustrate the concept, suppose that the estimate of
the average monthly income for a certain population reported in a survey is 900 United
States dollars, and that the actual average monthly income for members of this
population, obtained from a complete enumeration without errors of reporting and
processing, is US$ 850. Then, in this example, the error of the estimate would be US$
+50. In relation to household surveys, we are concerned with survey errors, that is to
say, errors of estimates based on survey data. According to Lyberg et. al. (1997),
“survey errors” can be decomposed in two broad categories: sampling and non-
sampling errors.
NB: In general, the sampling errors decrease as the sample size increases, whereas
non-sampling error increases as the sample size increases.
Recall that in multistage selection, the Primary and sometimes secondary stages of
selection involve geographical areas considered as clusters of households. In some
subsequent stage of selection, a list of households is obtained, or created, for a set
of relatively small geographical areas. At the last stage of selection, a list of persons
or residents in the household is created in each sampled area. Thus there are three
types of units that need to be considered when examining non-coverage errors in such
surveys: geographical units, households, and persons. These units may be
separate sources of non-response in household surveys.
Housing unit definitions are complex, in as much as they take into account
whether a physical structure is intended as living quarters, and whether
the persons living in the structure live and eat separately from others in
the same structure (as in multi-unit structures such as apartment
buildings). Living separately implies that the residents have direct access to
the living quarters from the outside of the structure, or from a shared lobby or
hallway. The ability to “eat separately” usually involves the presence of a place
to provide and prepare food, or the complete freedom of the residents to choose
the food they eat.
30
the seasonal residents usually live elsewhere, and should not be
counted as part of a household in the seasonal unit.
iv. The non-coverage problem in housing unit listing is made more difficult
by the temporal dimension. A housing unit may be unoccupied at the
time of listing, or under construction. If the survey is to be conducted at
some point in the future, these types of units may need to be included in
the listing.
Finally, within a sampled housing unit, listing of persons who are usual residents
is a part of the household listing process as well. Operational rules are required
to instruct interviewers regarding whom to include in the housing unit as a usual
resident. As in the case of housing units, most determinations are
straightforward. Most persons encountered are staying at the housing unit at
the time of contact, and it is their only place of residence. There are others who
are absent at the time of contact, but for whom the residence is an only
residence. However, there are persons for whom the housing unit is one of
several in which they live. A decision must be made in the field by part-time staff
about whether the sampled housing unit is the usual place where this person
resides.
31
differences between covered and non-covered persons (𝑌̅𝑐 − 𝑌̅𝑛𝑐 ), or must have a
𝑁𝑛𝑐
small proportion of the persons who are not covered by the survey ( ).
𝑁
These two types have quite different implications for survey results, and the methods
used to measure, reduce and report them, and to compensate for them, are in some
ways distinct as well.
32
5.9. Sources of non-response in household surveys
In household surveys, unit non-response can occur for several different kinds of units.
As is the case for non-coverage, non-response may occur for primary or secondary
sampling units. For example, a primary sampling unit might consist of a district or sub-
district in a country. Weather conditions or natural disasters may prevent survey
operations from being conducted in a district or sub-district that has been selected at
a primary, or secondary, stage of sampling. The unit is covered by the survey, but
during the survey period, it is not possible to collect data from any of the households
in the unit.
Non-response is more frequent at the household level. A listed housing unit chosen
for the sample may be found occupied, and an interview attempted. However, as the
interviewer visits the housing unit, several adverse events may prevent data collection.
For example:
A household member may refuse participation as an individual or as a
representative of the entire unit.
Although a housing unit is occupied, its residents may be away from home
during the entire survey period. In some developing countries, a considerable
problem is encountered with housing units clearly lived in but locked during the
entire data-collection period.
In many countries, although occupied housing units have individuals home at
the time of data collection, language may pose a barrier. A version of the
survey’s questionnaire may not have been translated into the language of the
household, or the interviewer may not speak the local language.
Person-level unit non-response also may occur. For surveys that allow proxy
reporting on survey questions, data can be collected from other household
members for persons in the household who are not at home at the time of
interview. For surveys, though, that require self-report for some or all questions,
a person who is not at home during the survey, refuses to participate, or has
another barrier (such as language) that precludes interviewing is a non-
respondent. Health conditions, whether permanent, such as hearing
impairment or blindness, or temporary, such as an episode of a severe acute
illness, may preclude an individual from responding as well.
33
As for non-coverage, the survey designer must either keep the non-response rate
small, or anticipate small differences between responding and non-responding
households and persons.
34