Data Collection Tools Lecture Notes
Data Collection Tools Lecture Notes
Page 1 of 21
Table of Contents
TOC \h \o "1-2"
Page 2 of 21
Learning Objectives
By the end of this topic, students should be able to:
1. Define data collection tools and explain their role within the overall research
process.
2. Describe the major categories of data collection tools used in community health
research, including questionnaires, interview schedules, focus group discussion
guides, observation checklists, record review forms, and measurement
instruments.
3. Distinguish between quantitative and qualitative data collection instruments and
identify which research designs each is best suited for.
4. Explain the criteria of validity and reliability as they apply to research instruments,
and describe practical methods of establishing both.
5. Discuss the process of developing, pretesting, and translating a data collection
tool for use within a specific community context such as Machakos County.
6. Identify common sources of error and bias associated with each type of data
collection tool, and describe strategies to minimize them.
7. Apply the principles learned to select an appropriate data collection tool for a
given community health research scenario.
Page 3 of 21
The selection of a data collection tool is never arbitrary. It flows directly from the
research objectives, the type of data required (quantitative, qualitative, or both), the
characteristics of the study population (literacy levels, language, cultural norms, age
distribution), the resources available to the researcher (time, funding, trained
personnel), and the practical and ethical constraints of the research setting. For
example, a study seeking to measure the prevalence of stunting among children under
five in Mwala Sub-County would primarily require anthropometric measurement tools,
whereas a study exploring why caregivers in the same sub-county delay seeking
treatment for childhood diarrhoea would be better served by in-depth interviews or focus
group discussions that can capture beliefs, fears, and decision-making processes that
numbers alone cannot reveal.
It is useful at the outset to distinguish between an instrument and a tool more broadly
construed, although in practice the two terms are frequently used interchangeably in
research methods literature. The instrument is typically the specific document or device
used (for instance, a structured questionnaire with thirty items), while the broader data
collection tool may also include the procedures, training, and protocols that govern how
that instrument is administered. For the purposes of this course, we will treat the terms
as synonymous, while emphasizing that the success of any tool depends not only on its
design but also on the rigor with which it is administered in the field.
Page 4 of 21
informant interview guides, require a trained researcher or research assistant to ask
questions and record responses, either face-to-face or by telephone. A third category
consists of observational tools, in which the researcher records what is seen or
measured directly, without relying on what a participant says, as in an observation
checklist used to assess hand hygiene practice at a health facility or a clinical
measurement tool used to record a child's weight and height.
A third dimension worth noting is the source of the data: primary data collection tools
gather new, firsthand information directly from study participants or through direct
observation, whereas secondary data collection tools, such as record review forms and
document analysis checklists, are used to extract information that already exists in
records, registers, or files. Many community health studies in practice combine multiple
types of tools within a single study design; for instance, a mixed-methods study on
uptake of the Covid-19 vaccine in Machakos Town Sub-County might use a structured
questionnaire to estimate the proportion of adults vaccinated, an observation checklist
to assess vaccination site readiness, and focus group discussions to explore community
hesitancy, thereby triangulating data from different instruments to produce a more
complete picture.
3. Questionnaires
Page 5 of 21
are easy to code and analyse statistically and reduce variability introduced by
differences in respondent expression, but they constrain respondents to predetermined
categories and may miss nuance. Open-ended questions, by contrast, allow
respondents to answer in their own words, providing richer detail at the cost of being
more time-consuming to code and analyse. Many community health questionnaires use
a hybrid approach, employing mostly closed-ended items for core variables (such as
age, household size, or vaccination status) while including a small number of open-
ended items to capture unanticipated information, such as 'In your own words, what
challenges have you faced in accessing the nearest health facility?'
Page 6 of 21
Machakos County Example
A study assessing knowledge, attitudes, and practices (KAP) regarding cervical cancer
screening among women aged 25 to 49 attending Machakos Level 5 Hospital might
use a structured questionnaire administered via tablet at the outpatient waiting area,
combining closed-ended items (e.g., 'Have you ever heard of cervical cancer
screening? Yes/No') with a handful of open-ended items asking women to describe, in
their own words, any barriers they perceive to attending screening.
Interview schedules also permit the interviewer to clarify a confusing question, probe for
more complete answers, and observe the respondent's demeanor, which can provide
valuable contextual information. However, this same interpersonal dynamic introduces
the risk of interviewer bias, in which the interviewer's tone, body language, or even
subtle rewording of questions inadvertently influences how a respondent answers. For
this reason, interviewers must be carefully trained to administer the schedule in a
standardized manner, reading questions exactly as worded and avoiding leading
prompts.
Page 7 of 21
community health surveys, such as the Kenya Demographic and Health Survey
(KDHS), where standardized data must be collected from thousands of households
across diverse counties, including Machakos, in a comparable manner.
Between these two extremes lies the semi-structured interview, which uses a guide
containing a core set of open-ended questions that are asked of every respondent,
supplemented by optional probes that the interviewer can use at their discretion to
explore interesting or unexpected responses in greater depth. Semi-structured
interviews are extremely common in health systems research because they balance the
comparability benefits of structure with the depth benefits of flexibility.
Page 8 of 21
respond not only to the moderator's questions but to one another's comments, which
can surface shared community norms, areas of consensus, and points of disagreement
that might never emerge in a one-on-one setting.
The composition of an FGD matters considerably. Researchers typically aim for relative
homogeneity within a group (for example, separate groups for men and women, or for
different age categories) to encourage open discussion, since participants may feel
inhibited discussing sensitive topics, such as family planning or stigmatized illnesses, in
the presence of individuals with substantially different social standing, such as elders or
local administrators. A skilled moderator is essential, both to keep the discussion
focused on the topic guide and to ensure that quieter participants are given the
opportunity to speak and that no single participant dominates the conversation.
FGDs are typically audio recorded with participant consent and later transcribed
verbatim for thematic analysis. A note-taker is usually present in addition to the
moderator, both to record non-verbal observations (such as group dynamics or
disagreement) and to serve as a backup in case of recording equipment failure.
Page 9 of 21
• Findings from an FGD reflect group-level rather than individual-level
perspectives, and cannot be used to estimate prevalence or attribute specific
views to specific individuals with confidence.
6. Observation Checklists
Observation as a data collection method involves the systematic watching and
recording of behaviours, conditions, processes, or events as they naturally occur, rather
than relying on a participant's self-report of those behaviours. An observation checklist
operationalizes this process by providing the observer with a structured list of specific
items to look for and record, often with predefined response categories (such as
'present/absent', 'adequate/inadequate', or a numeric rating scale).
Page 10 of 21
A significant methodological concern with observation is the Hawthorne effect, in which
the very act of being observed changes the behaviour of those being watched, often
making them behave in a more 'correct' or socially desirable manner than they would in
the observer's absence. Researchers attempt to minimize this effect through strategies
such as extended observation periods that allow novelty to wear off, or unobtrusive
positioning of the observer.
Common sources for record review in community health research include patient case
files, outpatient and inpatient registers, antenatal care (ANC) booklets and registers,
immunization registers, laboratory logbooks, death registers, and facility-level reports
such as monthly Kenya Health Information System (KHIS/DHIS2) summaries. Record
review is especially valuable for retrospective studies, such as a study examining trends
in malaria case notifications at Mwala Sub-County Hospital over the preceding five
years, since it allows researchers to access historical data that could not feasibly be
recreated through prospective data collection.
Page 11 of 21
The principal limitation of record review is that the researcher has no control over the
original quality, completeness, or accuracy of the records being reviewed; missing
entries, illegible handwriting, inconsistent coding practices between different clinicians,
and outright recording errors are common challenges, particularly in settings where
paper-based record keeping persists alongside the gradual rollout of electronic health
information systems. Researchers using this method must therefore develop clear
decision rules in advance for how to handle missing or ambiguous data, and should
ideally pilot the extraction form on a small sample of records before full-scale data
abstraction begins.
Page 12 of 21
checks (in which two observers independently measure the same participants and their
results are compared) are essential components of rigorous anthropometric and clinical
data collection.
Key informant interviews are particularly valuable for providing contextual or systemic
insight that complements data gathered from the general population, helping
researchers interpret patterns observed in survey data or triangulate findings across
different data sources and levels of the health system, from community level through
facility level to county administration.
Page 13 of 21
10. Validity and Reliability of Data Collection Tools
Regardless of which type of tool is selected, every data collection instrument used in
research must be evaluated against two foundational psychometric properties: validity
and reliability. These concepts are central to research methods because a tool that
lacks either property will produce data that cannot be trusted to accurately answer the
research question, no matter how sophisticated the subsequent statistical analysis may
be.
10.1 Validity
Validity refers to the extent to which an instrument actually measures the concept it is
intended to measure. A tool can be perfectly consistent (reliable) yet still measure the
wrong thing entirely, which is why validity and reliability, while related, are distinct
properties that must be assessed separately. Several specific types of validity are
commonly distinguished in research methods literature.
Face validity is a more basic and subjective form of validity referring to whether an
instrument appears, on its surface, to measure what it claims to measure, as judged by
laypersons or even by the researcher's own inspection. While necessary, face validity
alone is an insufficient basis for confidence in an instrument's overall validity.
Construct validity refers to the degree to which an instrument truly measures the
underlying theoretical construct it claims to measure, particularly important for abstract
concepts such as 'health-seeking behaviour', 'stigma', or 'self-efficacy' that cannot be
observed directly. Construct validity is often assessed through convergent validity
(demonstrating that the new instrument correlates as expected with other established
measures of the same or related constructs) and discriminant validity (demonstrating
Page 14 of 21
that the instrument does not correlate strongly with measures of unrelated constructs,
confirming it is measuring something distinct).
Criterion validity refers to the degree to which an instrument's results correlate with an
external, independently verified criterion or 'gold standard' measure of the same
concept. This is further divided into concurrent validity, where the instrument's results
are compared against a criterion measured at roughly the same point in time (for
example, comparing a rapid malaria diagnostic test against the gold-standard
microscopy result from the same blood sample), and predictive validity, where the
instrument's results are used to predict a future outcome (for example, assessing
whether scores on a depression screening tool predict subsequent clinical diagnosis of
depression).
10.2 Reliability
Reliability refers to the consistency, stability, and repeatability of an instrument's
measurements; a reliable instrument will produce the same or very similar results when
applied repeatedly under the same conditions, regardless of whether or not those
results are actually valid. Several specific approaches are used to assess different
forms of reliability.
Internal consistency reliability assesses whether the different items within a single
instrument that are intended to measure the same underlying construct produce
consistent results, and is most commonly quantified using Cronbach's alpha, a statistic
ranging from zero to one, with values above 0.7 generally considered acceptable for
research purposes; this approach is especially relevant for multi-item scales, such as a
ten-item scale measuring attitudes toward family planning.
Inter-rater (or inter-observer) reliability assesses the degree of agreement between two
or more independent observers or interviewers who use the same instrument to assess
the same subjects or events, and is particularly important for observation checklists and
clinical measurement tools where human judgment is involved in recording a result; this
Page 15 of 21
is commonly quantified using statistics such as Cohen's kappa for categorical data or
the intraclass correlation coefficient (ICC) for continuous data.
Parallel-forms reliability, less commonly used in routine community health research but
important in large-scale survey methodology, assesses whether two different but
theoretically equivalent versions of an instrument produce consistent results when
administered to the same group of respondents.
Applied Example
Before deploying a new MUAC measurement protocol across all sub-counties in
Machakos County, the research team might conduct an inter-rater reliability exercise in
which several research assistants independently measure the same group of fifteen
children, then calculate the intraclass correlation coefficient to confirm that the
assistants are taking sufficiently consistent measurements before they are deployed
independently across different health facilities.
When existing instruments are unavailable or do not adequately fit the specific local
context or research question, researchers develop new items, organizing them logically
(commonly proceeding from general demographic items, through the core substantive
content of the study, to any sensitive items, which are typically placed toward the end of
an instrument once rapport has been established). Item wording must avoid double-
barrelled questions (asking two things at once), leading or loaded language, jargon
Page 16 of 21
unfamiliar to the target population, and double negatives, all of which can introduce
measurement error.
11.2 Pretesting
Pretesting, sometimes called piloting, involves administering a draft instrument to a
small sample of individuals who are similar to the intended study population but who will
not be included in the actual study, in order to identify problems with question wording,
flow, length, comprehension, and overall administration time before the instrument is
used on a larger scale. Pretesting is an indispensable step that is sometimes skipped
under time or budget pressure, but doing so substantially increases the risk that flawed
instrument design is only discovered after costly large-scale data collection has already
taken place.
During pretesting, researchers pay close attention to several specific issues: whether
respondents interpret questions in the way intended by the researcher (sometimes
assessed through cognitive interviewing, in which respondents are asked to explain in
their own words what they understood a question to mean), whether any questions are
routinely skipped or met with confusion, whether the overall length of the instrument
leads to respondent fatigue, and whether response categories adequately capture the
range of likely answers. Following pretesting, the instrument is revised accordingly, and
in larger studies, a second round of pretesting may be conducted on the revised
version.
Page 17 of 21
discussion, often involving a third bilingual reviewer, before the translated instrument is
finalized. Beyond linguistic translation, cultural adaptation may also be necessary,
ensuring that examples, idioms, units of measurement, and culturally sensitive topics
are appropriately localized rather than merely translated word for word; for example, a
question about household assets used to assess socioeconomic status should
reference assets that are actually meaningful and common within rural Machakos
households, such as ownership of livestock, rather than assets more typical of an
entirely different setting.
• Recall bias occurs when respondents are unable to accurately remember past
events, behaviours, or exposures, and is a particular concern for questionnaires
and interviews asking about events occurring further in the past, such as a
caregiver's recall of a child's complete vaccination history if the vaccination card
has been lost.
• Social desirability bias occurs when respondents provide answers they believe
are more socially acceptable rather than entirely truthful answers, a concern that
affects self-reported behaviours around sensitive topics such as sexual practices,
substance use, or hygiene behaviours, and is one of the principal reasons
observation is sometimes preferred over self-report for such topics.
• Interviewer bias occurs when an interviewer's characteristics, tone, body
language, or subtle deviations from standardized wording systematically
influence how respondents answer, and is mitigated through rigorous interviewer
training and standardized administration protocols.
• The Hawthorne effect, introduced earlier in the context of observation, occurs
when the awareness of being observed itself changes the behaviour being
studied.
• Non-response bias occurs when individuals who decline to participate, or who do
not complete an instrument, differ systematically from those who do respond,
Page 18 of 21
potentially skewing study findings if non-response is associated with the
variables of interest.
• Instrument or measurement error refers to inaccuracies introduced by poorly
calibrated equipment, ambiguous item wording, or translation errors that distort
the true value of what is being measured.
Most of these sources of error can be substantially reduced, though rarely eliminated
entirely, through careful instrument design, thorough pretesting, rigorous interviewer
and observer training, standardized protocols for administration and measurement, and,
where appropriate, triangulation of data collected through multiple complementary tools.
Page 19 of 21
Tool Data Type Key Strength Key Limitation
14. Conclusion
The selection and design of an appropriate data collection tool is foundational to the
credibility of any community health research study. Researchers must align their choice
of tool, or combination of tools, with their specific research objectives, the nature of the
data required, the characteristics and constraints of the study population, and the
resources realistically available to them. Equally important is rigorous attention to the
validity and reliability of whatever instrument is chosen, achieved through careful
development grounded in existing literature, systematic pretesting within the actual
target population, and, where necessary, methodologically sound translation and
cultural adaptation.
Page 20 of 21
9. Explain the difference between a structured interview and an unstructured (in-
depth) interview, and describe one research scenario in which each would be the
more appropriate choice.
10. Discuss three strengths and three limitations of focus group discussions as a
qualitative data collection tool.
11. Define the Hawthorne effect and describe two practical strategies a researcher
could use to minimize its influence during structured observation.
12. Distinguish between validity and reliability, and explain why an instrument can be
reliable without being valid.
13. Describe the forward-and-back translation process and explain why it is preferred
over direct, single-step translation when adapting a research instrument into
Kikamba.
14. Identify and briefly explain three common sources of bias that can affect data
collected through self-administered questionnaires.
15. Design a brief outline of a mixed-methods data collection plan to investigate low
uptake of antenatal care services in a selected Machakos County sub-county,
specifying which tools you would use and why.
Page 21 of 21