Reliability and validity are concepts used to evaluate the quality of research.
They indicate how
well a method, technique. or test measures something. Reliability is about the consistency of a
measure, and validity is about the accuracy of a measure.
It’s important to consider reliability and validity when you are crea ng your research design,
planning your methods, and wri ng up your results, especially in quan ta ve research. Failing
to do so can lead to several types of research bias and seriously affect your work.
Reliability vs validity
Reliability Validity
What does it tell you? The extent to which the results can be reproduced The extent to which the results really
when the research is repeated under the same measure what they are supposed to
condi ons. measure.
How is it assessed? By checking the consistency of results across me, By checking how well the results correspond
across different observers, and across parts of the to established theories and other measures
test itself. of the same concept.
How do they relate? A reliable measurement is not always valid: the A valid measurement is generally reliable: if a
results might be reproducible, but they’re not test produces accurate results, they should
necessarily correct. be reproducible.
Understanding reliability vs validity
Reliability and validity are closely related, but they mean different things. A measurement can
be reliable without being valid. However, if a measurement is valid, it is usually also reliable.
What is reliability?
Reliability refers to how consistently a method measures something. If the same result can be
consistently achieved by using the same methods under the same circumstances, the
measurement is considered reliable.
You measure the temperature of a liquid sample several mes under iden cal condi ons. The
thermometer displays the same temperature every me, so the results are reliable.
A doctor uses a symptom ques onnaire to diagnose a pa ent with a long-term medical
condi on. Several different doctors use the same ques onnaire with the same pa ent but give
different diagnoses. This indicates that the ques onnaire has low reliability as a measure of the
condi on.
What is validity?
Validity refers to how accurately a method measures what it is intended to measure. If research
has high validity, that means it produces results that correspond to real proper es,
characteris cs, and varia ons in the physical or social world.
High reliability is one indicator that a measurement is valid. If a method is not reliable, it
probably isn’t valid.
If the thermometer shows different temperatures each me, even though you have carefully
controlled condi ons to ensure the sample’s temperature stays the same, the thermometer is
probably malfunc oning, and therefore its measurements are not valid.
If a symptom ques onnaire results in a reliable diagnosis when answered at different mes and
with different doctors, this indicates that it has high validity as a measurement of the medical
condi on.
However, reliability on its own is not enough to ensure validity. Even if a test is reliable, it may
not accurately reflect the real situa on.
The thermometer that you used to test the sample gives reliable results. However, the
thermometer has not been calibrated properly, so the result is 2 degrees lower than the true
value. Therefore, the measurement is not valid.
A group of par cipants take a test designed to measure working memory. The results are
reliable, but par cipants’ scores correlate strongly with their level of reading comprehension.
This indicates that the method might have low validity: the test may be measuring par cipants’
reading comprehension instead of their working memory.
Validity is harder to assess than reliability, but it is even more important. To obtain useful
results, the methods you use to collect data must be valid: the research must be measuring
what it claims to measure. This ensures that your discussion of the data and
the conclusions you draw are also valid.
Receive feedback on language, structure, and forma ng
How are reliability and validity assessed?
Reliability can be es mated by comparing different versions of the same measurement. Validity
is harder to assess, but it can be es mated by comparing the results to other relevant data or
theory. Methods of es ma ng reliability and validity are usually split up into different types.
Types of reliability
Different types of reliability can be es mated through various sta s cal methods.
Types of reliability
Type of What does it assess? Example
reliability
Test-retest The consistency of a A group of par cipants complete a ques onnaire designed
reliability measure across me: do you get to measure personality traits. If they repeat the
the same results when you repeat ques onnaire days, weeks or months apart and give the
the measurement? same answers, this indicates high test-retest reliability.
Interrater The consistency of a Based on an assessment criteria checklist, five examiners
reliability measure across raters or submit substan ally different results for the same student
observers: do you get the same project. This indicates that the assessment checklist has
results when different people low inter-rater reliability (for example, because the criteria
conduct the same measurement? are too subjec ve).
Internal The consistency of the You design a ques onnaire to measure self-esteem. If you
consistency measurement itself: do you get the randomly split the results into two halves, there should be
same results from different parts of a strong correla on between the two sets of results. If the
a test that are designed to measure two results are very different, this indicates low internal
the same thing? consistency.
Types of validity
The validity of a measurement can be es mated based on three main types of evidence. Each
type can be evaluated through expert judgement or sta s cal methods.
Types of validity
Type of validity What does it assess? Example
Construct validity The adherence of a measure A self-esteem ques onnaire could be assessed by
to exis ng theory and measuring other traits known or assumed to be related to
knowledge of the concept being the concept of self-esteem (such as social skills
measured. and op mism). Strong correla on between the scores for
self-esteem and associated traits would indicate high
construct validity.
Types of validity
Type of validity What does it assess? Example
Content validity The extent to which the A test that aims to measure a class of students’ level of
measurement covers all Spanish contains reading, wri ng and speaking
aspects of the concept being components, but no listening component. Experts agree
measured. that listening comprehension is an essen al aspect of
language ability, so the test lacks content validity for
measuring the overall level of ability in Spanish.
Criterion validity The extent to which the result of a A survey is conducted to measure the poli cal opinions of
measure corresponds to other voters in a region. If the results accurately predict the later
valid measures of the same outcome of an elec on in that region, this indicates that
concept. the survey has high criterion validity.
Face validity refers to the extent to which a test or measurement appears, on the surface, to
measure what it is intended to measure.
Face validity is a subjec ve assessment that evaluates whether a test, survey, or measurement
tool seems appropriate and relevant at first glance, without relying on sta s cal or theore cal
evidence. It is primarily concerned with public percep on and the impressions of test-takers,
experts, or stakeholders, rather than rigorous empirical valida on. A test with high face validity
looks like it measures the intended construct, while low face validity may confuse par cipants
or reduce their confidence in the assessment.
Comparison Summary
Validity Type Key Ques on Focus
Construct Are we measuring the right idea? Theory & Defini on
Content Does it cover the whole subject? Comprehensiveness
Face Does it look right to the user? Appearance/Relevance
Validity Type Key Ques on Focus
Criterion Does it match real-world results? Correla on with outcomes
To assess the validity of a cause-and-effect rela onship, you also need to consider internal
validity (the design of the experiment) and external validity (the generalizability of the results).
How to ensure validity and reliability in your research
The reliability and validity of your results depends on crea ng a strong research design,
choosing appropriate methods and samples, and conduc ng the research carefully and
consistently.
Ensuring validity
If you use scores or ra ngs to measure varia ons in something (such as psychological traits,
levels of ability or physical proper es), it’s important that your results reflect the real varia ons
as accurately as possible. Validity should be considered in the very earliest stages of your
research, when you decide how you will collect your data.
Choose appropriate methods of measurement
Ensure that your method and measurement technique are high quality and targeted to
measure exactly what you want to know. They should be thoroughly researched and based on
exis ng knowledge.
For example, to collect data on a personality trait, you could use a standardized ques onnaire
that is considered reliable and valid. If you develop your own ques onnaire, it should be based
on established theory or findings of previous studies, and the ques ons should be carefully and
precisely worded.
Use appropriate sampling methods to select your subjects
To produce valid and generalizable results, clearly define the popula on you are researching
(e.g., people from a specific age range, geographical loca on, or profession). Ensure that you
have enough par cipants and that they are representa ve of the popula on. Failing to do so
can lead to sampling bias and selec on bias.
Ensuring reliability
Reliability should be considered throughout the data collec on process. When you use a tool or
technique to collect data, it’s important that the results are precise, stable, and reproducible.
Apply your methods consistently
Plan your method carefully to make sure you carry out the same steps in the same way for each
measurement. This is especially important if mul ple researchers are involved.
For example, if you are conduc ng interviews or observa ons, clearly define how specific
behaviors or responses will be counted, and make sure ques ons are phrased the same way
each me. Failing to do so can lead to errors such as omi ed variable bias or informa on bias.
Standardize the condi ons of your research
When you collect your data, keep the circumstances as consistent as possible to reduce the
influence of external factors that might create varia on in the results.
For example, in an experimental setup, make sure all par cipants are given the same
informa on and tested under the same condi ons, preferably in a properly randomized se ng.
Failing to do so can lead to a placebo effect, Hawthorne effect, or other demand characteris cs.
If par cipants can guess the aims or objec ves of a study, they may a empt to act in
more socially desirable ways.
Where to write about reliability and validity in a thesis
It’s appropriate to discuss reliability and validity in various sec ons of
your thesis or disserta on or research paper. Showing that you have taken them into account in
planning your research and interpre ng the results makes your work more credible and
trustworthy.
Reliability and validity in a thesis
Sec on Discuss
Literature review What have other researchers done to devise and improve methods that are reliable and
valid?
Methodology How did you plan your research to ensure reliability and validity of the measures used?
This includes the chosen sample set and size, sample prepara on, external condi ons
and measuring techniques.
Results If you calculate reliability and validity, state these values alongside your main results.
Discussion This is the moment to talk about how reliable and valid your results actually were. Were
they consistent, and did they reflect true values? If not, why not?
Conclusion If reliability and validity were a big problem for your findings, it might be helpful to
men on this here.