SE 202
Presenter:
Amarah T. Ibra
MST – General Science
Summer 2025
VA L I D I T Y
&
RE L I AB I L I TY
Objectives
Explain what is meant by the term “validity” as it applies to the use of instruments in
educational research.
Name three types of evidence of validity that can be obtained and give an example of each type.
Explain what is meant by the term “correlation coefficient” and describe briefly the
difference between positive and negative correlation coefficients.
Explain what is meant by the terms “validity coefficient” and “reliability coefficient.”
Explain what is meant by the term “reliability” as it applies to the use of instruments in
educational research.
Explain what is meant by the term “errors of measurement.”
Explain briefly the meaning and use of the term “standard error of measurement.”
Describe briefly three ways to estimate the reliability of the scores obtained using a particular
instrument.
Describe how to obtain and evaluate scoring agreement.
Researchers use a number of procedures to ensure that the inferences
they draw, based on the data they collect, are valid and reliable.
Validity Reliability
refers to the refers to the
appropriateness,
appropriateness, consistency
consistency of scores
meaningfulness,
meaningfulness, or answers from one
correctness,
correctness, and administration of an
usefulness
usefulness of the instrument to another,
inferences a re and from one set of
items to another.
searcher makes.
VA L I D I T Y
Validation is the process of collecting and analyzing
evidence to support inferences.
Validity refers to the degree to which evidence supports
any inferences a researcher makes based on the data he
or she collects using a particular instrument.
VA L I D I T Y
An appropriate A meaningful A useful inference
inference would inference is one is one that helps
be one that is that says something researchers make
relevant—that about the meaning a decision related
is, related—to
of the information to what they
obtained through were trying to
the purposes of the use of an find out.
the study. instrument.
VA L I D I T Y
Validity, therefore, depends on the amount and
type of evidence there is to support the
interpretations researchers wish to make
concerning data they have collected.
Do the results of
the assessment
What kinds of
provide useful
evidence might a
information
researcher
about the topic
collect?
or variable being
measured?
three main types of evidence of validity
It refers to the content and format of the
1 Content-related evidence of validity instrument. The content and format must be
consistent with the definition of the variable and the
sample of subjects to be measured.
Essential Questions:
How appropriate is the content?
How comprehensive?
Does it logically get at the intended variable?
How adequately does the sample of items or questions represent the content to be
assessed?
Is the format appropriate?
three main types of evidence of validity
It refers to the content and format of the
1 Content-related evidence of validity instrument. The content and format must be
consistent with the definition of the variable and the
sample of subjects to be measured.
Essential Questions:
How appropriate is the content?
How comprehensive?
Does it logically get at the intended variable?
How adequately does the sample of items or questions represent the content to be
assessed?
Is the format appropriate?
1 Content-related evidence of validity
Example: Suppose a researcher desires to measure students’ ability to use
information that they have previously acquired. When asked what
she means by this phrase, she offers the following definition.
As evidence that students can use previously acquired in
formation, they should be able to:
1. Draw a correct conclusion (verbally or in writing) that is based
on information they are given.
2. Identify one or more logical implications that follow from a
given point of view.
3. State (orally or in writing) whether two ideas are identical,
similar, unrelated, or contradictory.
Here are three examples of the kinds of 2. Those who believe that increasing
questions she has in mind, designed to consumer expenditures would be the
produce each of the three types of best way to stimulate the economy
evidence listed above. would advocate
a. an increase in interest rates.
b. an increase in depletion
1. If A is greater than B, and B is allowances.
greater than C, then: c. tax reductions in the lower
a. A must be greater than C. income brackets.
b. C must be smaller than A. d. a reduction in government
c. B must be smaller than A. expenditures.
d. All of the above are true.
3. Compare the dollar amounts spent
by the U.S. government during the
past 10 years for ( a ) debt payments,
( b ) defense, and ( c ) social services.
Here are three examples of the kinds of
questions she has in mind, designed to
produce each of the three types of
evidence listed above.
given information
1. If A is greater than B, and B is
greater than C, then:
a. A must be greater than C.
b. C must be smaller than A.
c. B must be smaller than A. Students can draw a correct conclusion
d. All of the above are true. (verbally or in writing) that is based on
correct answer information they are given.
Although it could be considered
questionable, since students might view it
as somewhat tricky.
2. Those who believe that increasing
consumer expenditures would be the
point of view best way to stimulate the economy
would advocate
a. an increase in interest rates.
b. an increase in depletion
allowances.
c. tax reductions in the lower
Students can identify one or more income brackets.
logical implications that follow from d. a reduction in government
a given point of view. expenditures.
We would not rate the answers to 3 as 3. Compare the dollar amounts spent
valid, since students are not asked to by the U.S. government during the
contrast ideas, only facts. past 10 years for ( a ) debt payments,
( b ) defense, and ( c ) social services.
three main types of evidence of validity
It refers to the relationship between scores
2 Criterion-related evidence of validity obtained using the instrument and scores obtained
using one or more other instruments or measures
(often called a criterion).
Criterion is a second test or other assessment procedure presumed to measure the same
variable.
To obtain evidence of predictive validity, researchers allow a time interval to elapse
between administration of the instrument and obtaining the criterion scores.
On the other hand, when instrument data and criterion data are gathered at nearly the
same time, and the results are compared, this is an attempt by researchers to obtain
evidence of concurrent validity.
2 Criterion-related evidence of validity
A correlation coefficient, symbolized by the letter r, indicates the degree of relationship
that exists between the scores individuals obtain on two instruments.
When a correlation coefficient is used to describe the relationship between a set of
scores obtained by the same group of individuals on a particular instrument and their
scores on some criterion measure, it is called a validity coefficient.
***For example, a validity coefficient of +1.00 obtained by correlating a set of scores on
a mathematics aptitude test (the predictor) and another set of scores, this time on a
mathematics achievement test (the criterion), for the same individuals would indicate
that each individual in the group had exactly the same relative standing on both
measures.
2 Criterion-related evidence of validity
Gronlund suggests the use of an expectancy table as another way to depict criterion-
related evidence.
criterion
predictor categories
categories
three main types of evidence of validity
It refers to the nature of the psychological
3 Construct-related evidence of validity construct or characteristic being measured by the
instrument.
Steps involved in obtaining construct-related evidence of validity :
1 the variable being measured is clearly defined;
hypotheses, based on a theory underlying the variable, are formed
2 about how people who possess a lot versus a little of the variable
will behave in a particular situation;
and the hypotheses are tested both logically and
3
empirically.
3 Construct-related evidence of validity
Example: Suppose a researcher interested in developing a pencil-and-paper test
to measure honesty wants to use a construct-validity approach.
1 First, he defines honesty .
The researcher will next formulate a theory about how “honest” people
behave as compared to “dishonest” people. For example, he might
2 theorize that honest individuals, if they find an object that does not
belong to them, will make a reason able effort to locate the individual to
whom the object belongs.
3 Construct-related evidence of validity
Example: Suppose a researcher interested in developing a pencil-and-paper test
to measure honesty wants to use a construct-validity approach.
1 First, he defines honesty .
The researcher might hypothesize that individuals who score high on his
2 honesty test will be more likely to attempt to locate the owner of an
object they find than individuals who score low on the test.
3 Construct-related evidence of validity
Example: Suppose a researcher interested in developing a pencil-and-paper test
to measure honesty wants to use a construct-validity approach.
1 First, he defines honesty .
The researcher might hypothesize that individuals who score high on his
2 honesty test will be more likely to attempt to locate the owner of an
object they find than individuals who score low on the test.
The researcher then administers the honesty test,
3 separates the names of those who score high and those
who score low, and gives all of them an opportunity to be
honest through observation.
R E L I AB I L I TY
It refers to the consistency of scores or answers from
one administration of an instrument to another, and
from one set of items to another.
Errors of Measurement
Whenever people take the same test twice, they will seldom perform exactly the
same—that is, their scores or answers will not usually be identical. This may be
due to a variety of factors (differences in motivation, energy, anxiety, a different
testing situation, and so on), and it is inevitable. Such factors result in errors of
measurement.
Reliability estimates provide researchers with an idea of how much variation to
expect. Such estimates are usually expressed as another application of the
correlation coefficient known as a reliability coefficient.
A reliability coefficient expresses a relationship between scores of the same
individuals on the same instrument at two different times, or on two parts of
the same instrument.
The three best-known ways to obtain a reliability coefficient
Test – Retest Method
It involves administering the same test twice to the
same group after a certain time interval has elapsed. A
reliability coefficient is then calculated to indicate the
relationship between the two sets of scores obtained.
Reliability coefficients will be affected by the length of
time that elapses between the two administrations of
the test. The longer the time interval, the lower the
reliability coefficient is likely to be, since there is a
greater likelihood of changes in the individuals taking
the test.
The three best-known ways to obtain a reliability coefficient
Equivalent Forms Method
It is the use of two different but equivalent (also called
alternate or parallel ) forms of an instrument are
administered to the same group of individuals during the
same time period.
A reliability coefficient is then calculated between the
two sets of scores obtained. A high coefficient would
indicate strong evidence of reliability—that the two
forms are measuring the same thing.
The three best-known ways to obtain a reliability coefficient
Internal Consistency Methods
It require only a single administration of an instrument.
a. Split-half Procedure
The split-half procedure involves scoring two halves (usually odd
items versus even items) of a test separately for each person and then
calculating a correlation coefficient for the two sets of scores.
Spearman-Brown prophecy formula:
b. Kuder-Richardson Approaches
Perhaps the most frequently employed method for determining
internal consistency is the Kuder-Richardson approach , particularly
formulas KR20 and KR21.
Mean of the set of test scores
Number of items on the test
Standard deviation of the
set of test scores
***Formula KR20 does not require the assumption that all items are of equal difficulty,
although it is harder to calculate. Computer programs for doing so are commonly
available, however, and should be used whenever a researcher cannot assume that all
items are of equal difficulty.
ScoringAgreement
Scoring Agreement
Differences in the resulting scores with different administrators or scorers are still possible,
it is generally considered highly unlikely that they would occur. This is the case with instruments that
are susceptible to differences in administration, scoring, or both, such as essay evaluations.
In particular, instruments that use direct observation are highly vulnerable to observer
differences. Researchers who use such instruments are obliged to investigate and report the degree
of scoring agreement
What is desired is a correlation of at least .90 among scorers or agreement of at
least 80 percent.
THE STANDARD ERROR OF MEASUREMENT
(SEMeas)
The standard error of measurement (SEMeas) is an index
that shows the extent to which a measurement would vary
under changed circumstances (i.e., the amount of
measurement error ).
Standard deviation of
scores
SEM = SD 𝟏 − 𝒓𝟏𝟏
the reliability coefficient
appropriate to the
conditions that vary
Recap
Validity refers to the appropriateness, meaningful ness, correctness, and usefulness of any
inferences a researcher draws based on data obtained through the use of an instrument.
Evidence of Validity: content-related, criterion-related, construct-related
A criterion is a standard for judging; with reference to validity, it is a second instrument against
which scores on an instrument can be checked.
A validity coefficient is a numerical index representing the degree of correspondence between
scores on an instrument and a criterion measure.
An expectancy table is a two-way chart used to evaluate criterion-related evidence of validity.
Reliability refers to the consistency of scores or answers provided by an instrument.
Errors of measurement refer to variations in scores obtained by the same individuals on
the same instrument.
The test-retest method of estimating reliability involves administering the same instrument twice
to the same group of individuals after a certain time interval has elapsed.
Recap
The equivalent-forms method of estimating reliability involves administering two different, but
equivalent, forms of an instrument to the same group of individuals at the same time.
The internal-consistency method of estimating reliability involves comparing responses to
different sets of items that are part of an instrument.
Scoring agreement requires a demonstration that independent scorers can achieve satisfactory
agreement in their scoring.
The standard error of measurement is a numerical index of measurement error.
“Let us work together to achieve
much greater things”