students have an equal opportunity to
Validity of Tests demonstrate their knowledge and skills.
I. Introduction to Validity of Tests • Informed Decision-Making - Valid
assessments provide educators,
1. Definition of Validity administrators, and policymakers with
reliable information to make important
Validity is a fundamental concept in educational
decisions about curriculum, instruction, and
assessment that refers to the degree to which a
educational policies. Without validity, these
test measures what it is intended to measure.
decisions may be based on flawed or
As defined by Brown (1996, p. 231), validity is
misleading data.
"the degree to which a test measures what it
• Accountability - In an era of increased
claims, or purports, to be measuring." This
educational accountability, valid
definition underscores the critical nature of
assessments are essential for evaluating the
validity in ensuring that assessments accurately
effectiveness of educational programs,
reflect the knowledge, skills, or constructs they
teachers, and institutions. They provide a
are designed to evaluate.
solid foundation for measuring progress and
In the context of educational testing, validity identifying areas for improvement.
goes beyond mere accuracy. It encompasses • Student Motivation and Self-Assessment -
the appropriateness, meaningfulness, and When students perceive tests as valid
usefulness of the specific inferences made from measures of their knowledge and skills, they
test scores (Ornstein, 1990). A valid test are more likely to take them seriously and
provides results that can be interpreted with use the results for self-reflection and
confidence, allowing educators to make improvement.
informed decisions about student learning,
By ensuring that tests accurately measure what
curriculum effectiveness, and instructional
they are intended to measure, validity serves as
strategies.
the cornerstone of effective educational
assessment. It provides the foundation for
meaningful interpretation of test results and
2. Importance of Validity in Educational
supports the overall goal of improving teaching
Assessment
and learning outcomes.
The importance of validity in educational
assessment cannot be overstated. "If a test
does not measure what it is supposed to II. Four Fundamental Approaches to Validity of
measure, it is useless." This statement Tests
highlights the critical role that validity plays in
ensuring the effectiveness and utility of
educational assessments. 1. Construct Validity
Validity is considered the most central and A. Definition and examples
essential quality in the development,
interpretation, and use of educational systems. Construct validity is one of the fundamental
Its importance stems from several key factors: approaches to establishing the validity of a test.
It refers to the degree to which a test measures
• Accurate Measurement - Valid tests provide the theoretical construct or trait it claims to
an accurate representation of students' measure (Ornstein, 1990; Slavin, 1984). In
knowledge, skills, or abilities in the specific essence, construct validity establishes a link
area being assessed. This accuracy is crucial between the underlying psychological construct
for making informed decisions about we wish to measure and the visible
student progress, instructional performance we choose to observe.
effectiveness, and educational
interventions. Construct validity is particularly important
• Fairness - Validity contributes to the when dealing with abstract concepts or traits
fairness of assessments by ensuring that that are not directly observable. Some
test items are relevant and appropriate for examples of constructs that are often measured
the intended population and purpose. This in educational and psychological testing
helps to minimize bias and ensure that all include:
• Knowledge c) Convergent and Discriminant Validity:
• Intelligence Examine how the test results correlate with
• Interest and motivation other measures of the same construct
• Reading comprehension (convergent validity) and how they differ from
• Verbal ability measures of distinct constructs (discriminant
• Science and Technology Literacy validity).
For instance, when assessing "reading d) Factor Analysis: Use statistical techniques
comprehension," construct validity would like factor analysis to identify underlying
ensure that the test items actually measure a dimensions or factors that the test items are
student's ability to understand and interpret measuring.
written text, rather than just their vocabulary or
e) Expert Review: Engage subject matter
memory recall.
experts to review the test items and overall
structure to ensure they align with the
theoretical construct.
B. Questions to consider
f) Longitudinal Studies: Conduct studies over
When evaluating construct validity, there are
time to see if the test results predict future
several key questions that test developers and
performance or behaviors related to the
educators should consider:
construct.
- How would I indicate to others that I have this
g) Iterative Refinement: Continuously refine the
knowledge or possess this trait?
test based on empirical evidence and
- Does the test interrelate with other tests as a theoretical advancements in understanding the
measure of this construct should? construct.
These questions help ensure that the test items By employing these practices, test developers
and overall structure align with the theoretical can build a comprehensive body of evidence
understanding of the construct being supporting the construct validity of their
measured. For example, if developing a test for assessments. This ensures that the inferences
"science literacy," one might consider how a drawn from test scores accurately reflect the
science-literate individual would demonstrate underlying constructs they aim to measure,
their knowledge and skills in real-world thereby enhancing the overall quality and
situations. usefulness of the assessment in educational
contexts.
C. Best practices for measuring
2. Content Validity
Measuring construct validity is a complex
process, and there is "no single best way" to A. Definition and importance
establish it. Instead, the construct validity of a
Content validity refers to the extent to which a
test should be demonstrated through an
test adequately represents the content domain
accumulation of evidence. This approach aligns
it is intended to measure. It is "the degree to
with the modern view of validity as a unitary
which test items match some objective
concept supported by multiple lines of
criterion, such as content of a course or
evidence (Messick, 1995).
textbook, the skills to do a certain job, or
Some best practices for measuring and knowledge deemed essential for some
establishing construct validity include: purpose."
a) Theoretical Foundation: Ensure that the test The importance of content validity lies in its
is grounded in a clear theoretical understanding ability to ensure that a test comprehensively
of the construct being measured. covers the relevant subject matter or skills. For
instance, a high-validity Social Studies test
b) Multiple Measures: Use a variety of would cover the topics taught in a particular
assessment methods to measure the construct, grading period, comparing the knowledge, skills,
as different approaches can provide and values tested in the items to what was
complementary evidence. actually taught. This alignment between test
content and instructional content is crucial for and special conditions. Well-written
fair and accurate assessment of student performance objectives can substitute for a
learning. Table of Specifications by clearly outlining what
students should be able to do as a result of
instruction.
B. Techniques for defining test content
Two primary techniques are commonly used for
3. Applications to Formal Classroom
defining the intended content of a test: the
Assessments
Table of Specifications and Performance
(Instructional) Objectives. For formal classroom assessments, the
application of content validity principles is
1. Table of Specifications
crucial. The provided content suggests the
A Table of Specifications is a two-dimensional following guidelines:
chart that maps out the content areas and
• For short quizzes covering recent
cognitive levels or performance categories that
instruction, a formal Table of Specifications
a test is intended to measure. Its primary
may not be necessary. Teachers can rely on
purpose is to ensure that the test provides a
their mental conception of what is to be
representative sampling of the content and
measured, as long as they ensure alignment
cognitive processes emphasized in instruction.
with recent instructional content.
The typical structure of a Table of • For assessments covering several weeks of
Specifications includes content areas on one instruction, such as written exams and other
axis and categories of performance or cognitive summative evaluations, a formal Table of
levels on the other. The cells within the table Specifications or list of objectives should be
indicate the number or percentage of test used. This ensures comprehensive coverage
questions associated with each intersection of of the instructional content and appropriate
content and cognitive level. distribution of questions across content
areas and cognitive levels.
• The Table of Specifications or list of
To establish the numbers within the table, the objectives should be developed before
teacher must determine: constructing the test. This pre-planning
helps ensure that the content of the test is
1. The total number of questions to be appropriate and aligned with instructional
included in the test goals.
2. The number of questions for each content • As the test is completed, the table or list of
area objectives should be used as a checklist to
3. The number of questions for each cognitive verify that all intended content areas and
level or capability
cognitive levels are adequately represented.
By applying these techniques and
This distribution should reflect the emphasis considerations, educators can enhance the
placed on different topics and skills during content validity of their assessments, ensuring
instruction. The Table of Specifications serves that they accurately measure student learning
as a blueprint for test construction. It guides across the intended range of knowledge and
the test developer in creating a balanced skills.
assessment that accurately reflects the
instructional emphasis and desired learning
outcomes. 4. Other Procedures to Judge Test Content
Relevance
A. Face Validity
2. Performance (Instructional) Objectives
Face validity refers to the extent to which a
Performance objectives provide an alternative
test appears to measure what it claims to
or complementary approach to the Table of
measure, based on the subjective judgment of
Specifications. These objectives specify the
the test-takers or non-expert reviewers. Tests
content area through the behavior, situation,
questions are said to have face validity when
they appear to be relevant to the group being high curricular validity would include questions
examined. that accurately reflect the content and skills
outlined in that curriculum. It would not, for
For example, a math test that includes
instance, include advanced calculus problems
questions about solving equations and
that are beyond the scope of the 8th-grade
geometric problems would have high face
curriculum.
validity for measuring mathematical ability. In
contrast, if the same math test included Curricular validity is particularly important in
numerous questions about historical events, it standardized testing and in situations where
would have low face validity as a measure of assessments are used to evaluate the
mathematical skills. effectiveness of a curriculum. It helps ensure
that students are being tested on material they
have had the opportunity to learn, making the
Advantages: assessment fair and meaningful.
- Context Interpretation: Face validity can help
test-takers "use that 'context' to help interpret
C. Instructional Validity
the questions and provide more useful,
accurate answers." When a test appears Instructional validity, on the other hand, refers
relevant, examinees may approach it with more to how well the test items reflect what is
seriousness and engagement. actually taught in the classroom. This concept
recognizes that there can sometimes be a
- Acceptability: Tests with high face validity are
discrepancy between the intended curriculum
more likely to be accepted by test-takers,
and what is delivered through instruction.
potentially reducing test anxiety and increasing
motivation. For instance, due to time constraints, teacher
preferences, or other factors, certain topics in a
curriculum might receive more or less emphasis
Limitations: in actual classroom instruction. A test with high
instructional validity would accurately reflect
- Subjectivity: Face validity is based on
this reality, testing students primarily on the
subjective judgments and may not accurately
content and skills that were emphasized during
reflect the test's actual validity.
instruction.
- Potential Bias: As mentioned in the content,
Instructional validity is crucial for several
test-takers might "bend & shape their answers
reasons:
to what they think we want," potentially
introducing bias into their responses. • Fairness to students: It ensures that
students are primarily tested on material
- Lack of Scientific Rigor: Face validity does not
they have actually been taught, rather than
provide empirical evidence of a test's actual
on content that may have been in the
effectiveness in measuring the intended
curriculum but not covered in class.
construct.
• Feedback on instruction: By aligning closely
with what was taught, tests with high
instructional validity can provide valuable
B. Curricular Validity feedback on the effectiveness of
Curricular validity refers to how well test items instruction.
reflect the actual curriculum. In other words, it's • Identifying gaps: Comparing tests with high
the extent to which the content of a test aligns curricular validity to those with high
with the intended curriculum for a particular instructional validity can help identify gaps
subject or course. This type of validity is crucial between the intended curriculum and what
for ensuring that assessments are truly is actually being taught in classrooms.
measuring what students are expected to learn The relationship between curricular and
according to the established curriculum. instructional validity highlights the importance
For example, if a school district has a specific of alignment between curriculum, instruction,
mathematics curriculum for 8th grade that and assessment. Ideally, all three should be
emphasizes algebra and geometry, a test with closely aligned to ensure that students are
learning what they're supposed to learn and
that assessments accurately reflect both the
Predictive validity has several important
intended curriculum and the actual instruction.
applications in education:
• College Admissions: As mentioned,
3. Criterion Validity predictive validity is crucial in selecting
students who are likely to succeed in higher
Criterion validity refers to how well a test
education.
correlates with a criterion or outcome measure.
• Career Guidance: Tests with high predictive
It is often divided into two types: concurrent
validity can help guide students towards
validity and predictive validity.
careers where they are likely to excel.
A. Concurrent Validity • Early Intervention: Assessments with good
predictive validity can identify students who
Concurrent validity is the extent to which the
may need additional support to succeed
scores on a new test correlate with scores on a
academically.
well-established test of the same construct,
• Curriculum Planning: By predicting future
when both are administered at about the same
performance, educators can tailor curricula
time. It "measures how well a new test
to better prepare students for upcoming
compares to a well-established test."
challenges.
Example: To establish the concurrent validity of
The content notes that there is often a high
a new reading ability test, you could compare
correlation between GPA, college readiness
its results with those of a previously validated
tests, and educational success, indicating high
reading test. Both tests would be administered
predictive validity for these measures.
to the same group of students within a short
time frame, and the correlation between the
scores would be analyzed.
Concurrent validity is particularly useful when:
• Developing new assessment tools
• Evaluating the effectiveness of different
testing methods for the same construct
• Comparing the performance of different
groups on the same test
The content provides an example where
concurrent validity could be assessed by
comparing scores on a practical test and a
paper test for the same subject. If students
score well on both tests, it suggests good
concurrent validity.
B. Predictive Validity
Predictive validity refers to the extent to which
a test can predict future performance or
behavior. It is "the degree or extent to which
scores on a test can predict later behavior or
test scores."
Example: The content mentions college
admissions as an instance of predictive validity.
Admissions criteria such as Grade Point
Average (GPA) and college readiness scores are
used to predict a student's likely success in
higher education.