1
Validity
Validity is the most important property of any test or assessment. It answers the
question: "Are we measuring what we intend to measure?" A test can be reliable (consistent)
but invalid if it measures the wrong thing.
1. Face Validity
The extent to which a test appears to measure what it claims to measure, based on
superficial inspection. It's not a technical form of validity, but it affects how willing people are to
engage with the test.
Examples:
o A Driver's License Test: The test involves sitting in a car, driving on roads,
parallel parking, and following traffic rules. On its face, it clearly looks like it's
assessing driving ability. If the test involved solving a crossword puzzle, it would
lack face validity for assessing driving skill.
o A Depression Inventory: A questionnaire that asks, "In the last two weeks, how
often have you felt sad, lost interest in activities, or had low energy?" has high
face validity for depression. A questionnaire that only asks about favorite colors
would not.
2. Content Validity
The extent to which a test comprehensively samples the entire domain or construct is
supposed to measure. It asks: "Does this test cover all important aspects of the topic?"
Examples:
o A Final Exam for a History Course: If the course covers 10 chapters, a valid exam
should have questions from all 10 chapters. An exam that only has questions from
Psychological assessment (Bs 5th semester)
2
chapter 1 lacks content validity for the entire course. The syllabus acts as the
"blueprint" for ensuring content validity.
o A Test for Managerial Competence: A good test should sample all key areas:
leadership, communication, financial acumen, conflict resolution, and strategic
planning. A test that only measures financial skills has poor content validity for
the broad construct of "managerial competence."
3. Criterion-Related Validity
This type of validity compares the test scores to a concrete, external criterion. It has two
key subtypes:
A. Concurrent Validity
The test's ability to distinguish between groups who are known to differ on the
construct right now. The test and the criterion are measured at approximately the same time.
Examples
o You have a trusted, certified thermometer that reads 98.6°F. You test a new, fancy
thermometer. If the new one also reads 98.6°F, it has high concurrent validity with
the established "gold standard."
o A new 5-minute questionnaire is given to a group of people already diagnosed
with major depression and a group with no history of depression. If the scores
from the quick screen accurately distinguish the two groups at the present
moment, it has good concurrent validity.
B. Predictive Validity
The test's ability to forecast a future outcome or performance. The test is given now,
and the criterion is measured later.
Examples
Psychological assessment (Bs 5th semester)
3
The SAT/ACT for College Admissions: These tests are intended to have predictive
validity. Students take them in high school, and the scores are used to predict their
future academic performance (e.g., first-year college GPA). If students with high SAT
scores consistently get higher GPAs, the test has demonstrated predictive validity.
A battery of cognitive and personality tests is given to pilot trainees. The goal is to
predict which trainees will successfully complete flight school 2 years later. If the test
scores correlate strongly with later success or failure, the test has high predictive
validity.
4. Construct Validity
The overarching, most important type of validity. It is the extent to which a test accurately
measures an abstract theoretical construct. It's not established by a single study but is built up
over time by gathering multiple forms of evidence. Constructs include things like intelligence,
anxiety, self-esteem, or resilience.
How it's built (with examples):
1. Convergent Validity: The test scores correlate highly with other tests that are
supposed to measure the same or similar construct.
Example: A new "Resilience Scale" should yield scores that are positively
correlated with an established "Grit Questionnaire" and negatively
correlated with a "Hopelessness Scale." They are all tapping into related
ideas of psychological strength.
2. Discriminant Validity (Divergent Validity): The test scores do not correlate
strongly with tests that measure different, unrelated constructs.
Example: Scores on an "Emotional Intelligence" test should not be
strongly correlated with height or shoe size. If they were, we'd question
what the test is really measuring. It also should not be too highly
correlated with general intelligence (IQ), to prove it's a distinct construct.
Psychological assessment (Bs 5th semester)
4
3. Known-Groups Validity: The test can distinguish between groups that are
theoretically known to differ on the construct.
Example: A valid "Social Anxiety Scale" should yield significantly higher
scores in a group diagnosed with Social Anxiety Disorder compared to a
randomly selected group from the public.
Factors That Influence Validity
1) Poorly Defined Concept: You can't measure something accurately if you haven't clearly
defined what it is.
2) Length of the test: larger tests increase validity, but it shouldn’t be too large.
3) Bad Questions: Confusing or biased questions on a test make people answer incorrectly.
4) Incomplete Coverage: A test is invalid if it doesn't cover all parts of the topic it claims
to measure.
5) Wrong Difficulty: A test that is too easy or too hard for a group fails to show the real
differences between them.
6) Unfair Administration: If a test isn't given the same way to everyone, their scores
reflect the conditions, not their true ability.
7) Test-Taker's State: A person's temporary feelings, like anxiety or fatigue, can change
their score and make it invalid.
8) Faking Answers: When people answer questions to look good rather than being honest,
the test results are misleading.
9) Cultural Bias: A test is invalid if it requires specific cultural knowledge that not
everyone has.
10) Inconsistent Test: A test that gives different results each time cannot be trusted to
measure anything accurately.
Psychological assessment (Bs 5th semester)
5
11) Wrong Comparison Group: Interpreting your score by comparing it to the wrong group
of people leads to an invalid conclusion.
Summary Table: Validity in Daily Life & Psychology
Type of Psychological
Core Question Daily Life Example
Validity Example
Does it look like A depression survey
Face A driving test that
it measures the asking about mood and
Validity involves driving a car.
right thing? sleep.
A managerial test
Does it cover all A history exam with
Content assessing leadership,
parts of the questions from all
Validity finance, and
topic? chapters.
communication.
Can it identify A new thermometer A quick screen correctly
Criterion:
current groups matching a gold-standard identifying people
Concurrent
accurately? one. with/without depression.
Can it forecast A pilot aptitude test
Criterion: SAT scores predicting
future predicting flight school
Predictive first-year college GPA.
performance? success.
Is it truly Convergent: A new An IQ test correlating
Construct
measuring the fitness app's "effort" score with academic success
Validity
abstract correlating with heart (convergent) but not
Psychological assessment (Bs 5th semester)
6
Type of Psychological
Core Question Daily Life Example
Validity Example
theoretical rate. with extroversion
concept? Discriminant: The same (discriminant).
"effort"
score not correlating with
your phone's brand.
Validity is a matter of degree, not an all-or-none property. We say a test has "strong
evidence for validity" or "limited validity," and this is always in reference to a specific purpose
and population. A test valid for screening depression in adults may not be valid for diagnosing it
in children.
Psychological assessment (Bs 5th semester)