PSYCHOLOGICAL ASSESSMENT
PSYCHOMETRIC PROPERTIES OF A GOOD TEST: 1. Content validity.
VALIDITY o This is a measure of validity based on an
evaluation of the SUBJECTS, TOPICS, or
What is VALIDITY? CONTENT covered by the items in the test.
• a JUDGMENT or ESTIMATE of how well a test 2. Criterion-related validity.
measures what it purports to measure in a particular o This is a measure of validity obtained by
test evaluating the RELATIONSHIP OF SCORES
• in the language of psychological assessment, obtained on the test to scores on other tests or
validity is a term used in conjunction with the measures
MEANINGFULNESS of a test score—what the test 3. Construct validity.
score truly means. o This is a measure of validity that is arrived at by
▪ LEGITIMACY of the test executing a comprehensive analysis of
▪ There is BASIS/ EVIDENCE 1. how scores on the test RELATE TO
INFERENCE OTHER TEST SCORES and MEASURES,
• LOGICAL RESULT or DEDUCTION and
• Characterizations of the validity of tests and test 2. how scores on the test can be
scores are frequently phrased in terms such as UNDERSTOOD WITHIN SOME
“acceptable” or “weak.” THEORETICAL FRAMEWORK for
o These terms reflect a judgment about how understanding the construct that the test
adequately the test measures what it purports to was designed to measure.
measure.
• As a shorthand, assessors may refer to a particular TRINITARIAN VIEW
test as a “valid test.” • In this classic conception of validity, referred to as
• No test or measurement technique is “universally the trinitarian view (Guion, 1980)
valid” for all time, for all uses, with all types • it might be useful to visualize CONSTRUCT
of testtaker populations. VALIDITY as being “umbrella validity” because
• tests may be shown to be valid within what we would every other variety of validity falls under it.
characterize as REASONABLE o Why construct validity is the overriding variety of
BOUNDARIES of a contemplated usage. validity will become clear as we discuss what
makes a test valid and the methods and
VALIDATION procedures used in validation.
• process of GATHERING and EVALUATING • Trinitarian approaches to validity assessment are not
EVIDENCE about validity. mutually exclusive.
• Both the TEST DEVELOPER and the TEST USER • The trinitarian model of validity is not without its
may play a role in the validation of a test for a critics (Landy, 1986).
specific purpose
o It is the TEST DEVELOPER’S responsibility to TYPES OF VALIDITY
supply validity evidence in the test manual.
o It may sometimes be appropriate for TEST Face Validity
USERS to conduct their own validation studies • what a test appears to measure to the PERSON
with their own groups of testtakers. BEING TESTED than what the test actually means
• LOCAL VALIDATION STUDIES • a judgement concerning how relevant the test
o may yield insights regarding a particular items appear to be
population of testtakers as compared to the • the LEAST stringent type of validity, whether a
norming sample described in a test manual. test looks valid to test users, examiners and
o are absolutely necessary when the test user examinees
plans to alter in some way the FORMAT, • if a test definitely appears to measure what it
INSTRUCTIONS, LANGUAGE, or CONTENT of purports to measure “on the face of it,” then it could
the test. be said to be high in face validity.
• judgments about face validity are frequently thought
One way measurement specialists have traditionally of from the perspective of the testtaker, NOT the
conceptualized validity is according to THREE test user.
• a test that lacks face validity may still be RELEVANT
CATEGORIES: and USEFUL.
1
PSYCHOLOGICAL ASSESSMENT
o However, if the test is not perceived as relevant • Uncontaminated: Criterion contamination
and useful by TESTTAKERS, PARENTS, occurs if the criterion based on PREDICTOR
LEGISLATORS, and OTHERS, then negative MEASURES; the criterion used is a criterion
consequences may result. of WHAT IS SUPPOSED TO BE THE
• face validity may be MORE A MATTER OF PUBLIC CRITERION
RELATIONS than psychometric soundness. • indicates the TEST EFFECTIVENESS in
o Still, it is important nonetheless, and deserving of estimating an individual’s behavior in a particular
respect. situation
▪ Basically the first impression. • tells HOW WELL A TEST CORRESPONDS with
▪ Example: Exam is based on = napag-aralan a particular criterion
ba or hindi • judgment of how adequately a test score can be
▪ Example in real life: Before ordering to used to infer an individual’s MOST
Shopee, you are checking the reviews first. PROBABLE STANDING on some measure of
interest
Content Validity
• describes a judgement of how adequately a test Types of Criterion-Related Validity
sample behavior representative of the universe of Concurrent Validity
behavior that the test was DESIGNED TO SAMPLE • an index of the degree to which a TEST SCORE IS
• Issues arising from lack of content validity: RELATED TO SOME CRITERION measure
o Construct underrepresentation-Failure to obtained at the SAME TIME
capture IMPORTANT COMPONENTS of a o indicate the extent to which test scores may be
construct used to estimate an INDIVIDUAL’S PRESENT
o Construct-irrelevant variance- Happens when STANDING on a criterion.
scores are INFLUENCED BY FACTORS o Example: group of researchers explored
IRRELEVANT to the construct whether a test validated for use with ADULTS
could be used with ADOLESCENTS.
The quantification of content validity - The Beck Depression Inventory
• The measurement of content validity is important in (BDI; Beck et al., 1961, 1979; Beck &
EMPLOYMENT SETTINGS, where tests used to hire Steer, 1993),
and promote people are carefully scrutinized for their - REVISION; Beck Depression
relevance to the job, among other factors (Russell & Inventory-II (BDI-II; Beck et al., 1996)
Peterson, 1997). are self-report measures used to
Culture and the relativity of content validity identify symptoms of DEPRESSION
• Tests are often thought of as either VALID or NOT and quantify their SEVERITY.
VALID. - Although the BDI had been widely
• Example: A history test, either does or does not used with adults, questions were raised
accurately measure one’s knowledge of historical regarding its appropriateness for use
fact. with adolescents.
Cultural Relativity, History, and Test Validity - The findings suggested that the BDI
• POLITICS is another factor that may well play a part is valid for use with adolescents.
in perceptions and judgments concerning the
validity of tests and test items. Predictive Validity
• an index of the degree to which a test score
Criterion-Related Validity PREDICTS some criterion measure
• What is criterion? Example: entrance exam in college ( 1
• STANDARD against which a test or a test score test)
is EVALUATED - Through this, we will predict who will
• Characteristics of a criterion: excel in college.
• Relevant • Judgments of criterion-related validity, whether
- it is pertinent or applicable to the concurrent or predictive, are based on TWO TYPES
matter at hand OF STATISTICAL EVIDENCE: the validity
• Valid and Reliable coefficient and expectancy data.
- Valid- for the purpose for which it is
being used • VALIDITY COEFFICIENT
o CORRELATION COEFFICIENT that provides a
measure of the RELATIONSHIP between TEST
2
PSYCHOLOGICAL ASSESSMENT
SCORES AND SCORES on the criterion • An informed scientific idea DEVELOPED or
measure. HYPOTHESIZED to describe or explain a
• Incremental Validity – this type of validity is behavior; something built by MENTAL
related to predictive validity wherein it is defined SYNTHESIS.
as the degree to which an ADDITIONAL • UNOBSERVABLE, PRESUPPOSED TRAITS;
PREDICTOR EXPLAINS SOMETHING about something that the researcher thought to have
the criterion measure that is NOT explained by either HIGH OR LOW CORRELATION with other
predictors already in use variables
• value of including MORE THAN ONE • a judgment about the APPROPRIATENESS of
PREDICTOR depends on a couple of factors. inferences drawn from test scores regarding
individual standings on a variable called construct
• EXPECTANCY DATA • establishing construct validity involves both
o provide information that can be used in LOGICAL ANALYSIS and EMPIRICAL DATA
EVALUATING THE CRITERION-RELATED • Construct validity is like proving a theory through
VALIDITY of a test. evidences and statistical analysis
o an interval that may be seen as “passing,”
“acceptable,” and so on. Evidences of Construct Validity
o EXPECTANCY TABLE • Test is HOMOGENOUS, measuring a single
- shows the percentage of people within construct.
specified test-score intervals who • SUBTEST scores are correlated to the TOTAL
subsequently were placed in various test score.
categories of the criterion • COEFFICIENT ALPHA may be used as
o EXPECTANCY CHART homogeneity evidence.
- or Graphic Representation of an • SPEARMAN RHO can be used to correlate an
expectancy table. item to another item.
• PEARSON or POINT BISERIAL can be used to
Base Rates and Predictive Validity correlate an item to the total test score. (item-
• Base Rate - extent to which a particular trait, total correlation)
behavior, characteristic, or attribute EXISTS in the • TEST SCORE increases or decreases as a function
population of age, passage of time, or experimental
• Hit Rate - proportion of people a test accurately manipulation
IDENTIFIES AS POSSESSING OR EXHIBITING a • Some variable/construct are expected to change
particular trait, behavior, characteristic or attribute with age
• Example: Thinking that someone would failed the • PRETEST, POSTTEST DIFFERENCES
exam, then he did. • Difference of scores from pretest and
• Miss Rate - proportion of people the TEST FAILS posttest of a defined construct after
TO IDENTIFY as having, or not having, a particular careful manipulation would provide
characteristic or attribute validity
• Prediction is wrong/ did not occur • TEST SCORES differ from GROUPS
• False Positive - miss wherein the test predicted that • Also called a method of contrasted group
the testtaker did possess the particular characteristic • T-TEST can be used to test the difference of
or attribute being measured when in fact the groups
testtaker DID NOT • Test scores correlate with scores on other test in
• Example: An individual who take the PT was resulted accordance to what is predicted.
positive, but she is not pregnant. • Discriminant Validation
• False Negative - miss wherein the test predicted • Convergent Validity – a TEST
that the testtaker did not possess the particular CORRELATES highly with other variables
characteristic or attribute being measured when in with which it should correlate
fact the testtaker ACTUALLY DID • Divergent Validity – a TEST DOES NOT
• Example: An individual who take the PT was resulted CORRELATE significantly with variables from
negative, but after 3 months, her tummy gotten which it should differ
bigger. • Factor Analysis - shorthand term for a class
of MATHEMATICAL PROCEDURES
Construct Validity designed to identify factors or specific
• What is a construct? variables that are typically attributes
3
PSYCHOLOGICAL ASSESSMENT
Exploratory Factor Analysis - • Halo Effect – a type of rating error wherein the
ESTIMATING, or extracting factors rater views the object of the rating with
- Summarizing data/ factor EXTREME FAVOUR and tends to bestow
Confirmatory Factor Analysis - degree ratings inflated in a positive direction
to which a HYPOTHETICAL MODEL fits - An individual is pretty, therefore, you
the actual data would also think that she is smart.
- Generalizing data/ factor • Impression Management
• Cross-Validation - REVALIDATION of the - Example: when an man is courting a
test to a criterion based on another group woman, he would not tell his negative
different from the original group from which traits but he would tell his positive
the test was validated traits.
- After quiz = exchange paper = back to • Acquiescence and Non-acquiescence
the oener to check again or to cross - Answerable by YES or NO only. (hindi
validate. gitna, kahit may gitna)
Validity Shrinkage – DECREASE in • Faking-Good and Faking-Bad
validity after cross validation.
Co-validation – validation of more than Test Fairness
one test from the same group. • this is the extent to which a test is used in an
Co-norming – norming more than one IMPARTIAL, JUST and EQUITABLE way
test from the same group
Factors Influencing Test Validity
Test Bias • Appropriateness of the test
• a factor inherent in a test that SYSTEMATICALLY • Directions/Instructions
PREVENTS ACCURATE, IMPARTIAL • Reading Comprehension Level
MEASUREMENT • Item Difficulty
• Rating Error - a judgment resulting from the • Test Construction factors
INTENTIONAL OR UNINTENTIONAL MISUSE of • Length of Test
rating scales • Arrangement of Items
• Severity Error/Strictness Error – LESS • Patterns of Answer
THAN ACCURATE rating or error in
evaluation due to the rater’s tendency to
be overly critical
• Leniency Error/Generosity Error – a
rating error that occurs as a result of a
RATER’S TENDENCY to be too forgiving
and insufficiently critical
- Positive comments
• Central Tendency Error – a type of rating
error wherein the rater exhibits a GENERAL
RELUCTANCE to issue ratings at either a
positive or negative extreme and so all or most
ratings cluster in the MIDDLE of the rating
continuum
- Mid/ average rating
• Proximity Error – rating error committed due
to PROXIMITY/SIMILARITY of the traits being
rated
• Primacy Effect – “FIRST IMPRESSION”
affects the rating
• Contrast Effect – the PRIOR SUBJECT of
assessment affects the LATTER SUBJECT of
assessment
• Recency Effect – tendency to rate a person
based from RECENT RECOLLECTIONS
about that person