0% found this document useful (0 votes)
63 views8 pages

Validity: Types and Importance

The document discusses different types of validity for psychological tests, including content validity, criterion-related validity (which includes concurrent and predictive validity), and construct validity. It provides definitions and examples for each type. Content validity ensures test items adequately represent the domain being measured. Criterion-related validity examines the relationship between test scores and external outcomes. Concurrent validity predicts performance simultaneously, while predictive validity forecasts future performance.

Uploaded by

Tanya
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
63 views8 pages

Validity: Types and Importance

The document discusses different types of validity for psychological tests, including content validity, criterion-related validity (which includes concurrent and predictive validity), and construct validity. It provides definitions and examples for each type. Content validity ensures test items adequately represent the domain being measured. Criterion-related validity examines the relationship between test scores and external outcomes. Concurrent validity predicts performance simultaneously, while predictive validity forecasts future performance.

Uploaded by

Tanya
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

VALIDITY

Definition of validity, paraphrased from the influential Standards for Educational and
Psychological Testing (AERA, APA, & NCME, 1999):

A test is valid to the extent that inferences made from it are appropriate, meaningful, and
useful.

The validity of any measuring instruments depends upon the accuracy with which it
measures what it to be measured when compared with standard criterion. A test is valid
when the performance which it measures corresponds to the same performance as
otherwise independently measured or objectively defined.

The reliability of a test is determined by making reported measurements of the same facts
and validity is found by comparing the data obtained from the test with standard (and
sometimes arbitrary) measures. Since independent standards (that is, criteria) are hard to
get in mental measurement, the validity of a mental test can never be estimated as
accurately as can the validity of a physical instrument.

Validity is a relative term. A test is valid for a particular purpose; it is not generally valid.

Validity reflects an evolutionary, research based judgement of how adequately a test


measures the attribute it was designed to measure. Consequently, the validity of tests is not
easily captured by neat statistical summaries but is instead characterized on a continuum
ranging from weak to acceptable to strong.

Traditionally, the different ways of accumulating validity evidence have been grouped into
three categories:

• Content validity • Criterion-related validity • Construct validity

CONTENT VALIDITY

The concept of ‘content validity’ is employed in the selection of items for a test. Standard
educational achievement examination represents the consensus of many educators as to
what should a child of a given age or grade know about arithmetic, reading, spelling, history
and other subjects. A test of English history, for instance, would be valid if its content
consists of questions covering this area. The validation of content through competent
judgements is most satisfying under two conditions, (a) when the sampling of items is wide
and judicious and (b) when adequate standardisation groups are utilised.

Content validity is determined by the degree to which the questions, tasks, or items on a test
are representative of the universe of behavior the test was designed to sample. In theory,
content validity is really nothing more than a sampling issue (Bausell, 1986). The items of a
test can be visualized as a sample drawn from a larger population of potential items that
define what the researcher really wishes to measure. If the sample (specific items on the
test) is representative of the population (all possible items), then the test possesses content
validity. Content validity is a useful concept when a great deal is known about the variable
that the researcher wishes to measure. With achievement tests in particular, it is often
possible to specify the relevant universe of behaviors in advance.
Content validity is more difficult to assure when the test measures an ill-defined trait.

Evaluating content validity is carried out by either subjective or empirical methods.


Subjective methods typically involve asking experts to judge the relevance and
representativeness of the test items with regard to the domain being assessed (e.g.,
Hambleton, 1984). Empirical methods involve factor analysis or other advancedstatistical
procedures designed to show that the obtained factors or dimensions correspond to the
content domain (e.g., Davison, 1985). Not only should the test adequately cover the contents
of the domain being measured,but decisions must also be made about the relative
representation of specific aspects

Martuza (1977) and others have discussed statistical methods for determining the overall
content validity of a test from the judgments of experts. These methods tend to be very
specialized and have not been widely accepted. Nonetheless, their approaches can serve as
a model for a commonsense viewpoint on interrater agreement as a basis for content
validity. When two expert judges evaluate individual items of a test on the four-point scale
proposed in Figure 4.1, the ratings of each judge on each item can be dichotomized into
weak relevance (ratings of 1 or 2) versus strong relevance (ratings of 3 or 4). For each item,
then, the conjoint ratings of the two judges can be entered into the two-by-two agreement
table.

Messick (1989) suggests that content validity be discussed in terms of content relevance
and content coverage rather than as a category of validity, but his suggestion has not been
widely accepted as yet.

FACE VALIDITY
A test has face validity if it looks valid to test users, examiners, and especially the
examinees. Face validity is really a matter of social acceptability and not a technical form of
validity in the same category as content, criterion-related, or construct validity (Nevo, 1985).
From a public relations standpoint, it is crucial that tests possess face validity—otherwise
those who take the tests may be dissatisfied and doubt the value of psychological testing.
However, face validity should not be confused with objective validity, which is determined by
the relationship of test scores to other sources of information. In fact, a test could possess
extremely strong face validity—the items might look highly relevant to what is presumably
measured by the instrument—yet produce totally meaningless scores with no predictive
utility whatever.

A tear may have a great deal of face validity yet may not in fact be valid. Conversely, a test
may lack face validity but in reality be a valid measure of a particular variable. Clearly, face
validity is related to client rapport and cooperation,because ordinarily, a test that looks valid
will be considered by the client more appropriate and therefore taken more seriously than
one that does not. There are occasions, however, where face validity may not be desirable,
for example, in a test to detect “honesty” (see Nevo, 1985, for a review).

CRITERION RELATED VALIDITY

Criterion-related validity is demonstrated when a test is shown to be effective in estimating


an examinee’s performance on some outcome measure. In this context, the variable of
primary interest is the outcome measure, called a criterion. The test score is useful only
insofar as it provides a basis for accurate prediction of the criterion.

Experimentally, the validity of a test determined by finding the correlation between the test
and some independent criterion, may be an objective measure of performance, or a
quantitative measure such as a judgement of character or excellence in work done.
Intelligence tests were first to be validated against school grades/ratings for aptitude by
teachers, and other indices of ability. Personality, attitude and interest inventories are
validated in a variety of ways. The best way to check test prediction is evidence of validity,
provided that (a) the criterion was setup independently and (b) both the test and the criterion
are reliable. Criterion validity can be categorised into two types, that is, concurrent and
predictive. Concurrent validity involves prediction of an alternative method of measuring the
same characteristics of interest, while predictive validity attempts to show a relationship with
future behaviour. Both predictive and concurrent validities are accepted by deciding the
appropriate level of validity coeffi cient or correlation between a test score and some criterion
variable. The appropriate acceptance level depends upon the intended use of the test.

Characteristics of a good criterion

As noted, a criterion is any outcome measure against which a test is validated. In practical
terms, a criterion can be most anything. The choice of criteria is circumscribed, in part, by
the ingenuity of the test developer. However, criteria must be more than just imaginative;
they must also be reliable, appropriate, and free of contamination from the test itself. The
criterion must itself be reliable if it is to be a useful index of what the test measures. If you
recall the meaning of reliability—consistency of scores— the need for a reliable criterion
measure is intuitively obvious. After all, unreliable means unpredictable. An unreliable
criterion will be inherently unpredictable, regardless of the merits of the test.

A criterion measure must also be appropriate for the test under investigation. The Standards
for Educational and Psychological Testing sourcebook (AERA, APA, & NCME, 1985)
incorporates this important point as a separate standard: All criterion measures should be
described accurately, and the rationale for choosing them as relevant criteria should be
made explicit.

A criterion must also be free of contamination from the test itself.


If the screening test contains the same as the criterion, then the correlation between these
two measures will be artificially inflated. This potential source of error in test validation is
referred to as criterion contamination, since the criterion is “contaminated” by its artificial
commonality with the test. Criterion contamination is also possible when the criterion
consists of ratings from experts. If the experts also possess knowledge of the examinees’
test scores, this information may (consciously or unconsciously) influence their ratings.
When validating a test against a criterion of expert ratings, the test scores must be held in
strictest confidence until the ratings have been collected.

Statistical measure of validity

The validity coefficient is always less than or equal to the square root of the test reliability
multiplied by the criterion reliability. In other words, to the extent that the reliability of either
the test or the criterion (or both) is low, the validity coefficient is also diminished

CONCURRENT and PREDICTIVE VALIDITY

Two different approaches to validity evidence are subsumed under the heading of
criterion-related validity.

In concurrent validity, the criterion measures are obtained at approximately the same time
as the test scores.
In a concurrent validation study, test scores and criterion information are obtained
simultaneously. Concurrent evidence of test validity is usually desirable for achievement
tests, tests used for licensing or certification, and diagnostic clinical tests. An evaluation of
concurrent validity indicates the extent to which test scores accurately estimate an
individual’s present position on the relevant criterion. A test with demonstrated concurrent
validity provides a shortcut for obtaining information that might otherwise require the
extended investment of professional time. For example, the case assignment procedure in a
mental health clinic can be expedited if a test with demonstrated concurrent validity is used
for initial screening decisions. New and existing tests are often cited as evidence of
concurrent validity. This has a catch-22 quality to it—old tests validating a new test—but is
nonetheless appropriate if two conditions are met. First, the criterion (existing) tests must
have been validated through correlations with appropriate nontest behavioral data. In other
words, the network of interlocking relationships must touch ground with real-world behavior
at some point. Second, the instrument being validated must measure the same construct as
the criterion tests.

In predictive validity, the criterion measures are obtained in the future, usually months or
years after the test scores are obtained, as with the college grades predicted from an
entrance exam.
Predictive validity measures are used to estimate outcome measures obtained at a later
date. Predictive validity is particularly relevant for entrance examinations and employment
tests. Such tests share a common function—determining who is likely to succeed at a future
endeavor.

CONSTRUCT VALIDITY

A construct is a theoretical, intangible quality or trait in which individuals differ (Messick,


1995). Examples of constructs include leadership ability.
Constructs are inferred from behavior but are more than the behavior itself. In general,
constructs are theorized to have some form of independent existence and to exert broad but
to some extent predictable influences on human behavior. A test designed to measure a
construct must estimate the existence of an inferred, underlying characteristic (e.g.,
leadership ability) based on a limited sample of behavior. Construct validity refers to the
appropriateness of these inferences about the underlying construct. All psychological
constructs possess two characteristics in common: 1. There is no single external referent
sufficient to validate the existence of the construct; that is, the construct cannot be
operationally defined (Cronbach & Meehl, 1955). 2. Nonetheless, a network of interlocking
suppositions can be derived from existing theory about the construct (AERA, APA, & NCME,
1985). Construct validity pertains to psychological tests that claim to measure complex,
multifaceted, and theory-bound psychological attributes such as psychopathy, intelligence,
leadership ability, and the like. The crucial point to understand about construct validity is that
“no criterion or universe of content is accepted as entirely adequate to define the quality to
be measured” (Cronbach & Meehl, 1955). Thus, the demonstration of construct validity
always rests on a program of research using diverse procedures outlined in the following
sections. To evaluate the construct validity of a test, we must amass a variety of evidence
from numerous sources.
Construct validity is an umbrella term that encompasses any information about a particular
test; both content and criterion validity can be subsumed under this broad term. What makes
construct validity different is that the validity information obtained must occur within a
theoretical framework. If we wish to validate a test of intelligence, we must be able to specify
in a theoretical manner what intelligence is, and we must be able to hypothesize specific
outcomes.

Construct validity approach is much more complex than other forms of validity and is based
on the accumulation of data over a long period of time. Construct validity requires the study
of test scores in relation not only to variables that the test is intended to assess, but also in
relation to the study of those variables that have no relationship to the domain underlying the
instrument. Therefore, one builds a homothetic net or inferential definition of the
characteristics that a test is intended to assess. Another approach includes predications to
other tests that are tests which are assumed to measure the same underlying trait as well as
those that describe unrelated traits. Hence, we may find or predict that a specific intellectual
skill should have a moderate correlation with the test of general Intelligence Quotient (IQ),
little or no correlation with a measure of hypochondrias, and a strong correlation to another
test assessing the same intellectual skill. One should keep in mind the accuracy of the
original hypothesis. This hypothesis is related with the researcher’s comprehension of the
traits under study. One should be careful and not confuse a researcher’s misunderstanding
of either the intention of an instrument or the underplaying theory with the inefficiency of the
instrument itself.

Convergent and discriminant Validatiy

D. P. Campbell and Fiske (1959) and D. P. Campbell (1960) proposed that to show construct
validity, one must show that a particular test correlates highly with variables, which on the
basis of theory, it ought to correlate with; they called this convergent validity. They also
argued that a test should not correlate significantly with variables that it ought not to
correlate with, and called this discriminant validity. They then proposed an experimental
design called the multitraitmultimethod matrix to assess both convergent and discriminant
validity.

Convergent validity is demonstrated when a test correlates highly with other variables or
tests with which it shares an overlap of constructs. For example, two tests designed to
measure different types of intelligence should, nonetheless, share enough of the general
factor in intelligence to produce a hefty correlation (say, .5 or above) when jointly
administered to a heterogeneous sample of subjects. In fact, any new test of intelligence that
did not correlate at least modestly with existing measures would be highly suspect, on the
grounds that it did not possess convergent validity. Discriminant validity is demonstrated
when a test does not correlate with variables or tests from which it should differ. For
example, social interest and intelligence are theoretically unrelated, and tests of these two
constructs should correlate negligibly, if at all.

example- Suppose we have a true-false inventory of depression that we wish to validate.


Weneedfirst of all to find a second measure of depression that does not use a true-false or
similar format– perhaps a physiological measure or a 10-point psychiatric diagnostic scale.
Next, we need to find a different dimension than depression, which our theory suggests
should not correlate but might be confused with depression, for example, anxiety. We now
locate two measures of anxiety that use the same format as our two measures of
depression. We administer all four tests to a group of subjects and correlate every measure
with every other measure. To show convergent validity, we would expect our two measures
of depression to correlate highly with each other (same trait but different methods). To show
discriminant validity we would expect our true-false measure of depression not to correlate
significantly with the true-false measure of anxiety (different traits but same method). Thus
the relationship within a trait, regardless of method, should be higher than the relationship
across traits. If it is not, it may well be that test scores reflect the method more than anything
else
FACTORS AFFECTING VALIDITY

Validity of a psychological test may be influenced by a number of factors, some of which are
listed here. Group Differences The characteristics of a group of people on whom the test is
validated affect the criterion related validity. Differences among the group of people on
variables like sex, age and personality traits may affect the correlation coeffi cient between
the test and the selected criteria. Like reliability coefficient, the magnitude of the validity
coefficient depends on the degree of heterogeneity of the validation group on the test
variable. In a group having narrower range of test scores, that is, in a more homogeneous
group, the validity coeffi cient tends to be smaller. Since the size of a correlation coeffi cient
is a function of two variables, a narrowing of the range of the either of the predictors of the
criterion variable will tend to lower the validity coefficient.

Correction for Attenuation

Criterion Contamination

The validity of a test is also dependent upon the validity of the criterion itself as a measure of
the particular cognitive or affective characteristic of interest. Sometimes, the criterion is
contaminated or rendered invalid due to the method by which the criterion scores are
determined. Teachers have been known to test students’ scores on academic achievements
tests (AATs) before deciding what course grades to assign. Since AATs are also taken into
consideration by the admission offi ce to select students who are predicted to make
satisfactory grades. This method of assigning grades contaminates the criterion and, hence,
results in an inaccurate validity coeffi cient. Therefore, if AAT scores are to be used for
predicting grades, then grades should be arrived at independently without reference to AAT
scores.

RELIABILITY AND VALIDITY

Reliability and validity refer to different aspects of essentially the same thing, namely, test effi
ciency. Reliability is concerned with the stability of test scores. It does not go beyond the test
itself. Validity, on the other hand, implies evaluation in terms of an outside and independent
criteria. The purpose of a test is to fi nd a measure which will be an adequate and time
saving substitute for criterion measures, obtainable only after long intervals of time, for
example, school grades or performance records. In case of psychological tests, high
reliability and high validity is desirable. Or, we can say that it is the interaction between
reliability and validity that determines the desirability of a psychological test. For example,
see Figure 10.1. The most desirable test will be the test that carries a high reliability as well
as validity.

To be valid a test must be reliable. A highly reliable test is always a valid measure of some
function. To explain this further, if a test has a reliability coeffi cient of 0.81 and its index of
reliability is 0.90, then the test correlates 0.90 with the true measures that constitute the
criterion. However, a test may be theoretically valid and show little or no correlation with
anything else. For example, word cancellation test scores can be made highly reliable by
lengthening or repeating the test, so that the index of reliability becomes high. But the
correlation of these tests with such criteria as speed or accuracy are so low that they show
little practical validity.

You might also like