0% found this document useful (0 votes)
11 views5 pages

Understanding Test Validity in Education

Uploaded by

2257010228
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
11 views5 pages

Understanding Test Validity in Education

Uploaded by

2257010228
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

VALIDITY

By far the most complex criterion of an effective test—and arguably the most impor tant
principle—is validity, "the extent to which inferences made from assessment results are
appropriate, meaningful, and useful in terms of the purpose of the assess-ment" (Gronlund,
1998, p. 226). A valid test of reading ability actually measures reading ability—not 20/20
vision, nor previous knowledge in a subject, nor some other variable of questionable
relevance. To measure writing ability, one might ask students to write as many words as
they can in 15 minutes, then simply count the words for the final score. Such a test would be
easy to administer (practical), and the scoring quite dependable (reliable). But it would not
constitute a valid test of writing ability without some consideration of comprehensibility,
rhetorical discourse elements, and the organization of ideas, among other factors.
How is the validity of a test established? There is no final, absolute measure of validity, but
several different kinds of evidence may be invoked in support. In some cases, it may be
appropriate to examine the extent to which a test calls for performance that matches that of
the course or unit of study being tested. In other cases, we may be concerned with how well
a test determines whether or not students have reached an established set of goals or level
of competence. Statistical correlation with other related but independent measures is
another widely accepted form of evi-dence. Other concerns about a test's validity may focus
on the consequences— beyond measuring the criteria themselves-of a test, or even on the
test-taker's perception of validity. We will look at these five types of evidence below.
Content-Related Evidence
If a test actually samples the subject matter about which conclusions are to be drawn, and if
it requires the test-taker to perform the behavior that is being measured, it can claim content-
related evidence of validity, often popularly referred to as content validity (e.g., Mousavi,
2002; Hughes, 2003). You can usually identify content-related evidence observationally if
you can clearly define the achievement that you are measuring. A test of [tennis competency
[hat asks someone to run a tOO-yard dash obviously lacks content validity. If you are trying
to assess a person's ability to speak a second language in a conversational setting, asking
the learner to answer", pap~r-and-pencil multiple-choice questions requiring grammatical
judgments does ) not achieve content validity. A test that requires the learner actually to
speak within ' some sort of authentic context does. And ifa course has perhaps ten
objectives but only two are covered in a test, then content validity suffers. Consider the
following quiz on English articles for a high-beginner level of a conversation class (listening
and speaking) for English learners.
English articles quiz Directions: The purpose of this quiz is for you and me to find out how
well you know and can apply the rules of article usage. Read the-following passage and
write alan, the, or 0 (no article) in each blank. Last night, I had (1) a very strange dream.
Actually, it was (2)__ nightmare! You know how much I love (3) went to (4) San Francisco
zoo with (5) there, it was very dark, but (6) wanted to see (7) round and (9) zoos. Well, I
dreamt that I had a few friends. When we got the moon was:out, so we weren't afraid. I was
a monkey first, so we walked past (8) lions' cages to (10) monkey section. (The story
continues, with a total of 25 blanks to fill.)
The students had had a unit on zoo animals and had engaged in some open discus sions
and group work in which they had practiced articles, all in listening and speaking modes of
performance. In that this quiz uses a familiar setting and focuses on previouslrpracticed
language forms';-it-is--somewhat'contentvalid. The fact that it was administered in written
form, however, and required students to read the pas sage and write their responses makes
it quite low in content validity for a lis tening/speaking class. There are a few cases of highly
specialized and sophisticated testing instru ments that may have questionable content-
related evidence of validity. It is possible to contend,for example, that standard language
profici~ncy\tests, with their context reduced, academically oriented language and limited
stretches of discourse, l~ck cogtent vaJigity since they do not require the full spectrum of
communicative per formance on the part of the learner (see Bachman, 1990, for a full
discussion). ll1ere is good reasoning behind such criticism; nevertheless, what such
proficiency tests lack in content-related evidence they may gain in other forms of evidence,
not ,to mention practicality and reliability. Another way of understanding content validity is to
consider the difference between direct and indirect testing. Direct .testing involves the test-
taker in actu ally perfonning the target task. In an indirect test, learners are not performing
the task itself but rather a task that is related in some way. For example, ifyou intend to test
learners' oral production of syllable stress and your test task is to have learners mark (with
written accent marks) stressed syllables in a list of written words, you could, with a stretch of
logic, argue that you are indirectly testing their oral pro duction. A direct test of syllable
production would have to require that students actually produce. target_ words orally. The
mo~t feasible rule of thumb for achieving content Validity in classroom assessment is to test
performance directly. ConSider, for example, a listening! speaking class that is doing a unit
on greetings and exchanges that includes dis course for asking for personal information
(name, address, hobbies, etc.) with some form-focus on the verb to be, personal pronouns,
and question formation. The test on that unit should include all ofthe above discourse and
grammatical elements and involve students in the actual performance of listening and
speaking. What all the above examples suggest is that content is not the only type of evi
dence to-support the validityofa test,but classroom teachers have neither the·time nor the
budget to subject quizzes, midterms, and final exams to the extensive scrutiny of a full
construct valida~ion (see below). Therefore, it is critical that teachers hold content-related
evidence in high esteem in the process of defending the validity ofclassroom tests. Criterion-
Related Evidence A second form of evidence of the validity of a test may be found in what is
called criterion-related evidence, also referred to as criterion-related validity, or the extent to
which the "criterion" of the test has actually been [Link] will recall that in Chapter 1 it
was noted that most classroom-based assessment with teacher designed tests fits the
concept of criterion-referenced assessment. In such tests, \.. specified classroom
objectives are measured, and implied predetermined levels of performance are expected to
be·-·reached (SO·percent is considereda··minimal passing grade). In the case of teacher-
made classroom assessments, criterion-related evidence is best demonstrated through
a.,comparison of results of an assessment with results of some other measure of the same·
criterion. For example, in a course unit whose objective is for students to be able to orally
produce voiced and voiceless stops in all possible phonetic environments, the results of one
teacher's unit test might be compared with an independent assessment-possibly a
commercially produced test in a textbook-of the same phonemic profiCiency. A classroom
test designed to assess mastery of a point of grammar in communicative use will have
criterion validity if test scores are corroborated either by observed subsequent behavior or by
other communicative measures of the grammar point in question. Criterion-related evidence
usually falls into one of two categories: concurrent and predictive validity. A test has
concurrent validity ifits results are supported by other concurrent perfonnance beyond the
assessment itself. For example, the validity of a high score on the final exam of a foreign
language course will be substatitiated by aCtlJal proficiency in the language. The predictive
validity of an assessment
becomes"mportant in the case of placement tests, admissions assessment batteries,
language aptitude tests, and the like. The assessment criterion in such cases is not
to measure concurrent ability but to assess (and predict) a test-taker's likelihood of
future success.
','!'
"
,."
Construct..Related .Evidence
A third kind of evidence that can support validity, but one that does not playas large
a role for classroom teachers, is ,construct-related validity, commonly referred to as
construct Validity. A construct is any theory, hypothesis, or model that attempts to v'
explain [Link] in our universe of perceptions. Constructs mayor
may not be directly or empirically measured-their verification often requires infer
ential data. "Proficiency" and "communicative compet(!nce" are linguistic constructs;
"self-esieem" and "motiVation" are psychological constructs. VIrtUally every issue in
language learning and teaching involves theoretical constructs. In the field of assess
ment, construct validity ask.$, "Does this test actually tap into the theoretical con
struct as it has been defined?" Tests are, in a manner of speaking, 'operational t··
det,lnitions of constructs in that theyoperationalizethe entity that is being mea
sure<C(see Davidson, Hudson, & Lynch, 19B5)':"'~-"-"
. "'--'~-"'~~'-For most of the tests that you administer as a classroom teacher, a formal con
struct validation procedure may seem a daunting prospect. You will be tempted, per
haps, to run a quick content check and be satisfied with the test's validity. But don't
let the concept of construct validity scare you. An informal construct validation of
the use of virtually every classroom test is both essential and feasible.
Imagine, for example, that you have been given a procedure for conducting an
oral interview. The scoring analysis for the interview includes seyc;~l factors in the
f~al score:pronunciation~-fluenCY;··gramm~ticar'accu~cy,voc~.btit~ry u~¢,jindsocio
nfliUistic approprlateness':The' Jus@cafioIi'1oi-fliese'flve'factors lies in a theore'ticru
",~Q.P~,!~ct that claims those factors to be major components-of oral proficiency. So
ifyou were'asked to conduct an oral proficiency interview that evaluated only pro
nunciation and grammar, you could be justifiably suspicious about the construct
validity of that test. Likewise, let's suppose you have created a simple written vocab
ulary quiz, covering the content of a recent unit, that asks students to correctly
defme a set ofwords. Your chosen items may be a perfectly adequate sample ofwhat
was covered in the unit, but if the lexical objective of the unit waS the communica
tive use of vocabulary, then the writing of definitions certainly fails to match a construct of
communication language use. ,
Construct validity is a major issue in validating large-scale standardized tests of
proficiency. Because such tests must, for economic reasons, adhere to the principle
of practicality, and because they must sample a limited number of domains of language, they
may not be able to contain all the content of a particular field or skill. The
TOEFL®, for example, has until recently not attempted to sample oral production, yet
oral production is obviously an important part of academic success in a university course of
study. The TOEFL's omission of oral production content, however, is osten· sibly justified by
research that has shown positive correlations between oral production and the behaviors
(listening, reading, grammaticality detection, and writing) actually sampled on the TOEFL
(see Duran et al., 1985). Because of the crucial need to offer a fmancially affordable
proficiency test and the high cost of administering and scoring oral -production tests, the
omission of oral content from the TOEFL has been justified;1. as an economic necessity.
(Note: As this book goes to press, oral production tasks are being included in the TOEFL,
largely stemming from the demands of the professional community for authenticity and
content validity.) Consequential Validity As well as the above three widely accepted· forms of
evidence that may be introduced to support the validity of an assessment, two other
categories may be of some interest and utility in your own:quest for validating classroom
tests. Messick EI989), . Gronlund (1998), McNamara (2000), and Brindley (2001), among
others, underscore the potential importance of the consequences of using an assessment.
Consequential validity encompasses all the consequences of a test, including such consid \'
erations as its accuracy in measuring intended criteria, its. impact on the preparation of test-
takers, its effect on the learner, and the (intended and unintended) social consequences of a
test's interpretation and use. :As high-stakes assessment has gained ground in the last two
decades, one aspect of consequential validity has drawn special attention: the effect of test
preparation courses and manuals on performance. McNamara (2000, p. 54) cautions against
test results that may reflect socioeconomic conditions such as opportunities for coaching that
are "differentially available to the students being assessed (for example, because only some
families can afford coaching, or because children with more highly educated parents get help
from their parents)." The social consequences of large-scale, high-stakes assessment are
discussed in Chapter 6. Another important consequence of a test falls into the category of
washback, to be more fully discussed below. Gronlund (1998, pp. 209-210) encourages
teachers to consider the ~[Link] [Link],ld~A~'mQtiyation~ subsequent
p~rformance in a course, independent learning, study habits, and attitude toward school
work. Face Validity An important facet of consequential validity is the extent to which
"students view the assessment as fair, relevant, and useful for improving learning"
(Gronlund, 1998, p. 210), or what is popularly known as face validity. "Face validity refers to
the . degree to which a test looks right, and appears to measure the knowledge or abilities it
claims to measure, based on the subjective judgment of the examinees who take it, the
administrative personnel who decide on its use, and other psychometrically unsophisticated
observers" (Mousavi, 2002, p. 244). Sometimes students don't know what is being tested
when they tackle a test. They may feel, for a variety of reasons, that a test isn't testing what
it is "supposed" to test. Face validity means that the students perceive the test to be valid.
Face validity asks the question "Does the test, on the 'face' of it, appear from the leamer's
perspective to test what it is designed to test?" Face validity will likely be high if learners
encounter • a well-constructed, expected format with familiar tasks, • a test that is clearly
doable within the allotted time limit, • items that are clear and uncomplicated, • directions that
are crystal clear, • tasks that relate to their course work (content validity), and • a difficulty
level that presents a reasonable challenge. \./ Remember, face validity is not something that
can be empirically tested by a teacher or even by a testing expert. It is purely a factor of the
"·eye of the beholder"-how the test-taker, or possibly the test giver, intuitively perceives the
instrument. For this reason, some assessment experts (see Stevenson, 1985) view face
validity as a superficial factor that is dependent on the whim of the perceiver. The other side
of this issue reminds us that the psychological state of the learner (confidence, anxiety, etc.)
is an important ingredient in peak performance by a learner. Students can be distracted and
their anxiety increased ifyou "throw a curve" at them on a test. They need to have rehearsed
test tasks before the fact and feel comfortable with them. A classroom test is not the time to
introduce new tasks because you won't know ,if student difficulty is a factor of the task itself
or of the objectives you are testing. I onc<; administered a dictation test and a cloze test
(see Chapter 8 for a dis cuSsion of cloze tests) as a placement test for a group of learners of
English as a second language. Some learners were upset because such tests, on the face of
it, did not appear to them to test their true abilities in English. They felt that a multiple choice
grammar test would have [Link] appropriate format to use. A few claimed they didn't
perform well on the cloze and dictation because they were not accus tomed to these
formats. As it turned out, the tests served as superior instruments for placement, but the
students would not have thought so. Face validity was low, con tent validity was moderate,
and construct validity was very high. As already noted above, content validity is a very
imporcint ingredient in achieving face validity. If a test samples the actual content of what the
learner has achieved or expects to achieve, then face validity will be more likely to be
perceived. Validity is a complex concept, yet it is indispensable to the teacher's under
standing of what makes a good test. If in your language teaching you can attend to the
practicality, reliability, and validity of tests of language, whether those tests are classroom
tests related to a part of a lesson; fmal exams, or profiCiency tests, then you are well on the
way to making accurate judgments about the competence of the learners with whom you are
working.

You might also like