0% found this document useful (0 votes)
16 views3 pages

Understanding Test Validity Types

The document discusses different types of validity evidence that can be used to establish the validity of a test: 1. Content-related evidence examines how well a test samples the subject matter and requires test-takers to perform the behaviors being measured. 2. Criterion-related evidence looks at how test results correlate with other measures of the same criteria, through either concurrent or predictive validity. 3. Construct-related evidence examines the theoretical constructs a test aims to measure and how results correlate with these constructs, though this is more relevant for standardized tests than classroom assessments.

Uploaded by

Dina Ahmadi
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
16 views3 pages

Understanding Test Validity Types

The document discusses different types of validity evidence that can be used to establish the validity of a test: 1. Content-related evidence examines how well a test samples the subject matter and requires test-takers to perform the behaviors being measured. 2. Criterion-related evidence looks at how test results correlate with other measures of the same criteria, through either concurrent or predictive validity. 3. Construct-related evidence examines the theoretical constructs a test aims to measure and how results correlate with these constructs, though this is more relevant for standardized tests than classroom assessments.

Uploaded by

Dina Ahmadi
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

VALIDITY

By far the most complex criterion of an effective test-and arguably the most important
principle-is validity, "the extent to which inferences made from assessment results are
appropriate, meaningful and useful in terms of the purpose of the assessment.
How is the validity of a test established? There is no final, absolute measure of
validity, but several different kinds of evidence may be invoked in support. In some cases, it
may be appropriate to examine the extent to which a test calls for performance that matches
that of the course or unit of study being-tested.
There are 5 types of Validity :
1. Content-Related Evidence
If a test actually samples the subject matter about which conclusions are to be drawn,
and if it requires the test-taker to perform the behaviour that is being measured, it can claim
content-related evidence of validity, often popularly referred to as content validity. You can
usually identify content-related evidence observationally if you can clearly define the
achievement that you are measuring.
If you are trying to assess a person's ability to speak a second language in a
conversational setting, asking the learner to answer", paper-and-pencil multiple-choice
questions requiring grammatical judgments does ) not achieve content validity. A test that
requires the learner actually to speak within ' some sort of authentic context does. And if a
course has perhaps ten objectives but only two are covered in a test, then content validity
suffers.
The students had had a unit on zoo animals and had engaged in some open
discussions and group work in which they had practiced articles, all in listening and speaking
modes of performance. In that this quiz uses a familiar setting and focuses on previous
practiced language forms. it-is—somewhat content valid. The fact that it was administered in
written form, however, and required students to read the passage and write their responses
makes it quite low in content validity for a listening/speaking class.
Another way of understanding content validity is to consider the difference between
direct and indirect testing. Direct .testing involves the test-taker in actually performing the
target task. In an indirect test, learners are not performing the task itself but rather a task that
is related in some way. For example, if you intend to test learners' oral production of syllable
stress and your test task is to have learners mark (with written accent marks) stressed
syllables in a list of written words, you could, with a stretch of logic, argue that you are
indirectly testing their oral production. A direct test of syllable production would have to
require that students actually produce. target_ words orally.
2. Criterion-Related Evidence
A second form of evidence of the validity of a test may be found in what is called
criterion-related evidence, also referred to as criterion-related validity, or the extent to which
the "criterion" of the test has actually been reached.
In the case of teacher-made classroom assessments, criterion-related evidence is best
demonstrated through comparison of results of an assessment with results of some other
measure of the same· criterion. For example, in a course unit whose objective is for students
to be able to orally produce voiced and voiceless stops in all possible phonetic environments,
the results of one teacher's unit test might be compared with an independent assessment.
Criterion-related evidence usually falls into one of two categories: concurrent and
predictive validity. A test has concurrent validity if its results are supported by other
concurrent performance beyond the assessment itself. For example, the validity of a high
score on the final exam of a foreign language course will be substantiated by actual
proficiency in the language. The predictive validity of an assessment becomes important in
the case of placement tests, admissions assessment batteries, language aptitude tests, and the
like. The assessment criterion in such cases is not to measure concurrent ability but to assess
(and predict) a test-taker's likelihood of future success.
3. Construct - Related .Evidence
A third kind of evidence that can support validity, but one that does not play as large a
role for classroom teachers, is construct-related validity, commonly referred to as construct
Validity.
A construct is any theory, hypothesis, or model that attempts to v' explain observed
phenomena in our universe of perceptions. Constructs mayor may not be directly or
empirically measured-their verification often requires inferentail data. "Proficiency" and
"communicative competence" are linguistic constructs; "self-esteem" and "motivation" are
psychological constructs.
Construct validity is a major issue in validating large-scale standardized tests of
proficiency. Because such tests must, for economic reasons, adhere to the principle of
practicality, and because they must sample a limited number of domains of language, they
may not be able to contain all the content of a particular field or skill.
The TOEF for example, has until recently not attempted to sample oral production,
yet oral production is obviously an important part of academic success in a university 26
CHAPTER 2 Principles of Language Assessment course of study. The TOEFL's omission of
oral production content, however, is often obviously justified by research that has shown
positive correlations between oral production and the behaviours (listening, reading,
grammaticality detection, and writing).
4. Consequential Validity
Consequential validity encompasses all the consequences of a test, including such
considerations as its accuracy in measuring intended criteria, its. impact on the preparation
of test-takers, its effect on the learner, and the (intended and unintended) social
consequences of a test's interpretation and use.

5. Face Validity

An important facet of consequential validity is the extent to which "students view the
assessment as fair, relevant, and useful for improving learning" (Gronlund, 1998, p. 210),
or what is popularly known as face validity. "Face validity refers to the degree to which a
test looks right, and appears to measure the knowledge or abilities it claims to measure,
based on the subjective judgment of the examinees who take it, the administrative
personnel who decide on its use, and other psychometrics unsophisticated observers".

face validity is not something that can be empirically tested by a teacher or even by a
testing expert. It is purely a factor of the "·eye of the beholder"-how the test-taker, or
possibly the test giver, intuitively perceives the instrument. For this reason, some
assessment experts (see Stevenson, 1985) view face validity as a superficial factor
that is dependent on the whim of the perceiver.

I once administered a dictation test and a cloze test for a discussion of cloze tests as a
placement test for a group of learners of English as a second language. Some learners
were upset because such tests, on the face of it, did not appear to them to test their
true abilities in English. They felt that a multiple choice grammar test would have
been the appropriate format to use.

A few claimed they didn't perform well on the cloze and dictation because they were
not accustomed to these formats. As it turned out, the tests served as superior
instruments for placement, but the students would not have thought so.

Face validity was low, content validity was moderate, and construct validity was
very high.

Common questions

Powered by AI

Yes, a test can possess high construct validity even with moderate content validity and low face validity. Construct validity ensures the test measures theoretical constructs accurately, such as communicative competence, based on strong empirical foundations . Moderate content validity implies some adherence to the subject matter but not comprehensively. Low face validity, a subjective perception, does not necessarily impact the test's psychometric integrity. For language assessments, a test might appear unfair or not entirely cover all taught content, yet still reliably measure proficiency through validated construct relationships, as evidenced by dissatisfaction with but effectiveness of certain language tests .

There are five primary types of validity that contribute to assessing the effectiveness of a test: content-related, criterion-related, construct-related, consequential, and face validity. Content-related validity measures if the test samples the subject and the behavior being measured accurately . Criterion-related validity assesses the extent to which a test's results align with another measure, divided into concurrent and predictive validity . Construct-related validity evaluates if a test measures a theoretical construct like communicative competence or proficiency . Consequential validity considers the test's overall impact, including accuracy and social consequences . Face validity refers to the perceived appropriateness of the test by those involved, even though it cannot be empirically measured .

Content-related validity ensures a test samples the material it purports to cover, requiring test-takers to demonstrate the behavior being measured; this often involves direct testing of the actual task . Criterion-related validity, on the other hand, establishes a test's effectiveness by comparing its results with another measure, with focus on concurrent and predictive validity . For test design, content validity requires careful task selection that reflects actual objectives, whereas criterion validity demands alignment with reliable external measures to substantiate accuracy .

Consequential validity influences a test's effectiveness by addressing its accuracy in measuring intended criteria, its impact on preparation, and its social consequences . It includes unintended outcomes that could misalign test objectives and societal needs. Face validity affects perception by determining if stakeholders believe a test measures what it claims to. Although subjective and not empirically measured, face validity impacts acceptance and perceived fairness . Together, they influence how a test is received and whether its results are trusted, thereby affecting its overall utility.

A test with high construct validity but low face validity might be effective if it accurately measures underlying theoretical constructs despite appearing irrelevant or unfair to those taking it. Construct validity ensures the test assesses its intended skill or concept competently based on research and theory . However, its format or content might not be intuitively accepted by test-takers, as illustrated by learners dissatisfied with dictation and cloze tests despite their efficacy in placement . Thus, construct validity can override superficial perceptions in instances where empirical support is strong.

Low face validity in assessments can lead to test-takers perceiving the test as unfair or irrelevant, which can affect their motivation and engagement. In educational environments, this perception might result in students doubting their performance results or expressing dissatisfaction with test formats . An example of low face validity occurred when students disagreed with the use of cloze and dictation tests for English proficiency placement, believing a multiple-choice grammar test would better measure their abilities, impacting their acceptance of the results .

Predictive validity is crucial for admission tests and language aptitude assessments because it assesses the test's ability to forecast future success. For admission tests, predictive validity ensures the selected candidates are likely to succeed in future academic environments. Language aptitude tests utilize predictive validity to gauge whether candidates will effectively learn a language, guiding placement and instruction decisions . High predictive validity indicates that test results are reliable indicators of future performance, informing critical decisions in educational settings and admissions processes .

Construct-related validity impacts standardized language assessments by determining if the test adequately measures theoretical constructs like communicative competence. Large-scale tests may struggle to sample all content areas due to practicality, risking inadequate measurement of constructs . For example, while TOEFL omits oral production, research shows its components correlate with oral skills, thus maintaining its construct validity. However, gaps can result in less comprehensive assessments, necessitating careful research and validation .

Direct testing involves performing the actual task being measured, thus establishing strong content validity as it directly assesses the skills taught . Indirect testing, where tasks relate but do not replicate the target skill, can weaken content validity due to mismatches in expected performance and test format . For language assessments, this means designing tests that closely align with instructional goals, such as speaking to assess conversational skills directly rather than using unrelated multiple-choice questions. Direct testing is emphasized for authenticity in language use .

Indirect testing methods can challenge validity by failing to directly measure the skills intended, often leading to a mismatch between instructional goals and assessment focus. This can result in lower content validity, as test-takers perform tasks not representative of the skills they were taught, like assessing oral production through written tasks . This indirect approach might fail to capture the complexity of language use fully, affecting the authenticity and reliability of the test results in predicting true proficiency levels .

You might also like