0% found this document useful (0 votes)
7 views23 pages

Validity Coefficient: Importance and Examples

Validity refers to the appropriateness of inferences made from test scores and the degree to which evidence supports the intended interpretation of scores. Validity is inferred from multiple sources of evidence including content representativeness, criterion relationships, construct evidence, and consequences of using the assessment. Reliability provides consistency needed for validity and enables interpretation of results with greater confidence.

Uploaded by

Hafiz Rabbi
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views23 pages

Validity Coefficient: Importance and Examples

Validity refers to the appropriateness of inferences made from test scores and the degree to which evidence supports the intended interpretation of scores. Validity is inferred from multiple sources of evidence including content representativeness, criterion relationships, construct evidence, and consequences of using the assessment. Reliability provides consistency needed for validity and enables interpretation of results with greater confidence.

Uploaded by

Hafiz Rabbi
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

4 VALIDITY AND RELIABILITY

Studying this chapter should enable you to

[Link] between validity and reliability.


[Link] the essential features of the concept of validity.
[Link] how content-related evidence of validity is obtained.
[Link] factors that can lower the validity of achievement assessments.
[Link] procedures for obtaining criterion-related evidence of validity.
[Link] procedures for obtaining construct-related evidence of validity.
[Link] the role of consequences of using an assessment procedure
on its validity.
8. Describe the methods for estimating test reliability and the type of
information provided by each.
9. Describe how the standard error of measurement is computed and
interpreted.
10. Explain how to determine the reliability of a performance-based
assessment.

The two most important questions to ask about a test or other assessment pro-
cedure are (1) to what extent will the interpretation of the results be appropri-
ate, meaningful, and useful?, and (2) to what extent will the results be free from
errors? The first question is concerned with validity, the second with reliability.
An understanding of both concepts is essential to the effective construction,
selection, interpretation, and use of tests and other assessment instruments.
Validity is the most important quality to consider in the preparation and use of
assessment procedures. First and foremost, we want the results to provide a rep-
resentative and relevant measure of the achievement domain under consider-
ation. Our second consideration is reliability, which refers to the consistency of
our assessment results. For example, if we tested individuals at a different time,
or with a different sample of equivalent items, we would like to obtain approxi-
mately the same results. This consistency of results is important for two reasons.
(1) Unless the results are fairly stable, we cannot expect them to be valid.
47
48 Chapter 4 • Validity and Reliability

For example, if an individual scored high on a test one time and low another
time, it would be impossible to validly describe the achievement. (2) Consistency
of results indicates smaller errors of measurement and, thereby, more depend-
able results. Thus, reliability provides the consistency needed to obtain validity
and enables us to interpret assessment results with greater confidence.
Although it is frequently unnecessary to make elaborate validation and re-
liability studies of informal assessment procedures, an understanding of these
concepts provides a conceptual framework that can serve as a guide for more
effective construction of assessment instruments, more effective selection of stan-
dardized tests, and more appropriate interpretation and use of assessment results.

VALIDITY
Validity is concerned with the interpretation and use of assessment results. For
example, if we infer from an assessment that students have achieved the in-
tended learning outcomes, we would like some assurance that our tasks pro-
vided a relevant and representative measure of the outcomes. If we infer that
the assessment is useful for predicting or estimating some other performance,
we would like some credible evidence to support that interpretation. If we infer
that our assessment indicates that students have good “reasoning ability,” we
would like some evidence to support the fact that the results actually reflect
that construct. If we infer that our use of an assessment had positive effects
(e.g., increased motivation) and no adverse effects (e.g., poor study habits) on
students, we would like some evidence concerning the consequences of its use.
These are the kinds of considerations we are concerned with when considering
the validity of assessment results (see the summary in Figure 4.1).
The meaning of validity in interpreting and using assessment results can
be grasped most easily by reviewing the following characteristics.
1. Validity is inferred from available evidence (not measured).
2. Validity depends on many different types of evidence.
3. Validity is expressed by degree (high, moderate, low).
4. Validity is specific to a particular use.

Content Representativeness

Criterion Relationships
VALIDITY

Construct Evidence

Consequences of Using

FIGURE 4.1 Types of Considerations in Determining the Validity of Assessment Results.


Chapter 4 • Validity and Reliability 49

5. Validity refers to the inferences drawn, not the instrument.


6. Validity is a unitary concept.
7. Validity is concerned with the consequences of using the assessments.
Describing validity as a unitary concept is a basic change in how validity
is viewed. The traditional view that there were several different “types of valid-
ity” has been replaced by the view that validity is a single, unitary concept that
is based on various forms of evidence. The former “types of validity” (content,
criterion-related, and construct) are now simply considered to be convenient
categories for accumulating evidence to support the validity of an interpreta-
tion. Thus, we no longer speak of “content validity,” but of “content-related
evidence” of validity. Similarly, we speak of “criterion-related evidence” and
“construct-related evidence.”
For some interpretations of assessment results only one or two types of
evidence may be critical, but an ideal validation would include evidence from
all four categories. We are most likely to draw valid inferences from assessment
results when we have a full understanding of: (1) the nature of the assessment
procedure and the specifications that were used in developing it, (2) the rela-
tion of the assessment results to significant criterion measures, (3) the nature of
the psychological characteristic(s) or construct(s) being assessed, and (4) the
consequences of using the assessment. Although in many practical situations the
evidence falls short of this ideal, we should gather as much relevant evidence as
is feasible within the constraints of the situation. We should also look for the
various types of evidence when evaluating standardized tests (see Table 4.1).

Content-Related Evidence
Content-related evidence of validity is critical when we want to use perfor-
mance on a set of tasks as evidence of performance on a larger domain of tasks.
Let’s assume, for example, that we have a list of 500 words that we expect our
students to be able to spell correctly at the end of the school year. To test their
spelling ability, we might give them a 50-word spelling test. Their performance

TABLE 4.1 Basic Approaches to Validation

Type of Evidence Question to Be Answered

Content-Related How adequately does the sample of assessment tasks


represent the domain of tasks to be measured?
Criterion-Related How accurately does performance on the assessment
(e.g., test) predict future performance (predictive study)
or estimate present performance (concurrent study) on
some other valued measure called a criterion?
Construct-Related How well can performance on the assessment be
explained in terms of psychological characteristics?
Consequences How well did use of the assessment serve the intended
purpose (e.g., improve performance) and avoid adverse
effects (e.g., poor study habits)?
50 Chapter 4 • Validity and Reliability

on these words is important only insofar as it provides evidence of their ability


to spell the 500 words. Thus, our spelling test would provide a valid measure to
the degree to which it provided an adequate sample of the 500 words it repre-
sented. If we selected only easy words, only difficult words, or only words that
represented certain types of common spelling errors, our test would tend to be
unrepresentative and thus the scores would have low validity. If we selected a
balanced sample of words that took these and similar factors into account, our
test scores would provide a representative measure of the 500 spelling words
and thereby provide for high validity.
It should be clear from this discussion that the key element in content-
related evidence of validity is the adequacy of the sampling. An assessment is
always a sample of the many tasks that could be included. Content validation
is a matter of determining whether the sample of tasks is representative of the
larger domain of tasks it is supposed to represent.
Content-related evidence of validity is especially important in achieve-
ment assessment. Here we are interested in how well the assessment measures
the intended learning outcomes of the instruction. We can provide greater as-
surance that an assessment provides valid results by (1) identifying the learning
outcomes to be assessed, (2) preparing a plan that specifies the sample of tasks
to be used, and (3) preparing an assessment procedure that closely fits the set
of specifications. These are the best procedures we have for ensuring the as-
sessment of a representative sample of the domain of tasks encompassed by the
intended learning outcomes.
Although the focus of content-related evidence of validity is on the ad-
equacy of the sampling, a valid interpretation of the assessment results assumes
that the assessment was properly prepared, administered, and scored. Validity
can be lowered by inadequate procedures in any of these areas (see Box 4.1).
Thus, validity is “built in” during the planning and preparation stages and main-
tained by proper administration and scoring. Throughout this book we describe
how to prepare assessments that provide valid results, even though we do not
use the word validity as each procedure is discussed.

BOX 4.1 Factors That Lower the Validity of Assessment


Results

1. Tasks that provide an inadequate sample of the achievement to be assessed.


2. Tasks that do not function as intended, due to use of improper types of tasks,
lack of relevance, ambiguity, clues, bias, inappropriate difficulty, or similar
factors.
3. Improper arrangement of tasks and unclear directions.
4. Too few tasks for the types of interpretation to be made (e.g., interpretation by
objective based on a few test items).
5. Improper administration—such as inadequate time allowed and poorly con-
trolled conditions.
6. Judgmental scoring that uses inadequate scoring guides, or objective scoring
that contains computational errors.
Chapter 4 • Validity and Reliability 51

The makers of standardized tests follow these same systematic procedures


in building achievement tests, but the content and learning outcomes included
in the test specifications are more broadly based than those used in classroom
assessment. Typically, they are based on the leading textbooks and the recom-
mendations of various experts in the area being covered by the test. Therefore,
a standardized achievement test may be representative of a broad range of
content but be unrepresentative of the domain of content taught in a particular
school situation. To determine relevance to the local situation, it is necessary to
evaluate the sample of test items in light of the content and skills emphasized
in the instruction.
In summary, content-related evidence of validity is of major concern in
achievement assessment, whether you are developing or selecting the assess-
ment procedure. When constructing a test, for example, content relevance
and representativeness are built in by following a systematic procedure for
specifying and selecting the sample of test items, constructing high-quality
items, and arranging the test for efficient administration and scoring. In test
selection, it is a matter of comparing the test sample to the domain of tasks to
be measured and determining the degree of correspondence between them.
Similar care is needed when preparing and using performance assessments.
Thus, content-related evidence of validity is obtained primarily by careful,
logical analysis.

Criterion-Related Evidence
There are two types of studies used in obtaining criterion-related evidence of
validity. These can be explained most clearly using test scores, although they
could be used with any type of assessment result. The first type of study is
concerned with the use of test performance to predict future performance on
some other valued measure called a criterion. For example, we might use scho-
lastic aptitude test scores to predict course grades (the criterion). For obvious
reasons, this is called a predictive study. The second type of study is concerned
with the use of test performance to estimate current performance on some cri-
terion. For instance, we might want to use a test of study skills to estimate what
the outcome would be of a careful observation of students in an actual study
situation (the criterion). Since with this procedure both measures (test and cri-
terion) are obtained at approximately the same time, this type of study is called
a concurrent study.
Although the value of using a predictive study is rather obvious, a ques-
tion might be raised concerning the purpose of a concurrent study. Why would
anyone want to use test scores to estimate performance on some other measure
that is to be obtained at the same time? There are at least three good reasons for
doing this. First, we may want to check the results of a newly constructed test
against some existing test that has a considerable amount of validity evidence
supporting it. Second, we may want to substitute a brief, simple testing proce-
dure for a more complex and time-consuming measure. For example, our test
of study skills might be substituted for an elaborate rating system if it provided
a satisfactory estimate of study performance. Third, we may want to determine
52 Chapter 4 • Validity and Reliability

whether a testing procedure has potential as a predictive instrument. If a test


provides an unsatisfactory estimate of current performance, it certainly cannot
be expected to predict future performance on the same measure. On the other
hand, a satisfactory estimate of present performance would indicate that the test
might be useful in predicting future performance as well. This would inform us
that a predictive study would be worth doing.
The key element in both types of criterion-related study is the degree of
relationship between the two sets of measures: (1) the test scores and (2) the
criterion to be predicted or estimated. This relationship is typically expressed by
means of a correlation coefficient or an expectancy table.

CORRELATION COEFFICIENTS. A correlation coefficient (r) simply indicates


the degree of relationship between two sets of measures. A positive relationship
is indicated when high scores on one measure are accompanied by high scores
on the other; low scores on the two measures are similarly associated. A nega-
tive relationship is indicated when high scores on one measure are accompa-
nied by low scores on the other. The extreme degrees of relationship it is pos-
sible to obtain between two sets of scores are indicated by the following values:
1.00 = perfect positive relationship
.00 = no relationship
-1.00 = perfect negative relationship

When a correlation coefficient is used to express the degree of relation-


ship between a set of test scores and some criterion measure, it is called a
validity coefficient. For example, a validity coefficient of 1.00 applied to the
relationship between a set of aptitude test scores (the predictor) and a set of
achievement test scores (the criterion) would indicate that each individual in
the group had exactly the same relative standing on both measures, and would
thereby provide a perfect prediction from the aptitude scores to the achieve-
ment scores. Most validity coefficients are smaller than this, but the extreme
positive relationship provides a useful benchmark for evaluating validity coeffi-
cients. The closer the validity coefficient approaches 1.00, the higher the degree
of relationship and, thus, the more accurate our predictions of each individual’s
success on the criterion will be.
A more realistic procedure for evaluating a validity coefficient is to com-
pare it to the validity coefficients that are typically obtained when the two
measures are correlated. For example, a validity coefficient of .40 between a set
of aptitude test scores and achievement test scores would be considered small
because we typically obtain coefficients in the .50 to .70 range for these two
measures. Therefore, validity coefficients must be judged on a relative basis,
the larger coefficients being favored. To use validity coefficients effectively, one
must become familiar with the size of the validity coefficients that are typically
obtained between various pairs of measures under different conditions (e.g.,
the longer the time span between measures, the smaller the validity coefficient).
A number of statistical software programs and calculators are capable of
calculating a correlation coefficient between two sets of measures. However, it
Chapter 4 • Validity and Reliability 53

is not difficult to calculate a correlation using a simple calculator. It is simply


a matter of following a series of steps to arrive at the solution to the following
correlation coefficient formula.
N 1 g xy 2 - 1 g x 21 g y 2
r =
2 3 N1 g x 22 - 1 g x 2 2 4 3 N 1 g y2 2 - 1 g y 2 2 4
To illustrate the process, suppose we conduct a predictive study to determine
whether students’ scores on a paper-and-pencil achievement test predict future
performance on a related end-of-course project. Table 4.2 provides each stu-
dent’s score on both the test and the project.
You will note from the table that the test scores, the predictor, are labeled
x and the project scores, the criterion, are labeled y. As you will see, these la-
bels are necessary in order to calculate the correlation coefficient.
To arrive at the values to place into the correlation coefficient formula,
several calculations need to be made first. To make this task much easier, it is
helpful to create a new table by adding a few columns to Table 4.2. Table 4.3
shows the existing raw data, as well as derived data from several calculations.
Moving from left to right, the numerator of the correlation coefficient for-
mula begins with N, which refers to the number of students who completed
the assessments. In this example, N equals 10 students. The next value in the
numerator is the sum of x multiplied by y, or g xy. To obtain this number, we
must first multiply the x and y values for each student. For example, as shown
in Table 4.3, for student A the test score (x) of 15 was multiplied by the stu-
dent’s project score (y) of 16, and the product of 240 was placed in the last
column, labeled xy. This procedure was repeated for the remaining students.
The xy values were then added, resulting in the sum of xy equaling 1787. The
two remaining values to be placed in the numerator of the formula are the sum
of x, or g x, and the sum of y, or g y. As shown in the table, adding together
the test scores (x) results in g x = 140. Adding together the project scores (y)

TABLE 4.2 Student Scores on Achievement Test and


Class Project

Student Test Scores (x) Project Scores (y)


A 15 16
B 18 15
C 12 8
D 13 11
E 19 17
F 10 9
G 14 13
H 11 5
I 17 17
J 11 9
54 Chapter 4 • Validity and Reliability

TABLE 4.3 Raw and Derived Data

Student Test Scores (x) x2 Project Scores (y) y2 xy

A 15 225 16 256 240


B 18 324 15 225 270
C 12 144 8 64 96
D 13 169 11 121 143
E 19 361 17 289 323
F 10 100 9 81 90
G 14 196 13 169 182
H 11 121 5 25 55
I 17 289 17 289 289
J 11 121 9 81 99
Sums 1 g 2 140 2050 120 1600 1787

results in g y = 120. Thus, we have calculated all of the values needed for the
numerator of the formula.
Again, moving from left to right, the denominator of the correlation co-
efficient begins with N, which as explained earlier is 10. The next value to be
calculated is the sum of x squared, or g x2. To obtain this number, we must
first square the x values for each student. For example, as shown in Table 4.3,
the test score (x) for student A of 15 was squared, and the product of 225 was
placed in the third column, labeled x2. This procedure was repeated for the
remaining students. The x2 values were then added, resulting in the sum of x2
equaling 2050. Returning to the formula, the next value in the denominator is
the sum of x, or g x, calculated earlier to be 140. Also calculated earlier and in-
cluded in the denominator is the sum of y, or g y, which equals 120. Thus, the
only value in the denominator that remains unknown is the sum of y squared,
or g y2. As with the calculation of the g x2, we simply square the project scores
(y) for each student, place the product of each calculation in the fifth column of
the table labeled y2, and add the values in the column to determine the sum of
y squared. In this example, g y2 = 1600.
Now that all of the needed values have been calculated, we simply place
them in the correlation coefficient formula and conduct the calculations. The
result is a correlation coefficient of .89.

10117872 - 11402 11202


r = = .89
2 3 10120502 - 11402 2 4 3 10116002 - 11202 2 4
This validity coefficient indicates that the degree of relationship between the test
scores and project scores is relatively strong, although as noted earlier, the co-
efficient should ideally be compared with validity coefficients that are typically
obtained for these two measures. Given a relatively strong correlation, is it then
valid to infer that the test scores will predict performance on the end-of-course
Chapter 4 • Validity and Reliability 55

project? A relatively strong correlation would support such an inference. Of


course, a perfect positive correlation of 1.00 would provide the strongest predic-
tive evidence. But, these correlations are seldom if ever achieved.

EXPECTANCY TABLE. The expectancy table is a simple and practical means


of expressing criterion-related evidence of validity and is especially useful for
making predictions from test scores. The expectancy table is simply a two-
fold chart with the test scores (the predictor) arranged in categories down the
left side of the table and the measure to be predicted (the criterion) arranged
in categories across the top of the table. For each category of scores on the
predictor, the table indicates the percentage of individuals who fall within each
category of the criterion. An example of an expectancy table is presented in
Table 4.4.
Note in Table 4.4 that of those students who were in the above-average
group (stanines 7, 8, and 9) on the test scores, 43 percent received a grade of
A, 43 percent a B, and 14 percent a C. Although these percentages are based
on this particular group, it is possible to use them to predict the future perfor-
mance of other students in this science course. Hence, if a student falls in the
above-average group on this scholastic aptitude test, we might predict that he
or she has 43 chances out of 100 of earning an A, 43 chances out of 100 of
earning a B, and 14 chances out of a 100 of earning a C in this particular sci-
ence course. Such predictions are highly tentative, of course, due to the small
number of students on which this expectancy table was built. Teachers can
construct more dependable tables by accumulating data from several classes
over a period of time.
Expectancy tables can be used to show the relationship between any two
measures. Constructing the table is simply a matter of (1) grouping the scores
on each measure into a series of categories (any number of them), (2) plac-
ing the two sets of categories on a twofold chart, (3) tabulating the number
of students who fall into each position in the table (based on the student’s

TABLE 4.4 Expectancy Table Showing the Relationship Between Scholastic


Aptitude Scores and Course Grades for 30 Students in a Science Course

Grouped Scholastic
Aptitude Scores Percentage in Each Score Category Receiving
(Stanines) Each Grade

F D C B A
Above Average 14 43 43
(7, 8, 9)
Average 19 37 25 19
(4, 5, 6)
Below Average 57 29 14
(1, 2, 3)
56 Chapter 4 • Validity and Reliability

standing on both measures), and (4) converting these numbers to percentages


(of the total number in that row). Thus, the expectancy table is a clear way
of showing the relationship between sets of scores. Although the expectancy
table is more cumbersome to deal with than a correlation coefficient, it has the
special advantage of being easily understood by persons without knowledge
of statistics. Thus, it can be used in practical situations to clarify the predictive
efficiency of a test.

Construct-Related Evidence
The construct-related category of evidence focuses on assessment results as a
basis for inferring the possession of certain psychological characteristics. For
example, we might want to describe a person’s reading comprehension, rea-
soning ability, or mechanical aptitude. These are all hypothetical qualities, or
constructs, that we assume exist in order to explain behavior. Such theoretical
constructs are useful in describing individuals and in predicting how they will
act in many different specific situations. To describe a person as being highly
intelligent, for example, is useful because that term carries with it a series of
associated meanings that indicate what the individual’s behavior is likely to be
under various conditions. Before we can interpret assessment results in terms
of these broad behavior descriptions, however, we must first establish that the
constructs that are presumed to be reflected in the scores actually do account
for differences in performance.
Construct-related evidence of validity for a test includes (1) a description
of the theoretical framework that specifies the nature of the construct to be
measured, (2) a description of the development of the test and any aspects of
measurement that may affect the meaning of the test scores (e.g., test format),
(3) the pattern of relationship between the test scores and other significant
variables (e.g., high correlations with similar tests and low correlations with
tests measuring different constructs), and (4) any other type of evidence that
contributes to the meaning of the test scores (e.g., analyzing the mental process
used in responding, determining the predictive effectiveness of the test). The
specific types of evidence that are most critical for a particular test depend on
the nature of the construct, the clarity of the theoretical framework, and the
uses to be made of the test scores. Although the gathering of construct-related
evidence of validity can be endless, in practical situations it is typically neces-
sary to limit the evidence to that which is most relevant to the interpretations
to be made.
The construct-related category of evidence is the broadest of the three
categories. Evidence obtained in both the content-related category (e.g., rep-
resentativeness of the sample of tasks) and the criterion-related category (e.g.,
how well the scores predict performance on specific criteria) is also relevant
to the construct-related category because it helps to clarify the meaning of the
assessment results. Thus, the construct-related category encompasses a variety
of types of evidence, including that from content-related and criterion-related
validation studies (see Figure 4.2).
Chapter 4 • Validity and Reliability 57

CONSTRUCT-
RELATED
EVIDENCE

Content-Related VALIDITY
Studies OF
INFERENCES
Criterion-Related
Studies

Other Relevant
Evidence

FIGURE 4.2 Construct Validation Includes All Categories of Evidence.

The broad array of evidence that might be considered can be illustrated


by a test designed to measure mathematical reasoning ability. Some of the evi-
dence we might consider is the following:
1. Compare the sample of test tasks to the domain of tasks specified by the
conceptual framework of the construct. Is the sample relevant and repre-
sentative (content-related evidence)?
2. Examine the test features and their possible influence on the meaning of
the scores (e.g., test format, directions, scoring, reading level of items). Is
it possible that some features might distort the scores?
3. Analyze the mental process used in answering the questions by having
students “think aloud” as they respond to each item. Do the items require
the intended reasoning process?
4. Determine the internal consistency of the test by intercorrelating the test
items. Do the items seem to be measuring a single characteristic (in this
case mathematical reasoning)?
5. Correlate the test scores with the scores of other mathematical reasoning
tests. Do they show a high degree of relationship?
6. Compare the scores of known groups (e.g., mathematical majors and non-
majors). Do the scores differentiate between the groups as predicted?
7. Compare the scores of students before and after specific training in math-
ematical reasoning. Do the scores change as predicted from the theory
underlying the construct?
8. Correlate the scores with grades in mathematics. Do they correlate to a
satisfactory degree (criterion-related evidence)?
Other types of evidence could be added to this list, but it is sufficiently
comprehensive to make clear that no single type of evidence is adequate.
Interpreting test scores as a measure of a particular construct involves a com-
prehensive study of the development of the test, how it functions in a variety of
situations, and how the scores relate to other significant measures.
Assessment results are, of course, influenced by many factors other than
the construct they are designed to measure. Thus, construct validation is an
58 Chapter 4 • Validity and Reliability

attempt to account for all possible influences on the scores. We might, for ex-
ample, ask to what extent the scores on our mathematical reasoning test are in-
fluenced by reading comprehension, computation skill, and speed. Each of these
factors would require further study. Were attempts made to eliminate such fac-
tors during test development by using simple vocabulary, simple computations,
and liberal time limits? To what extent do the test scores correlate with measures
of reading comprehension and computational skill? How do students’ scores dif-
fer under different time limits? Answers to these and similar questions will help
us to determine how well the test scores reflect the construct we are attempting
to measure and the extent to which other factors might be influencing the scores.
Construct validation, then, is an attempt to clarify and verify the inferences to
be made from assessment results. This involves a wide variety of procedures and
many different types of evidence (including both content-related and criterion-re-
lated). As evidence accumulates from many different sources, our interpretations
of the results are enriched and we are able to make them with great confidence.

Consequence of Using Assessment Results


Validity focuses on the inferences drawn from assessment results with regard
to specific uses. Therefore, it is legitimate to ask, What are the consequences
of using the assessment? Did the assessment improve learning, as intended, or
did it contribute to adverse effects (e.g., lack of motivation, memorization, poor
study habits)? For example, assessment procedures that focus on simple learn-
ing outcomes only (e.g., knowledge of facts) cannot provide valid evidence of
reasoning and application skills, are likely to narrow the focus of student learn-
ing, and tend to reinforce poor learning strategies (e.g., rote learning). Thus, in
evaluating the validity of the assessment used, one needs to look at what types
of influence the assessments have on students. The following questions provide
a general framework for considering some of the possible consequences of
assessments on students.
1. Did use of the assessment improve motivation?
2. Did use of the assessment improve performance?
3. Did use of the assessment improve self-assessment skills?
4. Did use of the assessment contribute to transfer of learning to related
areas?
5. Did use of the assessment encourage independent learning?
6. Did use of the assessment encourage good study habits?
7. Did use of the assessment contribute to a positive attitude toward
schoolwork?
8. Did use of the assessment have an adverse effect in any of the above
areas?
Judging the consequences of using the various assessment procedures
is an important role of the teacher, if the results are to serve their intended
purpose of improving learning. Both testing and performance assessments are
most likely to have positive consequences when they are designed to assess
a broad range of learning outcomes, they give special emphasis to complex
Chapter 4 • Validity and Reliability 59

learning outcomes, they are administered and scored (or judged) properly,
they are used to identify students’ strengths and weaknesses in learning, and
the students view the assessments as fair, relevant, and useful for improving
learning.

RELIABILITY
Reliability refers to the consistency of assessment results. Would we obtain
about the same results if we used a different sample of the same type task?
Would we obtain about the same results if we used the assessment at a different
time? If a performance assessment is being rated, would different raters rate the
performance the same way? These are the kinds of questions we are concerned
about when we are considering the reliability of assessment results. Unless the
results are generalizable over similar samples of tasks, time periods, and raters,
we are not likely to have great confidence in them.
Because the methods for estimating reliability differ for tests and perfor-
mance assessments, these will be treated separately.

Estimating the Reliability of Test Scores


The score an individual receives on a test is called the obtained score, raw
score, or observed score. This score typically contains a certain amount of error.
Some of this error may be systematic error, in that it consistently inflates or low-
ers the obtained score. For example, readily apparent clues in several test items
might cause all students’ scores to be higher than their achievement would
warrant, or short time limits during testing might cause all students’ scores to
be lower than their “real achievement.” The factors causing systematic errors
are mainly due to inadequate testing practices. Thus, most of these errors can
be eliminated by using care in constructing and administering tests. Removing
systematic errors from test scores is especially important because they have a
direct effect on the validity of the inferences made from the scores.
Some of the error in obtained scores is random error, in that it raises
and lowers scores in an unpredictable manner. Random errors are caused by
such things as temporary fluctuations in memory, variations in motivation and
concentration from time to time, carelessness in marking answers, and luck in
guessing. Such factors cause test scores to be inconsistent from one measure-
ment to another. Sometimes an individual’s obtained score will be higher than
it should be and sometimes it will be lower. Although these errors are difficult
to control and cannot be predicted with accuracy, an estimate of their influence
can be obtained by various statistical procedures. Thus, when we talk about
estimating the reliability of test scores or the amount of measurement error in
test scores, we are referring to the influence of random errors.
Reliability refers to the consistency of test scores from one measurement
to another. Because of the ever-present measurement error, we can expect
a certain amount of variation in test performance from one time to another,
from one sample of items to another, and from one part of the test to another.
Reliability measures provide an estimate of how much variation we might
60 Chapter 4 • Validity and Reliability

expect under different conditions. The reliability of test scores is typically


reported by means of a reliability coefficient or the standard error of mea-
surement that is derived from it. Since both methods of estimating reliability
require score variability, the procedures to be discussed are useful primarily
with tests designed for norm-referenced interpretation.
As we noted earlier, a correlation coefficient expressing the relation-
ship between a set of test scores and a criterion measure is called a validity
coefficient. A reliability coefficient is also a correlation coefficient, but it indi-
cates the correlation between two sets of measurements taken from the same
procedure. We may, for example, administer the same test twice to a group,
with a time interval in between (test-retest method); administer two equivalent
forms of the test in close succession (equivalent-forms method); administer
two equivalent forms of the test with a time interval in between (test-retest
with equivalent forms method); or administer the test once and compute the
consistency of the response within the test (internal-consistency method).
Each of these methods of obtaining reliability provides a different type of in-
formation. Thus, reliability coefficients obtained with the different procedures
are not interchangeable. Before deciding on the procedure to be used, we
must determine what type of reliability evidence we are seeking. The four
basic methods of estimating reliability and the type of information each pro-
vides are shown in Table 4.5.

TEST-RETEST METHOD. The test-retest method requires administering the


same form of the test to the same group after some time interval. The time
between the two administrations may be just a few days or several years.
The length of the time interval should fit the type of interpretation to be made
from the results. Thus, if we are interested in using test scores only to group
students for more effective learning, short-term stability may be sufficient. On
the other hand, if we are attempting to predict vocational success or make
some other long-range predictions, we would desire evidence of stability over
a period of years.

TABLE 4.5 Methods of Estimating Reliability of Test Scores

Method Type of Information Provided

Test-retest method The stability of test scores over a given period of time.
Equivalent-forms method The consistency of the test scores over different forms of
the test (that is, different samples of items).
Test-retest with equivalent The consistency of test scores over both a time interval
forms and different forms of the test.
Internal-consistency The consistency of test scores over different parts of the
methods test.
Note: Scorer reliability should also be considered when evaluating the responses to supply-type items (for
example, essay tests). This is typically done by having the test papers scored independently by two scorers and
then correlating the two sets of scores. Agreement among scorers, however, is not a substitute for the methods of
estimating reliability shown in the table.
Chapter 4 • Validity and Reliability 61

Test-retest reliability coefficients are influenced both by errors within the


measurement procedure and by the day-to-day stability of the students’ re-
sponses. Thus, longer time periods between testing will result in lower reli-
ability coefficients, due to the greater changes in the students. In reporting
test-retest reliability coefficients, then, it is important to include the time inter-
val. For example, a report might state: “The stability of test scores obtained on
the same form over a three-month period was .90.” This makes it possible to
determine the extent to which the reliability data are significant for a particular
interpretation.

EQUIVALENT-FORMS METHOD. With this method, two equivalent forms of a


test (also called alternate forms or parallel forms) are administered to the
same group during the same testing session. The test forms are equivalent in
the sense that they are built to measure the same abilities (that is, they are
built to the same set of specifications), but for determining reliability it is also
important that they be constructed independently. When this is the case, the
reliability coefficient indicates the adequacy of the test sample. That is, a high
reliability coefficient would indicate that the two independent samples are
apparently measuring the same thing. A low reliability coefficient, of course,
would indicate that the two forms are measuring different behavior and that
therefore both samples of items are questionable.
Reliability coefficients determined by this method take into account errors
within the measurement procedures and consistency over different samples of
items, but they do not include the day-to-day stability of the students’ responses.

TEST-RETEST METHOD WITH EQUIVALENT FORMS. This is a combination of both


methods. Here, two different forms of the same test are administered with time
intervening. This is the most demanding estimate of reliability, since it takes into
account all possible sources of variation. The reliability coefficient reflects errors
within the testing procedure, consistency over different samples of items, and
the day-to-day stability of the students’ responses. For most purposes, this is
probably the most useful type of reliability, since it enables us to estimate how
generalizable the test results are over the various conditions. A high reliability
coefficient obtained by this method would indicate that a test score represents
not only present test performance but also what test performance is likely to be
at another time or on a different sample of equivalent items.

INTERNAL-CONSISTENCY METHODS. These methods require only a single ad-


ministration of a test. One procedure, the split-half method, involves scoring
the odd items and the even items separately and correlating the two sets of
scores. This correlation coefficient indicates the degree to which the two ar-
bitrarily selected halves of the test provide the same results. Thus, it reports
on the internal consistency of the test. Like the equivalent-forms method, this
procedure takes into account errors within the testing procedure and consis-
tency over different samples of items, but it omits the day-to-day stability of the
students’ responses.
62 Chapter 4 • Validity and Reliability

Since the correlation coefficient based on the odd and even items indi-
cates the relationship between two halves of the test, the reliability coefficient
for the total test is determined by applying the Spearman-Brown prophecy for-
mula. A simplified version of this formula is as follows:
2 * reliability for 1 > 2 test
Reliability of total test =
1 + reliability for 1 > 2 test
Thus, if we obtained a correlation coefficient of .60 for two halves of a test, the
reliability for the total test would be computed as follows:
2 * .60 1.20
Reliability of total test = = = .75
1 + .60 1.60
This application of the Spearman-Brown formula makes clear a useful
principle of test reliability: The reliability of a test can be increased by lengthen-
ing it. This formula shows how much reliability will increase when the length of
the test is doubled. Application of the formula, however, assumes that the test
is lengthened by adding items like those already in the test.
Another internal-consistency method of estimating reliability is by use of
the Kuder-Richardson Formula 20 (KR-20). Kuder and Richardson developed
other formulas but this one is probably the most widely used with standardized
tests. It requires a single test administration, a determination of the proportion
of individuals passing each item, and the standard deviation of the total set of
scores. The formula is not especially helpful in understanding how to interpret
the scores, but knowing what the coefficient means is important. Basically, the
KR-20 is equivalent to an average of all split-half coefficients when the test is
split in all possible ways. Where all items in a test are measuring the same thing
(e.g., math reasoning), the result should approximate the split-half reliability
estimate. Where the test items are measuring a variety of skills or content areas
(i.e., less homogeneous), the KR-20 estimate will be lower than the split-half
reliability estimate. Thus, the KR-20 method is useful with homogeneous tests
but can be misleading if used with a test designed to measure heterogeneous
content.
Internal-consistency methods are used because they require that the test
be administered only once. They should not be used with speeded tests, how-
ever, because a spuriously high reliability estimate will result. If speed is an im-
portant factor in the testing (that is, if the students do not have time to attempt
all the items), other methods should be used to estimate reliability.

STANDARD ERROR OF MEASUREMENT. The standard error of measurement


is an especially useful way of expressing test reliability, because it indicates
the amount of error to allow for when interpreting individual test scores. The
standard error is derived from a reliability coefficient by means of the following
formula:
Standard error of measurement = s 21 - rn
Chapter 4 • Validity and Reliability 63

where s = the standard deviation and rn = the reliability coefficient. In


applying this formula to a reliability estimate of .60 obtained for a test where
s = 4.5, the following results would be obtained.
Standard error of measurement = 4.521 - .60
= 4.52.40
= 4.5 * .63
= 2.8
The standard error of measurement shows how many points we must add
to, and subtract from, an individual’s test score in order to obtain “reasonable
limits” for estimating that individual’s true score (that is, a score free of error). In
our example, the standard error would be rounded to 3 score points. Thus, if a
given student scored 35 on this test, that student’s score band, for establishing
reasonable limits, would range from 32 (35 - 3) to 38 (35 + 3). In other words,
we could be reasonably sure that the score band of 32 to 38 included the stu-
dent’s true score (statistically, there are two chances out of three that it does).
The standard errors of test scores provide a means of allowing for error during
test interpretation. If we view test performance in terms of score bands (also
called confidence bands), we are not likely to overinterpret small differences
between test scores.
For the test user, the standard error of measurement is probably more use-
ful than the reliability coefficient. Although reliability coefficients can be used in
evaluating the quality of a test and in comparing the relative merits of different
tests, the standard error of measurement is directly applicable to the interpreta-
tion of individual test scores.

RELIABILITY OF CRITERION-REFERENCED MASTERY TESTS. As noted earlier, the


traditional methods for computing reliability require score variability (that is, a
spread of scores) and are therefore useful mainly with norm-referenced tests.
When used with criterion-referenced tests, they are likely to provide misleading
results. Since criterion-referenced tests are not designed to emphasize differenc-
es among individuals, they typically have limited score variability. This restricted
spread of scores will result in low correlation estimates of reliability, even if the
consistency of our test results is adequate for the use to be made of them.
When a criterion-referenced test is used to determine mastery, our primary
concern is with how consistently our test classifies masters and nonmasters. If
we administered two equivalent forms of a test to the same group of students,
for example, we would like the results of both forms to identify the same stu-
dents as having mastered the material. Such perfect agreement is unrealistic,
of course, since some students near the cutoff score are likely to shift from
one category to the other on the basis of errors of measurement (due to such
factors as lucky guesses or lapses of memory). However, if too many students
demonstrated mastery on one form but nonmastery on the other, our decisions
concerning who mastered the material would be hopelessly confused. Thus, the
reliability of mastery tests can be determined by computing the percentage of
consistent mastery-nonmastery decisions over the two forms of the test.
64 Chapter 4 • Validity and Reliability

FIGURE 4.3 Classification of 40 Students as Masters or Nonmasters on Two Forms


of a Criterion-Referenced Test.

The procedure for comparing test performance on two equivalent forms


of a test is relatively simple. After both forms have been administered to a
group of students, the resulting data can be placed in a two-by-two table like
that shown in Figure 4.3. These data are based on two forms of a 25-item test
administered to 40 students. Mastery was set at 80 percent correct (20 items),
so all students who scored 20 or higher on both forms of the test were placed
in the upper right-hand cell (30 students), and all those who scored below 20
on both forms were placed in the lower left-hand cell (6 students). The remain-
ing students demonstrated mastery on one form and nonmastery on the other
(4 students). Since 36 of the 40 students were consistently classified by the two
forms of the test, we apparently have reasonably good consistency.
We can compute the percentage of consistency for this procedure with
the following formula:
Masters 1 both forms 2 + Nonmasters 1 both forms 2
% Consistency = * 100
Total number in group
30 + 6
% Consistency = * 100 = 90%
40
This procedure is simple to use but it has a few limitations. First, two
forms of the test are required. This may not be as serious as it seems, however,
since in most mastery programs more than one form of the test is needed for
retesting those students who fail to demonstrate mastery on the first try. Second,
it is difficult to determine what percentage of decision consistency is necessary
for a given situation. As with other measures of reliability, the greater the con-
sistency, the more satisfied we will be, but what constitutes a minimum accept-
able level? There is no simple answer to such a question because it depends
on the number of items in the test and the consequences of the decision. If a
nonmastery decision for a student simply means further study and later retest-
ing, low consistency might be acceptable. However, if the mastery-nonmastery
decision concerns whether to give a student a high school certificate, as in
some competency testing programs, then a high level of consistency will be
demanded. Since there are no clear guidelines for setting minimum levels, we
will need to depend on experience in various situations to determine what are
reasonable expectations.
More sophisticated techniques have been developed for estimating the
reliability of criterion-referenced tests, but the numerous issues and problems
Chapter 4 • Validity and Reliability 65

BOX 4.2 Factors That Lower the Reliability of Test Scores


1. Test scores are based on too few items. (Remedy: Use longer tests or accumu-
late scores from several short tests.)
2. Range of scores is too limited. (Remedy: Adjust item difficulty to obtain larger
spread of scores.)
3. Testing conditions are inadequate. (Remedy: Arrange opportune time for ad-
ministration and eliminate interruptions, noise, and other disrupting factors.)
4. Scoring is subjective. (Remedy: Prepare scoring keys and follow carefully when
scoring essay answers.)

involved in their use go beyond the scope of this book. See Box 4.2 for factors
that lower reliability of test scores.

Estimating the Reliability of Performance Assessments


Performance assessments are commonly evaluated by using scoring rubrics that
describe a number of levels of performance, ranging from high to low (e.g.,
outstanding to inadequate). The performance for each student is then judged
and placed in the category that best fits the quality of the performance. The
reliability of these performance judgments can be determined by obtaining and
comparing the scores of two judges who scored the performances independent-
ly. The scores of the two judges can be correlated to determine the consistency
of the scoring, or the proportion of agreement in scoring can be computed.
Let’s assume that a performance task, such as writing sample, was
obtained from 32 students and two teachers independently rated the students’
performance on a four-point scale where 4 is high and 1 is low. The results of
the ratings by the two judges are shown in Table 4.6. The ratings for Judge 1
are presented in the columns and those for Judge 2 are presented in the rows.
Thus, Judge 1 assigned a score of 4 to seven students and Judge 2 assigned
a score of 4 to eight students. Their ratings agreed on six of the students and

TABLE 4.6 Classification of Students Based on Performance Ratings by Two


Independent Judges

Ratings by Judge 1

Row
Scores 1 2 3 4 Totals

4 2 6 8
Ratings by
Judge 2 3 3 7 1 11

2 2 6 8

1 5 5
Column Totals 7 9 9 7 32
66 Chapter 4 • Validity and Reliability

disagreed by one score on three of the students. The number of rating agree-
ments can be seen in the boxes on the diagonal from the upper right-hand
corner to the lower left-hand corner. The percentage of agreement can be
computed by adding the numbers in these diagonal boxes (6 + 7 + 6 + 5 =
24), dividing by the total number of students in the group (32), and multiply-
ing by 100.
24
Rater agreement = * 100 = 75%
32
By inspection, we can see that all ratings were within one score of each
other. The results also indicate that Judge 2 was a more lenient rater than
Judge 1 (i.e., gave more high ratings and fewer low ratings). Thus, a table of
this nature can be used to determine the consistency of ratings and the extent
to which leniency can account for the disagreements.
Although the need for two raters will limit the use of this method, it seems
reasonable to expect two teachers in the same area to make periodic checks on
the scoring of performance assessments. This not only will provide information
on the consistency of the scoring, but will provide the teachers with insight into
some of their rating idiosyncrasies. See Box 4.3 for factors that lower the reli-
ability of performance assessment.
The percentage of agreement between the scores assigned by indepen-
dent judges is a common method of estimating the reliability of performance
assessments. It should be noted, however, that this reports on only one type of
consistency—the consistency of the scoring. It does not indicate the consistency
of performance over similar tasks or over different time periods. We can obtain
a crude measure of this by examining the performance of students over tasks
and time, but a more adequate analysis requires an understanding of general-
izability theory, which is too technical for treatment here.

BOX 4.3 Factors That Lower the Reliability of


Performance Assessments
1. Insufficient number of tasks. (Remedy: Accumulate results from several assess-
ments. For example, several writing samples.)
2. Poorly structured assessment procedures. (Remedy: Define carefully the nature
of the tasks, the conditions for obtaining the assessment, and the criteria for
scoring or judging the results.)
3. Dimensions of performance are specific to the tasks. (Remedy: Increase gener-
alizability of performance by selecting tasks that have dimensions like those in
similar tasks.)
4. Inadequate scoring guides for judgmental scoring. (Remedy: Use scoring ru-
brics or rating scales that specifically describe the criteria and levels of quality.)
5. Scoring judgments that are influenced by personal bias. (Remedy: Check scores
or ratings with those of an independent judge. Receive training in judging and
rating, if possible.)
Chapter 4 • Validity and Reliability 67

Summary of Points
1. Validity is the most important quality to consider in assessment and is
concerned with the appropriateness, meaningfulness, and usefulness of
the specific inferences made from assessment results.
2. Validity is a unitary concept based on various forms of evidence (content-
related, criterion-related, construct-related, and consequences).
3. Content-related evidence of validity refers to how well the sample of tasks
represents the domain of tasks to be assessed.
4. Content-related evidence of validity is of major concern in achievement
assessment and is built in by following systematic procedures. Validity is
lowered by inadequate assessment practices.
5. Criterion-related evidence of validity refers to the degree to which assess-
ment results are related to some other valued measure called a criterion.
6. Criterion-related evidence may be based on a predictive study or a con-
current study and is typically expressed by a correlation coefficient or
expectancy table.
7. Construct-related evidence of validity refers to how well performance on
assessment tasks can be explained in terms of psychological characteris-
tics, or constructs (e.g., mathematical reasoning).
8. The construct-related category of evidence is the most comprehensive. It in-
cludes evidence from both content-related and criterion-related studies plus
other types of evidence that help clarify the meaning of the assessment results.
9. Consequences of using the assessment are also an important consider-
ation in validity—both positive and negative consequences.
10. Reliability refers to the consistency of scores (i.e., to the degree to which
the scores are free from measurement error).
11. Reliability of test scores is typically reported by means of a reliability coef-
ficient or a standard error of measurement.
12. Reliability coefficients can be obtained by a number of different meth-
ods (e.g., test-retest, equivalent-forms, internal-consistency) and each one
measures a different type of consistency (e.g., over time, over different
samples of items, over different parts of the test).
13. Reliability of test scores tends to be lower when the test is short, range of
scores is limited, testing conditions are inadequate, and scoring is subjective.
14. The standard error of measurement indicates the amount of error to allow
for when interpreting individual test scores.
15. Score bands (or confidence bands) take into account the error of mea-
surement and help prevent the overinterpretation of small differences
between test scores.
16. The reliability of criterion-referenced mastery tests can be obtained by
computing the percentage of agreement between two forms of the test in
classifying individuals as masters and nonmasters.
17. The reliability of performance-based assessments is commonly determined
by the degree of agreement between two or more judges who rate the
performance independently.
68 Chapter 4 • Validity and Reliability

References and Additional Reading


American Educational Research Association. Standards for Educational and Psychological
Testing (Washington, DC: AERA, 1999).
Linn, R. L., & Gronlund, N. E. Measurement and Assessment in Teaching, 8th ed. (Upper
Saddle River, NJ: Merrill/Prentice Hall, 2000).
Nitko, A. J., & Brookhart, S. M. Educational Assessment of Students, 5th ed. (Upper
Saddle River, NJ: Merrill/Prentice Hall, Pearson Education, 2007).
Oosterhoff, A. C. Classroom Applications of Educational Measurement, 3rd ed. (Upper
Saddle River, NJ: Merrill/Prentice Hall, 2001).

Self-Assessment
MULTIPLE CHOICE
DIRECTIONS: Select the best answer for each item by circling the corresponding letter.

1. The appropriateness and meaningfulness of the inferences we make from


assessment results refers to a test’s
(a) reliability
(b) validity
(c) objectivity
(d) difficulty
2. To obtain evidence of validity based on content considerations, you would examine the
(a) expectancy table
(b) size of the correlation coefficient
(c) type of criterion used
(d) table of specifications

3. Interpreting a student’s chances of success in college based on the Scholastic


Aptitude Test (SAT) requires a
(a) predictive study
(b) criterion study
(c) construct study
(d) concurrent study
4. External influences such as disruptions during testing that may lower the scores of
the students are referred to as
(a) systematic errors
(b) standard errors
(c) random errors
(d) reliability errors
5. Ensuring an assessment has a good representative sample of relevant test items is a
characteristic of
(a) validity
(b) referencing
(c) reliability
(d) weighting
Chapter 4 • Validity and Reliability 69

SUPPLY-TYPE ITEMS
DIRECTIONS: Using the test scores below, calculate the correlation coefficient between test
1 and test 2. Place your answers on the lines provided.

Student Test 1 Test 2


A 9 10
B 7 6
C 5 3
D 3 6
E 1 3
F 1 3
G 3 5
H 7 6
I 5 1
J 9 7

6. What is the correlation between test 1 and test 2?

SUPPLY-TYPE ITEMS
DIRECTIONS: Using the information provided in the paragraph that accompanies each of
the following test items, provide an answer on the lines provided.

7. Twenty students are given a 50-item multiple-choice test. To determine the internal
consistency of the test, the test has been split and the score based on odd items and
the score based on the even items for each student has been calculated. The cor-
relation between the two halves is .82.
What is the Spearman-Brown reliability coefficient for the test?
8. Thirty students are given a 100-item multiple-choice test. The standard deviation
of the raw scores is 11 and the Spearman-Brown reliability coefficient for the test
is .93.
What is the standard error of measurement for the test?
9. If one of the students in the question above received a raw score of 83, what would
be the confidence band around the raw score? ____________________

Restricted-Response Essay
10. The authors of your textbook discuss the factors that lower the validity of assess-
ment results. Describe each factor. Confine your answer to one page.

You might also like