We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF or read online on Scribd
‘3a + CHAPTER 6
been a data entry error, or the respondent could have been choosing
random responses. Although such interpretations are all possi
there is no way to choose among them without extra information,
Having the possibility of finding response patterns that are incon-
sistent with expectations is one of the most interesting advances in
measurement in the last two decades. Although it is bureaucratically
annoying to find such respondents, itis important from a perspec-
tive of understanding what the data tell us about the respondents.
high-fit cases such as that discussed earlier
in more detail to ensure that the measurements
ly, this creates the possibility of using the pattern
to establish a new class of respondents, those for
whom we should treat their estimates as suspicious.
6.3 RESOURCES
The debate about what is the best measurement model is broad and
deep: Just giving a comprehensive list of references would be ex-
hausting. Some entry into the literature might be gained from read-
ing the following: Andrich (in press), Bock (1977), Brennan (2001),
‘Traub (1997), and Wright (1977). An excellent source for discussion
and interpretations of misfit is Wright and Masters (1981).
6.4 EXERCISES AND ACTIVITIES
(following on from the exercises and activities in chaps. 1-5)
1. Read some of the sources listed in the “Resources” section, and
think about how the ideas expressed there are reflected in issues
that have arisen or ones you think mightarisc in developing your
instrument. Write down a brief summary of your thoughts.
Look back at the GradeMap output from Juan's study. Check for
item and person misfit following the procedures outlined in
Section 6.2—do you agree
‘Think through the steps outlined e:
oping your instrument. Write down notes about your plans.
. Share your plans and progress with others—discuss what you
and they are succeeding on and what problems have arisen.
Chapter 7
Reliability
7.0 CHAPTER OVERVIEW AND KEY CONCEPTS
measurement error
standard error of measurement
reliability coefficient
internal consistency reliability
test-retest reliability
alternate forms reliability
inter-rater reliability
whether the instrument does, whatever it does, with suffi-
cient consistency over individuals for the intended usage—
whether there is evidence for the reliability of the instru-
ment's usage. Traditionally, re
the instrument separate
T= aim of this chapter is to describe ways to investigate
ty. It is seen here as an integral
part of validity, but it is distinguished from the components that
make up the next chapter because (a) the reliability of an instru-
ment pertains to all of the validity characteristics, and (b) the tradi-
tion just mentioned.
139ia _ CHAPTER 7
7.1 MEASUREMENT ERROR
In creating a construct and realizing it through an instrument, the
measurer has assumed that each respondent who might be measured
has some amount of that construct and the amount is sufficiently mea-
surable to be useful. This is what was symbolized by @ in chapters 5
and 6—the respondent's location on the Wright map. When a respon-
dent actually gives a response and that response is scored, there are
‘many influences on that score besides @—all of these influences to-
gether mean that the estimated 0, labeled 6, will differ from the real 6
for an individual—and that difference is the measurement error—let
us call it €. Then we can write: 6= 6+ , which is analogous to the ex-
pression X = T + E from true score theory. There are many possible
sources of measurement error: (a) there are influences associated
with the individual respondent, such as their interest in the topic of
the instrument, their mood, and their health; (b) there are influences
associated with the conditions under which the instrument is being
responded to, such as the temperature of the room, the noisiness of
the environment, and the time of day; (c) there are influences associ-
ated with the specifics of the instrument, such as the selection of tems
and the style of presentation; and (d) there are influences associated
with scoring, such as the training of the raters and the consistency of
the raters. Note that there is nothing inherently wrong with these
‘errors—itis a normal and expected part of measuring—the term error
is used to mean the residual (cf. Eq. 6.4), that is, what is left unex:
plained after accounting for what the model (e.g., Eqs. 5.3, 5.5, etc.)
has explained. However, the measurer does want to avoid having a lot
of such error in the results.
‘There is no exhaustive and final way to classify all these potential
errors—that is their nature—they are, by definition, whatever is not
being modeled, and hence are not completely classifiable, Neverthe-
less, investigating their influence is important because an instru-
ment with little or no consistency across the different conditions
mentioned earlier will generally not be useful no matter how sound
the other parts of the argument are for its validity. One way to con-
‘ceptualize measurement error is to carry out a thought experiment,
sometimes called the brafnwasbing analogy. Imagine that the re-
spondent forgets that she or he has responded to an item (or set of,
items) immediately after making a response (that is the brainwash-
ing) and then repeatedly asks them to respond and give them scores
RELIABILITY > 147
under all possible combinations of the varied conditions, such as
those listed in the previous paragraph—one could then take the
‘mean across ll of these (possibly infinite number of) scores as giving
the true location for that respondent (i., 6). Then the variance of
the distribution of the observed locations, the variance of 6, is the
variance of the errors €. Of course it is unlikely that a respondent
would actually forget his or her responses, but this thought experi-
ment is one way to interpret the 6, the 6, and the ¢.
One index of measurement error for respondents, the standard er-
ror of measurement (sem(6)'), has already introduced in chapter 6
(Section 6.2.1). In using this index, the measurer is making use of the
brainwashing analogy by assuming each item is a “little instrument”
independent of the rest. When using a measurement on an individual
respondent, the sem is the most important tool for assessing the use-
fulness of that estimate of location. If the sem is too large, the mea-
surer will not be able to make intelligible interpretations of the
results. For example, in chapter 6 (Section 6.2.1), the 95% confidence
interval based on the sem was 2.36 logits wide—about 27% of the
width of the entire Wright map from maximum to minimum locations.
‘As was pointed out in the discussion there, this is certainly more infor-
‘mation than one had before getting the data from the respondent.
However, itis not accurate for individual usage—to see this, recall that
the confidence interval spans (for the second threshold) the range
from above “WalkOne” to “WalkMile,” a wide range of physical func-
tioning indeed. Thus, this short instrumentiis probably not very useful
for accurate clinical diagnosis of individuals, but it may well be useful
for initial screening or as a basis for group measures.
‘The sem(8) varies depending on the respondent's location.” This
is displayed for the PF-10 example in Fig. 7.1. The relationship is typi-
cally a “U" shape, with the minimum near the mean of the item
thresholds and the value increasing toward the extremes. The rea-
son for this can be seen by looking backat the IRF in Fig. 5.3 and not-
ing that the steeper the tangent’ to the IRF at a particular point, the
ako called the conditlonal standard error when the focusison the cas:
Note that the estimates ofsem(0) and inf) produced by GradeMap are calculated assum-
fem parameters are known either than estimated, which ean result in1a» CHAPTER 7
14
42
1
os
os
04
02
°
Standard Error
PP-10 logit)
FIG. 7.1. The standard error of measurement for the PF-10 instrument
(each dot represents a score)
more the item can contribute to finding the respondent's location
Yet the IRF is steepest at the item’s location, where the probability of
response is 0.50. Hence, there isa general conclusion: The closer the
respondentis to an item, the more the item can contribute to the es-
timation of the respondent ion. Now apply this to the situa-
tion for a typical instrument like the PF-10 (see Fig. 5.10). The
respondents in the middle will always have more items near them
than those at the extremes, hence the sem(6) will be smaller in the
item thresholds are distributed in any
uniform way over the construct—if the item threshold distribution
is bimodal, with a large distance between the modes, then the rela-
tionship between the respondent location and the sem(6) can be
more complex.
Another way to express this relationship is to use the Information
Gnf(), which is the reciprocal of the square of the sem(8) (Lord,
1980):
Inf(®) = 1/ sem(ay . aay
This index is used in calculating the sem(6) for hypothetical instru-
ments, which capitalizes on the feature that the information for the
whole instrumentis the sum of the information for each item, nf, ()
(Lord, 1980)
RELIABILITY = 143
Inf (®) = YInf,(8) 72)
This allows one to hypothesize that the information contribi
from a typical item is the mean of the information for the whole in-
strument:
Inf = inf) /1 73)
The equivalent of Fig. 7.1 in terms of information, is shown in Fig.
presented, true score theory assumes
that the graph in Fig, 7.1 is a horizontal straight line, and hence, so
would be the curve in the equivalent of Fig. 7.2.
These graphs are useful in designing an instrument. In the case of
the PF-10 scale, they show that the most sensitive part of the instru-
ment is from approximately ~2.0 to +2.0 logits. If this is indeed the
target range of the instrument, then that is a good thing. Loo!
back at the Wright map for PF-10 (Fig. 5.10), this corresponds to ap-
sroximately the range ofall the first thresholds (O'vs. 1&2) for all the
items except for three (i.¢., Bath, OneStair, WalkOne) and eight of
the second thresholds too (081 vs. 2). Thus, the instrument's range
of maximum sensitivity makes general sense with respect to the
item-response categories. However, the distribution of the respon-
35
Information
P10 lgits)
FIG. 7.2. The information for the PF-10 instrument.CHAPTER 7
dents in Fig. 5.10 shows that many respondents in this sample are
above 2.0 logits—hence, the instrument is not functioning optimally
for quite a large proportion of this sample. Of course it depends on
the ultimate purpose of the instrument—if itis to be used on similar
samples as this one, it probably should be augmented with more
items up at the VigAct end. Ifitis intended for a sample that is gener
ally lower in physical functioning than the current sample, then the
Current set of items will likely suffice. If the measurer wanted to look
carefully at a sample with low functioning, it would be best to add
‘new items at the low end (near Bath),
‘The shape of the graph is not the only important feature of Figs. 7.1
and 7.2. So too is the average height of the graph. Changing that
(down for 7.1 and up for 7.2) can result in increased consistency. The
‘most general way to accomplish this is to increase the number of items
(assuming they are of a similar nature as the existing ones). This will
almost always decrease the sem(6): The only situation where the mea-
surer might expect this not to result in greater consistency would be if
the new items were ofa diverse nature. One useful way to roughly ap-
Proximate the hypothetical effect of adding similar items is to: (a)
choose a location that makes a convenient reference
7.3 to estimate the contribution ofa typical item, (c) a
instrument Information using Eq. 7.2, and (d) convert
standard error of measurement using Equation 7.1
For example, suppose in the PF-10 example that the measurer
wished to know how much the sem(6) could be reduced by tripling
the number of items from 10 to 30. The minimum standard error of
measurement is 0.56 (hard to judge from Fig. 7.1, but see Appendix 2
for precise values), so the maximum information is 3.19 based on the
existing set of 10 items. Thus, the typical information contribution
by an item is 0.32. Hence, for 30 similar items, the maximum infor.
mation would be approximately 9.60. Then the minimum standard
error of measurement would be predicted to be approximately 0.32
ora little more than half (0.57) of the current minimum. Because of
the nature of the relationship, there will gencrally be diminishing re-
turns on investments in administering more items—such as in this
case where tripling the number of items is predicted to cut the
sem(@) to about half of what it was originally,
A second way to decrease the sem is to increase the standardiza-
tion of the conditions under which the instrument is delivered. The
that back to the
RELIABILITY > 145,
iihood of increasing consistency with this strategy, which is his-
torically quite common, must be balanced against the possibility of
decreasing the validity of the instrument by narrowing the descrip-
tive and construct-reference components of the items design. Aclas-
sic example of the perils ofthis strategy arose in the area of writing
assessment. Here it was discovered that one could increase the con.
sistency of scores on a writing test by adding multiple-choice items at
the expense of decreasing the actual writing that respondents did.
The logical conclusion of that observati
from the writing test and use only multiple-choice items. This was in.
deed what happened—at one point, many prominent writing tests
had no request for writing in them whatsoever. The response from
among educators who teach writing was one of horror—students
could pass the test without actually writing! After considerable d
bate, the situation has swung back to a point where some writing
tests now include only a single essay and deliver only a single score
which risks taking a student's measure on a sample of topics of size
one. Unfortunately, this is not a good situation either—the best
resolution lies in finding balance among the competing validity
demands, as discussed in the next chapter.
7.2 SUMMARIES OF MEASUREMENT ERROR
‘To develop quality-control indexes of consistency, the traditional ap-
Proach has been to find ways to compare how much of the observed
variance in respondent locations is attributable to the model as a pro-
Portion of the total variance. There are several ways to consider this
terms of: @) proportion of variance accounted for by the model, (b)
consistency over time, and (c) consistency over different sets of items
(Le., different forms). These constitute three different perspectives on
measurement error and are termed internal consistency, test—retest,
and alternate forms, respectively. Another issue that arises is consis.
fency between raters, and this is also discussed, The various summa-
ries of measurement error are summarized in Table 7.1
7.2.1 Internal Consistency Coefficients
‘The consistency coefficients described in this section are termed in.
ternal consistency coefficients. This is because the basis for their cal146 CHAPTER 7
TABLE 7.1
1s Measurement Ei
Summary of
Name
Internal consistency indicators
Kuder-Richardson 20/21
Used for ue score theory approach
ior dhoxomous responses
re score theory approach
Used fo pytomous responses
score theory approach
Cronbach's Aipha
Separation
Used when same respondents are measured
again
Alternate forms indi
Used when there are two sets of items
with a ctr
Inter Rater consistency indicators
Exact agreement proportion ‘compared to a reference.
omen proportion _Used when a rater is compared to areference,
‘Alternate forms co
culation is the information about variability that is contained in the
lata from a single administration of the instrument—effectively they
ire investigating the proportion of variance accounted for by the es-
imator of a respondent's location. This variance explained formula-
tion is familiar to many through its use in analysis of variance and
regression methods. Itis also. ‘applicable in the construct-ref-
erence approach adopted here: It can be used as a basis for calculat-
ration reliability (Wright & Masters, 1981), r. To
irst note that the observed total variance of the esti-
_ ReUABTUTTY Var
where @ is the mean estimated location over the respondents. In the
PE-10 example, the total variance is calculated to be 4.47. The vari-
ance accounted for by the errors can be calculated as the mean
‘square of the standard errors of measurement (MSE):
MSE = fe dsem(0,)" 5)
ted to be .67. Then the vari-
the difference berween
In the PF-10 example, the MSE is calet
ance accounted for by the model, Var(
these two:
Var(®) = Var(6) - Var(é) . (7.6)
‘Thus, for the PF-10 example, this variance works out to be 3.79. The
Proportion of variance accounted for by the model, ris then given by
1 = Var(8)/ Var(6) , a7
which gives a reliability coefficient of .85 for the PF-10 scale. Note
that this is not the only way to calculate a reliability estimate for these
\ta—other possibilities are discussed in chapter 9.
This value illustrates one of the shortcomings of
cients—the lack of any absolute standards for what is acceptable. Itis,
certainly true that a value of 0.90 is better than 0.84, but not so good
as 0.95. At what point should one reject the instrument? At what
point is it definitely acceptable? There are industry standards in
some areas of app)
point endorsed a
achievement tests used in schools for individual testing, but this
level has not been consistently applied. One reason that itis difficult
to set a single uniform acceptable standard is that instruments are
used for multiple purposes. A better approach is to consider each
type of application individually and develop specific standards based
on the context. For example, where an instrument is to be used to
make a single division into two groups (“pass/lail,” “positive/nega-
ive,” etc.), then a reliability coefficient may be quite misleading, us-
ing, as it does, data from the entire spectrum of the respondent
locations, It may be better to investigate false-positive and false-nega-
tive rates in a region near the cut location.148+ CHAPTER 7
y coefficient isan equivalent of the classical reliability
indexes (Kuder-Richardson 20 and 21 [Kuder & Richardson, 193
for dichotomous responses and coefficient alpha [Cronbach, 1951]
for polytomous responses), although it is calculated in this case in
the metric of the respondent locations rather than in the traditional
score metric. One can also calculate the expected score for each per-
ing Eq. 6.5 and use that to calculate an “expected score” rel
ability using the classical approach, but there is no particular
advantage to doing so.
7.2.2 Test-Retest Coefficients
As described in the previous section, there are many sources of mea
surement error that lic outside a single administration of a 7
‘ment. Each such source could be the basis for calculating a different
coeffic
y coefficient. In a test-retest reliability coef-
ficient, the measurer first arranges to have the same respondents
give responses to the questions twice, then the reliability coefficient
is calculated simply as the correlation between the two sets of re-
spondent locations. (In the classic approach, the same approach is,
applied to the raw scores.)
In observation of the brainwashing analogy, the test and retest
should be:
e by remembering the first,
but are genuinely responding to each item anew. This may be di
cult to achieve for some sorts of complex items, which may be q
gether for it to be reasonable to assume that there has beet
change. Obviously, this form of the reliability index will work better
where a stable construct is being measured with forgettable items,
as compared with a less stable construct being measured with
memorable items.
“Many would say tha forgetable tems were not good items, but here is a case where they
ae quite useful
149
7.2.3 Alternate Forms Coefficients
Another type of reliability coefficients the alternate forms reliability
coefficient. With this coefficient, the measurer arranges to develop
‘two sets of items for the instrument, each following the same series
of steps through the four building blocks as in chapters 2 through 5
The two alternate copies of the instrument are administered and cal-
ated, and then the two sets of locations are correlated to produce
the alternate forms rel coefficient. This coefficient is particu-
larly useful as a way to check that the use of the four building blocks
in chapters 2 through 5 has indeed resulted in an instrument that
represents the construct in a content-stable way. This approach can
be used for more than just calculating a reliability coefficient. For ex-
ample, it can be used to investigate the robustness of construct valid-
evidence: When linked using the technique in Appendix 9A, the
validity results can be compared using a Wright map.
Other classical consistency indexes have also been developed,
and they have their equivalents in the construct modeling approach.
For example, in the so-called split-balves y
instrument is split into two different (nonintersecting) but similar
parts, and the correl
The adjustment is a special case of the Spearman-Brown formu!
Lr
(78)
where / is the ratio of the number of items in the hypothetical testto
ifthe number of items were
the construct modeling ap-
ns of each
correlate the two and make the same adjustment.
‘These reliability coefficients can be calculated separately, and the
results are quite useful for understanding the consistency of the in-
's measures across each of the
practice, such influences will occur simultaneously, and it would be
better to have ways to investigate the influences simultaneously.
Such methods have indeed been developed: (a) generalizability the-{so + GUAPTER 7 _
ory (€.g., Shavelson & Webb, 1991) is an expansion of the analysis of
variance (ANOVA) approach mentioned earlier, and (b) facets analy-
(Linacre, 1989; Wilson & Hoskens, 2001) is an expansion of the
item-response modeling approach introduced previously.
7.3. INTER-RATER CONSISTENCY
Where the respondents’ responses are to be scored by raters, an-
other source of measurement error occurs—inconsistencies be-
tween the raters. There are many forms that such inconsistency can
take: (a) there are raters who do not fully accommodate the training,
and so never apply the scoring guides in a correct way; (b) there are
differences in rater severity—that is, some raters tend to score the
same responses higher or lower than others; (c) there are differences,
in raters use of the score categories, such as raters who use the ex-
tremes more often than others or not as often as others, as well as
more complex patterns; (d) there are raters who exhibit “halo ef
fects”—thatis, their scores are affected by recent scores; (¢) there are
raters who drift in their severity, their tendency to use extreme
scores, and so on; and (f) there are raters who are inconsistent with
themselves for a variety of reasons.
‘The most important steps to take to reduce rater inconsistency
are: a program of sound rater training, and a monitoring system that
helps both the administrators and raters know that they are keeping,
on track. A good training program includes:
1. background information on the concepts involved in the con-
struct;
2. an opportunity for the raters to examine and rate a large num-
ber and wide range of responses, including both examples that
are clearly within a category and examples that are not clear;
3. opportunities for the raters to discuss their ratings on specific
pieces of work, and justifications for those ratings, with their
fellow raters;
4, systematic feedback to the raters telling them how well they are
rating prejudged res;
5a system of rater ther results in a rater
being accepted as calibrated or being returned for further train-
ing and/or support.
Although a system like that just described constitutes a sound
foundation fora rater, ithas been found that they can soon drift awe
fromeven asound ing (see €.g., Wilson & Case, 2000). To de:
with this probl is important to have a monitoring program in
place also. There are essentially three ways to monitor the work of
the raters: (a) scatter prejudged responses among them, (b) re-rate
(by experts) some of their ratings, and (c) compare the records of (
of their) ratings to the ratings these in any
detail is beyond the scope of this volume (see, e.g., Wilson & C:
2000, for some specific procedures).
‘Once the ratings have been made, they need to be summarized in
ways that help the measurer see how consistent the raters have been.
‘There are ways to carry this out using the construct modeling ap-
proach (see ¢.g., Wilson & Case, 2000) and also using generaliz-
ability theory, but they are beyond the scope of this volume, so more
elementary methods are described. To apply these elementary meth-
ods, the first step is to gather a sample of ratings based on the same
responses for the raters under investigation. Then they are either (a)
compared to the ratings of an expert (or panel of experts) ot, where
that is not available, (b) compared to the mean ratings for the group.
In either case, these are referred to as the reference ratings.
Acomprehensive way to display the consistency of a rater with the
reference ratings is shown in Table 7.2. In this hypothetica
there are four score levels possible. The ratings for
the first column, and the reference ratings are displayed at the heads
ofthe next four columns. The number of cases of each possible pait
is recorded in the main body of the table—n,, being the number of re-
sponses scored s by rater r and t by the reference rating. The appro-
TABLE 7.2
Layout of Data for Checking Rater Consistency
ater’ Reference _ Ratings1+ CHAPTER 7
priate marginals are also recorded and labeled usin,
whether the row or column (or both) are summed. A directly inter-
pretable index of agreement is the proportion of exact agreement—
the proportion of responses in the leading diagonal of entries n,
Pesact = zn [es (7.9)
In cases where one wanted to control for the possibility that the
matching scores might have arisen by chance, an alternative index
called Coben's kappa is available (Cohen, 1960). A less rigorous index
of agreement is the proportion of responses in the same or adjacent
‘categories. This is not recommended when the number of categories
is small (as is the case in Table 7.1) because it can lead to overpositive
interpretations. The table can also be examined for various patterns:
(@) asymmetry of the diagonals would indicate differences in severity,
and (b) relatively larger or smaller numbers at either end could indi-
cate a tendency to the extremes or the middle. The table can also be
examined with chi-square methods or logilinear analysis (sce, ¢.8,,
for independence and other patterns. Note that
correlation coefficient can be a misleading way to examine the consis-
disguise differences in harshness between the raters.
7.4. RESOURCES
ic perspective on measure-
y can be found in Cronbach (1990). Included
‘examples of correlation-based reliability coeffi
t-retest and alternate forms, as well as an explanation
of how to calculate a correlation coefficient and a discussion on its in-
terpretation. Further discussion of the interpretation of errors under
the item-response modeling approach can be found in Lord (1980),
‘Wright and Stone (1979), and Wright and Masters (1981).
7.5 EXERCISES AND ACTIVITIES
(following on from the exercises and activities in chaps. 1-6)
RELIABILITY > 133
Look back at the GradeMap output you generated from Juan's
data. Check the standard errors of measurement for the stu.
dents in his study. Do they display the “U-shape” pattern men.
tioned earlier? Are they sufficiently small?
Locate the separation reliability and Cronbach's alpha in the
GradeMap output. How do they compare?
Try to locate a data set containing either test-retest or alternate
forms data and calculate a correlation coefficient to interpret as
a reliability coefficient.
Write down your plan for collecting reliability information
about your instrument.
Think through the steps outlined previously in the context of
Geveloping your instrument, and write down notes about your
plans,
Share your plans and progress with others—discuss what you
and they are succeeding on, and what problems have arisen.