0% found this document useful (0 votes)
18 views8 pages

Wilson Chapter 7

Chapter of a book

Uploaded by

edgar.moc.ley
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF or read online on Scribd
0% found this document useful (0 votes)
18 views8 pages

Wilson Chapter 7

Chapter of a book

Uploaded by

edgar.moc.ley
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF or read online on Scribd
‘3a + CHAPTER 6 been a data entry error, or the respondent could have been choosing random responses. Although such interpretations are all possi there is no way to choose among them without extra information, Having the possibility of finding response patterns that are incon- sistent with expectations is one of the most interesting advances in measurement in the last two decades. Although it is bureaucratically annoying to find such respondents, itis important from a perspec- tive of understanding what the data tell us about the respondents. high-fit cases such as that discussed earlier in more detail to ensure that the measurements ly, this creates the possibility of using the pattern to establish a new class of respondents, those for whom we should treat their estimates as suspicious. 6.3 RESOURCES The debate about what is the best measurement model is broad and deep: Just giving a comprehensive list of references would be ex- hausting. Some entry into the literature might be gained from read- ing the following: Andrich (in press), Bock (1977), Brennan (2001), ‘Traub (1997), and Wright (1977). An excellent source for discussion and interpretations of misfit is Wright and Masters (1981). 6.4 EXERCISES AND ACTIVITIES (following on from the exercises and activities in chaps. 1-5) 1. Read some of the sources listed in the “Resources” section, and think about how the ideas expressed there are reflected in issues that have arisen or ones you think mightarisc in developing your instrument. Write down a brief summary of your thoughts. Look back at the GradeMap output from Juan's study. Check for item and person misfit following the procedures outlined in Section 6.2—do you agree ‘Think through the steps outlined e: oping your instrument. Write down notes about your plans. . Share your plans and progress with others—discuss what you and they are succeeding on and what problems have arisen. Chapter 7 Reliability 7.0 CHAPTER OVERVIEW AND KEY CONCEPTS measurement error standard error of measurement reliability coefficient internal consistency reliability test-retest reliability alternate forms reliability inter-rater reliability whether the instrument does, whatever it does, with suffi- cient consistency over individuals for the intended usage— whether there is evidence for the reliability of the instru- ment's usage. Traditionally, re the instrument separate T= aim of this chapter is to describe ways to investigate ty. It is seen here as an integral part of validity, but it is distinguished from the components that make up the next chapter because (a) the reliability of an instru- ment pertains to all of the validity characteristics, and (b) the tradi- tion just mentioned. 139 ia _ CHAPTER 7 7.1 MEASUREMENT ERROR In creating a construct and realizing it through an instrument, the measurer has assumed that each respondent who might be measured has some amount of that construct and the amount is sufficiently mea- surable to be useful. This is what was symbolized by @ in chapters 5 and 6—the respondent's location on the Wright map. When a respon- dent actually gives a response and that response is scored, there are ‘many influences on that score besides @—all of these influences to- gether mean that the estimated 0, labeled 6, will differ from the real 6 for an individual—and that difference is the measurement error—let us call it €. Then we can write: 6= 6+ , which is analogous to the ex- pression X = T + E from true score theory. There are many possible sources of measurement error: (a) there are influences associated with the individual respondent, such as their interest in the topic of the instrument, their mood, and their health; (b) there are influences associated with the conditions under which the instrument is being responded to, such as the temperature of the room, the noisiness of the environment, and the time of day; (c) there are influences associ- ated with the specifics of the instrument, such as the selection of tems and the style of presentation; and (d) there are influences associated with scoring, such as the training of the raters and the consistency of the raters. Note that there is nothing inherently wrong with these ‘errors—itis a normal and expected part of measuring—the term error is used to mean the residual (cf. Eq. 6.4), that is, what is left unex: plained after accounting for what the model (e.g., Eqs. 5.3, 5.5, etc.) has explained. However, the measurer does want to avoid having a lot of such error in the results. ‘There is no exhaustive and final way to classify all these potential errors—that is their nature—they are, by definition, whatever is not being modeled, and hence are not completely classifiable, Neverthe- less, investigating their influence is important because an instru- ment with little or no consistency across the different conditions mentioned earlier will generally not be useful no matter how sound the other parts of the argument are for its validity. One way to con- ‘ceptualize measurement error is to carry out a thought experiment, sometimes called the brafnwasbing analogy. Imagine that the re- spondent forgets that she or he has responded to an item (or set of, items) immediately after making a response (that is the brainwash- ing) and then repeatedly asks them to respond and give them scores RELIABILITY > 147 under all possible combinations of the varied conditions, such as those listed in the previous paragraph—one could then take the ‘mean across ll of these (possibly infinite number of) scores as giving the true location for that respondent (i., 6). Then the variance of the distribution of the observed locations, the variance of 6, is the variance of the errors €. Of course it is unlikely that a respondent would actually forget his or her responses, but this thought experi- ment is one way to interpret the 6, the 6, and the ¢. One index of measurement error for respondents, the standard er- ror of measurement (sem(6)'), has already introduced in chapter 6 (Section 6.2.1). In using this index, the measurer is making use of the brainwashing analogy by assuming each item is a “little instrument” independent of the rest. When using a measurement on an individual respondent, the sem is the most important tool for assessing the use- fulness of that estimate of location. If the sem is too large, the mea- surer will not be able to make intelligible interpretations of the results. For example, in chapter 6 (Section 6.2.1), the 95% confidence interval based on the sem was 2.36 logits wide—about 27% of the width of the entire Wright map from maximum to minimum locations. ‘As was pointed out in the discussion there, this is certainly more infor- ‘mation than one had before getting the data from the respondent. However, itis not accurate for individual usage—to see this, recall that the confidence interval spans (for the second threshold) the range from above “WalkOne” to “WalkMile,” a wide range of physical func- tioning indeed. Thus, this short instrumentiis probably not very useful for accurate clinical diagnosis of individuals, but it may well be useful for initial screening or as a basis for group measures. ‘The sem(8) varies depending on the respondent's location.” This is displayed for the PF-10 example in Fig. 7.1. The relationship is typi- cally a “U" shape, with the minimum near the mean of the item thresholds and the value increasing toward the extremes. The rea- son for this can be seen by looking backat the IRF in Fig. 5.3 and not- ing that the steeper the tangent’ to the IRF at a particular point, the ako called the conditlonal standard error when the focusison the cas: Note that the estimates ofsem(0) and inf) produced by GradeMap are calculated assum- fem parameters are known either than estimated, which ean result in 1a» CHAPTER 7 14 42 1 os os 04 02 ° Standard Error PP-10 logit) FIG. 7.1. The standard error of measurement for the PF-10 instrument (each dot represents a score) more the item can contribute to finding the respondent's location Yet the IRF is steepest at the item’s location, where the probability of response is 0.50. Hence, there isa general conclusion: The closer the respondentis to an item, the more the item can contribute to the es- timation of the respondent ion. Now apply this to the situa- tion for a typical instrument like the PF-10 (see Fig. 5.10). The respondents in the middle will always have more items near them than those at the extremes, hence the sem(6) will be smaller in the item thresholds are distributed in any uniform way over the construct—if the item threshold distribution is bimodal, with a large distance between the modes, then the rela- tionship between the respondent location and the sem(6) can be more complex. Another way to express this relationship is to use the Information Gnf(), which is the reciprocal of the square of the sem(8) (Lord, 1980): Inf(®) = 1/ sem(ay . aay This index is used in calculating the sem(6) for hypothetical instru- ments, which capitalizes on the feature that the information for the whole instrumentis the sum of the information for each item, nf, () (Lord, 1980) RELIABILITY = 143 Inf (®) = YInf,(8) 72) This allows one to hypothesize that the information contribi from a typical item is the mean of the information for the whole in- strument: Inf = inf) /1 73) The equivalent of Fig. 7.1 in terms of information, is shown in Fig. presented, true score theory assumes that the graph in Fig, 7.1 is a horizontal straight line, and hence, so would be the curve in the equivalent of Fig. 7.2. These graphs are useful in designing an instrument. In the case of the PF-10 scale, they show that the most sensitive part of the instru- ment is from approximately ~2.0 to +2.0 logits. If this is indeed the target range of the instrument, then that is a good thing. Loo! back at the Wright map for PF-10 (Fig. 5.10), this corresponds to ap- sroximately the range ofall the first thresholds (O'vs. 1&2) for all the items except for three (i.¢., Bath, OneStair, WalkOne) and eight of the second thresholds too (081 vs. 2). Thus, the instrument's range of maximum sensitivity makes general sense with respect to the item-response categories. However, the distribution of the respon- 35 Information P10 lgits) FIG. 7.2. The information for the PF-10 instrument. CHAPTER 7 dents in Fig. 5.10 shows that many respondents in this sample are above 2.0 logits—hence, the instrument is not functioning optimally for quite a large proportion of this sample. Of course it depends on the ultimate purpose of the instrument—if itis to be used on similar samples as this one, it probably should be augmented with more items up at the VigAct end. Ifitis intended for a sample that is gener ally lower in physical functioning than the current sample, then the Current set of items will likely suffice. If the measurer wanted to look carefully at a sample with low functioning, it would be best to add ‘new items at the low end (near Bath), ‘The shape of the graph is not the only important feature of Figs. 7.1 and 7.2. So too is the average height of the graph. Changing that (down for 7.1 and up for 7.2) can result in increased consistency. The ‘most general way to accomplish this is to increase the number of items (assuming they are of a similar nature as the existing ones). This will almost always decrease the sem(6): The only situation where the mea- surer might expect this not to result in greater consistency would be if the new items were ofa diverse nature. One useful way to roughly ap- Proximate the hypothetical effect of adding similar items is to: (a) choose a location that makes a convenient reference 7.3 to estimate the contribution ofa typical item, (c) a instrument Information using Eq. 7.2, and (d) convert standard error of measurement using Equation 7.1 For example, suppose in the PF-10 example that the measurer wished to know how much the sem(6) could be reduced by tripling the number of items from 10 to 30. The minimum standard error of measurement is 0.56 (hard to judge from Fig. 7.1, but see Appendix 2 for precise values), so the maximum information is 3.19 based on the existing set of 10 items. Thus, the typical information contribution by an item is 0.32. Hence, for 30 similar items, the maximum infor. mation would be approximately 9.60. Then the minimum standard error of measurement would be predicted to be approximately 0.32 ora little more than half (0.57) of the current minimum. Because of the nature of the relationship, there will gencrally be diminishing re- turns on investments in administering more items—such as in this case where tripling the number of items is predicted to cut the sem(@) to about half of what it was originally, A second way to decrease the sem is to increase the standardiza- tion of the conditions under which the instrument is delivered. The that back to the RELIABILITY > 145, iihood of increasing consistency with this strategy, which is his- torically quite common, must be balanced against the possibility of decreasing the validity of the instrument by narrowing the descrip- tive and construct-reference components of the items design. Aclas- sic example of the perils ofthis strategy arose in the area of writing assessment. Here it was discovered that one could increase the con. sistency of scores on a writing test by adding multiple-choice items at the expense of decreasing the actual writing that respondents did. The logical conclusion of that observati from the writing test and use only multiple-choice items. This was in. deed what happened—at one point, many prominent writing tests had no request for writing in them whatsoever. The response from among educators who teach writing was one of horror—students could pass the test without actually writing! After considerable d bate, the situation has swung back to a point where some writing tests now include only a single essay and deliver only a single score which risks taking a student's measure on a sample of topics of size one. Unfortunately, this is not a good situation either—the best resolution lies in finding balance among the competing validity demands, as discussed in the next chapter. 7.2 SUMMARIES OF MEASUREMENT ERROR ‘To develop quality-control indexes of consistency, the traditional ap- Proach has been to find ways to compare how much of the observed variance in respondent locations is attributable to the model as a pro- Portion of the total variance. There are several ways to consider this terms of: @) proportion of variance accounted for by the model, (b) consistency over time, and (c) consistency over different sets of items (Le., different forms). These constitute three different perspectives on measurement error and are termed internal consistency, test—retest, and alternate forms, respectively. Another issue that arises is consis. fency between raters, and this is also discussed, The various summa- ries of measurement error are summarized in Table 7.1 7.2.1 Internal Consistency Coefficients ‘The consistency coefficients described in this section are termed in. ternal consistency coefficients. This is because the basis for their cal 146 CHAPTER 7 TABLE 7.1 1s Measurement Ei Summary of Name Internal consistency indicators Kuder-Richardson 20/21 Used for ue score theory approach ior dhoxomous responses re score theory approach Used fo pytomous responses score theory approach Cronbach's Aipha Separation Used when same respondents are measured again Alternate forms indi Used when there are two sets of items with a ctr Inter Rater consistency indicators Exact agreement proportion ‘compared to a reference. omen proportion _Used when a rater is compared to areference, ‘Alternate forms co culation is the information about variability that is contained in the lata from a single administration of the instrument—effectively they ire investigating the proportion of variance accounted for by the es- imator of a respondent's location. This variance explained formula- tion is familiar to many through its use in analysis of variance and regression methods. Itis also. ‘applicable in the construct-ref- erence approach adopted here: It can be used as a basis for calculat- ration reliability (Wright & Masters, 1981), r. To irst note that the observed total variance of the esti- _ ReUABTUTTY Var where @ is the mean estimated location over the respondents. In the PE-10 example, the total variance is calculated to be 4.47. The vari- ance accounted for by the errors can be calculated as the mean ‘square of the standard errors of measurement (MSE): MSE = fe dsem(0,)" 5) ted to be .67. Then the vari- the difference berween In the PF-10 example, the MSE is calet ance accounted for by the model, Var( these two: Var(®) = Var(6) - Var(é) . (7.6) ‘Thus, for the PF-10 example, this variance works out to be 3.79. The Proportion of variance accounted for by the model, ris then given by 1 = Var(8)/ Var(6) , a7 which gives a reliability coefficient of .85 for the PF-10 scale. Note that this is not the only way to calculate a reliability estimate for these \ta—other possibilities are discussed in chapter 9. This value illustrates one of the shortcomings of cients—the lack of any absolute standards for what is acceptable. Itis, certainly true that a value of 0.90 is better than 0.84, but not so good as 0.95. At what point should one reject the instrument? At what point is it definitely acceptable? There are industry standards in some areas of app) point endorsed a achievement tests used in schools for individual testing, but this level has not been consistently applied. One reason that itis difficult to set a single uniform acceptable standard is that instruments are used for multiple purposes. A better approach is to consider each type of application individually and develop specific standards based on the context. For example, where an instrument is to be used to make a single division into two groups (“pass/lail,” “positive/nega- ive,” etc.), then a reliability coefficient may be quite misleading, us- ing, as it does, data from the entire spectrum of the respondent locations, It may be better to investigate false-positive and false-nega- tive rates in a region near the cut location. 148+ CHAPTER 7 y coefficient isan equivalent of the classical reliability indexes (Kuder-Richardson 20 and 21 [Kuder & Richardson, 193 for dichotomous responses and coefficient alpha [Cronbach, 1951] for polytomous responses), although it is calculated in this case in the metric of the respondent locations rather than in the traditional score metric. One can also calculate the expected score for each per- ing Eq. 6.5 and use that to calculate an “expected score” rel ability using the classical approach, but there is no particular advantage to doing so. 7.2.2 Test-Retest Coefficients As described in the previous section, there are many sources of mea surement error that lic outside a single administration of a 7 ‘ment. Each such source could be the basis for calculating a different coeffic y coefficient. In a test-retest reliability coef- ficient, the measurer first arranges to have the same respondents give responses to the questions twice, then the reliability coefficient is calculated simply as the correlation between the two sets of re- spondent locations. (In the classic approach, the same approach is, applied to the raw scores.) In observation of the brainwashing analogy, the test and retest should be: e by remembering the first, but are genuinely responding to each item anew. This may be di cult to achieve for some sorts of complex items, which may be q gether for it to be reasonable to assume that there has beet change. Obviously, this form of the reliability index will work better where a stable construct is being measured with forgettable items, as compared with a less stable construct being measured with memorable items. “Many would say tha forgetable tems were not good items, but here is a case where they ae quite useful 149 7.2.3 Alternate Forms Coefficients Another type of reliability coefficients the alternate forms reliability coefficient. With this coefficient, the measurer arranges to develop ‘two sets of items for the instrument, each following the same series of steps through the four building blocks as in chapters 2 through 5 The two alternate copies of the instrument are administered and cal- ated, and then the two sets of locations are correlated to produce the alternate forms rel coefficient. This coefficient is particu- larly useful as a way to check that the use of the four building blocks in chapters 2 through 5 has indeed resulted in an instrument that represents the construct in a content-stable way. This approach can be used for more than just calculating a reliability coefficient. For ex- ample, it can be used to investigate the robustness of construct valid- evidence: When linked using the technique in Appendix 9A, the validity results can be compared using a Wright map. Other classical consistency indexes have also been developed, and they have their equivalents in the construct modeling approach. For example, in the so-called split-balves y instrument is split into two different (nonintersecting) but similar parts, and the correl The adjustment is a special case of the Spearman-Brown formu! Lr (78) where / is the ratio of the number of items in the hypothetical testto ifthe number of items were the construct modeling ap- ns of each correlate the two and make the same adjustment. ‘These reliability coefficients can be calculated separately, and the results are quite useful for understanding the consistency of the in- 's measures across each of the practice, such influences will occur simultaneously, and it would be better to have ways to investigate the influences simultaneously. Such methods have indeed been developed: (a) generalizability the- {so + GUAPTER 7 _ ory (€.g., Shavelson & Webb, 1991) is an expansion of the analysis of variance (ANOVA) approach mentioned earlier, and (b) facets analy- (Linacre, 1989; Wilson & Hoskens, 2001) is an expansion of the item-response modeling approach introduced previously. 7.3. INTER-RATER CONSISTENCY Where the respondents’ responses are to be scored by raters, an- other source of measurement error occurs—inconsistencies be- tween the raters. There are many forms that such inconsistency can take: (a) there are raters who do not fully accommodate the training, and so never apply the scoring guides in a correct way; (b) there are differences in rater severity—that is, some raters tend to score the same responses higher or lower than others; (c) there are differences, in raters use of the score categories, such as raters who use the ex- tremes more often than others or not as often as others, as well as more complex patterns; (d) there are raters who exhibit “halo ef fects”—thatis, their scores are affected by recent scores; (¢) there are raters who drift in their severity, their tendency to use extreme scores, and so on; and (f) there are raters who are inconsistent with themselves for a variety of reasons. ‘The most important steps to take to reduce rater inconsistency are: a program of sound rater training, and a monitoring system that helps both the administrators and raters know that they are keeping, on track. A good training program includes: 1. background information on the concepts involved in the con- struct; 2. an opportunity for the raters to examine and rate a large num- ber and wide range of responses, including both examples that are clearly within a category and examples that are not clear; 3. opportunities for the raters to discuss their ratings on specific pieces of work, and justifications for those ratings, with their fellow raters; 4, systematic feedback to the raters telling them how well they are rating prejudged res; 5a system of rater ther results in a rater being accepted as calibrated or being returned for further train- ing and/or support. Although a system like that just described constitutes a sound foundation fora rater, ithas been found that they can soon drift awe fromeven asound ing (see €.g., Wilson & Case, 2000). To de: with this probl is important to have a monitoring program in place also. There are essentially three ways to monitor the work of the raters: (a) scatter prejudged responses among them, (b) re-rate (by experts) some of their ratings, and (c) compare the records of ( of their) ratings to the ratings these in any detail is beyond the scope of this volume (see, e.g., Wilson & C: 2000, for some specific procedures). ‘Once the ratings have been made, they need to be summarized in ways that help the measurer see how consistent the raters have been. ‘There are ways to carry this out using the construct modeling ap- proach (see ¢.g., Wilson & Case, 2000) and also using generaliz- ability theory, but they are beyond the scope of this volume, so more elementary methods are described. To apply these elementary meth- ods, the first step is to gather a sample of ratings based on the same responses for the raters under investigation. Then they are either (a) compared to the ratings of an expert (or panel of experts) ot, where that is not available, (b) compared to the mean ratings for the group. In either case, these are referred to as the reference ratings. Acomprehensive way to display the consistency of a rater with the reference ratings is shown in Table 7.2. In this hypothetica there are four score levels possible. The ratings for the first column, and the reference ratings are displayed at the heads ofthe next four columns. The number of cases of each possible pait is recorded in the main body of the table—n,, being the number of re- sponses scored s by rater r and t by the reference rating. The appro- TABLE 7.2 Layout of Data for Checking Rater Consistency ater’ Reference _ Ratings 1+ CHAPTER 7 priate marginals are also recorded and labeled usin, whether the row or column (or both) are summed. A directly inter- pretable index of agreement is the proportion of exact agreement— the proportion of responses in the leading diagonal of entries n, Pesact = zn [es (7.9) In cases where one wanted to control for the possibility that the matching scores might have arisen by chance, an alternative index called Coben's kappa is available (Cohen, 1960). A less rigorous index of agreement is the proportion of responses in the same or adjacent ‘categories. This is not recommended when the number of categories is small (as is the case in Table 7.1) because it can lead to overpositive interpretations. The table can also be examined for various patterns: (@) asymmetry of the diagonals would indicate differences in severity, and (b) relatively larger or smaller numbers at either end could indi- cate a tendency to the extremes or the middle. The table can also be examined with chi-square methods or logilinear analysis (sce, ¢.8,, for independence and other patterns. Note that correlation coefficient can be a misleading way to examine the consis- disguise differences in harshness between the raters. 7.4. RESOURCES ic perspective on measure- y can be found in Cronbach (1990). Included ‘examples of correlation-based reliability coeffi t-retest and alternate forms, as well as an explanation of how to calculate a correlation coefficient and a discussion on its in- terpretation. Further discussion of the interpretation of errors under the item-response modeling approach can be found in Lord (1980), ‘Wright and Stone (1979), and Wright and Masters (1981). 7.5 EXERCISES AND ACTIVITIES (following on from the exercises and activities in chaps. 1-6) RELIABILITY > 133 Look back at the GradeMap output you generated from Juan's data. Check the standard errors of measurement for the stu. dents in his study. Do they display the “U-shape” pattern men. tioned earlier? Are they sufficiently small? Locate the separation reliability and Cronbach's alpha in the GradeMap output. How do they compare? Try to locate a data set containing either test-retest or alternate forms data and calculate a correlation coefficient to interpret as a reliability coefficient. Write down your plan for collecting reliability information about your instrument. Think through the steps outlined previously in the context of Geveloping your instrument, and write down notes about your plans, Share your plans and progress with others—discuss what you and they are succeeding on, and what problems have arisen.

You might also like