Designing Effective Measurement Items
Designing Effective Measurement Items
This chapter and the next tell the story about designing the item: the ways that the
measurer uses to stimulate the respondent to produce products and behaviors that can be
used as the basis for observations about the location of the respondents on the construct
in question. By “observation,” something more is meant than just seeing something and
recording it, or remembering it, or jotting down notes about it. In referring to it as a
special sort of observation that is generically called an item, it means that there exists (a)
a procedure, or design, which allows the observations to be made under a set of standard
conditions that span the intended range of the item contexts, and (b) a procedure for
classifying those observations into a set of standard categories. The first part is the topic
of this chapter, and the second is the topic of the next chapter. The measurement
instrument is then a set of these procedures (i.e., a set of items). First, the idea of an item
is developed in the section immediately below, and some typical types of items are
discussed. Then, a typology of items is introduced that is designed to show connections
among many different sorts of items. This is followed by a discussion of the important
facets of the whole-instrument design. After that, there are section devoted to (a) the
inclusion of a fairness perspective into the items design and (b) the incorporation of
human judgement into the items design. Finally, ways of developing items are discussed
in the last section. In the next chapter (Chapter 4), this understanding about the nature of
the item is complemented to include the way that the observations are recorded and
categorized.
Key concepts: item formats, participant observation, topics guide, constructed responses,
selected responses; facets of the items design, the construct facet, secondary facets;
universal design
Often the first inkling of an item comes in the form of an idea about a way to
reveal a particular characteristic (construct) of a respondent. The inkling can be quite
informal: A remark in a conversation, the way a student describes what they understand
about something, a question that prompts an argument, a particularly pleasing piece of
art, a newspaper article, a patient’s or client’s symptoms. The specific way in which the
measurer prompts an informative response from a respondent is crucial to the value of the
resulting measurements. In fact, in many, if not most, cases, the construct itself will not
be clearly defined until a relatively large set of items has been created and tried out with
respondents. Each new situation brings about the possibility of developing new and
different sorts of items or of adapting old ones.
1
Across many fields and across many topic areas within those fields, a rich variety
of types of items have been developed to deal with many different constructs and
situations. We have already seen two very different formats. In Chapter 1 (in Example
1) a constructed response type of item was closely examined: the Data Modeling MoV
“Piano width” item. Many examples of responses to that item were included in Table 1.1.
The LPS science argumentation assessment (Example 4), and the CUE interview
(Example 7) were additional examples of the constructed response type. In Chapter 2 we
branched out and examined selected response items as well: the PF-10 health survey
(Example 6) asked the question “Does your health now limit you in these activities?”
with respect to a range of physical activities but restricted the responses to a selected
response among: “Yes, limited a lot”, “Yes, limited a little”, and “No, not limited at all.”
The GEB Scale (Chapter 2, Example 3) was another example of the selected response
type, as it is an example of the familiar Likert-style item. These are examples of two ends
of a range of item formats that stretch from very open-format constructed responses to
very closed-format selected responses. In the following section we present a range of
item types that span across these two, and beyond (see Figure 3.5). Many other types of
items exist (e.g., see Brookheart & Nitko, 2018, for a large assortment from educational
assessment), and the measurer should be aware of both the specific types of items that
have previously been used in the specific area in which an instrument is being developed,
as well as item types that have been used in other related contexts.
Probably the most common type of item in the experience of most people is the
general constructed response item format that is commonly used in school classrooms
and many other settings around the world. The format can be expressed orally or in
writing (by hand or typing) or in other forms, such as concrete products, active
performances, or actions taken in a digital environment. The length of the response can
vary from a single number or word to lengthy essays, complex proofs, interviews,
extended performances or multi-part products. The item can be one that is produced
extemporaneously by the measurer (say, the teacher, or other professional) or it can be
the result of an extensive developmental process. This format is also used outside of
educational settings including workplaces, social settings, and in the everyday
conversational interchanges that we all experience. Typical sub-forms are the essay, the
brief demonstration, the product, and the short-answer format.
2
that the most commonly experienced format is not the most commonly published format.
As will become clear to the reader as they advance through the next several chapters, the
view developed here is that the constructed response format is the more basic format, and
the selected response format can be seen as an adapted version of it.
The relationship of the item to its construct is its most fundamental relationship 1.
Typically, the item is but one of many (often one from an infinite set) that could be used
to measure the construct. Paul Ramsden and his colleagues, writing about the assessment
of achievement in high school physics noted:
Educators are interested in how well students understand speed, distance
and time, not in what they know about runners or powerboats or people
walking along corridors2. Paradoxically, however, there is no other way
of describing and testing understanding than through such specific
examples. (Ramsden et al, 1993, p. 312; footnote added)
Similarly, consider the health measurement (PF-10) example described above. Here the
specific questions that are used are clearly neither necessary for defining the construct,
nor are they sufficient to encompass all the possible meanings of the concept of physical
functioning. Thus, the task of the measurer is to choose a finite set of items that does
indeed represent the construct in some reasonable way. As Ramsden hints, this is not the
straightforward task one might think it to be, on initial consideration.
Sometimes a measurer will feel the temptation to seek the “one true task,” the
“authentic item”, the single observation that will supply the mother lode of evidence
about the construct. Unfortunately, this misunderstanding, common amongst beginning
measurers, is founded on a failure to fully consider the need to establish sufficient levels
of validity and reliability for the instrument. Where one wishes to represent a wide range
of contexts in an instrument, it is better to have more items rather than less—this is
because (a) the instrument can then sample more of the content of a construct and more
of the situations where a student’s location on the construct might be displayed (see
Chapter 8 for more on this), and (b) because it can then generate more bits of information
about how a respondent stands with respect to the construct, which will yield greater
precision (see Chapter 7 for more on this). And both requirements need to be satisfied
within the time and cost limitations imposed on the measuring context.
The items design is the second building block in the Bear Assessment System
(BAS). It has already been introduced, lightly, in Chapter 1, and its relationship to the
other building blocks was illustrated there too—see Figure 3.1. It is the main focus of
this chapter.
1
Although an item may also be related to several different constructs, especially if the outcome space (see
Chapter 4) is different for each.
2
I.e., these are typical situations involved in high school physics items.
3
Figure 3.1 The Items Design, the second building block in the BEAR
Assessment System (BAS)
4
intended to relate items to a construct. This facet provides interpretational levels within
the construct—and hence it is called the construct facet. For example, the construct facet
in the Data Modeling MoV construct is provided in Figure 1.12. Thus, the construct facet
is essentially the content of the construct map—where an instrument is developed using a
construct map, the construct facet has already been established by that process.
However, beyond the specification of the construct map itself, the important issue
of how the categories of the item responses will be related to the construct map still needs
to be explicated. Each item can be designed to generate responses that span a certain
number of the waypoints on the construct map. Two is the minimum (otherwise the item
would not be useful), but beyond that, any number is possible up to the maximum
number of waypoints in the construct. With selected response items such as multiple-
choice items, this range is limited by the options that are offered. For example, the
observational DRDP item shown in Figure 2.14 has been designed to generate children’s
responses at 8 levels: from “Responds in basic ways to others” to “Compares own
preferences or feelings to those of others.” Thus, this item is polytomous (i.e., it generates
responses for greater than 2 waypoints). But in practice many selected response items are
dichotomous (i.e., generate responses for just two waypoints), such as typical multiple-
choice items in educational testing, where only one option is correct, and the others are
incorrect. In attitude scales, this distinction is also common: for some instruments using
Likert-style options, one might only ask for “Agree” versus “Disagree” but for others a
polytomous choice is offered, such as “Strongly Agree,” “Agree,” “Disagree” and
“Strongly Disagree,” etc. Although this choice can seem innocuous at the item design
stage, especially for selected response items using traditional option-sets, it is in fact
quite important, and will need special attention when we get to the fourth building block
in Chapter 5.
An example of how items match to waypoints was given in Chapter 1 for the
Data Modeling example item “Piano Width” where categories of responses were matched
to four waypoints in the MoV construct map (see Table 1.1).
A second example is from the LPS Argumentation Assessment (Example 4). The
construct map for this construct was given in Figure 2.9, and the relevant detail is shown
in Figure 3.2. The prompt for a bundle of LPS items is shown in Figure 3.3: the “Ice-to-
Water-Vapor” task. Three of the items developed to accompany this prompt (see items
4C, 5C and 6C in the left-hand panel of Figure 3.3). These are matched to the 0b, 0d and
1b waypoints shown in Figure 3.2, respectively, where they identify a claim, some
evidence and a warrant.
5
Figure 3.2 Detail of the Argumentation Construct Map (from Osborne et al, 2016)
Each item in an instrument not only plays a specific role in relating to at least two
waypoints, but the complete set of items should correspond meaningfully to the full range
of interest within the construct. Each item in an instrument not only plays a specific role
in relating to at least two waypoints, but the complete set of items should correspond
meaningfully to the full range of interest within the construct. If some waypoints are of
relatively greater interest (high end versus low end), then the items should be distributed
in a way that matches that preference.
6
Figure 3.3 The prompt for the “Ice-to-Water-Vapor” task.
Evan and Anna are wondering what happens when ice in a container is heated on a
hotplate for 10 minutes. The ice is -5°C. Evan thinks that the temperature will increase
like Graph A. Anna thinks it will be more like Graph B.
7
The most important part of the item’s design is in how the responses relate to the
waypoints. For every item all categories of the item’s responses must relate to specific
waypoints on the construct map. Faulty design here will make everything else harder to
get right. Not being sure of this at the beginning is a normal state of affairs, but the
development steps are designed to help the measurement designer to decide about this—
including redesigning construct maps and/or items when the construct facets do not
match up well with the categories of item responses3.
Having specified the construct component, a major task remains—to decide on all
the other characteristics that the set of items will need to have. Here the term “secondary
facets” is used, as these facets are used to describe an important design aspect of the item
set beyond the primary one of relation to the construct map. These are the facets that are
used to establish classes of items to populate the instrument—they are an essential part of
the basis for item generation and classification. Many types of secondary facets are
possible, but some are more common than others. For example, in the health assessment
(PF-10) example, the items are all self-report measures—this represents a decision to use
self-report as the only condition for the Response-Type facet for this instrument (and
hence to find items that people in the target population can easily respond to), and not to
use other possibilities, such as giving the respondents actual physical tasks to carry out.
Looking back to the “Ice-to-Water Vapor” items shown in Figure 3.2, one can see
an illustration of two conditions for the item format facet. As noted above, the left-hand
panel contains three constructed response items, and the right-hand panel contains three
selected response items directly related to those constructed response items. The selected
response items are examples of the “select and shift” style of items common in
technology-enhanced items (TEIs). In each case, the respondent must select a sentence
from Anna’s statement and shift it into the correct bin. As noted above the distinction
between selected response and constructed response is a fundamental one in items design.
In a broader sense, the distinction between the Response-Type facet conditions (SR, CR)
was not being used just to sample across types of student response—it was part of the
data collection design, where the (almost identical) SR and CR items were part of a study
of the empirical differences between written and selected responses modes for students.
In addition, the Response-Type facet had another function within the LPS project—the
selected response items were intended to be used in summative testing of students (at the
end of the teaching unit, the semester, etc.), whereas the constructed response questions
were intended for formative use within active teaching (i.e., as a context/prompt for
classroom discussions, for “in-class quizzes” when the teacher wanted to see what the
students were thinking, etc.). This is an example of the complexity of the concept of an
“item facet”: often it is seen as a frame for item sampling, but it can also be a planning
tool for deployment of different versions of an instrument for different purposes. In the
next section, we will also see this facet used as an item development strategy.
3
Note that the possibility of unclassifiable and missing response categories is being set aside here.
8
The most common type of design facet that one finds for items are content facets.
These refer to aspects of the subject matter of the instrument. In. one sense, the construct
map facet is an example of such. More common aspects of the content would be the
curriculum content categories in an academic test, or the job skills in a work placement
inventory: all depend heavily on the specific topics and contexts of the instrument’s
usage.
9
Figure 3.5. A content-by-process blueprint for an immunology test (from Raymond &
Grande, 2019)
There are many other potential facets for the items design, and typically, design
decisions will be made by the measurer to include some and not others. Sometimes these
decisions will be made on the basis of practical constraints on the instrument usage. For
example, such considerations were partly responsible for the design of the PF-10
(Example 6) when early on in the SF-36 development process it was deemed too time
consuming to have patients carry out actual physical functioning tasks. Sometimes such
decisions are made on the basis of historical precedents (also partly responsible for the
PF-10 design—it is based on research on an earlier, larger, instrument). And sometimes
on a practical basis because the realized item pool must have a finite set of facets, while
the potential pool has an infinite set of facets. Also note that, although these are couched
as decisions about the instrument, they are not entirely neutral to the idea of the construct
itself. While the underlying PF-10 construct might be thought of as encompassing many
different manifestations of physical functioning, the decision to use only a self-report
10
facet restricts its actual interpretation of the instrument (a) away from items that could
look beyond self-report, such as performance tasks, and (b) to items that are easy to self-
report.
Recall that the purpose of the item is to help determine the person’s location along
the construct of interest. Any list of facets used as a blueprint for a specific instrument
design plan must necessarily have some somewhat arbitrary limitations in the degree of
detail in the specifications—for example, none of the specifications mentioned so far
include “they must be in English”, yet this is indeed a feature of all of the mentioned
instruments. One of the most important ideas behind the items design is to decrease this
incidence of unspoken specifications by explicitly adopting a description of the item pool
quite early in instrument development. This initial items design will likely be modified
during the instrument development process, but that does not diminish the importance of
having an items design from very early in the instrument development. The generation of
at least a tentative items design should be one of the first steps (if not the first step) in
item generation. Items constructed before a tentative items design is developed should
primarily be seen as part of the process of developing the items design itself. Generally
speaking, it’s going to be easier to develop items from your design than to go backwards
and figure out a design based ona set of items. The items design can (and probably will
be) revised but having one in the first place makes the resulting item set much more
likely to be coherent.
One important way that different item formats can be characterized is by their
different amounts of pre-specification—that is by the degree to which the results from the
use of the instrument are developed before the instrument is administered to a respondent.
For example, open ended essays are less prespecified than true/false items. Of course,
when more is pre-specified before, less that has to be done after the response has been
made. Contrariwise, when there is little pre-specified (i.e., little is fixed before the
response is made) then more has to occur afterwards in order to match the response
categories to the construct map. This idea will be used as the basis for the matrix-like
classification of item types shown in Figure 3.5 which summarizes the whole story.
This typology is not only a way to classify the response style of items that exist in
research and practice. Its real strength lies in its nature as a guiding principle to the item
development process. I would argue that every instrument should go through a set of
developmental stages that will approximate the columns in Figure 3.6 through to the
desired level. Instrument development efforts that seek to skip some of these stages run
the risk of having to make more or less arbitrary decisions about item design at some
point in the development process. For example, deciding to create a Likert-type attitude
scale without first investigating the responses that people would choose to make to open-
ended prompts will leave the instrument with no defense against the criticism that the
11
pre-chosen Likert-style response format has distorted the measurement. The same sort of
criticism holds for traditional multiple choice achievement test items.
The item format with the lowest possible level of pre-specification would be one
where the administrator had not yet formulated ANY of the item format issues discussed
above, or even, perhaps the nature of the construct itself, the ultimate aim of the
instrument. What remains is the intent to observe. This type of very diffuse
instrumentation is exemplified by the participant observation technique (e.g., Emerson et
al., 2007; Ball, 1985; Spradley, 2016) common in qualitative studies. This is closely
related to the phenomenological interview technique or “informal conversational
interview” as described by Patton (1980):
Not only is it the case that the measurer (i.e., in this case usually called the “participant
observer”) may not know a priori the full purpose of the observation, but also “the
persons being talked with may not even realize they are being interviewed4” (Patton,
1980, p. 198). There are also participant observations where the timing and nature of the
observation has been planned in advance—for exampling, observing what a respondent
does at a particular point in a process.
4
Of course, not informing the respondents that they are part of a research study may well be a tactic that
needs an institutional review board (IRB) review.
12
3.3.2 Specifying (Just) the Topics
What if you know what topics you want to ask about, but no more than that?
The next level of pre-specification occurs when the aims of the instrument are
indeed pre-established—in the terms introduced above, one can call this the topics guide
format (i.e., in the second and third rows of the matrix). Patton (1980), in the context of
interviewing, labels this the “interview guide” approach—the guide consists of:
Description of
item components Specific items
Intent to
measure No Score Score
Item Format construct General Specific Guide Guide Responses
Note: “X” marks the column which encompasses the main activity at each level of Item
Format.
13
For Topics Guide, two levels of specificity can be distinguished. At the more
general level of specificity (second row of the matrix), the definition of the construct, and
the topics, are only specified to a summary level—this called the general topics guide
approach. Presumably, the full specification of these will occur after observations have
been made. At greater degree of specificity (third row of the matrix), the intended
complete set including the construct definition and the full set of topics is available
before administration—hence, this is the specific topic guide approach. The distinction
between these two levels is a matter of degree—one could have a very vague summary
and alternatively, there could be a more detailed summary that was nevertheless
incomplete.
The next level of pre-specification is the open-ended level (rows four and five of
the matrix). This includes the very common open-ended test and interview instruments
such as were mentioned at the beginning of this chapter—used by teachers and in
informal settings the world over. Here the items are determined before the administration
of the instrument, and are administered under standard conditions, including a pre-
determined order. In the context of interviewing, Patton (1980) has labeled this the
“standardized open-ended interview”. Like the previous level of item format, there are
two discernible levels within this category. At the first level (fourth row of the matrix),
the response categories are yet to be determined. Most tests that teachers make
themselves and use in their classrooms are at this level. At the second level (fifth row of
the matrix), the categories that the responses will be divided into are pre-determined—
call this the scoring guide level. Examples of the open-ended item format (also known as
constructed response items) have been given in Chapter 1 (Figure 1.6) and in this chapter
(Figure 3.4, left panel, and Figure 3.7).
14
Figure 3.7 A Student Responding to the Rainforest/Desert Task in the Cue interview
The interviewer explores the student’s thinking about the whether the plants in one
environment could live in other environment and why or why not, using the photos as
illustration and referent information. The interviewer asks:
Do you think these plants—that live in the rainforest—could survive in the desert too? If
S answers no: Why not? Any other reasons?
15
Figure 3.8 Coding Elements developed by the CUE Project
16
Figure 3.9. An example of a multiple-choice test item that would be a candidate for
ordered multiple-choice scoring.
This possibility can be built into multiple choice items right at the design stage
when a construct map has been used to design the instrument. By developing options that
each relate to different waypoints and ensuring that there are more than two waypoints
involved (say, three or four), then the options can be ordered with respect to the construct
map, resulting in an ordered multiple-choice (OMC) item (Briggs et al., 2006). This
development technique offers a way to improve the interpretability of outcomes from
traditional-looking multiple-choice instruments.
The GEB and PF-10 instruments described in previous chapters are examples
(Examples 3 and 6, respectively) of items with Likert-style response options. This item
format is equally as common in social science settings (and beyond) as multiple-choice
are in educational testing, so I will not give further examples. However, there is a related
response format that has been found to give better results, the Guttman-style item (Wilson
et al., 2021). In this format, the options are ordered according to the underlying
construct, and, when the instrument is being developed according to a construct map, the
most convenient way to do this is to design each option to match successive waypoints.
This is exactly what was done in the RIS example (Example 2—see the construct map in
Figure 2.4), where an initial set of Likert-style items was used to create a set of Guttman-
style items. An example is shown in Figure 3.10. Initially, each of these options was the
stem of a Likert-stle option (with choices from Strongly Agree to Strongly Disagree)—
changing the response option sin this way focusses the respondent on the construct under
measurement, using context-specific words to describe the relevant waypoints, and away
from merely reporting their agreement/disagreement level, which is at both vague and
very variable across respondents.
17
Figure 3.10 An example item from the RIS (Guttman response format items).
The above description of this sequence of formats has focused on what has to
happen before the item is administered. But, in fact the sequence of columns in the
matrix (as shown in Figure 3.6) also lays out what has to happen after the item is
administered if, indeed, the responses are to become the input to a measurement.
Essentially, in each row, when one looks to the right of the “X” (the point at which that
item format is administered), the steps remaining are indicated by the remaining columns.
Thus, texts that started as responses in a participant observation context (row one of the
matrix) would still have to go through the multiple stages of qualitative analysis
including the derivation of general and specific categories for the contents of the texts
(i.e., columns 2 and 3 of the matrix), then the development of coding and scoring guides
for the texts (columns 3 and 4 of the matrix), and the actual coding of them (column 5).
This same logic applies to each of the remaining rows, until row six, where there is no
further need for recoding, etc.— at this point each respondent has coded their own
response5.
5
Of course, the measurer may want to carry out further analysis and recoding after that.
18
At this level of specificity, and at each of the levels noted below, item
development should include a thoughtful consideration of the range of
participants/respondents and how their differences may alter how items are understood.
This should include consideration of respondents who are likely to be from different
waypoints of the construct, all the standard demographic groups (gender, ethnic-racial
identity, class, etc.) as well as specific groupings relevant to the construct under
measurement, such as, in the case of, say, the scientific argumentation construct
discussed above, students with different amounts of experience in science, and different
levels of interest in science.
When items are organized into instruments, there are also issues of instrument
format to consider. An important dimension of instrument design is the uniformity of the
formats within the instrument. An instrument can consist entirely of a single item format,
such as is typical in many standardized achievement tests where all are usually multiple-
choice items, and in many surveys, where Likert-type items are mostly used (though
sometimes with different response categories for different sections of the survey). But
6
But indeed, most refereed journal do not insist on such information. Lamentably often they recommend
deletion of such information to save on page length.
19
more complex mixtures of formats are also used. For example, the portfolio is an
instrument format common in the expressive and performance arts, and also in some
professional areas. This will consist of a sample of work that is relevant to the purpose of
the portfolio, and so may consist of responses to items of many sorts and may even be
structured in a variety of ways more or less freely by the respondent according to the
rules that are laid down. Tests may also be composed of mixed types—multiple choice
items as well as essays, say, or performance-tasks of various sorts. Surveys and
questionnaires may also be composed of different formats, true-false items, Likert-style
items, and short answer items. Interviews may consist of open-ended questions, as well
as forced choice sections. Care must be taken in devising complex designs such as these
—concern must be taken with respect to the time allowed for the different sections, the
extent of coverage of each, and way that the items of different formats are related to the
construct map(s).
A crucial step in the process of developing an instrument, and one unique to the
measurement of human beings, is that the measurer can ask the respondents what they are
thinking about when responding to the items. In Chapter 8 evaluative use of this sort of
information is presented as a major tool in gathering evidence for the validity of the
instrument and its use. In this section, formative use of this sort of information is seen as
a tool for improving the instrument, and in particular, the items of the instrument. There
are two major types of investigations of response processes, the “think-aloud” and the
“exit interview.” Other types of investigation may involve reaction time studies, eye
movement studies and various treatment studies, where, for example, the respondents are
given certain sorts of information before they are asked to respond.
In the think-aloud style of investigation, also called “cognitive labs” (AIR, 2000),
students are asked to talk aloud about what they are experiencing as their mental
activities while they are actually responding to the item. What the respondent says is
recorded, often what they do is being videotaped, and other characteristics may be
recorded, such as having their eye movements tracked. A professional should be at hand
to prompt such self-reports, and also to ask clarifying questions if necessary. Typically,
respondents need a certain amount of familiarization to know what it is that the
researcher is interested in, and to feel comfortable with the procedure. The results can
provide insights ranging from the very plain—such as “the respondent was not thinking
about the desired topic when responding”—to the very detailed, including evidence about
particular cognitive and metacognitive strategies that they are employing. Of course, the
scientific value of self-report information should be questioned (e.g., Brenner et al.,
2003). But, in this circumstance, the evidence from the think-aloud should be considered
more like forensic evidence in a criminal investigation—not directly indicating what has
really happened (mentally) but giving the measurer clues that can be linked to reasonable
hypotheses about what is going on in the hidden mental world of the respondent.
20
A sample think-aloud protocol is provided in Figure 3.11 which was developed as
a part of the LPS project (Example 4). Note that the preamble to the whole process,
designed to set the respondent at ease, and provide them with disclosure of the aims of
the exercise7. Then the interviewer demonstrates what they want the respondent to do
using a similar context (i.e., describing how the student would get to their school office),
for the respondent to become familiar with the procedures. Other familiarization
procedures are also commonly used, such as conducting a “dry run” with an item similar
to those in the actual study-set. The exercise then proceeds through each item to be
investigated. In this case, the cognitive lab was being recorded on a verbal recording
device, but it may also be video-taped, or recorded via a video-conferencing system
(although care must be taken with these to make sure that the respondent is not identified
in the stored data). One additional good idea to have a silent observer present during the
entire activity to assist in interpretation and evaluation of the results with the interviewer.
The data from these data sources are then coded to record the relevant aspects of
the information, and to allow accumulation and comparison across respondents. Each
development project will need to create its own version of a think-aloud coding sheet
according to the information they wish to glean from the think-aloud process. Some of
the relevant pieces of information that will be likely to valuable to record are:
(i) how accessible the questions were for the respondents,
(ii) whether the items appeared to be biased for (or against) certain subgroups (this, of
course, depends on their being representatives of these groups among the sampled
respondents, as noted above)
(iii) whether the respondents found the items to be simple and clear, and whether the item
had intuitive instructions and procedures
(iv) whether the text of the item was readable and comprehensible,
(v) whether the presentation of the items was legible, and
(vi) whether anything was left out.
7
The scope of this, of course, must be agreed to by your institutional IRB.
21
Figure 3.11 Sample script for a think-aloud investigation
22
Figure 3.10 Sample script for a think-aloud investigation (continued)
23
Figure 3.10 Sample script for a think-aloud investigation (continued)
24
An example of responses gathered during a think aloud session for the (early
version of an) item from the LPS project (shown in Figure 3.12) is shown in Figure 3.13.
The item relates to a different construct map used in the LPS project, for Ecology 8. Some
appreciation of the typical range of responses one gets from a think-aloud exercise can be
gained by reading through the various comments recorded in Figure 3.12—they range
from helpful corrections for poor wordings, to insights into how students are
misunderstanding the ideas involved, to insights into how students are misunderstanding
the representation of the ideas in the item figure, to student vocabulary deficits. In the
revisions to this item, the project:
(a) revised the preamble to read “The picture below shows the changes in an ecosystem
over time after a big fire. Scientists call this ecological succession”; and
(b) modified the representation to make clearer that the panels in the figure were views
(snapshots) of the same place over time (ie., by adding an element that is common
to all of them).
Succession
The picture below shows the process of ecological succession after a big fire,
which occurs after an ecosystem experiences a disturbance (like a big fire).
S12. What changes do you notice across the years in the figure above?
8
The details of that construct map are not needed to appreciate the observations about this example of
think-aloud notes, but in case the reader is interested one can look up that information via the Examples
Archive in Appendix A.
25
Figure 3.13 Example Notes from a Think-aloud Session
102, 106: Repetitive statement: “The picture below shows the process of
ecological succession after a big fire, which occurs after an ecosystem
experiences a disturbance (like a big fire).”
201: Student seemed to focus only on the most obvious change (bare forest to
smaller trees to taller trees), without examining the figure more closely for other
“smaller” changes. Question design may not facilitate close/detailed examination
of the figure. Maybe we can ask something like “Several patterns of change can
be observed in the above figure. Describe some of these patterns.”
202: The word “hardwood” might be unfamiliar to students. Student also said that
it was confusing why the vegetation would keep changing from one type of plant
to another, especially how pine would turn into hardwood (i.e., student does not
understand what a succession is). He thought that the pine trees might have died
out (because of a drought) then hardwood trees grew over the land. But student
knew that the grids were showing the same location.
202: Student did not seem to notice the smaller patterns/changes besides the
change in plant growth (bare to smaller trees to bigger trees), despite multiple
prompting.
110: Student thinks the plants over the years are the same plants just growing
bigger (less bushy to more bushy).
104 succession,
106 Disturbance, ecological (but the student said they weren’t that hard)
The exit interview is similar in aim to the think-aloud but is timed to occur after
the respondent has completed their item responses. It may be conducted after each item,
or after the instrument as a whole depending on whether the measurer judges that the
delay will or will not interfere with the respondent’s memory. The types of information
gained will be similar to those from the think-aloud, though generally it will not be so
detailed. This is not always a disadvantage, as sometimes it is exactly the results of a
respondent’s reflection which are desired. Thus, it may be that a data collection strategy
that involves both think-alouds and exit-interviews will be best.
26
Information from respondents can be used at several points along the instrument
development process as detailed in this book. Reflections on what the respondents say
can lead to wholesale changes in the idea of the construct, revisions of the construct facet,
the secondary facets, and specific items and item types, and changes in the outcome space
and scoring guides (to be described in the next Chapter). It is difficult to over-emphasize
the importance of including procedures for tapping into the insights of the respondents in
the instrument development process. A counterexample is useful here—in cognitive
testing of babies and young infants, the measurer cannot gain insights in this way, and
that has required the development of a whole range of specialized techniques to make up
for such a lack.
The question of whether the items in an instrument are fair to respondents should
be addressed at the design stage. According to the Meriam-Webster dictionary, fair
means: “marked by impartiality and honesty: free from self-interest, prejudice, or
favoritism.” At this point in instrument development, in the process of developing items,
it is a formative question; however, this issue will be revisited as a part of the evaluation
of the fairness evidence for an instrument in Chapter 8. The essential question is whether
there are important influences (besides the underlying construct that the instrument is
27
intended to measure) on a respondent’s reactions to specific items in an instrument, or
indeed the whole set of items. For example, consider the many well-established instances
in educational testing, as noted by a U.S. National Research Council (NRC) report ...
... in a written science assessment [or for any other subject matter] with open-
ended responses, is writing a target skill or an ancillary skill? Is the assessment
designed to make inferences about science knowledge, about written expression
of science knowledge, or about written expression of science knowledge in
English? The answers to these questions can assist with decisions about
accommodations, such as whether to provide a scribe to write answers or to
provide a translator to translate answers into English. If mathematics is required
to complete the assessment tasks, is mathematics computation a target skill or an
ancillary skill? Is the desired inference about knowing the correct equation to use
or about performing the calculations? (Here, the answers can guide decisions
about use of a calculator.) (NRC, 2006)
These are questions that can be asked for each respondent, but usually it is
couched in terms of how different groups of respondents will respond to the items. These
groups may differ from construct to construct and from situation to situation but, in a
given context, will generally correspond to typical demographic groups. Staying within
the educational testing context, these groups could be defined so as to include, say,
respondents from different gender, ethnic and racial groups, respondents with a different
language status (such as respondents whose native language is not the language of the
instrument), respondents with learning challenges (such as cognitive disabilities, learning
disabilities, or who are deaf or hard of hearing, etc.), or respondents with non-standard
citizen status (such as recent migrants, or visa-holders, etc.).
In addressing this issue in terms of the construct map and the development of
items that are sensitive to the construct, typically, the investigation of such effects on
demographic groups would be structured by assuming that there is (a) a reference group
(i.e., the group to whom the other groups will be compared), and (b) one or more focal
groups (i.e., groups for whom the fairness is in question). Then, the measurer will need to
consider the following three successively more complicated possible types of difference
in the way that the construct map and the items work between the focal and the reference
groups. These possibilities are:
(i) in terms of the measurements themselves, the respondents in the focal group are
essentially on the same construct map as the respondents in the reference group,
but they tend (on average) to be lower or higher than those from the reference
group (this is termed “differential impact”);
(ii) the respondents in the focal group are essentially on the same construct map, however
some items (or perhaps a class of items) in the instrument behave differently for
respondents in the focal group compared to their reference group peers who have
a matching level of on the construct map. This phenomenon has the generic label
“differential item functioning” or DIF9; see Gamerman et al., 2016, Kamata &
Vaughn, 2011; and Chapter 8); and
9
When this occurs for the aggregate of the whole test, this is termed “differential test functioning”)
28
(iii) the respondents in the focal group are on a different construct from those in the
reference group10.
Clearly, case (i) is the least complicated of the three—we might say that there is
measurement-fairness here. But it is important to distinguish between the fairness of the
measurement process itself (i.e., the instrument fairly maps from the student responses to
the eventual outcome), and real world fairness (i.e., there is no unfairness in the ways
that the individual respondents have gotten to this point of being measured). Examples of
real world unfairness night be where some individuals have had more resources devoted
to them in the forms of better nutrition, education or medical care. It might be naïve to
assume that, when no DIF is found that there is no unfairness that has occurred,
especially if there is a history of concerns over fairness in the specific contexts where the
instrument is typically used. To establish that case (i) does indeed hold, the remaining
two possibilities need to be examined, and eliminated as important effects.
Case (iii) is the most complex but relatively little research exists investigating this
possibility. Looking again at the case of English learners (ELs) learning mathematics, this
is at least partly due to the paucity of clearly established models of how ELs typically
learn and develop competence. But emerging research suggests that often ELs do not
follow the typical learning progressions—this should not be too surprising as our
understanding of these earning progressions are based primarily on research and models
for non-ELs (Sato et al., 2012). One would expect that, given individual differences in
educational history, socio-cultural background, and literacy and fluency in their native or
home language, ELs would interact with, academic content differently compared to non-
ELs (e.g., Solano-Flores & Trumbull, 2003). Yet a further layer of complexity is the
concern that ELs are a not a uniform group—there are substantial differences among ELs
in terms of the cognitive, cultural and linguistic resources as well as their prior
educational backgrounds that they bring to assessment tasks (Abedi, 2004).
10
This would be termed as a lack of factorial invariance across these groups.
29
On the other hand, some evidence suggests that learning progressions can be
appropriate tools for examining learning and development in students from different
cultural and linguistic backgrounds. In the area of early language and literacy, the
literature agrees that, despite opportunity gaps that lead to average differences between
groups, ELs have similar learning trajectories as their monolingual English-speaking
peers in areas of vocabulary development, grammar, phonemics, and writing (Hammer et
al., 2014). An empirical analysis, incorporating a learning progressions approach, found
that groups of young children in public early care and education programs who spoke
different languages at home had markedly similar trajectories of early language and
literacy development (Sussman et al., under review).
11
Further information on this topic is available from the National Center on Educational Outcomes
(NCEO) ([Link]
30
(f) designing for maximum readability and comprehensibility; and
(g) designing for maximum legibility.
Examples of each of these are given in Johnstone et al. (2006). In particular, given that
the reader is reading this book, the advice to develop “precisely defined constructs”
should be already at the top of the measurer’s list. But it is important, in light of the
discussion above, that the measurer be open to the possibility that the constructs may be
structured differently between reference and focal groups and investigate it thoroughly.
Apart from engaging in a focused literature search for evidence of this, the measurer
should also see that the cognitive lab methods described in Section 3.4 will be useful in
giving clues, assuming, of course, that the focal groups have been included among the
sampled respondents, which corresponds to recommendation (a) above.
A useful framework for universal design in a measurement context has been given
by Robert Dolan and his colleagues (Dolan et al, 2013). Nathaniel Brown has prepared a
summary of such issues (Brown, 2020) that is helpful in exploring the multifaceted nature
of the design issues, and this is included as Appendix 3B.
3.6 Resources
Within the area of educational achievement testing, there are several very useful
resources for types of items and methods to develop them such as: Brookheart & Nitko
(2018), Haladyna (1996, 1999), Nitko (1983), Osterlind (1998), and Roid and Haladyna
(1982). Many if not most other areas will have similar resources, and your informants
should be able to put you onto them. We recommend you consult with professors,
colleagues, and experts in your area of interest for more specific direction and support.
Such informants, especially those who have experience in instrument development, can
help explain specific issues that may arise, provide insight during the item development
process, and help critique your overall process.
1. Generate lots of types of items and several examples of each type. Be prolific.
2. Write down your initial items design based on the preceding activities (including the
activities listed for Chapters 1 and 2)—devote most of your attention to the construct
31
facet but listen to your informants about what should be the most important secondary
facets.
3. Give these draft items a thorough professional review at an “item panel” meeting—
where key informants and others you recruit to constructively critique the items generated
so far (see Appendix 3A).
4. Write down your interim Items Design based on the preceding activities, including
predicted numbers of item you (still) will need to generate.
5. Following the initial round of item generation and item paneling described in (1)
through (4), second or third rounds may be needed, and may also involve a return to
reconsider the construct definition or the definition of the other facets of the items design.
6. Enter your surviving items into the BASS Item Bank for the construct maps you are
working on.
7. Think through the steps outlined above in the context of developing your instrument
and write down notes about your plans.
8. Share your plans and progress with your informants and others—discuss what you are
succeeding on, and what problems have arisen.
32
Appendix 3A
33
4. How to carry out the Item Paneling.
(a) You will chair the meeting. Your aim, as chair, is to help each panelist contribute in
the most constructive way possible to creating the best set of items that well generate
responses that are related to the construct map. Each panelist needs to understand that
that is the aim. Panelists are to be as critical as they can, but with the aim of being
constructive as well. Disputes are to be left in the room at the end—the chair/item
developer will decide what to do with each comment.
(b) Order of business:
(i) Explain the purpose of the panel in case some panelists are new to the
procedure
(ii) Give panelists a brief overview of the construct map and the context for use of
the planned instrument, including a description of the intended
respondents and raters (if applicable). Invite comments and questions.
(iii) Systematically discuss the items to be paneled, keeping in mind as you go the
expected length of the meeting and the number of items to be paneled.
Before passing on to the next item, be sure that you are satisfied that you
are clear what are the panel’s recommended revisions for the current item.
(iv) After surveying all items ask for general comments on the item set and
especially whether the item set comprehensively represents the framework
for the variable.
(v) Collect any written notes about the items from the panelists
34
Appendix 3B
12
Brown, N. (2020). Supporting Construct-Irrelevant Actions. Research Report, Graduate School of
Education, Boston College.
35
36