0% found this document useful (0 votes)
6 views36 pages

Designing Effective Measurement Items

Uploaded by

Carlos - Tam
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOC, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views36 pages

Designing Effective Measurement Items

Uploaded by

Carlos - Tam
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOC, PDF, TXT or read online on Scribd

Chapter 3

The Items Design

3.0 Chapter Overview

This chapter and the next tell the story about designing the item: the ways that the
measurer uses to stimulate the respondent to produce products and behaviors that can be
used as the basis for observations about the location of the respondents on the construct
in question. By “observation,” something more is meant than just seeing something and
recording it, or remembering it, or jotting down notes about it. In referring to it as a
special sort of observation that is generically called an item, it means that there exists (a)
a procedure, or design, which allows the observations to be made under a set of standard
conditions that span the intended range of the item contexts, and (b) a procedure for
classifying those observations into a set of standard categories. The first part is the topic
of this chapter, and the second is the topic of the next chapter. The measurement
instrument is then a set of these procedures (i.e., a set of items). First, the idea of an item
is developed in the section immediately below, and some typical types of items are
discussed. Then, a typology of items is introduced that is designed to show connections
among many different sorts of items. This is followed by a discussion of the important
facets of the whole-instrument design. After that, there are section devoted to (a) the
inclusion of a fairness perspective into the items design and (b) the incorporation of
human judgement into the items design. Finally, ways of developing items are discussed
in the last section. In the next chapter (Chapter 4), this understanding about the nature of
the item is complemented to include the way that the observations are recorded and
categorized.

Key concepts: item formats, participant observation, topics guide, constructed responses,
selected responses; facets of the items design, the construct facet, secondary facets;
universal design

3.1 The Idea of an Item


What is the main purpose of an “item”?

Often the first inkling of an item comes in the form of an idea about a way to
reveal a particular characteristic (construct) of a respondent. The inkling can be quite
informal: A remark in a conversation, the way a student describes what they understand
about something, a question that prompts an argument, a particularly pleasing piece of
art, a newspaper article, a patient’s or client’s symptoms. The specific way in which the
measurer prompts an informative response from a respondent is crucial to the value of the
resulting measurements. In fact, in many, if not most, cases, the construct itself will not
be clearly defined until a relatively large set of items has been created and tried out with
respondents. Each new situation brings about the possibility of developing new and
different sorts of items or of adapting old ones.

1
Across many fields and across many topic areas within those fields, a rich variety
of types of items have been developed to deal with many different constructs and
situations. We have already seen two very different formats. In Chapter 1 (in Example
1) a constructed response type of item was closely examined: the Data Modeling MoV
“Piano width” item. Many examples of responses to that item were included in Table 1.1.
The LPS science argumentation assessment (Example 4), and the CUE interview
(Example 7) were additional examples of the constructed response type. In Chapter 2 we
branched out and examined selected response items as well: the PF-10 health survey
(Example 6) asked the question “Does your health now limit you in these activities?”
with respect to a range of physical activities but restricted the responses to a selected
response among: “Yes, limited a lot”, “Yes, limited a little”, and “No, not limited at all.”
The GEB Scale (Chapter 2, Example 3) was another example of the selected response
type, as it is an example of the familiar Likert-style item. These are examples of two ends
of a range of item formats that stretch from very open-format constructed responses to
very closed-format selected responses. In the following section we present a range of
item types that span across these two, and beyond (see Figure 3.5). Many other types of
items exist (e.g., see Brookheart & Nitko, 2018, for a large assortment from educational
assessment), and the measurer should be aware of both the specific types of items that
have previously been used in the specific area in which an instrument is being developed,
as well as item types that have been used in other related contexts.

Probably the most common type of item in the experience of most people is the
general constructed response item format that is commonly used in school classrooms
and many other settings around the world. The format can be expressed orally or in
writing (by hand or typing) or in other forms, such as concrete products, active
performances, or actions taken in a digital environment. The length of the response can
vary from a single number or word to lengthy essays, complex proofs, interviews,
extended performances or multi-part products. The item can be one that is produced
extemporaneously by the measurer (say, the teacher, or other professional) or it can be
the result of an extensive developmental process. This format is also used outside of
educational settings including workplaces, social settings, and in the everyday
conversational interchanges that we all experience. Typical sub-forms are the essay, the
brief demonstration, the product, and the short-answer format.

In counterpoint, probably the most common type of item in published instruments


is the selected response format. Some may think that this is the most common item
format, but that is because they are discounting the numerous everyday situations
involving constructed response item formats. The multiple-choice item is familiar to
almost every educated person and has had an important role in the educational
trajectories of many. The selected response (particularly the Likert-type response) format
is also very common in attitude scales, and surveys and questionnaires used in many
situations, such as in health, applied psychology and public policy settings, in business
settings such as employee and consumer ratings, and in governmental settings. The
responses are most commonly “Strongly Disagree” to “Strongly Agree”, but many other
response options are also found (as in the PF-10 example). It is somewhat paradoxical

2
that the most commonly experienced format is not the most commonly published format.
As will become clear to the reader as they advance through the next several chapters, the
view developed here is that the constructed response format is the more basic format, and
the selected response format can be seen as an adapted version of it.

The relationship of the item to its construct is its most fundamental relationship 1.
Typically, the item is but one of many (often one from an infinite set) that could be used
to measure the construct. Paul Ramsden and his colleagues, writing about the assessment
of achievement in high school physics noted:
Educators are interested in how well students understand speed, distance
and time, not in what they know about runners or powerboats or people
walking along corridors2. Paradoxically, however, there is no other way
of describing and testing understanding than through such specific
examples. (Ramsden et al, 1993, p. 312; footnote added)
Similarly, consider the health measurement (PF-10) example described above. Here the
specific questions that are used are clearly neither necessary for defining the construct,
nor are they sufficient to encompass all the possible meanings of the concept of physical
functioning. Thus, the task of the measurer is to choose a finite set of items that does
indeed represent the construct in some reasonable way. As Ramsden hints, this is not the
straightforward task one might think it to be, on initial consideration.

Sometimes a measurer will feel the temptation to seek the “one true task,” the
“authentic item”, the single observation that will supply the mother lode of evidence
about the construct. Unfortunately, this misunderstanding, common amongst beginning
measurers, is founded on a failure to fully consider the need to establish sufficient levels
of validity and reliability for the instrument. Where one wishes to represent a wide range
of contexts in an instrument, it is better to have more items rather than less—this is
because (a) the instrument can then sample more of the content of a construct and more
of the situations where a student’s location on the construct might be displayed (see
Chapter 8 for more on this), and (b) because it can then generate more bits of information
about how a respondent stands with respect to the construct, which will yield greater
precision (see Chapter 7 for more on this). And both requirements need to be satisfied
within the time and cost limitations imposed on the measuring context.

3.2 The Facets of the Items Design


What are the essential components of an item?

The items design is the second building block in the Bear Assessment System
(BAS). It has already been introduced, lightly, in Chapter 1, and its relationship to the
other building blocks was illustrated there too—see Figure 3.1. It is the main focus of
this chapter.

1
Although an item may also be related to several different constructs, especially if the outcome space (see
Chapter 4) is different for each.
2
I.e., these are typical situations involved in high school physics items.

3
Figure 3.1 The Items Design, the second building block in the BEAR
Assessment System (BAS)

One way to understand the items design is to see it as a description of the


theoretical population of items, called the item universe (Guttman, 1941), along with a
procedure for sampling the specific items to be included in the instrument. As such, the
instrument is the result of a series of decisions that the measurer has made regarding how
to represent the construct or, equivalently, how to stratify the “space” of items and then
sample from those strata. Some of those decisions will be principled ones relating to the
fundamental definition of the construct, and the research background of the construct.
Some will be practical, relating to the constraints of administration and usage. Some will
be rather arbitrary, being made to keep the item generation task within reasonable limits.
We can think of items as bring like gemstones, they can be viewed from a variety of
perspectives and/or they can be seen as having different characteristics. In this book we
will refer to the different characteristics of items as facets. Each facet of an item tells us
something important and contributes to its overall quality. Thus, whether the item is of a
selected response type, or a constructed response type would be one facet—this would be
a facet that divided the universe into disjoint subsets: (i) selected response items, (ii)
constructed response items and (iii) items not readily classifiable into either. Another
facet might be the readability of the text in the item—this would be a facet that associated
each item with a readability value (e.g., Flesch’s Reading Ease formula—Flesch, 1948)
but also it may need to include a subset for items that had no readability value (e.g.,
perhaps for a spoken item). For our purposes here, one can distinguish between two types
of facets of the universe of items that are useful in describing the item pool: (a) the
construct facet, and (b) the rest of the facets.

3.2.1 The Construct Facet


What is the most important facet?

The facet that is used to provide criterion-referenced interpretations along the


construct, from high to low, is essential and common to all items designs that are

4
intended to relate items to a construct. This facet provides interpretational levels within
the construct—and hence it is called the construct facet. For example, the construct facet
in the Data Modeling MoV construct is provided in Figure 1.12. Thus, the construct facet
is essentially the content of the construct map—where an instrument is developed using a
construct map, the construct facet has already been established by that process.

However, beyond the specification of the construct map itself, the important issue
of how the categories of the item responses will be related to the construct map still needs
to be explicated. Each item can be designed to generate responses that span a certain
number of the waypoints on the construct map. Two is the minimum (otherwise the item
would not be useful), but beyond that, any number is possible up to the maximum
number of waypoints in the construct. With selected response items such as multiple-
choice items, this range is limited by the options that are offered. For example, the
observational DRDP item shown in Figure 2.14 has been designed to generate children’s
responses at 8 levels: from “Responds in basic ways to others” to “Compares own
preferences or feelings to those of others.” Thus, this item is polytomous (i.e., it generates
responses for greater than 2 waypoints). But in practice many selected response items are
dichotomous (i.e., generate responses for just two waypoints), such as typical multiple-
choice items in educational testing, where only one option is correct, and the others are
incorrect. In attitude scales, this distinction is also common: for some instruments using
Likert-style options, one might only ask for “Agree” versus “Disagree” but for others a
polytomous choice is offered, such as “Strongly Agree,” “Agree,” “Disagree” and
“Strongly Disagree,” etc. Although this choice can seem innocuous at the item design
stage, especially for selected response items using traditional option-sets, it is in fact
quite important, and will need special attention when we get to the fourth building block
in Chapter 5.

An example of how items match to waypoints was given in Chapter 1 for the
Data Modeling example item “Piano Width” where categories of responses were matched
to four waypoints in the MoV construct map (see Table 1.1).

A second example is from the LPS Argumentation Assessment (Example 4). The
construct map for this construct was given in Figure 2.9, and the relevant detail is shown
in Figure 3.2. The prompt for a bundle of LPS items is shown in Figure 3.3: the “Ice-to-
Water-Vapor” task. Three of the items developed to accompany this prompt (see items
4C, 5C and 6C in the left-hand panel of Figure 3.3). These are matched to the 0b, 0d and
1b waypoints shown in Figure 3.2, respectively, where they identify a claim, some
evidence and a warrant.

5
Figure 3.2 Detail of the Argumentation Construct Map (from Osborne et al, 2016)

Each item in an instrument not only plays a specific role in relating to at least two
waypoints, but the complete set of items should correspond meaningfully to the full range
of interest within the construct. Each item in an instrument not only plays a specific role
in relating to at least two waypoints, but the complete set of items should correspond
meaningfully to the full range of interest within the construct. If some waypoints are of
relatively greater interest (high end versus low end), then the items should be distributed
in a way that matches that preference.

The organization of the item set into an instrument is usually summarized in a


representation called a blueprint, or sometimes an instrument “specification.” In the case
of th construct facet, this will take on a particularly simple format, by noting which items
relate to each waypoint. Of course, as other facets (as below) are added, this will become
more complex, eventually forming a multi-way matrix.

6
Figure 3.3 The prompt for the “Ice-to-Water-Vapor” task.

Evan and Anna are wondering what happens when ice in a container is heated on a
hotplate for 10 minutes. The ice is -5°C. Evan thinks that the temperature will increase
like Graph A. Anna thinks it will be more like Graph B.

Figure 3.4 The two sets of MoV items: CR and SR

7
The most important part of the item’s design is in how the responses relate to the
waypoints. For every item all categories of the item’s responses must relate to specific
waypoints on the construct map. Faulty design here will make everything else harder to
get right. Not being sure of this at the beginning is a normal state of affairs, but the
development steps are designed to help the measurement designer to decide about this—
including redesigning construct maps and/or items when the construct facets do not
match up well with the categories of item responses3.

3.2.2 The Secondary Design Facets


What else do we need to specify about an item?

Having specified the construct component, a major task remains—to decide on all
the other characteristics that the set of items will need to have. Here the term “secondary
facets” is used, as these facets are used to describe an important design aspect of the item
set beyond the primary one of relation to the construct map. These are the facets that are
used to establish classes of items to populate the instrument—they are an essential part of
the basis for item generation and classification. Many types of secondary facets are
possible, but some are more common than others. For example, in the health assessment
(PF-10) example, the items are all self-report measures—this represents a decision to use
self-report as the only condition for the Response-Type facet for this instrument (and
hence to find items that people in the target population can easily respond to), and not to
use other possibilities, such as giving the respondents actual physical tasks to carry out.

Looking back to the “Ice-to-Water Vapor” items shown in Figure 3.2, one can see
an illustration of two conditions for the item format facet. As noted above, the left-hand
panel contains three constructed response items, and the right-hand panel contains three
selected response items directly related to those constructed response items. The selected
response items are examples of the “select and shift” style of items common in
technology-enhanced items (TEIs). In each case, the respondent must select a sentence
from Anna’s statement and shift it into the correct bin. As noted above the distinction
between selected response and constructed response is a fundamental one in items design.
In a broader sense, the distinction between the Response-Type facet conditions (SR, CR)
was not being used just to sample across types of student response—it was part of the
data collection design, where the (almost identical) SR and CR items were part of a study
of the empirical differences between written and selected responses modes for students.
In addition, the Response-Type facet had another function within the LPS project—the
selected response items were intended to be used in summative testing of students (at the
end of the teaching unit, the semester, etc.), whereas the constructed response questions
were intended for formative use within active teaching (i.e., as a context/prompt for
classroom discussions, for “in-class quizzes” when the teacher wanted to see what the
students were thinking, etc.). This is an example of the complexity of the concept of an
“item facet”: often it is seen as a frame for item sampling, but it can also be a planning
tool for deployment of different versions of an instrument for different purposes. In the
next section, we will also see this facet used as an item development strategy.
3
Note that the possibility of unclassifiable and missing response categories is being set aside here.

8
The most common type of design facet that one finds for items are content facets.
These refer to aspects of the subject matter of the instrument. In. one sense, the construct
map facet is an example of such. More common aspects of the content would be the
curriculum content categories in an academic test, or the job skills in a work placement
inventory: all depend heavily on the specific topics and contexts of the instrument’s
usage.

One common way to express such content facets is through a content-by-process


matrix. An example is given in Figure 3.5. where the rows identify patient conditions,
while the columns are important immunology curriculum concepts. The check marks
show which pairings of patient condition and immunology concept is represented by an
item in the test. In the design of this instrument, the developers decided to include just
one item in the instrument for each such relevant pairing. This provides a formula for the
relative coverage of each topic in the instrument—i.e., for this instrument, each concept
is covered by as many immunology concepts it is associated with in the Figure.
Equivalently, each immunology concept is covered by as many patient conditions as it is
associated with in the Figure. Other facets may be used in the columns or rows of such a
matrix, such as the waypoints in a construct map, the cognitive levels from Bloom’s
Taxonomy, etc. (see other examples in Section 4.3). The plan for arraying items across
the facets can seem to be quite a minor decision and relatively harmless in its
implications. But it is an important design consideration, influencing many aspects of the
instrument’s performance, including the uncertainty that will be associated with
respondent estimated locations on the construct (see Section 6.2).

9
Figure 3.5. A content-by-process blueprint for an immunology test (from Raymond &
Grande, 2019)

There are many other potential facets for the items design, and typically, design
decisions will be made by the measurer to include some and not others. Sometimes these
decisions will be made on the basis of practical constraints on the instrument usage. For
example, such considerations were partly responsible for the design of the PF-10
(Example 6) when early on in the SF-36 development process it was deemed too time
consuming to have patients carry out actual physical functioning tasks. Sometimes such
decisions are made on the basis of historical precedents (also partly responsible for the
PF-10 design—it is based on research on an earlier, larger, instrument). And sometimes
on a practical basis because the realized item pool must have a finite set of facets, while
the potential pool has an infinite set of facets. Also note that, although these are couched
as decisions about the instrument, they are not entirely neutral to the idea of the construct
itself. While the underlying PF-10 construct might be thought of as encompassing many
different manifestations of physical functioning, the decision to use only a self-report

10
facet restricts its actual interpretation of the instrument (a) away from items that could
look beyond self-report, such as performance tasks, and (b) to items that are easy to self-
report.

Recall that the purpose of the item is to help determine the person’s location along
the construct of interest. Any list of facets used as a blueprint for a specific instrument
design plan must necessarily have some somewhat arbitrary limitations in the degree of
detail in the specifications—for example, none of the specifications mentioned so far
include “they must be in English”, yet this is indeed a feature of all of the mentioned
instruments. One of the most important ideas behind the items design is to decrease this
incidence of unspoken specifications by explicitly adopting a description of the item pool
quite early in instrument development. This initial items design will likely be modified
during the instrument development process, but that does not diminish the importance of
having an items design from very early in the instrument development. The generation of
at least a tentative items design should be one of the first steps (if not the first step) in
item generation. Items constructed before a tentative items design is developed should
primarily be seen as part of the process of developing the items design itself. Generally
speaking, it’s going to be easier to develop items from your design than to go backwards
and figure out a design based ona set of items. The items design can (and probably will
be) revised but having one in the first place makes the resulting item set much more
likely to be coherent.

3.3 Different Types of Item Responses


How open-ended versus closed form should an item be?

One important way that different item formats can be characterized is by their
different amounts of pre-specification—that is by the degree to which the results from the
use of the instrument are developed before the instrument is administered to a respondent.
For example, open ended essays are less prespecified than true/false items. Of course,
when more is pre-specified before, less that has to be done after the response has been
made. Contrariwise, when there is little pre-specified (i.e., little is fixed before the
response is made) then more has to occur afterwards in order to match the response
categories to the construct map. This idea will be used as the basis for the matrix-like
classification of item types shown in Figure 3.5 which summarizes the whole story.

This typology is not only a way to classify the response style of items that exist in
research and practice. Its real strength lies in its nature as a guiding principle to the item
development process. I would argue that every instrument should go through a set of
developmental stages that will approximate the columns in Figure 3.6 through to the
desired level. Instrument development efforts that seek to skip some of these stages run
the risk of having to make more or less arbitrary decisions about item design at some
point in the development process. For example, deciding to create a Likert-type attitude
scale without first investigating the responses that people would choose to make to open-
ended prompts will leave the instrument with no defense against the criticism that the

11
pre-chosen Likert-style response format has distorted the measurement. The same sort of
criticism holds for traditional multiple choice achievement test items.

3.3.1 Participant Observation


What is the least structured form of item?

The item format with the lowest possible level of pre-specification would be one
where the administrator had not yet formulated ANY of the item format issues discussed
above, or even, perhaps the nature of the construct itself, the ultimate aim of the
instrument. What remains is the intent to observe. This type of very diffuse
instrumentation is exemplified by the participant observation technique (e.g., Emerson et
al., 2007; Ball, 1985; Spradley, 2016) common in qualitative studies. This is closely
related to the phenomenological interview technique or “informal conversational
interview” as described by Patton (1980):

the researcher has no presuppositions about what of importance may be


learned by talking to people …The phenomenological interviewer wants
to maintain maximum flexibility to be able to pursue information in
whatever direction appears to be appropriate, depending on the
information that emerges from observing a particular setting or from
talking to one or more individuals in that setting. (pp. 198-199)

Not only is it the case that the measurer (i.e., in this case usually called the “participant
observer”) may not know a priori the full purpose of the observation, but also “the
persons being talked with may not even realize they are being interviewed4” (Patton,
1980, p. 198). There are also participant observations where the timing and nature of the
observation has been planned in advance—for exampling, observing what a respondent
does at a particular point in a process.

The degree of pre-specification of the participant observation item format is


located in the first row of the matrix in Figure 3.6 where “X” marks the column which
encompasses the main activity at each level of Item Format. This matrix emphasizes the
progressive increase in pre-specification as one moves from this initial participant
observation level to constructed response formats. Some may balk at considering a
technique like participant observation as an example of an “instrument” and including it
in a book on measurement. But it and the technique described in the next paragraph are
included here because (a) many of the techniques described in these chapters are
applicable to the results of such observations, (b) these techniques can be very useful
during an instrument development (more on this at the end of this section), and (c) the
techniques mark a useful starting point in thinking about the level of pre-specification of
types of item formats.

4
Of course, not informing the respondents that they are part of a research study may well be a tactic that
needs an institutional review board (IRB) review.

12
3.3.2 Specifying (Just) the Topics
What if you know what topics you want to ask about, but no more than that?

The next level of pre-specification occurs when the aims of the instrument are
indeed pre-established—in the terms introduced above, one can call this the topics guide
format (i.e., in the second and third rows of the matrix). Patton (1980), in the context of
interviewing, labels this the “interview guide” approach—the guide consists of:

a set of issues that are to be explored with each respondent before


interviewing begins. The issues in the outline need not be taken in any
particular order and the actual wording of questions to elicit responses
about those issues is not determined in advance. The interview guide
simply serves as a basic checklist during the interview to make sure that
there is common information that should be obtained from each person
interviewed. (p. 198)

Figure 3.6 Levels of Pre-specification for Different Item Formats

Description of
item components Specific items

Intent to
measure No Score Score
Item Format construct General Specific Guide Guide Responses

Participant X After After After After After


Observations

Topics guide Before X After After After After


(a) General

Topics guide Before Before X After After After


(b) Specific

Open-ended Before Before Before X After After

Scoring Guide Before Before Before Before X After

Fixed response Before Before Before Before Before X

Note: “X” marks the column which encompasses the main activity at each level of Item
Format.

13
For Topics Guide, two levels of specificity can be distinguished. At the more
general level of specificity (second row of the matrix), the definition of the construct, and
the topics, are only specified to a summary level—this called the general topics guide
approach. Presumably, the full specification of these will occur after observations have
been made. At greater degree of specificity (third row of the matrix), the intended
complete set including the construct definition and the full set of topics is available
before administration—hence, this is the specific topic guide approach. The distinction
between these two levels is a matter of degree—one could have a very vague summary
and alternatively, there could be a more detailed summary that was nevertheless
incomplete.

3.3.3. Open-ended Items


What is the most common form of item?

The next level of pre-specification is the open-ended level (rows four and five of
the matrix). This includes the very common open-ended test and interview instruments
such as were mentioned at the beginning of this chapter—used by teachers and in
informal settings the world over. Here the items are determined before the administration
of the instrument, and are administered under standard conditions, including a pre-
determined order. In the context of interviewing, Patton (1980) has labeled this the
“standardized open-ended interview”. Like the previous level of item format, there are
two discernible levels within this category. At the first level (fourth row of the matrix),
the response categories are yet to be determined. Most tests that teachers make
themselves and use in their classrooms are at this level. At the second level (fifth row of
the matrix), the categories that the responses will be divided into are pre-determined—
call this the scoring guide level. Examples of the open-ended item format (also known as
constructed response items) have been given in Chapter 1 (Figure 1.6) and in this chapter
(Figure 3.4, left panel, and Figure 3.7).

An interesting case of open-ended items are the structured interview questions


used in the CUE assessments. Recall that these are investigating students’ knowledge
about natural selection. In this interview, students were presented with a task illustrated
with a visual stimulus (See Figure 3.7). Each task related to one of the strands of the
Construct Map (given in Figure 2.8). For example, in the fit dimension the
“Rainforest/Desert Task,” which relates to the fit dimension begins with the interviewer
showing the student two large photographs, one of the deserts and one of the rainforests
(See Figure 3.7).

14
Figure 3.7 A Student Responding to the Rainforest/Desert Task in the Cue interview

The interviewer explores the student’s thinking about the whether the plants in one
environment could live in other environment and why or why not, using the photos as
illustration and referent information. The interviewer asks:

Do you think these plants—that live in the rainforest—could survive in the desert too? If
S answers no: Why not? Any other reasons?

If the student answers yes, then the interviewer asks:

Why? Any other reasons?


Are there any places in the world where these plants couldn’t live? Is there any other
reason they couldn’t live there?
Following this line of questioning, the interviewer probes the students’ thinking about the
reverse situation from desert to rainforest using the same question structure. The last part
of the item asked the student about how plants survive in the desert, essentially to explain
the fit between organism and this relatively extreme environment.
I’ve got another question for you about these desert plants. How come they can survive
here in the desert where there is so little water?
Children’s responses to this task reflected a broad range of sophistication, reflecting each
level within the fit dimension of the construct map (Metz et al., 2019). The relevant
codes, developed after many hours of coding and recoding effort by the project for this
type of task, are shown in Figure 3.8—these were augmented by extensive coding notes,
and all interviews were (at least) double-coded to check for consistency.

15
Figure 3.8 Coding Elements developed by the CUE Project

Live Where They Belong


Meeting Needs
Limiting factor X Differences in Organisms’ Structures
Survival/ risk value of trait [for the individual]
Those individuals with advantageous trait more likely to have offspring
Change in relative frequency (as in will be more of XXX)
Inheritance
Change in relative frequency from original generation to offspring generation
Accumulation of changes over many generations
Result of natural selection as organisms adapted to where they live
Other:
Answer to the question from student perspective but falls outside natural selection
concept
No Response/I Don’t Know
Uncodeable: 1. Silence, I don’t know, & Umm 2. Incomprehensible 3. Inaudible
Question Skipped

3.3.4 Selected Response Items


What is the most structured form of item?

The final level of specificity is the standardized fixed-response format typified by


the commonly occurring multiple-choice and Likert-style items (sixth and last row of the
matrix). Here the respondent chooses rather than generates a response to the item. As
mentioned above this is probably the most widely used item form in published
instruments.

Multiple-choice items are ubiquitous in educational assessment, and all should be


familiar with them already, but, just in case, a very modest example is shown in Figure
3.9. Under some circumstances, it can be interesting, even enlightening, to consider
alternative ways of scoring outcome categories. For example, in the case of multiple-
choice items, there are sometimes distractors that are found to be chosen by “better”
examinees than some other distractors (in the sense that the examinees obtained higher
scores on the instrument as a whole, or on some other relevant indicator). When this
difference is large enough, and when there is a way to interpret those differences with
respect to the construct definition, then it may make sense to try scoring these distractors
to reflect partial success. For example, consider the multiple-choice test item in Figure
4.1: A standard scoring scheme would be: A or C or D = 0; B = 1. Among these
distractors, it would seem reasonable to think that it would be possible to assign a
response C to a higher score than A or D, because Ghent is also in Belgium, and the other
two cities are not. Thus, an alternative hypothetical scoring scheme would be: A or D =
0; C = 1; B = 2. A similar analysis could be applied to any other outcome space where
the score levels are themselves meaningful. This can be informed by the analysis
described in Section 8.4.5.

16
Figure 3.9. An example of a multiple-choice test item that would be a candidate for
ordered multiple-choice scoring.

Q. What is the capital city of Belgium?


A. Amsterdam
B. Brussels
C. Ghent
D. Lille

This possibility can be built into multiple choice items right at the design stage
when a construct map has been used to design the instrument. By developing options that
each relate to different waypoints and ensuring that there are more than two waypoints
involved (say, three or four), then the options can be ordered with respect to the construct
map, resulting in an ordered multiple-choice (OMC) item (Briggs et al., 2006). This
development technique offers a way to improve the interpretability of outcomes from
traditional-looking multiple-choice instruments.

The GEB and PF-10 instruments described in previous chapters are examples
(Examples 3 and 6, respectively) of items with Likert-style response options. This item
format is equally as common in social science settings (and beyond) as multiple-choice
are in educational testing, so I will not give further examples. However, there is a related
response format that has been found to give better results, the Guttman-style item (Wilson
et al., 2021). In this format, the options are ordered according to the underlying
construct, and, when the instrument is being developed according to a construct map, the
most convenient way to do this is to design each option to match successive waypoints.
This is exactly what was done in the RIS example (Example 2—see the construct map in
Figure 2.4), where an initial set of Likert-style items was used to create a set of Guttman-
style items. An example is shown in Figure 3.10. Initially, each of these options was the
stem of a Likert-stle option (with choices from Strongly Agree to Strongly Disagree)—
changing the response option sin this way focusses the respondent on the construct under
measurement, using context-specific words to describe the relevant waypoints, and away
from merely reporting their agreement/disagreement level, which is at both vague and
very variable across respondents.

17
Figure 3.10 An example item from the RIS (Guttman response format items).

G4. Which statement best describes you?

(a) I don’t consider myself a part of a research community.


(b) I am beginning to feel like a part of a research community.
(c) I am a small part of a research community.
(d) I am a part of a research community.
(e) I am an important part of a research community.

The above description of this sequence of formats has focused on what has to
happen before the item is administered. But, in fact the sequence of columns in the
matrix (as shown in Figure 3.6) also lays out what has to happen after the item is
administered if, indeed, the responses are to become the input to a measurement.
Essentially, in each row, when one looks to the right of the “X” (the point at which that
item format is administered), the steps remaining are indicated by the remaining columns.
Thus, texts that started as responses in a participant observation context (row one of the
matrix) would still have to go through the multiple stages of qualitative analysis
including the derivation of general and specific categories for the contents of the texts
(i.e., columns 2 and 3 of the matrix), then the development of coding and scoring guides
for the texts (columns 3 and 4 of the matrix), and the actual coding of them (column 5).
This same logic applies to each of the remaining rows, until row six, where there is no
further need for recoding, etc.— at this point each respondent has coded their own
response5.

3.3.5 Steps in Item Development


From open-ended to selected response

As the motivation to create a new instrument is almost certainly that the


measurement designer wants to go beyond what was done in the past, it is important that
the measurer bring new sources of information to the development, beyond what will be
learnt from a literature review. One important source of information can be found through
exactly the sort of participant observation approach that has been described in the
previous section. The measurer should find situations where people who would be typical
respondents to the planned instrument could be observed and interviewed in the informal
mode of participant observation. That might include informal conversational interviews,
and collections of products, recordings of performances, etc. Information from these
processes is used to develop a richer and deeper background for the “theory of the
construct” that the measurer needs to establish the construct (i.e., the waypoints of the
construct map), and the contextual practices that are necessary to develop the secondary
facets of the instrument. The set of informants described in Section 1.8 would be of help
in this process, some as participants and some as observers.

5
Of course, the measurer may want to carry out further analysis and recoding after that.

18
At this level of specificity, and at each of the levels noted below, item
development should include a thoughtful consideration of the range of
participants/respondents and how their differences may alter how items are understood.
This should include consideration of respondents who are likely to be from different
waypoints of the construct, all the standard demographic groups (gender, ethnic-racial
identity, class, etc.) as well as specific groupings relevant to the construct under
measurement, such as, in the case of, say, the scientific argumentation construct
discussed above, students with different amounts of experience in science, and different
levels of interest in science.

Following the initial idea-building and background-filling work of the literature


review and the participant observations, the measurer should try an initial stab at the
items design topics guide. This is difficult to do in a vacuum of context, so, at the same
time, it is necessary to also develop some initial drafts of items. This is even true if the
plan is to leave the instrument at the topics guide level, as it is essential to try-out the
guides in practice (i.e., that means actually doing some interviews, etc.). The
development of the construct through the idea of construct map has already been
discussed in Chapter 2. The development of the other components will require insights
from the participant observation to know what to focus on, and how to express the
questions appropriately—some similar information may be gleaned from the literature
review, although usually such developmental information is not reported in refereed
journals6. The decision of whether to stop developing the topics guide at a summary
level, or whether to go on to the finer grained specific topics guide will depend on a
number of issues, such as the amount of training that the measurer will devote to the
administrators of the instrument, and the amount of time and effort that can be devoted to
the analysis. But, if the aim is for the finer level, then inevitably the coarser level will be
a step along the way.

Going on to an open-ended format will require either the generation of a set of


items or the development of a method for automatically generating them in a
standardized way. The latter is rather rare and quite specialized, so it will not be
addressed here (but see Williamson et al., 2006). Item development is a skill that is
partly science, partly engineering and partly art. The science lies in the construction
and/or discovery of theoretically sound and useful constructs, the engineering lies in the
development of sound specifications of the facets, and the art lies in making it work in
context. Every context is unique. If the aim is to develop fixed response items, then a
further step is needed. This step is discussed in the next chapter.

When items are organized into instruments, there are also issues of instrument
format to consider. An important dimension of instrument design is the uniformity of the
formats within the instrument. An instrument can consist entirely of a single item format,
such as is typical in many standardized achievement tests where all are usually multiple-
choice items, and in many surveys, where Likert-type items are mostly used (though
sometimes with different response categories for different sections of the survey). But
6
But indeed, most refereed journal do not insist on such information. Lamentably often they recommend
deletion of such information to save on page length.

19
more complex mixtures of formats are also used. For example, the portfolio is an
instrument format common in the expressive and performance arts, and also in some
professional areas. This will consist of a sample of work that is relevant to the purpose of
the portfolio, and so may consist of responses to items of many sorts and may even be
structured in a variety of ways more or less freely by the respondent according to the
rules that are laid down. Tests may also be composed of mixed types—multiple choice
items as well as essays, say, or performance-tasks of various sorts. Surveys and
questionnaires may also be composed of different formats, true-false items, Likert-style
items, and short answer items. Interviews may consist of open-ended questions, as well
as forced choice sections. Care must be taken in devising complex designs such as these
—concern must be taken with respect to the time allowed for the different sections, the
extent of coverage of each, and way that the items of different formats are related to the
construct map(s).

3.4 A Unique Feature of Human Measurement: Listening to the


Respondents
What can respondents tell us about the items?

A crucial step in the process of developing an instrument, and one unique to the
measurement of human beings, is that the measurer can ask the respondents what they are
thinking about when responding to the items. In Chapter 8 evaluative use of this sort of
information is presented as a major tool in gathering evidence for the validity of the
instrument and its use. In this section, formative use of this sort of information is seen as
a tool for improving the instrument, and in particular, the items of the instrument. There
are two major types of investigations of response processes, the “think-aloud” and the
“exit interview.” Other types of investigation may involve reaction time studies, eye
movement studies and various treatment studies, where, for example, the respondents are
given certain sorts of information before they are asked to respond.

In the think-aloud style of investigation, also called “cognitive labs” (AIR, 2000),
students are asked to talk aloud about what they are experiencing as their mental
activities while they are actually responding to the item. What the respondent says is
recorded, often what they do is being videotaped, and other characteristics may be
recorded, such as having their eye movements tracked. A professional should be at hand
to prompt such self-reports, and also to ask clarifying questions if necessary. Typically,
respondents need a certain amount of familiarization to know what it is that the
researcher is interested in, and to feel comfortable with the procedure. The results can
provide insights ranging from the very plain—such as “the respondent was not thinking
about the desired topic when responding”—to the very detailed, including evidence about
particular cognitive and metacognitive strategies that they are employing. Of course, the
scientific value of self-report information should be questioned (e.g., Brenner et al.,
2003). But, in this circumstance, the evidence from the think-aloud should be considered
more like forensic evidence in a criminal investigation—not directly indicating what has
really happened (mentally) but giving the measurer clues that can be linked to reasonable
hypotheses about what is going on in the hidden mental world of the respondent.

20
A sample think-aloud protocol is provided in Figure 3.11 which was developed as
a part of the LPS project (Example 4). Note that the preamble to the whole process,
designed to set the respondent at ease, and provide them with disclosure of the aims of
the exercise7. Then the interviewer demonstrates what they want the respondent to do
using a similar context (i.e., describing how the student would get to their school office),
for the respondent to become familiar with the procedures. Other familiarization
procedures are also commonly used, such as conducting a “dry run” with an item similar
to those in the actual study-set. The exercise then proceeds through each item to be
investigated. In this case, the cognitive lab was being recorded on a verbal recording
device, but it may also be video-taped, or recorded via a video-conferencing system
(although care must be taken with these to make sure that the respondent is not identified
in the stored data). One additional good idea to have a silent observer present during the
entire activity to assist in interpretation and evaluation of the results with the interviewer.

Four types of information are usually produced from these procedures:


(a) process records that show what a student does as they solve the item (i.e., verbal
and/or video recordings),
(b) products by the respondent (i.e., both the student responses to the item and student
jottings etc., while responding to the item),
(c) introspective reports from the respondent (i.e., respondent comments as they attempt
to respond), and
(d) retrospective information (i.e., respondent comments after they have completed the
item).
Regarding the fourth source of information, the prompts for this were not included in the
protocol, but the researchers noted that the interviewers ...
asked questions based on events that arose during the think-aloud protocol. For
example, we may have asked a process question (“How did you solve that?”)
when the student did not adequately verbalize. Or we may have asked a design
question (“Was there anything that confused you?”) when a student spent several
minutes on a sub-section of an item. (Johnstone, et al., 2006, p. 7)

The data from these data sources are then coded to record the relevant aspects of
the information, and to allow accumulation and comparison across respondents. Each
development project will need to create its own version of a think-aloud coding sheet
according to the information they wish to glean from the think-aloud process. Some of
the relevant pieces of information that will be likely to valuable to record are:
(i) how accessible the questions were for the respondents,
(ii) whether the items appeared to be biased for (or against) certain subgroups (this, of
course, depends on their being representatives of these groups among the sampled
respondents, as noted above)
(iii) whether the respondents found the items to be simple and clear, and whether the item
had intuitive instructions and procedures
(iv) whether the text of the item was readable and comprehensible,
(v) whether the presentation of the items was legible, and
(vi) whether anything was left out.
7
The scope of this, of course, must be agreed to by your institutional IRB.

21
Figure 3.11 Sample script for a think-aloud investigation

22
Figure 3.10 Sample script for a think-aloud investigation (continued)

23
Figure 3.10 Sample script for a think-aloud investigation (continued)

24
An example of responses gathered during a think aloud session for the (early
version of an) item from the LPS project (shown in Figure 3.12) is shown in Figure 3.13.
The item relates to a different construct map used in the LPS project, for Ecology 8. Some
appreciation of the typical range of responses one gets from a think-aloud exercise can be
gained by reading through the various comments recorded in Figure 3.12—they range
from helpful corrections for poor wordings, to insights into how students are
misunderstanding the ideas involved, to insights into how students are misunderstanding
the representation of the ideas in the item figure, to student vocabulary deficits. In the
revisions to this item, the project:
(a) revised the preamble to read “The picture below shows the changes in an ecosystem
over time after a big fire. Scientists call this ecological succession”; and
(b) modified the representation to make clearer that the panels in the figure were views
(snapshots) of the same place over time (ie., by adding an element that is common
to all of them).

Figure 3.12 An Early Version of an LPS Item

Succession

The picture below shows the process of ecological succession after a big fire,
which occurs after an ecosystem experiences a disturbance (like a big fire).

S12. What changes do you notice across the years in the figure above?

8
The details of that construct map are not needed to appreciate the observations about this example of
think-aloud notes, but in case the reader is interested one can look up that information via the Examples
Archive in Appendix A.

25
Figure 3.13 Example Notes from a Think-aloud Session

Extract from Interviewer Notes for “Succession”

102, 106: Repetitive statement: “The picture below shows the process of
ecological succession after a big fire, which occurs after an ecosystem
experiences a disturbance (like a big fire).”

201: Student seemed to focus only on the most obvious change (bare forest to
smaller trees to taller trees), without examining the figure more closely for other
“smaller” changes. Question design may not facilitate close/detailed examination
of the figure. Maybe we can ask something like “Several patterns of change can
be observed in the above figure. Describe some of these patterns.”

202: The word “hardwood” might be unfamiliar to students. Student also said that
it was confusing why the vegetation would keep changing from one type of plant
to another, especially how pine would turn into hardwood (i.e., student does not
understand what a succession is). He thought that the pine trees might have died
out (because of a drought) then hardwood trees grew over the land. But student
knew that the grids were showing the same location.

202: Student did not seem to notice the smaller patterns/changes besides the
change in plant growth (bare to smaller trees to bigger trees), despite multiple
prompting.

110: Student thinks the plants over the years are the same plants just growing
bigger (less bushy to more bushy).

These words were confusing in the Succession task

102, 105 Ecological succession,

104 succession,

106 Disturbance, ecological (but the student said they weren’t that hard)

The exit interview is similar in aim to the think-aloud but is timed to occur after
the respondent has completed their item responses. It may be conducted after each item,
or after the instrument as a whole depending on whether the measurer judges that the
delay will or will not interfere with the respondent’s memory. The types of information
gained will be similar to those from the think-aloud, though generally it will not be so
detailed. This is not always a disadvantage, as sometimes it is exactly the results of a
respondent’s reflection which are desired. Thus, it may be that a data collection strategy
that involves both think-alouds and exit-interviews will be best.

26
Information from respondents can be used at several points along the instrument
development process as detailed in this book. Reflections on what the respondents say
can lead to wholesale changes in the idea of the construct, revisions of the construct facet,
the secondary facets, and specific items and item types, and changes in the outcome space
and scoring guides (to be described in the next Chapter). It is difficult to over-emphasize
the importance of including procedures for tapping into the insights of the respondents in
the instrument development process. A counterexample is useful here—in cognitive
testing of babies and young infants, the measurer cannot gain insights in this way, and
that has required the development of a whole range of specialized techniques to make up
for such a lack.

Another issue that distinguishes measurement of humans from other sorts of


measurement is that the measurer is obliged to make sure that the items do not offend or
elicit personal information that might be detrimental to them, ask them to carry out
unlawful or harmful procedures, or unduly distress them. The steps described in the
previous section will prove very informative of such matters, and the measurer should
heed any information supplied by the respondents supply and make subsequent revisions
of the items based on these insights. But simply noting such comments is not sufficient.
There should be prompts that are explicitly aimed at addressing these issues, as the
respondents may think that such comments are “not wanted” in the think-aloud and exit
interview processes.

For example, to investigate whether items are offensive to potential respondents,


it is useful to assemble a group of people who are seen as representing a broad range of
potential respondents (these groups have various titles, such as a “community review
panel”). The specific community demographic categories that should be represented in
the group will vary depending on the instrument and its audience, but likely demographic
variables would be: age, gender, ethnic-racial identity, socio-economic status, language
status, etc., as well as specific groups relevant to the construct under measurement This
group is then asked to examine each item individually and the entire set of items as a
whole to recommend that items be deleted or amended on the grounds mentioned in the
previous paragraphs, or any other reasons that they feel are important. Of course, it is up
to the measurer to decide what to do with such recommendations, but they should have a
justification for not following any such suggestions.

3.5 Building-in Fairness through Design


How fair can an item be?

The question of whether the items in an instrument are fair to respondents should
be addressed at the design stage. According to the Meriam-Webster dictionary, fair
means: “marked by impartiality and honesty: free from self-interest, prejudice, or
favoritism.” At this point in instrument development, in the process of developing items,
it is a formative question; however, this issue will be revisited as a part of the evaluation
of the fairness evidence for an instrument in Chapter 8. The essential question is whether
there are important influences (besides the underlying construct that the instrument is

27
intended to measure) on a respondent’s reactions to specific items in an instrument, or
indeed the whole set of items. For example, consider the many well-established instances
in educational testing, as noted by a U.S. National Research Council (NRC) report ...
... in a written science assessment [or for any other subject matter] with open-
ended responses, is writing a target skill or an ancillary skill? Is the assessment
designed to make inferences about science knowledge, about written expression
of science knowledge, or about written expression of science knowledge in
English? The answers to these questions can assist with decisions about
accommodations, such as whether to provide a scribe to write answers or to
provide a translator to translate answers into English. If mathematics is required
to complete the assessment tasks, is mathematics computation a target skill or an
ancillary skill? Is the desired inference about knowing the correct equation to use
or about performing the calculations? (Here, the answers can guide decisions
about use of a calculator.) (NRC, 2006)

These are questions that can be asked for each respondent, but usually it is
couched in terms of how different groups of respondents will respond to the items. These
groups may differ from construct to construct and from situation to situation but, in a
given context, will generally correspond to typical demographic groups. Staying within
the educational testing context, these groups could be defined so as to include, say,
respondents from different gender, ethnic and racial groups, respondents with a different
language status (such as respondents whose native language is not the language of the
instrument), respondents with learning challenges (such as cognitive disabilities, learning
disabilities, or who are deaf or hard of hearing, etc.), or respondents with non-standard
citizen status (such as recent migrants, or visa-holders, etc.).

In addressing this issue in terms of the construct map and the development of
items that are sensitive to the construct, typically, the investigation of such effects on
demographic groups would be structured by assuming that there is (a) a reference group
(i.e., the group to whom the other groups will be compared), and (b) one or more focal
groups (i.e., groups for whom the fairness is in question). Then, the measurer will need to
consider the following three successively more complicated possible types of difference
in the way that the construct map and the items work between the focal and the reference
groups. These possibilities are:
(i) in terms of the measurements themselves, the respondents in the focal group are
essentially on the same construct map as the respondents in the reference group,
but they tend (on average) to be lower or higher than those from the reference
group (this is termed “differential impact”);
(ii) the respondents in the focal group are essentially on the same construct map, however
some items (or perhaps a class of items) in the instrument behave differently for
respondents in the focal group compared to their reference group peers who have
a matching level of on the construct map. This phenomenon has the generic label
“differential item functioning” or DIF9; see Gamerman et al., 2016, Kamata &
Vaughn, 2011; and Chapter 8); and

9
When this occurs for the aggregate of the whole test, this is termed “differential test functioning”)

28
(iii) the respondents in the focal group are on a different construct from those in the
reference group10.

Clearly, case (i) is the least complicated of the three—we might say that there is
measurement-fairness here. But it is important to distinguish between the fairness of the
measurement process itself (i.e., the instrument fairly maps from the student responses to
the eventual outcome), and real world fairness (i.e., there is no unfairness in the ways
that the individual respondents have gotten to this point of being measured). Examples of
real world unfairness night be where some individuals have had more resources devoted
to them in the forms of better nutrition, education or medical care. It might be naïve to
assume that, when no DIF is found that there is no unfairness that has occurred,
especially if there is a history of concerns over fairness in the specific contexts where the
instrument is typically used. To establish that case (i) does indeed hold, the remaining
two possibilities need to be examined, and eliminated as important effects.

With regards to case (ii), measurement researchers have investigated differential


item functioning with respect to many different focal groups and constructs. For example,
effects have been found with respect to construct-irrelevant language factors in items that
may prevent English learners (ELs) from demonstrating what they know and can do,
most notably in large-scale summative assessments in mathematics (e.g., Daro et al,
2019; Mahoney, 2008), but consistent findings have been rare. In a second example,
writing about the same situation, Lee and Randall (2011) have reviewed studies of DIF in
mathematics assessments between English Learners (ELs) and students who are not ELs,
focusing on how the complexity of the language used in the items (i..e., “language load”)
does influence DIF. They found that there were no consistent effects across the 8 studies
they located. They did find in their own additional study that a large proportion of the
items from the mathematics content area “Data Analysis, Statistics, and Probability”
showed DIF against ELs, but that the items in this area did not have much evidence of
language complexity, and they concluded by speculating that the DIF effects were more
likely due to differential educational exposure to these topics in their classrooms than to
item format effects.

Case (iii) is the most complex but relatively little research exists investigating this
possibility. Looking again at the case of English learners (ELs) learning mathematics, this
is at least partly due to the paucity of clearly established models of how ELs typically
learn and develop competence. But emerging research suggests that often ELs do not
follow the typical learning progressions—this should not be too surprising as our
understanding of these earning progressions are based primarily on research and models
for non-ELs (Sato et al., 2012). One would expect that, given individual differences in
educational history, socio-cultural background, and literacy and fluency in their native or
home language, ELs would interact with, academic content differently compared to non-
ELs (e.g., Solano-Flores & Trumbull, 2003). Yet a further layer of complexity is the
concern that ELs are a not a uniform group—there are substantial differences among ELs
in terms of the cognitive, cultural and linguistic resources as well as their prior
educational backgrounds that they bring to assessment tasks (Abedi, 2004).
10
This would be termed as a lack of factorial invariance across these groups.

29
On the other hand, some evidence suggests that learning progressions can be
appropriate tools for examining learning and development in students from different
cultural and linguistic backgrounds. In the area of early language and literacy, the
literature agrees that, despite opportunity gaps that lead to average differences between
groups, ELs have similar learning trajectories as their monolingual English-speaking
peers in areas of vocabulary development, grammar, phonemics, and writing (Hammer et
al., 2014). An empirical analysis, incorporating a learning progressions approach, found
that groups of young children in public early care and education programs who spoke
different languages at home had markedly similar trajectories of early language and
literacy development (Sussman et al., under review).

The implication of these complexities for measurement is that measurers need to


take a pro-active role during the design of items to investigate how the items interact with
the subgroups in the intended population of respondents. In particular, for ELs, measurers
should investigate national and cultural variations in how ELs access, engage, and
respond to assessment tasks which generic learning progressions may miss (Mislevy &
Duran, 2014). Measurers should also consider ELs’ educational trajectories, focusing on
their experiences in content domains as well as language domains. The related
implication for measurement is that research that is specifically aimed at developing
alternative construct maps for ELs and other focal groups is needed

3.5.1. Universal Design


When is the best time to be fair?

A guiding principle known as universal design, initially developed in design


professions such as architecture, can be adapted into the measurement context. The idea
is that products, say buildings—or measurements—should be designed so that a maximal
number of people can use them without the need for further modification—in other
words, they should be designed to eliminate unnecessary obstacles to access, and
unnecessary limitations on people’s success in using the products. In the case of
measurement practice this can have several sorts of manifestations. For example, if a test
is not of a speeded variety (i.e., the time limitation is not a part of the construct), then
offering extra time to those who need it because of specific known difficulties would not
provide those respondents with an unfair advantage over the rest of the respondents11.

To this end, Thompson et al. (2002) have proposed seven categories of


recommendations for examination of items and instruments to enhance their universal
design features:
(a) studying an inclusive sample of the respondent population;
(b) developing precisely defined constructs;
(c) designing accessible, non-biased items;
(d) designing items that are amenable to focal group accommodations;
(e) writing simple, clear, and intuitive instructions about the procedures;

11
Further information on this topic is available from the National Center on Educational Outcomes
(NCEO) ([Link]

30
(f) designing for maximum readability and comprehensibility; and
(g) designing for maximum legibility.
Examples of each of these are given in Johnstone et al. (2006). In particular, given that
the reader is reading this book, the advice to develop “precisely defined constructs”
should be already at the top of the measurer’s list. But it is important, in light of the
discussion above, that the measurer be open to the possibility that the constructs may be
structured differently between reference and focal groups and investigate it thoroughly.
Apart from engaging in a focused literature search for evidence of this, the measurer
should also see that the cognitive lab methods described in Section 3.4 will be useful in
giving clues, assuming, of course, that the focal groups have been included among the
sampled respondents, which corresponds to recommendation (a) above.

A useful framework for universal design in a measurement context has been given
by Robert Dolan and his colleagues (Dolan et al, 2013). Nathaniel Brown has prepared a
summary of such issues (Brown, 2020) that is helpful in exploring the multifaceted nature
of the design issues, and this is included as Appendix 3B.

3.6 Resources

Chapters 1 through 3 have provided the general necessary resources to create an


items design and generate items using the creativity, insight, and hard work of the
measurement developer. In terms of specific resources, there is far too wide a range of
potential types of constructs, areas of application, and item formats to even attempt to list
particular sources here. Nevertheless, the Exercises and Activities at the end of Chapters
1 and 2 should have directed the reader towards a better understanding of practices within
their area of interest including the background, current and past practices, and the
relevant range of item designs.

Within the area of educational achievement testing, there are several very useful
resources for types of items and methods to develop them such as: Brookheart & Nitko
(2018), Haladyna (1996, 1999), Nitko (1983), Osterlind (1998), and Roid and Haladyna
(1982). Many if not most other areas will have similar resources, and your informants
should be able to put you onto them. We recommend you consult with professors,
colleagues, and experts in your area of interest for more specific direction and support.
Such informants, especially those who have experience in instrument development, can
help explain specific issues that may arise, provide insight during the item development
process, and help critique your overall process.

3.7 Exercises and Activities

(following on from the exercises and activities in Chapters 1 and 2)

1. Generate lots of types of items and several examples of each type. Be prolific.

2. Write down your initial items design based on the preceding activities (including the
activities listed for Chapters 1 and 2)—devote most of your attention to the construct

31
facet but listen to your informants about what should be the most important secondary
facets.

3. Give these draft items a thorough professional review at an “item panel” meeting—
where key informants and others you recruit to constructively critique the items generated
so far (see Appendix 3A).

4. Write down your interim Items Design based on the preceding activities, including
predicted numbers of item you (still) will need to generate.

5. Following the initial round of item generation and item paneling described in (1)
through (4), second or third rounds may be needed, and may also involve a return to
reconsider the construct definition or the definition of the other facets of the items design.

6. Enter your surviving items into the BASS Item Bank for the construct maps you are
working on.

7. Think through the steps outlined above in the context of developing your instrument
and write down notes about your plans.

8. Share your plans and progress with your informants and others—discuss what you are
succeeding on, and what problems have arisen.

32
Appendix 3A

The Item Panel Meeting

1. How to prepare for the item paneling.


(a) For each item you generate, make sure
(i) you can explain its relationship to the framework,
(ii) that you can justify that it is appropriately expressed for the respondents,
(iii) that it is likely to generate the sort of information that you want, and
(iv) that the sorts of responses it elicits can be scored using your scoring guide.
(b) If possible, first try out the items in an informal, but hopefully informative, small
study using several of your informants or other likeley suspects. Ask them to take the
instrument, and to comment on what they thought of it.
(c) For each part of the framework that you have decided to measure, make sure that
there are a sufficient number of items (remembering that a 50% loss of items between
generation and final instrument is very likely). Note that you may well make too many
items to actually panel, so also indicate a subset that will definitely be discussed in the
panel, bearing in mind the most important different subtypes of items in your item bank.

2. Who should be at the Item Panel meeting?


The panel should be composed of the same profile of people as your informant group:
(a) where possible, some potential respondents,
(b) professionals, teachers/academics and researchers in the relevant areas, as well as
(c) people knowledgeable about measurement in general and/or measurement in the
specific area of interest, and
(d) other people who are knowledgeable and reflective about the area of interest and/or
measurement in that area such as administrators and policymakers, etc.
(e) You may need to have someone else dedicated to the role of note-taker, as it can be
difficult to chair the meeting as well as take notes.

3. What to supply to the panelists.


At least a week ahead send each panelist the following:
(a) The full construct map framework, including exemplars and other relevant materials,
along with suitable (but not overwhelming) background information;
(b) A description of how the instrument will be administered and the responses coded
and/or scored;
(c) A list of the items, with indications about how each relates to the waypoints in your
construct map, including quotes from respondents where they are available;
(d) Any other relevant information (remember, for judged items, each panelist should be
a reasonable facsimile of a judge).
Offer to discuss any and all of this with the panelists if they have questions or difficulties
understanding what they have been sent.

33
4. How to carry out the Item Paneling.
(a) You will chair the meeting. Your aim, as chair, is to help each panelist contribute in
the most constructive way possible to creating the best set of items that well generate
responses that are related to the construct map. Each panelist needs to understand that
that is the aim. Panelists are to be as critical as they can, but with the aim of being
constructive as well. Disputes are to be left in the room at the end—the chair/item
developer will decide what to do with each comment.
(b) Order of business:
(i) Explain the purpose of the panel in case some panelists are new to the
procedure
(ii) Give panelists a brief overview of the construct map and the context for use of
the planned instrument, including a description of the intended
respondents and raters (if applicable). Invite comments and questions.
(iii) Systematically discuss the items to be paneled, keeping in mind as you go the
expected length of the meeting and the number of items to be paneled.
Before passing on to the next item, be sure that you are satisfied that you
are clear what are the panel’s recommended revisions for the current item.
(iv) After surveying all items ask for general comments on the item set and
especially whether the item set comprehensively represents the framework
for the variable.
(v) Collect any written notes about the items from the panelists

5. What to do after the Panel meeting.


(a) Immediately after the meeting:
(i) Go over your notes (with the note-taker) so that you are clear about
recommended action for each item.
(ii) Follow-up any matters that were not clear in your review of the notes.
(b) Reflect upon the revisions recommended by the Panel members—bear in mind that
their recommendations are not necessarily correct—decide which to accept, which to
modify and which to follow up on.
(c) Carry out the revisions you have decided upon to items (extrapolating to items not
actually paneled), the construct map and other materials.
(d) Send the revised items (and revised construct map and background materials, if that
was recommended) to the panel for follow-up comments.
(e) Make further changes depending on their responses, and your own judgment.
(f) If you are not sure that there has been sufficient improvement, repeat the whole
exercise.

34
Appendix 3B

Supporting Construct-Irrelevant Actions12

12
Brown, N. (2020). Supporting Construct-Irrelevant Actions. Research Report, Graduate School of
Education, Boston College.

35
36

You might also like