Contexts
Contexts
Edited by
Diane Kelly-Riley
Ti Macklin, and Carl Whithaus
CONSIDERING STUDENTS,
TEACHERS, AND WRITING
ASSESSMENT: VOLUME 1,
TECHNICAL AND POLITICAL
CONTEXTS
PERSPECTIVES ON WRITING
Series Editors: Rich Rice and J. Michael Rifenburg
Consulting Editor: Susan H. McLeod
Associate Editors: Jonathan M. Marine, Johanna Phelps, and Qingyang Sun
Foreword. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . vii
Kathleen Blake Yancey and Brian Huot
Introduction to Volume 1, Technical and Political Contexts. . . . . . . . . . . . . . 3
Diane Kelly-Riley, Ti Macklin, and Carl Whithaus
Part 1. Technical Issues in the Assessment of Writing:
Reliability and Validity. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17
Retrospective. From Isolation to Integration: Technical Issues in the
Assessment of Writing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19
David H. Slomp
Chapter 1. Reframing Reliability for Writing Assessment. . . . . . . . . . . . . . . . 37
Peggy O’Neill
Chapter 2. Validity Inquiry of Race and Shared Evaluation Practices
in a Large-Scale, University-Wide Writing Portfolio Assessment. . . . . . . . . . . 65
Diane Kelly-Riley
Chapter 3. Three Interpretative Frameworks: Assessment of English
Language Arts-Writing in the Common Core State Standards Initiative. . . . . 93
Norbert Elliot, Andre A. Rupp, and David M. Williamson
Part 2. Politics and Public Policy of Large-Scale Writing
Assessment. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 123
Retrospective. The Politics and Public Policy of Large-Scale Writing
Assessment. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 125
Carolyn Calhoon-Dillahunt
Chapter 4. The Misuse of Writing Assessment for Political Purposes. . . . . . . 143
Edward M. White
Chapter 5. Issues in Large-Scale Writing Assessment: Perspectives
from the National Assessment of Educational Progress. . . . . . . . . . . . . . . . . 161
Arthur N. Applebee
Chapter 6. The Micropolitics of Pathways: Teacher Education, Writing
Assessment, and the Common Core. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 183
J. W. Hammond and Merideth Garcia
Chapter 7. Writing Assessment, Placement, and the Two-Year College. . . . 209
Christie Toth, Jessica Nastal, Holly Hassel, and
Joanne Baird Giordano
v
Contents
vi
FOREWORD
It’s a truism to note that writing assessment has come into its own during the
last several decades, and one of the factors propelling that growth is the Journal
of Writing Assessment (JWA). As the selected articles-now-chapters presented here
suggest, writing assessment is both more and different than what it seems to be.
While it can simply appear as a rudimentary exercise in evaluating writing, writ-
ing assessment, as the authors here have documented, researched, and theorized,
is at least twofold: (1) an exercise of considerable sophistication and complexity
operating within a context (2) that can overwhelm, and sometimes sabotage,
the exercise itself. These twin observations informed our goal when we created
JWA, a new journal focused on writing assessment that would circulate schol-
arship taking up questions about how to best assess writing as well as about the
contextual factors, often invisible, that shape and, too often, mis-shape writing
assessment. Put in the current vernacular, with JWA we hoped to make writing
assessment—and its many dimensions—transparent.
The articles in our first issue of JWA made this goal visible. In “Moving
Beyond Holistic Scoring through Validity Inquiry,” for instance, Peggy O’Neill
(2003) focused on validity, a key issue in writing assessment; her article is in-
cluded here. Turning to context, George Hillocks (2003) addressed the impact
of state assessments in his “How State Assessments Lead to Vacuous Thinking
and Writing.” Sandra Murphy did likewise, in her case looking not at the im-
pact of writing assessment on students, but rather at its impact on teachers in
one state; such teachers support students’ writing development as they practice
assessment within their classrooms.
That first issue of JWA concluded with an annotated bibliography; compiled
by Peggy O’Neill, Michael Neal, Ellen Schendel, and Brian Huot (2003), it too
spoke to JWA’s vision. Three bibliographic entries in particular articulate JWA’s
goal and its importance while forecasting the kinds of research, theory, and prac-
tice published in JWA during the last 14 years, as sampled in this edited collection.
The first bibliographic entry, Nicholas Lemann’s 1999 book The Big Test: The Secret
History of the American Meritocracy, details a social and cultural history of the SAT.
Although the stated purpose of the SAT was to change the college admissions
process by eliciting relevant information from college applicants so as to predict
their success in college, it also clearly intended to shift college admissions from one
based in legacies to one based in merit. The Lemann account also clarifies how the
SAT both succeeded and failed in that intention, demonstrating that assessment,
even when informed by the science of tests and measurements, is always contextu-
alized, always enacting a policy, whether visible or not.
A second item in the bibliography, O. Palmer’s College Board Report 42, “Six-
ty Years of English Testing,” (1960) argues that the science informing the College
Entrance Examination Board (CEEB) English testing contributes to such testing
“as a scientifically defensible practice” (O’Neill at al., 2003). Again, here too sci-
ence plays a role, not so much to forward a kind of democracy, however, but rather
to defend the practices of a growing assessment industrial complex. In the CEEB
model Palmer defends, both teachers and direct writing assessment are positioned
as opponents of CEEB, as “resistant to the scientific progress achieved in English
testing” (O’Neill et al., 2003). What teachers, rooted in the everyday of the human
classroom, may have understood better than measurement experts is how writing
assessment, regardless of the science, cannot be cleaved from the contexts and
complications accompanying it. As important, seeing students day in and day out,
teachers also understood how very contingent any decision based on assessment is.
In his 1994 “A Technological and Historical Consideration of Equity Issues
Associated with Proposals to Change the Nation’s Testing Policy,” George Ma-
daus seems to agree with teachers. Approaching what we might call the assess-
ment problem philosophically in this third bibliographic entry, with a view in-
formed by both phenomenology and practicality, Madaus observes that certain
principles define assessment. All evaluations, he notes, rely on samples of be-
havior; all evaluations make inferences “about a person’s probable performance
relative to the domain” (Madaus, 1994); and all assessments render decisions by
individual or institution. Moreover, the technologies don’t operate apart from
the culture of their origin. Instead, as
products of a culture, they often extend, shape, and reproduce
the same culture. The values that underlie testing are utili-
tarianism, economic competition, technological optimism,
objectivity, bureaucratic control and accountability, numer-
ical precision, efficiency, standardization, and conformity.
(O’Neill et al., 2003)
It’s worth noting that while such values, including standardization, confor-
mity, and economic competition, may locate the US, its testing industry, and its
schools, they are much less likely to be the values motivating teachers.
viii
Foreword
The articles in the two volumes of this edited collection carry these issues
of assessment and context forward, especially as they have been raised and con-
sidered over time. In Volume 1, the collection’s first section, Technical Issues in
the Assessment of Writing—Reliability and Validity, speaks to issues articulated by
both Peggy O’Neill and George Madaus, issues inherent in assessment that, as
both O’Neill and Madaus demonstrate, are not apart from larger human issues,
but are rather a part of them. The second section, Politics and Public Policy of
Large-Scale Writing Assessment, calls to mind the article by George Hillocks and
the history of college admissions provided by The Big Test. The third section,
Implications of Automated Scoring of Writing, questions how the evolution of au-
tomated essay scoring extends the dangerous logic of a “true” score as valid and
reliable across contexts. In Volume 2, the fourth section, Theoretical Evolutions—
Towards Fairness and Aspiring to Justice, again calls to mind the equity issues and
analysis developed by Madaus. And the fifth section, Students’ and Teachers’ Lived
Experiences, evokes the line of inquiry pursued by Sandra Murphy. As astute
readers have already noted, it’s also fair to observe that in this set of correspon-
dences between the introductory issue of JWA and the current collection’s chap-
ters, one section in the collection, Implications of Automated Scoring of Writing, is
left out: our first issue of JWA did not provide for the important questions about
writing assessment raised by digital technologies. Still, apprehending that they
were on the horizon, we made a start in the very next issue, courtesy of Michael
Williamson’s (2003) “Validity of Automated Scoring: Prologue for a Continuing
Discussion of Machine Scoring Student Writing.”
All of which is not to say that we anticipated all of the rich writing assessment
scholarship of the next decade and a half: our correspondences, of course, are
not predictive. But it is to say that the chapters here extend and elaborate what
we had hoped for in creating JWA, in the process refiguring continuing issues,
sounding new notes, and pointing us to new futures. For example, one chapter
argues that the divide between the educational measurement and the writing
assessment communities might be bridged with a “unified field of writing assess-
ment.” The construct of writing, another chapter explains, can no longer ignore
“the role of commonly available tools such as word processing software.” And yet
another chapter brings together science and the law in a shared inquiry into the
results and subsequent effects of writing assessment, employing a disparate im-
pact analysis framework contributing to a better, more human, more humane,
and more equitable assessment. Threaded throughout are the technical issues
and principles of writing assessment, the writing assessments themselves, and
the contexts in which they are embedded.
Remembering the recent history of writing assessment, focused on assess-
ments and their contexts as we prepare for a better writing assessment future, we
ix
Yancey and Huot
are very pleased to be learning from and with the authors included here. We feel
confident that you will be as well.
REFERENCES
Hillocks, G. (2003). How state assessments lead to vacuous thinking and
writing. Journal of Writing Assessment, 1(1), 5-21. [Link]
item/33k2v0w5
Lemann, N. (1999). The big test: The secret history of the American meritocracy. Farrar,
Straus and Giroux.
Madaus, G. F. (1994). A technological and historical consideration of equity issues
associated with proposals to change the nation’s testing policy. Harvard Educational
Review, 64(1), 76.
Murphy, S. (2003). That was then, this is now: The impact of changing assessment
policies on teachers and the teaching of writing in California. Journal of Writing
Assessment, 1(1), 23-45. [Link]
O’Neill, P. (2003). Moving beyond holistic scoring through validity inquiry. Journal of
Writing Assessment, 1(1), 47-65. [Link]
O’Neill, P., Neal, M., Schendel, E., & Huot, B. (2003). An annotated bibliography of
writing assessment. Journal of Writing Assessment, 1(2), 73-78. [Link]
org/uc/item/78q9v569
Palmer, O. (1960). Sixty years of English testing. College Board, 42, 8-14.
Williamson, M. M. (2003). Validity of automated scoring: prologue for a continuing
discussion of machine scoring student writing. Journal of Writing Assessment, 1(2),
85-104. [Link]
x
CONSIDERING STUDENTS,
TEACHERS, AND WRITING
ASSESSMENT: VOLUME 1,
TECHNICAL AND POLITICAL
CONTEXTS
INTRODUCTION TO
VOLUME 1, TECHNICAL AND
POLITICAL CONTEXTS
Diane Kelly-Riley
University of Idaho
Ti Macklin
Boise State University
Carl Whithaus
University of California, Davis
DOI: [Link] 3
Kelly-Riley, Macklin, and Whithaus
shifting political context of writing assessment and the rise of automated scor-
ing of writing. Volume 2 explores the evolutions in theory and practice related
to fairness and writing assessment and then the ways in which the people who
teach and learn in these spaces shape writing assessment practice.
TECHNICAL EVOLUTIONS
The first volume focuses on technical and political issues. The rise of local con-
siderations in writing assessment emerges most fully in the articles published
in JWA during the first two decades of the twenty-first century. Rather than
excluding the lived experiences of students and teachers, JWA has taken the lead
in documenting how contextual, situated, and localized forms of writing assess-
ment may provide fuller—more valid, reliable, and fairer—pictures of students’
writing. The history of this move valuing localized forms of writing assessment
has not been fully told. This movement reaches back to Edward White’s (1978)
early advocacy for direct assessment of students’ writing rather than a reliance on
indirect forms of writing assessment. It also echoes—perhaps even amplifies—
Kathleen Yancey’s (1999) and others’ work (Calfee & Perfumo, 1996; Elbow &
Belanoff, 1997; Hawisher & Selfe, 1997; Herman et al., 1993; Herter, 1991)
on writing portfolios in the 1990s reflecting on students’ emerging knowledge
about writing, their writing processes, and their development as writers. Since
its inception, JWA has published scholarship from the unique angle of how local
contexts inform writing assessments.
Revisions to the Standards for Educational and Psychological Testing (AERA,
APA & NCME, 2014) shifted discussions around the core educational measure-
ment constructs of validity and reliability and drove changes in these scholarly
areas. The Standards govern much of the thinking about standardized assess-
ments, particularly true within the psychometric and educational measurement
sides of writing assessment. The Standards is a living document open to revision,
and the changes between the 4th and 5th editions in 1999 shifted discussions
within writing assessment away from a singular focus on the importance of re-
liability to an understanding that validity is the most important consideration
in writing assessment systems and is situated in particular contexts. The pub-
lished discussions between Richard Haswell (1998) and Pamela Moss (1998)
foreground how debates around the concept of validity assumed an increasingly
important role in writing assessment. Once the focus of validity changed, teach-
ers had a clearer role in determining and contributing to meaningful assessment.
The increased emphasis on validity enabled teachers to push back against the
limitations of standardized tests, opening up a new area of research that involved
local contexts and faculty expertise. JWA’s establishment in 2003 provided a
4
Introduction
venue for writing teachers and educational researchers to explore the implica-
tions of considering local contexts on writing assessments.
PROGRAMMATIC IMPLICATIONS
A frequently told origin story of writing assessment in North American post-
secondary education points toward 1874 and the addition of an extemporane-
ous writing sample in the Harvard entrance examination. Norbert Elliot details
(2005) how the Harvard exam was used to place students into its curricula. More
than half of the students required remedial coursework and additional support
setting up the tension between assessment and instruction. Elizabeth A. Wright,
Suzanne Bordenlon, and S. Michael Halloran (2020), however, offer a corrective
to this historicizing of writing assessment. In “‘Available Means’ of Rhetorical
Instruction,” they take up Royster and Kirsch’s call to explore “the lessons taught
to those students unable to attend those schools for elite white men” (p. 245).
Wright, Bordenlon, and Halloran point out how late 19th-century rhetorical
education and writing instruction took place in a wide variety of secondary and
postsecondary educational contexts including Catholic institutions, women’s
colleges, historically Black universities and colleges, as well as within the often
repressive contexts of boarding schools for indigenous children (pp. 254-257).
Thus, there is a broader history of the structures and lasting impacts of writing
assessment yet to be explored.
Across all of these contexts, writing placement mechanisms grew more
profoundly as standardized tests became more widely available. Many of these
placement exams attempted to capture students’ readiness to enter postsecond-
ary study, but the means of the exams often did not correspond to the cur-
ricular realities in the classrooms. Haswell (2004) notes that the 1900’s “saw
testing firms grow ever more influential and departments of English grow ever
more divided between using ready made goods, running their own placement
examinations, or foregoing placement altogether.” (para. 3) Faculty in English
departments devised their own assessment systems. The English Equivalency
Exam (EEE) was used by the California State University and Colleges between
1973 and 1981; later, it was replaced by the English Placement Test in 1977 de-
veloped by Edward White and his colleagues. Haswell and Elliot (2017) observe
“the few scholars and test administrators who were using holistic scoring were
using all their energies to confront the problems of cost and scoring reliability,
as practical aspects of the large testing programs they were supervising.” (White,
1993, p. 82) The EEE went beyond that, as White (1984) himself declared in
his essay, “Holisticism.” The method of holistic scoring may have achieved some
pragmatic ends making “the direct testing of writing practical and relatively
5
Kelly-Riley, Macklin, and Whithaus
reliable” (White, 1984, p. 408), and it may have achieved some indirect social
ends, bringing “together English teachers to talk about the goals of writing in-
struction” ( p. 408), but beyond that “it embodies a concept of writing that is
responsible in the widest sense . . .” (p. 408). It was responsible for its product,
which was responsible for its advertised use.
This move toward localization continued in the late 1980s when an area of
research emerged from the lived experiences of teachers and students in compo-
sition courses in response to accreditation and accountability mandates. Moore,
O’Neill, and Crow (2016) detail this extensive history of compositionists “using
assessment to improve student learning before it was emphasized so much by
accreditors . . . [because they] understood the link between learning assessment
and teaching improvement before accreditors made the connection explicit”
(p. 20). Many of these teacher-researchers struggled with the day-to-day im-
plications of the theoretical constructs of validity and reliability. As they grap-
pled with these constructs in their contexts, new practices and research paths
emerged. Early examples are detailed by Moore et al. (2016) demonstrating the
field of composition’s historical response to external assessment mandates. The
first was Elbow and Belanoff’s (1997) portfolio system which replaced a mandat-
ed university proficiency exam. Another system, developed and implemented at
Washington State University, included an entry-level Writing Placement Exam
and junior Writing Portfolio developed by Richard Haswell and his colleagues
(2001). This program entwined formative writing assessment with disciplinarily
situated writing instruction across the entire undergraduate curriculum. At all
levels, writing teachers were involved in the assessments, and a comprehensive
writing center provided support for students, including required small group
sessions concurrently supporting students in upper-division disciplinary writing
courses for those who did not pass the mid-career assessment (Haswell, 2001).
The core educational measurement constructs of validity and reliability
continued to undergo major reconceptualization. In 1999, major revisions to
the Standards for Educational and Psychological Testing were jointly authored by
the American Psychological Association (APA), American Education Research
Association (AERA), and the National Council of Measurement in Education
(NCME)—professional organizations which guide and govern best practices
in assessment and measurement. In this revision, validity was cast as the most
important consideration above all. Now, tests or assessments were no longer
considered stand-alone entities that needed to adhere to standards of technical
qualities of reliability or validity. Instead, a major philosophical understanding
of assessment shifted to see these measurements in social contexts in which the
uses and interpretations of scores must be considered in each and every setting.
This was a revolutionary shift.
6
Introduction
The late 1990’s also saw significant educational reform in the US with assess-
ment playing a key role in these public and political arenas. During this time,
writing studies teachers pushed back against the standardized test movement
which attempted to represent and measure writing ability through knowledge of
grammar and other writing rules (Bloom et al., 1996). In standardized testing,
multiple choice test items were used as a way to measure the quality of students’
writing. Writing teachers and researchers resisted these indirect, decontextual-
ized forms of evaluating students’ writing abilities. From their positions in the
classroom, compositionists knew this evaluation did not serve the instructional
needs of either students or faculty, and they advocated for locally-developed
assessment measures attentive to classroom contexts and actual student learning
outcomes. Thus, portfolio assessment developed out of the work of postsecond-
ary writing teachers. This process is described in White et al.’s 1996 collection,
Assessment of Writing: Politics, Policies, Practices. By the mid 90s, ways of mea-
suring the construct of writing became more nuanced. Compositionists realized
that writing is socially situated and began to publish research findings support-
ing this position. Understanding the people who designed and participated in
the assessments and the multiple ways in which they were enacted across differ-
ent institutional sites became a key component of writing assessment.
The increased accountability context within educational settings in North
America resulted in innovative programmatic responses. The Council of Writing
Program Administrators (CWPA) started to discuss and collaborate on whether
“a pithy and effective list of objectives for writing [and] programs existed” (as
cited in Harrington et al., 2003, p. xv). These conversations among members
at all levels of expertise were enabled by many compositionists joining the then
newly created WPA-L email listserv in the late 1990s. The members of this
group recognized the multiple stakeholders who were invested in the outcomes
of first-year composition. This exigence resulted in the development of the WPA
Outcomes Statement for First-Year Composition (WPA OS) “a statement . . . plain
enough to speak to those outside the discipline, yet rooted in disciplinary lan-
guage enough to have status in the field” (Harrington et al., 2003, p. xvi). The
WPA OS is a consensus document detailing the expectations for first-year writ-
ing common to most postsecondary institutions in North America (Harrington
et al., 2001). Kathleen Yancey (2003) says that the WPA OS was intentional-
ly written as outcomes and not standards for performance that needed to be
achieved.
By framing and modeling curricular and assessment work as driven by facul-
ty and local contexts, the collaborators of the WPA OS also began to formalize
a new area of research. This new area of local programmatic response to as-
sessment had several offshoots as contextually situated responses to assessment
7
Kelly-Riley, Macklin, and Whithaus
8
Introduction
Gates and Lumina foundations, had strong legislative support to reshape Amer-
ican education into one focused on the preparation of workers to advance the
American economy. To measure their progress, these efforts partnered with test-
ing companies like Pearson and ETS. With complex assessments being imple-
mented on such a large scale, the possibilities of machine or computer scoring of
writing ascended to the forefront. The challenges for assessing learning through
writing remained and were amplified in these large-scale assessments.
These emerging curricular efforts recognized the importance of teaching
writing as situated within disciplinary genres from the beginning of school. As
accountability efforts moved from NCLB to the Smarter Balanced Assessment
Consortium (SBAC) and Partnership for Assessment of Readiness for College
and Careers (PARCC) national assessments, the initial response was to use tech-
nology, particularly the potentials for automated essay scoring, to support the
integration of writing across the K-12 curriculum. These large-scale assessments
have meant that writing assessment has taken a much more central role in ac-
countability efforts. The challenge remains to develop computer-based scoring
that represents the complexity of writing taught and assessed in the classroom.
9
Kelly-Riley, Macklin, and Whithaus
(2002, 2004, 2006) work, O’Neill suggests that by moving discussions of reli-
ability in writing assessment beyond inter-rater reliability, more nuanced, and
more accurate, forms of writing assessment can be developed. These emerging
forms of writing assessment might not only acknowledge but also account for—
in a psychometrically rigorous way—variations across readers and variations
across tests in the ways that Pamela Moss (1998) and Richard Haswell (1998)
recognize as hermeneutic or rhetorical practices.
Diane Kelly-Riley’s “Validity Inquiry of Race and Shared Evaluation Practic-
es in a Large-Scale, University-wide Writing Portfolio Assessment” (2011) ad-
vances the field’s understanding not only of the balance between reliability and
validity but also brings into the conversation vital contextual elements involving
race and racism. Her article takes on the question of race—and in more subtle
ways racism—by looking at the implementation of a locally-developed, con-
text-rich writing portfolio assessment system. Kelly-Riley’s article is a precursor
to the consideration of fairness and antiracist practices in writing assessment by
providing an empirical study that looked at how raters understand and apply
race in an assessment context. Writing assessment has struggled to develop an
operational definition of race. Race is often defined by government agencies that
collect data on race, but the experience in the writing classroom calls for more
nuanced representations of race.
O’Neill and Kelly-Riley’s work lead toward approaches outlined in Elliot
et al.’s “Three Interpretative Frameworks: Assessment of English Language
Arts-Writing in the Common Core State Standards Initiative” (2015). Elliot,
Rupp, and Williamson examine how standards-based definitions of validity,
reliability/precision, and fairness were integrated into the Smarter Balanced
Assessment Consortium (SBAC) and Partnership for Assessment of Readiness
for College and Career (PARCC) English Language Arts – writing assessments.
They encourage stakeholders to be informed consumers when interpreting and
using SBAC or PARCC scores about students’ writing. Their work foreshadows
a move within writing assessment research and practice encouraging stakehold-
ers (WPAs, students, teachers, and parents) to not just accept the scores from
large-scale state or national-level writing assessments at face value but to inte-
grate how they will be used, to examine their meaning and their use value.
10
Introduction
and Past President of the Two-Year College Association, synthesizes these major
educational reform movements and how they impact writing assessment schol-
arship and practices. This section highlights work by Edward M. White, Arthur
N. Applebee, Hammond and Garcia, and Toth et al. All of these authors antic-
ipate and wrestle with large scale writing assessment in terms of political and
policy issues. Political changes across educational reform movements have both
shaped and responded to assessment issues. The critiques of both placement in
two-year colleges and of AES have centered around how students’ writing must
be considered and evaluated as contextual, rather than stripped of context for
a placement decision afforded by the cost savings of having software evaluate a
piece of writing.
Edward M. White’s “The Misuse of Writing Assessment for Political Purpos-
es” (2005) and Arthur N. Applebee’s “Issues in Large-Scale Writing Assessment:
Perspectives from the National Assessment of Educational Progress” (2007) set the
stage for early political discussions around writing assessment. White argues that
many large-scale writing assessments are motivated “by political rather than educa-
tional, administrative, [or] professional concerns.” For White, No Child Left Be-
hind and its reliance on testing “without the resources and leadership for students
to achieve the skills they will be tested on” is a crucially flawed educational policy
and a misuse of writing assessments based on politicians’ misunderstanding of
what educational testing can tell us. He considers a wide range of mandated, large-
scale writing assessments ranging from required state-level testing of secondary
students through placement exams for incoming college students to graduation
requirements for college students. He suggests that the misuses of writing assess-
ments “[are derived] from an exaggerated, even a credulous misunderstanding, of
what particular kinds of assessments can accomplish.” Such observations continue
to underscore the misuse of assessments in educational settings.
In contrast to White’s critique of assessment as gatekeeping, Applebee’s “Is-
sues in Large-Scale Writing Assessment: Perspectives from the National Assess-
ment of Educational Progress” focuses on the contributions of the large-scale,
national-level programmatic assessment conducted through the National As-
sessment of Educational Progress (NAEP). Applebee suggests the NAEP writ-
ing assessment is valuable because it is not tied to the assessment of individual
students, but rather a way of looking at how students’ writing is developing and
comparing achievements in writing across states. White’s and Applebee’s works
are both polemic, but research like J. W. Hammond and Merideth Garcia’s show
the legacy of informed and principled approaches documenting the effects of
large-scale assessment on teachers.
Writing assessment, politics, and public policies in the first two decades of
the twenty-first century requires that we address the effects of No Child Left
11
Kelly-Riley, Macklin, and Whithaus
Behind (NCLB) and the Common Core State Standards (CCSS). Hammond
and Garcia’s “The Micropolitics of Pathways: Teacher Education, Writing As-
sessment, and the Common Core” (2017) and Toth et al.’s “ Writing Assess-
ment, Placement, and the Two-Year College” (2019) describe the impacts of
these initiatives on teachers and students. Hammond and Garcia take the Com-
mon Core State Standards (CCSS) as their point of departure. They examine
how teacher education programs frame the CCSS for their teachers-in-training.
Their works suggest postsecondary faculty, teachers, and teachers-in-training
“micropolitically interpret” the Common Core. In fact, Hammond and Garcia
suggest that writing teachers and writing teachers-in-training foreground their
own local writing assessments since teachers seem most focused on curriculum
and instruction issues and secondarily on CCSS and pathway-related reforms to
education. One of the key findings from Hammond and Garcia’s work is the val-
ue of adopting a micropolitical perspective when considering writing curricula,
instruction, and assessment.
The emphasis on learning pathways was championed in the educational re-
forms promoted through CCSS. As such, pathway-based reforms had a dra-
matic effect on community colleges, Toth et al.’s introduction to the Journal of
Writing Assessment’s Special Issue on Placement and Two-year Colleges takes up
the overlapping issues of educational reform and how writing assessments have
been used in placement decisions. In their ambitious and wide-ranging article,
Toth, Nastal, Hassel, and Giordano review the history of two-year colleges with-
in American higher education. They attend to the ways in which this history
and pathway-based educational reform movement intersects with models for
assessing and placing students within ESL, basic writing, or first-year composi-
tion courses. They then extend their discussion by turning to questions around
the validity and the uses for writing assessment and placement systems. Toth et
al.’s attention to sociocultural factors highlights the ways in which questions of
writing assessment are being looked at at a systems level rather than only at the
level of individual students.
12
Introduction
13
Kelly-Riley, Macklin, and Whithaus
REFERENCES
American Educational Research Association, American Psychological Association, &
National Council of Measurement in Education. (2014). Standards for educational
and psychological testing. American Educational Research Association.
Behm, N. N., Glau, G. R., Holdstein, D. H., Roen, D., & White, E. M. (Eds.).
(2013). The WPA Outcomes Statement: A decade later. Parlor Press.
Calfee, R. C., & Perfumo, P. (Eds.). (1997). Writing portfolios in the classroom: Policy
and practice, promise and peril. Routledge.
Elbow, P. & Belanoff, P. (1997). Reflections on an explosion: Portfolios in the ‘90s and
beyond. In K. B. Yancey & I. Weiser (Eds.), Situating portfolios: Four perspectives
(pp. 21-33). Utah State University Press.
Elliot, N. (2005). On a Scale: A social history of writing assessment in America. Peter
Lang.
Gallagher, C. (2007). Reclaiming assessment: A better alternative to the accountability
agenda. Heinemann.
14
Introduction
Harrington, S., Rhodes, K., Fischer, R. O., & Malenczyk, R. (2003). Introduction:
Celebrating and complicating the Outcomes Statement. In S. Harrington, K.
Rhodes, R. O. Fischer, & R. Malenczyk (Eds.), The outcomes book: Debate and
consensus after the WPA Outcomes Statement (pp. xv-xix). Utah State University Press.
[Link]
Harrington, S., Malenczyk, R., Peckham, I., Rhodes, K., & Yancey, K. B. (2001). WPA
outcomes statement for first-year composition. College English, 63(3), 321-325.
Haswell, R. H. (Ed.). (2001). Beyond outcomes: Assessment and instruction within a
university writing program. Ablex.
Haswell, R. H. (1998). Multiple inquiry in the validation of writing tests. Assessing
Writing, 5(1), 89-110.
Haswell, R. H. (2004). Post-secondary entry writing placement: A brief
synopsis, [Link]. [Link]
[Link]
Haswell, R. H. & Elliot, N. (2017). Innovation and the California State University
and Colleges English Equivalency Examination, 1973-1981: An Organizational
Perspective, Journal of Writing Assessment, 10(1), [Link]
item/7rt5v9p2
Hawisher, G. E., & Selfe, C. L. (1997). Wedding the technologies of writing portfolios
and computers. In K. B. Yancey & I. Weiser (Eds.), Situating portfolios: Four
perspectives (pp. 305-321). Utah State University Press.
Herman, J. L., Gearhart, M., & Baker, E. L. (1993). Assessing writing portfolios:
Issues in the validity and meaning of scores. Educational Assessment, 1(3), 201-224.
Herter, R. J. (1991). Writing portfolios: Alternatives to testing. English Journal, 80(1), 90.
Indiana University. (2014). IU to collaborate with high school teachers on Writing and
Reading Alignment Project. [Link]
[Link]
Lakoff, G. (2002). Moral politics: How liberals and conservatives think (2nd ed.).
University of Chicago Press.
Lakoff, G. (2004). Don’t think of an elephant! Know your values and frame the debate.
Chelsea Publishing.
Lakoff, G. (2006). Simple framing. Rockridge Institute. [Link]
org/projects/strategic/simple_framing
Moore, C., O’Neill, P., & Crow, A. (2016). Assessing for learning in an age of
comparability, In W. Sharer, T. A. Morse, M. F. Eble, & W. B. Banks (Eds.),
Reclaiming accountability: Improving writing programs through accreditation and
large-scale assessments (pp. 17-35). Utah State University Press.
Moss, P. (1998). Testing the test of the test: A response to “Multiple Inquiry in the
Validation of Writing Tests” [by Richard H. Haswell]. Assessing Writing 5(1), 111-
122.
Sharer, W., Morse, T. A., Eble, M. F., & Banks, W. P. (2016). Introduction:
Accreditation and assessment as opportunity. In W. Sharer, T. A., Morse, M. F. Eble,
& W. B. Banks (Eds.), Reclaiming accountability: Improving writing programs through
accreditation and large-scale assessments (pp. 3-13). Utah State University Press.
15
Kelly-Riley, Macklin, and Whithaus
White, E. M., Lutz, W. D., & Kamusikiri, S. (1996). Assessment of writing: Politics,
policies, practices. Modern Language Association.
White, E. M. (1978). Mass testing of individual writing: The California model. Journal
of Basic Writing, 1(4), 18-38. [Link]
Wright, E. A., Bordelon, S., & Halloran, M. (2020). “Available means of rhetorical
instruction”: “Broadening perspectives” on rhetorical education prior to 1900. In
J. J. Murphy & C. Thaiss, (Eds.), A short history of writing instruction: From ancient
Greece to the modern United States (pp. 244-271). Routledge.
Yancey, K. B. (1999). Looking back as we look forward: Historicizing writing
assessment. College Composition and Communication, 50(3), 483-503.
Yancey, K. B. (2003). Standards, outcomes, and all that jazz. In S. Harrington, K.
Rhodes, R. O. Fischer, & R. Malenczyk (Eds.), Debate and consensus after the WPA
Outcomes Statement (pp. 18-23). Utah State University Press.
16
PART 1.
FROM ISOLATION TO
INTEGRATION: TECHNICAL
ISSUES IN THE ASSESSMENT
OF WRITING
David H. Slomp
University of Lethbridge
DOI: [Link] 19
Slomp
interplay between the scholarship on writing assessment from within and across
the measurement and writing studies communities. In this commentary, I will
focus on the ways in which JWA’s legacy bridges the gap between educational
measurement and writing studies in three selected articles, and I will also explore
the implications for research and practice that emerge from dialogues between
these two fields. I begin, though, with an exploration of several tropes that have
shaped our thinking about the interplay between educational measurement and
writing studies communities.
20
Retrospective. From Isolation to Integration
21
Slomp
22
Retrospective. From Isolation to Integration
we need to attend to the values that underpin the field of writing studies. Her
argument echoes Onwuegbuzie and Leech (2005) who call for training future
researchers within a pragmatist tradition so that they are capable of navigating
both positivist and interpretivist models of research, drawing on and adapting
methods from within both traditions as the research warrants. O’Neill sums up
her position:
In determining reliability, many of us responsible for writing
assessments should collaborate as equal partners with col-
leagues who have the statistical expertise. Writing assessment
practitioners and scholars need to accept our responsibility
to develop and maintain writing assessments that are in-
formed by both language-based and psychometric theory and
research. We need to develop new methods for assessment
as well as for determining reliability and validity if current
methods do not work adequately for our purposes, as Parkes
(2007) argued. This may mean collaborating with others who
have different kinds of experiences and expertise, learning
more about psychometric theory and practices, and engaging
in difficult discussions with colleagues about what we value
and why it matters. (pp. 59-60)
She further argues that by focusing on our values, by continually bringing
these into the conversations about assessment design, appraisal, and use, we
can help to reframe reliability so that our pursuit of the values of accuracy, de-
pendability, stability, consistency, and precision in writing assessment can be
engineered to serve our students and our programs.
There is certainly evidence within the field of writing assessment to support
her claims. In North America, for writing assessment at the post-secondary level,
the response to this call has been evidenced in the uptake of communal writ-
ing assessment (Broad et al., 2009; Lindhardsen, 2020), community grading
(Shumake & Shah, 2017), contract grading (Litterio, 2016), and comparative
judgment (Sims et al., 2020) models of scoring: processes that rely on rigor-
ous discussion and documentation to demonstrate commitment to accuracy,
dependability, stability, consistency, and precision.
The broader value of O’Neill’s article is that it continues a tradition of ar-
guing for the role that composition studies can and should play in shaping the
discourses and practices surrounding writing assessment. Writing in Education
Measurement: Issues and Practice, Newton (2017), draws on the field of writ-
ing assessment—and indirectly on the scholarship in the field’s two major jour-
nals—to make the point that the measurement community needs to engage
23
Slomp
24
Retrospective. From Isolation to Integration
25
Slomp
models make explicit the link between validity, reliability, and the consequences
of assessment design and use.
Kelly-Riley’s (2011) study “Validity inquiry of race and shared evaluation
practices in a large-scale, university-wide writing portfolio assessment” demon-
strates how within an argument-based framework concerns for validity, reliabil-
ity, and ultimately fairness can be motivated by the revealed consequences of
an assessment’s use. She examines a portfolio-based assessment program that
had been in use for 20 years at an American university. The purpose of the
assessment was to identify students who needed additional support in their up-
per-divisional writing requirements. When an African American student ques-
tioned the pass rates of BIPOC students compared to those of white students,
unexamined questions related to fairness, reliability, and validity were brought
into focus. In response to the student’s question, Kelly-Riley’s study examined
how the assessment program in question might be unwittingly disadvantaging
students of color.
Drawing on Kane’s (2006) model of validation, Kelly-Riley links concerns
for consequences with questions of construct representation and issues of score
stability across populations of test-takers. While rightly cast as a validity study,
this paper examines the scoring inference—a reliability issue.
Her investigation revealed that this assessment program was in fact desig-
nating students of color as “needs work” more frequently than it did white stu-
dents who were more likely to receive a “pass” score. Analyzing the influence on
student scores of race, perceived demographic profiles of students, and scoring
criteria, she found, “race did not contribute to faculty raters’ functional defi-
nition of ‘good writing’ for any of the frameworks whether in the timed exam
format or for course papers” (p. 80). Instead, she found that “coherence, focus,
and correctness all contribute significantly to the functional definition of “good
writing” (p. 83). Additionally, she found that “large percentages of the variance
of writing quality are accounted for through the Demographic framework—
primarily through the rater’s perception of the writer’s intelligence and comfort
with writing” (p. 84). At the same time, however, there remain statistically sig-
nificant differences in performance by race on this assessment.
Kelly-Riley’s (2011) study demonstrates the value of localism in writing as-
sessment. As an administrator and professor she is well positioned to see first-
hand the impact of the assessment program she is investigating on the students
who walk through her door. In fact, it is the very questions and concerns raised
by students who were made different by the assessment, that prompted the focus
of her research. The power of Kelly-Riley’s study is that it leverages contempo-
rary validation frameworks to address these concerns and to advance local val-
ues of equity and opportunity. She positions her work in response to O’Neill’s
26
Retrospective. From Isolation to Integration
27
Slomp
and its stability across racial contexts. While this disparity in the opportunity
to learn may help explain the disparity in performance by BIPOC students in
Kelly-Riley’s study, deeper scrutiny of the assessment itself is likely necessary. In
particular, the writing construct underpinning the assessment, the scoring crite-
ria, and the operationalization of that criteria, likely needs to be more critically
examined. Cushman (2016) more succinctly made this point in her critique of
validity theory:
Fairness can address content of particular questions, but
it does little to adjust the overall ways in which validity
measures themselves, from the start, are based on colonial
difference that they help to create and maintain. . . . In this
instance, constructs will always be unrelated to the knowledg-
es and language practices of the peoples made different by the
construct and validity measures in the first place.
Cushman’s observations highlight the value of seeking out pluriversal un-
derstandings; of seeking out multiple and varied experiences and perspectives in
trying to understand how an assessment is functioning.
28
Retrospective. From Isolation to Integration
29
Slomp
30
Retrospective. From Isolation to Integration
Placement in the Two-Year College (Kelly-Riley & Whithaus, 2019) applies the
frameworks developed in the 2016 SI to the design and use of placement tests
in the Two-year college. In 2018, Pruchnic et al. advanced mixed methods ap-
proaches to collecting validity and reliability evidence designed to address the
concerns of both measurement specialists and writing studies professionals.
This work, however, remains uneven. Sprinkled through these same issues
are articles that continue to approach validity, for example, using dated models
and approaches: that speak of validating instruments rather than inferences and
decisions. I noted the same unevenness in how this concept was being handled
in articles published in Assessing Writing over the past decade:
One trend of concern across several of the papers published in
the past 10 years is the characterization of validation studies
as attempts to “establish” the validity of the assessments in
question. This language suggests a confirmation bias that was
not noticeable in the earlier validation studies published in
ASW. It is important to remember that we do not validate as-
sessments. Rather, we examine categories of evidence and then
use that evidence to form an interpretation and use argument
that is always contingent. Too often this contingency is not
expressed. (Slomp, 2019, p. 14)
As we draw on contemporary theories of validity, reliability, and fairness, to
assess the design implementation and use of locally developed assessment pro-
grams, a critical reflexive mindset remains important.
Looking forward, it is also important to recognize that measurement is not a
unified and monolithic discipline. Many scholars within this discipline, too, are
struggling with its roots and with its history. Stephen Sireci (2021), in his pres-
idential address to the National Council on Measurement in Education, for ex-
ample, called out the discipline for losing the public’s confidence in their work.
He cites four reasons for this: psychometric hypocrisy, psychometric censorship,
psychometric paralysis, and the discipline’s support for an educational culture of
distrust. Other measurement scholars I’ve cited in this paper—Newton, Mislevy,
Randal, Rupp, Oliveri—are but a few examples of scholars who are working to
take measurement in a more humanistic direction. Their work demonstrates
how collaborations with measurement scholars who share concern for the im-
pact of measurement both on diverse populations of students and educators,
and on systems of education, can support the development of a new generation
of writing assessment programs that focus first on the needs of students and
educators (Oliveri et al., 2021) for an example of such collaboration). The three
articles highlighted in this section offer a prescription for supporting this work.
31
Slomp
32
Retrospective. From Isolation to Integration
Writing and the Journal of Writing Assessment, we see a clear ethic of critical schol-
arship, and openness to exploring possibility, of listening to critique, and of
responding to it. Harnessing that ethic to a spirit of purposeful pluralism will
serve our discipline well as it innovates for the future.
Together these principles position us to approach our work with a sense of
humility and purpose, reminding us that the work of writing assessment re-
search should be done in the service of others. As our fields continue to evolve
purposeful, principled pluralism will be a key tool we can leverage to ensure that
writing assessments programs serve all students, promote quality learning, and
structure opportunity. In this spirit, educators—writing studies specialists—
need to increasingly insist on having a seat at the table, and they need to come
to that table equipped to engage with the measurement theories set before them,
while not neglecting to add to the conversation the insights and concerns of our
discipline, and in particular our enduring concern for the social consequences
of our assessment programs on students, educators, and systems of education.
REFERENCES
American Educational Research Association, American Psychological Association, &
National Council of Measurement in Education. (2014). Standards for educational
and psychological testing. American Educational Research Association.
Adler-Kassner, L., & O’Neill, P. (2010). Reframing writing assessment to improve
teaching and learning. Utah State University Press.
Bachman, L. & Palmer, A. (2010). Language assessment in practice: Developing language
assessments and justifying their use in the real world. Oxford University Press.
Baker-Bell. A. (2010). Playing with the stakes: A consideration of an aspect of the social
context of a gatekeeping writing assessment. Assessing Writing, 15(3), 133-153.
Barkaoui, K. (2010). Variability in ESL essay rating processes: The role of the rating
scale and rater experience. Language Assessment Quarterly, 7(1), 54-74.
Barkaoui, K. (2011). Effects of marking method and rater experience on ESL essay
scores and rater performance. Assessment in Education: Principles, Policy & Practice,
18(3), 279-293.
Behizadeh, N., & Engelhard Jr., G. (2011). Historical view of the influences of
measurement and writing theories on the practice of writing assessment in the
United States. Assessing Writing, 16(3), 189-211.
Brennan, R. L. (2001). An essay on the history and future of reliability from the
perspective of replications. Journal of Educational Measurement, 38(4), 295-317.
Broad, B. (2003). What we really value: Rubrics in teaching and assessing writing. Utah
State University Press.
Broad, B. (2016). This is not only a test: Exploring structured ethical blindness in the
testing industry. Journal of Writing Assessment, 9(1). [Link]
item/2bt3m3nf
33
Slomp
Broad, B., Adler-Kassner, L., Alford, B., Detweiler, J., Estrem, H., Harrington, S.,
McBride, M., Stalions, E., & Weeden, S. (Eds.). (2009). Organic writing assessment:
Dynamic criteria mapping in action. Utah State University Press.
Cushman, E. (2016). Decolonizing validity. Journal of Writing Assessment, 9(2). https://
[Link]/uc/item/0xh7v6fb
Dryer, D. B. (2012). At a mirror, darkly: The imagined undergraduate writers of ten
novice composition instructors. College Composition and Communication, 63(3),
420-452.
Elliot, N. (2005). On a scale: A social history of writing assessment in America. Peter Lang.
Elliot, N., & Perelman, L. (Eds.). (2012). Writing assessment in the 21st century: Essays
in honor of Edward M. White. Hampton Press.
Elliot, N., Rudniy, A., Deess, P., Klobucar, A., Collins, R., & Sava, S. (2016).
ePortfolios: Foundational measurement issues. Journal of Writing Assessment, 9(2).
[Link]
Elliot, N., Rupp, A. A., & Williamson, D. M. (2015). Three interpretative frame-
works: Assessment of English language arts-writing in the Common Core State
Standards initiative. Journal of Writing Assessment, 8(1). [Link]
item/4zb222xg
Goodwin, S. (2016). A Many-Facet Rasch analysis comparing essay rater behavior on
an academic English reading/writing test used for two purposes. Assessing Writing,
30, 21-31. [Link]
Huot, B. (2002). Re-articulating writing assessment. Utah State University Press.
Huot, B. (2003). Introduction. Journal of Writing Assessment, 1(2), 81-84. https://
[Link]/uc/jwa/1/2
Huot, B., & Yancey, K. (2003). Introduction. Journal of Writing Assessment, 1(1), 1-4.
Jølle, L. (2014). Pair assessment of pupil writing: A dialogic approach for studying the
development of rater competence. Assessing Writing, 20, 37-52.
Kane, M. T. (2006). Validation. In R. L. Brennan (Ed.), Educational measurement (4th
ed.) (pp. 17–64). American Council on Education; Praeger.
Kane, M. T. (2013). Validating the interpretation and uses of test scores. Journal of
Educational Measurement, 50(1), 1–73.
Kelly-Riley, D. (2011). Validity inquiry of race and shared evaluation practices in
a large-scale, university-wide writing portfolio assessment. Journal of Writing
Assessment, 4(1). [Link]
Kelly Riley, D., & Whithaus, C. (2016). Introduction to a special issue on a theory
of ethics for writing assessment. Journal of Writing Assessment, 9(1). https://
[Link]/uc/item/8nq5w3t0
Kelly-Riley, D., & Whithaus, C. (2019). Editors’ introduction: Special issue on
two-year college writing placement. Journal of Writing Assessment, 12(1). https://
[Link]/uc/item/7vg91466
Klein, J., & Taub, D. (2005). The effect of variations in handwriting and print on
evaluation of student essays. Assessing Writing, 10(2), 134-148.
Knoch, U. (2007). “Little coherence, considerable strain for reader”: A comparison between
two rating scales for the assessment of coherence. Assessing Writing, 12(2), 108-128.
34
Retrospective. From Isolation to Integration
35
Slomp
Sims, M. E., Cox, T. L., Eckstein, G. T., Hartshorn, K. J., Wilcox, M. P., & Hart, J.
M. (2020). Rubric rating with MFRM versus randomly distributed comparative
judgment: A comparison of two approaches to second‐language writing assessment.
Educational Measurement: Issues and Practice, 39(4), 30-40.
Sireci, S. G. (2021). NCME presidential address 2020: Valuing educational
measurement. Educational Measurement: Issues and Practice, 40(1), 7-16.
Slomp, D., Oliveri, M. E., Elliot, N. (2021). Afterword: Meeting the challenges of
workplace English communication in the 21st century. Journal of Writing Analytics
5(1), 342-370. [Link]
Slomp, D. H., (2019). Complexity, consequence, and frames: A quarter century of
research in assessing writing. Assessing Writing, 42(4), 1-17
Wang, J., Engelhard Jr, G., Raczynski, K., Song, T., & Wolfe, E. W. (2017). Evaluating
rater accuracy and perception for integrated writing assessments using a mixed-
methods approach. Assessing Writing, 33, 36-47.
Wind, S. A., & Engelhard Jr, G. (2013). How invariant and accurate are domain
ratings in writing assessment?. Assessing Writing, 18(4), 278-299.
Winke, P., & Lim, H. (2015). ESL essay raters’ cognitive processes in applying the
Jacobs et al. rubric: An eye-movement study. Assessing Writing, 25, 38-54.
Wiseman, C. S. (2012). Rater effects: Ego engagement in rater decision-making.
Assessing Writing, 17(3), 150-173.
Wolfe, E. M. (2005). Uncovering rater’s cognitive processing and focus using think-
aloud protocols. Journal of Writing Assessment, 2(1), 37–56. [Link]
uc/item/83b618ww
Zheng, Y., & Yu, S. (2019). What has been assessed in writing and how? Empirical
evidence from assessing writing (2000–2018). Assessing Writing, 42. [Link]
org/10.1016/[Link].2019.100421
36
CHAPTER 1.
Peggy O’Neill
Loyola University Maryland
Writing an essay about reliability and writing assessment presents several chal-
lenges. One comes from determining what we mean by writing assessment be-
cause as a field it encompasses teachers and researchers in K-12 education as
well as higher education. Some of these professionals are trained in education-
al measurement, but many others are trained primarily as literacy educators.
The field also includes test developers employed by testing companies, some of
whom may provide testing services for institutions, and government employees,
typically in departments of education, who work on assessments such as NAEP
or others. Another challenge concerns the very concept of reliability, which is
deeply embedded in statistical theories and methods. Many educators who teach
DOI: [Link] 37
O’Neill
writing and work in college writing assessment have been educated primarily
in humanities departments and are immersed in the subject of literacy educa-
tion; they are not psychometricians and are not experts in statistical theories and
methods, which seem to dominate approaches to reliability. Because of these
challenges, college writing assessment practitioners often side-step reliability to
some extent. They report instead, for example, a co-efficient about rater agree-
ment or percentages of samples needed to be scored by three or more readers,
but do not delve into the complexity of the issues associated with issues such as
calculating coefficients. Yet, reliability is an important component of writing
assessment that needs to be considered not just in its own right but also as part
of the validation process because it addresses consistency and generalizability,
among other values.
As writing assessment practitioners and scholars, we need to grapple with the
challenges associated with reliability by examining how it has been used in writ-
ing assessment scholarship, especially within the college composition communi-
ty, and how we can reframe it so that it both engages with what writing teachers
value and contributes appropriately to validation efforts. With the interest in
large-scale assessments (including writing assessment) and higher education in-
creasing, college writing faculty will need to address several issues, many of which
are related to reliability as well as validity (Adler-Kassner & O’Neill, 2010). One
such issue is automated scoring, which is generally critiqued by college composi-
tion professionals (Herrington & Moran, 2001; Ericsson & Haswell, 2006) but
which has found more support in the psychometric community (Williamson
2003). According to Williamson (2003):
Two things are certain. One, automated scoring programs can
replicate scores for a particular reading of student writing,
and this technology is reliable, efficient, fast, and cheap. Two,
automated scoring has been and will continue to be used in
various large-scale assessments of student writing. (p. 256)
Carl Whithaus (2005) also acknowledged the role of automated scoring in
large-scale testing and encouraged writing instructors not only to accept auto-
mated evaluation systems but also to integrate them (as well as other technolo-
gies) into their teaching (p. 13). Williamson, who doesn’t go as far as Whithaus
in supporting the use of automated evaluation, argued for a “productive alli-
ance” between those in educational measurement and those invested in teaching
writing (p. 101). To develop this kind of relationship, college writing instruc-
tors and program administrators need to “examine the research methodology or
social sciences as it impinges on assessment” and to “explore the potential for
collaborative research, not just within a social science or humanistic tradition”
38
Reframing Reliability for Writing Assessment
39
O’Neill
40
Reframing Reliability for Writing Assessment
when the Common Core State Standards Initiative identifies writing as one of the
key areas for “college and career readiness,” readers will understand this section
through the frame they have already have about writing. Parents, teachers, policy-
makers and assessment experts may not all share the same framework and so they
may understand the standards differently. The authors of the standards try to mit-
igate this situation by providing preliminary material that defines terms, explains
situations, and even articulates what is not covered by the standards. However
helpful that information is, it can also act as a away to reinforce what the reader
already thinks and believes because many of the associations and assumptions
work at unconscious levels. If we consider writing assessment, then, the same
theory applies. What seems practical and rational to writing teachers may seem
completely unreasonable or just wrong to policymakers or psychometricians,
who may approach the activity through completely different frames—different
values, experiences, assumptions, and world views. Specific technical terms as-
sociated with assessment, such as validity and reliability, will also be understood
differently depending on the frame surrounding them.
If we want to change a concept or redefine a concept, we need to consider the
frame that surrounds it and how that influences the way a term is understood.
Trying to change the term without taking into account the bigger picture will
not be successful, in Lakoff’s view, because much of what is evoked happens
automatically and unconsciously. Making visible the dominant associations and
assumptions so we can see the frame that currently in place is the first step in
trying to reframe writing assessment in general and reliability in particular.
41
O’Neill
recruits. In truth, the CEEB had been piloting the SAT for scholarship students
who needed to apply earlier and had found that the reliability and efficiency of
the SATs to be much superior to that of the essay examination. The development
of holistic scoring procedures in the 1960s, done by Educational Testing Service
researchers Godschalk, Swineford, and Coffman (1966), revitalized essay testing
because it provided a reliable way to score essays.
By the 1980s, holistically scored essays enjoyed widespread use for a variety
of writing assessments across educational levels, but especially in college. Ed-
ward White (1993) claimed: ‘[W]hen a university or college opens discussion of
the measurement of writing ability these days, the point of departure is usually a
holistically scored essay test” (p. 89). The holistic scoring of essay exams depend-
ed upon standardization of procedures for the test administration, of the tasks
and topics, and of the scoring. The holistic scoring sessions became, according
to White (1993), not just a method for scoring but also a means of professional
development as readers discussed anchor papers and practiced scoring samples
to internalize the scoring rubric so they could apply it in a consistent way. These
scoring sessions also required careful record keeping and checks for agreement
between two independent raters.
While White focused on the benefits of holistic scoring both in terms of
professional development and achieving acceptable reliability rates, Cherry and
Meyer (1993) critiqued the way reliability has been handled in writing assess-
ment. They explained that reliability “refers to how consistently a test measures
whatever it measures” (p. 110). The consistency of a measurement, Cherry and
Meyer (1993) explained, can come from the test design and administration,
the students, or the scoring. For essay testing, particular sources of error may
include the prompt—which may not produce reliable results—as well as the ad-
ministration and scoring of the essays. After reviewing research in direct writing
assessment from Starch and Elliot’s 1912 article, “Reliability of the grading of
high school work in English,” through several pieces in the mid to late 1980s,
Cherry and Meyer (1993) concluded that there have been four serious problems
with reliability as reported in writing research and evaluation (p. 116).
First, according to Cherry and Meyer (1993), reliability discussions (with
a few notable exceptions) have been limited to interrater reliability although
there are many other aspects of reliability that need to be considered. For ex-
ample, if students’ performances are not accurate in terms of their writing abil-
ities because of the prompt design, then results are not reliable no matter how
consistently raters apply the rubric and how much they agree with each other
(Hoetker, 1982).
Second, there has been confusion over reliability and validity in influential
studies of writing assessment. Cherry and Meyer (1993) critiqued Godshalk,
42
Reframing Reliability for Writing Assessment
43
O’Neill
also critiqued the method of calculating and reporting reliability found in the
literature, especially on more recent studies. They argued that interrater reliabil-
ity rates should be determined by statistical correlations and not the percentage
of agreement between the two independent raters. Hayes and Hatch (1999) ex-
plained that reliability calculated using a statistical correlation formula takes into
account the role of chance in the agreement rate while the percentage method
doesn’t. Depending on the scoring scale and the distribution of scores, chance
can account for a significant portion of agreement. For example, the fewer score
points on the rating scale, the greater the influence chance has on the agreement
rate; or, the more scores tend to cluster around certain scores, the more influence
chance has on the reliability measure.
Like Cherry and Meyer (1993), Hayes and Hatch (1999) noted that dif-
ferent methods for calculating reliability lead to different results, yet they also
found many researchers did not report the method for calculating reliability
correlations. Both Cherry and Meyer (1993) and Hayes and Hatch (1999) also
agreed that when researchers do not fully disclose how they determined reliabil-
ity estimates, it is difficult for readers to determine if the method is appropriate,
to compare reliability across studies, and to avoid confusion. Hayes and Hatch
(1999) concluded their essay with an acknowledgment that other methods exist
for measuring reliability, including generalizability measures, than those they
address although they don’t discuss them.
Both Cherry and Meyer (1993) and Hayes and Hatch (1999) framed reli-
ability in writing assessment using classical test theory. Shale (1996), however,
advocated using generalizability theory instead, arguing that it is more appro-
priate for addressing the issues associated with reliability in writing assessment
because it can address the multiple sources of error that can arise in a writing
assessment. In most writing assessments, Shale (1996) contended, reliability is
vague because it is only considered within the classical test theory, which was
developed for multiple-choice testing: “Considerable ambiguity arises because
the full sense of reliability as understood within the context of multiple-choice
testing does not transfer well to the world of essay testing” (p. 77). Shale ex-
plained that the consequences of considering reliability only in terms of classical
test theory has resulted in a “fixation on marker disagreement” which has led to a
distortions and limitations in writing assessment practices (p. 78). Shale (1996)
as with Cherry and Meyer (1993) and Hayes and Hatch (1999), also noted the
paucity of rigorous inquiry into reliability in writing assessment scholarship.
Reliability, how we should approach it, and what we mean by the term is still an
issue in college writing assessment as I discuss in more detail later.
While concerns about reliability of essay exams preoccupied writing as-
sessment scholars for a long time and, in effect framed writing assessment, the
44
Reframing Reliability for Writing Assessment
validity of essay testing was not seriously challenged because essay testing re-
quired students to write instead of completing multiple-choice items about lan-
guage conventions and grammar. White (1993) articulated the assumptions that
supported holistic scoring of essay exams: “It is a direct measure of writing, mea-
suring the real thing, and hence is more valid than indirect measures” such as
fill in the bubbles multiple choice exams and editing tests (p. 90). By the 1990s,
however, writing assessment scholars (as well as measurement theorists) began to
turn their attention to validity arguing that a portfolio of writing was preferable
to a single-sample, timed impromptu essay (Elbow and Belanoff, 1986).
The shift to validity began to take the focus away from reliability as a purely
statistical concept and to frame it as part of a validity argument, which addresses
both theoretical and quantitative, statistical evidence (Messick, 1989). Camp
(1993) addressed the tension between classical test theory and emerging theories
of writing and literacy. Camp (1993) argued: “Very likely we are seeing the signs
of a growing incompatibility between our views of writing and the constraints
necessary to satisfy the requirements of traditional psychometrics—in particular,
of reliability and validity narrowly defined” (p. 52). Camp (1993) explored this
tension, identifying some of the key factors that may need to be addressed to
develop writing assessments that take into account what we know about writing
as well as the principles of fairness, equity, and generalizability—concepts, she
explained, that are associated with reliability. The challenge, according to Camp,
has been to apply these principles in ways that lead far beyond the narrow focus
on score reliability and constricted definitions of validity that have characterized
earlier discussions of writing assessment (p. 68). At the time of Camp’s essay,
portfolios (like other performance assessments) were growing in popularity and
Camp concluded with a brief discussion of some portfolio projects.
Since the early 1990s, the popularity of portfolios in college writing pro-
grams has continued to spread for both teaching and assessment (although essay
testing also remained popular). Although writing portfolios seemed to be a sub-
stantive departure from impromptu essay testing, the discussion of reliability,
however, did not changed very much. The focus was still narrowly on interrat-
er reliability. As White (1993) looked to the future of portfolios, he identified
reliability of portfolio scoring as the major issue to deal with, which in effect,
continued to frame writing assessment in terms of reliability. At that time, he
recommended adapting many of the same procedures for portfolios that were
used for holistic scoring of essays: “At a minimum, each portfolio should receive
two independent scores, and reliability data should be recorded. While reliabil-
ity should not become the obsession for portfolio evaluation that it became for
essay testing, portfolios cannot become a serious means of measurement without
demonstrable reliability” (p. 105).
45
O’Neill
46
Reframing Reliability for Writing Assessment
We need to explore the concept more fully, considering it in light of what we know
about writing and learning to write, as well as psychometric theory, because as
Camp (1993) said, the principles that inform reliability are important.
Moss (1994), in fact, didn’t reject reliability outright. Rather she encouraged
assessment researchers and practitioners to explore it as a theoretical construct
in light of validity. She explained that “less standardized forms of assessment . . .
[such as portfolios] present serious problems for reliability, in terms of general-
izability across readers and tasks as across other facets of measurement” (p. 6).
Though carefully trained readers can achieve acceptable rates of reliability, Moss
(1994), an educational measurement theorist, argued that with “portfolios,
where tasks may vary substantially from student to student, and where multiple
tasks may be evaluated simultaneously, inter-reader reliability may drop below
acceptable levels for consequential decisions about individuals or programs” (p.
6). Moss concluded that “although growing attention to the consequences of
assessment use in validity research provides theoretical support for the move
toward less standardized assessment, continued reliance on reliability, defined as
quantification of consistency among independent observations, requires a signif-
icant level of standardization,” (p. 6). However, these less standardized forms of
assessment are often preferable “because certain intellectual activities” cannot be
documented through standardized assessments (p. 6).
Moss (1994) suggested that in educational assessment, we look beyond psy-
chometric theories and practices in cases where acceptable reliability rates are
difficult or impossible to achieve. She challenged the assessment community to
consider its definitions of reliability—and here we in writing assessment need
to remember that reliability is more than a quantification of consistency among
independent observations. Moss recommended a hermeneutic approach because
as a philosophical tradition, it values a “holistic and integrative approach to
interpretation of human phenomena” (p. 7). After summarizing the key per-
spectives of hermeneutics, Moss explained how this methodology would work:
A hermeneutic approach to assessment would involve holistic,
integrative interpretations of collected performances that seek
to understand the whole in light of its parts, that privilege read-
ers who are most knowledgeable about the context in which the
assessment occurs, and that ground those interpretations not
only in textual and contextual evidence available, but also in a
rational debate among the community of interpreters. (p. 7)
Critical features of this type of assessment include the recognition of dis-
agreement or difference in interpretations as evaluators bring their expertise and
experience to bear on the work. Positions of individual evaluators can change
47
O’Neill
as rational debate ensues, with the final decision coming out of consensus or
compromise. In supporting this approach in specific situations, Moss (1994)
reminded readers that reliability and objectivity are no guarantors of truth and
that they can, in fact, work against “critical dialogue” and can lead “to proce-
dures that attempt to exclude, to the extent possible, the values and contextu-
alized knowledge of the reader and that foreclose[s] on dialogue among readers
about specific performances being evaluated” (p. 9). Mislevy (2004), saw bene-
fits in Moss’s idea but also commented:
In assessment, as in other fields, difficulties arise when
novel problems appear and the usual heuristics fail. We now
envisage assessments that target inferences more subtle than
proficiency in a specified domain of tasks. . . . We must return
to first principles to establish the credentials of this evidence
. . . The hermeneutic tradition does offer insights into draw-
ing inferences from disparate masses of evidence, and we can
indeed learn much from dialectic between psychometrics and
hermeneutics. (p. 2)
He advises, though, that a first step is to acquire “a deeper understanding
of psychometric methods, an understanding of principles behind methods that
will not be found in common wisdom, familiar testing practices, or standard
textbook presentations” (p. 2).
Moss’s comments about a hermeneutical approach to complex performance
assessment echoed what writing assessment scholars praised about holistic scor-
ing sessions and alternative methods for evaluating student writing (whether
portfolios or essays). White (1993, 1994), who has been a stalwart supporter
of holistic scoring of student writing has often expounded on the benefits as-
sociated with norming and scoring sessions. Scholars, reporting on portfolio
assessments, made similar statements such as Hamp-Lyons and Condon (2000):
Instead of focusing on scores, readers spend time bringing
their reading processes into line with each other. They read
and discuss samples with an eye toward developing and refin-
ing a shared sense of values and criteria for scoring. In other
words, this method fosters a reading community in which re-
liability grows out of the readers’ ability to communicate with
each other, to grow closer in terms of the ways they approach
samples. (p. 133)
Although Hamp-Lyons and Condon (2000) addressed reliability in this way,
they still used more traditional reliability evidence to justify portfolio assessment:
48
Reframing Reliability for Writing Assessment
“The reliability obstacle, in some local contexts, has been overcome. Miami Uni-
versity’s reliability statistics, like Michigan’s, are within the .8 range of holistic
essay assessments . . .” (p.91). Their position echoed White’s (1993) concerns
about portfolios. However, Hamp-Lyons and Condon (2000) and others did
not address the concerns about reliability articulated by Cherry and Meyers
(1993), Shale (1996), and Hayes and Hatch (1999).
Other scholars pushed against the traditional holistic scoring approach de-
signing methods that privileged those most knowledgeable about the context,
that encouraged critical dialogue, and that used holistic and integrative judg-
ments. Smith (1992, 1993) found that placement decisions for students enter-
ing college composition were more reliable with an expert reader system than
when made via traditional holistic scoring procedures. In Smith’s system, readers
made decisions based on the most recent course they taught, either accepting or
rejecting the student for the course or rejecting. Haswell and his colleagues at
Washington State University (Haswell & Wyche, 1996; Haswell, 2001) devel-
oped a two-tiered expert reader system in which readers made the initial decision
of whether or not a student should start in the regular composition course—
the one most students take. A panel of expert readers made decisions for those
students who did not fit neatly into this course. In making their decisions, the
panel of readers could consult and discuss difficult cases instead of following
the standardized, objective procedures associated with holistic scoring. Writing
program administrators at the University of Cincinnati used a system of port-
folio assessment to replace the first-year composition essay exit exam (Roemer,
Shultz, & Durst, 1991; Durst, Roemer, & Shultz, 1994). The portfolio scoring
system used large group “norming” sessions in conjunction with trios of writ-
ing teachers who worked independently to determine if students met the basic
requirements to successfully exit the composition program. These alternative
systems were still interested in reliability but not in achieving acceptable rates
through the conventional approach to holistic scoring.
Others (e.g., Broad, 1994; Lowe & Huot, 1997; and Hester et al., 2007) chal-
lenged the traditional holistic scoring approach that characterized most portfo-
lio assessments. Broad (1994) and White (1993, 1995) represented the concerns
about reliability that circulated around the use of portfolios as a large-scale assess-
ment method, but writing assessment scholars as a field still did not interrogate the
concept of reliability. More recently, White (2005) noted the difficulty in reaching
acceptable reliability rates that has plagued portfolio assessments and proposed a
scoring method for portfolios “derived conceptually from portfolio theory, rather
than essay-testing theory” (p. 583), overturning his earlier position that portfoli-
os are basically just expanded essay tests (White, 1995). Although White (2005)
seemed to be advocating a method of portfolio evaluation distinct from holistic
49
O’Neill
scoring, he describes his approach, which focuses on the reflective letter or self-as-
sessment and clear statements of learning goals, this way:
Now we can speak sensibly of scoring, even holistic scoring,
of the reflective letter, which needs to meet certain quite
specific criteria. We are back to a single document, the basic
material for which holistic scoring was designed, and we can
usually agree on the quality of that document, though we may
disagree on the quality of the items in the portfolio that sup-
port that document. With some labor, we can come up with
a scoring guide and sample portfolios at various score points,
just as we can do with single essays. (p. 593)
In short, White’s new method was closely aligned with the old one and was
designed to streamline the portfolio scoring by focusing on a single text. Grant-
ed, he explained how the portfolio contents were used along with the writer’s
self-assessment, but he still framed of reliability in traditional conventional ways.
While Moss (1994) recognized that reliability standards, within the psycho-
metric tradition, are grounded in fairness to stakeholders, she contends that
from a hermeneutic perspective, reliability “can be criticized as arbitrarily au-
thoritarian and counterproductive” (pp. 9-10). In the end, Moss did not argue
for abandoning reliability but rather advocated that alternative approaches to
assessment theory and practice be considered when appropriate (p.10). Her po-
sition is especially relevant for those charged with writing assessments because
writing is a complex, multidimensional, contextually situated activity. Import-
ing psychometric theory and practices, especially in terms of reliability, may
undermine the very usefulness of a writing assessment’s results. However, psy-
chometric theory cannot be dismissed out of hand; instead, writing assessment
scholars and practitioners need to draw on language, literacy and psychomet-
ric theories as well as other interpretive traditions to design assessments. Some
scholars in college composition have done this (Smith, 1992, 1993; Haswell &
Wyche, 1994; Broad, 1994, 2003; Lowe & Huot, 1997; Huot, 2002) there are
still many assessment practitioners who conform to more narrow approaches,
relying on an interrater reliability statistic to demonstrate reliability as we saw
with Borrowman (1999).
REFRAMING RELIABILITY
Moss’s (1994) argument to reconsider reliability through alternative research
traditions appeals to those of us in writing assessment more comfortable with
literacy studies, literary theory, and qualitative research methods. However, it
50
Reframing Reliability for Writing Assessment
51
O’Neill
52
Reframing Reliability for Writing Assessment
53
O’Neill
54
Reframing Reliability for Writing Assessment
language merely evokes and reinforces the frame. Therefore, we need to be more
intentional and thoughtful about the language we use in discussing reliability
and writing assessment.
Lakoff (2006) explains that language choice is “vital” because “language
evokes frames—moral and conceptual frames” (p. 7). So far, we have allowed the
psychometric practitioners (and I would also argue conservative policymakers
and their constituencies) to frame reliability in ways that privilege their world-
view and support their values. We need to consider ways to reframe reliability so
that it evokes the values that literacy teachers hold and support in their research
about teaching, learning, and language. Thinking about reliability as a concept,
an issue, as well as the frame it evokes and how we can communicate more ef-
fectively about what we value, is a role that literacy educators are able to tackle
because it shifts the debate away from statistical methods and technical expertise
to the concept of reliability, the values it promotes, and the ways these values are
communicated (Parkes, 2007).
While few writing teachers and theorists are psychometricians or experts in
advanced statistics, many more are experts in language and literacy. We un-
derstand communication theory and language development. We know about
teaching, learning, and students. We have strong values and beliefs—such as
a belief that all children can learn, that all deserve access to quality education,
that context is critical in effective writing, and that writing assessment should
improve teaching and learning. Smith (1992, 1993) explored multiple aspects
of reliability in a series of ongoing studies that were, in effect, a process of val-
idating the locally-designed placement system he developed (O’Neill, 2003).
His goal was to make sure that students in his program were placed in the most
appropriate first-year writing course. Huot (2002) in arguing for a new theory
of writing assessment that values context, local control, rhetorical principles, and
accessibility considered reliability as part of the validation process.
Reframing a concept as ingrained and complex as reliability requires a com-
mitment because frames are developed overtime, unconsciously in most cases,
through repetition and reinforcement. Everyone has frames—and they are not
always theoretically consistent or compatible—although people are usually not
aware of them because they function at the unconscious level. To reframe an
issue, Lakoff (2006) explained, we need to be strategic. With reliability, we can
start by determining how it has been framed and then how we can reframe it
in ways that support our beliefs about teaching and learning. One place to start
is with the standard reference manuals in the field of psychometrics. To that
end, below are excerpts of basic explanations of reliability from the most re-
cent editions of two mainstream measurement reference manuals: The Standards
of Educational and Psychological Testing (AERA, APA, & NCME, 1999) and
55
O’Neill
Educational Measurement, 4th edition (ACE, 2006). From the Standards, here’s
the opening paragraph on the section “Reliability and Errors of Measurement”:
A test, broadly defined, is a set of tasks designed to elicit or a
scale to describe examinee behavior in a specified domain, or a
system for collecting samples of individual’s work in a partic-
ular area. Coupled with the devise is a scoring procedure that
enables the examiner to quantify, evaluate, and interpret the
behavior or work samples. Reliability refers to the consistency
of such measurements when the testing procedure is repeated
on a population of individuals or groups. (p. 25)
According to the glossary in the Standards (AERA, APA, & NMCE, 1999),
reliability is “the degree to which test scores for a group of test takers are con-
sistent over repeated applications of a measurement procedure and hence are
inferred to be repeatable for an individual test taker” (p.180). It also includes the
“degree to which scores are free of errors of measurement for a given group” (p.
180). Haertel (2006), in the fourth edition of Educational Measurement, opens
the chapter on reliability this way:
The concern of reliability is to quantify the precision of test
scores and other measurements . . . Like test validity, test score
reliability must be conceived relative to particular testing
purposes and contexts. The definition, quantification, and
reporting of reliability must each begin with considerations
of intended test uses and interpretations. However, whereas
validity is centrally concerned with the nature of the attri-
butes tests measure, reliability is concerned solely with how
the scores resulting from measurement procedure would be
expected to vary across replications of that procedure. Thus
reliability is conceived in more narrowly statistical terms than
is validity. (p. 65)
Both of these explanations highlight the statistical, technical apparatus that
typically frames reliability. In this frame, quantification and measurement are
invoked. Measurement implies a finite amount of something. This epistemol-
ogy is associated with objectivity that was the central to psychometrics in the
early and mid-twentieth century (Williamson, 1993, 1994). However, in these
excerpts, values of consistency and accuracy are also identified. Haertel (2006)
even acknowledged context as a value when he notes that it “must each begin
with considerations of intended test uses and interpretations” (p. 65) since these
aspects of an assessment will define, in part, the particular situation. And in
56
Reframing Reliability for Writing Assessment
fact, these are also values that were central to the development of psychomet-
rics. Parkes (2007) argued that reliability as a concept has been conflated with
its methodology and that what we need to do is remember that it is the values
that are primary. Camp (1993) made a similar point. The methods to demon-
strate reliability should not be more important that the values that reliability
represents. Parkes (2007) explained it this way:
The outcomes of the use of these tools—reliability coeffi-
cients, dependability coefficients, standard errors of mea-
surement, information functions, agreement indices—serve
as evidence of broader social and scientific values that are
critically important in assessment. So a reliability coefficient is
a piece of evidence that operationalizes the values of accuracy,
dependability, stability, consistency, or precision. In practice
and in rhetoric, however, the methodologies for evidence reli-
ability are often conflated with the social and scientific values
of reliability. (p. 2)
If the methods cannot produce the evidence needed to support reliability,
then we need to develop better methods. Parkes (2007) contended that reliabil-
ity, like validity, needs to be considered as an argument. According to Parkes
(2007), a reliability argument has six components, the first and most critical of
these is determining the social and scientific values clearly. He argued that in
constructing a reliability argument, assessment developers need to
1. Determine the social and scientific values (dependability, consistency,
etc.) that are most relevant and decide which ones are most important.
2. Articulate clear statements of the purpose and context of the assessment,
which includes making explicit the reasons the information is needed and
how it will be used.
3. Define “replication” in the particular context, specifically structural ver-
sus conceptual replication.
4. Determine the “tolerance” or level of reliability needed.
5. Collect the evidence from the assessment, which may include traditional
reliability data but it might also include other information such as narra-
tive evidence.
6. Pull all of the information together to make the judgment and explaining
how the evidence supports the final judgment. (pp. 6-7)
Parkes also emphasized that at the start, it is “easy to think of methods . . .
rather than values” first but that it is “critical to stay focused on the value itself ”
and to determine which value or values are more important than others (p. 6).
57
O’Neill
58
Reframing Reliability for Writing Assessment
CONCLUSION
Writing assessment scholars and practitioners have had significant influence in
promoting performance-based assessments as well as in developing methods for
scoring them (Lane & Stone, 2006). However, these assessment experts have not
always been experts in language and literacy but in psychometrics and educa-
tional measurement. In many ways, writing specialists have been content to as-
sign reliability and reliability methods to psychometricians, distancing ourselves
from it. Parkes’ (2007) contention, that reliability (like validity) needs to be con-
sidered as an argument, demands language and literacy experts to participate in
discussions of reliability because constructing the reliability argument requires
knowledge of more than psychometric statistics and methods. Reframing reli-
ability to emphasize our values about writing, teaching writing, and learning to
write will emphasize finding methods to build an effective reliability argument
instead of merely reporting reliability co-efficients, which scholars have demon-
strated to be problematic in writing assessment practice (Cherry & Meyer, 1993;
Hayes & Hatch, 1999).
In Parkes’ (2007) approach to reliability, writing assessment administrators
would need to explain how reliability is being determined, why this approach is
appropriate in the particular context, how specifically reliability is being calcu-
lated, the threshold for acceptable reliability and a justification for it, the limita-
tions of the reliability, and how reliability contributes to the overall validation
of the assessment’s results. In determining reliability, many of us responsible for
writing assessments should collaborate as equal partners with colleagues who
have the statistical expertise. Writing assessment practitioners and scholars need
to accept our responsibility to develop and maintain writing assessments that
are informed by both language-based and psychometric theory and research.
59
O’Neill
We need to develop new methods for assessment as well as for determining reli-
ability and validity if current methods do not work adequately for our purposes,
as Parkes (2007) argued. This may mean collaborating with others who have
different kinds of experiences and expertise, learning more about psychometric
theory and practices, and engaging in difficult discussions with colleagues about
what we value and why it matters.
By emphasizing values, we can begin to not only reframe reliability but
also build more collaborative relationships with the educational measurement
community. In writing assessment, this reframing can help writing teachers and
administrators discuss and negotiate appropriate writing assessments with in-
stitutional administrators and others in more nuanced and effective ways. We
must remember that validity and reliability connect to values such as accuracy,
consistency, fairness, responsibility, and meaningfulness that we share with oth-
ers, including psychometricians and measurement specialists. Focusing on these
values and working to develop methods for upholding them can lead to the
development of writing assessment methods that not only support teaching and
learning but also are supported by evidence-based and theoretically-informed ar-
guments. Over time, we will be able to shift the frame associated with reliability
away from statistical methods and calculations to values that these methods—as
well as methods not yet developed—should be supporting.
I believe we can be successful in our efforts to reframe reliability; after all,
we were instrumental in resisting the move away from essay exams made in the
1940s, insisting that student writing needed to be evaluated in writing assess-
ment. This position led to the development of holistic scoring and other meth-
ods for evaluating performance assessments (Huot & Neal, 2006; Lane & Stone,
2006). As scholars, teachers, and assessment practitioners, we need to engage in
thoughtful ways to reframe reliability so that our assessments serve students and
programs as they enact what we know about language and literacy.
REFERENCES
Allen, M. S. (1995). Valuing differences: Portnet’s first year. Assessing Writing, 2(1),
67-89.
American Educational Research Association, American Psychological Association, &
National Council of Measurement in Education. (1999). Standards for educational
and psychological testing. American Educational Research Association.
Belanoff, P., & Dickson, M. (Eds.). (1991). Portfolios: Process and product. Boynton/
Cook.
Black, L., Daiker, D., Sommers, J. & Stygall, G. (Eds.). (1994). New directions in
portfolio assessment: Reflective practice, critical theory and large-scale scoring. Boynton/
Cook.
60
Reframing Reliability for Writing Assessment
Black, L., Helton, E., & Sommers, J. (1994). Connecting current research on
authentic and performance assessment through portfolios. Assessing Writing, 1(1),
247-266.
Borrowman, S. (1999). Trinity of portfolio placement: Validity, reliability, and
curriculum reform. Writing Program Administration, 23(1-2), 7-27.
Broad, B. (2003). What we really value: Beyond rubrics in teaching and assessing writing.
Utah State University Press.
Broad, R. L. (1994). “Portfolio scoring”: A contradiction in terms. In L. Black, D.
Daiker, J. & Stygall, G. (Eds.). New directions in portfolio assessment: Reflective
practice, critical theory and large-scale scoring (pp. 263-277). Boynton/Cook.
Burke, K. (1966). Language as symbolic action. University of California Press.
Calfee, R. & Perfumo, P. (Eds.). (1996). Writing portfolios in the classroom. Lawrence
Erlbaum.
Camp, R. (1993). Changing the model for the direct assessment of writing. In M. M.
Williamson and B. A. Huot (Eds.), Validating holistic scoring for writing assessment:
Theoretical and empirical foundations (pp. 45-78). Hampton.
Cherry, R. D., & Meyer, P. R. (1993). Reliability issues in holistic assessment. In M.M.
Williamson and B. A. Huot (Eds.) Validating holistic scoring for writing assessment:
Theoretical and empirical foundations (pp. 109-141). Hampton.
Common Core State Standards Initiative. n.d. National Governors Association and
Council of Chief State School Officers. [Link]
Conference on College Composition and Communication. (Nov. 2006). Writing
assessment: A position statement (Rev. ed.). National Council of Teachers of
English. [Link]
Daiker, D. A., Sommers, J., & Stygall, G. (1996). Pedagogical implications of a
college placement portfolio. In E. M. White, W. D. Lutz, & S. Kamusikiri (Eds.),
Assessment of writing: Politics, policies, practices (pp. 257-270). Modern Language
Association.
Diederich, P. B. (1974). Measuring growth in English. National Council of Teachers of
English.
Durst, R. K., Roemer, M., & Schultz, L. (1994). Portfolio negotiations: Acts in speech.
In L. Black, D. Daiker, J. Sommers & G. Stygall (Eds.) Validating holistic scoring for
writing assessment: Theoretical and empirical foundations (pp. 286-300). Boynton/
Cook.
Elbow, P. & Belanoff, P. (1986). Staffroom interchange: Portfolios as a substitute for
proficiency examinations. College Composition and Communication, 37(3), 336-339.
Elliot, N. (2005). On a scale: A social history of writing assessment in America. Peter
Lang.
Faigley, L, Cherry R. D., Jolliffe, D. A., & Skinner, A. M. (1985). Assessing students’
knowledge and processes of composing. Ablex.
Gere, A. R. (1980). Written composition: Toward a theory of evaluation. College
English, 42(1), 44-48, 53-58.
Godshalk, F. I., Swineford, F. & Coffman, W. E. (1966). Measurement of writing
ability. (CEEB RM No. 6.) Educational Testing Service.
61
O’Neill
62
Reframing Reliability for Writing Assessment
LeMahieu, P. G., Eresh, J. T., & Wallace, R. C. (1992). Using student portfolios for a
public accounting. School Administrator, 49(11), 8-13.
LeMahieu, P. G., Gitomer, D., & Eresh, J. (1995). Portfolios in large scale assessment:
Difficult but not impossible. Educational Measurement: Issues and Practice, 14(3),
11-28.
Lowe, T. J., & Huot, B. (1997). Using KIRIS writing portfolios to place students in
first-year composition at the University of Louisville. Kentucky English Bulletin, 46,
46-64.
Lynne, P. (2004). Coming to Terms: A theory of writing assessment. Utah State University
Press.
Messick, S. (1989). Meaning and value in test validation: The science and ethics of
assessment. Educational Researcher, 18(2), 5-11.
Mislevy, R. J (2004). Can there be validity without “reliability?” Journal of Educational
and Behavioral Statistics, 29(2), 241-245.
Moss, P. A. (1994). Can there be validity without reliability? Educational Researcher,
23(4), 5-12.
Murphy, S. & Underwood, T. (2000). Portfolio practices: Lessons from schools, districts
and states. Christopher Gordon.
Nelson, A. (1999). Views from the underside: Proficiency portfolios in first-year
composition. Teaching English in the Two Year College, 26, 243-253.
Nystrand, M., Cohen, A., & Dowling, N. (1993). Addressing reliability problems in
the portfolio assessment of college writing. Educational Assessment, 1(1), 53-70.
O’Neill, P. (2003). Moving beyond holistic scoring through validity inquiry. Journal of
Writing Assessment, 1(1), 47-65. [Link]
Parkes, J. (2007). Reliability as argument. Educational Measurement: Issues and Practice,
26(4), 2-10.
Peckham, I. (2009). Online placement in first-year writing. College Composition and
Communication, 60(3), 517-540.
Penrod, Diane. (2005). Composition in convergence: The Impact of new media on writing
assessment. Lawrence Erlbaum.
Roemer, M., Schultz, L. M., & Durst, R. K. (1991). Portfolios and the process of
change. College Composition and Communication, 42(4), 445-469.
Shale, D. (1996). Essay reliability: Form and meaning. In E. M. White, W. D. Lutz,
& S. Kamusikiri (Eds.), Assessment of writing: Politics, policies, practices. (pp. 76-96).
Modern Language Association.
Smith, W. L. (1992). The importance of teacher knowledge in college composition
placement testing. In J. R. Hayes (Ed.), Reading empirical research studies: The
rhetoric of research. (pp. 289-316). Ablex.
Smith, W. L. (1993). Assessing the reliability and adequacy of using holistic scoring
of essays as a college composition placement program technique. In M. M.
Williamson and B. A. Huot (Eds.), Validating holistic scoring for writing assessment:
Theoretical and empirical foundations (pp. 142-205). Hampton Press.
Sommers, J., Black, L., Daiker, D., & Stygall, G. (1993). Challenges of rating
portfolios: What WPAs can expect. Writing Program Administration, 17(1-2), 7-29.
63
O’Neill
64
CHAPTER 2.
Diane Kelly-Riley
Washington State University
This article examines the intersections of students’ race with the eval-
uation of their writing abilities in a locally-developed, context-rich,
university-wide, junior-level writing portfolio assessment that relies
on faculty articulation of standards and shared evaluation practices.
This study employs sequential regression analysis to identify how faculty
raters operationalize their definition of good writing within this uni-
versity-wide writing portfolio assessment, and, in particular, whether
students’ race accounts for any of the variability in faculty’s assessment of
student writing. The findings suggest that there is a difference in student
performance by race, but that student race does not contribute to facul-
ty’s assessment of students’ writing in this setting. However, the findings
also suggest that faculty employ a limited set of the criteria published
by the writing assessment program, and faculty use non-programmatic
criteria—including perceived demographic variables—in their oper-
ationalization of “good writing” in this writing portfolio assessment.
This study provides a model for future validity inquiry of emerging
context-rich writing assessment practices.
DOI: [Link] 65
Kelly-Riley
66
Validity Inquiry of Race and Shared Evaluation Practices
67
Kelly-Riley
provide a sound scientific basis for the proposed score interpretations. It is the
interpretations of test scores required by proposed uses that are evaluated, not the
test itself. When test scores are used or interpreted in more than one way, each
intended interpretation must be validated. (AERA, p. 9) Kane (2006) asserted that
validation employs two kinds of argument. An interpretive
argument specifies the proposed interpretations and uses
of test results by laying out the network of inferences and
assumptions leading from the observed performances to the
conclusions and decisions based on the performances. The
validity argument provides an evaluation of the interpretive
argument. (p. 23)
The relevance of validity to writing assessment practitioners is apparent when
validity is understood as an ongoing argument to be made rather than a stat-
ic state to be achieved and justified. O’Neill (2003) contends that “validation
arguments are rhetorical constructs that draw from all the available means of
support” (p. 50). Huot and Schendel (1999) assert that validity and “assessment
must be discussed in the context of ethics, for the consequences of assessment
procedures are closely tied to the political and social contexts in which they take
place” (p. 40). O’Neill (2003) argues that such lines of inquiry and research
“[demonstrate] how systematic, ongoing validity research [function] to enhance
a particular local test and contributes—both theoretically and practically to the
scholarship of writing assessment” (p. 48). However, in spite of innovations and
implementations of new contextually-based college writing assessment practices,
systematic and rigorous validity inquiry into emerging college writing assess-
ment practices have been limited.
O’Neill notes the reductive tendency in composition studies to simplify va-
lidity to mean “honesty . . . accuracy . . . and rightness” (2003, p. 49) that lim-
its the complexity of the construct. There are many important theoretical calls
for the discipline to wrestle with validity issues contextually or hermeneutically
(Huot, 1996; Huot and Schendel, 1999; Moss, 1998a; Murphy, 2007; Inoue,
2007) and few forays of actual research and practice into validity inquiry in
college writing assessment (Smith, 1993; Williamson and Huot, 1993; Haswell,
1998a and 2000; Broad, 2000; O’Neill, 2003; Hester, O’Neill, Neal, Edging-
ton, & Huot, 2003; Elliot, Briller, & Joshi, 2007; Gere, Aull, Green and Porter,
2010). Researchers and scholars have neglected to conduct validity inquiries
of locally developed writing assessment practices and so have not documented
contributions or innovations these practices embody, and they fail to be atten-
tive to students who take the exams. Kane (2006) says “there are, potentially, a
large number of assumptions in any interpretive [validational] argument. We
68
Validity Inquiry of Race and Shared Evaluation Practices
take many of these assumptions for granted, at least until evidence to the con-
trary develops” (p. 23). To unearth some of these assumptions, previous scholars’
criticism of standardized testing helps articulate where to begin: “what kind of
proof do we have that students are wrong when they say, ‘I don’t belong in this
dummy class’?” (Elbow, 1996, p. 93). While Elbow originally leveled this ques-
tion at holistic or standardized tests, it is still relevant as a question for writing
assessment programs that employ shared evaluation practices—locally devel-
oped, context-rich, practices that rely on faculty articulation values—whether
via portfolios, direct-self placement, or other methods. Students who don’t meet
standards for writing tests face consequences that require completing additional
coursework, spending additional time, spending additional money (perhaps),
and dealing with the stigma of not passing the “test”. Moss (1995) cites Cron-
bach and argues that “when the anticipated consequences [of assessment] ‘im-
pinge on the rights and life chances of individuals’ (Cronbach, 1988, p. 6) . . .
the investigation of consequences becomes particularly salient” (p. 11).
Rigorous validity inquiry allows for in-depth investigation of issues that we
observe anecdotally—from student outrage at perceived unfair testing practices
to patterns of course enrollment that may have more students of color populating
the required writing support courses. Rigorous validity inquiry enables informed
practice in a setting and directly addresses concerns of power highlighted by Huot
and Williamson (1997) who note “assessment procedures [are] instruments of
power and control, revealing so-called theoretical concerns as practical and po-
litical” (p. 44). They “fear that unless we make explicit the important power re-
lationships in assessment, portfolios will fail to live up to their promise to create
important connections between teaching, learning and assessing” (p. 44). Such
a fear is applicable to any form of writing assessment that uses shared evaluation
practices, particularly as these issues relate to test fairness. Camilli (2006) asserts
while there are many aspects of fair assessment, it is generally agreed that tests
should be thoughtfully developed and that the conditions of testing should be
reasonable and equitable for all students . . . fairness issues are inevitably shaped
by the particular social context in which they are embedded. (p. 221)
Certainly, as Schmidt and Camara (2004) observe, there have been “per-
sistent score differences among racial groups” (p. 189) for a variety of standard-
ized tests. Similar studies for performance-based assessments are still inconclu-
sive but suggest that “subgroup gaps on traditional tests remain for [performance
based] assessments” (p. 193). Most of this research, though, has occurred at the
primary and secondary school level and not the college level.
Camilli (2006) states that “large differences are commonly encountered in
test scores among groups of different races and ethnicities, and it is important to
understand the extent to which these differences are artifacts of a test rather than
69
Kelly-Riley
70
Validity Inquiry of Race and Shared Evaluation Practices
and those of ethnicity with culture, the two concepts are not
clearly distinct from one another.
The APA Task Force on Diversity Issues at the Precollege and Undergraduate
Levels of Education in Psychology (1998) argued that:
“Race” has social meaning often accompanied by stereotyp-
ing; it suggests one’s status within the social system and intro-
duces power differences as people of different “races” interact
with one another. ‘Ethnicity,’ on the other hand, connotes
common culture and shared meaning. It includes feelings,
thoughts, perceptions, expectations, and actions of a group
resulting from shared historical experiences.
This study represents a starting point for this type of research, and hopefully
future studies can include more complex representations of race and ethnicity.
71
Kelly-Riley
widespread feeling existed among faculty about the poor quality of student writ-
ing. And, perhaps, the shift from holistically evaluating writing to relying on an
innovative system of evaluation represented a significant enough move to not im-
mediately surface new issues that embedded in the new methodology.
In response to the African American student’s question, I investigated students
of colors’ performances on the Writing Portfolio for Academic Year 2004-05 ac-
cording to the racial classifications collected by my institution. At that time, this
institution reported the demographic profile of undergraduate students as 76 per-
cent White; 1 percent American Indian Alaskan Native; 2 percent Black; 6 percent
Asian Pacific Islander (API); 4 percent Hispanic; 3 percent non-resident aliens; and
8 percent unknown. Students’ racial affiliation was obtained from this institution’s
Institutional Research Office by U.S. Census Bureau/ OMB categories, and then
was combined with students’ Writing Portfolio results. Tables 1.1 and 1.2 docu-
ment the difference in performance percentages for the impromptu exam portion
of the Writing Portfolio and the final review of the entire Writing Portfolio.
Tables 1.1 and 1.2 illustrate an unevenness in performance on the Writing
Portfolio by race. Simply examining the percentages does not indicate whether
these differences are significant. An analysis of variance was conducted on the
performances of students on the timed writing portion of the Writing Portfolio by
race. A random sample of 508 timed writing records were selected from the 5347
Writing Portfolio records recorded during AY 2004-2005. Students who spoke En-
glish as a second language were omitted from this analysis. An analysis of variance
showed that the difference in performance by race on the timed writing portion of
the Writing Portfolio was significant, F (4, 503)=6.032, p=.000. Post hoc analyses
using Tukey’s LSD for significance indicated that Black (M=1.58, SD=.496), API
(M=1.59, SD=.509), and Hispanic (M=1.7, SD=.462) students’ timed writing
performances were significantly lower than White students (M=1.84, SD=.550).
72
Validity Inquiry of Race and Shared Evaluation Practices
73
Kelly-Riley
74
Validity Inquiry of Race and Shared Evaluation Practices
75
Kelly-Riley
METHODS
This validity inquiry examines how raters functionally define good writing
through sequential regression analysis techniques that examine actual student
products submitted for the Writing Portfolio. These analyses are conducted
through three frameworks: the Writing Portfolio assessment criteria, an Al-
ternate set of writing criteria, and Demographic criteria. Each of these frame-
works are applied to the two distinct writing tasks—impromptu writing and
coursework written for regular undergraduate courses across the disciplines—
selected for inclusion in the Writing Portfolio as representative of the student’s
best writing.
This inquiry allows for more sophisticated statistical analysis of the factors
that account for the variability in the writing quality scores of the Writing
Portfolio using a finer grained instrument, the Writing Portfolio Differential
Scale for Writing and Demographic Information (see Appendix A). The Writ-
ing Portfolio Differential Scale was developed by this researcher to interpret
raters’ evaluation behaviors and determine the criteria they seemed to actually
use to evaluate writing; the criteria that seemed to carry more weight in their
evaluation process; and whether demographic features perceived about writers
accounted for any part of the evaluation results. Guiding questions for this
inquiry include:
1. What is the definition of “good writing” that faculty raters apply when
evaluating the Writing Portfolio?
2. Do faculty raters make demographic assumptions about students based
on their writing that effect the results?
3. Does this evaluation privilege forms of writing according to race?
Sequential regression analysis is used to assess the relationship between a
dependent variable (like writing quality) and several independent variables (like
criteria that comprise quality—focus, organization, or use of Standard American
English and so on) by entering variables in a specific order into regression equa-
tions to identify which variables account for the variability in—or the criteria
that comprise—the overall score (Tabachnick & Fidell, 2006). The methodol-
ogy for this study was piloted in an earlier project by the researcher in which
76
Validity Inquiry of Race and Shared Evaluation Practices
the Writing Portfolio Differential Scale was tested and the order of the criteria
variables were established for the regression analysis (Kelly-Riley, 2006).
The Writing Portfolio Differential Scale for Writing and Demographic In-
formation was adapted from the work of Piche, Rubin, Turner, and Michlin
(1978) and Osgood (1957). Piche et al. used the work of Osgood to examine
whether teachers evaluated Black elementary students’ writing differently from
their White counterparts. Osgood created semantic differential scales that “relate
to the functioning of representational processes in language behavior and hence
may serve as an index of these processes” (p. 9). Osgood’s work developed out
of experimental psychology to establish pairs that exist in what he called seman-
tic space, “which are assumed to represent a straight line function that passes
through the origin of this space, and a sample of such scale then represents a
multidimension space” (p. 25). His work attempts to quantify the complexity
inherent in measuring a construct like writing. Piche et al. (1978) developed
their scale items based on the research of Osgood (1957) and the application
of these scales by Williams, Whitehead, and Miller (1971) who examined re-
lationships of attitudes and children’s speech. Piche et al. examined teachers’
responses to their scale items by presenting teachers with different samples of
writing. Some of the samples contained inserted types of speech the researchers
characterized as African American Vernacular English (AAVE). Actual samples
of Black students’ writing were not used for their study. Instead, they used a
piece of writing, and added features identified as AAVE into the text.
In addition, Rubin and Williams-James (1997) examined the ways teachers
responded to international students’ writing using similar scales. These research-
ers created a text and inserted types of speech that appeared to be consistent with
writers from different nationalities. They did not use actual student products for
their evaluation. The instrumentations of these differential scales, however, set a
precedent to examine instructors’ impressions of students’ writing.
For this study, the Writing Portfolio Differential Scale contains three sep-
arate criteria frameworks: Writing Portfolio Criteria, the programmatic areas
articulated, published and evaluated by the Writing Program (Comprehension
of the Task, Focus, Organization, Support, and Proofreading) and two other
frameworks of criteria—Alternate Writing Criteria and Demographic Crite-
ria—which were previously used by Piche et al. and Rubin and Williams-James.
The Alternate Writing Criteria include Coherence, Use of Standard American
English, Logic, Grammar, Creativity, Level of Language Passivity, and Quality
of Writing; and the Demographic Criteria include the rater’s perception of the
writer in many areas: Strength of Writer, Intelligence, Socio-Economic Status,
Level of Cultural Advantage, Confidence, and Comfort as a Writer. This study
examined actual samples of student writing composed by actual students for
77
Kelly-Riley
78
Validity Inquiry of Race and Shared Evaluation Practices
Table 1.3. Variable Entry Order into the Sequential Regression Analysis
and Reliability Data for the Timed Writing analysis
Criteria Framework Variable Order of Entry Cronbach’s Alpha
Writing Portfolio Race, Focus, Proofreading, Support, Com- .8946
prehension of task, and Organization
Alternate Writing Race, Coherence, Logic, Creativity, Gram- .8902
mar, use of Standard American English,
Language passivity
Demographic Race, and raters’ perceptions of writers’ .7902
Confidence, Intelligence, Comfort with
writing, Socio-economic status, and Cultural
advantage
Note: N=150
Table 1.4 details the order of the criteria variables were entered into the se-
quential equation as well Cronbach’s alpha. The entry order of the variables are
slightly different based on the results from the stepwise analysis conducted by
Kelly-Riley (2006).
Table 1.4. Variable Entry Order into the Sequential Regression Analysis
and Reliability Data for the Course Paper analysis
Criteria Framework Order of Entry Cronbach’s Alpha
Writing Portfolio Race, Focus, Proofreading, Support, Organi- .8512
zation and Comprehension of task
Alternate Writing Race, Coherence, Logic, Grammar, Cre- .8647
ativity, Use of Standard American English,
Language passivity
Demographic Race, and raters’ perceptions of writers’ .8560
Comfort with writing, Intelligence, Confi-
dence, Socio-economic status, and Cultural
advantage.
Note: N=300
79
Kelly-Riley
RESULTS
The results from the sequential regression analyses for both types of writing—
impromptu and course paper submissions—suggest that the definition of good
writing is based more on variables of coherence, correctness, and confidence as
applied by faculty raters in this large scale Writing Portfolio assessment. Race
did not contribute significantly to faculty raters’ functional definition of “good
writing” for any of the frameworks whether in the timed exam format or for
the course papers. Surprisingly, faculty raters operationalize their assessment of
“good writing” based on criteria accounted more through non-programmatic
evaluation criteria of the Alternate Writing framework variables. In addition, a
high percentage of the writing scores—for impromptu writing as well as papers
written for courses—included demographic considerations of the writer. High-
er percentages of writing quality scores were accounted for through coherence
and grammar, part of the Alternate Writing framework. These variables over-
lap with focus and mechanics, which account for writing quality through the
Writing Portfolio framework, although the Writing Portfolio variables account
for a slightly lesser percentage of the writing quality score. A surprisingly high
percentage—nearly two thirds of the score—of writing quality is also accounted
for through raters’ perceptions of student writers’ intelligence and comfort with
writing. For each type of writing—impromptu exams and course paper submis-
sions—race did not contribute significantly to writing quality score through any
of the frameworks.
Table 1.5 details the separate regression equations that account for the vari-
ability in the impromptu exam analysis. The Alternate Writing Criteria account-
ed for the most variability in the impromptu writing score.
80
Validity Inquiry of Race and Shared Evaluation Practices
Tables 1.6, 1.7, and 1.8 provide detailed analysis of the significant regression
equations and the differences in the percentages of the variance explained for
each of the three separate criteria frameworks.
81
Kelly-Riley
Note. Each frame represents a separate regression equation with the variables in the order in which the
regression analysis specified. N=300 *p<.05. **p<.01.
Tables 1.10, 1.11 and 1.12 provide detailed analysis of the significant regres-
sion equations and the differences in the percentages of the variance explained
for each of the three separate criteria frameworks for the review of the course
papers.
82
Validity Inquiry of Race and Shared Evaluation Practices
83
Kelly-Riley
seem to apply idiosyncratic criteria that fall outside of the intended assessment.
Perhaps this disconnect can be explained by the explicit instructions in the rating
sessions for raters to reference their classroom writing experiences and expectations
and to be guided by the Writing Portfolio criteria in the assessment. They are asked
to operationalize the criteria as relevant to their disciplinary realities.
Similarly surprising, raters evaluate impromptu writing with slightly differ-
ent expectations than the writing done within courses. While Focus and Me-
chanics are included both Writing Portfolio criteria, Support is used by raters
to assess impromptu writing quality while Organization replaces it in the eval-
uation of the course papers. Likewise, this trend is observed in the Alternate
Writing criteria. Coherence and Grammar account for the variance in writing
quality for impromptu writing and course papers, but Creativity is important
in impromptu writing whereas Logic replaces it in the course paper writing.
These results suggest that faculty have different expectations for the two different
writing tasks included in the same Writing Portfolio. These results do not differ-
entiate between one set of criteria being better than the other; they only indicate
that faculty seem to view these tasks differently. Certainly, this interesting result
deserves further study.
The second research question examines whether the operationalized defi-
nition of good writing included demographic information. The findings sug-
gest that large percentages of the variance of writing quality are accounted for
through the Demographic framework—primarily through the rater’s perception
of the writer’s intelligence and comfort with writing. The variables of race, per-
ceived economic status, and perceived cultural advantage did not contribute
significantly to the writing quality score. While the two writing frameworks
have more obvious overlap, the demographic criteria seem to overlap with writ-
ing issues too. The demographic criteria that faculty use to account for writing
quality are based on variables that would be reasonable to identify a writer as
needing help: the student’s comfort level with writing, the student’s confidence
with writing, and the teacher’s perception of the student’s intelligence. The vari-
ables are not related to demographic features that are irrelevant to the classroom.
The third question examined whether the assessment process privileged
forms of writing according to race. The findings from this study suggest that
race is not a significant contributor to the faculty’s assessment of students’ writ-
ing for either the impromptu writing or papers written for courses. The results of
the sequential regression analyses suggest that race does not significantly account
for the variance in good writing. However, students’ performances by race on
the Writing Portfolio are significantly different like the studies conducted by
Schmidt and Camara (2004) and Breland et al. (2004), but the rating processes
used by faculty raters do not seem to the cause for these differences.
84
Validity Inquiry of Race and Shared Evaluation Practices
Such concern about the relationship between the rater and the writer is war-
ranted. Ball (1997) documented potential bias by readers for writers based on
dissimilar cultural backgrounds. Smitherman’s extensive research (highlighted in
Smitherman and Villanueva, 2003) has documented different linguistic struc-
tures of African American students and their implications in educational set-
tings. This study, though, found somewhat different results. The rating corps
used for this study represented a linguistically and culturally diverse set of facul-
ty—who were also representative of the regular Writing Portfolio rating corps—
attempting to address some of the concerns raised by Ball. Admittedly, this study
did not intend to address the specific relationship between rater and writer.
Furthermore, the extent to which Mechanics contributes to writing quality
is interesting in the light of Smitherman’s research, but this study included more
racial categories than Smitherman’s studies, which focused primarily on African
Americans. Such distinct differences between raters and writers with a multitude
of different backgrounds may not be as detectable as comparisons that look at
only two racial groups. Even though race was not a variable that accounted for
any writing quality in this study, some of Smitherman’s findings that connect
race and linguistic structure might seem supported by this study. Specifically,
Mechanics accounts for a great deal of the Writing Portfolio quality score. Me-
chanics accounts for 32.6% to the variance in impromptu writing and 21.4% of
the course papers. Overall, though, the Writing Portfolio criteria account for less
of the writing quality than the Alternate Writing criteria. In the Alternate crite-
ria, Grammar, while a significant contributor to quality, did not account for as
much as Mechanics, with 8.8% of the variance explained for impromptu writing
and 2.5% in the course papers. Issues of Coherence that go beyond Grammar
seem to be more important in raters’ assessment of students’ writing.
While the findings from this study suggest that race does not contribute
significantly to raters’ operationalization of good writing, it is disconcerting that
there are statistically significant differences in performances by race on the Writ-
ing Portfolio. While the reason may not be in how the raters evaluate student
writing, the subject requires further investigation. Schmidt and Camara (2004)
summarize the prevailing theories used to explain the gap in standardized test
performances by race: inequitable educational preparation, poverty, discrimi-
nation, poor educational opportunities, and lack of access to educational re-
sources. Studies such as these would be useful in large-scale performance-based
assessment programs, and studies examining the effectiveness of the structured
support required by these programs would be the next logical step in validity
research.
Validity, again, refers to the use and interpretation of test scores in a particu-
lar setting. What do these results mean for the use and interpretation of the test
85
Kelly-Riley
scores in the university-wide Portfolio context? The purpose of the Writing Port-
folio is to assess students’ readiness for the upper-division Writing in the Major
courses. In the rating sessions, faculty are overtly asked to draw on their class-
room experiences and expectations for the assessment situation. Perhaps this
request for raters to draw on pedagogical reference points helps explain the large
role that the Alternate Writing criteria play in accounting for writing quality.
These findings are consistent with Broad’s qualitative study that document the
frustration faculty felt in rubric-based assessments in first year writing programs.
Since the Writing Portfolio relies on the multitude of disciplinary definitions
of good writing, it’s important to have the starting point of common language
articulated in the Writing Portfolio criteria and to begin to fully acknowledge
the additional role that other non-programmatic criteria play. Additionally, in
this Writing Portfolio system, frustration levels are mitigated in that faculty don’t
have to agree about a static definition of writing; faculty simply place student
writers into three broad categories of placement: Pass, Pass with Distinction,
and Needs Work. The broadly defined functional placements mask the complex
process behind the rating behaviors. These rating behaviors need to be routinely
examined.
These findings suggest that writing assessment program administrators need
to play more of active role in looking at published program criteria, standards
for writing, and faculty enactment of these standards. The focused time allotted
to the evaluation of writing at the ends of the spectrum (weak or strong) in the
shared evaluation, expert-rater system does not translate into a systematic appli-
cation of the criteria of good writing as articulated through the programmatic
rubric. Locally developed writing assessment programs—whether portfolios or
directed self-placement or other mechanisms which rely on faculty articulation
of standards—need to compare the published criteria used by their programs to
the criteria used functionally through the rating process. This point is the “sig-
nificant site of power and knowledge” (O’Neill, 2003, p. 62) so often ignored
by compositionists.
A tension exists between the criteria the Writing Program articulates and pub-
lishes, and the actual multi-dimensional criteria enacted by faculty raters. The
ways in which programmatic criteria and disciplinary expectations intersect must
be examined further because they most certainly inform and reform each other in
a system that intends to be responsive to validity and reliability concerns. The ab-
sence or limited contribution of the some of the programmatic Writing Portfolio
criteria—comprehension of the task, organization, and support—to the writing
quality score points to a disjuncture. The findings suggest that these criteria—as
faculty use them—either contribute minimally to the quality score or not at all
for both the impromptu and course paper evaluations. This omission raises the
86
Validity Inquiry of Race and Shared Evaluation Practices
question as to whether raters—who are hired for these positions based on their ex-
tensive teaching expertise—don’t value or don’t know how to evaluate for these cri-
teria areas. While the program advertises and publishes specific criteria for evalua-
tion of Writing Portfolios, more than half of these criteria areas are not utilized by
raters in the evaluation setting. This omission questions the extent to which these
criteria are employed in classroom settings. The program administrators must be
aware of the tendency of raters to draw on personal pedagogical expectation, and
to move the raters toward the programmatic criteria particularly for decisions that
fall on the ends of the spectrum. Improved rater training and overt conversations
about this tendency in norming sessions might be a way to begin to further iden-
tify and address these issues.
Finally, research that includes more nuanced considerations of race and eth-
nicity into these large scale writing assessment practices need to be more com-
monplace. Educational research already has a robust agenda of research in stan-
dardized tests related to race and ethnicity, but most times, the standardized tests
are separated from the instructional or local context. Composition studies needs
to embrace a similar research agenda which considers the hermeneutically-orient-
ed assessment approaches that are rooted in local context. While this study only
examines the construct of good writing as applied by raters, there are many other
angles of necessary research and validity inquiry for students of colors’ experiences
in context-rich, locally-developed writing assessment programs. Given the more
mainstream position that college writing assessment methodologies have garnered
of late, such inquiry is important, timely, and vital—not only to examine the qual-
ity of the practices, but to ensure that such methodologies are not intentionally or
unintentionally leveling consequences for students—particularly those represent-
ed by small populations who may be easily overlooked.
REFERENCES
American Educational Research Association, American Psychological Association, &
National Council of Measurement in Education. (1999). Standards for Educational
and Psychological Testing. American Educational Research Association.
American Anthropological Association (1997). Response to OMB directive 15. Race
and ethnic standards for federal statistics and administrative reporting.
American Psychological Association Task Force on Diversity Issues at the Precollege
and Undergraduate Levels of Education in Psychology (1998). Enriching the focus
on ethnicity and race. Monitor, 29(3).
Ball, A. (1997). Expanding the dialogue on culture as a critical component when
assessing writing. Assessing Writing, 4(2), 169-202.
Bond, L. (1995). Unintended consequences of performance assessment: Issues of bias
and fairness. Educational Measurement: Issues and Practice, 14(4), 21-24.
87
Kelly-Riley
Breland, H., Kubota, M., Nickerson, K., Trapani, C., Walker, M. (2004). New SAT
writing prompt study: Analyses of group impact and reliability (Report No. 2004-1).
College Entrance Examination Board.
Broad, B. (2000). Pulling your hair out: Crises of standardization in communal writing
assessment. Research in the Teaching of English, 35(2), 213-260.
Camilli, G. (2006). Test fairness. In R. L. Brennan (Ed.), Educational Measurement (4th
ed) (pp. 221-256). American Council on Education/Oryx Press Series on Higher
Education.
Clary-Lemon, J. (2009). The racialization of composition studies: Scholarly rhetoric of
race since 1990. College Composition and Communication, 61(2), W1-W17.
Cronbach, L. (1988). Five perspectives on validity argument. In H. Wainer & H. I.
Braun (Eds.), Test validity (pp. 3-17). Erlbaum.
Elbow, P. (1996). Writing assessment in the 21st century: A utopian view. In L. Z.
Bloom, D. A. Daiker, & E. M. White (Eds.), Composition in the twenty-first century:
Crisis and change (pp. 83-100). Southern Illinois University Press.
Elliot, N., Briller, V., & Joshi, K. (2007). Portfolio assessment: Quantification and
community. Journal of Writing Assessment, 3(1), 5-30. [Link]
item/8nm1m6xc
Farr, M. & Nardini, G. (1996). Essayist literacy and sociolinguistic difference. In
E. M. White, W. D. Lutz & S. Kamusikiri (Eds.), Assessment of writing: Politics,
policies, practices (pp. 108-119). Modern Language Association.
Gere, A. R., Aull, L., Green, T., and Porter, A. (2010). Assessing the validity of directed
self-placement at a large university. Assessing Writing, 15(3), 154-176.
Griffee, D. (2002). Portfolio assessment: Increasing reliability and validity. The
Learning Assistance Review: The Journal of the Midwest College Learning Center
Association, 7(2), 5-17.
Hamp-Lyons, L. and W. Condon. (2000). Assessing the portfolio: Principles for practice,
theory and research. Hampton Press.
Haswell, R. (1998a). Multiple inquiry in the validation of writing tests. Assessing
Writing, 5(1), 89-109.
Haswell, R. (1998b). Rubrics, prototypes and exemplars: Categorization and systems
of writing placement. Assessing Writing, 5(2), 231-268.
Haswell, R. (2000). Documenting improvement in college writing: A longitudinal
approach. Written Communication, 17(3), 220-236.
Haswell, R. (Ed.). (2001). Beyond outcomes: Assessment and instruction within a
university writing program. Ablex.
Haswell, R. (2005). NCTE/CCCC’s recent war on scholarship. Written
Communication, 22(2), 198-223.
Haswell, R, & S. Wyche. (1996). A two-tiered rating procedure for placement essays.
In T. W. Banta (Ed.), Assessment in practice: Putting principles to work on college
campuses (pp. 204-207). Jossey-Bass.
Hester, V., O’Neill, P., Neal, M., Edgington, A., & Huot, B. (2007). Adding portfolios
to the placement process. In P. O’Neill (Ed.), Blurring boundaries: Developing
writers, researchers, and teachers (pp. 61-90). Hampton Press.
88
Validity Inquiry of Race and Shared Evaluation Practices
Huot, B. (1996). Toward a new theory of writing assessment. College Composition and
Communication, 47(4), 549-566.
Huot, B. & Schendel, E. (1999). Reflecting on assessment: Validity inquiry as ethical
inquiry. Journal of Teaching Writing, 17(1-2), 37-55.
Huot, B. & Williamson, M. M. (1997). Rethinking portfolios for evaluating writing:
Issues of assessment and power. In K. B. Yancey and I. Weiser (Eds.), Situating
portfolios: Four perspectives (pp. 43-56). Utah State University Press.
Inoue, A. (2007). Articulating Sophistic rhetoric as a validity heuristic for writing
assessment. Journal of Writing Assessment, 3(1), 31-54. [Link]
item/64n8z5mz
Kane, M. T. (2006). Validation. In R. L. Brennan (Ed.) Educational measurement (4th
ed.) (pp. 17-64). American Council on Education/Oryx Press Series on Higher
Education.
Kelly-Riley, D. (2006). A validity inquiry into minority students’ performances in a
large-scale writing portfolio assessment. (Doctoral Dissertation, Washington State
University).
LeMahieu, P. G., Gitomer, D. H., & Eresh, J. T. (1995). Portfolios in large-scale
assessment: Difficult but not impossible. Educational Measurement: Issues and
practice, 14(3), 11-28. [Link]
Lippi-Green, R. (1997). English with an accent: Language, ideology, and discrimination
in the United States. Routledge.
Lynne, P. (2004). Coming to terms: A theory of writing assessment. Utah University Press.
Moss, P. A. (1995). Themes and variations in validity theory. Educational Measurement:
Issues and Practice, 14(2), 5-13.
Moss, P. A. (1998a). Testing the test of the test: A response to “Multiple inquiry in the
validation of writing tests.” Assessing Writing, 5(1), 111-122.
Moss, P. A. (1998b). The role of consequences in validity theory.
Educational Measurement: Issues and Practice, 17(2), 6-12. [Link]
org/10.1111/j.1745-3992.1998.tb00826.x
Moss, P. A. (2007). Joining the dialogue on validity theory in educational research.
In P. O’Neill (Ed.), Blurring boundaries: Developing writers, researchers, and teachers
(pp. 91-100). Hampton Press.
Moss, P. A & Schutz, A. (2001). Educational standards, assessment and the search for
consensus. American Educational Research Journal, 38(1), 37-70.
Mountford, R. (1999). Let them experiment: Accommodating diverse discourse
practices in large-scale writing assessment. In C. R. Cooper & L. Odell (Eds.),
Evaluating writing: The role of teachers’ knowledge about text, learning, and culture
(pp. 366-396). National Council of Teachers of English.
Murphy, S. (2007). Culture and consequences: The canaries in the coal mine. Research
in the Teaching of English, 42(2), 228-244.
Omi, M. & Winant, H. (1994). Racial formation in the United States: From the 1960’s
to the 1990’s. Routledge.
O’Neill, P. (2003). Moving beyond holistic scoring through validity inquiry. Journal of
Writing Assessment, 1(1), 47-65. [Link]
89
Kelly-Riley
Osgood, C. E., Suci, G. J., & Tannenbaum, P. (1957). The Measurement of meaning.
University of Illinois Press.
Piche, G. L., Rubin, D. L., Turner, L. J. & Michlin, M. L. (1978). Teachers’ subjective
evaluations of standard and Black nonstandard English compositions: A study of
written language and attitudes. Research in the Teaching of English, 12(2), 107-118.
Rubin, D. L., & Williams-James, M. (1997). The impact of writer nationality on
mainstream teachers’ judgments of composition quality. Journal of Second Language
Writing, 6(2), 139-154.
Scharton, M. (1996). The politics of validity. In E. M. White, W. D. Lutz & S.
Kamusikiri (Eds.), Assessment of writing: Politics, policies, practices (pp. 53-75).
Modern Language Association.
Schmidt, A. E. & Camara, W. J. (2004). Group differences in standardized test scores
and other educational indicators. In R. Zwick (Ed.), Rethinking the SAT: The future
of standardized testing in university admissions (pp. 189-201). Routledge Falmer.
Shavelson, R. J. (1996). Statistical reasoning for the behavioral sciences (3rd ed.). Allyn
and Bacon.
Smith, W. L. (1993). Assessing the reliability and adequacy of using holistic scoring
of essays as a college composition placement technique. In M. M. Williamson and
B. A. Huot (Eds.), Validating holistic scoring for writing assessment: Theoretical and
empirical foundations (pp. 142-205). Hampton Press.
Smitherman, G. & Villanueva, V. (Eds.). (2003). Language diversity in the classroom:
from intention to practice. Southern Illinois University Press.
Tabachnick, B. & Fidell, L. S. (2006). Using multivariate statistics (5th ed.). Allyn and Bacon.
White, E. M. (2005). The scoring of writing portfolios: Phase 2. College Composition
and Communication, 56(4), 581-600.
Williams, F., Whitehead, J. L. & Miller, L. M. (1971). Attitudinal correlates of children’s
speech characteristics (USOE Project No. 0-0336). Center for Communication Research.
Williamson, M. M. & Huot, B. A. (1993). Validating holistic scoring for writing
assessment: Theoretical and empirical foundations. Hampton Press.
2. Focus
Unclear 1 2 3 4 5 6 Clear
90
Validity Inquiry of Race and Shared Evaluation Practices
3. Organization
Disorga- 1 2 3 4 5 6 Orga-
nized nized
4. Support
Not provided 1 2 3 4 5 6 Provided
5. Mechanics
Not Effective 1 2 3 4 5 6 Effective
9. Ungrammatical 1 2 3 4 5 6 Grammatical
[Link] 1 2 3 4 5 6 Imaginative
91
CHAPTER 3.
THREE INTERPRETATIVE
FRAMEWORKS: ASSESSMENT
OF ENGLISH LANGUAGE ARTS-
WRITING IN THE COMMON CORE
STATE STANDARDS INITIATIVE
Norbert Elliot
New Jersey Institute of Technology
Andre A. Rupp
Educational Testing Service
David M. Williamson
Educational Testing Service
DOI: [Link] 93
Elliot, Rupp, and Williamson
By the end of the nineteenth century in the United States, demand for universal
public education had become equated with assurance of participatory democ-
racy. In 1869-1870, 7.48 million students enrolled in kindergarten and grades
one through eight. By 1899-1900, that number had risen to 14.98 million.
This increase was accompanied by a dramatic rise in high school enrollment
as advanced education became necessary for better paying jobs. In 1869-1870,
80,000 students were enrolled in grades nine through twelve. In 1899-1900,
that number had risen to 519,000 (Snyder, 1993, p. 34, Table 8).
Accompanying this new influx of students were those who believed they
knew best how to shape the curriculum. Archetypal responses—the humanism
of Charles W. Eliot (1892), the developmentalism of G. Stanley Hall (1883),
the social efficiency of Joseph Mayer Rice (1893), and the social meliorism of
Lester Frank Ward (1883)—were to continue throughout the twentieth century
(Kliebard, 2004). Today, one may identify these enduring themes in the calls for
equity by Diane Ravitch (2010), the cognitive modeling of Howard Gardner
(2006), the emphasis on effective teaching by Bill and Melinda Gates (2015),
and the progressivist agenda of Arne Duncan (2015).
With enrollment projections for the school year 2015-2016 estimated at
49.8 million public elementary and secondary school students (Snyder & Dil-
low, 2015, p. 86, Table 203.10), these and other voices emerge to give council
on how best to spend a projected education budget of no less than $669 billion
(Snyder & Dillow, 2015, p. 58, Table 106.10). There is a loud roar of voices ac-
companying initiatives associated with the term “educational reform,” which has
become nearly deafening as the national debate has turned to the Common Core
State Standards Initiative (CCSSI) and associated state-led curricular guidelines
for a national school curriculum assessed by two consortia: the Smarter Balanced
Assessment Consortium (Smarter Balanced) and the Partnership for Assessment
of Readiness for College and Careers (PARCC).
As the most comprehensive effort in American history to leverage uniform
goal-based instruction, the CCSSI is designed to ensure that high school grad-
uates are prepared to take credit-bearing courses in two- or four-year college
programs or enter the workforce. At the present writing, forty-two states, the
District of Columbia, four territories, and the Department of Defense Educa-
tion Activity have adopted the CCSSI. Assessments in English language arts and
mathematics have taken place in the 2014-2015 school year, and preliminary
results are being released at the time of this writing.
94
Three Interpretative Frameworks
The development of the CCSSI and its assessment has been accompanied by
three categories of criticism: warnings of the dangers of neoliberalism; concerns
over the constraint of the writing construct; and fears that the achievement of
equity continues to elude educational reform. From their creation (in order to
enhance global competitiveness and workplace success) to their solicitation (in
order to encourage proposals for next generation assessment systems), the CCS-
SI have been informed by “a form of cultural politics and a set of economic
principles, policies, and practices devoted to handing over as much of social life
as possible to private interests” (Gallagher, 2011, p. 453).
Referencing this depiction of neoliberalism, Wilson has been critical of the
ways that such framing has diminished teacher agency (Shannon, Whitney, &
Wilson, 2014). In interacting with students and teachers, she argued, “you see
what matters, and you realize that these grand plans that Bill Gates has for how
it is that we’re going to improve education just don’t make any sense” (p. 299).
In similar fashion, Addison and McGee (2015) warned that the role of the Gates
Foundation compromises local efforts such as those sponsored by the National
Writing Project, “to gain compliance” with the CCSSI (p. 215). Concentrating on
the limits of construct representation following from the neoliberal policy climate,
Kristine Johnson (2015) found curricula based on the CCSSI “would focus almost
exclusively on expository/informational and fact-based argumentative writing,
with some narrative descriptive writing”—a “narrowing effect” that diminishes
coverage of the writing construct (p. 520). Applebee (2013) has also identified this
narrowing effect in his identification of four areas—separate emphasis on founda-
tional skills, grade-by-grade standards, absence of a developmental writing model,
and implementation issues—with “equal potential to distort curriculum” (p. 28).
While public debate swirls around societal impact, often absent are voices
of stakeholder groups directly involved with students: parents and guardians;
teachers and administrators; legislators; and workforce leaders. It is our aim in
this paper to suggest directions of inquiry for those stakeholder groups. Specif-
ically, we seek to empower these stakeholder groups by discussing how a deeper
understanding of the traditions, terminologies, and best practices of educational
measurement and writing assessment provide an excellent way to ask critical
questions about new curriculum and assessment initiatives.
Such strategies are needed to navigate a maze of complex debates in which
everything and its opposite both appear to be true. As researchers in writing
assessment (Elliot), cognitively-grounded diagnostic measurement (Rupp), as
well as automated scoring and modern psychometrics (Williamson), we are po-
sitioned to enter the controversial roar in a very precise way.
While we acknowledge and honor the ontological and axiological force of
voices interested in the social dimension of assessment, we focus in this paper
95
Elliot, Rupp, and Williamson
INTERPRETATIVE FRAMEWORK 1:
MULTIDISCIPLINARY RESEARCH ON WRITING
Part of the discipline of education, the field of educational measurement finds
its origin in 1892 with the founding of the American Psychological Association
(Fernberger, 1932) and the subsequent 1945 designation of Division 5, Eval-
uation and Measurement (Benjamin, 1997). Part of the discipline of English
language and literature, the field of writing assessment finds its origin with the
founding of the National Council of Teachers of English in 1911 (Lindemann,
2010) and the 2010 designation of Rhetoric and Composition/Writing Studies
as its own specialized field (Phelps & Ackerman, 2010).
Recent multidisciplinary research between educational measurement and
writing assessment has addressed the present landscape of writing assessment,
as well as methodology, consequence, and future directions for the field (Elliot
& Perelman, 2012). Clearly, the two fields have begun to influence each other;
the acknowledgment of mutually beneficial research agendas, for instance, has
resulted in recommendations for next-generation assessments to focus on social
and rhetorical knowledge, domain knowledge and conceptual strategies, writ-
ing processes, and knowledge of conventions (Sparks, Song, Brantley, & Liu,
2014). Such a multidisciplinary perspective provides a way to frame the CCSSI
96
Three Interpretative Frameworks
CONSTRUCT DEFINITION
A construct such as writing, which is the core focus of the definition and empir-
ical representation of models of student competence for CCSSI ELA-W assess-
ment, is generally defined rather broadly. Its description, however, should be as
concrete, comprehensive, and systemic as possible to be useful for instructional
guidance and assessment development. The operationalization of the way the
construct is measured through assessment tasks and their associated scoring rules
is a great leverage point for obtaining clarity about the boundaries of the con-
struct definition as targeted in an assessment.
Beginning with the protocol analyses of Flower and Hayes (1981), writing
has been understood as a complex process in which readers and writers con-
struct meaning through detailed, often internal, cognitive iterations concerning
variables such as discourse conventions, social context, language, purpose, and
knowledge. In negotiating meaning, writers create “webs of intention, carrying
out complex, individual, and socially bounded purposes, shaped by attitudes
and feelings, and other people” (Flower, 1994, p. 54). In recent iterations of the
model, attention has been drawn to the importance of source-based investiga-
tion, the design of visual content, and management of attention and motivation
(Hayes, 2012; Leijten, Van Waes, Schriver, & Hayes, 2014). As evidence of their
enduring presence, Beringer (2012) has documented the origin, traditions, and
future directions of cognitive perspectives on writing research. Based on con-
struct models derived from these perspectives, Deane and his colleagues (2015)
have recently developed a key practice framework linking ECD, scenario-based
assessment, and cognitively-based assessment in order to create English Lan-
guage Arts task sequences that support both instruction and assessment. So-
cial cognitive models are understood to yield high quality, specific information
about both the writing construct and its boundaries.
Informed by models of social cognition, the CCSSI ELA-W is designed
to specify performance-level objectives—knowledge descriptions that can be
mapped to grade levels. By these strategies, the CCSSI ELA-W models writing
from kindergarten through grade 12. That is, in the CCSSI ELA-W, the con-
struct is defined in actionable terms: “Students should demonstrate increasing
sophistication in all aspects of language use, from vocabulary and syntax to the
development and organization of ideas, and they should address increasingly
demanding content and sources” (CCSSI, 2015c). By extension, writing is also
viewed as part of the broader construct of ELA:
97
Elliot, Rupp, and Williamson
CONSTRUCT MEASUREMENT
While the CCSSI ELA-W is research-based, it is important to understand that
the conceptual model—the way the elements of writing are understood in their
relationship to each other within the given construct—was based on consensus
opinion. Distinct from construct definitions based on evidence from reflective
latent variable models (Graham, McKeown, Kiuhara, & Harris, 2012; Graham
& Perin, 2007; Hillocks, 1986; Rogers & Graham, 2008), this consensus defi-
nition is, in reality, a “stew” of elements that might or might not be empirically
related to each other (National Research Council, 2012). Put differently, as a
consensus model, the development and instantiation of the CCSSI has, so far,
been a state-led effort based on adoption, not on data collection. The means of
assessing students and the information resulting from that assessment are left to
98
Three Interpretative Frameworks
the discretion of the states as an activity distinct from the CCSSI—a very com-
plex task for individual states and collections of states.
The era of modern assessment has arguably been characterized by a focus
on creating writing tasks that are closely aligned with modern views of writing
from expert communities. In fact, without this involvement of the writing com-
munity it would be difficult to imagine how this new generation of assessment
would be different than the print-born bubble and booklet tests of the past. This
involvement has led to the use of digitally-delivered stand-alone writing tasks
and the embedding of writing activities in domain or profession-specific com-
plex performance tasks (Tucker, 2009). Designed to capture blended constructs,
integrated tasks incorporating content from source materials offer benefits such
as providing realistic, challenging activities, engaging students in writing respon-
sible to specific content, obviating practice effects associated with conventional
item types, evaluating language abilities consistent with integrated models of
literacy, and offering diagnostic value for instruction or self-assessment.
Challenges nevertheless remain. Cumming (2013) has noted that integrat-
ed writing tasks have associated risks. These include confounding measurement
of writing ability with abilities to comprehend source materials, merging as-
sessment and diagnostic information together in ineffective ways, and invoking
genres that are emerging and therefore difficult to score. As we discuss below,
navigating the complex system of tradeoffs when designing individual assess-
ments and systems of assessments over time for CCSSI ELA-W can be sub-
stantially facilitated, integrated, and scrutinized using the Standards and ECD
frameworks as guidance.
99
Elliot, Rupp, and Williamson
interpretations of test scores for the intended test uses” (p.1). However, while
Standards are designed for raising awareness and guiding decision-making about
assessment systems. However, while Standards are designed for raising awareness
and guiding decision-making about assessment systems at a high conceptual
level, the document is not designed to be step-by-step instructions of how to do
the necessary work on a day-to-day basis. That role falls to principled assessment
design frameworks like ECD, which we discuss in the next section.
Calls for increased assessment literacy such as those found in the Standards
(pp. 192-193) are not incidental to our purpose in this paper. Any fixed set of
curricular approaches or assessment methods yields particular kind of interpre-
tation and any such methodological exclusivity is inappropriate when dealing
with complex assessments such as the CCSSI-ELA-W. In fact, assessment of the
CCSSI-ELA-W is designed to generate the kinds of evidence needed to validate
multiple proposed interpretations and uses.
While the present version of the Standards is our concern here, the 4th re-
vision (1999) was the common referential point for both the Smarter Balanced
and PARCC consortia. Indeed, the five sources of validity evidence identified by
Sireci (2012) in his report of the Smarter Balanced research agenda—a report
to which we will turn later in order to establish the informed view of validity
used to support score interpretation and use (Kane, 2013, 2015) in the design
of the CCSSI ELA-W assessment—are taken directly from the 1999 version.
The Standards have played, and will continue to play, a significant role in the
development of assessments related to the CCSSI.
In their present form, the Standards are divided into three sections: founda-
tions, operations, and applications. By far, the foundations section is the most
significant in terms of assessment of the CCSSI ELA-W. It is here we find extended
discussion of the three overarching principles of validity, reliability/precision, and
fairness. Because these foundational concepts deeply inform Smarter Balanced and
PARCC assessment designs, a brief definition and discussion of each is warrant-
ed. Nevertheless, the concepts are not intended to be separated; rather, validity,
reliability/precision, and fairness are intended to be used in support of proposed
interpretation and use of scores associated with the CCSSI ELA-W assessment.
VALIDITY
In the Standards, validity is defined as the “degree to which accumulated evi-
dence and theory support a specific interpretation of test scores for a given use
of a test. If multiple interpretations of a test score for different uses are intend-
ed, validity evidence of each interpretation is needed” (p. 225). Although still
considered by many as an “up-or-down vote” or a simple “stamp of approval,”
100
Three Interpretative Frameworks
the 2014 edition is clear on the imprecision of such summary judgment: “State-
ments about validity should refer to particular interpretations and consequent
uses. It is incorrect to use the unqualified phrase ‘the validity of the test’” (p. 23).
While the origin of this characterization of validity may be found in the 1985
edition of the Standards, it is important to reflect on just how enduring the work
of Messick (1989) has become in his characterization of validity as “an integrated
evaluative judgement of the degree to which empirical evidence and theoretical
rationales support the adequacy and appropriateness of inferences and actions based
on test scores or other modes of assessment” (p. 13, emphasis in original). Equally
important is the work of Kane (2013) and his call for evidence-based interpreta-
tion and use arguments: “To validate an interpretation or use of test scores is to
evaluate the plausibility of the claims based on the test scores” (p. 1).
Validation therefore requires a clear statement of the claims inherent in the
proposed interpretations and uses of the test scores. “Public claims require pub-
lic justification” (Kane, 2013, p. 1). Influential in the development of the Stan-
dards and their manifestation in the assessment of the CCSSI ELA-W, Kane
(2015) has offered a two-step approach to validation:
First, the interpretation and use is specified as an interpretation/
use argument, which specifies the network of inferences and
assumptions leading from test performances to conclusions and
decisions based on the test scores. Second, the interpretation/
use argument is critically evaluated by a validity argument. (p.
4, emphasis in original). As a result of this orientation, validity
becomes a property of score interpretations—not as a property
of the assessment: “Once we adopt an interpretation, it can
make sense to talk about ‘the validity of a test’, but the ‘validity’
is relative to that interpretation” (Kane, 2015, p. 2).
This “flexible framework for validation,” as Kane terms it, is important in
that it allows for—indeed, encourages—multiple interpretations that may arise
from multiple groups. As Kane concludes, “[T]o restrict our conception of va-
lidity to one kind of interpretation seems unnecessary and would greatly limit
our ability to respond to the varied applications of test scores” (2015, p. 3).
RELIABILITY/PRECISION
Reliability/precision is defined as:
The degree to which test scores of a group of test takers are
consistent over repeated applications of a measurement proce-
101
Elliot, Rupp, and Williamson
102
Three Interpretative Frameworks
FAIRNESS
In the Standards fairness is defined as:
The validity of test score interpretations for intended use(s)
for individuals from all relevant subgroups. A test that is fair
minimizes construct-irrelevant variance associated with indi-
vidual characteristics and testing contexts that otherwise would
compromise the validity of scores for some individuals. (p. 219)
This section of the Standards has been expanded substantially over previous
revisions, with emphasis given to fairness for all examinees. Again, we see the pres-
ence of Messick (1989) who linked forms of validity with consequences related to
score use—an emphasis that has been maintained by Kane (2006, 2013).
Significantly, special attention is given in the Standards to the opportunity
to learn—“the extent to which individuals have had exposure to instruction or
knowledge that affords them the opportunity to learn the content and skills tar-
geted by the test” (p. 56). In an analysis consistent with this emphasis on exposure,
Pullin (2008) has highlighted connections among assessment, equity, and oppor-
tunity to learn, as both a reflection of the learning environment and a concept
demanding articulated connections between the assessment and the instructional
environment. Such characterizations afford identification and removal of barriers
to valid score interpretation for the widest possible range of individuals and sub-
groups, interpretative validity for examined populations, and the development of
suitable testing accommodations and safeguards to protect fair score usage.
Equally associated with fairness—and of special interest in terms of equity to
all stakeholders—is adherence to the principles of universal design. An approach
to assessment that strives to minimize construct distortion and maximize fairness
through uniform access for all intended examinees, universal design has been
identified in the Standards (2014) as a way to leverage fairness for all examinees
(p. 63). As Ketterlin-Geller (2008) has established, when student characteristics
are considered during the conceptualization, design, and implementation phase
of test development under principles of universal design (e.g., specifying content
and cognitive complexity in the test blueprint, as well as information about the
target and access skills), test performance of students with special needs is more
likely to reflect their construct knowledge. Furthermore, Mislevy et al. (2013)
103
Elliot, Rupp, and Williamson
INTERPRETATIVE FRAMEWORK 3:
EVIDENCE-CENTERED DESIGN (ECD)
As the discussions in the previous Standards section have made abundantly clear,
to build an evidentiary argument for assessment scores so that intended inter-
pretations and decisions comply with the Standards is a complex process. This
complex process is exemplified in the CCSSI assessment aim as it is identified by
Smarter Balanced: “The assessment system being developed by the Consortium
is designed to provide comprehensive information about student achievement
104
Three Interpretative Frameworks
that can be used to improve instruction and provide extensive professional de-
velopment for teachers” (Sireci, 2012, p. 4). As such, “the assessment system
focuses on the need to strongly align curriculum, instruction, and assessment, in
a way that provides valuable information to support educational accountability
initiatives” (p. 4).To help facilitate the construction of arguments supporting
such aims and to imbue the assessment ecosystem with appropriate character-
istics that support intended interpretations and decisions, a principled design
framework for practice such as ECD is needed. Proposed to make explicit the
evidentiary reasoning process of assessment interpretation and decision-making,
ECD helps organize assessment practices in ways that yield cohesive integrated
thinking about assessment aims, delivery capability, and justification of score
use. As such, ECD can be viewed as providing the “evidentiary grammar” for
evidence-based assessment arguments.
At its best, ECD is a powerful professional development tool that can help
interdisciplinary teams of experts (e.g., assessment developers, statisticians, in-
formation technology specialists, policy-makers, and other stakeholders) de-
velop common language, mental models, design artifacts, and best practices.
In addition, it can help such teams utilize these capacities to develop targeted
artifacts that move the assessment process forward in ways that best capture
the connected thinking underlying the design process. These goals are always
laudable and important, of course, but become especially important as the as-
sessments become more performance-oriented, more reliant on models of social
cognition, more responsive to correlates such as engagement or motivation, and
more situated within community practices. In short, ECD is highly relevant for
task-based CCSSI assessments of ELA-W.
Mislevy, Sternberg, and Almond (2003) identified five core structural/con-
ceptual elements for ECD and arrange them in what they term the conceptual
assessment framework: student models that characterize knowledge and skill;
task models that provide constructed response test items to elicit student knowl-
edge and skills; evidence models that provide a chain of inferential reasoning
from student test performance to knowledge and skill, with emphasis on scores
and their measurement; assembly models that specify how individual tasks are
combined to produce the final assessment; and presentation models that specify
how individual tasks are administered to students. In practice, spelling out these
different models means creating artifacts such as databases, spreadsheets, and
text files to document the key decisions that underlie the reasoning process.
Thus, a second layer in the day-to-day practice of assessment development
is putting the decisions captured in these artifacts into practice by setting up
a delivery, scoring, and reporting architecture, which Mislevy, Sternberg, and
Almond described as a four-process model of activity selection (the process of
105
Elliot, Rupp, and Williamson
106
Three Interpretative Frameworks
107
Elliot, Rupp, and Williamson
authors have paid close attention to score use—to the ways to compare student
performance across schools, districts and states, to measure growth across grade
levels, and to evaluate year-to-year changes. Because of the importance of such
comparisons and goal setting, the authors emphasized the need for “well-ar-
ticulated, cognitively-based constructs” based on the CCSSI, which should be
developed in order to establish the ordered claims and evidence requirements by
grade level.
Luecht and Camara noted that the ECD approach “may offer some advan-
tages over conventional item design and test specifications because such new
design approaches prioritize more explicit connections between items from task
models which are directly derived from evidence” (p. 15). Task models resulting
from ECD, as the report acknowledges, allow designers to control for content
through an emphasis on cognitive demand and yield greater efficiency in devel-
opment of the assessment over time.
As these three examples demonstrate, strategic use of Standards-based and
ECD frameworks at the planning stage yields a validity agenda and evidentiary
processes. In the next section, we provide some guiding questions for stakehold-
er networks that can help to raise awareness about what it means to translate
the different concepts in the Standards and ECD into thoughtful assessment
practice that supports meaningful interpretations and decisions.
108
Three Interpretative Frameworks
emphasis on networks and their logic proposed by Gallagher (2011); that is, the
questions we provide are intended to provide “analytic tools for understanding
how actors exercise power by virtue of their locations and relations” (p. 466, em-
phasis in original).
109
Elliot, Rupp, and Williamson
design decisions within the teaching and assessment ecosystem. We deeply be-
lieve that it is of value to connect the logic of educational measurement and
writing studies research with the logic of heuristics and biases research, if only
to remind everyone that complex ventures obligate us to think in complex ways.
In each table, we have used the Standards to generate a series of broad foun-
dational and operational questions that, in turn, are made specific by focusing
on specific facets of measurement. Because our focus is on an educational as-
sessment, we have integrated that application into the foundational and oper-
ational question and, hence, no additional table is provided for that section of
the Standards.
110
Three Interpretative Frameworks
111
Elliot, Rupp, and Williamson
Scores, Scales, Norms, If decisions regard- If cut scores have How have the How have the as-
Score Linking, and Cut ing placement and been established, are assessment devel- sessment developers
Scores: progression are to these scores to be opers demonstrated demonstrated that
“Test scores should be be made from the used for descriptive that scores have the norms and cut
derived in a way that assessment, have or decision-making been normed with scores established
supports the interpre- cut scores been purposes? student populations are congruent
tation of test scores for established for • How have assur- similar to those with workforce
the proposed uses of categories of student ances been made found at individual populations and
tests. Test developers performance? that the estab- schools or school employment needs?
and users should If cut scores have lishment of cut districts? • How have inter-
document evidence been established, scores does not • How have pretations been
of fairness, reliability, has the procedure undermine the differentiated established to
and validity of test been documented validity of score norms been help employers
scores for their pro- and communicated interpretations? established for interpret and use
posed use” (p. 102). in terms of both different gender, the established
technical specifi- race/ethnicity, norms and cut
cations and policy language, disabil- scores?
decisions? ity, economically
disadvantages,
grade, and age
groups?
Test Administration, How have the as- Because different How have resources How have test
Scoring, Reporting, and sessment developers stakeholder groups been leveraged to administration,
Interpretation: designed the digital may administer, ensure that the scoring, reporting,
“To support useful in- administration score, report, diverse stakeholder and interpretation
terpretations of score so that technical and interpret the groups needed to processes been
results, assessment disruptions do assessment, how administer, score, designed so that
instruments should not contribute to have procedures report, and interpret scores can be
have established construct-irrelevant been established to the assessment used to establish
procedures for test variance? ensure that score have the compe- connections with
administration, Have distinctions interpretation and tency and resources workplace needs?
scoring, reporting, been made between use are not compro- necessary to ensure How have
and interpretation. accommodations mised by failure of standardization? standardization
Those responsible for test takers based standardization? In cases of students processes resulted
for administering, on need and accom- How have assess- with disabilities or in the anticipation
scoring, reporting, modations based ment developers different language and removal of
and interpreting on misalignment demonstrated that backgrounds, how construct-irrelevant
should have sufficient between the digital- standardization will have nonstan- variance so that
training and supports ly-based assessment ensure that students dard models been scores from the
to help them follow and the print-based have the same ability established that will assessment can be
the established pro- curriculum? to demonstrate their allow these students used on a long-time
cedures. Adherence competency? to demonstrate basis?
to the established competence?
procedures should be
monitored, and any
material errors should
be documented and,
if possible, corrected”
(p. 114).
112
Three Interpretative Frameworks
Rights and Responsibil- How has the student How has the In order to protect If assessment scores
ities of Test Takers: been provided instructor provided students from are to be used to de-
“Test takers have the with accurate, free students with infor- potentially adverse termine workplace
right to adequate information about mation about the consequences, how competency, how
information to help the assessment? assessment, intend- has the legislative have assurances be
them prepare for a • As a means of ed score use, scoring process been used to established to assure
test so that the test reducing con- criteria, administra- delay justified score that students have
results accurately struct-irrelevant tive policy, available use? information about
reflect their standing variance, how of accommodations, • If the legislative how employers are
on the construct being has the student and confidentiality? process has been using scores?
assessed and lead to been provided • How have the used to delay • If assessment
fair and accurate score with practice students been score use, how scores are to be
interpretations. They access to the informed of have specific transferred to
also have the right to digital environ- their rights and determinations employers, how
protection of their ment in which the rights of been made have the data
personally identified the test will be their parents to regarding a range systems be de-
score results from un- administered? access assessment of decisions and a signed to assure
authorized access, use, results and be timeline for justi- confidentiality?
or disclosure. Further, protected from fied score use?
test takers have the re- unauthorized use
sponsibility to present of results?
themselves accurately
in the testing process
and to respect copy-
right in test materials”
(p. 133).
113
Elliot, Rupp, and Williamson
114
Three Interpretative Frameworks
During the imagined town meeting, attention might be drawn to the Smart-
er Balanced Consortium (2014b) document entitled “Interpretation and Use of
Scores and Achievement Levels” that we discussed in the previous section. Re-
call that scale scores and achievement level descriptors are identified in alignment
with the Standards in the document. Using the validity questions from Table 3.1,
teachers and administrators might focus on discussing the relationship between
test results and the curriculum in their classrooms, schools, and districts. Choic-
es in test design, administration, and reporting become critical as questions are
raised regarding the constructive alignment—the integrated instructional and as-
sessment systems and efforts used to map learning activities to outcomes (Biggs
& Tang, 2011)—that must be established among the individual student’s school,
the CCSSI ELA-W, and Smarter Balanced and PARCC assessments. Critically
discussing the implications of various decisions based on questions around con-
structive alignment would help establish a common understanding of the extent to
which the scores are faithful demonstrations of individual student ability.
Similarly, in using the questions to investigate sources of evidence related
to reliability, teachers and administrators would benefit by paying attention to
the concept of measurement precision and not just an overly simplistic single
descriptive statistic (Sireci, 2012). Estimates of score reliability (internal consis-
tency) and those based on examining students more than once (parallel forms)
thus become important sources of information to consider when determining
appropriate and less appropriate interpretations of scores.
For students, guardians, teachers, and administrators, questions of what con-
stitutes appropriate score interpretation and use would be especially relevant
in light of the disaggregated information about student performance obtained
from the Smarter Balanced field test that was administered between March and
June 2014 (Smarter Balanced, 2014a). The test revealed clear performance dif-
ferences among key student subgroups that allow for a critical discussion of how
these differences are related to potential differences in opportunities to learn.
Specifically, at the Grade 11 level, 40.9 percent of total students examined (n
= 31,018) met the cut score of Level 3 (or above) in achievement levels ranging
from Level 1 (novice) to Level 4 (advanced). Among American Indian/Alas-
kan Native students (n = 777), 26.6 percent passed; Asian students (n = 2,334)
passed at 54.1 percent; Black/African American students (n = 2,552) passed at
21.2 percent; Hispanic/Latino students (n = 10,041) passed at 32.4 percent; Na-
tive Hawaiian/Other Pacific Islander students (n = 195) passed at 32.8 percent;
White/Caucasian students (n = 16,020) passed at 46.2 percent; Multi-ethnic/
Multi-racial students (n = 889) passed at 45.1 percent. Among those enrolled in
an Individualized Education Program (n = 2,084), 9.0 percent passed; among
those classified as Limited English Proficient/English language learners (n =
115
Elliot, Rupp, and Williamson
1,767), 5.7 percent passed; among those classified under special program en-
rollment preventing discrimination based on disability (n = 366), 36.1 percent
passed; among those classified as Economically Disadvantaged students (n =
13,962), 32.6 percent passed (Smarter Balanced, 2014a, p. 12).
The literature associated with opportunity to learn is a particularly rich
framework for advancing instructional equity among student groups (Moss,
Pullin, Gee, Haertel, & Young, 2008). In terms of the fairness questions raised
in Table 3.1, using scores as a way to promote opportunity to learn can help in
identification of barriers to success and creation opportunities to foster educa-
tional advancement. Making Standards-based conceptual and empirical connec-
tions among issues around validity, reliability/precision, and fairness through
the lens of opportunity to learn is, we believe, an especially powerful logic that
can be used to guide discussion of assessment results.
Because the continuum among school, college, and workplace writing ap-
pears to exhibit more disjuncture than congruence (Burstein, Elliot, & Molloy,
in press; Melzer, 2014), Table 3.2 might be used to call attention to the espe-
cially difficult generalization inference between academic and workplace writing
established by the CCSSI ELA-W. Because the CCSSI specifically identifies both
academic and workplace readiness, it is reasonable for post-secondary academic
and workplace leaders to ask questions that allow them to obtain more clarity on
critical assessment design, delivery, and scoring decision. Moreover, it is import-
ant that the ensuing discussions are used to elucidate any remaining ambiguities
around how performance certification decisions should be informed by scores
from CCSSI ELA-W assessments. In terms of the report “Interpretation and Use
of Scores and Achievement Levels” that we discussed in the previous section,
questions of score use become especially important in light of the fact that paral-
lel operational definitions and frameworks are still under development for career
readiness (Smarter Balanced Consortium, 2014b, p. 2). Present at the imagined
town meeting, academic and workplace leaders could certainly highlight issues
regarding the learning continuum.
Legislators will want to attend to both the intended and unintended conse-
quence of the CCSS ELA-W in terms of validity evidence and factors external
to the assessment. Determination of score use is especially important in the case
of value-added methods used to make inferences about teacher performance,
especially when current research reveals that the scores resulting from such pro-
cedures may be systematically biased in favor of some instructors and against
others (Haertel, 2013). In anticipating legal issues associate with CCSSI ELA-W
assessment, stakeholders will find the empirical techniques associated with quan-
tifying disparate impact equally useful (Poe, Elliot, Cogan, & Nurudeen, 2014)
so that they can meaningfully help to advance opportunities to learn.
116
Three Interpretative Frameworks
CONCLUSION
As these examples from our town hall thought experiment illustrate, while the
questions in Table 3.1 and Table 3.2 are not meant to be exhaustive, they might
prove useful for three reasons. First, because their phrasing is informed by the
program of research begun by Tversky and Kahneman (2011), it is possible
that such questions might act as a bridge between the kinds of evidence-based,
argumentative logic that assessment designers employ in ECD (Mislevy, Stein-
berg, & Almond, 2003) and the availability, representativeness, and adjustment
involved in heuristic reasoning that other assessment stakeholders may use in
decision-making (Gilovich & Griffin, 2002). Bridging the logic of the assess-
ment developer and the logic of the assessment user is a worthy goal that might
be served by attention to decision-making under uncertainty. Tables 3.1 and
3.2 contribute to our desire to help stakeholders ask principled questions about
assessment design, score use, and consequences. Second, attention to diverse
reasoning processes is inherent in the social cognitive view of writing that in-
forms the CCSSI ELW-W and its assessment. As Gilovich & Griffin (2002)
have observed, the heuristic reasoning program fits well with our present un-
derstanding of how the mind works. Third, the imagined town meeting as the
forum for deliberative discussion suggests the need for the development of what
Rawls (2001) has referred to as overlapping consensus. The aim of reasonable
pluralism is a worthy goal that may be achieved if common referential frames are
established of the kinds we have suggested here.
The concepts we have presented in this paper are complex, and the chal-
lenges we have identified are real and must be addressed. We believe that our
collective logic can be guided by interpretative frameworks such as the three pre-
sented here that speak to core issues associated with advancement of opportunity
to learn. As present curricular and assessment innovations merge to produce
information about student performance, many questions nevertheless remain.
Especially notable are questions regarding the relationship between assessment
and opportunity structure. Future work must turn to questions left unanswered
here.
REFERENCES
Addison, J., & McGee, S. J. (2015). To the core: College composition classrooms in
the age of accountability, standardized testing, and Common Core State Standards.
Rhetoric Review, 34(2), 200-218.
Almond, R. G., Steinberg, L. S., & Mislevy, R. G. (2002). A four process architecture
for assessment delivery, with connections to assessment design. Educational Testing
Service. [Link]
117
Elliot, Rupp, and Williamson
118
Three Interpretative Frameworks
Flower, L., & Hayes, J. R. (1981). A cognitive process theory of writing. College
Composition and Communication, 32(4), 365–387.
Gallagher, C. W. (2011). Being there: (Re)making the assessment scene. College
Composition and Communication, 63(3), 450-476.
Gardner, H. (2006). Five minds for the future: Leadership for the common good. Harvard
Business School Press.
Gates, B., & Gates, M. (2015). College-ready education. [Link]
org/What-We-Do/US-Program/College-Ready-Education
Gilovich T. & Griffin, D. (2002). Introduction—Heuristics and biases: Then and
now. In T. Gilovich, D. Griffin, & D, Kahneman (Eds.), Heuristics and biases: The
psychology of intuitive judgment (pp. 1-18). Cambridge University Press.
Graham, S. (2006). Writing. In P. Alexander & P. Winne (Eds.), Handbook of
educational psychology (pp. 457-478). Erlbaum.
Graham, S., McKeown, D., Kiuhara, S. A., Harris, K. R. (2012). A meta-analysis of
writing instruction for students in the elementary grades. Journal of Educational
Psychology, 104, 879-896.
Graham, S., & Perin, D. (2007). A meta-analysis of writing instruction for adolescent
students. Journal of Educational Psychology, 99(3), 445-476.
Hall, G. S. (1883). The contents of children’s minds. Princeton Review, 11, 249-272.
Haertel, E. H. (2013). Reliability and validity of inferences about teachers based on
student test scores. [William H. Angoff 14th memorial lecture]. National Press
Club, Washington, DC. [Link]
Hayes, J. R. (2012). Modeling and remodeling writing. Written Communication, 29(3),
369-388.
Hillocks, G. (1986). Research on written composition: New directions for teaching.
National Council of Teachers of English.
Johnson, K. (2013). Beyond standards: Disciplinary and national perspectives on
habits of mind, College Composition and Communication, 64(3), 517-541.
Kahneman, D. (1973). Attention and effort. Prentice-Hall.
Kahneman, D. (2011). Thinking fast and slow. Farrar, Straus and Giroux.
Kane, M. T. (2006). Validation. In R. L. Brennan (Ed.), Educational measurement (4th
ed.). (pp. 17-64). American Council on Education and Praeger.
Kane, M. T. (2013). Validating the interpretation and uses of test scores. Journal of
Educational Measurement, 50(1), 1-73.
Kane, M. T. (2015). Explicating validity. Assessment in Education: Principles, Policy &
Practice, 23(2), 198-211.
Ketterlin-Geller, L. R. (2008). Testing student with special needs: A model for under-
standing the interaction between assessment and student characteristics in a universal-
ly designed environment. Educational Measurement: Issues and Practice, 27(3), 3-16.
Kliebard, H. M. (2004). The struggle for the American curriculum, 1893-1958 (3rd ed.).
Routledge.
Leijten, M., Van Waes L., Schriver, K., & Hayes, J. R. (2014). Writing in the
workplace: Constructing documents using multiple digital sources. Journal of
Writing Research, 5(3), 285-337.
119
Elliot, Rupp, and Williamson
Lindemann, E. (Ed.). (2010). Reading the past, writing the future: A century of American
literacy education and the National Council of Teachers of English. National Council
of Teachers of English.
Lord, F. M. (1980). Applications of item response theory to practical testing problems.
Erlbaum.
Luecht, R. M., & Camara, W. J. (2011). Evidence and design implications required
to support compatibility claims. [Link]
[Link]
Melzer, D. (2014). Assignments across the curriculum: A national study of college writing.
Utah State University Press.
Messick, S. (1989). Validity. In R. L. Linn (Ed.), Educational measurement (3rd ed.) (pp.
13-103). American Council on Education and Macmillan.
Mislevy, R. J. (2007). Validity by design. Educational Researcher, 36(8), 463-469.
Mislevy, R. J. (2008). How cognitive science challenges the educational measurement
tradition. [Link]
Mislevy, R. J., Almond, R. G., & Lukas, J. F. (2004). A brief introduction to evidence-
centered design (ETS Research Report-03-16). Educational Testing Service. https://
[Link]/Media/Research/pdf/[Link]
Mislevy, R. J., Haertel, G., Cheng, B., Ructtinger, L., DeBarger, A., Murray, E.,
Rose, D., Gravel, J., Colker, A. M., Rustein, D., & Vendlinski. T. (2013). A
“conditional” sense of fairness in assessment. Educational Research and Evaluation:
An International Journal on Theory and Practice, 19(2-3), 121-140.
Mislevy, R. J., Steinberg, L. S., & Almond, R. G. (2003). On the structure of
educational assessment (with discussion). Measurement: Interdisciplinary Research
and Perspective, 1, 3-62.
Moss, P. A., Pullin, D. C., Gee, J. P., Haertel, E. H. & Young, L. J. (Eds.), (2008).
Assessment, equity, and opportunity to learn. Cambridge University Press.
National Research Council. (2012). Education for life and work: Developing transferable
knowledge and skills in the 21st century. Committee on Defining Deeper Learning
and 21st Century Skills, J. W. Pellegrino, & M. L. Hilton (Eds.). Board on Testing
and Assessment and Board on Science Education, Division of Behavioral and Social
Sciences and Education. The National Academies Press.
Phelps, L. W., & Ackerman, J. M. (2010). Making the case for disciplinarity
in rhetoric, composition, and writing studies: The visibility project. College
Composition and Communication, 18(1), 180-215.
Poe, M., Elliot, N., Cogan, J. A., & Nurudeen, T. G. (2014). The legal and the
local: Using disparate impact analysis to understand the consequences of writing
assessment. College Composition and Communication, 65(4), 588-611.
Pullin, D. C. (2008). Assessment, equity, and opportunity to learn. In P. A. Moss, D.
C. Pullin, J. P. Gee, E. H. Haertel, & L. J. Young (Eds.), Assessment, equity, and
opportunity to learn (pp. 333-351). Cambridge University Press.
Ravitch, D. (2010). The death and life of the great American school system: How testing
and choice are undermining education. Basic Books.
Rawls, J. (2001). Justice as fairness: A restatement. Harvard University Press.
120
Three Interpretative Frameworks
121
PART 2.
Carolyn Calhoon-Dillahunt
Yakima Valley College
For more than 150 years, standardized testing has been a part of the U.S. ed-
ucation system. Almost from the outset, standardized testing was inextricably
linked to writing assessment and, thus, to writing instruction and, ultimately, to
writing as a discipline. Early concerns about the “problem” of student writing
revealed by standardized assessments resulted in increased attention to writing
and writing instruction for teachers, for schools, and, eventually, for policymak-
ers. As a result, for good and bad, writing (granted, often defined and assessed
in reductive ways) holds a position of primacy in assessment and in educational
policy, a position that garners attention and resources, but also scrutiny and
intrusion.
In this section introduction, I briefly trace the history of large-scale writing
assessment and how it has been entwined with politics and policymaking, sit-
uating the specific essays featured in Part Two of this collection in the “reform
and accountability era” of large-scale standardized testing. From there, I discuss
core themes around which these distinct articles coalesce: the policy intentions
for and resulting uses and misuses of large-scale writing assessment in the 2000s;
the consequences of mandated writing standards and high stakes writing assess-
ments on curriculum, teachers and teaching, and students; and the possibilities
enabled through some large-scale writing assessments.
placement into “remedial” writing coursework, starting with Ivy League schools
in the late 1800s (Haswell, 2004). These early forays into writing assessment as
gatekeeping planted the seeds of both basic writing and near universal first-year
writing requirements in postsecondary study.
By the turn of the century, the founding of the College Entrance Examina-
tion Board meant that admissions testing and thus writing assessment became
“outsourced,” and assessment became a professional, prolific, and profitable in-
dustry separate from the institutions that relied on their results (Huot, O’Neill,
& Moore, 2010). In Before Shaughnessey, Ritter (2009) observes that accessibility
of higher education, increasingly available to the masses after WWI and even
more so with the GI Bill post-WWII, shifted the focus of writing assessment.
Writing assessment became preoccupied with surface-level correctness, and re-
mediation was prescribed to resolve students’ perceived lack of preparation for
college-level writing. Over the course of the 20th century, writing placement
also became increasingly disconnected from writing curriculum, as many in-
stitutions, especially open-admissions institutions, shifted from locally scored
timed writing exams to externally scored standardized indirect writing assess-
ments (Haswell, 2004).
In the 20th century, standardized testing expanded to assess proficiency, apti-
tude, intelligence, and more. However, according to Rosales and Walker (2021),
“since their inception almost a century ago, the tests have been instruments
of racism and a biased system,” founded on the pseudo-science, eugenics, and
grounded in white racial habitus (Inoue, 2015). Nowhere is this racism more
apparent than in standardized writing assessments. The purposes for such testing
grew beyond simple gatekeeping for university admissions to diagnosing deficits,
measuring skill sets, and predicting future performance. As a result, standardized
testing was increasingly tied to educational decision-making (NEA, 2020), with
the results of a single measure–generally an indirect measure embedded in White
language and culture supremacy–being used to classify, rank, track, and exclude
students. These approaches disproportionately affected historically underserved
students, particularly students of color. Political support of large-scale testing as
an important educational tool was sealed with the passage of the 1965 Elemen-
tary and Secondary Education Act. The first national assessment, the National
Assessment of Academic Progress (NAEP), addressed in Applebee’s article in this
section, was administered in 1969.
In the later 20th century, alarming reports of an impending literacy crisis, a
crisis of “mediocracy,” and its implications for the U.S. economy, such as News-
week’s ”Why Johnny Can’t Write” (Sheils, 1975), A Nation at Risk (National
Commission on Excellence in Education, 1983), and Time for Results (National
Governors Association, 1985), led to calls for reform and accountability. These
126
The Politics and Public Policy of Large-Scale Writing Assessment
calls for action resulted in a range of state-level policy solutions. One common
action was increased implementation of statewide standards and assessment of
students, from elementary to secondary-level, often in the form of direct assess-
ments of writing and other basic skills. These standards and assessments were
designed to impact curriculum and instruction and frequently were developed
in response to employer demands, but with the influence of disciplinary ex-
perts. For instance, Sandra Murphy’s (2003) Journal of Writing Assessment article,
“That Was Then, This Is Now: The Impact of Changing Assessment Policies on
Teachers and the Teaching of Writing in California,” describes the California
Assessment Program. This program developed in the early-1980s and was re-
garded as cutting edge for its focus on direct writing assessment. Murphy (2003)
notes that half the states also were conducting direct writing assessments by the
mid-1980s.
The essays in this section were published during a new era of large-scale
assessment focused on educational “accountability.” These approaches assumed
test scores and high stakes could be used to raise standards. Literacy and writing
remained key areas of concern and focus. By the late 1990s, many legislatures
were moving toward holding schools and teachers accountable for improving
students’ performance on state-delineated standards, such as California’s 1999
Public Schools Accountability Act (Murphy, 2003); however, the passage of
the No Child Left Behind Act (NCLB) of 2001 mandated regular state-wide
standardized testing coupled with financial performance-based penalties and re-
wards to push educational reform.
Although problems with one-dimensional accountability–and accountability
resting entirely on the test scores of “hapless students” (White, 2005, p. 148)—
were evident early on in K-12 education, this high stakes, testing-centered ap-
proach to educational accountability quickly “trickled up” to higher education.
The 2006 Spellings Commission Report, which called for improving “accessibil-
ity, affordability, and accountability” in higher education, resulted in the 2008
Higher Education Opportunity Act, ushering in a wave of new accountability
measures, increased federal regulation and data reporting requirements and a
greater federal oversight role in institutional accreditation (Eaton, 2008).
Under the Obama administration, the accountability movement accelerated
and increasingly gravitated toward the neoliberal economic policies of “paying
for performance,” what Toth, Sullivan, and Calhoon-Dillahunt (2016) describe
as “a dubious method of improving educational outcomes through financial
penalties and rewards already well-tested (and failing) in K-12 reform efforts”
(p. 392). In elementary and secondary education, Race to the Top competi-
tive grants, funded through the 2009 American Recovery and Reinvestment
Act, helped propel states toward adopting the newly-minted Common Core
127
Calhoon-Dillahunt
State Standards (CCSS), which Hammond and Garcia (2017) studied in their
piece in this section. The English Language Arts and Mathematics Common
Core, initiated by the Council of Chief State School Officers (CCSSO) and
the National Governors Association (NGA) with the support of Achieve, Inc.,
were taken up by nearly every state (CCSS Initiative, 2022), often alongside the
PARCC or Smarter Balanced online tests designed to measure these standards.
According to Adler-Kassner (2017), this accountability age has been driv-
en by increased external influence on educational standards and outcomes by
lawmakers, influential corporations, and many groups and actors that make up
the reform-minded Educational Industrial Complex (EIC), who tell the story
of “The Problem with American Education and How to Fix It” (p. 320). Toth
et al. (2019) note, “over the last few decades, calls among both state and federal
policymakers to improve student retention and degree completion have increas-
ingly been framed as a matter of institutional ‘accountability’” (p. 2). Accord-
ing to Calhoon-Dillahunt (2018), the EIC’s solutions “privilege proficiency and
efficiency (aka ‘success’ and ‘completion’) over learning and development” and
their view of ‘accountability’ is market-oriented, with ‘value’ measured almost
exclusively in economic terms” (p. 281). As a result, developmental and first-
year writing are primary targets in “the EIC’s quest to streamline and economize
higher education” (Calhoon-Dillahunt, 2018, p. 281). In the past decade, some
states–Florida and Connecticut, for instance–have intruded into policies that
were once institutionally determined, such as placement and developmental ed-
ucation, and most states have enacted performance-based funding policies in an
attempt to drive reform.
ACCOUNTABILITY CONSEQUENCES AT
STATE AND NATIONAL LEVELS
The four chapters in this section are situated directly in the reform and account-
ability era. While the scale of “large-scale” and the policy implications—local,
state, or national—vary with each assessment studied, the chapters together ex-
amine the intentions, politics, and misperceptions behind externally imposed
writing standards and high stakes writing assessments and the resulting material
and policy ramifications of these reform and accountability efforts.
In “The Misuse of Writing Assessment for Political Purposes,” Edward M.
White (2005) identifies three focal areas of writing assessment that have been
shaped by politics and public policy: high school proficiency testing, college
placement, and mid-career assessments in colleges. The latter, “junior” writing
assessments, which are addressed only in White’s piece, are comparable to high
school proficiency testing in many ways. The remainder of the collection of
128
The Politics and Public Policy of Large-Scale Writing Assessment
articles focus primarily on one of two significant and long-standing types of as-
sessments White describes: secondary-level writing proficiency assessments and
college writing placement testing.
In addition to White, Arthur N. Applebee and co-authors J. W. Hammond
and Meredith Garcia all address K-12 writing proficiency testing and standards
at the state and national level. Applebee’s “Issues in Large Scale Writing Assess-
ment: Perspectives from the National Assessment of Educational Progress,” and
Hammond’s and Garcia’s “The Micropolitics of Pathways: Teacher Education,
Writing Assessment, and the Common Core” detail national writing standards
and writing assessments and their consequences broadly. Applebee (2007) dis-
cusses the National Assessment of Educational Progress (NAEP), a congressio-
nally mandated assessment across multiple subject areas, including a writing
assessment, given to a representative sample of elementary and secondary stu-
dents across the country. Applebee documents issues with large-scale writing
assessments and the ways disciplinary expertise has been leveraged to improve
the test and its utility. Hammond and Garcia (2017), on the other hand, focus
on the Common Core State Standards (CCSS). Rather than analyzing the large-
scale assessments associated with CCSS, PARCC, or Smarter Balanced (SBAC),
they study how teachers navigate these common national standards in their own
local contexts.
Along with White, co-authors Christie Toth, Jessica Nastal, Holly Hassel,
and Joanne Giordano interrogate college writing placement in the age of high
stakes. In “Introduction: Writing Placement, Assessment, and the Two-Year Col-
lege,” which is part of a JWA special issue on two-year college writing placement,
Toth et al. (2019) outline how two-year college writing placement has become
a particular target for educational reformers, which has resulted in a reconsider-
ation of the role of placement and common placement practices.
Collectively, these four chapters coalesce around three core themes:
• The intentions behind and (mis)use of mandated writing standards
and assessments for accountability purposes.
• The consequences of large-scale, high stakes writing assessments on
curriculum, teachers, and students.
• Positive outcomes and spaces for possibility among some large-scale
writing assessments and the policy implications.
129
Calhoon-Dillahunt
“accountability era.” In their articles in this section, the authors share that in-
tentions behind common writing standards and standardized assessments often
seem reasonable and even laudable. For instance, White (2005) asserts that it
is entirely logical to expect high school students to demonstrate a certain level
of reading and writing skill upon graduating. High school writing standards
and accompanying writing proficiency tests are promoted as a way to prepare
students for postsecondary writing. Hammond and Garcia (2017) describe how
definitions of “preparedness” became codified in the Common Core State Stan-
dards, enabling measurement of this elusive idea of “college and career readi-
ness.” According to the Common Core State Standards Initiative (2021) website,
a consistent, nationwide set of standards can be used to articulate and measure
student progress and to ensure students have acquired the necessary skills and
knowledge to achieve success in postsecondary education and the workforce.
Standardized testing, then, is viewed by policymakers and others involved in ed-
ucation reform as a way of raising standards and monitoring progress. According
to the 2004 National Commission on NAEP report, a “high school diploma
was no longer the culminating degree for most students” (Applebee, 2007, p.
86). Applebee also observes that about half of high school students who con-
tinued on to college were placed into developmental education, suggesting that
many students were graduating from high school underprepared to do the sort
of writing required in higher education. Thus, assessing 12th graders’ readiness
for college, military, and career seems essential.
According to White (2005) and Toth et al. (2019), in some ways, place-
ment testing aligns with intentions for high school writing proficiency testing,
ensuring students are “ready” to do college work. The theory behind placement
assessments is to match students to appropriate coursework, which allows col-
lege writing programs to maintain high standards in first year writing while
providing support for underprepared students before or as part of their first-year
writing coursework (White, 2005). In their article, Toth et al. (2019) share Will-
ingham’s 1974 algorithm for understanding the role of placement assessments,
a logic still pervasive in placement and developmental writing today. This log-
ic suggests that, by identifying students with poor writing skills and matching
those students to coursework designed to improve those skills, student learning
and retention in writing courses will be improved.
Holding institutions accountable for student learning and achievement is also
reasonable, according to White (2005): “it is wholly appropriate for politicians
and citizens to inquire into whether the schools are accomplishing established
goals” (p. 25). After all, states and local taxpayers, in particular, invest heavily
in education, and they should expect students to graduate with the knowledge
and skills needed for postsecondary pursuits. However, as Toth, Sullivan, and
130
The Politics and Public Policy of Large-Scale Writing Assessment
131
Calhoon-Dillahunt
learn to write. Linda Adler-Kassner (2017) argues that “this lament, this story
that students ‘can’t write,’ works from the premise that writing is ‘just writing.’
It’s a thing that writers bang out. It is constituted of words that are clear, that
mean the same thing to everyone, that are easily accessible and need only to be
plugged into forms” (p. 317). Toth et al. (2019) describe the foundational logic
of traditional writing placement in much the same way; it’s built on the notion
that such writing skills are attainable, measurable, and relevant to subsequent
college-level writing coursework and that assessing these generic skills–and plac-
ing students accordingly–will lead to improved writing. In describing the devel-
opment of the revised framework for the 2011 NAEP writing assessment, Ap-
plebee (2007) references a range of scholars who have challenged the “traditional
emphasis on writing as a generic skill, taught primarily in English language arts
or composition classes, and assessable through generic writing tasks detached
from particular disciplinary or socially constituted contexts” (p. 163), yet the
myths that “writing is just writing” and that “good writing” can be measured by
a single test and without regard to context persist.
Raising the stakes on writing assessments and at the same time basing such
assessments on fundamental misunderstandings about writing, assessment, and
accountability has led to misuse rather than reform. For instance, the perception
of writing as a generic skill has led to assessment tools that are often built to
prioritize ease of measurement rather than achievement of higher order skills,
resulting in assessments that focus on editing skills or formulaic writing tasks
(Applebee, 2007). According to Toth et al. (2019), “The widespread reliance
on commercially produced [writing placement] tests that measure a very limit-
ed construct of writing has prioritized knowledge of Edited American English
conventions at the expense of any other outcome, primarily because these are
the skills that can be easily measured” (p. 219). Thus, the tools that determine
whether consequences will be meted out do not capture the lofty goals of the
reform movement, and they are also biased against historically marginalized and
minoritized students by design, essentially ensuring that the schools that serve
such students will be penalized. These misuses are costly, in all senses.
In some cases, the high stakes assessments work against the very reforms they
are trying to institute, case in point, high school writing proficiency testing. As
several authors in this section articulate, the intentions behind large-scale high
school writing assessments are to raise standards and increase student proficiency
in writing for their postsecondary pursuits, as writing is a perceived “problem”
despite the fact that high school graduation rates are over 85 percent and about
two-thirds of those students enroll in postsecondary education after high school
(U.S. Department of Education, 2021). To “inspire” students and teachers to
take these standard-raising writing assessments seriously, many states tie earning
132
The Politics and Public Policy of Large-Scale Writing Assessment
CONSEQUENCES
Attaching penalties and rewards to student performance on single assessment
measures in order to drive educational reform and accountability policies has
had far-reaching repercussions. The authors in this section address the negative
consequences that have resulted from the use of mandated standards and high
stakes writing assessments in three particular areas: curriculum, teachers and
teaching, and students.
Impact on Curriculum
One of the most well-studied consequences of high stakes standardized testing is
its impact on curriculum. Sandra Murphy (2003) notes high stakes assessments
do not just measure achievement; they define it. Several authors in this section
observed the ways that such assessments narrow, constrain, and distort writing
133
Calhoon-Dillahunt
curriculum. Applebee (2007) argues that attempts to shape curriculum and as-
sessment around abstract notions of “career and college readiness” have generally
resulted in “a system of curriculum and assessment that focused on basic skills or
on generic workplace tasks (e.g., business letter format) that easily degenerated
into formulas with little real-world relevance” (p. 167).
The curricular impact of high stakes assessments can also be seen in postsec-
ondary writing placement. According to Toth et al. (2019), “[i]n the nation’s
open-admissions two-year colleges, where students enter from a wide range of
academic trajectories and often have not taken any kind of admissions exam,
placement assessment is nearly universal” (p. 215), and the use of commercial
placement products predominates. One of the results of this sort of placement
mechanism is that most two-year colleges offer multiple levels of pre-college
writing courses, which may be similarly disconnected from first-year writing
curriculum, focused instead on the “basic skills” developmental writers seeming-
ly lack, and which sometimes prohibit students from accessing other college-lev-
el courses outside of English. On the other end of the spectrum, some colleges
may exempt high performing students from the first-year writing requirement
altogether, which suggests that first-year writing curriculum is not about intro-
ducing students to a discipline, but, instead, teaching generic “writing” skills.
Writing assessments that are disconnected from a college’s first-year writ-
ing curriculum provide limited utility for authentic placement, but they send
powerful messages about how the institution views and values writing. Toth et
al. (2019) recognize that writing placement “is not a neutral action” (p. 218);
it communicates particular values and ideologies that affect how students, local
high schools, and others perceive writing, and as a result, it can impact both
high school curriculum and perceptions about the role of developmental and
first-year writing on college campuses. Simultaneously, commercial placement
tests also fail to communicate anything particular about a writing program, the
theory that underlies its curriculum, and the practices it values; such assessment
instead perpetuate the narrow conceptions of writing many students bring with
them from high school and the commonly held notion that first-year writing is
a course they need to “get out of the way.” Additionally, writing curriculum is
impacted, negatively and positively, by current reform movements that seek to
limit and accelerate developmental writing offerings (Toth et al., 2019).
134
The Politics and Public Policy of Large-Scale Writing Assessment
Impact on Students
While the studies included in this set of articles don’t address the impact of
high stakes testing on students directly, the implications are clear: students bear
the brunt of the consequences of standardized writing assessments. There is a
long history of using writing assessments to gatekeep and rank students, and the
consequences are even greater for students, especially historically underserved
students, when assessments are tied to diplomas for college-level access. White
(2005) argues that “Each of these assessments [high school proficiency exams,
placement tests, mid-career writing assessments] represents a gate through which
students must pass if they are to gain access to the privileges and enhanced sala-
ries of college graduates, and so they carry a particular social weight along with
their academic importance” (p. 145). The negative impacts of accountability
policies and high stakes assessments previously described, from penalizing al-
ready under-resourced schools to narrowing the curriculum and reducing teach-
er agency and professionalization, also affect the quality of education students
receive.
Toth et al. (2019) discuss most directly the impact standardized assessments
have had on students in the context of placement. The authors cite Haswell’s work
135
Calhoon-Dillahunt
136
The Politics and Public Policy of Large-Scale Writing Assessment
assessments, especially when the stakes for such assessments remain relatively
low for students, teachers, and institutions.
Because, as Applebee (2007) indicates, NAEP also served as a model for
many state-developed assessments, NAEP’s conscientiously designed and the-
oretically grounded assessment in writing had reverberating and likely positive
effects on other large-scale writing assessments. Granted, the NAEP assessment,
which appears to have been largely replaced at the high school-level by the
CCSS-connected Smarter Balanced and PARCC assessments, still struggles with
its intended goal of assessing student writing in ways that inform “preparedness
for postsecondary endeavors,” likely impossible to measure within a single, stan-
dardized assessment. However, the results have provided a fertile ground for
study, on a large scale, which has enabled the field of writing studies to evolve.
Of course, mandated writing standards and large-scale writing assessments
largely remain externally directed and developed. However, Hammond and
Garcia (2017) remind us that policies have to be put into practice: “Standards
. . . are never as autonomous or agentive as sometimes imagined; they are largely
contingent on interpretation and implementation by the very actors they are
intended to coordinate and perhaps constrain” (p. 184); indeed, they continue,
“reforms put in place are seldom as stable and standardized as intended” (Ham-
mond & Garcia, 2017, p. 186). The fact that policy is not determinative, is
“not so easily tamed,” means that policy requires support and buy-in to be fully
enacted. Policy implementation is also negotiated and navigated within partic-
ular contexts: “Homogenizing educational projects like the CCSS are always
alloyed with heterogeneous local perspectives, assumptions, and aims. While
perhaps obscured by standardizing efforts, local differences are not erased by
them.” (Hammond and Garcia, 2017, p. 185). These mediated spaces are places
of possibility, enabling the tools of policy implementation to be productively
adapted and providing agency for those involved in their implementation.
In their study of student teachers, mentor teachers, and field instructors at
three midwestern high schools, Hammond and Garcia (2017) observed that,
while all teachers involved in their study utilized CCSS in some way in their cur-
riculum development, they used and assessed the standards in different ways and
for their own purposes, tied to their own local contexts. Study participants tend-
ed to curate and even “retrofit” the standards, rather than adopt them outright,
which enabled the participants to select and prioritize the outcomes that fit
their curriculum and goals and their students’ needs as well as to use low stakes,
classroom-based assessment practices to determine mastery. The study revealed
that, instead of finding CCSS restrictive, the participating teachers tended to use
the standards as a rhetorical tool, “as a medium for managing communication
with stakeholders and—by extension—signaling professional participation in
137
Calhoon-Dillahunt
the collective enterprise of American education” (p. 4). Some found the CCSS
provided a common language for teachers, students, and parents to facilitate
teaching and learning in the discipline, and others found using this “profession-
al lingua franca” validated their work to external audiences, whether adminis-
trators, community members, or policymakers. Hammond and Garcia’s study
reveals that, while policymakers may devise standards, teachers are the ones who
enact them; the possibilities of educational reforms are tied to teacher buy-in
and teacher agency to implement such reforms in context.
Further, their study suggests that teacher agency in determining and design-
ing curriculum and assessment in context facilitates “professional accountabili-
ty,” as described by Linda Darling-Hammond. According to Darling-Hammond
(1989), “[p]rofessional accountability” requires that teachers are knowledgeable
and engaged practitioners, who participate collectively in all aspects of teaching
and learning, including assessment and local decision-making. Professional ac-
countability has much more potential to drive positive and lasting change than
the “carrot and stick” approaches associated with “accountability era” reforms.
Hammond and Garcia’s (2017) work reveals that when teachers have agency
in curricular decisions and when they are not threatened with punitive con-
sequences, teachers often view imposed standards and large-scale assessments
favorably.
CONCLUSION
Education reform’s “accountability” turn has often been framed in terms of “val-
ue added,” with value defined–and “accountability” enforced–through neoliber-
al economic ideologies. Ravitch argues this competitive, market-based approach
is wrong for public schools, which should function collaboratively and should
share what works with others (Inskeep, 2010). In the “reform and accountabili-
ty” era, large-scale writing assessments have often enabled these competitive and
punitive policies. However, Rose (2012) asserts that “our philosophy of educa-
tion—our guiding rationale for creating schools—has to include the intellectu-
al, social, civic, moral, and aesthetic motives as well. If these further motives are
not articulated, they fade from public policy, from institutional mission, from
curriculum development” (p. 185). Because it’s connected to policy, mission,
and curriculum–and, in fact, should emerge from these areas, writing assessment
is foundational to how we articulate and ascertain “value” in education, and the
future direction of writing assessment should consider “value-added” from the
broader perspective Rose identifies.
To this end, the chapters in this section suggest a range of possibilities
for future research. White’s, Applebee’s, and Hammond and Garcia’s work all
138
The Politics and Public Policy of Large-Scale Writing Assessment
recognize the critical role of teachers in education reform and reveal the impor-
tance of teacher engagement with standards and assessments and of assessments
emerging from and shaping curriculum. As teachers are enactors of reform
policies, more attention should be directed toward understanding the impact
of education policies and large-scale assessments on their practice and the role
professionalization and “professional accountability” plays in facilitating edu-
cational reform. Such research may reveal that investing in the changemakers,
teachers, rather than investing in large-scale assessment tools may yield better
results. Additionally, few studies talk to students about the ways in which they are
experiencing “accountability” reforms, particularly how such policies and high
stakes assessments affect their development and self-perceptions as writers and
their conceptions of writing.
Toth, Nastal, Hassel, and Giordano’s work highlights the importance of as-
sessing assessment tools. The work of researchers that questioned the validity,
reliability, and predictability of commonly used commercial placement tests has
resulted in many institutions abandoning such tests in favor of local alternatives
or reducing the stakes by using such tests as one consideration, among others,
for placement. These studies also led to revisions in commercial products them-
selves, often including a direct assessment of writing, albeit computer-scored.
Not only is it important to assess validity and reliability in large-scale writing
assessments, Toth et al. remind us of the importance of assessing the fairness of
writing assessment tools and methodologies, especially in large-scale and high
stakes assessments. Given that standardized writing assessments are rooted in
White Language Supremacy and ableism, studying the consequences of writing
assessments, in particular the disparate impacts of such assessments, can pro-
vide direction for how to redesign and even reimagine writing assessment tools
that attend to local contexts and value diverse students. Toth et al. argue–and
I agree–that two-year colleges are important spaces in which to conduct this
research, as two-year colleges serve diverse students and communities and, with
their open admissions policies, often serve as the primary access point for post-
secondary education for the least advantaged students.
Finally, writing assessment research is one key way to change the public nar-
rative around writing and to help policymakers develop informed solutions to
the educational problems they are trying to solve. Writing researchers and schol-
ars can contribute by asking different questions that counter the predominant
failure-driven narrative. For instance, how can writing assessments provide evi-
dence that student writing isn’t a “problem” and instead highlight the rich and
rhetorically conscious ways students language and compose in classrooms with
professionalized teachers developing curriculum appropriate to local contexts
and students’ needs? How can large-scale writing assessments account for the
139
Calhoon-Dillahunt
varied ways students demonstrate proficiency and success, for instance, in con-
sidering multiple measures instead of single assessments? How do lower stakes
assessments provide more meaningful information and yield more positive re-
sults? How can large-scale writing assessments provide evidence of “college and
career-readiness” by centering rhetorical dexterity and situated language practic-
es instead of facility with Edited American English?
In addition to researching in ways that change the dominant discourse around
writing, writing researchers and scholars can also practice their own rhetorical
dexterity by sharing writing research in accessible ways with public audiences
and policymakers. In other words, it is incumbent upon writing researchers to
“[find] ways to communicate our expertise to those outside of our discipline
and [seek] opportunities to participate in public conversations about literacy
education” (Calhoon-Dillahunt, 2015). Future writing researchers can also take
a page from two-year college teacher-scholar-activists who view engagement in
educational policy as a professional responsibility, which requires “undertak[ing]
the public work of defending educational access, teaching for democratic par-
ticipation, and advocating for practices and policies grounded in disciplinary
knowledges” (Toth, Sullivan, & Calhoon-Dillahunt, 2019).
REFERENCES
Adler-Kassner, L. (2017). Because writing is never just writing. College Composition and
Communication, 69(2), 317-340.
Applebee, A. (2007). Issues in large-scale writing assessment: Perspectives from the
National Assessment of Educational Progress. Journal of Writing Assessment, 3(2),
81-98. [Link]
Calhoon-Dillahunt, C. (2015). Finding our public voice. In D. Cambridge & P.
Lambert Stock (Eds.). Structural kindness: Essays on literacy education in honor of
Kent D. Williamson (pp. 163-169). National Council of Teachers of English.
Calhoon-Dillahunt, C. (2018). Returning to our roots: Creating the conditions and
capacity for change. College Composition and Communication, 70(2), 273-293.
Common Core State Standard Initiative. (2022). Frequently asked questions. http://
[Link]/about-the-standards/frequently-asked-questions/
Darling-Hammond, L. (2007). Evaluating ‘No Child Left Behind’. SCOPE: Stanford
Center for Opportunity Policy in Education. [Link]
blog/873
Eaton, J. S. (2008). The Higher Education Opportunity Act of 2008: What does
it mean and what does it do? Inside Accreditation. [Link]
education-opportunity-act-2008-what-does-it-mean-and-what-does-it-do
Hammond, J. W. & Garcia, M. (2017). The micropolitics of pathways: Teacher
education, writing assessment, and the Common Core. Journal of Writing
Assessment, 10(1). [Link]
140
The Politics and Public Policy of Large-Scale Writing Assessment
141
CHAPTER 4.
Edward M. White
University of Arizona
As I detail in a College English article, I first became involved with writing as-
sessment as a result of political interference with the teaching of first-year com-
position (White, 2001a). In that article, I point out how I stumbled into the
field of assessment more than 30 years ago as one of several English department
chairs trying to protect our first-year composition programs from being defined
by a demeaning test that the Cal State system chancellor wanted us to use to
further his political career. Every year since, I have been involved in one way or
another with the political dimension of assessment, a perspective that is usually
oppressive, insensitive, disrespectful, and manipulative to teachers and students.
I look back on three decades of struggling to live with such misuse of writing
assessment, even as I have stressed in my scholarship over the last three decades
the importance of teacher involvement and understanding of assessment as a
professional responsibility, indeed one with undoubted political ramifications.
Political figures love assessment because it allows them to posture about edu-
cation and pretend to themselves and to others that they are improving edu-
cation by measuring a simplified version of it. Teachers generally dislike and
distrust assessment, because it almost inevitably narrows and often reduces what
they do to simple numbers that will be used against their students and them.
Meanwhile, those of us actually teaching writing use assessment of one sort or
another all of the time in our classrooms (Huot, 2002; White, 2006). How
else, for example, can we teach self-assessment and revision? Regardless of the
centrality of assessment to the teaching of writing, we are forever fending off
the efforts of politicians and testing companies to use assessment improperly,
to prove that our students are not learning, and that we are at fault. Although I
agree that teachers and writing program administrators (WPAs) are responsible
for assessing those programs, the current assessment climate often makes teach-
ers, students, and WPAs accountable to ill-conceived, poorly constructed, and
misused assessments. No wonder that the very mention of assessment is enough
to send many teachers racing from the room, even if it sends them back to their
offices—to continue responding to this week’s set of papers.
In this article, I focus on writing assessment in its political definition, not
as the form of professionalism that allows us to do our jobs with our students.
This is an important distinction because the mandated assessments from those
ignorant of what we do have little or nothing to do with our teaching or our
students. One canny reviewer of the MLA book I edited with two others enti-
tled Assessment of Writing: Politics, Policies, Practices (White, Lutz, & Kamusikiri,
1996) wrote that it should really have been titled Assessment of Writing: Politics,
Politics, Politics. So I am going to follow his advice here, attending solely to the
politics of writing assessment, an aspect of the field that is, unfortunately, its
most prominent and unexamined face. We do need to assess our students’ work
to help them improve and to assess our programs to see if they are doing what we
expect them to do. But we also must dispute the view that testing, particularly
testing using nationally normed tests, can determine if we are teaching well and
responsibly.
I intend to look at three places where writing assessment is most prominently
misused: the high school writing assessments, now afflicting students seeking
144
The Misuse of Writing Assessment for Political Purposes
their diplomas in all but two states; placement testing, the usual sorting of first-
year students into those supposedly ready for regular college work and those
who are not; and, finally, mid-career assessments, required of college students as
they move from the sophomore year to the junior year in an attempt to ensure
that such students will have a certain level of ability at reading and writing, at
least enough to placate their major professors in college and their employers after
graduation. Each of these assessments represents a gate through which students
must pass if they are to gain access to the privileges and enhanced salaries of
college graduates, and so they carry a particular social weight along with their
academic importance. In other words, each of these tests carry significant con-
sequences or high stakes. In each case, I examine the political reasons why these
assessments are set in motion and point to the inner contradictions that make
it quite impossible for them ever to accomplish their vaguely stated purposes—
which leads to a certain amount of thrashing about to identify the problems
and possible solutions. Ultimately, I believe we need to reconstruct the stage for
writing assessment, and I hope my discussion can begin this important work.
We could thus cast this discussion as a study of violations of test validity, using
modern definitions of validity that extend beyond score correlations into the en-
tire context of a testing program, including consequences for test takers and any-
thing else that affects the decisions made on behalf of a measure. But in a short
article focusing on political issues, I focus specifically on the inherent problems
and contradictions these programs represent and allude to some effective ways to
approach the political goals in a responsible way. It bears mentioning that if test
users and developers adhered to current conceptions of validity summarized in
the most recent Standards for Educational and Psychological Testing (1999) most
of the problems I explore in this article would not exist.
145
White
146
The Misuse of Writing Assessment for Political Purposes
to Vacuous Thinking and Writing.” The Murphy study compares the effects of a
careful test in 1988, designed largely by teachers, with a commercial standard-
ized test given in 2001; the results of the later test showed a clear “narrowing and
fragmentation of the curriculum” (p. 40). The Hillocks study looks closely at
statewide tests in Texas and Illinois, concluding that they “work against the goal
of learning how to think critically and argue persuasively” (p. 20).
In addition to the scholarly evidence for the unfortunate effects of these
politically directed tests on students, teachers, and learning, I can add a per-
sonal experience, from my graduate course in writing research in California,
one of the states where the SAT-9 was a high-stakes test, determining budgets
and “success” for high schools. One of my students, a fine high school teacher,
told me of her confrontation with the school principal, at a teachers’ meeting.
He had distributed the SAT-9 scores, which were down, and then informed the
teachers that everything they did in class must be directed to improving those
scores. My student, emboldened by my course, spoke out: “I’m an English
teacher. Are you saying that I can’t teach reading and writing because they’re
not on the test?” She spoke mournfully of his reply: “He pointed his finger at
me and told me very forcefully that I was not to waste class time on reading and
writing or I’d be fired!”
To the obvious contradiction of a senior-level high school test undermin-
ing the curriculum so that it can be passed by eighth graders, we need to add
the further problem of college entrance. Shouldn’t such a test serve for college
placement? Well, logically yes. But in practice, almost half of the graduating
high school seniors are not heading for college, so why should their high school
diplomas depend on a college entrance measure? Besides, the test is in fact de-
signed for eighth graders. Furthermore, it is quite possible that the best high
school classes in both English and math are more demanding and set higher
standards than the usual first-year college courses in those subjects, so we have
no clear definition of what college-level proficiency means beyond particular
college practice. National tests, one might imagine, pose a kind of definition;
but these range from the relatively strict standards of the Advanced Placement
Program to the most minimal multiple choice scores embodied by the General
Examinations of the College-Level Examination Program, both administered by
the Educational Testing Service, serving consumers at all levels; test criteria and
standards move lower still as we look at the products of less professional testing
firms. Because we have no reference point for the definition of “college-level”
performance from such varied test criteria, we cannot take solace from national
tests without national curricula, which nobody really wants. Thus, the stage is
set for a continuing muddle, with the writing assessment asked to solve unsolv-
able problems and to assure everyone that all can be made well if only teachers
147
White
worked harder and the administration cracked down on the worst slackers and
we tested students often enough.
To be sure, the issue of school accountability is neither trivial nor superficial.
It is wholly appropriate for politicians and citizens to inquire into whether the
schools are accomplishing established goals. But if they were serious about the
matter, this accountability would not rest entirely on the hapless students taking
more or less relevant tests. Genuine questions about school accountability would
ask about the school environment (does it support learning and is it a supportive,
well-maintained, and pleasant place?), teachers and administrators (are they well
trained and well paid, the kind of people who should be entrusted with students?),
and parents (are they respected as partners in student learning, do they partici-
pate?), as well as student test data; but these matters refer to political responsibility
for schools in ways that do not allow the politicians to point fingers at others in
nice sound bytes. So only the students are assessed, on the cheap and irresponsibly,
and these student tests are assumed to represent the status of schools.
But in fact, nobody really pays much attention to the entire operation, aside
from the politicians, pointing with pride to their efforts to raise standards, and
the students, forced by punishments or induced by free doughnuts or some
other bribe to take meaningless tests. The colleges and universities universally
ignore the high school tests, preferring to use tests designed for college admis-
sion, and usually, sensibly, preferring their own placement procedures, tailored
to their own students. (But that is probably going to change; see the following
section of this article.) And high school graduates seem to read and write about
as well or as badly as they did before all of these tests were instituted, despite test
scores rigged to show improvement, because those actual proficiencies depend
on the parents, teachers, and the school environment, the key ingredients in any
education. It is not hard to imagine more constructive uses for the vast sums
now being spent on testing, to very little purpose, in this sad pretense at school
accountability.
148
The Misuse of Writing Assessment for Political Purposes
labeling, and retrograde employment practices) and the popular right (object-
ing to the use of university resources for those defined as not ready for univer-
sity work). When we think systematically about placement into the first-year
writing course, we encounter a tangle of academic, professional, political, and
social issues that makes it difficult to decide on an appropriate course of action
in general or at our own institutions. Again, as with high school proficiency
tests, we find that political motives and naïveté about assessment normally lead
to meaningless or destructive tests, useful primarily for political posturing and
jockeying for funding.
The least satisfactory method of placement—and the most common in
American colleges—is by means of some multiple-choice testing of editing skills,
a quick impromptu writing sample, or some combination of both. The problems
with this kind of assessment have become obvious. The multiple-choice test of
editing skills does not require the production of text and so measures skills not
directly related to the first-year writing course. Edgington, Ware, Tucker, and
Huot (2005) report that more than 250 students placed in remedial courses
through the COMPASS test (an untimed editing exercise on computers) were
also placed by a writing sample into the regular first-year writing course, and all
these students chose the higher placement. More than 70% of these students
received an A or B in the course, and more than 90% of these students received
at least a C. The indirect relation of such tests to writing is in much dispute and
seems particularly weak for students from homes that do not speak the school
dialect. Although a written impromptu placement test is certainly a better op-
tion than tests that do not contain any writing at all, we already have several
examples of portfolio placement programs that are accurate, reliable, and afford-
able (Hamp-Lyons & Condon, 2000; Hester, Neal, O’Neill, & Huot, 2005;
Willard-Traub et al., 1999; [Link] On the other
hand, as recently as a decade ago, at least half of all respondents to a national sur-
vey on placement indicated that they were using something other than student
writing to make placement decisions (Huot, 1994). With the validity of these
placement decisions so questionable, one must ask why they dominate Amer-
ican higher education. There are numbers of answers, of course, but political
considerations are certainly behind most of them. I became convinced of this,
a few years ago, when I tried to convince the writing directors of the California
State University system to replace their outdated English Placement Test (EPT;
whose development and implementation I administered in 1975-1977) with a
more modern and more valid portfolio requirement. “Keep your hands off our
EPT,” they said, unified for once. “All of our financing depends on those scores.”
I may be surprising some readers, because I have, for some decades been
a strong advocate of placement testing, based on the theoretical arguments
149
White
150
The Misuse of Writing Assessment for Political Purposes
the only escape is to leave college altogether. And meanwhile, everyone knows
that such untested matters as social class, finances, motivation, self-confidence,
reading experience, and family responsibilities play a large role in student success
in every writing class. In other words, large-scale placement tests, which tend to
measure editing skills on other people’s prose or impromptu fluency on a writing
topic about which there is little time to think, do not allow for the same kind
of decision making into every college’s writing program. They measure only a
small component of what is needed for student success, and they cannot be re-
sponsive to the program into which they are placing students. They tend to be
a social-sorting mechanism, useful for political posturing, but of limited use for
students, teachers, or institutions.
So, how can we place students into a well-designed series of college writing
classes, including a variety of basic writing instruction, that will lead to student
and teacher satisfaction and to as much student success as possible? Clearly, the
first step is for each college or university to design well-defined writing courses
that are appropriate for its own student body, including some clear sense of
what a student should be able to demonstrate in order to profit from a particular
course. This is a crucial activity that large-scale placement testing, with its built-
in illusion that all college programs are the same, has allowed most colleges to
avoid. For them, it is cheaper and easier to let the tests place students, to staff
the writing courses with part-time help whose voices on curricular matters will
not be heard, to hope that whatever such teachers do in class will be minimally
respectable, and (in too many cases) to wish that the students in need of extra
help will blame themselves for their weak preparation and just go away quietly,
after surrendering their tuition dollars. Regardless of what placement procedures
an institution uses, there must be a systematic, rigorous program of validity
inquiry in which placement decisions are studied from a variety of perspectives
including but not limited to student success in the course and teacher and stu-
dent satisfaction with the placement procedures.
One interesting and important innovation in placement shifts the proposed
solution from assessing students’ writing, editing, or grammar or vocabulary
knowledge to an enhanced form of counseling. Part of the attractiveness of Di-
rected Self-Placement (DSP) is that it proposes a way through this tangle, one
that might keep the advantages of placement yet avoid the disadvantages of
placement testing. The idea is deceptively simple. In place of testing students,
the institution puts its efforts into informing students about the demands and
expectations of the composition courses available to them and how they can
meet the writing requirement. Then the student makes an informed choice,
and takes full responsibility for that choice, instead of more or less grudgingly
accepting test results and institutional placement. DSP assumes that students
151
White
will be mature enough to choose the course that is right for them, if they have
enough information and pressure to choose wisely. DSP also assumes that there
may be many reasons besides test performance for students to choose more or
less demanding writing courses in their first year of college. And—perhaps the
most perilous assumption of all—DSP depends on the institution clearly de-
fining the requirements and proposed outcomes of its different writing courses,
maintaining consistency in those definitions, and then communicating them to
entering students. For DSP to be effective, the institution must develop some
means of making that information meaningful to young students, generally be-
mused by the mass of lectures, warnings, greetings, and exhortations offered in
the weeks before the opening of classes (Royer & Gilles, 2003).
Of course, DSP is no panacea, although its promise is encouraging. Like many
other solutions to educational problems, DSP offers new problems in place of old.
Yet, the new problems are those that postsecondary education should be meet-
ing anyway: helping students take responsibility for their own learning, replacing
reductive placement testing with sound counseling, developing clear curricular
guidelines and outcomes, and becoming less paternal and more, shall we say, avun-
cular. At heart, DSP, like the concept of placement itself, is a conservative propos-
al, one that maintains the first-year writing requirement as an essential introduc-
tion to college-level writing, thinking, and problem solving. DSP is an answer to
those unwisely calling for an end to college writing requirements as unnecessary in
modern times of technological and vocational revolution. At the same time, DSP
proposes a radical solution to the persistent problems of over testing, negative
labeling, and student alienation from required coursework.
Will it work? That is, will it be able to convince those inside and outside of
academe that it is meeting the political goals of assessment when it avoids assess-
ment entirely? At this point, nobody really knows. Maybe entering college stu-
dents are not really able to make wise course decisions; perhaps communicating
with entering students about their choices is too difficult; maybe the curriculum
is in too much disarray to become transparent. Many institutions will need to
revamp their counseling procedures for new students to make DSP possible and
such change is exceedingly difficult. All kinds of unforeseen problems lurk be-
hind the implementation of DSP, perhaps most pointedly a shift in perception
of who should be responsible for academic decisions. The critiques of DSP are
appearing along with the encomiums, even in the Royer and Gilles book. But
the concept is promising enough for widespread trials—now under way every-
where one looks—and we need to gather information about what happens, as
concept becomes procedure at real institutions.
But, as we may expect, a simple and crude political solution to the issue
of placement stands ready to replace existing local placement experiments and
152
The Misuse of Writing Assessment for Political Purposes
abort the promise of DSP. Both of the major American college aptitude testing
institutions, the College Board, and the American College Testing Service, have
added short impromptu writing tests to their admissions testing programs in
2005. Because most students bound for 4-year colleges and universities take one
of these tests, almost every admissions office will now have ready-made place-
ment information at hand, paid for by the student rather than the college, and
buttressed by an imposing set of comparative statistics. It will not matter that on
many, perhaps most campuses, the information will be useless or worse; it will
be politically difficult, if not impossible, to resist using it to place students. So
we can anticipate that local placement procedures and the high promise of DSP
will fade away in short order.
What is wrong with using national scores on a short piece of impromptu
writing to place students in college writing courses? Think for a moment of
devising a writing topic appropriate for the privileged students applying to Dart-
mouth and for the struggling residents of inner-city blighted neighborhoods;
consider attempting to score such an examination—or, worse still, attempting
to program a computer to score such an examination—with some regard for
the diversity of its examinees; consider trying to understand the results when
comparing students who grew up in homes using the school dialect to those
for whom other dialects or even other languages were used at home. Locally
administered placement tests, locally scored, have been able to deal with these
problems in various ways, but all those accommodations will probably now be
swept away with one universal score, based on national norms. Perhaps most
damaging will be the effects of the new tests on the college composition cur-
riculum (oh yes, that), now more or less tailored to the students who wind up
sitting in actual classrooms. If we think of the essential purpose of placement,
to match particular students to a particular curriculum at a particular campus,
it becomes preposterous to even imagine that a single common test score can
be used to make accurate, consequential decisions for more than 2 million stu-
dents entering a variety of institutions. And because tests inevitably define their
subjects, think of the high school students for whom writing will increasingly
become narrow test preparation.
An additional cruel twist still awaits. Although the commercial firms devising
and scoring these written tests are busy recruiting battalions of human readers
to score them, does anyone doubt that those humans will shortly be replaced by
computers, now moving rapidly into the scoring of writing? A grim satire looms:
student computers writing out prose to be read by scoring computers, in turn
placing the students into composition sections increasingly taught in computer
centers by computer-based instruction. The economy and efficiency is stunning:
Neither students nor teachers will need to write or read, or even show up on
153
White
campus. Of course, I exaggerate here for effect, and I’m not dismissing the very
real and important role computer technology can play in the teaching of writing.
On the other hand, my exaggeration has its point. We can emphasize technology
at the expense of creating suitable environments for teaching and learning.
Leaving the futuristic satire for the present, we must agree that it will be a
bold institution indeed willing to budget its own placement procedures, for its
own students, in the face of the scores that will be arriving at no additional cost
to the college. Where will we find the political will to fight such a battle? We can
expect an impressive marketing campaign, arguing that the vexatious problem of
coping with individual students and a broad writing curriculum has now been
solved. We must hope that institutions and faculty will resist such false solutions
and the mechanistic future they preshadow. As this essay goes to press it is heart-
ening to observe that several members of the WPA listserv report some resistance
to using the new ACT or SAT writing tests for placement purposes.
154
The Misuse of Writing Assessment for Political Purposes
The problems with the rising junior exams are not as severe as they are with
the high school tests, at least on the surface. A faculty can often agree on what
a student should be able to demonstrate in order to succeed in upper division
courses: the ability to read texts of moderate difficulty and write about them
clearly enough to show that understanding; the ability to assert some kind of
idea and develop it coherently for a few pages; the ability to use source material
to support an assertion rather than to substitute for one; and ability to edit writ-
ten work so it is reasonably free from distracting or embarrassing errors. Sounds
easy. But as various departments begin to consider their special needs, more
criteria start to appear: the ability to write about scientific or technical matters
so a nontechnical reader can understand; the ability to use technology to write
and revise; the ability to integrate data and charts into an argument; and so on.
Thus, the creation of a responsible test becomes either so complicated and
wide ranging as to be very expensive and time-consuming, or so simple that it
loses all credibility. As always, the national testing firms are prominent in the
market with their multiple-choice tests, which few faculty respect, if they can
even be cajoled into evaluating the instruments. Usually, the English department
is told to manage the thing somehow and the rest of the faculty wash their hands
of the matter. Meanwhile, about half of the students (those who can be forced
or cajoled into taking the test) fail it, no matter what it is. They have been coun-
seled to get first-year writing courses “out of the way,” and have written little or
nothing in their other lower division courses, so they struggle to remember how
to do whatever is called for.
If the creation of the rising junior test is difficult and expensive, the scoring
of it is more so. Large institutions wind up with hundreds, sometimes thousands
of tests to grade and little money for paying graders. More than one such test has
been abandoned for lack of money to pay readers (the University of Arizona’s
UDWPE, for example) and on some campuses absurd multiple-choice tests have
been used as a way to keep the shell of the requirement in effect on the cheap (as
one Texas university does). But even when the scoring is supported, by student
fees or otherwise, the standards for scoring become a vexatious issue. Can we
really expect the students in math or agriculture or physical education to come
up to the same standards we might expect of English or history majors? To what
degree should we tailor the writing topics and test standards as well as the criteria
for scoring to the student’s major? It is difficult to harmonize such matters as
the preference for brevity and clarity in the sciences with the taste for complex-
ity, metaphor, and wit in the humanities, especially when English faculty end
up being responsible for constructing and scoring the tests. Even more vexing
for scoring is the ambiguity behind the assessment’s purpose: Is the test really a
minimum proficiency exam, designed to catch only students whose writing is
155
White
156
The Misuse of Writing Assessment for Political Purposes
CONCLUSION
I want to be explicit here that I am not making a case against writing assessment.
We will be better teachers of writing if we know how to assess our students’ work
responsibly, and our students will learn how to revise their work if they learn
from us how to assess their own work. Furthermore, careful and responsible as-
sessment of writing beyond the classroom is professionally important, as we have
learned from much experience; if we do not meet the academic and political
demand for writing assessment at various levels, others will happily take on that
task, whether they know anything about the matter or not. Keith Rhodes (in
conversation) has named my little proverb on this matter “White’s first law of as-
sessodynamics”: Assess thyself or assessment will be done unto thee. Indeed, in some
ways, the misuses of writing assessment I have been discussing are symptoms of
our own failures to accept this responsibility. I am, in short, a strong supporter
of the responsible uses of writing assessment.
But what I have been dealing with in this article is the misuse of writing
assessment. In some ways, this misuse derives from an exaggerated, even a credu-
lous misunderstanding, of what particular kinds of assessments can accomplish.
In other ways, it merely reflects an all-too-American view that competition is a
positive value and that it is good for society to have a few winners and many
losers. In still other cases, it embodies a devious way to avoid difficult problems
by substituting a test score—any old score from any old test—as a pseudo an-
swer to such hard social problems as the meaning of a high school diploma or a
college degree, or even for whom the doors of opportunity should swing open or
157
White
shut. This fast, easy and mis-use of assessment is an important part of the Bush
Administration’s No Child Left Behind Legislation that emphasizes an elaborate
testing, standards and accountability program without the resources and lead-
ership for students to achieve the skills they will be tested on. We must guard
against the misuse of assessment while at the same time we promote a climate of
responsibility in writing instruction and writing program administration. Just as
administrators, politicians and the private sector urge us to be more accountable,
we must also hold these people and the testing companies to the very principles
of validity that should drive all test use.
REFERENCES
American Educational Research Association. (1999). Standards for educational and
psychological testing. American Educational Research Association.
Edgington, A., Ware, K., Tucker, M., & Huot, B. (2005). The road to mainstreaming:
One program’s successful but cautionary tale. In C. Handa & S. McGee (Eds.),
Discord and direction: The postmodern WPA. Utah State University Press.
Hamp-Lyons, L., & Condon, W. (2000). Assessing the portfolio: Principles for practice,
theory and research. Hampton Press.
Haswell, R. (Ed.). (2001). Beyond outcomes: Assessment and instruction within a
university writing program. Ablex.
Hester, V., Neal, M., O’Neill, P., & Huot, B. (2005). Adding portfolios to the
placement process: A longitudinal study. In P. O’Neill (Ed.), Blurring boundaries:
Research and teaching beyond a discipline. Hampton Press.
Hillocks, G. (2003). How state assessments lead to vacuous thinking and writing.
Journal of Writing Assessment, 1(1). [Link]
Huot, B. (1994). A survey of college and university writing placement practices. WPA:
Writing Program Administration, 17(3), 49-67.
Huot, B. (2002). Toward a new discourse of assessment for the college writing
classroom. College English, 65(2), 163-180.
Lutz, W. (1996). Legal issues in the practice and politics in the assessment of writing.
In E. M. White, W. Lutz, & S. Kamusikiri (Eds.), Assessment of writing: Politics,
policies, practices (pp.33-34). Modern Language Association.
McNenny, G. (Ed.). (2001). Mainstreaming basic writers: Politics and pedagogies of
access. Erlbaum.
Murphy, S. (2003). That was then, this is now: the impact of changing assessment
policies on teachers and the teaching of writing in California. Journal of Writing
Assessment, 1(1). [Link]
Phipps, R. (Ed.). (1998). College remediation: What it is, what it costs, what’s at stake.
The Institute for Higher Education Policy.
Royer, D., & Gilles, R. (Eds.). (2003). Directed self-placement: Principles and practices.
Hampton Press.
158
The Misuse of Writing Assessment for Political Purposes
159
CHAPTER 5.
ISSUES IN LARGE-SCALE
WRITING ASSESSMENT:
PERSPECTIVES FROM THE
NATIONAL ASSESSMENT OF
EDUCATIONAL PROGRESS
Arthur N. Applebee
University at Albany, SUNY
This chapter reviews the development of the framework for the 2011
National Assessment of Educational Progress in writing. An issue pa-
per commissioned by the National Assessment Governing Board is used
to consider a number of continuing issues in large-scale assessment of
writing, including the definition of the domain of writing tasks, which
tasks should actually be assessed at which grade levels, the relationship
of the assessment to postsecondary demands, the role of commonly avail-
able tools such as word processing software in the construct of writing
achievement, the specification and measurement of achievement, the
development of appropriate topics for writing, the issue of time for writ-
ing, and accommodations for English learners, students with disabili-
ties, and low achievers.
162
Issues in Large-Scale Writing Assessment
163
Applebee
Texas, for example, requires writing for “various audiences and purposes,” in a
variety of forms, including “business, personal, literary, and persuasive texts.”
California instead treats these generalized purposes as part of “writing strate-
gies,” and specifies a variety of specific genres to be assessed (e.g., at Grade 11,
fictional, autobiographical, or biographical narrative; responses to literature; re-
flective com-positions; historical investigation reports; and job applications and
resumes.)
There are other alternatives. College entrance exams from the College Board
and ACT both assume that good writing is a generic skill, at least in academic
contexts; the College Board, for example, advises that high scores will go to
“essays that insightfully develop a point of view with appropriate reasons and
examples and use language skillfully” (College Board, 2008).
From the Australian genre-theory perspective, Martin and Rothery point
in another direction, with a list of schooled nonfiction genres: recount, report,
procedure, explanation, persuasion, and discussion. Their listing, like others
from the Australian group, introduces terminology unfamiliar to American
readers, and also collapses their original insights about the situated nature of
genre knowledge into a generic set of “school” genres that are not all that distant
from Britton et al.’s (1975) and Moffett’s (1968) subcategories of informational
or expository writing.
THE OUTCOME
Lacking a widely accepted way to resolve these problems in definition and cat-
egorization, the committees developing the 2011 NAEP framework (NAGB,
2007) proposed organizing the assessment around three broad purposes for writ-
ing that are closely related to the distinctions made in earlier assessments:
1. to persuade, in order to change the reader’s point of view or affect the
reader’s action;
2. to explain, in order to expand the reader’s understanding; and
3. to convey experience, real or imagined.
These represent an attempt to clarify and elaborate the categories of persuade,
inform, and narrate in the previous assessment. The framework also attempts to
separate purposes from the ways they are carried out, noting that there are a wide
variety of strategies for thinking and writing that writers may use in addressing
these purposes, including the traditional modes of narration and description, as
well as processes such as analyzing and interpreting, and organizational strategies
such as compare and contrast. Taking this notion of choices available to writers
even further, the 2011 framework recommends that students in Grades 8 and
164
Issues in Large-Scale Writing Assessment
12 be allowed to choose the particular genre or form in which they will respond
(e.g., letter, essay, brochure), rather than having the form of the response dictat-
ed by the writing prompt.
The purposes embodied in this proposal, despite the changes in terminology,
will provide an easy transition for other assessments that look to NAEP for guid-
ance. The proposal also acknowledges that there is at present no widely accepted
alter-native in either theory or practice.
165
Applebee
THE OUTCOME
Most large-scale assessments have too few separate writing items to have a wide
range of task difficulty. The NAEP framework for 2011 similarly recommends
a focus on tasks that will encourage all students to write at some length, rather
than including some unusually easy or unusually difficult tasks.
The NAEP framework for 2011 emphasizes the importance of writing for
a wide range of purposes at all grade levels, including some attention to each of
three broad purposes included in the assessment framework. In recognition of
the shifting demands of the curriculum, however, the framework places some-
what more emphasis on writing to explain and to persuade in the upper grades,
and correspondingly less emphasis on writing to convey experience (primarily
storytelling and personal experience essays). For all three purposes, the frame-
work recommends increasingly abstract content and more distant audiences in
the upper grades.
The specific types of writing to be emphasized at different grades warrants
careful consideration in any large-scale assessment. Curriculum has a tendency
to narrow around the types that are assessed, often coupled with unintended
effects on what counts as writing well (Hillocks, 2002). Assessments that have
to rely on a limited number of tasks at a given grade level might do well to con-
sider designs that sample from a larger range of possible tasks at each grade level
assessed, rather than focusing on one or two types.
166
Issues in Large-Scale Writing Assessment
At the same time, 45%-55% of entering freshmen are unprepared for college
work, as reflected in placements in remedial coursework during their first year
in college.
Lacking any other national standard for measuring preparedness, the Com-
mission recommended that new NAEP frameworks for the 12th grade be ori-
ented toward assessing preparedness for the challenges of college, workplace
training, and the military. At the same time, the Commission noted that there is
little consensus on what “preparedness” means, and that validating measures of
preparedness is likely to require extensive follow-up studies exploring how stu-
dents at various achievement levels do in various post-high school contexts. The
NAGB’s Assessment Development Committee has endorsed this emphasis on
12th-grade preparedness, while noting that the issue is complex and the message
that NAEP will send in this regard is very important.
The history of attempts to shape curriculum and assessment around pre-
paredness for future life or work is not a happy one (Applebee, 1974). Past
attempts to inventory necessary skills have tended to converge on simple skills
that are easy to itemize (spelling, punctuation) rather than higher-level skills
(e.g., thoughtful argument and use of evidence) that virtually everyone cites as
essential goals of education. The result was usually a system of curriculum and
assessment that focused on basic skills or on generic workplace tasks (e.g., busi-
ness letter format) that easily degenerated into formulas with little real-world
relevance.
The most extensive recent effort to relate high school achievement to pre-
paredness both for college study and for the workplace is the American Diploma
Project (2004). Drawing on studies of the skills needed in high-performance,
high-growth jobs, as well as the requirements for college-level tasks, the Amer-
ican Diploma Project report emphasizes higher-level skills such as expressing
ideas clearly and persuasively, and producing high quality writing resulting from
careful planning, drafting, and meaningful revision. The report also includes
extensive benchmarks meant to indicate the level of achievement appropriate for
high school graduation. The 10 benchmarks for writing cover a wide range, from
planning, drafting, and revising; to selecting language appropriate for purpose,
audience, and context; to writing well-structured academic essays and work-re-
lated texts; to using appropriate software programs. Benchmarks under other
headings also refer to writing tasks, however, including benchmarks labeled as
research, logic, informational text, media, and literature. Although the over-
all emphasis remains on higher-level accomplishments, the benchmarks show
some of the problems of earlier attempts, with appropriate citation of print or
electronic sources emerging as a benchmark at the same level of importance as
writing an academic essay.
167
Applebee
THE OUTCOME
The issue of how current performance relates to the demands of future contexts
is an important consideration in the development of any assessment. The new
NAEP framework addresses this issue by stressing the continuity of skills that
will be needed in postsecondary contexts, rather than by emphasizing particular
postsecondary types of writing: Good writing at all levels entails appropriate
development of ideas, logical organization, language facility, and use of conven-
tions—all shaped by purpose and audience. Postsecondary contexts also empha-
size effective analysis, interpretation, and problem-solving, which is reflected in
the 2011 framework in a gradual increase in writing to explain and to persuade
at Grades 8 and 12.
168
Issues in Large-Scale Writing Assessment
writing on a computer in school and at home; on the other hand, those who are
not used to writing on a computer will either be handicapped by poor keyboard-
ing skills, or if they compose by hand by the greater length of essays produced by
their computer-using peers.
The research base on the effects of word processors on assessment results is
slim and not particularly convincing; arguments that paper-and-pencil tests un-
derestimate achievement of students who are used to writing on word processors
treat writing as though it were being evaluated against an external, fixed standard
(e.g., Russell & Plati, 2002), when in fact writing rubrics ordinarily reflect the
circumstances of production. Rather than an overall increase in performance, a
switch to a computerized assessment including word processing software is more
likely to lead to changes in the benchmarks at each level in the scoring rubric to
reflect the advantages accrued from the new format.
The most extensive study of the effects of computerizing a writing assessment
is NAEP’s 2002 study of writing online (Sandene et al., 2005). This special study
compared performance on two NAEP writing tasks (one informative and one per-
suasive) at the eighth-grade level, when given as part of the regular paper-and-pen-
cil assessment or given in a special Web- or laptop-based format that also included
simple word processing tools. The detailed results show a number of topic-specific
differences in performance across formats, but are generally encouraging. There
were no equity-related differences in essay quality, although there was a 1% higher
response rate for the paper-and-pencil version of one task. Males also wrote signifi-
cantly longer responses on computer than on the paper-and-pencil version of one
task, but their essays were not rated significantly higher.
Students with more hands-on computer skill (as measured by typing speed,
error rate, and ability to use word processing tools) did better on both of the com-
puter based writing tasks; the correlation between their overall writing score and
the measure of computer skill was .42; even after adjusting for paper-and-pencil
writing achievement, computer skill still accounted for about 11% of the variation
in computer-based measures of writing achievement. The “hands-on” comput-
er familiarity measure, however, had a significant literacy component that may
account for much of this relationship. Other measures of computer experience,
including frequency of completing various kinds of writing assignments on a com-
puter, were unrelated to computer-based writing achievement.
Overall, the authors of the NAEP writing online study conclude that ag-
gregated scores from online assessment do not differ significantly from pa-
per-and-pencil results, although results for individual students may do so.
Although school-level data have recently suggested that equity issues in com-
puter access have been reduced, at the student level issues of access have not
been completely resolved. In 2003, for example, there were fewer computers
169
Applebee
THE OUTCOME
In designing the 2011 NAEP framework, the committees decided that writing
in the 21st century will be computer-based. This is already how most students
write, and it is certainly an expectation for writing in the workplace and in post-
secondary education. Thus, the 2011 writing framework calls for assessing the
ability to write using word processing software at Grades 8 and 12. The frame-
work calls for students to write using “commonly available tools,” including the
various writing and editing tools widely available in commercial word process-
ing programs. The framework also calls for a computer-based assessment to be
phased in at Grade 4 over the life of the framework, as access to and experience
with word processing becomes more widespread in the elementary grades.
Computer-based assessment seems almost an inevitable response to the fre-
quency and scale of mandated assessments in all areas of the curriculum. For writ-
ing assessment, developers will need to consider how advances in computer use
and availability are impacting writing instruction, and what this means for defi-
nitions of what it means to write well. If equity issues can be resolved, a comput-
er-based assessment has a number of advantages in measuring writing achievement
and in providing accommodations to students who need them (see section on
accommodations). Equity issues, however, are much more acute for assessments
that report individual scores than they are for NAEP, whose results can serve policy
development by highlighting issues of access without penalizing individuals.
170
Issues in Large-Scale Writing Assessment
171
Applebee
(Breland, Camp, Jones, Morris, & Rock, 1987; Godshalk, Swineford, & Coff-
man, 1966); however, they have long been resisted by the community of writing
educators because of their impact on curriculum and instruction. Such short-an-
swer formats divert the focus of instruction away from student experiences with
writing extended text.
Thus, another benefit of computer-based analyses of features of writing is
the ability to derive these measures from samples of extended writing rather
than from short answer or multiple-choice formats. This could provide a richer
portrait of writing achievement without sacrificing the emphasis on the creation
of complete texts.
THE OUTCOME
For the 2011 NAEP, the framework development committees have recommend-
ed a focused holistic scoring system with components tailored to the three pur-
poses to be assessed (to explain, to persuade, and to convey experience). Raters
will be trained to attend to the development of ideas, to organization, and to
language facility and use of conventions, all as appropriate and relevant to the
purpose and audience of each task.
The new framework also envisions a “Profile of Student Writing” that would
examine in more detail each of these three components. The profile will rely to
the extent possible on measures that can be computed automatically from the
word processed writing samples, but will also include analytic scoring of a sub-
sample of student writing for features that cannot be derived from computerized
text analyses.
172
Issues in Large-Scale Writing Assessment
At the same time, writing plays a role in virtually all of the other subject area
assessments in NAEP. Both short and extended constructed responses comprise
major sections of the current assessments in science, history, geography, civics, and
reading, as well as the frameworks for new assessments in economics and foreign
languages. Rubrics in these assessments bear little similarity to the rubrics in the
writing assessment, however, often emphasizing listing of specific content rather
than the construction of an argument or explanation. This creates an artificial
separation of writing from content knowledge. As Hillocks (2002) pointed out
in his critique of state writing assessments, one of the biggest problems in many
assessments is the lack of a substantive content base on which to base the writing.
Without a content base, much of the writing that results is formulaic and shallow .
THE OUTCOME
For the 2011 NAEP, the framework committees have recommended a continued
focus on generic, easily accessible content, including short reading passages, visual
stimuli, or graphics. The one major change in the content of the writing tasks is
the recommendation that students at Grades 8 and 12 be allowed to choose the
genre or form they consider most appropriate to the audience and purpose speci-
fied in the prompt. The framework recommends pilot-testing items in a variety of
formats (with form specified, without form specified, and with a choice of forms
specified) in order to better understand the interaction between purpose, choice of
genre or form, and student performance in an assessment context.
Other large-scale assessments vary in the degree to which they rely on gener-
ic, easily accessible content. Although many use items very similar to those in
NAEP, others, such as New York State, base writing on extended reading pas-
sages, or include at least some classroom-based writing as part of the assessment
(Kentucky, Vermont). A more general issue for assessment developers is whether
it would be useful to increase the content load of student writing prompts, and
if so, how this could be done within current assessment frameworks or through
extensions of them. One possibility, particularly if writing and other assessments
become computerized, would be through the adoption of some common met-
rics for assessing quality of writing across assessments in different content areas.
173
Applebee
minutes for completing forms to nearly 30 minutes on some essay tasks. The move
to BIB spiraling in the 1984 assessment reduced the maxi-mum time to 15 min-
utes. Beginning with the 1992 assessment, this was increased to 25 minutes (with
a subset of 50-minute writing tasks that was eliminated in the 2002 assessment).
Two issues usually dominate discussions of writing time: Do the results mis-
represent overall writing achievement because students have too little time to
write? And does the limited time allowed penalize some groups of students,
particularly those whose classrooms have emphasized an extended process of
writing and revision? (Conversely, will extended time frustrate lower achieving
students and exacerbate achievement gaps?)
The issue of time has been driven by a tension between the constraints of
assessment and the conventional wisdom on instruction. One of the accomplish-
ments of the writing process movement in instruction was to remind teachers and
students that writing takes place over time—that there are identifiable strategies
for generating ideas, drafting, revising, editing, and sharing that shape and reshape
a final written text. During the past 30 years of writing assessment, the propor-
tion of teachers claiming to emphasize process-oriented approaches to writing in-
struction has risen sharply; by 1998 it was central to the instruction of 70% of
fourth- grade teachers surveyed, and used to supplement instruction by another
28% (Applebee & Langer, 2006). Comparable figures were reported by Grade 8
and 12 students in the 2002 assessment. (Background questions and grade levels
at which they are asked vary from assessment to assessment so there is no single set
of data on which to draw.) Given the constraints of large-scale assessment, NAEP
has always emphasized that the writing assessment focuses on first-draft writing (as
do the College Board and ACT in their college entrance examinations).
Given the overall design of the assessment, when NAEP has included
50-minute tasks the trade-off has been these tasks have not been scalable. (With
a 50-minute prompt, each student completes only one task, so interrelationships
among tasks cannot be determined.) In 1998,the results of these longer tasks
do not seem to have even been reported. Previous NAEP studies of the im-
pact of additional time have yielded mixed results. One special study compared
11th-graders’ performance on a persuasive writing task given in 16- or 50-min-
ute time blocks but mixed together for scoring with identical rubrics. As com-
mon sense might suggest, the students who had more time for writing scored
higher—although the gain was less than might have been expected: 45.4% pro-
duced adequate or better responses in 50 minutes, compared with 33.8% in
the 16-minute format. The benefits of extra time were not equally distributed
among students, however; the extra time made little difference to the weak-
er writers, increasing the performance gap between the two groups(Applebee,
Langer, & Mullis, 1989). The 1992 assessment reported results for 50-minute as
174
Issues in Large-Scale Writing Assessment
THE OUTCOME
For the moment at least, the NAEP writing assessment remains constrained
by a 25- or 30-minute format. Other large-scale assessments have the option
to explore formats that go well beyond this, however, and to investigate the
effects of variations in time and administrative procedures on student perfor-
mance. New York State, for example, provides substantive material for students
to read and write about, using extended, 3-hour time blocks. Kentucky pairs
classroom-based writing with an assigned task, and also insists that some of the
classroom-based writing come from subject areas other than English. Hillocks
(2002) commented favorably on both of these assessments in his critical look
at the quality of writing elicited by various approaches to writing assessment at
the state level.
175
Applebee
THE OUTCOME
The 2011 NAEP writing framework recommends typical accommodations such
as large-print booklets, extended time, or one-on-one testing when needed. It
also emphasizes item-development procedures that will ensure that every item is
presented in a simple and clear format accessible to all students.
A move to a computer-based administration for NAEP and other large-scale
assessments opens up the possibility of more tailored accommodations in the fu-
ture, however. By taking advantage of the computer platform, future assessments
might be able to individualize such factors as reading load and vocabulary level
in ways that are not possible with paper-and-pencil assessment booklets. Assess-
ment developers need to continue to give serious consideration to the effects of
accommodations for poor readers and English-language learners on the overall
content of the assessment, and look for alternatives that might provide a richer
array of assessment options for all students.
176
Issues in Large-Scale Writing Assessment
CONCLUSION
The framework for the NAEP writing assessment has evolved significantly over
the years, in the nature of the writing prompts, in the time available for each task,
and in its emphasis on rhetorical features such as audience and purpose. The issues
considered in developing a new framework for the 2011 writing assessment have
no easy answers, but the changes recommended for 2011 represent an updating
that reflects recent changes in scholarship and practice, and that will also return
NAEP to its position as a leader in assessment practice and assessment technology.
The most significant change involves the movement from paper-and-pencil book-
lets to a computer-based assessment, which carries with it potential changes in
many different aspects of the assessment: in the underlying construct that is being
assessed, in possibilities for analyzing the writing samples and reporting on student
performance, and in adaptive testing. The challenges will be large, but the oppor-
tunity for improving our understanding of student performance is equally large.
States and other groups developing writing assessments will have to confront
similar issues in the design of their own assessments, though the particular an-
swers they reach will vary in response to their varying purposes and constraints.
REFERENCES
Applebee, A. N. (1974). Tradition and reform in the teaching of English: A history.
National Council of Teachers of English.
Applebee, A. N. (2005). NAEP 211 writing assessment: Issues in developing a framework
and specifications. National Assessment Governing Board.
Applebee, A. N., & Langer, J. A. (2006). The state of writing instruction in America’s
schools: What existing data tell us (A report to the National Writing Project and the
College Board). Center on English Learning and Achievement.
Applebee, A. N., Langer, J. A., & Mullis, I. V. S. (1989). Understanding direct writing
assessments: Reflections on a South Carolina writing study. Educational Testing Service.
Applebee, A. N., Langer, J. A., Mullis, I. V. S., Latham, A. S., & Gentile, C. A. (1994).
NAEP 1992 writing report card. U.S. Government Printing Office for the National
Center for Education Statistics, Office of Educational Research and Improvement,
U.S. Department of Education.
Bereiter, C., & Scardamalia, M. (1987). The psychology of written composition. Erlbaum.
Breland, H. M., Camp, R., Jones, R. J., Morris, M. M., & Rock, D. (1987). Assessing
writing skill. College Entrance Examination Board.
Britton, J., Burgess, T., Martin, N., McLeod, A., & Rosen, H. (1975). The development
of writing abilities. Macmillan Education.
College Board. (2008). Strategies for success on the SAT essay. [Link]
com/student/testing/sat/pre_one/essay/[Link]
Cope, B., & Kalantzis, M. (Eds.). (1993). The powers of literacy: A genre approach to
teaching writing. University of Pittsburgh Press.
177
Applebee
178
Issues in Large-Scale Writing Assessment
179
Applebee
FRAMEWORK REFERENCES
Norris. E. L. (Ed.). (1969). Writing objectives. Committee on Assessing the Progress of
Education.
180
Issues in Large-Scale Writing Assessment
181
CHAPTER 6.
THE MICROPOLITICS OF
PATHWAYS: TEACHER EDUCATION,
WRITING ASSESSMENT, AND
THE COMMON CORE
J. W. Hammond
University of Michigan
Merideth Garcia
University of Michigan
184
The Micropolitics of Pathways
185
Hammond and Garcia
and how reformers have sought to manage them. In the past century, new assess-
ment technologies, including writing scales, rubrics, normed holistic scoring,
and automated essay scoring, have emerged in response to the perceived prob-
lem of heterogeneity (i.e., unreliability) across teacher assessments of student
writing (Elliot, 2005). New pathway-related reforms have likewise proliferated,
promising increased consistency, commonness, and standardization.
Still, educational complexity is not so easily tamed; the pathways that re-
forms put in place are seldom as stable and standardized as intended. This is as
true for postsecondary reforms as it is for those primarily targeting K-12 edu-
cation. To give one recent, community college-focused example, Bailey, Jaggars,
and Jenkins (2015) suggested student outcomes can be raised through adoption
of what they call “guided pathways,” which provide students with directive guid-
ance and a more focused curriculum—using faculty and advisors to coordinate
(or guide) students “instead of letting students find their own paths through
college” (p. 16). We might think of the guided pathways approach as something
of a spiritual successor to the CCSS—at least to the extent that both reforms
propose to manage the complexity of the curricular paths students take. Finding
much promise in the guided pathways idea, Rose (2016) nevertheless reminded
us of “how messy and unpredictable the process of reform can be” (para. 12),
noting that reforms relying on articulation between faculty members can run
into particular challenges: “faculty can have quite different beliefs about con-
cepts like ‘improving students’ lives.’ And some of these differing beliefs can
present resilient barriers to change” (para. 18). Reform initiatives can only stan-
dardize so much; where their pathways lead is always partly contingent on the
assumptions and aims of the teachers who maintain them.
Rose (2016) underscores that the politics of pathways—our overt contestation
over the paths structured for students—can be complicated or confounded by
the ways educators interpret and engage with reform initiatives, something Blase
(2005) has called the micropolitics of educational change. The term “micropolitics”
has been used in education research to account for the heterogeneity, dissensus,
and complexity at the core of education work. In Achinstein’s (2002) words, “Mic-
ropolitical theories . . . spotlight individual differences, goal diversity, conflict, uses
of informal power, and the negotiated and interpretive nature of organizations” (p.
423). Adopting a micropolitical perspective sensitizes us to the idea that educator
behavior is not fully shaped and determined by the structures educators participate
in; instead, educators partly shape those structures through “the use of formal and
informal power . . . to achieve their goals” (Blase, 1991, p. 11; see also Achinstein,
2002; Blase, 2005). This interpretive influence of teachers touches virtually every
aspect of educational practice. The complex process of socializing new teachers is
micropolitical (Kelchtermans & Ballet, 2002), as is the messy act of collaboration
186
The Micropolitics of Pathways
(Achinstein, 2002; Adamson & Walker, 2011) and the often-overlooked strategic
work of interpreting standards and reforms (Blase, 2005; März & Kelchtermans,
2013; see also Dover, Henning, & Agarwal-Rangnath, 2016).
Yet despite disciplinary understandings that teachers are necessary media-
tors of educational change (Blase, 2005; Gallagher, 2011), and that teachers’
beliefs and perceptions affect their teaching (Hillocks, 1999), teacher interpre-
tation and negotiation of the CCSS remains understudied. Work of this kind
is essential, for “there are more scholars theorizing about the CCSS than those
who are actually collecting and analyzing data from teachers who are respon-
sible for implementing the standards” (Ajayi, 2016, p. 3). To date, larger-scale
research on teacher perceptions of the CCSS suggests teachers hold broadly pos-
itive views of the CCSS (Matlock et al., 2016), and of the writing and language
standards specifically (Troia & Graham, 2016). Even several years into CCSS
adoption and implementation, teachers report widespread unfamiliarity with
the ELA CCSS-related assessments (Troia & Graham, 2016; also Ajayi, 2016);
they also hold conflicted views that those assessments are “more rigorous than
their prior state writing tests” but “fail to address important aspects of writing
development and do not accommodate the needs of students with diverse writ-
ing abilities” (Troia & Graham, 2016, p. 1740; Murphy & Haller, 2015; Ruec-
ker et al., 2015). Perhaps understandably, in trying to take the general measure
of emerging teacher engagements with the CCSS, this existing scholarship has
focused on broad patterns in teacher perceptions of the CCSS, seldom digging
deeper into the messiness of these perceptions—or how teachers micropolitically
engage with and locally instantiate the CCSS.
METHODS
Participants
Participants worked together at three different high school sites in professional
triads composed of field instructors, mentor teachers, and student teachers at
each site. Participants and sites associated with them were assigned pseudonyms
beginning with the same letter, chosen to alliteratively signal and clarify rela-
tionships. Sites and participants associated with Triad A all begin with “A”—
Amanda, Anne, and Alicia at Allendale High; Triad B—Barbara, Brenda, and
Brandon, at Bardstown High; and Triad C—Caleb, Cathy, and Cal at Clayville
High (Appendix, Table 6.1).
Field Instructors. Recruitment began at Midwestern University by email-
ing field instructors in its ELA teacher education program. Three instructors
expressed interest in participating—Amanda, Barbara, and Caleb (Appendix A,
187
Hammond and Garcia
Table 2). All had previously been secondary ELA teachers. We asked them to
recommend participants from student and mentor teacher pairs in their cohorts.
Field instructors facilitated communication between the university and mentor
teachers, supported student teachers in a weekly course relevant to placement
experiences, conducted at least three classroom observations of student teachers,
and attended beginning- and end-of-semester meetings with student and men-
tor teachers. They completed evaluations required for teacher certification, and
frequently wrote recommendation letters for students’ applications to teaching
jobs and graduate programs.
Mentor Teachers. This study includes three mentor teachers—Anne, Bren-
da, and Cathy (Appendix A, Table 6.3)—from among those recommended by
our field instructors. Mentor teachers opened their classrooms to student teach-
ers and field instructors, providing student teachers the opportunity to observe
instruction daily and, for part of the year, to take responsibility for two or more
classes. They guided student teachers in preparing lessons according to school
and state requirements, and helped student teachers apply abstract content and
procedural knowledges to real workplaces. Mentor teachers completed two for-
mal evaluations of student teacher performance for inclusion in the student
teacher’s certification application.
Student Teachers. We draw on data from three student teachers—two (Ali-
cia and Brandon) enrolled in the undergraduate teacher certification program,
and one (Cal) in a Master’s level certification program (Appendix A, Table 6.4).
The Master’s program placed students for the entire school year, while the un-
dergraduate program placed students for one semester. Student teachers in both
programs observed mentor teachers daily, coordinating with them to plan and
enact instructional units (usually spanning four to six weeks) in at least two
classes. They submitted unit plans to their field instructors for feedback and
evaluation, and scheduled their field instructors’ observations to showcase devel-
oping instructional skills.
School Sites
Secondary school sites were located in the same state as Midwestern Universi-
ty, a public Research I university whose teacher education program is accred-
ited by the Teacher Education Accreditation Council. While a CCSS adoptee
throughout data collection and the writing of this article, this state articulated
its standards to and through a standardized test other than the Partnership for
Assessment of Readiness for College and Careers (PARCC) and Smarter Bal-
anced Assessment Consortium (SBAC) assessments. During the data collection
period, Midwestern University hosted 22 secondary-level student teachers and
188
The Micropolitics of Pathways
partnered with a number of secondary school sites, including the three repre-
sented in our study.
Allendale High. Alicia described Allendale High as having a “relatively ho-
mogeneous” student population, and Anne explained “we are about 1200 stu-
dents, 9 through 12. We serve primarily suburban, upper-middle class or afflu-
ent families.” She added, “primarily we’re a school full of white students, but we
do pull from a lot of other populations,” and that “Allendale tends to pull from
students whose parents have found a way to land in the neighboring city and get
themselves into the district.” State information indicates only 9.7% of the test-
ing population scored not-proficient on the statewide standardized assessment
test given to the 268 11th-graders enrolled in Allendale during the 2014-2015
school year.
Bardstown High. Brenda said that Bardstown High has “about 1900 stu-
dents there, so it’s large . . . and it’s pretty homogeneous,” serving a “mostly
white” and “middle to middle-upper class” student population. She explained
that parents selected Bardstown because “the scores are very high here . . . last
year we had the number one AP scores in the state.” Brandon concurred that,
“It’s one of the best schools in the state, and it’s probably, I would say, probably
considered one of the best schools in the Midwest for public schools.” State
information indicates 12.1% of the testing population scored not-proficient on
the statewide standardized assessment test given to the 464 11th-graders en-
rolled in Bardstown during the 2014-2015 school year.
Clayville High. Cathy described Clayville High as “a small alternative ed-
ucation setting with at-risk students in an urban setting. We have about 235
students total that range in age from 14 to 25.” In Cal’s account, Clayville pri-
marily served students who “have been kicked out or for other disruptive reasons
have left their high school, and they are now here. It’s really homogeneous. 99%,
just about, African American. All are high needs, high trauma.” State informa-
tion indicates that 80.6% of the testing population scored not-proficient on the
statewide standardized assessment test given to the 44 11th-graders enrolled in
Clayville during the 2014-2015 school year.
189
Hammond and Garcia
data iteratively and collaboratively to tease out nuanced differences within and
across participant responses. This analytic approach was supplemented with
memos and notes shared between and reviewed by both researchers. At all
analysis stages, we sought to document and learn from the diversity evident in
participant accounts, rather than evaluate their comparative merits and omis-
sions. Consequently, our work does not account for the myriad effects the
CCSS might, in reality, have had—on pathways, curricula, and assessments—
beyond those participants raised. Evaluating teacher perspectives and casting
our analytic focus beyond them are crucially important projects, but they are
not ours here.
FINDINGS
Sensitive to the intended pathway-consolidating function of the CCSS, partic-
ipants described the CCSS as having the potential to put teachers and students
across the country (in Barbara’s words) “on the same page”—a phrase used also
on the CCSS’s official webpage: “With students, parents, and teachers all on the
same page [emphasis added] and working together toward shared goals, we can
ensure that students make progress each year and graduate from high school
prepared to succeed in college, career, and life” (Common Core State Standards
Initiative, “Read the Standards” n.p.). Yet while each of our participants report-
ed using the CCSS in some way in their curricular planning, none held identical
perceptions of the CCSS, and none described using it (or locally assessing it)
in quite the same way. Importantly, none of our participants reported that the
CCSS fully determined the educational pathways their own students traveled
down. Instead, participants reported micropolitically interpreting and repurpos-
ing the CCSS—drawing strategically on the standards to supplement and sup-
port the local pathways they already had in mind for students.
Cathy, for instance, asserted that standards themselves are—without local
curation, negotiation, and interpretation—improper guides for the educational
pathways traveled by students. “I think they’re [the CCSS] too restrictive,” she
told us, adding:
I think standards in general are too restrictive. The needs of
students change based on the environment the students live in
and the environment that they’re going to be going into. If it’s
a college prep school, standards might be, you know, a little
bit, there should be higher expectations. Students that are just
going to go out into the world, they just want to find jobs, and
they just want their high school diploma so that they can have
190
The Micropolitics of Pathways
191
Hammond and Garcia
by providing teachers with the terms necessary to voice what they were al-
ready doing. She told us, “The CCSS seems to allow quality teachers access to
the language that describes the skills that they were likely already teaching in
their classroom all along.” Underpinning this perception was the idea that the
standards represented little more than what good teachers are always already
doing in their classrooms—albeit, perhaps, without the vocabulary to make
their positive practices known to stakeholders. “My classmates and I, generally,
agreed that it [the CCSS] is a document that only suggests skills and lessons
that ‘good teaching’ should have anyway,” she wrote. This sense was shared by
field instructors Barbara and Caleb, the former of whom reported that, among
teachers, “nobody’s bothered by it [the CCSS]”; nodding to the communicative
uses of the standards, she asked, “what’s the big deal? It’s nice that everybody’s
on the same page.”
As a general touchstone for talking about “good teaching,” the common
language afforded by the CCSS was micropolitically useful to our participants,
who drew on it to warrant their instructional decisions. Brandon, for his part,
described the CCSS as micropolitically “beneficial” as a kind of professional lin-
gua franca, enabling him to enter disciplinary conversations about professional
expectations and practices. He confided:
When I got to student teaching, when my mentor teacher
started talking about Common Core, I wasn’t like, deer in
the headlights, or anything like that. I could discuss it with
them. I talked to the principal a lot, and we had departmental
meetings and things like that, and I wasn’t just sitting there
completely with a blank stare on my face. I could contribute
to the conversation . . . .
The communicative uses of the CCSS were evident also in Brandon’s lesson
planning. When he and his mentor teacher developed a unit plan for Huckle-
berry Finn, they started with themes that they wanted to emphasize—“such as
friendship, love, and trust, empathy. Those are the things that again, I just think
they’re so important for kids to learn about”—then moved to the final assess-
ments they would locally implement: a group presentation on banned books, a
multiple-choice test on Huckleberry Finn, and a portfolio “compiled of a bunch
of things” students had composed throughout the unit. After the text, themes,
and summative assessment were settled, Brandon and his mentor “went through
all the Common Core Standards” and matched them to lessons where “they’ll
fit in.” In other words, Brandon strategically curated the standards, selecting
those that best matched his goals and assessments to signal compliance. Cal, too,
discussed the CCSS in terms of its communicative uses. He relied on the CCSS
192
The Micropolitics of Pathways
not to define the curricular path he paved for students, but instead to signal to
outsiders that this path was an appropriate, standards-approved one.
Cal’s sense, though, was less that the CCSS opened professional doors than
that it closed opportunities for professional censure. Speaking of the CCSS, Cal
described the negative assumptions that might accompany failure to draw on the
vocabulary of the CCSS: “if you don’t necessarily have it embedded into, like, you
know your lesson plans and everything you’re doing, then you’re not necessarily an
effective teacher.” Like Brandon and Alicia, Cal’s facility with the CCSS provided
him a way to signal professional growth to the field instructor who supervised his
lesson planning. As Cal tells us, though, this lesson planning was a subtle rhetor-
ical task: “while he wants to make sure that I know them [the standards], he also
wants to make sure that I don’t know them, if that makes like any sense. It’s like,
‘Use them, but don’t necessarily be pigeon-holed by them.’” Instead of letting the
CCSS narrow his curriculum, Cal mined the CCSS for pieces that were relevant,
“tak[ing] like bits and like two or three of them, varying on the grade level, and
apply[ing] them for my unit plan on a weekly basis.” What his description touched
on was a pattern in the way our participants engaged with the CCSS: They rhetor-
ically used the standards, and actively resisted being used by them.
Cal’s mentor teacher, Cathy, also described curating the standards. She said
of the CCSS, “After reading them, they’re pretty straightforward, and they can
be kind of twisted however you need to use them.” Cathy went so far as to sug-
gest that rhetorical engagement with the CCSS—strategically selecting, inter-
preting, and negotiating them—ought to be “a mandatory class,” because such
training would facilitate pedagogical self-awareness and develop a capacity to
communicate about the educational pathways locally maintained in the class-
room. Cathy argued teachers do not need to demonstrate equal fidelity to all
educational standards, but she noted:
It’s important to know what you don’t like. And it is im-
portant to be able to explain why. . . . I think every teacher
needs to be educated to the point of being experts on these
[the standards] because that’s the only way we can get around
them if we need to.
No classroom is an island; educational pathways are the shared jurisdiction
of multiple stakeholders and, as such, must be negotiated. Using the standards
as a shared local language secured for our participants a kind of self-determi-
nation that could only come with persuasively communicating with outside
stakeholders.
A Common Language for Improving Pedagogy. Indeed, like many educa-
tors (including one of the present writers), Cathy is both a teacher and a parent.
193
Hammond and Garcia
194
The Micropolitics of Pathways
different groups of teachers related the same essential practices and processes.
Importantly, Brenda’s description of the CCSS here positioned it less as a reform
of practice than of what we call that practice—with teachers adopting a new,
shared convention for existing staples of their local work.
This terminological shift alone was one that Brenda believed beneficial. She
held that educational pathways already in place would be more easily navigable
by students, who would “have a language that transfers from teacher A to teach-
er B.” This affordance of the CCSS was one voiced regularly in the interviews
we conducted, with participants noting that a shared vocabulary provided ed-
ucational pathways at least superficially an increased kind of intelligibility and
coherence for students. When Brenda discussed movements between the class-
rooms of teachers A and B, she was thinking specifically of vertical course align-
ment, with students advancing from “English 11A” to the next course (“English
11B”) in a sequence. However, other participants (like Anne and Caleb) noted
the benefits of this shared vocabulary for students who, in Anne’s words, “have
to be mobile.” As Caleb put it, an easily-navigable lateral movement from school
to school is essential “so you can live in a country where people move a lot.”
195
Hammond and Garcia
skill that I don’t know if our students have.” Rather than calibrating her classes
to CCSS-aligned large-scale tests, though, Anne reported another (more local)
way the CCSS was instantiated: formative classroom assessment. “I don’t really
prioritize essay grading. Instead, I try to prioritize whatever small chunks I can
do and give them [students] feedback on immediately,” Anne admitted. Such
a trade-off was consistent with one of her personally-held “goals as a teacher[,
which] is to get feedback to [her] kids in a more meaningful and timely way.”
She took CCSS-adoption as an occasion to develop a local system for appraising
student writing—in small chunks, rather than essays—that reflected her own
aims and beliefs as an educator.
Having observed an SBAC practice implementation, Brenda thought the
test “was fabulous. I thought it was amazing.” Praising the test for its explicit
alignment, Brenda regretted the state’s decision to pursue a different—poten-
tially less-explicitly aligned—testing system, stating, “I was really sad that we
didn’t go to it.” Even voicing this support, Brenda described her classroom not
as caught in the thrall of large-scale standardized test preparation, but instead as
backwards planned (Wiggins & McTighe, 2005) against meaningful, locally-de-
veloped summative goals. Brenda’s school developed and implemented local
CCSS-aligned “common final exams in the core classes” like English. In Brenda’s
own classes, essays were assessed against a rubric targeting “three focus correction
areas per paper”—areas that “come from the Writing Core.” Brenda’s writing
rubric was reframed, not narrowed, to reflect the CCSS’s commonly-shared vo-
cabulary: “to be honest, I don’t think it’s [the rubric] anything extraordinarily
different [from past, pre-CCSS rubrics], but the wording and the language is
going to match the wording and the language on the Common Core. In fact, it
might even list the strand.” In this way, the CCSS was micropolitically leveraged
to signal (rather than determine) local writing assessment priorities—marking,
in new terminology, the pathway Brenda had already set.
Cathy, too, reported that state-mandated large-scale tests had “gotten much
more difficult” in response to the CCSS, but faulted the CCSS for misalignment
to her students’ needs: “I really think it focuses too much . . . towards the testing.
I think that impacts students a lot, because our kids, . . . they need to be ready
for the real world, and Common Core does not always address those needs.”
Departing from other mentor teachers, Cathy said of the CCSS, “Yes, it’s raising
the bar to a higher level, but sometimes students need more than that.” Here,
“raising the bar” was equated to something decidedly less than what students
needed. Cathy identified the “need to write an argumentative essay” as some-
thing that, perhaps, “isn’t really important for their [her student’s] life needs.
Can they write a resumé? That’s more important. Do they know how to look
up a job application and fill that out? That’s more important for these kids.” In
196
The Micropolitics of Pathways
this rendition, the CCSS pulled too far in the direction of what were perceived
to be postsecondary writing needs at the expense of more immediately valuable
(professional) writing-related skills.
“I don’t look at the standards as standards, I look at them as suggestions,”
Cathy claimed; “They’re a good place to start, but they can go either way. You
can advance them, take it to the next step, or you can cut things out.” In a way,
our study’s mentor teachers all practiced what Cathy preached here—micropo-
litically mediating the CCSS in the service of preexisting commitments and be-
liefs regarding assessment. They used the CCSS as license to increase and explore
their preferred formative assessment strategies (Anne), leveraged the CCSS’s vo-
cabulary to validate assessment practices they believed effective (Brenda), and
strategically determined which standards were emphasized, how, and to what
ends (Cathy).
Student Teachers. Echoing their mentor teachers, student teachers discussed
negotiating and interpreting the CCSS by means of local assessment. Asked if
the CCSS dictated or shaped classroom evaluations of student performance,
Alicia stressed the multiple ways (beyond standardized testing) assessment could
locally instantiate the CCSS:
it is all about the type of interpretation you take toward the
CCSS. I believe that our assessments and classroom evalua-
tions of student performance should be loosely based off of
the skills that are offered by the CCSS. However, I do not
think that this means that teachers need to merely provide a
standardized test assessing these skills. Quite the opposite, in
fact. I believe that teachers should provide final assessments
that ask students to use a wide range of tasks the CCSS focus-
es on (e.g., using evidence to support a claim, determining the
central idea of a text, etc.).
Consistent with this line of thinking, Brandon speaks favorably about
state-mandated standardized testing and its ability to help “comparison through-
out the US education to be a little more accurate,” but was careful to claim that
such testing ought not drive curriculum, because “if you master the Common
Core Standards, you’re going to do well on standardized testing.” Instead, when
he discussed the local meaning the CCSS had for guiding assessment, Brandon
thought not of state-mandated large-scale tests, but instead—like Brenda—talk-
ed of teachers in a department sharing common tests “based on the Common
Core Standards.” Appropriate assessment was a local matter; as a baseline for
describing “good teaching,” the CCSS provided local actors the vocabulary nec-
essary to collaboratively develop (and discuss) shared local tests.
197
Hammond and Garcia
In stark contrast to Alicia and Brandon, Cal believed the CCSS was “dumb-
ing down education,” and worried formative and individualized assessments
might be crowded out because “people could be kind of pigeon-holed to only
have a couple of assessments that actually show the students are following the
Common Core standard.” Crucially, Cal’s concerns related less to his classroom
than those of other teachers more “willing to kind of take it [the CCSS] by law.
. . .” By contrast, Cal’s classes prioritized individualized writing assessment (“dif-
ferent benchmarks” indexed to individual students), privileging feedback-rich
writing instruction grounded in the “notion of writing as rewriting”—all while
carving out space for preparing students to create resumés (a local need identi-
fied by Cathy). Even under the CCSS, essay writing and assessment could serve,
for Cal, as tools for teaching deeper life skills. “‘Listen,’” he exhorted an imag-
ined audience of his students:
“Your writing can be something that can be extremely superb,
but it is something . . . that you have to be willing to work
on, meaning that you need a work ethic for [it], and it has to
be something that you have to realize that you have to accept
criticism for and seek it out for that.” And, hopefully, again,
[students will] kind of retain that to their life [sic].
Moving between the field instructors at Midwestern University and the men-
tor teachers at their high school placements, student teachers micropolitically
negotiated the CCSS in the process of developing assessments they believed best
supported student learning. In response to the CCSS’s standardizing potential,
they insisted on the need for multiple measures of student progress (Alicia);
imagined that collaboratively-developed, standards-aligned assessments would
naturally prepare students for success on large-scale standardized tests (Bran-
don); and advocated individualized, rather than standardized), assessment, tai-
lored to student needs and preexisting pathways (Cal).
Field Instructors. Field instructors had limited direct knowledge of the
CCSS-related state-mandated large-scale standardized tests and instead reported
on what they gleaned secondhand in placement sites. Interestingly, though, all
three field instructors expressed some form of support for the CCSS’s large-scale
standardized tests—their perspectives underpinned by a more general sense that
the CCSS represented little more than (in Caleb’s words) “a nice minimum”
good teachers always already meet: “I remember the first time I read through
them [the CCSS], . . . my feeling was, ‘Well if you’re not doing these things,
what the heck are you doing in English class?’ These are just the things that you
should be doing.” Broadly speaking, the field instructors thought of the CCSS as
aligning with their own micropolitical sense of what was normal or appropriate
198
The Micropolitics of Pathways
for “good” teaching. With this understanding in mind, field instructors expected
the effects of CCSS-aligned standardized testing to be positive or unobtrusive—
altering local practices only in cases of curricular negligence or ineptitude.
Within this context, Caleb referenced the idea that tests drive curriculum—a
familiar concern in writing assessment scholarship (e.g., Hillocks, 2002)—but
regarded this possibility as a feature rather than a flaw. Comparing the sample
CCSS-aligned large-scale tests with previous state tests he had observed as a high
school teacher, he told us:
This [CCSS-related assessment] seems more difficult, so
[much] more rigorous, more focused on critical thinking and
synthesis of information, and I know that a lot of times, tests
drive curriculum, so, I mean, I think the hope is that curricu-
lum will become more rigorous than they were—than it was
under [previous] state standards.
Importantly, though, when Caleb spoke about tests driving curriculum, he was
not envisioning rote test preparation; his sense was, to the contrary, that quali-
ty instruction was already aligned to (and preparatory for) CCSS-aligned large-
scale tests. Considering himself “ideologically aligned” with what he considered
the CCSS’s emphasis on teaching with “big questions” in mind, Caleb claimed
“teaching those types of lessons well ensures that kids are going to do fine on the
assessment.” Caleb attributed the controversy over the CCSS large-scale tests to
“misconceptions” in the wake of a weak introduction: “the way it [the CCSS] was
sort of rolled out and implemented didn’t really promote a lot of clarity, and I think
some parents are refusing to let their kids test, and some states are opting out.”
For their parts, Barbara and Amanda—neither of whom had seen a CCSS-
aligned sample test—expressed regret that politics had complicated CCSS-
aligned testing. Barbara’s central complaint about the new testing regime seemed
to be that state-level political forces had been unwilling to commit to standards
and tests long enough for schools to gauge requirements and prepare adequate-
ly—a kind of politics of pathways that exchanged student futures for political
pride. The state, she argued,
was reluctant at first to go with the Core, and it’s like, ‘Every-
body else is on board, why . . .?’ The legislature again, feeling
like, ‘We have to be autonomous. We don’t need to have the
Core. We can have our own guidelines.’ It’s like, ‘Why?’ The
same thing with the testing . . . .
Amanda, by contrast, regretted that for all the CCSS’s promised common-
ness, its official large-scale tests remained plural, a promise of commonness
199
Hammond and Garcia
DISCUSSION
When this study’s participants communicated with one another about instruc-
tion and assessment, they invoked the CCSS as an articulating document. How-
ever, beneath the veneer of unity provided by the CCSS, we found substantive
disagreement both within and across teacher groups, indicating that pathway-re-
lated reforms and the consensus they seek to impose are always fraught with
local, micropolitical dissensus. For example, one might expect that field instruc-
tors would share an orientation toward the CCSS, using it to evaluate student
teachers’ lesson plans in a state that had adopted the CCSS. Yet among our field
instructors, Amanda insisted that every lesson plan build on a specific standard,
Barbara encouraged student teachers to plan around skills and match standards
to them retroactively, and Caleb reported that students attended more closely
to selecting texts than standards when planning. Consider, too, micropolitical
dissensus within Caleb’s Clayville triad: Where Caleb viewed the CCSS as intro-
ducing more rigor to the curriculum, Cal saw it as “dumbing down education,”
and Cathy explained that students needed more than what the standards offered.
More generally, participants reported developing locally-meaningful lessons and
assessments, and strategically curating the “common core” of standards to match
200
The Micropolitics of Pathways
201
Hammond and Garcia
202
The Micropolitics of Pathways
203
Hammond and Garcia
REFERENCES
Achinstein, B. (2002). Conflict amid community: The micropolitics of teacher
collaboration. Teachers College Record, 104(3), 421-455.
Adamson, B., & Walker, E. (2011). Messy collaboration: Learning from a learning
study. Teaching and Teacher Education, 27(1), 29-36.
Addison, J. (2015). Shifting the locus of control: Why the common core state
standards and emerging standardized tests may reshape college writing classrooms.
Journal of Writing Assessment, 8(1). [Link]
Addison, J. & McGee, S. J. (2015). To the core: College composition classrooms in
the age of accountability, standardized testing, and common core state standards.
Rhetoric Review, 34(2), 200-218.
Ajayi, L. (2016). High school teachers’ perspectives on the English language arts
common core state standards: An exploratory study. Educational Research for Policy
and Practice, 15(1), 1-25.
Bailey, T. R., Jaggars, S. S., & Jenkins, D. (2015). Redesigning America’s community
colleges: A clearer path to student success. Harvard University Press.
Blase, J. (1991). The micropolitical perspective. In J. Blase (Ed.), The politics of life in
schools: Power, conflict, and cooperation (pp. 1-18). Corwin Press.
Blase, J. (2005). The micropolitics of educational change. In A. Hargreaves (Ed.),
Extending educational change: International handbook of educational change (pp. 264-
277). Springer.
Bridges-Rhoads, S., & Van Cleave, J. (2016). #theStandards: Knowledge, freedom, and
the common core. Language Arts, 93(4), 260-272.
Clark-Oates, A., Rankins-Robertson, S., Ivy, E., Behm, N., & Roen, D. (2015).
Moving beyond the common core to develop rhetorically based and contextually
sensitive assessment practices. Journal of Writing Assessment, 8(1). https://
[Link]/uc/item/9325025k
Common Core State Standards Initiative. (n.d.). Frequently asked questions. http://
[Link]/wp-content/uploads/[Link]
204
The Micropolitics of Pathways
Common Core State Standards Initiative. (n.d.). What parents should know. http://
[Link]/what-parents-should-know/
Common Core State Standards Initiative. (n.d.). Read the standards. [Link]
[Link]/read-the-standards/
Dover, A. G., Henning, N., Agarwal-Rangnath, R. (2016). Reclaiming agency: Justice-
oriented social studies teachers respond to changing curricular standards. Teaching
and Teacher Education, 59, 457-467.
Elliot, N. (2005). On a scale: A social history of writing assessment in America. Peter Lang.
Elliot, N. (2016). A theory of ethics for writing assessment. Journal of Writing
Assessment, 9(1). [Link]
Gallagher, C. W. (2011). Being there: (Re)making the assessment scene. College
Composition and Communication, 62(3), 450-476.
Hillocks, G., Jr. (1999). Ways of thinking, ways of teaching. Teachers College Press.
Hillocks, G., Jr. (2002). The testing trap: How state writing assessments control learning.
Teachers College Press.
Jacobson, B. (2015). Teaching and learning in an “audit culture”: A critical genre
analysis of common core implementation. Journal of Writing Assessment, 8(1).
[Link]
Kelchtermans, G., & Ballet, K. (2002). The micropolitics of teacher induction.
A narrative-biographical study on teacher socialisation. Teaching and Teacher
Education, 18(1), 105-120.
Kliebard, H. M. (2004). The struggle for the American curriculum (3rd ed.). Routledge.
Lazarín, M. (2014). Testing overload in America’s schools. [Link]
org/wp-content/uploads/2014/10/[Link]
März, V., & Kelchtermans, G. (2013). Sense-making and structure in teachers’
reception of educational reform. A case study on statistics in the mathematics
curriculum. Teaching and Teacher Education, 29, 13-24.
Matlock, K. L., Goering, C. Z., Endacott, J., Collet, V. J., Denny, G. S., Jennings-
Davis, J., & Wright, G. P. (2016). Teachers’ views of the common core state
standards and its implementation. Educational Review, 68(3), 291-305.
Murphy, A. F., & Haller, E. (2015). Teachers’ perceptions of the implementation of
the literacy common core state standards for English language learners and students
with disabilities. Journal of Research in Childhood Education, 29(4), 510-527.
Poe, M. (2008). Genre, testing, and the constructed realities of student achievement.
College Composition and Communication, 60(1), 141-152.
Poe, M., & Inoue, A. B. (2016). Toward writing assessment as social justice: An idea
whose time has come. College English, 79(2), 119-126.
Rose, M. (2016). Reassessing a redesign of community colleges. Inside Higher Ed.
[Link]
pathways-model-restructuring-two-year-colleges
Ruecker, T., Chamcharatsri, B., Saengngoen, J. (2015). Teacher perceptions of
the impact of the common core assessments on linguistically diverse high
school students. Journal of Writing Assessment, 8(1). [Link]
item/0pq793rq
205
Hammond and Garcia
Stein, Z. (2016). Social justice and educational measurement: John Rawls, the history of
testing, and the future of education. Routledge.
Troia, G. A., & Graham, S. (2016). Common core writing and language standards and
aligned state assessments: A national survey of teacher beliefs and attitudes. Reading
and Writing, 29(9), 1719-1743.
Wiggins, G., & McTighe, J. (2005). Understanding by design (2nd ed.). ASCD.
206
The Micropolitics of Pathways
207
Hammond and Garcia
• What do you know about the organization that created the CCSS?
• Have you read the document? In what form (online, printed, con-
densed, complete)?
• Have you received any training or professional development in using
the CCSS? If so, describe what you took away from that experience.
• Beliefs and Attitudes Regarding the CCSS
• What value (if any) do you think the CCSS have for classroom teachers?
• What concerns (if any) do you have about how the CCSS might affect
classroom teachers?
• What value (if any) do you think the CCSS have for students?
• What concerns (if any) do you have about how the CCSS might affect
students?
• Briefly describe any additional ways in which you think the CCSS
might be valuable.
• Briefly describe any additional concerns you have about the CCSS and
its effects.
• Assessment and CCSS
• How do you evaluate (in class) whether students have met the CCSS?
• Do you think classroom evaluations of student performance are
shaped or dictated by the CCSS? How?
• What do you know about the state-wide tests in development for
measuring the CCSS?
• Have the state-wide tests for evaluating student progress changed in
response to the CCSS? How?
• Have you ever implemented standards other than the Common Core?
If so, do the CCSS seem the same or different from previous stan-
dards? Explain.
• Have the procedures evaluating you as a teacher been shaped or dictat-
ed by the CCSS? How?
• CCSS and Relevant Social Groups
• *How would you describe your field instructor’s knowledge of the CCSS?
• *How would you describe your mentor teacher’s knowledge and im-
plementation of the CCSS?
• How useful do you think the CCSS are for new teachers versus experi-
enced teachers?
• Do you think the CCSS play a different role in the education of students
in different kinds of programs (like advanced placement, regular classes,
or remedial classes)? If so, can you walk me through the differences?
• What else would you like me to know about your thoughts on the
CCSS?
208
CHAPTER 7.
WRITING ASSESSMENT,
PLACEMENT, AND THE
TWO-YEAR COLLEGE
Christie Toth
University of Utah
Jessica Nastal
Prairie State College
Holly Hassel
North Dakota State University
Joanne Baird Giordano
Salt Lake Community College
210
Writing Assessment, Placement, and the Two-Year College
2013; Lovas, 2002; Nist & Raines, 1995; Toth & Sullivan, 2016), including
emerging conversations about writing assessment, fairness, and social justice.
This dynamic may be shifting. A 2016 special issue of College English on writ-
ing assessment as social justice, edited by Poe and Inoue, featured two essays
focusing on community college students (Alexander, 2016; Naynaha, 2016).
Chapters in Poe, Inoue, and Elliot’s collection Writing Assessment, Social Justice,
and the Advancement of Opportunity also begin to address these gaps (Moreland,
2018; Toth, 2018a; 2018b). However, many of these studies demonstrate little
or no engagement with the scholarly literature in two-year college writing stud-
ies, and none were written by two-year college English faculty. While scholars
at all institution types can advance this important scholarly conversation, the
authors of this special issue of the Journal of Writing Assessment believe it is es-
sential that two-year college faculty participate as knowledge-makers as well as
beneficiaries of writing assessment research. Local context matters, and studies
conducted at two-year college sites by two-year college faculty can directly in-
form institutional work and improve student experiences and outcomes. These
studies can also make distinctive and important contributions to the broader
scholarly conversation about writing assessment.
The underrepresentation of two-year colleges in the writing assessment liter-
ature is an urgent ethical issue given the racial, ethnic, and socioeconomic diver-
sity of two-year college students. Nationwide, students of color attend commu-
nity colleges at disproportionately high rates: These institutions enroll 56% of
Native American undergraduates, 52% of Hispanic/Latinx students, and 43%
of African American students (American Association of Community Colleges,
2017). Likewise, many “minority-serving”—or New Majority—institutions
(e.g., historically or predominantly Black colleges, Hispanic-serving institutions,
and tribally-controlled colleges) are primarily associate-granting. Two-year col-
lege students are more likely than students at selective-admissions institutions to
come from low-income or working-class backgrounds and/or be among the first
generation of their family to attend college. They are also more likely to be older/
returning students, parents, veterans, immigrants or refugees, and/or students
with disabilities (Cohen et al., 2014). These groups of students have long been
systemically underrepresented, underserved, discouraged, and disadvantaged in
postsecondary education, reflecting and reproducing broader structures of social
inequality in the United States. Given these demographic realities, the scholarly
conversation about writing assessment, social justice, and the advancement of
opportunity must explicitly attend to two-year college contexts. Further, it must
do so with an awareness of the distinctive conditions of teaching and admin-
istering writing in these settings, including the missions and student popula-
tions served, constraints on institutional resources, writing instructors’ varying
211
Toth, Nastal, Hassel, and Giordano
AN ERA OF REFORM
Community college researchers and reformers often invoke low and inequitable
degree completion rates as a major motivation for enacting change (e.g., Bailey,
Jeong, & Cho, 2010; Barnett & Reddy, 2017; Scott-Clayton, Crosta, & Bel-
field, 2014; Zaback, Carlson, Laderman, & Mann, 2016). In 2016, only 39% of
students who enrolled at two-year colleges earned any kind of credential within
six years, and nationally, just 16% of entering two-year college students go on
to earn a bachelor’s degree (Shapiro et al., 2016). There are also unjust racial
disparities in these completion rates: Only 33% of Hispanic/Latinx students
and 26% of African American students who enroll at two-year colleges earn a
credential within six years, and just 11% of Hispanic/Latinx students and 9% of
African Americans who begin at two-year colleges eventually complete bachelor
degrees (Shapiro et al., 2017). Few argue that there is no need for reform; rather,
debates hinge on the nature, goals, and underlying ideologies of those reforms.
As Sullivan (2008, 2017) has reminded us, measuring “student success” at
open admissions institutions is a complex endeavor. Not all two-year college
students aspire to transfer or even earn degrees: Many are pursuing two-year
vocational, technical, or para-professional certifications, or they may be “test-
ing the waters” to see whether college is for them; others are dual-enrollment/
early college high school students or “reverse transfers” who have already attend-
ed four-year institutions and, for a variety of reasons, stopped out or changed
their goals. Degree-seeking students may also shift their aspirations as they gain
exposure to and experience with postsecondary education, and many students
find themselves facing financial pressures, life crises, or family and community
responsibilities that take priority over schooling, at least temporarily (Griffiths
& Toth, 2017; Sullivan, 2008, 2017). Furthermore, longstanding federal mea-
sures of completion rates have penalized community colleges by not including
part-time students or those who transfer to four-year-institutions in their success
metrics. When the Department of Education revised these criteria in 2017, it
found the 8-year combined graduation and transfer rate for community college
students was 60% (Carey, 2017).
Although they face many limitations and constraints, local and compara-
tively affordable open admissions two-year colleges provide a crucial point of
entry to students who would otherwise be unable to access (or re-access) pub-
lic postsecondary education. Many of these students are not making “market”
212
Writing Assessment, Placement, and the Two-Year College
choices between two- and four-year institutions, but rather between two-year
colleges or no college at all, or between two-year colleges and for-profit institu-
tions that may leave them deep in debt with unimproved employment prospects
(Toth, Calhoon-Dillahunt, & Sullivan, 2016). To the extent that writing assess-
ment—whether for placement, in the classroom, or as a requirement for exiting
required course sequences—functions to support or undermine student success
at two-year colleges, it plays a key role in either opening or foreclosing access to
learning, credentials, and, ultimately, socioeconomic mobility for some of the
least advantaged students in our postsecondary system.
Over the last few decades, calls among both state and federal policymakers to
improve student retention and degree completion have increasingly been framed
as a matter of institutional “accountability.” As Toth et al. (2016) have observed,
accountability measures often fail to acknowledge that “the academic playing
field is not level. An institution’s record of ‘success’ is largely shaped by its student
demographics and resources. The performance metrics are stacked in favor of
selective colleges and universities, particularly the most elite among them” (p.
401). This dynamic makes mounting pressures for performance-based funding
problematic. Perversely, such policies risk punishing under-resourced institutions
that serve under-resourced students by further denying them resources. They also
incentivize heretofore open admissions institutions to begin refusing entry to stu-
dents deemed unlikely to succeed (Toth et al., 2016), determinations typically
made based on those students’ performance on admissions or placement tests. In
this situation, placement assessment and other forms of standardized testing can
function to deny access—again, often the only available access to public postsec-
ondary education—to already disadvantaged students. Thus, the stakes of writing
assessment in the context of the accountability “movement” are high.
In recent years, the problem of degree completion at two-year colleges has at-
tracted the attention of mega-philanthropies like the Lumina and Gates founda-
tions, as well as higher education researchers who have made use of the influx of
funding from such organizations. These parties have been a driving force behind
many proposed policy reforms. Perhaps the most influential researchers have been
those associated with the Community College Research Center (CCRC) at Colum-
bia University’s Teachers College. Over the last decade, the CCRC has produced
a number of high-profile publications arguing that one major cause of departure
prior to degree completion is the amount of time many two-year college students
spend in developmental courses before they can enroll in credit-bearing college-lev-
el coursework (e.g., Bailey et al., 2010; Jaggars & Stacey, 2014): During the first
decade of the twenty-first century, 68% of two-year college students enrolled in at
least one developmental course (Chen, 2016). These researchers have found that, for
many students, the costs of the time and resources spent in developmental courses
213
Toth, Nastal, Hassel, and Giordano
214
Writing Assessment, Placement, and the Two-Year College
THEORIZING PLACEMENT
Placement is a writing assessment process unique to postsecondary education in
the United States (Haswell, 2004). While other countries use proficiency testing
for institutional admissions, many U.S. colleges use placement assessments once
students have already been admitted. In the nation’s open-admissions two-year
colleges, where students enter from a wide range of academic trajectories and
often have not taken any kind of admissions exam, placement assessment is
nearly universal. The rationale for placement hinges on the following argument:
1. Placement testing identifies students with the weakest writing abilities.
2. In order to boost those abilities, placement tests funnel students into spe-
cific classes or sections where instruction can be more manageable and
students can learn better.
3. Therefore, placement testing leads to improved student learning, reten-
tion, and completion.
This rationale is predicated on the algorithmic, decision-tree approach to
placement advanced by Willingham (1974) more than four decades ago (Figure
7.1). This binaristic, decontextualized model has become the tacit theory under-
girding most writing placement.
215
Toth, Nastal, Hassel, and Giordano
216
Writing Assessment, Placement, and the Two-Year College
217
Toth, Nastal, Hassel, and Giordano
218
Writing Assessment, Placement, and the Two-Year College
219
Toth, Nastal, Hassel, and Giordano
may seem like a “good enough” number for some, Smith (1993) argued, “For the
students and for the teachers, ‘very few’ is too many” (p. 192). This may be partic-
ularly true at open admissions two-year colleges, where underplacement into de-
velopmental courses can lengthen time to degree or discourage students from per-
sisting or even enrolling (Adams, 1993; Bailey, 2009; Bailey et al., 2010; Henson
& Hern, 2019; Nastal, 2019), while overplacement might increase the possibility
of student failure, costing them time and tuition dollars and potentially resulting
in academic probation or suspension.
According to Scott-Clayton (2012), high-stakes, single-score placement tests
were being used by 92% of two-year colleges at the beginning of the decade. As the
articles in this special issue of JWA demonstrate, we are only beginning to attend to
what Messick (1989) called the social consequences of these longstanding placement
practices. From the perspective of racial and socioeconomic equity, those conse-
quences are often profoundly troubling. As Morris, Greve, Knowles, and Huot
(2015) noted in their recent overview of book-length studies of writing assessment:
While there is little or no scholarship focused specifically on
two-year college writing assessment, it is important to rec-
ognize the important influence writing assessment can have
for students’ educational opportunities, especially at two-year
colleges, which enroll the majority of postsecondary under-re-
sourced students. (p. 120)
Furthermore, they argued, “Writing assessment can also be a critical issue for
two-year college identity and legitimacy” (Morris et al., 2015, pp. 120–121). Over
the course of our own careers, we have heard university-based colleagues speak
dismissively of community colleges on the basis of their purportedly uncritical
placement and “remediation” practices. Thus, the visible disconnect between writ-
ing assessment theory and on-the-ground placement practice has consequences for
the reputations of two-year colleges, their instructors’ professional status within
the discipline, and the perceived value of the education their students receive.
In sum, we know the consequences of writing placement based on decontex-
tualized algorithmic thinking and limited construct representation can be dire. It
sends inaccurate and counter-productive messages about what we value in college
writing; it appears to misplace students at unacceptable and often inequitable
rates; it fails to assess key capacities necessary for college success; and it does not
provide information about what kinds of supplementary supports might benefit
students—something that contextualized, nonbinaristic measures with broader
construct representation can offer (Hassel & Giordano, 2015). At two-year insti-
tutions, the consequences of poor placement practices are not simply a matter of
how many credit-bearing writing courses a student will need to complete. In an
220
Writing Assessment, Placement, and the Two-Year College
221
Toth, Nastal, Hassel, and Giordano
222
Writing Assessment, Placement, and the Two-Year College
assumptions about students rather than from reflective latent variable mod-
els validated under field-test conditions” was not explicitly attended to in the
Standards, and as a result, students and practitioners may encounter technically
sound assessment practices whose social consequences have been ignored.
The 2016 special issue contributors drew on foundational texts in social sci-
ence, education, and assessment to arrive at their call to use writing assessment
as a means of achieving social justice “as a principle of fairness so opportunities
do not merely exist, but rather, so each individual has a fair chance to secure
such opportunities” (Slomp, 2016b). Elliot (2016) identified “fairness” in writ-
ing assessment as “the identification of opportunity structures created through
maximum construct representation under conditions of constraint—and the
toleration of constraint only to the extent to which benefits are realized for the
least advantaged.” This rethinking of fairness in terms of opportunity structures
has powerful implications for two-year colleges, which have a mission to provide
access to educational opportunity for the “least advantaged.”
As we have discussed, the commercial exams that have long dominated two-
year college writing placement typically offer inadequate representation of local
constructs of college writing. They also reproduce language and literacy ideologies
that advantage students from White, middle-class communities. While we have
long tolerated such constraints in the name of efficiency at often under-resourced
open admissions institutions, it is now clear that those constraints have, in fact,
harmed the least advantaged. Through systematic misplacement, particularly un-
derplacement that delays enrollment in college-level courses, we have reduced those
students’ likelihood of degree completion. In the process, we have also sent them
negative, destructive messages about their capacities as writers and learners and
about the value of the rhetorical and literacy practices in their out-of-school com-
munities. These disparate, adverse impacts are neither fair nor, in many cases, legal
(Klausman et al., 2016; Poe & Cogan, 2016; Poe et al., 2014). The educational
policy shifts of the last decade have created an opportunity to rethink “business as
usual” (Klausman et al., 2016, p. 139) in two-year college writing placement. We
will need all of the critical tools emerging from the field of writing assessment to
reform these processes in ways that advance opportunity and social justice.
223
Toth, Nastal, Hassel, and Giordano
224
Writing Assessment, Placement, and the Two-Year College
225
Toth, Nastal, Hassel, and Giordano
more extensive track record for DSP in open admissions settings than the scholarly
literature has suggested. She finds that DSP offers a promising alternative to man-
datory placement at two-year colleges, but that it also presents distinctive consid-
erations for implementation that warrant deeper theorization and further research.
The special issue concludes with a collaboratively authored “Forum” that dis-
cusses how contributors see the special issue affirming, extending, and/or com-
plicating the principles articulated in the 2016 TYCA white paper. This polyvo-
cal conversation surfaces shared convictions as well as points of contention and
unresolved questions that suggest areas for future activism, policy-making, and
research. Informed by the articles in this special issue and critical questions raised
in the body of the forum, Elliot, Poe, and Nastal offer a roadmap for two-year
college placement reform that synthesizes the principles of the TYCA white pa-
per with additional theoretical insights from the writing assessment and educa-
tional measurement literature. This document is designed to help facilitate local
conversations about placement reform among faculty, administrators, and other
stakeholders at two-year colleges. With this critical and theoretically-grounded yet
practical resource for making institutional change, Elliot et al. offer a milestone
example of cross-sector alliance in writing assessment that helps equip two-year
college English faculty to assert professional authority in local policy decisions.
Taken together, the pieces in this special issue model the kind of critical reform-
er role that two-year college faculty can take on. We believe faculty at open admis-
sions institutions need to be participants in conversations about writing assessment
and social justice, and these articles demonstrate that two-year college faculty have
much to offer those discussions. In addition to contributing to disciplinary knowl-
edge, their efforts can provide colleagues at other two-year colleges with valuable
insight and precedent for pursuing reform at their own institutions. Finally, these
articles suggest that cross-sector scholarly alliances can strengthen our collective ef-
forts to pursue more equitable approaches to writing assessment: approaches that
honor open admissions students and the rhetorical resources of our communities.
In sum, we hope this special issue persuades readers at all institution types that two-
year colleges are important sites for making knowledge about writing assessment
and for putting that knowledge to work as social justice-oriented praxis.
REFERENCES
Adams, P. D. (1993). Basic writing reconsidered. Journal of Basic Writing, 12(1), 22–36.
Adams, P. D., Gearhart, S., Miller, R., & Roberts, A. (2009). The Accelerated learning
program: Throwing open the gates. Journal of Basic Writing, 28(2), 50–69.
Alexander, J. (2016). Queered writing assessment. College English, 79(2), 202–205.
American Association of Community Colleges. (2017). 2017 FactsSheet. American Association
of Community Colleges. [Link]
226
Writing Assessment, Placement, and the Two-Year College
227
Toth, Nastal, Hassel, and Giordano
Casazza, M. E., & Silverman, S. L. (2013). Meaningful access and support: The path to
college completion. Council of Learning Assistance and Developmental Education
Association. [Link]
Caswell, N. I. (2018). Queering assessment: Fairness, affect, and the impact on LGBTQ
writers. In M. Poe, A. B. Inoue, & N. Elliot (Eds.), Writing assessment, social justice,
and the advancement of opportunity (pp. 353-378). The WAC Clearinghouse; Univer-
sity Press of Colorado. [Link]
Chen, X. (2016). Remedial course taking at U.S. public 2- and 4-year institutions: Scope,
experiences, and outcomes (U.S. Department of Education No. NCES 2016-405).
National Center of Education Statistics. [Link]
Cho, S.-W., Kopko, E., Jenkins, D., & Jaggars, S. S. (2012). New evidence of success for community
college remedial English students: Tracking the outcomes of students in the Accelerated Learning Program
(CCRC Working Paper No. 53). Community College Research Center, Columbia University.
Cohen, A. M., Brawer, F. B., & Kisker, C. B. (2014). The American community college
(6th ed.). John Wiley & Sons.
Conference on College Composition and Communication Executive Committee.
(2009). Writing assessment: A position statement. [Link]
positions/writingassessment
Council of Learning Assistance and Developmental Education Associations. (n.d.).
College access (Policy Statement). Council of Learning Assistance and Developmental
Education Associations.
Council of Writing Program Administrators. (2014). WPA outcomes statement for first-
year composition. [Link]
Cushman, E. (2016). Decolonizing validity. Journal of Writing Assessment, 9(1). https://
[Link]/uc/item/0xh7v6fb
Elliot, N. (2015). Validation: The pursuit. College Composition and Communication,
66(4), 668–687.
Elliot, N. (2016). A theory of ethics for writing assessment. Journal of Writing
Assessment, 9(1). [Link]
Faigley, L., Cherry, R., Jolliffe, D., & Skinner, A. (1985). Assessing writers’ knowledge
and processes of composing. Ablex Publishing Corporation.
Gallagher, C. W. (2007). Reclaiming assessment: A better alternative to the accountability
agenda. Heinemann Educational Books.
Gomes, M. (2018). Writing assessment and responsibility for colonialism. In M.
Poe, A. B. Inoue, & N. Elliot (Eds.), Writing assessment, social justice, and the
advancement of opportunity (pp. 201-226). The WAC Clearinghouse; University
Press of Colorado. [Link]
Goudas, A. M., & Boylan, H. R. (2012). Addressing flawed research in developmental
education. Journal of Developmental Education, 36(1), 2–13.
Goudas, A. M., & Boylan, H. R. (2013). A brief response to Bailey, Jaggars, and Scott-
Clayton. Journal of Developmental Education, 36(3), 28–32.
Griffiths, B. (2017). Professional autonomy and teacher-scholar-activists in two-year
colleges: Preparing new faculty to think institutionally. Teaching English in the Two-
Year College, 45(1), 47-68.
228
Writing Assessment, Placement, and the Two-Year College
Griffiths, B., & Toth, C. (2017). Rethinking “class”: Poverty, pedagogy, and two-year
college writing programs. In W. Thelin & G. Carter (Eds.), Class in the composition
classroom: Pedagogy and the working class (pp. 231-257). Utah State University Press.
Hammond, J. W. (2018). Toward a social justice historiography for writing assessment.
In M. Poe, A. B. Inoue, & N. Elliot (Eds.), Writing assessment, social justice, and the
advancement of opportunity (pp. 41-70). The WAC Clearinghouse; University Press
of Colorado. [Link]
Harrington, S. (2005). Learning to ride the waves: Making decisions about placement
testing. WPA: Writing Program Administration, 28(3), 9–29.
Hassel, H., & Giordano, J. B. (2011). First-year composition placement at open-
admission, two-year campuses: Changing campus culture, institutional practice,
and student success. Open Words: Access and English Studies, 5(2), 29–39.
Hassel, H., & Giordano, J. B. (2013). Occupy writing studies: Rethinking college
composition for the needs of the teaching majority. College Composition and
Communication, 65(1), 117–139.
Hassel, H., & Giordano, J. B. (2015). The blurry borders of college writing:
Remediation and the assessment of student readiness. College English, 78(1), 56–80.
Hassel, H., Klausman, J., Giordano, J. B., O’Rourke, M., Roberts, L., Sullivan, P., &
Toth, C. (2015). TYCA white paper on developmental education reforms. Teaching
English in the Two-Year College, 42(3), 227–243.
Haswell, R. (2004). Post-secondary entrance writing placement: A brief synopsis of re-
search. [Link]. [Link]
Haswell, R. (2005). Post-secondary entrance writing placement. [Link]. http://
[Link]/profresources/[Link]
Henson, H. & Hern, K. (2019). Let them in: Increasing access, completion, and
equity in English placement policies at a two-year college in California. Journal of
Writing Assessment, 12(1). [Link]
Hern, K. (2012). Acceleration across California: Shorter pathways in developmental
English and math. Change: The Magazine of Higher Learning, 44(3), 60–68.
Herrington, A., & Moran, C. (2001). What happens when machines read our
students’ writing? College English, 63(4), 480–499.
Herrington, A., & Moran, C. (2012). Writing to a machine is not writing at all. In N.
Elliot & L. Perelman (Eds.), Writing assessment in the 21st century: Essays in honor of
Edward M. White (pp. 219–232). Hampton Press.
Hodara, M., Jaggars, S. S., & Karp, M. M. (2012). Improving developmental education
assessment and placement: Lessons from community colleges across the country (CCRC
Working Paper No. 51). Community College Research Center, Columbia University.
Huddleston, E. M. (1954). Measurement of writing ability at the college entrance level: Objec-
tive vs. subjective testing techniques. Journal of Experimental Education, 22(3), 165–213.
Hughes, K. L., & Scott-Clayton, J. E. (2011). Assessing developmental assessment in
community colleges. Community College Review, 39(4), 327–351.
Inoue, A. B. (2009a, Summer). Self-assessment as programmatic center: The first-
year writing program and its assessment at California State University, Fresno.
Composition Forum, 20.
229
Toth, Nastal, Hassel, and Giordano
230
Writing Assessment, Placement, and the Two-Year College
Nist, E. A., & Raines, H. H. (1995). Two-year colleges: Explaining and claiming our
majority. In J. Janangelo & K. Hansen (Eds.), Resituating writing: Constructing and
administering writing programs (pp. 59–70). Boynton/Cook.
Patthey-Chavez, G. G., Dillon, P. H., & Thomas-Spiegel, J. (2005). How far do they
get? Tracking students with different academic literacies through community college
remediation. Teaching English in the Two-Year College, 32(3), 261–277.
Perelman, L. (2012). Mass-market writing assessments as bullshit. In N. Elliot & L.
Perelman (Eds.), Writing assessment in the 21st century: Essays in honor of Edward M.
White (pp. 425–438). Hampton Press.
Poe, M., & Cogan, J. A. (2016). Civil rights and writing assessment: Using the
disparate impact approach as a fairness methodology to evaluate social impact.
Journal of Writing Assessment, 9(1). [Link]
Poe, M., Elliot, N., Cogan, J. A., & Nurudeen, T. G. (2014). The legal and the
local: Using disparate impact analysis to understand the consequences of writing
assessment. College Composition and Communication, 65(4), 588–611.
Poe, M., & Inoue, A. B. (2016). Toward writing as social justice: An idea whose time
has come. College English, 79(2), 119–126.
Scott-Clayton, J. (2012). Do high-stakes placement exams predict college success?
(Working Paper No. 41). Columbia University.
Scott-Clayton, J., & Belfield, C. (2015). Improving the accuracy of remedial placement
(Research Overview). Columbia University.
Scott-Clayton, J., Crosta, P. M., & Belfield, C. R. (2014). Improving the targeting of
treatment: Evidence from college remediation. Educational Evaluation and Policy
Analysis, 36(3), 371–393.
Shapiro, D., Dundar, A., Huie, F., Wakhungu, P., Yuan, X., Nathan, A., & Hwang,
Y. (2017). Completing college: A national view of student attainment rates by race
and ethnicity- Fall 2010 cohort (Signature Report No. 12b). National Student
Clearinghouse Research Center.
Shapiro, D., Dundar, A., Wakhungu, P., Yuan, X., Nathan, A., & Hwang, Y. (2016).
Completing college: A national view of student attainment rates- Fall 2010 cohort
(Signature Report No. 12). National Student Clearinghouse Research Center.
Slomp, D. (2016a). An integrated design and appraisal framework for ethical writing
assessment. Journal of Writing Assessment, 9(1). [Link]
item/4bg9003k
Slomp, D. (2016b). Ethical considerations and writing assessment. Journal of Writing
Assessment, 9(1). [Link]
Smith, W. L. (1993). Assessing the reliability and adequacy of using holistic scoring
of essays as a college composition placement technique. In M. M. Williamson &
B. A. Huot (Eds.), Validating holistic scoring for writing assessment: Theoretical and
empirical foundations (pp. 142–205). Hampton Press.
Stein, Z. (2016). Social justice and educational measurement: John Rawls, the history of
testing, and the future of education. Routledge.
Sullivan, P. (2008). Measuring “success” at open admissions institutions: Thinking
carefully about this complex question. College English, 70(6), 618–632.
231
Toth, Nastal, Hassel, and Giordano
232
PART 3.
IMPLICATIONS OF AUTOMATED
SCORING OF WRITING
RETROSPECTIVE.
IMPLICATIONS OF AUTOMATED
SCORING OF WRITING
Laura Aull
University of Michigan
Popular notions of automated scoring are often oversimplified, and grim. They
bring to mind product-oriented treatments of writing. They summon images
of machines replacing teachers. They conjure inhumane outsourcing—an un/
necessary evil, the equivalent of getting a robot on the phone who cannot un-
derstand you and leaves you desperately annunciating, operator!
The implications of being misunderstood are far more serious than an unsuc-
cessful phone call, of course. Scoring algorithms are unable to parse some students’
creative ideas because they are in language deemed nonstandardized by schools that
reinscribe prescriptive and oppressive histories (Hammond, 2019; Perryman-Clark,
2013). Automated scoring that is most able to focus on machine-readable text does
not focus on ‘languaging’: the mental processes of meaning-making that surround
the text produced (Ivanič, 2004). In both of these examples, automated scoring
belies the practices and principles most of us support in our writing courses.
Yet we cannot ignore automated scoring any more than we can ignore any
widespread writing assessment approach. It is part of student writing today, and
it touches on all the major themes in this collection: technical issues; evolving
ideas about writing; teachers’ and students’ lived experiences; policy; and em-
bedded concerns about reliability, validity, bias, and fairness. Automated scoring
requires our engagement, even as it can be hard to know what to think, between
media representations of automated scores, outsourced automated tools, high
stakes for students and instructors, and the varied demands on writing students,
educators, and administrators to use automated scoring tools.
This essay strives to offer an overall look at automated scoring in writing
assessment over the past two decades, particularly how it constructs student
writing and writers and how writing educators might engage with it. To do so, I
draw on three articles from the Journal of Writing Assessment that illustrate writ-
ing educators’ critical engagement with automated scoring:
1. Validity of Automated Scoring: Prologue for a Continuing Discussion of
Machine Scoring Student Writing by Michael Williamson (2003)
236
Implications of Automated Scoring of Writing
237
Aull
Constructing Writing
Computer algorithms are written by humans and carried out by computers;
they are limited to (and enhanced by, depending on your perspective) what
computers can do. Any given automated scoring algorithm implies what
matters most, constructing writing according to what it measures, such as
mechanical choices, length (Perelman), text-matching, and/ or standardized
written academic English (Canzonetta and Kannan). In turn, it implies that
certain aspects of writing don’t count—the things not measured, if the auto-
mated score is the only score used.
All the articles emphasize that like any kind of writing, academic writing is
an integration of many processes, and this is part of concerns about automated
scoring. Using the same automated scoring tool on each one—or comparing
automated scores across them—belies the situated rhetorical action entailed in
each one. The articles underscore that automated scoring does not construct
writing as situated language use: it cannot account for writing as rhetorical
action (Perelman), writing as situated literacy (Williamson), and writing and
source use as culturally-specific action (Canzonetta & Kannan). The use of dif-
ferent automated tools, in turn, emphasizes different conceptions of assessment
validity, which is also always situated (Williamson, 2003).
Constructing writing through design and use of automated scoring can also
point to the lack of a clear writing construct. This is a problem Perelman delineates
in his critique of the Shermis and Hamner study: “Without [any explicit construct
of writing], it is, of course, impossible to judge the validity of any measurement”
(para. 6). An illustration Perelman offers is the use of multiple constructed re-
sponse tasks compared in the same way, e.g., some that require understanding and
incorporation of included reading texts, and some that do not.
Constructing Assessment
Williamson traces ideas about validity with attention to the situated nature of
literacy. Williamson’s attention to the “complexity of validity inquiry” (p. 259)
illuminates differences in conceptions of assessment “between English Studies
and educational measurement, the difference between social science and hu-
manistic disciplines” (p. 260). To date in 2004, Williamson argued, many re-
searchers in English Studies subscribed to “an older notion of validity,” thereby
“unwittingly missing” more contextualized, less rigid conceptions of validity for
writing assessment (p. 262). In other words, early questions about validity fo-
cused on a given test, and whether it measured what it purported to measure.
In those cases, assessment is constructed as a process of consistent measurement,
238
Implications of Automated Scoring of Writing
regardless, for instance, of the validity of the writing construct, or the particular
abilities emphasized in an assessment task.
Assessment tools construct writing assessment through what they do not
evaluate as well. Deane et al., all ETS researchers, (2013) show that “[automated
scoring] systems do not explicitly evaluate the validity of reasoning, the strength
of evidence, or the accuracy of information” (para. 2). They illustrate how this
poses a risk because it can mean a disconnect between what assessors value and
what an automated tool can measure. In such cases, the scoring makes an inter-
pretive argument about a piece of writing based only on partial evidence about
the piece of writing.
Perelman shows that the scoring part of automated scoring is important not
only in what it interprets but also in how it is represented. His article reviews a
study that represents its findings as though they suggest that automated scorers
are as accurate as human raters. Yet as Perelman shows, the automated scores
were rounded to integer values in ways that favored the automated scores. Perel-
man writers, “Essays scores, be they holistic, trait, or analytical, always are con-
tinuous variables, not discrete variables (integers), even though graders almost
always have to give integer values as scores” (para. 14). Interrogating automated
scoring and representations thereof have high stakes for how assessment is con-
structed vis-à-vis policy decisions. The paper Perelman critiques, for instance,
was sent to the Partnership for Assessment of Readiness of College and Careers
and the Smarter Balanced Assessment Consortium. Thus Perelman argues that
automated scoring demands rigorous statistical analysis and offers important
information in light of potential policy decisions.
239
Aull
public partnerships (para. 11). These strategies construct writing as rule mastery,
writers as needing to be regulated, and instructor-assessors as regulators.
In sum, the articles remind us that vis-à-vis the question what does automated
scoring do?, we can answer: it constructs writing, writers, and assessors in par-
ticular ways. As a writing assessment tool, any automated scoring tool is consti-
tutive; all assessment activities are complex and value-laden activities, whether
baked into an algorithm or required of a human (Aull, 2017). Thus the authors
remind us that, as is any writing assessment approach, automated scoring is an
opportunity to see what we prioritize in writers and writing—a chance to under-
stand and question what that is and what’s left out.
240
Implications of Automated Scoring of Writing
241
Aull
writing, like all human communication, is not that it is true or false, correct or
incorrect, but that it is an action, that it does something in the world” (para. 4).
Humans more easily read and write in this sociocultural way than machines. Yet
Perelman’s point reminds us that human and automated scorers alike can belie
this sociocultural conceptualization when they focus on error-hunting rather
than situated meaning-making.
It follows that like writing and literacy, writing assessment and its validity are
situated. We can see in Williamson’s 2003 article that ideas about validity need-
ed expanding at the time: he cautions against seeing validity only as reliability.
In other words, he decries the notion that validity means “a test has to measure
what it purports to measure,” because validity is situated “in a particular use of
a test, in a particular context, at a particular time” (p. 267). This emphasis on
use anticipates Kane’s work on interpretation and use arguments that are part of
any test: scores represent inferences drawn from assessments. Those scores (and
use of those scores, such as admissions or course placement) make interpretive
arguments (Kane, 2013). Likewise, recent notions of reliability have expanded,
calling into question prevailing standards built on narrow testing constructs,
moving instead toward reliability based on the measurement, conditions, and
objectives of complex writing performances (Ross & LeGrand, 2017).
242
Implications of Automated Scoring of Writing
243
Aull
244
Implications of Automated Scoring of Writing
All three articles illuminate the risks of keeping writing and measurement special-
ists apart. First, Williamson notes a lack of cross-talk between “writing teachers”
and the “assessment community,” the two stakeholder groups most implicated in
the scoring of writing. He explores “beliefs and assumptions held by each side” (p.
254), including disciplinary differences entailing different goals: more humanistic
approaches emphasizing social context in college writing studies; more scientific
approaches emphasizing aggregated patterns for assessment professionals.
Separation between the stakeholder groups, Williamson notes, means a
lack of “productively learn[ing] to talk to each other about automated scor-
ing.” It means insufficient questioning of key assessment concepts such as va-
lidity, which becomes a “taken-for-granted ubiquity” in the lived experiences
of students and teachers by becoming a “totalizing” and “naturalized concept
and invisible instrument of rigor.” Describing the early 21st century, Williamson
writes, “English teacher response to automated scoring has been limited and . . .
does not refer to any of the evidence presented by the developers of automated
scoring programs” (p. 254).
Alternatively, Williamson describes, debate and understanding across the
groups can challenge those in English Studies to address basic procedural is-
sues of social science with questions of validity. It can challenge those in social
science to consider validity as a situated construct, one that must observe the
same situatedness that literacy theorists have been articulating for some time.
Ultimately, the groups can come together around “the shared goal of moving
toward more reliable and efficient ways to measure educational achievement and
writing ability” (p. 254). At minimum, Williamson writes, “automated scoring
is an incredible research opportunity through which we can explore the many
different ways student writing can be read, valued, and sanctioned” (p. 254).
All three articles are examples of how productive examination of automated
scoring is facilitated by engaging educational measurement and writing edu-
cation. Perelman’s article exemplifies the possibility of multiple methods and
perspectives coming together. He uses statistical tests to expose unsupported
claims in Shermis and Hamner’s paper in ways valued by social scientists, and he
attends to constructed response task as a situated written action in ways valued
by those in English studies. Perelman’s critique also shows the importance of
assessment and writing researchers’ engagement with publicly-rendered claims
about writing and scoring. Shermis and Hamner’s study findings, critiqued by
Perelman, circulated widely: following Shermis’s presentation of their findings at
the National Council on Measurement in Education’s annual meeting, the study
245
Aull
was cited in Inside Higher Ed and The New Scientist, and by a press release from
the University of Akron.
Canzonetta and Kannan add additional emphasis, calling for cross-talk about
the role of large corporations in global automated assessment services. In so doing,
they take up Grabill’s call for more attention to automated writing technologies by
composition and rhetoric scholars, stressing that “[g]lobally, millions of students
are subjected to writing technologies that writing experts did not design” (p. 296).
Canzonetta and Kannan specifically outline plagiarism detection tools (PDSs),
using [Link]’s success as an example of corporate influence in U.S. univer-
sities. They analyze how PDSs constitute instructors (presumed to be members
of the “[Link] educational community”) as preservers of ethical and moral
standards, positioned antagonistically against students, and assumed to be consis-
tent across institutions and geographic locations. They call for direct engagement,
so that there is greater understanding of the global cultural work of automated
plagiarism and assessment tools.
In sum, the articles underscore that writing specialists have a responsibility to
engage critically with automated scoring. The alternative, they imply, constitutes
risks and missed opportunities. They bring to mind White’s urging: “Assessment
of writing can be a blessing or a curse, a friend or a foe, an important support
for our work as teachers or a major impediment to what we need to do for our
students. Like nuclear power, say, or capitalism, it offers enormous possibilities
for good or ill, and, furthermore, it often shows both its benign and destructive
faces at the same time” (White, 1994, p. 137). In none of these metaphors is
there an option for writing educators to leave alone automated scoring as part
of writing assessment.
246
Implications of Automated Scoring of Writing
close, I consider related possibilities for automated scoring, for making it a site
of investigating rather than damning language difference, a site of collective ex-
ploration rather than a site of top-down design and individual mastery.
Decolonizing Validity
The articles by Williamson, Perelman, and Canzonetta and Kannan concep-
tualize automated scoring as a site for ongoing inquiry into available tools and
decisions about their use. They call for understanding writing, writing assess-
ment, and validity as situated in rhetorical situations. They show how the use of
automated scoring tools can entail a cyclical approach to validity: if we define
validity narrowly and measure it in narrow tests, we learn only about those nar-
row conceptualizations. If validity is a matter of narrowly-defined, consistent
scoring, then an automated scorer can measure length, mechanics, and lexical
cohesion in a timed writing task, for instance, and be valid. Alternatively, if
validity is instead a matter of fairness defined as equitable distribution of scores
across different student groups, then a valid test must do very different things.
Cushman argues that to date, the “concept of validity creates the colonial
difference as it maintains social, epistemic, and linguistic hierarchies.” It does so
by “identifying what is objective and what is evidence,” and by hiding its social
construction: validity “is a naturalized concept and invisible instrument of rigor
that totalizes the realities of students and researchers” (para. 7). Drawing on
Tiostanova and Mignolo’s (2012) phrasing (“dwelling in the borders” in order
to ‘”change the construct” itself ), Cushman offers an alternative conception of
validity, one in which “[d]welling in the borders begins with the knowledge,
languages, histories, and practices understood and valued by the people who
live these realities” (para. 25). In this conceptualization, validity is collectively
constructed and navigated.
In turn, validity evidence tools work “not as a way to maintain, protect, con-
form to, confirm, and authorize the current systems of assessment and knowledge
making, but rather as a way to better understand difference in and on its own
terms.” In other words, validity could be seen not as a way to hold individuals to
one set of metrics determined by an external group—not as “a way to maintain,
protect, conform to, confirm, and authorize the current systems of assessment and
knowledge making.” Rather, validity could be seen in terms of descriptive power:
what it helps us learn about difference. In this approach, validity measures do
not “mak[e] [one] experience into a universal one, the baseline against which all
Others are tested and their knowledges and languages are deemed deficit to” (para.
26). A valid assessment approach would thus be one that “seek[s] to identify un-
derstandings in and on the terms of the peoples who experience them” (para. 26).
247
Aull
Exploring Difference
Williamson underscores that “fluent adult reading” expects different views from
different readers (Williamson p. 265). Generally, automated scoring and in-
ter-rater agreement expect the opposite: they expect “convergent reading.” Like-
wise, Perelman demonstrates the dangers of disparate approaches to resolving
difference. And like other studies focused on agreement between human scorers,
and/or between human and machine scoring, Deane et al. prioritize what Wil-
liamson calls convergence: the smaller the difference, the better (and “ideal for
operational use” are very small differences in scores inferred from reading).
Here we see a good example of Cushman’s point: in this case, disparate reading
is, in a sense, invalid; validity rests on agreement in reading and inferences. What
if we could construct writing, and reading one’s own and other’s writing, not as
a site of deciding whether it was right or wrong, but of exploring the inferences
drawn from machines and from humans? What would it take to create that world?
Canzonetta and Kannan underscore that this is not easy, because rhetorics of stan-
dardization and consistency are beneficial for Turnitin’s business model.
Formative Assessment
Canzonetta and Kannan describe that “Formative assessment necessitates that
teachers respond to students’ needs, personalities, struggles, and strengths; and
get to know them apart from their writing.” Ultimately, Canzonetta and Kannan
caution against automated plagiarism tools’ role in formative assessment, but
they do point to that role as a site for critical investigation: “it is important to
critically interrogate Turnitin’s rhetorics of formative assessment, which obscure
the company’s cooptation of student data and potential to undermine writing
program goals” (para. 27). This is all the more important because the message
from plagiarism tools can promote formative uses that ultimately “aim to quell
critique and breed a compliant, submissive population of students” instead of
a more student-centered invitation of the students’ active questioning of ideas
about plagiarism and writing.
One way that automated scoring could be used would be in formative reports
for students’ use. Following Cushman, these reports could be a site of exploring
248
Implications of Automated Scoring of Writing
difference. What range of responses were there? What did these differences achieve?
How did they respond to or change the rhetorical situation of the task?
CONCLUSION
In the sections above, we’ve seen automated scoring framed as the use of algo-
rithms to evaluate aspects of writing and as a site for exploring ideas about writ-
ing and writers embedded in any given approach to writing assessment. We’ve
seen that automated scoring constructs writing, writers, and the practices of
249
Aull
REFERENCES
Allen, L., Crossley, S., Kyle, K., & McNamara, D. S. (2014). The importance of
grammar and mechanics in writing assessment and instruction: Evidence from
data mining. Grantee Submission. [Paper Presentation] International Conference on
Educational Data Mining.
Aull, L. L. (2015). First-year university writing: A corpus-based study with implications
for pedagogy: Springer.
Aull, L. L. (2017). Tools and tech: A new forum. Assessing Writing. 33, A2-A7.
Aull, L. L. (2020). How students write: A linguistic analysis. MLA.
Aull, L. L. (2021). What is “Good Writing?”: Metadiscourse as civil discourse. Journal
of Teaching Writing 36(1), 37-60.
Aull, L. L. (2024). You can’t write that . . . 8 myths about correct writing. Cambridge
University Press.
Canzonetta, J., & Kannan, V. (2016). Globalizing plagiarism & writing assessment: a
case study of Turnitin. The Journal of Writing Assessment, 9(2). [Link]
org/uc/item/5vq519dr
Crossley, S. A., Bradfield, F., & Bustamante, A. (2019). Using human judgments
to examine the validity of automated grammar, syntax, and mechanical errors in
writing. Journal of Writing Research, 11(2), 251-270.
Crossley, S. A., Kyle, K., & McNamara, D. S. (2015). To Aggregate or not? Linguistic
features in automatic essay scoring and feedback systems. Journal of Writing
Assessment, 8(1). [Link]
Cushman, E. (2016). Decolonizing validity. Journal of Writing Assessment, 9(1). https://
[Link]/uc/item/0xh7v6fb
Deane, P., Williams, F., Weng, V., & Trapani, C. S. (2013). Automated essay scoring in
innovative assessments of writing from sources. Journal of Writing Assessment, 6(1),
40-56. [Link]
Elliot, N. (2005). On a scale: A social history of writing assessment in America. Peter Lang.
250
Implications of Automated Scoring of Writing
251
Aull
252
CHAPTER 8.
VALIDITY OF AUTOMATED
SCORING: PROLOGUE FOR A
CONTINUING DISCUSSION
OF MACHINE SCORING
STUDENT WRITING
Michael Williamson
Indiana University of Pennsylvania
Writing assessment has developed along two separate lines, one centered
in professional organizations for writing teachers and the other centered
in professional organizations for the broader assessment community. As
the controversy about automated scoring continues to develop, it is im-
portant for writing teachers and researchers to become fluent in the
discourse of the broader assessment community. Continuing to label the
work of the broader assessment community as positivist and continuing
to ignore it will only result in a continuing sense of defeat as automated
assessment is adopted more widely. On the other hand, an examination
of the literature on educational assessment will reveal that the theoreti-
cal base for assessment is quite consistent with the principles adopted by
the writing assessment community.
254
Validity of Automated Scoring
English, Page’s original work drew a response similar to Anson and Herrington
and Moran from Macrorie (1969). On the other hand, Coombs (1969) was
skeptical, but not entirely dismissive of the potential demonstrated in Project
Essay Grade. However, automated scoring does not seem to have been whole-
heartedly embraced by anyone in English Studies publishing in typical outlets,
such as College English or Research in the Teaching of English.
On the other hand, a recently burgeoning literature on automated scoring
has appeared in the literature typically examined by the broader assessment com-
munity, much of it suggesting that automated scoring does have valid applica-
tions for the assessing of writing.
255
Williamson
VALIDITY
As early as 1951, validity was defined by Edward Cureton in the first edition of
what would become a periodic definition of the state of the art in educational
measurement, Educational Measurement.
The essential question of test validity is how well a test does the job it is
employed to do. The same test may be used for several different purposes, and
its validity may be high for one, moderate for another, and low for a third
(p. 621).
256
Validity of Automated Scoring
257
Williamson
Stalnaker (1951) labels achievement testing. Tests like statistical operations are
conducted to make informed educational judgments. The simplicity of distinc-
tion between a procedural and conceptual understanding of validity is not al-
ways as clear and separate as it might seem. The fundamental nature of validity
can be rendered confusing by educational researchers themselves.
While the definition of validity seems simple and straightforward, there are
several different types of validity that are relevant in the social sciences. Each
of these types of validity takes a somewhat different approach in assessing the
extent to which a measure measures what it purports to (Carmines & Zeller,
1980, p. 17).
Broad (2003) labels one stance in writing assessment “positivist,” a stance
that can be traced to Berlin’s (1984) history of writing instruction. Positivism
as a theoretical approach to the philosophy of science certainly characterizes
early psychometric theory and its attempt to define psychology and educational
and psychological assessment as a science. Guilford (1954) traces the emergence
of statistical investigation in psychology and grounds his approach to the field
in mathematics, as well as statistical inquiry, “The progress and maturity of a
science are often judged by the extent to which it has succeeded in the use of
mathematics” (p. 1). Gulliksen (1950) specifically limits his description of men-
tal testing to those defined by quantitative methods, while specifically noting the
difference between statistics and mathematics. In Guilford’s terms, mathematics
is a “universal language that any discipline may use with power and convenience”
(p. 1). That this movement toward the use of mathematics and quantification
may be positivist is one that deserves larger exploration in the literature of the
field. However, there is an interesting contrast to what may be perceived as the
problem of quantification in writing assessment.
As early as the 1950s, at least, such issues as validity were seen less as defined
by the results of a statistical test than as a matter of disciplinary disputation,
the assembling of evidence, not the simple results of a statistical test (Cureton,
1951). In a related example, in discussing educational evaluation, one of the
primary applications of educational measurement, Cooley and Lohnes (1976),
both eventually to become president of the American Educational Research As-
sociation, suggest that the scrutiny of the field and not objectivity is the issue.
Moss (1998) calls her response essay to a study of writing assessment validation,
“The Test of the Test.” For Moss, validation is a practice in turning the gaze to-
ward the construct of the assessment itself. It is a form of reflective practice, or
as Ellen Schendel (1999) claims, “social action.”
What tends to keep researchers honest is the publicly available record of
what they did and what they found, and not a godlike objectivity which some
people seem to feel those doing evaluations should exhibit. Scientists doing basic
258
Validity of Automated Scoring
research know that if their work is to have any value whatsoever, it will be closely
read and critically examined by their colleagues in the field (Cooley & Lohnes,
1976, p. 2).
These and other perspectives of validity are rooted in the ideas of Cronbach
(1988, 1989) and Messick (1989). Cronbach (1989) characterizes validity as a
form of disciplinary argumentation, one that is never finished and that evolves
with each new use of an assessment in a new locale: “Validation is a lengthy, even
endless process” (p. 151). Such a definition is supported by Cureton (1951) and
Anastasi (1976) as well. It is this definition that leads Huot (2002) to character-
ize assessment as a continuing form of research. Thus, writing assessment should
be viewed as a continuing examination of the available tools for assessment, as
they are used for making new decisions. New developments will inevitably bring
new tools, all of them requiring validity inquiry of their own.
Smith (1993) is probably the first researcher in writing assessment who fully
reflects the complexity of validity inquiry. Although his work is some of the first
substantive research that looks at the validity and not reliability of a writing
assessment (Huot, 1996), ironically, he eschewed the word validity because he
wanted to avoid any baggage associated with such a term. He used accuracy of
placement as the goal for his placement testing program at the University of
Pittsburgh. With collaborators, he designed a series of studies on the proce-
dures that structure the way teachers make decisions based upon their reading
of placement essays. Each of the studies led to a modification of the procedures
that allowed a stronger claim to the validity of the assessment, the accuracy of
placement of students in the writing program. This not only demonstrates more
accurate placement of students over time, but it also led to a modification of the
scoring procedures themselves. The end result was a less costly system because
the reading and decision making were rooted in the context about which the
teachers were expert.
The notion of validity as argument and the nature of professional judgment
is related to Bleich’s (1975) view of interpretive communities and Kuhn’s (1996)
view of the way that science changes through changes in the worldview of the
members of the discipline. The meaning of a text, be it a poem or a validity
inquiry, lies with the community of readers in the field and their intertextual
experiences with the field. Such a position reflects a more postmodern view than
the positivism cited as the basis for psychometric theory.
An additional consideration for validity is the impact of the assessment
(Messick, 1989). The consequences of decisions made on behalf of a test is a core
concern for validity inquiry because the use of a test may impact what is learned
and how that learning takes place. This concern for the impact of a test is one
of the ethical bases for validity theory. Thus, validity inquiry must examine how
259
Williamson
260
Validity of Automated Scoring
261
Williamson
ARTIFICIAL INTELLIGENCE
Automated scoring is based in the technology of AI, and claims to bring the rel-
ative efficiency of automation to scoring essays. These two concepts need to be
defined as part of the process of validity inquiry. AI is a research paradigm built
around several sciences. The primary goal of the emergent paradigm has been
the simulation of human intelligence and behavior in the electronic system of a
computer. Developments in each of these sciences, from linguistics to psychology
and mathematics to computer science, have allowed a nearly continuous devel-
opment of demonstrations of intelligent machines. The emergent technologies
have resulted in a variety of applications that both enhance and simulate human
performance in a variety of fields. Thus, it seems that the use of such technolo-
gies would inevitably lead to their application in English studies. The first such
application—Project Essay Grade—was seen by its developers as a method of
relieving writing teachers of the burden of grading, leading also to more objec-
tive grades (Ajay, Tillett, & Page, 1973; Page, 1966, 1967a, 1967b, 1995; Page
& Fisher, 1968). After an initial ambiguous response (Coombs, 1969; Macrorie;
1969), the concept of computer grading seems to have had little attention from
researchers in composition and rhetoric for some time (Huot, 1996).
The development of the personal computer in the 1980s led to an outburst
of enthusiasm for the use of computers in the writing classroom. The cutting
edge of the field of computers and composition was initially defined by the
seminal work of Hugh Burns (1979) with rhetorical invention and the rapid
growth of word processing, among other business and personal applications.
Burns’ work reflected the early applications of artificial intelligence to En-
glish Studies. His work demonstrated the programming theories of artificial
intelligence pioneered by Joseph Weizenbaum in the development of Eliza, a
computer program designed to simulate the psychotherapeutic interviews of
Carl Rogers. Eliza was considered to be a failure because the program did not
meet Turing’s criterion for a computing machine simulating human behavior,
262
Validity of Automated Scoring
AUTOMATION
Automated scoring—the use of computers to simulate holistic ratings of En-
glish essays—is quite accurately described as automation in the original sense
of the word—the use of technology to relieve humans of repetitive work, work
that taxes the limits of our abilities. It is, simply, the performance of tasks by
machines, tasks that were originally performed by skilled humans, made skilled
humans more productive, or created less skilled work from more complex work.
Early automation is represented by the agricultural machines that first improved
tilling the soil and subsequently harvesting. The original Luddites of 1811-1812
were weavers in England, members of a craft guild who attempted to destroy the
newly invented machinery that left fewer jobs for unskilled workers. Mechanical
263
Williamson
264
Validity of Automated Scoring
265
Williamson
essays from the same writers, intended as a form of portfolio assessment (Breland
et al., 1987). The earlier study suggested that multiple-choice tests of grammar
predict a student writer’s performance more accurately than an essay when it
is scored using a holistic procedure by two or three raters. Consequently, the
relatively cheaper and more efficient indirect approach was justified because it
could predict an individual’s score on a criterion with the greater precision and
accuracy than a writing sample. The claim for the validity of indirect assessment
is based on a form of criterion validity known as concurrent validity that com-
pares an examinee’s performance on two different valued measures. Veal and
Hudson (1983) dispute that result in another study of the use of holistic scoring
using state assessment data from Georgia in which students’ performances on
multiple-choice tests of usage and grammar do match well with a holistic score.
Breland et al. (1987), stipulating that direct writing assessment is more effi-
cient and less costly, demonstrated that one essay read two or three times could
attain the reliability of indirect measures. Their criterion definition of writing
was six essays from each writer. They conclude that the best approach to writing
assessment is a combination of both direct and indirect assessment because the
two work together to provide both a broader and more reliable picture.
Although psychometric theory clearly supports the need for studying va-
lidity in particular applications of a test, in practice multiple-choice tests were
considered adequate when used “off the shelf ” by educational institutions. Thus,
although the theory was suggesting the need for more study of assessment pro-
cedures in particular applications, conventional wisdom allowed for their use as
ready made instruments for student, teacher, and program evaluation. This was
equally true in the use of holistic scoring. For the most part, writing assessments
used holistic scoring without much examination of the validity of its actual use,
because the understanding of assessment theory prevalent in the field was that
a test using writing is more valid on its face and in its content than any form of
indirect test (Yancey, 1999).
White (1994) recounted the political struggles involved with the adoption of
direct assessment. However, the extant theory in measurement could have been
used to support the argument against indirect assessments had more writing
assessment developers, like Veal and Hudson (1983), used the theory to argue
their position. Hence, with greater fluency in the theory that was used, writing
teachers and researchers would likely have been able to develop assessments that
could be demonstrated to have the same kinds of properties that were valued in
the validation of indirect assessment. One good example is the study by Breland
et al. (1987), which suggests that holistic scoring of a writing portfolio leads to
more accurate predictions than the score of any single essay in the portfolio.
As early as Terman’s 1916 book on the measurement of intelligence, statistical
266
Validity of Automated Scoring
267
Williamson
validation data for any of those examinations because they do not have relevant
local data to determine the suitability of each examination for the decision to
admit or deny admission to an applicant to a particular program. Validation
data, such as national norms and performance of students with self-reported
characteristics such as GPA are frequently part of these examinations. But, the
only place to determine the validity of admissions decisions is within the insti-
tution using the scores. In the case of the SAT, most admissions departments use
the scores in formulas to predict such things as first-year GPA. Similarly, ETS
reports the success of similar predictions for a number of schools as part of their
validation research.
The GRE is now scored by one human rater and eRater. For the most part,
ETS has been examining the accuracy of eRater in predicting holistic scores
from human raters. Their research suggests that eRater is able to predict the
scores of six raters with greater accuracy than two human raters. The question,
then, is, are the eRater scores any more or less accurate than the scores provided
by the two human raters typically used? If the criterion is the more raters the
better, then the answer is obviously, yes. The science of psychometrics depends
on the sheer magnitude of neumbers in order to statistically prove anything. A
traditional direct writing assessment like holistic scoring generates a single score,
technically a one item test. Because reliability is greatly improved by the number
of scores, it is easy to see how subtly and quickly the question can turn to reli-
ability. In the case of Smith’s (1993) accuracy of placement, accuracy focuses on
the decision and the underlying principle that all decisions are not equal. eRater,
however, focuses on the predictive power of one set of procedures compared to
another. For validity, the real question for eRater is whether the scores help make
better decisions about students than the current procedures used by a particular
college or university.
For those of us who use traditional holistic scoring procedures, the answer
is likely to be that it does, because eRater is going to provide more stable scores
than two holistic raters. However, the real test of the validity of eRater may lie in
a comparison with procedures like Smith’s that focus on the expert knowledge
of teachers who determine whether the student who wrote the essay belongs in
their course or the one above or below it. In this case, it is not clear that one pro-
cedure has an advantage over the other because there has never been an attempt
to examine the relative value of eRater compared to the expert placement model
defined by Smith.
Because the immediate question of the validity of automated scoring turns
on reliability, as Huot (2002) asserts, reliability has always been the focus of the
debate about writing assessment. Thus, the question of which assessment pro-
vides the best judgment of a student’s placement into a writing program has still
268
Validity of Automated Scoring
not been answered. As various new assessments have been created (e.g., Broad,
2003; Haswell, 2001; Murphy & Underwood, 1998; Royer & Gilles, 2003),
there has been a pressing need to document that these assessments promote
valid and reliable educational decisions about students, teachers, and programs.
Unfortunately, systematic and rigorous attention is not always given to things
like consequences for various participants in the assessment.
For placement, the study of the validity of writing assessment should be
focused, like Smith’s, on the decision about the best course for a student to
enter the writing program at a particular college. Writing exit examination val-
idation research should be focused on a decision about a student’s mastery of
the curriculum, for both college and school students. Furthermore, there is little
reporting of validation research in the assessment literature, in part, I suspect,
because writing assessment is a field marginalized by most writing teachers and
researchers. Most teachers, with good reason, fear any use of assessment, because
assessment has become highly politicized by federal and state government, as
well as by local school boards and administrators.
269
Williamson
should be how AI might augment the teaching of writing in the future. Explicit
views of the future are not of much value, particularly because the likelihood of
automation replacing some aspects of teaching writing is already evident, as we
have been seeing, the continuing use of electronic technology to compliment or
replace some of the work of teachers. As we have also experienced, there will be
those who claim that computers allow for greater efficiency, justifying increasing
the numbers of students working with individual teachers. It seems clear that
computers are here to stay in English Studies, even if only as word processors
to make the production of paper text easier and as communication devices to
connect writers to one another for responding. We have to expect that the future
will also hold some developments that can help us and some that can be hurtful.
Some developments will be faddish, oversold by developers and producers of the
technology, whereas others will enter our toolbox with the potential to help stu-
dents learn if used properly. My answer to the problem of automated assessment
is precisely the last point. Its potential suggests that it might have some value in
writing classrooms, but it is not clear what that may be. Second, if it does have
value, it will take continuing study understand the consequences and to estab-
lish the value through validity inquiry.
I am suggesting a stance on automated assessment that can best be charac-
terized as carefully directed critique toward the developers of automated assess-
ment. Because Pennsylvania has adopted automated assessment and the results
of that automation will be used to determine funding for school districts, there is
no question it is being used in regulatory ways. Why should we expect anything
different? Assessment has been used as a gate keeper for as long as assessment has
resulted in excluding some and including others in schooling.
Out-of-hand or outright rejection of automated assessment, a blanket con-
demnation, can only be self-serving. More importantly, we need to examine the
use of automated scoring as we would any other assessment, according to the cri-
teria of the most current theories on validating educational assessment. Arguing
that theories of literacy do not justify the use of automated assessment, is similar
to earlier arguments that indirect assessment does not have content validity. This
argument is not going to be compelling with an educational measurement audi-
ence, not to mention policymakers and regular citizens. Furthermore, without
an understanding of the common language of assessment as it is grounded in the
social sciences research methodology, we will find that our righteous indigna-
tion, our hermeneutic arguments about the meaning of new types of assessment,
are met by a wondering stare, at best, and a dismissive glare, at worst.
What I am arguing we do is to study automated assessment in order to ex-
plicate the potential value for teaching and learning, as well as the potential
harm. The theory of the developers can itself be used as a ground for validity
270
Validity of Automated Scoring
CONCLUSION
Writing assessment in American education has two professional groups with de-
veloped bodies of theory and practice. The first group, whose primary interest is
assessment, is the membership of the two professional organizations, the Amer-
ican Education Research Association (AERA) and the American Psychological
Association (APA). They far outnumber the members of the second group, the
membership of the National Council of Teachers of English and College Com-
position and Communication. For a number of years, APA and AERA were
loosely allied through members with dual memberships. More recently, recog-
nizing their common concerns and shared field, they began to work together.
The result is a clearly defined statement of definitions and standards for test
development and validation (AERA, 1999). Although the measurement com-
munity is not inherently hostile to the concerns of writing teachers, its members
will be looking for the kinds of evidence articulated in the standards, applying
the technology of validation research to the discussion of implementing auto-
mated scoring. Furthermore, their direct involvement with public education, as
the primary source for assessment tools, lends them a strong voice in the federal,
state, and local politics of assessment.
The contrasts between English Studies and educational assessment are many,
running beyond concepts or methodology. The common ground is also quite
large. One important point of comparison lies in the question of what consti-
tutes important research in the two fields. In English, researchers are typically
expected to demonstrate their mastery of the field in publications that are au-
thored by a single individual. In assessment, as in most scientific fields, import-
ant research can only be conducted by a team of people, each contributing to
the conceptualization and execution of the study. If it is time to examine the
research methodology or social sciences as it impinges on assessment, it may
also be time to explore the potential for collaborative research, not just within
either a social science or humanistic tradition (see Huot, 2002, for a discussion
of a unified field of writing assessment). If we continue to espouse outmoded
271
Williamson
REFERENCES
American Educational Research Association, American Psychological Association, &
National Council of Measurement in Education. (1999). Standards for educational
and psychological testing. American Educational Research Association.
272
Validity of Automated Scoring
Ajay, H. B., Tillett, P. I., & Page, E. B. (1973). Analysis of essays by computer (AEC-
II). Final report to the National Center for Educational Research and Development
(Project No. 80101), p. 231.
Anastasi, A. (1976). Psychological testing (4th ed.). Macmillan.
Anson, C. R. (2003). Responding to and assessing student writing: The uses and limits
of technology. In P. Takayoshi & B. Huot (Eds.), Teaching writing with computers
(pp. 234- 246). Houghton Mifflin.
Berlin, J. A. (1984). Writing instruction in nineteenth-century American colleges.
Southern Illinois University.
Blakesley, D. (2003). Directed self-placement in the university. In D. Royer & R.
Gilles (Eds.), Directed self-placement: Principles and practices (pp. 31-48). Hampton
Press.
Bleich, D. (1975). Readings and feelings: An introduction to subjective criticism. National
Council of Teachers of English.
Breland, H. M., Camp, R, Jones, R. J., Morris, M. M., & Rock, D. A. (1987).
Assessing writing skill. College Entrance Examination Board Research Report No.
11. Educational Testing Service.
Broad, B. (2003). What we really value: Beyond rubrics in teaching and assessing writing.
Utah State University Press.
Burns, H. L. (1979). Stimulating rhetorical invention in English composition through
computer-assisted instruction. Dissertation Abstracts International, DAI-A 40/70, p.
3734, January 1980, DAI Order number AAT 7928268.
Burstein, J. C. (2003). The E-rater® Scoring Engine: Automated essay scoring with
natural language processing. In M. D. Shermis & J. C. Burstein (Eds.), Automated
essay scoring: A cross-disciplinary perspective (pp. 113-121). Erlbaum.
Carmines, E. G., & Zeller, R. A. (1980). Reliability and validity assessment. Sage
Publications.
Cooley, W. W., & Lohnes, P.R. (1976). Evaluation research in education: Theory,
principles, and practice. Irvington.
Coombs, D. H. (1969). Review of The Analysis of Essays by Computer, by Ellis B. Page
and Dieter H. Paulus. Research in the Teaching of English, 3(2), 222-228.
Cooper, C. R., & Odell, L. (1977). Evaluating writing: Describing, measuring, judging.
National Council of Teachers of English.
Cronbach, L. J. (1971). Test validation. In R. L. Thorndike (Ed.), Educational
measurement (2nd ed.) (pp. 443-507). American Council on Education.
Cronbach, L. J. (1988). Five perspectives on test validity argument. In H. Wainer
(Ed.), Test validity (pp. 3-17). Erlbaum.
Cronbach, L. J. (1989). Validity after thirty years. In R. Linn (Ed.), Intelligence:
Measurement theory and public policy (pp. 147-171). University of Illinois Press.
Cureton, E. (1951). Validity. In E. F. Lindquist (Ed.), Educational measurement (pp.
621-694). American Council on Education.
Darling-Hammond, L., & Youngs, P. (2002). Defining “highly qualified teachers”:
What does “scientifically-based research” actually tell us? Educational Researcher,
31(9), 13-25.
273
Williamson
Elbow, P., & Yancey, K. B. (1994). On the nature of holistic scoring: An inquiry
composed on email. Assessing Writing, 1(1), 91-108.
Gottshalk, F. I., Swineford, F., & Coffman, W. (1966). The measurement of writing
ability. College Entrance Examination Board Research Monograph N. 6.
Educational Testing Service.
Guilford, J. P. (1954). Psychometric methods. New York: McGraw-Hill.
Gulliksen, H. (1950). Theory of mental tests. Wiley.
Haswell, R. (Ed.). (2001). Beyond outcomes: Assessment and instruction in a university
writing program. Ablex.
Herrington, A., & Moran, C. (2001). What happens when machines read our
students’ writing? College English, 63(4), 480-499.
Huot, B. (1993). The influence of holistic scoring procedures on reading and rating
student essays. In M. M. Williamson & B. A. Huot (Eds.), Validating holistic scoring
for writing assessment (pp. 206-236). Hampton Press.
Huot, B. (1996). Computers and assessment: Understanding two technologies.
Computers and Composition, 13(2), 231-244.
Huot, B. (2002). (Re)Articulating writing assessment for teaching and learning. Utah
State University Press.
Johnson-Laird, P. N. (1977). Mental models: Towards a cognitive science of language,
inference, and consciousness. Harvard University Press.
Kuhn, T. S. (1996). Structure of scientific revolutions (3rd ed.). University of Chicago Press.
Macrorie, K. (1969). Review of The Analysis of Essays by Computer by Ellis B. Page
and Dieter H. Paulus. Research in the Teaching of English, 3(2), 228-236.
Messick, S. (1989). Test validity. In R. Linn (Ed.), Educational measurement (3rd ed.)
(pp. 13- 103). American Educational Research Association; National Council of
Measurement in Education.
Moss, P. (1998). The role of consequences in validity theory. Educational measurement:
Issues and practices, 17(2), 6-12.
Murphy, S., & Underwood, T. (1998). Interrater reliability in a California middle school
English/language arts portfolio assessment program. Assessing writing, 5(2), 201-230.
Page, E. B. (1966. The imminence of grading essays by computer. Phi Delta Kappan,
238-243.
Page, E. B. (1967a). Grading essays by computer: Progress report. Proceedings of the
1966 Invitational Conference on Testing (pp. 87-100). Educational Testing Service.
Page, E. B. (1967b). Statistical and linguistic strategies in the computer grading of
essays. Proceedings of the Second International Conference on Computational
Linguistics. Grenoble, France.
Page, E. B. (1985). Computer grading of student essays. In T. Husén & Postlethwaite
(Eds.), International Encyclopedia of Educational Research (pp. 944-946). Pergamon.
Page, E. B. (1993. New computer grading of student prose, using a powerful grammar
checker. [Paper presentation]. Annual meeting of the North Carolina Association
for Research in Education. Greensboro, NC.
Page, E. B. (1995). Computer grading of essays: A different kind of testing? [Invited
address] American Psychological Association, Divisions 5, 7, 15, 16.
274
Validity of Automated Scoring
Page, E. B., Fisher, G. A., & Fisher, M. A. (1968). Project Essay Grade: A FORTRAN
program for statistical analysis of prose. British journal of mathematical and statistical
psychology, 21(1), 139.
Page, E. B., Tillett, P. I., & Ajay, H. B. (1989). Computer measurement of subject-
matter essay tests: Past research and future promise. Proceedings of the First Annual
Meeting of the American Psychological Society, Alexandria, VA.
Pula, J., & Huot, B. (1993). A model of background influences on holistic raters.
In M. M. Williamson & B. A. Huot (Eds.), Validating holistic scoring for writing
assessment (pp. 237- 265). Hampton Press.
Royer, D., & Gilles, R. (2003). Directed self-placement: Principles and practices.
Hampton Press.
Schendel, E. (1999). Exploring the theories and consequences of self-assessment
through ethical inquiry. Assessing writing, 6(2), 199-227.
Smith, W. L. (1993). Assessing the reliability and adequacy of using holistic scoring
of essays as a college composition placement technique. In M. M. Williamson &
B. A. Huot (Eds.) Validating holistic scoring for writing assessment (pp. 142-205).
Hampton Press.
Stalnaker J. M. E. (1951). The essay type of examination. In E. F. Lindquist (Ed.),
Educational Measurement (pp. 495-532). American Council on Education.
Sun Tsu. (1994). The art of war. Westview.
Terman, L. M. (1916). The measurement of intelligence. Houghton Mifflin.
Veal, R. A., & Hudson, S. A. (1983). Direct and indirect measures for the large-scale
evaluation of writing. Research in the Teaching of English, 17(3), 285-296.
White, E. M. (1994). Teaching and assessing writing: Recent advances in understanding,
evaluating, and improving student performance. Jossey-Bass.
Williamson, M. M. (1993). An introduction to holistic scoring: The social, theoretical,
and historical context for writing assessment. In M. M. Williamson & B. A. Huot
(Eds.), Validating holistic scoring for writing assessment (pp. 1-43). Hampton Press.
Williamson, M. M. (1994). The worship of efficiency: Untangling theoretical and
practical consideration in writing assessment. Assessing writing, 1(2), 147-174.
Wolfe, E. W. (1997). The relationship between essay reading style and scoring
proficiency in a psychometric scoring system. Assessing writing, 4(1), 83-106.
Yancey, K. B. (1999). Looking back as we look forward: Historicizing writing
assessment. College Composition and Communication, 50(3), 483-503.
275
CHAPTER 9.
CRITIQUE OF MARK D.
SHERMIS AND BEN HAMNER,
“CONTRASTING STATE-OF-THE-
ART AUTOMATED SCORING
OF ESSAYS: ANALYSIS”
Les C. Perelman
MIT
On April 16, 2012, Mark D. Shermis, Dean of the School of Education at the
University of Akron, presented a paper at the annual meeting of the Nation-
al Council on Measurement in Education on “Contrasting State-of-the-Art in
Automated Scoring of Essays: Analysis.” Despite its fairly nondescript title, the
paper claimed that machines graded essays as well as expert human raters, a
claim that was publicized in various press releases and newspaper articles. A press
release from the University of Akron, for example, stated, “A direct comparison
between human graders and software designed to score student essays achieved
virtually identical levels of accuracy, with the software in some cases proving to
be more reliable, a groundbreaking study has found” (Man and Machine, 2012).
A headline in Inside Higher Ed read, “A Win for the Robo-Readers,” and the
story included statements such as the following:
The study, funded by the William and Flora Hewlett Founda-
tion, compared the software-generated ratings given to more
than 22,000 short essays, written by students in junior high
schools and high school sophomores, to the ratings given to the
same essays by trained human readers. The differences, across a
number of different brands of automated essay scoring software
(AES) and essay types, were minute. (Kolowich, 2012)
Even the venerable British publication, The New Scientist, reported “The es-
say marks handed out by the machines were statistically identical to those from
the human graders, says [Jaison] Morgan. ‘The result blew away everyone’s ex-
pectations,’ he says.” (Giles, 2012) Yet these reports and other statements can
best be characterized as unsubstantiated overstatement. The study, however, em-
ploys an inconsistent and questionable methodology that favors the machines
over the human graders. Even with these biased procedures and results, the data
still give some, but lacking the full test essay sets, inconclusive indication that in
actual assessments of writing, human scorers were more reliable than machines.
The study derived from the Automated Student Assessment Prize (ASAP), a
competition sponsored by the William and Flora Hewlett Foundation, to assess
the efficacy of automated scoring engines. The competitions involved evaluat-
ing essays from statewide assessments that had already been scored by human
readers. Phase One, which dealt with “long-form constructed responses” (al-
though over half of the responses were essentially paragraphs) had two parts.
The first involved scoring engines developed by nine testing companies such as
the Educational Testing Service, Pearson Knowledge Technologies, and CTB/
McGraw-Hill. The second competition was an open contest among software
developers. The Shermis and Hamner paper reports on only the first part, the
performance of the nine vendors.
278
Critique
279
Perelman
20% each. Consequently, of the total sample size of 22,029 papers, only 4,343
papers in eight different essay sets comprised the actual test sets, ranging from
304 to 601 papers each.
However, half the essay sets and over half the aggregate number of papers
in the test set were not evaluated on any construct connected to writing. The
study defines the four essay sets #3-#6 as source-based writing assignments.
Source-based writing assessments measure student writing in response to spe-
cific texts or data. A prominent example of this kind of assessment is the Docu-
ment-Based Question in the Advanced Placement Language and Composition
Examination (Perelman, 2008). Students are given a passage or a short essay
and then asked to write an argument or analysis about it. Unlike the rubrics
that govern the scoring of essay-sets in the Shermis study, the rubrics for these
AP Examinations emphasize writing skills such as organization, argument, and
expression as well as a student’s mastery of content. Essay sets #3-#6, on the
other hand, contain prompts and rubrics that are not based on document or
source-based writing, however, but are content-based or content dependent
exercises that are scored solely on the understanding of content rather than
any assessment of writing ability.
Two of these essay sets, #3 and #4, are focused solely on literary analy-
sis. Essay set #3 consists of responses to a prompt based on “Rough Road
Ahead” by Joe Kurmaskie: “Write a response that explains how the features
of the setting affect the cyclist. In your response, include examples from the
essay that support your conclusion.” Essay set #4 consists of responses to a
prompt based on “Winter Hibiscus” by Minfong Ho. The prompt repeats the
last paragraph of the story and then asks students to “Write a response that
explains why the author concludes the story with this paragraph. In your re-
sponse, include details and examples from the story that support your ideas”
(see Appendix A).
The rubrics, based on a scale from 0-3, are identical for each of these two
essay sets (#3 and #4). The score of 3, the highest score, is defined in each of the
rubrics by the following language:
Score 3: The response demonstrates an understanding of the complexities of
the text.
• • Addresses the demands of the question
• • Uses expressed and implied information from the text
• • Clarifies and extends understanding beyond the literal
The rubrics and other materials for essay sets #5 and #6 explicitly define
them as reading tests, while defining a different scale for writing tests. (See Ap-
pendix A; Kaggle-Data 2012.) The prompt for essay set #6 is based on an excerpt
280
Critique
discussing the obstacles to putting a mooring mast for dirigibles on top of the
Empire State Building:
Based on the excerpt, describe the obstacles the builders of the
Empire State Building faced in attempting to allow dirigibles
to dock there. Support your answer with relevant and specific
information from the excerpt.
Immediately following the rubric, which focuses on ability to understand the
content, not on ability to write, were these scoring notes.
The obstacles to dirigible docking include:
• Building a mast on top of the building
• Meeting with engineers and dirigible engineers
• Transmitting the stress of the dirigible all the way down the building;
the frame had to be shored up to the tune of $60,000
• Housing the winches and other docking equipment
• Dealing with flammable gases
• Handling the violent air currents at the top of the building
• Confronting laws banning airships from the area
• Getting close enough to the building without puncturing
I took my first and last programming class in 1967, learning Fortran IV.
Even with my rusty recollection of an antique programming language, I am fair-
ly confident that I could construct a program that could also do very well scor-
ing these essays simply by counting strings of key words and phrases, including
synonyms. Such a program, however, would not in any way be assessing writing.
281
Perelman
282
Critique
#3, #4, #5, and #6, all compute the RS as the higher of the two scores.2 This
procedure, followed by four essay sets - half of the total number - and containing
55% of the essays in the aggregate data sample, skews many of the measures used
in favor of the AES scores.
Before going into a more technical explanation, the bias produced by using
these two different measures can be best illustrated by a hypothetical example. A
company is hiring an additional reader to score essays. It has two applicants for
the position and will select the applicant that has the greatest reliability in scor-
ing. A reader who already works for the company has scored all of the essays. The
first applicant scores the essays and her reliability is determined by comparing her
scores to those of the first reader. The second applicant, however, is told before
scoring the essays that his reliability will be determined by comparing his score to
the higher of the two previous readers’ scores if the scores differ. He realizes that his
chances improve dramatically simply by always selecting the higher score in any
case in which he is wavering between two scores. He does so, scores more reliably,
and gets the job. Clearly, the procedure was biased in favor of the second applicant.
This example is completely analogous to the procedure used in the study for essay
sets #3, #4, #5, and #6, which are biased towards the machines. If such a proce-
dure as described in the scenario were actually implemented in the real world, it
would clearly be an unfair hiring practice. Similarly, this practice used in the Sher-
mis and Hamner study unfairly biases and therefore invalidates half the results.
Essays scores, be they holistic, trait, or analytical, always are continuous vari-
ables, not discrete variables (integers), even though graders almost always have
to give integer values as scores. The report recognizes this fact in the observation
on page 24 that values for the Pearson r “might have been higher except that the
vendors were asked to predict integer values only.” Each reader has to select a
single integer value even though some essays might be on the border between two
adjacent integers. Some 3’s on a 4-point-scale might be very high 3’s bordering on
a 4, while other 3’s may be very low 3’s bordering on a 2. Significantly, some of
the training materials for the essay sets included essays scores with plus and minus
signs. In the terminology of Classical Test Theory, the True Score might be 3.3 or
2.8. Consequently, adjacent agreement in the correct direction between two read-
ers (e.g. one rater gives an essay a score of 3 and the second rater gives the essay a
score of 4) will more closely approximate a True Score of 3.4 than two scores of 3.
Resolving scores merely by selecting the higher one ignores the continuous
nature of the scores being measured and penalizes human raters while giving
AES algorithms a substantial advantage by allowing them to optimize agreement
with the RS by rounding up just like the example of the second job applicant.
In the case of the essay that has a True Score of 3.4, for example, there are four
likely pairs of scores that would be produced by two human raters: 3-3, 3-4,
283
Perelman
4-3, and 4-4. Note that in three of these four cases, selecting the higher score
makes 4 the resolved score, and that in two of these three instances one of the
two reader scores will be lower than the resolved score. This unjustified score bias
can be observed in the data in the tables at the end of the report. Table 9.1 and
Table 4 of the study (Appendix B) display the means for the five score sets that
ether combine rater scores to compute RS (#1, #7, and #8) or use a single score
as the RS (#2A and #2B). In these essay sets, the resolved score does not bias the
results against the human readers. Significantly, the means of the human reader
scores match the means of the RS much more closely than those of the means of
the machine scores. Table 9.2 displays the means for the essay sets that used the
higher human rater score as the resolved score. The contrast in the differences
between the human reader means and the resolved score means in the two tables
is striking and provides a powerful illustration of how computing the RS as the
higher scores skews the results against human readers.
Table 9.1. Test Set Mean for Resolved Scores=Sum of Scores or Single Score
H1 H2 RS Diff. Avg. Human Range of AES Range of Diff. of
Rater Means from Mean Scores AES Mean Scores
RS Means from RS Means
1 8.61 8.62 8.62 -0.01 8.49-8.80 -0.13-0.18
2A — 3.39 3.41 -0.02 3.33-3.41 -0.08-0.00
2B — 3.34 3.32 0.02 3.18-3.37 -0.14-0.05
7 20.02 20.24 20.13 0.00 19.46-20.05 -0.67- -0.08
8 36.45 36.70 36.67 -0.09 37.04-37.79 .037-1.12
Separating the essay sets into two groups, those that use a single human score
or a sum of two human scores to compute the resolved score and those that use
the higher score as the resolved score, present two very different sets of values of
the metrics used in the study. Tables 9.1 & 9.2 also demonstrate the substantial
difference for means.
Table 9.2. Test Set Mean for Resolved Scores=Higher Human Rater Scores
H1 H2 RS Diff. Avg. Human Range of AES Range of Diff. of
Rater Means from Mean Scores AES Mean Scores
RS Means from RS Means
3 1.79 1.73 1.90 0.14 1.84-1.95 -0.06-0.05
4 1.38 1.40 1.51 0.12 1.34-1.57 -0.17-0.06
5 2.31 2.35 2.51 0.18 2.44-2.54 -0.07-0.03
6 2.57 2.58 2.75 0.18 2.54-2.83 -0.04-0.08
284
Critique
Similar distinctions can be shown in the other tables. Indeed, in the five
measures of agreement, exact agreement (Table 8), exact and adjacent agree-
ment (Table 10), Kappas (Table 12), Quadratic Weighted Kappas (Table 14)
and the Pearson r (Table 16), the human raters in the group of essay sets
clearly outperform the AES engines in the first three and have mixed results
for the Quadratic Weighted Kappa and Pearson r. Curiously, for the Quadrat-
ic Weighted Kappa (Table 14) the relationship of the two groups is inverted
- human raters in two of the four essay sets that use the higher score as the
resolved score (#3 and #4) as well as score sets #2A and #2B outperform the
AES engines while AES engines outperform human raters in the other essay
sets. This anomaly may partially be an artifact of the Quadratic Weighted
Kappa measuring correspondence not between two raters, as is its intended
use, but between a rater (i.e., the machine score) and the artificial construct
of the resolved score as higher of the two scores. Another possible explanation
is offered by Brenner & Kliebsch (1996) who noted that that quadratically
weighted kappa coefficients tend to increase with larger scales while unweight-
ed kappa coefficients decrease. They noted that “variation of the quadratically
weighted kappa coefficient with the number of categories appears to be stron-
gest in the range from two to five categories” (p. 201). As displayed in Table
3 of the report, the scales for essay sets #3 and #4 consisted of a scale of four
(0-3), while essay sets #5 and #6 consisted of a scale of five (0-4). With the
exception of the four point scale for score set #2B, all the other essay sets had
scales greater than five. For essay set #1 the range of the rubric was 1-6 and
the range of the resolved score was 2-12. For scoring set #2A, the range was
1-6; for scoring set #7, the range of the rubric was 0-12, and the range of the
resolved score was 0-24. For essay set #8, the range of the rubric was 0-30, and
the range of the resolved score was 0-30.
The confusion between human scores and resolved score is found through-
out the text. The report states, for example, on page 22, “all vendor engines
generated predicted means within 0.10 of the human mean for Essay set #3
which had a rubric range of 0-3.” The report, however, is referring to the
mean of the resolved score not the mean of the human raters, which were,
in actuality, lower than the resolved score by 0.11 and 0.17 respectively. (See
Table 9.2)
The standard method for comparing the reliability of machine scores to
human scores is to compare the reliability of the machine scores to each of the
two human scores and then compare those scores to reliability of the human
scorers to each other (McCurry, 2010). In McCurry’s study, as in many others,
humans clearly outperformed machines. Yet the Shermis and Hamner study
instead chose to use different variables for humans and machines.
285
Perelman
The two readers’ individual scores compared to the resolved score (H1 and
H2) are consistently higher than those of the machine scores for all of the
metrics displayed in all of the tables (Appendix B). This phenomenon could
well be an artifact of the individual reader score being a contributing element
to the resolved score. However, of the nine score sets, the two scores of H2, the
second human reader, for #2A and #2B are completely independent of the re-
solved score because reader H1 defined the resolved score and H2’s scores were
used only for computing grading reliability. Consequently, in essay sets #2A
and #2B the human reader score and the machine scores are compared to the
same measure. That the human rater in essay sets #2A and #2B outperformed
all of the machines in every metric except for one machine in Pearson r correla-
tion offers some evidence that the high individual reader scores compared to
the resolved score are not solely an artifact of their being a part of the whole.
As shown in Table 9.3, #2A, which measured ideas, content, organization,
style, and voice, had an exact agreement value of 0.76, compared to the range
of machine values of 0.55-0.70. Its Kappa was 0.62, compared to the range of
machine values of 0.30-0.51. Its Quadratic Weighted Kappa was 0.80, com-
pared to the range of machine values of 0.62-0.74. And its Pearson r was 0.73,
compared to the range of machine values of 0.62-0.74. Similarly, #2B, which
measured conventions of grammar, usage, punctuation, and spelling, had an
exact agreement value of 0.73 compared to the range of machine values of
0.55-0.69. Its Kappa was 0.56, compared to the range of machine values of
0.27-0.49. Its Quadratic Weighted Kappa was 0.76, compared to the range
of machine values of 0.62-0.74. And its Pearson r was 0.76, compared to the
range of machine values of 0.55-0.71. Significantly, the prompt in essay set #2
was a traditional argumentative prompt.
Table 9.3. Essay Set #2-H2 Score Compared to Resolved Score vs.
Machine Scores
Metric 2A 2B
H2 2A Range of H2 2B Range of
Machine Scores Machine Scores
Exact Agreement 0.76 0.55-0.70 0.73 0.55-0.69
Kappa 0.62 0.30-0.51 0.56 0.27-0.47
Quadratic 0.80 0.62-0.74 0.76 0.62-0.74
Weighted Kappa
Pearson r 0.73 0.62-0.74 0.76 0.55-0.71
286
Critique
In sum, the use of the artificially inflated resolved scores skews any mean-
ingful analysis. This is particularly serious in the case of the quadratic weighted
kappa, which is meant to compare the scores of two autonomous readers, not
a reader score and an artificially resolved score. (Sim & Wright, 2005). In a
subsequent report (Morgan, Shermis, Van Deventer, & Vander Ark, 2013), this
same team uses the quadratic weighted kappa as the single measure of “the con-
cordance between hand scores and machine scores” (p. 11), apparently unaware
that they were measuring the concordance between resolved scores and machine
scores and that in addition to all the other problems associated with employing
resolved scores, the quadratic weighted kappa is an inappropriate measure for
such a comparison.
The design of the study allows random chance to produce some seeming-
ly impressive machine scores. Having a pair of readers compete against nine
scoring engines is, in essence, like running multiple T-tests or any other kind
of multiple comparisons. An occurrence can appear significant but might just
be a lucky random occurrence. Any single high machine score among the nine
scores by nine different vendors compared by five different metrics - that is 405
individual measures - could possibly be a random anomaly, or to put it in more
colloquial terms, a lucky guess, especially since the size of the individual test
essay sets were relatively small, ranging from 304 to 601. Statisticians have long
known the dangers of producing what is called a Type I Error or False Positive
when there are multiple comparisons without any overall testing of the entire
model. When deciding if the difference between two variables is significant or
possibly due to random chance, the standard statistical practice is to require that
the probability of the difference being a product of random chance less than
5% or 1 in 20. But with repeated instances or comparisons, the probability of
producing one or more statistically significant events increases. The chance of
rolling two dice and getting two sixes is one in thirty six or 2.7%, but if I roll
the dice thirty times, there is over a 50% chance I will roll two sixes. Although
the case of comparisons in the study is slightly different, the basic analogy holds.
The comparison of 405 measures to the resolved scores will produce some high
correlations merely by chance.
The standard methodology to prevent these kinds of errors is to perform a
test of the model as a whole. Unfortunately, no such tests were performed in
the Shermis and Hamner study. Indeed, although various claims were made
in the paper, no tests of statistical significance were reported by the authors.
Instead, the authors present impressionistic assertions such as “In general, per-
formance on kappa was slightly less with the exception of essay prompts #5
& #6. On these data sets, the AES engines, as a group, matched or exceeded
287
Perelman
human performance” (p. 23). There are no parameters given on what constitut-
ed matching or exceeding human performance.
Eleven months after the paper was presented and widely publicized, the lead
author was quoted in the press as stating that he did not perform a regression
analysis or any other statistical tests on the data in his study because that was one
of conditions imposed upon him by major vendors of essay grading software,
including McGraw-Hill and Pearson (Rivard, 2013). Such conditions were not
disclosed in the Methods Section of the original paper, even though disclosures
of such externally imposed restraints is standard practice in academic publica-
tions and especially in empirical studies such as this one.
288
Critique
Exact agreement is summarized in Table 9.4. The report aggregates the ranges
of agreement for the two human readers H1H2 among all eight essay sets and all
nine rows of data, stating on page 22 that “The human exact agreements ranged
from 0.28 on essay set #8 to 0.76 for essay set #2.” The report then states, “the
predicted machine score and had a range from 0.07 on essay set #2 [sic] to 0.72 on
essay sets #3 and #4. An inspection of the deltas on Table 9 shows that machines
performed particularly well on essay sets #5 and 6, two of the source-based essays.”
The report ignores how human scorers performed better than the machines
for most of the essay sets. Of the nine scores, the human rater agreement coeffi-
cients exceeded the top score of the machines in six of them, tying in a seventh.
In essay set # 1 both readers performed .17 better than the best performing
machine. In essay set #2A, the single “read-behind” reader performed .06 better
than the best performing machine. In essay set #2B, the single “read-behind”
reader performed .04 better than the best performing machine. The next four es-
say sets are content-based reading tests. For essay sets #3 and #4, the agreement
of the two readers outperforms all but one of the machines and ties that one.
The report also makes the careless error of incorrectly attributing the 0.07 exact
agreement to essay set #2 instead of to essay set #7.
289
Perelman
Table 9.5 summarizes the Kappa scores. On page 23, the report states that
“in general, performance on kappa was slightly less with the exception of essay
prompts #5 & #6. On these essay sets, the AES engines, as a group, matched or
exceeded human performance.” While this last claim is true for essay set #5, it
was not true for essay set #6, where the value for H1H2 fell right in the middle
of the machine scores. Moreover, the machine performance was not “slightly”
lower than human performance measured by H1H2, it was substantially lower
for all essay sets except 5 & 6 as can be observed simply by comparing H1H2
with the median and range values of the machine scores in Table 9.5.
Tables 9.6 and 9.7 summarize the scores on the quadratic weighted kappa
and the Pearson r. As mentioned previously, the machines do better on the qua-
dratic weighted kappa except for score sets #2A and #2B and the literary analysis
questions, essay sets #3 and #4. The performance of H1H2, the comparison of
the two readers’ scores, is mixed against the machine scores.
290
Critique
These results, as stated previously, may simply be the artifact of using differ-
ent measures for machines and human readers as well as the improper use of the
quadratic weighted kappa.
CONCLUSION
The study’s numerous and substantial defects clearly undermine its conclusions.
Only three of the eight essay sets used in the study contained scores that as-
sessed students’ ability to write more than a paragraph, and only one of the five
other essay sets contained scores that were concerned with writing ability at all.
Even more disturbing was that, with the exception of essay set #2, the study did
not measure the correspondence between human readers and machine scores
but used different measures for human and machine reliability that artificially
inflated machine performance in half the essay sets. For the one essay set, #2,
in which the study directly compared human and machine reliability, human
readers were clearly more reliable than all of the machines for both of the writing
scores contained in this essay set. Moreover, the study failed to follow standard
statistical practice to guard against false positives and also made its assertions in
the absence of any statistical tests, only based on the impressions of the authors.
Consequently, Professor Shermis and Mr. Hamner should consider formally re-
tracting all versions of this study in print or, at a minimum, respond in print to
the criticisms enumerated in this article. Even with the flawed overall design of
the study, further and rigorous statistical analysis of data may yield some inter-
esting and extremely important information. Moreover, there are pressing policy
decisions that argue for further analysis of these data. This paper has been report-
ed to both the Partnership for Assessment of Readiness of College and Careers
291
Perelman
and the Smarter Balanced Assessment Consortium. The data and conclusions in
this report may inform decisions by these two consortia about the use of auto-
mated essay scoring in the high stakes testing connected to the Common Core
Standards, therefore, it is imperative that the authors publicly post the raw test
set data from this study for rigorous statistical analysis.
NOTE
Although the adjudication rules given for the essay set descriptions for Essay
Sets #3 and #4 do not mention it, examination of the training set revealed that,
like Essay Sets #5 and #6, the resolved score was computed by taking the higher
of two adjacent scores. There were no sets of scores in the training sample for
Essay Sets #3 and #4 that contained pairs of scores that differed by more than
one point, and no third rater scores. Consequently, four of the data sets from,
at most, two states computed the resolved score by taking the higher score if the
two rater scores were not identical. The authors mention, on page 9, instances
in which the higher of the two scores in one essay set (#5) was not the resolved
scores. In the two instances I identified, the two readers’ scores were not adjacent
and the resolved score was probably an adjudicated score.
REFERENCES
Austin, J. L. (1962). How to do things with words. Harvard University Press.
Brenner, H., & Kliebsch, U. (1996). Dependence of Weighted Kappa Coefficients on
the Number of Categories. Epidemiology , 7(2), 199-202.
Giles, J. (2012). AI graders get top marks for scoring essay questions. The new scientist,
2861. [Link]
[Link]
Grice, H. P. (1989). Studies in the way of words. Harvard University Press.
Kaggle. (2012). Data—The Hewlett Foundation Automated Essay Scoring. http://
[Link]/c/asap-aes/data
Kolowich, S. (2012). A Win for the Robo-Readers. Inside Higher Ed. [Link]
[Link]/news/2012/04/13/large-study-shows-little-difference-between-
human-and-robot-essay-graders
Man and machine: Better writers, better grades. (2012). The University of
Akron News. [Link]
dot?newsId=40920394-9e62-415d-b038-15fe2e72a677
McCurry, D. (2010). Can machine scoring deal with broad and open writing. Assessing
Writing, 15(2), 118-129.
Morgan, J., Shermis, M. D., Van Deventer, L., & Vander Ark, T. (2013). Automated
Student Assessment Prize: Phase 1 & Phase 2. [Link]
uploads/2013/02/[Link]
292
Critique
APPENDIX A
Prompts, Rubrics, and Other Materials
From The Hewlett Foundation: Automated Essay Scoring [Data and train-
ing material]. [Link] Copyright 2012 by Kag-
gle. Reprinted with permission.
APPENDIX B
Selected Tables
Source: Mark D. Shermis & Ben Hamner, “Contrasting State of the Art
Automated Scoring of Essays: Analysis.” [Link]
paper/Contrasting-state-of-the-art-automated-scoring-of-Hamner-Shermis/
cad818cdb3b8bd2e2837431618578268548209e1
293
CHAPTER 10.
GLOBALIZING PLAGIARISM
AND WRITING ASSESSMENT:
A CASE STUDY OF TURNITIN
Jordan Canzonetta
Syracuse University
Vani Kannan
Syracuse University
There is nothing immutable about the cheating culture that now exists
in many educational settings worldwide. On the contrary, we know the
values of students can be changed when institutions invest in the right strate-
gies. This has happened in areas related to diversity, gender relations, and
substance abuse—both in the U.S. and overseas. So far, though, promot-
ing integrity has not commanded adequate attention or resources. This
session will explore key drivers of the cheating culture and outline what
it will take to dismantle that culture. It will examine cases where education
institutions have changed how young people think and behave—and how
these lessons can be applied to promoting integrity.
– [Link], 3rd Annual Plagiarism Education Week (emphasis added)
In the keynote address at the 2016 Computers and Writing conference, Jeff
Grabill argued automated writing technologies need to be at the forefront of
disciplinary conversations and actions within the field of composition and rhet-
oric. His speech marks a clear exigence: Globally, millions of students are sub-
jected to writing technologies that writing experts did not design. Grabill argued
disciplinary action is urgent because “students whose community and home lan-
guages are not mainstream are being given bad robots”; because Turnitin is the
most popular writing technology deployed globally; and because so many of
these programs advance “writing as a fundamentally individualized activity in-
volving a student, a computer, and an algorithm” (2016). Popular automated as-
sessment programs have been decried by writing experts because they “align with
the narrow view of writing that was dominant in the more recent era of testing
and accountability, a view that is increasingly thrown into question. New tech-
nologies . . . are for the most part being used to reinforce old practices” (Vojak et
al., 2011, p. 99). Further, these programs fail to use technology that promotes an
understanding of core concepts writing experts believe about writing: “that it is
a socially-situated practice; that it is a functionally and formally diverse activity;
and that it is increasingly multimodal” (Vojak et al., 2011, p. 108).
Grabill’s keynote emerges in a kairotic moment in higher education, as
for-profit assessment companies like Turnitin expand their global reach and be-
gin to deploy “formative” and “summative” writing assessment programs. We
adopt NCTE’s definition of formative assessment: “the lived, daily embodiment
of a teacher’s desire to refine practice based on a keener understanding of current
levels of student performance, undergirded by the teacher’s knowledge of possi-
ble paths of student development within the discipline and of pedagogies that
support such development” (NCTE, 2013b, p. 2). Summative assessment, then,
for the purposes of our framework, refers to “final evaluative judgment” of stu-
dent writing (NCTE, 2013b, p. 2). However, we should mention that Turnitin’s
use of these terms does not appear to align with NCTE’s definitions.
Turnitin’s artificial intelligence for writing assessment, a program called
“adaptive technology,” is now marketed as a cutting-edge product for assessing
student writing. The “Turnitin Scoring Engine” website claims the platform can
“Us[e] your previously-graded sample essays . . . [to identify] patterns to grade
new writing like your own instructors would. Give the Engine a set of samples,
and it will accurately score an unlimited number of new essays quickly and
reliably” (“Turnitin Scoring Engine,” n.d).1 This scoring engine offers to mimic
the behavior of teachers by using algorithmic technology to analyze a teacher’s
prompts and grading comments to produce an evaluative response to student
writing (“Features: Overview,” n.d). Thus, Turnitin’s “intelligent assessment”
1 Because Turnitin is in the process of testing its new assessment platforms, the company’s
technology, language, and website are constantly changing. Thus, the information we refer to may
appear on the website under different headings or may have been otherwise altered.
296
Globalizing Plagiarism and Writing Assessment
alleges to grade papers like humans can on categories of “lexical, syntactic, and
stylistic features of writing, such as word choice and genre conventions. It uses
these features to assess content mastery and genre awareness (“Turnitin Scoring
Engine,” n.d). According to Grabill, such corporate assessment programs are in-
fluencing vast student populations—as Turnitin boasts, “30 million” students—
across the globe (“Homepage,” n.d).
Turnitin’s success in the U.S. is deeply connected to corporate influence in
U.S. universities, heavy reliance on contingent labor, a culture of standardized
testing, hegemonic cultural expectations about writing and authorship, and the
complex web of material factors that shape writing assessment (Chatterjee &
Maira, 2014; Giroux, 2007; Herrington & Moran, 2001; Vie, 2013a; Vojak
et al., 2011). We have three central concerns in this article: Turnitin’s institu-
tionalized plagiarism detection, its move to writing assessment, and its global
expansion. Prominent and respected organizations in the field of composition
and rhetoric, including the CCCC Intellectual Property Committee [CCCC-
IP], the Council of Writing Program Administrators [CWPA], and the Na-
tional Council of Teachers of English [NCTE], have aligned themselves against
the detrimental pedagogical practices advanced by Turnitin (CCCC-IP, 2006;
CWPA, 2003; NCTE, 2013a). Of particular concern is that PDSs demonize
nonnative English speakers and “unwittingly construct international students
as plagiarists” (Hayes & Introna, 2005, p. 55). This important scholarship asks
the discipline to pay particular attention to the rhetorical construction of the
student-plagiarist by PDSs, and the values ascribed to plagiarism, authorship,
and intellectual property. Additionally, now that Turnitin offers an assessment
platform, plagiarism detection technology must be understood in conjunction
with such platforms, as they are now (or will be) packaged and sold together.
This move toward “scalable” assessment, as Grabill suggested, has global im-
plications; from Turnitin’s inception, it has linked integrity, values, and honesty
to its global community of users:
[Link] is currently helping high school teachers and
university professors everywhere bring academic integrity
back into their classrooms . . . We encourage any educator
who values academic honestly to help us take a stand against
online cheating and become a member of the [Link]
educational community. (“About Us,” March 31, 2001)
Although the company now adopts more nuanced rhetorical approaches to
sell their product, this original language is likely still familiar to those who teach,
work, and study in educational institutions. This familiarity is part of its in-
sidiousness—it situates instructors (presumed to be members of the “Turnitin.
297
Canzonetta amd Kannan
METHODS
To attend to these questions, we first offer an overview of Turnitin’s plagia-
rism detection software, mapping the company’s movement towards writing
298
Globalizing Plagiarism and Writing Assessment
299
Canzonetta amd Kannan
300
Globalizing Plagiarism and Writing Assessment
the U.S. have led efforts to both petition against and sue the company, citing
concerns about intellectual property. In a 2007 case, students at McLean High
School in McLean, VA, and Desert Vista High School in Phoenix, AZ, filed
a lawsuit against Turnitin (Zimmerman, 2007). The events that led up to the
eventual filing of the lawsuit in March 2007 began in September of 2006, when
a group of students at McLean High School circulated a petition to oppose the
mandatory submission of their work to a newly adopted [Link]: “[t]he
petition, which garnered 1,190 student signatures of the approximately 1800
students that attend the school requested that the mandate to submit work to
Turnitin be removed and that an ‘opt-out’ option be allowed” (Zimmerman,
2007).While students did not win the case, their work to contest Turnitin’s use
of student intellectual property, and the call for the student choice to “opt-out”
of Turnitin (mirroring movements to “opt-out” of standardized testing) drew
attention to the negative impact of PDSs, and the corporatization of education
more broadly, on students. Unfortunately, neither these lawsuits nor repeated
criticisms of PDSs have impacted Turnitin’s widespread adoption by educational
institutions, but the company has shifted its marketing rhetoric from “catching
plagiarists” to “meet[ing] exigencies” in our field to both deflect criticism and
respond to the labor crisis in higher education (Vie, 2013b).
In the current iteration of the website, the word plagiarism only appears on
the main page twice (in smaller text than other language on the page) under
subheadings; this is a departure from its early website iterations, which fore-
ground anti-plagiarism zeal (“About Us,” Wayback Machine, March 31, 2001;
“Homepage,” n.d). Despite Turnitin’s move towards broader writing assessment
technologies, it still uses problematic plagiarism detection software. Its plagia-
rism detection “tool” can only provide students and teachers with a report con-
taining percentages of text that corresponds to various sources on the Internet,
sources in its database, and periodicals, journals and publications, and cannot
infallibly identify plagiarism (“FAQ,” n.d; Purdy, 2009, pp. 65-67).With Turni-
tin’s increased presence in global writing assessment technology, PDSs become
more problematic when we consider the effects they have on nonnative English
speakers. Hayes and Introna (2005) suggest PDSs may inhibit some ELL stu-
dents who are trying to participate in the writing process, but are stymied in
their attempts because the detective component of the programs “limit[s] the
opportunities and time that students have to learn how to write in the new
western, not to mention subject specific, educational context” (p. 67). The use of
PDSs at the onset of the composing process implies students have higher stakes
for writing in new cultural contexts. Without having the chance to learn about
new practices in those environments, students are discouraged from taking risks,
“experiment[ing],” or “observ[ing]” (p. 67).
301
Canzonetta amd Kannan
Current PDS platforms, then, are shaping educational space so that students
are castigated for departing from Edited American English (EAE) and western
ideals about singular authorship, as Introna and Hayes (2011) explained:
Plagiarist practices are often the outcome of many complex
and culturally situated influences . . . [E]ducators need to ap-
preciate these differing cultural assumptions if they are to act
in an ethical manner when responding to issues of plagiarism
among international students. (p. 215)
Originality/singularity is not globally accepted as the primary theory of au-
thorship; not all students are asked to produce original work, and imitation can
often be a staple in some writing processes (Hayes & Introna, 2005, p. 59).
Thus, Turnitin’s emphasis on originality/singularity elides a complex cultural un-
derstanding of plagiarism and authorship.
These underlying ideologies of original/singular authorship were laid bare and
explicitly connected to culture in Turnitin’s “Plagiarism Education Week” event
“Copy/Paste/Culture.” Held April 20-24, 2015, the conference was marketed as
investigating “how current global trends are affecting our values, especially those
related to education, and proposing strategies on how we can address these chal-
lenges. #integrity2015.” The conference focused on how to dismantle the “culture
of plagiarism,” variously described as a “mindset” of narcissism and entitlement
(Hoyt, 2015). As the conference description shows, “our values” are presumed to
align with western constructions of authorship. Indeed, something as banal and
familiar as the hashtag “integrity”—a word that students and teachers are likely
used to seeing mobilized in discussions of plagiarism—immediately connects in-
tellectual property to character, and by extension, plagiarism to poor character.
The “Plagiarism Across Europe and Beyond” conference proceedings echo
these stark character judgments, and explicitly situate them in terms of a geo-
graphic binary of west and nonwest, including designations of “high trust” versus
“low trust” societies and populations (Burkatzki, Platje, & Gerstlberger, 2013, p.
171). In this framework, it becomes the duty of the west (and PDSs) to counter
tolerance towards plagiarism, export knowledge, and modernize culture. Through
this mapping of nonwest, the proceedings constitute and consolidate geographic
sites for corporate/state-level plagiarism detection intervention, with the assump-
tion that Turnitin possesses the correct values of authorship. Howard (1999) ex-
plained such rhetorics are largely related to archaic constructions of plagiarism,
and don’t allow much space for cultural variance in writing processes:
For the past century and more, [western] academic textual
values have been relatively unified, ascribing four properties to
302
Globalizing Plagiarism and Writing Assessment
303
Canzonetta amd Kannan
304
Globalizing Plagiarism and Writing Assessment
305
Canzonetta amd Kannan
306
Globalizing Plagiarism and Writing Assessment
of data on their own classes, since the parties that control the
spin put on this information will have the last word in every
forum. (p. 180)
Thus, it is important to critically interrogate Turnitin’s rhetorics of formative
assessment, which obscure the company’s cooptation of student data and poten-
tial to undermine writing program goals.
Furthermore, Deborah Harris Moore (2013) contends that the fear caused by
surveillance can be disempowering to students: “Using fear as a deterrent . . . is
unethical because it forces students into behaviors based on their perceived pow-
erlessness . . . [S]tudents may see [this technology] as an all-seeing, determining,
and surveying mechanism” (pp. 110-111). After the McLean High School lawsuit,
this culture of surveillance now appears to be taken for granted by many students,
who, according to instructors, view Turnitin as either an “arbitrary hoop” to jump
through to submit their papers, or as a “psychological deterrent” and “authority”
on plagiarism (Canzonetta, 2014, pp. 21-33). Turnitin’s database was initially de-
signed for this purpose—to deter students from plagiarizing by invoking its vast,
national collection of student writing (Zimmerman, 2007).
Beyond serving as a deterrent to plagiarism, Turnitin has seized the opportuni-
ty to exploit the current labor crisis in higher education.2 As Herrington and Mo-
ran (2001) noted, “when human labor is in crisis, we often turn toward technology
to mitigate human stress and loss of funding to alleviate insufficient staffing” (p.
220). Indeed, the company has positioned Revision Assistant as an ally and re-
source for overworked teachers, arguing that it “takes many of the challenges of
continuous feedback out of the teaching equation, such as the pressure on instruc-
tors to provide consistent, timely feedback for all of their students . . . teachers are
provided with a better picture of each student’s progress when making a final as-
sessment” (“Features: Overview,” n.d). By offering a tool to lighten workloads and
the pressures of promptly returning students’ work with feedback (“Customers,”
n.d), the company appeals to administrators whose instructional staffs are either
overburdened or understaffed; for those who may not share composition and rhet-
oric’s critiques of PDSs, Turnitin is proffered as a solution to the complex problem
that grading writing presents. The artificial intelligence Turnitin is testing claims
to be for students, and for teachers who need more time; it instead appears to be
a band-aid for upper-level university administrators who would rather put money
into a technological “panacea,” as Marsh (2004) wrote, than contend with hiring
more faculty. Instead of learning about students, Turnitin’s formative assessment
2 In the U.S. in 2012-2013 academic year, approximately 76% of higher education’s instruc-
tional staff consisted of contingent laborers (Curtis & Thornton, 2013, p. 8).
307
Canzonetta amd Kannan
model learns teachers and their behaviors, assesses generic writing processes, and
supplies an automated response to a perceived problem. Considering the contin-
gent positions that many writing instructors occupy, and the money-saving im-
perative of corporatizing universities, Turnitin’s formative assessment model poses
a major threat for agency and autonomy within writing programs. The data pro-
duced through this program could have serious implications for instructors’ job
security if students aren’t achieving scores administrations approve of—scores that
could be set and established by Turnitin.
What, then, are the implications of these moves in light of Turnitin’s expan-
sion abroad? Rhetorical links between adaptability, assessment, plagiarism, and
pedagogy are visible in the “Plagiarism Across Europe and Beyond” conference
proceedings, and Turnitin is cited by many presenters as a positive pedagogical
tool that offers opportunities for teachers to craft formative assessment peda-
gogies that directly result in lowered instances of plagiarism. Indeed, formative
assessment is implicitly used to justify the use of Turnitin (Meacheam & Faifua,
2015, p. 45). Our analysis of the conference proceedings reveals a particular
emphasis on rhetorics of integrity and consistency, linking western values of au-
thorship with standardization across institutions and geographies. Of particular
note is a reference in a keynote address to the monetary investment (€ 300,000)
the European Union designated for the project Impact of Policies for Plagiarism
in Higher Education Across Europe (IPPHEAE), conducted between 2010 and
2013. In this discussion, presenters asked:
What impact did the project have on national and institu-
tional policies for academic dishonesty and plagiarism? What
evidence is there that policies for academic integrity in higher
education in different parts of Europe are fit for purpose?
How can institutions be sure their policies are effective and
being applied consistently? What more needs to be done?
(Glendinning, 2015, p. 7)
Through this neoliberal rhetoric of fitness (Dingo, 2012), we see a clear call
for uniformity in coping with plagiarism—a pedagogical problem that, as com-
position and rhetoric scholarship shows, is highly contextual and occurs on a
“continuum,” not in a vacuum (Sutherland-Smith, 2008, p. 8). Similarly, pre-
sentations in both 2013 and 2015 advocated worldwide implementation of an
“ANTIPLAG system” that has been adopted in Slovakia and is now enforced
there by law:
the SK ANTIPLAG system (a central repository of theses and
dissertations, a plagiarism detection system, a comparative
308
Globalizing Plagiarism and Writing Assessment
309
Canzonetta amd Kannan
Interestingly, in the proceedings, calls for consistency are paired with presen-
tations calling for contextual understandings of plagiarism, incorrectly suggesting
that the conference represents a fair debate and echoing Turnitin’s uptake of dis-
ciplinary critiques of PDSs: “Every single instance of plagiarism is unique and
requires careful examination of all the circumstances and facts, but universal stan-
dards on the systematic level also should exist and serve as a prevention of plagia-
rism and other types of research misconduct” (Vasiljevienė & Jurčiukonytė, 2015,
p. 164). However, these gestures mean little when set alongside the framing of the
conference and its broad geographical consolidation, “across Europe and beyond.”
Considering Turnitin’s partial sponsorship of this conference, paired with their
new initiative to implement formative assessment technology, we have to con-
sider that rhetorics of standardization and consistency are beneficial for Turnitin’s
business model, and promises of contextual specificity will be necessary in order
to persuade fields like composition and rhetoric to adopt its assessment program.
CONCLUSION
Scholarship in composition and rhetoric defines plagiarism as a highly contextu-
al, case-by-case pedagogical issue. Turnitin’s assessment platform could be used
to execute standardized assessment of writing “across Europe and beyond,” as the
conference title indicates. Corporate PDSs, thus, have the potential to standard-
ize student writing itself, potentially on a global scale. Turnitin has set the stage
for and monopolized the plagiarism detection market—the end results of which
are promoting singular, original conceptions of authorship globally. Now more
than ever, it is time for rhetoricians and compositionists to “use our own ped-
agogies and technologies . . . [and] fix our gaze on the millions of learners who
are being taught with technologies made by people who know very little about
writing and learning to write” (Grabill, 2016). Scholars within the field have
begun to develop new writing technologies, such as Eli Review, a program cre-
ated by writing experts—Grabill, Hart-Davidson, and McLeod—at Michigan
State University (“About Eli Review”). Other scholars have endorsed emergent
technologies; for example, Les Perelman, famed debunker of the robo-graders
supports the technology WriteLab (Berdik, 2015). However, the discipline still
faces the problem of globalized models for standardized plagiarism detection
and writing assessment.
Who will benefit from globalized programs like Turnitin’s, and who will be
left out? Herrington and Moran (2001) warned:
The marketing muscle of these testing companies, and the
concurrent expansion of the computer-as-reader of students’
310
Globalizing Plagiarism and Writing Assessment
REFERENCES
Please note that all citations that link to Turnitin in this article are no longer live or
retrievable.
Barlow, L., Liparulo, S. P., & Reynolds, D. W. (2007). Keeping assessment local:
The case for accountability through formative assessment. Assessing writing, 12(1),
44-59.
Berdik, C. (Sept., 2015). A critic’s second thoughts on robo-grading. Bostonglobe.
com.
Bretag, T. (2015). Enacting academic integrity—it takes courage. In Plagiarism Across
Europe and Beyond 2015 Conference Proceedings (p. 6). Brno: Mendel University.
Burkatzki, E., Platje, J. & Gerstlberger, W. (2013). Cultural differences regarding
expected utilities and costs of plagiarism between high-trust and low-trust
societies—preliminary results of an international survey study. In Plagiarism Across
Europe and Beyond 2013 Conference Proceedings (pp. 171-191). Brno: Mendel
University.
Canzonetta, J. (2014). Plagiarism detection services: Instructors’ perceptions and uses
in the first-year writing classroom (Master’s thesis). [Link]
docview/1553839828
311
Canzonetta amd Kannan
312
Globalizing Plagiarism and Writing Assessment
Herrington, A., & Moran, C. (2001). What happens when machines read our
students’ writing? College English 63(4), 480-499.
Hesford, W. & Schell, E. E. (2008). Introduction: Configurations of transnationality:
Locating feminist rhetorics. College English, 70(5), 461-470.
Homepage. (2016). [Link].
Howard, R. (1999). Standing in the shadow of giants: Plagiarists, authors, collaborators. Ablex.
Howard, R. (2000). Sexuality, textuality: The cultural work of plagiarism. College
English, (62)4, 473-491
Hoyt, S. M. (2015). Copy/Paste/Culture: Plagiarism education week at K-state
libraries. K-State Today, [Link]
php?id=19646
Introna, L. D., & Hayes, N. (2011). On sociomaterial imbrications: What plagiarism
detection systems reveal and why it matters. Information and organization, 21(2),
107-122. [Link]
Janssens, K. & Tummers, J. (2015). A pilot study on students’ and lecturer’s perspective
on plagiarism higher professional education in Flanders. In Plagiarism across Europe
and beyond 2015 conference proceedings (pp. 12-23). Brno: Mendel University.
Kannan, V. (2014). Rhetorics of song: Critique, persuasion, and education in Woody
Guthrie and Martin Hoffman’s “deportees” (Master’s thesis). UMI Dissertation
Services from ProQuest.
Kokkinaki, A., Iacovidou, M. & Demoliou, C. (2015). Students’ perceptions on
plagiarism and relevant policies in Cyprus. In Plagiarism across Europe and beyond
2015 conference proceedings (pp. 192-200). Brno: Mendel University.
Kravjar, J. (2015). SK antiplag is bearing fruit. In Plagiarism across Europe and beyond
2015 conference proceedings (pp. 147-163). Brno: Mendel University.
Kravjar, J. & Noge, J. (2013). Strategies and responses to plagiarism in Slovakia. In
Plagiarism across Europe and beyond 2013 conference proceedings (pp. 201-215).
Brno: Mendel University.
Krokoscz, M. & Putvinskis, R. (2013). Analysis of the perceptions of undergraduate
students in business administration on the occurrence of academic plagiarism in
Brazil. In Plagiarism across Europe and beyond 2013 conference proceedings (pp. 281-
282). Brno: Mendel University.
LightSide Labs. (2015). [Link]
Marsh, Bill. (2004). [Link] and the scriptural enterprise of plagiarism detection.
Computers and Composition, 21(4), 427-438.
Meacheam, D. & Faifua, D. (2015). Perspectives on turnitin use in an Australian
setting. In Plagiarism across Europe and beyond 2015 conference proceedings (pp. 37-
53). Brno: Mendel University.
Moore, D. H. (2013). Instructors as surveyors, students as criminals: Turnitin and the
culture of suspicion. In M. Donnely & R. Ingalls (Eds.), Critical conversations about
plagiarism (pp. 101-118). Parlor Press.
National Council of Teachers of English. (2013a). (2013). Resolutions & sense of the
house motions. Resolution 3. National Council of Teachers of English. [Link]
[Link]/cccc/resolutions/2013
313
Canzonetta amd Kannan
314
Globalizing Plagiarism and Writing Assessment
315
EDITORS AND RETROSPECTIVE
CONTRIBUTORS
Laura Aull is Associate Professor and Writing Program Director at the Univer-
sity of Michigan, where she teaches English linguistics and writing pedagogy.
She is editor of the Assessing Writing Tools & Tech Forum, and she is the author,
most recently, of How Students Write: A Linguistic Analysis and the forthcoming
book You Can’t Write That: 8 Myths about Correct English.
Carolyn Calhoon-Dillahunt, a former CCCC Chair and TYCA Chair,
teaches writing at Yakima Valley College, an open admissions Hispanic-serv-
ing Institution. She also helps coordinate program and institutional assessment
within the Arts & Sciences division of the college and is engaged in departmen-
tal and college-wide equity work. Her scholarly interests center on pedagogy,
assessment, and education policy. She has published articles in The WPA Journal,
TETYC, and CCC and has co-authored chapters in New Directions for Commu-
nity Colleges and the recently published collection, Writing Placement in Two-
Year Colleges: The Pursuit of Equity in Postsecondary Education.
Brian Huot has been a full time writing teacher and writing program admin-
istrator since 1980. Currently he is Professor of English at Kent State University.
He is past chair of the College Section Committee and Member of the NCTE
Executive Committee (2006-2008) and a current member of the Council of Writ-
ing Program Administrator Executive Board. He is a contributing scholar to the
literature on the teaching and assessing of writing and has served as consultant for
various institutions. He is currently a member of the NCTE Consulting Network.
Diane Kelly-Riley is Professor of English and Vice Provost for Faculty at the
University of Idaho. She studies writing assessment theory and practice, validity
theory, race and writing assessment, public humanities and multimodal compo-
sition. She was editor of the Journal of Writing Assessment from 2011-2022. She
published Improving Outcomes: Disciplinary Writing, Local Assessment and the
Aim of Fairness with Norbert Elliot (MLA, 2021).
Ti Macklin is the Director of First-Year Writing at Boise State University
where she teaches courses in composition and rhetoric. Her research interests
lie largely in First-Year Writing and writing assessment with a particular focus
on assessment at the individual, classroom, and programmatic levels. Her most
recent work examines the experiences of graduate and undergraduate students
in first-year writing. She served on the editorial staff of the Journal of Writing
Assessment for nine years.
317
Editors and Retrospective Contributors
318
CONSIDERING STUDENTS, TEACHERS, AND
WRITING ASSESSMENT
The editors and authors in this edited collection, available in two volumes,
consider the increasing importance of students’ and teachers’ lived experiences
within the development and use of writing assessments. Presenting key
work published in the Journal of Writing Assessment since its founding in
2003, the collection explores five major themes: technical psychometric
issues; politics and public policies shaping large scale writing assessments;
automated scoring of writing; fairness; and the lived experiences of humans
involved in assessment ecologies. The book also provides reflections from
leading writing assessment scholars who examine how these themes continue
to shape current and future directions in writing assessment.
Diane Kelly-Riley is Professor of English and Vice Provost for Faculty
at the University of Idaho. She studies writing assessment theory and
practice, validity theory, race and writing assessment, public humanities and
multimodal composition. She was editor of the Journal of Writing Assessment
from 2011-2022. Ti Macklin is the Director of First-Year Writing at Boise
State University, where she teaches courses in composition and rhetoric. Her
research interests lie largely in first-year writing and writing assessment. Her
most recent work examines the experiences of graduate and undergraduate
students in first-year writing. She served on the editorial staff of the Journal
of Writing Assessment for nine years. Carl Whithaus is a Professor of Writing
and Rhetoric at the University of California, Davis. He studies the impact
of information technology on literacy practices, writing assessment, and
writing in the sciences and engineering. His books include Multimodal
Literacies and Emerging Genres (University of Pittsburgh Press, 2013) and
Teaching and Evaluating Writing in the Age of Computers and High-Stakes
Testing (Erlbaum, 2005).
Perspectives on Writing
Series Editors: Rich Rice and J. Michael Rifenburg
ISBN 978-1-64215-216-6