Placement Testing
Placement Testing
This report describes a reliability and validity study of the placement test which
is used at the University of Surrey as a means of identifying students who may
require English language support whilst studying at undergraduate or postgraduate
levels. The English Language Institute is charged with testing all incoming stu-
dents, irrespective of their primary language or subject specialization. F[ir and
accurate assessment of student abilities, and refening individuals to appropriate
language support courses (the in-ressional programme), is an essential support
service to all academic departments. The goal of placement testing is to reduce
to an absolute minimum the number of students who may face problems or even
fail their academic degrees because of poor language ability or study skills. This
study looks at the administrative and logistic constraints upon what can be done,
and assesses the usefulness of the placement test developed within this context.
I Introduction
It has recently been noted that although placement testing is probably
one of the most widespread uses of tests within institutions, there is
relatively little research literature relating to the reliability and val-
idity-of such measures (Wall, Claptram and A-lderson, 1994):'Publi-
cations which deal with placement tests frequently provide qualitative
assessments of instruments (Goodbody, 1993), or are concerned with
the placement of linguistic minority students in programmes which
are not related to language teaching (Schmitz and DelMas, 1990; Tru-
man, 1992). Wall, Clapham and Alderson ( 1994) offer one of the
few empirical studies of a placement test which is designed to screen
students entering a British university for deficiencies in English lang-
uage skills, which might impede their progress in undergraduate or
postgraduate studies. They investigate face validity (student percep-
tions of the test), content validity (asking whether tutors thought that
test content represented programme content), construct validity
(through a correlational study), concurrent validity (with self-assess-
ment, and the assessment of tutors) and reliability. The main problem
they discovered in attempting to validate the University'of Lancaster
placement test was in finding appropriate external criteria to conduct
a concurrent study.
II Background
The English Language Institute (ELI) has a number of roles within
the University of Surrey, ranging from its support function in provid- l
ing English language and study skills courses for students at the uni-
versity, English for Academic Purposes (EAP) for overseas students
who intend to study in an English-medium tertiary institution, and an
MA in linguistics (TESOL) for teachers of English, in distance-learn-
ing mode. As part of its service function related to the in-sessional
programme, the ELI is charged by the university to administer an
English language test to all students entering the university on taught
courses (undergraduate and postgraduate, English as primary langu-
age speakers, and speakers of other languages). The purpose of the
test is to identify students whose lack of language skills or ability
to communicate may cause problems in their academic work with
their departments.
----A-major-eonstraint in assessing studems is -tha[ togeth-erwith-
organization (handing out and collecting papers, giving verbal
instructions), the ELI must complete testing within one hour for each
test administration. This means that the test cannot exceed 45 minutes,
and must therefore be restricted in its length and content. Secondly,
all scripts must be processed within five days of the day of the test,
and results disseminated. In 1994, 1619 students took the test. The
number of markers available, and the number of scripts which each
can process in this time frame, also dictates that much of the paper
be objectively scored. However, no language test can use only one
response format, or sample only one construct, if it is to be considered
valid on even prima-facie grounds (APA, 1985: 73; 75). For this
reason, a writing component is included.
The placement test acts as a screening device to reduce the number
of students who attend an oral interview. It is those who are invited
to come to the ELI for an oral interview who have the greater prob-
ability of requiring English language support. These students are
Glenn Fulcher 115
are 'popular' in that the passages are written for the interested
(intelligent) layperson, not for the specialist.
2 Pilot study
The new format was piloted during the summer of 1994, using 67
students attending the ELI Summer School English for Academic Pur-
poses (EAP) programme. All items in Sections 2 and 3 of the test
were studied using classical item analysis. Only items with a facility
index between 0.3 and 0.8 were retained, and items with a discrimi-
nation index of less than 0.3 were either abandoned, or rewritten and
piloted for a second time. The point biserial correlation for all remain-
ing items was above 0.4. Four pilot versions of the test were originally
wriffen, and 75Vo of all questions discarded, leaving one operational
test with items drawn from all four pilot tests. This is a reminder that,
no matter how experienced one may be in test development, there is
always a need for pretesting all items before tests become operational,
and decisions are taken on the basis of test results.
In Section l, writing samples were collected, graded by six tutors,
and features of performance at each level of a rating scale established.
Prototypical samples from each band level were then used in the train-
of raters prior to the operationalization of the test in October 1994.
ing
IV Method
The final version of the test was used operationally with the univer-
sity's entire intake on taught (undergraduate or postgraduate) courses
in October 1994. The- main-studies were all carried out using thls
population.
I Reliability
In order to establish estimates of the reliability and validity of the
. placement test, a number of approaches were taken. In the investi-
gation of reliability, correlation coefficients, means and standard devi-
ations (inter- and intrarater reliability), were established for rating
patterns on Section I of the test. For Sections 2 and 3, it was decided
to fit a logistic model. The reason for this decision was essentially
because of the need, within a university setting, to have multiple
forms of a test, and for the interpretation of scores on these forms to
be comparable.
This implies that forms must be equated in some way. Using a
logistic model, it is possible to create multiple parallel forms, using
'anchor' items from previous form(s) which can be used to calibrate
Glenn Fulcher ll7
new items in later forms, even if this is difficult for short tests. In
the first attempt to fit a Rasch model to the data, it became clear
that there were two distinct populations within the total test-taking
population. These distinct populations are those whose primary langu-
age is English, and those whose primary language is other than
English. Consider Table 1, which is a simple )( test of significance
of primary language classification and referrals, on the basis of the
test scores (for a discussion of cut scores for referrals, see Section
VI below). Using Yates' correction for a 2x2 gnd, )(=516.61, a
highly significant result, indicating that these are certainly separate
populations. [t is not possible, unfortunately, to isolate differences
between referrals or test scores between learners whose primary lang-
uage is not English, because of the small n size of many of the pri-
mary languages represented by the 614 overseas students tested.
The question which arose, therefore, was on which population to
standardize the objective components of the test. It was finally
decided to standardize on the group which did not have English as a
primary language, as this is the population which is more likely to
require English language support. Those speakers of English as a pri-
mary language (73 in this sample) whose score profiles are similar
to those of the referrals of the norming population will still be ident-
ified by the test, and remedial action can be taken.
The Rasch model allows the test developer to calibrate both item
difficulty and learner ability to the same scale (Crocker and Algina,
1986: 340-41), measured in logits. This was done using the program
RASCAL (Assessment Systems Corporation, 1994) for Sections 2
and 3 of the test separately. This was done as there was no theoretical
reason to stl-spect ttratrhscorfihined-scores of Section 2 and 3 would
represent a unidimensional scale. The disadvantage of this approach,
however, is that Section 2 consists of only ten items, and Section 3
of eight items. Although n = 614, the low number of items inevitably
reduces test reliability. This could not be avoided, because of the time
constraints as described above.
Tabfe 1 X table to compare primary and nonprimary English language speakers to test
whether they belong to the same test-taking populaUon
Rater 2 0.92
Rater 3 0.87 0.94
Rater 4 0.75 0.83 0.93
1V Retiability o
I Section I: writing
To establish the reliability of the assessment of writing samples, 20
essays were selected from the population, and were marked by four
tutors. After a period of two weeks, the tutors were then asked to re-
mark a subset of six samples. This allows the calculation of inter-
and intrarater reliability. Table 2 shows the results of the inter-rater
reliability study. It can bg_seen 1[4gelgglqgplbetwg-e1 raters is well
within acceptable reliability ranges for this type of test, with an aver-
age Pearson product moment correlation of 0.87. The average band
awarded was 5.46, with a standard deviation of l.2L Table 3 shows
that the four raters did not differ significantly from this in their indi-
vidual grade profiles.
In the intrarater reliability study, the average reliability coefficient
was 0.69, somewhat lower than the inter-rater coefficients, but still
not so low as to cause undue worry. This figure indicated a need for
further rater training before the second operational testing session, in
2 Section 2: structure
Rasch scaling was conducted using RASCAL (Assessment Systems
Corporation, 1994). Table 4 shows the questions in order of difficulty,
from the easiest to the most difficult, in logits, together with the stan-
dard error associated with the difficulty estimate, and the fit )(
statistic.
From Table 4 it can be seen that there are three misfitting items.
That is, they do not meet the criterion of unidimensionality in this
subtest, and must therefore be removed, and replaced by other items
in rhis form of the test. It is instructive to return to misfitting iterns
to attempt to provide a linguistic rationale for why they misfit. In this
case, it is particularly enlightening to look at item 5, because of the
very large misfit statistic. Item 5 is:
Eac
696
uJo
Ability
Figure I Test characteristic curve for Section 2 (structure)
122 An English language placement test
3 3: reading comprehension
Section
Table 5 shows the results of the Rasch analysis for Section 3. Only
item 17 was found to misfit, and this item was therefore removed
c
o
l'
E
L
€g
Ability
Figure 2 Test information curve for Section 2 (structurel
Glenn Fulcher 123
from this form of the test. With the test centred on item difficulty,
mean item difficulty was 0.00, and the standard deviation of difficulty
1.48. Average ability was {.01, with a standard deviation of 1.22.
The test characteristic curve for Section 3 is presented in Figure 3.
The curve in Figure 3 is not as steep as that in Figure 2, but this is
only to be expected, with only eight questions in the subtest. It is
difficult to achieve reasonable discrimination with so few items in a
subtest. Similarly, the reliability coefficient (equivalent of KR20) is
only 0.59. The test information curve in Figure 4 is much flatter than
the curve in Figure [Link] an ideal world, the test should be lengthened,
as the discussion of this subtest in section VI shows.
The test has been found to be reasonably reliable for its purpose.
However, the reliability coefficients do not meet the levels which
E'E
coF e [Link]
696
ulo
g
o
(!
E
o
c
Ability
Figure4 Test characteristic curve for Section 3 (reading)
n- ry4:-g
rtt(l rttd)
-
where n is the number of times the test length must be increased with
items or raters to achieve the desired reliability, rttd is the desired
level of reliability and rtt is the present observed reliability of the
it would
test. For Section 2, to achieve a reliability coefficient of 0.80,
be necessary to increase test length by 2.34 (28 questions), and for
Section 3, 2.77 (22 questions). This would essentially double the
administration time of the test!
Compromises need to be made between desired levels of test
reliability and the lack of time for testing within a large educational
organization. Such decisions cannot be taken by the test designers
alone, but need to be discussed at all levels within the institution
in relation to general testing policy, including the assessment of the
consequences of the unreliability which the institution is prepared
to tolerate
Glenn Fulcher 125
VI Validity
1 Correlational and principal components analysis
The test was designed to measure three separate abilities: writing,
English structure and reading. If this is actually the case, the corre-
lation coefficients between the three sections should be modest, and
any attempt to factor the matrix should show that subtests load on
different factors. Table 6 provides the correlation matrix in the bottom
triangle, and the reliability coefficients of the subtests in the diagonal.
The upper-right triangle presents the correlation coefficients corrected
for attenuation (taking unreliability of measures into account).
Although all are significant (to be expected, in subtests which are all
language related), variance overlap is not so high that we would wish
to challenge the view that each subtest is tapping some unique ability
as well as a common ability. It should also be noted that each of
the two writing tasks also appears to be tapping different skills to
some degree.
The correlation mahix was subjected to principal components
analysis (PCA), in order to discover if the tasks were loading on
different factors. Eigen values for the extraction of factors were set
at 0.6, on the assumption that some of the factors in which we are
interested are indeed, as Farhady (1983: 18-19) argues, less than
unity. Three factors were extracted, ild rocated using the Varimax
technique. The results are presented in Table 7. The three-factor sol-
ution is clearly interpretable in relation to the correlation matrix. Fac-
tor I is a writing factor, factor 2 is a reading factor and factor 3 is a
structure factor, which the argumentative essay also loads on. How-
ever, it must be Ctrbssed ih-at-thii-is nof iuffibiant evidence upon which
to claim construct validity. Exploratory analysis of this type is sugges-
tive of further avenues for investigating construct validity, but can
never be sufficient to meet the requirement for convergent and diver-
gent evidence.
cases 1184
No. 1184 1029 1 159 1 184
Minimum 0 _l) 4.64__ -2.77 0
Maximum I I 2.95 2.74 33
Range 9 I 3.59 5.51 3i!
Mean 6.92 6.85 1.85 0.63 26.57
sD 2.14 1.01 0.88 1.10 3.25
Standard enor0.06 0.029 0.03 o.03 0.10
x01234567890 T
y 10 2 2 13 36 1U 196 32 0 0 0 435
n00004302417091692110 1184
r 10 2 2 13 /t{t 174 437 741 169 21 10 1619
x01234567890 T
y3206847 151 168A2 100 435
n00021232556941781912 1184
T 32 0 6 10 48 174 423 716 179 19 12 1619
x0 1 23456789 10 o T
yI 2 12 16 29 39 Tf 98 80 56 17 0 435
n0 0 0 0 1 10 74 230 375 338 146 10 1184
T9 2 12 16 30 49 151 3?8 455 394 163 10 1619
128 An English language placement test
Table 14
Refenals by scores on Reading
x012345678OT
y162369109111762542043s
n O 10 40 151 274 361 239 92 15 11 1184
T 16 33 109 260 385 437 255 96 17 11 1619
and reading, respectively. Students who did not answer a question are
counted as 'omitting' (O). Score ranges where numbers are high-
lighted in bold represent the areas in which there is potential for
errprs of judgeme-nt to be made yhq [Link] referrals.
Ttre cut score for tasks I and 2 (Section l) was 6, for Section 2,
a raw score of 6 (or ability level 0.38) and Section 3, a raw score of
4 (or ability level 0.02). The overall cut score was 22. Raters wetb
asked to consider the evidence from each individual section in making
a decision regarding referral. For example, if the overall score was
22+, and on only one of the tasks did the score fall below the cut
point, then they were probably not to refer the student. This accounts
for some of the nonreferrals in band 5 on both the writing tasks, a
score of 5 on Section 2 and a score of 2 or 3 on Section 3, as well
as those who were referred within a number of bands above the cut
score.
In the case of Sections 2 and 3, however, w€ are able to calculate
the standard error of the cut score, because of the methodology which
has been used. The standard error associated with an ability level of
[Link] is 0.852, meaning that we can be 95Vo confident that a scaled
- cut score of 0.02 is 0.02+ 1.67, or anywhere ffim -J35To l.S9. That
is a raw score from 3 to 7 . The standard eror increases as one moves
further from the mean! In Section 3, the standard error associated
with an ability level of 0.38 is 0.727, which means that we can be
95Vo confident that a score of 0.38 is 0.38 + I.43, somewhere between
-1.05 to 1.81, or in raw score terms, anywhere from 3 to 6. However,
in practice, we can see that the potential for error is much higher in
Section 3 (Table 14) than in any other subtest, as it fails to diicrimi-
nate between sfudents as a result of its length.
It is only when these calculations are made that the reliability coef-
ficients take on meaning. It highlights the fact that decision-making
on cut scores involves elTor, and that sensitive interpretation of scores
on all tasks is required before referring or not referring a student to an
English language support programme. However, this evidence lends
support to the current practice of interviewing all referred students,
to ensure that some students do not needlessly attend the in-sessional
programme if it is not necessary. It also highlights the importance of
Glenn Fulcher 129
3 Concurrent validiry
Concurrent validity was investigated by considering the association
of scores on the placement test, with scores on the TOEFL. A total
of 33 students had taken the TOEFL test immediately prior to arrival
at the university, and this proximity allows some degree of compari-
son between the results of the two tests. Table 15 gives the correlation
coefficients between the TOEFL scores of the students, the total score
on the placement test and each of the components of the placement
test.
From this strength of association between the total placement test
score and TOEFL, the best prediction of the placement test score is:
The new estimate for the lowest TOEFL score required to follow a
course of study without English language support would be 499,
which appears to be inappropriate on experiential grounds. In con-
clusion to this section, only moderate association was discovered
between the placement test and the TOEFL. This is most likely asso-
ciated with the small sample size, but may also be a factor of the
difference in content and pu{pose between the two tests. Nevertheless,
it is hoped that concurrent validity may be further investigated as
additional data are gathered over a number of years, producing a
much larger sample of students from whom estimates may be made.
4 Content validiry
Content validity was investigated by asking three informants, one
each from the Faculties of Human Studies, Engineering and Physics,
to 'eyeball' test items. In'only one case did an informant feel parti-
cularly strongly about a test item, which is given here in full:
Text
Hints of a revolution in superconductivity have been found by French
researchers who claim to have discovered a substance that loses all electrical
resistance at only 20 degrees Celsius below zero. Superconductors today need
to be cooled to -135"C. Superconductors are materials whose resistance to an
electrical current disappears completely. Superconducting magnets could make
trains levitate and drive ships, and superconducting cables could produce a
much more efficient national grid network. Superconductivity was first noted
in metals cooled to within a few degrees of absolute zeto (-273"C). Then in
1986 ceramics that could superconduct at much higher temperatures caused a
storrn in the scientific community, holding out the hope that other materials
would superconduct at room temperatures. Until now, the best efforts have
----succeeded-only-in producing a material that superconducts-rvhen [Link]-_
liquid nitrogen - which is commercially useful.
Question
Research into superconductivity is being undertaken because
(a) the use of liquid nitrogen is not practical in'nonrommercial applications
(b) the railway systems of the world wish to levitate their trains
(c) ceramics is a current topic of discussion within the French scientific com-
munity
(d) superconductors only operate at low temperatures
From the pilot study, the item statistics given in Table 16 were
obtained, ild the item included as a result. However, the expert judge
from the Faculty of Physics argued that this item would be unfair to
anyone taking the test with scientific training. The reason for under-
taking research into superconductivity, he argued, was a combination
of (a), (b) and (c). Proposition (d), which he recognized as the
required answer, is true, but without the fact that liquid nitrogen is
impracticd (a), it would not matter that (d) was true. Further, if prop-
osition (b) about the industrial applications were not important, then
Glenn Fulcher 131
(O O) F- Ol O,
rrru?e?
oooo(>
ttrtt
a, rC)Al(DO
o
o rqqqq
ooooo
qt
o
o
C)t-(IrgrO
6 ol--ol9
E ooooo
o
= c
.9
=ct ?
FE oo@to@
olr9u?9
i6 ooooo
.G'
bE
r(\l(I)rtO
=E
6
o
o
.cr
c
6
o-
co
6
c
.E
'Ex
o E€
[Link]
(t
E
ao
6
E
ctl
L
o E
E .9
o E
o
o. lo
o !o
(L o
lt'
o
.E
th
.F
(6 E
o
.=
.tt
lj
E
o
.!P
G' o
o -c
a o
(o il
o
ct
c o
a;
o
I
E
o (o o
F a
132 An English language placement test
No response Very unfair Unfair Fair Quite fair Very lair Total
10 71
Of the 14 respondents who thought that the test lvas unfair, five
provided no reason for their view. The responses for the other nine
students were as follows:
It is clear from these responses that dissatisfaction was, for the most
part, not a function of the test itself, but of the restrictions which are
imposed on the testing process by the time and facilities available for
134 An English language placement test
l,lo response 1 0 1
Very unfair 3 (75%l 1 4
Unfair 2 (20Phl 8 10
Fair 6 (17Vol I 35
Quite fair 6 (357.) 11 17
Very fair 2 (W"l 2 4
Total n egv"l 51 71
Glenn Fulcher 135
VII Discussion
I Ingistic and administrational constraints
ry large institution there will be-logistieandadffirstratio-n-al -
-Wftlrin
constraints. What is not often recognized, however, is that these con-
straints lead to limitations on testing, which have a direct impact on
reliability of score interpretation. It is therefore important that insti-
tutions make policy decisions which take this relationship into
account. It should never be assumed that decisions taken by those in
administration do not have academic impact. In the case of the Uni-
versity of Surrey, I practical compromise has been reached in the
form of a two-tier testing system, with the less precise placement test
screening out students from the oral interview stage of the process.
This has clearly had a negative impact upon the usefulness of the
reading subtest, which needs to be lengthened.
Vm Conclusion
We are aware that there are other approaches to developing placement
tests, many of which are effective where the numbers of students
__taking the tr$_qlg small 4nd the constraints in local _administration
allow the use of lengthy instruments which include oral components
(Paltridge, 1992). Each approach must be sensitive to the constraints
imposed by, and the information requirernents of, unique institutions.
Similarly, when conducting score gain studies in the future, criterion-
based approaches such as that suggested by Brown (1989) will be
considered, especially if multiple forms of the current test are avail-
able. We nevertheless believe that the current test fulfils its purpose
as well as can be expected within the context described in this article.
The methodology used in developing the evaluation of this place-
ment test was based on Wall, Clapham and Alderson (1994), as the
first article in the field of language testing to address the issue of
the assessment of placement instruments in any depth. However, the
methodology used in this study includes additional aspects which
would appear to be worthy of investigation in the particular context
of placement testing. It is hoped that the approach adopted, building
on Wall, Clapham and Alderson's pioneering work, may be of use
to others in the evaluation of their placement tests.
138 An English language plncement test
LK Rderences
z\t-62.
Pophanr, WJ. l9?8: Criterion-referenced mcasuremznt. Englewood Cliffs,
NJ: hentice-Hall.
Slchrdq C.C. and DelMes, RC. 1990: Determining the validity of place.
ment exams for developmental college anicula. Applied Measurement
in Educarton 4, 37-52.
Shohamy, E. 1983: The stability of oral proficiency ass€ssment in the oral
interview procedure. Language lzaraing 33, 527 -40.
1988: A proposed framework for testing the oral language of second
foreign language leatms. Studies in Second language Acquisition lO,
- t65-79.
1990: Ianguage testing priorities: a different perspecive. Foreign Lan-
guage Awals 23, t85-94.
-Shohamy, E., Reves, T. and Bejarano, Y, 1986: Intoducing a new comprc-
hensive test of oral proficienry. English Language Teaching Journal
40,2t2-20.
Glenn Fulcher 139