0% found this document useful (0 votes)
12 views27 pages

Placement Testing

This document describes a study examining the reliability and validity of an English language placement test used at a university. It outlines the test format, development process including piloting, and methodology used to evaluate reliability and validity through statistical analysis and rater judgments. The goal is to accurately place students in language support programs to help them succeed academically.

Uploaded by

mathpix2525
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
12 views27 pages

Placement Testing

This document describes a study examining the reliability and validity of an English language placement test used at a university. It outlines the test format, development process including piloting, and methodology used to evaluate reliability and validity through statistical analysis and rater judgments. The goal is to accurately place students in language support programs to help them succeed academically.

Uploaded by

mathpix2525
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

An English language placement test:

issues in reliability and validity


Glenn Fulcher English Language lnstitute, University of
Surrey

This report describes a reliability and validity study of the placement test which
is used at the University of Surrey as a means of identifying students who may
require English language support whilst studying at undergraduate or postgraduate
levels. The English Language Institute is charged with testing all incoming stu-
dents, irrespective of their primary language or subject specialization. F[ir and
accurate assessment of student abilities, and refening individuals to appropriate
language support courses (the in-ressional programme), is an essential support
service to all academic departments. The goal of placement testing is to reduce
to an absolute minimum the number of students who may face problems or even
fail their academic degrees because of poor language ability or study skills. This
study looks at the administrative and logistic constraints upon what can be done,
and assesses the usefulness of the placement test developed within this context.

I Introduction
It has recently been noted that although placement testing is probably
one of the most widespread uses of tests within institutions, there is
relatively little research literature relating to the reliability and val-
idity-of such measures (Wall, Claptram and A-lderson, 1994):'Publi-
cations which deal with placement tests frequently provide qualitative
assessments of instruments (Goodbody, 1993), or are concerned with
the placement of linguistic minority students in programmes which
are not related to language teaching (Schmitz and DelMas, 1990; Tru-
man, 1992). Wall, Clapham and Alderson ( 1994) offer one of the
few empirical studies of a placement test which is designed to screen
students entering a British university for deficiencies in English lang-
uage skills, which might impede their progress in undergraduate or
postgraduate studies. They investigate face validity (student percep-
tions of the test), content validity (asking whether tutors thought that
test content represented programme content), construct validity
(through a correlational study), concurrent validity (with self-assess-
ment, and the assessment of tutors) and reliability. The main problem
they discovered in attempting to validate the University'of Lancaster
placement test was in finding appropriate external criteria to conduct
a concurrent study.

Language Testing 1997 14 (2) 113-138 0265-5322(97)LT127OA @ 1997 Arnold


ll4 An English language placement test

The methodology used to ascertain the reliability and validity of


the test in this study is similar to that of the Lancaster study in many
ways, and is described in detail below. However, this study does
attempt to exand on their methodology for the evaluation of place-
ment tests within a university context in four areas. These are:
. the use of pooled judgements in establishing cut scores for place-
ment;
o the use of additional statistical means to analyse test data;
o the consideration of the need to develop parallel forms; and
o the use of student questionnaires to investigate face validity.

II Background
The English Language Institute (ELI) has a number of roles within
the University of Surrey, ranging from its support function in provid- l
ing English language and study skills courses for students at the uni-
versity, English for Academic Purposes (EAP) for overseas students
who intend to study in an English-medium tertiary institution, and an
MA in linguistics (TESOL) for teachers of English, in distance-learn-
ing mode. As part of its service function related to the in-sessional
programme, the ELI is charged by the university to administer an
English language test to all students entering the university on taught
courses (undergraduate and postgraduate, English as primary langu-
age speakers, and speakers of other languages). The purpose of the
test is to identify students whose lack of language skills or ability
to communicate may cause problems in their academic work with
their departments.
----A-major-eonstraint in assessing studems is -tha[ togeth-erwith-
organization (handing out and collecting papers, giving verbal
instructions), the ELI must complete testing within one hour for each
test administration. This means that the test cannot exceed 45 minutes,
and must therefore be restricted in its length and content. Secondly,
all scripts must be processed within five days of the day of the test,
and results disseminated. In 1994, 1619 students took the test. The
number of markers available, and the number of scripts which each
can process in this time frame, also dictates that much of the paper
be objectively scored. However, no language test can use only one
response format, or sample only one construct, if it is to be considered
valid on even prima-facie grounds (APA, 1985: 73; 75). For this
reason, a writing component is included.
The placement test acts as a screening device to reduce the number
of students who attend an oral interview. It is those who are invited
to come to the ELI for an oral interview who have the greater prob-
ability of requiring English language support. These students are
Glenn Fulcher 115

'referred' for oral testing. A small number of students are informed


after the interview that they do not need to attend the in-sessional
courses. It is acknowledged (see section VI below) that all tests con-
tain error, and it is during the interview procedure that we attempt to
identify students who have been misclassified by the test. Most stu-
dents are referred to a programme, although attendance is optional.
Lists of all referred students are sent to heads of department, and later
in the year heads are also sent an update of student attendance at
courses, and a qualitative assessment of individual student progress
for those who did attend.

III Test design


I Test format
Prior to October lgg4: the test contained one piece of writing, and
was marked subjectively. It became clear that the ELI did not know
whether this was a reliable or valid placement instrument, but it was
considered inappropriate on other grounds. A single essay title may
be biased in favour of some students and against others, and any
single task of this nature is unlikely to elicit an adequate sample upon
which decisions can be made (APA, t985: 75; Upshur, l97l: 47;
Van Weeren, 1981: 57; Shohamy, Reves and Bejerano, 1986; Sho-
hamy, 1983; 1988; 1990). A new test was devised, with the follow-
ing format:

Section I:Essay 1: descriptive (no choice of essay title).


---------Essay -2: argumentative (one title to-te seleeted-from--
three options).
Section 2: Stnrcture of English (10 items).
Section 3: Reading comprehension (8 items).

Essay titles for Section I were screened carefully to avoid bias in


favour of, or against, students from any particular discipline, or cul-
ture. Essay I is of a general nature, whilst ttre titles for Essay 2 arc
an attempt to reflect the interests of the three faculties within the
university: Human Studies, Engineering and Science. In Section 2,
item content was selected for known difficulty in the in-sessional pro-
gramme grammar revision course. The context of the items is general
university life, on the grounds that subject-specific contexts may, once
again, disadvantage some subsections of the test-taking population.
In Section 3, six texts were selected, for their general academic inter-
est. Of these, three were drawn from the humanities, and three from
the sciences. However, the journals from which the texts were taken
116 An English language placement test

are 'popular' in that the passages are written for the interested
(intelligent) layperson, not for the specialist.

2 Pilot study
The new format was piloted during the summer of 1994, using 67
students attending the ELI Summer School English for Academic Pur-
poses (EAP) programme. All items in Sections 2 and 3 of the test
were studied using classical item analysis. Only items with a facility
index between 0.3 and 0.8 were retained, and items with a discrimi-
nation index of less than 0.3 were either abandoned, or rewritten and
piloted for a second time. The point biserial correlation for all remain-
ing items was above 0.4. Four pilot versions of the test were originally
wriffen, and 75Vo of all questions discarded, leaving one operational
test with items drawn from all four pilot tests. This is a reminder that,
no matter how experienced one may be in test development, there is
always a need for pretesting all items before tests become operational,
and decisions are taken on the basis of test results.
In Section l, writing samples were collected, graded by six tutors,
and features of performance at each level of a rating scale established.
Prototypical samples from each band level were then used in the train-
of raters prior to the operationalization of the test in October 1994.
ing

IV Method
The final version of the test was used operationally with the univer-
sity's entire intake on taught (undergraduate or postgraduate) courses
in October 1994. The- main-studies were all carried out using thls
population.

I Reliability
In order to establish estimates of the reliability and validity of the
. placement test, a number of approaches were taken. In the investi-
gation of reliability, correlation coefficients, means and standard devi-
ations (inter- and intrarater reliability), were established for rating
patterns on Section I of the test. For Sections 2 and 3, it was decided
to fit a logistic model. The reason for this decision was essentially
because of the need, within a university setting, to have multiple
forms of a test, and for the interpretation of scores on these forms to
be comparable.
This implies that forms must be equated in some way. Using a
logistic model, it is possible to create multiple parallel forms, using
'anchor' items from previous form(s) which can be used to calibrate
Glenn Fulcher ll7
new items in later forms, even if this is difficult for short tests. In
the first attempt to fit a Rasch model to the data, it became clear
that there were two distinct populations within the total test-taking
population. These distinct populations are those whose primary langu-
age is English, and those whose primary language is other than
English. Consider Table 1, which is a simple )( test of significance
of primary language classification and referrals, on the basis of the
test scores (for a discussion of cut scores for referrals, see Section
VI below). Using Yates' correction for a 2x2 gnd, )(=516.61, a
highly significant result, indicating that these are certainly separate
populations. [t is not possible, unfortunately, to isolate differences
between referrals or test scores between learners whose primary lang-
uage is not English, because of the small n size of many of the pri-
mary languages represented by the 614 overseas students tested.
The question which arose, therefore, was on which population to
standardize the objective components of the test. It was finally
decided to standardize on the group which did not have English as a
primary language, as this is the population which is more likely to
require English language support. Those speakers of English as a pri-
mary language (73 in this sample) whose score profiles are similar
to those of the referrals of the norming population will still be ident-
ified by the test, and remedial action can be taken.
The Rasch model allows the test developer to calibrate both item
difficulty and learner ability to the same scale (Crocker and Algina,
1986: 340-41), measured in logits. This was done using the program
RASCAL (Assessment Systems Corporation, 1994) for Sections 2
and 3 of the test separately. This was done as there was no theoretical
reason to stl-spect ttratrhscorfihined-scores of Section 2 and 3 would
represent a unidimensional scale. The disadvantage of this approach,
however, is that Section 2 consists of only ten items, and Section 3
of eight items. Although n = 614, the low number of items inevitably
reduces test reliability. This could not be avoided, because of the time
constraints as described above.

Tabfe 1 X table to compare primary and nonprimary English language speakers to test
whether they belong to the same test-taking populaUon

Not referred Relened

Non-English primary language 252 362 614


English primary language 902 73 1005
Toal 1184 435 1619
I 18 An English language placement test
2 Validiry
a Correlation and principal components analysis.' Construct val-
idity was assessed using correlation, and a principle components
analysis. Each of the sections of the test were designed to measure
different aspects of English language proficiency, and so the factorial
structure of the test is an issue.

b Analysis of cut scores across referred and nonreferred stu-


dents: Cut scores for the test as a whole, and each section of the
test, were established after the pilot study. Scores were considered in
relation to lecturers' assessments of whether students with certain pro-
files were ready for study in a tertiary-medium institution. Cut scores
were established using this 'pooled judgements' technique (Popham,
1978: 165). Groups of essays at a range of scores were presented
randomly to lecturers who were asked to decide whether this was an
acceptable piece of work for undergraduate or postgraduate work in
the university. Discussion was acceptable, and the cut point for each
essay established at the point where the highest agreement in making
dichotomous judgements was reached. However, the results require
empirical investigation to ensure that the cut scores are in fact leading
to appropriate decisions on whether to refer students to language sup
port prognammes.

c Concurrent validity: Thirty-three students from the overseas


student population taking the test had recently taken the Test of
English as a Foreign I-anguage (TOEFL). Although this number is
small, it is nevertheless possible to begin to investigate the relation-
ship between-this-test and-auniversity placement test. As more data -
are gathered in future years, it should be possible to establish stable
concuffent validity statistics with a number of major tests which are
currently used for entrance purposes.

d Content validity: Content validity was investigated by


requesting three subject specialists to comment on questions set in
the placement test. One was drawn from each of the three faculties
within the university: Human Studies, Engineering and Physics. Only
one informant suggested that the test was not content valid in one
particular field.

e Feedback from students: It is clearly important to obtain quali-


tative feedback from students on the opemtion of the test. All tests
have consequences for the test-takers and for the institutions which
base decisions on scores. [f the test is not perceived to be fair by the
test-takers and score users, the role of the placement test within the
Glewt Fulcher I 19
Table 2 lnter-rater reliability - Pearson product mornenl correlations

Rater 1 Rater 2 Rater 3

Rater 2 0.92
Rater 3 0.87 0.94
Rater 4 0.75 0.83 0.93

institution is compromised. Including a qualitative study of this nature


therefore relates not to the technical qualities of a testing instrument,
or to accurate decision-making, but to one aspect of the social conse-
quences of testing for the institution.

1V Retiability o

Reliability of subtests was initially calculated during the pilot study,


and recalculated during the first operational testing. The following
figures relate to the operational version of the test, with n=614.

I Section I: writing
To establish the reliability of the assessment of writing samples, 20
essays were selected from the population, and were marked by four
tutors. After a period of two weeks, the tutors were then asked to re-
mark a subset of six samples. This allows the calculation of inter-
and intrarater reliability. Table 2 shows the results of the inter-rater
reliability study. It can bg_seen 1[4gelgglqgplbetwg-e1 raters is well
within acceptable reliability ranges for this type of test, with an aver-
age Pearson product moment correlation of 0.87. The average band
awarded was 5.46, with a standard deviation of l.2L Table 3 shows
that the four raters did not differ significantly from this in their indi-
vidual grade profiles.
In the intrarater reliability study, the average reliability coefficient
was 0.69, somewhat lower than the inter-rater coefficients, but still
not so low as to cause undue worry. This figure indicated a need for
further rater training before the second operational testing session, in

Table 3 lnter-rater reliability - Variation in means and standard deviations

Rater 1 Rater 2 Rater3 Rater4 Average

Mean 5.50 5.33 5.50 5.50 5./16


sD 1.00 1.O3 1.64 1.64 1.21
nA An English language placement test

October 1995. The intrarater reliability coefficients were: rater l:


0.83; rater 2: A.57; rater 3: 0.68; and rater 4: A.7O (Pearson product
moment correlations).

2 Section 2: structure
Rasch scaling was conducted using RASCAL (Assessment Systems
Corporation, 1994). Table 4 shows the questions in order of difficulty,
from the easiest to the most difficult, in logits, together with the stan-
dard error associated with the difficulty estimate, and the fit )(
statistic.
From Table 4 it can be seen that there are three misfitting items.
That is, they do not meet the criterion of unidimensionality in this
subtest, and must therefore be removed, and replaced by other items
in rhis form of the test. It is instructive to return to misfitting iterns
to attempt to provide a linguistic rationale for why they misfit. In this
case, it is particularly enlightening to look at item 5, because of the
very large misfit statistic. Item 5 is:

She's always other students on her course, even though she


is very busy herself.
a. help
b. helped
c. used to help
d. helping

,With hindsight,it appears ob-vious that-there-are two pssible keys


to this item, but this was not spotted during pretesting and item
revision. Such mistakes highlight the importance for thorough pretest-
ing and post hoc analysis of all test items, on tests where important

Table 4 Rasch analysis of Section 2

Difficulty Standard enor

7 -1.370 0.129 7.5U


2 -1.085 0.120 14.279
1 4.770 0.111 5.137
6 -{.668 0.109 9.298
8 -,0.499 0.105 14.466 misflt
10 -0.@5 0.099 ?2.ffi5 mlsfit
g 0.128 0.096 9.597
I 0.598 0.093 8.055
4 0.745 0.093 8.0s5
5 3.006 0.124 45.036 misfit
Glenn Fulcher l2l
decisions are being made. It cannot be emphasized enough that how-
ever experienced an item writer or test designer someone may be, it
is not enough to rely on 'eyeballing' tests or test items.
With the test centred on item difficulty, mean item difficulty was
0.00, and the standard deviation of difficulty 1.26. Average ability
was 0.89, with a standard deviation of 1.15. The test characteristic
curve for Section 2 is presented in Figure l. This plots estimated
proportion of test items correct as a function of the ability of the
student on the latent trait (structure of English), and may be inter-
preted as a nonlinear regression curve for relating raw scores to the
latent trait.
When using a Rasch model, it is possible to look at reliability in
rwo ways. First, we may calculate a reliability coefficient which is
the equivalent of KR20 (an estimate of internal consistency) used in
classical test alralysis and, secondly, we may talk about test infor- *
mation. The reliability coefficient may be understood as the degree
to which the test characteristic curve in Figure I may be relied upon
as a translation of raw scores to latent trait scores. This section of
the test contains only l0 items, a decision taken because of time con-
straints as discussed above. As reliability is related to test length,
although it was hoped that reliability coefficients would be adequate,
they were not expected to be very high. The reliability coefficient for
Section 2 was 0.63, which is below what would be required for a
high-stakes test. However, this may be reasonable for a placement
test of this size.
1.00

Eac
696
uJo

-3.0 -2.O -1.0 0.0 1.0 2.o 3.0

Ability
Figure I Test characteristic curve for Section 2 (structure)
122 An English language placement test

The amount of information which a test provides diffen according


to region on the latent trait scale. Where the test characteristic curve
is steepest, there is more discrimination amongst test-takers, and
hence the test is providing greater information. Test information for
Section 2 of this test is presented in Figure 2. Figure 2 indicates that
Section 2 of the test is providing the most information from -1.0 to
0.5 on the latent trait scale, which is precisely what was required of
this placement test. It is not necessary to discriminate finely between
students at the higher end of the latent continuum. Similarly, below
a certain ability level, fine discrimination is not necessary. What is
crucial in this kind of testing is providing the maximum amount of
information around the region of the scale where decisions are being
made regarding whether students below this score should attend
English support programmes. This result is therefore welcomed, as
the test does appedr to be providing the most reliable information at
precisely the point at which it is required.
We may conclude, with some certainty, that for a subtest with only
ten items, Section 2 is providing adequate information for the pur-
poses to which the test is being put.

3 3: reading comprehension
Section
Table 5 shows the results of the Rasch analysis for Section 3. Only
item 17 was found to misfit, and this item was therefore removed

c
o
l'
E
L
€g

-3.0 -2.0 -1.0 0.0 1.0 2.0 3.0

Ability
Figure 2 Test information curve for Section 2 (structurel
Glenn Fulcher 123

Table 5 Rasch analysis of Section 3

Difficulty Standard error

11 -2.252 o.127 10.190


14 -1.501 0.108 10.459
12 -o.730 0.097 4.660
16 4.260 0.094 3.983
17 0.016 0.093 14.989 mlslll
18 1.436 0.106 5.032
13 1.569 0.109 9.180
15 1.723 0.112 8.240

from this form of the test. With the test centred on item difficulty,
mean item difficulty was 0.00, and the standard deviation of difficulty
1.48. Average ability was {.01, with a standard deviation of 1.22.
The test characteristic curve for Section 3 is presented in Figure 3.
The curve in Figure 3 is not as steep as that in Figure 2, but this is
only to be expected, with only eight questions in the subtest. It is
difficult to achieve reasonable discrimination with so few items in a
subtest. Similarly, the reliability coefficient (equivalent of KR20) is
only 0.59. The test information curve in Figure 4 is much flatter than
the curve in Figure [Link] an ideal world, the test should be lengthened,
as the discussion of this subtest in section VI shows.
The test has been found to be reasonably reliable for its purpose.
However, the reliability coefficients do not meet the levels which

E'E
coF e [Link]

696
ulo

-3.0 -2.O -1.0 1.0 2.O

Figure 3 Test characteristic curye for Section 3 {readingl


124 An English language placement test

g
o
(!
E
o
c

-2.0 -1.0 0.0 1.0

Ability
Figure4 Test characteristic curve for Section 3 (reading)

would be required for a high-stakes test, and this appears tobea


function of test length. The relationship between reliability and test
length may be described as:

n- ry4:-g
rtt(l rttd)
-
where n is the number of times the test length must be increased with
items or raters to achieve the desired reliability, rttd is the desired
level of reliability and rtt is the present observed reliability of the
it would
test. For Section 2, to achieve a reliability coefficient of 0.80,
be necessary to increase test length by 2.34 (28 questions), and for
Section 3, 2.77 (22 questions). This would essentially double the
administration time of the test!
Compromises need to be made between desired levels of test
reliability and the lack of time for testing within a large educational
organization. Such decisions cannot be taken by the test designers
alone, but need to be discussed at all levels within the institution
in relation to general testing policy, including the assessment of the
consequences of the unreliability which the institution is prepared
to tolerate
Glenn Fulcher 125

VI Validity
1 Correlational and principal components analysis
The test was designed to measure three separate abilities: writing,
English structure and reading. If this is actually the case, the corre-
lation coefficients between the three sections should be modest, and
any attempt to factor the matrix should show that subtests load on
different factors. Table 6 provides the correlation matrix in the bottom
triangle, and the reliability coefficients of the subtests in the diagonal.
The upper-right triangle presents the correlation coefficients corrected
for attenuation (taking unreliability of measures into account).
Although all are significant (to be expected, in subtests which are all
language related), variance overlap is not so high that we would wish
to challenge the view that each subtest is tapping some unique ability
as well as a common ability. It should also be noted that each of
the two writing tasks also appears to be tapping different skills to
some degree.
The correlation mahix was subjected to principal components
analysis (PCA), in order to discover if the tasks were loading on
different factors. Eigen values for the extraction of factors were set
at 0.6, on the assumption that some of the factors in which we are
interested are indeed, as Farhady (1983: 18-19) argues, less than
unity. Three factors were extracted, ild rocated using the Varimax
technique. The results are presented in Table 7. The three-factor sol-
ution is clearly interpretable in relation to the correlation matrix. Fac-
tor I is a writing factor, factor 2 is a reading factor and factor 3 is a
structure factor, which the argumentative essay also loads on. How-
ever, it must be Ctrbssed ih-at-thii-is nof iuffibiant evidence upon which
to claim construct validity. Exploratory analysis of this type is sugges-
tive of further avenues for investigating construct validity, but can
never be sufficient to meet the requirement for convergent and diver-
gent evidence.

Table 6 Correlation matrix showing assochtion between tasks

Writing task 1 Writing task 2 Structure Reading

Writing task 1 0.87 0.49 0.36 o.44


Writing task 2 0.43 0.87 0.61 o.42
Structure o.27 0.45 0.6:t o.44
Reading o.17 0.30 0.27 0.59
126 An English language placemcnt test

Table 7 Varimax-rotated factor solution

Factor 1 Factor 2 Factor 3

Writing 1 [Link]) 0.049 0.099


Wdting 2 0552 0.210 0.579
Structure 0.087 o.110 0.942
Reading 0.087 o.9&l 0.145
Per cent ol total variance 30.572 25.615 31.363
etqlained by rotated
@mponents

2 Analysis of cut scores across referred and nonreferred students


Cut scores for referrals were decided on the basis of evidence from
the pilot study, and implemented in an operationalruse of the test. It
was therefore imperative that the judgements made be considered in
the light of real outcomes. Tables 8 and 9 give the descriptive stat-
istics for nonreferrals and referrals respectively, by task and total test
score. The first thing to ensure is that the cut scores established in
the pilot study did in fact divide the test-taking population into two

Table 8 Descdptive statistics for nonrefenals by task and total

Writing 1 Writing 2 Structure-L Reading-L Total (raw)

cases 1184
No. 1184 1029 1 159 1 184
Minimum 0 _l) 4.64__ -2.77 0
Maximum I I 2.95 2.74 33
Range 9 I 3.59 5.51 3i!
Mean 6.92 6.85 1.85 0.63 26.57
sD 2.14 1.01 0.88 1.10 3.25
Standard enor0.06 0.029 0.03 o.03 0.10

Table 9 Descriptive statistics lor referrals by task and total

Writing 1 Writing 2 Struc{ure-L Reading-L Total (raw)

No. cases 435 435 408 416 43s


Minimum O 0 -2.67 -2.77 2
Maximum 7 I 2.95 2.74 29
Range 7 I 5.62 5.51 27
Mean 5.31 4.92 0.88 -0.35 20.24
sD 1.24 1.66 1.21 1.13 4.13
Standard enorO.06 0.08 0.06 0.06 0.20
Glenn Fulcher 127
Table 1O Probability of real ditferences on subtests between referals and nonreferrals
with 1 df (Bartlett tesl lor homogeneity ol group variances)

Sublest Writing 1 Writing 2 Structure-L Reading-L Total (raw)

Xp)e pX p)ep )ep


15.4.8 0.0@ 178.6 0.000 67.38 0.000 0.52 0.471 38.73 0.000

Table 11 Flefenals by scores on Writing 1

x01234567890 T
y 10 2 2 13 36 1U 196 32 0 0 0 435
n00004302417091692110 1184
r 10 2 2 13 /t{t 174 437 741 169 21 10 1619

statistically distinct groups, and this is shown to be the case in Table


l0 for Sections I and 2 of the test, and the total test score.
Section 3, the reading subtest, fails to discriminate adequately
between referrals and nonreferrals. Although Figure 4 shows that
information is being provided at the appropriate point on the scale,
this is not enough to make decisions. It is clear that in the case of
this subtest, the use of only eight items seriously affects the validity
of the subtest for its intended purpose.
Showing that there is a statistical difference between referrals and
nonreferrals in Sections I and 2 of the test does not, in itself, tell us
that the cut score has been adequately placed. In the following four
tables,(Tables I I - t 4) are the-numbers of-students-sho-are-referred
(y), the number not referred (n) by raw score (x), and the total num-
ber of students with each score, for the two writing tasks, stntcture

Table 12 Refenals by scores on Writing 2

x01234567890 T
y3206847 151 168A2 100 435
n00021232556941781912 1184
T 32 0 6 10 48 174 423 716 179 19 12 1619

Table 13 Refenals by scores on Structure

x0 1 23456789 10 o T
yI 2 12 16 29 39 Tf 98 80 56 17 0 435
n0 0 0 0 1 10 74 230 375 338 146 10 1184
T9 2 12 16 30 49 151 3?8 455 394 163 10 1619
128 An English language placement test

Table 14
Refenals by scores on Reading

x012345678OT
y162369109111762542043s
n O 10 40 151 274 361 239 92 15 11 1184
T 16 33 109 260 385 437 255 96 17 11 1619

and reading, respectively. Students who did not answer a question are
counted as 'omitting' (O). Score ranges where numbers are high-
lighted in bold represent the areas in which there is potential for
errprs of judgeme-nt to be made yhq [Link] referrals.
Ttre cut score for tasks I and 2 (Section l) was 6, for Section 2,
a raw score of 6 (or ability level 0.38) and Section 3, a raw score of
4 (or ability level 0.02). The overall cut score was 22. Raters wetb
asked to consider the evidence from each individual section in making
a decision regarding referral. For example, if the overall score was
22+, and on only one of the tasks did the score fall below the cut
point, then they were probably not to refer the student. This accounts
for some of the nonreferrals in band 5 on both the writing tasks, a
score of 5 on Section 2 and a score of 2 or 3 on Section 3, as well
as those who were referred within a number of bands above the cut
score.
In the case of Sections 2 and 3, however, w€ are able to calculate
the standard error of the cut score, because of the methodology which
has been used. The standard error associated with an ability level of
[Link] is 0.852, meaning that we can be 95Vo confident that a scaled
- cut score of 0.02 is 0.02+ 1.67, or anywhere ffim -J35To l.S9. That
is a raw score from 3 to 7 . The standard eror increases as one moves
further from the mean! In Section 3, the standard error associated
with an ability level of 0.38 is 0.727, which means that we can be
95Vo confident that a score of 0.38 is 0.38 + I.43, somewhere between
-1.05 to 1.81, or in raw score terms, anywhere from 3 to 6. However,
in practice, we can see that the potential for error is much higher in
Section 3 (Table 14) than in any other subtest, as it fails to diicrimi-
nate between sfudents as a result of its length.
It is only when these calculations are made that the reliability coef-
ficients take on meaning. It highlights the fact that decision-making
on cut scores involves elTor, and that sensitive interpretation of scores
on all tasks is required before referring or not referring a student to an
English language support programme. However, this evidence lends
support to the current practice of interviewing all referred students,
to ensure that some students do not needlessly attend the in-sessional
programme if it is not necessary. It also highlights the importance of
Glenn Fulcher 129

increasing test length when possible, within the administrative and


logistic constraints of large organizations.
One further point does need to be made. There is as yet no way
of discovering whether nonreferrals have mistakenly been so categor-
ized, unless they self-refer, or are referred by their subject tutors.
These numbers should, however, be low, as the policy of the ELI is:
if in doubt, refer. It is easier to correct such an error at a later stage.

3 Concurrent validiry
Concurrent validity was investigated by considering the association
of scores on the placement test, with scores on the TOEFL. A total
of 33 students had taken the TOEFL test immediately prior to arrival
at the university, and this proximity allows some degree of compari-
son between the results of the two tests. Table 15 gives the correlation
coefficients between the TOEFL scores of the students, the total score
on the placement test and each of the components of the placement
test.
From this strength of association between the total placement test
score and TOEFL, the best prediction of the placement test score is:

Placement test - -34.73 + 0.1 (TOEFL)


Using the cut score established for the placement test, an estimated
TOEFL score of 555 would be required before a student could follow
a university course without the need for additional English language
support. However, a correlation of 0.64 is clearly not high enough to
--E able-t6takC th-eie kinds of decisions without the use of-the paCF
ment test itself. This was confirmed by an examination of the scatter
plot of the scores of the 33 students, which revealed two outliers with
scores of 5 l0 and 577 on TOEFL, and placement test scores of 12
and 9 respectively. Once these students are removed, the TOEFL-
TOTAL correlation is only 0.56, providing the best prediction at:

Placement test = -5.4 + 0.6 (TOEFL)

Table 15 Conelation between TOEFL and the university placement test

Placement test Section1 Section2 3


Section Section 4
(writing) (writing) (structure) (reading)

0.34 (n/s) 0.64


130 An English language placement test

The new estimate for the lowest TOEFL score required to follow a
course of study without English language support would be 499,
which appears to be inappropriate on experiential grounds. In con-
clusion to this section, only moderate association was discovered
between the placement test and the TOEFL. This is most likely asso-
ciated with the small sample size, but may also be a factor of the
difference in content and pu{pose between the two tests. Nevertheless,
it is hoped that concurrent validity may be further investigated as
additional data are gathered over a number of years, producing a
much larger sample of students from whom estimates may be made.

4 Content validiry
Content validity was investigated by asking three informants, one
each from the Faculties of Human Studies, Engineering and Physics,
to 'eyeball' test items. In'only one case did an informant feel parti-
cularly strongly about a test item, which is given here in full:
Text
Hints of a revolution in superconductivity have been found by French
researchers who claim to have discovered a substance that loses all electrical
resistance at only 20 degrees Celsius below zero. Superconductors today need
to be cooled to -135"C. Superconductors are materials whose resistance to an
electrical current disappears completely. Superconducting magnets could make
trains levitate and drive ships, and superconducting cables could produce a
much more efficient national grid network. Superconductivity was first noted
in metals cooled to within a few degrees of absolute zeto (-273"C). Then in
1986 ceramics that could superconduct at much higher temperatures caused a
storrn in the scientific community, holding out the hope that other materials
would superconduct at room temperatures. Until now, the best efforts have
----succeeded-only-in producing a material that superconducts-rvhen [Link]-_
liquid nitrogen - which is commercially useful.
Question
Research into superconductivity is being undertaken because
(a) the use of liquid nitrogen is not practical in'nonrommercial applications
(b) the railway systems of the world wish to levitate their trains
(c) ceramics is a current topic of discussion within the French scientific com-
munity
(d) superconductors only operate at low temperatures
From the pilot study, the item statistics given in Table 16 were
obtained, ild the item included as a result. However, the expert judge
from the Faculty of Physics argued that this item would be unfair to
anyone taking the test with scientific training. The reason for under-
taking research into superconductivity, he argued, was a combination
of (a), (b) and (c). Proposition (d), which he recognized as the
required answer, is true, but without the fact that liquid nitrogen is
impracticd (a), it would not matter that (d) was true. Further, if prop-
osition (b) about the industrial applications were not important, then
Glenn Fulcher 131

(O O) F- Ol O,
rrru?e?
oooo(>
ttrtt

a, rC)Al(DO
o
o rqqqq
ooooo
qt
o
o
C)t-(IrgrO
6 ol--ol9
E ooooo
o
= c
.9
=ct ?
FE oo@to@
olr9u?9
i6 ooooo

.G'
bE
r(\l(I)rtO
=E
6
o
o
.cr
c
6
o-

co
6
c
.E
'Ex
o E€
[Link]
(t
E
ao
6
E
ctl
L
o E
E .9
o E
o
o. lo
o !o
(L o
lt'
o
.E
th
.F
(6 E
o
.=
.tt
lj
E
o
.!P
G' o
o -c
a o
(o il

o
ct
c o
a;
o
I
E
o (o o
F a
132 An English language placement test

there would be no motivation (other than academic interest, which


was not given as an option) to do the research. The main problem
with using option (d) as the key, for this informant, was that although
the text indicates that it is true, there is no evidence to suggest that
it is the prime motive for the research interest. Most scientists are not
interested in developing high-temperature superconductors just for the
sake of having something that operates at high temperatures, but
because they wish to levitate trains, and the like.
Although it is possible for language testers and applied linguists to
write good items which are relevant to the disciplines of the students
who are taking placement tests, this exercise has provided clear evi-
dence that it is important for subject specialists to be consulted
regarding the way in which they would interpret texts as 'insider'
readers. Failure to do this may lead to building bias into test items
at the construction phase, which may not be detected in the pretesting
phase without conducting time-consuming differential item function
studies.

5 Feedback from students


After the tests had taken place, a questionnaire was circulated to all
students who had been referred for English language support. Of the
l
435 questionnaires distributed,T were completed and returned. This
is a response rate of 16.32Vo, indicating that the following results
must be treated with some caution.
The questionnaire requested students to indicate, first, whether they
thought the in-sessional placement test was a 'fair' test of their ability
to opera ejn English il-au_ae_ademic,-context, measured on a fiv.e.
point Likert scale, and serondly whether they had attended any of
the English courses within the in-sessional programme at the English
Language Institute. There were open-ended prompts to discover why
students thought the test was 'unfair', if they did, how it could be
improved, ood to find out why referred students had not attended an
English course, if they had not. The questionnaire was circulated to
all referred students two weeks after the date of the placement test.
By this time, referred students were expected to have enrolled in
appropriate courses, and so the results of the questionnaire could be
compared directly with course attendance.
a Perception of 'fairness': Table 17 indicates that, generally, stu-
dents did perceive the in-sessional placement test to be fair, with a
mean score of 3.1 on the Likert scale. The raw scores are presented
in Table 18, to show that of the 7l respondents, one individual did
not answer the question, and only 14 thought that the test was either
'very unfair' or 'unfair'.
Glenn Fulcher 133

Table 17 Student perception ol test faimess

Minimum Maxirnum Standard enor

71 3.1 0.92 0.11

Table 18 Flesponses to the question

No response Very unfair Unfair Fair Quite fair Very lair Total

10 71

Of the 14 respondents who thought that the test lvas unfair, five
provided no reason for their view. The responses for the other nine
students were as follows:

Student /: Greek More time was needed to complete the


test
Student 2: British The testing environment was poor: the
room was too hot and there was no air
conditioning
Student 3: Bosnian The test should have been based on
materials which students had already
studied
Student 4: British More time was needed to complete the
test,-ffid thE-qre-stiohs -wore' ambiguous'
Student 5: Japanese There should be a listening component
Student 6: French The use of multiple-choice questions
should be avoided, as these are not
capable of testing proficiency in English
Student 7: Norwegian The questions were ambiguous. The
student also commented that the
questionnaire was badly designed, and he
could not understand any of the questions
Student 8: Japanese The invigilator turned up very late on the
day of the test, and test administration
was poor
Student 9: Singaporean There should be a vocabulary cornponent

It is clear from these responses that dissatisfaction was, for the most
part, not a function of the test itself, but of the restrictions which are
imposed on the testing process by the time and facilities available for
134 An English language placement test

testing. The complaint of 'ambiguity' in the questions can be taken


as a comment on the distracters in the multiple-choice questions. [n
a number of the questions the distracters operated extremely well, and
some students who got the answers incorrect complained bitterly.
Students who reported that they found the test fair, quite fair or
very fair, also echoed some of these views. Thirteen students
expressed the need for more time to answer the questions, but at least
half these also thought that the test should be longer too. [n particular,
three students requested additional writing tasks, five students
requested a speaking test (which is practically impossible for the
entire test-taking population) and four students requested a listening
test. One student wished to have a subject-specific test module
(business). The only other observation which was echoed, was that of
the poor test-taking environment, particularly the lack of ventilation in
rooms allocated for the test. There is some concern over the appropri-
ateness of rooms allocated for testing, particularly for those students
whose primary language is not English, as it is well known that
environmental conditions can have a negative impact on test scores
(Ascher, 1990), and those who supply rooms should be made awtue
of this for future administrations.

b Attendance on the in-sessional programme: The response to the


question of whether students had attended English language pro-
grammes was a simple yes/no choice. The results are compiled in
Table 19 by their response to the first question (Ql = perception of
fairness of the test). A statistical link cannot be tested, because the
nurnber of entries in cells is less than the proportion which would
generally be consideretl-aceeptaltle to-Fmfide reliatlle p values.
Of the 20 students who were referred to English language support,
did not attend an English programme and replied to the questionnaire,
18 provided reasons for nonattendance:

Table 19 The number of students attending English language programmes by response


to the question on the perceived laimess of the placement test

Response to Ql Did not attend Did attend Total

l,lo response 1 0 1
Very unfair 3 (75%l 1 4
Unfair 2 (20Phl 8 10
Fair 6 (17Vol I 35
Quite fair 6 (357.) 11 17
Very fair 2 (W"l 2 4
Total n egv"l 51 71
Glenn Fulcher 135

. Not enough time because of subject-specific study commitments


(eight students).
o Claimed they did not know about the in-sessional programme
(three students).
. After the oral interview, they were not required to attend (three
students).
. Does not know how to use the library, BDd spends all free time
looking for subject-specific materials (one student).
o Attending English classes is a waste of time (three students).
ln the last case, there appear to be two different attinrdes. In the case
of one student whose primary language was English, he was
'offended' at the suggestion that he may need additional support. In
the other two cases, the number of years of English as a second langu-
age study w?p seen to 'exempt' them from the need for further langu-.
age study. Cftre student wrote: 'I don't think because is useful except-
essay writing. I spent too many yeius learning English.'
Attendance at in-sessional programme courses is generally very
high, and the number of students referring themselves is higher than
the number of referrals who do not attend. Total attendances for the
in-sessional programme in the L994195 academic year reached 5ll
students. For an optional programme, this seems to be a reasonably
satisfactory state of affairs.

VII Discussion
I Ingistic and administrational constraints
ry large institution there will be-logistieandadffirstratio-n-al -
-Wftlrin
constraints. What is not often recognized, however, is that these con-
straints lead to limitations on testing, which have a direct impact on
reliability of score interpretation. It is therefore important that insti-
tutions make policy decisions which take this relationship into
account. It should never be assumed that decisions taken by those in
administration do not have academic impact. In the case of the Uni-
versity of Surrey, I practical compromise has been reached in the
form of a two-tier testing system, with the less precise placement test
screening out students from the oral interview stage of the process.
This has clearly had a negative impact upon the usefulness of the
reading subtest, which needs to be lengthened.

2 The need for pretesting


This study has highlighted the importance of rigorously pretesting
all items for inclusion on placement tests. Even with trialling, some
t 36 An English language placement test
unsuitable items will survive. Without pretesting and post hoc analy-
sis, no institution could be sure that its placement tests were providing
better information than tossing a coin 1619 times. It is important that
institutions such as universities know the estimates of eror in their
placement procedures, in fairness to the students, and to improve
screening and language support provision

3 Equating test forms


The methodology used in this study will allow forms of the test to
be equated, using anchor items. This means that test results will be
comparable from year to year, and the English Language Institute
will be able to monitor the English language ability, as defined by
the test, of incoming students in the future. This will allow it to alert
the university authorities to sudden or gradual changes in English
language support requirements. It is very common for teachers and
students to make comparisons across tests or test formats. In many
cases this is not possible, but when forms are equated, it is both poss-
ible and beneficial.
Section I tasks cannot be equated in the same way as Sections 2
and 3. In the case of Section I it will be necessary to link writing
prompts using expert judgements, ensuring that prototypical answers
awarded at each band of the rating scale are available for rater train-
ing. From form to form, much more care will have to be taken with
the interpretation of Section l, even though its reliability in the form
discussed here is higher than Sections 2 or 3. This type of equating is
essentially judgmental, and one of the less exacting forms of linkage.
the OaSe of Sections 2 and 3, a Stfong form of tEStE-qUatingmay
be attempted, involving strong statistical linking.
-fn
4 Future research
In the next academic y€ff, a further form of the test will be produced,
and equated with the first form, using a single-group design with con-
current validation (Petersen, Kolen and Hoover, 1989: 256). In prin-
ciple, this could be done each year until enough forms have been
developed to ensure test security even if one form were compromised.
As scores on all forms would be strictly comparable (Linn, 1993:85)
it would not matter which form students take. This would add to
test security, while maintaining score interpretability. However, this
project may be somewhat ambitious. Equating tests is much more
difficult with short tests than longer tests, as the burden of information
provided by each individual item is much higher in shorter tests
(Linn, 1993: 88).
Glenn Fulcher 137

If, however, we are able to produce multiple forms, it also becomes


possible to conduct score gain studies. It is a truism that in British
universities there is no evidence to suggest that after a certain amount
of time on any given programme, a student will 'improve' by 'x' ,
whatever the unit '.r' is in terms of ability. Indeed, score gain studies,
even if conducted, would be meaningless, without a clear understand-
ing of what 'r' is. In this case, the unit of measurement is the scale
which has been established in the October 1994 study. This is arbi-
trary, but person and item free (Masters, 1990). As such, score gain
studies may be conducted with new groups of students over different
timescales, following different courses. If successful, this would allow
the English Language Institute to say to a department that a particular
student would require (given error) from 'a' to'D' months of langu-
age tuition to reach a level at which he or she would, with p prob-
ability, be able to cope with an academic course with (or without)
English language support.
This kind of information is not currently available, but in principle
could be, as the result of a careful development of research within
the context of specific language programmes of large educational
institutions.

Vm Conclusion
We are aware that there are other approaches to developing placement
tests, many of which are effective where the numbers of students
__taking the tr$_qlg small 4nd the constraints in local _administration
allow the use of lengthy instruments which include oral components
(Paltridge, 1992). Each approach must be sensitive to the constraints
imposed by, and the information requirernents of, unique institutions.
Similarly, when conducting score gain studies in the future, criterion-
based approaches such as that suggested by Brown (1989) will be
considered, especially if multiple forms of the current test are avail-
able. We nevertheless believe that the current test fulfils its purpose
as well as can be expected within the context described in this article.
The methodology used in developing the evaluation of this place-
ment test was based on Wall, Clapham and Alderson (1994), as the
first article in the field of language testing to address the issue of
the assessment of placement instruments in any depth. However, the
methodology used in this study includes additional aspects which
would appear to be worthy of investigation in the particular context
of placement testing. It is hoped that the approach adopted, building
on Wall, Clapham and Alderson's pioneering work, may be of use
to others in the evaluation of their placement tests.
138 An English language plncement test

LK Rderences

APA 1985: Standards for educational and psychological rasdng. Wash-


ingon, DC: Arnerican Psychological Association.
[Link], C. 1991J: Assessing bilingual swdents tor placenunt and instruaion.
ERIC Digest 65, 8D322n3. New York: ERIC Clearinghouse on
Urban Education.
Ass€ssment Systems Corporation 1994: MSCAL 3.5, Rasch analysis pro-
8rarz. Minnesota, MN: ASC.
Brown, JJ). 1989: lmproving ESL placement tests using two perspectives.
TESOL Suarterb 23, 65-83.
Crocker, L and Algin& J. 1986: Introduction to classical and n dem test
theory. Chicago,Il-: Holt, Rinehart & Winston.
Farhady, H. 1983: On the plausibility of the unitary language profcierrcy
facton In Oller, J.W,, {ior, Issues in language tef,ring, Rowley, MA:
l-28.
Newbury House, I
.
Goodbody, M.W. 1993: bning the students choose: a placement procedure
for a pre-sessional course. In Blue, G.M., di|cr, Itnguage, leaming
and success: studying rhrough English, L,ondon: Modem English Pub-
lications and the British Council, 49-57.
Lin4 RJ^ 1993: Linking results of distinct assessmqtts. Arylicd Measure-
ment in Hucation 6, 83-102.
Maslers, GJ{. 1990: Psychometric aspects of individual tneasurcmenl In
de Jong, H.A.L. and Stevenson, D.K., editors, IndiviAnhzinS thc
assessment of language abilities, Clevdon: Multilingual Mauen,
56-70.
Paltridge, B. 1992: EAP placement testing: an integrated approarh. English
lor Specifc Purposes ll,243-68.
Jctaseorl{S"I(olen,}[Link] Hoover, ILD. 1989: Scaling norming-and---.-
ditor, &lucational mcaswemmt, New York:
equating, In Linn, R.L.,
National Council on Measurement in E<lucation and Maqnillan, ..

z\t-62.
Pophanr, WJ. l9?8: Criterion-referenced mcasuremznt. Englewood Cliffs,
NJ: hentice-Hall.
Slchrdq C.C. and DelMes, RC. 1990: Determining the validity of place.
ment exams for developmental college anicula. Applied Measurement
in Educarton 4, 37-52.
Shohamy, E. 1983: The stability of oral proficiency ass€ssment in the oral
interview procedure. Language lzaraing 33, 527 -40.
1988: A proposed framework for testing the oral language of second
foreign language leatms. Studies in Second language Acquisition lO,
- t65-79.
1990: Ianguage testing priorities: a different perspecive. Foreign Lan-
guage Awals 23, t85-94.
-Shohamy, E., Reves, T. and Bejarano, Y, 1986: Intoducing a new comprc-
hensive test of oral proficienry. English Language Teaching Journal
40,2t2-20.
Glenn Fulcher 139

Tnrman, \ry.L. 1992: College placement testing. AMATYC Review 13,


58-64.
Utrxhur, J.A. l97l: Objective evaluation of oral proficiency in the ESOL
classroom. TESOL Quarterly 5, 47 -59.
van Weeren, J. l98l: Testing oral proficiency in everyday situations. In
Klein-Braley, C. and Stevenson, D.K., editors, Practice ard problems
in language testing. Volume /, Frankfurt: Bern, 54-59.
Wall, D., Clapham, C. and Alderson, J.C. 1994: Evaluating a placement
test. Innguage Testing lL, 321-344.

You might also like