0% found this document useful (0 votes)
14 views22 pages

TOEFL iBT Scores and Academic Success

This study explores the relationship between TOEFL iBT scores and academic performance, specifically GPA, among 2594 students from 10 U.S. universities. While the correlation between TOEFL scores and GPA was not strong, higher TOEFL scores generally indicated higher GPAs, suggesting that TOEFL iBT scores provide additional predictive validity for non-native English speakers beyond other admissions tests. The study discusses the complexities of predicting academic success and the limitations of using GPA as a criterion for such predictions.

Uploaded by

Yewande Emmanuel
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
14 views22 pages

TOEFL iBT Scores and Academic Success

This study explores the relationship between TOEFL iBT scores and academic performance, specifically GPA, among 2594 students from 10 U.S. universities. While the correlation between TOEFL scores and GPA was not strong, higher TOEFL scores generally indicated higher GPAs, suggesting that TOEFL iBT scores provide additional predictive validity for non-native English speakers beyond other admissions tests. The study discusses the complexities of predicting academic success and the limitations of using GPA as a criterion for such predictions.

Uploaded by

Yewande Emmanuel
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

430368

2012
LTR29310.1177/0265532211430368Cho and BridgemanLanguage Testing

/$1*8$*(
Article 7(67,1*

Language Testing

Relationship of TOEFL
29(3) 421­–442
© The Author(s) 2012
Reprints and permission:
iBT® scores to academic [Link]/[Link]
DOI: 10.1177/0265532211430368
performance: Some evidence [Link]

from American universities

Yeonsuk Cho and Brent Bridgeman


Educational Testing Service, USA

Abstract
This study examined the relationship between scores on the TOEFL Internet-Based Test
(TOEFL iBT®) and academic performance in higher education, defined here in terms of grade
point average (GPA). The academic records for 2594 undergraduate and graduate students were
collected from 10 universities in the United States. The data consisted of students’ GPA, detailed
course information, and admissions-related test scores including TOEFL iBT, GRE, GMAT, and
SAT scores. Correlation-based analyses were conducted for subgroups by academic status and
disciplines. Expectancy graphs were also used to complement the correlation-based analyses by
presenting the predictive validity in terms of individuals in one of the TOEFL iBT score subgroups
belonging to one of the GPA subgroups. The predictive validity expressed in terms of correlation
did not appear to be strong. Nevertheless, the general pattern shown in the expectancy graphs
indicated that students with higher TOEFL iBT scores tended to earn higher GPAs and that the
TOEFL iBT provided information about the future academic performance of non-native English
speaking students beyond that provided by other admissions tests. These observations led us to
conclude that even a small correlation might indicate a meaningful relationship between TOEFL
iBT scores and GPA. Limitations and implications are discussed.

Keywords
academic performance, English proficiency, grade point average, international students, predictive
validity, TOEFL

The Test of English as a Foreign Language (TOEFL®) is one of the English proficiency
tests that non-native English speaking (NNES) students may take to demonstrate their
language proficiency when applying to English-medium colleges and universities.
TOEFL scores or similar information are often considered by many institutions in

Corresponding author:
Yeonsuk Cho, Educational Testing Service, Rosedale Road, MS 07-R, Princeton, NJ 08541, USA
Email: ycho@[Link]

Downloaded from [Link] at PENNSYLVANIA STATE UNIV on September 18, 2016


422 Language Testing 29(3)

determining whether a prospective NNES applicant has met, as Chalhoub-Deville and


Deville (2006) put it, ‘a linguistic threshold that enables them to approach academic
work in English in a meaningful manner’ (p. 520). Given the widespread use of TOEFL
scores for admissions purposes, a number of studies have been carried out to find empiri-
cal evidence to support such a use of test scores.
TOEFL underwent several major improvements to better reflect language use in an
academic context, leading to the introduction of the TOEFL Internet-Based Test (TOEFL
iBT®) worldwide in 2006. As the TOEFL iBT is quite a different English proficiency test
from the earlier versions of the TOEFL, new evidence is sought to support the use of
TOEFL iBT scores for the same purpose in admissions decisions. Using data from 10
universities in the United States, this study revisits an old question in the TOEFL context
– what is the relationship between performance on an English proficiency test and aca-
demic performance? Before presenting the results of the current investigation, we will
briefly summarize the predictive validity evidence related to the predecessors of TOEFL
iBT and discuss some general issues raised around predictive validity studies of lan-
guage tests for academic performance that are also relevant to the present study.

Predictive studies examining the validity of TOEFL1


Undoubtedly, English language proficiency is a critical factor for the academic perfor-
mance of NNES students in a setting where English is used for teaching and learning. For
this reason, many institutions of higher education in English-speaking countries require
prospective NNES students to demonstrate English proficiency with a standardized test
score as one of the admissions requirements. Logically, such a mandate has generated a
great deal of research interest in the relationship between performance on language tests
and future academic performance.
Graham (1987) has given a good summary of predictive validation studies of lan-
guage tests. She compared 18 studies, many of which focused on TOEFL, and catego-
rized them according to their results. Her review revealed inconsistent results across the
studies. Graham concluded that the research findings did not provide ‘clear-cut answers’
for admissions officers. Nevertheless, the lack of consistent evidence, she added, did not
contradict the importance of language proficiency in academic settings because the
nature of the relationship between language proficiency and academic success is com-
plex and difficult to demonstrate, in part due to a lack of adequate data.
More studies have been conducted subsequently, but using small samples in a limited
number of academic fields. Ayers and Quattlebaum (1992) analyzed the TOEFL and
Graduate Record Examination® (GRE) data for 67 Asian graduate engineering students
from one institution. The study demonstrated that GRE quantitative scores were the best
predictor of GPA for the engineering students in the sample: GRE scores explained about
10% of the variance of GPA. TOEFL was moderately correlated with the GRE verbal
(r = .63), but its correlation with GPA was almost negligible (r = .05). Similar results
were found in another study by Neal (1998). GRE quantitative scores were the best pre-
dictor of GPA for 47 international students in the study who completed a master’s pro-
gram in science or engineering (r = .33), but a negative correlation between TOEFL and
GPA was observed (r = −.14). The poor correlations between TOEFL and GPA in the two

Downloaded from [Link] at PENNSYLVANIA STATE UNIV on September 18, 2016


Cho and Bridgeman 423

studies are in contrast with the result of a study by Ayers and Peters (1977), as cited in
Ayers and Quattlebaum (1992), in which a much larger correlation was observed for a
sample of 50 Asian graduate students in physical science or engineering programs (r =
.40, p < .01). These three studies focused on the predictive validity of TOEFL for stu-
dents in master’s-level science or engineering fields, but like the earlier studies summa-
rized in Graham (1987), the results from these three studies are inconsistent.
The value of TOEFL in predicting academic performance has also been investigated
in contexts outside North America. Vinke and Jochems (1993) examined the relationship
between TOEFL and academic success, using a sample of 90 Indonesian students who
were studying engineering in a postgraduate program in the Netherlands where English
was the language of instruction. One of the criteria for academic success were the aver-
age scores on seven written qualifying exams required as part of the degree requirement.
TOEFL scores showed a moderate correlation with the average exam scores (r = .51).
Even stronger evidence was found when comparing the passing rates for the qualifying
exam of two TOEFL subgroups in the same study: 74% for students with TOEFL scores
below 450, compared to 98% for those with TOEFL scores above 450.
Al-Musawi and Al-Ansari (1999) examined the predictive validity of TOEFL in com-
parison with that of the First Certificate in English (FCE). FCE is one of the English
proficiency examinations developed by the University of Cambridge ESOL examina-
tions. The test is claimed to be at the level B2 of the Council of Europe Common
European Framework of Reference for Language (CEFR), and it measures four skills
and explicit knowledge of grammar and vocabulary. The study included the academic
records of 86 English major undergraduate students in Bahrain. Sub-scores on both
TOEFL and FCE were used as predictor variables, and overall GPA and GPA in English
courses were used as measures of academic success in a stepwise regression. Despite
some moderate correlations with overall GPA, none of the TOEFL sub-scores contrib-
uted to the prediction of overall GPA when the FCE scores were already entered into the
equation. Similar results were observed in the prediction of average grades earned in
English courses. Based on the results, the researchers concluded that FCE might be a
better predictor of academic success of language learners in EFL contexts than TOEFL.
However, their interpretation should be taken in light of the analytical approach used in
the study. Because both FCE and TOEFL are measures of English proficiency, the sub-
scores on the two tests would have been redundant, creating a collinearity problem in the
analysis. It is known that the selection of variables in a stepwise regression is affected by
the correlations among the predictors. Also, results of a stepwise regression from one
sample are hard to replicate in another sample.
More recently, using logistic regression, Van Nelson, Nelson, and Malone (2004) tried
to determine which combination of variables, including TOEFL, could predict the aca-
demic success of international students. The data came from 866 students in master’s-
degree programs across various academic fields. Both final GPA and degree completion
were used as criterion variables. For the logistic regression analysis, GPA was dichoto-
mized by dividing the students into two groups – a group of students with a GPA of 3.5
and above, and the other with students who received a GPA below 3.5. The results were
mixed; when degree completion was used as a criterion, TOEFL was not a good predictor
of academic success. However, it did contribute to the prediction of the final GPA.

Downloaded from [Link] at PENNSYLVANIA STATE UNIV on September 18, 2016


424 Language Testing 29(3)

Why is it so difficult to predict academic success with a


language test?
As shown in the studies discussed above, research findings are mixed, which makes it
difficult to draw a definitive conclusion about the value of TOEFL or similar tests in
predicting academic performance. Predictive validity studies are difficult to design and
interpret for a number of reasons.
First, a fundamental issue is that there is no one-to-one correspondence between lan-
guage and academic performance. Language is a crucial factor in learning, but it is only
one of many factors. Evidence and common sense suggest that lack of proficiency may
interfere with NNES students’ academic ability when English is a medium of instruction.
However, being linguistically proficient does not guarantee academic success. There are
many other factors such as motivation, learning strategies, and quantitative skills that
also contribute to one’s academic performance. One obvious example that attests to this
is that not all native speakers are academically successful. By the same logic, there is no
reason to expect that all NNES students with high levels of English proficiency would
necessarily have high GPAs. In this regard, showing that language does impact academic
performance is quite challenging.
The lack of appropriate criterion variables is another issue in conducting predictive
validity research for language tests as well as admissions tests (e.g. Hartnett &
Willingham, 1980; Graham, 1987; Heil & Aleamoni, 1974; Zwick, 2002). GPA has been
pointed out to be of limited value even in the literature surrounding predictive validity
studies of admission tests – such as the Scholastic Aptitude Test (SAT) – that, in fact, are
designed to predict academic performance. The relationship between admissions test
scores and GPA is hard to demonstrate also because of the effects of other factors on
GPA. Despite its limitation, GPA is by far the most widely used criterion mainly because
it is easy to obtain (Kuncel, Crede, & Thomas, 2007). Hartnett and Willingham (1980)
discuss other indicators of academic success and the drawbacks associated with them.
The purpose of TOEFL or similar assessments is quite different from that of other
admission tests in that language tests are designed to determine a ‘linguistic threshold’ for
learning academic content while other admissions tests focus on a broad range of aca-
demic reasoning skills. However, because language tests are used in the same context as
admissions tests, the relationship between English proficiency and academic performance
is of interest to test users – especially admission officers – and of relevance in supporting
the use of test scores for high-stakes admissions decisions. Thus, even though there may
be no direct correspondence between language proficiency and academic performance
beyond a certain level of language proficiency, many validity studies of language tests
continue to use the latter as a criterion (e.g. Fox, 2004; Watt & Roessingh, 2001, cited in
Fox, 2004; Woodrow, 2006). Alternative criteria that are more in line with language tests
in terms of construct, such as students’ self-evaluations of language abilities and faculty
ratings of students’ language abilities have been suggested and included in some predic-
tive validity studies (e.g. Powers, Kim, & Weng, 2008; Sinharay et al., 2009; Woodrow,
2006). Nevertheless, alternative criteria also complicate the interpretation of results
because of other factors influencing these measures. Subjectivity, for instance, is an obvi-
ous, complicating factor. Faculty evaluations of students’ English proficiency may be

Downloaded from [Link] at PENNSYLVANIA STATE UNIV on September 18, 2016


Cho and Bridgeman 425

influenced by the faculty members’ own experiences, such as their individual fields or
prior exposure to NNES students. Similarly, as noted by Ross (1998), respondents’ experi-
ences and variation in the interpretation of a self-assessment question affect how they
answer questions. This is particularly an issue when respondents are from diverse back-
grounds and cultures, as is the case with the population of TOEFL test takers.
Whether the criteria are appropriate or not, range restriction is mentioned in almost
every predictive validity study because the validity evidence is typically expressed in
terms of correlation, and when the data do not represent a full range – that is, there is no
information regarding how those who did not get selected would have performed after-
ward – a correlation is underestimated. Although some statistical adjustment can be
made to correct correlations for range restriction (Ree, Carretta, Earles, & Albert, 1994;
Sackett & Yang, 2000; Wiberg & Sundstrom, 2009), some issues have been noted with
correction methods. For example, possible sign changes after correction, and specifica-
tion errors in selecting an appropriate correction formula can occur (Ree et al., 1994).
The lack of criterion-related information for those not included in a study is a problem in
designing and interpreting predictive validity studies when the real question of interest is
how performance on a test relates to the future performance of all test takers.
A number of researchers expressed the view that correlation is inherently difficult to
interpret even for experienced social scientists (Schrader, 1965; Rosenthal & Rubin,
1982; Sackett, Borneman, & Connelly, 2008). Recognizing these drawbacks of using
correlations as a validity index in predictive validity research, Schrader (1965) suggested
using an expectancy table as a concrete way to represent test validity. An expectancy
table is a simple way to summarize ‘the relation between two (or more) variables by stat-
ing the probability that individuals who belong to each of a set of subgroups defined on
the basis of one (or more) variables will belong to each of a set of subgroups defined on
the basis of another variable’ (Schrader, 1965, p. 29). One advantage of this approach is
that it is easy to explain results to a lay audience. More recently, Sackett et al. (2008)
stated that the magnitude of even small correlation coefficients is not well understood,
and pointed out the benefit of a similar approach: ‘converting correlations to differences
in odds of success results both in a readily interpretable metric and in a positive picture
… in short, there is a long history of expressing the value of a test in a metric more read-
ily interpretable than percentage of variance accounted for’ (p. 216).
The current study investigated the predictive validity of TOEFL iBT for academic
performance with both correlation-based analyses and expectancy graphs as analytic
methods. In addition, because the reading and writing skills assessed in the TOEFL iBT
overlap somewhat with the skills assessed by admissions tests that are already required
(e.g. SAT and GRE), admissions officers may be interested in finding out whether
TOEFL iBT provides unique information that cannot be obtained from the other admis-
sions tests. We thus asked the following two research questions:

1. What is the relationship between TOEFL iBT scores and future academic perfor-
mance as defined by GPAs? Is the relationship the same across fields of study?
2. Does TOEFL iBT provide additional information beyond what other admissions-
related tests can in predicting GPAs? Is the relationship the same across fields of
study?

Downloaded from [Link] at PENNSYLVANIA STATE UNIV on September 18, 2016


426 Language Testing 29(3)

Method
Data source
Initially, 40 institutions in the United States with large numbers of international students
according to Open Doors Online ([Link] were contacted for
participation in the study. The following information was requested for full-time, degree-
seeking, non-transfer students who had TOEFL iBT scores and at least a first-year GPA
from the current institution:

1. demographic information that does not reveal student identity;


2. grade point average (GPA), preferably broken down by year;
3. scores on admissions-related tests including TOEFL iBT, GRE/GMAT, SAT/ACT;
4. for undergraduate students, detailed course-level information including a course
title, credit hours, grade, and the semester taken.

Ten universities agreed to participate in the study. Seven schools provided informa-
tion for both undergraduate and graduate students whereas the other three sent data
only for graduate students. Course-specific grades were available for four schools.
Geographically, half of the schools were in the Midwest, and the others were in various
regions. Except for one private university, all were large public universities. According
to the admission selectivity published in U.S. News, seven schools in our study sample
were regarded as ‘more selective,’ one ‘most selective,’ one ‘selective,’ and one ‘less
selective.’ A total of 2594 students – 1850 graduate students and 744 undergraduate stu-
dents – were represented in the data.

Data preparation and analysis


Coding
Majors and individual courses were coded into four broad categories: business (BU),
humanities and arts (HA), sciences and engineering (SE), and social sciences (SS).
Majors and courses that could not be categorized with the four categories or were not
known were coded accordingly as other (OT) or unknown (UN).

Variables
The following variables were included in the study.

a. GPA: An overall GPA. Types of GPA information varied across schools: some
schools sent cumulative GPAs while others provided year-by-year GPAs or
course information that allowed the calculation of a GPA. The distinction among
the different types of GPA information was, not made, however.
b. GPA_BU, GPA_HA, GPA_SE, and GPA_SS: Discipline-specific GPAs are
weighted average grades earned in courses in the same academic category. They
were computed only for undergraduate students whose individual course

Downloaded from [Link] at PENNSYLVANIA STATE UNIV on September 18, 2016


Cho and Bridgeman 427

information was available. During the subgroup analysis by majors, the discipline-
specific GPAs were used as a criterion instead of overall GPAs.
c. TOEFL_IBT: A total TOEFL iBT score.
d. SAT_RW: A combined score from the SAT reading and writing sections for under-
graduate students.
e. SAT_M: A SAT math score.
f. GRE_V: A GRE verbal score for non-business graduate students. GRE also
reports a writing score separate from a verbal score. Although it was preferable to
use a combined score of verbal and writing scores, GRE_V was used because
writing scores were missing for many students in the study sample.
g. GRE_Q: A GRE quantitative score.
h. GMAT: A total GMAT score for business graduate students. GMAT reports sub-
scores for the verbal and math sections. However, the sub-score information was
not available for many students in the study sample.

Analyses
Predictive validity was examined in terms of correlation and probability. Correlation and
hierarchical multiple regression analyses were performed. These analyses were then
complemented by expectancy graphs. Because of obvious differences in grading stand-
ards and test scores, analyses were conducted separately for academic levels, different
disciplines and institutions. Results were then aggregated across institutions.

Correlation-based analyses
Simple correlations were computed to address Research Question 1. To answer Research
Question 2, SAT_RW was entered as a single predictor of GPA or disciple-specific GPA
(Model 1), and TOEFL iBT was then added (Model 2) for undergraduate students.
Furthermore, although it was not the focus of our research, because quantitative skills
are considered an important element in academic performance and measured accord-
ingly in many admissions tests, SAT_M was subsequently added (Model 3). The regres-
sion analyses were done in a similar manner for graduate students. However, for
business graduate students, only two regression models were compared because sub-
scores of the GMAT were not available. A difference in squared multiple correlation
(R2) between the first two regression models was used as an index of incremental valid-
ity of TOEFL iBT – that is, whether TOEFL iBT provides unique information beyond
the language measures of other admission-related tests. In running the hierarchical mul-
tiple regression models, collinearity among the predictors was checked by examining
the two collinearity indices, tolerance and variation-inflation factor (VIF), and no sig-
nificant collinearity was present.
A weighted average of simple correlations and squared multiple correlations (R2)
was then used to aggregate the results of correlation-based analyses across institu-
tions. Sample sizes, which varied across disciplines and institutions, were used to
give more weight to the results based on the larger samples. In order to avoid the capi-
talization of chance factors that could artificially influence regression estimates in

Downloaded from [Link] at PENNSYLVANIA STATE UNIV on September 18, 2016


428 Language Testing 29(3)

small samples, an adjusted R2 was computed within each institution prior to comput-
ing a weighted average.

Expectancy graphs
For expectancy graphs, the students in the study were first assigned to one of three
subgroups (i.e. top 25%, middle 50%, and bottom 25%) according to their relative
standing on GPA and test scores within institution and within academic status (and
further within academic discipline for graduate students). This was done to accommo-
date to different academic standards across academic levels, disciplines, and schools.
For example, a GPA subgroup that a graduate student with a computer science major
belonged to was determined with respect to graduate peers in science and engineering
(SE) majors within the same university. Academic discipline, however, was not taken
into consideration in assigning undergraduate students to groups because (1) for many
undergraduate students in the study, majors were unknown, but more importantly (2)
their first years of university education were less likely to reflect academic perfor-
mance within a single discipline.
Once the subgroups were created, expectancy graphs were drawn by cross-tabulating
TOEFL iBT and GPA subgroups for Research Question 1. Expectancy graphs were also
used to address Research Question 2, but with some modification. In this part of the
analysis, the incremental validity of TOEFL iBT was examined by focusing on students
whose performance on other admissions tests (i.e. GRE_V, GMAT, or SAT_RW) was in
the middle 50% of a score distribution – that is, those in the top and bottom 25% of
GRE_V were not included in the analysis of expectancy graphs. This was done for both
conceptual and logistical reasons.
Test takers who receive extremely low or high scores on TOEFL iBT are also likely
to receive extremely low or high scores on the language measure of an admissions test,
whereas there is more uncertainty in the middle score range – that is, the relationship of
two similar tests is less certain when you look at the test scores in the middle range. Thus,
TOEFL iBT scores might be particularly useful in predicting the academic performance
of students whose GRE verbal scores are in the middle score range. Furthermore, because
of extremely rare occurrences of students who did very well on GRE_V or SAT_RW but
who received very low scores on TOEFL iBT, it was not practically possible to demon-
strate the incremental validity of TOEFL iBT for students in the top and bottom 25%
GRE_V or SAT_RW score groups. Therefore, expectancy graphs were drawn using the
students whose test performance on GRE_V or SAT_RW belonged to the middle 50%
group. To be parallel to this approach, the students in the top 25% and bottom 25%
GMAT groups were also excluded from expectancy graphs.

Results
Descriptive statistics
The data included the academic records of a total of 2594 students consisting of 1850
graduate students and 744 undergraduate students. Table 1 presents the summary

Downloaded from [Link] at PENNSYLVANIA STATE UNIV on September 18, 2016


Table 1. Descriptive statistics of GPA and admission test scores

GPA TOEFL iBT GRE_V / SAT_RWa GMAT

Status School n M (SD) n M (SD) n M (SD) n M (SD)


Graduate A 213 3.65 (.28) 213 100 (11.07) 150 490 (132.37) 6 555 (92.88)
Cho and Bridgeman

B 118 3.52 (.42) 118 94 (11.36) 60 383 (98.26) 10 551 (70.47)


C 236 3.69 (.29) 236 99 (13.33) 130 465 (129.47) 38 621 (75.13)
D 157 3.74 (.34) 157 98 (15.43) 120 458 (138.83) 9 651 (81.16)
E 214 3.57 (.41) 214 101 (11.92) 136 497 (134.01) 20 603 (69.88)
F 160 3.66 (.38) 160 103 (10.89) 123 499 (135.45) 4 643 (47.84)
G 250 3.64 (.34) 250 99 (15.09) 206 465 (137.24) 35 635 (86.75)
H 300 3.62 (.36) 300 94 (11.66) 140 385 (100.62) 71 537 (84.33)
I 159 3.64 (.29) 159 105 (10.09) 79 481 (129.98) 54 687 (45.13)
J 43 3.40 (.54) 43 100 (10.03) 20 391 (77.22) 19 586 (88.71)
Undergraduate A 85 3.30 (.58) 85 91 (16.42) 76 966 (179.91)
B 15 3.29 (.79) 15 84 (16.63) - -
C 200 3.08 (.70) 200 96 (12.62) 174 1079 (145.05)
D 175 3.06 (.67) 175 96 (10.51) 156 1026 (135.93)
E 107 2.98 (.74) 107 82 (15.55) 63 887 (181.00)
F - - - - - -
G - - - - - -
H 27 3.33 (.56) 27 85 (21.68) 26 966 (169.90)

Downloaded from [Link] at PENNSYLVANIA STATE UNIV on September 18, 2016


I - - - - - -
J 135 3.15 (.80) 135 86 (13.71) 28 956 (111.89)
a GRE Verbal for graduate schools, and SAT Reading plus SAT Writing for undergraduate schools.
429
430 Language Testing 29(3)

statistics for GPA and test scores by individual school. As expected in a restricted sample
of enrolled students, the average TOEFL scores were higher and the standard deviations
were smaller than those reported in the summary statistics of all examinees ([Link]
org/toefl/research/test_score_data_summary). On average, the graduate students had
higher TOEFL iBT scores than the undergraduate students, and this was true for all insti-
tutions. In addition, the summary statistics showed that graduate students tend to receive
higher grades than undergraduate students. A comparison of the statistics indicated that
there was little variability across schools in terms of average GPA. Average test scores
varied somewhat across schools.
With respect to academic discipline, science and engineering (SE) majors were most
popular (47%) among the graduate students, followed by business (BU) majors (26%),
social sciences (SS) majors (16%), and humanities and arts (HA) majors (10%). SE
majors were also popular among the undergraduates (26%). About 23% of the under-
graduate students had HA majors. However, majors were unknown or not declared for
44% of the undergraduate students.
The cut-points for the subgroups are summarized in Table 2. The cut-points varied
considerably between undergraduate and graduate students, and across academic

Table 2. Ranges of the cut-scores for the subgroups based on test scores and GPAs

Group

Status Variable Bottom 25% ≤ Top 25% ≥


Graduate GPA
BU 2.86 – 3.68 3.59 – 4.00
HA 2.00 – 3.82 3.96 – 4.00
SE 3.18 – 3.60 3.79 – 4.00
SS 2.42 – 3.76 3.72 – 4.00
TOEFL iBT
BU 88 – 107 90 – 114
HA 88 – 103 107 – 112
SE 84 – 97 109 – 111
SS 86 – 100 103 – 112
GRE_V/GMAT
BU 455 – 660a 590 – 720a
HA 300 – 450 410 – 670
SE 320 – 420 430 – 600
SS 275 – 373 400 – 690

Undergraduate GPA 2.63 – 3.06 3.56 – 3.90


GPA_BU 2.21 – 3.37 3.79 – 4.00
GPA_HA 2.91 – 3.39 3.82 – 3.86
GPA_SE 2.43 – 2.74 3.50 – 3.77
GPA_SS 2.00 – 3.12 3.60 – 4.00
TOEFL iBT 69 – 88 93 – 106
SAT_RW 760 – 980 1000 – 1180
a For business majors, these values represent scores on GMAT.

Downloaded from [Link] at PENNSYLVANIA STATE UNIV on September 18, 2016


Cho and Bridgeman 431

disciplines and institutions. For example, in one school, a graduate business student
whose GPA was below 3.68 was assigned to the bottom 25% GPA group, whereas a GPA
of 3.59 or above was considered the top 25% performance in another school. We would
like to emphasize at this point that being in the bottom 25% group in the current study
did not necessarily constitute academic failure. In fact, as these cut-scores clearly indi-
cate, students in adjacent subgroups were not much different in terms of ability being
considered.

What is the relationship between TOEFL iBT scores and future


academic performance as defined by GPAs? Is the pattern the same
across fields of study?
Tables 3 and 4 present the correlation coefficients of TOEFL iBT with GPA for graduate
and undergraduate students, respectively, along with the weighted average correlations
(rw) in the last row of each table. Both observed correlations (robs) and correlations cor-
rected for range restriction (rcor) are presented, and the results in this section are dis-
cussed in terms of observed correlation. Corrections were made using the Thorndike
Case 2 formula and assuming the standard deviations of TOEFL iBT scores reported in
the TOEFL iBT score summary for 2007 as unrestricted population parameters. Thorndike
Case 2 formula is used when the range restriction is due to direct selection – that is, indi-
viduals are selected based only on a selection measure, and the variance of the unre-
stricted population is known only for the selection measure (Sackett & Yang, 2000).
Comparisons of the weighted average correlations across academic disciplines at the

Table 3. Observed and corrected correlations between TOEFL iBT and GPA for graduate
students in four major categories

School BU HA SE SS
(n=413) (n=186) (n=959) (n=283)

n robs rcor n robs rcor n robs rcor n robs rcor


A 38 .28 .52 32 .24 .46 126 .13 .27 17 .14 .28
B 11 −.03 −.05 - - - 92 .17 .34 13 .34 .59
C 46 .16 .28 34 .26 .42 107 .33 .51 49 .41 .61
D - - - 23 .17 .25 72 −.12 −.18 61 .08 .12
E 65 .35 .58 14 .30 .52 101 .24 .43 33 .23 .42
F 19 −.29 −.54 14 .20 .39 112 .14 .28 15 .12 .25
G 30 .03 .05 28 .34 .48 145 .27 .39 44 .19 .28
H 120 .32 .55 18 -.07 -.14 136 .07 .13 25 .39 .64
I 69 .40 .71 23 .38 .68 56 .39 .69 11 .44 .74
J 15 .31 .60 - - - 12 −.14 −.31 15 .35 .65

Mw .26 .29 .24 .41 .17 .28 .25 .35


Note: Major categories are business (BU), humanities and arts (HA), science and engineering (SE), and social
sciences (SS).

Downloaded from [Link] at PENNSYLVANIA STATE UNIV on September 18, 2016


432 Language Testing 29(3)

Table 4. Observed and corrected correlations between TOEFL iBT and GPA for undergraduate
students in four course categories

GPA GPA_BU GPA_HA GPA_SE GPA_SS


(n=744) (n=92) (n=435) (n=445) (n=391)

School n robs rcor n robs rcor n robs rcor n robs rcor n robs rcor
A 85 .13 .20 28 .23 .33 78 .10 .15 85 .12 .18 81 .05 .08
B 15 −.33 −.47 - - -   9 .02 .03 10 -.59 -.74 11 .23 .34
C 200 .17 .33 24 .29 .51 189 .03 .06 198 .18 .35 165 .26 .46
D 175 .18 .40 - - - - - - - - - - - -
E 107 .35 .51 - - - - - - - - - - - -
F - - - - - - - - - - - - - - -
G - - - - - - - - - - - - - - -
H 27 .31 .35 - - - 27 .33 .37 27 .29 .33 22 .49 .54
H - - - - - - - - - - - - - - -
J 135 .14 .24 40 -.06 -.10 132 .25 .42 125 .16 .29 112 .34 .54
Mw .18 .33 .12 .19 .13 .20 .15 .27 .25 .41
Note: Course categories are business (BU), humanities and arts (HA), science and engineering (SE), and social
sciences (SS).

graduate level indicated that the predictive power of TOEFL iBT was comparable but
small across BU, HA, and SS, ranging between .24 and .26 (i.e. 6–7% of the variance in
GPA). SE majors showed the lowest average correlation (rw= .17) – that is, about 3%
explained variance in GPA. The weighted average correlation based on the whole group
of graduate students aggregated across the four academic disciplines was rw= .20, indi-
cating that TOEFL iBT could explain about 4% of the variance in GPA.
Similar results were observed at the undergraduate level (Table 4). TOEFL iBT
explained about 3% of the variance in GPA for undergraduate students (rw= .18). The
weighted average correlations between TOEFL iBT and discipline-specific GPAs
ranged between .13 and .25. Although discipline-specific GPAs have the advantage of
grouping courses with similar characteristics, thus allowing us to compare the relation-
ship between TOEFL iBT scores and GPA across disciplines, they are less reliable than
overall GPA because they are based on a smaller number of courses, compared to
overall GPA.
The same data were then analyzed using expectancy graphs. Due to the space limita-
tion, we present two graphs in the main body of the paper – one graph representing a
general trend, and the other graph that is less consistent with the others. Figure 1 was
drawn using the aggregated sample of the graduate students across academic disciplines
and institutions. Group memberships were determined based on the performance data
within an institution (and further within an academic discipline for graduates). The
graphs show what percentage of students within each of three TOEFL iBT groups
received a GPA in each of the three GPA categories (i.e. the bottom 25%, middle 50%,
and top 25% of GPA). The three bars on each graph represent the three TOEFL iBT score

Downloaded from [Link] at PENNSYLVANIA STATE UNIV on September 18, 2016


Cho and Bridgeman 433

Graduate (n=1841)
100%
16%
26%
33%
80%
Top 25% GPA
60% 50%
Mid 50% GPA
48%
40% 51% Boom 25% GPA

20%
34%
26%
16%
0%
Boom 25% iBT Mid 50% iBT Top 25% iBT

Figure 1. Percentage of the graduate students earning top 25%, middle 50%, and bottom 25%
GPA by TOEFL iBT score groups

groups. A percentage within a bar indicates what percentage of students within each
TOEFL iBT group belonged to one of the three GPA groups indicated by different shades.
Even though the average correlation was small (rw=.20), the graph indicates that grad-
uate students with relatively high TOEFL scores earned higher overall GPAs. Figure 1
shows that 34% of the graduate students in the low TOEFL iBT score group received a
GPA in the bottom 25% while 16% earned a GPA in the top 25%. The opposite pattern is
shown in the high TOEFL iBT group; 16% of the graduate students in the high TOEFL
iBT group received a GPA in the bottom 25% range, and 33% received a GPA in the top
25%. These results suggest that there was a much greater chance for students in the high
TOEFL iBT group to earn a top 25% GPA, and also that the chance of earning a bottom
25% GPA decreases substantially for the high TOEFL iBT group. This pattern was
observed in the other expectancy graphs for the graduate subgroups by academic disci-
pline (Appendix A).
Although not shown, all the graphs for the aggregated undergraduate student group
and all the subgroups by disciplines and disciplines-specific GPAs followed the same
pattern shown in Figure 1, except for GPA_BU (Figure 2). Figure 2 shows the relation-
ship between TOEFL iBT scores and GPA in undergraduate-level business courses. It
shows that the middle TOEFL iBT group showed a larger percentage of students per-
forming above the top 25% GPA_BU than the top 25% TOEFL iBT group (46% vs.
30%). However, a comparison of the percentages of students receiving the bottom 25%
GPA indicates that the chance of receiving a bottom 25% GPA in business courses is
much smaller for students with high TOEFL iBT scores than those with low TOEFL iBT
scores. This result, to some extent, supports the positive relationship between TOEFL
iBT and undergraduate GPA in business courses. Furthermore, it should be pointed out
that GPA_BU had a fairly small sample size compared to the others (see Table 5).

Downloaded from [Link] at PENNSYLVANIA STATE UNIV on September 18, 2016


434 Language Testing 29(3)

GPA_BU (n=92)
100%

31% 30%
80%
46%

60% Top 25% GPA_BU


39%
40% 25% 57% Mid 50% GPA_BU

Boom 25% GPA_BU


20%
31% 29%
13%
0%
Boom 25% iBT Mid 50% iBT Top 25% iBT

Figure 2. Percentage of undergraduate students earning top 25%, middle 50%, and bottom 25%
GPAs in business courses (GPA_BU) by TOEFL iBT score groups

Does TOEFL iBT provide additional information beyond what other


admission-related tests can in predicting GPAs? Is the relationship the
same across fields of study?
The adjusted R2 of each regression model and changes in the adjusted R2 between the mod-
els are presented in Tables 5 and 6 for graduate students and Table 7 for undergraduate stu-
dents. Corrections for range restriction could not be made in this section because a covariance
matrix for unrestricted population was not available for multivariate corrections. The results
of the analyses are discussed in terms of the weighted averages of observed R2. It should be
noted that the results were based on extremely small samples in many cases.

Table 5. Adjusted R2 for graduate students with business majors

School n Model 1a Model 2b


A 6 .81 .80
B 10 .40 .32
C 37 −.01 −.04
D − − −
E 20 −.02 .33
F 4 −.49 .89
G 30 −.01 −.05
H 69 .01 .05
I 54 .24 .26
J 15 −.04 .11
Mw 245 .08 .14
aModel 1 – Predictors: (Constant), GMAT
bModel 2 – Predictors: (Constant), GMAT, TOEFL_IBT
Dependent variable: GPA

Downloaded from [Link] at PENNSYLVANIA STATE UNIV on September 18, 2016


Cho and Bridgeman 435

Table 6. Summary of adjusted R2 for graduate students with non-business majors

Major School n Model 1a Model 2b Model 3c


HA A 18 −.03 .07 .07
B − − − −
C 12 .04 .05 .35
D 15 −.06 .26 .25
E 7 −.19 −.48 −.87
F 6 −.03 .15 .89
G 17 .06 .19 .14
H 5 .25 .07 .48
I 19 .00 .11 .12
J − − − −
Mw 99 .00 .09 .16

SE A 115 .00 .01 .00


B 53 .06 .05 .15
C 85 .09 .12 .25
D 55 .06 .04 .02
E 84 .01 .12 .28
F 103 −.01 .02 .07
G 143 .00 .06 .09
H 105 .01 .00 .00
I 48 .02 .07 .13
J 12 .22 .14 .03
Mw 803 .02 .05 .10

SS A 11 .06 .28 .31


B 6 .36 .90 .94
C 24 −.03 .03 −.01
D 49 .00 −.01 .04
E 24 −.04 −.04 .01
F 13 .12 .03 .15
G 44 .01 .00 −.01
H 16 −.02 .12 .08
I 11 .16 .24 .14
J 7 −.17 −.09 −.20
Mw 205 .01 .06 .07
a Model 1 – Predictors: (Constant), GRE_V
b Model 2 – Predictors: (Constant), GRE_V, TOEFL_IBT
c Model 3 – Predictors: (Constant), GRE_V, TOEFL_IBT, GRE_Q

Dependent variable: GPA

For business graduate students, the result of regression model 1 indicated an average
of 8% of the explained variance in GPA (Table 5). When both GMAT and TOEFL iBT
scores were included as predictors, about 14% of the variance in GPA was explained. The
difference in the amount of variance between the models indicated that TOEFL iBT could

Downloaded from [Link] at PENNSYLVANIA STATE UNIV on September 18, 2016


436 Language Testing 29(3)

Table 7. Adjusted R2 for undergraduate students

Variable School n Model 1a Model 2b Model 3c


GPA A 76 −.01 −.02 −.02
C 174 .00 .00 .00
D 156 .01 .04 .10
E 63 .20 .18 .17
H 26 −.04 .12 .13
J 28 −.03 −.05 .02
Mw 523 .02 .03 .05
GPA_BU A 25 .09 .05 .04
C 19 .08 .22 .20
H − − − −
J 13 −.09 .06 .04
Mw 57 .01 .02 .02
GPA_HA A 70 −.01 −.02 −.04
C 165 −.01 −.01 .07
H 26 .00 .08 .06
J 27 .20 .21 .20
Mw 288 .01 .02 .05
GPA_SE A 76 −.01 −.02 .01
C 172 .00 .00 −.01
H 26 −.02 −.01 .25
J 28 −.04 −.07 .07
Mw 302 −.01 −.01 .03
GPA_SS A 74 −.01 −.02 −.02
C 145 .03 .05 .04
H 21 .00 .18 .31
J 23 .20 .33 .35
Mw 263 .03 .06 .07
a Predictors: (Constant), SAT_RW
b Predictors: (Constant), SAT_RW, TOEFL_IBT
c Predictors: (Constant), SAT_RW, TOEFL_IBT, SAT_MATH

Dependent variable: GPA_SS

explain an additional 6% of the variance in GPA for graduate students majoring in busi-
ness. For graduate students in the other academic disciplines, three regression models
were compared, as presented in Table 6. The adjusted R2 of regression model 1 indicated
that GRE verbal scores as a single predictor was not an effective predictor of GPA for
international graduate students in HA, SE, and SS. The differences in adjusted R2 between
regression models 1 and 2 showed that TOEFL iBT scores could account for an additional
9%, 3%, and 5% of the variance in GPA for HA, SE, and SS, respectively. Adding GRE
quantitative scores to the regression equation (model 3) slightly improved the prediction

Downloaded from [Link] at PENNSYLVANIA STATE UNIV on September 18, 2016


Cho and Bridgeman 437

Graduates with GRE Verbal scores


within mid 50% (n=545)
100%
13%
30% 29%
80%

60% 55%
Top 25% GPA
46%
40% 56% Mid 50% GPA
Boom 25% GPA
20%
33%
24%
15%
0%
Boom 25% iBT Mid 50% iBT Top 25% iBT
(n=77) (n=309) (n=159)

Figure 3. Percentage of graduate students in the middle 50% of GRE scores earning top 25%,
middle 50%, and bottom 25% GPA by TOEFL iBT score groups

Undergraduates with SAT RW scores


within mid 50% (n=256)
100%
14% 19% 23%
80%
37%
60%
55% 52% Top 25% GPA
40%
Mid 50% GPA
20% 49% Boom 25% GPA
27% 25%
0%
Boom 25% iBT Mid 50% iBT Top 25% iBT
(n=43) (n=157) (n=56)

Figure 4. Percentage of undergraduate students in the middle 50% of SAT_RW scores earning
top 25%, middle 50%, and bottom 25% GPA by TOEFL iBT score groups

of GPA in the three academic disciplines, but the total amount of the explained variance
was still fairly small: an average of 16% for HA, 10% for SE, and 7% for SS.
At the undergraduate level, combined SAT reading and writing scores (SAT RW)
explained 2% of the variance of overall GPA, and TOEFL iBT added an additional 1%
of the variance (Table 7). Using discipline-specific GPAs as a criterion did not improve
the results for the undergraduate students. In general, SAT_RW as a single predictor
explained a very small amount of the variance in the discipline-specific GPAs, and add-
ing TOEFL iBT scores resulted in little or no improvement in the adjusted R2 (i.e. 0–3%
in the weighted average adjusted R2 change).
The incremental predictive validity of TOEFL iBT was examined graphically using the
expectancy graphs. Due to the space limitation, we present only two graphs: Figure 3 for

Downloaded from [Link] at PENNSYLVANIA STATE UNIV on September 18, 2016


438 Language Testing 29(3)

graduate students and Figure 4 for undergraduate students. Figure 3 shows that among
those students whose GRE_V scores were in the middle 50% group, students were likely
to have a higher GPA when they had a better TOEFL iBT score. The percentage of stu-
dents earning a top 25% GPA was more than twice as large for the students in the top 25%
TOEFL iBT group (29%) as for those in the bottom 25% group (13%). In addition, the
percentage of students earning a bottom 25% GPA decreased to 15% in the top 25%
TOEFL iBT group from 33% in the bottom 25% TOEFL iBT group.
Similar results were observed in Figure 4. Among the undergraduate students whose
SAT_RW were within the middle 50%, the percentage of students receiving a bottom
25% GPA was double for the bottom 25% TOEFL iBT group (49%) than it was for the
top 25% TOEFL iBT group (25%). This pattern was also found in the expectancy graphs
for GPA_SE and GPA_SS (not shown), suggesting that among those whose SAT_RW
scores were not extremely high or low, TOEFL iBT scores could provide additional
unique information in the prediction of GPA earned in SE and SS courses. The pattern
illustrated in Figures 3 and 4 was not always observed in other graphs of subgroups by
disciplines and discipline-specific GPAs. Less consistent results were found in GPA_
HA, and more extreme ones in GPA_BU. Inconsistent patterns seemed partly due to an
extremely small number of students in one of the TOEFL iBT groups.

Conclusions and discussion


This research investigated whether and to what extent TOEFL iBT scores could predict
NNES students’ future academic performance as measured by GPA. Academic records
of 2594 students collected from 10 universities in the United States were analyzed
using correlation-based analyses and expectancy graphs. Overall, the predictive valid-
ity correlation coefficients of TOEFL iBT were fairly small, with the average weighted
correlation being rw = .16 for the group of graduate students, and rw = .18 for under-
graduate students – that is, 3% of the variance of GPA explained for both groups in the
study sample. When the data were analyzed by academic disciplines, the results
improved slightly for some disciplines at the graduate level. The incremental validity
of TOEFL iBT expressed in the amount of explained variance was also shown to be
fairly small. Similarly, none of the admissions tests of interest in the study showed
strong correlations with GPA. GRE verbal and combined SAT reading and writing
scores explained about 2–3% of the variance in GPA, and TOEFL iBT scores accounted
for an additional 3–4%.
Heeding the warning that the predictive validity expressed in terms of correlation is
not easily interpretable and that even a small correlation can indicate a meaningful rela-
tionship (Schrader, 1965; Rosenthal & Rubin, 1982; Sackett, Borneman, & Connelly,
2008), the relationship between TOEFL iBT and GPA was examined also in terms of
probability. Expectancy graphs demonstrated that the chance of being in the top 25%
GPA category doubled for both undergraduate and graduate students when their TOEFL
iBT scores were in the top 25%, in comparison with those in the bottom 25% TOEFL
group. Also, results indicated that the percentage of students receiving a bottom 25%
GPA was considerably smaller for students with relatively high TOEFL scores. There
was also an indication that TOEFL iBT provided additional information beyond what
other admissions tests could offer. Among the students whose combined SAT reading

Downloaded from [Link] at PENNSYLVANIA STATE UNIV on September 18, 2016


Cho and Bridgeman 439

and writing scores were in the middle 50%, the percentage of students receiving a bottom
25% GPA decreased drastically between the bottom and top 25% TOEFL iBT groups
(49% vs. 25%). Similarly, among the students whose GRE verbal scores were in the mid-
dle range, the percentage dropped between the two TOEL iBT subgroups, from 33% to
15%. Although students who scored extremely well or poor on admissions tests had to be
excluded from the analysis using expectancy graphs, the patterns observed in the expec-
tancy graphs provide some evidence for the incremental validity of TOEFL iBT. The
admissions tests such as SAT include measures of verbal ability, but TOEFL iBT still
seems to provide unique and additional information about NNES students’ language
ability, for example, speaking, that is relevant to academic performance.
In summary, in reconciling the results of the correlation-based analyses and expec-
tancy graphs, we believe that even small correlations or seemingly trivial amounts of
variance explained may be an indication of a meaningful relationship between two vari-
ables. The results provide some evidence that TOEFL iBT scores predict the academic
performance of NNES students as measured by GPA.
Some limitations of the study need to be acknowledged. First, because the study
sample was limited to a small number of four-year institutions in the United States, its
findings cannot be generalized to educational contexts that were not represented in the
study and more studies are needed to see whether similar results are observed in data
from different contexts. TOEFL iBT scores are used worldwide by various types of insti-
tutions including community colleges and vocational and professional schools. It would
be prudent to replicate the study in different contexts.
Second, range restriction was still an issue in the study. Without knowing the future
academic outcome of those who are not admitted, thus not included in the study, it is dif-
ficult to accurately assess the relationship between language proficiency and academic
performance.
Furthermore, the lack of variability in GPA, especially at the graduate level, should be
considered in interpreting the results of the expectancy graphs. As shown in the summary
of the cut-points used to create the subgroups for analysis, the different subgroups were
not in fact much different from each other in terms of their performance on those varia-
bles. In many cases, the students assigned to the bottom 25% GPA group in the study
would be considered academically successful. For instance, in one institution, being in
the bottom 25% GPA group meant having a GPA of 3.82 or below. This small difference
may not be meaningful in practice.
Finally, and most importantly, the study would have been strengthened had other indi-
cators of academic performance been included. Some progress has been made in this
regard in other studies. For example, in evaluating the predictive validity of CAEL for
academic performance, Fox (2004) gathered multiple sources of evidence over time to
measure NNES students’ academic performance, including average grades, EAP teacher
evaluation of NNES students’ language ability, attendance rates, and comments from
field-specific professors. Longitudinal studies including other types of criterion informa-
tion, similar to Fox’s study, are needed in the TOEFL context.
Despite these limitations, we believe that the current study contributes to the literature
on the predictive validity research of language tests. This is the first large-scale study that
documents the predictive validity of TOEFL iBT for both graduate and undergraduate stu-
dents from multiple institutions. Research findings previously reported in the literature

Downloaded from [Link] at PENNSYLVANIA STATE UNIV on September 18, 2016


440 Language Testing 29(3)

were based on very small sample sizes or a sample within a single institution, making it
difficult to compare results across studies. The data from multiple institutions in this study
allowed us to show a general pattern concerning the validity of TOEFL iBT in predicting
academic performance. The study is also important methodologically in that it introduces an
alternative way to evaluate the utility of a language test in predicting academic success.

Acknowledgements
This research was funded by the TOEFL program at Educational Testing Service. The authors
would like to thank the participating institutions and individuals who assisted with data collection.
The authors also thank Nan Kong, who assisted with data manipulation. The views expressed in
this publication do not necessarily reflect those of the TOEFL program. The authors are responsi-
ble for all the statements and errors in this publication.

Note
1. TOEFL has so far been administered in three versions: paper-based (PBT), computer-based
(CBT), and Internet-based (iBT) version. Unless noted otherwise, TOEFL without version
indication in the current section refers to the PBT version.

References
Al-Musawi, N. M., & Al-Ansari, S. H. (1999). Test of English as a Foreign Language and First
Certificate of English tests as predictors of academic success for undergraduate students at the
University of Bahrain. System, 27, 389–399.
Ayers, J. B., & Quattlebaum, R. F. (1992). TOEFL performance and success in a master’s program
in engineering. Educational and Psychological Measurement, 52, 973–975.
Chalhoub-Deville, M., & Deville, C. (2006). Old, borrowed, and new thoughts in second language
testing. In R. L. Brennan (Ed.), Educational measurement (4th ed.). Westport, CT: Praeger.
Fox, J. (2004). Test decisions over time: Tracking validity. Language Testing, 21(4), 437–465.
Graham, J. G. (1987). English language proficiency and the prediction of academic success.
TESOL Quarterly, 21(2), 505–521.
Hartnett, R. T., & Willingham. W. W. (1980). The criterion problem: What measure of success in
graduate education? Applied Psychological Measurement, 4(3), 281–291.
Heil, D. K., & Aleamoni, L. M. (1974). Assessment of the proficiency in the use and understand-
ing of English by foreign students as measured by the Test of English as a Foreign Language.
(ERIC Document Reproduction Service No. ED 093 948).
Institute for International Education (2007). Open doors: Report on international educational
exchange. Retrieved October 23, 2007 from [Link]
Kuncel, N. R., Crede, M., & Thomas, L. L. (2007). A meta-analysis of the predictive validity of
the Graduate Management Admission Test (GMAT) and undergraduate grade point average
(UGPA) for graduate student academic performance. Academy of Management Learning &
Education, 6(1), 51–68.
Neal, M. E. (1998). The predictive validity of the GRE and TOFEL exams with GGPA as the
criterion of graduate success for international graduate students in science and engineering.
(ERIC Document Reproduction Service No. ED424294).
Powers, D. E., Kim, H.-J., & Weng, V. Z. (2008). The redesigned TOEIC (listening and reading)
test: Relations to test-taker perceptions of proficiency in English (ETS Research Report no.
RR-08-56). Princeton, NJ: ETS.
Ree, M. J., Carretta, T. R., Earles, J. A., & Albert, W. (1994). Sign changes when correction for
range restriction: A note on Pearson’s and Lawley’s selection formula. Journal of Applied
Psychology, 79(2), 298–301.

Downloaded from [Link] at PENNSYLVANIA STATE UNIV on September 18, 2016


Cho and Bridgeman 441

Rosenthal, R., & Rubin, D. B. (1982). A simple, general purpose display of magnitude of experi-
mental effect. Journal of Educational Psychology, 74, 166–169.
Ross, S. (1998). Self-assessment in second language testing: A meta-analysis and analysis of expe-
riential factors. Language Testing, 15(1), 1–20.
Sackett, P. R., Borneman, M. J., & Connelly, B. S. (2008). High-stakes testing in higher education
and employment: Appraising the evidence for validity and fairness. American Psychologist,
63(4), 215–227.
Sackett, P. R., & Yang, H. (2000). Correction for range restriction: An expanded typology. Jour-
nal of Applied Psychology, 85(1), 112–118.
Schrader, W. B. (1965). A taxonomy of expectancy tables. Journal of Educational Measurement,
2, 29–35.
Sinharay, S., Powers, D. E., Feng, Y., Saldivia, L., Giunta, A., Simpson, A., & Weng, V. (2009).
Appropriateness of the TOEIC® Bridge test for students in three countries of South America.
Language Testing, 26(4), 589–619.
Van Nelson, C., Nelson, J. S., & Malone, B. G. (2004). Predicting success of international graduate
students in an academic university. College and University Journal, 80(1), 19–27.
Vinke, A. A., & Jochems, W. M. G. (1993). English proficiency and academic success in interna-
tional postgraduate education. Higher Education, 26, 275–285.
Wiberg, M., & Sundstrom, A. (2009). A comparison of two approaches to correction of restriction
of range in correlation analysis. Practical Assessment, Research & Evaluation, 14(5), 1–9.
Woodrow, L. (2006). Academic success of international postgraduate education students and the
role of English proficiency. University of Sydney Papers in TESOL, 1, 51–70.
Zwick, R. (2002). Fair game? The use of standardized admissions tests in higher education.
New York: RoutledgeFalmer.

Appendix A: Expectancy graphs for the graduate subgroups by academic disciplines


(percentage of the graduate students earning top 25%, middle 50%, and bottom 25%
GPA by TOEFL iBT score groups)

BU (n=413)
100%
15% 22%
80% 37%
60% 50% Top 25% GPA
51%
40% 47% Mid 50% GPA
20% 36% Boom 25% GPA
27% 16%
0%
Boom 25% Mid 50% Top 25%
iBT iBT iBT

HA (n=186)
100%
13%
80% 28% 28%
60% 51% Top 25% GPA
46%
40% 61% Mid 50% GPA
20% 36% Boom 25% GPA
26%
0% 11%
Boom 25% Mid 50% Top 25%
iBT iBT iBT

Downloaded from [Link] at PENNSYLVANIA STATE UNIV on September 18, 2016


442 Language Testing 29(3)

SE (n=959)
100%
17% 26%
80% 33%
60% 51% Top 25% GPA
48%
40% 49% Mid 50% GPA
20% 32% Boom 25% GPA
26% 18%
0%
Boom 25% Mid 50% Top 25%
iBT iBT iBT

SS (n=283)
100%
18% 27%
80% 30%

60% 47% Top 25% GPA


48%
40% 58% Mid 50% GPA
20% 35% Boom 25% GPA
25%
0% 12%
Boom 25% Mid 50% Top 25%
iBT iBT iBT

Figures A1–A4

Downloaded from [Link] at PENNSYLVANIA STATE UNIV on September 18, 2016

You might also like