TOEFL iBT Scores and Academic Success
TOEFL iBT Scores and Academic Success
2012
LTR29310.1177/0265532211430368Cho and BridgemanLanguage Testing
/$1*8$*(
Article 7(67,1*
Language Testing
Relationship of TOEFL
29(3) 421–442
© The Author(s) 2012
Reprints and permission:
iBT® scores to academic [Link]/[Link]
DOI: 10.1177/0265532211430368
performance: Some evidence [Link]
Abstract
This study examined the relationship between scores on the TOEFL Internet-Based Test
(TOEFL iBT®) and academic performance in higher education, defined here in terms of grade
point average (GPA). The academic records for 2594 undergraduate and graduate students were
collected from 10 universities in the United States. The data consisted of students’ GPA, detailed
course information, and admissions-related test scores including TOEFL iBT, GRE, GMAT, and
SAT scores. Correlation-based analyses were conducted for subgroups by academic status and
disciplines. Expectancy graphs were also used to complement the correlation-based analyses by
presenting the predictive validity in terms of individuals in one of the TOEFL iBT score subgroups
belonging to one of the GPA subgroups. The predictive validity expressed in terms of correlation
did not appear to be strong. Nevertheless, the general pattern shown in the expectancy graphs
indicated that students with higher TOEFL iBT scores tended to earn higher GPAs and that the
TOEFL iBT provided information about the future academic performance of non-native English
speaking students beyond that provided by other admissions tests. These observations led us to
conclude that even a small correlation might indicate a meaningful relationship between TOEFL
iBT scores and GPA. Limitations and implications are discussed.
Keywords
academic performance, English proficiency, grade point average, international students, predictive
validity, TOEFL
The Test of English as a Foreign Language (TOEFL®) is one of the English proficiency
tests that non-native English speaking (NNES) students may take to demonstrate their
language proficiency when applying to English-medium colleges and universities.
TOEFL scores or similar information are often considered by many institutions in
Corresponding author:
Yeonsuk Cho, Educational Testing Service, Rosedale Road, MS 07-R, Princeton, NJ 08541, USA
Email: ycho@[Link]
studies are in contrast with the result of a study by Ayers and Peters (1977), as cited in
Ayers and Quattlebaum (1992), in which a much larger correlation was observed for a
sample of 50 Asian graduate students in physical science or engineering programs (r =
.40, p < .01). These three studies focused on the predictive validity of TOEFL for stu-
dents in master’s-level science or engineering fields, but like the earlier studies summa-
rized in Graham (1987), the results from these three studies are inconsistent.
The value of TOEFL in predicting academic performance has also been investigated
in contexts outside North America. Vinke and Jochems (1993) examined the relationship
between TOEFL and academic success, using a sample of 90 Indonesian students who
were studying engineering in a postgraduate program in the Netherlands where English
was the language of instruction. One of the criteria for academic success were the aver-
age scores on seven written qualifying exams required as part of the degree requirement.
TOEFL scores showed a moderate correlation with the average exam scores (r = .51).
Even stronger evidence was found when comparing the passing rates for the qualifying
exam of two TOEFL subgroups in the same study: 74% for students with TOEFL scores
below 450, compared to 98% for those with TOEFL scores above 450.
Al-Musawi and Al-Ansari (1999) examined the predictive validity of TOEFL in com-
parison with that of the First Certificate in English (FCE). FCE is one of the English
proficiency examinations developed by the University of Cambridge ESOL examina-
tions. The test is claimed to be at the level B2 of the Council of Europe Common
European Framework of Reference for Language (CEFR), and it measures four skills
and explicit knowledge of grammar and vocabulary. The study included the academic
records of 86 English major undergraduate students in Bahrain. Sub-scores on both
TOEFL and FCE were used as predictor variables, and overall GPA and GPA in English
courses were used as measures of academic success in a stepwise regression. Despite
some moderate correlations with overall GPA, none of the TOEFL sub-scores contrib-
uted to the prediction of overall GPA when the FCE scores were already entered into the
equation. Similar results were observed in the prediction of average grades earned in
English courses. Based on the results, the researchers concluded that FCE might be a
better predictor of academic success of language learners in EFL contexts than TOEFL.
However, their interpretation should be taken in light of the analytical approach used in
the study. Because both FCE and TOEFL are measures of English proficiency, the sub-
scores on the two tests would have been redundant, creating a collinearity problem in the
analysis. It is known that the selection of variables in a stepwise regression is affected by
the correlations among the predictors. Also, results of a stepwise regression from one
sample are hard to replicate in another sample.
More recently, using logistic regression, Van Nelson, Nelson, and Malone (2004) tried
to determine which combination of variables, including TOEFL, could predict the aca-
demic success of international students. The data came from 866 students in master’s-
degree programs across various academic fields. Both final GPA and degree completion
were used as criterion variables. For the logistic regression analysis, GPA was dichoto-
mized by dividing the students into two groups – a group of students with a GPA of 3.5
and above, and the other with students who received a GPA below 3.5. The results were
mixed; when degree completion was used as a criterion, TOEFL was not a good predictor
of academic success. However, it did contribute to the prediction of the final GPA.
influenced by the faculty members’ own experiences, such as their individual fields or
prior exposure to NNES students. Similarly, as noted by Ross (1998), respondents’ experi-
ences and variation in the interpretation of a self-assessment question affect how they
answer questions. This is particularly an issue when respondents are from diverse back-
grounds and cultures, as is the case with the population of TOEFL test takers.
Whether the criteria are appropriate or not, range restriction is mentioned in almost
every predictive validity study because the validity evidence is typically expressed in
terms of correlation, and when the data do not represent a full range – that is, there is no
information regarding how those who did not get selected would have performed after-
ward – a correlation is underestimated. Although some statistical adjustment can be
made to correct correlations for range restriction (Ree, Carretta, Earles, & Albert, 1994;
Sackett & Yang, 2000; Wiberg & Sundstrom, 2009), some issues have been noted with
correction methods. For example, possible sign changes after correction, and specifica-
tion errors in selecting an appropriate correction formula can occur (Ree et al., 1994).
The lack of criterion-related information for those not included in a study is a problem in
designing and interpreting predictive validity studies when the real question of interest is
how performance on a test relates to the future performance of all test takers.
A number of researchers expressed the view that correlation is inherently difficult to
interpret even for experienced social scientists (Schrader, 1965; Rosenthal & Rubin,
1982; Sackett, Borneman, & Connelly, 2008). Recognizing these drawbacks of using
correlations as a validity index in predictive validity research, Schrader (1965) suggested
using an expectancy table as a concrete way to represent test validity. An expectancy
table is a simple way to summarize ‘the relation between two (or more) variables by stat-
ing the probability that individuals who belong to each of a set of subgroups defined on
the basis of one (or more) variables will belong to each of a set of subgroups defined on
the basis of another variable’ (Schrader, 1965, p. 29). One advantage of this approach is
that it is easy to explain results to a lay audience. More recently, Sackett et al. (2008)
stated that the magnitude of even small correlation coefficients is not well understood,
and pointed out the benefit of a similar approach: ‘converting correlations to differences
in odds of success results both in a readily interpretable metric and in a positive picture
… in short, there is a long history of expressing the value of a test in a metric more read-
ily interpretable than percentage of variance accounted for’ (p. 216).
The current study investigated the predictive validity of TOEFL iBT for academic
performance with both correlation-based analyses and expectancy graphs as analytic
methods. In addition, because the reading and writing skills assessed in the TOEFL iBT
overlap somewhat with the skills assessed by admissions tests that are already required
(e.g. SAT and GRE), admissions officers may be interested in finding out whether
TOEFL iBT provides unique information that cannot be obtained from the other admis-
sions tests. We thus asked the following two research questions:
1. What is the relationship between TOEFL iBT scores and future academic perfor-
mance as defined by GPAs? Is the relationship the same across fields of study?
2. Does TOEFL iBT provide additional information beyond what other admissions-
related tests can in predicting GPAs? Is the relationship the same across fields of
study?
Method
Data source
Initially, 40 institutions in the United States with large numbers of international students
according to Open Doors Online ([Link] were contacted for
participation in the study. The following information was requested for full-time, degree-
seeking, non-transfer students who had TOEFL iBT scores and at least a first-year GPA
from the current institution:
Ten universities agreed to participate in the study. Seven schools provided informa-
tion for both undergraduate and graduate students whereas the other three sent data
only for graduate students. Course-specific grades were available for four schools.
Geographically, half of the schools were in the Midwest, and the others were in various
regions. Except for one private university, all were large public universities. According
to the admission selectivity published in U.S. News, seven schools in our study sample
were regarded as ‘more selective,’ one ‘most selective,’ one ‘selective,’ and one ‘less
selective.’ A total of 2594 students – 1850 graduate students and 744 undergraduate stu-
dents – were represented in the data.
Variables
The following variables were included in the study.
a. GPA: An overall GPA. Types of GPA information varied across schools: some
schools sent cumulative GPAs while others provided year-by-year GPAs or
course information that allowed the calculation of a GPA. The distinction among
the different types of GPA information was, not made, however.
b. GPA_BU, GPA_HA, GPA_SE, and GPA_SS: Discipline-specific GPAs are
weighted average grades earned in courses in the same academic category. They
were computed only for undergraduate students whose individual course
information was available. During the subgroup analysis by majors, the discipline-
specific GPAs were used as a criterion instead of overall GPAs.
c. TOEFL_IBT: A total TOEFL iBT score.
d. SAT_RW: A combined score from the SAT reading and writing sections for under-
graduate students.
e. SAT_M: A SAT math score.
f. GRE_V: A GRE verbal score for non-business graduate students. GRE also
reports a writing score separate from a verbal score. Although it was preferable to
use a combined score of verbal and writing scores, GRE_V was used because
writing scores were missing for many students in the study sample.
g. GRE_Q: A GRE quantitative score.
h. GMAT: A total GMAT score for business graduate students. GMAT reports sub-
scores for the verbal and math sections. However, the sub-score information was
not available for many students in the study sample.
Analyses
Predictive validity was examined in terms of correlation and probability. Correlation and
hierarchical multiple regression analyses were performed. These analyses were then
complemented by expectancy graphs. Because of obvious differences in grading stand-
ards and test scores, analyses were conducted separately for academic levels, different
disciplines and institutions. Results were then aggregated across institutions.
Correlation-based analyses
Simple correlations were computed to address Research Question 1. To answer Research
Question 2, SAT_RW was entered as a single predictor of GPA or disciple-specific GPA
(Model 1), and TOEFL iBT was then added (Model 2) for undergraduate students.
Furthermore, although it was not the focus of our research, because quantitative skills
are considered an important element in academic performance and measured accord-
ingly in many admissions tests, SAT_M was subsequently added (Model 3). The regres-
sion analyses were done in a similar manner for graduate students. However, for
business graduate students, only two regression models were compared because sub-
scores of the GMAT were not available. A difference in squared multiple correlation
(R2) between the first two regression models was used as an index of incremental valid-
ity of TOEFL iBT – that is, whether TOEFL iBT provides unique information beyond
the language measures of other admission-related tests. In running the hierarchical mul-
tiple regression models, collinearity among the predictors was checked by examining
the two collinearity indices, tolerance and variation-inflation factor (VIF), and no sig-
nificant collinearity was present.
A weighted average of simple correlations and squared multiple correlations (R2)
was then used to aggregate the results of correlation-based analyses across institu-
tions. Sample sizes, which varied across disciplines and institutions, were used to
give more weight to the results based on the larger samples. In order to avoid the capi-
talization of chance factors that could artificially influence regression estimates in
small samples, an adjusted R2 was computed within each institution prior to comput-
ing a weighted average.
Expectancy graphs
For expectancy graphs, the students in the study were first assigned to one of three
subgroups (i.e. top 25%, middle 50%, and bottom 25%) according to their relative
standing on GPA and test scores within institution and within academic status (and
further within academic discipline for graduate students). This was done to accommo-
date to different academic standards across academic levels, disciplines, and schools.
For example, a GPA subgroup that a graduate student with a computer science major
belonged to was determined with respect to graduate peers in science and engineering
(SE) majors within the same university. Academic discipline, however, was not taken
into consideration in assigning undergraduate students to groups because (1) for many
undergraduate students in the study, majors were unknown, but more importantly (2)
their first years of university education were less likely to reflect academic perfor-
mance within a single discipline.
Once the subgroups were created, expectancy graphs were drawn by cross-tabulating
TOEFL iBT and GPA subgroups for Research Question 1. Expectancy graphs were also
used to address Research Question 2, but with some modification. In this part of the
analysis, the incremental validity of TOEFL iBT was examined by focusing on students
whose performance on other admissions tests (i.e. GRE_V, GMAT, or SAT_RW) was in
the middle 50% of a score distribution – that is, those in the top and bottom 25% of
GRE_V were not included in the analysis of expectancy graphs. This was done for both
conceptual and logistical reasons.
Test takers who receive extremely low or high scores on TOEFL iBT are also likely
to receive extremely low or high scores on the language measure of an admissions test,
whereas there is more uncertainty in the middle score range – that is, the relationship of
two similar tests is less certain when you look at the test scores in the middle range. Thus,
TOEFL iBT scores might be particularly useful in predicting the academic performance
of students whose GRE verbal scores are in the middle score range. Furthermore, because
of extremely rare occurrences of students who did very well on GRE_V or SAT_RW but
who received very low scores on TOEFL iBT, it was not practically possible to demon-
strate the incremental validity of TOEFL iBT for students in the top and bottom 25%
GRE_V or SAT_RW score groups. Therefore, expectancy graphs were drawn using the
students whose test performance on GRE_V or SAT_RW belonged to the middle 50%
group. To be parallel to this approach, the students in the top 25% and bottom 25%
GMAT groups were also excluded from expectancy graphs.
Results
Descriptive statistics
The data included the academic records of a total of 2594 students consisting of 1850
graduate students and 744 undergraduate students. Table 1 presents the summary
statistics for GPA and test scores by individual school. As expected in a restricted sample
of enrolled students, the average TOEFL scores were higher and the standard deviations
were smaller than those reported in the summary statistics of all examinees ([Link]
org/toefl/research/test_score_data_summary). On average, the graduate students had
higher TOEFL iBT scores than the undergraduate students, and this was true for all insti-
tutions. In addition, the summary statistics showed that graduate students tend to receive
higher grades than undergraduate students. A comparison of the statistics indicated that
there was little variability across schools in terms of average GPA. Average test scores
varied somewhat across schools.
With respect to academic discipline, science and engineering (SE) majors were most
popular (47%) among the graduate students, followed by business (BU) majors (26%),
social sciences (SS) majors (16%), and humanities and arts (HA) majors (10%). SE
majors were also popular among the undergraduates (26%). About 23% of the under-
graduate students had HA majors. However, majors were unknown or not declared for
44% of the undergraduate students.
The cut-points for the subgroups are summarized in Table 2. The cut-points varied
considerably between undergraduate and graduate students, and across academic
Table 2. Ranges of the cut-scores for the subgroups based on test scores and GPAs
Group
disciplines and institutions. For example, in one school, a graduate business student
whose GPA was below 3.68 was assigned to the bottom 25% GPA group, whereas a GPA
of 3.59 or above was considered the top 25% performance in another school. We would
like to emphasize at this point that being in the bottom 25% group in the current study
did not necessarily constitute academic failure. In fact, as these cut-scores clearly indi-
cate, students in adjacent subgroups were not much different in terms of ability being
considered.
Table 3. Observed and corrected correlations between TOEFL iBT and GPA for graduate
students in four major categories
School BU HA SE SS
(n=413) (n=186) (n=959) (n=283)
Table 4. Observed and corrected correlations between TOEFL iBT and GPA for undergraduate
students in four course categories
School n robs rcor n robs rcor n robs rcor n robs rcor n robs rcor
A 85 .13 .20 28 .23 .33 78 .10 .15 85 .12 .18 81 .05 .08
B 15 −.33 −.47 - - - 9 .02 .03 10 -.59 -.74 11 .23 .34
C 200 .17 .33 24 .29 .51 189 .03 .06 198 .18 .35 165 .26 .46
D 175 .18 .40 - - - - - - - - - - - -
E 107 .35 .51 - - - - - - - - - - - -
F - - - - - - - - - - - - - - -
G - - - - - - - - - - - - - - -
H 27 .31 .35 - - - 27 .33 .37 27 .29 .33 22 .49 .54
H - - - - - - - - - - - - - - -
J 135 .14 .24 40 -.06 -.10 132 .25 .42 125 .16 .29 112 .34 .54
Mw .18 .33 .12 .19 .13 .20 .15 .27 .25 .41
Note: Course categories are business (BU), humanities and arts (HA), science and engineering (SE), and social
sciences (SS).
graduate level indicated that the predictive power of TOEFL iBT was comparable but
small across BU, HA, and SS, ranging between .24 and .26 (i.e. 6–7% of the variance in
GPA). SE majors showed the lowest average correlation (rw= .17) – that is, about 3%
explained variance in GPA. The weighted average correlation based on the whole group
of graduate students aggregated across the four academic disciplines was rw= .20, indi-
cating that TOEFL iBT could explain about 4% of the variance in GPA.
Similar results were observed at the undergraduate level (Table 4). TOEFL iBT
explained about 3% of the variance in GPA for undergraduate students (rw= .18). The
weighted average correlations between TOEFL iBT and discipline-specific GPAs
ranged between .13 and .25. Although discipline-specific GPAs have the advantage of
grouping courses with similar characteristics, thus allowing us to compare the relation-
ship between TOEFL iBT scores and GPA across disciplines, they are less reliable than
overall GPA because they are based on a smaller number of courses, compared to
overall GPA.
The same data were then analyzed using expectancy graphs. Due to the space limita-
tion, we present two graphs in the main body of the paper – one graph representing a
general trend, and the other graph that is less consistent with the others. Figure 1 was
drawn using the aggregated sample of the graduate students across academic disciplines
and institutions. Group memberships were determined based on the performance data
within an institution (and further within an academic discipline for graduates). The
graphs show what percentage of students within each of three TOEFL iBT groups
received a GPA in each of the three GPA categories (i.e. the bottom 25%, middle 50%,
and top 25% of GPA). The three bars on each graph represent the three TOEFL iBT score
Graduate (n=1841)
100%
16%
26%
33%
80%
Top 25% GPA
60% 50%
Mid 50% GPA
48%
40% 51% Boom 25% GPA
20%
34%
26%
16%
0%
Boom 25% iBT Mid 50% iBT Top 25% iBT
Figure 1. Percentage of the graduate students earning top 25%, middle 50%, and bottom 25%
GPA by TOEFL iBT score groups
groups. A percentage within a bar indicates what percentage of students within each
TOEFL iBT group belonged to one of the three GPA groups indicated by different shades.
Even though the average correlation was small (rw=.20), the graph indicates that grad-
uate students with relatively high TOEFL scores earned higher overall GPAs. Figure 1
shows that 34% of the graduate students in the low TOEFL iBT score group received a
GPA in the bottom 25% while 16% earned a GPA in the top 25%. The opposite pattern is
shown in the high TOEFL iBT group; 16% of the graduate students in the high TOEFL
iBT group received a GPA in the bottom 25% range, and 33% received a GPA in the top
25%. These results suggest that there was a much greater chance for students in the high
TOEFL iBT group to earn a top 25% GPA, and also that the chance of earning a bottom
25% GPA decreases substantially for the high TOEFL iBT group. This pattern was
observed in the other expectancy graphs for the graduate subgroups by academic disci-
pline (Appendix A).
Although not shown, all the graphs for the aggregated undergraduate student group
and all the subgroups by disciplines and disciplines-specific GPAs followed the same
pattern shown in Figure 1, except for GPA_BU (Figure 2). Figure 2 shows the relation-
ship between TOEFL iBT scores and GPA in undergraduate-level business courses. It
shows that the middle TOEFL iBT group showed a larger percentage of students per-
forming above the top 25% GPA_BU than the top 25% TOEFL iBT group (46% vs.
30%). However, a comparison of the percentages of students receiving the bottom 25%
GPA indicates that the chance of receiving a bottom 25% GPA in business courses is
much smaller for students with high TOEFL iBT scores than those with low TOEFL iBT
scores. This result, to some extent, supports the positive relationship between TOEFL
iBT and undergraduate GPA in business courses. Furthermore, it should be pointed out
that GPA_BU had a fairly small sample size compared to the others (see Table 5).
GPA_BU (n=92)
100%
31% 30%
80%
46%
Figure 2. Percentage of undergraduate students earning top 25%, middle 50%, and bottom 25%
GPAs in business courses (GPA_BU) by TOEFL iBT score groups
For business graduate students, the result of regression model 1 indicated an average
of 8% of the explained variance in GPA (Table 5). When both GMAT and TOEFL iBT
scores were included as predictors, about 14% of the variance in GPA was explained. The
difference in the amount of variance between the models indicated that TOEFL iBT could
explain an additional 6% of the variance in GPA for graduate students majoring in busi-
ness. For graduate students in the other academic disciplines, three regression models
were compared, as presented in Table 6. The adjusted R2 of regression model 1 indicated
that GRE verbal scores as a single predictor was not an effective predictor of GPA for
international graduate students in HA, SE, and SS. The differences in adjusted R2 between
regression models 1 and 2 showed that TOEFL iBT scores could account for an additional
9%, 3%, and 5% of the variance in GPA for HA, SE, and SS, respectively. Adding GRE
quantitative scores to the regression equation (model 3) slightly improved the prediction
60% 55%
Top 25% GPA
46%
40% 56% Mid 50% GPA
Boom 25% GPA
20%
33%
24%
15%
0%
Boom 25% iBT Mid 50% iBT Top 25% iBT
(n=77) (n=309) (n=159)
Figure 3. Percentage of graduate students in the middle 50% of GRE scores earning top 25%,
middle 50%, and bottom 25% GPA by TOEFL iBT score groups
Figure 4. Percentage of undergraduate students in the middle 50% of SAT_RW scores earning
top 25%, middle 50%, and bottom 25% GPA by TOEFL iBT score groups
of GPA in the three academic disciplines, but the total amount of the explained variance
was still fairly small: an average of 16% for HA, 10% for SE, and 7% for SS.
At the undergraduate level, combined SAT reading and writing scores (SAT RW)
explained 2% of the variance of overall GPA, and TOEFL iBT added an additional 1%
of the variance (Table 7). Using discipline-specific GPAs as a criterion did not improve
the results for the undergraduate students. In general, SAT_RW as a single predictor
explained a very small amount of the variance in the discipline-specific GPAs, and add-
ing TOEFL iBT scores resulted in little or no improvement in the adjusted R2 (i.e. 0–3%
in the weighted average adjusted R2 change).
The incremental predictive validity of TOEFL iBT was examined graphically using the
expectancy graphs. Due to the space limitation, we present only two graphs: Figure 3 for
graduate students and Figure 4 for undergraduate students. Figure 3 shows that among
those students whose GRE_V scores were in the middle 50% group, students were likely
to have a higher GPA when they had a better TOEFL iBT score. The percentage of stu-
dents earning a top 25% GPA was more than twice as large for the students in the top 25%
TOEFL iBT group (29%) as for those in the bottom 25% group (13%). In addition, the
percentage of students earning a bottom 25% GPA decreased to 15% in the top 25%
TOEFL iBT group from 33% in the bottom 25% TOEFL iBT group.
Similar results were observed in Figure 4. Among the undergraduate students whose
SAT_RW were within the middle 50%, the percentage of students receiving a bottom
25% GPA was double for the bottom 25% TOEFL iBT group (49%) than it was for the
top 25% TOEFL iBT group (25%). This pattern was also found in the expectancy graphs
for GPA_SE and GPA_SS (not shown), suggesting that among those whose SAT_RW
scores were not extremely high or low, TOEFL iBT scores could provide additional
unique information in the prediction of GPA earned in SE and SS courses. The pattern
illustrated in Figures 3 and 4 was not always observed in other graphs of subgroups by
disciplines and discipline-specific GPAs. Less consistent results were found in GPA_
HA, and more extreme ones in GPA_BU. Inconsistent patterns seemed partly due to an
extremely small number of students in one of the TOEFL iBT groups.
and writing scores were in the middle 50%, the percentage of students receiving a bottom
25% GPA decreased drastically between the bottom and top 25% TOEFL iBT groups
(49% vs. 25%). Similarly, among the students whose GRE verbal scores were in the mid-
dle range, the percentage dropped between the two TOEL iBT subgroups, from 33% to
15%. Although students who scored extremely well or poor on admissions tests had to be
excluded from the analysis using expectancy graphs, the patterns observed in the expec-
tancy graphs provide some evidence for the incremental validity of TOEFL iBT. The
admissions tests such as SAT include measures of verbal ability, but TOEFL iBT still
seems to provide unique and additional information about NNES students’ language
ability, for example, speaking, that is relevant to academic performance.
In summary, in reconciling the results of the correlation-based analyses and expec-
tancy graphs, we believe that even small correlations or seemingly trivial amounts of
variance explained may be an indication of a meaningful relationship between two vari-
ables. The results provide some evidence that TOEFL iBT scores predict the academic
performance of NNES students as measured by GPA.
Some limitations of the study need to be acknowledged. First, because the study
sample was limited to a small number of four-year institutions in the United States, its
findings cannot be generalized to educational contexts that were not represented in the
study and more studies are needed to see whether similar results are observed in data
from different contexts. TOEFL iBT scores are used worldwide by various types of insti-
tutions including community colleges and vocational and professional schools. It would
be prudent to replicate the study in different contexts.
Second, range restriction was still an issue in the study. Without knowing the future
academic outcome of those who are not admitted, thus not included in the study, it is dif-
ficult to accurately assess the relationship between language proficiency and academic
performance.
Furthermore, the lack of variability in GPA, especially at the graduate level, should be
considered in interpreting the results of the expectancy graphs. As shown in the summary
of the cut-points used to create the subgroups for analysis, the different subgroups were
not in fact much different from each other in terms of their performance on those varia-
bles. In many cases, the students assigned to the bottom 25% GPA group in the study
would be considered academically successful. For instance, in one institution, being in
the bottom 25% GPA group meant having a GPA of 3.82 or below. This small difference
may not be meaningful in practice.
Finally, and most importantly, the study would have been strengthened had other indi-
cators of academic performance been included. Some progress has been made in this
regard in other studies. For example, in evaluating the predictive validity of CAEL for
academic performance, Fox (2004) gathered multiple sources of evidence over time to
measure NNES students’ academic performance, including average grades, EAP teacher
evaluation of NNES students’ language ability, attendance rates, and comments from
field-specific professors. Longitudinal studies including other types of criterion informa-
tion, similar to Fox’s study, are needed in the TOEFL context.
Despite these limitations, we believe that the current study contributes to the literature
on the predictive validity research of language tests. This is the first large-scale study that
documents the predictive validity of TOEFL iBT for both graduate and undergraduate stu-
dents from multiple institutions. Research findings previously reported in the literature
were based on very small sample sizes or a sample within a single institution, making it
difficult to compare results across studies. The data from multiple institutions in this study
allowed us to show a general pattern concerning the validity of TOEFL iBT in predicting
academic performance. The study is also important methodologically in that it introduces an
alternative way to evaluate the utility of a language test in predicting academic success.
Acknowledgements
This research was funded by the TOEFL program at Educational Testing Service. The authors
would like to thank the participating institutions and individuals who assisted with data collection.
The authors also thank Nan Kong, who assisted with data manipulation. The views expressed in
this publication do not necessarily reflect those of the TOEFL program. The authors are responsi-
ble for all the statements and errors in this publication.
Note
1. TOEFL has so far been administered in three versions: paper-based (PBT), computer-based
(CBT), and Internet-based (iBT) version. Unless noted otherwise, TOEFL without version
indication in the current section refers to the PBT version.
References
Al-Musawi, N. M., & Al-Ansari, S. H. (1999). Test of English as a Foreign Language and First
Certificate of English tests as predictors of academic success for undergraduate students at the
University of Bahrain. System, 27, 389–399.
Ayers, J. B., & Quattlebaum, R. F. (1992). TOEFL performance and success in a master’s program
in engineering. Educational and Psychological Measurement, 52, 973–975.
Chalhoub-Deville, M., & Deville, C. (2006). Old, borrowed, and new thoughts in second language
testing. In R. L. Brennan (Ed.), Educational measurement (4th ed.). Westport, CT: Praeger.
Fox, J. (2004). Test decisions over time: Tracking validity. Language Testing, 21(4), 437–465.
Graham, J. G. (1987). English language proficiency and the prediction of academic success.
TESOL Quarterly, 21(2), 505–521.
Hartnett, R. T., & Willingham. W. W. (1980). The criterion problem: What measure of success in
graduate education? Applied Psychological Measurement, 4(3), 281–291.
Heil, D. K., & Aleamoni, L. M. (1974). Assessment of the proficiency in the use and understand-
ing of English by foreign students as measured by the Test of English as a Foreign Language.
(ERIC Document Reproduction Service No. ED 093 948).
Institute for International Education (2007). Open doors: Report on international educational
exchange. Retrieved October 23, 2007 from [Link]
Kuncel, N. R., Crede, M., & Thomas, L. L. (2007). A meta-analysis of the predictive validity of
the Graduate Management Admission Test (GMAT) and undergraduate grade point average
(UGPA) for graduate student academic performance. Academy of Management Learning &
Education, 6(1), 51–68.
Neal, M. E. (1998). The predictive validity of the GRE and TOFEL exams with GGPA as the
criterion of graduate success for international graduate students in science and engineering.
(ERIC Document Reproduction Service No. ED424294).
Powers, D. E., Kim, H.-J., & Weng, V. Z. (2008). The redesigned TOEIC (listening and reading)
test: Relations to test-taker perceptions of proficiency in English (ETS Research Report no.
RR-08-56). Princeton, NJ: ETS.
Ree, M. J., Carretta, T. R., Earles, J. A., & Albert, W. (1994). Sign changes when correction for
range restriction: A note on Pearson’s and Lawley’s selection formula. Journal of Applied
Psychology, 79(2), 298–301.
Rosenthal, R., & Rubin, D. B. (1982). A simple, general purpose display of magnitude of experi-
mental effect. Journal of Educational Psychology, 74, 166–169.
Ross, S. (1998). Self-assessment in second language testing: A meta-analysis and analysis of expe-
riential factors. Language Testing, 15(1), 1–20.
Sackett, P. R., Borneman, M. J., & Connelly, B. S. (2008). High-stakes testing in higher education
and employment: Appraising the evidence for validity and fairness. American Psychologist,
63(4), 215–227.
Sackett, P. R., & Yang, H. (2000). Correction for range restriction: An expanded typology. Jour-
nal of Applied Psychology, 85(1), 112–118.
Schrader, W. B. (1965). A taxonomy of expectancy tables. Journal of Educational Measurement,
2, 29–35.
Sinharay, S., Powers, D. E., Feng, Y., Saldivia, L., Giunta, A., Simpson, A., & Weng, V. (2009).
Appropriateness of the TOEIC® Bridge test for students in three countries of South America.
Language Testing, 26(4), 589–619.
Van Nelson, C., Nelson, J. S., & Malone, B. G. (2004). Predicting success of international graduate
students in an academic university. College and University Journal, 80(1), 19–27.
Vinke, A. A., & Jochems, W. M. G. (1993). English proficiency and academic success in interna-
tional postgraduate education. Higher Education, 26, 275–285.
Wiberg, M., & Sundstrom, A. (2009). A comparison of two approaches to correction of restriction
of range in correlation analysis. Practical Assessment, Research & Evaluation, 14(5), 1–9.
Woodrow, L. (2006). Academic success of international postgraduate education students and the
role of English proficiency. University of Sydney Papers in TESOL, 1, 51–70.
Zwick, R. (2002). Fair game? The use of standardized admissions tests in higher education.
New York: RoutledgeFalmer.
BU (n=413)
100%
15% 22%
80% 37%
60% 50% Top 25% GPA
51%
40% 47% Mid 50% GPA
20% 36% Boom 25% GPA
27% 16%
0%
Boom 25% Mid 50% Top 25%
iBT iBT iBT
HA (n=186)
100%
13%
80% 28% 28%
60% 51% Top 25% GPA
46%
40% 61% Mid 50% GPA
20% 36% Boom 25% GPA
26%
0% 11%
Boom 25% Mid 50% Top 25%
iBT iBT iBT
SE (n=959)
100%
17% 26%
80% 33%
60% 51% Top 25% GPA
48%
40% 49% Mid 50% GPA
20% 32% Boom 25% GPA
26% 18%
0%
Boom 25% Mid 50% Top 25%
iBT iBT iBT
SS (n=283)
100%
18% 27%
80% 30%
Figures A1–A4