0% found this document useful (0 votes)
11 views21 pages

Feedback Timing in L2 Vocabulary Learning

This document summarizes a research article that studied the effects of feedback timing on second language vocabulary learning. The study compared immediate feedback, which was provided right after responses, to delayed feedback, which was provided after all items were practiced. 98 Japanese students learned English-Japanese word pairs and were tested immediately, 1 week, and 4 weeks later. The study aimed to determine the optimal timing of feedback and addressed limitations of previous research by controlling for lag to test and manipulating practice frequency. The results suggested feedback timing may have little effect on vocabulary learning when lag to test is controlled, regardless of error frequency during learning.

Uploaded by

Noy Pabico
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
11 views21 pages

Feedback Timing in L2 Vocabulary Learning

This document summarizes a research article that studied the effects of feedback timing on second language vocabulary learning. The study compared immediate feedback, which was provided right after responses, to delayed feedback, which was provided after all items were practiced. 98 Japanese students learned English-Japanese word pairs and were tested immediately, 1 week, and 4 weeks later. The study aimed to determine the optimal timing of feedback and addressed limitations of previous research by controlling for lag to test and manipulating practice frequency. The results suggested feedback timing may have little effect on vocabulary learning when lag to test is controlled, regardless of error frequency during learning.

Uploaded by

Noy Pabico
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Language[Link]

com/
Teaching Research

Effects of feedback timing on second language vocabulary learning: Does


delaying feedback increase learning?
Tatsuya Nakata
Language Teaching Research published online 27 July 2014
DOI: 10.1177/1362168814541721

The online version of this article can be found at:


[Link]

Published by:

[Link]

Additional services and information for Language Teaching Research can be found at:

Email Alerts: [Link]

Subscriptions: [Link]

Reprints: [Link]

Permissions: [Link]

Citations: [Link]

>> OnlineFirst Version of Record - Jul 27, 2014

What is This?

Downloaded from [Link] at University of Otago Library on September 7, 2014


541721
research-article2014
LTR0010.1177/1362168814541721Language Teaching ResearchNakata

LANGUAGE
TEACHING
Article RESEARCH

Language Teaching Research

Effects of feedback timing on


1­–20
© The Author(s) 2014
Reprints and permissions:
second language vocabulary [Link]/[Link]
DOI: 10.1177/1362168814541721
learning: Does delaying [Link]

feedback increase learning?

Tatsuya Nakata
Victoria University of Wellington, New Zealand

Abstract
Feedback, or information given to learners regarding their performance, is found to facilitate
second language (L2) learning. Research also suggests that the timing of feedback (whether it is
provided immediately or after a delay) may affect learning. The purpose of the present study was
to identify the optimal feedback timing for L2 vocabulary learning. This study differs from previous
feedback timing studies in two important respects. First, unlike some previous studies, feedback
timing was not confounded with lag to test (interval between the last encounter with a given item
and the posttest). Second, in order to test the view that delayed feedback may be particularly
effective when learners make few errors during learning, the present study manipulated the
frequency of practice to influence learning phase performance. In this study, 98 Japanese college
students studied 16 English–Japanese word pairs. Immediate feedback was given immediately
after each response, whereas delayed feedback was withheld until all target items were practised.
Learning was measured by posttests administered immediately, 1 week, and 4 weeks after the
treatment. Results suggested that when lag to test is controlled, feedback timing may have little
effect on L2 vocabulary learning regardless of the frequency of errors during learning.

Keywords
Delayed feedback, timing of feedback, vocabulary learning

I Introduction
Research suggests that feedback, which is defined as information given to learners
regarding their performance, facilitates second language (L2) learning (e.g., Lee, 2013;

Corresponding author:
Tatsuya Nakata, School of Linguistics and Applied Language Studies, Victoria University of Wellington, PO
Box 600, Wellington 6140, New Zealand.
Email: tatsuya.nakata2@[Link]

Downloaded from [Link] at University of Otago Library on September 7, 2014


2 Language Teaching Research 

Li, 2010; Lyster, Saito, & Sato, 2013). The role of corrective feedback, feedback pro-
vided in response to learner errors, has attracted particular attention from researchers
investigating conversational interaction (e.g., Li, 2010; Lyster et al., 2013; Rolin-Ianziti,
2010) and writing (e.g., Evans, Hartshorn, McCollum, & Wolfersberger, 2010; Lee,
2013). Not only SLA (second language acquisition) researchers but also cognitive psy-
chologists have examined the effects of feedback on learning (e.g., Butler, Karpicke, &
Roediger, 2007; Kulik & Kulik, 1988; Metcalfe, Kornell, & Finn, 2009). In cognitive
psychology, feedback typically refers to the provision of the correct answer following a
learner’s response. Note that while corrective feedback is provided only in response to
learner errors (e.g., Li, 2010; Lyster et al., 2013), feedback can be given after successful
as well as unsuccessful performance.
The cognitive psychology literature suggests that the timing of feedback may affect
learning. For instance, suppose that the learner was asked to translate a Swahili word
chakura (‘food’) into English. Would it be more effective to give the correct answer
immediately after the response (immediate feedback) or 1 week after (delayed feed-
back)? Empirical studies have yielded mixed results regarding the effects of immediate
and delayed feedback (e.g., Butler et al., 2007; Kulik & Kulik, 1988; Metcalfe et al.,
2009). Furthermore, previous studies on feedback timing may suffer from several limita-
tions (see below). Due to the inconsistent findings and limitations of the previous
research, it is not clear whether immediate or delayed feedback is more effective for L2
learning. The present study compared the effects of immediate and delayed feedback on
L2 learning while addressing the limitations of previous studies. Findings of this study
may be useful because they may allow us to determine whether immediate or delayed
feedback should be used in order to optimize L2 learning.

II  Review of literature


1  Theoretical background
The timing of feedback is regarded as a factor that may affect L2 acquisition (e.g.,
DeKeyser, 2007; Evans et al., 2010; Lee, 2013; Quinn, 2013; Sheen, 2012). DeKeyser
(2007), for instance, notes that feedback should not be too immediate or too delayed.
Evans et al. (2010) and Lee (2013) argue that written corrective feedback should be pro-
vided immediately, although their claims are not based on empirical evidence. Some
cognitive psychology studies, however, suggest that delaying feedback may increase
learning (e.g., Butler et al., 2007; Carpenter & Vul, 2011; Roediger, Agarwal, Kang, &
Marsh, 2010). The superiority of delayed over immediate feedback is referred to as the
‘delay-retention effect’ (e.g., Metcalfe et al., 2009; Mory, 2004).
The delay-retention effect may be accounted for by the distributed practice effect and
interference theory. The distributed practice effect refers to the phenomenon where larger
spacing leads to better long-term retention than shorter spacing or no spacing at all (e.g.,
Cepeda et al., 2009; Cepeda, Pashler, Vul, Wixted, & Rohrer, 2006; Cepeda, Vul, Rohrer,
Wixted, & Pashler, 2008). Because delayed feedback is given after a greater delay than
immediate feedback, the distributed practice effect predicts that delaying feedback may
facilitate learning (e.g., Butler et al., 2007; Metcalfe et al., 2009). Interference theory

Downloaded from [Link] at University of Otago Library on September 7, 2014


Nakata 3

also predicts an advantage of delayed over immediate feedback (e.g., Butler et al., 2007;
Carpenter & Vul, 2011; Mory, 2004). Suppose learners made errors (e.g., *chakura =
‘drink’) during learning. (Note that throughout this study, ‘errors’ refer only to incorrect
responses [i.e., errors of commission] and do not include blank responses [i.e., errors of
omission]; Metcalfe et al., 2009.) When the correct answer (chakura = ‘food’) is given
immediately after the response, learners may confuse their error with the correct response
and might learn false information (e.g., Butler et al., 2007; Carpenter & Vul, 2011; Mory,
2004). In contrast, when feedback is given after a delay, learners’ errors may be forgotten
by the time they receive feedback. As a result, their error is less likely to interfere with
the correct response, which may increase learning (e.g., Butler et al., 2007; Carpenter &
Vul, 2011; Mory, 2004).
The theory of errorless learning (e.g., Skinner, 1954), on the other hand, predicts an
advantage of immediate feedback (e.g., Butler et al., 2007; Metcalfe et al., 2009).
According to this theory, immediate feedback may be more effective because unless feed-
back is given immediately after the response, errors might be consolidated, and learners
might learn false information. Note that predictions based on interference theory and the
theory of errorless learning both rest on the assumption that learners make errors during
learning. If learners produce only few errors, both theories may be irrelevant (see below).

2  Empirical evidence
Let us now consider empirical evidence regarding the effects of immediate and delayed
feedback. Because none of the previous feedback timing studies examined L2 learning
(unpublished conference presentations – such as Quinn, 2013; Sheen, 2012 – will not be
discussed in detail in this study), this section will review previous studies that investi-
gated the learning of materials other than L2 such as first language (L1) vocabulary or
reading materials. Empirical studies have yielded mixed results regarding the effects of
immediate and delayed feedback. Kulik and Kulik (1988), for instance, conducted a
meta-analysis of 53 experiments on feedback timing and found that although 26 of them
observed the delay-retention effect, the other 27 failed to do so. At least three explana-
tions have been offered for the inconsistent results. First, Kulik and Kulik (1988) point
out that studies conducted in laboratory settings are more likely to observe a significant
delay-retention effect than those conducted in actual classroom settings. Butler et al.
(2007) and Roediger et al. (2010) speculate that the setting of the experiment (laboratory
or classroom) may interact with the delay-retention effect probably because participants
may process feedback differently in laboratory and classroom studies. Specifically,
Butler et al. and Roediger et al. point out that delayed feedback may not be studied as
carefully as immediate feedback in classroom studies. Laboratory studies, in contrast,
usually require learners to study feedback for a fixed amount of time in both immediate
and delayed feedback conditions, ensuring that both types of feedback are processed
equally carefully. Because learners may pay more attention to delayed feedback in labo-
ratory studies than in classroom studies, laboratory studies may be more likely to pro-
duce a significant delay-retention effect (Butler et al., 2007; Roediger et al., 2010). A
possible interaction between the delay-retention effect and experimental settings may
partially account for the mixed results of previous studies.

Downloaded from [Link] at University of Otago Library on September 7, 2014


4 Language Teaching Research 

Second, Metcalfe et al. (2009) point out that the effects of immediate and delayed
feedback may be conditional upon the frequency of errors produced during learning. As
noted above, interference theory predicts an advantage of delayed over immediate feed-
back, whereas the theory of errorless learning predicts a superiority of immediate feed-
back. It should be noted that both predictions are based on the assumption that learners
make errors during learning and may be irrelevant when only few errors are produced.
The distributed practice effect, in contrast, may be observed for both correct and incor-
rect responses. As a result, when learners produce few errors during learning, a signifi-
cant delay-retention effect may be observed due to the distributed practice effect. If
learners make many errors, in contrast, the beneficial effects of delaying feedback (larger
spacing and less interference) might be offset by the risk of not correcting an error imme-
diately, and a delay-retention effect may not be observed (Metcalfe et al., 2009). The
inconsistent findings of previous studies may be attributed in part to a possible interac-
tion between the delay-retention effect and the frequency of errors during learning.
Two experiments conducted by Metcalfe et al. (2009) suggest that the delay-retention
effect may interact with the frequency of errors during learning. In their Experiment 1,
27 American grade school children studied 72 low-frequency English words. At the
beginning of the treatment, participants were presented with a target word followed by
its definition. After all target words were introduced, participants practised retrieval.
More specifically, they were presented with a definition and asked to type the corre-
sponding target word. Delayed feedback was provided 1 or 4 days after the response,
whereas immediate feedback was given on the same day. In Experiment 2, 20 Columbia
University students studied 75 low-frequency English words. Immediate feedback was
given on the same day as the initial treatment session, whereas delayed feedback was
given 3.85 days after on average. Metcalfe et al. found a larger delay-retention effect in
their Experiment 1 than in their second experiment. Metcalfe et al. argue that the results
might have been caused by a difference in the frequency of errors during learning.
Participants in their second experiment produced more errors (61% of all responses) dur-
ing the treatment than in their Experiment 1 (40%). Because delayed feedback may be
particularly effective when learners make few errors during learning, Metcalfe et al.’s
Experiment 1 might have produced a larger delay-retention effect than in their Experiment
2. One limitation of their study, though, is that although their first experiment was con-
ducted with grade school children, the participants in Experiment 2 were university stu-
dents. As a result, the results of their experiments may be at least partly attributed to the
difference in the age of participants rather than differential learning phase performance.
Third, the inconsistent results of existing studies may be partially due to methodologi-
cal differences. Previous studies differ in the operationalization of immediate and delayed
feedback (Roediger et al., 2010). For instance, in Carpenter and Vul (2011), immediate
feedback was given immediately after each response, and delayed feedback was given 3
seconds after. In Phye and Baller (1970), in contrast, immediate feedback was provided
after 30 minutes, and delayed feedback was provided 2 days after the response. This
implies that what is classified as immediate feedback in some studies may qualify as
delayed feedback in others. Earlier studies also differ in the materials. The materials used
by previous studies include L1 vocabulary, word-number pairs (e.g., right-12), trigram
pairs (e.g., NHK-RCX), L2-trigram pairs (e.g., chakura-NHK), face-name pairs, reading

Downloaded from [Link] at University of Otago Library on September 7, 2014


Nakata 5

materials, motor skills, programming languages, mathematics, chemistry, and psychol-


ogy (see Kulik & Kulik, 1988, for a review). These methodological differences could
partially be responsible for the inconsistent results of previous studies as well.

3  Limitations of previous studies


Previous feedback timing studies not only report inconsistent results but also suffer from
at least three limitations. First, none of the previous feedback timing studies examined
L2 learning. Thus, it is unclear to what extent their findings may be applicable to L2
learning. Second, some earlier studies did not control for lag to test. Lag to test refers to
an interval between the last encounter with a given item and the test (e.g., Cepeda et al.,
2008; Metcalfe et al., 2009; Rohrer, Taylor, Pashler, Wixted, & Cepeda, 2005). For
instance, if a posttest is given 24 hours after the last encounter with a given item (lag to
test is 24 hours) instead of 1 hour after (lag to test is 1 hour), memory performance will
naturally be worse. In some previous studies, delayed feedback was associated with
greater lag to test than immediate feedback. Let me illustrate this point by using Butler
et al. (2007, Experiment 2) as an example. Butler et al. examined the effects of immedi-
ate and delayed feedback on the retention of reading materials. Their experiment was
conducted over a period of 8 days. On Day 1, 40 American undergraduate students read
12 passages and answered multiple-choice comprehension questions. Feedback was
given immediately after each response on Day 1 in the immediate feedback condition,
while it was provided on Day 2 in the delayed feedback condition. The posttest was
administered on Day 8. Butler et al. (2007) found that delayed feedback led to a higher
posttest score (70%) than immediate feedback (60%). Based on their finding, Butler et
al. argue that delaying feedback may increase long-term retention.
Metcalfe et al. (2009), however, point out that Butler et al.’s (2007) finding may be at
least partly attributed to lag to test rather than feedback timing per se. Specifically, in the
immediate feedback condition in Butler et al., participants received feedback on Day 1,
and the posttest was conducted on Day 8. Hence, there was a lag of 7 days between feed-
back and the posttest. In the delayed feedback condition, however, there was a lag of only
6 days between feedback and the posttest. The confounding of feedback timing and lag
to test is problematic because a shorter lag to test generally leads to better memory per-
formance than a longer lag (e.g., Cepeda et al., 2008; Metcalfe et al., 2009; Rohrer et al.,
2005). Feedback timing was confounded with lag to test in other existing studies as well
(e.g., Kulhavy & Anderson, 1972; O’Neill, Rasor, & Bartz, 1976; Swindell & Walls,
1993).
Third, as described above, Metcalfe et al. (2009) observe that delayed feedback may
be particularly effective when learners make few errors during learning. This suggests
that in order to obtain a comprehensive picture regarding the effects of immediate and
delayed feedback, it may be useful to manipulate the frequency of errors during learning.
Yet, a possible relationship between the delay-retention effect and learning phase perfor-
mance has not been explored thoroughly in the existing literature. Metcalfe et al.’s (2009)
study constitutes the only exception. However, the results of their study may be at least
partly attributed to the difference in the age of participants (i.e., grade school vs. college
students) rather than the difference in the proportions of errors per se. It would be useful

Downloaded from [Link] at University of Otago Library on September 7, 2014


6 Language Teaching Research 

to examine a possible relationship between the delay-retention effect and learning phase
performance without confounding the frequency of errors with the age of participants.

III  Present study


This study differs from previous feedback timing studies in three important respects.
First, because none of the previous studies on feedback timing examined L2 learning, it
is unclear to what extent their findings may be applicable to L2 learning. With this limita-
tion in mind, the present study compared the effects of immediate and delayed feedback
on L2 vocabulary learning. Vocabulary rather than syntax or phonology was chosen as a
focus of this study because L1 vocabulary research has observed a significant delay-
retention effect (Metcalfe et al., 2009), suggesting that feedback timing may affect
vocabulary learning. In contrast, because none of the previous feedback timing studies
looked into the learning of syntax or phonology either in L1 or L2, it is not yet clear
whether the learning of these aspects may be affected by feedback timing. Hence, the
present study examined L2 vocabulary learning. Investigating the effects of feedback
timing on L2 vocabulary learning may also have pedagogical value because previous
SLA (e.g., Ellis & He, 1999; Lyster et al., 2013) and cognitive psychology studies
(e.g., Metcalfe & Kornell, 2007; Pashler, Cepeda, Wixted, & Rohrer, 2005) have shown
that feedback may increase vocabulary learning.
Second, because feedback timing was confounded with lag to test, the results of ear-
lier studies may be at least partly attributed to lag to test rather than feedback timing per
se. In order to address this limitation, immediate and delayed feedback in this study were
controlled for lag to test. Third, in order to test the view that delayed feedback may be
particularly effective when learners make few errors during learning (Metcalfe et al.,
2009), the present study set out to examine a possible relationship between the delay-
retention effect and learning phase performance. Unlike Metcalfe et al. (2009), where the
frequency of errors during learning was confounded with the age of participants, the
effects of these two factors were isolated in this study.
In order to control feedback timing and the frequency of errors during learning, this
study was conducted in a computer-assisted language learning (CALL) environment. At
the beginning of the treatment, Japanese learners of English were presented with an L2
(English) target word together with its L1 (Japanese) meaning (e.g., mane = たてがみ).
From the second encounter, learners practised retrieval. In other words, participants were
presented with a Japanese word and asked to type the corresponding English translation
(e.g., たてがみ = ____?). Retrieval was followed by either immediate or delayed feed-
back. In the former, feedback was provided immediately after each retrieval attempt. In
the latter, feedback was withheld until all target items were practised. The frequency of
errors during learning was manipulated by using four levels of retrieval frequency: one,
three, five, and seven. Retrieval frequency refers to the number of retrieval attempts
(e.g., たてがみ = ____?) during the learning phase. For instance, if learners practise
retrieval five times, the retrieval frequency is five. Because repeated retrieval may lead
to more successful learning phase performance than fewer retrievals (e.g., practising
retrieval seven times may lead to better learning phase performance than practising
retrieval three times), manipulating retrieval frequency may allow us to investigate a

Downloaded from [Link] at University of Otago Library on September 7, 2014


Nakata 7

Table 1.  Design of the present study.

Retrieval frequency Feedback timing (within-participant variable)


Subgroups
(between-participant variable)  Item set A Item set B
Retrieval 1 X (n = 12) Immediate Delayed
(n = 24) Y (n = 12) Delayed Immediate
Retrieval 3 X (n = 12) Immediate Delayed
(n = 24) Y (n = 12) Delayed Immediate
Retrieval 5 X (n = 12) Immediate Delayed
(n = 24) Y (n = 12) Delayed Immediate
Retrieval 7 X (n = 13) Immediate Delayed
(n = 26) Y (n = 13) Delayed Immediate

Note. Each item set consisted of 8 items.

possible relationship between the feedback timing effect and learning phase perfor-
mance. The research question of this study is as follows: Is delayed feedback more effec-
tive than immediate feedback for L2 vocabulary learning when lag to test is controlled
and learners make few errors during learning?

IV Method
1  Participants
The participants were 98 first-year engineering students (aged 15–16) from three EFL
classes at a technical college in Gifu, Japan. They had been studying English for at least
3 years. Students were given participant information sheets and asked to sign consent
forms if they chose to participate.

2  Experimental design
There were three independent variables in the current study. The first independent vari-
able was the timing of feedback: immediate and delayed. The second independent vari-
able was the retrieval frequency (i.e., number of retrieval attempts during the learning
phase): one, three, five, and seven. The third independent variable was the retention
interval (i.e., interval between the treatment and posttest): immediate, 1-week delayed,
and 4-week delayed posttests. Retrieval frequency was a between-participant variable,
and the timing of feedback and retention interval were within-participant variables. The
dependent variable was the number of correct responses on the posttest.
Table 1 summarizes the design of the current study. The participants were randomly
assigned by a computer program to one of the four groups: Retrieval 1, 3, 5, and 7.
Although no data were available regarding their English proficiency, an analysis of
learning phase performance suggested that the four retrieval frequency groups might not
have differed significantly from each other in their ability to learn L2 vocabulary

Downloaded from [Link] at University of Otago Library on September 7, 2014


8 Language Teaching Research 

(see Nakata, 2013, for details). The Retrieval 7 group consisted of 26 participants, and
each of the remaining three groups consisted of 24 participants.1 The imbalance in the
number of participants was caused by the absence of participants. The participants in
each group were randomly divided into two subgroups: Subgroups X and Y. Sixteen
target word pairs were also divided into two sets of eight items: Sets A and B (see below).
The two subgroups of participants in each group studied both sets of word pairs under
different feedback conditions (immediate or delayed), thus counterbalancing the effects
of target items (Table 1).

3  Target and filler items


Sixteen English–Japanese word pairs (e.g., mane – たてがみ) were used as target items.
Words that are outside the most frequent 9,000 word families in the British National
Corpus frequency lists (Nation, 2006) were chosen because the target items needed to be
unfamiliar to participants. The 16 word pairs were divided into Sets A and B so that the
learning difficulty would be distributed as evenly as possible, Set A: billow, gouge, grig,
jibe, levee, loach, toupee, and urn; Set B: apparition, citadel, dally, husk, mane, mirth,
rue, and warble. Learning difficulty was operationally defined as the pretest and posttest
scores in a similar previous experiment, where another group of 95 Japanese college
students studied the same 16 English–Japanese word pairs (Nakata, 2013). Although the
two sets may not be completely equivalent in their difficulty, it was judged that a possible
difference, if any, might not have a major effect on the results of the present study because
effects of target items would be counterbalanced across participants (Table 1).
The number of target items was set to 16 (eight items per feedback type) based on the
results of pilot studies conducted with 53 Japanese learners of English with a similar
learning profile as the participants in this study. One may argue that the relatively small
number of items may reduce the potential of showing a difference between immediate
and delayed feedback. Despite this potential disadvantage, it was decided to set the num-
ber of target items to 16 for two reasons. First, previous studies on word-pair learning
found significant differences between conditions using a relatively small number of
items per condition such as two, four, or six (e.g., Cull, Shaughnessy, & Zechmeister,
1996; Karpicke & Roediger, 2007). These studies suggest that even if the number of
target items is set to eight per condition, we may still be able to observe a significant
feedback timing effect provided that such an effect exists. Second, even if only 16 target
items are used, it may not considerably decrease the probability of detecting a significant
effect of feedback timing because the cell size is relatively large (98). For the above two
reasons, the number of target items was set to 16.
Three additional items were used as filler items: tyro, valor, and lava. They were
chosen based on the same criterion as the target items. Filler items were studied and
tested like target word pairs, but were excluded from analysis. To ensure that the filler
items would be treated in the same way as the target items, the participants were not
informed about any differences between items. Filler items were studied at the beginning
and end of the treatment and used as primacy and recency buffers (e.g., Karpicke &
Roediger, 2007). In other words, when items are presented in series, the first and last
several items tend to be remembered better than the middle ones, a phenomenon known

Downloaded from [Link] at University of Otago Library on September 7, 2014


Nakata 9

as serial position effects. The primacy and recency buffers were included to reduce these
effects on target items. The same target and filler items were used during the pretest,
treatment, and posttest.

4  Procedure
The experiment was conducted during regular class hours with a computer program
developed by the author. The experiment consisted of three sessions.

a  Session 1.  Participants received explanations about the computer program and prac-
tised using it with three sample word pairs. After the practice, productive and receptive
pretests were given in that order. The pretest was followed by the treatment, where par-
ticipants studied 19 English–Japanese word pairs (including three filler items) using a
computer program. After the treatment, participants answered 10 two-digit additions
(e.g., 53 + 49 = ?) as a filler task. The immediate posttest was given after the filler task.

b  Sessions 2 and 3.  In order to measure retention, the delayed posttest was administered
1 and 4 weeks after the treatment.

5  Treatment
Table 2 presents the overview of the treatment in the present study. The treatment con-
sisted of the initial presentation, retrieval phases (e.g., R1, R2, and R3 in Table 2), delayed
feedback phases (e.g., D1, D2, and D3 in Table 2), and final review. In the initial presenta-
tion, the English and Japanese words were presented simultaneously for 8 seconds per
word pair (e.g., mane = たてがみ). Each word pair was presented only once in the initial
presentation. The initial presentation was followed by a series of alternating retrieval and
delayed feedback phases. There were one, three, five, and seven sets of retrieval and
delayed feedback phases for the Retrieval 1, 3, 5, and 7 groups, respectively.
In the retrieval phase, participants were presented with a Japanese word and asked to
type the corresponding English translation (e.g., たてがみ = ____?). Each word pair
appeared only once in each retrieval phase. Participants were allowed to take as much
time as they needed to type a response. For items assigned to the immediate feedback
condition, feedback was provided to the participants immediately after each response.
The target English word, Japanese translation, and learners’ response were given in the
feedback window. Feedback also indicated whether the response was correct, partially
correct, or incorrect. Partially correct responses (e.g., apparation, appartion, and appli-
tion for apparition) were defined as those that would be awarded 0.75 using a lexical
production scoring protocol (e.g., Barcroft, 2007). Delayed feedback was not given until
the end of each retrieval phase (e.g., D1, D2, and D3 in Table 2), where feedback for all
eight delayed feedback items was presented one at a time. Both immediate and delayed
feedback were shown for 5 seconds per response.
The retrieval and delayed feedback phases were followed by the final review, where
the English and Japanese words were presented simultaneously for 5 seconds per word
pair. The final review was included to control for lag to test (see Section II) in the

Downloaded from [Link] at University of Otago Library on September 7, 2014


10 Language Teaching Research 

Table 2.  Overview of the treatment.

Group Initia1 presentation Retrieval and delayed feedback phasesa Final review
Retrieval 1 Initia1 presentation R1 D1 Final review
Retrieval 3 Initia1 presentation R1 D1 R2 D2 R3 D3 Final review
Retrieval 5 Initia1 presentation Rl Dl R2 D2 R3 D3 R4 D4 R5 D5 Final review
Retrieval 7 Initia1 presentation Rl Dl R2 D2 R3 D3 R4 D4 R5 D5 R6 D6 R7 D7 Final review

Note. a R: retrieval phase; D: delayed feedback phase.

immediate and delayed feedback items. As each retrieval phase was followed by a
delayed feedback phase (Table 2), without the final review, the delayed feedback items
would be clustered at the end of the treatment and have a shorter interval to the posttest
than the immediate feedback items. The inclusion of the final review ensured that the
immediate and delayed feedback items were controlled for lag to test. Because it may be
common for learners to review what they will be tested on shortly before a test (Kornell,
2009), the inclusion of the final review may also reflect authentic learning and increase
ecological validity. The three filler items were used as primacy and recency buffers (e.g.,
Karpicke & Roediger, 2007) and studied at the beginning and end of the treatment in all
four groups.

6  Dependent measures
a Pretest.  Immediately before the treatment, productive and receptive tests were given
in that order as the pretest. In the former, participants were presented with a Japanese
word and asked to type the corresponding English translation (e.g., たてがみ = ____?).
In the latter, they translated target English words into Japanese (e.g., mane = ____?). The
order of items in the test was randomized anew for each participant to minimize the
potential of an order effect. In the productive pretest, it was necessary to prevent partici-
pants from providing synonyms for a target word because if participants produced hair
for the target word mane, for instance, it would not be clear whether or not they were
familiar with mane. In order to prevent learners from providing synonyms, one letter in
the target word and the number of letters in the word were given as a hint (e.g., _ _ n _
for mane) together with the Japanese translation in the productive pretest. The hints were
determined so as to minimize effects that they may have on performance on the receptive
pretest, which was administered immediately after the productive pretest (see Nakata,
2013, for the protocol to determine the hints).

b Posttest.  Immediate, 1-week delayed, and 4-week delayed posttests were adminis-
tered in the present study. At each test administration, productive and receptive tests
were given in that order. Although the target items were practised in a productive format
during the treatment, learning was measured by receptive as well as productive tests.
This is because research suggests that administering multiple types of tests may be useful
because it may give us a more comprehensive picture of lexical development than a sin-
gle test (see Webb, 2012, for a review). The posttests were exactly the same as the

Downloaded from [Link] at University of Otago Library on September 7, 2014


Nakata 11

pretests except that the hints (e.g., _ _ n _ for mane) were not given in the productive
posttest. This is because in the posttest, learners were instructed to produce only English
words that were studied during the treatment and informed that giving a synonym for
target words would be marked as incorrect. The delayed posttests were administered
without prior notice so that participants would not review the target words during the
period between the treatment and delayed posttests. As in the pretest, the order of items
in the test was randomized anew for each participant to minimize the potential of an
order effect. The immediate and the two delayed posttests were exactly the same except
the item order.

7  Scoring
Responses on the pretest and posttest were scored using the following two procedures:
strict and sensitive. In the productive test, in the strict scoring method, only responses
without any misspellings were scored as correct. In the sensitive scoring procedure,
responses that would be awarded 0.75 using a lexical production scoring protocol (e.g.,
Barcroft, 2007) were also scored as correct (e.g., apparation, appartion, and applition
for apparition). In the receptive test, in the strict scoring method, responses were scored
as incorrect if (1) they were of a wrong part of speech (e.g., 後悔 [noun] for rue) or (2)
an intransitive verb was given for a transitive verb (e.g., 調和させる [transitive] for jibe)
and vice versa. In the sensitive scoring system, the above two kinds of responses were
both marked as correct.

V Results
1  Pretest
None of the participants exhibited prior knowledge of any of the target words on the
productive pretest. When collapsed across the four retrieval frequency groups, the aver-
age receptive pretest score (SDs in parentheses) was 0.04 (0.20) and 0.07 (0.26) out of 8
with strict scoring and 0.05 (0.22) and 0.08 (0.28) out of 8 with sensitive scoring in the
immediate and delayed feedback conditions, respectively. These findings suggest that
participants had little or no prior knowledge of the target items.

2  Learning phase data


First, let us investigate how much spacing intervened between the retrieval attempt and
feedback in the delayed feedback condition. The retrieval attempt and delayed feedback
were separated by 95.61 (45.76), 91.05 (21.32), 93.88 (15.89), and 90.72 (21.79) seconds
on average in the Retrieval 1, 3, 5, and 7 groups, respectively (SDs in parentheses). When
collapsed across the four groups, the mean interval duration was 92.77 (28.12) seconds.
No statistically significant difference was found among the four groups in their average
spacing, F (3, 97) = 0.17, p = .919, producing no effect size (η2 < .01). As the difference
is relatively small, it may be possible to assume that the four retrieval frequency groups
had roughly equivalent spacing between the retrieval attempt and delayed feedback.

Downloaded from [Link] at University of Otago Library on September 7, 2014


12 Language Teaching Research 

Table 3.  Average number of correct responses on the productive posttest (standard
deviations in parentheses).

Immediate 1 week 4 weeks

Scoring Group Immediate Delayed Immediate Delayed Immediate Delayed


feedback feedback feedback feedback feedback feedback
Strict Retrieval 1 4.83 4.63 2.33 2.08 2.00 1.88
  (n = 24) (2.28) (2.12) (2.04) (2.00) (1.91) (1.96)
  Retrieval 3 5.25 5.58 2.21 2.46 2.04 2.00
  (n = 24) (2.44) (2.10) (1.84) (2.34) (1.83) (1.72)
  Retrieval 5 7.75 7.75 4.08 3.75 3.75 3.92
  (n = 24) (0.44) (0.61) (1.93) (1.82) (2.36) (2.38)
  Retrieval 7 7.58 7.65 4.73 4.73 3.69 4.12
  (n = 26) (1.03) (0.80) (2.09) (2.20) (2.20) (2.16)
  Total 6.38 6.43 3.37 3.29 2.89 3.00
  (n = 98) (2.17) (2.05) (2.24) (2.33) (2.23) (2.29)
Sensitive Retrieval 1 5.71 5.63 3.25 2.88 2.79 2.92
  (n = 24) (2.22) (2.14) (2.27) (2.31) (2.26) (2.43)
  Retrieval 3 5.92 6.46 2.75 3.33 2.79 3.13
  (n = 24) (2.28) (1.82) (2.15) (2.35) (1.98) (1.80)
  Retrieval 5 7.92 7.83 5.63 5.04 5.13 5.42
  (n = 24) (0.28) (0.48) (1.95) (2.16) (2.07) (2.32)
  Retrieval 7 7.81 7.81 5.65 5.42 5.31 5.31
  (n = 26) (0.63) (0.57) (2.10) (1.98) (2.43) (2.20)
  Total 6.86 6.95 4.35 4.19 4.03 4.21
  (n = 98) (1.89) (1.70) (2.48) (2.43) (2.48) (2.47)

Note. The maximum score is 8 for each cell.

In order to examine a possible relationship between the delay-retention effect and


learning phase performance, the frequency of errors during learning was analysed. (Note
that ‘errors’ here refer only to incorrect responses [i.e., errors of commission] and do not
include blank responses [i.e., errors of omission].) The analysis showed that 36.2%
(22.3%), 24.4% (17.6%), 18.1% (8.8%), and 16.6% (14.0%) of responses during the
treatment were errors (SDs in parentheses) in the Retrieval 1, 3, 5, and 7 groups, respec-
tively. The difference among the four groups was statistically significant, F (3, 97) =
7.28, p < .001, η2= .04. The results suggest that repeated retrieval led to better learning
phase performance as intended. This enables us to test the view that delayed feedback
may be particularly effective when learners make few errors during learning (e.g.,
Metcalfe et al., 2009). The results also indicate that the proportion of errors in this study
(16.6%–36.2%) was lower than in Metcalfe et al.’s (2009) experiments (Experiment 1:
40%, Experiment 2: 61%). Hence, this study might produce a larger delay-retention
effect compared with Metcalfe et al.

3  Posttest performance
Tables 3 and 4 summarize the immediate and delayed posttest results as a function of
feedback timing and retrieval frequency. Cronbach’s alpha was .84 or higher (.84–.91)

Downloaded from [Link] at University of Otago Library on September 7, 2014


Nakata 13

Table 4.  Average number of correct responses on the receptive posttest (standard deviations
in parentheses).

Immediate 1 week 4 weeks

Scoring Group Immediate Delayed Immediate Delayed Immediate Delayed


feedback feedback feedback feedback feedback feedback
Strict Retrieval 1 5.54 5.88 5.21 4.88 4.75 4.83
  (n = 24) (1.93) (1.94) (2.06) (2.19) (1.92) (1.79)
  Retrieval 3 5.63 5.75 5.13 5.00 4.71 4.88
  (n = 24) (2.18) (1.87) (2.15) (2.41) (1.92) (2.25)
  Retrieval 5 7.17 7.17 6.96 6.54 6.96 6.67
  (n = 24) (0.82) (1.01) (1.08) (1.18) (1.12) (1.49)
  Retrieval 7 7.23 7.50 7.15 7.08 6.54 6.81
  (n = 26) (1.07) (0.81) (0.97) (1.23) (1.88) (1.83)
  Total 6.41 6.59 6.13 5.90 5.76 5.82
  (n = 98) (1.77) (1.65) (1.88) (2.04) (2.00) (2.06)
Sensitive Retrieval 1 5.67 6.17 5.29 5.04 4.88 5.00
  (n = 24) (1.97) (1.88) (2.07) (2.26) (1.98) (1.84)
  Retrieval 3 5.71 6.08 5.13 5.08 4.75 4.96
  (n = 24) (2.24) (1.98) (2.15) (2.41) (1.96) (2.29)
  Retrieval 5 7.50 7.50 7.13 7.00 7.00 6.92
  (n = 24) (0.72) (0.78) (0.95) (1.18) (1.06) (1.35)
  Retrieval 7 7.62 7.65 7.31 7.23 6.69 7.00
  (n = 26) (0.64) (0.69) (0.84) (1.24) (1.87) (1.90)
  Total 6.64 6.87 6.23 6.11 5.85 5.99
  (n = 98) (1.79) (1.60) (1.88) (2.10) (2.02) (2.10)

Note. The maximum score is 8 for each cell.

for all dependent measures, indicating good reliability. The productive and receptive test
scores were analysed by a three-way mixed design 2 (feedback timing: immediate /
delayed) × 4 (retrieval frequency group: 1 / 3 / 5 / 7) × 3 (retention interval: immediate /
1-week delayed / 4-week delayed) ANOVA. As some items were answered correctly on
the receptive pretest, the pretest scores were subtracted from the posttest scores, and
gains were analysed when examining the receptive test results. Table 5 shows the results
of the ANOVAs. (In the present study, retrieval frequency was manipulated in order to
examine whether the proportion of errors during learning, which may be a function of
retrieval frequency, might interact with the feedback timing effect [see Section III].
Because the question of whether repeated retrieval may increase learning is outside the
scope of this study, the main effect of retrieval frequency will not be investigated.) The
table indicates the following three things. First, the main effect of feedback timing was
not significant on either the productive or receptive posttest regardless of the scoring
procedure. This suggests that the timing of feedback had little effect on learning when
collapsed across the retrieval frequency groups and the retention intervals. Second, the
interaction between feedback timing and the retention interval was significant on the

Downloaded from [Link] at University of Otago Library on September 7, 2014


14

Table 5.  Results of three-way ANOVAs for the posttest scores.

Productive posttest Receptive posttest


  Strict scoring Sensitive scoring Strict scoring Sensitive scoring

  df F p partial η2 df F p partial η2 df F p partial η2 df F p partial η2


Feedback timing 1, 94 0.06 .814 .00 1, 94 0.17 .678 .00 1, 94 0.11 .741 .00 1, 94 0.22 .641 .00
Feedback timing × 3, 94 0.78 .511 .02 3, 94 2.00 .119 .06 3, 94 1.51 .217 .05 3, 94 0.70 .553 .02
Retrieval frequency
Feedback timing × 2, 188 0.60 .550 .01 2, 188 2.07 .129 .02 2, 188 4.87 .009 .05 2, 188 3.69 .027 .04
Retention interval
Feedback timing × 6, 188 0.64 .702 .02 6, 188 0.93 .472 .03 6, 188 0.33 .921 .01 6, 188 0.87 .518 .03
Retrieval frequency ×
Retention interval

Downloaded from [Link] at University of Otago Library on September 7, 2014


Language Teaching Research 
Nakata 15

Table 6.  Results of simple main effect of feedback timing on the receptive posttest.

Strict scoring Sensitive scoring


Retention
interval F p partial η2 CI of diff d F p partial η2 CI of diff d
Immediate 1.52 .221 .02 [–0.40, 0.10] 0.09 2.43 .123 .02 [–0.44, 0.05] 0.12
1 week 4.81 .031 .05 [0.03, 0.51] 0.14 1.61 .208 .02 [–0.08, 0.40] 0.08
4 weeks 0.05 .827 .00 [–0.30, 0.25] 0.02 0.71 .403 .01 [–0.37, 0.16] 0.06

Note. CI of diff: 95% confidence intervals of difference in the mean gains from the pretest to the posttest
between the immediate and delayed feedback conditions. df = (1, 97). Effect sizes (d) of 0.20, 0.50, and 0.80
are indicative of small, medium, and large effects, respectively (Cohen, 1988).

receptive posttest with both strict and sensitive scoring, but not on the productive posttest
irrespective of the scoring system. Third, none of the other interactions involving feed-
back timing were significant on any of the dependent variables.
As the interaction between feedback timing and the retention interval proved signifi-
cant on the receptive posttest, the simple main effect of feedback timing was tested to
investigate where the significant differences lay. Testing the simple main effect of feed-
back timing allows us to examine whether any difference existed between immediate and
delayed feedback at each retention interval. The results of the simple main effect tests are
summarized in Table 6. When testing the simple main effect, the type I error rate was
controlled using Ryan’s method for multiple comparisons. The table shows that when
collapsed across the four retrieval frequency groups, immediate feedback significantly
outperformed delayed feedback with strict scoring on the 1-week delayed receptive post-
test (p = .031). However, despite statistical significance, only a very small effect size was
observed (d = 0.14, partial η2 = .05), and the difference in the mean gains between imme-
diate (6.09) and delayed feedback (5.83) was small. The finding is also supported by the
relatively narrow confidence intervals of difference in the mean gains: [0.03, 0.51],
which indicates that there is a 95% chance that the difference in the mean gains between
the two feedback conditions lay somewhere between 0.03 and 0.51. The simple main
effect of feedback timing was not significant in all other cases (p ≤ .123), yielding very
small effect sizes (0.02 ≤ d ≤ 0.12, partial η2 ≤ .02). Overall, although a statistically sig-
nificant effect was found, given the small effect size and narrow confidence intervals of
difference, it might be reasonable to assume that feedback timing had little effect on
posttest results irrespective of the retrieval frequency or retention interval. The statistical
significance may be partially due to the relatively large cell size (98).

VI Discussion
The purpose of this study was to identify the optimal feedback timing for L2 vocabulary
learning. Unlike some previous studies, the immediate and delayed feedback conditions in
this study were controlled for lag to test. In order to examine a possible relationship between
the delay-retention effect and the frequency of errors during learning, retrieval frequency
was also manipulated. The results of this study suggested that when lag to test is controlled,
feedback timing may have little effect on learning regardless of the frequency of errors

Downloaded from [Link] at University of Otago Library on September 7, 2014


16 Language Teaching Research 

during the learning phase. Although the experimental settings in this study were closer to
those in laboratory studies, which are more likely to find the superiority of delayed over
immediate feedback than classroom studies (Butler et al., 2007; Kulik & Kulik, 1988;
Roediger et al., 2010; see Section II), a significant delay-retention effect was not observed.
Taken together, the results suggest that the benefits of delayed feedback at the intervals
used in this study may be limited as far as L2 vocabulary learning is concerned.
It should be noted that the immediate posttest scores of the Retrieval 5 and 7 groups
neared the ceiling in both feedback conditions (Tables 3 and 4). Hence, the lack of a
significant feedback timing effect in these two groups on the immediate posttest may be
partly ascribed to a possible ceiling effect. On the delayed posttests, however, neither a
ceiling nor floor effect was observed. The lack of statistical significance on the delayed
posttest, therefore, seems to indicate that feedback timing may have little effect on long-
term retention. Because from a pedagogical perspective, scores on the delayed posttest
may be more important than those on the immediate posttest, the present study may
nonetheless have pedagogical value despite the possible ceiling effect in the Retrieval 5
and 7 groups on the immediate posttest.

1  Theoretical implications
As discussed in Section II, there exist conflicting views about the effectiveness of imme-
diate and delayed feedback. On one hand, delayed feedback is considered more effective
because it may introduce larger spacing as well as cause less interference than immediate
feedback. On the other hand, according to the theory of errorless learning, feedback
needs to be given immediately after retrievals because otherwise, learners’ errors might
be consolidated. The present study did not find any significant difference between the
two types of feedback. The lack of a significant feedback timing effect was caused pos-
sibly because the beneficial effects of delaying feedback (larger spacing and less interfer-
ence) might have been offset by the risk of not correcting an error immediately. It should
be noted, however, that a significant delay-retention effect was not observed in this study
although the proportion of errors (16.6%–36.2%) was lower than in Metcalfe et al.’s
(2009) experiments (40%–61%). The interaction between feedback timing and the
retrieval frequency group was not significant either although repeated retrieval was asso-
ciated with better learning phase performance. The findings seem to be inconsistent with
the observation that when learners make few errors during learning, delayed feedback
may be more effective because of the distributed practice effect (Metcalfe et al., 2009).
The results may suggest that the findings of Metcalfe et al.’s experiments may be at least
partly attributed to the difference in the age of participants (i.e., grade school vs. college
students) rather than the difference in the frequency of errors per se.
Alternatively, the conflicting results might be due in part to at least two methodologi-
cal differences between the present study and Metcalfe et al. (2009). First, while Metcalfe
et al. investigated the learning of L1 vocabulary, the present study looked into L2 vocab-
ulary learning. Second, although delayed feedback was given after a delay of 1 day or
longer in Metcalfe et al., it was provided 92.77 seconds on average after the response in
this study (see Section V.2). Because larger spacing generally leads to better long-term
retention than shorter spacing (distributed practice effect; e.g., Cepeda et al., 2006, 2008,

Downloaded from [Link] at University of Otago Library on September 7, 2014


Nakata 17

2009), a significant delay-retention effect might have been observed in this study if the
delayed feedback condition had used much larger spacing. Future research may provide
delayed feedback after a longer delay in order to explore this possibility.

2  Pedagogical implications
Pedagogically, the results imply that either immediate or delayed feedback may be used
in L2 teaching and learning. Because learners generally prefer immediate over delayed
feedback, the use of immediate feedback may have a positive effect on learners’ motiva-
tion and might be more desirable. Karpicke, Smith, and Grimaldi (2009), for instance,
surveyed 103 American college students and found that when learning from paper-based
flashcards, 91% of them confirm the correct answer immediately after retrieval attempts,
which is equivalent to receiving immediate feedback. Feedback timing may also be
determined based on practical considerations. For instance, immediate feedback may be
used when learning vocabulary from paper-based flashcards because it may be easier to
implement manually than delayed feedback. Immediate feedback may also be preferable
when correcting written work because it may speed up the revision process, creating
more opportunities for learners to revise their writing and receive teacher feedback
(e.g., Evans et al., 2010). In conversational interaction, however, immediate feedback
may be avoided because providing corrective feedback during communicative activities
may potentially inhibit learners’ willingness to speak up (Rolin-Ianziti, 2010). Delayed
feedback may also be more feasible for tests that require marking by teachers.
Nonetheless, because this study investigated only L2 vocabulary learning in a CALL
environment, the issue of whether the results of this study also apply to the learning of
other aspects of L2 (e.g., syntax or phonology) or other instructional activities (e.g., face-
to-face interaction or writing) awaits future research.

VII  Conclusions and directions for further research


The purpose of this study was to identify the optimal feedback timing for L2 vocabulary
learning while controlling the lag to test and manipulating the frequency of errors during
learning. The results of this study suggested that when lag to test is controlled, feedback
timing may have little effect on learning regardless of the frequency of errors during the
learning phase. At the same time, the present study may suffer from some limitations.
First, because both immediate and delayed feedback conditions in this study included the
final review, it is possible that the effects of feedback timing might have been overshad-
owed by the final review. Future research may manipulate the presence or absence of the
final review in order to explore this possibility. (The author is grateful to an anonymous
reviewer for pointing this out.) Another limitation may be a possible ceiling effect on the
immediate posttests in the Retrieval 5 and 7 groups (Tables 3 and 4). The nonsignificant
feedback timing effect in these two groups on the immediate posttest may be partly
ascribed to a possible ceiling effect. Third, the retention interval was a within-participant
variable in this study, and each participant sat posttests at three retention intervals: imme-
diately, 1 week, and 4 weeks after the treatment. Because correct responses in the pro-
ductive test were used as cues in the receptive test, and correct responses in the receptive

Downloaded from [Link] at University of Otago Library on September 7, 2014


18 Language Teaching Research 

test were used as cues in the productive test, earlier tests might have affected perfor-
mance on later tests. In future research, in order to reduce potential learning effects
from posttests, the retention interval may be manipulated between participants (e.g.,
Cepeda et al., 2008; Karpicke & Roediger, 2007; Rohrer et al., 2005).

Acknowledgements
This article is based on part of the author’s doctoral dissertation, which was submitted to Victoria
University of Wellington in 2013. I am very grateful to Stuart Webb, Paul Nation, Rod Ellis, Jan
Hulstijn, and Kazuya Saito for their invaluable advice and Tomohiro Tsuchiya for his cooperation
with data collection.

Funding
This work was supported by Faculty Research Grants from Victoria University of Wellington
(grant number 109605).

Note
1. In order to counterbalance the effects of target items across participants, each retrieval fre-
quency group needs to consist of an equal number of participants from Subgroups X and Y
(Table 1). The Retrieval 7 group, however, consisted of 13 Subgroup X and 14 Subgroup Y
participants. One Subgroup Y participant was excluded from analysis so that the Retrieval 7
group would have an equal number of participants (13) from both subgroups. The participant
was excluded so that it would minimize effects on the mean posttest scores of the subgroup
from which the student was dropped (for details, see Nakata, 2013).

References
Barcroft, J. (2007). Effects of opportunities for word retrieval during second language vocabulary
learning. Language Learning, 57, 35–56.
Butler, A.C., Karpicke, J.D., & Roediger, H.L. (2007). The effect of type and timing of feedback
on learning from multiple-choice tests. Journal of Experimental Psychology: Applied, 13,
273–281.
Carpenter, S.K., & Vul, E. (2011). Delaying feedback by three seconds benefits retention of face–
name pairs: The role of active anticipatory processing. Memory & Cognition, 39, 1211–1221.
Cepeda, N.J., Coburn, N., Rohrer, D., Wixted, J.T., Mozer, M.C., & Pashler, H. (2009). Optimizing
distributed practice: Theoretical analysis and practical implications. Experimental Psychology,
56, 236–246.
Cepeda, N.J., Pashler, H., Vul, E., Wixted, J.T., & Rohrer, D. (2006). Distributed practice in verbal
recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132, 354–380.
Cepeda, N.J., Vul, E., Rohrer, D., Wixted, J.T., & Pashler, H. (2008). Spacing effects in learning:
A temporal ridgeline of optimal retention. Psychological Science, 19, 1095 –1102.
Cohen, J. (1988). Statistical power analysis for the behavioral sciences. 2nd edition. Hillsdale, NJ:
Lawrence Erlbaum.
Cull, W.L., Shaughnessy, J.J., & Zechmeister, E.B. (1996). Expanding understanding of the
expanding-pattern-of-retrieval mnemonic: Toward confidence in applicability. Journal of
Experimental Psychology: Applied, 2, 365–378.
DeKeyser, R. (2007). Situating the concept of practice. In R. DeKeyser (Ed.), Practice in a sec-
ond language: Perspectives from applied linguistics and cognitive psychology (pp. 1–18).
Cambridge, UK: Cambridge University Press.

Downloaded from [Link] at University of Otago Library on September 7, 2014


Nakata 19

Ellis, R., & He, X. (1999). The roles of modified input and output in the incidental acquisition of
word meanings. Studies in Second Language Acquisition, 21, 285–301.
Evans, N.W., Hartshorn, K.J., McCollum, R.M., & Wolfersberger, M. (2010). Contextualizing
corrective feedback in second language writing pedagogy. Language Teaching Research, 14,
445–463.
Karpicke, J.D., & Roediger, H.L. (2007). Expanding retrieval practice promotes short-term reten-
tion, but equally spaced retrieval enhances long-term retention. Journal of Experimental
Psychology: Learning, Memory, & Cognition, 33, 704–719.
Karpicke, J.D., Smith, M.A., & Grimaldi, P.J. (2009). Learning with flashcards: You’re prob-
ably doing it wrong. Poster presented at the Eighty-first Annual Meeting of the Midwest
Psychological Association, Chicago, IL, USA.
Kornell, N. (2009). Optimising learning using flashcards: Spacing is more effective than cram-
ming. Applied Cognitive Psychology, 23, 1297–1317.
Kulhavy, R.W., & Anderson, R.C. (1972). Delay-retention effect with multiple-choice tests.
Journal of Educational Psychology, 63, 505–512.
Kulik, J.A., & Kulik, C.-L.C. (1988). Timing of feedback and verbal learning. Review of
Educational Research, 58, 79–97.
Lee, I. (2013). Research into practice: Written corrective feedback. Language Teaching, 46,
108–119.
Li, S. (2010). The effectiveness of corrective feedback in SLA: A meta-analysis. Language
Learning, 60, 309–365.
Lyster, R., Saito, K., & Sato, M. (2013). Oral corrective feedback in second language classrooms.
Language Teaching, 46, 1–40.
Metcalfe, J., & Kornell, N. (2007). Principles of cognitive science in education: The effects of
generation, errors, and feedback. Psychonomic Bulletin and Review, 14, 225–229.
Metcalfe, J., Kornell, N., & Finn, B. (2009). Delayed versus immediate feedback in children’s and
adults’ vocabulary learning. Memory & Cognition, 37, 1077–1087.
Mory, E.H. (2004). Feedback research revisited. In D.H. Jonassen (Ed.), Handbook of research
on educational communications and technology. 2nd edition (pp. 745–783). Mahwah, NJ:
Lawrence Erlbaum.
Nakata, T. (2013). Optimising second language vocabulary learning from flashcards (Unpublished
doctoral dissertation). Victoria University of Wellington, New Zealand.
Nation, I.S.P. (2006). How large a vocabulary is needed for reading and listening? The Canadian
Modern Language Review, 63, 59–82.
O’Neill, M., Rasor, R.A., & Bartz, W.R. (1976). Immediate retention of objective test answers as a
function of feedback complexity. The Journal of Educational Research, 70, 72–75.
Pashler, H., Cepeda, N.J., Wixted, J.T., & Rohrer, D. (2005). When does feedback facilitate
learning of words? Journal of Experimental Psychology. Learning, Memory & Cognition,
31, 3–8.
Phye, G., & Baller, W. (1970). Verbal retention as a function of the informativeness and
delay of informative feedback: A replication. Journal of Educational Psychology, 61,
380–381.
Quinn, P. (2013). The effects of altering the timing of corrective feedback. Paper presented at the
American Association for Applied Linguistics Conference, Dallas, TX, USA.
Roediger, H.L., Agarwal, P.K., Kang, S.H.K., & Marsh, E.J. (2010). Benefits of testing memory:
Best practices and boundary conditions. In G.M. Davies, & D.B. Wright (Eds.), New frontiers
in applied memory (pp. 13–49). Brighton, UK: Psychology Press.
Rohrer, D., Taylor, K., Pashler, H., Wixted, J.T., & Cepeda, N.J. (2005). The effect of overlearning
on long-term retention. Applied Cognitive Psychology, 19, 361–374.

Downloaded from [Link] at University of Otago Library on September 7, 2014


20 Language Teaching Research 

Rolin-Ianziti, J. (2010). The organization of delayed second language correction. Language


Teaching Research, 14, 183–206.
Sheen, Y. (2012). The timing of corrective feedback and L2 learning. Paper presented at the
Second Language Research Forum Conference, Pittsburgh, PA, USA.
Skinner, B.F. (1954). The science of learning and the art of teaching. Harvard Educational Review,
24, 86–97.
Swindell, L.K., & Walls, W.F. (1993). Response confidence and the delay retention effect.
Contemporary Educational Psychology, 18, 363–375.
Webb, S.A. (2012). Depth of vocabulary knowledge. In C. Chapelle (Ed.), Encyclopedia of Applied
Linguistics (pp. 1656–1663). Oxford, UK: Wiley-Blackwell.

Author biography
Tatsuya Nakata received a PhD from Victoria University of Wellington. His research interests are
vocabulary acquisition and CALL. His research has appeared in the Encyclopedia of Applied
Linguistics, Language Teaching Research, and Computer Assisted Language Learning. He is a
recipient of the EuroSLA Doctoral Award and the EUROCALL Research Award.

Downloaded from [Link] at University of Otago Library on September 7, 2014

Common questions

Powered by AI

Important methodological considerations include controlling for the lag to test, ensuring consistent intervals between feedback and post-testing across conditions, accounting for the setting of the experiment (laboratory vs. classroom), and manipulating the frequency and context of errors made during the learning phase. These factors can significantly influence observed outcomes and should be carefully controlled to accurately assess the effects of feedback timing .

Three primary explanations for inconsistent findings in studies comparing immediate and delayed feedback include (1) the differences in experimental settings (laboratory vs. classroom), which may lead participants to process feedback differently, (2) confounding factors such as the lag to test, where delayed feedback is improperly equated to shorter intervals, and (3) the potential variance in learner error rates, where studies did not thoroughly manipulate or assess the effect of errors during learning .

The two main theories that support the superiority of delayed feedback over immediate feedback are the distributed practice effect and interference theory. The distributed practice effect suggests that larger spacing between learning and feedback leads to better long-term retention than shorter spacing. Delayed feedback, due to its inherent time gap, facilitates this effect, thus enhancing learning . Interference theory suggests that delayed feedback helps in avoiding confusion between errors and correct responses by the time feedback is given, reducing the likelihood of learning false information .

Cognitive psychologists view feedback following learner responses as crucial for both correcting errors and reinforcing correct performances. Feedback serves to provide the correct answer after a learner's response, which is especially pertinent for language learning. According to the theory of errorless learning, immediate feedback is essential to prevent the consolidation of errors. On the other hand, interference theory supports delayed feedback to reduce the confusion between errors and correct information .

Empirical studies suggest that lag to test, setting of the study (laboratory versus classroom), and the quantity of learner errors can confound interpretations of feedback timing effects. For example, differences in posttest timing relative to feedback delivery (lag to test) were not always controlled, leading to discrepancies in retention outcomes. Additionally, laboratory settings may emphasize feedback effects differently than classroom settings. The error rate during learning phases can also affect how feedback timing impacts learning outcomes .

Metcalfe et al.'s findings imply that delayed feedback may be more effective when learners make fewer errors during the learning phase. This suggests that delayed feedback could enhance retention by minimizing interference from incorrect responses, particularly in contexts where errors are infrequent. Furthermore, experiments should consider the frequency and types of errors when evaluating feedback timing to ensure accurate conclusions about their respective impacts on learning .

The setting of an experiment, whether it is a laboratory or classroom, can influence the effectiveness of feedback timing. Studies suggest that the delay-retention effect is more frequently observed in laboratory settings compared to classroom settings. This difference may arise because participants could process feedback differently depending on the setting, with laboratory environments potentially offering more controlled conditions that highlight the effects of delayed feedback .

The 'delay-retention effect' posits that delayed feedback can be more beneficial than immediate feedback due to mechanisms like the distributed practice effect and interference theory. The distributed practice effect suggests that feedback delivered after a delay provides greater spacing practice, which enhances long-term retention. Interference theory supports delayed feedback by suggesting that temporal distance from errors helps to minimize the confusion between incorrect initial responses and correct information given later, thereby facilitating better retention .

Immediate feedback might prevent the consolidation of learner errors by correcting them instantly, thereby avoiding reinforcing incorrect knowledge. Conversely, delayed feedback may allow learners' initial errors to be forgotten by the time correction occurs, thus theoretically preventing interference with the correct information provided later. However, if errors are remembered, delayed feedback could inadvertently reinforce misunderstandings due to a lack of immediate correction .

The timing of feedback might have limited effects when retrieval frequency is manipulated because the interaction between feedback timing and retention interval might not significantly impact learning outcomes. When retrieval frequency varies, it can alter the likelihood of errors and interactions with the feedback timing, but empirical results have shown that there's no significant main effect of feedback timing on learning, regardless of retrieval frequency .

You might also like