2.
LOOKING AT INTERLANGUAGE DATA
30
gather additional data to test whether or not
Thus, we may want to potential for a given type of data set
the
these descriptive rules exhaustanalyze data, either our own or data from
In generalthen, when we always ask this
question: What else
literature, we want to
the published lite bythe data presented? We no
that is not shown
do we want to findout datacan be collected.
turnto ways in which
2.3. DATA COLLECTION
we have selected here to discuss is only a vo
what
We point out that data-collection further point out that
methods. We
small number of their
although not all, of second language research methods have
much, other disciplines, notably linguistics
origins in research methods from
language acquisition, sociology, and psychology. It is also impor.
child research questions will often
tant to be aware of the fact that particular
lead to a particular research nmethodology. language
There are two predominant ways of collecting data in second
research: one is longitudinal and the other, cross-sectional.
Longitudinal studies are generally case studies (although not always,
as we willsee later), with data being collected from a single learner (or
at least a small number of learners) over a prolonged period of time. The
frequency of data collection varies. However, samples of a learner's lan
guage are likely to be collected weekly, biweekly,or monthly.
With regard to longitudinal studies, there are four characteristics
to be discussed: (a) number of subjects and time frame of data collec
tion, (b) amount of descriptive detail, (c) type of data, and (d) type of
analysis.
Typical of longitudinal studies is the detail provided on a learner's
speech, on the setting in which the speech event occurred, and on other
details relevant to the analysis of the data (e.g.. other
ticipants and their relationship with the learner). The conversational par
following 1s a
description one longitudinal study, reported in Hakuta (1974a).
of
The data come from a longitudinal study of
of English as a second language by a the untutored acquisition
five-year-old
Her family came tothe United States
for a
Japanese girl[Uguisu
her father was a visiting scholar at period of two years Wnte
in North Cambridge in a Harvard. and thev took residence
p. 287) working-class neighborhood. (Hakuta,1*
Hakuta went on to describe that the
her primary source of children Uguisu played witn
of her school
activity, language
saying thatinput." He also included a description
she went to " kindergarten for
DN
31
two hours every day, and later
in English syntax. Most ot her elementary school, but with no tutoring
class at school" (Hakuta, 1974a, [Link]
287).
friends were in her same
In most longitudinal studies
data come from spontaneous (particularly those that are case studies),
speech. This does not mean that the
researcher does not set up aconversation to generate
data. It simply means that longitudinal studies do not afitparticular type of
into the experi
mental paradigm (tobe discussed) of control group, experimental
counterbalancing, and so forth. An important methodological question group,
that arises in connection with spontaneous speech data collection is:
How
can aparticular type of data be generated through spontaneous speech?
While there cannot be a 100% guarantee that certain designated interlan
guage forms will appear, the researcher can ask certain types of questions
in the course of data collection that will likely lead to specific structures.
For example, if someone were interested in the interlanguage develop
ment of the past tense, learners could be asked during each recording ses
sion to tell about something that happened to them the previous day.
Analyses of data obtained through longitudinal studies (and particu
larly in case studies) are often in the form of descriptive qualitative com
ments or narrative expositions. While quantification of data may not be
the goal of such studies, the researcher may report the frequency of occur
rence of some form. In the reporting of results from longitudinally col
lected data, there are likely to be specific examples of what a learner said
and how their utterances are to be interpreted.
This type of data is highly useful in determining developmental trends
influences
as well as in interpreting various social constraintsand input
a major
(see Chapter 10) on the learner's speech. On the other hand,
longitudinal study
drawback concerns the time involved. Conducting a
well as in tran
requires time in collecting data at regular intervals, as extensive detail
scription of the speech, which is ideally accompanied bythe speech event
on the social, personal, and physical setting in which
of generalizability.
took place.A second drawback is related to the lack number of learn
in the
Given that longitudinal studies are often limited results. It is difficult to
generalize the
ers investigated, it is difficult to
know with any degree of certainty whether the results obtained are appli
only to the one or two learners studied, or whether they are indeed
cable difficulty with spon
characteristicof a wide range of subjects. Another
taneously produced longitudinal data, and
perhapsthe most serious one,
there is no way of probing their
is that when learners produce a form,
they have produced spontaneously
knowledge any further than whatparticularly
section 2.1). This is the case if the researchers
(see data sets in researchers have not gen
themselves have not collected the data or if the gathering infor
to
erated specific hypotheses and are not predisposed
32 INTERLANGUAGE DATA
2. LOOKING
AT
forms of speech. For example, if in a particular set
mation about specific only produces the present
learner tense
of Spontaneously elicited data, a is allthat learner knows? We
of verbs, does that mean that
that
actually present, cannot
basis of what is
interpret data only on theforms means lack of knowledge of forms
because
we
do not know ifabsence of
data-collection method involves cross-sectional stud-
of
fouridentifiable learners and that
tvpe
Asecond characteristics are
ies. Here, too, there are
(a) number of
time generally
frame of data
with such studies:
associated
(c) descriptive detail, and (d) analysis of data.
ollection.(b)tvpe of data, gathered
Across-sectional study generally consists of data
the idea
from
being that
alarge
point in time, We are
number oflearners at a single which isused to piece together actal
able to see aslice of development,
development.
case studies, which are based primarily on spontaneous speech
Unlike based on controlled out.
cross-sectionaldata are often (but not always)
the format is one in which a researcher is attempting to gath.
put. That is, hypothesis. The data, then, come
er databased on a particular researchprespecified task.
from learners´ performance on some what we have seen
The type of background information differs from
with longitudinal studies. Participants are not identified individually,
certain amount of
nor is detailed descriptive information provided. A
in Table 2.5.
background data is likely to be presented in tabular form, aslearners,
Because cross-sectional data involve large numbers of there
istypically an experimental format to the research, both in design and
in analysis. Results tend to be more quantitative and less descriptive than
in longitudinal studies, with statistical analyses and their interpretation
being integral parts of the research report.
One can use a cross-sectional design to create a pseudolongitudinal
study. In such a study, the emphasis, like that of a longitudinal study, 15
on language change (i.e., acquisition), with data being collected at a sin
gle point in time, but with different proficiency levels represented. For
example, if one is investigating the acquisition of the progressive,
one would want to know not just what learners can do at a particula
Point in time (because the question involves acquisition and not static
TABLE 2.5
Typical Data Presentation in aCross-Sectional Design
Language No. of participants Gender Age Proficiency
Arabic 24
13F;11M 22-26 8Beg/8lnt/8Adv.
Spanish 24
12F;12M 23-28 12Beg/12Adv
Japanese 24
11F;13M 21-23 20Beg/4Adv
23. DATA COLLECTION 33
knowledge), but also what happens over a period of time. One way of
gatheringsuch data is through a longitudinal study, carefully noting
every instance in which the progressive is and is not used (see data set
I). Another way of gathering information about linguistic development
would be totake alarge group of learners, at three specified proficiency
levels -let's say, beginner, intermediate, and advanced-and give each
groupthe same test. The assumption underlying this method is that com-
paring these three groupS would yield results similar to what would be
found if we looked at a single individual over time. The extent to which
controversial.
this assumption is warranted is
One advantage to across-sectional approach is the disadvantage of
the
longitudinal data: Given that there are large numbers of learners in
wider
former, it is more likely that the results can be generalized to a
group. The disadvantage is that, at least in the second language acquisi-
about the learners
tion literature, there is often no detailed information
production was
themselves and the linguisticenvironment in which
appropriate inter
elicited. Both types of information may be central to an much a problem
not so
pretation of the data. This criticism, of course, is
resultshave been report
with the research approach, asit is with the way
ed in the literature.
often associated with descrip
As noted earlier, longitudinal data are
and pseudolongitudinal data,
tive (or qualitative) data. Cross-sectionalwith quantitative or statistical
associated
on the other hand, are often
conduct statistical analyses on longi
measures. However, one can easily descriptive analyses of cross
tudinaldata and one can easily provide
mistake to assume that longitudinal
sectional data. It is furthermore a
be able to put together a profile of
datacannot be generalized. One maystudies.
learners based on many longitudinalone type of data-collection procedure
Why would a researcher select choice is
What is most important in understanding this
over another?
of the relationship between a research question and
the understanding While there may not always be a 1:1relationship,
research methodology. questions and certain kinds of external pres
there are certain kinds of selecting one type of approach to research
sures that would lead one into information about
example, one wanted to gather
over another. If, for learn to apologize in asecond language, one
how non-native speakers period of time, noting instances of apolo
over a setting). On
could observe learners experiment or in a naturalistic
controlled
gizing (either in a
could use across-sectional approach by setting up a
the other hand,one
groups of second language speakers what they
situation and asking large
forces production, the former waits until it hap
wouldsay. The latter more
argue that the former is "better" in that it
pens. While many would wait for
clear that one might have to
dccuratelyreflects reality, it is also
2. LOOKING AT
34
INTERLANGUAGE DATA
aconsiderable amount of time before getting any information that
be useful in answering the original research question. Thus, the would
approach.
cies of the situation lead a researcher to a particularparadigmatic exigen-
It wouldlbe a mistake to think of any of these bound.
anes as rigid: it would also be a mistake to associate longitudinal stud-
ies with naturalistic data collection. One can conduct alongitudinal
Ispeakers; one can also collect data study
with large numbers of
longitudinally
using an evpenmental format. In astudy on relative clauses, Gass (1979a,
points in time (at
197b) gathered data from 17learners at six typical
study itselfsatisfied the definition of monthly
intervals). Thus the
tudinal. However, iit did not satisfy the definition of a case study, as it longi-
did not involve detailed descriptions of spontaneous(which speech. On the
involved
other hand, given the experimental nature of the study
forced production of relative clauses), it more appropriately belongs in
the category of cross-sectional. Inother words, the categories we have
described are only intended to be suggestive. They do not constitute rigid
categories; there is much flexibility in categorizing research as being of
one type or another.
We next take a look at two studies to give an idea of the range of data
that has been lookedat in second language acquisition.
Firstis astudy by Kumpf (1984), who was interested in understand
ing how nonnative speakers expressed temporality in English (see also
Chapter 6). One way to gather such information is to present learners
with sentences (perhaps with the verb form deleted) and ask them to fill
in the blank with the right tense. This, however, would not give infor
mation about how that speaker uses tense in a naturalistic environment.
Only a long narrative would give that information. Following is the text
produced by the native speaker of Japanese in Kumpf'sstudy. The learn
er is awoman who had lived in the United States for 28 vears at the time
of taping. For the purposes of data collection, she was asked to
a narrative account. produce
First time Tampa have a tornado
Was about seven forty-five
come to.
Bob go to work, nIwas inna
bathroom.
And...a ...tornado come shake
Door was flyin open, I was scared. everything.
Hanna was sittin in window
Hanna is a little dog.
French poodle.
Icall Baby.
Anyway, she never wet bed, she never wet
anywhere.
2.3. DATA COLLECTION 35
But she was so scared and crying' run to the bathroom, come to me, and
she tinkle, tinkle, tinkle all over me.
She was so scared.
Isee sonmebody throwin abrick onna trailer
wind was blowin so hard
ana light.. outside strvet light was on
oh l was eally scared.
Anden seond stop
Sol try to opendoor
lould not open
Isay, "Oh, my God. What's happen?"
Ilook window. Awning was gone. (pp. 135-136)
With regard to temporality, there are a few conclusions that Kumpf
draws from these data. One is that there is a difference between scene
setting information (i.e., that which provides background information to
functions are
the story) and information about the action-line. These two
reflectedin the use or lack of use of the verb to be with the progressive.
In the scene-setting descriptions, descriptive phrases (wind was blowin,
door was flyin open, Hanna wassittin in window), the past form of to be is
apparent. However,
How when this speaker refers to specific events, no form
brick onna trailer).
of the verb to be was used (somebody throwin
frequency with which certain
A second finding from this study is the
tense. The copula (to be) is tensed 100%
types of verbs are marked with
time; verbs expressing the habitual past (used to) are tensed 63%
of the (e.g., try) 60% of the time.
time, and continuous action verbs
of the elicited through a controlled obser
Could this information have been
The first set of results (determining the differences
vation procedure? action-based information), probably not; the
between scene-setting and
(frequency of verb tenses), probably yes. In the first instance,
second set
an experimental paradigm that would have elicit
it is difficult to imagine the second, one could more easily imagine set
ed such information. In sentences) in which the same
using isolated
ting up a situation (evenobtained.
results would have been to one speaker, one would like to know
limited
Because these data are phenomenon or not. Results from studies such
whether this isa general attempting to force production from larger
by
as this can be verifiedHowever, the fact thateven one speaker made
numbers of learners. verb to be and its nonuse suggests that
the
distinction between the use of
generalization. One question at the forefront of much
that
this is a possible ILacquisition research is: Are the language systems sys
second language what is found in natural language
with
learners create consistent
36 2. LOOKING AT INTERLANGUAGE DATA
tems? That is, what are the boundaries of human languages? Given the
primacy of questions such as these, the fact of a singleindividual creat-
verb atoparticular
vto ditfeentiate between (in
l generalization this case using or not using the
two discourse functions) is enough
to provide initial answers.
Tet'sonsider a study that yathers data within an experimental para.
whowere concerned with the
digm. This isone by Gass and Ard (984),
various meanings of the progres
knowledge that learners have about
database came from responses by 139 learners to four differ.
[Link] asked to judge the acceptability
the first task, learners were
ent tasks In meanings of the progressive, as in 2.
ontaining the various
ot sentenes
55 and 2-56:
York tomorrow.
(2-55)John is traveling to New tomorrow.
(2-56) John travels to New York
task, the sentences were enmbedded in short conversations
In the second
mother in a hurry.
(2-57) Mary: Ineed to senda package to my
Jane: Where does she live?
Mary: In New York. traveling to
Jane: Oh, in that case John can take it. John is
New York tomorrow.
In the third task, there were again isolated sentences, although these were
in groups of five. Again, acceptability judgments were asked for.
(2-58) The ship sailed to Miami tomorrow.
(2-59) The ship is sailing to Miami tomorrow.
(2-60) The ship will sail to Miami tomorrow.
(2-61) The ship sails to Miamni tomorrow.
(2-62) The ship has sailed to Miami tomorrow.
In the fourth task, the subjects were given a verb form and asked to write
as long asentence as possible including that form.
What was found was that there was an order of preference of differ
ent meanings for the progressive. For example, most learners ordered
the various meanings of the progressive so that the most common was
the progressive to express the present (John is smoking American cigarettes
now); the next was the progressive to express futurity (John is traveling to
New York tomorrow); the next to express present time with
ception (Dan is seeing better now); the next with verbs such asverbs of per
connect (1he
2.4. DATA ELICITATION 37
new bridge is connecting Detroit and Windsor), and finally, with the copula
(Mary is leing in Chicago now). 1The authors used this information to gain
information about the development of meaning, including both proto
typical meanings and more extended meanings. Throughspontaneous
speech alone (whether a case study or not), this would not have been
possible. Only aforced-choice data task would elicit the relevant infor
mation. One should also note that controlled observations of sponta
neousspeech may underestimate the linguistic knowledge of a learner,
particularly in those cases where the task is insensitive to the linguistic
structure being elicited or is too demanding.
2.4. DATA ELICITATION
are numerous ways of eliciting second language data. As we men
There in other disciplines.
tioned earlier, many, but not all, have their origins
suggestiveof the
This section is not intendedto be inclusive. Rather, it is
and data-elicitation methods that have been used in sec
kinds of data
ond language studies.
Tests
2.4.I. Standardized Language
often used as asource for second language
is not test is
Thistype of instrument most common type of standardized work
primarily because the exception is the
(An
data
and does not yield productive [Link] 5).
objective 1992] discussed in
of Ard & Homburg (1983, language tests are often used as
standardized given research
On the other hand,proficiencylevels. For example, in aTOEFL (Test of
gauges for measuring those who have a
study, advanced
learners may be
score above acertain level. Even with
Language) accepted cutoff point
English as aForeignhowever, there is no absolute fact, one difficulty
standardized tests, beginner, and so forth.
In
advanced, intermediate,
because there is no accepted cutoff
another's
SLAstudies is thatcategory may correspond to litera
for
in comparing
one researcher's advanced(1994), based on asurvey of the (a)
point, Thomas proficiency:
intermediate category. common ways of assessing first semester of
ture, has
identified four
(b) institutionalstatus (e.g., (d) standardized
impressionistic judgments,
specific
research-designed test,
proficiency, the field
second-year French), (c) ways to measure studies. This is
Because there are SO manydifficulty in comparing
tests, considerable
with (see Chapter 4), in which
SLA isleft acquisition
of child language
unlike the field of