CHAPTER
The Experimental Method
Any timeyou use phrases like: "On average, I cycle about 100 miles a
week"or"We can expect a lot ofrain atthis time ofyear"or"The earlier
you start revising, the betteryou are likely to do inthe exam"you are
making a statistical statement, even though youmay have performed no
calculations. (Rowntree, 1981, p. 13)
INTRODUCTION AND OVERVIEW
In this chapter, we will look more closelyat the experimental method, one of the
two 'pure' research paradigms (Grotjahn, 1987). In Grotjahn's terms, this para
digm involves (1) experimental designs, (2) quantitative data, and (3) statistical
analyses. The experimental method is basically a collection of research designs,
guidelines for using them, principles and procedures for determining statistical
significance, and criteria for determiningthe quality of a study. The experimen
tal method is part of the psychometric tradition, and it is also referred to as the
scientific method. For some researchers, the experimental method is the premier
method, all others being 'ground clearing' operations, that is, preliminary data
collection and interpretation exercises to prepare por a formal experiment.
We will begin this chapter by adding to our earlier discussion of possible
confounding variables. Then we will add more research designs to those you
have read about in earlier chapters. We will systematize this discussion by ana
lyzing and exemplifying the research designs, dividing them into classes, and
explaining their relationships to one another. T[hen we will use an extended
example to look at the issue of extrapolating from samples to populations.
This extrapolation is based on the logic of the normaldistribution, whichwill be
discussed as well.
83
As you read this material, keep in mind that different forms of" research have
different cultures. The experimental method has one ofthemost strictly codified
sets ofvalues and procedures of any of the main methods we will study. It also in
volves a fair amount of jargon, which can sometimes be a bit intimidating. But
just imagine that you are learning new vocabulary, as you would when entering
any new culture.
REFLECTION
What do you picture when you read the phrase the experimental method}
What images does it evoke for you?
In this section, we will build on key concepts that were introduced earlier.
These concepts included samples, populations, variables, reliability, and validity.
As we saw earlier, experiments are generally conducted in order to test the
strength of relationships between variables. We also saw that when the re
searcher is testing the influence of one variable on another, the variable doing
the influencing is called the independent variable, while the one being influ
enced is called the dependent variable. For example, in a study of the effect of
two different methods for teaching grammar, the teaching method would be the
independent variable, and the students' performance on a test of grammar
knowledge would be the dependent variable.
In Chapter 3, we discussed confounding variables—those factors that might
negatively influence the interpretation of your results. In the experimental
method, one of the researcher's key goals is to control and systematically manip
ulate variables in order to determine cause-and-effect relationships. This goal
has such a high value in the culture of the experimental method that people have
written extensively about the things that can go wrong. These types of con
founding variables are also called extraneous variables or threats to -validity.
QUALITY CONTROL ISSUES:
THREATS TO INTERNAL VALIDITY
Quality control in the experimental method is largely a matter of understanding
the many things that cango wrongand taking stepsto preventor minimize those
threats. Many of these safeguards are embodied in the various research designs
described here. Aswe saw in Chapter 3, there are threats to both internal and ex
ternal validity. We will revisit these issues now in more detail.
The threats to internal validity can be divided into three categories(Tuckman,
1999). These are sources of bias based on experience, participants, and instrumen
tation, lixperience bias factors are those "based on what occurs within a research
study as it progresses" (p. 134). Participant bias is a result of the "characteristics of
84 EXPLORING SECOND LANGUAGE CLASSROOM RESEARCH
thepeople onwhom thestudy isconducted" (ibid.). And instrumentation bias has to
do with "the way the dataare collected" (ibid.).
Threats to Internal Validity Based on Experience
There are three experience bias factors: history, testing, and expectancy. In this
context, history refers to events—things that happen during an experiment,
which may influence the results. For example, if classes in a study are disrupted
due to naturaldisasters or political unrest, the research will be affected. History
can also have an unintended influence on the outcomes of the treatment. Imag
ine you were teaching Japanese to secondary school students in an English-
speaking country and running anexperiment inwhich one group gotto see films
about Japan during class and one group did not. Butif all the students go to see
some popular new action film aboutJapanese samurai warriors outside ofschool,
the treatment couldbe compromised bythat eventsince both groupswouldhave
been exposed to the film.
Testing (also called the practice effect) refers to jhe fact that taking apre-test
may influence the subjects' performance on a post-test. That is, in addition to
learning from thetreatment, learners may dobetteron the post-test because the
pre-test alerted themto what was being investigated in the study.
The third issue, expectancy, is an interesting psychological problem. Tuckman
(1999) explains it this way:
A treatment mayappear to increase learning effectiveness as compared
to that of a control or comparison group, not because it really boosts
effectiveness but becauseeither the experimenter or the subjectsbelieve
that it does and behave according to this expectation, (p. 135)
These two threats are called researcher expectancy and subject expectancy,
respectively.
There are steps researchers can take to overcome or minimize these prob
lems. For example, if you use a design with a pre-test, it is important that the
post-test be a different form of the test, instead of the same test the subjects en
countered at the beginning of the study. (We will read about other safeguards
later.)
Threats to InternalValidity Based on Participants
There are five participant bias factors, andyouhave already read about some of
[Link] are (1) selection, (2) maturation, (3) statistical regression, (4) exper
imentalmortality, and (5) the interactive combinations of factors.
Selection as a threat to internal validity is the idea that somehow the groups
to be compared turn out to be different before the treatment. Random selection
and random assignment are used to combat this problem. The logic here is that
randomly selecting subjects and then randomly assigning them to different
conditions distributes 'contaminating' participant factors across both the
experimental and the control group. You would thus be able to argue that any
Chapter^ TheExperimental Method 85
differences observed in terms of the dependent variable are due to the experi
mental treatment because the othervariables thatmight have had an effect pre
sumably exist in equal quantities in both the experimental and control groups,
and therefore cancel one another out.
After randomly selecting your subjects from the population, you can use
pre-test data elicited from all the subjects in the experiment before you assign
them to groups. This step allows you to make sure, for instance, that the inter
mediate learners in two different groups are at roughly the same level of lan
guage development to begin with. Of course, using a pre-test introduces the
possibility of the testing threat, so there are some trade-offs in the decisions
you must make.
Maturation refers to the normal development people undergo whether or
not theyare receiving a treatmentof some kind. This threat is particularly rele
vant in longitudinal studies involving children. If we see syntactic development
in five-year-olds taught with a certain method for a school year, can we be sure
that that development was due to the treatment, or was it due to the normal lin
guistic changes that small children experience in their first language, or both?
This problem is addressed through the use of a control group, which is at the
same developmental level asthe treatmentgroup and goes through the same ex
periences (exceptfor the treatment itself) for the same period of time.
Statistical regression is the name of a tendency for people's test scores to
change whether or not the knowledge, skill, or ability being measured truly
changes. Tuckman (1999) gives as an example of the situation in which
a group of students take an IQ test, and only the highest third and the
lowest third are selected for the experiment, eliminating the middle
third. Statistical processes would create a tendency for the scores on any
post-test measurement of the high IQ students to decrease toward the
mean, while the scores of the low IQ students would increase toward the
mean. Thus, the groups would differ less in the post-test results, even
without experiencing any experimental treatment, (p. 136)
Tuckman explains that this pattern happens because "chance factors are more
likely to contribute to extreme scores than to average scores, and such factors
are unlikely to reappear during a second testing" (ibid.). Using subjects who
represent an entire range of abilitylevels in your designis one wayto avoid this
problem.
Experimental mortality (or just mortality) is the problem of losing subjects
from the study. It canbe especially worrisome if the groups end up beingof quite
different sizes because people dropped out. To deal with this threat, researchers
often try to recruit more people for a study than they may actually need. Re
searchers must sometimes also try to locate subjects who took the pre-test and
experienced the treatment (or were in the control group) but then moved away
or were absent when the post-test was administered.
As the name suggests, the intei-active combination offactors happens when
more than one threat is presentin a study. For example, ifyouconducta studyof
86 EXPLORING SECOND LANGUAGE CLASSROOM RESEARCH
reading readiness and choose two intact groups for a study—say, two classes of
kindergartners at two different schools—you may end up with children from
widely divergent socioeconomic backgrounds. It is possible that the more afflu
ent children have better nutrition and more opportunities to be read to than do
the poorer children. This is, in part, a selection issue, but if those children are
developing at different rates as a result oftheir nutrition, then you may also have
a maturational issue. Careful planning isneeded to avoid these sorts of problems.
Threats to Internal Validity Based on Instrumentation
The term instrumentation refers to the "measurement or observation proce
dures used during an experiment" (Tuckman, 1999, p. 137). Instrumentation
includes tests, questionnaires, observation systems, elicitation devices, audio-
and videotaping—in short, any means of collecting [Link] threat of instru
mentation is also sometimes called instability of measures (Brown, 1988, p. 39),
because it occurs if the measurement or recording processes change during the
experiment. I
Instrumentation is related to reliabilitysince it involves consistency in data
gathering. To minimize this threat, data collection procedures must "remain
constant across time aswellas constant aa-oss groups (or conditions)" (Tuckman, 1999,
p. 138). The mostlikely instrumentation problemin classroom research involves
studies with human observerstaking notes or using coding systems during class
room interaction. If the observers are not consistent as they collect data, the data
willnot accurately represent the eventsbeing observed. Observer training, rater
training, and the carefulpilotingof all questionnaires and data collection devices
are the best safeguards against instrumentation problems.
REFLECTION
InChapters 1,2, and 3, find two examples ofinstruments used inresearch
that mightbe subjectto the instrumentation threat.
QUALITY CONTROL ISSUES:
THREATS TO EXTERNAL VALIDI
In Chapter 3, we saw that external validity (or generalizability) is the extent to
which the findings of an experimentwillgeneralizeto nonexperimentalcontexts.
There are four issues of concern about external validity: (1) the reactive effects of
testing, (2) the reactive effects of experimental arrangements, (3) the interaction
effects of selection bias, and (4) multiple-treatment interference (Tuckman,
1999).
The external validity threat called reactive effects of testing is related to the
testing threat (or practice effect) to internal validity. In this situation, the presence
Chapter 4 TheExperimental Method 87
of a pre-test may give the effects of the treatment a boost. Then, when the out
comes of the experiment are transferred to the real world where no pre-test is
involved, that extra boost will be lacking. This problem is also applicable toatti
tude questionnaires (ibid.). Ifa researcher administers a questionnaire at thestart
of an experiment in order to see if the subjects' attitudes change as a result of the
treatment, the questionnaire maysensitize the subjects to the fact that attitude is
an issue in the study. As a result, they may respond more positively to the ques
tionnaire when it is administered following the treatment, leading the researcher
to conclude that the treatment improved students' attitudes. However, that im
provement may not be present (or may not beas pronounced) in nonexperimen
tal conditions where there is no pre-test attitude questionnaire.
The reactive effects ofexperimental arrangements threat is a very interesting
problem. This idea refers to the fact thatsometimes just knowing that one is in
an experiment is enough to cause a difference that may be captured by the de
pendent variable, whether or not one is in the treatment group! You may come
across the term the Hawthorne effect as a label for this threat because of a famous
experimentcarried out at the Western ElectricCompany in Hawthorne. Illinois.
in the 1920s:
The researchers wanted to determine the effects of changes in the phys
ical characteristics of the work environment as well as in incentive rates
and rest periods. They discovered, however, that production increased
regardless of the conditions imposed, leading them to conclude diat the
workers were reacting to their role in the experiment and the impor
tance placed on them by management, (ibid., p. 140)
In other words, simply being included in a stud)' can influence subjects' behav
ior. If you conduct observational research in a class other than your own, the
teacher may say to you, "The students were on their best behavior because you
were here," or "The students were really rowdy because you were here." These
are examples of the reactive effects of experimental arrangements in language
classroom research.
REFLECTION
Think about the times that you have observed a class or have been ob
served when you were teaching a class. Did any reactive effects occur?
What were they? WTio was affected? What might be done to counteract
such problems when an observer visits a language class?
The threat known as the interaction effects of selection bias occurs when the
sample in an experiment is not really representative of the population from
which it was drawn. This is a major issue for language classroom researchers be
causeit hingesupon firstdefining the populationwe wish to study and then upon
88 EXPLORING SECOND LANGUAGE CLASSROOM RESEARCH
sampling fromthat population appropriately. For example, assume youare inter
ested in cognitive style and you wish to investigate the effects of analytic versus
holistic teaching styles on language learners. You conducta studywith a hundred
students who are divided into two groups, one oft which is taught analytically
while the other group is taught holistically. Unfortunately, after you finish your
study, you read a research report about left-handed people tending to be more
holistically oriented and right-handed people tending to be more analytically
oriented. When you checkyour records, you see that only one of the hundred
subjects was left-handed. This is unfortunate because left-handed people make
up about thirteen percent of the population. So, left-handed people have been
underrepresentedin your sample.
A way to cope with this threat is a process called stratified random sampling.
This term means that before we select people from the population to be in the
sample, wedetermine whatthe relevant characteristics of the population are (like
handedness) and make sure the levels (strata) in the sample represent the popu
lation appropriately. In this case, we would make sure that the sample included
about thirteen percent left-handed people. If there were enough left-handed
people included in the sample, you could even buildin handedness as a modera
tor variable to check its effects in your research on holisticand analyticteaching
styles.
Finally, there is a threat known as midtiple-tredtment interaction. This threat
is hard to manage in classroom research, especially in second language settings
where students have access to the target language outside of class. It can be a
problem in foreign language settings, too. Imagine that you are using an intact
groups design, in which your 9 A.M. class of secondary school French students
is taught with the traditional materials, butyousupplement the lessons foryour
1 P.M. class with recordings of French popular music. The students in the 1 P.M.
class respond enthusiastically and even share the French music recordings with
their friends in the 9 A.M. class. In effect, the comparison group has now gotten
the treatment!
RESEARCH DESIGNS IN THE EXPERIMENTAL
METHOD
In order to deal with these threats, the experimental method includes many dif
ferent research designs to counteract the possible confounding variables that
could influence the internal and external validity of a study. The various designs
have different strengths and weaknesses. Anyone who chooses to do an experi
ment must balance the focus of the research question against the time and re
sources available for conducting the study in order to choose the best design.
In this section, we will review some research designs discussed in Chapters
1, 2, and 3 (the true experimental designs, the intact groups design, and two in
the ex post facto class—correlation and criterion groups designs). We will also
introduce some other designs that are used in the experimental method. To show
Chapter4 The Experimental Method 89
you the differences among these designs, we will start with a research situation
and develop it through a series of evolving scenarios. (There are manymore re
search designs in the experimental method. We are just describing some of the
most important ones here.)
Scenario 1: The One-Shot Case Study Design
Suppose you are an EFL program administrator. You want to determine whether
or not the students taking the TOEFL preparation course in your program are
benefiting from the course. You decide to conducta studyin which the twenty
students in the course are tested on the TOEFL at the end of the fifteen-week
term.
REFLECTION
When you have completed this study, what information will you have
about the students? What information willyou lack?
This research design is called a one-shot case study. (This phrase does not
mean the same thing as the term case study does in naturalistic inquiry—a point
we will explore in Chapter 6.) The one-shot case study is a weakdesign because
of the problems inherent in the interpretation of the results. Since there is no
pre-test, we don't know how proficient the students were in English at the begin
ning of the TOEFL preparation class. As a result, we can't really say, on solid
empirical grounds, whether the course helped them (though the students and
teacher may be sure that it did). And since only one set of data is available, no
comparisons are possible.
Scenario 2: The One-Group Pre-Test Post-Test Design
You still want to determine whether or not the students taking the TOEFL
preparation course are benefitingfrom the class, but you realizesome things that
could be done to improve the study. You decide to add a pre-test to your design,
so the twenty students in the course are tested on the TOEFL at the beginning
and at the end of the fifteen-week term.
REFLECTION
What informationwillyou have at the end of this study that you wouldn't
have had after conducting the one-shot case study described in Scenario 1?
This design is called a one-group pre-test post-test design. It is a weak design,
but it is an improvement over the one-shot case study because you can at least
90 EXPLORING SECOND LANGUAGE CLASSROOM RESEARCH
tell if the students made progress by comparing their scores at the beginning of
the class with their scores at the end. (Hence the name pre-test post-test design.)
Comparing the pre-test and post-test scores allows you to determine whether
they made progress during (but not necessarily because of) the TOEFL prepa
ration course. (Perhaps they got extra tutoring outside of class or had some
English-speaking friends.)
The difference between the pre-test scores and the post-test scores is called
the gain scores. If the students' post-test scores are higher than their pre-test
scores, then we can conclude that their English (as measured by the TOEFL) has
improved. Butbe forewarned: There are alsonegative gains scores—when the post-
test scores are lower than the pre-test scores. This situation can be discouraging
for both teachers and students. It may mean the students have experienced some
loss of proficiency, or that they didn't test as well the second time, or that there is
just some flux in the measurement process. (We will return to this point later.)
REFLECTION
What steps could you take to refine this design so that you could confi
dently say that the students' measured improvement was in fact due to the
TOEFL preparation course?
Scenario 3: The Intact Groups Design
Suppose you wonder whether or not the TOEFL preparation course helps the
students increase their TOEFL scores. You conduct a study in which the twenty
students in the course are tested on the TOEFL at the end of the fifteen-week
term. You will compare their end-of-term TOEFL scores to those of twenty
similar students in another class that meets at the same time of day and for the
same number of hours every week. However, the students in that class are not
studying specifically to prepare for the TOEFL.
ACTION
Answer these questions about the study described above.
1. What is die research question or hypothesis for this study?
2. What is the independent variable and how many levels does it have?
3. What is the dependent variable?
4. What are the two control variables?
5. Identify one problem inherent in this study.
6. What are two things that could be done to improve this study?
Compare your answers with those of a classmate or colleague.
Chapter 4 TheExperimental Method 91
As you will recall from Chapter 2, this research design is called an intact
groups design. It is known as a weak design because of the problems inherent in
the interpretation of the results, but it isstronger than the one-shot casestudy or
the one-group pre-test post-test design.
REFLECTION
What are the weaknesses inherent in the intact groups design? (Think back
to our discussions of randomization in Chapters 2 and 3.)
The main problems with the intact groups design stem from the fact that the
subjects in the groups being compared were not randomly selected from the
population, nor were they randomly assigned to groups. (Hence, the name "in
tact groups design.") Without randomization (and without a pre-test), we cannot
be certain that the groups being compared were identical (or at least quite simi
lar) to begin with. Perhaps the students in one group are more motivated than
those in the other group, or have greater language aptitude or higher language
proficiency to start with. As a result, we cannot be sure that any differences we
find are truly due to the treatment (the TO FIT preparation course). Therefore,
when we usean intact groups design we must be conservative when we report the
results.
These three designs—the one-shot case study, the one-group pre-test post-
test design, and the intact groups design—all belong to the pre-experimentalclass
of designs. They are pre-experimental in that they lack some of the defining
characteristics of the true experimental designs.
ACTION
Decide which of the three designs discussed above is being used in each of
the following research situations.
1. An Arabic teacher wanted to know whether pronunciation exercises im
proved her students' pronunciation of difficultwords. She gave the stu
dents a pronunciation test before doing the pronunciation exercises, and
then she tested them again after the class.
2. A teacher used two different methods to teach vocabulary with her two
intermediate French classes. One group got a list of randomly ordered,
unrelated words to memorize. The second group got the same vocabu
lary lists, but in addition, the teacher created jazz chants using the
words. At die end of the term, both classes were tested over the vocab
ulary presented in the course.
92 EXPLORING SECOND LANGUAGE CLASSROOM RESEARCH
3. A teacher wanted to see if listening to Spanish radio programs would
help the pronunciation of his beginning Spanish students. He consis
tently assigned three hours per week of listening to Spanish radio for
homework. At the end of the course, the teacher asked a native Spanish
speaker to judge the students' pronunciation.
Compare your ideas with those of a classmate or colleague.
Scenario 4: The Nonequivalent Control (Comparison) Groups Design
Suppose you are an EFL program administrator who wants to determine
whether or not the students taking the TOEFL preparation course in your pro
gram are benefiting from the course. Vou conduct a study in which the twenty
students in the course are tested on the TOEFL at the beginning and at the end
of the fifteen-week term. You will compare their TOEFL scores to those of
twenty similar students in another class that meets at the same time and for the
same number of hours each week. These students are not studying for the
TOEFL, but they are also tested at the beginning and end of the fifteen-week
term with the same pre-test and post-test that the TOEFL preparation students
take.
REFLECTION
What element has been added to this scenario that was not present in Sce
nario 3, which described the intact groups design?
ACTION
Reread the paragraph above and answer these questions.
1. What is the research question or hypothesis for this study?
2. What is die independent variable and how many levels does it have?
3. What is the dependent variable?
4. What are the two control variables?
5. Identify one problem inherent in this study.
6. What are two things that could be done to improve this study?
Compare your answers with those of a classmate or colleague.
This research design is called a nonequivalent control groups design, or—more
properly— a nonequivalent comparison groups design. (The name comes from the
Chapter 4 The Experimental Method 93
fact that the groups were not randomly sampled so we cannot claim they are
conceptually equal.) This design isnot as strong as the trueexperimental designs
because it lacks randomization, but it is stronger than the one-shot case study,
the one-group pre-test post-test design, or the intactgroups design.
The benefitof the nonequivalent comparison groups design over the intact
groups design is the data from the pre-test. Those data allow you to saywhether
or not the groups were identical (or quitesimilar) at the beginning of the term.
You can also compare the groups' gain scores instead of just their post-test
scores.
However, because the groups being compared were intact (i.e., they were
not randomly sampled or randomly assigned), there are still limitations on the
claims you can make. In fact, that is whywe prefer the term comparison groups in
the design's name rather than control groups. By definition, a control group is one
made up of people randomly selected from the population and randomly as
signed to the groupsin the study. In addition, in the strongestdesigns—those in
the true experimental class—which group serves as the control group is often
randomly determined, perhaps by the flip of a coin. This step is a further safe
guardto ensure that there are no known preexisting differences that mightinflu
ence the outcome of the study.
Scenario 5: The Time Series Design
At this point, in order to introduce two different research designs, we want to
change our focus a bit. Let's look at the TOEFL preparation course from the
point of view of the teacher. In this situation, there is only one group of stu
dents—avery familiarsituation in language classroom research.
Imagine that you are the teacher for this course and that there are twenty
students. Youwant to help the students prepare for the TOEFL. Given the im
portance of academic vocabulary on the TOEFL, at the end of every week you
give the students a twenty-item quiz that tests the vocabularyfor that week. The
average scores on the quizzes tend to be around 13 or 14 points each week.
During the eighth week of the semester, you administer a practice TOEFL.
Several students are unhappily surprised by their low scores. The experience of
taking the practice TOEFL seems to motivate them, and, thereafter, they exert
more effort in studying for the weeklyvocabulary quizzes. In the weeksafter the
practice TOEFL, you notice that the quiz scores are higher than they were be
fore the practice exam, with the average ranging between 16 and 18 points. The
average scores on the weeklyvocabulary quiz are indicated by the asterisks (*) in
Figure 4.1. The vertical gray bar indicates the administration of the practice
TOEFL.
This scenario is an example of the time series design. In this design there is
only one group, so comparison with another group is not possible. However, it
is possible to compare the group's average scores before the practice TOEFL,
which we can think of as a treatment, with their average scores after the practice
TOEFL. We can use this design, for instance, to help answer the question of
whether or not administering the practice TOEFL motivated the students to
94 EXPLORING SECOND LANGUAGE CLASSROOM RESEARCH
20 -|
19-
18- * * *
* * *
fl 17"
0
* 15-
14- * * * *
13- * * * i:
12- i i i i i i i 1 l i i i i i i
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15
Week
FIGURE 4.1 Average scores on weekly twenty-point vocabulary quizzes
before and after the administrationof the practiceTOEFL
FIGURE 4.2 Possible outcomes in a time series design (Tuckman,
1999, p. 169)
study harder. From the data above, it appears that that effort resulted in higher
vocabulary quiz scores.
In the time series design, the group under investigation serves as its own
control. That is, before the treatment, the students are functioning as a control
group, but after the treatment, they are analogous to an experimental group.
This situation is far from ideal, but it can be informative. We need to be careful,
though, about interpreting the results, as shown in Figure 4.2. In this figure,
T stands for time.
Chapter4 The Experimental Method 95
The inference that the treatment caused an effect is most justified in cases 1 and
2 above. It is least justified in eases 3 and 4 (Tuckman, 1999).
REFLECTION
Why does Tuckman (1999) assert that cases 3 and 4 suggest that the treat
ment has not caused an effect?
These four sets ol data all provide different information about the possible
efteets of the treatment. The first line in Figure 4.2 tells us that the scores were
higher after the treatment than they were before, and that they remained consis
tently [Link] second line tells us that the scores improvedafter the treatment
but that the improvement tapered off later. The third line indicates that the
scores were higher after the treatment and continued to increase, but we cannot
say that the treatment caused this improvement because the scores had already
begun to increase before the treatment. It appears that this developmental trend
might have continued with or without the treatment being administered. Finally,
in the fourth line, the scores are somewhat erratic. They improve immediately
alter the treatment and then drop back down again, but the scoreswere relatively
high at one point before the treatment as well.
Scenario 6: The Equivalent Time Samples Design
.Another design that is used in contexts where there is no comparison group or
control group is called the equivalent time samples design. It is related to the time
series design. To see how this design works, we will use another scenario related
to the TOEFL preparation course.
Once again, please imagine that you are the teacher for the course. You have
implemented the practice of giving weekly vocabulary quizzes to encourage the
students to study the academic vocabulary covered in class. You read a research
article about the importance ol providing a meaningful context for the study of
vocabulary and you decide to test the author's ideas. For every other week of the
fifteen-week term, starting with the first week, you provide a story at the begin
ning ol the week that oilers a clear, memorable context for the vocabulary the
class will study that week. On alternate weeks, you simply provide the week's vo
cabulary list without contextualizing the vocabulary items in a story. At die end
of every week you administer a twenty-point quiz. The results are depicted as a
bar graph in Figure 4.3.
REFLECTION
What can you infer from the results depicted in Figure 4.3?
96 EXPLORING SECOND LANGUAGE CLASSROOM RESEARCH
20 -|
19-
18-
17-
n
{16
a.
15-
14-
13-
12 T T t—n r—n r—n
3 4 5 8 9 10 11 12 13 14 15
Week
FIGURE 4.3 Average scores on weekly twenty-point vocabulary
quizzes for weeks with contextualization (gray bars) and
weeks without contextualization (white bars)
Once again, when we work with the equivalent time samples design, we say
that the group serves as its own control. There is no separate control or compar
ison group, but we are able to compare the scores on the dependent variable (the
twenty-point vocabulary quiz) for the weeks when contextualization was pro
vided with the vocabulary scores for the weeks when it was not. In effect, the
contextualization is the treatment, and it is alternately provided and withheld
from the same group of students. This design is stronger than the time series de
sign because multiple comparisons are possible. That is, since the treatment has
been given and withheld several times, we have a better chance of detecting its
effect (if any).
These three designs (the nonequivalent comparison group design, the time
series design, and the equivalent time samples design) all belong to a class called
the quasi-experimental designs. This class is characterized by (1) the possibility of
making comparisons on the dependent variable, but also by (2) the lack of a
randomly sampled and randomly assigned control group.
REFLECTION
Think of a research question that you could address with a time series de
sign and one that could be addressed using the equivalent time samples
design. Remember that these are designs that can be used when no con
trol group (or even a comparison group) is available. Given that fact,
what could you confidently say about the possible outcomes of your two
studies?
Chapter 4 The Experimental Method 97
ACTION
Identify the particular quasi-experimental design involved in each of the
following situations. Compare your ideas with those of a colleague or
classmate.
1. An EFL teacher in Turkey wanted to see if having a party where only
English was spoken would have an impact on the conversational fluency
of her students. The party was arranged for the middle of the semester.
Every week before the party, she recorded a brief conversation with
each student. After the party, she continued to record the weekly
conversations to try to determine whether the party had had an impact
on the students' conversational fluency.
2. A teacher of Swedish as a second language wondered whether imple
menting drama techniques would improve her students' pronunciation.
The teacher tested the pronunciation of her three intermediate speak
ing classes at the beginning of the term. With the 9 a.m. group, she
used just regular conversation practice. With the 11 a.m. class, she used
role plays in addition to conversation practice. And with the 2 p.m.
class, she used role plays and conversation practice, but she also had the
students perform scripted plays in Swedish. At the end of the term, she
tested all the students again to see whether their pronunciation had
improved and, if so, whether those gains differed across the three
groups.
3. An EFL teacher in Shanghai wanted to determine the effect of translat
ing new vocabulary items into Chinese for the students (instead of just
talking about their meanings in English). She gave vocabulary quizzes
every week and tried the translation approach to vocabulary teaching
everyother week. At the end of the semester, she compared the student'
quiz averages for the weeks following the translation approach with
those from the weeks in which only English explanations were used.
Scenario7: Post-Test Only Control Group Design
We are now moving into a different class of designs—the true experimental de
signs. These are characterized by (1) random selection of subjects from the pop
ulation, (2) random assignment of subjects to groups, and (3) the presence of an
actual control group.
Let's return to the situation about the TOEFL preparation class. Please
imagine once more that you are an EFL program administrator who wants to
determine whether taking the popular '[Link] preparation course in your pro
gram helps the students. You conduct a study in which the twentystudents in the
course are tested on the TOEFL at the end of the semester. You will compare
98 EXPLORING SECOND LANGUAGE CLASSROOM RESEARCH
their TOEFL scores to those of twenty similar students who are tested at the
same time but who are not studying for the TOEFL.
In order to make sure the two groups of students are as similar as possible,
you randomly draw forty names from the list of over a hundred students who
wish to take the TOEFL preparation course. Then you randomly assign twenty
of those names to one group and twenty to another group. Finally, you flip a coin
to see which group will enroll in the TOEFL preparation class this term and
which group will wait until next term. This second group is placed in a grammar
review course, which meets at the same time and for the same number of hours
as the TOEFL preparation course.
ACTION
Answer the following questions about this study.
1. What is the research question or hypothesis for this study?
2. What is the independent variable and how many levels does it have?
3. What is the dependent variable?
4. Name a problem inherent in this study.
5. What is one thing that could be done to improve this study?
Compare your ideas with those of a classmate or colleague.
This research design iscalled the post-test only control group design. It is one of
the true experimental designs because of the presence of a control group and the
random selection and random assignment of subjects.
Scenario8: Pre-Test Post-Test Control Group Design
You have probably already realized from the name of this design what additional
change could be made. If we simply add a pre-test to the study in Scenario 7, we
will have a pre-test post-test control group design. The independent variable
and its levels remain the same.
REFLECTION
Now what is the dependent variable? What is one possible threat to valid
ity inherent in this design?
The pre-test post-test control group design is one of the true experimental de
signs because of the presence of a control group and the random selection and
random assignment of subjects. In one sense, it is stronger than the post-test
only control group design because you can measure improvement (through the
gain scores). However, it is also susceptible to the testing threat.
Chapter4 The ExperimentalMethod 99
COMPARISON OF MAIN RESEARCH DESIGNS
IN THE EXPERIMENTAL METHOD
So far in this chapter, we have studied eightresearch designs. In Chapter 3, we
discussed the criterion groups design and the correlation design, which are the
two members of the expostfacto class. The fourmain classes of designs and the
various specific designs that comprise them are depicted in Table 4.1.
If youstart with the upperleft box in Table 4.1 anddraw a large Z through
the four boxes, you will have a way of remembering the increasing power of
these designs. That is, the pre-experimental designs are the weakest, followed
by the quasi-experimental designs and then the ex post facto designs. The true
experimental designs are the strongest. Within the culture of experimental re
search, this relative strength is a function of increasing control over variables.
The stronger designs are those with the greatest internal and external validity.
Three designs are marked with asterisksin Table 4.1. The addition of one or
more moderator variables in these designs makes themfactorial. For example, a
factorial post-test only control group design is one in which there are one or more
moderator variables in the study. Theoretically, it would also be possible to add
a moderator variable to an intact groups design or a nonequivalent comparison
groups design, but this is rarely done because adding a moderator variable adds
complexity to the statistical analysis. Since these designs are relativelyweak, it's
hardly worth the extra effort to add a moderator variable.
Another wayto contrast these designsis to list their defining characteristics.
Table 4.2 does so by answering the followingyes/no questions:
Column 1: Does the design involve more than one group of subjects?
Column 2: Does the design involve administering a treatment?
Column 3: Does the design involve a randomly selected, randomly as
signedcontrol group to comparewith the experimental group(s)?
TABLE 4.1 Major designs and classes of research designs in
experimental research (after Shavelson, 1981, pp. 30-44)
Pre-Experimental Class Quasi-Experimental Class
1. One-Shot Case Study 1. NonequivalentControl (or
Comparison) Groups Design
2. One-Group Pre-Test Post-Test Design 2. Time Series Design
3. Intact Groups Design 3. EquivalentTime Samples Design
ExPost Facto Class True Experimental Class
1. Criterion Groups Design* 1. Post-test Only Control Group Design*
2. Correlation Design 2. Pre-test Post-test Control Group
Design*
100 EXPLORING SECOND LANGUAGE CLASSROOM RESEARCH
TABLE 4.2 Comparison often experimental research designs
Design 1 i
3 4 5 6 7
One-shot case study No Yes No No No No No
One-group pre-test
post-test No Yes No No Yes No No
Intact groups Yes Yes No Yes No No No
Nonequivalent
comparison groups Yes Yes No Yes Yes No No
Time series No Yes No No Yes No No
Equivalent time samples No Yes No No Yes No No
Criterion groups Yes No No Yes No Yes No
Correlation No No No No No Yes No
Post-test only control
group Yes Yes Yes No No Yes Yes
Pre-test post-test control
group Yes Yes Yes No Yes Yes Yes
Column 4: Does the design involvesome other kind of comparison group?
Column 5: Is a pre-test given?
Column 6: Is random selection used to constitute the sample?
Column 7: Is random assignment used to constitute the groups?
There are some caveats to remember when interpreting this table. First, it is
important to distinguish between the presence of an actual control group (Col
umn 3) and some other sort of comparison group (Column 4). Secondly, in cor
relation studies, the statistics used to perform the correlation analysesare always
based on two (or more) sets of data from one group of people. However, some
studies include different sets of correlation statistics if correlations are sought in
more than one group. (We will deal more with the statistics used in correlation
designs in Chapter 13.)
ACTION
With a classmate or colleague, talk through the yes/no responses in
Table 4.2. Make sure you understand how these seven defining character
istics can help you identify these research designs in the experimental
method.
Chapter 4 The Experimental Method 101
RESEARCH DESIGN ISSUES AND INFERENTIAL
STATISTICS
To review these concepts, let's consider asituation inwhich an experiment might
be an appropriate way ofgathering data. Imagine thatyou have developed some
innovative listening materials based on authentic radio and television programs.
You have used these materials with good results in your secondary school EFL
classes. Although you feel strongly thatyour innovative materials aresuperior to
the school's traditional listening program, your colleagues are skeptical. Your
challenge is to test the possible superiority of your materials. You have several
options here. You could obtain the opinions of the students through surveys and
questionnaires. Alternatively (or additionally), you could ask a sympathetic col
league to observe your classes and make an observational record of the teaching
and learning thatgoes on. However, you may feel that these steps areunlikely to
sway your more skeptical colleagues, who are only likely to be convinced by
superior test scores.
Yourfirst thought is to test your students' listening comprehension at the end
of the semester, and,assuming that the results are favorable, presentthesefindings
to your colleagues. However, you come across the following criticism of such an
approach (which you now recognize as the one-shot case studydesign):
Much research in education todayconformsto a design in whicha single
group is studied only once, subsequent to some agent or treatment pre
sented to cause change. Such studies might be diagrammed as follows:
XO
IX = the treatment administered to the subjects, and O = the observation.]
[Unfortunately]... such studies havesuch a total absence of control as to
be of almost no scientific value ... It seems well-nigh unethical... to
allow as theses or dissertations in education, case studies of this nature
(i.e., involving a single group observed at one time only). (Campbell and
Stanley, 1963, pp. 176-177)
If you are convinced by this argument,your next inclination might be to test
two of your classes—one taught with the innovative materialsand one taught with
the traditional approach. (Doing so would involve using the intact groups design.)
However, you quickly realize that it is no good simply testing the students at
the end of the semester and comparing their scores becausethe groups might not
have been at the same level to begin with. The solution would seem to be to test
both groups at the beginning and at the end of the semester. Then, if the group
that has been taught through the innovative materials makesgreater gains in lis
tening comprehension than the traditional group, you can presumably ascribe
the superior results to the innovation. (Here you will recognize the nonequiva
lent comparison groups design.)
Your research design is becoming more rigorous, but it is not yet rigorous
enough to allow you to make claims that there is a causal relationship between
102 EXPLORING SECOND LANGUAGE CLASSROOM RESEARCH
the independent variable (the innovative materials) and the dependent variable
(the students' listening comprehension test results). There is always thepossibility
that factors other than the innovative materials are responsible for any observed
differences in the scores.
REFLECTION
Make a list of possible issues that might be responsible for differences in
the groups'scores in the study described above, Do these issues affect the
internal validity or the external validity of the stkdy?
There are many possible influences that can affect the outcome of a study
such as this one. If different teachers are involved in teaching the different
groups, then it could be the teachers rather than the materials that make a differ
[Link] one teacherworkswith both groups,you will havecontrolled for teacher
style as a factor, but the teacher's enthusiasm for (or boredom with) one type of
materials could influence the results. Even the time of day at which a classis held
can affect learning outcomes. These issues weaken the internal validity of the
study because it is not possible to state categorically that the treatment brought
about any differences observed in the students' test scores.
While factors such as those mentioned above may impinge on research out
comes, participant factors (such asthe selection threat)are the most pervasive. For
example, you may have happened to select a group of fast-track or high-aptitude
smdents as the recipients of the experimental authentic materials, and a group of
slow learners that used the traditional materials. In order to guard against the
possibility that factors such as age, motivation, or aptitude might influence the
research outcomes, sound experimentaldesign in the psychometrictradition sug
geststhat you assign subjectsrandomlyto the control and experimental conditions.
Using randomization puts you in a better position to argue that any ob
served differences on the end-of-course test are due to the innovative materials
because possibleconfounding variablesthat might have had an effect (such as in
telligence and aptitude) are presumably evenly distributed in the experimental
and control groups. You can also test both groups of students before the experi
ment just to make sure that the groups really are the same, though this step in
troduces the possibility of the testing threat. (Doing so would entail the use of
the pre-test post-test control group design.)
Unfortunately, in ongoing programs, it is not always practical to rearrange
students and randomly assign them into different groups or classes. In many
schools, if an experiment is to be conducted, it will have to be with classes to
which students have been preassigned. That is why the intact groups design and
the nonequivalentcomparison groups design are so often used in languageclass
room research—because they involve groups that were already established (by
means other than randomization) without the researcher's control. In these
circumstances, while the internal validity of the experiment is weakened, the
study may still be worthwhile.
Chapter4 The Experimental Method 103
REFLECTION
Imagine that you are able to carry out the experiment described above. You
randomly assign ten final-year secondary school students to the control
group and ten to the experimental group. A 100-point listeningpre-test in
dicates that the groups are at the same level of proficiency. You teach both
groups for a semester, using the innovative authentic materials with the
experimental group and the traditional materials with the control group.
At the end of the semester, the groups are retested with an alternate form
of" the 100-point listening test. You calculate the averages:
Control Group: Experimental Group:
Post-test average: 80 85
The experimental group has thus outscored the control group.
Are you entitled to claim that the innovative materials are superior to the
traditional materials? If so, why? If not, why not?
The answer to the question is "Not yet!" You have selected a sample, or sub
set, of all the possible students in the final year of secondary school as your ex
perimental subjects. II you retested them again tomorrow, or if you selected a
different group of subjects and tested them, it's highly unlikely that you would
get exactly the same scores. The students might be more tired (or more ener
getic), or the weather could affect their performance. In short, a whole range of
factors could be responsible for test score variation. What you need to decide is
whether the variation in scores between the control and experimental groups
might have happened by chance, or whether the differences were a result of the
experimental treatment. In order to do this, you must make inferences based on
statistical procedures.
FROM SAMPLES TO POPULATIONS: THE LOGIC OF
INFERENTIAL STATISTICS
The aim of this section is to introduce you to the logic of statistical inference.
The information presented here will probably not equip you to carry out your
own statistical analyses, but it should help you to understand and appreciate the
logic behind the statistical procedures that enable researchers to make claims
about an entire population based on a sample or subset of subjects from that
population.
In most research, it is not possible to collect data from the entire population
in which we are interested. Consider your investigation of the authentic materi
als in secondary school EFL classes. Although not impossible, it would be ex
tremely time-consuming and cumbersome to obtain data on all the secondary
104 EXPLORING SECOND LANGUAGE CLASSROOM RESEARCH
school EFL students in your country, state, province, county, or prefecture. (In
fact, that's what national examination boards and international testing companies
try to do.) Normally, someone who wanted to carry out such an investigation
would select a sample of students (say 20, 100. or 2,000 or more) from the wider
population and test them.
However, a problem immediately arises. The problem has to do with decid
ing the extent to which the data obtained from the sample are representative of
the population as a whole and, in fact, what that population is. In everyday life,
overgeneralizations arc common (witness the prevalence of "dumb blond"
jokes). But overgeneralizations also occur in research when investigators fail to
recogni/.e that their subjects are not fully representative of the population being
investigated.
Here is an anecdote from Rowntree (1981) that illustrates the complicated
issue of matching samples and populations:
During the Second World War, gunners in bombers returning from
raids were asked from which direction they were most frequently
attacked by enemy fighters. The majority answer was "from above and
behind." (p. 23)
REFLECTION
Why might it have been unwise to assume, supposing you were a gunner,
that this claim would be true of attacks on gunners in general?
Rowntree provides the following answer to the question:
The risk of a false generalization lay in the fact that the researcher was
able to interview only the survivors of attacks. It could well be that
attacks from below and behind were no less frequent, but did not get
represented in the sample because (from the enemy's point of view) they
were successful, (ibid.)
I Ie goes on to say that this issue of matching samples and populations is a para
dox of sampling:
A sample is misleading unless it is representative of the population; but
how can we tell it is representative unless we already know what we need
to know about the population and therefore have no need of samples!
The paradox cannot be completely resolved; some uncertainty must re
main. Nevertheless, our statistical methodology enables us to collect
samples that are likely to be as representative as possible. This allows us
to exercise proper caution and avoid over-generalization, (ibid.)
lb help protect against overgeneralization, researchers use several procedures
based on descriptive and inferential statistics. (Explaining all these procedures is
beyond the scope of this book, but we will deal with some in Chapter 13.)
Chapter 4 The Experimental Method 105
TABLE 4.3 EFL students' scores on a listening comprehension test
(n = 20)
ControlGroup Experimental
Student ID Scores Student ID Group Scores
C-l 80 E-l 85
C-2 82 E-2 87
C-3 78 E-3 83
C-4 77 E-4 82
C-5 83 E-5 88
C-6 80 E-6 85
C-7 76 E-7 81
C-8 84 E-8 89
C-9 75 E-9 84
C-10 85 E-10 86
Mean 80 85
DescriptiveStatistics
To understand the logic behind the procedures that enable extrapolation from
samples to populations, you need to be familiar with a number of statistical con
cepts. The two most important of these are the mean and standard deviation.
These are two of the descriptive statistics—so labeled because they describe the
sample.
For experimental researchers, two particularly interesting features of nu
merical data sets are the extent to which individual items in the data set are sim
ilar and the extent to which they differ or are dispersed. The most important
measure of similarity is the numerical average, or mean (symbolized by a capital
X with a horizontal bar above it and called X-bar). The average is obtained by
adding the individual scores together and dividing the sum by the total number
of scores. To illustrate, let's look at the post-test scores from the students in the
control and experimental groups in the (hypothetical) study about innovative
listening materials.
The scores presented in Table 4.3 are simply listed in order of the students'
identification codes. They are not organized in any particular fashion. It can be
quite useful to rank order these scores, in order to see existing patterns more
clearly. When we rank order these scores, we find the data in Table 4.4.
The lowest score in this entire data set is 75 and the highest score is 89. The
difference between the highest score and the lowest score is the range. This is
one of the descriptive statistics. (Ranking the scores in this way is useful because
it allows us to see the range quickly.) The range in the control group was 75 to
85 and the range in the experimental group was 81 to 89. So, just by looking at
106 EXPLORING SECOND LANGUAGE CLASSROOM RESEARCH
TABLE 4.4 Ranking of EFL students' scores on a listening-
comprehension test (n = 20)
Ranking of Control Ranking of Experimental
Control Group Experimental Group
Group Scores Scores Group Scores Scores
1 85 1 89
2 84 2 88
3 83 3 87
4 82 4 86
5.5 80 5.5 85
5.5 SO 5.5 85
7 78 7 84
8 77 8 83
9 76 9 82
10 75 10 si
Average SO 85
the difference in the ranges, we can see that the experimental group did better.
However, range is also reported in terms of its absolute value. That is, we some
times subtract the lowestscore from the highest score. Using this procedure, we
can saythat die range for the control group is ten, and the range for the experi
mental group is eight. We cannot tell which group did better—we can only tell
that there was more variability in the control group's scores than in the experi
mental group's scores.
REFLECTION
In Table 4.4, in the columns that provide the ranking of the two groups'
scores, there are two places where the rank is given as 5.5—one for the
control group and one for the experimental [Link] do you think these
ranks are given as 5.5 instead of 5 and 6?
The answer to the question in the reflection box above has to do with prin
ciples of decision making. Let's use the prize money in a golf tournament as a
metaphor. One golfer is the first place winner, and he receives a check for
S100,000. Two golfers are tied for second place. The prize money for second
place is $50,000, and the prize money for third place is $30,000. What is a fair
way to decide which golfer was second and which was third, since their scores
were tied? The answer is to add the prize money for second place and third place
Chapter4 The ExperimentalMethod 107
and divide thattotal amount by two (for the two tied golfers). So, instead offlip
ping a coin to decide who gets $50,000 andwho gets $30,000, we add these two
sums and get $80,000. We then divide that amount by two and the two second
place contestants each receive $40,000.
The same logic applies in a set of ordinal data when you have tied ranks.
Lookat the ranks that the scores would have covered, had theynot beentied. In
Table 4.4, these are the fifth and sixth ranks. Since the tied scores are identical,
instead of calling one "fifth" and the other "sixth" they are both assigned the
rank of 5.5—halfway between the fifth and sixth ranks.
Frequency Polygons
Anotherway to look at these test datais to see how many studentsobtainedeach
particular score. Let's start with the scores of all twenty students combined. We
canlistthe score values across the bottomaxis of a chartand the numberof peo
ple who obtained each particular score on the vertical axis. The termfrequency
here refers to how often each score was obtained. In Figure 4.4, each asterisk
shows how many people out of the twenty subjects in the study received each
possible score value.
If you were to draw bars down from each of these asterisks to the horizontal
axis, you would have a bar graph, or histogram. If you were to draw a line con
nectingthese asterisks, you would have whatis called afrequency polygon—a chart
of the frequency with which each score value was obtained. Notice that no one
got a score of [Link] this point the line connecting the asterisks would drop down
to the horizontal axis, to indicate that no one received that score.
The frequency polygon is a very important conceptual tool as well as an
informative visual aid. In fact, the frequency polygon is the basis for many of
the most important statistical procedures that we use in language classroom
research.
0 i i i i t i i i i i i i I I I t"
75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90
Score
FIGURE 4.4 Frequency of 20 EFL students' scores on a listening
comprehension test
108 EXPLORING SECOND LANGUAGE CLASSROOM RESEARCH
ACTION
In Figure 4.4, we combined the scores of the control and experimental
group, but you can also draw a frequency polygon that contrasts two or
moresets of data. Usingthe chartframework below, plot the scores forthe
control and experimental groups separately. Remember that for anyscore
that did not appear in the data, the frequency is zero.
3-1
1-
—i 1 1 1 1 1 1 1 1 1 1 1 1 1 i i
75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90
Score
With very small samples like this, frequency polygons can be nearly flat. But
an interesting thing happens when large data sets are plotted on a frequency
polygon. Imagine for a moment that ninety-five students actually took this
100-point test. If we had that many scores to enter in the frequency polygon, it
might look something like this:
i 1 1 1 1 i 1 i i i i i i i r
75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90
Score
FIGURE 4.5 Frequency of EFL students' scores on a listening-
comprehension test (n = 95)
Ifyou connect the asterisks in Figure 4.5, you can see the rough shape of a
bell emerging. In fact, a frequency polygon with this shape is called a bell curve or
Chapter4 The Experimental Method 109
a bell-shaped curve because it is typically high in the middle with sloping sides ta
peringoff to tails on the leftand right. It is alsocalled the normal distribution be
cause it is such a common pattern when variables are measured in large groups
ofpeople. Thatis tosay, the characteristic being measured (whether it is height,
intelligence, language aptitude, etc.) isdistributed normally throughout thepop
ulation. Relatively few people score very low on the measurement and relatively
few people score very high. Most people score somewhere in the middle of the
range.
Keep in mind that "we never get a completelynormal data distribution. The
normal distribution is an idealized concept" (I latch and Lazaraton, 1991,
p. 194). But with large samples, the pattern does appear, and this fact leads to
some interesting opportunities for analyzing data. In very large data sets, the
normal distribution is predictable. The shape of the bell is smooth and regular,
and its sides are very symmetrical, like this:
FIGURE 4.6 Standard normal distribution (downloaded from
[Link] on August 14, 2006)
The vertical line drawn straight down from the apex of the bell in Figure 4.6
represents the midpoint or median—defined as the middle score in a data set. In
the normal distribution, it also represents the mean—the average—and the
vwdc—the most frequently obtained score in that data set. That the line repre
sents the median makes sense because the bell is symmetrical. The fact that it
represents the mode also makes sense because the hells apex (the high point)
shows the most frequent score. These three descriptive statistics are collectively
referred to as the measures ofcentral tendency because they all provide information
about the tendency of scores to cluster in the middle range of the curve.
The three other descriptive statistics are the range and standard deviation
discussed above, and also variance. (We will discuss variance in Chapter 13.) To
gether these make up the measures ofdispersion because thev provide information
110 EXPLORING SECOND LANGUAGE CLASSROOM RESEARCH
about the variability in a set of scores—how dispersed the scores in the data set
are. The standard deviation is the most important measure of dispersion. It tells
us die average amount by which scores in the data set vary from the mean. In
other words, it tells ushow spread out the scores are. In Figure 4.6, the numbers
running along the horizontal axis (from negative four to four) represent standard
deviations.
REFLECTION
Why is the number zero directly under the line that represents the mean,
the median, and the mode in Figure 4.6? (This issue is a matterof logic and
definition rather than of mathematics.)
Calculating the standard deviations for the scores in Table 4.3 gives us the
values reported in the last row of Table 4.5. (We won't go into the formula for
calculating the standard deviation here. We will work with it in Chapter 13.)
We said earlier that the ranges for the scores of the control and experimen
tal groups were ten and eight, respectively. Can you see how the range is re
flected in the standard deviations reported in Table4.5? Sincestandard deviation
is an index of how spread out thescores arein a given data set, it makes sense that
where there is a wider range there will he a bigger standard deviation.
TABLE 4.5 Scores, means, and standard deviations of two groups of
EFL students on a listening comprehension test (n = 20)
Control Group Experimental
Student ID Scores Student ID Group Scores
C-l SO E-l 85
C-2 82 E-2 87
C-3 78 F.-3 83
C-4 77 I-.-4 82
C-5 83 E-5 88
C-6 80 E-6 85
C-7 76 E-7 81
C-8 84 E-8 89
C-9 75 E-9 84
C-10 85 K-10 86
Mean 80 85
Range 10 8
(75-85) (81-89)
Standard 3.25 2.58
deviation
Chapter? The Experimental Method 111
REFLECTION
To put the concept of standard deviation in a practical context, thinkabout
thesituation where you are about to start teaching three different classes of
intermediate English students. For all three classes, the mean score on the
program's 100-point placement exam was 70 points. But the standard devi
ations for the three classes were 15, 10 and 5 points. What do these values
tell you about the composition of the three classes? As a teacher, what can
you expect as a result of this variability?
Inferential Statistics
Atthis point, we want to remind you of some things thatyou already know about
percentages. We want to build on your existing knowledge and confidence to in
troduce a new concept.
ACTION
Look at the pie charts below and decide what percentage of the area of
each circle is indicated by the various sections.
Circle 1 Circle 2 Circle 3
We are quite sure that you recognize the percentages represented by these
divisions. In Circle 1, die sections of the bisected pie chart represent 50% and
50%. In Circle 2, the pie chart is divided into four equal portions. Each segment
represents 25% of the whole circle. And in Circle 3, we have three equal
sections, each of which represents 33M% of the whole circle. You recognize these
percentage values because you have seen charts like this ever since you started
school.
We can use the image of the normal distribution to indicate percentages,
too. As noted above, when the bell curve is bisected by the vertical line repre
senting the mean, the median, and the mode, 50% of the scores fall above that
line and 50% fall below it. (Remember the frequency polygons where you
connected the asterisks. A bell curve is just a symmetrical frequency polygon.)
112 EXPLORING SECOND LANGUAGE CLASSROOM RESEARCH
—3a —2a —la \± la 2a 3<r
FIGURE 4.7 The areas under the normal curve, as indicated by-
standard deviations (downloaded on August 13, 2006
from Wikipedia)
The image of the bell curve can also be divided in predictable, recognizable
ways.
When you see a chart of the normal distribution, it usually has numbers
written on the horizontal axis. These numbers range from negative four to pos
itive four, as mentioned above, to indicate the location of the standard deviations
in the diagram. But, for ease of interpretation, such diagrams also often have
vertical lines drawn that represent the standard deviations. These lines help us
see percentage divisions, as shown in Figure 4.7. Oust think of this as a different
shaped pie chart—one with which you may not be very familiar at this time.)
Here and elsewhere the Greek symbol a (the lower-case sigma) stands for
standard deviation. The symbol /x (the Greek letter /////) represents the popula
tion mean.
REFLECTION
Look at Figure 4.7. Can you see why we said that 68.2% of the scores fall
within one standard deviation of the mean? (Another way to saythis is that
68.2% of the scores fall in the area between one standard deviation below
the mean and one standard deviation above the mean.)
ACTION
Look up normal distribution in a statistics text or on the Internet. There
will probably be more examples and more detail than we have been able to
provide.
Chapter 4 The Experimental Method 113
Means and standards deviations are important when it comes to comparing
different sets of interval data, such as test scores from a control group and an ex
perimental group. We can also think about the standard deviation and the mean
of a population. Imagine you are reviewing reading test scores for nine-year-olds
in 100 different primary schools. Those scores are likely to be reported as
schoolwide averages because looking at individual scores for all the nine-year-
old pupils at 100 schools would be a daunting task. Even looking at the individ
ual class means for even' school would be very time-consuming, since there
could be five to ten classes of nine-year-olds at even,' school. Seeing the scores
reported as schoolwide averagesis economical and it also allows us to make com
parisons across the various schools.
If wewere to draw a frequencypolygon using the 100schools' readingaverages
as the data to be entered in die polygon, we would once again get a bell-shaped
curve. This time the data points in die frequency polygon would not represent
individual students' scores but rather the mean reading scores of the 100 schools.
We have already discussed the concept of a population. Statistically, a pop
ulation is defined in terms of means and standard deviations. If the means and
standards deviations for two sets of test scores are quite similar, then the sub
jects can be said to be drawn from the same population. If they are very differ
ent, then they are drawn from different populations. The key question here is,
"I low different do they have to be for us to be confident in claiming that they
come from different populations?" The phrase statistically significant has to do
with this question of how different a set of scores must be (or how strong a
correlation must be) in order to consider the difference (or the correlation)
important and trustworthy.
Over time, researchers using the experimental method have agreed that re
sults can be considered statistically significant if there is a less than 5% chance
that they are wrong—that is, if the results are due to chance rather than to an
actual relationship between the variables being investigated. (There are some
situations where more stringency is required and the level is set at a 1% chance
instead.) Another way to say this is that researchers generally want to have 95%
(or 99%) confidence that their results are trustworthy before they reject the null
hypothesis andaccept the alternative hypothesis. So, when you read that a differ
ence or a correlation was statistically significant, it means that the values
obtained met the statistical requirements for having confidence that the results
were not due to chance.
REFLECTION
The mean, median, mode, standard deviation, range, and variance are the
descriptive statistics. What do you think is meant by inferential statistics?
Let's return to our hypothetical investigation into the relative effectiveness
of the innovative versus traditional listening materials. On the end-of-the-
semester test, we discovered that the mean score for the traditional (control)
114 EXPLORING SECOND LANGUAGE CLASSROOM RESEARCH
group was 80, and the mean score for the innovative (experimental) group was
85. The control group continues to represent the population fromwhich it was
drawn. What we want to know now is whether, through our experimental treat
ment, we have 'created' a different population—roughly defined as "listeners
taught through innovative materials based on authentic input." In order to an
swerthis question, we need to knownot only the mean for each group, but also
the standard deviation. The reason is that the further an individual score is from
the group mean, the less likely it is to occur by chance, and statistical tools can
tell us fairly precisely how likely or unlikely this is.
How likely? Here's where the standard deviation comes in. Please take our
word for this temporarily (or read ahead in Chapter 13).Asshown in Figure 4.7,
for any given set of normally distributed scores, 68% of all scoreswill be within
one standard deviation of the mean, 95% of scores will be within two standard
deviations, and 99% will be within three standard deviations.
When there is a large difference between the means for the control and ex
perimental groups, we can say that the difference is statisticallysignificant. How
large a difference is determined by the standard deviations in the bell-shaped
curve. For instance, if the mean for the control group falls between two and
three standard deviations below the mean for the experimental group, we can say
that there is less than a 5% possibility that this difference would occur by chance.
If our level of confidence is set at 5%, then we can say that the difference is sta
tistically significant. Therefore, we can conclude that the innovative materials
were significantly superior to the traditional materials in enhancing listening
skills. Keep in mind, however, that there is still a 5% possibility that the differ
ence could have occurred by chance.
Given that the control group remains 'untreated,' it represents the popula
tion of interest. The experimental group represents a changed version of that
original population—a group that benefited (we hope!) from the treatment. And
this is what is meant by inferential statistics: We use the data from the sample to
make inferences about what would happen to the population if the treatment
were implementedthere (i.e., if it were generalized). We do that, quite often, by
comparing the means and standard deviations of the sample groups to one an
other, using particular statistical formulae.
Now let's go back to our data in which the control group mean was 80 and
the experimental group mean was 85. We can use an inferential statistic called a
t-test (the t here always being printed in lower case)to see if these differences are
[Link] t-test is speciallydesigned for comparing two means
of small groups where the means are based on interval data. (We will learn more
about t-tests and other inferential statistics in Chapter 13.)And in fact, when we
conduct the t-test on these data, we find that there was a statisticallysignificant
difference between the control and experimental groups. Using the innovative
listening materials did improve the students' scores on the listening comprehen
sion test. Now you can confidently show your skeptical colleagues your results!
In this section, we have oversimplified things somewhat. (If you are really
curious about these issues, you can read ahead in Chapter 13.) Our purpose
here is simply to provide a very basic introduction to the logic behind statistical
Chapter4 The Experimental Method 115
inferences and to show you the basis upon which researchers working in the ex
perimental tradition make their claims for significance.
We willuse thesestatistical concepts in future chapters to illustrate the ways
classroom researchershave analyzed their data. At this point, we willsummarize
a study that illustrates many issues related to research design and threats to
validity.
A SAMPLE STUDY
Many years ago, Kathi Bailey was involved as an observer in a process-product
study, the results of which were never published. We willdescribe it here because
it is a beautiful example of a formal experiment in language teaching and because
there is much to be learned from the way the project evolved.
The research project was conducted at a large military school in the United
States,which had a program in Russian as a foreign language. The students were
young adults. The study was designed to determine whether the school's regular
method of teaching Russian was more effective, or whether Suggestopedia was
more effective.
Suggestopedia is a language teaching method that was originally developed
by Lozanov(1979; 1982). It uses musicto relax the students, who are given new
target culture identities in the class. There is little or no formal error correction
and students learn by hearing and seeing lengthy dialogs in the target language.
The method developsspeaking fluency, reading skills,and vocabulary. Suggesto
pedia normally requires extensive teacher training, and there is a Suggestopedia
institute that provides that training.
The military school's usual teaching method, in contrast, emphasized gram
matical accuracy. There were frequent tests, and error treatment was regularly
used to help the students improve their grammar, pronunciation, word knowl
edge, etc.
The researcher set up a classic textbook experiment. He was able to ran
domly select a group of forty people from the incoming Russian students. All
fortywere true beginners in Russian. He then randomly assigned thesestudents
to two groups,eachof whichcomprised two Russian classes. There were ten stu
dents in each class. Two classes were taught with Suggestopedia and two with the
regular teaching method. (See Figure 4.8.) The twoSuggestopedia classes were
taught by certified instructors who had been trained in the method. The other
classes were taught by regular employees at the school—all experienced Russian
teachers.
The students were all true beginners of the Russian language, so no pre-test
was necessary. But the researcher did administer an attitude questionnaire to all
forty students before the classes began, so that he could compare their attitudes
before and after the Russian course. At the end of the course, all forty students
were given the attitude questionnaire again, and their Russian ability was also
tested.
116 EXPLORING SECOND LANGUAGE CLASSROOM RESEARCH
Experimental
Sample
Control Group: Experimental Group:
Regular Teaching Suggestopedia
Method (n = 20) (n = 20)
r \ f
Class 1 Class 2 Class 3 Class 4
(n = 10) (n = 10) (n = 10) (n = 10)
\ J v J v J V
FIGURE 4.8 The control and experimental groups in the sample study
ACTION
Answer the following questions based on what you know about this study.
1. What were the likely research questions in this study?
2. What is the independent variable and how many levels does it have?
3. What are the dependent variables? (Hint: There are two.)
4. What is one control variable?
5. What is the design of this study? (Hint: This is a trick question. Think
about the dependent variables.)
6. Can you anticipate any possible threats to validity in this situation?
Compare your ideas with those of a classmate or colleague.
The researcher clearly understood the differences among process studies,
product studies, and process-product studies. (See Chapter 1.) For this reason,
he arranged for trained classroom observers to be present (one at a time) in some
of the Suggestopedia classes and some of the regular classes. The observers were
to take notes on the sessions so that if any significant differences were found
between the results yielded by the two teaching methods, those outcomes could
be linked to the actual teaching and learning processes documented by the
observers.
Several interesting problems arose as the study progressed. We report these
problems with great respect for the researcher, who had set up a well-designed
experiment.
Chapter 4 TheExperimental Method 117
First, the Suggestopedia teachers—who were visitors from out of town—
looked at the fifteen-week curriculum and said that they could cover that amount
of material in ten weeks instead of the fifteen weeks the regular classes would
take. Therefore, the Suggestopedia classes ended five weeks before the regular
classes, even though they had covered the same curriculum. At that point, the
students who had been in the Suggestopedia classes were tested and then inte
grated into the regular classes.
Secondly, due to the normal scheduling patterns at the school, the students
taught with the regular method had many different teachers during a typical day.
The two Suggestopedia classes each stayed widi one teacher for the entire day.
There were never substitutes in the Suggestopedia classes, but due to adminis
trative requirements of meetings and testing, substitute teachers were frequent
in the regular classes. (In fact, one day when Kathi Bailey entered a classroom
quite early, before the regular teacher had arrived, a student said, "Oh, no! Not
another one!" When the observer introduced herself and asked him what he
meant, he said he had thought she was another substitute teacher.) In addition,
one of the teachers using the regular method found it too stressful to be observed
so often and chose to drop out of the study.
On a different occasion, when Bailey was observing a Suggestopedia class,
another visitor was present. That person was introduced as a representative of
the Suggestopedia training program. When it was time for a break, the students
all left the classroom. That observer approached the teacher and said, "What are
you doing? This isn't Suggestopedia!" The teacher replied, "I know, but what
can I do? The students want grammar rules and error correction."
The students in both the control and experimental groups were young mil
itary personnel, who lived together in large dormitories. They had been assured
at the outset of the experiment that their performance on the post-test would in
no way influence their subsequent job postings. But the students in the
Suggestopedia classes apparently doubted this promise, and—knowing that
they would face the school's regular accuracy-oriented testing at the end of the
experiment—several of them began to study at night in the dormitories with
their friends who were in the regular classes. In fact, the two groups of students
regularly exchanged their Russian class materials.
REFLECTION
What threats to validity can you identify in the description above?
PAYOFFS AND PITFALLS
As the sample study illustrates, there are many pitfalls in using the experimental
method to conduct language classroom research. Many of these problems
stem from the difficulty of controlling all the possible confounding variables
118 EXPLORING SECOND LANGUAGE CLASSROOM RESEARCH
associatedwith research on human subjects. Human beings have agency, desires,
anxieties, and goalsthat are often far beyondthe abilityof the researcherto man
agein anycomprehensive way. As a result, even in the strongestresearch designs,
threats to validity sometimes arise that compromise the interpretation of the
results. In classroom research,
the securityof isolatingvariables and defining diem operationally, a secu
rity obtained by laboratory-like experiments and statistical inferences, is
largelylost, as the researcher is forced to look for determinants of learn
ing in the fluid dynamics of real-time contexts, (van Lier, 1998, p. 157)
Indeed, classroom research "entails a very large number of human and institu
tional factorsthat can affectresearch designand outcomesin manyunforeseen and
unforeseeable ways. It is not for the timid" (Rounds and Schachter, 1996,p. 108).
There is also a more philosophical problem with the experimental method.
For some people, if a phenomenon is not measurable, it is not worth studying.
For them, the psychometric tradition, and the experimental method in particu
lar, may be the only valuable ways of conducting research. However,in an effort
to quantify phenomena of interest we may miss important issues. We may not
investigate key variables because they are not easilyquantified, or we may focus
on trivial issuesthat are easyto quantify. Furthermore, the data collected in lan
guage classroom research are often collectedfrom people, but the need to quan
tify and the widespread use of group averages sometimes make it seem that
individual learners are represented only by test scores. And the seeming dehu-
manization of participants is also noticeable when researchers talk about the
sample in a study as "experimental subjects."
Another problem relates to the issue of objectivity and subjectivity. The ex
perimental method emphasizesobjectivity in hopes of counteracting the threats
of researcher and subject expectancy. As a result, teachers (and learners) have
typically not been seen as potential collaboratorsin languageclassroomresearch
conducted with the experimental method. Teachers have had important roles,
but often primarily as deliverers of a particular treatment in an experiment.
However, for teachers, it is sometimes very difficult to be simply a treatment de
liverer when students' needs and desires run counter to the prescribed treatment.
(Remember the Suggestopedia teacher's comment, "I know, but what can I do?
The students want grammar rules and error correction.")
There are also several payoffs associated with the experimental method.
First, it is a well-documented, highly codified approach to conducting educa
tional research with well-developed quality control procedures. There are many
textbooks available and it is relatively easy to locate courses if you wish to get fur
ther training.
Secondly, the experimental method is an internationally recognized way of
conducting research. If you conduct an experiment in Jakarta and publish your
findings in Prospect, the TESOL journal of Australia, readers in Germany, Iran,
India, Canada, and Brazilwillall understand the report (provided they have been
trained in the language and culture of the experimental method).
Chapter4 TheExperimentalMethod 119
Third, there are clear criteria for interpreting the outcomes in such re
search. The conceptof statistical significance provides the field withways to un
derstand whether a correlation is powerful enough or whether a difference in
scoresis big enough to warrant generalizing the resultsof the study.
Fourth, because of the emphasis on operational definitions and controlling
variables, it is sometimes possible to make comparisons across studies conducted
at different sites. The widelyunderstood research designs and numerous statisti
cal procedures allow researchers to replicate studies conducted by other people.
Finally, because it has high prestige in many contexts, using the experimen
tal method may enable researchers to obtain grant money or get their reports
publishedin venuesthat are not as accessible to those trying to publishaction re
search reports or the findings of naturalistic inquiry. (Of course, the research has
to be done well in order to be published. Simply using a prestigious research
method is no guarantee that your report will be accepted.)
CONCLUSIONS
The experimental method has been very important, historically, in all sorts of
research, in both the physical and social sciences. It has often been used in lan
guage classroom research, but with varied success, since it is so difficult to con
trol all the possible confounding variables that can arise in research with real
people. It is, however, a valuable approach to understanding teaching and
learning, and it has influenced many other approaches, as we will see in future
chapters.
QUESTIONS AND TASKS
1. Identify the specific design in each of the following research situations.
A. At the beginning of an advanced composition course for international
collegestudents, the instructor had each student write an in-classessay
about the differences between the educational systems of their host
country and their home countries. After ten weeks, the teacher had the
students write on the same topic. Then the two sets of essays were
mixed together at random and rated by another teacher who did not
know about the experiment.
B. A Greek teacher in Australia wanted to test the hypothesis that exposing
his students to the cultural setting of the target languagewould enhance
their language [Link] firstthree weeks of the class he tested them
everyweek on basic grammar points from the lessonsdiscussed in class.
The fourth week, he took them to a Greek festival, complete with food,
music, dancing, and traditional Greek clothing. The following three
weeks, he gave them weekly tests on the grammar lessons.
C. A French teacher in a secondary school wanted to know if using
new vocabulary words in a hands-on experience would help students
120 EXPLORING SECOND LANGUAGE CLASSROOM RESEARCH
remember them. So, she performed an experiment using her two be
ginning French classes as the groups to be compared. One group met
in the home economics classroom kitchen and made crepes while
learning the words for the ingredients. The other groupmet in class as
usual, read about crepes, and made vocabulary lists of the ingredients.
Afterthis experiment, the teacher tested both classes with a vocabulary
quiz to see which group did better.
D. The teacher of an elementary Spanish class wanted to know how well
the students would perform at the end iof the course since he had
changed to a newtextbook. He used the students' final exam scores at
the end of the semester as an indication of the success of the new text
book.
E. A language teacher wanted to determine the relationship between oral
proficiency and grammatical accuracy of his students. (He knew that
his teachingmethod stressedspeaking skills but that his students would
be faced with standardized tests of grammatical accuracy.) He designed
a study in which he could plot the students' scores on an oral profi
ciencyinterview against their scores on a 100-point grammar test.
E An ESL teacher in Scotland wanted to compare the listening compre
hension of his students from Asia, Latin America, Europe, and Africa.
He administered a test of English listening comprehension and then
computed the mean scores for the students from these four regions.
G. The ESL teacher in Scotland wondered whether those students who
had traveled in the United Kingdom had better listening comprehen-
sion than those who hadn't. For this reason, he added U.K. travel expe
rience as a moderator variable in his study comparing the listening
comprehension scores of students from Asia, Latin America, Europe,
and Africa.
H. A German language teacher wanted to know if viewing videotapes
about various aspects of German culture would increase her students'
vocabulary acquisition. So, she taught her German course for sixteen
weeks with the curriculum divided into four four-week modules. She
gave the students a vocabulary test at the end of every second week. In
the third week of each module, she showed a German film about art,
theatre, music, sports,etc. At the end of the semester, she comparedthe
four average scores from the weeks with films and the four average
scores from the weeks without films.
2. What is/are the research questions for each of the situations described
above? Can you state the hypothesis (or hypotheses) for each situation?
3. Not all the designs we have studied are represented in the paragraphs
above. Which ones are not represented here? Choose one of those missing
designs and write a brief scenario that provides sufficient detail for your
colleaguesor classmates to be able to identify which specificdesign you are
describing. Exchange your paper with a classmate or colleague.
Chapter4 The ExperimentalMethod 121
4. Turn back to Chapter 3 andreread the summary of Sato's (1982) investiga
tion of the turn-taking behaviors of the Asian and non-Asian students in
twoESL classes. How might the different datacollection procedures used
in those two classes lead to the instrumentation threat in her study?
5. Lookback at the vignettes of the sample studies in Chapters 2 and 3. Iden
tifythe threats to validity in those studies, usingthe vocabulary introduced
here.
6. The increasing strengths of the research designs enable researchers to cope
with the threats to validity, but such threats are often present. For each de
signdiscussed above, identify the specific threat(s) that mayinfluence your
interpretation of the findings.
7. Read the following situation and decide how you could design a study to
address your friend's request.
You havea friend who has been teachingEnglishinJapan for many
years. Your friend noticed that those of her students who often sang
English songs at karaoke clubs seemed to have better pronunciation,
better listeningskills, better speakingfluency, and better Englishvocab
ulary skills than those students who did not participate in karaoke
singing.
Your friend began to use karaoke singing regularly as an activity in
her English conversation classes for adult students. This experience
was so successful that she created a Web site, called "Karaoke Corner."
The Web site providesmusic and lyricsso students can sing along with
English songs in the privacy of their own homes. The Web site also
offers follow-up activities with the English song lyrics—cloze passages,
vocabulary quizzes, and reviews of grammar structures used in the
songs. The Web site has been running for over three years, and your
friend now has about four thousand clients (both adults and teenagers)
who visit the site regularly and pay to sing along in English. Yourfriend
also has two employeeswho request formal copyright permission to use
the songs and who regularly update the Web site with new songs and
activities.
Now your friend wants to expand the market for "Karaoke Corner"
by showinghowsuccessful its regularusers are in improvingtheir Eng
lish skills. She wants you to collect some data that would encourage
companies to pay for "Karaoke Corner" subscriptions for their employ
ees. Since she knows you are learning about conducting research, she
asks you to design a study that would provide evidencethat would show
potential corporate clients the effectiveness of the products and services
offered by the "Karaoke Corner" Web site.
A. Pose the research question(s) you wish to ask. You could write a formal
hypothesis if you wish.
122 EXPLORING SECOND LANGUAGE CLASSROOM RESEARCH
B. Identify the independent variable in this context. How many levels of
the independent variable will you use in your investigation?Why?
C. Determine what the dependent variable would be.
D. Do you want to incorporate a moderator variable? If so, what is it?
E. What are some possible confounding variablesyou should be aware of
in carrying out this research?
E What are some possible control variables you would like to impose?
G. Think about how you could operationallydefine all those variables.
H. Select a research design from among those that we have discussed.
I. Identify the strengths and weaknesses of the study as planned.
J. Identify the practical difficulties in carrying out this study as planned.
8. We noted at the beginning of this chapter that sometimes the jargon asso-
ciated with the experimental method can be a bit confusing. This is partly
the case because there are many synonyms. Look at the two columns of
words below. Draw a line matching each item in the left column with its
synonym in the right column.
Average Experimental method
Bar graph Practice effect
Normal distribution Generalizability
Testing threat Threats to validity
External validity Histogram
Confounding variables Mean
Scientific method Bell (shaped) curve
Hawthorne effect Reactiveeffects of experimental
arrangements
SUGGESTED READINGS
For more information on research designs, see Mitchell and Jolley (1988),
Shavelson (1981, 1996), and Tuckman (1999).
J. D. Brown (1988) provides a very good description of the normal distribu
tion. He also provides a clear introduction to the descriptive statistics (1988;
2001).
Chapter4 TheExperimental Method 123