Active Learning's Impact on Biology Learning
Active Learning's Impact on Biology Learning
Article
Submitted July 25, 2011; Revised August 24, 2011; Accepted September 2, 2011
Monitoring Editor: Daniel J. Klionsky
Previous research has suggested that adding active learning to traditional college science lectures
substantially improves student learning. However, this research predominantly studied courses
taught by science education researchers, who are likely to have exceptional teaching expertise. The
present study investigated introductory biology courses randomly selected from a list of prominent
colleges and universities to include instructors representing a broader population. We examined the
relationship between active learning and student learning in the subject area of natural selection.
We found no association between student learning gains and the use of active-learning instruction.
Although active learning has the potential to substantially improve student learning, this research
suggests that active learning, as used by typical college biology instructors, is not associated with
greater learning gains. We contend that most instructors lack the rich and nuanced understanding of
teaching and learning that science education researchers have developed. Therefore, active learning
as designed and implemented by typical college biology instructors may superficially resemble
active learning used by education researchers, but lacks the constructivist elements necessary for
improving learning.
394
a wide variety of science disciplines (Ruiz-Primo et al., 2011), (Gregory et al., 2011), 2) it is conceptually challenging for stu-
including biology (Jensen and Finley, 1996; Udovic et al., 2002; dents to learn (Bishop and Anderson, 1990; Nehm and Reilly,
Knight and Wood, 2005; Freeman et al., 2007; Nehm and Reilly, 2007; Gregory, 2009), and 3) well-developed instruments ex-
2007; Haak et al., 2011), physics (Shaffer and McDermott, 1992; ist to measure conceptual understanding of natural selection
Crouch and Mazur, 2001; Deslauriers et al., 2011), and chem- (e.g., Anderson et al., 2002; Nehm and Reilly, 2007; Nehm and
istry (Wright, 1996; Naiz et al., 2002). There have been so many Schonfeld, 2008). We contacted a total of 88 introductory biol-
papers documenting this trend that it is now widely accepted ogy instructors, sending at least three emails over the course
that students taught with active learning will learn substan- of a month and following up with at least one phone call.
tially more than students taught the same material with direct We made a final attempt 4 to 6 mo later to contact instructors
instruction. who did not respond to our initial queries.
However, a close review of the literature supporting the Of the 88 instructors from 77 institutions we invited to par-
effectiveness of active learning reveals a serious limitation: ticipate in our study, 33 (38%) agreed to participate fully; these
Most of the active-learning courses studied to date were instructors are hereafter referred to as “fully participating in-
taught by instructors who had science education research ex- structors.” These instructors taught 29 different courses at 28
perience. However, a close review of the literature supporting institutions in 22 states. Two of these institutions were private
the effectiveness of active learning revels a serious limita- and 26 were public. Seven institutions were on the U.S. News
tion: Most of the active learning courses studied to date were & World Report Best Colleges 2009 list. If an instructor declined
taught by instructors who had science education research ex- to participate in our study, we asked him or her to complete a
perience (e.g., Hake, 1998a,b; Knight and Wood, 2005; Deslau- survey describing his or her course and teaching methods so
riers et al., 2011). By science education research, we mean they we could account for nonresponse bias. We were able to col-
published papers on science education, received funding for lect data from an additional 22 instructors, which represents
education research, or attended conferences on science edu- 25% of the entire random sample of instructors and 44% of
cation research. We expect education researchers to have a the instructors who did not agree to fully participate in this
rich and nuanced understanding of their field. This expertise study. These instructors are hereafter referred to as “partially
may improve an instructor’s effectiveness in many ways, in- participating instructors.” Compared with previous research
cluding his or her ability to use active learning (Pollock and assessing the relationship between active-learning instruc-
Finkelstein, 2008; Turpen and Finkelstein, 2009). This limita- tion and student learning gains, our sample is the largest
tion has been recognized (Hake, 1998a; Pollock and Finkel- sample of instructors and institutions and the only sample
stein, 2008), but the implications of this potential problem randomly selected from a broad population of college science
have not been explored. In particular, we are concerned the instructors.
impressive learning gains documented in the active-learning
literature may not be representative of what typical instruc-
tors are likely to obtain.
The goal of this research was to address that gap by study- Assessing Learning
ing the relationship between the use of active-learning in- In each fully participating course, we assessed how much
struction and how much students learned about natural se- students learned about natural selection. We assessed learn-
lection in a random sample of introductory college biology ing by testing students near the beginning (pretest) and end
courses from around the United States. (posttest) of the term, using two instruments that measure
conceptual understanding of natural selection. First, we used
the Conceptual Inventory of Natural Selection—Abbreviated
(CINS-abbr), a 10-question, multiple-choice test (see sample
METHODS questions in Supplemental Material 1; Anderson et al., 2002;
Anderson, 2003; Fisher et al., unpublished data.). The ques-
Sample tions on this instrument are nearly identical to those on an
The goal of our sampling design was to infer results to in- original concept inventory with well-established reliability
troductory biology courses at major colleges and universities (Anderson et al., 2002). Each distracter, or wrong answer, was
throughout the United States. Thus, we began with a list of the designed to appeal to students who hold common miscon-
two largest public colleges and universities from each of the ceptions about natural selection. The content and face validity
50 states, plus a list of the 50 top-ranked colleges and univer- of these questions has been established for both the original
sities from the U.S. News & World Report Best Colleges 2009 CINS and the CINS-abbr (Anderson et al., 2002; Fisher et al.,
ranking; some institutions were on both lists. From this com- unpublished data.). Second, students completed one open-
bined list of 144 institutions, we randomly selected 77. We ended question from a set of five questions developed by
then contacted instructors at these institutions to ask them to Bishop and Anderson (1990) and later revised by Nehm and
participate during one of three consecutive semesters in 2009 Reilly (2007) to measure college biology majors’ understand-
and 2010. ing of natural selection. This set of open-ended questions
In each school, we sought out introductory biology courses was designed to assess student understanding across differ-
that taught natural selection and were designed for biol- ent levels of Bloom’s taxonomy (Nehm and Reilly, 2007). We
ogy majors. To identify appropriate courses and course in- used a question at Bloom’s “application” level, which tests a
structors, we used information gathered on institution web- student’s ability to apply knowledge to a novel question. This
sites and from biology department staff. We chose to survey set of questions tends to be more difficult for students than
courses teaching natural selection, because 1) it is a mech- CINS questions (Nehm and Schonfeld, 2008). The question
anism of evolution and therefore a core concept in biology we used, hereafter called the “cheetah question,” was:
Effect size (Cohen’s d for repeated measures)a tp [2(1 – r)/(n)]1/2 tp = t statistic from Student’s paired t
r = correlation between pre/post
n = students completing pre/post
Average normalized gainb (Post − Pre)/(10 − Pre)c Post = mean course posttest score
Pre = mean course pretest score
Percent change (Post − Pre)/Pre Post = mean course posttest score
Pre = mean course pretest score
Raw change Post − Pre Post = mean course posttest score
Pre = mean course pretest score
a See Dunlap et al. (1996).
b See Hake (1998a).
c This is the equation for the CINS-abbr. The cheetah question equation would be (Post-Pre)/(9 − Pre).
Cheetahs (large African cats) are able to run faster than between courses using online testing and those using paper
60 miles per hour when chasing prey. How would a testing (see full results in the Supplemental Material).
biologist explain how the ability to run fast evolved In some courses, students earned nominal course credit for
in cheetahs, assuming their ancestors could run only
20 miles per hour?
logging into the online test, but actually completing test ques-
tions was voluntary in all classes. To look for differences be-
To score student responses to the cheetah question, we tween test performance in courses in which students earned
developed, piloted, refined, and applied a coding rubric credit and courses in which students did not, we again used
(Supplemental Material 2). Biology experts (T.M.A. and independent samples t tests. We found only one significant
S.T.K.) designed the rubric after reviewing a rubric previ- difference between courses in which students earned credit
ously developed for the cheetah question (Nehm and Reilly, and those in which students did not, and it was in the oppo-
2007). Our rubric gave more weight to three concepts we site direction than would be expected if awarding credit led to
felt were core concepts a student must understand in order increased participation or performance. Students in courses
to understand natural selection: the existence of phenotypic in which credit was not awarded had significantly higher
variation within a population, the heritability of that varia- scores on the posttest CINS-abbr than students in courses
tion, and differential reproductive success among individu- that awarded credit (p = 0.03; see full results in the Supple-
als. We gave less weight to three additional concepts we felt mental Material). Because awarding credit was not associated
were representative of more advanced understanding: the with improved test performance or learning gains as would
causes of variation, a change in the distribution of individ- be predicted, we did not include this as a variable in further
ual traits within a population, and change taking place over analysis.
many generations. We designed this coding rubric to be sen-
sitive to developing understanding, while allowing room for
students to demonstrate more advanced understanding. To Calculating Learning Gains
establish interrater reliability (IRR), two researchers (T.M.A. To decide how best to calculate how much students learned
and C.A.C.) independently scored a random sample of 210 about natural selection, we examined the intercorrelations
responses. IRR was measured using Pearson’s correlation. among pre- and posttest scores and four possible calculations
There was a strong correlation between the total essay score of learning gains: effect size (Cohen’s d), average normalized
awarded by the two researchers (r = 0.93, p < 0.0001). The re- gain, percent change, and raw change (Table 1). The calcu-
searchers then independently scored the remaining responses lations of learning gains were highly intercorrelated (see the
to the cheetah question using the coding rubric. Supplemental Material), with Pearson’s r ranging from 0.79–
Due to the large number of students included in this study 0.99 (all p values < 0.001; p was calculated using the Holm-
(more than 8000), we scored a subsample of student responses Bonferonni method to account for error associated with mul-
to the cheetah question from each course. We randomly se- tiple comparisons). Although average normalized gain is a
lected ∼50 students from each course and scored responses commonly used estimator of learning gains in research on ac-
from students who completed both pre- and posttest cheetah tive learning (Hake, 1998a; Crouch and Mazur, 2001; Knight
questions. The mean subsample size was 42 students (SD = and Wood, 2005), it was strongly correlated with mean course
12). For analyses of learning gains on the cheetah question pretest scores on the CINS-abbr (r = 0.51, p = 0.012; Supple-
described below, we excluded three courses whose subsam- mental Material). This correlation means that courses with
ple included fewer than 20 students and one course in which high average pretest scores receive relatively higher normal-
pre- and posttest responses could not be matched by student. ized gains than courses with lower average pretest scores. Ad-
Students completed the CINS-abbr and the cheetah ques- ditionally, percent change was strongly negatively correlated
tion on paper or online. To test for differences between stu- with pretest scores on the cheetah question (r = –0.62, p =
dent performance on paper versus performance using online 0.003; Supplemental Material), resulting in a bias in the other
instruments, we used independent samples t tests. We found direction. We ultimately chose to calculate learning gains us-
no significant differences in mean test scores or learning gains ing Cohen’s d for a repeated measures design (Dunlap et al.,
Table 2. Percent of instructors reporting how often they use specific active-learning exercises
More than once Once per Once per Never (or almost
Exercise per class (%) class (%) week (%) never) (%)
Activities in which students use data to answer 5.7 2.9 34.3 57.1
questions while working in small groups
Student discussions in pairs or small groups to answer a 17.1 17.1 22.9 42.9
question
Individual writing activities that require students to 0 3.1 28.1 68.8
evaluate their own thinkinga
Clicker questions that test conceptual understanding 34.3 11.4 5.7 48.6
Classroom-wide interactions that require students to 8.6 20.0 37.1 34.3
apply principles presented in class to a novel question
Other small group activities 5.7 8.6 25.7 60.0
an = 32; for all others n = 33.
1996). No calculation of learning gains is without problems, ming the frequencies they reported for all six categories of
so we also repeated the analyses described in Data Analysis, exercises. To do so, we assumed each course met three times
using each of the four calculations of learning gains. If all per week and counted “Once per week” as once per week,
analyses produced similar results, we would feel confident “Once per class” as three times per week, and “More than
that the way we chose to quantify student learning was not once per class” as six times per week. This variable may un-
impacting our overall results. derestimate the use of active learning by excluding other ex-
ercises an instructor was using to promote active learning
and by limiting “More than once per week” to only six ex-
Surveying Teaching Methods and Course Details
ercises per week, so we also asked instructors to report their
We gathered details on each course from the instructor and general use of any active learning by asking how often they
the students. An online survey (Supplemental Material 3) used exercises meeting Hake’s (1998a) definition of interac-
was used to gather data from fully participating instructors, tive engagement (another commonly used term for active
as well as partially participating instructors. The instructor learning):
survey solicited information about the course, the instruc-
tor’s teaching experience and teaching methods, and the stu- activities designed at least in part to promote concep-
dents’ backgrounds. We also surveyed students during the tual understanding through interactive engagement of
posttest about their instructor’s teaching methods and their students in heads-on (always) and hands-on (usually)
perceptions of the course (Supplemental Material 4). activities which yield immediate feedback through dis-
To corroborate self-report data from instructors, the instruc- cussion with peers and/or instructors.
tor and his or her students answered an identical question
about the instructor’s use of active learning (Supplemental Finally, we asked instructors how many active-learning ex-
Material 3, question 8; Supplemental Material 4, question 3). ercises they used during the section of the course dedicated
Student reports of active learning agreed with instructor re- to teaching natural selection. For all three questions, we de-
ports. Agreement between instructor responses and the most scribed exercises instead of using common names (e.g., peer
common student response (i.e., the mode) in each course was instruction, think–pair–share), so instructors would not have
calculated using Cohen’s kappa, which indicated substan- previous associations with the exercises described.
tial agreement (κ = 0.69; Viera and Garrett, 2005). Therefore, We found instructors’ reports of using specific active-
we used instructor reports of active learning for all further learning exercises during the course were strongly correlated
analyses. with their reports of using active learning in teaching natu-
Previous research has typically categorized instructors’ ral selection (r = 0.52, p = 0.002). We therefore chose to use
methods as either “active learning” or “traditional lectures,” instructor reports about active learning use throughout the
but as active-learning methods have become more widely course for further analyses. In contrast to our expectations, in-
used, this categorization no longer adequately captures structors provided more conservative estimates of their gen-
the variation among instructors’ teaching methods. We ap- eral use of any active learning (as defined by Hake, 1998a)
proached the problem of measuring an instructor’s use of than their estimates of their use of specific active-learning
active learning by asking several questions and examining exercises (Figure 1). For example, instructors who reported
the relationships among instructor responses to these ques- using general active-learning methods just once per week re-
tions. We asked instructors three questions about their use ported a mean of 3.33 specific exercises per week. Ultimately,
of active learning in the lecture portion of their course. First, we decided to quantify an instructor’s use of active learning
we asked instructors to report how often they used specific as the weekly frequency with which they used specific active-
active-learning exercises (described in Table 2) previously learning exercises, because this quantification allowed us to
shown to be effective (Ebert-May et al., 1997; Crouch and capture more variability among instructor methods. How-
Mazur, 2001; Andrews et al., 2011; Deslauriers et al., 2011). ever, we also conducted statistical analyses with the more
We then created a continuous variable describing an instruc- general report of active learning to assure results remained
tor’s weekly use of these active-learning exercises by sum- the same.
Figure 1. Comparison of instructor reports of their weekly use of specific active-learning exercises and instructor reports of their general
use of active-learning exercises as defined by Hake (1998a). The line in the middle of the box represents the median weekly frequency of
active-learning use for instructors in the group. The top of the box represents data points in the 75th percentile and the bottom of the box
represents data points in the 25th percentile. The space within the box is called the interquartile range (IQR). Whiskers represent the lowest
and highest data points no more than 1.5 times the IQR above and below the box. Data points not included in this range are represented
as circles.
Data Analysis gains, as well as a model that replaced the continuous weekly
We used data gathered from the instructor survey to compare frequency of specific active-learning exercises with the more
fully participating instructors with partially participating in- general categorical use of any active learning, to see if results
structors to check for selection bias resulting from nonre- remained the same. We checked assumptions for our models
sponse. We looked for differences using independent sam- using QQ plots and plots of fitted values versus residuals.
ples t tests and Fisher’s exact tests. We found no differences Assumptions were met for all linear regression models.
between the two groups of instructors, suggesting nonre- Many factors affect how much students learn in a course, so
sponse did not cause a selection bias. There were no sig- we used data gathered from the instructor survey and student
nificant differences in the mean number of specific active- survey to control for variation in learning gains due to fac-
learning exercises used per week, mean class size, mean tors other than the use of active learning. We included several
teaching experience, mean class time dedicated to teaching continuous control variables in our linear models, including
natural selection, or mean attendance rates. Neither were the number of years an instructor had taught college biol-
there differences in type of institution (public or private), ogy, hours of class time devoted to teaching natural selection,
the list their institution came from (large public institutions proportion of students who attended class regularly, propor-
or most prestigious institutions), the instructor’s position, or tion of students who completed both the pre- and posttest,
the frequency with which they used general active-learning and class size. Student responses to questions about how
methods (see full results in the Supplemental Material). difficult they found the course compared with previous sci-
To answer our question of interest—is active-learning in- ence courses and how interesting they found the course were
struction positively associated with student learning gains coded numerically and also included as continuous control
in typical college biology courses?—we used general linear variables. Students chose from a Likert scale (Supplemental
regression models. We used one model with effect size of Material 4), which we then coded from one to five, where one
learning on the CINS-abbr as the response variable and one corresponded to “Very uninteresting” and “Much less diffi-
model with effect size of learning on the cheetah question cult” and five corresponded to “Very interesting” and “Much
as the response variable. We used two models, because four more difficult.” We then calculated means for each course.
courses had insufficient data to analyze learning gains on We used indicator variables to include categorical control
the cheetah question, but had complete CINS-abbr data. Us- variables in our models. We included a factor accounting for
ing two models allowed us to avoid unnecessarily excluding the presence or absence of nonmajors in a course. We also in-
these courses from all analyses. Additionally, we examined cluded a two-level factor for the instructor’s position: tenure
linear regression models using other calculations of learning track or non-tenure track. Last, we included two factors to
Table 3. Instructor reports of the frequency with which they use Table 4. Descriptive statistics for course pre- and posttest scores on
active-learning exercisesa the CINS-abbr and the cheetah question
DISCUSSION
We have shown that even though instructors of introductory
college biology courses are using active learning, students in
many of their courses have learned very little about natural
selection. Notably, students in most courses were no more
successful in applying their knowledge of natural selection
to a novel question at the end of the course than they were at
the beginning of the course. The absence of a relationship be-
tween active learning and student learning is in stark contrast
to a large body of research supporting the effectiveness of ac-
tive learning. We attribute this contrast to the fact that we
studied a different population of instructors. We randomly
sampled college biology faculty from a list of major universi-
ties. Therefore, instructors using active learning in our study
represent the range of science education expertise among in-
troductory college biology instructors using these methods.
In contrast, most of the faculty using active learning in previ-
ous studies had backgrounds in science education research.
The expertise gained during research likely prepares these in-
structors to use active learning more effectively (Pollock and
Finkelstein, 2008; Turpen and Finkelstein, 2009).
Specifically, it is possible that a thorough understanding
of, commitment to, and ability to execute a constructivist
approach to teaching are required to successfully use active
learning (NRC, 2000). Constructivism—the theory that stu-
dents construct their own knowledge by incorporating new
ideas into an existing framework—likely permeates all as-
Figure 2. Relationship between learning gains (Cohen’s d) and the pects of education researchers’ instruction, including how
number of active-learning exercises an instructor used per week.
The number of active-learning exercises per week was calculated
they use active learning. Without this expertise, the active-
by summing the number of times per week instructors reported learning exercises an instructor uses may have superficial
using all of the exercises described in Table 2. (A) Learning gains similarities to exercises described in the literature, but may
on the CINS-abbr (n = 33). (B) Learning gains on the cheetah lack constructivist elements necessary for improving learning
question (n = 29). (NRC, 2000). For example, our results suggest that addressing
common student misconceptions may lead to higher learning
gains. Constructivist theory argues that individuals construct
selection (Sinatra et al., 2008; Kalinowski et al., 2010). Because new understanding based on what they already know and
common misconceptions are used as distracters in CINS-abbr believe (Piaget, 1973; Vygotsky, 1978; NRC, 2000), and what
questions, we would expect courses in which misconceptions students know and believe at the beginning of a course is
were directly targeted to have higher learning gains on this often scientifically inaccurate (Halloun and Hestenes, 1985;
instrument. That said, misconceptions seem to be the largest Bishop and Anderson, 1990; Gregory, 2009). Therefore, con-
barrier to understanding that students face when learning structivist theory argues that we can expect students to retain
natural selection (Bishop and Anderson, 1990; Gregory, 2009), serious misconceptions if instruction is not specifically de-
so a test that measures the extent to which students reject mis- signed to elicit and address the prior knowledge students
conceptions is likely to be a reliable measure of their overall bring to class.
understanding (Nehm and Schonfeld, 2008). Further research A failure to address misconceptions is just one example
will be necessary to determine the relationship between how of how active-learning instruction may fall short. Instructors
students learn natural selection and how an instructor ad- may fail to achieve the potential of active learning in the de-
dresses common misconceptions about natural selection. sign or implementation of exercises, or both. There are many
Table 5. Results of linear models examining the relationship between student learning gains (Cohen’s d) and active-learning instruction
* p < 0.05
possible ways that active-learning exercises could be poorly nected to other material in the course, such that students fail
designed. For example, questions used in an exercise may to see important relationships among concepts (NRC, 2000).
only require students to recall information, when higher- It is also possible that the active-learning exercises used to
order cognitive processing (e.g., application) is required to discuss fundamental theories may not be sufficiently inter-
fully grasp scientific concepts (Crowe et al., 2008). Alter- esting to students to motivate them to participate (Boekaerts,
natively, questions posed to students could be poorly con- 2001).
Table 6. Comparisons between the direction and significance of the association between explanatory variables in the CINS-abbr linear
model and different calculations of learning gains as the response variable
Linear model coefficient Effect size Average normalized gain Percent change Raw change
Intercept −* −* − −
Weekly active learning − − − −
Instructor position (tenure track) − − + −
Students regularly attending (%) − + − −
Hours spent on natural selection − − + +
Class size − − + −
Years of teaching experience + + − +
Students pre/posttest (%) + + + +
Misconceptions (explained)a +* + +* +
Misconceptions (active learning +* + + +
otherwise)b
Course difficulty (student-rated)c +* +* + +
Student interest in course +* + +* +*
Nonmajors (absent)d + + +* +
(−) indicates a negative association with learning in the model and (+) indicates a positive association with learning.
a Two-level factor: Instructor did or did not explain why misconceptions are incorrect.
b Two-level factor: Instructor did or did not use active-learning exercises and otherwise make a substantial effort toward correcting
misconceptions.
c Relative to past science courses the student had taken.
d Two-level factor: Presence or absence of nonbiology majors in the course.
*p < 0.05
Figure 3. Relationship between four different calculations of learning gains on the CINS-abbr and the number of active-learning exercises an
instructor used per week. The CINS-abbr was scored out of 10 points, so a raw change of one is equivalent to earning one more point on the
posttest than on the pretest. Overall, these graphs are very similar; there is no evidence of a positive relationship between learning gains and
the use of active-learning instruction, no matter how we calculate learning gains.
On the other hand, regardless of how well an active- aware of their own erroneous ideas (Crouch et al., 2004).
learning exercise is designed, an instructor must make many Furthermore, instructors may display any number of sub-
implementation decisions that will ultimately affect the suc- tle behaviors or attitudes that influence the extent to which
cess of the exercise. For example, a think–pair–share dis- students participate in active-learning exercises, and thereby
cussion may not be effective if the instructor does not al- affect how much students learn (Turpen and Finkelstein, 2009,
low students enough time to think about a question (Allen 2010).
and Tanner, 2002). Or, an instructor may solicit only one an- Our results corroborate research showing that college sci-
swer from the class and therefore fail to expose the range ence teachers are incorporating active-learning methods but
of ideas held by students. In addition, an instructor may are often doing so ineffectively. A recent study compared
not ask students to predict the outcome of a demonstration college biology instructor’s self-reports of teaching with ex-
or thought experiment and therefore fail to make students pert observations of the instructor’s teaching and found that
the instructors felt they were using reform methods, but ex- ACKNOWLEDGMENTS
perts did not confirm this (Ebert-May et al., 2011). Similarly,
Support for this study was provided by NSF-CCLI 0942109. We thank
a national survey of teaching practices in college physics the instructors and course coordinators who dedicated time to this
courses found that 63.5% of instructors reported using think– study: Theresa Theodose, Heather Henter, Brian Perry, Carla Hass,
pair–share discussions, but 83% of the instructors who used Alan L. Baker, Clyde Herreid, Norris Armstrong, Farahad Dastoor,
this method did not use it as suggested by researchers Michael R. Tansey, Denise Woodward, Kristen Porter-Utley, Leana
(Henderson and Dancy, 2009). Mounting evidence suggests Topper, Susan Piscopo, Rogene Schnell, Brent Ewers, John Longino,
that somewhere in the communication between science ed- Waheeda Khalfan, Denise Kind, Scott Freeman, Jon Sandridge, Re-
bekka Darner, Peter Houlihan, Jacob Krans, Wyatt Cross, Peter Dunn,
ucation researchers and typical college science instructors, Don Waller, Scott Solomon, Benjamin Normark, Jason Flores, Teena
elements of evidence-based methods and curricula crucial to Michael, Drew Joseph, Dmitri Petrov, Dustin Rubenstein, and oth-
student learning are lost. ers who wish to remain anonymous. We also thank the students in
The results of this study have three implications for educa- participating courses. Finally, we thank Tatiana Butler, Megan Higgs,
tion researchers across science disciplines. First, we need to Scott Freeman, and two anonymous reviewers for assistance with re-
build a better understanding of what makes active-learning search, analysis, and manuscript revision. This research was exempt
from the requirement of review by the Institutional Review Board,
exercises effective by rigorously exploring which elements are no. SK082509-EX.
necessary and sufficient to improve learning (e.g., see Crouch
et al., 2004; Smith et al., 2009, 2011; Perez et al., 2010). Sec-
ond, we need to develop active-learning exercises useful for
a broad population of instructors. Third, we need to identify REFERENCES
what training and ongoing support the general population of
Allen D, Tanner K (2002). Answers worth waiting for: one second is
college science faculty and future faculty need to be able to
hardly enough. Cell Biol Educ 1, 3–5.
effectively use active learning, taking into account obstacles
instructors will face, including individual, situational, and Allen D, Tanner K (2005). Infusing active learning into the large-
enrollment biology class: seven strategies, from the simple to com-
institutional barriers to reform (Henderson, 2005; Henderson plex. Cell Biol Educ 4, 262–268.
and Dancy, 2007, 2008).
Our results also have two important implications for in- Anderson DL (2003). Natural selection theory in non-majors biology:
instruction, assessment, and conceptual difficulty. PhD Dissertation,
structors. First, no one can assume that they are teaching San Diego, CA: University of California and San Diego State Univer-
effectively just because they are using active learning. There- sity.
fore, instructors need to carefully assess the effectiveness
Anderson DL, Fisher KM, Norman GJ (2002). Development and eval-
of their instruction to determine whether active learning is uation of the conceptual inventory of natural selection. J Res Sci Teach
reaching its potential. There are a growing number of re- 39, 952–978.
liable and valid multiple-choice and essay tests that assess
Andrews TM, Kalinowski ST, Leonard MJ (2011). “Are humans
student knowledge (Anderson et al., 2002; Baum et al., 2005; evolving?” A classroom discussion to change student misconcep-
Nehm and Reilly, 2007; Smith et al., 2008; Nadelson and tions regarding natural selection. Evol Educ Outreach 4, 456–466.
Southerland, 2009). We recommend using these tests in a
Angelo TA, Cross KP (1993). Classroom Assessment Techniques: A
pre/posttest design to assess the effectiveness of instruction, Handbook for College Teachers, 2nd ed., San Francisco, CA: Jossey-
as well as using formative assessments to monitor learn- Bass.
ing throughout instruction (Angelo and Cross, 1993; Marrs
Baum DA, Smith SD, Donovan SSS (2005). The tree-thinking chal-
and Novak, 2004). Second, instructors should assume stu- lenge. Science 310, 979–980.
dents enter science courses with preexisting ideas that im-
Bishop B, Anderson C (1990). Student conceptions of natural selection
pede learning and that are unlikely to change without in-
and its role in evolution. J Res Sci Teach 27, 415–427.
struction designed specifically for that purpose (NRC, 2000).
To replace students’ misconceptions with a scientifically ac- Boekaerts M (2001). Context sensitivity: activated motivational be-
liefs, current concerns and emotional arousal. In: Motivation in
cepted view of the world, instructors need to elicit miscon-
Learning Contexts: Theoretical and Methodological Implications, ed.
ceptions, create situations that challenge misconceptions, and S. Volet and S. Järvelä, Oxford, UK: Elsevier Science.
emphasize conceptual frameworks, rather than isolated facts
Bonwell CC, Eison JA (1991). Active Learning: Creating Excitement
(Hewson et al., 1998; Tanner and Allen, 2005; Kalinowski et al., in the Classroom, ASHE-ERIC Higher Education Report No. 1, Wash-
2010). ington, DC: George Washington University School of Education and
Our study revealed that active learning was not as- Human Development.
sociated with student learning in a broad population of Boyer Commission on Educating Undergraduates in the
introductory college biology courses. These results imply ac- Research University (1998). Reinventing Undergraduate Ed-
tive learning is not a quick or easy fix for the current defi- ucation: A Blueprint for America’s Research Universities.
ciencies in undergraduate science education. Simply adding [Link] (accessed 19 July 2011).
clicker questions or a class discussion to a lecture is unlikely to Crouch CH, Fagen AP, Callan JP, Mazur E (2004). Classroom
lead to large learning gains. Effectively using active learning demonstrations: learning tools or entertainment? Am J Phys 72,
requires skills, expertise, and classroom norms that are fun- 2004.
damentally different from those used in traditional lectures. Crouch CH, Mazur E (2001). Peer instruction: ten years of experience
Appreciably improving student learning in college science and results. Am J Phys 69, 970–976.
courses throughout the United States will likely require re- Crowe A, Dirks C, Wenderoth MP (2008). Biology in bloom: imple-
forming the way we prepare and support instructors and the menting Bloom’s Taxonomy to enhance student learning in biology.
way we assess student learning in our classrooms. CBE Life Sci Educ 7, 368–381.
Deslauriers L, Schelew E, Wieman C (2011). Improved learning in a McConnell DA, et al. (2006). Using conceptests to assess and im-
large-enrollment physics class. Science 332, 862–864. prove student conceptual understanding in introductory geoscience
courses. J Geoscience Educ 54, 61–68.
Dunlap WP, Cortina JM, Vaslow JB, Burke MJ (1996). Meta-analysis
of experiments with matched groups or repeated measures designs. McKeachie WJ, Pintrich PR, Lin Y-G, Smith DAF, Sharma R (1990).
Psychol Methods 1, 170–177. Teaching and Learning in the College Classroom: A Review of the Re-
search Literature, 3rd ed., Ann Arbor: University of Michigan Press.
Ebert-May D, Brewer C, Allred S (1997). Innovation in large lectures—
teaching for active learning. Bioscience 47, 601–608. Nadelson LS, Southerland SA (2009). Development and preliminary
evaluation of the measure of understanding of macroevolution: in-
Ebert-May D, Derting TL, Hodder J, Momsen JL, Long TM,
troducing the MUM. J Exp Educ 78, 151–190.
Jardeleza SE (2011). What we say is not what we do: effective
evaluation of faculty development programs. Bioscience 61, 550– Naiz M, Aguilera D, Maza A, Liendo G (2002). Arguments, contradic-
558. tions, resistances, and conceptual change in students’ understanding
of atomic structure. Sci Educ 86, 505–525.
Fisher K, Williams KS, Lineback JE, Anderson D (in prep.). Concep-
tual Inventory of Natural Selection—Abbreviated (CINS-abbr). National Research Council (NRC) (1997). Science Teaching Reconsid-
ered: A Handbook, Washington, DC: National Academies Press.
Freeman S, O’Connor E, Parks JW, Cunningham M, Hurley D, Haak
D, Dirks C, Wenderoth MP (2007). Prescribed active learning in- NRC (2000). How People Learn: Brain, Mind, Experience, and School,
creases performance in introductory biology. CBE Life Sci Educ 6, Washington, DC: National Academies Press.
132–139.
NRC (2003) Improving Undergraduate Instruction in Science,
Gregory E, Ellis JP, Orenstein AN (2011). A proposal for a common Technology, Engineering, and Mathematics: Report of a Workshop,
minimal topic set in introductory biology courses for majors. Am Biol Washington, DC: National Academies Press.
Teach 73, 16–21.
NRC (2004). BIO2010: Transforming Undergraduate Education for
Gregory TR (2009). Understanding natural selection: essential con- Future Research Biologists, Washington, DC: National Academies
cepts and common misconceptions. Evol Educ Outreach 2, 156– Press.
175.
National Science Foundation (1996) Shaping the Future: New Expe-
Haak DC, HilleRisLambers J, Pitre E, Freeman S (2011). Increased riences for Undergraduate Education in Science, Mathematics, Engi-
structure and active learning reduce the achievement gap in intro- neering, and Technology. Report of the Advisory Committee to the
ductory biology. Science 332, 1213–1216. NSF Directorate for Education and Human Resources, Washington,
DC: National Science Foundation.
Hake RR (1998a). Interactive-engagement versus traditional meth-
ods: a six-thousand-student survey of mechanics test data for intro- Nehm RH, Reilly L (2007). Biology majors’ knowledge and miscon-
ductory physics courses. Am J Phys 66, 64–74. ceptions of natural selection. BioScience 57, 263–272.
Hake RR (1998b). Interactive engagement methods in introductory Nehm RH, Schonfeld IS (2008). Measuring knowledge of natural
physics mechanics courses. Am J Phys 66, 1. selection: a comparison of the CINS, and open-response instrument,
and an oral interview. J Res Sci Teach 45, 1131–1160.
Halloun IA, Hestenes D (1985). The initial knowledge state of college
physics students. Am J Phys 53, 1043–1055. Nelson CE (2008). Teaching evolution (and all of biology) more effec-
tively: strategies for engagement, critical reasoning, and confronting
Handelsman, J et al. (2005). Scientific teaching. Science 23, 521–511. misconceptions. Integr Comp Biol 48, 213–225.
Henderson C (2005). The challenges of instructional change under the Perez KE, Stauss EA, Downey N, Galbraith A, Jeanne R, Cooper
best of circumstances: a case study of one college physics instructor. S (2010). Does displaying the class results affect student discussion
Am J Phys 73, 778–786. during peer instruction? CBE Life Sci Educ 9, 133–140.
Henderson C, Dancy MH (2007). Barriers to the use of research- Piaget J (1973). The Language and Thought of the Child, London:
based instructional strategies: the influence of both individual and Routledge and Kegan Paul.
situational characteristics. Phys Rev PER 3, 020102.
Pollock SJ, Finkelstein ND (2008). Sustaining educational reforms in
Henderson C, Dancy MH (2008). Physics faculty and educational introductory physics. Phys Rev PER 4, 010110.
researchers: divergent expectations as barriers to the diffusion of
innovations. Am J Phys 76, 79–91. Ruiz-Primo MA, Briggs D, Iverson H, Talbot R, Shepard LA (2011).
Impact of undergraduate science course innovations on learning.
Henderson C, Dancy MH (2009). Impact of physics education re- Science 331, 1269–1270.
search on the teaching of introductory quantitative physics in the
United States. Phys Rev PER 5, 020107. Shaffer PS, McDermott LC (1992). Research as a guide for curricu-
lum development: an example from introductory electricity. Part II:
Hewson PW, Beeth ME, Thorley NR (1998). Teaching for conceptual Design of instructional strategies. Am J Phys 60, 1003–1013.
change. In: International Handbook for Science Education, ed. BJ
Fraser and KG Tobin, London: Kluwer Academic. Sinatra GM, Brem SK, Evans M (2008). Changing minds? Implications
of conception change for teaching and learning biological evolution.
Jensen MS, Finley FN (1996). Changes in students’ understanding of Evol Educ Outreach 1, 189–195.
evolution resulting from different curricular and instructional strate-
gies. J Res Sci Teach 33, 879–900. Smith MK, Wood WB, Adams WK, Wieman C, Knight JK, Guild N,
Su TT (2009). Why peer discussion improves student performance
Kalinowski ST, Leonard MJ, Andrews TM (2010). Nothing in evolu- on in-class concept questions. Science 323, 122–124.
tion makes sense except in the light of DNA. CBE Life Sci Educ 9,
Smith MK, Wood WB, Knight JK (2008). The genetics concept assess-
87–97.
ment: a new concept inventory for gauging student understanding
Knight JK, Wood WB (2005). Teaching more by lecturing less. Cell of genetics. CBE Life Sci Educ 7, 422–430.
Biol Educ 4, 298–310.
Smith MK, Wood WB, Krauter K, Knight JK (2011). Combining
Marrs KA, Novak G (2004). Just-in-time teaching in biology: creating peer discussion with instructor explanation increases student learn-
an active learner classroom using the internet. Cell Biol Educ 3, 49– ing from in-class concept questions. CBE Life Sci Educ 10, 55–
61. 63.
Tanner K, Allen D (2005). Approaches to biology teaching and tive learning in an introductory biology course. BioScience 52, 272–
learning: understanding wrong answers–teaching toward concep- 281.
tual change. Cell Biol Educ 4, 112–117.
Viera AJ, Garrett JM (2005). Understanding interobserver agreement:
Turpen C, Finkelstein ND (2009). Not all interactive engagement is the kappa statistic. Fam Med 37, 360–363.
the same: variations in physics professors’ implementation of peer
Vygotsky LS (1978). Mind in Society: The Development of the Higher
instruction. Phys Rev PER 5, 020101.
Psychological Processes, Cambridge, MA: Harvard University
Turpen C, Finkelstein ND (2010). The construction of different class- Press.
room norms during peer instruction: students perceive differences.
Wright JC (1996). Authentic learning environment in analyti-
Phys Rev PER 6, 020123.
cal chemistry using cooperative methods and open-ended lab-
Udovic D, Morris D, Dickman A, Postlethwait J, Wetherwax P oratories in large lecture courses. J Chem Educ 73, 827–
(2002). Workshop biology: demonstrating the effectiveness of ac- 832.