Chapter 5-6 Exercise Packet
Chapter 5-6 Exercise Packet
ONE-WAY ANOVA
5.6 Exercises
Conceptual Exercises
5.1 Two-sample t-test? True or False: When there are only two groups, an ANOVA model is
equivalent to a two-sample pooled t-test.
5.2 Random selection. True or False: Randomly selecting units from populations and perform-
ing an ANOVA allow you to generalize from the samples to the populations.
5.3 No random selection. True or False: In datasets in which there was no random selection,
you can still use ANOVA results to generalize from samples to populations.
5.4 Independence transformation? True or False: If the dataset does not meet the indepen-
dence condition for the ANOVA model, a transformation might improve the situation.
5.5 Comparing groups, two at a time. True or False: It is appropriate to use Fisher’s LSD
only when the p-value in the ANOVA table is small enough to be considered significant.
Exercises 5.6–5.8 are multiple choice. Choose the answer that best fits.
5.6 Why ANOVA? The purpose of an ANOVA in the setting of this chapter is to learn about:
5.7 Not a condition? Which of the following statements is not a condition of the ANOVA
model?
5.8 ANOVA plots. The two best plots to assess the conditions of an ANOVA model are:
5.9 Which sum of squares? Match the measures of variability with the sum of square terms:
5.10 Student survey. You gather data on the following variables from a sample of 75 under-
graduate students on your campus:
• Major
• Sex
• Class year (first year, second year, third year, fourth year, other)
a. Assume that you have a quantitative response variable and that all of the above are possible
explanatory variables. Classify each variable as quantitative or categorical. For each categor-
ical variable, assume that it is the explanatory variable in the analysis and determine whether
you could use a two-sample t-test or whether you would have to use an ANOVA.
b. State three research questions pertaining to these data for which you could use ANOVA.
(Hint: For example, one such question would be to investigate whether average sleep times
differ for students with different political inclinations.)
5.11 Car ages. Suppose that you want to compare the ages of cars among faculty, students,
administrators, and staff at your college or university. You take a random sample of 200 people
who have a parking permit for your college or university, and then you ask them how old their
primary car is. You ask several friends if it’s appropriate to conduct ANOVA on the data, and you
obtain the following responses. Indicate how you would respond to each.
262 CHAPTER 5. ONE-WAY ANOVA
a. “You can’t use ANOVA on the data because there are four groups to compare.”
b. “You can’t use ANOVA on the data because the response variable is not quantitative.”
c. “You can’t use ANOVA on the data because the sample sizes for the four groups will probably
be different.”
d. “You can do the calculations for ANOVA on the data, even though the sample sizes for the
four groups will probably be different, but you can’t generalize the results to the populations
of all people with a parking permit at your college/university.”
5.12 Comparing fonts. Suppose that an instructor wants to investigate whether the font used
on an exam affects student performance as measured by the final exam score. She uses four different
fonts (times, courier, helvetica, comic sans) and randomly assigns her 40 students to one of those
four fonts.
category variable data
a. Identify the explanatory and response variables in this study. randomlized experiment
c. Even though the subjects in this study were not randomly selected from a population, it’s
still appropriate to conduct an analysis of variance. Explain why.
5.13 Comparing fonts (continued). Refer to Exercise 5.12. Determine each of the entries
that would appear in the “degrees of freedom” column of the ANOVA table.
5.14 All false. Reconsider the previous exercise. Now suppose that the p-value from the ANOVA
F-test turns out to be 0.003. All of the following statements are false. Explain why each one is
false.
a. The probability is 0.003 that the four groups have the same mean score.
b. The data provide very strong evidence that all four fonts produce different mean scores.
c. The data provide very strong evidence that the comic sans font produces a different mean
score than the other fonts.
d. The data provide very little evidence that at least one of these fonts produces a different
mean score than the others.
e. The data do not allow for drawing a cause-and-effect conclusion between font and exam score.
f. Conclusions from this analysis can be generalized to the population of all students at the
instructor’s school.
5.6. EXERCISES 263
5.15 Can this happen? You conduct an ANOVA to compare three groups.
a. Is it possible that all of the residuals in one group are positive? Explain how this could
happen, or why it cannot happen.
b. Is it possible that all but one of the residuals in one group is positive? Explain how this could
happen, or why it cannot happen.
c. If you and I are two subjects in the study, and if we are in the same group, and if your score
is higher than mine, is it possible that your residual is smaller than mine? Explain how this
could happen, or why it cannot happen.
d. If you and I are two subjects in the study, and if we are in different groups, and if your score
is higher than mine, is it possible that your residual is smaller than mine? Explain how this
could happen, or why it cannot happen.
5.16 Schooling and political views. You want to compare years of schooling among American
adults who describe their political viewpoint as liberal, moderate, or conservative. You gather a
random sample of American adults in each of these three categories of political viewpoint.
b. What additional information do you need about these three samples in order to conduct
ANOVA to determine if there is a statistically significant difference among these three means?
c. What additional information do you need in order to assess whether the conditions for ANOVA
are satisfied?
5.17 Schooling and political views (continued). Now suppose that the three sample sizes in
the previous exercise are 25 for each of the three groups, and also suppose that the standard devi-
ations of years of schooling are very similar in the three groups. Assume that all three populations
do, in fact, have the same standard deviation. Suppose that the three sample means turn out to
be 11.6, 12.3, and 13.0 years.
a. Without doing any ANOVA calculations, state a value for the standard deviation that would
lead you to reject H0 . Explain your answer, as if to a peer who has not taken a statistics
course, without resorting to formulas or calculations.
b. Repeat part (a), but state a value for the standard deviation that would lead you to fail to
reject H0 .
264 CHAPTER 5. ONE-WAY ANOVA
a. Suppose you have 10 observations from Group 1 and 10 from Group 2. What are the degrees
of freedom that would be used to calculate a typical pooled two-sample confidence interval
to compare the means of Groups 1 and 2?
b. Suppose you have 10 observations each from Groups 1, 2, and 3. What are the degrees
of freedom that would be used to calculate a confidence interval based on Fisher’s LSD to
compare the means of Groups 1 and 2?
Guided Exercises
5.19 Life spans. The World Almanac and Book of Facts lists notable people of the past in various
occupational categories; it also reports how many years each person lived. Do these sample data
provide evidence that notable people in different occupations have different average lifetimes? To
investigate this question, we recorded the lifetimes for 973 people in various occupation categories.
Consider the following ANOVA output:
Source DF SS MS F P
Occupation 2749 0.000
Error 968 195149 202
Total 972 206147
a. Fill in the three missing values in this ANOVA table. Also show how you calculate them.
b. How many different occupations were considered in this analysis? Explain how you know.
5.20 Meth labs. Nationally, the abuse of methamphetamine has become a concern, not only
because of the effects of drug abuse, but also because of the dangers associated with the labs that
produce them. A stratified random sample of a total of 12 counties in Iowa (stratified by size of
county—small, medium, or large) produced the following ANOVA table relating the number of
methamphetamine labs to the size of the county. Use this table to answer the following questions:6 .
6
Data from the Annie E. Casey Foundation, KIDS COUNT Data Center, [Link]
5.6. EXERCISES 265
Source DF SS MS F
type 37.51
Error
Total 70.60
5.21 Palatability.7 A food company was interested in how texture might affect the palatability
of a particular food. They set up an experiment in which they looked at two different aspects of
the texture of the food: the concentration of a liquid component (low or high) and the coarseness
of the final product (coarse or fine). The experimenters randomly assigned each of 16 groups of 50
people to one of the four treatment combinations. The response variable was a total palatability
score for the group. For this analysis you will focus only on how the coarseness of the product
affected the total palatability score. The data collected resulted in the following ANOVA table.
Use this table to answer the following questions:
Source DF SS MS F
Coarseness
Error 6113
Total 16722
5.22 Child poverty. The same dataset used in Exercise 5.20 also contained information about
the child poverty rate in those same Iowa counties. Below is the ANOVA table relating child poverty
rate to type of county.
7
Data explanation and link can be found at [Link]
266 CHAPTER 5. ONE-WAY ANOVA
Source DF SS MS F P
type 2 0.000291 0.000145 0.22 0.810
Error 9 0.006065 0.000674
Total 11 0.006356
a. Give the hypotheses that are being tested in the ANOVA table both in words and in symbols.
b. Given in Figure 5.13 is the dotplot of the child poverty rates by type of county. Does this
dotplot raise any concerns for you with respect to the use of ANOVA? Explain.
5.23 Palatability (continued). In Exercise 5.21, you analyzed whether the coarseness of the
product had an effect on the total palatability score of a food product. In the same experiment,
the level of concentration of a liquid component was also varied between a low level and a high
level. The following ANOVA table can be used to analyze the effects of the concentration on the
total palatability score:
a. Give the hypotheses that are being tested in the ANOVA table both in words and in symbols.
5.6. EXERCISES 267
b. Given in Figure 5.14 is the dotplot of the palatability score for each of the two concentration
levels. Does this dotplot raise any concerns for you with respect to the ANOVA? Explain.
5.24 Fantasy baseball. A group of friends who participate in a “fantasy baseball” league became
curious about whether some of them take significantly more or less time to make their selections in
the “fantasy draft” through which they select players.8 The table at the end of this exercise reports
the times (in seconds) that each of the eight friends (identified by their initials) took to make their
24 selections in 2008 (the data are also available in the datafile FantasyBaseball):
a. Produce boxplots and calculate descriptive statistics to compare the selection times for each
participant. Comment on what they reveal. Also identify (by initials) which participant took
the longest and which took the shortest time to make their selections.
b. Conduct a one-way ANOVA to assess whether the data provide evidence that averages as far
apart as these would be unlikely to occur by chance alone if there really were no differences
among the participants in terms of their selection times. For now, assume that all condi-
tions are met. Report the ANOVA table, test statistic, and p-value. Also summarize your
conclusion.
c. Use Fisher’s LSD procedure to assess which participants’ average selection times differ signif-
icantly from which others.
8
Data provided by Allan Rossman.
268 CHAPTER 5. ONE-WAY ANOVA
Round DJ AR BK JW TS RL DR MF
1 42 35 49 104 15 40 26 101
2 84 26 65 101 17 143 43 16
3 21 95 115 53 66 103 113 88
4 99 41 66 123 6 144 16 79
5 25 129 53 144 6 162 113 48
6 89 62 80 247 17 55 369 2
7 53 168 32 210 7 37 184 50
8 174 47 161 164 5 36 138 84
9 105 74 25 135 14 118 102 163
10 99 46 60 66 13 112 21 144
11 30 7 25 399 107 17 55 27
12 91 210 69 219 7 65 62 1
13 11 266 34 436 75 27 108 76
14 93 7 21 235 5 53 23 187
15 20 35 26 244 19 120 94 19
16 108 61 13 133 25 13 90 40
17 95 124 9 68 5 35 95 171
18 43 27 9 230 5 52 72 3
19 123 26 13 105 6 41 32 18
20 75 58 50 103 13 38 57 86
21 18 11 10 40 8 88 20 27
22 40 10 119 39 6 51 46 59
23 33 56 20 244 6 38 13 41
24 100 18 27 91 11 23 31 2
5.25 Fantasy baseball (continued).
a. In Exercise 5.24, part (a), you produced boxplots and descriptive statistics to assess whether
an ANOVA model was appropriate for the fantasy baseball selection times of the various
members of the league. Now produce the normal probability plot of the residuals for the
ANOVA model in Exercise 5.24 and comment on the appropriateness of the ANOVA model
for these data.
b. Transform the selection times using the natural log. Repeat your analysis of the data and
report your findings.
5.26 Fantasy baseball (continued). Continuing with your analysis in Exercise 5.25, use Fisher’s
LSD to assess which participants’ average selection times differ significantly from which others.
5.27 Fantasy baseball (continued). Reconsider the data from Exercise 5.24. Now disregard
the participant variable, and focus instead on the round variable. Perform an appropriate ANOVA
analysis of whether the data suggest that some rounds of the draft tend to have significantly longer
selection times than other rounds. Use a transformation if necessary. Write a paragraph or two
5.6. EXERCISES 269
5.28 Fenthion. Fenthion is a pesticide used against the olive fruit fly in olive groves. It is toxic
to humans, so it is important that there be no residue left on the fruit or in olive oil that will be
consumed. One theory was that, if there is residue of the pesticide left in the olive oil, it would
dissipate over time. Chemists set out to test that theory by taking a random sample of small
amounts of olive oil with fenthion residue and measuring the amount of fenthion in the oil at 3
different times over the year—day 0, day 281 and day 365.9
a. Two variables given in the dataset Olives are f enthion and time. Which variable is the
response variable and which variable is the explanatory variable? Explain.
b. Check the conditions necessary for conducting an ANOVA to analyze the amount of fenthion
present in the samples. If the conditions are met, report the results of the analysis.
c. Transform the amount of fenthion using the exponential. Check the conditions necessary for
conducting an ANOVA to analyze the exponential of the amount of fenthion present in the
samples. If the conditions are met, report the results of the analysis.
5.29 Hawks. The dataset on hawks was used in Example 5.9 to analyze the length of the tail
based on species. Other response variables were also measured in Hawks. We now consider the
weight of the hawks as a response variable.
a. Create dotplots to compare the values of weight between the three species. Do you have any
concerns about the use of ANOVA based on the dotplots? Explain.
b. Compute the standard deviation of the weights for each of the three species groups. Do you
have any concerns about the use of ANOVA based on the standard deviations? Explain.
5.30 Blood pressure. A person’s systolic blood pressure can be a signal of serious issues in
their cardiovascular system. Are there differences between average systolic blood pressure based
on smoking habits? The dataset Blood1 has the systolic blood pressure and the smoking status
of 500 randomly chosen adults.10
a. Perform a two-sample t-test, using the assumption of equal variances, to determine if there
is a significant difference in systolic blood pressure between smokers and nonsmokers.
b. Compute an ANOVA table to test for differences in systolic blood pressure between smokers
and nonsmokers. What do you conclude? Explain.
9
Data provided by Rosemary Roberts and discussed in “Persistence of Fenthion Residues in Olive Oil,” Chaido
Lentza-Rizos, Elizabeth J. Avramides, and Rosemary A. Roberts, (January 1994), Pest Management Science, 40(1):
63–69.
10
Data were used as a case study for the 2003 Annual Meeting of the Statistical Society of Canada. See
[Link]
270 CHAPTER 5. ONE-WAY ANOVA
c. Compare your answers to parts (a) and (b). Discuss the similarities and differences between
the two methods.
5.31 North Carolina births. The file NCbirths contains data on a random sample of 1450 birth
records in the state of North Carolina in the year 2001. This sample was selected by John Holcomb,
based on data from the North Carolina State Center for Health and Environmental Statistics. One
question of interest is whether the distribution of birth weights differs among mothers’ racial groups.
For the purposes of this analysis, we will consider four racial groups: white, black, Hispanic, and
other (including Asian, Hawaiian, and Native American). Use the variable MomRace, which gives
the races with descriptive categories. (The variable RaceMom uses only numbers to describe the
races.)
a. Produce graphical displays of the birth weights (in ounces) separated by mothers’ racial group.
Comment on both similarities and differences that you observe in the distributions of birth
weight among the races.
b. Report the sample sizes, sample means, and sample standard deviations of birth weights for
each racial group.
c. Explain why it’s not sufficient to examine the four sample means, note that they all differ,
and conclude that all races do have birth weight distributions that differ from each other.
5.32 North Carolina births (continued). Return to the data discussed in the previous
exercise.
a. Comment on whether the conditions of the ANOVA procedure and F-test are satisfied with
these data.
b. Conduct an ANOVA. Report the ANOVA table, and interpret the results. Do the data
provide strong evidence that mean birth weights differ based on the mothers’ racial group?
Explain.
5.33 North Carolina births (continued). We return to the birth weights of babies in North
Carolina one more time.
a. Apply Fisher’s LSD to investigate which racial groups differ significantly from which others.
Summarize your conclusions, and explain how they follow from your analysis.
b. This is a fairly large sample so even relatively small differences in group means might yield
significant results. Do you think that the differences in mean birth weight among these racial
groups are important in a practical sense?
5.6. EXERCISES 271
5.34 Blood pressure (continued). The dataset used in Exercise 5.30 also measured the size
of people using the variable Overwt. This is a categorical variable that takes on the values 0 =
Normal, 1 = Overweight, and 2 = Obese. Is the mean systolic blood pressure different for these
three groups of people?
a. Why should we not use two-sample t-tests to see what differences there are between the means
of these three groups?
b. Compute an ANOVA table to test for differences in systolic blood pressure between normal,
overweight, and obese people. What do you conclude? Explain.
c. Use Fisher’s LSD to find any differences that exist between these three groups’ mean systolic
blood pressures. Comment on your findings.
5.35 Salary. A researcher wanted to know if the mean salaries of men and women are different.
She chose a stratified random sample of 280 people from the 2000 U.S. Census consisting of men
and women from New York State, Oregon, Arizona, and Iowa. The researcher, not understanding
much about statistics, had Minitab compute an ANOVA table for her. It is shown below:
Source DF SS MS F P
sex 1 8190848743 8190848743 12.45 0.000
Error 278 1.82913E+11 657958980
Total 279 1.91103E+11
Open-Ended Exercises
5.36 Hawks (continued). The dataset on hawks was used in Example 5.9 to analyze the length
of the tail based on species. Other response variables were also measured in the Hawks data.
Analyze the length of the culmen (a measurement of beak length) for the three different species
represented. Report your findings.
5.37 Sea slugs.11 Sea slugs, common on the coast of southern California, live on vaucherian
seaweed. The larvae from these sea slugs need to locate this type of seaweed to survive. A study
was done to try to determine whether chemicals that leach out of the seaweed attract the larvae.
Seawater was collected over a patch of this kind of seaweed at 5-minute intervals as the tide was
coming in and, presumably, mixing with the chemicals. The idea was that as more seawater came
in, the concentration of the chemicals was reduced. Each sample of water was divided into 6 parts.
Larvae were then introduced to this seawater to see what percentage metamorphosed. Is there a
difference in this percentage over the 6 time periods? Open the dataset SeaSlugs, analyze it, and
report your findings.
5.38 Auto pollution.12 In 1973 testimony before the Air and Water Pollution Subcommittee of
the Senate Public Works Committee, John McKinley, president of Texaco, discussed a new filter
that had been developed to reduce pollution. Questions were raised about the effects of this filter
on other measures of vehicle performance. The dataset AutoPollution gives the results of an
experiment on 36 different cars. The cars were randomly assigned to receive either this new filter
or a standard filter and the noise level for each car was measured. Is the new filter better or worse
than the standard? The variable Type takes the value 1 for the standard filter and 2 for the new
filter. Analyze the data and report your findings.
5.39 Auto pollution (continued). The experiment described in Exercise 5.38 actually used
12 cars, each of three different sizes (small = 1, medium = 2, and large = 3). These cars were,
presumably, chosen at random from many cars of these sizes.
a. Is there a difference in noise level among these three sizes of cars? Analyze the data and
report your findings.
b. Regardless of significance (or lack thereof), how must your conclusions for Exercise 5.38 and
this exercise be different and why?
11
Data explanation and link can be found at [Link] and then clicking on “Course
datasets.”
12
Data explanation and link can be found at [Link]
6.6. EXERCISES 311
6.6 Exercises
Conceptual Exercises
Exercises 6.1–6.6. Factors and levels. For each of the following studies, (a) give the response;
(b) the name of the two factors; (c) for each factor, tell whether it is observational or experimental
and the number of levels it has; and (d) tell whether the study is a complete block design.
6.1 Bird calcium. Ten male and 10 female robins were randomly divided into groups of five.
Five birds of each sex were given a hormone in their diet; the other 5 of each sex were given a
control diet. At the end of the study, the researchers measured the calcium concentration in the
plasma of each bird.
6.2 Crabgrass competition. In a study of plant competition,9 two species of crabgrass (Digitaria
sanguinalis (D.s.) and Digitatria ischaemum (D.i.)) were planted together in a cup. In all, there
were twenty cups, each with 20 plants. Four cups held 20 D.s. each; another four held 15 D.s.
and 5 D.i., another four held 10 of each species, still another four held 5 D.s. and 15 D.i., and the
last four held 20 D.i. each. Within each set of four cups, two were chosen at random to receive
nutrients at normal levels; the other two cups in the set received nutrients at low levels. At the
end of the study, the plants in each cup were dried and the total weight recorded.
6.3 Behavior therapy for stuttering. Thirty-five years ago the journal Behavior Research and
Therapy 10 reported a study that compared two mild shock therapies for stuttering. There were
18 subjects, all of them stutterers. Each subject was given a total of three treatment sessions,
with the order randomized separately for each subject. One treatment administered a mild shock
during each moment of stuttering, another gave the shock after each stuttered word, and the third
treatment was a control, with no shock. The response was a score that measured a subject’s
adaptation.
6.4 Noise and ADHD. It is now generally accepted that children with attention deficit and
hyperactivity disorder tend to be particularly distracted by background noise. About 20 years ago
a study was done to test this hypothesis.11 The subjects were all second-graders. Some had been
diagnosed as hyperactive; the other subjects served as a control group. All the children were given
sets of math problems to solve, and the response was their score on the set of problems. All the
children solved problems under two sets of conditions, high noise and low noise. (Results showed
that the controls did better with the higher noise level, whereas the opposite was true for the
hyperactive children.)
9
Katherine Ann Maruk (1975), “The Effects of Nutrient Levels on the Competitive Interaction between Two
Species of Digitaira,” unpublished master’s thesis, Department of Biological Sciences, Mount Holyoke College.
10
D. A. Daly and E. B. Cooper (1967), “Rate of Stuttering Adaptation under Two Electro-shock Conditions,”
Behavior Research and Therapy, 5(1):49–54.
11
S. Zentall and J. Shaw (1980), “Effects of Classroom Noise on Performance and Activity of Second-grade Hyper-
active and Control Children,” Journal of Educational Psychology, 72(6):630–840.
312 CHAPTER 6. MULTIFACTOR ANOVA
6.5 Running dogs. In a study conducted at the University of Florida,12 investigators compared
the effects of three different diets on the speed of racing greyhounds. (The investigators wanted to
test the common belief among owners of racing greyhounds that giving their dogs large doses of
vitamin C will cause them to run faster.) In the University of Florida study, each of 5 greyhounds
got all three diets, one at a time, in an order that was determined using a separate randomization
for each dog. (To the surprise of the scientists, the results showed that when the dogs ate the diet
high in vitamin C, they ran slower, not faster.)
6.6 Fat rats. Is there a magic shot that makes dieting easy? Researchers investigating appetite
control measured the effect of two hormone injections, leptin and insulin, on the amount eaten by
rats.13 Male rats and female rats were randomly assigned to get one hormone shot or the other.
(The results showed that for female rats, leptin lowered the amount eaten, compared to insulin; for
male rats, insulin lowered the amount eaten, compared to leptin.)
6.7 F-ratios. Suppose you fit the two-way main effects ANOVA model for the river data of
Example 6.2, this time using the log concentration of copper (instead of iron) as your response,
and then do an F-test for differences between rivers.
a. If the F-ratio is near 1, what does that tell you about the differences between rivers?
6.8 Degrees of freedom. If you carry out a two-factor ANOVA (main effects model) on a
dataset with Factor A at four levels and Factor B at five levels, with one observation per cell, how
many degrees of freedom will there be for:
a. Factor A?
b. Factor B?
c. interaction?
d. error?
6.9 More degrees of freedom. If you carry out a two-factor ANOVA (with interaction) on a
dataset with Factor A at four levels and Factor B at five levels, with three observations per cell,
how many degrees of freedom will there be for:
a. Factor A?
b. Factor B?
c. interaction?
12
(July 20, 2002),“Antioxidants for Greyhounds? Not a Good Bet” Science News, 162(2):46.
13
(July 20, 2002),“Gender Differences in Weight Loss,” Science News, Vol. 162(2):46.
6.6. EXERCISES 313
d. error?
6.11 Fill in the blank. If your dataset has two factors and you carry out a one-way ANOVA, ig-
noring the second factor, your SSE will be too (small, large) and you will be
(more, less) likely to detect real differences than would a two-way ANOVA.
6.12 Fill in the blank, again. If you have two-way data with one observation per cell, the only
model you can fit is a main effects model and there is no way to tell whether interaction is present.
If, in fact, there is interaction present, your SSE will be too (small, large) and you will
be (more, less) likely to detect real differences due to each of the factors.
6.13 Interaction. Is interaction present in the following data? How can you tell?
Heart Soul
Democrats 2, 3 10, 12
Republicans 8, 4 11, 8
6.14 Interaction, again. Is interaction present in the following data? How can you tell?
Exercises 6.15–6.18 are True or False exercises. If the statement is false, explain why it is false.
6.15 If interaction is present, it is not possible to describe the effects of a factor using just one
set of estimated main effects.
6.16 In a randomized complete block study, at least one of the factors must be experimental.
6.17 The conditions for the errors for the two-way additive ANOVA model are the same as the
conditions for the errors of the one-way ANOVA model.
314 CHAPTER 6. MULTIFACTOR ANOVA
6.18 The conditions for the errors of the two-way ANOVA model with interaction are the same
as the conditions for the errors of the one-way ANOVA model.
Guided Exercises
6.19 Burning calories. If you really work at it, how long does it take to burn 200 calories on an
exercise machine? Does it matter whether you use a treadmill or a rowing machine? An article14 in
Medicine and Science in Sports and Exercise reported average times to burn 200 calories for men
and for women using a treadmill and a rowing machine for heavy exercise.
The results:
6.20 Drunken teens, part 1. A survey was done to find the percentage of 15-year-olds, in each
of 18 European countries, who reported having been drunk at least twice in their lives. Here are
the results, for boys and girls, by region. (Each number is an average for 6 countries.)
Male Female
Eastern 24.17 42.33
Northern 51.00 51.00
Continental 24.33 33.17
Draw an interaction plot, and discuss the pattern. Relate the pattern to the context. (Don’t just
say “The lines are parallel, so there is no interaction,” or “The lines are not parallel, so interaction
is present.”)
6.21 Happy face: interaction. Researchers at Temple University15 wanted to know the
following: If you work waiting tables and you draw a happy face on the back of your customers’
14
Steven Swanson and Graham Caldwell (2001) “An Integrated Biomechanical Analysis of High Speed Incline and
Level Treadmill Running,” Medicine and Science in Sports and Exercise, 32(6):1146–1155.
15
B. Rind and P. Bordia (1996), “Effect on Restaurant Tipping of Male and Female Servers Drawing a Happy Face
on the Backs of Customers’ Checks,” Journal of Social Psychology, 26:215–225.
6.6. EXERCISES 315
checks, will you get better tips? To study this burning question at the frontier of science, they
enlisted the cooperation of two servers at a Philadelphia restaurant. One was male, the other
female. Each server recorded his or her tips for their next 50 tables. For 25 of the 50, following
a predetermined randomization, they drew a happy face on the back of the check. The other 25
randomly chosen checks got no happy face. The response was the tip, expressed as a percentage of
the total bill. The averages for the male server were 18% with a happy face, 21% with none. For
the female server, the averages were 33% with a happy face, 28% with none.
a. Regard the dataset as a two-way ANOVA, which is the way it was analyzed in the article.
Name the two factors of interest, tell whether each is observational or experimental, and
identify the number of levels.
b. Draw an interaction graph. Is there evidence of interaction? Describe the pattern in words,
using the fact that an interaction, if present, is a difference of differences.
6.22 Happy face: ANOVA (continued). A partial ANOVA table is given below. Fill in the
missing numbers.
Source df SS MS F
Face (Yes/No)
Gender (M/F) 2,500
Interaction 400
Residuals 100
Total 25,415
6.23 River iron. This is an exercise with a moral: Sometimes, the way to tell that your model is
wrong requires you to ask, “Do the numbers make sense in the context of the problem?” Consider
the New York river data of Example 6.2, with iron concentrations in the original scale of parts per
million:
Grasse Oswegatchie Raquette St. Regis Mean
Upstream 944 860 108 751 665.75
Midstream 525 229 36 568 339.50
Downstream 327 130 30 350 209.25
Mean 598.7 406.3 58.0 556.3 404.83
b. Obtain a normal probability plot of residuals. Is there any indication from this plot that the
normality condition is violated?
c. Obtain a plot of residuals versus fitted values. Is there any indication from the shape of
this plot that the variation is not constant? Are there pronounced clusters? Is there an
unmistakable curvature to the plot?
316 CHAPTER 6. MULTIFACTOR ANOVA
d. Finally, look at the leftmost point, and estimate the fitted value from the graph. Explain
why this one fitted value strongly suggests that the model is not appropriate.
6.24 Iron deficiency. In developing countries, roughly one-fourth of all men and half of all
women and children suffer from anemia due to iron deficiency. Researchers16 wanted to know
whether the trend away from traditional iron pots in favor of lighter, cheaper aluminum could be
involved in this most common form of malnutrition. They compared the iron content of 12 samples
of three Ethiopian dishes: one beef, one chicken, and one vegetable casserole. Four samples of each
dish were cooked in aluminum pots, four in clay pots, and four in iron pots. Given below is a
parallel dotplot of the data.
m = meat
8
p = poultry
v = vegetable
6
Iron content of food
m
m
m
m p
4
p
p
v
v
m
m p p
p
m m v
m p
m
2
m v
v
v v
m v
v
v
0
Describe what you consider to be the main patterns in the plot. Cover the usual features keeping in
mind that in any given plot, some features deserve more attention than others: How are the group
averages related (to each other and to the researchers’ question)? Are there gross outliers? Are
the spreads roughly equal? If not, is there evidence that a change of scale would tend to equalize
spreads?
6.25 Alfalfa sprouts. Some students were interested in how an acidic environment might affect
the growth of plants. They planted alfalfa seeds in 15 cups and randomly chose five to get plain
water, five to get a moderate amount of acid (1.5M HCl), and five to get a stronger acid solution
(3.0M HCl). The plants were grown in an indoor room so the students assumed that the distance
from the main source of daylight (a window) might have an affect on growth rates. For this reason,
they arranged the cups in five rows of three, with one cup from each Acid level in each row. These
are labeled in the dataset as Row: a = farthest from the window through e = nearest to the
16
A. A. Adish, et al. (1999), “Effect of Food Cooked in Iron Pots on Iron Status and Growth of Young Children:
A Randomized Trial,” The Lancet 353:712–716.
6.6. EXERCISES 317
window. Each cup was an experimental unit and the response variable was the average height of
the alfalfa sprouts in each cup after four days (Ht4). The data are shown in the table below and
stored in the Alfalfa file:
Treatment/Cup a b c d e
water 1.45 2.79 1.93 2.33 4.85
1.5 HCl 1.00 0.70 1.37 2.80 1.46
3.0 HCl 1.03 1.22 0.45 1.65 1.07
a. Find the means for each row of cups (a, b, ..., e) and each treatment (water, 1.5HCl, 3.0HCl).
Also find the average and standard deviation for the growth in all 15 cups.
b. Construct a two-way main effects ANOVA table for testing for differences in average growth
due to the acid treatments using the rows as a blocking variable.
c. Check the conditions required for the ANOVA model.
d. Based on the ANOVA, would you conclude that there is a significant difference in average
growth due to the treatments? Explain why or why not.
e. Based on the ANOVA, would you conclude that there is a significant difference in average
growth due to the distance from the window? Explain why or why not.
6.26 Alfalfa sprouts (continued). Refer to the data and two-way ANOVA on alfalfa growth
in Exercise 6.25. If either factor is significant, use Fisher’s LSD (at a 5% level) to investigate which
levels are different.
6.27 Unpopped popcorn. Lara and Lisa don’t like to find unpopped kernels when they make
microwave popcorn. Does the brand make a difference? They conducted an experiment to compare
Orville Redenbacher’s Light Butter Flavor versus Seaway microwave popcorn. They made 12
batches of popcorn, 6 of each type, cooking each batch for 4 minutes. They noted that the microwave
oven seemed to get warmer as they went along so they kept track of six trials and randomly chose
which brand would go first for each trial. For a response variable, they counted the number of
unpopped kernels and then adjusted the count for Seaway for having more ounces per bag of
popcorn (3.5 vs. 3.0). The data are shown in Table 6.8 and stored in Popcorn.
Brand/Trial 1 2 3 4 5 6
Orville Redenbacher 26 35 18 14 8 6
Seaway 47 47 14 34 21 37
a. Find the mean number of unpopped kernels for the entire sample and estimate the effects (α1
and α2 ) for each brand of popcorn.
318 CHAPTER 6. MULTIFACTOR ANOVA
b. Run a two-way ANOVA model for this randomized block design. (Remember to check the
required conditions.)
c. Does the brand of popcorn appear to make a difference in the mean number of unpopped
kernels? What about the trial?
6.28 Swahili attitudes.17 Hamisi Babusa, a Kenyan scholar, administered a survey to 480
students from Pwani and Nairobi provinces about their attitudes toward the Swahili language. In
addition, the students took an exam on Swahili. From each province, the students were from 6
schools (3 girls’ schools and 3 boys’ schools), with 40 students sampled at each school, so half of the
students from each province were males and the other half females. The survey instrument contained
40 statements about attitudes toward Swahili and students rated their level of agreement on each.
Of these questions, 30 were positive questions and the remaining 10 were negative questions. On
an individual question, the most positive response would be assigned a value of 5, while the most
negative response would be assigned a value of 1. By summing (adding) the responses to each
question, we can find an overall Attitude Score for each student. The highest possible score would
be 200 (an individual who gave the most positive possible response to every question). The lowest
possible score would be 40 (an individual who gave the most negative response to every question).
The data are stored in Swahili.
a. Investigate these data using P rovince (Nairobi or Pwani) and Sex to see if attitudes toward
Swahili are related to either factor or an interaction between them. For any effects that are
significant, give an interpretation that explains the direction of the effect(s) in the context of
this data situation.
b. Do the normality and equal variance conditions look reasonable for the model you chose in
(a)? Produce a graph (or graphs) and summary statistics to justify your answers.
c. The Swahili data also contain a variable coding the school for each student. There are 12
schools in all (labeled A, B, ... , L). Despite the fact that we have an equal sample size
from each school, explain why an analysis using School and P rovince as factors in a two-way
ANOVA would not be a balanced complete factorial design.
Open-Ended Exercises
6.29 Mental health and the moon. For centuries, people looked at the full moon with some
trepidation. From stories of werewolves coming out, to more crime sprees, the full moon has gotten
the blame. Some researchers18 in the early 1970s set out to actually study whether there is a
“full-moon” effect on the mental health of people. The researchers collected admissions data for
17
Thanks to Hamisi Babusa, visiting scholar at St. Lawrence University for the data.
18
The original discussion of the study appears in S. Blackman and D. Catalina (1973), “The Moon and the
Emergency Room,” Perceptual and Motor Skills 37:624–626. The data can also be found in Richard J. Larsen and
Morris L. Marx (1986), Introduction to Mathematical Statistics and Its Applications, Prentice-Hall:Englewood Cliffs,
NJ.
6.6. EXERCISES 319
the emergency room at a mental health hospital for 12 months. They separated the data into rates
before the full moon (mean number of patients seen 4–13 days before the full moon), during the full
moon (the number of patients seen on the full moon day), and after the full moon (mean number of
patients seen 4–13 days after the full moon). They also kept track of which month the data came
from since there was likely to be a relationship between admissions and the season of the year.
The data can be found in the file MentalHealth. Analyze the data to answer the researcher’s
question.
Supplementary Exercises
When group standard deviations are very unequal (Smax /Smin is large), you can sometimes find a
suitable transformation as follows:
Step 1 : Compute the average and standard deviation for each group.
Step 4 : Compute p = 1 − slope. p tells us the transformation. For example, p = 0.5 means the
square root, p = 0.25 means the fourth root, and p = −1 means the reciprocal. (For technical
reasons, p = 0 means the logarithm.)
6.30 Simple illustration. This exercise was invented to show the method described above at
work using simple numbers. Consider a dataset with four groups and three observations per group:
Notice that you can think of each set of observed values as m − s, m, and m + s. Find the ratio
s/m for each group and notice that the “errors” are constant in percentage terms. For such data,
a transformation to logarithms will equalize the standard deviations, so applying Steps 1–4 should
show that transforming is needed, and that the right transformation is p = 0.
a. Compute the means and standard deviations for the four groups. (Don’t use a calculator.
Use a short cut instead: Check that if the response values are m − s, m, and m + s, then the
mean is m and the standard deviation is s.)
320 CHAPTER 6. MULTIFACTOR ANOVA
b. Compute Smax /Smin . Is a transformation called for? Plot log10 (s) versus log10 (m), and fit a
line by eye. (Note that the fit is perfect: The right transformation will make Smax /Smin = 1
in the new scale.)
d. Use a calculator to transform the data and compute new group means and standard devia-
tions.
e. Check Smax /Smin . Has changing scales made the standard deviations more nearly equal?
6.31 Sugar metabolism. The plot below shows a scatterplot of the log(s) versus log(ave) for
the data from this chapter’s case study:
3.0
2.5
2.0
log(SDs)
1.5
1.0
0.5
0.0
0 1 2 3
log(Aves)
a. Fit a line by eye to all eight points, estimate the slope, and compute p = 1 − slope. What
transformation is suggested?
b. If you ignore the two outliers, the remaining six points lie very close to a line (see the figure
below). Estimate its slope and compute p = 1 − slope. What transformation is suggested?
3.0
2.5
2.0
log(SDs)[1:6]
1.5
1.0
0.5
0.0
0 1 2 3
log(Aves[1:6])
6.6. EXERCISES 321
6.32 Diamonds. Here are the means, standard deviations, and their logs for the diamond data
of Example 5.7.
a. Plot log(s) versus log(ave) for the four groups. Do the points suggest a line?
For many two-way datasets, the additive (no interaction) model does not fit when the response is
in the original scale. For some of these datasets, however, there is a transformed scale for which
the additive model does fit well. When such a transformation exists, you can find it as follows:
Step 1 : Fit the additive model to the cell means, and write the observed values as a sum:
6.33 Sugar metabolism. Figure 6.16 shows the plot of residuals versus comparison values for the
data from this chapter’s case study. Fit a line by eye and estimate its slope. What transformation
is suggested?
6.34 River iron. Here are the river iron data in the original scale:
4
Residuals from Additive Model
2
0
−2
−4 −5 0 5
Comparison Values
Figure 6.16: Graph of the residuals versus the comparison values for the sugar metabolism study
A decomposition of the observed values gives a grand average of 404.83; site effects of 260.92,
−65.33 and −195.58; and river effects of 193.83, 1.50, −346.83, and 151.50. The residuals are
c. Fit a line by eye and estimate its slope. What transformation is suggested?