Understanding Variables and Data Analysis
Understanding Variables and Data Analysis
experience suggests that many students misunderstand these terms. Teaching tip:
Wikipedia ([Link]) has a good summary and some links for more
1
information on Kristen Gilbert.
Solutions
••• In-Class Activities
Activity 1-1: Student Data
a. Answers will vary.
b. Answers will vary.
c. No, every word did not contain the same number of letters.
d. The variable is the number of letters in each word.
e. • How many hours you slept in the past 24 hours: quantitative
• Whether you slept for at least 7 hours in the past 24 hours: binary categorical
• How many states you have visited: quantitative
• Handedness: binary categorical, unless you classify “ambidextrous,” in which
case it is not binary.
• Day of the week on which you were born: categorical
• Gender: binary categorical
• Average study time per week: quantitative
• Score on the first exam in this course: quantitative
f. No, neither average height of students in the class nor percentage of students in the
class who have used a cell phone today can legitimately be considered variables when
the observational units are the students in your class. Both of these are numbers
that provide summary information about the class as a whole. They do not vary
from student to student.
g. If you record the average student height or percentage of student cell phone usage
by class taught at your school, these would become legitimate variables. Now these
numbers would (potentially) take on different values from class to class. The
observational units are no longer the students in your class, but rather all classes
taught at your school.
6 Topic 1: Data and Variables
1
b. One variable is whether Gilbert worked on the shift. This variable is categorical and
binary. The other variable is whether a patient died on the shift. This variable is
also categorical and binary.
Solutions
••• In-Class Activities
Activity 2-1: Penny Thoughts
2
a. This is a binary categorical variable.
b. Answers will vary by class. One example is 15/18 or .833 voted to retain the
penny.
c. Answers will vary, but for the class in part b, 3/18 or .167 voted to abolish the
penny.
d. Answers will vary, but for the class in part b, the bar graph is shown here:
.5
.4
.3
.2
.1
0
Retain Abolish
Response
e. For the class in part c, more than 80% favors retaining the penny, but the vote is
not unanimous. A significant proportion of the class (.167) favors abolishing the
penny.
0 5 10 15 20 25 30 35 40 45 50
Number of States Visited
2
c. Students in the regular section tended to score more points than those in the
sports section. Scores in the regular section appear to be centered around 340
(85% of the possible points), whereas those in the sports section are centered
around 310–320 points (a bit less than 80% of the possible points). Scores in the
sports section are more spread out than those in the regular section. Students
in the sports section had the six lowest scores, all less than 260 points, but that
section also had the highest overall score, greater than 390 points.
d. Students in the regular section tended to score more points than those in
the sports section. Most students in the regular section scored between
300–380 points, with a center of approximately 340 points. In contrast, many
students in the sports section scored less than 300 points, and the center was
approximately 310–320 points.
e. No, some sports students scored more points than some regular students. The
statistical tendency means that a typical student in the regular section scored
more points than a typical student in the sports section.
f. The proportions are found by dividing the counts by 29 for the regular section
and by 28 for the sports section. These proportions are
Regular section: .552 good .379 fair .069 poor
Sports section: .250 good .536 fair .214 poor
g. The bar graphs follow:
Good 1
Fair .9
Proportion of Students
.8
Poor
.7
.6
.5
.4
.3
.2
.1
0
Regular Sports
Section
h. The bar graphs reveal similar results to the dotplots: Students in the regular
section tended to score higher than those in the sports section. More than half of
the regular students were in the good category, compared to only one-fourth of
the students in the sports section. At the other extreme, only 6.9% of the regular
students did poor work, compared to 21.4% of sports students.
i. You cannot draw a cause-and-effect conclusion between the type of section and
student performance. You will study these issues again in the next topic, but one
key is that students self-selected which section to take. Perhaps those who chose to
take the sports section had lower academic aptitude than those who selected the
20 Topic 2: Data and Distributions
regular section, or perhaps students were sleepier in the sports section because it
met earlier in the day.
Rating 1 2 3 4 5 6 7 8 9
Tally (Count) 0 0 1 0 5 6 11 6 6
You might point out that asking for the explanatory and response variables puts this
question one additional step beyond what students have already begun to practice
extensively. Parts b and c then deal with both of the big issues in this topic: whether a
cause-and-effect conclusion can be drawn and to what population the results can be
generalized. You might summarize that these are the two questions students should
learn to ask about every statistical study. These are very different issues, due to two
problems (confounding and bias, respectively), and therefore require different remedies.
Those remedies are the subjects of the next two topics.
Solutions
••• In-Class Activities
Activity 3-1: Elvis Presley and Alf Landon
a. Population: all adult Americans (record company interest)
Sample: those listening to the radio who called in
b. No, 56% is probably not an accurate reflection of the opinions of all adult
Americans on this issue. People who chose to call in (who took the time and
were willing to spend the money) probably felt differently and more strongly
about the issue than other adult Americans. The timing (on the anniversary of
Elvis’ death) could also have influenced the opinions of those who called. You
also have no indication of how widely distributed across the country the radio
stations were (perhaps there could be bias if the stations tended to be mostly
from the south).
c. Population: all Americans eligible to vote in 1936
Sample: the 2.4 million who returned the questionnaires
d. The Literary Digest’s prediction was in error because its sampling method was
biased. By sampling people who owned vehicles and telephones in 1936, Literary
Digest was sampling from a subset of the population who during that period
tended to be wealthy. Historically, the wealthy have tended to support the
Republican candidate (conservative), whereas those without money have tended
to vote Democrat (for social change). Thus, the pollsters contacted primarily
Activity 3-2 31
Republican voters, but on election day, there was a heavy Democratic turn out.
Furthermore, those who chose to respond were probably more dissatisfied with
the incumbent (Roosevelt) than those who chose not to respond.
e. • The 56% of callers who believed that Elvis was alive: statistic
• The 57% of voters who indicated they would vote for Alf Landon: statistic
• The 63% of votes who actually voted for Franklin Roosevelt: parameter
f. • The proportion of students in your class who use instant-messaging or text-messaging
on a daily basis: statistic
• The proportion of students at your school who use instant-messaging or text-
3
messaging on a daily basis: parameter
• The average number of hours students at your school spent watching television last
week: parameter
• The average number of hours students in your class slept last night: statistic
g. • The proportion of voters who voted for President Bush in the 2004 election:
parameter
• The proportion of voters surveyed by CNN who voted for John Kerry in the 2004
election: statistic
• The proportion of voters among your school’s faculty members who voted for Ralph
Nader in the 2004 election: parameter (assuming the population is all of your
school’s voting faculty members)
• The average number of points scored in a Super Bowl game: parameter (assuming
the population is all Super Bowl games)
h. A categorical variable leads to a parameter or statistic that is a proportion; a
quantitative variable leads to a parameter or statistic that is an average.
3
less sleep might differ in some other way that could account for the increased
rate of obesity. For example, amount of exercise could be a confounding variable.
Perhaps children who exercise less have more trouble sleeping, in which case
exercise would be confounded with sleep. You have no way of knowing whether
the higher rate of obesity is due to less sleep or less exercise, or both, or due to
some other variable that is also related to both sleep and obesity.
c. The population from which these children were selected is apparently all children
aged 5–10 in primary schools in the city of Trois-Rivières. These Quebec children
might not be representative of all children in this age group worldwide, so you
should be cautious about generalizing that a relationship between sleep and
obesity exists for children around the world.
Solutions
••• In-Class Activities
Activity 4-1: Sampling Words
a. Answers will vary. The answers given here are one example.
b. Here are completed tables:
Number of Letters 5 5 7 4 5
Number of Letters 3 4 4 7 6
c. Dotplot:
3 4 5 6 7
Number of Letters per Word
j. If you use this method, you would also be likely to select too many long words in
your sample because the long words take up more space on the page and therefore
have a greater chance of being selected when you blindly point to a location.
k. No, increasing the sample size will not make up for the biased sampling method.
You would still tend to overrepresent the long words.
l. You need to employ a truly random method to select the words. You could
write each word on the same size slip of paper, put each slip in a hat, mix them
thoroughly, and then draw ten slips from the hat.
4
1 2 3 4 5
Word Length 3 4 3 1 1
d. The distribution is much closer to being centered at 4.29 and has a smaller horizontal
spread than the previous one did (though the latter is not always the case).
e. The sample averages are roughly split evenly on both sides of 4.29.
f. Yes, random sampling appears to have produced unbiased estimates of the average
word length in the population.
1 2 3 4 5
Number of Letters 3 5 4 3 6
b. Both distributions are roughly bell-shaped, centered at about 4.29 words, with
a horizontal spread from about 2 to 7 words.
c. Yes, these distributions seem to have similar variability.
d. No, not much changed when you sampled from the larger population.
participate might differ systematically in some ways from those who were included.
Nevertheless, the researchers did use randomness to select their sample, and they
probably obtained as representative a sample as reasonably possible.
d. Perhaps mothers in those groups were in a lower economic class and therefore
less likely to have phones in the first place, or perhaps they had to work so their
children were in daycare.
e. These comparisons address the issue of bias, not precision. The sampling method
was slightly biased with regard to the mother’s race and age and the infant’s birth
weight.
f. These percentages are statistics because they are based on the sample.
g. The large sample size produces high precision. This means that the sample
statistics are likely to be close to their population counterparts. For example, the
4
population proportion of infants who sleep on their backs should be close to the
sample proportion who sleep on their backs.
h. The sample size for subgroups is smaller than for the whole group, so the sample
results would be less precise.
Solutions
5
••• In-Class Activities
Activity 5-1: Testing Strength Shoes
a. No, this is anecdotal evidence based on just two of your friends. The friend who
wears the strength shoe may be much more athletic in general than your other
friend, or may be taller, etc., which would explain why he or she is able to jump
farther.
b. Explanatory: whether the individual Type: binary categorical
wears strength shoes
Response: length of jump Type: quantitative
c. No, you cannot legitimately conclude that strength shoes cause longer jumps
because the subjects self-selected which type of shoe they would wear, so your
results are from an observational study. There could be something else different
about people who choose to wear the strength shoes. For example, males might be
more likely to choose the strength shoe than females.
d. Randomly assign six of the subjects to each group.
e. You could flip a coin for each subject. If it lands heads up, that subject will wear
ordinary shoes; otherwise, the subject will wear strength shoes. Continue to flip
the coin until you have six subjects in the ordinary (or strength) shoe group. If
you have not filled both groups, the remaining subjects should all be placed in
the unfilled group. Technically, this is not random assignment, even though
this approach is used quite frequently. A more correct approach would be to
number each subject and put the numbers in a bag, and then the first six numbers
drawn out of the hat are assigned to the “strength shoe” group and the others are
assigned to the “ordinary shoe” group.
62 Topic 5: Designing Experiments
b. No, you probably will not get the exact same assignment of subjects to the groups
nor the same difference in proportions.
c. Answers will vary as this is a prediction.
d. The distribution should be centered at zero. It may or may not be what the
students predicted.
e. Random assignment does not always balance out the gender variable exactly, but it
does tend to balance out gender between the two groups. You can see this because
the dotplot is centered at zero, and zero is the difference that occurred most
frequently. Occasionally, there were differences as large or small as .8333.
f. This distribution is roughly symmetric, centered near zero, ranging from roughly
5.5 to 5.5. This indicates that the randomization also tends to balance out the
heights between the two groups.
g. Answers will vary; the question asks for student expectation.
h. Both dotplots (for gene and x-variable) are roughly symmetric and centered at
zero. This indicates that random assignment tends to balance out variables even
when they are unseen or unrecorded.
5
i. By randomly assigning the subjects to the two groups, you have (hopefully) balanced
out all other potentially confounding variables, making the only difference between
the two groups the type of shoe. Therefore, if you then find that the strength shoe
group jumps substantially farther, on average, than the ordinary shoe group, you
would be able to conclude that the increase was due to the strength shoe because
there should not be any other explanation for a difference between the two groups.
c. The instructor randomly decided which grouping of letters each student would
receive. This was important because it prevented self-selection and controlled for
confounding variables. You should not expect any differences between the groups
prior to the treatment.
d. The students were blind to the fact that there were two different groupings of
letters given out initially. They were unaware that you were trying to compare
the effect of these familiar and unfamiliar groupings, so they could not
unintentionally influence the results.
e. Answers will vary. Here is one representative set of answers.
JFK
4 8 12 16 20 24 28
Group
JFKC
4 8 12 16 20 24 28
Number of Letters Correct
f. Yes, these data appear to support the conjecture that those who receive the letters
in convenient three-letter chunks tend to memorize more letters. The center of
this plot is about six letters higher than for the JFCK plot.
g. Yes, because this was a well-designed, randomized controlled experiment you
could legitimately conclude that the grouping of letters into familiar chunks
caused the higher scores. Because you randomly assigned the students to each type
of grouping, there should have been roughly an equal number of good memorizers
in both the JFK and JFKC groups, so the randomization controlled for this
potentially confounding variable.
Solutions
••• In-Class Activities
Activity 6-1: Government Spending
a. Too little: 398/646 .616 About right: 198/646 .307
Too much: 50/646 .077
b. Here is the bar graph for the marginal distribution of the variable opinion about
federal spending on the environment:
Environment Spending
1
.9
.8
.7
Proportion
.6
.5
.4
.3
.2
.1
0
Too Little About Right Too Much
6
Respondents’ Opinions
c. Assuming the sample is representative, you can say that Americans tend to believe
that the government is spending too little on the environment. More than 60%
of the sample felt this way, whereas less than 8% felt that the government is
spending too much. About 30% of this sample felt that government spending on
the environment is right where it should be.
d. The proportion of liberal respondents who say the federal government spends too
little on the environment is 127/155 or .819.
e. The proportion of liberal respondents who say the federal government spends the
right amount on the environment is 27/155 or .174.
f. The proportion of liberal respondents who say the federal government spends too
much on the environment is 1/155 or .006.
g. Yes, these three proportions add up to 1.00 if you don’t use the rounded
proportions.
h. Here is the conditional distribution for the “liberal” column of the table:
Environment Spending
100
Too much
90
About right
80
Too little
70
Percentage
60
50
40
30
20
10
0
Liberal Moderate Conservative
Political Viewpoint
j. Yes, the distribution of spending opinion seems to differ among the three political
groups. The more liberal the group, the more likely they are to believe that the
government is not spending enough on the environment. Whereas 82% of the
liberals in the sample feel the government is spending too little money, 62% of the
moderates and only 48% of the conservatives believe this; on the other hand, 14%
of conservatives believe the government is spending too much on the environment
compared to less than 1% (.6%) of the liberals and 7% of the moderates.
l. Approximately 35% of the liberals in the sample believe the government spends
too much on the space program.
m. There is a difference—but not a big difference—in how the three political groups
feel about government spending on the space program. Roughly 40% of each
group stated the government is spending too much, roughly 10% stated the
government is spending too little, and roughly 50% stated the government is
spending about the right amount.
n. Government spending is closest to being independent of political viewpoint with
the issue of the space program. You can tell because the breakdown of the three
segments in the segmented bar graph is nearly identical across the three political
parties, whereas they are quite distinct across the three political parties in the
environmental spending graph.
o. Answers will vary. Here is an example graph for an issue where opinion about
government spending is perfectly independent of political viewpoint:
60
50
40
30
20
10
0
Liberal Moderate Conservative
Political Viewpoint
Activity 6-3 89
The goal is for the three segments to have the same breakdown across the
three groups (though not necessarily being an equal size within each political
party).
p. The proportion who identify themselves as liberal is 127/398 or .319.
q. No, the proportion in part p is less than half the value found in part d (.819).
HIV-infected 40 13 53
6
d. For the placebo group, 40/183 or .219. For the AZT group, 13/180 or .072. Yes,
these calculations are consistent with the segmented bar graph.
e. The difference is .219 – .072 or .147 (or .146 if using more than three decimal
places for the proportions). This difference does not appear to be terribly
large.
f. The ratio is .219/.072 or 3.04 (or 3.03 if using more than three decimal places for
the proportions). The risk of a baby being born HIV-infected is more than three
times greater for those whose mothers were in the placebo group.
g. Yes, you can legitimately conclude that AZT is the cause of the three-fold
reduction in HIV-infection rate compared to the placebo group because this
was a well-designed, randomized experiment. These results can be cautiously
generalized to HIV-positive pregnant women (you don’t have much information
about how the women in the sample were selected).
c. Answers will vary by class. Here is one representative set of answers from a class of
78 students:
.6
.5
.4
.3
.2
.1
0
Olympic Medal Nobel Prize Academy Award
Preferred Achievement
More than 45% of this class would prefer to win a Nobel prize, whereas less than
20% aspire to win an Academy Award. Just over one-third of the class would like
their greatest lifetime achievement to be winning an Olympic medal.
d. Using the example data from part c gives these counts:
Male Female
Olympic Medal 11 16
Nobel Prize 12 24
Academy Award 1 14
e. With this information you would only be able to fill in the totals; you need
to know the information on both gender and preferred lifetime achievement
simultaneously in order to fill in each cell of the table.
6
No regular position 100
Back 90
Side 80
70
Stomach
Percentage
60
50
40
30
20
10
0
1992 1993 1994 1995 1996
Year
e. The segmented bar graph reveals that the distribution of infant sleep positions
changed considerably over this five-year period. The percentage of parents who
placed infants on their stomachs declined dramatically, from more than two-
thirds (70%) to less than one-fourth (24%) in these five years. The proportions
of infants placed on their backs and sides both increased during this period.
Because this is an observational study and not an experiment, you cannot say
that the recommendation or promotional campaign caused these changes, but
(assuming the sample is representative) it is nonetheless heartening to find
that parents were generally changing their habits and placing infants in safer
positions to sleep.
118 Topic 7: Displaying and Describing Distributions
4. Set the viewing window for the histogram. To do this, press the p key and
enter numbers so that the plot fits the screen properly, making sure you use the Ã
key to express the negative sign, not the subtraction key.
5. Press the s key, followed by the r key, and then the right and left
arrows to obtain information about your histogram. To reduce or increase the
number of subintervals, change the Xscl setting in the Window screen.
Solutions
••• In-Class Activities
Activity 7-1: Matching Game
Consider the following seven variables:
a. These are all quantitative variables (except perhaps “jersey number;” even though
it is numerical, it does not make sense to determine the “average jersey number”).
b. 1. Variable: point values of letters in the board game Scrabble
Explanation: The values are regularly spaced (point values are whole numbers)
with most letters being worth few points.
2. Variable: prices of properties on the Monopoly game board
Explanation: The values are very regularly spaced with two properties having
exactly the same prices.
3. Variable: annual snowfall amounts for a sample of cities around the United States
Explanation: Most cities have little or no snow, but some have quite a bit.
4. Variable: jersey numbers of Cal Poly football players in 2006
Explanation: There are virtually no repeats and the distribution covers nearly
all the values from 1–99.
5. Variable: blood pressure measurements for a sample of healthy adults
Explanation: The values are fairly symmetric, with both a high and low outlier.
6. Variable: weights of rowers on the 2004 U.S. men’s Olympic team
Explanation: There is one low outlier (coxswain) and two weight classes.
7. Variable: quiz percentages for a class of statistics students (quizzes were quite
straightforward for most students)
Explanation: Most of the values are high (near 80–90) with a few very low
outliers.
c. Dotplots 1 and 3 have a similar shape (they are skewed to the right), whereas
dotplots 6 and 7 are both skewed to the left with outliers.
b. Cipollone is the apparent outlier at 120 lbs. His job (the coxswain) is to call out
the cadence to the rowers—he does not row himself. So it is important that he not
add excess weight to the boat.
c. There is a cluster of rowers whose weights seem to be no more than 160 lbs. These
rowers are all involved in “lightweight” events, which require them to weigh
below a certain amount on race day. The upper cluster are not involved in the
lightweight events and have no upper limit on their weights.
7
h. One-quarter of the monarchs reigned for fewer than 9.5 years. One-quarter of the
monarchs reigned for more than 34 years.
9998888550 0 033344556889999
43322 1 00014567
982100 2 136
0 3
0 4
5
6 6
leaf unit one percentage point
b. Western states tend to have higher percentage growths than eastern states. The
distributions of percentage growth for both the eastern and western states are
both skewed to the right. The percentages in the 26 eastern states display less
variability than the western states, ranging from near 0% to about 26%, with
a center of about 10%. The percentages in the 24 western states range from a
minimum near 0% to a maximum 30%, with a high outlier at 40% and a center
of about 13%. There is also one western state with an extremely high population
growth rate (Nevada)—more than 66%.
120 Topic 7: Displaying and Describing Distributions
300
Number of People
250
200
150
100
50
0
0 20 40 60 80
Age of Diabetes Diagnosis
140
120
Number of People
100
80
60
40
20
0
0 20 40 60 80
Age of Diabetes Diagnosis
50
40
Number of People
30
20
10
0
0 12 24 36 48 60 72 84
Age of Diabetes Diagnosis
Activity 7-7 121
g. The first graph appears to have a uniform shape; the second is slightly skewed to
the right. But neither is a legitimate histogram of the Olympic rowers’ weights
because the variable (weight) is not on the horizontal axis—it is on the vertical
axis! For example, it would be very misleading to say the “center” of the weights
was around Dan Berry’s weight—he actually has one of the greatest weight values.
The rowers’ names are of less interest than the pattern in the weight values to the
behavior of the weights.
0 68
1 00000555555588
2 00000000122555556668 leaf unit .1 mile
3 000002224458
4 00000566
5 00055668
6 0000
7 004
8
9 5
b. The distribution of hike distances is sharply skewed to the right, indicating there
are many hikes on the short side and only a few longer hikes. A typical hike is
between 2 and 3 miles. Most hikes are between 1 and 6 miles, but two hikes are
less than a mile and a few are more than 6 miles. The longest hike is 9.5 miles,
which is a bit unusual and could be considered an outlier because this hike is
7
more than two miles longer than the next longest hike (7.4 miles). Many hikes
have a reported distance that is a multiple of a whole number or a half number
of miles.
18
16
14
12
Frequency
10
8
6
4
2
0
0 20 40 60 80 100
Age at Death (in years)
Activity 8-1 139
SetUpEditor. Press Ö and then press Õ and ¿. You can now edit any of
the entries using the arrow keys to highlight a value. To delete an entry, press the
{ key.
4. Now use 1-VarStats to recompute the mean and median.
The SetUpEditor feature can set up more than one list within the editor (up to 20 in
the order you specify). On your Home screen, after SetUpEditor, enter one or more
lists separated by commas. One nice feature is that if a list is archived, SetUpEditor will
automatically unarchive the list and place it in the editor.
Other useful features in the STAT menu include EDIT, SORTA(, SORTD(, and
CLRLIST. The SORTA( and SORTD( features are definitely worth exploring with your
students. These two features will sort a given list either in ascending or descending
order, respectively. The most useful capability of these two features is when you have
two or more dependent lists, for example, a data list and its frequency list. If you sort
the data list in either ascending or descending order, you can couple the frequency
list with the data list so you do not break the dependency. So, for example, you would
enter SORTA(datalist,frequencylist) on your Home screen if you wanted to sort your
data list in ascending order. This dependency coupling can be done with any number
of lists.
Solutions
••• In-Class Activities
Activity 8-1: Sleeping Times
a. Observational units: students in a statistics course
Explanatory variable: section of the course Type: categorical
Response variable: sleeping time in hours Type: quantitative
b. The centers are not all similar: The center for Section 1 is noticeably less than the
8
center for Section 2, which is less than the center for Section 3.
c. The center for Section 1 is about 6 hours. The peak of the distribution is about
6 hours. Approximately half of the times are less than this, and approximately
half are more.
d. The mean is 106.25/17 or 6.25 hours.
e. The median is 6.25 hours. You calculate (171)/2 9. The ninth observation,
counting from either end, is 6.25 hours.
f. Section 2 mean: 7.000 Section 3 mean: 7.523
Explanation: The mean for Section 2 should be less because there are many
observations at 8 hours in Section 3, but not in Section 2, and the classes behave
similarly at 6 hours and less.
g. Section 2 median: 7.0 Section 3 median: 7.5
h. No, based on this example, the mean and median of a dataset do not always equal
each other.
i. The mode of the section variable is Section 3. More students were enrolled in this
section than in any other section. Perhaps that is because it was offered much later
in the morning!
140 Topic 8: Measures of Center
c. When the distribution is skewed to the left, the mean is less than the median.
When the distribution is skewed to the right, the mean is greater than the
median. In symmetric distributions, the mean and median are similar.
8
variables that explain the strong association here.
customers when there are more than two choices. It could be the case that
100 customers chose chocolate, 90 chose vanilla, and 80 chose strawberry.
The mode is chocolate, but 170/270 .63, so 63% prefer a flavor other than
chocolate.
students have successfully rid themselves of common misconceptions; you might also
consider using this activity as a group quiz. Activity 9-22 is similar to 8-22 in that it
asks students to create examples of hypothetical data that satisfy various properties.
Note that Activity 9-23 introduces two additional terms: midrange and midhinge.
Possible Assignments:
Day 1: Activity 9-7 on standard deviations; Activity 9-18 on IQRs; Activity 9-19
if you did the memory experiment here or in Topic 5; and Activity 9-21, which asks
students to reflect on information from a study. If you did part a of Activity 9-25 in
class, then assign Activity 9-25 part b for homework.
Day 2: Activity 9-8 on the IQR, SD, and the empirical rule; Activity 9-11 on
z-scores; and Activity 9-14 on the empirical rule.
Technology Ideas
Activity 9-22: Hypothetical Exam Scores (Fathom)
In this activity, students need to create hypothetical datasets of ten exam scores to
obtain certain properties. A combination of Fathom dynamic graphs and summary
tables can help ease the computational burdens as students test ideas.
Part a of this activity asks for a dataset where less than half the exam scores fall
within one standard deviation of the mean. To explore this property, drag an empty
case table from the shelf and create a new attribute called exam with ten hypothetical
scores as an initial guess. Create a dotplot of the exam scores and choose Graph | Plot
Value to show the values of mean( ) – s( ) and mean( ) + s( ), thus
clearly showing where “within one standard deviation of the mean” lies.
Drag the data points around to try to get at least six outside the boundaries. The
mean and standard deviation, and thus the boundaries, are automatically recomputed
as you move the points. To keep track of those values, create a summary table, drag the
exam attribute to show the mean, and then choose Summary | Add Formula to also
show s( ). Once a reasonable solution is found, the student may need to edit the data
to make all the exam scores integers.
For other parts of this activity, show Q1( ) and Q3( ) on the plot, or display
IQR( ) in the summary table. For a sneaky challenge, ask students to try to get all ten
exam scores more than one standard deviation away from the mean—it’s impossible!
Solutions
•••
9
In-Class Activities
Activity 9-1: Baseball Lineups
a. Observational units: baseball players
Explanatory variable: team Type: binary categorical
Response variable: age Type: quantitative
b. Here are the comparative dotplots:
Yankees
Team
22 23 24 25 26 27 28 29 30 31 32 33 34 35
Tigers
22 23 24 25 26 27 28 29 30 31 32 33 34 35
Age (in years)
156 Topic 9: Measures of Spread
The average age of both teams seems to be about the same (about 30 years), but
the spreads are quite different. The 2006 Tigers are much closer together in age
than the 2006 Yankees.
c. Yankees Mean: 29.7 years Median: 31.5 years
Tigers Mean: 30 years Median: 29.5 years
The centers of these distributions are relatively similar.
d. No; although the centers are the same, the spreads are very different, with the
Yankees having the youngest and oldest players in the two distributions and not
much consistency in the ages of their players.
e. The Yankees’ lineup appears to have greater variability in its ages.
f. Oldest: 35 years Youngest: 22 years Difference: 13 years
g. Oldest: 34 years Youngest: 25 years Difference: 9 years
h. Lower quartile: 28 years Upper quartile: 32 years IQR: 4 years
i. The Yankees have the greater age range and greater IQR. These values are
consistent with the answer to part e.
j. The average age of the starting lineups on both teams is about 30 years, but the
Tigers’ ages are fairly tightly clustered from 28–34 years, with the exception of
one player (Granderson) who is only 25 years old. In comparison, the Yankees’
ages range from a low of 22 years to a high of 35 years and also have a larger
interquartile range. It is difficult to judge the shape with these small sample sizes,
but the distribution of the ages for the Tigers appears more symmetric, whereas
the distribution of ages for the Yankees is more skewed to the left.
I. Rodriguez 34 34 30 4 4 16
Casey 32 2 2 4
Perez 33 3 3 9
Inge 29 1 1 1
Guillen 30 0 0 0
Monroe 29 1 1 1
Granderson 25 5 5 25
Gomez 28 2 2 4
Young 32 2 2 4
Robertson 28 2 2 4
Total 300 0 22 68
Activity 9-3 157
b. This sum in the “deviation from mean” column is zero, which makes sense
because the positive deviations from the mean “cancel out” the negative deviations
from the mean.
c. See the table in part a. The sum of the absolute deviations is 22 years.
d. The mean of the absolute deviation is 22/10 or 2.2 years.
e. See the table in part a. The sum of the squared deviations is 68 years2.
f. 68/9 7.56 years2
g. 2.749 years
h. The standard deviation for the Tiger lineup’s ages is 2.749 years. The standard
deviation for the Yankees lineup’s ages is 4.62 years.
As expected, the Yankees’ standard deviation is larger because the ages tend to be
located farther from the average age.
i. Answers will vary by student expectation. This change will definitely affect the
range and the standard deviation, but it should have little or no effect on the IQR
because you are changing only an extreme value (endpoint).
j. See table in part k.
k. Here is the completed table:
l. See table in part k. These results demonstrate that the IQR is resistant, but the
range or standard deviation is not. You know this because the value of the IQR
does not change when the value of the outlier changes, whereas the range and
standard deviation are affected dramatically.
9
Activity 9-3: Value of Statistics
a. Answers will vary by student prediction. Many students will pick class F, focusing
incorrectly on the irregularity in the heights of the bars.
b. Answers will vary by student prediction. Many students will incorrectly predict
that class J has the most variability because more of the possible data values appear
in the histogram. Students again may incorrectly see class H as having more
variability because they are looking at the differences in the heights of the bars.
c. Here is the completed table:
Range 6 8 8 8 8
d. According to these measures of spread, class G has more variability than class F,
because class G has more data values farther from the mean.
e. According to these measures of spread, class I has the most variability and class H
has the least variability. Class I has more of its data values at the extremes (and
far from the center), whereas most of class H’s observations are close to the mean.
Class J is in between the two.
f. Class F has more bumpiness in its histogram, but has less variability than class G.
g. Class J has the greatest number of distinct values but does not have the most
variability among classes H, I, and J.
h. No, based on the two previous questions, variability does not measure either
bumpiness or variety; variability measures spread from the center (mean). A
distribution can be very “bumpy” without having a great deal of variability, and
vice versa. It is more important to consider the overall tendency for data values to
be far from the center.
i. Many answers are possible, but all ten values need to be the same so the standard
deviation is zero.
j. Only one answer is possible: {1, 1, 1, 1, 1, 9, 9, 9, 9, 9}. This dataset maximizes
the distances of observations from the mean and has a standard deviation of 4.22.
Any other combination will have a smaller standard deviation. (Note: If you did
not balance the 1s and 9s, the mean would shift away from 5 and would put the
more frequent values closer to the mean.)
9
d. The ordered difference in couple’s ages are
7, 5, 5, 2, 1, 1, 0, 0, 1, 1, 1, 1, 1, 2, 2, 3, 3, 3, 3, 5, 7, 8, 10, 15
The median is the average of the 12th and 13th ordered values: (1 1)/2 or 1 year.
The mean is the sum of these differences divided by 24, which turns out to be
45/24 or 1.9 years.
Notice that the mean of the age differences is equal to the difference in mean ages
between husbands and wives: 1.9 35.7 33.8. But this property does not quite
hold for the median.
e. The quartiles are 0.5 and 3, so the IQR is 3.5 years. The standard deviation of these
age differences is 4.8 years. The IQR of the differences and the standard deviations of
the differences calculated here are less than the individual IQRs (19.5 and 17.5) and
the individual standard deviations (14.56 and 13.56) calculated in part b.
160 Topic 9: Measures of Spread
f. To be within one standard deviation of the mean is to be within 1.9 4.8 years,
which means between –2.9 and 6.7 years. Seventeen of the age differences fall
within this interval, which is a proportion of 17/24 or .708, or 70.8%. This
percentage is quite close to 68%, which is what the empirical rule predicts.
Because the distribution of the age differences does look fairly symmetric and
mound-shaped, this outcome is not surprising.
g. The mean and median indicate that, on average, people marry someone within
a couple years of their own age. More importantly, the measures of spread are
fairly small for the differences, much smaller than for individual ages. This result
suggests that there is not much variability in the differences, which suggests that
people do tend to marry people of similar ages.
h. The differences have less variability because even though people get married from
their teens to seventies (and beyond), they tend to marry people within a few years
of their own age.
1 1134579
2 1345568
3 56778
4 00248
5 012468
6 59
7 6
10 5
11
12 4
b. The median is 37 people; the upper quartile is 52; the lower quartile is 23; and
the interquartile range is 29 people.
c. The mean is 40.83 people and the standard deviation is 25.25 people.
d. You calculate 40.83 25.25 [15.58, 66.08], so 26/35 or 0.743 of the students’
results fall within one standard deviation of the mean.
172 Topic 10: More Summary Measures and Graphs
Solutions
••• In-Class Activities
Activity 10-1: Natural Selection
a. Observational units: adult male sparrows
Explanatory variable: total length Type: quantitative
Response variable: whether the sparrow survived Type: binary categorical
b. This is an observational study because Bumpus simply recorded this information
about the sparrows; he did not impose any treatment on them.
c. Median: 159 mm
Lower quartile: 158 mm Upper quartile: 160 mm
Minimum: 153mm Maximum: 166 mm
Activity 10-1 173
d. Median: 159 mm
Lower quartile: 158 mm Upper quartile: 160 mm
Minimum: 153 mm Maximum: 166 mm
e. The sparrows that died tended to be longer than the sparrows that survived.
Seventy-five percent of those that died were at least 161 mm long, but 75% of those
that survived were shorter than 160 mm. The typical length for the sparrows that
died was 162 mm, and the typical length was only 159 mm for those that survived.
f. The following boxplots display the distribution of lengths for the sparrows that
survived and died:
Died
Status
Survived
Outliers would be outside the interval [155, 163], so there are three low outliers
(153, 154, and 154) and two high outliers (165 and 166).
h. The following modified boxplots display the lengths for the surviving sparrows
and for the sparrows that died:
Died *
Status
Survived * ** * *
k. Yes, it is clear from this study that shorter sparrows were more likely to survive
the storm. This makes no claim about why they were more likely to survive the
storm—only that the shorter sparrows tended to survive more frequently than did
the longer sparrows.
b. The following boxplots display the distribution of calorie amounts for the three
brands:
Dreyer’s
Dreyer’s ice cream appears to have significantly fewer calories than the other two
brands, as well as less variability among its calorie amounts. All of the Dreyer’s
ice cream flavors have fewer than 200 calories, whereas only 25% of the Ben &
Jerry’s flavors have 220 or fewer calories. There is a great deal of variability in
the number of calories for the Ben & Jerry flavors as they range from 110 to
360 calories, whereas (excluding outliers) the Cold Stone Creamery flavors range
from 360 to 440 calories.
c. The serving sizes may not be the same for all three brands. This would make it
difficult to compare the calories as given.
d. You could convert the Cold Stone Creamery serving from 170 grams to the
comparable measure of volume in 1/2 cups.
e. Divide each of the Cold Stone Creamery listings by 170 grams, then multiply by
73 grams per 1/2 cup.
f. You calculate new value (old value)/170 * 73.
g. Here is the five-number summary for Cold Stone’s calorie amounts using the “per
half cup” scale: min 55.82, Q L 154.59, median 167.47, Q U 171.76,
max 188.94.
h. The following boxplots display the three distributions of calorie amounts:
Dreyer’s
Once you adjust the ice cream servings so that they all have the same serving size,
Cold Stone Creamery still has several flavors that are low outliers (meaning that
these flavors have an unusually small amount of calories per serving). Excluding
the outliers, the Dreyer’s flavors are generally the lowest in calorie content,
followed by the Cold Stone Creamery flavors, which have a very narrow spread
(from only about 155–189 calories per 1/2 cup); at least 75% of the Ben & Jerry
flavors have more calories than either of the other two brands.
The fan cost index ranges from a low of about $120 to a high of about $220,
with an outlier of Boston at $287.84. The average FCI is $171, and the median
is slightly lower at $166. The standard deviation is $35.05, whereas the IQR
is $47.76.
c. Answers will vary.
d. Answers will vary.
e. Answers will vary.
f. Highest team: NY Mets Value: $4.75
Lowest team: Baltimore, Milwaukee, and Kansas Value: $2.00
g. The number of ounces in a “small” soda or beer is not the same in every park.
h. Highest team: LA Dodgers Value: $0.354/oz
Lowest team: Pittsburgh Value: $0.1125/oz
Advanced Compact
Camera Type
Compact
Subcompact * **
Super-Zoom
There is considerable overlap in prices among these four groups, but some camera
types (e.g. advanced compact) do tend to cost more than other types.
b. Advanced compact cameras tend to cost the most, followed by super-zoom
cameras. Compact and subcompact camera prices are similar, but the subcompact
cameras cost a bit more on average than the compact cameras.
c. Advanced compact cameras have the most spread in terms of prices, again
followed by super-zoom cameras. But compact cameras have more spread in prices
than do subcompact cameras, which have the least variability in prices.
Activity 10-6 177
d. The boxplots of camera ratings are shown in the following the table. The five-
number summaries, as reported by the software package Minitab, are as shown here:
Advanced Compact 63 69 70 73 78
Compact 62 65 71 73.5 76
Advanced Compact
Camera Type
Compact
Subcompact
Super-Zoom
50 55 60 65 70 75 80
Rating Score
The super-zoom cameras tend to have the highest ratings, and the subcompact
ones tend to rate the lowest. The subcompact cameras also have the most
variability in ratings.
e. Even though the advanced compact cameras have the highest median price by far,
their median rating is surpassed by both super-zoom and compact cameras. In
fact, compact cameras have the second-highest median rating despite having the
second-lowest median price.
11
Technology Ideas
Activity 11-18: Dice-Generated Ice Cream Prices (Fathom)
This activity asks students to use technology to simulate (1000 times) the prices for an ice
cream special that arise from rolling two dice and following the larger result with the smaller
result to create a two-digit price. Here’s one way to accomplish this simulation using Fathom:
1. Drag a new blank collection from the shelf, choose Collection | New Cases from
the menu, and enter 1000 to create that many cases (rows) in the collection.
2. Drag an empty table from the shelf and create two new attributes, Die1 and
Die2, to contain the random dice rolls for each price. Select the Die1 column and
choose Edit | Edit Formula to enter the formula randomInteger(1,6) to
simulate the roll of a die. Repeat for the Die2 column.
3. Create a third attribute called Price to contain the simulated prices. Use the formula
editor again and enter a formula that begins with if(Die1>Die2). When the
if( ) structure is entered in a Fathom formula, the program creates two question
marks. Replace the first (top) question mark with the computation when the
condition is true (10*Die1+Die2), and replace the second (bottom) question
mark with the computation when the condition is false (10*Die2+Die1).
4. Use the values that are computed for the Price attribute to create a dotplot of the
1000 simulated prices and compute the mean.
Solutions
••• In-Class Activities
Activity 11-1: Random Babies
Answers will vary. Here is one representative set of answers:
a. One mother received her own baby.
b. Here is a record of the random “dealing”:
Number of Repetitions 1 2 3 4 5
Number of Matches 1 1 2 1 2
204 Topic 11: Probability
1 5 5 5 1
2 2 10 7 .7
3 4 15 11 .7333
4 5 20 16 .8
5 4 25 20 .8
6 4 30 24 .8
7 4 35 28 .8
8 0 40 28 .7
9 3 45 31 .6889
10 3 50 34 .6800
11 2 55 36 .6545
12 1 60 37 .6167
13 3 65 40 .6154
14 3 70 43 .6143
15 4 75 47 .6267
16 4 80 51 .6375
17 4 85 55 .6471
18 3 90 58 .6444
19 5 95 63 .6632
20 2 100 65 .6500
d. The following graph displays the cumulative proportion of trials with at least one
match vs. the cumulative number of trials.
1.2
1.0
Trials with at Least One Match
Cumulative Proportion of
0.8
0.6
0.4
0.2
0
5 15 25 35 45 55 65 75 85 95
Cumulative Number of Trials
Activity 11-1 205
11
e. The proportion of trials that result in at least one mother getting the correct baby
fluctuates more at the beginning of this process.
f. Yes, the relative frequency appears to be settling down and approaching one
particular value. Answers will vary as to what that value is, but it should be in the
ballpark of .627.
g. Here is a completed table with counts and proportions recorded:
Count 34 33 28 0 5 100
l. Yes, these simulation results are reasonably consistent with the class results.
m. Pr(at least one match) 1 .359 .641.
n. Yes, this graph appears to be fluctuating less as more trials are performed,
approaching a limiting value:
p. It is not impossible to get four matches, but it seems very rare/unlikely because it
appears to have happened only about 39 out of 1000 times.
q. No, results of zero, one, or two matches do not seem to be unlikely. Each would
appear to occur more than 25% of the time; a zero match and one match even
occur more than one-third of the time each.
11
c. Answers will vary. The following were obtained using row 31 of the table.
Random Digit 3 4 2 8 2 5 0 7
Gender B G G G G B G B
Number of Girls 1 2 1 1
Tally (Count) 4 5 11
e. No, it does not appear that the probability of each of these outcomes is 1/3. It
appears that the probability of having one child of each gender is about twice that
of having two girls or two boys.
f. You can obtain better empirical estimates of these probabilities if you simulate
more families.
g. Here is a completed table of simulation results:
h. No, it does not appear that the probability of each of these outcomes is 1/3. It
appears that the probability of having one girl and one boy each is about .5,
whereas the probability of having two girls or of two boys is about .25.
i. Two girls: Pr(GG) 1/4 .25 Two boys: Pr(BB) 1/4 .25
One of each: Pr(BG or GB) 1/2 .5
Yes, these probabilities are reasonably close to the empirical estimates from the
class simulations.
f. A sample of size 12 is more likely to contain at least one-third senior citizens. This
result makes sense because only 20% of the population is senior citizens, so it is
easier to select 4 of 12 jurors who are senior citizens than it is to select 25 of 75.
There should be more sampling variability (easier to get an unlucky result) with
the smaller sample size.
g. Based on the simulations, you see that a sample size of 75 is more likely to contain
between 15% and 25% (11.25 and 21) senior citizens than a sample size of 25.
This result makes sense because with the larger sample size, the sample proportion
should be close to the population proportion more often.
h. The previous questions show that a larger sample size is more likely to produce a
sample proportion that is close to the population proportion, and smaller sample sizes
are more likely to produce samples with greater variability from sample to sample.
11
k. This expected value makes sense because with three women randomly assigned
among two groups, you would expect half of them to be assigned to each group
in the long run, and half of 3 is 1.5.
Solutions
••• In-Class Activities
Activity 12-1: Body Temperatures and Jury Selection
a. Both distributions are roughly symmetric and mound-shaped.
b. The following sketch approximates the general shape in the two histograms:
c. The dashed curve (C) has a mean of 50 and a standard deviation of 10. The dashed
curve (B) has a mean of 70 and a standard deviation of 10. The solid curve (A) has a
mean of 70 and a standard deviation of 5. Here is the labeled drawing:
10 20 30 40 50 60 70 80 90 100 110
Exam Score
d. Yes, these predictions are quite close to each other and to what the empirical rule
would predict.
e. The z-score for the value 97.5 in the distribution of body temperatures is z
(97.5 98.249)/.733 or 1.02.
f. The z-score for the value 11.5 in the distribution of number of senior citizens in
the jury pool is z (11.5 14.921)/3.336 or 1.025.
g. The two z-scores are almost identical.
h. Both z-scores indicate that the observations are just over one standard deviation
below their respective means.
Activity 12-2 227
12
a. The following graph shows the shaded region whose area corresponds to the
probability that a baby will have a low birth weight:
1000 1500 2000 2500 3000 3500 4000 4500 5000 5500
Birth Weight (in grams)
99.9 99.9
99 99
90 90
Percentage
Percentage
50 50
10 10
1 1
0.1 0.1
50 100 150 200 30 60 90 120
Systolic Blood Pressure Diastolic Blood Pressure
99.9
99
90
Percentage
50
10
1
0.1
30 60 90 120
Pulse Rate
These probability plots confirm that the systolic pressures were unlikely to have
come from a normal distribution, and the diastolic pressures could quite possibly
have come from a normal distribution.
Females
Males
10 15 20 25 30 35 40
Footprint Length (in cm)
Activity 12-4 229
b. Because you are dealing with a male footprint, the z-score is (22 – 25)/4 or –0.75.
Using Table II (or technology), the probability of the footprint being less than
22 centimeters is .2266. So, roughly 22.66% of men have a footprint smaller than
22 centimeters and would be misclassified as female.
12
Females
Males
10 15 20 25 30 35 40
z= -0.75
Footprint Length ( in cm)
c. Now you are dealing with a female footprint, so the z-score is (22 – 19)/3 or 1.00.
Using Table II (or technology), the probability of the footprint being longer than
22 centimeters is 1 – .8413 or .1587. This indicates that about 16% of females
would be mistakenly identified as male.
Females
Males
10 15 20 25 30 35 40
z = 1.00
Footprint Length (in cm)
Females
Males
.08
10 15 20 25 30 35 40
z = -1.41
Footprint Length (in cm)
e. Using this new cutoff value of 19.36 centimeters, the z-score for a female footprint
is (19.36 – 19)/3 or 0.12. The probability of a female footprint being longer than
19.36 centimeters is 1 – .5478 or .4522 (using Table II or technology).
Females
Males
.4522
10 15 20 25 30 35 40
z = 0.12
Footprint Length (in cm)
13
to practice distinguishing between the population, sample, and sampling distributions,
perhaps asking them to sketch all three distributions. Activities 13-5 and 13-6 provide
practice with the terms parameter and statistic. Note that Activities 13-9, 13-10, 13-11,
and 13-13 ask students to perform simulations. The remaining homework activities ask
students to work with the CLT more mathematically.
Possible assignment: Activities 13-5 (parameter vs. statistic; it is not necessary to assign
all 15 parts), 13-9 (requires a simulation), and 13-12 or 13-16 (thinking about plausibility)
Technology Ideas
Activity 13-9: Pet Ownership (Fathom)
Part d of this activity asks students to use technology to simulate the random selection
of 500 samples, each of size 200 from a population of households where 25% have cats.
Although this can be done within the Reece’s Pieces applet, students need to ignore the
candy context. To accomplish this simulation using Fathom, follow these steps:
1. Create a new collection and choose Collection | New Cases to specify 500 cases
for the samples.
2. Label the first attribute in this collection as Cats and choose Edit | Edit Formula
to enter the formula randomBinomial(200,0.25) to create the simulated
counts of cats in each sample.
3. Create a second attribute Phat to contain the proportion for each sample, and in
the formula editor, enter Cats/200 as the sample proportion.
4. Counting how many sample proportions are within .05 of the population propor-
tion can be a bit tricky because attempting to sort the column will cause Fathom
to recalculate the random values. Instead, double-click on the collection to show its
inspector, and on the Measures panel, define a measure named Close and enter the for-
mula 500-Count(Phat.20)-Count(Phat.30) to compute the count.
These simulation steps can also be applied to Activities 13-10, 13-11, and 13-13.
Solutions
••• In-Class Activities
Activity 13-1: Candy Colors
Answers will vary. Here is one representative set of answers.
a. The count and proportion of each color in this sample are recorded in the table:
Count 13 7 5
h. The observational units in this graph are the samples of 25 candies. The variable
being measured from unit to unit is the proportion of the sample that is colored
orange.
i. This dotplot of the sample proportions is symmetric and mound-shaped
(roughly normal), with a center of .5, and with most of the sample
proportions falling between .4 and .6 (min .36, max .68). The standard
deviation is .08.
j. Based on the sample results from this class, a reasonable guess for would be .5.
k. Most estimates would be reasonably close to , but a very few estimates
would be way off. You can see this from the dotplot. Most of the class results
are the same (near .5), but a few of the class results are quite extreme (far
from .5).
l. If each student had taken samples of size 10 instead of size 25, you would expect
more variability (greater horizontal spread) in the dotplot.
m. If each student had taken samples of size 75 instead of size 25, you would expect
less variability (less horizontal spread) in the dotplot.
13
.1 .2 .3 .4 .5 .6 .7 .8
∏
Mean = .448 SD = .099 ^p
d. Yes, the distribution appears roughly normal, centered at about .45, with a
standard deviation of about .1.
e. A normal curve seems to model the simulated sample proportions very well.
f. Mean of p̂ values: .449 Standard deviation of p̂ values: .100
g. Roughly speaking, more sample proportions are close to .45 than are far away
from it.
h. Here is the completed table:
q. The applet reports that 336/500 or 95% of the sample proportions are within
.114 of .45.
r. About 95% of the students’ intervals would contain the actual population
proportion of .45.
s. Theoretical mean of p̂ values: .45
________
(.45)(.55)
Theoretical standard deviation of p̂ values: ________ .0995 .1
25
t. Theoretical mean of p̂ values: .45
________
(.45)(.55)
Theoretical standard deviation of p̂ values: ________ .057
75
u. No, the normal model does not summarize this distribution well. This is not a
contradiction to the Central Limit Theorem because n 25(.1) 2.5 10.
(.5)(.5)
Spread: ______ .0449
124
c. Yes, the histogram does appear to be consistent with what the CLT predicts. It is
bell-shaped, centered at about .5, and extends from about .5 3(.0449) or .3653
to about .5 3(.0449) or .6347.
d. p̂ 80/124 .645
e. Yes, it would be very surprising to observe such a sample proportion (.645) if 1/2
of all kissing couples lean their heads to the right; this sample proportion never
occurred in 1000 simulations.
f. The z-score for the observed sample proportion is z (.645 .5)/.0449 3.23.
g. Yes, this is a very surprising z-score; Pr(Z 2.33) .0099. If 1/2 of all kissing
couples lean their heads to the right, you would see a sample result as or more
extreme than .645 in less than 1% of random samples.
(.667)(.333)
__________ .042
124
Activity 13-5 253
13
a surprising result at all. Therefore, the sample data provide no reason to doubt
that the population proportion of kissing couples who lean their heads to the
right equals 2/3.
d. The value .645 is pretty far along the lower tail of the second histogram.
This indicates that the observed sample proportion would rarely occur if the
population proportion were equal to 3/4. Further evidence of this result is
provided by the rather large negative z-score:
.645
z ____________
__________ .645 .750 2.69
.750 __________
.39
(.750)(.250)
__________
124
Therefore, the sample data provide fairly strong evidence that the population
proportion of kissing couples who lean their heads to the right is not 3/4 (because
it would be rather surprising to find a sample proportion so far from this
population proportion by chance alone).
e. A reasonable estimate of the population proportion is the sample proportion
.645. An estimate of the standard deviation of p̂ would then be
__________
(.645)(.355)
__________ .043
124
Doubling this standard deviation gives .086. The interval is, therefore, .645
.086, which runs from .559 to .731. Notice that 1/2 and 3/4 are not in this
interval, but 2/3 is. The interval is consistent with the earlier analysis of the
plausibility of the values 1/2, 2/3, and 3/4 for the population proportion of kissing
couples who lean their heads to the right.
4. To calculate the mean of this sample, choose the Measures panel in the inspector
for the sample collection and define a new measure called XBar with the formula
mean(Change _ Amount).
5. To save the sample means for repeated samples, select the sample collection box
and choose Collection | Collect Measures from the top menu. This creates
a third collection with, by default, the sample means (XBar) for five different
samples. Drag a new table from the shelf to view the contents of this collection.
6. Double-click on the new collection to show its inspector, and on the Collect
Measures panel, turn off animation by unchecking the Animation on box, check
Replace existing cases, and change the number of measures to collect from 5 to
500. Press the Collect More Measures button to perform the 500 simulations.
14
7. Use the Xbar attribute in the Measure from Sample of Change collection to
create a dotplot, normal quantile plot, and summary statistics to describe the
distribution of the sample means.
Solutions
••• In-Class Activities
Activity 14-1: Coin Ages
a. Observational units: pennies Variable: age
Quantitative or categorical? quantitative
b. These values are parameters, represented by symbols (mean) and (standard
deviation).
c. No, the distribution of ages does not follow a normal distribution; it is strongly
skewed to the right.
d. Answers will vary. One example (using row 48 of the Random Digits Table, three
digits at a time) is coins numbered 788, 929, 977, 718, 049, which have ages 20,
28, 35, 17, and 2, respectively. Note that we would delete any three-digit numbers
that correspond to a coin that has already been selected (sampling without
replacement). You are also free to read the table vertically, backward, selecting a
new row, etc. In order for someone to evaluate your procedure though, include
enough details so your method is clear.
0 5 10 15 20 25 30 35 40 45 50 55 60
Penny Age
f. The following table gives the sample mean age for each sample:
1 2 3 4 5
_
Sample Mean (x ) 20.4 18.8 1.2 18 10.6
g. No, you do not get the same value for the sample mean all five times. This reveals
the phenomenon of sampling variability. This variable is quantitative, whereas the
variable in the Candy Colors activity (Activities 13-1 and 13-2) was binary categorical.
__ __
h. Mean of x values: 13.8 years SD of x values: 7.99 years
__
i. The mean of x values is reasonably close to the population mean ( 12.264 years).
The standard deviation is less than the population standard deviation of 9.613 years.
j. The following dotplot displays the sample means:
0 4 8 12 16 20 24 28
Mean = 13.05 SD = 9.95
Activity 14-2 271
Yes, this dotplot does resemble the distribution of ages in the population; it is
definitely skewed to the right. Note: Any observations to the right of the applet
window are piled in the last bin.
b. The mean is 13.05 years and the standard deviation is 9.95 years. The shape
is somewhat skewed to the right.
c. A sketch of the results displayed in the dotplot follow (500 sample means
with n 5):
14
0 4 8 12 16 20 24 28
SD = 4.21 Mean = 12.32 x–
Now the shape of the distribution is much more symmetric and less skewed. The
center is 12.32 years and the standard deviation is 4.2 years.
d. The mean of the sampling distribution is very close to the population mean, but
the standard deviation of the sampling distribution is much smaller.
e. A sketch of the results displayed in the dotplot follow (500 sample means
with n 25):
0 4 8 12 16 20 24 28
Mean=12.27 SD=1.94
The shape is even more symmetric, and the spread is much smaller. (The center
has remained the same.)
f. The mean of the sampling distribution is very close to the population mean, but
the standard deviation of the sampling distribution is much smaller.
272 Topic 14: Sampling Distributions: Means
g. A sketch of the results displayed in the dotplot follow (500 sample means
with n 50):
0 4 8 12 16 20 24 28
Mean=12.21 SD=1.29
The shape is even more symmetric, and the spread is even smaller, but the center
has remained at about 12.2 years.
h. Once again the mean of the sampling distribution is very close to the population
mean, but the standard deviation of the sampling distribution is much smaller.
Sample Size Mean of Sample Means SD of Sample Means Shape of Sample Means
80
70
60
Frequency
50
40
30
20
10
0
0.00 0.14 0.28 0.42 0.56 0.70 0.84 0.98
Sample Mean
k. Yes, this probability plot indicates the distribution is definitely well modeled by a
normal curve:
99.9
99
95
90
80
70
Percentage 60
50
40
30
14
20
10
5
0.1
0.3 0.4 0.5 0.6 0.7
Sample Mean
__ ___
l. You calculate /n 9.613/50 1.359 years. This value is reasonably close
to the standard deviation of the 500 simulated sample means (1.29).
__ ___
m. You calculate /n 28.866/30 4.08¢. This is reasonably close to the
simulated standard deviation of 5¢.
f. No, a sample mean of $857 would not be at all surprising as it is located close to
the center of the distribution and there is a great deal of area to the right of this
value in the plot shown in part e.
_
g. The CLT says that that sampling distribution of x would be approximately
____
normal with mean $800 and standard deviation $250/ 922 $8.23. Because
the sample size is large, your answer does not depend on the shape of the
distribution of expected expenditures in the population. Here is a sketch for this
distribution:
This time a sample mean of $857 would be very surprising because this value does
not appear on the normal curve modeling the behavior of the sample means.
____
h. The standard deviation should be $250/ 922 $8.23.
Activity 14-4 275
14
larger than 30), you do not need to know whether the population distribution
of heights is normal because the Central Limit Theorem tells you the shape of
the sampling distribution of the sample mean will be approximately normal
with this sample size, regardless of the shape of the original population
distribution.
d. The CLT establishes that the sampling distribution of the sample mean height
____ is
approximately normal, with mean 69 inches and standard deviation 3/ 100
0.3 inches.
0.3
e. Doubling the standard deviation of the sample mean gives you 0.6 inches, so
the sample mean height among the CEOs would have to be at least 69.6 inches
to persuade the psychologist that, on average, CEOs are indeed taller than the
average adult male.
f. If the sample size were 30, the normal approximation should still be valid, and
___
the standard deviation of the sampling distribution would increase to 3/ 30
0.548 inches. Doubling this standard deviation gives you 1.096 inches, so the
sample mean height would have to be at least 70.096 inches to be persuasive. The
smaller sample size produces more sampling variability, hence a larger standard
deviation. As a result, the cut-off value needed for a persuasive sample mean
height is larger.
276 Topic 14: Sampling Distributions: Means
n = 100
0.548
n = 30
67 68 69 70 71
Average Height in Sample (in inches)
100
80
Frequency
60
40
20
0
80 90 100 110 120 130 140 150
Sample Mean (in bpm)
This distribution appears normally distributed with mean 108.19 bpm and
standard deviation 11.58 bpm.
d. Yes, the distribution appears roughly normal in spite of the small sample size.
This is because the population itself was so close to being normally distributed.
288 Topic 15: Central Limit Theorem
good example of thinking through the effect of increasing sample size and the danger
of automatically thinking that a larger sample size produces a smaller probability.
Instead, you can point out that if an interval includes the population parameter, then
larger samples produce larger probabilities. If an interval is in one of the tails of the
distribution, larger samples produce smaller probabilities. For other kinds of intervals,
the issue is less clear.
Possible assignment: Activity 15-6 reminds students to identify the type of variable;
Activity 15-9 reviews proportions and statistical significance; and Activity 15-11 works
with means and statistical confidence.
Solutions
••• In-Class Activities
Activity 15-1: Smoking Rates
a. The symbol represents the proportion .209.
b. No; in general, the sample result will not equal .209 exactly because of sampling
variability.
c. The CLT predicts the sampling distribution of p̂ will be approximately normal,
centered at .209 with a standard deviation equal to
__________
(.209)(.791)
√
__________
100
.04066.
d. The following graph shows the shaded area that corresponds to a sample
proportion exceeding .25:
have a sample proportion greater than .25. The following sketch displays these
results:
15
Sample Proportion of Smokers (n = 400)
h. You calculate z (.25 .209)/.0203 2.02; Pr(Z 2.02) .0217. Yes, this
probability has decreased as predicted.
i. No, the population size of the United States did not enter into the calculations.
j. The previous calculations would not change in any way.
15
5.105
Now you can use the Normal Probability Calculator applet or the
Standard Normal Probabilities Table to find the probability of interest. The z-score
corresponding to a sample mean weight of 159.574 pounds is (159.574 167)/5.105
1.45. The probability of the weight being less than 159.574 pounds is found from the
table to be .0736, so the probability of exceeding this weight is 1 .0736 .9264. It’s
not surprising the boat capsized with 47 passengers!
Technology Ideas
Activity 16-3: Generation M (Fathom)
Part n of this activity asks students to compute 95% confidence intervals for
proportions based on sample proportions of boys and girls who have a television in
their bedrooms. This calculation may be done with an estimate box in Fathom to
allow students to quickly generate the intervals needed for this comparison, using just
the summary counts given for each sample. To get the actual counts for each gender,
students need to multiply the sample proportion given by the sample size and round to
the nearest integer, so that 0.72 * 996 or 718 boys and 0.64 * 1036 or 663 girls in the
samples have a television in their bedrooms.
1. Starting with a new Fathom document (no need to enter any data), drag an empty
estimate box from the shelf.
2. Click on Empty Estimate to show a menu of types of confidence intervals and
choose Estimate Proportion to set up the confidence interval.
3. Any of the blue values in the estimate box may be edited to reflect the actual
sample values, so change “10 out of 20” to 718 out of 996 for the sample of boys
to produce the desired 95% confidence interval.
16
4. Choose Estimate | Verbose from the main Fathom menu to turn on and off
concise vs. verbose versions of the results.
5. Edit the blue values in the estimate box to find an interval based on the 663 out
of 1036 girls in the sample.
Solutions
••• In-Class Activities
Activity 16-1: Generation M
a. The observational units are American youth ages 8–18.
b. The variable is whether each youth has a television in his or her bedroom (binary
categorical).
c. The value .68 is a statistic, which can be denoted by the symbol p̂.
318 Topic 16: Confidence Intervals: Proportions
i. You calculate .68 2(.01035) .68 .0207 (.6593, .7007). Consider the
value of the population parameter, , to be somewhere in this interval.
j. You do not know for sure whether the actual value of is contained in this
interval (as in part e).
.9800
.01 .01
–2.33 0 2.33
–z* z-values z*
.9500
.025 .025
–1.96 0 1.96
–z* z-values z*
16
z* 1.96
l. The midpoints of both intervals are the same (.68), but the 99% confidence
interval is wider; it has a greater margin-of-error (.267) vs. (.203).
m. .50: not plausible .75: not plausible Two-thirds: plausible
Explanation: Both .50 and .75 do not seem to be plausible values for as they
are not contained in either the 95% or 99% confidence intervals. However,
.67 is contained in both confidence intervals, so it seems to be a plausible
value for .
___________
n. Boys: .72 (1.96) √ .72(.28)996 .72 (1.96)(.0142) (.692, .748)
____________
Girls: .64 (1.96) √ .64(.36)1036 .64 (1.96)(.0149) (.611, .669)
o. These intervals do seem to indicate that there is a difference in the population
proportion of boys and girls who have a television in their bedrooms. You are 95%
confident the population proportion of boys with a television in their bedrooms
is at least .69 (and no more than .75), whereas you are 95% confident that the
proportion of girls with a television in their bedrooms is between .61 and .67.
There is no overlap in these intervals; the values of that are plausible values for
boys are not plausible values for girls.
p. The margins-of-error for these intervals are .028 (boys) and .029 (girls). These are
greater than the margin-of-error based on the entire sample, which makes sense
because the entire sample is roughly twice as large as the single gender samples,
so you would expect it to have less sampling variability and therefore a smaller
margin-of-error.
q. You set the margin-of-error formula equal to .01 and then solve for the sample
size n, as follows:
_________
(.68)(.32) 2
1.96 8359.322; n 8,360
n (.68)(.32) ____
√ ________ .01
.01 1.96 n
r. Answers will vary by student expectation, but the required sample size will
increase.
s. To determine the sample size, you calculate
_________
(.68)(.32) 2
2.576 14439.448; n 14,440
n (.68)(.32) _____
√ ________ .01
.01 2.576 n
d. The intervals that fail to capture have midpoints that are fairly far away (more
than 2 standard deviations) from .45 (either on the low end of the scale or the
high end).
e. No; if you had taken a single sample in a real situation, you would have no way of
knowing whether the true value of was contained in your interval because you
would not know what was; you are not guaranteed that the constructed interval
will capture the value of .
f. In the following simulation, the proportion of intervals that succeed in capturing
is 9521000 95.2%:
16
g. Yes, this percentage is close to 95%. Yes, it should be close to 95% because these
are 95% confidence intervals; the procedure should be “successful” 95% of the
time. This simulation reveals that the phrase “95% confidence” indicates that
your method of creating confidence intervals is successful in capturing the true
population parameter () 95% of the time and that it fails 5% of the time in the
long run (over many, many intervals).
h. Answers will vary by student prediction, but the intervals will become less wide
(more narrow) as the sample size increases.
322 Topic 16: Confidence Intervals: Proportions
This is reasonably close to the percentage from part f. The noticeable difference
about these intervals is that they are not as wide as those that were generated with
samples of 75 candies. [Example interval: (.341, .499).]
j. Answers will vary by student prediction, but students should predict that the
length of the intervals will shorten and the success rate will decrease when the
confidence level is changed to 90%.
k. In the following simulation, the proportion of intervals that succeed in capturing
is 9021000 90.2%.
Activity 16-6 323
16
This is not particularly close to the percentages from parts f and i, but it shouldn’t
be because you changed the confidence level. The noticeable differences about
these intervals are that they are not as wide as the previous intervals [example
interval: (.334, .466)], nor are they as successful in capturing .
90
80
Number of Couples
70
60
50
40
30
20
10
0
Right Left
Direction that Kissing
Couple Leans
c. A 95% confidence interval for the population proportion of all couples who lean
to the right is
_______________
(.645)(1 .645)
.645 1.96
√ ______________
124
which is .645 .084, or (.561, .729). You are 95% confident that the population
proportion of all kissing couples who lean to the right is somewhere between
.561 and .729. This “95% confidence” means that if you were to take many
random samples and generate a 95% confidence interval (CI) from each, then
in the long run, 95% of the resulting intervals would succeed in capturing the
actual value of the population proportion, in this case the proportion of all kissing
couples who lean their heads to the right.
d. A 90% CI is
_______________
(.645)(1 .645)
.645 1.645
√ ______________
124
which is .645 .071, or (.574, .716).
A 99% CI is
_______________
(.645)(1 .645)
.645 2.576
√ ______________
124
which is .645 .111, or (.534, .756).
The higher confidence level produces a wider confidence interval. All of these
intervals have the same midpoint: the sample proportion .645.
e. Because none of these intervals includes the value .5, it does not appear to be
plausible that 50% of all kissing couples lean to the right. In fact, all of the
intervals lie entirely above .5, so the data suggest that more than half of all kissing
couples lean to the right. The value 23 is quite plausible for this population
proportion because .667 falls within all three confidence intervals. The value
34 is not very plausible because only the 99% CI includes the value .75; the 90%
CI and 95% CI do not include .75 as a plausible value.
f. The sample size condition is clearly met, as np̂ 80 is greater than 10, and
n(1 p̂) 44 is also greater than 10. But the other condition is that the sample
be randomly drawn from the population of all kissing couples. In this study, the
couples selected for the sample were those who happened to be observed in public
Activity 16-9 325
places while the researchers were watching. Technically, this is not a random
sample, and so you should be cautious about generalizing the results of the
confidence intervals to a larger population.
(.49)(.51)
Video game player margin-of-error: (1.96)
_________
√ ________
2032
.0217
(.31)(.69)
Computer margin-of-error: (1.96)
√ ________
2032
.0201
b. Yes, the margin-of-error also depends on the confidence level and on the sample
proportion. You can tell this from the above results because all four devices have
the same sample size and confidence level, but the sample proportions differ, as do
16
the margins-of-error. The more similar the sample proportions (similar distance
from .5), the more similar the margins-of-error, all other factors being equal.
c. The video game player produces the greatest margin-of-error and the CDtape
player produces the smallest.
d. Answers will vary by student conjecture. The sample proportion that produces the
greatest margin-of-error is p̂ .5.
Technology Ideas
Activity 17-2: Flat Tires (Fathom)
As an alternative to using the Test of Significance Calculator applet
in part b of this activity, students can calculate the details of the test of a proportion,
using a test box in Fathom. The procedure is very similar to constructing a confidence
interval with an estimate box based on the summary counts, as described in the
previous topic.
1. Drag down an empty test box from the shelf, and select the Empty Test label and
change it to Test Proportion.
2. As with the estimate box, edit the blue values to change “10 out of 20” to 24 out
of 74 to reflect those students who chose the right-front tire in the sample.
3. In the sentence or expression that describes the alternative hypothesis, change the
default “is not equal to” to is greater than and change the hypothesized value
from “.5” to .25.
4. The verbose version of the test output shows the p-value at the end of a
long sentence that carefully describes what the p-value measures. Choosing
Test | Verbose to remove the check mark gives more concise output.
5. To show a graphical display of the p-value as an area in the tail of a normal
distribution approximating the distribution of sample proportions when the
null hypothesis is true, choose Test | Show p_hat Distribution.
Solutions
••• In-Class Activities
Activity 17-1: Flat Tires
a. 14 or .25
b. This value is a parameter because it represents the overall process, not simply
results that you observed, and is represented by .
c. .25
d. The sampling distribution of the sample proportion will be______________
approximately normal,
with mean equal to .25, and standard deviation equal to √ .25(1 .25)73
.0507. Here is a sketch of the sampling distribution:
Activity 17-1 341
e. The conditions necessary for the CLT to be valid are that n 10 and
n(1 ) 10. These conditions are met (73(.25) 10 and 73(.75) 10).
However, you also need to believe that the sample is representative of the
larger population process. This is less clear, but there may not be any reason to
believe that this professor’s class would behave substantially differently on this
issue than college students in general, or you may want to think more carefully
about how you define the population in this activity (e.g., students of similar
age and major).
17
34 .466. Yes, this
f. To determine the sample proportion, you calculate p̂ __
73
sample proportion is greater than 14.
g. The following graph displays the shaded area:
.0000101
.25 .466
Sample Proportion of Right-Front Tires
h. This sample result would be surprising if there were nothing special about the
right-front tire. With a p-value of approximately zero, you would consider this
sample result so surprising that you would conclude the right-front tire would be
chosen by more than one-fourth of the population.
i. Let represent the proportion of the population who will select the right-front
tire when asked this “flat tire” question.
j. H0: .25
k. Ha: .25
.466 .25
l. The test statistic is z _______________
______________ .466 .5 4.26.
________
√ .25(1 .25)73 .0507
17
a. Population parameter of interest: Let represent the proportion of times a spun
tennis racquet will land “up.”
b. H0: .5 (A spun tennis racquet will land “up” half the time.)
H a: .5 (A spun tennis racquet will not land “up” half the time.)
c. Technical conditions: As long as there was nothing unusual about the spinning
process, you will consider these data a “random” sample, and 100(.5)
100(1 .5) 50 10, so the technical conditions for the validity of this test
procedure are satisfied.
.46 .5 0.80.
d. The test statistic is z ________
_______
(.5)(.5)
√ ______
100
e. p-value 2 Pr(Z 0.80) 2 .2119 .4238
344 Topic 17: Tests of Significance: Proportions
Here is a sketch of the standard normal curve with the p-value shaded:
.212 .212
.46 .5 .54
Sample Proportion of Spins Landing “Up”
f. This p-value is very large (.4238 .05). Therefore, do not reject the null
hypothesis at the .05 significance level.
g. Conclusion in context: You do not have statistical evidence that would allow you
to conclude that a spun tennis racquet will fail to land “up” half the time.
Sample # “right-
Size front”
p̂ z statistic p-value ␣ .10? ␣ .05? ␣ .01? ␣ .001?
d. When the sample size is small, a sample result of .30 is not statistically
significant at any level. But, as the sample size increases, this result becomes
more significant—meaning that it becomes more unlikely that you would obtain
a sample result of .30 (or more extreme) if the population parameter is actually
.25 as you use larger and larger samples.
Activity 17-5 345
17
.55 .60 .65 .70 .75 .80 .85 .90 .95
Sample Proportion of Games with Big Bang
.3 .4 .5 .6 .7
Sample Proportion of Games with Big Bang
j. This p-value is not small at all, suggesting that the sample data are quite
consistent with Marilyn’s hypothesis that half of all games contain a big bang.
The sample data provide no reason to doubt Marilyn’s hypothesized value for .
k. A 95% confidence interval for (the population proportion of games that
contain a big bang) is given by
________
p̂(1 pˆ )
√
p̂ z* ________
n
with z* 1.96, which is
___________
(.495)(.505)
√
.495 1.96 __________
95
which is .495 .101, which is the interval from .394 to .596. Therefore, you are
95% confident that between 39.4% and 59.6% of all major-league baseball games
contain a big bang. The grandfather’s claim (75%) is not within this interval
or even close to it, which explains why it was so soundly rejected. Marilyn’s
conjecture (50%) is well within this interval of plausible values, which is consistent
with it not being rejected.
18
significance.
Solutions
••• In-Class Activities
Activity 18-1: Generation M
a. This is a statistic because it is a number that represents a sample.
b. For a 99% confidence interval, you calculate .68 (2.576)(.010348)
.68 (.02666) (.6533, .7066).
c. The values .70 and .6667 are in this interval; .65 and .707 are not.
d. A significance test should reject the hypothesis that .65, but .7 is contained
in the interval, so this could be a plausible value for . A significance test would
not reject the hypothesis that .7.
e. Using Minitab’s Test and CI for One Proportion:
Test of p 0.65 vs p not 0.65
Sample X N Sample p 99% CI Z-Value P-Value
1 1382 2032 0.680118 (0.653465, 0.706771) 2.85 0.004
Using the normal approximation.
18
The alternative hypothesis is that this player has improved and is now better than
a .250 hitter. In symbols, Ha: .250.
b. A Type I error would be deciding that the player has improved his batting
performance when, in fact, he is still batting no better than .250.
c. A Type II error would be failing to realize that the player has improved.
d. Answers will vary. The following is a representative set:
This distribution is roughly normal, centered at about 7.5 hits, and extends from
about 1 hit to about 17 hits.
368 Topic 18: More Inference Considerations
g. 28200 14%
h. No, it does not appear very likely that a .333 hitter will be able to establish that
he is better than a .250 hitter in 30 at-bats. Based on this simulation, he had only
about a 14% chance of establishing his improvement (performing well enough to
convince the manager that his success rate was now greater than .250).
i. Power .14
j. The following is based on one representative running of the applet:
Based on this simulation, a player would need at least 33 hits (out of 100) in order
for the probability of a .250 hitter to do that well by chance alone to be less than
.05. The approximate power of this test is 123200 or .615.
k. Answers will vary by student expectation, but the test will be more powerful if the
player improves to a .400 hitter. It should be easier to detect the improvement to
.400 because it is farther away from .250 than .333 is.
Activity 18-6 369
The simulation confirms this result; the approximate power is 108200 or .54
(using 30 at-bats).
l. Answers will vary by student expectation, but the test will be more powerful if
you use a higher significance level.
The simulation confirms this result; the approximate power is 52200 or .26
(using 30 at-bats, alternative value of .333 and the significance level .10).
m. i. The magnitude of the difference between the hypothesized value 0 and the
18
particular alternative value of
ii. The significance level
2. Rather than editing the blue values in the estimate box directly to reflect the
summary statistics from the sample, drag the Body_Temp attribute to the
“Attribute (Numeric)” label at the top of the estimate box to find the relevant
statistics automatically (mean, standard deviation, and sample size) and compute
the 95% confidence interval for the mean. If desired, choose Estimate | Verbose
to display a less wordy version of the output.
3. Select the blue “95%” and change it to 90% and then 99% to compute the other
intervals needed for the comparison in part g.
Solutions
••• In-Class Activities
Activity 19-1: Christmas Shopping
a. The amount expected to be spent on Christmas presents in 1999 is a quantitative
variable.
19
b. The value $857 is a statistic because it is a number that describes a sample. This
_
statistic is represented by x.
c. The parameter is the average (mean) amount expected to be spent by all American
adults on Christmas presents in 1999. This parameter is represented by .
d. You do not know the value of the , but it is more likely to be close to $857 than
to be far from it.
_ __ ____
e. The standard deviation of the sample mean x is √n 250√922 $8.23.
f. This interval estimate works out to be $857 2($8.23) $857 $16.47
($840.53, $873.47).
g. The sample standard deviation, s, is a reasonable substitute for .
388 Topic 19: Confidence Intervals: Means
h. Answers will vary. The following is from one representative running of the applet:
You find 96% of the intervals succeed in capturing the value of . Your total
should be roughly 95%.
.9500
19
.025 –t* 0 t* .025
T, df = 9
f. For a 95% confidence interval, t* 2.045. This value is less than the previous t*
value, which makes sense because you have increased the sample size, decreasing
the uncertainty in estimating by s.
g. For a 90% confidence interval t* 1.699; for a 99% confidence interval t*
2.759. The t* for a 99% confidence interval is greater, which is appropriate
because this interval claims more confidence (more certainty). In order to be more
confident, the interval will need to be wider.
h. With 100 degrees of freedom, t* 1.984.
25
20
Frequency
15
10
0
96.75 97.50 98.25 99.00 99.75 100.50
Body Temperature (in °F)
99.9
99
95
90
80
70
Percentage
60
50
40
30
20
10
5
0.1
95 96 97 98 99 100 101
Body Temperature (in ⴗF)
19
90% confidence interval is the most narrow (width .2131), followed by the
95% confidence interval (width .2545), whereas the 99% confidence interval
is the widest (width .3363).
h. It does not appear that 98.6 is a plausible value for the mean body temperature
for the population of all healthy adults because this value is not contained in any
of the confidence intervals.
i. If the sample size had been only 13, but the sample mean and standard deviation
had been the same, the 95% CI would be much wider (though it would have the
same center) because (i) the t* value used to create the interval would be much
greater and (ii) the standard error would be greater (the square root of 13 is much
smaller than the square root of 130).
j. For a 95% confidence interval
___ with 12 degrees of freedom, you calculate
98.249 (2.179) (.733)√13 (97.807, 98.692)°F.
As predicted, the midpoint is still 98.249, but the width is much greater (.88597).
In fact, with this interval, 98.6 would be a plausible value for the mean body
temperature for the population of all healthy adults.
392 Topic 19: Confidence Intervals: Means
3 30 6.6 0.825
1 10 6.6 0.825
2 10 6.6 1.597
4 30 6.6 1.597
You are 90% confident the average amount of sleep per night obtained by all
students at the school is between 6.45 and 7.51 hours.
e. Seven of the 40 sleep times fall within this interval. This is 17.5%.
f. No, this percentage is not close to 90%, but there is no reason that it should be.
Your interval is designed to estimate the average sleep time. It is not telling you
anything about the individual sleep times. They may or may not fall within this
interval.
20
16
Frequency
12
0
.00 .02 .04 .06 .08 .10 .12 .14 .16 .18 .20
19
Ratio of Backpack Weight to Body Weight
c. Calculating a 99% CI for the population mean by hand using the formula
_ s__
x t* ___
√n
e. The first condition is that the sample be randomly selected from the population.
This is not literally true in this case because the student researchers did not obtain
a list of all students at the university and select randomly from that list, but they
did try to obtain a representative sample. The second condition is either that the
population of weight ratios is normal or that the sample size is large. In this case,
the sample size is large (n 100, which is greater than 30), so this condition is
satisfied even though the distribution of ratios in the sample is somewhat skewed
(and so presumably is the population).
f. You do not expect 99% of the sample, nor 99% of the population, to have
a weight ratio between .0673 and .0869. You are 99% confident that the
population mean weight ratio is between these two endpoints. In fact, only 18 of
the 100 students in the sample have a weight ratio in this interval.
Technology Ideas
Activity 20-2: Sleeping Times (Fathom)
Part e of this activity asks students to use technology to calculate test statistics and
p-values for four different samples of hypothetical sleep times.
1. Open the [Link] file, drag an empty test box from the shelf, and
change the type of test to Test Mean.
2. Drag one of the samples to the top row of the test box, and edit the blue value
in either hypothesis to change the hypothesized mean from “0” to 7. Select
the test box or press Õ to accept the change. You might want to choose
Test | Verbose from the main Fathom menu to see more concise results.
3. Drag each of the other samples to the same test box in turn to find the results
needed to summarize each test in the table in part b.
Solutions
As an alternative, you can explore the Draw feature.
20
••• In-Class Activities
Activity 20-1: Basketball Scoring
a. If the rule changes had no effect on scoring, would have the value 183.2. This
is the null hypothesis.
b. If the rule changes had the desired effect on scoring, then 183.2, which
would be the alternative hypothesis.
408 Topic 20: Tests of Significance: Means
Scoring definitely seems to have increased over the previous season’s mean of
183.2 points per game. In only 4 of these 25 games were fewer than 183 points
scored, and the center of this dotplot is now about 196 points per game.
_
d. Sample mean x : 195.88 points Sample standard deviation s: 20.27 points
e. Yes, the sample mean is in the direction specified in the alternative hypothesis
(greater than 183.20 points).
f. Yes, it is possible to have gotten such a large sample mean even if the new rules had
no effect on scoring.
g. H0: 183.2
Ha: 183.2
h. Technical conditions: The sample size is not large (n 25 30), so the
population must follow a normal distribution. A probability plot of the data
provides evidence that these sample data are not arising from a normal population.
Still, the data are reasonably symmetric and the sample size is moderately large, so
this condition could be considered met with caution.
99
95
90
80
70
Percentage
60
50
40
30
20
10
5
1
150 175 200 225 250
Points Scored
However, the data are not a simple random sample gathered from the population
as the data are all the NBA games played from December 10–12, 1999. These
games occurred relatively early in the season and are probably not representative of
scoring over the course of the entire season, especially with respect to a new rule
change. So the technical conditions for the validity of this t-test have not been met.
195.88 183.2
i. The test statistic is t _____________
___ 3.13.
20.27√25
j. Here is a sketch of the t-distribution:
Activity 20-2 409
df = 24
.00227
0 3.13
t-values
m. If the average number of points per game for all NBA games in this season were still
183.2 points, there is only a .0023 chance that you would find a random sample
of 25 games with a mean of at least 195.88 points. Because finding a sample as
extreme as this one is so unlikely by chance alone, you conclude that the mean
number of points per game this season has increased; it is no longer 183.2 points.
n. Yes, you would reject the null hypothesis at the .10, .05, .01, and .005 levels
because the p-value is less than each of these significance levels.
20
o. If these data had been a random sample from a normal population, you would have
very strong statistical evidence that the mean points per game in the 1999–2000
season were greater than in the previous season. However, you would not be able to
conclude that the rule change caused the average point increase because this is not
a randomized experiment.
The null hypothesis is that the mean sleep time of the population is 7 hours. In
symbols, the null hypothesis is H0: 7.0 hours.
The alternative hypothesis is that the mean sleep time of the population is not
7 hours. In symbols, the alternative hypothesis is Ha: 7.0 hours.
6.981 ___7 0.06.
The test statistic is t _________
1.981√40
Using Table III with 39 degrees of freedom, p-value 2 .20 .40.
Using Minitab, p-value 2 .476231 .952462.
Because the p-value is not small, do not reject H0.
You do not have any statistical evidence to suggest that the mean sleep time of all
students at your school differs from 7.0 hours.
b. Technical conditions: The sample size is large (40 30), but the sample was
not randomly selected because it consisted of only the students in your class. It
might not be representative of students at your school with regard to sleep hours,
as students in a statistics class may tend to have similar majors and may tend to
study and sleep more or less than the typical student. Here is the completed table:
Sample Number Sample Size Sample Mean Sample SD Test Statistic p-value
c. Sample 1 would produce a smaller p-value because it has the smaller standard deviation.
d. Sample 3 would produce a smaller p-value because it has the larger sample size.
e. See the table following part b.
f. Only sample 3 gives enough evidence to reject the null hypothesis at the .05 level.
g. Answers will vary by student conjecture.
8
7
6
Frequency
5
4
3
2
1
0
.6 .7 .8 .9
Width-to-Length Ratio
Activity 20-4 411
99
95
90
80
70
Percentage
60
50
40
30
20
10
5
1
.3 .4 .5 .6 .7 .8 .9 1.0
Width-to-Length Ratio
.6695 ___
.618 2.05.
The test statistic is t ___________
.0925√20
The p-value is 2 Pr(T 2.05).
Using Table III with 19 degrees of freedom, 1.729 2.05 2.093, so
2 .025 p-value 2 .05 or .05 p-value .10. You would reject
H0 at the .10 significance level.
Using the applet, the p-value .0539 .10. You would reject H0 at the
.10 significance level.
20
You can conclude that the average width-to-length ratio for all beaded rectangles
made by the Shoshoni Indians is not the golden ratio.
The variable measured here is the amount of television the student watches in a
typical week, which is quantitative. The parameter is the mean number of hours of
television watched per week among the population of all third- and fourth-graders.
This population mean is denoted by . The question asked about watching an average
of two hours of television per day, so convert that to be 14 hours per week.
The null hypothesis is that third-and fourth-graders in the population watch an
average of 14 hours of television per week (H0: 14). The alternative hypothesis is that
these children watch more than 14 hours of television per week on average (Ha: 14).
Check technical conditions:
• The sample of children was not chosen randomly; they all came from two schools
in San Jose. You might still consider these children to be representative of third-
and fourth-graders in San Jose, but you might not be willing to generalize to a
broader population.
• The sample size is large enough (198 is far greater than 30) that the second
condition holds regardless of whether the data on television watching follow a
normal distribution. You do not have access to the child-by-child data in this case,
so you cannot examine graphical displays; however, the large sample size assures
you that this condition may be considered satisfied.
_
Test statistic: The sample size is n 198; the sample mean is x 15.41 hours;
and the sample standard deviation is s 14.16 hours. The test statistic is
15.41 ____
14 1.401
t __________
14.16√198
indicating the observed sample mean lies 1.401 standard errors above the conjectured
value for the population mean. Using Table III and the 100 degrees of freedom line
(rounded down from the actual number of degrees of freedom of 198 1 197)
reveals the p-value (probability to the right of t 1.401) to be between .05 and .10.
Technology calculates the p-value more exactly to be .081.
Test decision: This p-value is not less than the .05 significance level. The sample
data, therefore, do not provide sufficient evidence to conclude that the population
mean is greater than 14 hours of television watching per week.
Conclusion in context: This conclusion stems from realizing that obtaining a
sample mean of 15.41 hours or greater would not be terribly uncommon when the
population mean is really 14 hours per week. If you had used a greater significance level
(such as .10), which requires less compelling evidence in order to reject a hypothesis,
then you would have concluded that the population mean exceeds 14 hours per week.
Activity 21-1 441
21
2. Edit the blue values to reflect the counts and sample sizes for each of the two
groups (740 out of 1000 in one sample and 240 out of 1000 in the other).
3. Select the blue text, “is not equal to” in the alternative hypothesis statement ( Ha )
and change it to is greater than for the lower tail test.
4. (Optional) For a more informative display, change the generic names
“FirstAttribute” and “SecondAttribute” to 1992 and 1996, respectively, and
change “Category” to Stomach as a reminder of what the proportions measure.
To confirm the interval calculation, follow these steps:
1. Drag a new estimate box from the shelf and change the type of estimation to
Difference of Proportions.
2. As in the steps above, edit the counts to show those found in the two samples.
Solutions
••• In-Class Activities
Activity 21-1: Friendly Observers
a. Explanatory: whether the observer shares Type: binary categorical
the prize
Response: whether the participants beat Type: binary categorical
the threshold time
b. This is an experiment because the researchers randomly assigned the subjects to
the control and treatment groups.
c. The sample proportions of success for each group are p̂A .25; p̂B .6667.
The following segmented bar graph displays the results:
442 Topic 21: Comparing Two Proportions
100
Did not beat threshhold
90
Beat threshhold
80
70
Percentage
60
50
40
30
20
10
0
A: Observer B: No Sharing
Shares Prize of Prize
Sharing of Prize
Repetition # 1 2 3 4 5
0 1 2 3 4 5 6 7 8 9 10 11
Number of Successes Assigned to Group A
i. The most common values for the number of successes are 5 and 6. This result
makes sense because if the cards are randomly assigned to groups A and B, then
half of the 11 successes should end up in group A. So, typically you expect 5 or 6
of the successes in group A.
j. Forty repetitions were performed by this class. Five of these repetitions gave a
result of 3 or fewer successes in group A. This represents a proportion of .125.
k. Based on this class’s simulated results, it does not appear unlikely for random
assignment to produce a result as extreme as (or more extreme than) the actual
sample when the observer has no effect on subjects’ performance.
Activity 21-2 443
21
l. No, in light of your answer to the previous question, the data do not provide
reasonably strong evidence in support of the researchers’ conjecture; this type of
sample result appears to happen in about 12–13% of random assignments when
the observer has no effect on subjects’ performance.
m. Yes, the number of successes randomly assigned to group A varies. The five values
that occur in this running of the applet are 6, 5, 5, 8, and 7.
n. In this simulation, the resulting distribution is approximately normal, centered at
about 5.4 successes, with a standard deviation of 1.327 successes.
o. Here are the simulation results:
.24 .70
d. The test statistic is z ______________________
_____________________ 20.61.
1
1 _____
√ (
(.47)(.53) _____
1000 1000 )
Yes, this is a very large z-score (in absolute value).
e. p-value Pr(Z 20.61) .000
f. If the proportion of infants who sleep on their stomachs is the same in 1996 as
it was in 1992, the probability that you would find a difference in sample results
as or more extreme than this study by random sampling alone is essentially 0,
which means it would just about never happen. Because you did find this sample
difference, you have very strong evidence that the null hypothesis was an incorrect
conjecture, and the proportion of infants who sleep on their stomachs in 1996 was
less than it was in 1992.
g. Reject the null hypothesis at the .01 significance level because the p-value is
smaller than .01.
h. This sample difference provides extremely strong evidence that the proportion
of all households that place infants to sleep on their stomachs decreased between
1992 and 1996.
___________________
(.70)(.30) (.24)(.76)
1000 √
i. For a 95% CI, you calculate (.70 .24) (1.96) ________ ________ ,
1000
which is .46 (1.96) (.0198), which is .46 .0388 (.4212, .4988).
j. You are 95% confident the difference in population proportions is between
.420 and .499. Because this interval does not contain 0, you are 95% confident
the population proportions are not the same. Because the values in the interval are
strictly positive, you are 95% confident that the percentage of infants who sleep
on their stomachs has decreased (between 42 and 50 percentage points) in the
years from 1992 to 1996.
k. Here are the applet results:
Activity 21-4 445
21
l. Answers will vary by student expectation. ___________________
(.70)(.30) (.24)(.76)
√
m. For a 95% CI, you calculate (.24 .70) (1.96) ________ ________
1000 1000
(.4988, .4212). The endpoints are reversed and multiplied by 1 in this
interval. The interval has the same width as the previous interval, however, and its
midpoint is the negative of that of the previous one.
with their appearance, and this difference of 10 people would not seem to be
significant (it could have arisen simply from random sampling variability).
c. Answers will vary, but one example is if 1000 men and 1000 women were
surveyed. In this case, you would not expect much sampling variability by chance,
and the observed difference would seem much more convincing of a difference in
the populations.
d. Here is the completed table:
Significant at:
e. The larger the sample size, the more statistically significant a difference in
proportions will be. A large difference may not be significant with a small sample
size, but even a small difference may be statistically significant when the sample
size is large.
21
Activity 21-6: Nicotine Lozenge
a. The explanatory variable is the type of lozenge (nicotine or placebo). The response
variable is whether the smoker successfully abstains from smoking for the year. Both
variables are categorical and binary.
b. This is an experiment because the researchers assigned the subjects to take a
particular kind of lozenge (nicotine or placebo).
c. The null hypothesis is that the nicotine lozenge is no more (or less) effective than
the placebo, where effectiveness is measured by the proportion of smokers who
successfully abstain from smoking for a year. The alternative hypothesis is that the
nicotine lozenge is more effective than the placebo, meaning a higher proportion
of smokers (who are interested in and might potentially use such a product)
would successfully quit with a nicotine lozenge than with a placebo lozenge. In
symbols, the hypotheses are H0: nicotine placebo vs. Ha: nicotine placebo, where
represents the population proportion of smokers who successfully abstain for a
year if given the nicotine lozenge or the placebo.
d. The 2 2 table is shown here:
The following segmented bar graph shows those smokers taking the nicotine
lozenge had a higher success rate (proportion) in this study than those smokers
taking the placebo lozenge, almost twice as high (.179 vs. .096). But the graph
also reveals that in both groups, many more smokers resumed smoking than were
able to abstain successfully.
100
Resumed smoking
90
Successfully abstained
80
70
Percentage
60
50
40
30
20
10
0
Nicotine Placebo
Type of Lozenge
Solutions
••• In-Class Activities
Activity 22-1: Close Friends
a. This is an observational study because the researcher simply observed the gender
of each subject—he/she did not randomly assign gender to subjects.
b. Explanatory: gender Type: binary categorical
Response: number of “close friends” Type: quantitative
c. The null hypothesis is men and women tend to mention the same average number
of close friends in response to this question. In symbols, H0: m f .
The alternative hypothesis is men and women differ in the average number of close
friends they tend to mention in response to this question. In symbols, Ha: m f .
d. The two-sample z-test procedure from Topic 21 does not apply in this situation
because you are testing hypotheses about population means rather than population
proportions. The one-sample t-test procedure from Topic 20 does not apply
because you are comparing two means in this situation, not testing a claim about
a single mean.
e. These values are statistics because they describe samples, not populations.
f. The following boxplots compare the distributions of the number of close friends:
Male
Female
0 1 2 3 4 5 6
Number of Close Friends
g. There are very few differences in the distributions of the number of close friends
mentioned between males and females. The males have a slightly lower mean,
median, and lower quartile, but the upper quartiles and maximums are identical
to those of females.
h. Yes, it would be possible to obtain sample means this far apart even if the
population means were equal.
1.861 2.089 2.45.
i. The test statistic is t _______________
______________
1.777 2 _____
______ 1.76 2
√ 654 813
j. Using 500 degrees of freedom, 2.586 2.45 2.334, so 2 .005
p-value 2 .01, which means .01 p-value .02.
k. The correct interpretation of the p-value is “The p-value is the probability of
getting sample data so extreme if, in fact, males and females have the same mean
number of close friends in the populations.”
l. Yes, this p-value is small enough to reject the null hypothesis at the ␣ .05
significance level (p-value .02 .05).
Activity 22-2 473
m. No, the observed difference in sample means is not statistically significant at the
␣ .01 significance level (p-value .01).
n. Technical conditions:
22
i. The data are a random sample broken into two distinct groups (see page 418).
ii. Both sample size are large (greater than 30).
Because both sample sizes are large, you do not need to worry about the strong
skewness in the sample data. It does not provide any reason to doubt the validity
of this test.
o. For a 95% CI for f m with 500 degrees of freedom, you calculate (2.089
1.861) (1.965)(0.0929) (0.045, 0.411). This interval is entirely positive (and
does not include zero), which means you can be 95% confident that the mean
number of close friends that women have is between .045 and .411 greater than
the mean number of close friends that men have.
p. Here are the applet results:
Alex: Route 1 10 28 6
.155
Alex: Route 2 10 32 6
d. H0: 1 2 Ha: 1 2
28 32 1.49.
The test statistic is t _________
________
2
6 62
___ ___
√ 10 10
474 Topic 22: Comparing Two Means
Barb: Route 1 10 25 6
.002 (Minitab)
Barb: Route 2 10 35 6 .0047 (applet)
Carl: Route 1 10 28 3
.008 (Minitab)
Carl: Route 2 10 32 3 .0154 (applet)
Donna: Route 1 40 28 6
.004 (Minitab)
Donna: Route 2 40 32 6 .0049 (applet)
22
symbols, Ha: JFK JFKC.
Answers will vary by class. The following is one representative set of answers.
e. The following visual displays display the data:
JFK
4 8 12 16 20 24 28
Group
JFKC
4 8 12 16 20 24 28
Number of Correctly Memorized Letters
JFK
JFKC
0 5 10 15 20 25 30
Number of Correctly Memorized Letters
Yes, these data appear to support the conjecture that those who receive the letters
in convenient three-letter chunks tend to correctly memorize more letters. The
center of this plot is about six letters higher than for the JFKC plot, whereas the
spreads are not very different.
f. The data arise from random assignment of subjects to two treatment groups, but the
sample sizes are not quite large enough (n 27 and 26, respectively). Probability
plots indicate that the JFKC (inconvenient chunks) distribution is likely to be
normal, but the JFK (three-letter chunk) distribution is skewed left and not normal.
99
95
90
80
70
Percentage
60
50
40
30
20
10
5
1
0 10 20 30 40
JFK
476 Topic 22: Comparing Two Means
99
95
90
80
70
Percentage 60
50
40
30
20
10
5
1
10 0 10 20 30
JFKC
Therefore, for these data, the results of the test of significance should be
interpreted with caution.
16 10.88 3.03.
g. The test statistic is t ______________
_____________
6.45 2 _____
_____ 5.86 2
√27 26
Using Table III with 25 degrees of freedom, 2.787 3.03 3.450, so
.001 p-value .005. Using Minitab, the p-value is .002. Using the applet,
the p-value is .0028.
With a p-value less than .05, reject H0. You have strong statistical evidence that
the average number of letters correctly memorized is higher under the convenient
three-letter condition than under the inconvenient condition. If there were no
difference in the average number of letters correctly memorized under these two
conditions, you would see a result as extreme or more extreme than in this study
by chance alone in only about .2% of random assignments.
h. The confidence interval formula is: (16 10.88) t*(1.69). Using Table III
with 25 degrees of freedom, t* 1.708, so the confidence interval is (2.23,
8.01). Using Minitab, the confidence interval is (2.28, 7.95). Using the applet,
the interval is (2.23, 8.01). You are 90% confident that the convenient three-
letter grouping helps individuals memorize between 2.3 and 8.0 more letters, on
average, than the inconvenient grouping.
i. Because this is a well-designed experiment in which the subjects were randomly
assigned to the two groups, you can conclude that the convenient three-letter
grouping did cause higher memory scores on average.
b. The explanatory variable is whether the waitress included her name as part of
her greeting; this variable is categorical and binary. The response variable is the
amount of the tip, which is quantitative.
22
c. The null hypothesis is that there is no effect of the waitress using her name in
her greeting. In other words, the null hypothesis says that the population mean
tip amount would be the same when she uses her name as when she does not.
The alternative hypothesis is that using her name has a positive effect, that the
population mean tip amount would be greater when she uses her name than when
she does not. In symbols: H0: name no name vs. Ha: name no name.
d. A Type I error would mean that the waitress decides that using her name helps
when it really doesn’t, so she would waste the minimal effort of giving her name
and reap no benefit from it. A Type II error means the waitress decides using her
name is not helpful even though it actually is helpful, so she would not bother
to give customers her name and would lose out on that benefit. Because the cost
of giving her name is minimal, losing out on potential tips is probably more of a
concern.
5.44 3.49
e. The test statistic is t ________________
_______________ 1.95 4.19.
_____
(1.75) 2
(1.13) 2 0.466
√ ______
20
______
20
Using Table III with 19 degrees of freedom reveals that this test statistic is off the
chart, so the p-value is less than .0005.
f. Because the p-value is less than .05, reject the null hypothesis at the ␣ .05 level.
Indeed, you would also reject the null hypothesis at the ␣ .01 and even at the
␣ .001 levels. The data provide very strong evidence that giving her name as
part of her greeting does lead to higher tips on average.
g. You do not have enough information to check the technical conditions
thoroughly. You do know that the parties were randomly assigned to one group or
the other. But the sample sizes are not large, so you should check whether the tip
data could reasonably have come from normal distributions; however, you only
have the summary statistics, not the actual tip amounts from each party, so you
cannot check this condition. You should ask the waitress to provide the actual
party-by-party tip amounts to help you assess the shape of the distribution of tip
amounts.
h. A 95% confidence interval for name no name is
_______________
(1.75) 2 (1.13) 2
√
(5.44 3.49) 2.093 ______ ______
20 20
which is 1.95 2.093(0.466), which is 1.95 0.97, which is the interval (0.98,
2.92). You can be 95% confident that the waitress would earn between $0.98 and
$2.92 more per party, on average, with a $23.21 check, by giving her name as part
of her greeting (assuming that the tip amounts are roughly normally distributed).
i. Conclude a causal link between the waitress giving her name and receiving
higher tips on average. Random assignment would have assured that the only
difference between the groups was whether the party was given the waitress’s
name. Because the group who was told her name gave significantly higher tips
on average (p-value .0005), you can attribute that to being told her name in
greeting unless the waitress gave better service to customers to whom she gave
478 Topic 22: Comparing Two Means
her name. The confidence interval enables you to say more: that giving her
name to customers increases the waitress’ tips by an average of about
1–3 dollars per dining party at Sunday brunch in this restaurant. But you
must be cautious about generalizing this result to other waitresses because only
one waitress participated in this study. Even for this particular waitress, you
should be cautious about generalizing the results to customers beyond those
who partake of Sunday brunch at that particular Charlie Brown’s restaurant
in southern California. You should also remember that these p-value and
confidences interval calculations are only valid if the tip amounts roughly
follow a normal distribution.
3. Rather than finding the mean and standard deviation of the two lists, select Data
for Inpt and insert HUSB as List1 and WIFE as List2 .
4. Leave the frequencies as 1 and Pooled as No.
5. Select Calculate and press Õ.
6. As an alternative, use the Draw feature of the 2-SampTTest.
23
Solutions
••• In-Class Activities
Activity 23-1: Marriage Ages
a. The null hypothesis is that the population mean marriage age for husbands is the
same as the mean marriage age for their wives. In symbols, H0: husbands wives.
The alternative hypothesis is that the population mean marriage age for husbands is
greater than the mean marriage age for their wives. In symbols, Ha: husbands wives.
35.71 33.83 0.46.
The test statistic is t ________________
_______________
14.56 2 ______
______ 13.56 2
√ 24 24
Using Table III with 23 degrees of freedom, 0.46 0.858, so the p-value is off
the chart. This means the p-value .20. Using Minitab, the p-value is .323.
Using the applet with 23 degrees of freedom, the p-value is .3226.
Do not reject H0 at the ␣ .05 significance level (.32 .05). You do not have
statistical evidence that the population mean marriage age for husbands is greater
than for wives.
b. For a 90% confidence interval: (35.71 33.83) t*(4.06). Using Table III with
23 degrees of freedom, t* 1.714, so the confidence interval is (5.08, 8.84).
Using Minitab, the interval is (4.95, 8.70). Using the applet, the interval is
(5.08, 8.84).
c. Most of the values are above the y x line, which indicates a clear tendency for
husbands to be older than their wives.
d. The husband’s and wife’s ages appear to be related. People tend to marry people
in the same age group: younger people tend to marry younger people, and older
people tend to marry older people.
e. For a 90% CI with ___23 degrees of freedom, you calculate 1.875
(1.714) ( 4.812√ 24 ) 1.875 .98225 (.1914, 3.5586). You are 90%
confident the population mean difference in ages between husbands and
their wives is between .19 and 3.6 years.
f. This interval is entirely positive, indicating that there is a difference in the
population mean ages of husbands and their wives. The incorrect interval in part
b contained both positive and negative values (and zero). The midpoint of this
interval is 1.875 years and its width is 3.367 years. The midpoint of the previous
(incorrect) interval was also 1.875 years, but the width was 13.65 years.
g. The null hypothesis is that the population mean difference in ages between
husbands and their wives is zero. In symbols, H0: d 0. The alternative
504 Topic 23: Analyzing Paired Data
hypothesis is that the population mean difference in ages between husbands and
their wives is greater than 0. In symbols, Ha: d 0.
1.875___ 1.91.
The test statistic is t __________
4.812√ 24
Using Table III with 23 degrees of freedom, 1.714 1.91 2.069, so .025
p-value .05. Using the applet with 23 degrees of freedom, the p-value is .0344.
Reject H0 at the ␣ .05 significance level (.0344 .05). You have moderate
statistical evidence the population mean difference in ages between husbands and
their wives is greater than zero.
h. This is the opposite of the (incorrect) conclusion you drew in part a.
i. Technical conditions: This is not a simple random sample, but you have no
reason to suspect that this sample is not representative of at least marriages in
Cumberland County, Pennsylvania. The sample size of pairs is not large (n
24 30), but a probability plot of the sample differences indicates that it is
plausible the sample comes from a normally distributed population.
99
95
90
80
70
Percentage
60
50
40
30
20
10
5
1
10 5 0 5 10 15 20
Difference (in years)
j. The paired analysis produces such a different conclusion from the independent-
samples analysis because it reduces the variability so much. The standard
deviation of the differences in ages is much smaller than the standard deviations
of the ages. This makes the observed difference in the mean ages seem larger (less
likely to happen by chance). You can also see that this reduces the denominator of
the test statistic creating a larger value and lowering the p-value.
k. Yes, the researcher was wise to gather paired data. Because the husband’s age and
the wife’s age were related, pairing was very helpful for estimating the population
mean age difference among married couples.
Chocolate Chip
56 70 84 98 112 126 140
23
Melting Time (in seconds)
Chocolate Chip
_
n x s Min QL Median Qu Max
Chocolate Chip 16 79.9 23.3 55.0 56.3 77.5 95.0 140
Peanut 16 93.8 27.5 45.0 71.3 100.0 113.8 148
Butter Chip
It appears that for these students, peanut butter chips tended to take a little longer
than chocolate chips to melt (mean 93.8 vs. 79.9 seconds; almost all five numbers
in five number summary are larger).
d. The null hypothesis is that the population mean melting time for both chips
is the same. In symbols, H0: chocolate peanut butter. The alternative hypothesis is
that the population mean melting time for both chips is not the same. In symbols,
Ha: chocolate peanut butter.
Technical conditions: The data are from a completely randomized experiment.
However, the sample sizes are not large (16 and 16), but probability plots do not
provide strong evidence that the data come from nonnormal distributions (though
there are some outliers with the chocolate chip melting times).
99
95
90
80
70
Percentage
60
50
40
30
20
10
5
1
20 40 60 80 100 120 140
Chocolate Chip Melting Time (in seconds)
506 Topic 23: Analyzing Paired Data
99
95
90
80
70
Percentage 60
50
40
30
20
10
5
1
20 40 60 80 100 120 140 160
Peanut Butter Chip Melting Time (in seconds)
The null hypothesis is that the population mean difference in melting times
between chocolate and peanut butter chips is 0. In symbols, H0: d 0. The
alternative hypothesis is that the population mean difference in melting times
between chocolate and peanut butter chips is not 0. In symbols, Ha: d 0.
8.8___ 2.27.
The test statistic is t _________
22.0√ 32
Activity 23-4 507
23
8.8 (1.696)( 22.0√ 32 ) (15.4, 2.2)
Because the p-value of .03 is less than .05, you have some statistical evidence that
there is a difference in the mean melting time between the two kinds of chips.
You are 90% confident the melting time of the peanut butter chips is between 2.2
and 15.4 seconds longer than that of the chocolate chips, on average.
g. The conclusions for these two studies were quite different. In the first study, you had
no statistical evidence of a difference in the chip melting time (p-value .135), and
in the paired data study, you had some evidence of a difference (p-value .03).
Pairing the data reduced the variability in the data, as the standard deviation of
the differences was 22.0 seconds, whereas the standard deviations of the individual
melting times were both more than 23 seconds.
c. The null hypothesis is that the population mean time required to wake up to the
conventional smoke alarm is the same as the population mean time required to
wake up to the personalized alarm. In symbols, H0: conventional personalized.
The alternative hypothesis is the population mean time required to wake up to
the conventional smoke alarm is greater than the population mean time required
to wake up to the personalized alarm. In symbols, Ha: conventional personalized.
d. Wake each child on two different nights—once with the personalized alarm and
once with the conventional smoke alarm. Use randomization to decide the order
of the alarms. Record the time required for the child to wake each time, and then
calculate the difference in the times (for instance, conventional personalized) for
each child. Use the one-sample t-test to determine whether the average of these
differences is greater than zero.
e. Let d represent the average of the population differences in the times required to
wake a child (conventional alarm personalized alarm).
H0: d 0 vs. Ha: d 0
f. You should recommend the matched-pairs design because it would have less
variability. The times required to wake the children might vary quite a bit, but
the difference in the times required to wake the children using the two methods
should vary much less.
Three of these differences are negative, and seven are positive. Therefore, in most
pairs, the woman lasted longer until fatigue set in than the man did. The mean
of these differences is 895 seconds, so the women in the sample outlasted the men
by almost 15 minutes on average. The standard deviation is 1148 seconds. The
distribution of differences appears to be a bit skewed to the right.
e. Because of the small sample size (n 10 pairs), the distribution of differences
must be approximately normal in order for the technical conditions to be satisfied.
But you have already seen that the distribution is a bit skewed to the right. This
skewness is also seen in a normal probability plot, which does not look very linear:
Activity 23-5 509
99
95
90
80
70
Percentage
60
23
50
40
30
20
10
5
1
4000 3000 2000 1000 0 1000 2000 3000 4000 5000
Difference (in seconds)
You also lack either a randomly selected sample or random assignment to treatment
groups, and therefore the paired t-test conditions were not met. We should stop
the analysis at this point, but we will continue with the calculations for the sake of
practice, remembering that we cannot take the results very seriously.
f. You are testing H0: d 0 vs. the alternative Ha: d 0, where d represents the
population mean of the differences in times until fatigue between healthy young
adult men and women. The test statistic is
895 ___ 2.465
t _________
1148√ 10
The p-value, from Table III (t-Distribution Critical Values) with 9 degrees of
freedom and remembering the alternative hypothesis is two-sided, is between
2(.01) and 2(.025), which means the p-value is between .02 and .05. Technology
gives a p-value of .036. This p-value is fairly small, so you would reject the null
hypothesis at the ␣ .05 level. The sample data provide fairly strong evidence
that the mean time until fatigue differs between men and women.
g. A 90% confidence interval for d is
1148
895 1.833 _____
___
√ 10
which is 895 665.4, which gives the interval (229.6, 1560.4). This interval means
you can be 90% confident that women take between 230 seconds and 1560 seconds
longer, on average, to reach fatigue than men do (between 426 minutes).
h. This is not a randomized experiment, so even if the test were valid, you would
not be able to attribute the longer time to fatigue for women to any particular
cause. Moreover, because the subjects volunteered, you should be cautious about
generalizing the results even to all healthy young adults in the area of the study.
Finally, the technical conditions of the paired t-test were not met because of the
small sample size and skewed distribution of differences, so you cannot take the
inference results very seriously. Researchers might want to repeat the study with
large sample sizes.
Activity 24-1 529
6. For the built-in calculator 2 GOF-Test, you can explore the Draw feature as an
alternative. (The CHIGOF program does not have the Draw feature.)
24
1. Create a new (empty) collection, and choose Collection | New Cases to add
486 cases to the collection. Drag a case table from the shelf to view the (empty)
collection.
2. Name the first attribute in the collection Ball and use the formula editor to enter
the formula RandomPick(1,2,3,4) to create a sample in this column of ball
numbers where each of the four digits is equally likely.
3. Drag a test box from the shelf and change the type of test to Goodness of Fit.
Drag the Ball attribute to the top slot, labeled “Attribute (categorical),” to produce
the goodness-of-fit test results needed to answer parts a–c of Activity 24-12.
To collect the results for tests on 1000 such simulated samples in Fathom, continue
with these steps:
4. Select the test box and choose Test | Collect Results as Measures from the main
Fathom menu to create a new collection with the results (including the chi-square
statistic and p-value) from five simulations. Drag a new table from the shelf to
view these results.
5. Double-click the new Measures from Test of Collection 1 box to show its
inspector. On the Collect Measures tab, uncheck the “Animation on” box
to turn off animation, check Replace existing cases, and change the number
of measures to 1000. Click Collect More Measures to perform the 1000
simulations and tests.
6. To obtain the count for part f of the simulated test statistics that exceed the test
statistic from the original data in Activity 24-11, select the test statistic attribute
in the new Measures collection and choose Table | Sort Descending from the
main Fathom menu to reorder the 1000 simulated p-values. Compare this count
out of 1000 to the p-value for the original test.
Solutions
••• In-Class Activities
Activity 24-1: Birthdays of the Week
a. Observational units: writers
Variable: days of the weeks on which the birthdays fall Type: categorical
530 Topic 24: Goodness-of-Fit Tests
Percentage
.12
.10
.08
.06
.04
.02
0
Monday Tuesday Wednesday Thursday Friday Saturday Sunday
Day of Week
There does not appear to be a great difference in the proportion of people born on
any given day of the week.
c. These values ( Tu, We, . . . , Su ) describe the proportion of all people who were
born on Tuesday, Wednesday,… Sunday. The values are parameters because they
describe an entire population.
d. The values are M Tu … Su 17 .1429.
e. You would expect the count of Monday birthdays to be (17) 147 21.
f. See the middle row of the following table.
(O E)
________
2
.7619 1.1905 .0476 .1905 .1905 1.7143 .7619 4.857
E
.598
24
0 4.587
k. If all seven days of the week are equally likely to be a person’s birthday, the
probability that you would obtain a test statistic this large (or larger) by random
chance alone is at least .2.
l. With the larger p-value, do not reject H0 at the ␣ .10, ␣ .05, or ␣ .01
levels.
m. Yes; all the expected counts are 21, which is greater than 5. However, the random
sampling condition is not met. These are the birthdays of “noted writers of the
present,” which is not a random sample of the population of U.S. citizens.
n. You have no statistical evidence against the null hypothesis that the seven days of
the week are all equally likely to be a person’s birthday (at least for the population
of famous writers).
Observed 17 26 22 23 19 15 25 147
Count
(O E)
________
2
2.2959 .0918 .2551 .0918 1.235 .6173 13.270 17.857
E
f. You have strong statistical evidence that the null hypothesis is not true; that is,
that one or more of the population proportions is not equal to the values that you
proposed in the null hypothesis. (So, one or more of M 16, Tu 16, W
16, Th 16, F 16, Sa 112, Su 112 is not true.) If all of these
were true, you would see a sample result (chi-square test statistic) as large or larger
by random chance alone somewhere between .5% and 1% of the time (of random
samples from such a population). Because this would occur so rarely by random
chance, you conclude the null hypothesis is false.
e. Technical conditions: The expected counts are at least five in both categories
(smallest 31). The couples selected for the sample were those who happened
to be observed in public places while the researchers were watching. This is not a
random sample, so you should be cautious about generalizing the results of this
test to a larger population.
f. Considering these data to be a representative sample, you have very strong
statistical evidence that H0 is false. That is, you are confident right .75.
g. You calculate (2.7)2 7.29 7.27. The X 2-test statistic appears to be
(roughly) the square of the z-test statistic. Without rounding the z-test statistic
value as much, z 2.696 and z 2 7.2688, and the two test statistics are
identical.
24
h. The p-values are roughly the same (.0069 .007). Without rounding, the values
are identical.
.3
.25
.2
.15
.1
.05
0
1 2 3 4 5 6 7 8 9
Leading Digit
This graph does appear to be consistent with the Benford probabilities, insofar
as the most common leading digit in this sample is 1 by a wide margin, followed
by 2, and then gradually tapering off from there. Note: Even though the leading
digits are numbers, you can consider this variable to be categorical, because the
number simply separates the countries into groups. That’s why a bar graph is
appropriate rather than a histogram and also why a chi-square goodness-of-fit test
is appropriate.
b. The null hypothesis is H0: 1 .301, 2 .176, 3 .125, 4 .097, 5
.079, 6 0.067, 7 .058, 8 .051, and 9 .046. In other words, the
null hypothesis says the Benford probability model is correct for the data. The
alternative hypothesis simply says this model is not correct, meaning at least one
of the hypothesized Benford probabilities is wrong. In symbols, you can write
Ha: at least one i differs from its hypothesized value.
c. The expected count in the “1” category is 194(.301) 58.394 countries. This
is fairly close to the observed count of 55 countries with a leading digit of 1.
Calculating the expected counts for the other categories in the same way produces
these results:
534 Topic 24: Goodness-of-Fit Tests
1 55 58.394
2 36 34.144
3 22 24.250
4 23 18.818
5 15 15.326
6 14 12.998
7 12 11.252
8 12 9.894
9 5 8.924
All of these expected counts are greater than five, so that technical condition
is satisfied. If you consider the population to be all numbers appearing in
the almanac, you do not literally have a random sample of numbers from the
almanac, but these are likely to be representative of population values.
d. The contribution to the test statistic from the “1” category is
(55 58.394) 2
_____________ 0.197
58.394
Calculating the other contributions in the same way produces these results:
(O E)
________
2
Leading Digit Observed Count Expected Count
E
1 55 58.394 0.197
2 36 34.144 0.100
3 22 24.250 0.209
4 23 18.818 0.929
5 15 15.326 0.007
6 14 12.998 0.077
7 12 11.252 0.050
8 12 9.894 0.448
9 5 8.924 1.725
The test statistic equals 3.742, the sum of the nine values in the last column. (You
might get a slightly different, more accurate answer if you carry more than
three decimal places in your intermediate calculations.) Looking at Table IV
with 8 degrees of freedom (one less than the nine categories), this test statistic
is way off the chart to the left. Therefore, the p-value is much greater than .2.
Activity 24-6 535
24
0 5 10 15 20 25 30
3.742
Chi-Square Values (df = 8)
e. This is a very large p-value, much greater than all common significance levels, so fail
to reject the null hypothesis that Benford’s probabilities adequately model these data.
f. The large p-value (based on the small value of the chi-square test statistic) from
this chi-square analysis reveals the sample data are extremely consistent with the
Benford probabilities. If Benford’s model were correct, it would not be the least
bit surprising to observe sample data as found with these nations’ populations.
In other words, the sample data provide no reason to doubt Benford’s model
describes the distribution of leading digits in the almanac.
g. An accountant or IRS agent could apply a chi-square test to see how well the
numbers in a tax return follow Benford’s probabilities. If the p-value turns out to
be very small, that suggests the leading digits of numbers on that tax return differ
substantially from what Benford’s probabilities predict. This certainly does not
prove the tax return is fraudulent, but it might suggest the numbers were made
up and not legitimate. A tax return whose leading digits do not follow Benford’s
probabilities might be worth a closer look.
3. Fill in the counts as they appear in the two-way table. The expected counts and
results for the chi-square test are updated as you enter the values.
4. (Optional) For more informative labeling of the categories or attributes, edit any
of the blue text or table headings (e.g., “FirstAttribute” or “RowCategory1”) to
better match the categories shown in the original table.
2. Arrow over to the EDIT menu and press Õ to select matrix [A].
25
Note: The table at the beginning of this activity has three categories and three years, so
your matrix needs three rows and three columns (i.e., the dimensions should be 3 3).
3. Now enter the numbers from the table into the matrix, using the arrow keys to
move the cursor.
4. Once matrix [A] is completed, press Õ and arrow over to the TESTS menu and
select 2-Test.
5. Select CALCULATE and then press Õ. The expected values are automatically
entered into the matrix [B].
Solutions
••• In-Class Activities
Activity 25-1: Pursuit of Happiness
a. The observational units are the 4409 American adults who were surveyed.
b. Explanatory: year Type: categorical
Response: happiness level Type: categorical
c. This is an observational study because the researchers did not randomly assign the
subjects to the years.
d. A segmented bar graph is the appropriate graphical display for these data.
e. The proportions of Americans who considered themselves very happy in each of
these three years were very similar (about 30%).
f. The two-sample z-test from Topic 21 is not appropriate for testing whether
these population proportions differ significantly because there are three years to
consider—not just two.
g. H0: 1972 1988 2004, where 1972 represents the proportion of all American
adults who considered themselves happy in 1972, and so on.
h. For the three years combined, the proportion of respondents who were very happy
was 1403/4409 or .318.
554 Topic 25: Inference for Two-Way Tables
l. For the “less than very happy” people in 1988, you calculate (968 999.5)2/
999.5 .993.
m. The test statistic is X 2 5.065.
n. Large values of the test statistic would constitute evidence against the null
hypothesis that the three populations have the same proportions of very happy
Americans because if the proportions are the same, then the observed and
expected values in each year will be very similar, and each year’s contribution
to the test statistic will be small, so the overall test statistic value will be small.
However, if one or more of the proportions is not the same, then there will be a
large difference between the observed and expected values in that year, and this
difference will make a large contribution to the test statistic.
o. Using Table IV with (2 1) (3 1) 2 degrees of freedom, 4.61 5.065
5.99, so .05 p-value .10.
p. These sample data are a bit unlikely to have occurred by chance alone if the
population proportions of very happy people had been identical for these three
years, but you would fail to reject the null hypothesis at the ␣ .05 level.
q. You have weak statistical evidence that the population proportions of very happy
people were not identical for these three years. (Because these were random
samples, you are safe in generalizing this conclusion to the populations of all
American adults in each year.)
d. “1988/not too happy” has the largest contribution (16.927) to the test statistic.
The observed count (136) is less than the expected count (193.18).
“1972/not too happy” has the next largest contribution (13.458) to the test
statistic. The observed count (265) is greater than the expected count (211.63).
e. You have strong statistical evidence that the population distributions of happiness
levels were not the same for all three years. In 1972, the sample proportion of
those who were not too happy was substantially greater than you would expect if
the population proportions were equal in all three years, and in 1988, the sample
proportion of those who were not too happy was substantially less than you would
expect.
25
Response: whether the customer left a tip Type: binary categorical
c. This is an experiment because the waiter randomly assigned the customers to
treatment groups.
d. This study involves random assignment to groups.
e. H0: The population proportions of potential customers who would leave a tip is
the same regardless of the type of card they receive.
Ha: The population proportion of potential coffee bar customers who would leave
a tip is affected by the type of card they receive.
Using Minitab, the test statistic is X 2 9.953 with 2 degrees of freedom and the
p-value is .007.
As the p-value is .007, which is less than .10, reject H0 at the .10 significance level.
The “joke/left a tip” cell makes the largest contribution to the test statistic, and
the observed count is much greater than the expected count. So, you conclude
that potential coffee bar customers are not equally likely to leave a tip regardless
of the type of card they receive—leaving a joke card is more likely to generate a
tip than leaving an advertisement card or no card.
f. Because this was an experiment where the only difference between the groups
should be the type of card the customers received, you can conclude that the
joke card caused customers to be more likely to leave a tip. (Note also that the
technical conditions for this test to be valid were satisfied—the subjects were
randomly assigned to treatment groups, and the expected counts in each cell
of the table were at least five.)
g. You should not generalize your conclusion to all waiters and waitress in all food
and drink establishments. This experiment was carried out in only one coffee bar
with one waiter.
b. In Activities 25-1 and 25-2, the researchers took independent random samples
from three populations (1972, 1988, and 2004). In this activity, the researchers
took one random sample and then classified the subjects by both explanatory and
response variables.
c. H0: Political inclination and self-reported happiness level are independent in the
population of adult Americans.
Ha: Political inclination is related to self-reported happiness level.
Here is a bar graph:
60
50
40
30
20
10
0
Liberal Moderate Conservative
Political Inclination
Using Minitab, the test statistic is X 2 22.454 with 4 degrees of freedom and
p-value .000.
With such a small p-value, reject H0 and conclude political inclination and
self-reported happiness level are related in the population of adult Americans.
The largest contributions to the test statistic are made by the “conservative/very
happy” cell, in which the observed count (187) is greater than the expected
count (152.88), and the “conservative/not too happy” cell, in which the observed
count (48) is less than the expected count (66.06).
H0: Type of pet and feelings of closeness to pet are independent in the population
of adult Americans.
Ha: Type of pet is related to feelings of closeness to pet.
Technical conditions: The observations arose from a random sample of the
population, and the expected counts are at least five for each cell in the table.
Activity 25-6 557
Using Minitab, the test statistic is X 2 49.63 with 1 degree of freedom and the
p-value .000.
With such a small p-value, reject H0 and conclude type of pet and feeling close to
the pet are related in the population of adult Americans. The largest contribution
to the test statistic is made by the “cat/does not feel close to pet” cell, in which the
observed count (110) is greater than the expected count (66.57).
b. The null hypothesis is that the proportion of the population of all dog owners
who feel close to their pets is the same as the proportion of all cat owners who feel
close to their pets. In symbols, the null hypothesis is H0: dog cat.
The alternative hypothesis is the proportion of the population of all dog owners
who feel close to their pets is not the same as the proportion of all cat owners who
feel close to their pets. In symbols, the alternative hypothesis is Ha: dog cat.
Technical conditions: The number of successes and failures in each group is at
least five, and the data were randomly selected.
25
.93988 .83988
The test statistic is z ___________________________
__________________________ 7.045.
1 ____
.9031(1 .9031) _____ 1
√ 1181 687
p-value 2 Pr(Z 7.045) .0000
With such a small p-value, reject H0 at any commonly used significance level.
You have very strong statistical evidence that the population proportion of dog
owners who feel close to their pets is not the same as the population proportion
of cat owners who feel close to their pets.
c. The p-values are the same in both cases.
d. (7.045) 2 49.63; 49.63 X 2
60
Every day
50
40
30
20
10
0
Male Female
Gender
558 Topic 25: Inference for Two-Way Tables
This graph reveals that a higher proportion of men than women read the
newspaper every day (191/422 .453 for men vs. 167/484 .345 for women). A
higher proportion of women than men read the paper a few times a week (.211 for
men vs. .275 for women) and once a week (.126 for men vs. .167 for women).
c. The null hypothesis asserts newspaper reading frequency is independent of
gender. In other words, the null hypothesis says the proportional distribution
of the various reading frequency categories is identical for men and women
in the population of all American adults. The alternative hypothesis says
newspaper reading frequency is related to gender, which means these population
distributions are not the same for men and women.
The expected counts appear in parentheses in the table:
For an example of one of these calculations, the expected count for the “every
day/male” cell is found by taking (358)(422)/906 166.75.
Technical conditions: All of these expected counts are greater than five (smallest
33.07). The subjects were randomly selected from the population of adult
Americans, so the technical conditions are satisfied.
The chi-square test statistic, computed as
(O E )2
________
E
turns out to be
X 2 3.526 3.075
2.006 1.749
1.420 1.238
0.002 0.002
0.000 0.000 13.020
Comparing this value to a chi-square distribution with (5 1)(2 1)
4 degrees of freedom reveals that the p-value is slightly greater than .01 (13.02 is
just below 13.28 in Table IV). This is less than .05, so the test decision is to reject
the null hypothesis at the .05 significance level. The data provide fairly strong
evidence that newspaper reading frequency is related to gender.
In particular, the table cells in the top row contribute the most to the test statistic
calculation. This finding indicates more men than expected read the newspaper
Activity 25-7 559
every day and fewer women than expected read the newspaper every day. The
sample proportions of men and women who read the newspaper every day are
191/422 .453 for men and 167/484 .345 for women.
25
40
30
20
10
0
Liberal Moderate Conservative
Political Inclination
This graph reveals that a very high proportion (.819) of liberals in this sample feel
the govenernment is spending too little on the envirnoment.
b. No; although the observed count in this category is only 1, it is the expected
count that needs to be at least five in order for the chi-square test to be valid.
c. A chi-square test in this setting would be one of independence because the data
were collected by one random sample from a population and then classified by
both the independent and response variables.
d. H0: Political affiliation and feelings about government spending on the
environment are independent variables in the population of adult Americans.
Ha: Political affiliation and feelings about government spending on the
environment are related.
Technical conditions: The observations arose from a random sample of the
population, and the expected counts are at least five for each cell in the table.
Using Minitab, the test statistic is X 2 52.117 with 4 degrees of freedom and
p-value .000. Using Table IV, 52.117 20.00, so p-value .0005.
With a p-value less than .025, reject H0 at the .025 significance level. You can
conclude that political affiliation and feelings about government spending on the
environment are related in the population of adult Americans.
e. The largest contributions to the test statistic are made by the “liberal/too little
cell,” in which the observed count (127) is greater than the expected count (95.5);
the “conservative/too much” cell, in which the observed count (32) is also greater
than the expected count (18.27); and the “liberal/too much” cell, in which the
observed count (1) is less than the expected count (12).
Activity 26-1 591
3. Press q and then Æ to display the graph (do not forget to turn off the other plots).
To create a labeled scatterplot, which is necessary for part b, you have two choices.
Choice 1:
a. Divide these data between men and women (four lists total: two for men and
two for women).
b. Use Plot1 and Plot2 simultaneously using different symbols.
Choice 2:
c. Create a new list (call it GNDR for gender) and enter a 0 if the corresponding
data is from a male student and a 1 if the data is from a female student (your
calculator only accepts numeric values in lists).
d. Download the program named LBLSCAT and use SPAN as the X-LIST, HGHT as
the Y-LIST, and GNDR as the CATEGORICAL VARIABLE LIST. You will need to type
in names for the different categories (i.e., you could use M for male and F for
female).
Solutions
26
••• In-Class Activities
Activity 26-1: House Prices
a. The observational units are the houses.
b. Two variables reported: price and size. Both are quantitative variables.
c. Yes, the data suggest that bigger houses tend to cost more than smaller ones, as
the house sizes tend to increase down the list as well.
d. Bottom left: 2130 Beach St. Top right: 833 Creekside Dr.
e. The point corresponding to the house at 845 Pearl Drive is circled on the
scatterplot below:
650
Price (in thousands of dollars)
600
550
500
450
400
350
300
f. Yes, the scatterplot reveals a positive relationship between a house’s size and price.
The upward trend of the graph indicates that, in general, the larger the house, the
greater the price.
g. Direction: positive
Strength: moderate
Form: linear
h. Answers will vary. One example is 2545 Lancaster Dr. (1030 sq ft, $344,720) and
415 Golden West Pl. (883 sq ft, $359,500). (The smaller house has the greater
price.)
Letter of Scatterplot D G A H C E I F B
75
65
60
55
17 18 19 20 21 22 23
Handspan (in inches)
26
Gender 75
Female
Male
Height (in inches)
70
65
60
55
17 18 19 20 21 22 23
Handspan (in inches)
Yes, men and women tend to differ with regard to handspan. Men tend to have
longer handspans than women and less variability in their handspans. This result
is indicated by the red points clustered fairly close together on the upper end of
the horizontal scale, whereas the black points are spread widely on the lower to
middle end of the horizontal scale.
Yes, men and women tend to differ with regard to height. Most of the men in this
class are taller than most of the women, although there are some fairly tall women
in this class. You can tell because the black points (women) range from the low to
the highest end of the vertical scale, whereas the red points are primarily clustered
toward the upper end of the vertical scale.
Yes, the association between handspan and height appears to be different between
men and women. There is a much stronger linear relationship between these
variables for the women than there is for the men. There is actually a very weak,
positive association been height and handspan for the men.
c. The scatterplot of height vs. foot length follows:
594 Topic 26: Graphical Displays of Association
75
65
60
55
19 20 21 22 23 24 25 26 27 28
Foot Length (in inches)
There is a moderately strong, positive, linear relationship between height and foot
length. The scatterplot shows one observation (21 in, 55 cm) that does not seem to
fit the overall trend. In general, the taller the person, the longer her or his feet.
The labeled scatterplot of height vs. foot length follows:
Gender 75
Female
Male
Height (in inches)
70
65
60
55
19 20 21 22 23 24 25 26 27 28
Foot Length (in inches)
As with handspans, most of the men’s feet are longer than most of the women’s
feet. However, there is one man of medium height with (relatively) short feet.
Both genders display a moderate, positive, linear association between height and
foot length, although the association is somewhat stronger among the women.
80
Life Expectancy (in years)
70
60
50
40
26
70
60
50
40
30
20
10
10 20 30 40 50 60 70 80 90 100
Maintenance Outsourced (percentage)
c. The scatterplot reveals a fairly strong positive association between these variables.
Airlines that outsource a higher percentage of their maintenance tend to have a
greater percentage of delays caused by the airline, as compared to airlines that
outsource a smaller percentage of their maintenance. The form of the association
is clearly nonlinear, as the relationship appears to follow a curved pattern.
d. No, you cannot conclude from these data that outsourcing maintenance causes
airlines to experience more delays. This is an observational study, not an
experiment. Many other variables could explain the observed association between
outsourcing and delays. For example, perhaps less financially successful airlines
tend to outsource more of their maintenance and also to have more delays.
Solutions
••• In-Class Activities
27
Activity 27-1: Car Data
Letter of D G A H C E I F B
Scatterplot
Correlation
.907 .721 .450 .244 .081 .235 .510 .889 .994
Coefficient
180,000
140,000
120,000
100,000
80,000 Hawaii
60,000
The scatterplot shows a weak positive linear association with one outlier (Hawaii).
c. Answers will vary by student guess.
d. The correlation coefficient is .334.
e. Yes, one of the states appears to be unusual. Hawaii has an unusually large
median housing value ($272,700) and the salary of Hawaii’s governor is in the
lower half of the distribution.
f. The new correlation coefficient is .576. The revised scatterplot follows:
200,000
Governor’s Salary (in dollars)
180,000 Hawaii
160,000
140,000
120,000
100,000
80,000
60,000
300,000
Governor’s Salary (in dollars)
Hawaii
250,000
200,000
150,000
100,000
200,000
100,000
50,000
Hawaii
0
27
b. Answers will vary by student guess.
c. The correlation coefficient is .743.
d. Yes, the value of the correlation coefficient is fairly high even though the
association between the variables is not linear.
e. No, the fairly high value of the correlation coefficient is not evidence of a cause-
and-effect relationship between the variables. There are many confounding
variables that could explain the association, as discussed in Activity 26-6.
Repetition Number 1 2 3 4 5 6 7 8 9 10
Your Guess .6 .7 .7 .8 .85 .7 .15 .9 .85 .1
Actual Value .69 .87 .75 .87 .76 .74 .11 .84 .95 .37
.8
.6
.4
.2
Guess
0
.2
.4
.6
.8
1
1 .8 .6 .4 .2 0 .2 .4 .6 .8 1
Actual (Correlation : .99)
The correlation coefficient r is .99. Most students will be surprised that their
guesses were so consistent.
e. The scatterplot of your errors vs. the actual values follows:
.8
.6
.4
.2
Error
.2
.4
.6
.8
1
1 .8 .6 .4 .2 0 .2 .4 .6 .8 1
Actual
.8
.6
.4
.2
Error
0
.2
.4
.6
.8
1
0 4 8 12
Trial
In this example, there is no real evidence that the accuracy of the guesses changed
over time.
g. The correlation would be 1.00.
h. The correlation would be 1.00 in this case also.
i. No; if the correlation is 1.00, it does not mean that you guessed perfectly
27
every time. It means your guesses are consistent: either consistently correct or
consistently incorrect by the same amount and in the same direction. You cannot
use the correlation coefficient to tell whether you guess correctly, but you can use
it to tell whether your guesses change appropriately with the size and direction of
the actual correlation.
100
90
Exam 2 Score
80
70
60
50
70 75 80 85 90 95 100
Exam 1 Score
100
90
Exam 2 Score
80
70
60
50
60 65 70 75 80 85 90
New Exam 1 Score
f. The scatterplot of new exam 2 score vs. new exam 1 score follows:
200
190
180
70 75 80 85 90 95 100
New Exam 1 Score
100
(hypothetical exam 1 score ⴙ 10)
Hypothetical Exam 2 Score
90
27
80
70
60
50 60 70 80 90
Hypothetical Exam 1 Score
90
(twice hypothetical exam 1 score)
Hypothetical Exam 2 Score
80
70
60
50
25 30 35 40 45
Hypothetical Exam 1 Score
350
300
250
Draft Number
200
150
100
50
0
0 50 100 150 200 250 300 350
Sequential Date
350
300
250
Draft Number
200
150
100
50
0
0 50 100 150 200 250 300 350
Sequential Date
Activity 27-9 619
The correlation coefficient is .014, which is very close to 0. This value indicates
there is no evidence of association between draft number and sequential date,
suggesting the lottery process was fair and random in 1971. The mixing
mechanism was greatly improved after the anomaly with the 1970 results was
spotted.
27
e. In class C, students who scored less than 50 on exam 1 also scored less than
50 on exam 2, although some scores increased and some decreased. Students who
scored 70 and greater on exam 1 also performed this well on exam 2, but some
scores increased and some decreased. There were no students who scored
between 50 and 70 on either exam.
f. The correlation coefficient is .954. This value is deceptively high. The two
clusters of scores, which have no particular pattern within each cluster, form a
line (any two points will form a line) that causes an artificially inflated overall
correlation coefficient.
Solutions
••• In-Class Activities
Activity 28-1: Heights, Handspans, and Foot Lengths
a. There is a moderate positive linear association between height and foot length in
this scatterplot. Answers will vary by student.
The following is one representative set of answers.
ˆht 38.293 1.048 foot size.
b. The equation reported is (predicted) heig
c. No, everyone in the class did not obtain this same line.
d. Answers will vary.
e. The SAE value is 56.46. Yes, someone had a smaller value (55.11).
Activity 28-1 635
ˆht 37.503 1.074 foot size; the SSE is 238.19. Yes, someone
g. The equation is heig
had a smaller value (235.28). A graph follows:
28
636 Topic 28: Least Squares Regression
i. Using the least squares regression line, predicted height is 38.302 1.033 28
or 67.226 inches. Yes, this prediction seems reasonable based on the scatterplot.
j. Using the least squares regression line, predicted height is 38.302 1.033 29
or 68.259 inches.
k. These predictions differ by 1.033 inches, which is the slope of the least squares line.
l. The least squares line would predict a height of 38.302 inches for a person with a
0 cm foot length. No, this does not make any sense because a foot length of 0 cm
is outside the range of plausible values.
m. Using the least squares regression line, predicted height is 38.302 1.033
45 or 84.79 inches, which is more than 7 feet tall. You should not consider this
prediction as reliable as the one for a person with a 28 cm foot length, because for
this prediction (84.79 in) you do not have much data for someone more than
7 feet tall. You should not feel comfortable assuming the same relationship will
hold for such extreme observations.
n. No, the least squares regression line does not move much at all when you change
this student’s height.
o. Changing the height of the student with the shortest or longest foot has a much
more drastic effect on the regression line. Points on the end of the regression line
are clearly more influential than points near the middle.
_
p. SSE (y ) 475.75
q. Here is the percentage change in the SSE:
475.75 235 50.6%
100% ____________
475.75
Activity 28-2 637
^
Price ⴝ 265,222 ⴙ 168.6 Size
650,000
600,000
28
Price (in dollars)
550,000
500,000
450,000
400,000
350,000 Greatest negative residual
300,000
The house with the greatest negative residual is located at 2545 Lancaster Drive.
i. Each increase of one square foot in the house size increases the predicted house
price by $168.60.
j. You calculate 100 $168.60 or $16,860.
k. The predicted price for a house with an area of 0 square feet is $265,222. This
value makes no sense in this context because there is no such thing as a house
with zero area.
l. The percentage of variability in house prices explained by the least squares
regression line with size is 60.8% as r2 correlation coefficient2 (.78)2.
m. No, it would not be reasonable to use this regression line to predict the price of a
3500 ft2 house because the given house sizes are from about 500–2000 ft2. This
638 Topic 28: Least Squares Regression
would be extrapolation; you have no idea whether the given association continues
beyond this range of house sizes.
n. No, it would not be reasonable to use this regression line to predict the price of a
2000 ft2 house in Canton, New York. This regression line was created with data
collected solely from one region in California. The housing market in New York
is very different, and you cannot expect this regression line to make accurate
predictions for any market other than Arroyo Grande, California.
0.50
0.25
0.00
Residual
–0.25
–0.50
–0.75
–1.00
Yes, the residual plot reveals a pattern that suggests the linear model is not
appropriate.
e. The scatterplot of trot speed vs. log(body mass) follows:
^
Trot Speed ⴝ 0.8209 ⴙ 0.4637 Log(Body Mass)
Trot Speed (in meters per second)
2.0
1.5
1.0
0.5
0.0
–2 –1 0 1 2
Log(Body Mass)
Yes, the association between trot speed and log10(body mass) appears to be fairly
linear.
Activity 28-4 639
0.3
0.2
0.1
0.0
Residual
–0.1
–0.2
–0.3
–0.4
–0.5
–0.6
Yes, the points in this plot are fairly randomly scattered, indicating the linear
model is reasonable for the transformed data.
h. 10 kilograms: predicted trot speed 0.8209 0.4637 log(10) 1.2846 m/sec
100 kilograms: predicted trot speed 0.8209 0.4637 log(100) 1.7483 m/sec
i. The difference between these predictions is 0.4637 m/sec slope of regression
line.
28
b. These scatterplots follow:
180
160
140
Price (in dollars)
120
100
80
60
40
20
0
180
160
140
The scatterplot of price vs. pages reveals a fairly strong positive linear association.
The scatterplot of price vs. year also indicates a positive association, but the
association is much weaker and not very linear. Two unusual textbooks from
the early 1970s, with very low prices, appear to be outliers, because they differ
substantially from the pattern of the other textbooks, and potentially influential
observations.
c. Number of pages appears to be a much better predictor of price than year. The
relationship is much stronger and also more linear.
d. The equation of this least squares line is predicted price 3.42 0.147 pages. It
is shown on the scatterplot here:
180
160
140
Price (in dollars)
120
100
80
60
40
20
0
0 200 400 600 800 1000 1200
Pages
75
50
25
Residual
0
25
50
This residual plot reveals no obvious pattern, which suggests the least squares line
is a reasonable model for the relationship between price and pages.
i. The equation of this least squares line is predicted price 4969 2.516 year,
with r 2 .186. The least squares line is shown on the scatterplot here:
180
160
140
Price (in dollars)
120
100
80
60
40
28
20
0
100
50
Residual
50
This plot reveals that the middle years (19902000) tend to have negative
residuals. This pattern suggests a linear model is not very appropriate for
predicting price from year of publication. A transformation might lead to a more
appropriate linear model.
670 Topic 29: Inference for Correlation and Regression
Solutions
••• In-Class Activities
Activity 29-1: Studying and Grades
a. Observational units: students at UOP
Explanatory variable: study hours per week Type: quantitative
Response variable: GPA Type: quantitative
b. The scatterplot of GPA vs. study hours follows:
4.0
3.5
GPA
3.0
2.5
2.0
0 1 2 3 4 5 6 7 8 9
Study Hours per Week
–1.12 –0.1 –0.08 –0.06 –0.04 –0.02 0.0 0.02 0.04 0.06 0.08 0.1 0.12
0.0
29
would happen by random chance alone less than .1% of the time if there were no
association between GPA and hours studied in the population.
r. Using t* 2.000 with 60 degrees of freedom, a 95% CI for  is
0.0894 2 (0.02771), which is (0.034, 0.145).
s. You are 95% confident the slope of the population regression line between GPA
and hours studied is somewhere between 0.034 and 0.145 points/hour.
t. You have strong statistical evidence of a positive association between GPA and
hours studied at UOP. You are 95% confident that an additional hour of study
corresponds to an average increase of 0.034 to 0.145 points in overall student
GPA. You are not necessarily drawing a cause-and-effect relationship, however, as
this is an observational study and not an experiment.
672 Topic 29: Inference for Correlation and Regression
99.9
99
95
90
14 80
70
Percentage
12 60
10 50
40
Frequency
8 30
6 20
4 10
2 5
0
1
–0.8 –0.4 0.0 0.4 0.8
Residual
0.1
–2 –1 0 1 2
Residual
The distribution of residuals appears roughly normal. These plots do not reveal
any marked features suggesting nonnormality.
b. The scatterplot of residual vs. study hours follows:
1.0
0.5
Residual
0.0
–0.5
–1.0
0 1 2 3 4 5 6 7 8 9
Study Hours per Week
This residual plot does not reveal a strong pattern (curvature). The variability of
the residuals does not differ substantially at various x-values (although it is a little
different at the lowest and highest x-values than it is at the middle x-values).
The alternative hypothesis is that larger houses in this population tend to have
greater purchase prices, or the slope of the population regression line between
these two variables is positive. In symbols, Ha:  0.
168.6 5.29.
The test statistic is t _____
31.88
Using Table III with 18 degrees of freedom, p-value .0005. Using Minitab, the
p-value is .000.
With the small p-value, reject H0 at any commonly used significance level, and
conclude there is very strong statistical evidence of a positive linear association
between house size and price in this population.
The residual plot follows:
100,000
50,000
Residual
–50,000
–100,000
A residual plot does not reveal any strong curvature and shows that the standard
deviations of the y-values can be reasonably considered the same at each x-value.
The data are from a simple random sample, so technical conditions 1, 2, and 4 are
met. A histogram and probability plot indicate the residuals are not beautifully
normal, but the nonnormality is probably not strong enough to convince you that
technical condition 3 has been violated.
6 99
5
29
95
Frequency
4
90
3
80
2
70
Percentage
1 60
0 50
40
0 30
0
00
00
00
0
0
00
00
00
00
,0
,0
,0
,0
20
5,
0,
5,
25
75
50
00
–7
–5
–2
–1
10
Residual
5
1
–200,000 –100,000 0 100,000 200,000
Residual
The value zero is not in this interval, which is consistent with the test result in
part b where you found that zero is not a plausible value for the population slope
coefficient.
180
160
140
Price (in dollars)
120
100
80
60
40
20
0
The equation of this line is predicted price 3.42 0.147 (pages). The value
of r 2 is .677, indicating that 67.7% of the variability in textbook prices can be
explained by the number of pages in the texts.
b. The slope coefficient is b 0.147, and technology reports the standard error to be
0.019. The slope indicates that the predicted price increases by $0.147 on average
for each additional page in the book. The standard error measures the variability
in the sample slopes from repeated random samples of 30 textbooks from this
population.
c. The population of interest in this study is all textbooks in the Cal Poly Bookstore
in November 2006.
d. Let  represent the slope of the least squares line for predicting textbook price
from number of pages in the population. Then the null hypothesis is H0:  0,
meaning there is no association between textbook price and number of pages in the
population. The alternative is Ha:  0, meaning there is a positive association
between these variables in the population. The test statistic is
0.147 7.74
b _____
t _____
SE(b) 0.019
(Without rounding, technology reports the test statistic to be 7.65.) Comparing
this to the t-distribution using Table III with 30 2 or 28 degrees of freedom
reveals that the p-value is much less than .0005. Such a small p-value provides
extremely strong evidence of a positive relationship between a textbook’s price
and its number of pages in the population of all textbooks in that bookstore at
that time.
e. Using Table III, the t* critical value for 90% confidence with 28 degrees of
freedom is 1.701. A 90% confidence interval for the population slope  is
0.147 1.701(0.019), which is 0.147 0.032, which is the interval from 0.115
through 0.179. You can be 95% confident that in the population of all textbooks
in that bookstore, the predicted price of a textbook increases between 11.5 and
17.9 cents for each additional page.
f. A 99% confidence interval for the population slope  has the same midpoint,
29
namely the sample slope 0.147. But the 99% interval is wider than the 90%
interval in order to achieve higher confidence of capturing the population slope
coefficient.
g. Technical conditions: First, Shaffer and Kaplan did take a random sample of
textbooks, so that condition is satisfied as long as you restrict your population to
the Cal Poly Bookstore in November 2006. To check the normality condition,
consider the following histogram and normal probability plot of the residuals.
7
6
5
Frequency
4
3
2
1
0
80 40 0 40 80
Residual
676 Topic 29: Inference for Correlation and Regression
99
95
90
80
Percentage
70
60
50
40
30
20
10
5
1
150 100 50 0 50 100 150
Residual
75
50
25
Residual
25
50