Variability in Quiz Scores by Class
Variability in Quiz Scores by Class
How many hours you have slept in the past 24 hours Q N/A
Dot
Whether you have slept for at least 7 hours in the
C Yes
past 24 hours Bar
1
Watch Out!
Answer the following questions based on the Watch Out on page 6 in your textbook.
1. It is important to distinguish between an observational unit and variable. It is helpful to fill in the blanks for the
sentence below with the vocabulary mentioned above:
Suppose your teacher is interested in how many text messages each student in this class has sent today. Complete the
sentence below for this particular scenario:
Is the summary number, which would be the average number of text messages for the class, the variable in this
scenario? Explain. (Hint: Does “average number of text messages” represent a question that can be asked of each
observational unit?) No. We wish to focus on the texts of an individual student. The summary number, or average
number of texts, is not considered a variable from student to student. You don’t ask one student what their
average number of text messages sent is; you ask one student for a total number and then you calculate the
average from all the students’ numbers.
2
Activity 2-3
a. This is a quantitative variable.
b. Answers will vary, but here is an example:
Watch Out!
Read the Watch Out on page 18 and answer below:
1. What does each dot in a dot plot represent? A different observational unit
2. What type of variable is displayed in dot plots? Quantitative
What type of variable is displayed in bar graphs? Categorical
Now read page 21 in the textbook. When are dotplots and bar graphs most illuminating? When used to compare the
distribution of a variable between two or more groups.
b. The type of law variable is categorical and binary (although New Hampshire has no law so if
you didn’t say binary that’s fine). The percentage usage variable is quantitative.
c. Answers will vary. The typical usage percentage for a primary-type seatbelt law state is about
87%; for states with a secondary-type law the typical usage percentage appears to be about
80%.
d. No, a state with a primary law does not always have a higher usage percentage than a state
with a secondary law; for example, Tennessee has a primary law, but its usage percentage is
80.2%, whereas West Virginia has a secondary law but a higher usage percentage of 89.6%.
e. Yes, states with a primary law tend to have higher usage percentages than states with a
secondary law. You see this in the dotplots because most of the dots for the primary law states
are clustered at the high percentage values (from 85-97%), whereas most of the secondary law
states have percentages that are less than 83%.
3
f. Yes, the data seem to support the contention that tougher laws lead to more seatbelt usage, but
you cannot draw a definite cause-and-effect conclusion. There might be hidden or confounding
variables that explain the association between primary seatbelt laws and increased seatbelt
usage. For example, perhaps states with primary laws also tend to have lower highway speed
limits than states with secondary laws. This is not necessarily a cause-and-effect situation.
Activity 2-5 Key: February Temperatures
a. San Luis Obispo tended to have the highest temperatures that month (its temperatures are all
clustered from 50°F–90°F, with a significant chunk of the temperatures falling above 75°F), and
Lincoln tended to have the lowest temperatures, (with many temperatures below 45°F), whereas
Sedona‘s temperatures ranged from about 48°F–68°F.
b. Sedona had the most day-to-day consistency in its high temperatures that month (the
temperatures all stayed between 48-68°F), and Lincoln had the least consistency because its
temperatures ranged from roughly 10°F all the way to about 75°F.
Tendency vs. Consistency
Now read page 22 in the textbook.
1. Define tendency and consistency in your own words.
Tendency refers to the center of the distribution; consistency (or variability) to the spread in the distribution. You
should always look for both.
2. State the difference, in your own words, between tendency and consistency.
Tendency is what a typical observational unit is and consistency is how likely a typical observational unit will be that.
b. Answers will vary from class to class, but here are some sample answers.
c. Yes; seven was chosen more often than any other value.
d. Twenty-nine students (29/35, or 0.829) gave a response greater than 5; One student (1/35, or 0.029) gave a response
less than 5.
e. The vast majority of this class, (more than 80%), feel that statistics is important to society. If fact, more than 65% of the
class feel that statistics is very important to society. About 14% of the class is neutral about the important of statistics, and
only 1 of the 35 students in this group believe that statistics is unimportant to society.
4
Exercise 2-10: Value of Statistics
i. Class C
ii. Class D
iii. Class E
iv. Class A
v. Class B
a. These are all quantitative variables, with the exception of the football jerseys.
3. Variable: G. Annual snowfall amounts for a sample of cities around the U.S.
Explanation: Most cities have little or no snow, but some have quite a bit.
4. Variable: C. Jersey numbers of Cal Poly State University football players in 2006
Explanation: There are virtually no repeats and the distribution covers almost all the values from 1–
99.
7. Variable: F. Quiz percentages for a class of statistics students (quizzes were quite straightforward for most students)
Explanation: Most of the values are high (near 80–90) with a few very low outliers.
c. Dotplots 1 and 3 have a similar shape (they are skewed to the right) whereas dotplots 6 and 7 are
skewed to the left with one low outlier.
Then answer the following questions using the information in the textbook in and around Activity 7-1 (read pp 129-130):
1. What is the most important aspect of a set of data to notice and describe? Center of distribution.
2. What is another term for variability? Spread or consistency
3. What is the second important feature that should be discussed? Spread
5
5. If a distribution’s left side is roughly a mirror image of the right side, what shape is it
said to have? Sketch an example. Symmetric.
6.
7. If the tail of a distribution extends to the larger
values on the right side of the axis, what shape is it
said to have? Sketch an example. Skewed right.
8. If the tail of a distribution extends to the smaller values on the left side of the axis,
what shape is it said to have? Sketch an example. Skewed left.
10. What are data values that differ markedly from the pattern established by the vast majority of the data? Outliers
11. What is the arithmetic average and how is it found? It is found by adding up the values for each of the
observational units and dividing by the number of values. It is the balance point of the distribution.
12. What is the value of the middle observation and how is it found? Median. Once the n values have been
arranged in numerical order, the median of an odd number of data values is located in position (n + 1)/2.
The median of an even number of values is defined to be the arithmetic (mean) of the middle two values, on
either side of position (n + 1)/2.
13. Can mean and median be applied to categorical variables? No.
14. Can mean or median be parameters or statistics or both? Either, depending on whether the data constitute a
population or a sample. What symbol is used to represent a population mean (parameter)? µ What symbol is
used to represent a sample mean (statistic)? x
17. What type of measure is relatively unaffected by the presence of outliers in a distribution? a resistant measure
Watch Out!
Turn to page 160 and answer the questions based on the Watch Out.
19. Is center a property or measure? Property. Mean and median are measures of center.
20. What two ways do we measure center? Mean and median.
Which is preferable? Neither; both have their own properties and strengths.
21. What features are not indicated by center? Spread, shape, clusters, and outliers.
6
Activities 7-2, 8-3 & 10-21 Rowers’ Weights
a. Give the name of the team member whose weight makes him an apparent outlier. Suggest an explanation for why his
weight differs so substantially from the others.
McElhenney, the coxswain, is the apparent outlier at 121 lbs. His job is to call out the cadence to the rowers; he
does not row himself. It is important that he not add excess weight to the boat. (Mentioned in part d.)
b. Suggest an explanation for the clusters and gaps apparent in the distribution. (Hint: Consider the events in which the
cluster of less heavy rowers competed.)
There is a cluster of four rowers whose weights seem to be around 160 lbs. These rowers are all involved in LW or
“lightweight” events which require them to weigh below a certain amount on race day. The remaining rowers in
the upper cluster of weights are not involved in the lightweight events, and therefore have no upper limit on their
weights.
c. Download the list RWGHT linked above and find the mean and median. Record them in the first column of the table
below. Why does it make sense that the mean and median differ in this direction?
Explanation: The distribution is skewed left with an extreme low outlier (the coxswain) so it makes sense that the
mean will be less than the median.
Whole Team Without Coxswain With Max Weight at 329 With Max Weight at 2229
Mean 197.96 201.17 201.96 277.96
Median 205 207 205 205
d. McElhenney, the coxswain, is the team member who calls out rowing instructions and keep the rowers in sync.
Confirm algebraically that he is an outlier. Are there any additional outliers?
1.5(217 – 187.5) = 44.25; 187.5 – 44.25 =143.25 McElhenney is a lower outlier
1.5(217 – 187.5) = 44.25; 217 + 44.25 =261.25 There are no upper outliers
e. Predict what will happen to the mean if you remove the coxswain from this analysis. Also predict what will happen to
the median. Explain. Answers will vary, but removing the lower outlier will mean that for the remaining rowers
the average weight should be higher. Removing the minimum data value will move the median up to the next value
or cause it to be the average of the next two values; the median could stay the same numerical value depending on
the values though.
f. Remove McElhenney from your list and recalculate the mean and median. Record the values in the second column of
the table. Did these values change as you had predicted? Answers will vary.
g. Now consider what would happen if you put McElhenney back in and increased the weight of the heaviest rower.
How would this change the mean and median? Explain. Answers will vary, but making the heaviest rower weigh
more should bring up the average weight for every rower. Moving the maximum value up higher will not change
the median because the middle values are still the same/in the same position.
h. Increase Newlin’s weight from 229 to 329. Recalculate the mean and median and record them in the third column of
the table. Did they change as you predicted? Answers will vary.
i. Suppose the typist’s finger slipped when recording Newlin’s weight and entered 2229. Predict how this error will
affect the mean and median. Answers will vary, but the mean should increase a lot more while the median should
still stay the same.
j. Increase Newlin’s weight to 2229 and recalculate the mean and median. Record them in the last column of the table.
k. A measure whose value is relatively unaffected by the presence of outliers in a distribution is said to be resistant.
Based on these calculations, is the mean and/or median resistant? Explain why this conclusion makes sense; base your
argument on how each measure is calculated.
The median is resistant. This conclusion makes sense because to find its value you only use the middle position in
the data, not the ends where outliers would occur. The mean is not resistant, which also makes sense - to find the
mean you use the numerical values of all of the data. Outliers will pull the mean toward them – potentially
affecting the value of the mean drastically. The only reason the median changed in the second column is that the
number of values changed when we removed McElhenney.
l. In principle, is there any limit to how much the mean can increase or decrease simply by changing one of the values in
the distribution?
The mean only has to stay between the minimum and maximum values in the dataset. (And so, if one of these
values changes without bound, so does the mean.)
7
m. Create a boxplot of the data and reproduce the boxplot below. Be sure to include appropriate labels.
Notes on how the paragraph is written: The context is that this is a distribution of weights of rowers in pounds and then
the measurement units, pounds, are included throughout. Most of the data is on the right, so the shape is left-skewed
which is one reason to use resistant measures to describe the distribution. There is also an outlier, which is another reason
to use resistant measures. Outlier calculations are included to show how we know there is an outlier; we can’t just see it
on the graph. Once the shape and outliers are described, a measure of center should be picked and its value should be
given with measurement units; the median is used here since it is resistant. A measure of spread should also be picked
and its value should be given with measurement units; the IQR is used here since it is resistant. At the end, sum up what
has been learned about a typical observational unit in the distribution or make a comment about something else that would
be useful to know to conclude your paragraph.
Watch Out!
Answer the following questions from the Watch Out on page 202.
1. Do you know how many data points are in each quarter of a boxplot?
No, unless you know the sample size.
Turn to page 133 in your book and examine the chart in order to answer parts a, b and c.
a. Consider the western states. Would it be reason able to use 0.1 as the leaf unit? Explain. (Hint: How many stems
would there be? How many leaves would be on each stem?)
No, it would not be reasonable to use 0.1 as the leaf unit because this would require 33 stems and
many of the stems would have one or no leaves on them. No stem would have more than three leaves.
b. Still considering the western states, would it be reasonable to use 1 as the leaf unit? Explain.
It would be more reasonable to use 1 as the leaf unit, but this would create only four stems and the “0” stem would
contain too many leaves.
c. If we use 1 as the leaf unit and create a split stemplot for the western states, how many stems will we have?
If we create a split stemplot for the western states, we will have seven stems. Two each for 0, 1, 2 and only one for 3
since there are no values above 35.
d. Complete the following side-by-side stemplot to compare the distributions of population growth between eastern and
western states by filling in the states beginning with the letter N. [Hints: Use the stems provided here. An underscore (_)
8
indicates where you should fill in a leaf value. Use single digit stems by truncating the tenths digit after the decimal point.
Alaska’s percentage appears as 1|1 ignoring the 0.4.]
e. Examine the chart again and calculate the median percentage change for the eastern and western states. Comment on
how the medians compare.
For the western states, the median is 8.6%. For the eastern states, the median is 4.7%
Truncating (using the numbers in the stemplot) digits: For the western states, the median is 8.5%. For the eastern
states, the median is 4.5%.
These medians indicate that the population in the western states is growing, on average, about 4
percentage points faster than in the eastern states.
f. Based on the shapes of the distributions for population growth in eastern and western states, how would you expect the
means to compare to the medians? Explain. Because both distributions are somewhat skewed to the right, you should
expect the means to be larger than the medians.
g. How would you expect the mean percentage change in the western states to change if you removed Nevada from the
analysis? Explain. You should expect the mean to decrease significantly if you remove the high outlier, Nevada,
from the analysis because the mean is not resistant to outliers.
h. Produce boxplots (carefully enter the data from the
stemplots or from p 133 of the textbook) for comparing the
distributions of percentage changes in population between
eastern and western states. Comment on what the boxplots
reveal.
The western states have a high outlier of 32.3% growth
(Nevada), so any comparison should use resistant
measures. The western states have a median of 8.6%
growth while the eastern states have a median of 4.7%
growth. These boxplots reveal that the western states tend to have a higher percentage of population growth than
the eastern states, although there was also greater variability among the percentages in the western states.
i. Using all of the information you have collected, compare and contrast the distributions of population growth between
eastern and western states.
The distributions of percentage of population growth for both the eastern and western states are skewed to the
right. For the western states, there is at least one high outlier (6.9-1.5(8.05), 14.95+1.5(8.05)). There are no outliers
in the eastern states (3-1.5(7.7), 10.7+1.5(7.7)). The median of the western states is 8.6% and the median of the
eastern states is 4.7%. The IQR of the western states is 8.05% and the IQR of the eastern states is 7.7%. Western
states tend to have higher percentage population growth rates than do the eastern states, but there is more
variability in population growth in the western states. It would be interesting to investigate the very high growth in
the outlier state, Nevada.
Notes on how the paragraph is written: The context is the population growth in eastern and western states, measured by
percentage, and both distributions are skewed to the right with more observations on the left side. The skew, even in one
distribution, is a reason to use resistant measures. The outlier calculations are shown, and remember different technology
makes boxplots differently, so in some cases you might find one high outlier and in other cases two. It could also depend
if you use the full values given in the textbook or the truncated values (the decimal part is not included). Either method is
9
fine; we are trying to describe a typical value, which is not an outlier, but knowing there are outliers means we should use
resistant measures. With skewed distributions and/or at least one outlier, the median and IQR are used for both
distributions (their values are given with the measurement units, %). Then the median values are used to compare the
tendency of the distributions and the IQR values are used to compare the variability of the distributions. For a
comparative paragraph, this is the conclusion since your goal is to compare typical values between the distributions.
Go to page 135 in the textbook and read the information about the diabetes study. Download the data list to complete this
activity.
Notes on how the paragraph is written: The context is the age Americans are diagnosed with diabetes, not how old they
currently are. The distribution is left-skewed because most Americans are diagnosed when they are older. Using the
outlier calculations, at least the minimum value of 1 year old is an outlier. The skew and the outlier mean resistant
measures are best to describe a typical American’s age when they are diagnosed. The median and IQR are given (values
with measurement units). The last sentence is what we now know about a typical American’s age when they are
diagnosed with diabetes.
f. In the histogram given on p 136 in the textbook, the bin width is 5. Use technology and the given data list to create a
histogram and paste it below. What bin width did the technology use automatically or did it let you choose the bin width
(and what did you choose)? How does changing the bin width/using a different bin width change the appearance of the
histogram?
Answers may vary.
Stapplet used a bin width of 10, but there is a way to change the
bin width easily. The bin width of 10 has a similar shape to the
histogram given in the textbook.
10
g. Now change your bin width to 20 and paste the histogram below. Finally change the bin width to 2 and paste the
histogram below.
Bin width of 20:
Bin width of 2:
h. Which of the four histograms (the one the technology produced initially, the bin width of 2, the bin width of 5 or the bin
width of 20) do you think provides the most informative display? Explain. Neither the histogram with a bin width of
20, nor the histogram with a bin width of 2 provides an informative display of these data. In both cases, we cannot
see the true shape of the distribution. With these data, the histogram created using a bin width of 5 seems to
provide the most informative histogram. It provides enough information for you to see both clusters without
adding too much clutter to the graph.
Watch Out!
Turn to pages 137-138 and answer the questions from the watch outs.
a. Here is the completed table (make sure you used the data lists, not just the five flavors in the textbook…):
11
b. The following boxplots display the distribution of calorie amounts for the three brands:
Dreyer’s Ice Cream appears to have significantly fewer calories than the other two brands, as well as less variability
among its calorie amounts. All of the Dreyer ice cream flavors have fewer than 200 calories, whereas only 25% of the Ben
& Jerry’s flavors have 220 or fewer calories. There is a great deal of variability in the number of calories of the Ben &
Jerry flavors as they range from 110 to 360 calories, whereas (excluding outliers) the Cold Stone Creamery flavors range
from 360 to 440 calories.
c. The serving sizes may not be the same for all three brands. This would make it difficult to compare the calories as
given.
d. You could convert the Cold Stone Creamery serving from 170 grams to the comparable measure of volume in ½ cups.
e. Divide each of the Cold Stone Creamery listings by 170 grams, then multiply by 73 grams per ½ cup.
f. You calculate new value = (old value)/170 × 73.
g. Here is the five-number summary for Cold Stone’s calorie amounts using the “per half cup” scale: min = 55.82,
QL = 154.59, median = 167.47, QU = 171.76, max = 188.94. All of the values are much lower than the original and there is
less variability.
h. The following boxplots display the three distributions of calorie amounts:
Once you adjust the ice creams so that they all have the same
serving size, Cold Stone Creamery still has several flavors that
are low outliers (meaning that these flavors have an unusually
small amount of calories per serving). Excluding the outliers,
the Dreyer’s flavors are general lowest in calorie content,
followed by the Cold Stone Creamery flavors, which have a
very narrow spread (from only about 155 – 189 calories per ½
cup); at least 75% of the Ben & Jerry flavors have more
calories than either of the other two brands.
a. The dotplot for the variable mileage corresponds to boxplot I. This plot is strongly skewed to the
right with a couple of high outliers.
b. The dotplot for the variable year of manufacture corresponds to boxplot III. This plot is skewed to the left.
c. The dotplot for the variable price corresponds to boxplot II. This plot has long tails on both sides,
although the left tail is much longer than the right.
12
Video 6 Follow-Up Questions:
3. a. Data Set B has the larger standard deviation because it has more values farther from the mean.
3. b. Data Sets C and D have the same standard deviation – they have the same number of values on either side of
the mean.
3. c. Data Set E has the larger standard deviation because it has more values farther from the mean.
4. A is 1.124; B is 1.451; C and D are 1.026; E is 1.589
Watch Out!
Turn to page 178 and answer the following questions based on the Watch Out
1. If there is an odd number of observations, is the median included in the upper half of the data, the lower half or neither?
Do not include the median in either the upper or lower half.
3. Can range and interquartile range be used with categorical data? No.
b. See table.
Class F Class G Class H Class I Class J
c. See table.
IQR 2 3 0 8 4
13
d. According to these measures of spread, because the IQR and SD are both higher, class G has more variability
than class F. That means class G has more data values farther from the mean,
e. According to these measures of spread, class I has the most variability (highest IQR and SD) and class H has the
least variability (lowest IQR and SD). Class I has more of its data values at the extremes and far from the center,
whereas most of class H’s observations are close to the mean. The spread of class J is in between the spreads of
class I and class H.
f. Class F has more bumpiness or unevenness in its graph, but has less variability (lower values for measures of
spread) than class G.
g. Class J has the greatest number of distinct values but does not have the most variability (highest values for
measures of spread) among classes H, I and J.
h. Based on the two previous questions, variability does not measure either bumpiness nor variety/distinct values;
variability measures spread from the center (mean). A distribution can be very ‘bumpy’ without having a great
deal of variability and vice versa. It is more important to consider the overall tendency for data values to be near
or far from the center in determining variability.
i. Many answers are possible, but all 10 values need to be the same so that the standard deviation is zero.
j. Only one answer is possible: {1, 1, 1, 1, 1, 9, 9, 9, 9, 9}. This dataset maximizes the distances of the observations
from the mean and has a standard deviation of 4.22. Any other combination will have a smaller standard deviation.
(Note: if you did not balance the 1s and 9s, the mean would shift away from 5 and would put the more frequent
values closer to the mean.)
Exercise 9-13:
a. Answers will vary. Students should consider that moving the entire dataset two values to the right should
also move the center by this amount.
b. The mean and median values have increased by 2 years to 16.89 years and 18 years respectively.
c. Answers will vary. Students should realize that adding two to each value in the dataset will not change how
the data vary.
d. The IQR and SD did not change. They are 18.5 years and 10.69 years respectively.
e. Answers will vary.
f. The mean is now 29.78 years and the median is now 32 years. The IQR is 37 years and the SD is 21.92
years. All of these values have doubled.
(We will return to Topic 9 in Unit 4; you are only responsible for content through Activity 9-3 right now.)
Extra Practice with Topic 9: ([Link]
Problems begin on book p 190: 9-10
14
SOCS Problem 1:
Notes on how the paragraph was written: The context is chest sizes of Scottish militiamen, in inches, from 1846 which is
all important to mention. While there are outliers, with such a symmetric graph, the non-resistant measures can be used.
They are similar to the resistant measures, which is another reason why they can be used (and is also making the graph
symmetric). But since there are outliers, using resistant measures would also be fine. The mean and standard deviation
are given (values with measurement units) to describe the chest size of a typical Scottish militiaman from 1846.
SOCS Problem 2:
SD 10.755% 10.307%
The AP Statistics test scores for the First class are left-skewed and the Last class scores are roughly symmetric and
mound-shaped. For the First class, there is one outlier at the lower end and no outliers at the upper end (73-
1.5(13), 86+1.5(13)). For the Last class, there are no outliers (68-1.5(13.5), 81.5+1.5(13.5)). The median of the
scores in the First class of the day is 82% and the median of the scores in the Last class of the day is 74%. The
IQR for the First class is 13% and the IQR for the Last class is 13.5%. Generally, the test scores from the First
class of the day are higher than those from the Last class. Since the two classes have very similar variability, the
First class is consistently better; at least on this test.
Notes on how the paragraph was written: The scores are for the same AP Stats test, but for two different classes.
You could use % or points as the measurement units since it isn’t clear. The First class is skewed and has a
lower outlier, so resistant measures should be used to compare. The outlier calculations are shown for both
classes; you can’t just look at the distribution to decide if there are outliers. The medians and IQRs are given
15
(values with measurement units). Then the distributions are compared; what is similar or different about what a
typical score tends to be and what is similar or different about the variability in scores.
SOCS Problem 3:
(a) What do the stems and leaves represent in the stemplot? Have the data been rounded?
Stems = thousands, leaves = hundreds. The data have been rounded to the nearest $100.
The distribution of tuitions and fees at the 81 colleges and universities in Michigan in 1999 in the 1999 -
2000 school year is skewed strongly to the right. There are no outliers (1800-1.5(8800), 10600+1.5(8800)).
The median cost for tuition and fees is $4500 and the IQR is $8800. While most colleges and universities
have costs around $4500, there is a lot of variability and many options for under $2000.
Notes on how the paragraph was written: While Kalamazoo was mentioned twice in the problem, these are all
the colleges and universities in Michigan. Kalamazoo just happens to have the cheapest and most expensive,
but we are interested in the typical cost so this isn’t really important. Mentioning other details like the year are
important. There are no outliers, but there is a very strong skew, so resistant measures are best. The median
and IQR are given (with values and measurement units) and then used to describe the costs for a typical school
in Michigan.
SOCS Problem 4:
The distribution of test scores on Test 1 in Mr. Carey’s BC Calculus class is skewed left whereas the
distribution of test scores on Test 2 are roughly symmetric and mound shaped. There is at least one low
outlier on Test 1 (80-1.5(12.5), 92.5+1.5(12.5)), but Test 2 doesn’t have any outliers (66-1.5(16.5),
82.5+1.5(16.5)). The median of Test 1 is 86 points and the median of Test 2 is 75 points. The IQR of Test
1 is 12.5 points and the IQR of Test 2 is 16.5 points. Students tended to do consistently better on Test 1;
maybe it was easier.
Notes on how the paragraph was written: The context is scores on two different tests using points. One
distribution is skewed, and one isn’t, but when comparing you should use the same measures for both
distributions, so resistant measures are best. There is also at least one outlier on Test 1, so this also points to
using resistant measures for both to compare. The medians and IQRs are given (with values and measurement
units). The last sentence compares what we learned about the difference in test scores on the two tests.
1. The distribution of the percentage of residents aged 65 or older in each of the 50 states is slightly skewed to the
left. There is one outlier in each tail of the distribution (11.5-1.5(2.2), 13.7+1.5(2.2)). There are gaps of 2% in the
distribution before the outliers, Alaska and Florida, occur. The IQR is 2.2% and the median is 12.75%. It makes
sense that states with more extreme temperatures would be outliers, with a higher percent living in the warmer
Florida and a lower percent living in the colder Alaska, but most states have about 13% of their population aged
65 and older with little variability.
Notes on how the paragraph was written: Include the context, percentage of residents 65 and older by state, throughout the
paragraph. With values, use %. The shape is slightly skewed; symmetric and/or mound-shaped would also be fine. Show
outlier calculations, don’t just assume the values at either end are outliers. Mention any gaps if they seem significant;
these gaps are about as big as the IQR so that seems significant. Give the resistant measures (value with units) because of
16
the outliers and slight skew, then use the measures to talk about the typical percentage of residents aged 65 and older.
While the outliers maybe interesting, they are not typical.
2.
The distribution for female life expectancy for 15 regions of the world is left-skewed while the male
distribution is more symmetric. There are no outliers in the female distribution (70.3-1.5(12),
82.3+1.5(12)) or the male distribution (64.1-1.5(11.9), 76.1+1.5(11.9)). The median of the female life
expectancy is 73.9 years with an IQR of 12 years. The median of the male life expectancy is 67 years with
an IQR of 11.9 years. In these 15 regions of the world, females tend to consistently live longer. It would
be interesting to look at a scatterplot of this data to see if the regions are similar for females and males or
if putting the data together and looking at females and males separately is different.
Notes on how the paragraph was written: There are life expectancies, given in years, for females and males in
15 regions of the world being compared. With so few observations, dotplots or stemplots should be used
instead of histograms. Boxplots are especially useful in comparing though. Without any outliers, resistant
measures are still used due to the skew in the female distribution. And resistant measures are used for both
distributions so they can be compared. The medians and IQRs are given (values with units) and then they are
used to compare the tendency and consistency of the life expectancies.
Turn to page 571 and fill in the blanks from the textbook definitions:
Three aspects of the association between quantitative variables are direction, strength, and form. Direction refers
to whether greater values of one variable tend to occur with greater values of the other variable (positive
association) or with smaller values of the other variable (negative association). The strength of the
association indicates how closely the observations follow the relationship between the variables.
In other words, the strength of the association reflects how accurately you could predict the value of one variable
based on the value of the other variable. The form of the association can be linear, or it can follow some more
complicated pattern.
17
Activity 26-1: House Prices
e. Direction: positive
Strength: moderate
Form: linear
b.
There appears to be a strong, positive, curved association.
c. This is a ridiculous argument. Simply sending televisions to countries such as Haiti will not lead to an increase
in life expectancy for its inhabitants.
d. No, this is a very good example of two variables being strongly associated without having a cause-and-effect
relationship between them. Having a television does not increase life expectancy. There are several other
variables that effect both life expectancy and number of televisions per 1000 people.
e. Many answers are possible. One possibility is wealth of the nation. Being a wealthier nation will often mean
there are more televisions per 1000 people and better health care which leads to longer life expectancy.
18
Watch Out
Turn to page 572 and answer the questions from the “Watch-Out”:
a. There does appear to be a very strong relationship between Raleigh’s average temperature and
the month number.
b. The correlation coefficient is 0.257. This indicates a weak positive association.
c. The correlation coefficient is so close to zero despite the very strong relationship because the
relationship is not linear. The correlation coefficient measures the strength of a linear
relationship.
Watch Out!
Turn to page 599 and answer the questions using the “Watch-Out”:
19
Extra Practice with Topic 27: ([Link]
Problems begin on book p 598: 27-3 Lists: LFEXP, TVPER, 27-4, 27-13
20
21
Extra Practice with Topic 28: ([Link]
Problems begin on book p 637: 28-11 (use technology to find the equation in part c), 28-27 Lists: ANN, PAULA, ROUND
22