Introductory Statistics First Edition (2021)
Instructor Edition
In-Class Activity 16.B
ANOVA for Regression (OPTIONAL)
Overview and student objectives
In-Class Activity Length: 25 minutes
Prior In-Class Activity: In-Class Activity 16.A, “What’s in the Price of a House?”
Next In-Class Activity: In-Class Activity 16.C, “Capital Bikeshare Rentals” (25 minutes)
Constructive Perseverance Level: 2
Content Learning Outcomes: DP.3.D, DP.4.B
DCMP Learning Goals: Communication, Reasoning
Overview
In this in-class activity, students will extend what they have learned about analysis of
variance (ANOVA) by using multiple representations including scatterplots,
mathematical formulas, and statistical terminology to describe the partitioning of
variation in a regression context. Students will then revisit the coefficient of
determination and inference for the regression slope, viewing these concepts through
the lens of ANOVA. Finally, students will discuss the factors that affect the value of the
ANOVA F-statistic in a regression context.
Objectives
Students will understand:
• Sums of squares can be used to partition variation in a regression context.
• The coefficient of determination (𝑅! ) and inference for the regression slope can
be described in terms of sums of squares.
Students will be able to:
• Explain what is measured by SSRegression, SSResiduals, and SSTotal in a
regression context.
• Discuss the factors that affect the value of the F-statistic in a regression context.
Suggested resources and preparation
Materials and technology
• Computer, projector, document camera
• Preview Assignment 16.B
• Student Pages for In-Class Activity 16.B
• Practice Assignment 16.B
• Access to the DCMP Data Analysis Tools
Copyright © 2021, The Charles A. Dana Center at The University of Texas at Austin
Introductory Statistics First Edition (2021)
Instructor Edition
• Calculator
Prerequisite assumptions
Students should be able to:
• Use an ANOVA table to organize sums of squares and calculate an F-statistic.
• Interpret the coefficient of determination (𝑅! ) in context.
• Use appropriate symbols to represent sample means, actual values, and
predicted values.
Making connections
This activity:
• Connects back to the coefficient of determination, one-way ANOVA, and
inference for the regression slope.
• Connects forward to multiple linear regression.
Background context
None
Suggested instructional plan
Frame the activity (3 minutes)
Structure • Compared to some of the other in-class activities in this curriculum,
this activity is more concept-driven and less context-driven, so it’s
especially important that students have time to wrestle with
conceptual questions in groups before answers are provided. To
make time for these group discussions, you may need to use direct
instruction to introduce a few definitions, formulas, and conventional
representations.
• Although sums of squares and the F-test involve multi-step
calculations, highlighting conceptual interpretations throughout the
lesson is recommended. This also provides an alternate point of
entry for students who are less fluent with calculations.
Question 1
• Have students work in pairs to recall the purpose of the ANOVA
Think-Pair- table.
Share
• Ask students to share their responses to Question 1.
• The mechanical details of the calculations may be front of mind for
students, but make sure that this discussion emphasizes the
purpose of ANOVA (in particular, the partitioning of variability and
inference).
Copyright © 2021, The Charles A. Dana Center at The University of Texas at Austin
Introductory Statistics First Edition (2021)
Instructor Edition
• Transition to the in-class activity by briefly discussing the Objectives
for the activity.
Activity flow (17 minutes)
Questions 2–4
• Introduce sums of squares in a regression context (SSTotal,
Whole-Class SSRegression, and SSResiduals) and answer Questions 2–4 as a
large group.
• In-Class Activity 14.A (one-way ANOVA) did not include formulas
with summation notation, so some students may be seeing the
symbol Σ for the first time.
• Help students visualize the deviations that are being summed by
marking them on the scatterplots (below the objectives on the
student pages).
• Encourage students to describe what each sum of squares
measures in their own words.
• To support students’ understanding of what each sum of squares
measures, prompt them to imagine extreme scenarios where the
sums of squares are equal to 0.
Question 5
Pairs • Have students work in pairs to answer Question 5, which involves
sketching three scatterplots that satisfy the given criteria:
o SSTotal = 0
o SSRegression = 0 but SSTotal ≠ 0
o SSResiduals = 0 but SSTotal ≠ 0
• While the students are working, walk around the room to help pairs
that are stuck and check their work.
• If pairs are stuck, bring them back to the scatterplots below the
objectives and remind them that we are summing up squared
deviations. What would it look like if those deviations were all equal
to 0?
• If pairs have used the same graph twice, make sure they notice the
requirement that SSTotal ≠ 0 for the last two graphs.
• Ask a few students to share their work with the whole class, either by
reproducing their plots on the board or projecting their work using
the document camera.
Copyright © 2021, The Charles A. Dana Center at The University of Texas at Austin
Introductory Statistics First Edition (2021)
Instructor Edition
• Encourage students who are presenting to share their reasoning,
reinforcing what each sum of squares measures by putting it in their
own words.
Brief • Explain what it means to “partition” the variation in the data into two
Discussion parts: the part that is explained by the regression model
(SSRegression) and the part that remains unexplained
(SSResiduals).
• You may want to provide a few synonyms for “partition” in case this
word is unfamiliar to students.
• Show students how to express the coefficient of determination, 𝑅2,
using sums of squares.
• The formulas are written in terms of a proportion, but remind
students that 𝑅2 can also be expressed as a percentage.
• There are two formulas given for 𝑅2. The formula using
SSRegression will probably be intuitive for students if you’ve already
described SSRegression as the “explained” variability. For the
formula using SSResiduals, a conceptual explanation is
recommended: 1 minus the unexplained variability is another way of
describing the same thing, since the explained and unexplained
variability sum up to the total variability (100% or 1). You could also
show the connection between the two formulas algebraically.
Question 6
• Have students revisit the extreme examples they sketched in the
Pairs previous question and find 𝑅2 for these datasets.
• Some pairs will be able to answer this question very quickly based
on the formulas, but you want to make sure their small group
discussions go deeper than that. Warn them in advance that you will
be asking for conceptual/intuitive justifications in addition to
justifications based on the formulas.
• While the students are working, walk around the room to help groups
that are stuck and check their work.
• If pairs are stuck, reframe what it means for variability to be
“explained.” If they knew the value of 𝑋, would it be helpful for
predicting 𝑌?
• For groups that finish quickly, prompt them to rehearse their
conceptual/intuitive explanations.
• Ask a few groups of students to share their answers and
justifications with the whole class (ideally groups that haven’t had a
chance to present yet).
Copyright © 2021, The Charles A. Dana Center at The University of Texas at Austin
Introductory Statistics First Edition (2021)
Instructor Edition
Question 7
• Show how sums of squares can be organized in an ANOVA table in
Whole Class the context of the relationship between neighborhood income and
organic food access.
• We recommend getting straight to the example in context rather than
walking through the ANOVA table two separate times (once without
context and once with context).
• For the sake of time, it is sufficient for you to show the linear
regression tool (i.e., data analysis tool) on the projector. It is not
necessary for each student to use the tool themselves.
• As you fill in each cell of the ANOVA table, explain where the value
comes from by referencing the “formulas” given in the table above.
• In particular, point out that 𝑝 = the number of predictors. Up until
now, the students have only used simple linear regression models (𝑝
= 1), but that will change when they get to multiple linear regression
in In-Class Activity 17.A.
Question 8
• Show how to use the F Distribution to calculate the P-value in
Question 8.
• Encourage students to include detailed sketches on their pages. It
may be helpful to use notation similar to the data analysis tool so
students will remember how to use the data analysis tool later on.
• Since the F-statistic for these data is so large it’s “off the charts,” the
P-value will be extremely small. You may want to demonstrate how
the tool works for smaller F-statistics so they will be familiar with how
to calculate a P-value that doesn’t round to 0.
Questions 9 and 10
Pairs • Give students time to answer Questions 9 and 10 with partners.
• Using the data analysis tool, show that as the F-statistic gets larger,
the P-value gets smaller, indicating stronger evidence of an
association.
• Make sure students discuss the factors that affect the strength of
evidence (Question 10).
• While the students are working, walk around the room to help groups
that are stuck and check their work.
• If groups are stuck, prompt them to consider extreme examples. If
the regression line were almost perfectly flat, would that provide
stronger evidence of an association? What would provide stronger
Copyright © 2021, The Charles A. Dana Center at The University of Texas at Austin
Introductory Statistics First Edition (2021)
Instructor Edition
evidence, a sample size of 10 or a sample size of 10,000? Would an
almost perfect line of points be likely to occur by chance?
• For groups that finish quickly, prompt them to rehearse multiple
explanations—some using formulas and some using intuition.
Wrap-up/transition (5 minutes)
Wrap-up • Call on students to share their answers and justifications for
Questions 9 and 10.
• Prompt multiple explanations—some based on formulas and some
based on intuition.
• Since this discussion is being used a wrap-up, make sure to
reinforce key terms/concepts whenever possible. For example, the
slope of the regression line is related to SSRegression, the sample
size is related to MSResiduals, and the spread of the points around
the line is related to SSResiduals. (Sample size is actually related to
SSResiduals, SSRegression, and SSTotal as well because when we
have more data points, we are adding up more terms in the SS.)
• Emphasize precise communication and the appropriate use of
statistical terminology to describe the strength of evidence of an
association.
• Note that the F-statistic is equal to the t-statistic squared, and the P-
values will be the same as long as the t-test for the slope is two-
sided. It is impossible to conduct a one-sided test using ANOVA
because everything is squared, so we have lost information about
the direction of the relationship.
• Have students refer back to the Objectives for the activity and
check the ones they recognize.
Transition
• In the next in-class activity, students will use a bikeshare context to
interpret confidence intervals and prediction intervals using predicted
values from a linear regression equation.
Suggested assessment, assignments, and reflections
• Give Practice Assignment 16.B.
• Give the preview assignments, if any, for the activities you plan to complete in
the next class meeting.
Copyright © 2021, The Charles A. Dana Center at The University of Texas at Austin
Introductory Statistics First Edition (2021)
Instructor Edition
In-Class Activity 16.B ANOVA for Regression ANSWERS
Earlier in this course, you learned how to
conduct a one-way ANOVA for scenarios
that involve comparing more than two
groups. In this in-class activity, you’ll extend
what you learned about ANOVA to a
regression context. Using these new tools,
you will consider the relationship between
neighborhood income and organic food
access.
Credit: iStock/MEDITERRANEAN
1) What was the purpose of the ANOVA
table when comparing more than two groups?
Answers will vary.
Sample answer: The ANOVA table allowed us to break up the total variation,
SSTotal, into two parts: SSGroup (variation between groups) and SSError (variation
within groups). These sums of squares were then used to calculate the F-statistic
and carry out inference for more than two means.
Objectives for the activity
You will understand:
¨ Sums of squares can be used to partition variation in a regression context.
¨ The coefficient of determination (𝑅! ) and inference for the regression slope can be
described in terms of sums of squares.
You will be able to:
¨ Explain what is measured by SSRegression, SSResiduals, and SSTotal in a
regression context.
¨ Discuss the factors that affect the value of the F-statistic in a regression context.
[Continued on the next page.]
Copyright © 2021, The Charles A. Dana Center at The University of Texas at Austin
Introductory Statistics First Edition (2021)
Instructor Edition
Sum of Squares Sum of Squares Sum of Squares
Total Regression Residuals
'(𝑦 − 𝑦+)! '(𝑦- − 𝑦+)! '(𝑦 − 𝑦-)!
In a regression context, the ANOVA table includes three different sums of squares:
SSTotal, SSRegression, and SSResiduals.
2) What does SSTotal measure? Mark the previous scatterplots to show the deviations
between the actual values of the response variable (𝑦) and the mean response (𝑦+).
Sample plots:
Answers will vary.
Sample answer: SSTotal measures how spread out the values of the response
variable are around their mean.
3) What does SSRegression measure? Mark the previous scatterplot to show the
deviations between the predicted values of the response variable (𝑦-) and the mean
response (𝑦+).
Copyright © 2021, The Charles A. Dana Center at The University of Texas at Austin
Introductory Statistics First Edition (2021)
Instructor Edition
Answers will vary. (Sample plot shown in answer to Question 2.)
Sample answer: SSRegression measures how steep the regression line is. In other
words, it measures how much variability has been explained.
4) What does SSResiduals measure? Mark the previous scatterplot to show the
deviations between the actual values of the response variable (𝑦) and the predicted
values of the response variable (𝑦-).
Answers will vary. (Sample plot shown in answer to Question 2.)
Sample answer: SSResiduals measures how spread out the values of the response
variable are around the regression line (the size of the residuals). In other words, it
measures how much variability remains unexplained.
To better understand what each sum of squares measures, let’s imagine extreme
scenarios where the sums of squares are equal to 0.
5) Fill in the following table by sketching three scatterplots that satisfy the given criteria.
SSTotal = 0 SSRegression = 0 SSResiduals = 0
(but SSTotal ≠ 0) (but SSTotal ≠ 0)
Sample plots:
Copyright © 2021, The Charles A. Dana Center at The University of Texas at Austin
Introductory Statistics First Edition (2021)
Instructor Edition
An ANOVA is a way to “partition” the variation in the data. In other words, it divides the
total variation into two parts: the part that is explained by the regression model
(SSRegression) and the part that remains unexplained (SSResiduals).
𝑆𝑆𝑇𝑜𝑡𝑎𝑙 = 𝑆𝑆𝑅𝑒𝑔𝑟𝑒𝑠𝑠𝑖𝑜𝑛 + 𝑆𝑆𝑅𝑒𝑠𝑖𝑑𝑢𝑎𝑙𝑠
In In-Class Activity 6.C, you learned that the coefficient of determination, 𝑅2, is
interpreted as the percentage of variation in the response variable that can be explained
by the linear relationship with an explanatory variable. This quantity can be expressed
using the sums of squares. Note that 𝑅2 can be expressed as a percentage or as a
proportion.
𝑣𝑎𝑟𝑖𝑎𝑡𝑖𝑜𝑛 𝑒𝑥𝑝𝑙𝑎𝑖𝑛𝑒𝑑 𝑆𝑆𝑅𝑒𝑔𝑟𝑒𝑠𝑠𝑖𝑜𝑛 𝑆𝑆𝑅𝑒𝑠𝑖𝑑𝑢𝑎𝑙𝑠
𝑅! = = =1−
𝑡𝑜𝑡𝑎𝑙 𝑣𝑎𝑟𝑖𝑎𝑡𝑖𝑜𝑛 𝑆𝑆𝑇𝑜𝑡𝑎𝑙 𝑆𝑆𝑇𝑜𝑡𝑎𝑙
6) Revisit the extreme examples that you sketched in Question 5.
Part A: When SSRegression = 0, 𝑅2 = _____.
Answer: 0
Part B: When SSResiduals = 0, 𝑅2 = _____.
Answer: 1
Sums of squares can be organized in an ANOVA table. The following table provides the
information necessary to calculate an F-statistic in the context of regression. Note that
𝑛 = sample size and 𝑝 = number of predictors. In simple linear regression, 𝑝 = 1.
Source Df Sum sq Mean sq F value
""#$%&$''()* ,"#$%&$''()*
Regression 𝑝 𝑆𝑆𝑅𝑒𝑔𝑟𝑒𝑠𝑠𝑖𝑜𝑛 𝑀𝑆𝑅𝑒𝑔𝑟𝑒𝑠𝑠𝑖𝑜𝑛 = + ,"#$'(-./0'
""#$'(-./0'
Residuals 𝑛−1−𝑝 𝑆𝑆𝑅𝑒𝑠𝑖𝑑𝑢𝑎𝑙𝑠 𝑀𝑆𝑅𝑒𝑠𝑖𝑑𝑢𝑎𝑙𝑠 = *121+
Total 𝒏−𝟏 𝑺𝑺𝑻𝒐𝒕𝒂𝒍
Copyright © 2021, The Charles A. Dana Center at The University of Texas at Austin
Introductory Statistics First Edition (2021)
Instructor Edition
A statistics student from San Antonio, Texas completed a project to explore whether
there is a relationship between neighborhood income and access to organic items at
local grocery stores. Specifically, the student counted the number of organic vegetables
offered at 37 H.E.B. grocery stores. She then cross-referenced these data with the
average household incomes (in dollars) for the zip codes where the stores are located.1
7) Enter the grocery store data into the DCMP Linear Regression tool at
[Link] Select “Average income
in zip code” as the explanatory (𝑋) variable and “Number of organic items offered” as
the response (𝑌) variable. Under “Regression Options,” click the box to show the
ANOVA table. Use the information from the tool to fill out the following table.
Source Df Sum sq Mean sq F value
Regression 1 17175.1 175175.1 56.3
Residuals 35 10673.4 305.0
Total 36 27848.4
Answer: Noted above in red.
The ANOVA table will also include a P-value, which tells the probability of obtaining an
F-statistic as large or larger than the one in the sample if the null hypothesis was true.
An ANOVA F-test can be used to test the population slope for simple linear regression,
the same scenario where you used a t-test in In-Class Activity 16.A:
𝐻3 : 𝛽 = 0 vs. 𝐻4 : 𝛽 ≠ 0, where 𝛽 = the population slope relating the number of organic
items offered and the average income in zip code
To model the values of the F-statistic that would occur if the null hypothesis was true
and the assumptions for inference were met, you will use an F Distribution with 𝑑𝑓2 = 𝑝
and 𝑑𝑓! = 𝑛– 1– 𝑝.
8) Use the DCMP F Distribution tool at [Link] to
calculate a P-value that measures the evidence of an association between the
number of organic items offered and the average household income. Include a
sketch of the F Distribution. (You may assume that a linear model is appropriate and
all assumptions for inference are met.)
1
Scenario adapted from Skew the Script: [Link]
Copyright © 2021, The Charles A. Dana Center at The University of Texas at Austin
Introductory Statistics First Edition (2021)
Instructor Edition
Sample plot:
Answer: The image shows the F Distribution with df1 = p = 1 and df2 = n – 1 – p = 37
– 1 – 1 = 35. The value from our sample, F = 56.32, is larger than the values we
would expect to occur by chance, and the P-value rounds to 0.
9) At the 𝛼 = 0.05 significance level, do these data provide sufficient evidence of an
association between the number of organic items offered and the average
household income in the neighborhood? State your conclusion in context.
Answer: Yes; the P-value is approximately 0, which is much smaller than 𝛼 = 0.05.
These data provide very strong evidence of an association between the number of
organic vegetables offered and the average household income in the neighborhood.
Suppose you had conducted a t-test for the slope instead of an F-test for the slope in
this scenario. The value of the t-statistic would have been 7.50, the square root of the F-
statistic. The P-value for the two-sided t-test would be the same as the P-value for the
F-test.
10) As the F-statistic gets larger, the P-value gets smaller, indicating stronger evidence
of an association. Answer Parts A through C to understand which factors affect the
strength of evidence.
Part A: As the slope of the regression line gets steeper, the evidence of an
association gets ______ (stronger/weaker).
Answer: stronger
Part B: As the spread of the points around the regression line increases, the
evidence of an association gets ______ (stronger/weaker).
Answer: weaker
Copyright © 2021, The Charles A. Dana Center at The University of Texas at Austin
Introductory Statistics First Edition (2021)
Instructor Edition
Part C: Assuming that the slope and the spread of the points around the regression
line stay about the same, as the sample size increases, the evidence of an
association gets ______ (stronger/weaker).
Answer: stronger
Copyright © 2021, The Charles A. Dana Center at The University of Texas at Austin
Introductory Statistics First Edition (2021)
Instructor Edition
Copyright © 2021, The Charles A. Dana Center at The University of Texas at Austin