0% found this document useful (0 votes)
32 views14 pages

Analyzing Quantitative Data Distributions

Module 4 focuses on the distribution of quantitative data, emphasizing the description of shape, center, and spread of data through various graphical representations. It includes activities comparing graphs of hypothetical exam scores, analyzing dot plots and histograms, and distinguishing between categorical and quantitative variables. The module aims to develop statistical thinking by encouraging students to analyze data distributions and understand variability.

Uploaded by

gabeaguzman2006
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
32 views14 pages

Analyzing Quantitative Data Distributions

Module 4 focuses on the distribution of quantitative data, emphasizing the description of shape, center, and spread of data through various graphical representations. It includes activities comparing graphs of hypothetical exam scores, analyzing dot plots and histograms, and distinguishing between categorical and quantitative variables. The module aims to develop statistical thinking by encouraging students to analyze data distributions and understand variability.

Uploaded by

gabeaguzman2006
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Gabe Guzman

Module 4 – Distributions of Quantitative Data

Module 4.1 – Distributions of Quantitative Data: Introduction

Learning Goal: For the distribution of a quantitative variable, describe the overall pattern
(shape, center, and spread) and striking deviations from the pattern.
Specific Learning Objectives:

 Develop a way to describe and distinguish graphs of a quantitative variable.


 Identify reasonable explanations for what might explain the differences seen in different
data sets.
Set-up/Warm-up: Statisticians collect data in order to study and analyze groups. They focus on
the group’s data in aggregate, instead of focusing on individuals. For a statistician creating and
comparing graphs is often the first step in analyzing data.
In this activity you will compare and contrast the graphs of hypothetical sets of exam scores. In
other words, these graphs show made-up data. These hypothetical data sets have been
constructed to help you begin to “see” like a statistician. By comparing these data sets, you will
begin to develop an informal understanding of the key features of a graph that statisticians use
to describe data. We will develop these ideas more formally in future lessons.
During this activity do your best to describe what you see. Jot down notes to capture your
thinking as you go. You can use your notes during our class discussion of this activity. This
activity does not require you to remember anything or to apply previous knowledge.
For this module, we’ll be developing the following 3 things about the distribution of data:

Symmetric: The right and left sides of the graph are mirror images of each other.
Skewed to the right: The graph has more spread in the upper half than the lower half. There may be
outliers to the right, so the graph has a longer tail to the right.
Shape
Skewed to the left: The graph has more spread in the lower half than the upper half. There may be
outliers to the left, so the graph has a longer tail to the left.
Uniform: The graph is shaped like a rectangle. Each value occurs with the same frequency.

Center A representative number or interval of numbers that best describes the distribution of the data

Spread The smallest data value to the biggest data value

Page 17
1) Compare and contrast the 3 graphs shown
at the right. How are the graphs similar?
How are they different? What is the most
distinctive feature that distinguishes these
three graphs from each other?

The graphs are similar in shape and


spread, however each center is
different.

For each graph, pick a single exam score to summarize the overall performance of the
students. In other words, summarize each set of data with one number.

87
Set A___________ 89
Set B __________ 91
Set C___________

The average score for Set A is 87.4. What do you think the average score is for Set B and for
Set C? (See if you can answer this without doing any calculations.)

Set B 89.4 Set C 91.4

Which, if any, of the statements below is a reasonable explanation for the differences in the
graphs? Why?

 The graphs represent different classes. Different groups of students will obviously
perform differently on an exam.

 The graphs represent a single class after the teacher adjusted the grades. The teacher
realized that some of the exam questions were not well written, so she adjusted the
grades by adding points to the original exam scores. most probable

 There are no differences in the graphs. They look the same.

Page 18
2) Compare and contrast the 3 graphs shown
at the right. worked together
a. How are the graphs similar? How are
they different? What is the most
distinctive feature that distinguishes
these three graphs from each other?
Similar shape and center but the
spread is different.

b. Thinking about center: The average


exam score is the same for each set of
data (D, E and F). What do you think
the average is?

75

c. Thinking about spread: Spread is a description of the variability we see in the data.
70 80
For exam set D the scores vary from a low of __________ to a high of _________.
60 90
For exam set E the scores vary from a low of __________ to a high of _________.
50 100
For exam set F the scores vary from a low of __________ to a high of _________.

d. Which of the statements below, if any, is a reasonable explanation for the differences in
the graphs? Why?

 The graphs represent a single class after the teacher adjusted the grades. The teacher
realized that some of the exam questions were not well written, so she adjusted the
grades by adding points to the original exam scores.
 The graphs represent 3 different classes with teachers that have different grading
standards. One is an easy grader. One is a really hard grader.
 The differences could be explained by whether the teachers allowed the students to
work together on the exam and how much time the students were given to finish the
exam.

most probable

Page 19
3) Compare and contrast the 3 graphs
shown at the right. Skewed left
a. In each of these 3 graphs, the Range:
average score is the same. The 5080
average score is 75. In what other = 30
ways are the graphs similar? How
are they different? What is the
most distinctive feature that
distinguishes these three graphs Range:
Symmetrical
from each other? 6090
= 30
Graphs have similar center and
spread. The shapes are different.

Skewed right Range:


6797
30

b. What might explain the differences in these graphs?

Teacher might be an easy or hard grader. Or the exam may have been easy, hard, or normal
difficulty level.

c. Thinking about shape: How would you describe the shapes (symmetric, skewed left,
skewed right) of these three distributions of exam scores?

Skewed left, symmetric, and skewed right

Page 20
Module 4.2 Dot Plots

Learning Goal: For the distribution of a quantitative variable, describe the overall pattern
(shape, center, and spread) and striking deviations from the pattern.
Specific Learning Objective:

 Distinguish between categorical and quantitative variables.


 Identify graphs that represent the distribution of a quantitative variable.
 Analyze the distribution of a quantitative variable using a dotplot. Describe the shape,
give a general estimate of center, and determine the overall range.
Overview:
In this activity, you will practice analyzing the distributions of quantitative variables using
descriptions of shape, center and spread. We will focus on dotplots.
1) Here is a partial spreadsheet of 2011-2012 data for a set of hatchback cars.

Car make and model City miles per gallon Drive EPA size class Engine
Acura ZDX 16 mpg All wheel drive Sport utility 6 cylinder
Audi TTS 21 mpg All wheel drive Midsize car 4 cylinder
Chevrolet Aveo 27 mpg Front wheel drive Compact car 4 cylinder

a) Who are the individuals described by this data?


The Car make and model

b) How many variables are shown in the spreadsheet?

Four Names or groups

c) Which variables are categorical variables?


"Drive" and "EPA size class"
Numbers

d) Which variables are quantitative variables?

"City miles per gallon and "Engine"

e) Describe what the values in the spreadsheet tell us about the Chevrolet Aveo.

Chevy Aveo has largest city mpg and has the smallest engine

Page 21
2) The graphs shown are a case-
value graph and a dotplot of the
2011-2012 data for a set of
hatchback cars. Use one or both
graphs to answer the questions
below. Jot down notes or draw
on the graphs to show how you
determined your answers.

a) What is the best city miles per gallon (mpg) for this group of hatchbacks? 43 mpg Dot Plot

What model of hatchback gets the highest city mpg? Lexus CT 200h Case Value

What is the worst city mpg for this group of hatchbacks? 16 mpg Dot Plot

What model of hatchback gets the worst city mpg? Acura ZDX Case Value

For each of the above questions, indicate which graph (case-value or dot plot) was the
easiest to use to answer the question.

b) How many hatchbacks get 25 mpg in the city? Which graph was the easiest to use to
answer this question?

2 hatchbacks and Dot Plot was easiest to use.

c) What are the mpg rates that occur most frequently for hatchbacks? Which graph was
the easiest to use to answer this question?

26 mpg, 27 mpg, and 30 mpg. The Dot Plot was easiest.

d) If you had to pick one mpg rate to represent this data, what would it be? Why did you
choose that value?

Around 27 mpg.

Page 22
3) The case-value graph and the dotplot are two different graphs of the same data. What do
you see as the advantages and disadvantages of each type of graph?

I think the advantages of the case value graph are you get to focus on each individual car
while the advantages of the dot plot are that its more simplified with it
focusing more on the mpg. This is true for the questions as well since case value graph
was easiest to use for questions about certain models while the dot plot was easiest to use
for questions about mpg.
4) Statisticians make graphs to summarize data so they prefer to use graphs that show the
distribution of the data. A statistical distribution is defined as “an arrangement of the values
of a variable showing their observed frequency of occurrence.” Here is another definition of
a statistical distribution: “a representation that shows the possible values of a variable and
how often the variable takes those values.” Which graph, the case-value graph or the
dotplot shows that distribution of the variable MPGCity?

Dotplot

5) For each of the following dotplots, draw a smooth curve outlining the distribution, and the n
describe the shape of the distribution using course vocabulary (ex: symmetric, skewed to
the left, skewed to the right).

Symmetrical

Skewed right

Page 23
6) Suppose that almost everyone does well on the first exam with a few people (who did not
study) performing poorly. What is the shape of the distribution of exam scores?

Skewed left.

7) Thinking like a statistician: describing shape, center and spread.

a) Thinking about shape: Our shape descriptions (symmetric, skewed to the left, skewed to
the right) do not fit this distribution. How would you describe this distribution’s shape?

As partly skewed right with a little bit of symmetry.

b) Thinking about center: Previously you chose one value to represent the distribution of
mpg ratings for hatchbacks. What was that value?

27 mpg

To represent the distribution, Ann decided to calculate the average mpg rating using the
mean (or average). She got 34 mpg. Do you think Ann made a mistake? How can you tell
without doing any calculations?
She made a mistake because there are only 3 other cars above 34
mpg while there are at least 20 cars below 30 mpg meaning an
average of 34 mpg is improbable.
c) Thinking about spread: Spread is a description of the variability we see in the data.
16
Here the city mpg for these hatchbacks varies from a low of _______mpg to a high of
43
_______mpg.
What is the overall range for the mpg values (max-min)? 431627
18
Typical hatchbacks have a city mpg that varies from about _______mpg to about
31
_______mpg.
d) Thinking about deviations from the pattern: Are there any unusually larger or unusually
small mpg rates? If so, what are the unusual values?

41 and 43 mpg

Page 24
8) Here are dotplots of the sugar content (grams per serving) for some adult cereals and child
cereals. Compare the two distributions by comparing shapes, estimates of center, and
spreads.
Skewed right
Center = 6
Spread = 014
Range = 14

Skewed left
Center = 10
Spread = 115
Range = 14

Page 25
Module 4.3 - Histograms

Learning Goal: For the distribution of a quantitative variable, describe the overall pattern
(shape, center, and spread) and striking deviations from the pattern.
Specific Learning Objective:

 Distinguish between categorical and quantitative variables.


 Identify graphs that represent the distribution of a quantitative variable.
Analyze the distribution of a quantitative variable using a histogram. Describe shape, give a
general estimate of center and the overall range, and calculate relevant percentages.
Set-up/Warm-up: In this activity you will again practice analyzing the distributions of
quantitative variables using descriptions of shape, center and spread. This is the same type of
thinking you did previously with dot plots, but this time the data will be summarized in a
histogram.
This histogram shows the distribution of exam scores for 15 students in a Biology class.
a) How would you describe the
shape of this distribution of
exam scores? (Use course
vocabulary)

Roughly symmetrical

b) Give an interval that describes typical grades on this exam.

Around 60 to 80

c) The range is the largest value minus the smallest value (Range = Max – Min). What is the
largest the range could be? What is the smallest the range could be?

10040 = 60 9050 40

d) What percentage of the students made a D on the exam (a grade of 60-69%)?

4/15 0.2667 = 26.67%

e) What percentage of the students passed the exam with a 70 or better?


8/15 0.5333 = 53.33%
Page 26
Group Work:
1) Which of the graphs below is a histogram? How do you know?

This graph is the histogram because it


gives rough estimates and looks like the previous
histogram provided.

2) The following is a histogram


indicating the distribution of
scores on the Spring 2012
Module 1 Checkpoint for 82
students in Math 27 classes.

a) How would you describe


the shape of this
distribution of quiz
scores? (Use course
vocabulary.)

Roughly skewed left

b) Give an interval that describes typical performance on this quiz.

80104

Page 27
c) For each of the following questions, answer the question if the histogram provides
enough information to answer it. If not, write "not enough information" and explain
why.

i. What percentage of students scored below 80%?


10/82 0.1219 = 12.19%

ii. How many students made an A (scored a 90% or higher)?


not enough information

iii. What is the lowest grade on the Module 1 Checkpoint?

Somewhere between 4048

iv. What percentage of the students aced the quiz (a score of 100%)?

not enough information

v. What is the average (mean) quiz score?

not enough information

vi. Did the majority of students pass the quiz (70% or better)?

Yes

3) Match the following descriptions to the histograms I-III.

a) Scores on an easy exam for a class of students.


b) Scores on a hard exam for a class of students.
c) Number of siblings for a large sample of US adults.
d) Exact volume of soda in a one-liter bottle for a case of 24 bottles.
e) Dates on the pennies I have in my car ashtray.
f) Weights for a large sample of newborn babies.

Page 28
4) The data graphed in these histograms describes 43 elementary school children. The
variable is “percent of body weight carried in the school backpack.” A child who weighs
60 pounds and carries 9 pounds has a variable value of 15% since 9÷60=0.15=15%. The
American Chiropractic Association (ACA) recommends that children carry no more than
10% of their body weight.

a) Of the 3rd graders, how many are following the ACA recommendation?

b) Of the 3rd graders what percentage is following the ACA recommendation?

7/21 0.3333 = 33.33%

c) Of the 5th graders, what percentage is following the ACA recommendation?

7/22 0.3181 = 31.81%

d) Of all the children in this study, what percentage is NOT following the ACA
recommendation?

29/43 0.6744 = 67.44%

e) Of the 5th graders who are NOT following the ACA recommendation, what
percentage are carrying more than 25% of their body weight?
3/22 0.1363 = 13.63%
Page 29
5) This data comes from a survey of 228 students enrolled at Carnegie Mellon University in
Pennsylvania. Students were asked if they preferred to sit in the back (B), front (F) or
middle (M) of the classroom. 51 reported a preference for sitting at the front, 131 prefer
to sit in the middle, 46 prefer the back. They also reported their college GPA.

a) Do students who sit at the front of the class tend to have higher GPAs compared to
students who sit at the back of the class? Calculate percentages to support your
answer.

b) Do students who sit at the back of the class tend to have lower GPAs compared to
students who sit at the front or middle?

c) To answer the questions in (a) and (b), it is better to use percentages than counts.
Why is this?

Page 30

Common questions

Powered by AI

Both the dotplot and case-value graph of hatchbacks exhibit common center and spread characteristics, such as similar ranges with the lowest mpg being around 16 to the highest around 43, indicating a 27 mpg range. The typical central value selected was around 27 mpg, reflecting the cluster of common values .

The shapes of the three exam score distributions are skewed left, symmetric, and skewed right. These differences could be explained by variations in grading difficulty; for instance, a hard exam might result in a left-skewed distribution if most students score poorly, whereas a normal or easy exam might result in a symmetric or right-skewed distribution if several students score well .

Recognizing striking deviations is crucial for understanding anomalies that can distort overall data interpretation. For example, unusual mpg values like 41 and 43 might indicate outliers that significantly affect assessments of central tendency and spread, impacting data-driven decisions .

Percentage calculations allow for a normalized comparison across groups of different sizes. For instance, while raw numbers might obscure trends due to varying group sizes, percentages clearly indicate tendencies such as students who sit at the front having higher GPAs compared to those at the back .

Using an incorrect summary statistic like the mean inappropriately, as in the mpg scenario, can misrepresent the data center. When most data are below a few outlying high values, as seen with mpg where only three are above 34 with twenty below 30, this can lead to inaccurate interpretations affecting subsequent analysis and decisions .

A dotplot shows the distribution of the variable 'MPGCity' more effectively as it provides a clear visualization of the frequency of mpg values, while the case-value graph allows for an individual focus on each car model. Dotplots are advantageous for summarizing overall trends, whereas case-value graphs are better for identifying specific data points .

Selecting a representative mpg rate is important for summarizing the central tendency of hatchback data. A value of 27 mpg was chosen to represent the data, based on it being a mode that frequently appears in the distribution, providing an appropriate measure of central tendency without calculating mean or median .

Exam difficulty can be inferred from the score distribution shape; a symmetric distribution suggests a balanced level of difficulty, whereas skewed distributions might indicate an exam that was either too easy (left-skewed) or too hard (right-skewed), affecting students' performance patterns .

A histogram depicts the distribution shape by showing the frequency of score intervals. The distribution of exam scores is roughly symmetrical, and a typical grade interval is around 60 to 80, indicating where most students' performance falls .

The dotplot for adult cereals is skewed right with a center of 6 grams and a spread from 0-14 grams. In contrast, child cereals have a left-skewed distribution with a center around 10 grams, and a similar spread of 1-15 grams. These characteristics indicate that child cereals generally have higher sugar content with less variation .

You might also like