0% found this document useful (0 votes)
7 views41 pages

Organizing and Graphing Data Techniques

This document discusses the organization and graphical representation of qualitative and quantitative data, including methods for creating frequency distribution tables, bar graphs, and pie charts. It includes case studies illustrating survey results on political ideologies and banking knowledge among Millennials. The chapter emphasizes the importance of descriptive statistics in summarizing and analyzing large data sets.

Uploaded by

hamima2005
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views41 pages

Organizing and Graphing Data Techniques

This document discusses the organization and graphical representation of qualitative and quantitative data, including methods for creating frequency distribution tables, bar graphs, and pie charts. It includes case studies illustrating survey results on political ideologies and banking knowledge among Millennials. The chapter emphasizes the importance of descriptive statistics in summarizing and analyzing large data sets.

Uploaded by

hamima2005
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

2

C HA P TE R

Spencer Platt/Getty Images, Inc.


Organizing and Graphing Data
2.1 Organizing and Graphing What is your political ideology? Do you classify yourself a consistently liberal person or a consistently
Qualitative Data conservative person, a mostly liberal person or a mostly conservative person, or are you someone
Case Study 2–1 Ideological who belongs to a group called “mixed”? Pew Research Center conducted a national survey of 10,013
Composition of the U.S. adults in 2014 to find the political views of adults in the United States. To see the results of this study,
Public, 2014 see Case Study 2–1.

Case Study 2–2 Millennials’


Views on Their Level of In addition to thousands of private organizations and individuals, a large number of U.S. government
Day-to-Day Banking agencies (such as the Bureau of the Census, the Bureau of Labor Statistics, the National Agricultural
Knowledge Statistics Service, the National Center for Education Statistics, the National Center for Health Statis-
2.2 Organizing and Graphing tics, and the Bureau of Justice Statistics) conduct hundreds of surveys every year. The data collected
Quantitative Data from each of these surveys fill hundreds of thousands of pages. In their original form, these data sets
Case Study 2–3 Car Insurance may be so large that they do not make sense to most of us. Descriptive statistics, however, supplies
Premiums per Year in the techniques that help to condense large data sets by using tables, graphs, and summary mea-
50 States
sures. We see such tables, graphs, and summary measures in newspapers and magazines every day.
Case Study 2–4 Hours At a glance, these tabular and graphical displays present information on every aspect of life. Conse-
Worked in a Typical Week
quently, descriptive statistics is of immense importance because it provides efficient and effective
by Full-Time U.S. Workers
methods for summarizing and analyzing information.
Case Study 2–5 How Many
Cups of Coffee Do You This chapter explains how to organize and display data using tables and graphs. We will learn
Drink a Day? how to prepare frequency distribution tables for qualitative and quantitative data; how to construct
2.3 Stem-and-Leaf Displays bar graphs, pie charts, histograms, and polygons for such data; and how to prepare stem-and-leaf
2.4 Dotplots displays.

36
2.1 Organizing and Graphing Qualitative Data 37

2.1 Organizing and Graphing Qualitative Data


This section discusses how to organize and display qualitative (or categorical) data. Data sets are
organized into tables and displayed using graphs. First we discuss the concept of raw data.

2.1.1 Raw Data


When data are collected, the information obtained from each member of a population or sample
is recorded in the sequence in which it becomes available. This sequence of data recording is
random and unranked. Such data, before they are grouped or ranked, are called raw data.

Raw Data Data recorded in the sequence in which they are collected and before they are
processed or ranked are called raw data.

Suppose we collect information on the ages (in years) of 50 students selected from a
university. The data values, in the order they are collected, are recorded in Table 2.1. For
instance, the first student’s age is 21, the second student’s age is 19 (second number in the
first row), and so forth. The data in Table 2.1 are quantitative raw data.

Table 2.1 Ages of 50 Students

21 19 24 25 29 34 26 27 37 33
18 20 19 22 19 19 25 22 25 23
25 19 31 19 23 18 23 19 23 26
22 28 21 20 22 22 21 20 19 21
25 23 18 37 27 23 21 25 21 24

Suppose we ask the same 50 students about their student status. The responses of the
students are recorded in Table 2.2. In this table, F, SO, J, and SE are the abbreviations for
freshman, sophomore, junior, and senior, respectively. This is an example of qualitative (or
categorical) raw data.

Table 2.2 Status of 50 Students

J F SO SE J J SE J J J
F F J F F F SE SO SE J
J F SE SO SO F J F SE SE
SO SE J SO SO J J SO F SO
SE SE F SE J SO F J SO SO

The data presented in Tables 2.1 and 2.2 are also called ungrouped data. An ungrouped data
set contains information on each member of a sample or population individually. If we rank the
data of Table 2.1 from lowest to the highest age, they will still be ungrouped data but not raw data.

2.1.2 Frequency Distributions


The Gallup polling agency recently surveyed randomly selected 1015 adults aged 18 and over
from all 50 U.S. states and the District of Columbia. These adults were asked, “Please tell me
how concerned you are right now about each of the following financial matters, based on your
current financial situation—are you very worried, moderately worried, not too worried, or not
worried at all.” Among a series of financial situations, one such situation was not having enough
money to pay their normal monthly bills. Table 2.3 lists the responses of these adults. The Gallup
report contained the percent of adults belonging to each category, which we have converted to
38 Chapter 2 Organizing and Graphing Data

numbers in the table. In this table, the variable is how much are adults worried about not having
enough money to pay normal monthly bills. The categories representing this variable are listed in
the first column of the table. Note that these categories are mutually exclusive. In other words,
each of the 1015 adults belongs to one and only one of these categories. The number of adults
who belong to a certain category is called the frequency of that category. A frequency distribu-
tion exhibits how the frequencies are distributed over various categories. Table 2.3 is called a
frequency distribution table or simply a frequency table.

Table 2.3 Worries About Not Having Enough


Money to Pay Normal Monthly Bills

Variable Response Number of Adults Frequency column

Very worried 162


Moderately worried 203
Category Not too worried 305 Frequency
Not worried at all 325
Others 20

Sum = 1015
Source: Gallup Poll.

Frequency Distribution of a Qualitative Variable A frequency distribution of a qualitative variable


lists all categories and the number of elements that belong to each of the categories.

Example 2–1 illustrates how a frequency distribution table is constructed for a qualitative
variable.

EX AM PLE 2 –1 What Variety of Donuts Is Your Favorite?


A sample of 30 persons who often consume donuts were asked what variety of donuts is their
Constructing a frequency
favorite. The responses from these 30 persons are as follows:
distribution table for
qualitative data. glazed filled other plain glazed other
frosted filled filled glazed other frosted
glazed plain other glazed glazed filled
frosted plain other other frosted filled
filled other frosted glazed glazed filled
Construct a frequency distribution table for these data.

Solution Note that the variable in this example is favorite variety of donut. This variable
has five categories (varieties of donuts): glazed, filled, frosted, plain, and other. To prepare
a frequency distribution, we record these five categories in the first column of Table 2.4.
Then we read each response (each person’s favorite variety of donut) from the given infor-
mation and mark a tally, denoted by the symbol |, in the second column of Table 2.4 next
to the corresponding category. For example, the first response is glazed. We show this in
the frequency table by marking a tally in the second column next to the category glazed.
© Jack Puccio/iStockphoto
Note that the tallies are marked in blocks of five for counting convenience. Finally, we
record the total of the tallies for each category in the third column of the table. This column is
called the column of frequencies and is usually denoted by f. The sum of the entries in the fre-
quency column gives the sample size or total frequency. In Table 2.4, this total is 30, which is the
sample size.
2.1 Organizing and Graphing Qualitative Data 39

Table 2.4 Frequency Distribution of Favorite Donut Variety

Donut Variety Tally Frequency ( f )

Glazed 8
Filled 7
Frosted 5
Plain 3
Other 7

Sum = 30

2.1.3 Relative Frequency and Percentage Distributions


The relative frequency of a category is obtained by dividing the frequency of that category by
the sum of all frequencies. Thus, the relative frequency shows what fractional part or proportion
of the total frequency belongs to the corresponding category. A relative frequency distribution
lists the relative frequencies for all categories.

Calculating Relative Frequency of a Category


Frequency of that category
Relative frequency of a category =
Sum of all frequencies

The percentage for a category is obtained by multiplying the relative frequency of that
category by 100. A percentage distribution lists the percentages for all categories.

Calculating Percentage
Percentage = (Relative frequency) · 100%

E X A MP L E 2 –2 What Variety of Donuts Is Your Favorite?


Determine the relative frequency and percentage distributions for the data in Table 2.4. Constructing relative frequency
and percentage distributions.
Solution The relative frequencies and percentages from Table 2.4 are calculated and listed in
Table 2.5. Based on this table, we can state that 26.7% of the people in the sample said that glazed
donut is their favorite. By adding the percentages for the first two categories, we can state that
50% of the persons included in the sample said that glazed or filled donut is their favorite. The
other numbers in Table 2.5 can be interpreted in similar ways.

Table 2.5 Relative Frequency and Percentage Distributions


of Favorite Donut Variety

Donut Variety Relative Frequency Percentage

Glazed 8/30 = .267 .267(100) = 26.7


Filled 7/30 = .233 .233(100) = 23.3
Frosted 5/30 = .167 .167(100) = 16.7
Plain 3/30 = .100 .100(100) = 10.0
Other 7/30 = .233 .233(100) = 23.3

Sum = 1.000 Sum = 100%


CASE STUDY 2–1
IDEOLOGICAL IDEOLOGICAL COMPOSITION OF THE U.S. PUBLIC, 2014
COMPOSITION OF
THE U.S. PUBLIC, Consistently
12%
2014 liberal

Mostly liberal 22%

Mixed 39%

Mostly conservative 18%

Consistently
9%
conservative

Data source: Pew Research Center

Pew Research Center conducted a national survey of 10,013 adults January 23 to March 16, 2014, to find the
political views of adults in the United States. As the above bar chart shows, 12% of the adults polled said
that they were consistently liberal, 22% indicated that they were mostly liberal, and so on. In this survey, Pew
Research Center also found that, overall, the percentage of Americans who indicated that they were consist-
Source: Pew Research Center, June,
2014 Report: Political Polarization in the ently conservative or consistently liberal has increased from 10% to 21% during the past two decades. Note
American Public. that in this chart, the bars are drawn horizontally.

Notice that the sum of the relative frequencies is always 1.00 (or approximately 1.00 if the
relative frequencies are rounded), and the sum of the percentages is always 100 (or approxi-
mately 100 if the percentages are rounded). ◼

2.1.4 Graphical Presentation of Qualitative Data


All of us have heard the adage “a picture is worth a thousand words.” A graphic display can reveal
at a glance the main characteristics of a data set. The bar graph and the pie chart are two types
of graphs that are commonly used to display qualitative data.

Bar Graphs
To construct a bar graph (also called a bar chart), we mark the various categories on the horizon-
tal axis as in Figure 2.1. Note that all categories are represented by intervals of the same width. We
mark the frequencies on the vertical axis. Then we draw one bar for each category such that the
height of the bar represents the frequency of the corresponding category. We leave a small gap
between adjacent bars. Figure 2.1 gives the bar graph for the frequency distribution of Table 2.4.

Figure 2.1 Bar graph for the


frequency distribution of 9
Table 2.4. 8
7
6
Frequency

5
4
3
2
1
0
Glazed Filled Frosted Plain Other
Donut variety
40
CASE STUDY 2–2
MILLENNIALS’ VIEWS
EWS ON THEIR LEVEL OF DAY
DAY-TO-DAY BANKING KNOWLEDGE
TO DAY BANK MILLENNIALS’
Not at all VIEWS ON THEIR
knowledgeable
1%
Extremely
knowledgeable
LEVEL OF DAY-
Not very 19% TO-DAY BANKING
knowledgeable
5% KNOWLEDGE
Somewhat Very
knowledgeable knowledgeable
35% 40%

Data source: TD Bank: The Millennial Financial Behaviors & Needs Survey

TD Bank conducted a poll of Millennials (aged 18–34) January 28 to February 10, 2014, with the main focus
on understanding their banking behaviors and preferences. As shown in the pie chart, 19% of the Millennials
said they were extremely knowledgeable about day-to-day banking, 40% said they were very knowledgeable,
Source: TD Bank: The Millennial
35% said they were somewhat knowledgeable, 5% admitted not to be very knowledgeable, and 1% men- Financial Behaviors & Needs.
tioned that they were not knowledgeable at all. February 2014.

Bar Graph A graph made of bars whose heights represent the frequencies of respective categories
is called a bar graph.

The bar graphs for relative frequency and percentage distributions can be drawn simply by
marking the relative frequencies or percentages, instead of the frequencies, on the vertical axis.
Sometimes a bar graph is constructed by marking the categories on the vertical axis and the
frequencies on the horizontal axis. Case Study 2–1 presents such an example.

Pareto Chart
To obtain a Pareto chart, we arrange (in a descending order) the bars in a bar graph based on their
heights (frequencies, relative frequencies, or percentages). Thus, the bar with the largest height
appears first (on the left side) in a bar graph and the one with the smallest height appears at the
end (on the right side) of the bar graph.

Pareto Chart A Pareto chart is a bar graph with bars arranged by their heights in descending
order. To make a Pareto chart, arrange the bars according to their heights such that the bar with
the largest height appears first on the left side, and then subsequent bars are arranged in descend-
ing order with the bar with the smallest height appearing last on the right side.

Figure 2.2 shows the Pareto chart for the frequency distribution of Table 2.4. It is the same bar
chart that appears in Figure 2.1 but with bars arranged based on their heights.

Pie Charts
A pie chart is more commonly used to display percentages, although it can be used to display
frequencies or relative frequencies. The whole pie (or circle) represents the total sample or popu-
lation. Then we divide the pie into different portions that represent the different categories.
41
42 Chapter 2 Organizing and Graphing Data

9
8
7
6

Frequency
5
4
3
2
1
0
Glazed Filled Other Frosted Plain
Donut variety
Figure 2.2 Pareto chart for the frequency distribution of
Table 2.4.

Pie Chart A circle divided into portions that represent the relative frequencies or percentages
of a population or a sample belonging to different categories is called a pie chart.

Figure 2.3 shows the pie chart for the percentage distribution of Table 2.5.

Figure 2.3 Pie chart for the percentage distribution


of Table 2.5.

Other
23.3% Glazed
26.7%

Plain
10.0%

Filled
Frosted 23.3%
16.7%

EXE R CI S E S
CON CE PTS AND PROCEDURES a. Prepare a frequency distribution table.
b. Calculate the relative frequencies and percentages for all
2.1 Why do we need to group data in the form of a frequency table? categories.
Explain briefly. c. What percentage of the elements in this sample belong to
category Y?
2.2 How are the relative frequencies and percentages of categories
d. What percentage of the elements in this sample belong to
obtained from the frequencies of categories? Illustrate with the help of
category N or D?
an example.
e. Draw a pie chart for the percentage distribution.
2.3 The following data give the results of a sample survey. The f. Make a Pareto chart for the percentage distribution.
letters Y, N, and D represent the three categories.
A P P LIC ATIO NS
D N N Y Y Y N Y D Y 2.4 In the past few years, many states have built casinos and many more
Y Y Y Y N Y Y N N Y are in the process of doing so. Forty adults were asked if building casinos
N Y Y N D N Y Y Y Y is good for society. Following are the responses of these adults, where G
Y Y N N Y Y N N D Y stands for good, B indicates bad, and I means indifferent or no answer.
2.2 Organizing and Graphing Quantitative Data 43

B G B B I G B I B B c. What percentage of the respondents mentioned vegetables


G B B G B B B G G I and fruits, poultry, or cheese?
B G B B I G G G B B d. Make a Pareto chart for the relative frequency distribution.
I G B B B G G B B G 2.6 The following data show the method of payment by 16 cus-
tomers in a supermarket checkout line. Here, C refers to cash,
a. Prepare a frequency distribution table. CK to check, CC to credit card, D to debit card, and O stands for
b. Calculate the relative frequencies and percentages for all other.
categories.
c. What percentage of the adults in this sample said building C CK CK C CC D O C
casinos is good? CK CC D CC C CK CK CC
d. What percentage of the adults in this sample said building
casinos is bad or were indifferent? a. Construct a frequency distribution table.
e. Draw a bar graph for the frequency distribution. b. Calculate the relative frequencies and percentages for all
f. Draw a pie chart for the percentage distribution. categories.
g. Make a Pareto chart for the percentage distribution. c. Draw a pie chart for the percentage distribution.
2.5 A [Link] survey asked residents of Japan to 2.7 In a 2013 survey of employees conducted by Financial Finesse
name their favorite pizza topping. The possible responses included Inc., employees were asked about their overall financial stress
the following choices: pig-based meats, for example, bacon or levels. The following table shows the results of this survey (www.
ham (PI); seafood, for example, tuna, crab, or cod roe (S); vegeta- [Link]).
bles and fruits (V); poultry (PO); beef (B); and cheese (C). The
following data represent the responses of a random sample of 36 Financial Stress Level Percentage of Responses
people.
No financial stress 14
V PI B PI V PO S PI V S V S Some financial stress 63
PI S V V V PI S S V PI C V
High financial stress 18
V V C V S PO V PI S PI PO PI
Overwhelming financial stress 5
a. Prepare a frequency distribution table.
b. Calculate the relative frequencies and percentages for all a. Draw a pie chart for this percentage distribution.
categories. b. Make a Pareto chart for this percentage distribution.

2.2 Organizing and Graphing Quantitative Data


In the previous section we learned how to group and display qualitative data. This section explains
how to group and display quantitative data.

2.2.1 Frequency Distributions


Table 2.6 gives the weekly earnings of 100 employees of a large company. The first column lists
the classes, which represent the (quantitative) variable weekly earnings. For quantitative data,
an interval that includes all the values that fall on or within two numbers—the lower and upper
limits—is called a class. Note that the classes always represent a variable. As we can observe, the
classes are nonoverlapping; that is, each value for earnings belongs to one and only one class. The
second column in the table lists the number of employees who have earnings within each class.
For example, 4 employees of this company earn $801 to $1000 per week. The numbers listed in
the second column are called the frequencies, which give the number of data values that belong
to different classes. The frequencies are denoted by f.
For quantitative data, the frequency of a class represents the number of values in the data set
that fall in that class. Table 2.6 contains six classes. Each class has a lower limit and an upper
limit. The values 801, 1001, 1201, 1401, 1601, and 1801 give the lower limits, and the values
1000, 1200, 1400, 1600, 1800, and 2000 are the upper limits of the six classes, respectively. The
data presented in Table 2.6 are an illustration of a frequency distribution table for quantitative
data. Whereas the data that list individual values are called ungrouped data, the data presented in
a frequency distribution table are called grouped data.
44 Chapter 2 Organizing and Graphing Data

Table 2.6 Weekly Earnings of 100 Employees


of a Company
Variable Weekly Earnings Number of Employees Frequency column
(dollars) f

801 to 1000 4
1001 to 1200 11
Third class 1201 to 1400 39 { Frequency of
the third class
1401 to 1600 24
1601 to 1800 16
1801 to 2000 6
Lower limit of Upper limit of
the sixth class the sixth class

Frequency Distribution for Quantitative Data A frequency distribution for quantitative data
lists all the classes and the number of values that belong to each class. Data presented in the form
of a frequency distribution are called grouped data.

The difference between the lower limits of two consecutive classes gives the class width.
The class width is also called the class size.

Finding Class Width


To find the width of a class, subtract its lower limit from the lower limit of the next class. Thus:
Width of a class = Lower limit of the next class − Lower limit of the current class

Thus, in Table 2.6,


Width of the first class = 1001 − 801 = 200
The class widths for the frequency distribution of Table 2.6 are listed in the second column of
Table 2.7. Each class in Table 2.7 (and Table 2.6) has the same width of 200.
The class midpoint or mark is obtained by dividing the sum of the two limits of a class by 2.

Calculating Class Midpoint or Mark


Lower limit + Upper limit
Class midpoint or mark =
2

Thus, the midpoint of the first class in Table 2.6 or Table 2.7 is calculated as follows:

801 + 1000
Midpoint of the first class = = 900.5
2

The class midpoints for the frequency distribution of Table 2.6 are listed in the third column of
Table 2.7.
2.2 Organizing and Graphing Quantitative Data 45

Table 2.7 Class Widths and Class Midpoints for Table 2.6

Class Limits Class Width Class Midpoint

801 to 1000 200 900.5


1001 to 1200 200 1100.5
1201 to 1400 200 1300.5
1401 to 1600 200 1500.5
1601 to 1800 200 1700.5
1801 to 2000 200 1900.5

2.2.2 Constructing Frequency Distribution Tables


When constructing a frequency distribution table, we need to make the following three major
decisions.

(1) Number of Classes


Usually the number of classes for a frequency distribution table varies from 5 to 20, depending
mainly on the number of observations in the data set.1 It is preferable to have more classes as the
size of a data set increases. The decision about the number of classes is arbitrarily made by the
data organizer.

(2) Class Width


Although it is not uncommon to have classes of different sizes, most of the time it is preferable
to have the same width for all classes. To determine the class width when all classes are the same
size, first find the difference between the largest and the smallest values in the data. Then, the
approximate width of a class is obtained by dividing this difference by the number of desired
classes.

Calculation of Class Width


Largest value − Smallest value
Approximate class width =
Number of classes

Usually this approximate class width is rounded to a convenient number, which is then used
as the class width. Note that rounding this number may slightly change the number of classes
initially intended.

(3) Lower Limit of the First Class or the Starting Point


Any convenient number that is equal to or less than the smallest value in the data set can be used
as the lower limit of the first class.
Example 2–3 illustrates the procedure for constructing a frequency distribution table for
quantitative data.

1
One rule to help decide on the number of classes is Sturge’s formula:
c = 1 + 3.3 log n
where c is the number of classes and n is the number of observations in the data set. The value of log n can be obtained
by using a calculator.
46 Chapter 2 Organizing and Graphing Data

EX AM PLE 2 –3 Values of Baseball Teams, 2015


Constructing a frequency
The following table gives the value (in million dollars) of each of the 30 baseball teams as esti-
distribution table for mated by Forbes magazine (source: Forbes Magazine, April 13, 2015). Construct a frequency
quantitative data. distribution table.

Values of Baseball Teams, 2015

Value Value
Team (millions of dollars) Team (millions of dollars)

Arizona Diamondbacks 840 Milwaukee Brewers 875


Atlanta Braves 1150 Minnesota Twins 895
Baltimore Orioles 1000 New York Mets 1350
Boston Red Sox 2100 New York Yankees 3200
Chicago Cubs 1800 Oakland Athletics 725
Chicago White Sox 975 Philadelphia Phillies 1250
Cincinnati Reds 885 Pittsburgh Pirates 900
Cleveland Indians 825 San Diego Padres 890
Colorado Rockies 855 San Francisco Giants 2000
Detroit Tigers 1125 Seattle Mariners 1100
Houston Astros 800 St. Louis Cardinals 1400
Kansas City Royals 700 Tampa Bay Rays 605
Los Angeles Angels of Anaheim 1300 Texas Rangers 1220
Los Angeles Dodgers 2400 Toronto Blue Jays 870
Miami Marlins 650 Washington Nationals 1280

Solution In these data, the minimum value is 605, and the maximum value is 3200. Suppose
we decide to group these data using six classes of equal width. Then,

3200 − 605
Approximate width of each class = = 432.5
6

Now we round this approximate width to a convenient number, say 450. The lower limit of the
first class can be taken as 605 or any number less than 605. Suppose we take 601 as the lower
limit of the first class. Then our classes will be
601–1050, 1051–1500, 1501–1950, 1951–2400, 2401–2850, and 2851–3300
We record these five classes in the first column of Table 2.8.

Table 2.8 Frequency Distribution of the Values of Baseball Teams, 2015

Value of a Team Number of Teams


(in million $) Tally ( f)

601–1050 16
1051–1500 9
1501–1950 1
1951–2400 3
2401–2850 0
2851–3300 1

Σ f = 30
2.2 Organizing and Graphing Quantitative Data 47

Now we read each value from the given data and mark a tally in the second column of
Table 2.8 next to the corresponding class. The first value in our original data set is 840, which
belongs to the 601–1050 class. To record it, we mark a tally in the second column next to the
601–1050 class. We continue this process until all the data values have been read and entered in
the tally column. Note that tallies are marked in blocks of five for counting convenience. After
the tally column is completed, we count the tally marks for each class and write those numbers
in the third column. This gives the column of frequencies. These frequencies represent the
number of baseball teams with values in the corresponding classes. For example, 16 of the
teams have values in the interval $601–$1050 million.
Using the ∑ notation (see Section 1.7 of Chapter 1), we can denote the sum of frequencies
of all classes by ∑ f. Hence,

∑f = 16 + 9 + 1 + 3 + 0 + 1 = 30

The number of observations in a sample is usually denoted by n. Thus, for the sample data,
∑f is equal to n. The number of observations in a population is denoted by N. Consequently, ∑ f
is equal to N for population data. Because the data set on the values of baseball teams in Table 2.8
is for all 30 teams, it represents a population. Therefore, in Table 2.8 we can denote the sum of
frequencies by N instead of ∑ f. ◼

Note that when we present the data in the form of a frequency distribution table, as in
Table 2.8, we lose the information on individual observations. We cannot know the exact value
of any team from Table 2.8. All we know is that 16 teams have values in the interval $601–$1050
million, and so forth.

2.2.3 Relative Frequency and Percentage Distributions


Using Table 2.8, we can compute the relative frequency and percentage distributions in the
same way as we did for qualitative data in Section 2.1.3. The relative frequencies and percent-
ages for a quantitative data set are obtained as follows. Note that relative frequency is the same
as proportion.

Calculating Relative Frequency and Percentage


Frequency of that class f
Relative frequency of a class = =
Sum of all frequencies ∑ f
Percentage = (Relative frequency) · 100%

Example 2–4 illustrates how to construct relative frequency and percentage distributions.

E X A MP L E 2 –4 Values of Baseball Teams, 2015


Calculate the relative frequencies and percentages for Table 2.8. Constructing relative frequency
and percentage distributions.
Solution The relative frequencies and percentages for the data in Table 2.8 are calculated and
listed in the second and third columns, respectively, of Table 2.9.
Using Table 2.9, we can make statements about the percentage of teams with values
within a certain interval. For example, from Table 2.9, we can state that about 53.3% of the
baseball teams had estimated values in the interval $601 million to $1050 million in April
2015. By adding the percentages for the first two classes, we can state that about 83.3% of the
baseball teams had estimated values in the interval $601 million to $1500 million in April
2015. Similarly, by adding the percentages for the last three classes, we can state that about
48 Chapter 2 Organizing and Graphing Data

Table 2.9 Relative Frequency and Percentage Distributions


of the Values of Baseball Teams

Value of a Team Relative


(in million $) Frequency Percentage

601–1050 16/30 = .533 53.3


1051–1500 9/30 = .300 30.0
1501–1950 1/30 = .033 3.3
1951–2400 3/30 = .100 10.0
2401–2850 0/30 = .000 0.0
2851–3300 1/30 = .033 3.3

Sum = .999 Sum = 99.9%

13.3% of the baseball teams had estimated values in the interval $1951 million to $3300 million
in April 2015. ◼

2.2.4 Graphing Grouped Data


Grouped (quantitative) data can be displayed in a histogram or a polygon. This section describes
how to construct such graphs. We can also draw a pie chart to display the percentage distribution
for a quantitative data set. The procedure to construct a pie chart is similar to the one for qualita-
tive data explained in Section 2.1.4; it will not be repeated in this section.

Histograms
A histogram can be drawn for a frequency distribution, a relative frequency distribution, or
a percentage distribution. To draw a histogram, we first mark classes on the horizontal axis
and frequencies (or relative frequencies or percentages) on the vertical axis. Next, we draw
a bar for each class so that its height represents the frequency of that class. The bars in a
histogram are drawn adjacent to each other with no gap between them. A histogram is called
a frequency histogram, a relative frequency histogram, or a percentage histogram
depending on whether frequencies, relative frequencies, or percentages are marked on the
vertical axis.

Histogram A histogram is a graph in which classes are marked on the horizontal axis and the
frequencies, relative frequencies, or percentages are marked on the vertical axis. The frequencies,
relative frequencies, or percentages are represented by the heights of the bars. In a histogram,
the bars are drawn adjacent to each other.

Figures 2.4 and 2.5 show the frequency and the percentage histograms, respectively, for
the data of Tables 2.8 and 2.9 of Sections 2.2.2 and 2.2.3. The two histograms look alike
because they represent the same data. A relative frequency histogram can be drawn for the
relative frequency distribution of Table 2.9 by marking the relative frequencies on the vertical
axis.
In Figures 2.4 and 2.5, we used class midpoints to mark classes on the horizontal axis. How-
ever, we can show the intervals on the horizontal axis by using the class limits instead of the class
midpoints.
CASE STUDY 2–3
CAR INSURANCE PREMIUMS PER YEAR IN 50 STATES CAR INSURANCE
AUT
INSU
O PREMIUMS PER
$800 to $1024 20% RAN xxxx
IN GOD
TRUSTWE

YEAR IN 50 STATES
x
x CE xxxx xxxxxxxxxxxx
D

xxx
xxxx xxxxx xxxx x
xxxx xxxxxxxx xxxx
xx x
x
D

WE
IN GODST xxxx
TRU
xxxx
xxxx
x x
xxxx
$1025 to $1249 34%
xxxx
xxxx xxxxxxxx
xxxx
xxxx
x
xxxx xxxx
xxxx x
xxxx
xxxx
x

$1250 to $1474 18%

$1475 to $1699 18%


IN GOD
TRUSTWE

$1700 or higher 10%


IN GOD
TRUSTWE

Data source: [Link]

The above histogram shows the percentage distribution of annual car insurance premiums in 50 states. The
data used to make this distribution and histogram are based on estimates made by [Link]. They col-
lected data from six large insurance companies in 10 ZIP codes for each state. The rates were obtained “for
the same full-coverage policy for the same driver—a 40-year-old man with a clean driving record and good
credit.” The rates used in this histogram are the averages “for the 20 best-selling vehicles in the U.S.” As the
histogram shows, in 20% of the states the car insurance rates were in the interval $800 to $1024, and so on. Source: [Link], April 13,
Note that the last class ($1700 or higher) has no upper limit. Such a class is called an open-ended class. 2015.

16 60
14
50
12
40
Frequency

10
Percent

8 30
6
20
4
10
2
0 0
825.5 1275.5 1725.5 2175.5 2625.5 3075.5 825.5 1275.5 1725.5 2175.5 2625.5 3075.5
Value (million $) Value (million $)
Figure 2.4 Frequency histogram for Table 2.8. Figure 2.5 Percentage distribution histogram for
Table 2.9.

Polygons
A polygon is another device that can be used to present quantitative data in graphic form. To
draw a frequency polygon, we first mark a dot above the midpoint of each class at a height equal
to the frequency of that class. This is the same as marking the midpoint at the top of each bar in
a histogram. Next we include two more classes, one at each end, and mark their midpoints. Note
that these two classes have zero frequencies. In the last step, we join the adjacent dots with
straight lines. The resulting line graph is called a frequency polygon or simply a polygon.
A polygon with relative frequencies marked on the vertical axis is called a relative frequency
polygon. Similarly, a polygon with percentages marked on the vertical axis is called a percentage
polygon.
49
50
CASE STUDY 2–4
Chapter 2 Organizing and Graphing Data

HOURS WORKED IN HOURS WORKED IN A TYPICAL WEEK BY FULL-TIME U.S. WORKERS


A TYPICAL WEEK 10
11 12 Less than 40 hours
12 1 1
BY FULL-TIME U.S. 10
11 2
3
9
8
2
12
60 or more 11 12 1 11 12 1
WORKERS 9
8 5
4
12 1
7
6 5 4
3
11 18% 18% 10
2
10 2
9 3
10 2
7 6 11
39
10 2 8
50 to 59 7 4 8 4
3 6 5 7 5
9
4
921%
3 6
8
7 5
6
8 40 to 49
53%
4
7 6 5

Data source: [Link]

The above pie chart shows the percentage distribution of hours worked in a typical week by full-time workers
in the United States. The data are based on a recent Gallup poll of 1271 workers. As the numbers in the pie
Source: [Link], chart show, 8% of these workers said they work for less than 40 hours a week, 53% work for 40 to 49 hours
August 29, 2014. a week, and so on. As you can observe, two of the classes are open-ended classes in this chart.

Polygon A graph formed by joining the midpoints of the tops of successive bars in a histogram
with straight lines is called a polygon.

Figure 2.6 shows the frequency polygon for the frequency distribution of Table 2.8.

16
14
12
Frequency

10
8
6
4
2
0
375.5 825.5 1275.5 1725.5 2175.5 2625.5 3075.5 3525.5
Value (million $)
Figure 2.6 Frequency polygon for Table 2.8.

For a very large data set, as the number of classes is increased (and the width of classes is
decreased), the frequency polygon eventually becomes a smooth curve. Such a curve is called a
frequency distribution curve or simply a frequency curve. Figure 2.7 shows the frequency curve
for a large data set with a large number of classes.
50
2.2 Organizing and Graphing Quantitative Data 51

Frequency
x
Figure 2.7 Frequency distribution curve.

2.2.5 More on Classes and Frequency Distributions


This section presents two alternative methods for writing classes to construct a frequency distri-
bution for quantitative data.

Less-Than Method for Writing Classes


The classes in the frequency distribution given in Table 2.8 for the data on values of baseball
teams were written as 601–900, 901–1200, and so on. Alternatively, we can write the classes in a
frequency distribution table using the less-than method. The technique for writing classes shown
in Table 2.8 is used for data sets that do not contain fractional values. The less-than method is
more appropriate when a data set contains fractional values. Example 2–5 illustrates the less-than
method.

E X A MP L E 2 –5 Federal and State Tax on Gasoline as


of April 1, 2015
Based on the information collected by American Petroleum Institute, Table 2.10 lists the total of Constructing a frequency
federal and state taxes (in cents per gallon) on gasoline for each of the 50 states as of April 1, distribution using the
2015 ([Link]). less-than method.

Table 2.10 Total Federal and State Tax on Gasoline as of April 1, 2015

State Gasoline Tax State Gasoline Tax

Alabama 39.3 Maryland 48.7


Alaska 29.7 Massachusetts 44.9
Arizona 37.4 Michigan 51.5
Arkansas 40.2 Minnesota 47.0
California 66.0 Mississippi 37.2
Colorado 40.4 Missouri 35.7
Connecticut 59.3 Montana 46.2
Delaware 41.4 Nebraska 44.9
Florida 54.8 Nevada 51.6
Georgia 44.9 New Hampshire 42.2
Hawaii 62.1 New Jersey 32.9
Idaho 43.4 New Mexico 37.3
Illinois 52.5 New York 62.9
Indiana 51.3 North Carolina 54.7
Iowa 50.4 North Dakota 41.4
Kansas 42.4 Ohio 46.4
Kentucky 44.4 Oklahoma 35.4
Louisiana 38.4 Oregon 49.5
Maine 48.4 Pennsylvania 70.0
(Continued)
52 Chapter 2 Organizing and Graphing Data

State Gasoline Tax State Gasoline Tax

Rhode Island 51.4 Vermont 48.9


South Carolina 35.2 Virginia 40.8
South Dakota 48.4 Washington 55.9
Tennessee 39.8 West Virginia 53.0
Texas 38.4 Wisconsin 51.3
Utah 42.9 Wyoming 42.4

Construct a frequency distribution table. Calculate the relative frequencies and percentages for all
classes.

Solution The minimum value in the data set of Table 2.10 is 29.7, and the maximum value is
70. Suppose we decide to group these data using five classes of equal width. Then,
70 − 29.7
Approximate class width = = 8.06
5
We round this number to a more convenient number—say 9—and take 9 as the width of each
class. We can take the lower limit of the first class equal to 29.7 or any number lower than 29.7.
If we start the first class at 27, the classes will be written as 27 to less than 36, 36 to less than 45,
and so on. The five classes, which cover all the data values of Table 2.10, are recorded in the first
column of Table 2.11. The second column in Table 2.11 lists the frequencies of these classes.
A value in the data set that is 27 or larger but less than 36 belongs to the first class, a value that
is 36 or larger but less than 45 falls into the second class, and so on. The relative frequencies and
percentages for classes are recorded in the third and fourth columns, respectively, of Table 2.11.
Note that this table does not contain a column of tallies.

Table 2.11 Frequency, Relative Frequency, and Percentage Distributions of the Total Federal
and State Tax on Gasoline

Federal and State Tax


(in cents) Frequency Relative Frequency Percentage

27 to less than 36 5 .10 10


36 to less than 45 21 .42 42
45 to less than 54 16 .32 32
54 to less than 63 6 .12 12
63 to less than 72 2 .04 4

Sum = 50 Sum = 1.00 Sum = 100

Note that in Table 2.11, the first column lists the class intervals using the boundaries, and not
the limits. When we use the less than method to write classes, we call the two end-points of a
class the lower and upper boundaries. For example, in the first class, which is 27 to less than
36, 27 is the lower boundary and 36 is the upper boundary. The difference between the two
boundaries gives the width of the class. Thus, the width of the first class is 36 − 27 = 9. All
classes in Table 2.11 have the same width, which is 9. ◼
A histogram and a polygon for the data of Table 2.11 can be drawn the same way as for the
data of Tables 2.8 and 2.9.

Single-Valued Classes
If the observations in a data set assume only a few distinct (integer) values, it may be appropriate
to prepare a frequency distribution table using single-valued classes—that is, classes that are
made of single values and not of intervals. This technique is especially useful in cases of discrete
data with only a few possible values. Example 2–6 exhibits such a situation.
CASE STUDY 2–5
HOW MANY CUPS OF COFFEE DO YOU DRINK A DAY? HOW MANY CUPS
4 or more cups OF COFFEE DO YOU
DRINK A DAY?
10%
3 cups
9%
0 cups
36%

2 cups
19%

1 cup
26%

Data source: Gallup poll of U.S. adults aged 18 and older conducted July 9–12, 2012

In a Gallup poll conducted by telephone interviews on July 9–12, 2012, U.S. adults of age 18 years and older
were asked, “How many cups of coffee, if any, do you drink on an average day?” According to the results of
the poll, shown in the accompanying pie chart, 36% of these adults said that they drink no coffee (repre-
sented by zero cups in the chart), 26% said that they drink one cup of coffee per day, and so on. The last
class is open-ended class that indicates that 10% of these adults drink four or more cups of coffee a day.
Source: [Link]
This class has no upper limit. Since the values of the variable (cups of coffee) are discrete and the variable 156116/Nearly-Half-Americans-Drink-
assumes only a few possible values, the first four classes are single-valued classes. [Link].

E X A MP L E 2 –6 Number of Vehicles Owned by Households


The administration in a large city wanted to know the distribution of the number of vehicles Constructing a frequency
owned by households in that city. A sample of 40 randomly selected households from this city distribution using single-valued
produced the following data on the number of vehicles owned. classes.

5 1 1 2 0 1 1 2 1 1
1 3 3 0 2 5 1 2 3 4
2 1 2 2 1 2 2 1 1 1
4 2 1 1 2 1 1 4 1 3
Construct a frequency distribution table for these data using single-valued classes.

Solution The observations in this data set assume only six distinct values: 0, 1, 2, 3, 4, and 5.
Each of these six values is used as a class in the frequency distribution in Table 2.12, and these
six classes are listed in the first column of that table. To obtain the frequencies of these classes,
© Jorge Salcedo/iStockphoto
the observations in the data that belong to each class are counted, and the results are recorded in
the second column of Table 2.12. Thus, in these data, 2 households own no vehicle, 18 own one
vehicle each, 11 own two vehicles each, and so on.
The data of Table 2.12 can also be displayed in a bar graph, as shown in Figure 2.8. To con-
struct a bar graph, we mark the classes, as intervals, on the horizontal axis with a little gap
between consecutive intervals. The bars represent the frequencies of respective classes.
The frequencies of Table 2.12 can be converted to relative frequencies and percentages the
same way as in Table 2.9. Then, a bar graph can be constructed to display the relative frequency
or percentage distribution by marking the relative frequencies or percentages, respectively, on the
vertical axis.
53
54 Chapter 2 Organizing and Graphing Data

Table 2.12 Frequency Distribution of the


18
Number of Vehicles Owned
15
Number of

Frequency
12
Vehicles Owned Households ( f )
9
0 2
6
1 18
2 11 3

3 4 0
0 1 2 3 4 5
4 3 Vehicles owned
5 2
Figure 2.8 Bar graph for Table 2.12.
∑f = 40

2.2.6 Cumulative Frequency Distributions


Consider again Example 2–3 of Section 2.2.2 about the values of baseball teams. Suppose we
want to know how many baseball teams had values of $1500 million or less in 2015. Such a question
can be answered by using a cumulative frequency distribution. Each class in a cumulative
frequency distribution table gives the total number of values that fall below a certain value. A
cumulative frequency distribution is constructed for quantitative data only.

Cumulative Frequency Distribution A cumulative frequency distribution gives the total number
of values that fall below the upper boundary of each class.

In a cumulative frequency distribution table, each class has the same lower limit but a dif-
ferent upper limit. Example 2–7 illustrates the procedure for preparing a cumulative frequency
distribution.

EX AM PLE 2 –7 Values of Baseball Teams, 2015


Constructing a cumulative
Using the frequency distribution of Table 2.8, reproduced here, prepare a cumulative frequency
frequency distribution table. distribution for the values of the baseball teams.

Value of a Team Number of Teams


(in million $) ( f)

601–1050 16
1051–1500 9
1501–1950 1
1951–2400 3
2401–2850 0
2851–3300 1

Solution Table 2.13 gives the cumulative frequency distribution for the values of the baseball
teams. As we can observe, 601 (which is the lower limit of the first class in Table 2.8) is taken as
the lower limit of each class in Table 2.13. The upper limits of all classes in Table 2.13 are the
same as those in Table 2.8. To obtain the cumulative frequency of a class, we add the frequency
of that class in Table 2.8 to the frequencies of all preceding classes. The cumulative frequencies
are recorded in the second column of Table 2.13.
2.2 Organizing and Graphing Quantitative Data 55

Table 2.13 Cumulative Frequency Distribution of Values


of Baseball Teams, 2015

Class Limits Cumulative Frequency

601–1050 16
601–1500 16 + 9 = 25
601–1950 16 + 9 + 1 = 26
601–2400 16 + 9 + 1 + 3 = 29
601–2850 16 + 9 + 1 + 3 + 0 = 29
601–3300 16 + 9 + 1 + 3 + 0 + 1 = 30

From Table 2.13, we can determine the number of observations that fall below the upper
limit of each class. For example, 26 teams were valued between $601 and $1950 million. ◼
The cumulative relative frequencies are obtained by dividing the cumulative frequencies
by the total number of observations in the data set. The cumulative percentages are obtained by
multiplying the cumulative relative frequencies by 100.

Calculating Cumulative Relative Frequency and Cumulative Percentage


Cumulative frequency of a class
Cumulative relative frequency =
Total observations in the data set
Cumulative percentage = (Cumulative relative frequency) · 100%

Table 2.14 contains both the cumulative relative frequencies and the cumulative percentages
for Table 2.13. We can observe, for example, that 90% of the teams were valued between $601
and $1800 million.

Table 2.14 Cumulative Relative Frequency and Cumulative Percentage


Distributions for Values of Baseball Teams, 2015

Cumulative Cumulative
Class Limits Relative Frequency Percentage

601–1050 16/30 = .5333 53.33


601–1500 25/30 = .8333 83.33
601–1950 26/30 = .8667 86.67
601–2400 29/30 = .9667 96.67
601–2850 29/30 = .9667 96.67
601–3300 30/30 = 1.000 100.00

2.2.7 Shapes of Histograms


A histogram can assume any one of a large number of shapes. The most common of these shapes
are
1. Symmetric
2. Skewed
3. Uniform or rectangular
A symmetric histogram is identical on both sides of its central point. The histograms shown
in Figure 2.9 are symmetric around the dashed lines that represent their central points.
56 Chapter 2 Organizing and Graphing Data

Frequency

Frequency
Variable Variable
Figure 2.9 Symmetric histograms.

A skewed histogram is nonsymmetric. For a skewed histogram, the tail on one side is
longer than the tail on the other side. A skewed-to-the-right histogram has a longer tail on the
right side (see Figure 2.10a). A skewed-to-the-left histogram has a longer tail on the left side
(see Figure 2.10b).
Frequency

Frequency
Variable Variable
(a) (b)
Figure 2.10 (a) A histogram skewed to the right. (b) A histogram skewed to the left.

A uniform or rectangular histogram has the same frequency for each class. Figure 2.11 is
an illustration of such a case.

Figure 2.11 A histogram with uniform


distribution.
Frequency

Variable

Figures 2.12a and 2.12b display symmetric frequency curves. Figures 2.12c and 2.12d show
frequency curves skewed to the right and to the left, respectively.
Frequency

Frequency

Variable Variable
(a) (b)
Frequency

Frequency

Variable Variable
(c) (d)
Figure 2.12 (a), (b) Symmetric frequency curves. (c) Frequency curve skewed to
the right. (d) Frequency curve skewed to the left.
2.2 Organizing and Graphing Quantitative Data 57

2.2.8 Truncating Axes


Describing data using graphs gives us insights into the main characteristics of the data. But graphs,
unfortunately, can also be used, intentionally or unintentionally, to distort the facts and deceive the
reader. The following are two ways to manipulate graphs to convey a particular opinion or impression.
1. Changing the scale either on one or on both axes—that is, shortening or stretching one or
both of the axes.
2. Truncating the frequency axis—that is, starting the frequency axis at a number greater than
zero.
Suppose 400 randomly selected adults were asked whether or not they are happy with their
jobs. Of them, 156 said that they are happy, 136 said that they are not happy, and 108 had no
opinion. Converting these numbers to percentages, 39% of these adults said that they are happy,
34% said that they are not happy, and 27% had no opinion. Let us denote the three opinions by
A, B, and C, respectively. The following table shows the results of this survey.

Opinion Percentage

A 39
B 34
C 27

Sum = 100

Now let us make two bar graphs—one showing the complete vertical axis and the second using a
truncated vertical axis. Figure 2.13 shows a bar graph with the complete vertical axis. By looking
at this bar graph, we can observe that the opinions represented by three categories in fact differ
by small percentages. But now look at Figure 2.14 in which the vertical axis has been truncated
to start at 25%. By looking at this bar chart, if we do not pay attention to the vertical axis, we may
erroneously conclude that the opinions represented by three categories vary by large percentages.

40.0
40
37.5
Percentage

Percentage

30 35.0

20 32.5
30.0
10
27.5
0 25.0
A B C A B C
Category Category
Figure 2.13 Bar graph without Figure 2.14 Bar graph with truncation
truncation of the vertical axis. of the vertical axis.

When interpreting a graph, we should be very cautious. We should observe carefully whether
the frequency axis has been truncated or whether any axis has been unnecessarily shortened or
stretched.

EXE RC I S E S
CONCEPT S AND PROCEDURES
2.9 How are the relative frequencies and percentages of classes
2.8 Briefly explain the three decisions that have to be made to group obtained from the frequencies of classes? Illustrate with the help of an
a data set in the form of a frequency distribution table. example.
58 Chapter 2 Organizing and Graphing Data

2.10 Three methods—writing classes using limits, using the less- 2.12 A data set on money spent on lottery tickets during the past year
than method, and grouping data using single-valued classes—were by 200 households has a lowest value of $1 and a highest value of $1167.
discussed to group quantitative data into classes. Explain these three Suppose we want to group these data into six classes of equal widths.
methods and give one example of each. a. Assuming that we take the lower limit of the first class as
$1 and the width of each class equal to $200, write the class
APPLI CAT IO NS limits for all six classes.
2.11 A local gas station collected data from the day’s receipts, record- b. Find the class midpoints.
ing the gallons of gasoline each customer purchased. The following 2.13 The following data give the one-way commuting times (in
table lists the frequency distribution of the gallons of gas purchased by minutes) from home to work for a random sample of 50 workers.
all customers on this one day at this gas station.
23 17 34 26 18 33 46 42 12 37
Gallons of Gas Number of Customers 44 15 22 19 28 32 18 39 40 48
16 11 9 24 18 26 31 7 30 15
0 to less than 4 31 18 22 29 32 30 21 19 14 26 37
4 to less than 8 78 25 36 23 39 42 46 29 17 24 31
8 to less than 12 49
a. Construct a frequency distribution table using the classes 0–9,
12 to less than 16 81 10–19, 20–29, 30–39, and 40–49.
16 to less than 20 117 b. Calculate the relative frequency and percentage for each class.
20 to less than 24 13 c. Construct a histogram for the percentage distribution made in
part b.
a. How many customers were served on this day at this gas station? d. What percentage of the workers in this sample commute for
b. Find the class midpoints. Do all of the classes have the same 30 minutes or more?
width? If so, what is this width? If not, what are the different e. Prepare the cumulative frequency, cumulative relative fre-
class widths? quency, and cumulative percentage distributions using the
c. Prepare the relative frequency and percentage distribution table of part a.
columns.
d. What percentage of the customers purchased 12 gallons or Exercises 2.14 to 2.19 are based on the following data
more? The following table includes partial data for a health fair held at a
e. Explain why you cannot determine exactly how many cus- local mall by community college nursing students. The data are
tomers purchased 10 gallons or less. provided for 30 male and 30 female participants who stopped by
f. Prepare the cumulative frequency, cumulative relative fre- the health fair booth. Data includes participant’s age to the nearest
quency, and cumulative percentage distributions using the year, body weight in pounds, and blood glucose level measured in
given table. mg/dL.

Blood Blood
Age Age Weight Weight Glucose Glucose
Participant (Males) (Females) (Males) (Females) (Males) (Females)

1 55 23 197 111 140 124


2 24 38 137 137 106 105
3 45 25 139 136 93 121
4 38 47 261 105 92 109
5 49 23 169 208 117 97
6 41 46 201 117 77 86
7 54 34 176 244 128 131
8 32 35 203 234 95 143
9 30 36 210 258 149 134
10 59 58 168 156 131 90
11 55 58 223 211 130 124
12 21 46 248 163 106 108
13 52 26 151 99 108 110
14 61 28 220 162 143 144
15 56 53 211 253 76 81
2.2 Organizing and Graphing Quantitative Data 59

Blood Blood
Age Age Weight Weight Glucose Glucose
Participant (Males) (Females) (Males) (Females) (Males) (Females)

16 55 33 104 119 124 109


17 52 62 262 231 89 81
18 33 35 122 258 107 85
19 34 31 136 138 103 135
20 33 29 156 139 100 105
21 31 20 222 232 98 97
22 27 34 174 140 112 132
23 34 50 230 236 121 113
24 48 47 167 234 105 149
25 25 44 262 149 81 99
26 47 21 234 196 134 149
27 49 55 146 202 113 75
28 22 43 214 229 93 85
29 35 53 130 220 124 126
30 29 31 232 190 76 148

2.14 a. Construct a frequency distribution table for ages of male d. What percentage of the female participants have weights less
participants using the classes 20–29, 30–39, 40–49, 50–59, than 161 lbs?
and 60–69. e. Compare the histograms for Exercises 2.16 and 2.17 and
b. Calculate the relative frequency and percentage for each mention the similarities and differences.
class.
2.18 a. Construct a frequency distribution table for blood glucose
c. Construct a histogram for the frequency distribution of part a.
levels of male participants using the classes 75–89, 90–104,
d. What percentage of the male participants are younger
105–119, 120–134, and 135–149.
than 40?
b. Calculate the relative frequency and percentage for each
2.15 a. Construct a frequency distribution table for ages of female class.
participants using the classes 20–29, 30–39, 40–49, 50–59, c. Construct a histogram for the percentage distribution of
and 60–69. part b.
b. Calculate the relative frequency and percentage for each d. What percentage of the male participants have blood glucose
class. levels more than 119?
c. Construct a histogram for the frequency distribution of part a. e. Prepare the cumulative frequency, cumulative relative fre-
d. What percentage of the female participants are younger than 40? quency, and cumulative percentage distributions using the
e. Compare the histograms for Exercises 2.14 and 2.15 and table of part a.
mention the similarities and differences.
2.19 a. Construct a frequency distribution table for blood glucose
2.16 a. Construct a frequency distribution table for weights of male levels of female participants using the classes 75–89, 90–104,
participants using the classes 91–125, 126–160, 161–195, 105–119, 120–134, and 135–149.
196–230, and 231–265. b. Calculate the relative frequency and percentage for each
b. Calculate the relative frequency and percentage for each class.
class. c. Construct a histogram for the percentage distribution of
c. Construct a histogram for the relative frequency distribution part b.
of part b. d. What percentage of the female participants have blood
d. What percentage of the male participants have weights less glucose levels more than 119?
than 161 lbs? e. Compare the histograms for Exercises 2.18 and 2.19 and
mention the similarities and differences.
2.17 a. Construct a frequency distribution table for weights of female
f. Prepare the cumulative frequency, cumulative relative fre-
participants using the classes 91–125, 126–160, 161–195,
quency, and cumulative percentage distributions using the
196–230, and 231–265.
table of part a.
b. Calculate the relative frequency and percentage for each
class. 2.20 The following table lists the number of strikeouts per game
c. Construct a histogram for the relative frequency distribution (K/game) for each of the 30 Major League baseball teams during the
of part a. 2014 regular season.
60 Chapter 2 Organizing and Graphing Data

Team K/game Team K/game Team K/game

Arizona Diamondbacks 7.89 Houston Astros 7.02 Philadelphia Phillies 7.75


Atlanta Braves 8.03 Kansas City Royals 7.21 Pittsburgh Pirates 7.58
Baltimore Orioles 7.25 Los Angeles Angels 8.28 San Diego Padres 7.93
Boston Red Sox 7.49 Los Angeles Dodgers 8.48 San Francisco Giants 7.48
Chicago Cubs 8.09 Miami Marlins 7.35 Seattle Mariners 8.13
Chicago White Sox 7.11 Milwaukee Brewers 7.69 St. Louis Cardinals 7.54
Cincinnati Reds 7.96 Minnesota Twins 6.36 Tampa Bay Rays 8.87
Cleveland Indians 8.95 New York Mets 8.04 Texas Rangers 6.85
Colorado Rockies 6.63 New York Yankees 8.46 Toronto Blue Jays 7.40
Detroit Tigers 7.68 Oakland Athletics 6.68 Washington Nationals 7.95
Data source: [Link].

a. Construct a frequency distribution table. Take 6.30 as the lower 2.22 The following table gives the frequency distribution for the
boundary of the first class and .55 as the width of each class. numbers of parking tickets received on the campus of a university
b. Prepare the relative frequency and percentage distribution during the past week by 200 students.
columns for the frequency distribution table of part a.
2.21 The following data give the number of turnovers (fumbles and Number of Tickets Number of Students
interceptions) made by both teams in each of the football games
played by a university during the 2014 and 2015 seasons. 0 59
1 44
2 3 1 1 6 5 3 5 5 1 5 2 1 2 37
5 3 4 4 5 8 4 5 2 2 2 6
3 32

a. Construct a frequency distribution table for these data using 4 28


single-valued classes.
b. Calculate the relative frequency and percentage for each class. Draw two bar graphs for these data, the first without truncating the
c. What is the relative frequency of games in which there were frequency axis and the second by truncating the frequency axis. In the
4 or 5 turnovers? second case, mark the frequencies on the vertical axis starting with
d. Draw a bar graph for the frequency distribution of part a. 25. Briefly comment on the two bar graphs.

2.3 Stem-and-Leaf Displays


Another technique that is used to present quantitative data in condensed form is the stem-and-
leaf display. An advantage of a stem-and-leaf display over a frequency distribution is that by
preparing a stem-and-leaf display we do not lose information on individual observations. A stem-
and-leaf display is constructed only for quantitative data.

Stem-and-Leaf Display In a stem-and-leaf display of quantitative data, each value is divided


into two portions—a stem and a leaf. The leaves for each stem are shown separately in a display.

Example 2–8 describes the procedure for constructing a stem-and-leaf display.

EX AM PLE 2 –8 Scores of Students on a Statistics Test


Constructing a stem-and-leaf
The following are the scores of 30 college students on a statistics test.
display for two-digit numbers.
75 52 80 96 65 79 71 87 93 95
69 72 81 61 76 86 79 68 50 92
83 84 77 64 71 87 72 92 57 98

Construct a stem-and-leaf display.


2.3 Stem-and-Leaf Displays 61

Solution To construct a stem-and-leaf display for these scores, we split each score into two
parts. The first part contains the first digit of a score, which is called the stem. The second part
contains the second digit of a score, which is called the leaf. Thus, for the score of the first stu-
dent, which is 75, 7 is the stem and 5 is the leaf. For the score of the second student, which is 52,
the stem is 5 and the leaf is 2. We observe from the data that the stems for all scores are 5, 6, 7,
8, and 9 because all these scores lie in the range 50 to 98. To create a stem-and-leaf display, we
draw a vertical line and write the stems on the left side of it, arranged in increasing order, as
shown in Figure 2.15.

Stems Figure 2.15 Stem-and-leaf display.

5 2 Leaf for 52
6
7 5 Leaf for 75
8
9

After we have listed the stems, we read the leaves for all scores and record them next
to the corresponding stems on the right side of the vertical line. For example, for the first
score we write the leaf 5 next to stem 7; for the second score we write the leaf 2 next
to stem 5. The recording of these two scores in a stem-and-leaf display is shown in
Figure 2.15.
Now, we read all scores and write the leaves on the right side of the vertical line in the
rows of corresponding stems. The complete stem-and-leaf display for scores is shown in
Figure 2.16.

5 2 0 7 Figure 2.16 Stem-and-leaf


6 5 9 1 8 4 display of test scores.
7 5 9 1 2 6 9 7 1 2
8 0 7 1 6 3 4 7
9 6 3 5 2 2 8

By looking at the stem-and-leaf display of Figure 2.16, we can observe how the data
values are distributed. For example, the stem 7 has the highest frequency, followed by stems 8,
9, 6, and 5.
The leaves for each stem of the stem-and-leaf display of Figure 2.16 are ranked (in increas-
ing order) and presented in Figure 2.17.

5 0 2 7 Figure 2.17 Ranked


6 1 4 5 8 9 stem-and-leaf display of
7 1 1 2 2 5 6 7 9 9 test scores.
8 0 1 3 4 6 7 7
9 2 2 3 5 6 8 ◼

As already mentioned, one advantage of a stem-and-leaf display is that we do not lose infor-
mation on individual observations. We can rewrite the individual scores of the 30 college students
from the stem-and-leaf display of Figure 2.16 or Figure 2.17. By contrast, the information on
individual observations is lost when data are grouped into a frequency table.
62 Chapter 2 Organizing and Graphing Data

EX AM PLE 2 –9 Monthly Rents Paid by Households


The following data give the monthly rents paid by a sample of 30 households selected from a
Constructing a stem-and-leaf
display for three- and four-digit
small town.
numbers.
880 1081 721 1075 1023 775 1235 750 965 960
1210 985 1231 932 850 825 1000 915 1191 1035
1151 630 1175 952 1100 1140 750 1140 1370 1280
Construct a stem-and-leaf display for these data.

Solution Each of the values in the data set contains either three or four digits. We will take the
first digit for three-digit numbers and the first two digits for four-digit numbers as stems. Then
we will use the last two digits of each number as a leaf. Thus for the first value, which is 880, the
stem is 8 and the leaf is 80. The stems for the entire data set are 6, 7, 8, 9, 10, 11, 12, and 13. They
are recorded on the left side of the vertical line in Figure 2.18. The leaves for the numbers are
recorded on the right side.

6 30 Figure 2.18 Stem-and-leaf


7 21 75 50 50 display of rents.
8 80 50 25
9 65 60 85 32 15 52
10 81 75 23 00 35
11 91 51 75 00 40 40
12 35 10 31 80
13 70 ◼

Sometimes a data set may contain too many stems, with each stem containing only a few
leaves. In such cases, we may want to condense the stem-and-leaf display by grouping the stems.
Example 2–10 describes this procedure.

EX AM PLE 2 –10 Number of Hours Spent Working on Computers


by Students
Preparing a grouped
The following stem-and-leaf display is prepared for the number of hours that 25 students spent
stem-and-leaf display.
working on computers during the past month.

0 6
1 1 7 9
2 2 6
3 2 4 7 8
4 1 5 6 9 9
5 3 6 8
6 2 4 4 5 7
7
8 5 6
Mark Harmel/Stone/Getty Images Prepare a new stem-and-leaf display by grouping the stems.

Solution To condense the given stem-and-leaf display, we can combine the first three rows, the
middle three rows, and the last three rows, thus getting the stems 0–2, 3–5, and 6–8. The leaves
for each stem of a group are separated by an asterisk (*), as shown in Figure 2.19. Thus, the leaf
6 in the first row corresponds to stem 0; the leaves 1, 7, and 9 correspond to stem 1; and leaves 2
and 6 belong to stem 2.
0–2 6 * 1 7 9 * 2 6 Figure 2.19 Grouped stem-and-
3–5 2 4 7 8 * 1 5 6 9 9 * 3 6 8 leaf display.
6–8 2 4 4 5 7 * * 5 6
2.3 Stem-and-Leaf Displays 63

If a stem does not contain a leaf, this is indicated in the grouped stem-and-leaf display by two
consecutive asterisks. For example, in the stem-and-leaf display of Figure 2.19, there is no leaf
for 7; that is, there is no number in the 70s. Hence, in Figure 2.19, we have two asterisks after the
leaves for 6 and before the leaves for 8. ◼
Some data sets produce stem-and-leaf displays that have a small number of stems relative to
the number of observations in the data set and have too many leaves for each stem. In such cases,
it is very difficult to determine if the distribution is symmetric or skewed, as well as other char-
acteristics of the distribution that will be introduced in later chapters. In such a situation, we can
create a stem-and-leaf display with split stems. To do this, each stem is split into two or five parts.
Whenever the stems are split into two parts, any observation having a leaf with a value of 0, 1, 2,
3, or 4 is placed in the first split stem, while the leaves 5, 6, 7, 8, and 9 are placed in the second
split stem. Sometimes we can split a stem into five parts if there are too many leaves for one stem.
Whenever a stem is split into five parts, leaves with values of 0 and 1 are placed next to the first
part of the split stem, leaves with values of 2 and 3 are placed next to the second part of the split
stem, and so on. The stem-and-leaf display of Example 2–11 shows this procedure.

E X A MPLE 2–11
Consider the following stem-and-leaf display, which has only two stems. Using the split stem
Stem-and-leaf display with
procedure, rewrite this stem-and-leaf display. split stems.

3 1 1 2 3 3 3 4 4 7 8 9 9 9
4 0 0 0 1 1 1 1 1 1 2 2 2 2 2 3 3 6 6 7

Solution To prepare a split stem-and-leaf display, let us split the two stems, 3 and 4, into two
parts each as shown in Figure 2.20. The first part of each stem contains leaves from 0 to 4, and
the second part of each stem contains leaves from 5 to 9.

3 1 1 2 3 3 3 4 4
3 7 8 9 9 9
4 0 0 0 1 1 1 1 1 1 2 2 2 2 2 3 3
4 6 6 7
Figure 2.20 Split stem-and-leaf display.

In the stem-and-leaf display of Figure 2.20, the first part of stem 4 has a substantial number
of leaves. So, if we decide to split stems into five parts, the new stem-and-leaf display will look
as shown in Figure 2.21.

3 1 1
3 2 3 3 3
3 4 4
3 7
3 8 9 9 9
4 0 0 0 1 1 1 1 1 1
4 2 2 2 2 2 3 3
4
4 6 6 7
Figure 2.21 Split stem-and-leaf display.

There are two important properties to note in the split stem-and-leaf display of Figure 2.21.
The third part of split stem 4 does not have any leaves. This implies that there are no observations
in the data set having a value of 44 or 45. Since there are observations with values larger than 45,
we need to leave an empty part of split stem 4 that corresponds to 44 and 45. Also, there are no
observations with values of 48 or 49. However, since there are no values larger than 47 in the
data, we do not have to write an empty split stem 4 after the largest value. ◼
64 Chapter 2 Organizing and Graphing Data

EXE R CI S E S
CON CE PTS AND PROCEDURES 1, 2, 3, and 4, and the second part should contains the leaves
5, 6, 7, 8, and 9.
2.23 Briefly explain how to prepare a stem-and-leaf display for a c. Which display (the one in part a or the one in part b) provides
data set. You may use an example to illustrate. a better representation of the features of the distribution?
2.24 What advantage does preparing a stem-and-leaf display have over Explain why you believe this.
grouping a data set using a frequency distribution? Give one example.
2.28 The following data give the taxes paid (rounded to thousand
2.25 Consider the following stem-and-leaf display. dollars) in 2014 by a random sample of 30 families.

2–3 18 45 56 * 29 67 83 97 11 17 35 3 15 9 21 13 5 19
4–5 04 27 33 71 * 23 37 51 63 81 92 5 12 8 16 10 8 12 6 14 18
6–8 22 36 47 55 78 89 * * 10 41 8 12 5 3 14 28 38 18 22 15

Write the data set that is represented by this display. a. Prepare a stem-and-leaf display for these data. Arrange the
leaves for each stem in increasing order.
b. Prepare a split stem-and-leaf display for these data. Split each
APPLI CAT IO NS
stem into two parts. The first part should contain the leaves
2.26 The National Highway Traffic Safety Administration collects 0 through 4, and the second part should contain the leaves
data on fatal accidents that occur on roads in the United States. The fol- 5 through 9.
lowing data represent the number of vehicle fatalities for 39 counties in
2.29 The following data give the one-way commuting times (in
South Carolina for 2012 ([Link]/States).
minutes) from home to work for a random sample of 50 workers.
4 48 9 9 31 22 26 17
20 12 6 5 14 9 16 27 23 17 34 26 18 33 46 42 12 37
3 33 9 20 68 13 51 13 44 15 22 19 28 32 18 39 40 48
48 23 12 13 10 15 8 1 16 11 9 24 18 26 31 7 30 15
2 4 17 16 6 52 50 18 22 29 32 30 21 19 14 26 37
25 36 23 39 42 46 29 17 24 31
Prepare a stem-and-leaf display for these data. Arrange the leaves for
Construct a stem-and-leaf display for these data. Arrange the leaves
each stem in increasing order.
for each stem in increasing order.
2.27 The following data give the times (in minutes) taken by 50 stu-
dents to complete a statistics examination that was given a maximum 2.30 The following data give the money (in dollars) spent on text-
time of 75 minutes to finish. books during the Fall 2015 semester by 35 students selected from a
university.
41 28 45 60 53 69 70 50 63 68
37 44 42 38 74 53 66 65 52 64 565 728 870 620 345 868 610 765 550
26 45 66 35 43 44 39 55 64 54 845 530 705 490 258 320 505 957 787
38 52 58 72 67 65 43 65 68 27 617 721 635 438 575 702 538 720 460
64 49 71 75 45 69 56 73 53 72 840 890 560 570 706 430 968 638

a. Prepare a stem-and-leaf display for these data. Arrange the a. Prepare a stem-and-leaf display for these data using the last
leaves for each stem in increasing order. two digits as leaves.
b. Prepare a split stem-and-leaf display for the data. Split each b. Condense the stem-and-leaf display by grouping the stems as
stem into two parts. The first part should contains the leaves 0, 2–4, 5–6, and 7–9.

2.4 Dotplots
One of the simplest methods for graphing and understanding quantitative data is to create a
dotplot. As with most graphs, statistical software should be used to make a dotplot for large data
sets. However, Example 2–12 demonstrates how to create a dotplot by hand.
Dotplots can help us detect outliers (also called extreme values) in a data set. Outliers
are the values that are extremely large or extremely small with respect to the rest of the data
values.
2.4 Dotplots 65

Outliers or Extreme Values Values that are very small or very large relative to the majority of
the values in a data set are called outliers or extreme values.

E X A MP L E 2 –1 2 Ages of Students in a Night Class


A statistics class that meets once a week at night from 7:00 PM to 9:45 PM has 33 students. Creating a dotplot.
The following data give the ages (in years) of these students. Create a dotplot for these data.

34 21 49 37 23 22 33 23 21 20 19
33 23 38 32 31 22 20 24 27 33 19
23 21 31 31 22 20 34 21 33 27 21

Solution To make a dotplot, we perform the following steps.


Step 1. The minimum and maximum values in this data set are 19 and 49 years, respectively.
First, we draw a horizontal line (let us call this the numbers line) with numbers that cover the
given data as shown in Figure 2.22. Note that the numbers line in Figure 2.22 shows the values
from 19 to 49.

19 21 23 25 27 29 31 33 35 37 39 41 43 45 47 49
Figure 2.22 Numbers line.

Step 2. Next we place a dot above the value on the numbers line that represents each of the ages
listed above. For example, the age of the first student is 34 years. So, we place a dot above 34 on
the numbers line as shown in Figure 2.23. If there are two or more observations with the same
value, we stack dots above each other to represent those values. For example, as shown in the data
on ages, two students are 19 years old. We stack two dots (one for each student) above 19 on the
numbers line, as shown in Figure 2.23. After all the dots are placed, Figure 2.23 gives the complete
dotplot.

19 21 23 25 27 29 31 33 35 37 39 41 43 45 47 49
Ages of Students
Figure 2.23 Dotplot for ages of students in a statistics class.

As we examine the dotplot of Figure 2.23, we notice that there are two clusters (groups) of
data. Eighteen of the 33 students (which is almost 55%) are 19 to 24 years old, and 10 of the 33
students (which is about 30%) are 31 to 34 years old. There is one student who is 49 years old
and is an outlier. ◼
66 Chapter 2 Organizing and Graphing Data

EXE R CI S E S
CON CE PTS AND PROCEDURES 23 17 34 26 18 33 46 42 12 37
44 15 22 19 28 32 18 39 40 48
2.31 Briefly explain how to prepare a dotplot for a data set. You may 16 11 9 24 18 26 31 7 30 15
use an example to illustrate. 18 22 29 32 30 21 19 14 26 37
2.32 What are the benefits of preparing a dotplot? Explain. 25 36 23 39 42 46 29 17 24 31
2.33 Create a dotplot for the following data set.
Create a dotplot for these data.
1 2 0 5 1 1 3 2 0 5
2 1 2 1 2 0 1 3 1 2 2.37 The following table, which is based on Consumer Reports
tests and surveys, gives the overall scores (combining road-test and
reliability scores) for 28 brands of vehicles for which they had
APPLI CAT IO NS enough data (USA Today, February 25, 2015). Create a dotplot for
these data.
2.34 The National Highway Traffic Safety Administration collects
data on fatal accidents that occur on roads in the United States. The fol-
lowing data represent the number of vehicle fatalities for 39 counties in
Brand Overall Score Brand Overall Score
South Carolina for 2012 ([Link]/States).
4 48 9 9 31 22 26 17 Acura 65 Kia 68
20 12 6 5 14 9 16 27 Audi 73 Lexus 78
3 33 9 20 68 13 51 13 Buick 69 Lincoln 59
48 23 12 13 10 15 8 1
Cadillac 58 Mazda 75
2 4 17 16 6 52 50
Chevrolet 59 MBW 66
Make a dotplot for these data. Chrysler 54 Mercedes-Benz 56
2.35 The following data give the times (in minutes) taken by 50 stu- Dodge 52 MiniCooper 46
dents to complete a statistics examination that was given a maximum Fiat 32 Nissan 59
time of 75 minutes to finish.
Ford 53 Porsche 70
41 28 45 60 53 69 70 50 63 68 GMC 61 Scion 54
37 44 42 38 74 53 66 65 52 64
Honda 69 Subaru 73
26 45 66 35 43 44 39 55 64 54
38 52 58 72 67 65 43 65 68 27 Hyundai 64 Toyota 74
64 49 71 75 45 69 56 73 53 72 Infiniti 59 Volkswagen 60
Jeep 39 Volvo 65
Create a dotplot for these data.
2.36 The following data give the one-way commuting times (in
minutes) from home to work for a random sample of 50 workers.

USES AND MISUSES...

G R AP H I C A LLY S PE AKI NG … Number of


Imagine a high school in Middle America, if you will. We shall call it Sport Tickets Sold
Central High School. In Central High there are a fair number of very
talented athletes who play soccer, baseball, football, and basketball. Soccer 23
At the annual fundraiser for the school, the cheer squad sells advance Baseball 30
tickets to the home games so that the community can be entertained Football 63
by the local school athletes.
Basketball 59
As it happens, the cheer squad was able to sell a total of 175
tickets at this year’s fundraising event. The ticket sales for the various In order to show the results of the ticket sales to the Central High
sports were as follows: community, some of the students got together and decided to make
Glossary 67

an attractive graph to display the ticket sales in the local paper. As A more conservative approach would have been to use a stand-
dutiful students, they carefully investigated types of charts that might ard bar chart. The following bar graph was constructed using the
work, and ultimately decided to arrange the graph by sport on the same information as in the pictogram above.
horizontal-axis and the number of tickets sold on the vertical-axis. To
make the graph a bit more vivid, they embellished it using a pictogram
(that is, a graph made with pictures). They used the following graph to
70
publish in the local paper.

Number of tickets sold


60

50
Number of tickets sold

70
60 40
50
30
40
30 Football Basketball 20
20
Baseball 10
10 Soccer
0 0
Soccer Baseball Football Basketball
Sport
Sport
The graph was very well received by the readers of the local paper,
and was very attractive to the eye. Further, the football and basket- Obviously, this graph does not have the same eye appeal as the
ball teams and their supporters were happy to see just how much one reported in the newspaper; however, it is much clearer to an
more popular their respective sports were compared to soccer and observer. It is obvious from this graph that the number of soccer tick-
baseball. However, there is a major problem with the graph. This ets and the number of baseball tickets were quite close, as were the
graph had the intent of showing the number of tickets sold. This number of football and the number of basketball tickets. It is also clear
means that we should focus on the height of the balls for each sport, that the number of football tickets is about three times larger than the
which, when read horizontally across to the vertical-axis, do in fact number of soccer tickets and that the number of football tickets is
match the number of tickets reported to have been sold in the table about two times larger than the number of baseball tickets. These rela-
above. But the eye is easily deceived. What is immediately clear is tionships were much more difficult to discern in the pictogram.
that we tend to focus on the relative sizes of the balls, not the heights In 1954, Darrell Huff wrote a fascinating and entertaining book
that they represent. In that context, it seems that a huge number of titled How to Lie with Statistics (Huff, Darrell, How to Lie with Statis-
basketball tickets were sold compared to the other sports because tics, 1954, New York: W. W. Norton). That book contains many illus-
the eye focuses immediately on the circumference or the area of the trations of how graphs can be used to deceive—either intentionally or
circle represented by the basketball, which appears much larger than unintentionally—those doing a visual interpretation of data. As Mr. Huff
the other balls. In fact, comparing soccer to basketball, the area of states in his book, “The crooks already know these tricks; honest
the basketball in the graph is 6.6 times larger than that of the soccer men must learn them in self-defense.” In that regard, it is sometimes
ball, whereas the tickets sold for basketball were about 2.6 times useful to learn how not to do something in order to understand the
more than the soccer tickets sold. importance of doing it correctly.

Glossary
Bar graph A graph made of bars whose heights represent the fre- Cumulative frequency The frequency of a class that includes all values
quencies of respective categories. in a data set that fall below the upper boundary or limit of that class.
Class An interval that includes all the values in a (quantitative) data Cumulative frequency distribution A table that lists the total number
set that fall within two numbers, the lower and upper limits of the class. of values that fall below the upper boundary or limit of each class.
Class boundary The lower and upper numbers of a class interval Cumulative percentage The cumulative relative frequency multi-
in less-than method. plied by 100.
Class frequency The number of values in a data set that belong to Cumulative relative frequency The cumulative frequency of a
a certain class. class divided by the total number of observations.
Class midpoint or mark The class midpoint or mark is obtained Frequency distribution A table that lists all the categories or
by dividing the sum of the lower and upper limits (or boundaries) of classes and the number of values that belong to each of these catego-
a class by 2. ries or classes.
Class width or size The difference between the two boundaries of a class Grouped data A data set presented in the form of a frequency
or the difference between the lower limits of two consecutive classes. distribution.
68 Chapter 2 Organizing and Graphing Data

Histogram A graph in which classes are marked on the horizontal Relative frequency The frequency of a class or category divided by
axis and frequencies, relative frequencies, or percentages are marked the sum of all frequencies.
on the vertical axis. The frequencies, relative frequencies, or percent- Skewed-to-the-left histogram A histogram with a longer tail on
ages of various classes are represented by the heights of bars that are the left side.
drawn adjacent to each other.
Skewed-to-the-right histogram A histogram with a longer tail on
Outliers or Extreme values Values that are very small or very large the right side.
relative to the majority of the values in a data set.
Stem-and-leaf display A display of data in which each value is
Pareto chart A bar graph in which bars are arranged in decreasing divided into two portions—a stem and a leaf.
order of heights. Symmetric histogram A histogram that is identical on both sides
Percentage The percentage for a class or category is obtained by of its central point.
multiplying the relative frequency of that class or category by 100. Ungrouped data Data containing information on each member of a
Pie chart A circle divided into portions that represent the relative sample or population individually.
frequencies or percentages of different categories or classes. Uniform or rectangular histogram A histogram with the same
Polygon A graph formed by joining the midpoints of the tops of frequency for all classes.
successive bars in a histogram by straight lines.
Raw data Data recorded in the sequence in which they are collected
and before they are processed.

Supplementary Exercises
2.38 The following data give the political party of each of the first a. Construct a frequency distribution table. Take 32 as the lower
30 U.S. presidents. In the data, D stands for Democrat, DR for Demo- limit of the first class and 6 as the class width.
cratic Republican, F for Federalist, R for Republican, and W for b. Calculate the relative frequency and percentage for each class.
Whig. c. Construct a histogram for the frequency distribution of part a.
d. On what percentage of these 40 days did this student send 44
F F DR DR DR DR D D W W or more text messages?
D W W D D R D R R R e. Prepare the cumulative frequency, cumulative relative fre-
R D R D R R R D R R quency, and cumulative percentage distributions.
2.41 The following data give the number of orders received for a
a. Prepare a frequency distribution table for these data. sample of 30 hours at the Timesaver Mail Order Company.
b. Calculate the relative frequency and percentage distributions.
c. Draw a bar graph for the relative frequency distribution and a
34 44 31 52 41 47 38 35 32 39
pie chart for the percentage distribution.
28 24 46 41 49 53 57 33 27 37
d. Make a Pareto chart for the frequency distribution.
30 27 45 38 34 46 36 30 47 50
e. What percentage of these presidents were Whigs?
2.39 The following data give the number of television sets owned by a. Construct a frequency distribution table. Take 23 as the lower
40 randomly selected households. limit of the first class and 7 as the width of each class.
b. Calculate the relative frequencies and percentages for all
1 1 2 3 2 4 1 3 2 1 classes.
3 0 2 1 2 3 2 3 2 2 c. For what percentage of the hours in this sample was the number
1 2 1 1 1 3 1 1 1 2 of orders more than 36?
2 4 2 3 1 3 1 2 2 4 d. Prepare the cumulative frequency, cumulative relative fre-
quency, and cumulative percentage distributions.
a. Prepare a frequency distribution table for these data using 2.42 The following data give the amounts (in dollars) spent on
single-valued classes. refreshments by 30 spectators randomly selected from those who
b. Compute the relative frequency and percentage distributions. patronized the concession stands at a recent Major League Baseball
c. Draw a bar graph for the frequency distribution. game.
d. What percentage of the households own two or more televi-
sion sets?
4.95 27.99 8.00 5.80 4.50 2.99 4.85 6.00
2.40 The following data give the number of text messages sent on 9.00 15.75 9.50 3.05 5.65 21.00 16.60 18.00
40 randomly selected days during 2015 by a high school student: 21.77 12.35 7.75 10.45 3.85 28.45 8.35 17.70
19.50 11.65 11.45 3.00 6.55 16.50
32 33 33 34 35 36 37 37 37 37
38 39 40 41 41 42 42 42 43 44 a. Construct a frequency distribution table using the less-than
44 45 45 45 47 47 47 47 47 48 method to write classes. Take $0 as the lower boundary of the
48 49 50 50 51 52 53 54 59 61 first class and $6 as the width of each class.
Advanced Exercises 69

b. Calculate the relative frequencies, and percentages for all classes. c. Draw a histogram for the frequency distribution.
c. Draw a histogram for the frequency distribution. d. Make the cumulative frequency, cumulative relative frequency,
d. Prepare the cumulative frequency, cumulative relative fre- and cumulative percentages distributions.
quency, and cumulative percentage distributions. 2.44 The following data give the number of text messages sent on
2.43 The following table lists the average one-way commuting 40 randomly selected days during 2015 by a high school student.
times (in minutes) from home to work for 30 metropolitan areas around
the world with more than 1 million residents (urbandemographics. 32 33 33 34 35 36 37 37 37 37
[Link]). 38 39 40 41 41 42 42 42 43 44
44 45 45 45 47 47 47 47 47 48
Average Average 48 49 50 50 51 52 53 54 59 61
Commute Length Commute Length
City (minutes) City (minutes) Prepare a stem-and-leaf display for these data.

Barcelona 24.2 Paris 33.7 2.45 The following data give the number of orders received for a
sample of 30 hours at the Timesaver Mail Order Company.
Belém 31.5 Paulo 42.8
Belo Horizonte 34.4 Porto Alegre 27.7 34 44 31 52 41 47 38 35 32 39
28 24 46 41 49 53 57 33 27 37
Berlin 31.6 Recife 34.9
30 27 45 38 34 46 36 30 47 50
Boston 28.9 Rio de Janeiro 42.6
Brasilia - DF 34.8 Salvador 33.9 Prepare a stem-and-leaf display for these data.
Chicago 30.7 San Francisco 28.7 2.46 The following data give the number of text messages sent on 40
Curitiba 32.1 Santiago 27.0 randomly selected days during 2015 by a high school student.
Fortaleza 31.7 Seattle 26.9 32 33 33 34 35 36 37 37 37 37
London 37.0 Shanghai 50.4 38 39 40 41 41 42 42 42 43 44
44 45 45 45 47 47 47 47 47 48
Los Angeles 28.1 Stockholm 35.0
48 49 50 50 51 52 53 54 59 61
Madrid 33.0 Sydney 34.0
Milan 26.7 Tokyo 34.5 Create a dotplot for these data.
Montréal 31.0 Toronto 33.0 2.47 The following data give the number of orders received for a
sample of 30 hours at the Timesaver Mail Order Company.
New York 34.6 Vancouver 30.0
34 44 31 52 41 47 38 35 32 39
a. Prepare a frequency distribution table for these data using the 28 24 46 41 49 53 57 33 27 37
less-than method to write classes. Use 22 minutes as the lower 30 27 45 38 34 46 36 30 47 50
boundary of the first class and 6 minutes as the class width.
b. Calculate the relative frequencies and percentages for all classes. Create a dotplot for these data.

Advanced Exercises
2.48 The following frequency distribution table gives the age distri- c. How can you change the frequency distribution so that the
bution of drivers who were at fault in auto accidents that occurred resulting histogram gives a clearer picture?
during a 1-week period in a city. 2.49 Suppose a data set contains the ages of 135 autoworkers ranging
from 20 to 53 years.
Age (years) f a. Using Sturge’s formula given in footnote 1 in section 2.2.2,
find an appropriate number of classes for a frequency distribu-
18 to less than 20 7 tion for this data set.
20 to less than 25 12 b. Find an appropriate class width based on the number of
classes in part a.
25 to less than 30 18
30 to less than 40 14 2.50 Stem-and-leaf displays can be used to compare distributions
for two groups using a back-to-back stem-and-leaf display. In such a
40 to less than 50 15 display, one group is shown on the left side of the stems, and the other
50 to less than 60 16 group is shown on the right side. When the leaves are ordered, the
60 and over 35 leaves increase as one moves away from the stems. The following
stem-and-leaf display shows the money earned per tournament entered
for the top 30 money winners in the 2008–09 Professional Bowlers
a. Draw a relative frequency histogram for this table. Association men’s tour and for the top 21 money winners in the 2008–
b. In what way(s) is this histogram misleading? 09 Professional Bowlers Association women’s tour.
70 Chapter 2 Organizing and Graphing Data

Women’s Men’s d. Does either of the tours appears to have any outliers? If so,
what are the earnings levels for these players?
8 0
2.51 Statisticians often need to know the shape of a population to
8871 1 make inferences. Suppose that you are asked to specify the shape of
65544330 2 334456899 the population of weights of all college students.
840 3 03344678 a. Sketch a graph of what you think the weights of all college
students would look like.
52 4 011237888
b. The following data give the weights (in pounds) of a random
21 5 9 sample of 44 college students (F and M indicate female and
6 9 male, respectively).
5 7
8 7 123 F 195 M 138 M 115 F 179 M 119 F 148 F 147 F
9 5 180 M 146 F 179 M 189 M 175 M 108 F 193 M 114 F
179 M 147 M 108 F 128 F 164 F 174 M 128 F 159 M
The leaf unit for this display is 100. In other words, the data used 193 M 204 M 125 F 133 F 115 F 168 M 123 F 183 M
represent the earnings in hundreds of dollars. For example, for the 116 F 182 M 174 M 102 F 123 F 99 F 161 M 162 M
women’s tour, the first number is 08, which is actually 800. The 155 F 202 M 110 F 132 M
second number is 11, which actually is 1100.
a. Do the top money winners, as a group, on one tour (men’s or
women’s) tend to make more money per tournament played i. Construct a stem-and-leaf display for these data.
than on the other tour? Explain how you can come to this con- ii. Can you explain why these data appear the way they do?
clusion using the stem-and-leaf display. c. Construct a back-to-back stem-and-leaf display for the data
b. What would be a typical earnings level amount per tourna- on weights, placing the weights of the female students to
ment played for each of the two tours? the left of the stems and those of the male students to the
c. Do the data appear to have similar spreads for the two tours? right of the stems. Does one gender tend to have higher
Explain how you can come to this conclusion using the stem- weights than the other? Explain how you know this from the
and-leaf display. display.

Self-Review Test
1. Briefly explain the difference between ungrouped and grouped 4. Thirty-six randomly selected senior citizens were asked if their
data and give one example of each type. net worth is more than or less than $200,000. Their responses are
given below, where M stands for more, L represents less, N means do
2. The following table gives the frequency distribution of times (to
not know or do not want to tell.
the nearest hour) that 90 fans spent waiting in line to buy tickets to a
rock concert.
M M L L L N M N L L M M
L L N L L M M L N L L L
Waiting Time M M L M L L M M L N L N
(hours) Frequency
a. Make a frequency distribution for these 36 responses.
0 to 6 5
b. Using the frequency distribution of part a, prepare the relative
7 to 13 27 frequency and percentage distributions.
14 to 20 30 c. What percentage of these senior citizens claim to have net
21 to 27 20 worth less than $200,000?
d. Make a bar graph for the frequency distribution.
28 to 34 8 e. Make a Pareto chart for the frequency distribution.
f. Draw a pie chart for the percentage distribution.
Circle the correct answer in each of the following statements, which
are based on this table. 5. Forty-eight randomly selected car owners were asked about their
a. The number of classes in the table is 5, 30, 90. typical monthly expense on gas. The following data show the responses
b. The class width is 6, 7, 34. of these 48 car owners.
c. The midpoint of the third class is 16.5, 17, 17.5.
d. The lower boundary of the second class is 6.5, 7, 7.5. $210 160 430 255 176 135 221 359 380 405 391 477
e. The upper limit of the second class is 12.5, 13, 13.5. 333 209 267 121 357 87 167 95 347 487 302 545
f. The sample size is 5, 90, 11. 351 256 492 277 245 367 159 187 253 287 456 64
g. The relative frequency of the second class is .22, .41, .30. 76 166 304 444 193 479 188 148 53 327 234 110

3. Briefly explain and illustrate with the help of graphs a symmetric a. Construct a frequency distribution table. Use the classes
histogram, a histogram skewed to the right, and a histogram skewed to 50–149, 150–249, 250–349, 350–449, and 450–549.
the left. b. Calculate the relative frequency and percentage for each class.
Technology Instructions 71

c. Construct a histogram for the percentage distribution made in d. What percentage of these customers spent less than $140 at
part b. this grocery store?
d. What percentage of the car owners in this sample spend $350 e. Make a histogram for the frequency distribution.
or more on gas per month? f. Prepare the cumulative frequency, cumulative relative frequency,
e. Prepare the cumulative frequency, cumulative relative frequency, and cumulative percentage distributions using the table of part a.
and cumulative percentage distributions using the table of part a.
7. Construct a stem-and-leaf display for the following data, which
6. Thirty customers from all customers who shopped at a large give the times (in minutes) that 24 customers spent waiting to speak
grocery store during a given week were randomly selected and their to a customer service representative when they called about problems
shopping expenses was noted. The following data show the expenses with their Internet service provider.
(in dollars) of these 30 customers.
12 15 7 29 32 16 10 14 17 8 19 21
89.20 145.23 191.54 45.36 67.98 123.67 4 14 22 25 18 6 22 16 13 16 12 20
187.57 56.43 102.45 158.76 134.07 67.49
8. Consider this stem-and-leaf display:
212.60 165.75 46.50 111.25 64.54 23.67
135.09 193.46 87.65 133.76 156.28 88.64 3 0 3 7
65.90 120.50 55.45 91.54 153.20 44.39 4 2 4 6 7 9
5 1 3 3 6
a. Prepare a frequency distribution table for these data. Use the 6 0 7 7
classes as 20 to less than 60, 60 to less than 100, 100 to less 7 1 9
than 140, 140 to less than 180, and 180 to less than 220.
Write the data set that was used to construct this display.
b. What is the width of each class in part a?
c. Using the frequency distribution of part a, prepare the relative 9. Make a dotplot for the data on typical monthly expenses on gas
frequency and percentage distributions. for the 48 car owners given in Problem 5.

Mini-Projects
Note: The Mini-Projects are located on the text’s Web site, [Link]/college/mann.

Decide for Yourself


Note: The Decide for Yourself feature is located on the text’s Web site, [Link]/college/mann.

TE C HNO L OGY
Chapter 2
IN S TRU C TION S
Note: Complete TI-84, Minitab, and Excel manuals are available for download at the textbook’s Web site, [Link]/college/mann.

TI-84 Color/TI-84
The TI-84 Color Technology Instructions feature of this text is written for the TI-84 Plus C
color graphing calculator running the 4.0 operating system. Some screens, menus, and func-
tions will be slightly different in older operating systems. The TI-84 and TI-84 Plus can perform
all of the same functions but will not have the “Color” option referenced in some of the menus.
TI-84 calculators do not have the options/capabilities to make the bar graph, pie chart, stem-
and-leaf display, and dotplot.

Creating a Frequency Histogram for Example 2–3 of the Text


1. Enter the data from Example 2–3 of the text into L1. (See Screen 2.1.)
2. Select 2nd > Y = (the STAT PLOT menu).
3. If more than one plot is turned on, select PlotsOff and then press ENTER to turn off the
plots. Screen 2.1
72 Chapter 2 Organizing and Graphing Data

4. From the STAT PLOT menu, select Plot 1. Use the following settings (see Screen 2.2):
• Select On to turn the plot on.
• At the Type prompt, select the third icon (histogram).
• At the Xlist prompt, enter L1 by pressing 2nd > 1.
• At the Freq prompt, enter 1.
• Select BLUE at the Color prompt.
Screen 2.2
(Skip this step if you do not have the TI-84 Color calculator.)
5. Select ZOOM > ZoomStat to display the histogram. This function will automatically
choose window settings and class boundaries for the histogram. (See Screen 2.3.)
6. Press TRACE and use the left and right arrow keys to see the class boundaries and
frequencies for each class. (See Screen 2.4.)
7. To manually change the window settings and the class boundaries, press WINDOW and use
the following settings:
Screen 2.3 • Type 600 at the Xmin prompt. Xmin is the extreme left value of the window and the
lower boundary value of the first class.
• Type 2700 at the Xmax prompt. Xmax is the extreme right value of the window.
• Type 300 at the Xscl prompt. Xscl is the class width.
• Type -3 at the Ymin prompt. Ymin is the extreme bottom value of the window.
• Type 15 at the Ymax prompt. Ymax is the extreme top value of the window.
• Type 3 at the Yscl prompt. Yscl is the distance between tick marks on the y-axis.
Screen 2.4 • Press GRAPH to see the histogram with the new settings.

Minitab
The Minitab Technology Instructions feature of this text is written for Minitab version 17. Some
screens, menus, and functions will be slightly different in older versions of Minitab.

Creating a Bar Graph for Example 2–1 of the Text


Bar Graph for Ungrouped Data
1. Enter the 30 data values of Example 2–1 of the text into column C1.
2. Select Graph > Bar Chart.
3. Use the following settings in the resulting dialog box:
• Select Counts of unique values, select Simple, and click OK.
4. Use the following settings in the new dialog box:
• Type C1 in the Categorical variables box, and click OK.
5. The bar graph will appear in a new window. (See Screen 2.5.)
Screen 2.5

Bar Graph for Grouped Data


1. Enter the grouped data from Example 2–1 as shown in Table 2.4 of the text. Put the Donut
Variety into column C1 and the Frequency into column C2.
Technology Instructions 73

2. Select Graph > Bar Chart.


3. Use the following settings in the resulting dialog box:
• Select Values from a table, select Simple, and click OK.
4. Use the following settings in the new dialog box:
• Type C2 in the Graph variables box, C1 in the Categorical variable box, and click OK.
5. The bar graph will appear in a new window. (See Screen 2.5.)
To make a Pareto chart, after step 4 above, click on Chart Options in the same dialog box.
Select Decreasing Y in the next dialog box. Click OK in both dialog boxes. The Pareto chart
will appear in a new window.

Creating a Pie Chart for Example 2–1 of the Text


Pie Chart for Ungrouped Data
1. Enter the 30 data values of Example 2–1 of the text into column C1.
2. Select Graph > Pie Chart.
3. Use the following settings in the resulting dialog box:
• Select Chart counts of unique values, and type C1 in the Categorical
variables box.
• To label slices of the pie chart with percentages or category names, select
Labels > Slice Labels. Choose the options you prefer and click OK.
4. The pie chart will appear in a new window. (See Screen 2.6.) Screen 2.6

Pie Chart for Grouped Data


1. Enter the grouped data from Example 2–1 as shown in Table 2.4 of the text. Put the Donut
Variety into column C1 and the Frequency into column C2.
2. Select Graph > Pie Chart.
3. Use the following settings in the resulting dialog box:
• Select Chart values from a table, type C1 in the Categorical variable box, and type C2
in the Summary variables box.
• To label slices of the pie chart with percentages or category names, select Labels > Slice
Labels. Choose the options you prefer, and click OK.
4. The pie chart will appear in a new window. (See Screen 2.6.)

Creating a Frequency Histogram for Example 2–3 of the Text


1. Enter the 30 data values of Example 2–3 of the text into column C1.
2. Select Graph > Histogram, select Simple, and click OK.
3. Use the following settings in the resulting dialog box:
• Type C1 in the Graph variables box.
• If you wish to add labels, titles, or subtitles to the histogram, select Labels and choose the
options you prefer, and click OK.
74 Chapter 2 Organizing and Graphing Data

4. The histogram will appear in a new window. (See Screen 2.7.)


5. To change the class boundaries double-click on any of the bars in the
histogram and the Edit Bars dialog box will appear. Click on the
Binning tab. Now choose one of the following two options:
• To set the class midpoints, select Midpoint under Interval Type,
select Midpoint/Cutpoint positions under the Interval Definition
box, and then type the class midpoints (separate each value with a
space). Click OK.
• To set the class boundaries, select Cutpoint under Interval Type,
select Midpoint/Cutpoint positions, and then type the left boundary
for each interval and the right endpoint of the last interval (separate
Screen 2.7 each value with a space). Click OK.

Creating a Stem-and-Leaf Display for Example 2–8 of the Text


1. Enter the 30 data values of Example 2–8 of the text into column C1.
2. Select Graph > Stem-and-Leaf.
3. Type C1 in the Graph variables box and click OK.
4. The stem-and-leaf display will appear in the Session window. (See Screen 2.8.)
5. If there are too many stems, you can specify an Increment for each branch of the
stem-and leaf display in the box next to Increment of the above dialog box. For example,
the stem-and-leaf display shown in Screen 2.8 has an increment of size 5. In other words,
the stem-and-leaf display in Screen 2.8 is a split stem-and-leaf display with each stem
split in two.

Screen 2.8

Creating a Dotplot for Example 2–12 of the Text


1. Enter the 33 data values from Example 2–12 of the text into column C1.
2. Name the column Age.
3. Select Graph > Dotplot.
4. Select Simple from the One Y box and click OK.
5. Type C1 in the Graph variables box and click OK.
6. The dotplot will appear in a new window. (See Screen 2.9.)

Screen 2.9

Excel
The Excel Technology Instructions feature of this text is written for Excel 2013. Some screens,
menus, and functions will be slightly different in older versions of Excel. To enable some of the
advanced Excel functions, you must enable the Data Analysis Add-In for Excel.
Excel does not have the capability to make a stem-and-leaf display and dotplot.
Technology Instructions 75

Creating a Bar Graph for Example 2–1 of the Text


1. Enter the grouped data from Example 2–1 as shown in Table 2.4 of the text.
2. Highlight the cells that contain the donut varieties and the frequencies. (See
Screen 2.10.)
3. Click INSERT and then click the Insert Column Chart icon in the Charts group.
4. Select the first icon in the 2–D Column group in the dropdown box.
5. A new graph will appear with the bar graph.

Screen 2.10

Creating a Pie Chart for Example 2–1 of the Text


1. Enter the grouped data from Example 2–1 as shown in Table 2.4 of the text.
2. Highlight the cells that contain the donut varieties and the frequencies. (See Screen 2.10.)
3. Click INSERT and then click the Insert Pie Chart icon in the Charts group.
4. Select the first icon in the 2–D Pie group in the dropdown box.
5. A new graph will appear with the pie chart. (See Screen 2.11.)
Note: You may click in the text box and change “Chart Title” to an appropriate title or you may select the
text box and delete it.

Creating a Frequency Histogram for Example 2–3 of the Text Screen 2.11

1. Enter the data from Example 2–3 of the text into cells A1 to A30.
2. Determine the class boundaries and type the right boundary for each class in cells B1 to B8.
In this example, we will use the right boundary values of 750, 1000, 1250, 1500, 1750, 2000,
2250, and 2500, which are different from the ones used in Example 2–3.
3. Click DATA and then click Data Analysis from the Analysis group.
4. Select Histogram from the Data Analysis dialog box and click OK.
5. Click in the Input Range box and then highlight cells A1 to A30.
6. Click in the Bin Range box and then highlight cells B1 to B8.
7. In the Output options box, select New Worksheet Ply and check the Chart Output check
box. Click OK.
8. A new worksheet will appear with the frequency distribution and the corresponding histogram.
(See Screen 2.12.)

Screen 2.12
76 Chapter 2 Organizing and Graphing Data

TECHNOLOGY ASSIGNMENTS
TA2.1 In the past few years, many states have built casinos, and a. Make a histogram for these data.
many more are in the process of doing so. Forty adults were asked if b. Prepare a stem-and-leaf display for these data.
building casinos is good for society. Following are the responses of
c. Make a dotplot for these data.
these adults, where G stands for good, B indicates bad, and I means
indifferent or no answer. TA2.7 Refer to Data Set VII on McDonald’s menu items that accom-
B G B B I G B I B B panies this text (see Appendix A). Create a dotplot for the sodium
G B B G B B B G G I content of the menu items listed in column 6.
B G B B I G G G B B TA2.8 Refer to Data Set X on Major League Baseball hitting that
I G B B B G G B B G accompanies this text (see Appendix A). Construct a dotplot for the
number of home runs hit listed in column 7.
a. Construct a bar graph and pie chart for these data.
b. Arrange the bars of the bar graph of part a in decreasing order TA2.9 Refer to Data Set VII on McDonald’s menu items that accom-
to prepare a Pareto chart. panies this text (see Appendix A).

TA2.2 Refer to Data Set XII on coffeemaker ratings that accompa- a. Construct a histogram for the data on total carbs listed in
nies this text (see Appendix A). Create a bar graph and pie chart for column 5, allowing technology to choose the width of the
the variable listed in Column 7. intervals for you.
b. Construct another histogram for the data on total carbs listed
TA2.3 Refer to Data Set V on the highest grossing movies of 2014
in column 5, but use intervals that are half the width of the
that accompanies this text (see Appendix A). Construct a histogram
histogram in part a.
for the opening weekend gross earnings listed in Column 5.
c. Construct another histogram for the data on total carbs listed
TA2.4 Refer to Data Set I on the prices of various products in differ- in column 5, but use intervals that are twice the width of the
ent cities across the United States that accompanies this text (see histogram in part a.
Appendix A). Construct a histogram for the prices of 1 gallon of regu-
d. Comment on which histogram you believe gives the best
lar unleaded gas listed in column 13.
understanding of the data and why you feel this way.
TA2.5 Refer to Data Set XII on coffeemaker ratings that accompa-
TA2.10 Refer to Data Set IV on the Manchester Road Race that
nies this text (see Appendix A). Create a stem-and-leaf display for the
accompanies this text (see Appendix A).
overall scores of coffeemakers listed in column 3.
a. Create a bar graph for the gender of the runners listed in
TA2.6 The following data give the one-way commuting times (in
column 6.
minutes) from home to work for a random sample of 50 workers.
b. Create a histogram for the net time to complete the race listed
23 17 34 26 18 33 46 42 12 37 in column 3.
44 15 22 19 28 32 18 39 40 48
16 11 9 24 18 26 31 7 30 15
18 22 29 32 30 21 19 14 26 37
25 36 23 39 42 46 29 17 24 31

You might also like