Click to edit Master title style
Quantitative
Methods
Module 1.3
Click to
Some Common
edit Master
Distribution
title style
Shapes
Symmetrical
Skewed to the right Skewed to the left when the right and left
when the left tail of the tail of the histogram
when the right tail of the
histogram is longer than histogram is longer than appears to be mirror
the left tail the right tail images of each other
Click to edit Master title style
Frequency Polygon,
Cumulative Distribution,
Ogive
Click to editPolygons
Frequency Master title style
another graphical display that can be used to depict a frequency
distribution
How to Construct a Frequency Polygon?
1. We plot a point above each class midpoint at a height equal to the
frequency of the class – the height can also be the class relative
frequency or class percent frequency if so desired.
2. Then we connect the points with line segments.
Click to edit
Example: Comparing
Master title
Twostyle
Grade Distributions
The table below lists the scores earned on the first exam by the 40
students in a statistics course.
Because exam scores are often reported by using 10-point grade
ranges (for instance, 80 to 90 percent), we define the following
classes:
30 < 40, 40 < 50, 50 < 60, 60 < 70, 70 < 80, 80 < 90, 90 < 100
Classes Class Frequenc Percent
Midpoint y
30 < 40 35 1 1 ÷ 40 ⋅ 100 = 2.5%
40 < 50 45 1 2.5%
Method of left inclusion
50 < 60 55 3 7.50%
- include the left boundary point in
60 < 70 65 13 32.50%
the class and not the right
70 < 80 75 3 7.50%
boundary point in the class
80 < 90 85 9 22.50%
Example:
90 < 100 95 10 25%
The score 50 will be included in the
Total: 40 Total: 100%
third class (50-60).
Click to edit
Example: Comparing
Master title
Twostyle
Grade Distributions
A Percent Frequency Polygon of the Exam Scores
Classes Class Percent
Midpoint
30 < 40 35 2.5%
40 < 50 45 2.5%
50 < 60 55 7.50%
60 < 70 65 32.50%
70 < 80 75 7.50%
80 < 90 85 22.50%
90 < 100 95 25%
Total:
100%
Click to editDistribution
Cumulative Master title and
styleOgives
Constructing a Cumulative Distribution
• Record for each class the number of measurements that are
less than the upper boundary of the class
Example: Payment Time Data
Click to edit Master title style
A Frequency Distribution, Cumulative Frequency Distribution, Cumulative Relative Frequency
Distribution, and Cumulative Percent Frequency Distribution for the Payment Time Data
Column (3) gives the cumulative frequency for each class.
The cumulative frequency for class 10 < 13 is the number of payment time less than 13.
The cumulative frequency for class 13 < 16 is the number of payment time less than 16. (3 + 14 = 17)
The cumulative frequency for class 16 < 19 is the number of payment time less than 19. (3 + 14 + 23 = 40)
⋮
Click to edit Master title style
Column (4) gives the cumulative relative frequency for each class, which is obtained by summing the
relative frequencies of all classes representing values less than the upper boundary of the class.
For example, the cumulative relative frequency for class 13 < 16 is 0.2615. (That is 17/65.)
Column (5) gives the cumulative percent frequency for each class, which is obtained by summing the
percent frequencies of all classes representing values less than the upper boundary of the class.
For example, the cumulative percent frequency for class 10 < 13 is 4.62%. (That is .0462 × 100.)
Sample Interpretation:
60 of the 65 payment times are 24 days or less, or, equivalently, 92.31 percent of the payment times (or a
fraction of .9231 of the payment times) are 24 days or less.
Click to edit Master title style
Ogive
A Percent Frequency Ogive of the
• An ogive is a graph of a cumulative Payment Times
distribution.
How to Construct a Frequency Ogive:
• Plot the points above each upper class
boundary at a height equal to the cumulative
frequency of the class.
• Connect the plotted points with line segments
Click to edit Master title style
Dotplot
ClickPlots
Dot to edit Master title style
Scores Earned on the First Exam by
the 40 students in a Statistics Course
• a very simple graph that can be used to summarize
a data set
How to make a Dot Plot?
1. Draw a horizontal axis that spans the range of the
measurements in the data set.
2. Place dots above the horizontal axis to represent
the measurements.
Dot Plot of Scores on Exam 1
The two dots above the score of 90 tells
us that two students received 90 on the
exam.
Click to edit Master title style
• Dot plots are useful for detecting outliers, which are unusually large or small observations
that are well separated from the remaining observations.
• How we handle an outlier depends on its cause. If the outlier results from a measurement
error or an error in recording or processing the data, it should be corrected. If such an
outlier cannot be corrected, it should be discarded. If an outlier is not the result of an error
in measuring or recording the data, its cause may reveal important information.
• For example, the dot plot for exam 1 indicates that the score 32 seems unusually low. This
outlying exam score of 32 convinced the instructor that the student needed a tutor.
Click to edit Master title style
Stem and Leaf
Displays
Click to edit Master
Stem-and-Leaf Displays
title style
• This kind of graph places the measurements in order
from smallest to largest, and allows the analyst to
simultaneously see all of the measurements in the data
set and see the shape of the data set’s distribution.
Click to edit Master title style
Example: A Sample of 50 Mileages for a New Midsize Model
29 8
30 8 1 843558776
31 7630384940127457489564
32 10451518437273
33 30
29 8
30 13455677888
31 0012334444455667778899
32 01112334455778
33 03
Click to edit Master title style
• When constructing a stem-and-leaf display, there are no rules that
dictate the number of stem values (rows) that should be used.
• If you feel that the display has collapsed the mileages too closely
together, we can stretch the display by assigning each set of leading
digits to two or more rows. This is called splitting the stems.
• In the previous example,
Click to edit Master title style
Contingency
Tables
Click to edit Master title style
• Crosstabulation is a process that classifies data on two
dimensions. This process results in a table that is called
a contingency table.
• This table consists of rows and columns – the rows
classify the data according to one dimension and the
columns classify the data according to a second
dimension. Together, the rows and columns represent
all possibilities (or contingencies).
Example
Click to edit Master title style
The Brokerage Firm Case: Studying Client Satisfaction
An investment broker sells several kinds of investment
products—a stock fund, a bond fund, and a tax-deferred annuity.
The broker wishes to study whether client satisfaction with its
products and services depends on the type of investment product
purchased. To do this, 100 of the broker’s clients are randomly
selected from the population of clients who have purchased
shares in exactly one of the funds. The broker records the fund
type purchased by each client and has one of its investment
counselors personally contact the client. When contacted, the
client is asked to rate his or her level of satisfaction with the
purchased fund as high, medium, or low. The resulting data are
given in the next slide.
Results of a Customer Satisfaction Survey Given to 100
Click to edit Master title style
Randomly Selected Clients Who Invest in One of Three Fund
Types—a Bond Fund, a Stock Fund, or a Tax-Deferred Annuity
A Contingency Table of Fund Type versus
Level of Client Satisfaction
A Contingency Table of Fund Type versus
Click to edit Master title style Level of Client Satisfaction
• The classification categories for the two variables In the table above, these frequencies
are defined along the left and top margins of the tell us that 15 clients invested in the
table. bond fund and reported a high level
• The three row labels—bond fund, stock fund, and of satisfaction, 4 clients invested in
tax deferred annuity— define the three fund the stock fund and reported a
categories and are given in the left table margin. medium level of satisfaction, and so
• The three column labels— high, medium, and low— forth.
define the three levels of client satisfaction and are
given along the top table margin. • The row totals provide a frequency
• Each row and column combination, that is, each distribution for the different fund types.
fund type and level of satisfaction combination, • The column totals provide a frequency
defines what we call a “cell” in the table. distribution for the different satisfaction
• The counts in the cells are called the cell levels
frequencies.
Click to edit Master title style
By dividing the row totals by the total of 100 clients surveyed, we can obtain relative
frequencies; and by multiplying each relative frequency by 100, we can obtain percent
frequencies. That is, we can obtain the frequency, relative frequency, and percent frequency
distributions for fund type as follows:
We see that 30 percent of the clients
invested in the bond fund, 30 percent
invested in the stock fund, and 40 percent
invested in the tax deferred annuity.
By dividing the column totals by the total of 100 clients surveyed, we can obtain
relative frequencies, and by multiplying each relative frequency by 100, we can
obtain percent frequencies. That is, we can obtain the frequency, relative frequency,
and percent frequency distributions for level of satisfaction as follows:
We see that 40 percent of all clients
reported high satisfaction, 40 percent
reported medium satisfaction, and 20
percent reported low satisfaction.
Click to edit Master title style
• One good way to investigate relationships (such as fund type
and level of satisfaction) is to compute row percentages and
column percentages.
Row Percentages for Each Fund Type
We see that 50 percent of bond fund investors report high satisfaction, while 40
percent of these investors report medium satisfaction, and only 10 percent
report low satisfaction.
Click to edit Master title style
Scatter Plots
Click to edit Master title style
• A simple graph that can be used to study the relationship between
two variables is called a scatter plot.
• A scatter plot can reveal various kinds of relationships.
A Positive Linear Relationship Little or No Linear Relationship A Negative Linear Relationship
Example:
Click to edit
Values
Master
of Advertising
title style Expenditure and Sales
Volumes for Ten Sales Regions
Suppose that a marketing manager wishes to investigate the relationship between the
sales volume (in thousands of units) of a product and the amount spent (in units of
$10,000) on advertising the product. To do this, the marketing manager randomly
selects 10 sales regions having equal sales potential. The manager assigns a different
level of advertising expenditure for January 2011 to each sales region as shown in the
table below.
Values of Advertising Expenditure (in $10,000s) and
Sales Volume (in 1000s) for Ten Sales Regions
ClickTo to edit
construct theMaster title
scatterplot, we place style
the variable advertising
expenditure (denoted 𝑥) on the horizontal axis and we place the
variable sales volume (denoted 𝑦) on the vertical axis.
A Scatter Plot of Sales Volume versus
Advertising Expenditure
Conclusion:
The scatter plot shows that there is a positive
relationship between advertising expenditure
and sales volume—that is, higher values of
sales volume are associated with higher levels
of advertising expenditure.