0% found this document useful (0 votes)
6 views117 pages

Chapter 1 Statistics

Chapter 1 of 'Statistics for Business and Economics' focuses on describing data through graphical methods. It covers key concepts such as population vs. sample, parameter vs. statistic, and the distinction between descriptive and inferential statistics, along with various types of data and levels of measurement. The chapter also includes practical applications using Rstudio to create and interpret various types of graphs for categorical and numerical data.

Uploaded by

saheliano.halleb
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views117 pages

Chapter 1 Statistics

Chapter 1 of 'Statistics for Business and Economics' focuses on describing data through graphical methods. It covers key concepts such as population vs. sample, parameter vs. statistic, and the distinction between descriptive and inferential statistics, along with various types of data and levels of measurement. The chapter also includes practical applications using Rstudio to create and interpret various types of graphs for categorical and numerical data.

Uploaded by

saheliano.halleb
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Statistics for

Business and Economics

Chapter 1

Describing Data: Graphical

Ch. 1-1
Chapter Goals
After completing this chapter, you should be able to:
 Explain how decisions are often based on incomplete information
 Explain key definitions:
 Population vs. Sample

 Parameter vs. Statistic

 Descriptive vs. Inferential Statistics

 Describe random sampling


 Explain the difference between Descriptive and Inferential statistics
 Identify types of data and levels of measurement
 Applications with Rstudio:
Ch. 1-2
Chapter Goals
(continued)
 Create and interpret graphs to describe categorical
variables:
 frequency distribution, bar chart, pie chart, Pareto diagram
 Create a line chart to describe time-series data
 Create and interpret graphs to describe numerical
variables:
 frequency distribution, histogram, ogive, stem-and-leaf display
 Construct and interpret graphs to describe relationships
between variables:
 Scatter plot, cross table
 Describe appropriate and inappropriate ways to display
data graphically
Ch. 1-3
1.1
Dealing with Uncertainty

Everyday decisions are based on incomplete information

Consider:

 Will the job market be strong when I graduate?


 Will the price of Yahoo stock be higher in six months than
it is now?
 Will interest rates remain low for the rest of the year if
the federal budget deficit is as high as predicted?

Ch. 1-4
Dealing with Uncertainty
(continued)

Numbers and data are used to assist decision making

 Statistics is a tool to help process, summarize, analyze,


and interpret data

Ch. 1-5
1.2
Key Definitions

 A population is the collection of all items of interest or


under investigation
 N represents the population size

 A sample is an observed subset of the population


 n represents the sample size

 A parameter is a specific characteristic of a population


 A statistic is a specific characteristic of a sample

Ch. 1-6
Population vs. Sample

Population Sample

a b cd b c
ef gh i jk l m n gi n
o p q rs t u v w o r u
x y z y

Values calculated using Values computed from


population data are called sample data are called
parameters statistics Ch. 1-7
Examples of Populations

 Names of all registered voters in the United States

 Incomes of all families living in Daytona Beach

 Annual returns of all stocks traded on the New York Stock Exchange
 Grade point averages of all the students in your university

Ch. 1-8
Random Sampling

Simple random sampling is a procedure in which

 each member of the population is chosen strictly by


chance,
 each member of the population is equally likely to be
chosen,
 every possible sample of n objects is equally likely to be
chosen

The resulting sample is called a random sample

Ch. 1-9
Descriptive and Inferential Statistics

Two branches of statistics:


 Descriptive statistics
 Graphical and numerical procedures to summarize and process data

 Inferential statistics
 Using data to make predictions, forecasts, and estimates to assist decision making

Ch. 1-10
Descriptive Statistics

 Collect data
 e.g., Survey

 Present data
 e.g., Tables and graphs

 Summarize data
X i
 e.g., Sample mean = n
Ch. 1-11
Inferential Statistics
◼ Estimation
◼ e.g., Estimate the population
mean weight using the sample
mean weight
◼ Hypothesis testing
◼ e.g., Test the claim that the
population mean weight is 140
pounds

Inference is the process of drawing conclusions or


making decisions about a population based on
sample results Ch. 1-12
Types of Data

Data

Categorical Numerical

Examples:
◼ Marital Status
◼ Are you registered to Discrete Continuous
vote?
◼ Eye Color
(Defined categories or Examples:
groups) Examples: ◼ Weight
◼ Number of Children ◼ Voltage
◼ Defects per hour (Measured characteristics)
Ch. 1-13
(Counted items)
Measurement Levels

Differences between
measurements, true Ratio Data
zero exists
Quantitative Data

Differences between
measurements but no Interval Data
true zero

Ordered Categories
(rankings, order, or Ordinal Data
scaling)
Qualitative Data

Categories (no
ordering or direction) Nominal Data
Ch. 1-14
Graphical
1.3
Presentation of Data

 Data in raw form are usually not easy to use for decision making

 Some type of organization is needed

Table
Graph

 The type of graph to use depends on the variable being summarized

Ch. 1-15
Graphical
Presentation of Data
(continued)
 Techniques reviewed in this chapter:

Categorical Numerical
Variables Variables

• Frequency distribution • Line chart


• Bar chart • Frequency distribution
• Pie chart • Histogram and ogive
• Pareto diagram • Stem-and-leaf display
• Scatter plot

Ch. 1-16
Tables and Graphs for
Categorical Variables

Categorical
Data

Tabulating Data Graphing Data

Frequency
Distribution Bar Pie Pareto
Table Chart Chart Diagram

Ch. 1-17
The Frequency
Distribution Table
Summarize data by category

Example: Hospital Patients by Unit


Hospital Unit Number of Patients

Cardiac Care 1,052


Emergency 2,245
Intensive Care 340
Maternity 552
Surgery 4,630
(Variables are
categorical) Ch. 1-18
Bar and Pie Charts

 Bar charts and Pie charts are often used for qualitative
(category) data

 Height of bar or size of pie slice shows the frequency or


percentage for each category

Ch. 1-19
Bar Chart Example

Hospital Number
Unit of Patients

Cardiac Care 1,052


Emergency 2,245 Hospital Patients by Unit
5000
Intensive Care 340
Maternity 552 4000

patients per year


Surgery 4,630

Number of
3000

2000

1000

Cardiac

Emergency

Intensive

Surgery
Maternity
Care

Care
Ch. 1-20
Pie Chart Example

Hospital Number % of
Unit of Patients Total
Hospital Patients by Unit
Cardiac Care 1,052 11.93
Emergency 2,245 25.46 Cardiac Care
12%
Intensive Care 340 3.86
Maternity 552 6.26
Surgery 4,630 52.50

Emergency
Surgery 25%
53%

Intensive Care
(Percentages 4%
are rounded to Maternity
the nearest 6%
Ch. 1-21
percent)
Pareto Diagram

 Used to portray categorical data

 A bar chart, where categories are shown in descending order of frequency

 A cumulative polygon is often shown in the same graph


 Used to separate the “vital few” (20%) from the “trivial many” (80%)

Ch. 1-22
WHEN TO USE A PARETO CHART

 When analyzing data about the frequency of problems or causes in a process


 When there are many problems or causes and you want to focus on the most
significant
 When analyzing broad causes by looking at their specific components

23
Pareto Diagram Example
Example: 400 defective items are examined
for cause of defect:

Source of
Manufacturing Error Number of defects
Bad Weld 34
Poor Alignment 223
Missing Part 25
Paint Flaw 78
Electrical Short 19
Cracked case 21
Total 400
Ch. 1-24
Pareto Diagram Example
(continued)

Step 1: Sort by defect cause, in descending order


Step 2: Determine % in each category

Source of
Manufacturing Error Number of defects % of Total Defects
Poor Alignment 223 55.75
Paint Flaw 78 19.50
Bad Weld 34 8.50
Missing Part 25 6.25
Cracked case 21 5.25
Electrical Short 19 4.75
Ch. 1-25
Total 400 100%
Pareto Diagram Example
(continued)
Step 3: Show results graphically
Pareto Diagram: Cause of Manufacturing Defect
60% 100%
% of defects in each category

90%

cumulative % (line graph)


50%
80%

70%
(bar graph)

40%

60%

30% 50%

40%

20%
30%

20%
10%

10%

0% 0%
Poor Alignment Paint Flaw Bad Weld Missing Part Cracked case Ch. 1-26
Electrical Short
1.4
Graphs for Time-Series Data

 A line chart (time-series plot) is used to show the values of a variable over
time

 Time is measured on the horizontal axis

 The variable of interest is measured on the vertical axis

Ch. 1-27
Line Chart Example

Magazine Subscriptions by Year

350

300
Thousands of subscribers

250

200

150

100

50

0
1990

1991

1992

1993

1994

1995

1996

1997

1998

1999

2000

2001

2002

2003

2004

2005

2006
Ch. 1-28
1.5 Graphs to Describe
Numerical Variables

Numerical Data

Frequency Distributions Stem-and-Leaf


and Display
Cumulative Distributions

Histogram Ogive

Ch. 1-29
Frequency Distributions

What is a Frequency Distribution?


 A frequency distribution is a list or a table …
 containing class groupings (categories or ranges within which the data
fall) ...
 and the corresponding frequencies with which data fall within each class
or category

Ch. 1-30
Why Use Frequency Distributions?

 A frequency distribution is a way to summarize data


 The distribution condenses the raw data into a more useful
form...
 and allows for a quick visual interpretation of the data

Ch. 1-31
Class Intervals
and Class Boundaries

 Each class grouping has the same width


 Determine the width of each interval by

largest number − smallest number


w = interval width =
number of desired intervals

◼ Use at least 5 but no more than 15-20 intervals


◼ Intervals never overlap
◼ Round up the interval width to get desirable
interval endpoints
Ch. 1-32
Frequency Distribution Example

Example: A manufacturer of insulation randomly selects 20 winter days and


records the daily high temperature

24, 35, 17, 21, 24, 37, 26, 46, 58, 30,

32, 13, 12, 38, 41, 43, 44, 27, 53, 27

Ch. 1-33
Frequency Distribution Example
(continued)

 Sort raw data in ascending order:


12, 13, 17, 21, 24, 24, 26, 27, 27, 30, 32, 35, 37, 38, 41, 43, 44, 46,
53, 58

 Find range: 58 - 12 = 46
 Select number of classes: 5 (usually between 5 and 15)
 Compute interval width: 10 (46/5 then round up)
 Determine interval boundaries: 10 but less than 20, 20 but less than 30, . .
. , 60 but less than 70
 Count observations & assign to classes

Ch. 1-34
Frequency Distribution Example
(continued)
Data in ordered array:
12, 13, 17, 21, 24, 24, 26, 27, 27, 30, 32, 35, 37, 38, 41, 43, 44, 46, 53, 58

Relative
Interval Frequency Percentage
Frequency
10 but less than 20 3 .15 15
20 but less than 30 6 .30 30
30 but less than 40 5 .25 25
40 but less than 50 4 .20 20
50 but less than 60 2 .10 10
Total 20 1.00 100
Ch. 1-35
Histogram

 A graph of the data in a frequency distribution is called a histogram


 The interval endpoints are shown on the horizontal axis
 the vertical axis is either frequency, relative frequency, or percentage
 Bars of the appropriate heights are used to represent the number of
observations within each class

Ch. 1-36
Histogram Example

Interval Frequency
Histogram : Daily High Tem perature
10 but less than 20 3
20 but less than 30 6 7 6
30 but less than 40 5
40 but less than 50 4
6 5
50 but less than 60 2 5 4

Frequency
4 3
3 2
2
1 0 0
(No gaps 0
between 0 0 1010 2020 30 30 40 40Ch.50 50 60 60 70
1-37
bars) Temperature in Degrees
Histograms in Excel

1
2
Select Data Tab
Click on Data Analysis

Ch. 1-38
Histograms in Excel
(continued)

3
Choose Histogram

(
Input data range and bin
range (bin range is a cell
4 range containing the upper
interval endpoints for each class
grouping)

Select Chart Output


and click “OK”
Ch. 1-39
Questions for Grouping Data
into Intervals

 1. How wide should each interval be?


(How many classes should be used?)

 2. How should the endpoints of the intervals


be determined?
 Often answered by trial and error, subject to
user judgment
 The goal is to create a distribution that is
neither too "jagged" nor too "blocky”
 Goal is to appropriately show the pattern of
variation in the data
Ch. 1-40
How Many Class Intervals?

 Many (Narrow class intervals) 3.5


3
 may yield a very jagged distribution with gaps from 2.5
empty classes

Frequency
2
1.5
 Can give a poor indication of how frequency varies 1
across classes 0.5
0

4
8
12
16
20
24
28
32
36
40
44
48
52
56
60
More
Few (Wide class intervals)
Temperature

12
 may compress variation too much and yield a blocky 10
distribution 8

Frequency
 can obscure important patterns of variation. 6
4

0
0 30 60 More
Temperature
(X axis labels are upper class endpoints)
Ch. 1-41
The Cumulative
Frequency Distribuiton

Data in ordered array:


12, 13, 17, 21, 24, 24, 26, 27, 27, 30, 32, 35, 37, 38, 41, 43, 44, 46, 53, 58

Cumulative Cumulative
Class Frequency Percentage
Frequency Percentage

10 but less than 20 3 15 3 15


20 but less than 30 6 30 9 45
30 but less than 40 5 25 14 70
40 but less than 50 4 20 18 90
50 but less than 60 2 10 20 100
Total 20 100 Ch. 1-42
The Ogive
Graphing Cumulative Frequencies
Upper
interval Cumulative
Interval endpoint Percentage
Less than 10 10 0
10 but less than 20 20 15
20 but less than 30 30 45 Ogive: Daily High Temperature
30 but less than 40 40 70
40 but less than 50 50 90 100

Cumulative Percentage
50 but less than 60 60 100
80
60
40
20
0
10 20 30 40 50 60
Interval endpoints
Ch. 1-43
Stem-and-Leaf Diagram

 A simple way to see distribution details in a data set

METHOD: Separate the sorted data series


into leading digits (the stem)
and
the trailing digits (the leaves)

Ch. 1-44
Example

Data in ordered array:


21, 24, 24, 26, 27, 27, 30, 32, 38, 41

 Here, use the 10’s digit for the stem unit:

Stem Leaf
◼ 21 is shown as 2 1
◼ 38 is shown as 3 8

Ch. 1-45
Example
(continued)
Data in ordered array:
21, 24, 24, 26, 27, 27, 30, 32, 38, 41

 Completed stem-and-leaf diagram:

Stem Leaves
2 1 4 4 6 7 7
3 0 2 8
4 1

Ch. 1-46
Using other stem units

 Using the 100’s digit as the stem:

 Round off the 10’s digit to form the leaves

 613 would become 6 1Stem Leaf


 776 would become 7 8
 ...

 1224 becomes 12 2

Ch. 1-47
Using other stem units
(continued)

 Using the 100’s digit as the stem:


 The completed stem-and-leaf display:
Data:
Stem Leaves
613, 632, 658, 717, 6 136
722, 750, 776, 827, 7 2258
841, 859, 863, 891, 8 346699
894, 906, 928, 933,
9 13368
955, 982, 1034,
1047,1056, 1140, 10 356
1169, 1224 11 47
Ch. 1-48
12 2
1.6 Relationships Between Variables

 Graphs illustrated so far have involved only a single variable


 When two variables exist other techniques are used:

Categorical Numerical
(Qualitative) (Quantitative)
Variables Variables

Cross tables Scatter plots


Ch. 1-49
Scatter Diagrams

 Scatter Diagrams are used for paired


observations taken from two
numerical variables
 The Scatter Diagram:

 one variable is measured on the


vertical axis and the other variable is
measured on the horizontal axis

Ch. 1-50
Scatter Diagram Example

Volume Cost per


Cost per Day vs. Production Volume
per day day
23 125 250
26 140
200

Cost per Day


29 146
150
33 160
38 167 100
42 170 50
50 188
0
55 195
0 10 20 30 40 50 60 70
60 200
Volume per Day
Ch. 1-51
Scatter Diagrams in Excel

1 Select the Insert tab


2 Select Scatter type from
the Charts section

3 When prompted, enter the data range, desired legend, and


desired destination to complete the scatter diagram Ch. 1-52
Cross Tables

 Cross Tables (or contingency tables) list the number of observations for
every combination of values for two categorical or ordinal variables

 If there are r categories for the first variable (rows) and c categories
for the second variable (columns), the table is called an r x c cross
table

Ch. 1-53
Cross Table Example

 4 x 3 Cross Table for Investment Choices by Investor


(values in $1000’s)
Investment Investor A Investor B Investor C Total
Category
Stocks 46.5 55 27.5 129
Bonds 32.0 44 19.0 95
CD 15.5 20 13.5 49
Savings 16.0 28 7.0 51
Total 110.0 147 67.0 324
Ch. 1-54
Graphing
Multivariate Categorical Data
(continued)

 Side by side bar charts

C o m p arin g In vesto rs

S avings

CD

B onds

S toc k s

0 10 20 30 40 50 60

Inves tor A Inves tor B Inves tor C


Ch. 1-55
Side-by-Side Chart Example
 Sales by quarter for three sales territories:
1st Qtr 2nd Qtr 3rd Qtr 4th Qtr
East 20.4 27.4 59 20.4
West 30.6 38.6 34.6 31.6
North 45.9 46.9 45 43.9

60

50

40
East
30 West
North
20

10

0
1st Qtr 2nd Qtr 3rd Qtr 4th Qtr Ch. 1-56
1.7
Data Presentation Errors

Goals for effective data presentation:


 Present data to display essential information

 Communicate complex ideas clearly and accurately

 Avoid distortion that might convey the wrong message

Ch. 1-57
Data Presentation Errors
(continued)

 Unequal histogram interval widths


 Compressing or distorting the vertical axis
 Providing no zero point on the vertical axis
 Failing to provide a relative basis in comparing data
between groups

Ch. 1-58
Chapter Summary

 Reviewed incomplete information in decision making


 Introduced key definitions:
 Population vs. Sample
 Parameter vs. Statistic
 Descriptive vs. Inferential statistics
 Described random sampling
 Examined the decision making process

Ch. 1-59
Chapter Summary
(continued)
 Reviewed types of data and measurement levels
 Data in raw form are usually not easy to use for decision
making -- Some type of organization is needed:
 Table  Graph

 Techniques reviewed in this chapter :


◼ Line chart
 Frequency distribution ◼ Frequency distribution
◼ Histogram and ogive
 Bar chart
◼ Stem-and-leaf display
 Pie chart ◼ Scatter plot
 Pareto diagram ◼ Cross tables and
side-by-side bar charts
Ch. 1-60
Chapter Goals
(continued)
After completing this chapter, you should be able to:
 Create and interpret graphs to describe categorical
variables:
 frequency distribution, bar chart, pie chart, Pareto diagram
 Create a line chart to describe time-series data
 Create and interpret graphs to describe numerical
variables:
 frequency distribution, histogram, ogive, stem-and-leaf display
 Construct and interpret graphs to describe relationships
between variables:
 Scatter plot, cross table
 Describe appropriate and inappropriate ways to display
data graphically
Ch. 1-61
1.1
Dealing with Uncertainty

Everyday decisions are based on incomplete information

Consider:

 Will the job market be strong when I graduate?


 Will the price of Yahoo stock be higher in six months than
it is now?
 Will interest rates remain low for the rest of the year if
the federal budget deficit is as high as predicted?

Ch. 1-62
Dealing with Uncertainty
(continued)

Numbers and data are used to assist decision making

 Statistics is a tool to help process, summarize, analyze,


and interpret data

Ch. 1-63
1.2
Key Definitions

 A population is the collection of all items of interest or


under investigation
 N represents the population size

 A sample is an observed subset of the population


 n represents the sample size

 A parameter is a specific characteristic of a population


 A statistic is a specific characteristic of a sample

Ch. 1-64
Population vs. Sample

Population Sample

a b cd b c
ef gh i jk l m n gi n
o p q rs t u v w o r u
x y z y

Values calculated using Values computed from


population data are called sample data are called
parameters statistics Ch. 1-65
Examples of Populations

 Names of all registered voters in the United States

 Incomes of all families living in Daytona Beach

 Annual returns of all stocks traded on the New York Stock Exchange
 Grade point averages of all the students in your university

Ch. 1-66
Random Sampling

Simple random sampling is a procedure in which

 each member of the population is chosen strictly by


chance,
 each member of the population is equally likely to be
chosen,
 every possible sample of n objects is equally likely to be
chosen

The resulting sample is called a random sample

Ch. 1-67
Descriptive and Inferential Statistics

Two branches of statistics:


 Descriptive statistics
 Graphical and numerical procedures to summarize and process data

 Inferential statistics
 Using data to make predictions, forecasts, and estimates to assist decision making

Ch. 1-68
Descriptive Statistics

 Collect data
 e.g., Survey

 Present data
 e.g., Tables and graphs

 Summarize data
X i
 e.g., Sample mean = n
Ch. 1-69
Inferential Statistics
◼ Estimation
◼ e.g., Estimate the population
mean weight using the sample
mean weight
◼ Hypothesis testing
◼ e.g., Test the claim that the
population mean weight is 140
pounds

Inference is the process of drawing conclusions or


making decisions about a population based on
sample results Ch. 1-70
Types of Data

Data

Categorical Numerical

Examples:
◼ Marital Status Discrete Continuous
◼ Are you registered to
vote?
◼ Eye Color Examples: Examples:
(Defined categories or ◼ Number of Children ◼ Weight
groups) ◼ Defects per hour ◼ Voltage
(Counted items) (Measured characteristics)
Ch. 1-71
Measurement Levels

Differences between
measurements, true Ratio Data
zero exists
Quantitative Data

Differences between
measurements but no Interval Data
true zero

Ordered Categories
(rankings, order, or Ordinal Data
scaling)
Qualitative Data

Categories (no
ordering or direction) Nominal Data
Ch. 1-72
Graphical
1.3
Presentation of Data

 Data in raw form are usually not easy to use for decision making

 Some type of organization is needed

Table
Graph

 The type of graph to use depends on the variable being summarized

Ch. 1-73
Graphical
Presentation of Data
(continued)
 Techniques reviewed in this chapter:

Categorical Numerical
Variables Variables

• Frequency distribution • Line chart


• Bar chart • Frequency distribution
• Pie chart • Histogram and ogive
• Pareto diagram • Stem-and-leaf display
• Scatter plot

Ch. 1-74
Tables and Graphs for
Categorical Variables

Categorical
Data

Tabulating Data Graphing Data

Frequency
Distribution Bar Pie Pareto
Table Chart Chart Diagram

Ch. 1-75
The Frequency
Distribution Table
Summarize data by category

Example: Hospital Patients by Unit


Hospital Unit Number of Patients

Cardiac Care 1,052


Emergency 2,245
Intensive Care 340
Maternity 552
Surgery 4,630
(Variables are
categorical) Ch. 1-76
Bar and Pie Charts

 Bar charts and Pie charts are often used for qualitative
(category) data

 Height of bar or size of pie slice shows the frequency or


percentage for each category

Ch. 1-77
Bar Chart Example

Hospital Number
Unit of Patients

Cardiac Care 1,052


Emergency 2,245 Hospital Patients by Unit
5000
Intensive Care 340
Maternity 552 4000

patients per year


Surgery 4,630

Number of
3000

2000

1000

Cardiac

Emergency

Intensive

Surgery
Maternity
Care

Care
Ch. 1-78
Pie Chart Example

Hospital Number % of
Unit of Patients Total
Hospital Patients by Unit
Cardiac Care 1,052 11.93
Emergency 2,245 25.46 Cardiac Care
12%
Intensive Care 340 3.86
Maternity 552 6.26
Surgery 4,630 52.50

Emergency
Surgery 25%
53%

Intensive Care
(Percentages 4%
are rounded to Maternity
the nearest 6%
Ch. 1-79
percent)
Pareto Diagram

 Used to portray categorical data

 A bar chart, where categories are shown in descending order of frequency

 A cumulative polygon is often shown in the same graph


 Used to separate the “vital few” from the “trivial many”

Ch. 1-80
Pareto Diagram Example
Example: 400 defective items are examined
for cause of defect:

Source of
Manufacturing Error Number of defects
Bad Weld 34
Poor Alignment 223
Missing Part 25
Paint Flaw 78
Electrical Short 19
Cracked case 21
Total 400
Ch. 1-81
Pareto Diagram Example
(continued)

Step 1: Sort by defect cause, in descending order


Step 2: Determine % in each category

Source of
Manufacturing Error Number of defects % of Total Defects
Poor Alignment 223 55.75
Paint Flaw 78 19.50
Bad Weld 34 8.50
Missing Part 25 6.25
Cracked case 21 5.25
Electrical Short 19 4.75
Ch. 1-82
Total 400 100%
Pareto Diagram Example
(continued)
Step 3: Show results graphically
Pareto Diagram: Cause of Manufacturing Defect
60% 100%
% of defects in each category

90%

cumulative % (line graph)


50%
80%

70%
(bar graph)

40%

60%

30% 50%

40%

20%
30%

20%
10%

10%

0% 0%
Poor Alignment Paint Flaw Bad Weld Missing Part Cracked case Ch. 1-83
Electrical Short
1.4
Graphs for Time-Series Data

 A line chart (time-series plot) is used to show the values of a variable over
time

 Time is measured on the horizontal axis

 The variable of interest is measured on the vertical axis

Ch. 1-84
Line Chart Example

Magazine Subscriptions by Year

350

300
Thousands of subscribers

250

200

150

100

50

0
1990

1991

1992

1993

1994

1995

1996

1997

1998

1999

2000

2001

2002

2003

2004

2005

2006
Ch. 1-85
1.5 Graphs to Describe
Numerical Variables

Numerical Data

Frequency Distributions Stem-and-Leaf


and Display
Cumulative Distributions

Histogram Ogive

Ch. 1-86
Frequency Distributions

What is a Frequency Distribution?


 A frequency distribution is a list or a table …
 containing class groupings (categories or ranges within which the data
fall) ...
 and the corresponding frequencies with which data fall within each class
or category

Ch. 1-87
Why Use Frequency Distributions?

 A frequency distribution is a way to summarize data


 The distribution condenses the raw data into a more useful
form...
 and allows for a quick visual interpretation of the data

Ch. 1-88
Class Intervals
and Class Boundaries

 Each class grouping has the same width


 Determine the width of each interval by

largest number − smallest number


w = interval width =
number of desired intervals

◼ Use at least 5 but no more than 15-20 intervals


◼ Intervals never overlap
◼ Round up the interval width to get desirable
interval endpoints
Ch. 1-89
Frequency Distribution Example

Example: A manufacturer of insulation randomly selects 20 winter days and


records the daily high temperature

24, 35, 17, 21, 24, 37, 26, 46, 58, 30,

32, 13, 12, 38, 41, 43, 44, 27, 53, 27

Ch. 1-90
Frequency Distribution Example
(continued)

 Sort raw data in ascending order:


12, 13, 17, 21, 24, 24, 26, 27, 27, 30, 32, 35, 37, 38, 41, 43, 44, 46,
53, 58

 Find range: 58 - 12 = 46
 Select number of classes: 5 (usually between 5 and 15)
 Compute interval width: 10 (46/5 then round up)
 Determine interval boundaries: 10 but less than 20, 20 but less than 30, . .
. , 60 but less than 70
 Count observations & assign to classes

Ch. 1-91
Frequency Distribution Example
(continued)
Data in ordered array:
12, 13, 17, 21, 24, 24, 26, 27, 27, 30, 32, 35, 37, 38, 41, 43, 44, 46, 53, 58

Relative
Interval Frequency Percentage
Frequency
10 but less than 20 3 .15 15
20 but less than 30 6 .30 30
30 but less than 40 5 .25 25
40 but less than 50 4 .20 20
50 but less than 60 2 .10 10
Total 20 1.00 100
Ch. 1-92
Histogram

 A graph of the data in a frequency distribution is called a histogram


 The interval endpoints are shown on the horizontal axis
 the vertical axis is either frequency, relative frequency, or percentage
 Bars of the appropriate heights are used to represent the number of
observations within each class

Ch. 1-93
Histogram Example

Interval Frequency
Histogram : Daily High Tem perature
10 but less than 20 3
20 but less than 30 6 7 6
30 but less than 40 5
40 but less than 50 4
6 5
50 but less than 60 2 5 4

Frequency
4 3
3 2
2
1 0 0
(No gaps 0
between 0 0 1010 2020 30 30 40 40Ch.50 50 60 60 70
1-94
bars) Temperature in Degrees
Histograms in Excel

1
2
Select Data Tab
Click on Data Analysis

Ch. 1-95
Histograms in Excel
(continued)

3
Choose Histogram

(
Input data range and bin
range (bin range is a cell
4 range containing the upper
interval endpoints for each class
grouping)

Select Chart Output


and click “OK”
Ch. 1-96
Questions for Grouping Data
into Intervals

 1. How wide should each interval be?


(How many classes should be used?)

 2. How should the endpoints of the intervals be


determined?

 Often answered by trial and error, subject to


user judgment
 The goal is to create a distribution that is
neither too "jagged" nor too "blocky”
 Goal is to appropriately show the pattern of
variation in the data
Ch. 1-97
How Many Class Intervals?

 Many (Narrow class intervals) 3.5


3
 may yield a very jagged distribution with gaps from 2.5
empty classes

Frequency
2
1.5
 Can give a poor indication of how frequency varies 1
across classes 0.5
0

4
8
12
16
20
24
28
32
36
40
44
48
52
56
60
More
Few (Wide class intervals)
Temperature

12
 may compress variation too much and yield a blocky 10
distribution 8

Frequency
 can obscure important patterns of variation. 6
4

0
0 30 60 More
Temperature
(X axis labels are upper class endpoints)
Ch. 1-98
The Cumulative
Frequency Distribuiton

Data in ordered array:


12, 13, 17, 21, 24, 24, 26, 27, 27, 30, 32, 35, 37, 38, 41, 43, 44, 46, 53, 58

Cumulative Cumulative
Class Frequency Percentage
Frequency Percentage

10 but less than 20 3 15 3 15


20 but less than 30 6 30 9 45
30 but less than 40 5 25 14 70
40 but less than 50 4 20 18 90
50 but less than 60 2 10 20 100
Total 20 100 Ch. 1-99
The Ogive
Graphing Cumulative Frequencies
Upper
interval Cumulative
Interval endpoint Percentage
Less than 10 10 0
10 but less than 20 20 15
20 but less than 30 30 45 Ogive: Daily High Temperature
30 but less than 40 40 70
40 but less than 50 50 90 100

Cumulative Percentage
50 but less than 60 60 100
80
60
40
20
0
10 20 30 40 50 60
Ch. 1-
Interval endpoints 100
Stem-and-Leaf Diagram

 A simple way to see distribution details in a data set

METHOD: Separate the sorted data series


into leading digits (the stem)
and
the trailing digits (the leaves)

Ch. 1-
101
Example

Data in ordered array:


21, 24, 24, 26, 27, 27, 30, 32, 38, 41

 Here, use the 10’s digit for the stem unit:

Stem Leaf
◼ 21 is shown as 2 1
◼ 38 is shown as 3 8

Ch. 1-
102
Example
(continued)
Data in ordered array:
21, 24, 24, 26, 27, 27, 30, 32, 38, 41

 Completed stem-and-leaf diagram:

Stem Leaves
2 1 4 4 6 7 7
3 0 2 8
4 1

Ch. 1-
103
Using other stem units

 Using the 100’s digit as the stem:

 Round off the 10’s digit to form the leaves

 613 would become 6 1Stem Leaf


 776 would become 7 8
 ...

 1224 becomes 12 2

Ch. 1-
104
Using other stem units
(continued)

 Using the 100’s digit as the stem:


 The completed stem-and-leaf display:
Data:
Stem Leaves
613, 632, 658, 717, 6 136
722, 750, 776, 827, 7 2258
841, 859, 863, 891, 8 346699
894, 906, 928, 933,
9 13368
955, 982, 1034,
1047,1056, 1140, 10 356
1169, 1224 11 47
Ch. 1-
12 2 105
1.6 Relationships Between Variables

 Graphs illustrated so far have involved only a single variable


 When two variables exist other techniques are used:

Categorical Numerical
(Qualitative) (Quantitative)
Variables Variables

Cross tables Scatter plots


Ch. 1-
106
Scatter Diagrams

 Scatter Diagrams are used for paired


observations taken from two
numerical variables
 The Scatter Diagram:

 one variable is measured on the


vertical axis and the other variable is
measured on the horizontal axis

Ch. 1-
107
Scatter Diagram Example

Volume Cost per


Cost per Day vs. Production Volume
per day day
23 125 250
26 140
200

Cost per Day


29 146
150
33 160
38 167 100
42 170 50
50 188
0
55 195
0 10 20 30 40 50 60 70
60 200
Volume per Day
Ch. 1-
108
Scatter Diagrams in Excel

1 Select the Insert tab


2 Select Scatter type from
the Charts section

3 When prompted, enter the data range, desired legend, and


desired destination to complete the scatter diagram Ch. 1-
109
Cross Tables

 Cross Tables (or contingency tables) list the number of observations for
every combination of values for two categorical or ordinal variables

 If there are r categories for the first variable (rows) and c categories
for the second variable (columns), the table is called an r x c cross
table

Ch. 1-
110
Cross Table Example

 4 x 3 Cross Table for Investment Choices by Investor


(values in $1000’s)
Investment Investor A Investor B Investor C Total
Category
Stocks 46.5 55 27.5 129
Bonds 32.0 44 19.0 95
CD 15.5 20 13.5 49
Savings 16.0 28 7.0 51
Total 110.0 147 67.0 324
Ch. 1-
111
Graphing
Multivariate Categorical Data
(continued)

 Side by side bar charts

C o m p arin g In vesto rs

S avings

CD

B onds

S toc k s

0 10 20 30 40 50 60

Inves tor A Inves tor B Inves tor C


Ch. 1-
112
Side-by-Side Chart Example
 Sales by quarter for three sales territories:
1st Qtr 2nd Qtr 3rd Qtr 4th Qtr
East 20.4 27.4 59 20.4
West 30.6 38.6 34.6 31.6
North 45.9 46.9 45 43.9

60

50

40
East
30 West
North
20

10

0
1st Qtr 2nd Qtr 3rd Qtr 4th Qtr Ch. 1-
113
1.7
Data Presentation Errors

Goals for effective data presentation:


 Present data to display essential information

 Communicate complex ideas clearly and accurately

 Avoid distortion that might convey the wrong message

Ch. 1-
114
Data Presentation Errors
(continued)

 Unequal histogram interval widths


 Compressing or distorting the vertical axis
 Providing no zero point on the vertical axis
 Failing to provide a relative basis in comparing data
between groups

Ch. 1-
115
Chapter Summary

 Reviewed incomplete information in decision making


 Introduced key definitions:
 Population vs. Sample
 Parameter vs. Statistic
 Descriptive vs. Inferential statistics
 Described random sampling
 Examined the decision making process

Ch. 1-
116
Chapter Summary
(continued)
 Reviewed types of data and measurement levels
 Data in raw form are usually not easy to use for decision
making -- Some type of organization is needed:
 Table  Graph

 Techniques reviewed in this chapter :


◼ Line chart
 Frequency distribution ◼ Frequency distribution
◼ Histogram and ogive
 Bar chart
◼ Stem-and-leaf display
 Pie chart ◼ Scatter plot
 Pareto diagram ◼ Cross tables and
Ch. 1-
side-by-side bar charts117

You might also like