0% found this document useful (0 votes)
6 views37 pages

STA2023 Chapter 2 Notes Spring 2023

Chapter 2 of STA2023 focuses on descriptive statistics, specifically frequency distributions and their graphical representations. It covers the construction of frequency distributions, histograms, frequency polygons, and cumulative frequency polygons, along with examples and steps for using a TI-84 calculator. Additionally, it discusses stem-and-leaf plots, dot plots, and methods for graphing qualitative data such as Pareto and pie charts.

Uploaded by

megan.harold15
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views37 pages

STA2023 Chapter 2 Notes Spring 2023

Chapter 2 of STA2023 focuses on descriptive statistics, specifically frequency distributions and their graphical representations. It covers the construction of frequency distributions, histograms, frequency polygons, and cumulative frequency polygons, along with examples and steps for using a TI-84 calculator. Additionally, it discusses stem-and-leaf plots, dot plots, and methods for graphing qualitative data such as Pareto and pie charts.

Uploaded by

megan.harold15
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

STA2023 - Chapter 2

Descriptive Statistics
(2.1) Frequency Distribution and Their Graphs
Organizing Quantitative Data Using a Frequency Distribution
• Data collected in original form is called __________________.
• A _________________________ is the body of raw data in a table form, using
classes (intervals) and frequencies.
• _____________of a class is the number of data entries in that class.
• ___________________is the least number that can belong to the class.
• ___________________is the greatest number that can belong to the class.
• __________________ is the distance between any two consecutive classes.
• To find the _____________, get the difference between the lower limits of two
consecutive classes or the upper limits of two consecutive classes.
• _________ is the difference between greatest and smallest values of data set

EXAMPLE 1 The ages of people involved in a study related to the ages of the 50
the wealthiest in the world are listed below.

Raw Data Set Frequency Distribution

a) The number of classes is____


b) The class width is __________________________________________
c) The smallest value of the data set is ____ and the largest value is _____
d) The range is __________________
e) The lower class limits are _________________________and the upper class
limits are___________________________________
1
Page

f) The number of values belong to the 5th class is ________.


Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Guidelines for Constructing a Frequency Distribution from a Data Set
1) Determine the number of classes (it should be between 5 and 20).
2) Find the range = Largest Value – Smallest Value
3) Find the class Width = Range ÷ number of classes (Round up to nearest whole number).
Normally 3.2 would round to be 3, but in rounding up, it becomes 4.
4) If Range ÷ number of classes = a whole number (no remainder), then you can either add
one to the number of classes or add one to the class width.
5) The smallest value is the lower limit of the 1st class.
6) To find the remaining lower limits, add the class width to the 1st lower limit and keep doing
this based on the number of classes.
7) Classes can’t overlap. So, find the 1st upper limit, then add the class width to the 1st upper
limit and keep doing this based on the number of classes.
8) Make a tally mark for each data entry that belongs to the appropriate class.
9) Count the tally marks to find the total frequency ( f ) for each class.

EXAMPLE 2 The following data represents the record of high temperatures for
each of the 50 states. Construct a grouped frequency distribution for
the data using seven classes.

Class Tally Frequency, f

∑f =
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Finding Midpoints, Relative Frequency, and Cumulative Frequency
We can include several additional features that will help provide a better
understanding of the data and help to graph the data in different ways.

Lower class limit + Upper class limit


Midpoint =
2
𝑓
Relative frequency = portion or % of the data that falls in a class =
∑𝑓
Cumulative frequency = the sum of frequencies of a class and all previous classes.

EXAMPLE 3 Use the data set in example 2 to fill in the blanks

Class Midpoint Relative Cumulative


Frequency, f
frequency frequency
100 - 104 2
105 - 109 8
110 - 114 18
115 - 119 13
120 - 124 7
125 - 129 1
130 - 134 1

Finding Class Boundaries


• Class boundaries are numbers that separate classes without forming gaps between
them. Class boundaries are used to graph a histogram.
• If the data entries are integers such as 9, 5, 7, and 30;
o Lower class boundary = lower class limit – 0.5
What about if
o Upper class boundary = upper class limit + 0.5
the data
• If the data entries are in tenth such as 7.80 and 8.40;
entries are in
o Lower class boundary = lower class limit – 0.05
o Upper class boundary = upper class limit + 0.05 hundredth?

Class Limits Class Boundaries


100 - 104
105 - 109
110 - 114
115 - 119
120 - 124
3

125 - 129
Page

130 - 134
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Constructing Graphs Using TI 84 Calculator
1) Constructing a Histogram
The histogram is a graph that displays the
data by using vertical touched bars of
various heights to represent the frequencies
of the classes. The class boundaries are
represented on the horizontal axis.

EXAMPLE 4 Construct a histogram to visualize the data in the example 2


that represents the record of high temperatures for each of the 50 states.
Steps of constructing a Histogram from Grouped Data:
1) Press [STAT], [ENTER] to insert the lower-class
boundaries in L1, then move to L2 to insert frequencies.
2) Press [STATPLOT], make sure no equations there.
3) Press [2ND] [STATPLOT] [ENTER] to turn Plot 1 ON.
Make sure that all other plots are OFF.

4) Highlight the 3rd graph for Histogram, press [ENTER], for XList: L1 and for Freq: L2
5) Press [WINDOW] and fix the values:
• Xmin: Smallest Class boundary or a bit smaller
• Xmax: Largest Class boundary or a bit larger
• XScl: Class Width = difference of 2 consecutive lower boundaries or 2 lower limits
• Ymin: Smallest frequency = set to 0 or a bit smaller
• Ymax: Largest frequency or a bit larger
• YScl: 1
• Xres: 1
4

6) Press [GRAPH] to display the histogram.


Page

7) To obtain the frequency of each class, press [TRACE], followed by ◄ or ►


Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Steps of constructing a Histogram from Raw Data:
1) Press [STAT], [ENTER] to insert the data
values in L1
2) Press [STAT] → CALC [ENTER] ↓ List: L1
CALCULATE [ENTER] to find the smallest and
largest values of this data set.
3) Press [STATPLOT], make sure no equations
there.
4) Press [2ND] [STATPLOT] [ENTER] to turn Plot 1 ON. Make sure that all other plots are OFF.
5) Highlight the 3rd graph for Histogram, press [ENTER], for XList: L1 and for Freq: 1
6) Press [WINDOW] and fix the values:
• Xmin: Smallest data value or a bit smaller
• Xmax: Largest data value or a bit larger
• XScl: Class Width = Range / # classes (Let say # of classes is 7)
• Ymin: Smallest frequency = set to 0 or a bit smaller
• Ymax: Largest frequency (make a guess based on the number of values. You may need to
change it several times until you see the top of every bar of the histogram.
• YScl: 1
• Xres: 1
7) Press [GRAPH] to display the histogram.
8) To obtain the frequency of each class, press [TRACE], followed by ◄ or ►

EXAMPLE 5 Answer the following questions using the constructed histogram in


example 4 that represents the record of high temperatures for each of the 50 states.

a) Describe the shape of this data distribution.

b) How many states have high temperatures between 109.5 and 119.5?

c) How many states have the highest temperatures?

d) How many states have the lowest temperatures?

e) Which class represents the mode of this distribution?

f) What is the frequency of the 3rd class?


5
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
2) Constructing a Frequency Polygon
• The frequency polygon is a graph that displays the
data by using lines that connect points plotted for the
frequencies at the class midpoints.

• Frequency Polygon is attached to the x-axis before first class and after last class.

EXAMPLE 6 Construct a frequency polygon to represent the data in example 2.

Steps of constructing a Frequency Polygon:


1) Construct a frequency distribution table.
2) To make the graph be attached to the x-axis before first
class and after last class, add two fake midpoints with a
zero frequency for each as follows:
o Smallest Class midpoint – class width

o Largest Class midpoint + class width

3) Press [STAT], [ENTER] to insert the midpoints in L1 then move to L2 to insert frequencies.
4) Press [STATPLOT] and make sure no equations there
5) Press [2ND] [STATPLOT] [ENTER] to turn Plot 1 ON. Make sure that all other plots are OFF.
6) Highlight the 2nd graph for Line graph. press [ENTER], for XList: L1 and for Freq: L2
7) Press [WINDOW] and fix the values:
• Xmin: Small fake Class midpoint or a bit smaller
• Xmax: Large fake Class midpoint or a bit larger
• XScl: Class Width = difference of 2 consecutive midpoints or 2 consecutive lower limits
• Ymin: Smallest frequency = set to 0 or a bit smaller
• Ymax: Largest frequency or a bit larger
• YScl: 1
• Xres: 1
8) Press [GRAPH] to display the histogram.
6

9) To obtain the frequency of each class, press [TRACE], followed by ◄ or ►


Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
3) Constructing a Cumulative Frequency Polygon (Ogive)
• The ogive is a graph that represents the
cumulative frequencies for the classes in a
frequency distribution.
• The upper-class boundaries are
represented on the horizontal axis.

EXAMPLE 7 Construct an Ogive to represent the data in example 2.

Steps of constructing an Ogive:


1) Construct a frequency distribution table
2) Press [STAT], [ENTER] to insert the upper-class
boundaries in L1 and the cumulative frequencies in L2
3) Press [STATPLOT] and make sure no equations there
4) Press [2ND] [STATPLOT] [ENTER] to turn Plot 1 ON.

5) Highlight the 2nd graph for Line graph., press [ENTER], XList: L1 and Freq: L2
6) Press [WINDOW] and fix the values:
Xmin: Smallest Class boundary or a bit smaller
Xmax: Largest Class boundary or a bit larger
XScl: Class Width = difference of 2 consecutive lower boundaries or 2 lower limits
Ymin: Smallest cumulative frequency = set to 0 or a bit smaller
Ymax: Largest cumulative frequency or a bit larger
YScl: 1
Xres: 1
7) Press [GRAPH] to display the histogram.
8) To obtain the frequency of each class, press [TRACE], followed by ◄ or ►

EXAMPLE 8 Using the constructed ogive in example 7,

a) What is the sample size?

b) How many states with a temperature that is lower than 124.5 degree?
7
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
(2.2) More Graphs and Displays
4) Constructing a Stem-and-Leaf Plot
• A stem and leaf plot is a data plot that uses part
of a data value as the stem
(such as tens or hundreds) and part of the data
value as the leaf (such as ones) to form groups or
classes.
• It has the advantage over the grouped frequency distribution of retaining the
_____________ while showing them in graphic form.
• If you count the __________, you will know how many data values there are.
• In the stem and leaf plot above, values are listed from _________to________
• There are _____ values. Smallest value is _____ and largest value is ______.
• ____ is the group that has no values but must be there to ________________
• The group ______ has more values than the other groups.
• The number ______ is the most occurred one.

EXAMPLE 1 At an outpatient testing center, the number of cardiograms


performed each day for 20 days is shown. Construct a stem and leaf plot for the
data. Give your interpretation.

25 31 20 32 13 Title: _____________________
14 43 02 57 23 Stem Leaf
36 32 33 32 44

32 52 44 51 45

Key: ___|___=
Interpretation:
8
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
5) Construction a Dot Plot
• A dot plot is a statistical graph in which each
data value is plotted as a point (dot) above the
horizontal axis.
• Dot plots are useful for showing how

values are distributed, and for finding extremely high or low data values (outlier)
• In the dot plot above, ____ is the most occurred number, and ____ is an outlier.
• ____ and ____ are occurred three times, _____ is occurred two times, and
____, _____, ____, are occurred only once, excluding the outlier.
• The horizontal scale used is appropriate because ______________________

EXAMPLE 2 Construct and analyze a dot plot from the data.

Steps:
1) Choose an appropriate horizontal scale
according to the lowest and highest data
values.
2) Plot the values using dots.
3) Plot more dots for repeated values accordingly.
4) Locate any outlier.

Interpretation:
9
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Graphing Qualitative Data Set

Pareto Chart:
• It is a vertical bar graph in which the height of each bar
represents frequency or relative frequency.
• The bars are placed in order of decreasing height, with
the tallest bar on the left.
• Such positioning helps highlight essential data and used frequently in business.

Pie Chart:
• It is a circle divided into sectors that represent
categories.
• The area of each sector is proportional (%) to the
frequency of each category.
• The pie chart shows the relationship of the parts to the
whole.

Paired Data Sets for Quantitative Data


• Paired data in statistics, often referred to as ordered pairs, refers to two
variables in the individuals of a population that are linked together to determine
the correlation between them.
• For a data set to be considered paired data, both data values must be attached
or linked to one another and not considered separately.

Examples of Paired Data Examples of Unpaired Data

• __________________________________________________________________________ • _______________________________________________________________________________

__________________________________________________________________________ _______________________________________________________________________________



__________________________________________________________________________
_______________________________________________________________________________

__________________________________________________________________________
_______________________________________________________________________________



__________________________________________________________________________
_______________________________________________________________________________

__________________________________________________________________________
10

_______________________________________________________________________________
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Graphing Paired Data Sets
• A Scatter plot is a graph of ____________ of data values
that used to determine if a ___________ (relationship)
exists between the two variables.
• The correlation can be ________, ________, or ____
• The correlation can also be ________or __________
• The graph, above, shows_______________________ as an independent variable
and_______________________ as a dependent variable.
• The two variables have __________, ___________ correlation because
_________________________________________________________________

EXAMPLE 3 Professor Lee wanted to see if there was a relationship between


the number of absences and the final grades of the students in STA2023.
A random sample of 7 students shows the following information:

a) Which variable is the independent variable?


b) Which variable is the dependent variable?
c) Draw a scatter plot.
d) Using the scatter plot, what type of relationship, if any, exists?
e) What can you conclude from the scatter plot?

11
Page

Try to use the TI-84 Plus calculator, as instructed at the end of this chapter, to do this example

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
• Time Series Chart can be used to visualize trends in
numerical values over ________
• The horizontal axis is used to plot the date or time
increments, and the vertical axis is used to plot the values
of the variable that you are measuring.
• By doing this, each point on the graph corresponds to a
_____ and a ________ quantity.
• A straight line connects the ______ on the graph in the order in which they occur.

The table lists the number of motor vehicle thefts (in millions) in the United States
for the years 2005 through 2015. Construct a time series chart either by hand or
using TI-84 Plus calculator, as instructed at the end of this chapter.
Then give your interpretation.

Interpretation:
12
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
(2.3) Summarize Data Using Measures of Central Tendency
Mean – Median – Mode – Mean of Grouped data – Weighted Mean

Central tendency:
• It is a value that represents the central entry of a data set.
• Generally, it is measured by the ________, ________, and _________

Outlier
• It is the value that is numerically distant from the rest of the data value causing
a gap in the distribution.
• For the set 20, 21, 24, 26, 27, 65, the Outlier is _____________________

Mean:
Sum of entries Sum
• It is the ____________ of a data set: = .
Number of entries Count

• Mean ______ affected by ____________, hence ________________________


• For the set 20, 26, 40, 36, 23, 42, 35, the mean is _______________________
∑𝑥 ∑𝑥
• Population Mean "mu": 𝜇 = and Sample Mean "x bar": 𝑥̅ =
𝑁 𝑛
• ____________ is a numerical description of a population characteristics.
• ____________ is a numerical description of a sample characteristics.

Median:
• It is the _________value (if the # of values is odd) or the average of ________
________________(if the # of values is even) when the data entries are in
_____________.
• Median __________ affected by __________, hence ____________________
• To determine the median POSITION after the data values are ordered, we use
the following formula:
Count of Values
+ 1/2
2
• So, for the set 713, 300, 618, 595, 311, 401, and 292, the median is in the
_____ position when values are in order because _______________________.
13

Hence, the median is _______. _____________________________________


Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Mode:
• It is the value that ___________ the most.
• Mode _________ affected by _________, hence ____________________
• There may be ______, _______, or ________ modes.
• Data set with one mode is called ________, with two modes is called ______,
and with more than two modes is called __________
• For the set 18.0, 14.0, 34.5, 10, 11.3, 10, 12.4, 10, the mode is __________
• For the following responses of a sample of audience, the mode is ________

EXAMPLE 1 Calculate the mean, median, and mode of the following data sets

12 14 16 15 13 14 15 18 16 16 12 16 15 17

EXAMPLE 2 Which measure of central tendency does best represent the following
data set? 14
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Mean of a Frequency Distribution for Grouped Data
For the data presented in a frequency distribution, you can estimate the mean

̅ = ∑(𝑓 ∙ 𝑚)
as shown: 𝑥 ∑ 𝑓
𝒙 is the estimate of the mean from “grouped” data, a.k.a. frequency distribution
𝒇 stands for the frequency in each class
𝒎 stands for the midpoint of each class
∑𝒇 stands for the total frequencies, a.k.a. the sample size 𝑛
∑(𝒇 ∙ 𝒎) stand for the sum of the products of each midpoint and its frequency

EXAMPLE 3 This is a frequency distribution of miles run per week. Find the mean.
Class Boundaries Frequency (f) Class Midpoint (m) (f *m)
5.5 – 10.5 1 ∑(𝑓 ∙ 𝑚)
𝑥̅ =
10.5 – 15.5 2 ∑𝑓
15.5 – 20.5 3
20.5 – 25.5 5
25.5 – 30.5 4
30.5 – 35.5 3
35.5 – 40.5 2
∑𝑓 = ∑(𝑓 ∙ 𝑚) =

• An alternative way using the TI-84 Plus calculator,

15
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Weighted Mean
• It is a type of mean that considers an additional factor, and it is used when the
values have different levels of weight.

∑(𝑥 ∙ 𝑤) Sum of the product of the entires and weights


𝑥= =
∑𝑤 Sum of the weights

EXAMPLE 4 A student received the following grades. Find the corresponding GPA.

16
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Shapes of Distributions of Data Values

Discrete Distribution Continuous Distribution Observations

Left or Negatively Skewed Left or Negatively Skewed


____________________________________________________________

____________________________________________________________

____________________________________________________________

____________________________________________________________

____________________________________________________________

____________________________________________________________

Bell Shape - Binomial Bell Shape - Normal ____________________________________________________________

____________________________________________________________

____________________________________________________________

____________________________________________________________

____________________________________________________________

____________________________________________________________

Right or Positively Skewed Right or Positively Skewed ____________________________________________________________

____________________________________________________________

____________________________________________________________

____________________________________________________________

____________________________________________________________

____________________________________________________________
17 Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
(2.4) Summarize Data Using Measures of Variation
Range – Variance – Standard Deviation
Coefficient of Variation – Empirical Rule – Chebyshev’s Theorem

Variation
• It is the amount of _________ the
values away from the _______ value.
• Smaller value = Less variation
• Larger value = More variation

Range
• Range measures the largest _________ between any two values in the data set.
• In a data set, the range R = largest value − smallest value
• It is sensitive to outliers, and it ignores how data are distributed.
• The range of 1,1,1,1,2,2,2,2,3,3,3,3,4,5 is: _____________________
• The range of 1,1,1,1,2,2,2,2,3,3,3,3,4,120 is: _____________________
• The range of 7, 8, 9, 10, 11, 12 is: _____________________
• The range of 7, 8, 9, 10, 11, 12, 12, 12 is: _____________________

EXAMPLE 1 Two experimental brands of outdoor paint are tested to see how long
each will last before fading. Six cans of each brand create a small population.
The results (in months) are shown. Find the mean and range of each group.

18
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Variance and Standard Deviation
• The deviation of value, in a data set, is the __________ between this value
and the mean.
• Variance measures the average deviations (___________) around the mean
in squared units. So, the variance units are _________ from the data set ☹.

• To overcome this problem, take the _____________ of the variance to get the
Standard Deviation 𝝈, which has the same units of measure as the data set.
• As the values get farther from the ________, the value of 𝝈 _________
• Values lying more than _______ standard deviations from the ________ are
considered unusual. Values lying more than _______ standard deviations from
the ________ are considered very unusual (______________).

• Variance and Standard Deviation never be ________, but they can be ___ if no
variation at all in the data set. It happens when all entries have the ____ value.
∑(𝒙 −𝝁)𝟐 ∑(𝒙 −𝝁)𝟐
• Population Variance is 𝝈𝟐 = and Population S.D is 𝝈 = √𝝈𝟐 = √
𝑵 𝑵

∑(𝒙 −𝒙 )𝟐 ∑(𝒙 −𝒙 )𝟐
• Sample Variance is 𝒔𝟐 = and Sample S.D is 𝒔 = √𝒔𝟐 = √
𝒏−𝟏 𝒏−𝟏

EXAMPLE 2 Sample office rental rates (in dollars per square foot per year) are
listed. Find the mean rental rate, variance, and standard deviation.
18 27 21 14 20 20 24 11
16 07 12 22 10 15 21 34
23 13 38 16 18 30 15 30 19
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Standard Deviation of a Frequency Distribution for Grouped Data
To estimate the sample standard deviation for grouped data,
Use TI-84 Plus calculator as instructed below.

EXAMPLE 3 This is a frequency distribution of miles run per week.


Find the mean and standard deviation.
Class Frequency Class 1. Find the midpoint of each class.
Boundaries (f) Midpoint (m) 2. Press “STAT” then press “ENTER”
3. Enter the midpoint values in L1
5.5 – 10.5 1
4. Enter the frequency values in L2
10.5 – 15.5 2
15.5 – 20.5 3 5. Press STAT → Calc. Press ENTER
6. In FreqList, press 2nd (2) for L2
20.5 – 25.5 5
7.  Calculate, then press ENTER
25.5 – 30.5 4
For TI-83's and older TI 84"s:
30.5 – 35.5 3 in step # 6, press 2ND (1), 2ND (2)
35.5 – 40.5 2 So, you are telling the calculator to use the lists
L1 and L2. Finally, press the ENTER button again.

Coefficient of Variation
• The coefficient of variation CV is used to compare the ____________ (a.k.a.
variability or riskiness) for data sets with different units or different means.
• Data set with a ________ CV has a greater spread (a.k.a. more variable or risky).
𝑆
• CV is expressed as a percentage using this formula: 𝐶𝑉 = ∙ 100
𝑥

EXAMPLE 4 Stock A had an average price of $50 last year with $5 standard
deviation. Stock B had an average price of $100 with $5 standard deviation.
Which stock is less risky to buy?

EXAMPLE 5 The mean of the number of sales of cars over 3 months is


87, and the standard deviation is 5. The mean of the commissions is $5225, and the
standard deviation is $773. Which one is more variable?
20
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Empirical Rule (or 68 – 95 – 99.7 Rule)
• The empirical rule is a statistical rule that applies
only to a normal (a.k.a. symmetric) distribution.
• The empirical rule is stating that,
o about 68% of a data fall within ±𝟏 standard deviations
o about 95% of a data fall within ±𝟐 standard deviations
o about 99.7% of a data fall within ±𝟑 standard deviations

EXAMPLE 6 Using the Empirical rule, if the mean is 50.5 and the standard
deviation is 1.05, find the interval in which at least 95% of the data will lie.

EXAMPLE 7 The mean speed of a sample of vehicles is 67 miles per hour, with a
standard deviation of 4 miles per hour. Estimate the percent of vehicles whose
speeds are between 63 miles and 71 miles per hour.

21
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
EXAMPLE 8 In a conducted survey, the sample mean height of women in the
U.S.A (ages 20 – 29) was 64.2 inches, with a sample standard deviation of 2.9
inches. Estimate the percent of women whose height are between 64.2 and 67.1
inches.

Chebychev’s Rule

• Chebychev’s Rule applies to all ____________, especially if the _________ is not


bell-shaped or if it is unknown.
• According to the Chebychev’s Rule, the portion of any data set lying within k
1
standard deviations of the mean is at least 1 − , for 𝑘 > 1
𝑘2

• When k = 2; at least ____________________ of the


data lie within___ standard deviation of the mean.
• When k = 3; at least _____________________of the
data lie within ___ standard deviation of the mean.
22
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
EXAMPLE 9 The mean score on a Statistics exam is 82 points, with a standard
deviation of 3 points. Apply Chebychev’s Rule to the data using k = 4.
Interpret the results.

EXAMPLE 10 You are conducting a survey on the number of pets per household in
your area. From a sample with n = 40, the mean number of pets per household is 2
pets, and the standard deviation is 1 pet. Using Chebychev’s Rule, determine at least
how many of the households have 0 to 4 pets.

23
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
1) The increases (in cents) in cigarette taxes for 17 states in a 6-month period are
60, 20, 40, 40, 45, 12, 34, 51, 30, 70, 42, 31, 69, 32, 8, 18, 50.
a) Find the range.
b) Find the mean.
c) Find the variance and standard deviation.
d) According to Empirical Rule, use the mean and standard deviation to find
the interval in which at least 68% of the data will lie.
2) The mean number of runs per game scored by the Chicago Cubs during the
2016 World Series was 3.86, with a standard deviation of 3.36 runs. Apply
Chebychev’s Rule to the data using k = 2. Interpret the results.

24
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
(2.5) Summarize Data Using Measures of Position
Standard Scores – Quartiles – Percentiles

• Measures of Position are used to locate the relative position of a value in the data.
• The position of a data value can be measured using the Standard Score
(a.k.a. Z-Score), Quartiles, and Percentiles.

Standard Score (Z - Score)


• The Standard Score (or Z - Score): tells how many standard deviations a data
value is above or below the mean. The larger the absolute value of the Z-score,
the farther the data value is from the mean.

• Positive z-score means ___________________ • A data value is considered an extreme

• Negative z-score means ___________________ outlier if its Z-score is less than ______ or
greater than ________
• Zero z-score means _______________________

Value − Mean 𝑥−𝜇


• To convert the actual value to Z-Score use: 𝑍 = =
Standard Deviation 𝜎

EXAMPLE 1 A student scored 65 on a calculus test that had a mean of 50 and a


standard deviation of 10; she scored 30 on a history test with a mean of 25 and a
standard deviation of 5. Compare her relative positions on the two tests.

25
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Percentiles
• Percentiles express the __________ of Example: You are the fourth tallest
data falls below a specific value to person in a group of 20.
show the ______ of this value in the 80% of people are shorter than you.
data set.

• The percentile of a data value x,

number of data values below 𝑥


= ∙ 100 Means: you are at the 80th percentile.
total number of data values

EXAMPLE 2 The ogive, at the right, shows a total


of 10,000 people visited the shopping mall over 12
hours.
a) Estimate the 30th percentile (when 30% of the
visitors had arrived).

b) Estimate what percentile of visitors had arrived


after 11 hours.

EXAMPLE 3 A teacher gives a 20-point test to 10 students. The scores are


18, 15,12, 6, 8, 2, 3, 5, 20, 10. Find the percentile rank of a score of 12.

26
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Quartiles
• Quartiles are the values that split an ordered
data set into _________ (___ equal groups).
• Steps to find the three quartiles:
Step 1: Order values from lowest to highest.
Step 2: Find the median of the data values.
This is the Q2 value.
Step 3: Find the median of the data values that fall below Q2. This is the Q1 value
Step 4: Find the median of the data values that fall above Q2. This is the Q3 value

EXAMPLE 4 Find Q1, Q2, and Q3 for the data set: 15, 13, 6, 5, 12, 50, 22, 18

EXAMPLE 5 Find Q1, Q2, and Q3 for the following data set using TI-84 calculator.

44 30 38 23 20 29 19 44 29 17 45 39
18 43 45 39 24 44 26 34 20 35 30 36

1. Press “STAT” then press “ENTER”


2. Enter values into L1 (press ENTER after inserting each value)
27

3. Press STAT → Calc. Press ENTER  Calculate, then press ENTER


Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Interquartile Range (IQR)
• The IQR (also called mid-spread )
measures the spread in the middle
50% of the data. IQR = Q3 – Q1
• IQR is used to identify outliers:
A data value less than the lower fence
Q1 – 1.5(IQR) or greater than the
upper fence Q3 + 1.5(IQR) is considered an outlier.

EXAMPLE 6 Find the interquartile range (IQR) and any outliers for the data set.
44 30 38 23 20 29 19 44 29 17 45 39
18 43 45 39 24 44 26 34 20 35 30 36

EXAMPLE 7 Find the interquartile range (IQR) and any outliers for the data set.
22 25 22 24 20 24 19 22 29 21
21 20 23 25 23 23 21 25 23 22
28
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Five Number Summary and Boxplot
• Five-Number Summary contains five values of the data set:
__________, _________, ___________, ___________, _________
• Boxplot is a graphical display of the data based on the Five-Number Summary:

• Boxplot help describes the __________ , __________ , _________ of data.

• Left rectangle is _____ • Left rectangle is • Right rectangle is


to Right rectangle. ______________ ______________
• Left tail is _______ the • Left tail is __________ • Right tail is __________
to Right tail • Q2 – Q1 ____ Q3 – Q2 • Q2 – Q1 ____ Q3 – Q2
• Q2 – Q1 ____ Q3 – Q2 • Median line is _______ • Median line is _______
• Median line is ______ _____________ _____________
_________________ • Median _______ Mean • Median _______ Mean
• Median _______ Mean • Most of data values are • Most of data values are
• Data values are to the _______ of the to the _______ of the
distributed_________ mean. mean.
29

around the center.


Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
30
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
EXAMPLE 8 Use the following box plot to identify the five number summary.

EXAMPLE 9 Use the following box plots to describe the shape of the distributions.

a) b)

c) d) 31
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
EXAMPLE 10 The number of hours spent studying per day by a sample of 28
students are:

2 8 7 2 3 3 3 2 2 7 8 3 5 1 1 2 6 1 5 7 3 8 5 3 3 7 6 2

a) Find the five number summary and draw a box plot that represent the data.

b) About 75% of the students studied no more than how many hours per day?

c) What percent of the students studied more than 3 hours per day?

d) You randomly select one student from the sample. What is the likelihood that
the student studied less than 2 hours per day? Write your answer as percent.
32
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
EXAMPLE 11 The lengths of songs played at two different concerts are shown.

a) Describe the shape of each distribution. Which concert has less variation in
song lengths?

b) Which distribution is more likely to have outliers? Explain.

c) Which concert do you think has a standard deviation of 16.3? Explain.

d) Can you determine which concert lasted longer? Explain.


33
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
1) What is a z score?

2) Which of these exam grades has a better relative position?


a. A grade of 82 on a test with a mean = 85 and S.D = 6

b. A grade of 56 on a test with a mean = 60 and S.D = 5

3) For the data set 5, 12, 16, 25, 32, 38:


a. Find Q1, Q2, Q3

b. Find the percentile rank for 16

4) The data for a random sample is 270, 180, 250, 290, 130, 260, 340, 310.
Using boxplot, Describe the distribution of this data.

34
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
35
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
36
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
37
Page

Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson

You might also like