STA2023 Chapter 2 Notes Spring 2023
STA2023 Chapter 2 Notes Spring 2023
Descriptive Statistics
(2.1) Frequency Distribution and Their Graphs
Organizing Quantitative Data Using a Frequency Distribution
• Data collected in original form is called __________________.
• A _________________________ is the body of raw data in a table form, using
classes (intervals) and frequencies.
• _____________of a class is the number of data entries in that class.
• ___________________is the least number that can belong to the class.
• ___________________is the greatest number that can belong to the class.
• __________________ is the distance between any two consecutive classes.
• To find the _____________, get the difference between the lower limits of two
consecutive classes or the upper limits of two consecutive classes.
• _________ is the difference between greatest and smallest values of data set
EXAMPLE 1 The ages of people involved in a study related to the ages of the 50
the wealthiest in the world are listed below.
EXAMPLE 2 The following data represents the record of high temperatures for
each of the 50 states. Construct a grouped frequency distribution for
the data using seven classes.
∑f =
Page
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Finding Midpoints, Relative Frequency, and Cumulative Frequency
We can include several additional features that will help provide a better
understanding of the data and help to graph the data in different ways.
125 - 129
Page
130 - 134
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Constructing Graphs Using TI 84 Calculator
1) Constructing a Histogram
The histogram is a graph that displays the
data by using vertical touched bars of
various heights to represent the frequencies
of the classes. The class boundaries are
represented on the horizontal axis.
4) Highlight the 3rd graph for Histogram, press [ENTER], for XList: L1 and for Freq: L2
5) Press [WINDOW] and fix the values:
• Xmin: Smallest Class boundary or a bit smaller
• Xmax: Largest Class boundary or a bit larger
• XScl: Class Width = difference of 2 consecutive lower boundaries or 2 lower limits
• Ymin: Smallest frequency = set to 0 or a bit smaller
• Ymax: Largest frequency or a bit larger
• YScl: 1
• Xres: 1
4
b) How many states have high temperatures between 109.5 and 119.5?
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
2) Constructing a Frequency Polygon
• The frequency polygon is a graph that displays the
data by using lines that connect points plotted for the
frequencies at the class midpoints.
• Frequency Polygon is attached to the x-axis before first class and after last class.
3) Press [STAT], [ENTER] to insert the midpoints in L1 then move to L2 to insert frequencies.
4) Press [STATPLOT] and make sure no equations there
5) Press [2ND] [STATPLOT] [ENTER] to turn Plot 1 ON. Make sure that all other plots are OFF.
6) Highlight the 2nd graph for Line graph. press [ENTER], for XList: L1 and for Freq: L2
7) Press [WINDOW] and fix the values:
• Xmin: Small fake Class midpoint or a bit smaller
• Xmax: Large fake Class midpoint or a bit larger
• XScl: Class Width = difference of 2 consecutive midpoints or 2 consecutive lower limits
• Ymin: Smallest frequency = set to 0 or a bit smaller
• Ymax: Largest frequency or a bit larger
• YScl: 1
• Xres: 1
8) Press [GRAPH] to display the histogram.
6
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
3) Constructing a Cumulative Frequency Polygon (Ogive)
• The ogive is a graph that represents the
cumulative frequencies for the classes in a
frequency distribution.
• The upper-class boundaries are
represented on the horizontal axis.
5) Highlight the 2nd graph for Line graph., press [ENTER], XList: L1 and Freq: L2
6) Press [WINDOW] and fix the values:
Xmin: Smallest Class boundary or a bit smaller
Xmax: Largest Class boundary or a bit larger
XScl: Class Width = difference of 2 consecutive lower boundaries or 2 lower limits
Ymin: Smallest cumulative frequency = set to 0 or a bit smaller
Ymax: Largest cumulative frequency or a bit larger
YScl: 1
Xres: 1
7) Press [GRAPH] to display the histogram.
8) To obtain the frequency of each class, press [TRACE], followed by ◄ or ►
b) How many states with a temperature that is lower than 124.5 degree?
7
Page
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
(2.2) More Graphs and Displays
4) Constructing a Stem-and-Leaf Plot
• A stem and leaf plot is a data plot that uses part
of a data value as the stem
(such as tens or hundreds) and part of the data
value as the leaf (such as ones) to form groups or
classes.
• It has the advantage over the grouped frequency distribution of retaining the
_____________ while showing them in graphic form.
• If you count the __________, you will know how many data values there are.
• In the stem and leaf plot above, values are listed from _________to________
• There are _____ values. Smallest value is _____ and largest value is ______.
• ____ is the group that has no values but must be there to ________________
• The group ______ has more values than the other groups.
• The number ______ is the most occurred one.
25 31 20 32 13 Title: _____________________
14 43 02 57 23 Stem Leaf
36 32 33 32 44
32 52 44 51 45
Key: ___|___=
Interpretation:
8
Page
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
5) Construction a Dot Plot
• A dot plot is a statistical graph in which each
data value is plotted as a point (dot) above the
horizontal axis.
• Dot plots are useful for showing how
values are distributed, and for finding extremely high or low data values (outlier)
• In the dot plot above, ____ is the most occurred number, and ____ is an outlier.
• ____ and ____ are occurred three times, _____ is occurred two times, and
____, _____, ____, are occurred only once, excluding the outlier.
• The horizontal scale used is appropriate because ______________________
Steps:
1) Choose an appropriate horizontal scale
according to the lowest and highest data
values.
2) Plot the values using dots.
3) Plot more dots for repeated values accordingly.
4) Locate any outlier.
Interpretation:
9
Page
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Graphing Qualitative Data Set
Pareto Chart:
• It is a vertical bar graph in which the height of each bar
represents frequency or relative frequency.
• The bars are placed in order of decreasing height, with
the tallest bar on the left.
• Such positioning helps highlight essential data and used frequently in business.
Pie Chart:
• It is a circle divided into sectors that represent
categories.
• The area of each sector is proportional (%) to the
frequency of each category.
• The pie chart shows the relationship of the parts to the
whole.
• __________________________________________________________________________ • _______________________________________________________________________________
__________________________________________________________________________ _______________________________________________________________________________
•
•
__________________________________________________________________________
_______________________________________________________________________________
__________________________________________________________________________
_______________________________________________________________________________
•
•
__________________________________________________________________________
_______________________________________________________________________________
__________________________________________________________________________
10
_______________________________________________________________________________
Page
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Graphing Paired Data Sets
• A Scatter plot is a graph of ____________ of data values
that used to determine if a ___________ (relationship)
exists between the two variables.
• The correlation can be ________, ________, or ____
• The correlation can also be ________or __________
• The graph, above, shows_______________________ as an independent variable
and_______________________ as a dependent variable.
• The two variables have __________, ___________ correlation because
_________________________________________________________________
11
Page
Try to use the TI-84 Plus calculator, as instructed at the end of this chapter, to do this example
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
• Time Series Chart can be used to visualize trends in
numerical values over ________
• The horizontal axis is used to plot the date or time
increments, and the vertical axis is used to plot the values
of the variable that you are measuring.
• By doing this, each point on the graph corresponds to a
_____ and a ________ quantity.
• A straight line connects the ______ on the graph in the order in which they occur.
The table lists the number of motor vehicle thefts (in millions) in the United States
for the years 2005 through 2015. Construct a time series chart either by hand or
using TI-84 Plus calculator, as instructed at the end of this chapter.
Then give your interpretation.
Interpretation:
12
Page
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
(2.3) Summarize Data Using Measures of Central Tendency
Mean – Median – Mode – Mean of Grouped data – Weighted Mean
Central tendency:
• It is a value that represents the central entry of a data set.
• Generally, it is measured by the ________, ________, and _________
Outlier
• It is the value that is numerically distant from the rest of the data value causing
a gap in the distribution.
• For the set 20, 21, 24, 26, 27, 65, the Outlier is _____________________
Mean:
Sum of entries Sum
• It is the ____________ of a data set: = .
Number of entries Count
Median:
• It is the _________value (if the # of values is odd) or the average of ________
________________(if the # of values is even) when the data entries are in
_____________.
• Median __________ affected by __________, hence ____________________
• To determine the median POSITION after the data values are ordered, we use
the following formula:
Count of Values
+ 1/2
2
• So, for the set 713, 300, 618, 595, 311, 401, and 292, the median is in the
_____ position when values are in order because _______________________.
13
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Mode:
• It is the value that ___________ the most.
• Mode _________ affected by _________, hence ____________________
• There may be ______, _______, or ________ modes.
• Data set with one mode is called ________, with two modes is called ______,
and with more than two modes is called __________
• For the set 18.0, 14.0, 34.5, 10, 11.3, 10, 12.4, 10, the mode is __________
• For the following responses of a sample of audience, the mode is ________
EXAMPLE 1 Calculate the mean, median, and mode of the following data sets
12 14 16 15 13 14 15 18 16 16 12 16 15 17
EXAMPLE 2 Which measure of central tendency does best represent the following
data set? 14
Page
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Mean of a Frequency Distribution for Grouped Data
For the data presented in a frequency distribution, you can estimate the mean
̅ = ∑(𝑓 ∙ 𝑚)
as shown: 𝑥 ∑ 𝑓
𝒙 is the estimate of the mean from “grouped” data, a.k.a. frequency distribution
𝒇 stands for the frequency in each class
𝒎 stands for the midpoint of each class
∑𝒇 stands for the total frequencies, a.k.a. the sample size 𝑛
∑(𝒇 ∙ 𝒎) stand for the sum of the products of each midpoint and its frequency
EXAMPLE 3 This is a frequency distribution of miles run per week. Find the mean.
Class Boundaries Frequency (f) Class Midpoint (m) (f *m)
5.5 – 10.5 1 ∑(𝑓 ∙ 𝑚)
𝑥̅ =
10.5 – 15.5 2 ∑𝑓
15.5 – 20.5 3
20.5 – 25.5 5
25.5 – 30.5 4
30.5 – 35.5 3
35.5 – 40.5 2
∑𝑓 = ∑(𝑓 ∙ 𝑚) =
15
Page
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Weighted Mean
• It is a type of mean that considers an additional factor, and it is used when the
values have different levels of weight.
EXAMPLE 4 A student received the following grades. Find the corresponding GPA.
16
Page
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Shapes of Distributions of Data Values
____________________________________________________________
____________________________________________________________
____________________________________________________________
____________________________________________________________
____________________________________________________________
____________________________________________________________
____________________________________________________________
____________________________________________________________
____________________________________________________________
____________________________________________________________
____________________________________________________________
____________________________________________________________
____________________________________________________________
____________________________________________________________
____________________________________________________________
17 Page
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
(2.4) Summarize Data Using Measures of Variation
Range – Variance – Standard Deviation
Coefficient of Variation – Empirical Rule – Chebyshev’s Theorem
Variation
• It is the amount of _________ the
values away from the _______ value.
• Smaller value = Less variation
• Larger value = More variation
Range
• Range measures the largest _________ between any two values in the data set.
• In a data set, the range R = largest value − smallest value
• It is sensitive to outliers, and it ignores how data are distributed.
• The range of 1,1,1,1,2,2,2,2,3,3,3,3,4,5 is: _____________________
• The range of 1,1,1,1,2,2,2,2,3,3,3,3,4,120 is: _____________________
• The range of 7, 8, 9, 10, 11, 12 is: _____________________
• The range of 7, 8, 9, 10, 11, 12, 12, 12 is: _____________________
EXAMPLE 1 Two experimental brands of outdoor paint are tested to see how long
each will last before fading. Six cans of each brand create a small population.
The results (in months) are shown. Find the mean and range of each group.
18
Page
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Variance and Standard Deviation
• The deviation of value, in a data set, is the __________ between this value
and the mean.
• Variance measures the average deviations (___________) around the mean
in squared units. So, the variance units are _________ from the data set ☹.
• To overcome this problem, take the _____________ of the variance to get the
Standard Deviation 𝝈, which has the same units of measure as the data set.
• As the values get farther from the ________, the value of 𝝈 _________
• Values lying more than _______ standard deviations from the ________ are
considered unusual. Values lying more than _______ standard deviations from
the ________ are considered very unusual (______________).
• Variance and Standard Deviation never be ________, but they can be ___ if no
variation at all in the data set. It happens when all entries have the ____ value.
∑(𝒙 −𝝁)𝟐 ∑(𝒙 −𝝁)𝟐
• Population Variance is 𝝈𝟐 = and Population S.D is 𝝈 = √𝝈𝟐 = √
𝑵 𝑵
∑(𝒙 −𝒙 )𝟐 ∑(𝒙 −𝒙 )𝟐
• Sample Variance is 𝒔𝟐 = and Sample S.D is 𝒔 = √𝒔𝟐 = √
𝒏−𝟏 𝒏−𝟏
EXAMPLE 2 Sample office rental rates (in dollars per square foot per year) are
listed. Find the mean rental rate, variance, and standard deviation.
18 27 21 14 20 20 24 11
16 07 12 22 10 15 21 34
23 13 38 16 18 30 15 30 19
Page
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Standard Deviation of a Frequency Distribution for Grouped Data
To estimate the sample standard deviation for grouped data,
Use TI-84 Plus calculator as instructed below.
Coefficient of Variation
• The coefficient of variation CV is used to compare the ____________ (a.k.a.
variability or riskiness) for data sets with different units or different means.
• Data set with a ________ CV has a greater spread (a.k.a. more variable or risky).
𝑆
• CV is expressed as a percentage using this formula: 𝐶𝑉 = ∙ 100
𝑥
EXAMPLE 4 Stock A had an average price of $50 last year with $5 standard
deviation. Stock B had an average price of $100 with $5 standard deviation.
Which stock is less risky to buy?
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Empirical Rule (or 68 – 95 – 99.7 Rule)
• The empirical rule is a statistical rule that applies
only to a normal (a.k.a. symmetric) distribution.
• The empirical rule is stating that,
o about 68% of a data fall within ±𝟏 standard deviations
o about 95% of a data fall within ±𝟐 standard deviations
o about 99.7% of a data fall within ±𝟑 standard deviations
EXAMPLE 6 Using the Empirical rule, if the mean is 50.5 and the standard
deviation is 1.05, find the interval in which at least 95% of the data will lie.
EXAMPLE 7 The mean speed of a sample of vehicles is 67 miles per hour, with a
standard deviation of 4 miles per hour. Estimate the percent of vehicles whose
speeds are between 63 miles and 71 miles per hour.
21
Page
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
EXAMPLE 8 In a conducted survey, the sample mean height of women in the
U.S.A (ages 20 – 29) was 64.2 inches, with a sample standard deviation of 2.9
inches. Estimate the percent of women whose height are between 64.2 and 67.1
inches.
Chebychev’s Rule
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
EXAMPLE 9 The mean score on a Statistics exam is 82 points, with a standard
deviation of 3 points. Apply Chebychev’s Rule to the data using k = 4.
Interpret the results.
EXAMPLE 10 You are conducting a survey on the number of pets per household in
your area. From a sample with n = 40, the mean number of pets per household is 2
pets, and the standard deviation is 1 pet. Using Chebychev’s Rule, determine at least
how many of the households have 0 to 4 pets.
23
Page
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
1) The increases (in cents) in cigarette taxes for 17 states in a 6-month period are
60, 20, 40, 40, 45, 12, 34, 51, 30, 70, 42, 31, 69, 32, 8, 18, 50.
a) Find the range.
b) Find the mean.
c) Find the variance and standard deviation.
d) According to Empirical Rule, use the mean and standard deviation to find
the interval in which at least 68% of the data will lie.
2) The mean number of runs per game scored by the Chicago Cubs during the
2016 World Series was 3.86, with a standard deviation of 3.36 runs. Apply
Chebychev’s Rule to the data using k = 2. Interpret the results.
24
Page
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
(2.5) Summarize Data Using Measures of Position
Standard Scores – Quartiles – Percentiles
• Measures of Position are used to locate the relative position of a value in the data.
• The position of a data value can be measured using the Standard Score
(a.k.a. Z-Score), Quartiles, and Percentiles.
• Negative z-score means ___________________ outlier if its Z-score is less than ______ or
greater than ________
• Zero z-score means _______________________
25
Page
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Percentiles
• Percentiles express the __________ of Example: You are the fourth tallest
data falls below a specific value to person in a group of 20.
show the ______ of this value in the 80% of people are shorter than you.
data set.
26
Page
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Quartiles
• Quartiles are the values that split an ordered
data set into _________ (___ equal groups).
• Steps to find the three quartiles:
Step 1: Order values from lowest to highest.
Step 2: Find the median of the data values.
This is the Q2 value.
Step 3: Find the median of the data values that fall below Q2. This is the Q1 value
Step 4: Find the median of the data values that fall above Q2. This is the Q3 value
EXAMPLE 4 Find Q1, Q2, and Q3 for the data set: 15, 13, 6, 5, 12, 50, 22, 18
EXAMPLE 5 Find Q1, Q2, and Q3 for the following data set using TI-84 calculator.
44 30 38 23 20 29 19 44 29 17 45 39
18 43 45 39 24 44 26 34 20 35 30 36
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Interquartile Range (IQR)
• The IQR (also called mid-spread )
measures the spread in the middle
50% of the data. IQR = Q3 – Q1
• IQR is used to identify outliers:
A data value less than the lower fence
Q1 – 1.5(IQR) or greater than the
upper fence Q3 + 1.5(IQR) is considered an outlier.
EXAMPLE 6 Find the interquartile range (IQR) and any outliers for the data set.
44 30 38 23 20 29 19 44 29 17 45 39
18 43 45 39 24 44 26 34 20 35 30 36
EXAMPLE 7 Find the interquartile range (IQR) and any outliers for the data set.
22 25 22 24 20 24 19 22 29 21
21 20 23 25 23 23 21 25 23 22
28
Page
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
Five Number Summary and Boxplot
• Five-Number Summary contains five values of the data set:
__________, _________, ___________, ___________, _________
• Boxplot is a graphical display of the data based on the Five-Number Summary:
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
30
Page
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
EXAMPLE 8 Use the following box plot to identify the five number summary.
EXAMPLE 9 Use the following box plots to describe the shape of the distributions.
a) b)
c) d) 31
Page
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
EXAMPLE 10 The number of hours spent studying per day by a sample of 28
students are:
2 8 7 2 3 3 3 2 2 7 8 3 5 1 1 2 6 1 5 7 3 8 5 3 3 7 6 2
a) Find the five number summary and draw a box plot that represent the data.
b) About 75% of the students studied no more than how many hours per day?
c) What percent of the students studied more than 3 hours per day?
d) You randomly select one student from the sample. What is the likelihood that
the student studied less than 2 hours per day? Write your answer as percent.
32
Page
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
EXAMPLE 11 The lengths of songs played at two different concerts are shown.
a) Describe the shape of each distribution. Which concert has less variation in
song lengths?
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
1) What is a z score?
4) The data for a random sample is 270, 180, 250, 290, 130, 260, 340, 310.
Using boxplot, Describe the distribution of this data.
34
Page
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
35
Page
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
36
Page
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson
37
Page
Awad 2020 – FGCU - All content adapted from Larson, Elementary Statistics, Seventh Edition, Pearson