MODULE 1 TOPICS
Introduction to Statistics
Tables and Statistical Graphs
Measures of Center
Measures of Variation
Standard Scores and the Empirical Rule
Percentiles and Quartiles
Instructor-created lecture videos aligned with guided notes are posted by topic in
Canvas. You should budget your time to ensure the following tasks are completed by the
end of each week: (1) Print the guided notes, (2) complete the guided notes as you view
the lecture videos, and (3) complete the homework assignments in My Math Lab.
Introduction to Statistics
_______________________ are collections of observations, such as measurements, genders, or survey responses
_______________________ the science of planning studies and experiments, obtaining data, and then
organizing, summarizing, presenting, analyzing, interpreting, and drawing conclusions based on the data.
A _______________________ is the complete collection of all measurements or
data that are being considered.
A _______________________ is a representative subcollection of members
selected from a population
Polls, studies and surveys collect data from a small part (sample) of a larger group (population) so that we can
learn something about the larger group. This is a common and important goal of statistics: Learn about a large
group by examining data from some of its members.
1. The Gallup polling company collected data from 1000 adults in the United States. Results showed that
65% of the respondents worried about identity theft.
a. What is the population?
b. What is the sample?
A numerical measurement describing some characteristic of a population is called a _____________________
A numerical measurement describing some characteristic of a sample is called a _______________________
2. The Gallup polling company collected data from 1000 adults in the United States. Results showed that
65% of the respondents worried about identity theft. Does value 65% represent a parameter or a
statistic? Why?
Quantitative vs. Categorical Data
• Quantitative (or numerical) data consists of numbers representing counts or measurements.
• Categorical (or qualitative) data consists of names or labels (representing categories).
3. Determine whether the data described below are quantitative or qualitative and explain why.
a. The weights of supermodels
b. The ages of respondents
c. The college majors for a sample of students
d. Shirt numbers of professional athletes uniforms
Observational Study vs. Experiment
Statistical methods are driven by the data that we collect. We typically obtain data from two distinct sources:
observational studies and experiments.
In an ____________________________, the researcher observes and measures specific characteristics
without attempting to modify the individuals being studied.
In an ____________________________, the researcher applies a ____________________________ and then
observes its effects on the individuals.
4. Determine whether the description corresponds to an observational study or an experiment.
a. Gallup surveyed 250 voters and asked who they would vote for in the upcoming election.
b. A clinical trial tested a new medication on 100 adults with high blood pressure.
Sampling Methods
______________________ sample – use results that are easy to collect.
Example:
______________________ sample – individuals choose whether they wish to participate.
Example:
______________________ sample – members from the population are selected in such a way that each
individual member in the population has an equal chance of being selected.
Example:
______________________ sample – select a starting point and then select every nth member of the
population.
Example:
______________________ sample – divide the population into at least two different groups that share the
same characteristics, then draw a sample from each subgroup.
Example:
______________________ sample – divide the population into sections (or clusters), then randomly select
some of those clusters, and choose all members from selected clusters.
Example:
5. Management at a retail store is concerned about the possibility of drug use by the employees and
decide to have a sample of employees undergo a drug test. Several plans for choosing the sample are
proposed. Identify the sampling method used in each situation.
a. There are four employee classifications: supervisors, full-time clerks, part-time clerks, maintenance
staff. Randomly select ten people from each category.
b. Choose the fourth person that arrives to work for each shift.
c. Choose everyone who is in the cafeteria at 12:15 pm.
d. There are four employee classifications: supervisors, full-time clerks, part-time clerks, and
maintenance staff. Randomly select an employee classification and test all the people who work in
that classification.
e. Each employee has an identification number. Randomly choose 40 numbers.
f. Send a message out to all employees via email and ask them to come for the drug test if they time.
Frequency Distributions
A frequency distribution is a table that shows how data are distributed among several categories along with
the number (or ________________________) of data values in each of them.
A relative frequency distribution gives the _______________________________ for each class.
1. The scores on a statistics exam are given below. Summarize the data in a frequency distribution.
72, 87, 78, 41, 93, 81, 70, 67, 77, 71, 76, 84, 62, 73, 85, 68, 99, 66, 78, 65
Exam Scores Frequency
40-49
50-59
60-69
70-79
80-89
90-99
2. Using percentages, what are the relative frequencies of the five classes?
Exam Scores Relative Frequency
40-49
50-59
60-69
70-79
80-89
90-99
Note: The relative frequencies may not add up to exactly 100% due to rounding.
• The class ________________________ is the difference between two consecutive lower-class limits or
two consecutive lower-class boundaries.
• The class ________________________ are the values in the middle of the classes and can be found by
adding the lower-class limit to the upper-class limit and dividing the sum by 2
• The class ________________________ are the numbers used to separate classes, but without the
gaps created by class limits.
3. Identify the class width, class midpoints, and class boundaries for the given frequency distribution.
Exam Scores Frequency
40-49 1
50-59 0
60-69 5
70-79 8
80-89 4
90-99 2
Class width: _______
Class midpoints: _______, _______, _______, _______, _______, _______
Class boundaries: _______, _______, _______, _______, _______, _______, _______
Histogram
A histogram is a graph consisting of bars of equal width drawn adjacent to each other (unless there are gaps in
the data).
• The horizontal scale represents the classes of quantitative data values
• The vertical scale represents the frequencies.
• The heights of the bars correspond to the frequency values.
4. Construct a histogram for the frequency distribution.
Exam Scores Frequency
40-49 1
50-59 0
60-69 5
70-79 8
80-89 4
90-99 2
Describing the Shape of a Distribution
A distribution of data is ______________________ if it has a “bell” shape. Characteristics of the bell shape:
o The frequencies increase to a maximum, and then decrease, and
o symmetry, with the left half of the graph roughly a mirror image of the right half.
A distribution of data is ______________________ if it is not symmetric and extends more to one side than
the other.
A distribution of data is ______________________ if the frequencies remain nearly constant.
Stem and Leaf Plot
Represents quantitative data by separating each value into two parts: the stem (such as the leftmost digit) and
the leaf (such as the rightmost digit).
4. The data represents the heights of eruptions by a geyser. Use the heights to construct a stemplot.
Identify the two values that are closest to the middle when the data are sorted in order from lowest to
highest
61 37 50 90 80
50 40 70 50 68
75 57 57 66 69
60 74 70 48 85
Dotplot
Consists of a graph in which each data value is plotted as a point (or dot) along a scale of values. Dots
representing equal values are stacked.
5. The data represents the volumes of a generic soda brand. Use the volumes to construct a dotplot.
Are there any outliers?
80 75 70 75 50
80 65 65 75 85
70 70 70 65
Time Series
Data that have been collected at different points in time: time-series data
Scatterplot
A plot of paired (x, y) quantitative data with a horizontal x-axis and a vertical y-axis. Used to determine
whether there is a relationship between the two variables.
Example For a sample of individuals, measure waist circumference (x) and arm circumference (y).
Bar Graph
Uses bars of equal width to show frequencies of categorical, or qualitative, data. Vertical scale represents
frequencies or relative frequencies. Horizontal scale identifies the different categories of qualitative data.
A multiple bar graph has two or more sets of bars and is used to compare two or more data sets.
Pareto Chart
A bar graph for qualitative data, with the bars arranged in descending order according to frequencies or
relative frequencies (highest to lowest)
Pie Chart
A graph depicting qualitative data as slices of a circle, in which the size of each slice is proportional to
frequency count
Measures of Center: Mean, Median, Mode, and Midrange
The mean is found by adding the data values and dividing by the total number of data values
Notation
individual data value _________ the sum of a set of values _________
sample size _________ population size _________
sample mean _________ population mean _________
1. Find the mean of the first five counts for Chips Ahoy regular cookies: 22 chips, 22 chips, 26 chips,
24 chips and 23 chips.
The median is the middle value of a data set when the data values are arranged in order.
2. Find the median of the five chip counts: 22, 22, 26, 24, 23
3. Find the median of the first six chip counts: 22, 22, 26, 24, 23, 27
The mode is the value(s) that occur with the greatest frequency.
4. Find the mode of the data set 22, 22, 26, 24, 23
5. Find the mode of the data set 22, 22, 22, 23, 23, 23, 24, 24, 26, 27
6. Find the mode of the data set 22, 23, 24, 26, 27
7. Find the mode of the Ice cream preferences. vanilla, chocolate, chocolate, strawberry, vanilla,
vanilla, chocolate, strawberry, strawberry, chocolate
The midrange is the value midway between the maximum and minimum values of a data set.
8. Find the midrange of the five chip counts: 22, 22, 26, 24, 23
9. Here are the hourly wages of some employees: $9 $8 $19 $11 $8
$12 $100 $19 $8 $10
Use your calculator to find the mean, median, mode, and midrange.
Which of the statistics above best describes the center of this distribution? Why?
Weighted Mean
Formula: Mean =
10. A student earned grades of B, B, A, C, and D. Those courses had these corresponding numbers of credit
hours: 4, 5, 1, 5, 4. The grading system assigns quality points to letter grades as follows: A=4, B=3, C=
2, D=1, and F=0. Compute the grade point average (GPA) and round the result to two decimal places.
11. A student earned grades of 84, 78, 84, and 72 on her four regular tests. She earned a grade of 78 on
the final exam and 86 on her class projects. Her combined homework grade was 87. The four regular
tests count for 40% of the course grade, the final exam counts for 30%, the project counts for 10%, and
homework counts for 20%. What is her weighted mean grade?
Measures of Variation: Range, Variance, Standard Deviation
The dotplots below show the test score distributions for three college algebra classes.
Give the mean and median for each.
The range is the difference between the maximum data value and the minimum data value.
1. Consider the depths (in meters) of the five Great Lakes
Erie 60
Huron 230
Ontario 240
Michigan 280
Superior 400
Find the range of depths.
Why could this value be misleading?
*Plan of attack* – Find a statistic that gives the “average distance” of data values from the mean
2. Calculate the standard deviation of the depths of the five Great Lakes.
Erie 60
Huron 230
Ontario 240
Michigan 280
Superior 400
Properties of the Standard Deviation
The standard deviation measures how much the values _______________________________________.
The smaller the standard deviation, the __________________ spread out the data values are from the mean
The larger the standard deviation, the __________________ spread out the data values are from the mean
Can the standard deviation be negative? _____ Zero? _____
Formulas: Variance Standard Deviation
Population σ 2
=
∑ (x − µ ) 2
σ=
∑ (x − µ ) 2
N N
Sample s 2
=
∑ (x − x )2
s=
∑ (x − x )2
n −1 n −1
A data value more than 2 standard deviations either above or below the mean is considered “unusual”
minimum “usual” value = ________________________________
maximum “usual” value = ________________________________
3. The five-year rates of return for a sample of two types of stocks are shown
Financial Stocks Energy Stocks
16.1 18.6 4.0 13.9 17.9 2.5 9.7 21.7 8.3 5.1
3.3 8.4 21.3 34.2 16.6 7.7 23.7 5.6 6.1 9.7
Calculate the mean and standard deviation for both types of stock and interpret the results.
4. A statistics professor finds that the times (in seconds) required to complete a short quiz have a mean
of 180 sec and a standard deviation of 30 sec. Is a time of 90 sec considered unusual? Why or why
not?
Z scores
A z-score (standard score) is the number of standard deviations that a data value is above/below the mean.
data value − mean
Formula: z= z= z=
standard deviation
Z-scores are ____________________
Round Z-scores to ___________________
Most z-scores are between _______ and _______
1. Suppose a male patient measured his pulse rate to be 48 beats per minute. Is his pulse rate unusual if
the mean adult male pulse rate is 67.3 beats per minute with a standard deviation of 10.3?
2. The combined math and verbal scores on the SAT have a mean of 1000 and a standard deviation of
200. The scores on the ACT have a mean of 20 and a standard deviation of 5. Kristen scores 1380 on
the SAT and 31 on the ACT. On which test did Kristen perform better on?
The Empirical Rule
For data sets having a distribution that is approximately normal (bell shaped), the following properties apply:
• About ________ of all values fall within 1 standard deviation of the mean.
• About ________ of all values fall within 2 standard deviations of the mean.
• About ________of all values fall within 3 standard deviations of the mean.
3. Heights of men are normally distributed with a mean of 69 inches and a standard deviation of 3 inches.
a. 68% of men will have a height between ___________ and ___________
b. 95% of men will have a height between ___________ and ___________
c. 99.7% of men will have a height between ___________ and ___________
d. What percent of men have a height between 69 inches and 72 inches?
Percentiles
Divide a data set into 100 groups, each containing about 1% of the data values.
Notation: Pn denotes the nth percentile
Pn = X means n % of the data values are less than X.
Formula: percentile value of X = (round to nearest whole number)
1. A sample of twenty-five sweater prices are given in the stemplot. Find the percentile corresponding to
a sweater that costs $42.
2. A sample of SAT scores is given below:
460 650 750 800 870
890 920 950 970 990
1000 1030 1060 1080 1130
1140 1160 1200 1350 1500
Find the SAT score corresponding to the
67th percentile 50th percentile
79th percentile 90th percentile
Quartiles
Divide a data set into 4 groups, each containing about 25% of the data values.
Notation: First quartile ___________
Second quartile ___________
Third quartile ___________
Interquartile Range ______________________
3. Find the quartiles and interquartile range for the SAT score data.
460 650 750 800 870
890 920 950 970 990
1000 1030 1060 1080 1130
1140 1160 1200 1350 1500
Boxplots and the 5-Number Summary
For a set of data, the 5-number summary consists of ______________________________________________
A boxplot is a graph that displays the 5-number summary.
4. Find the 5-number summary for the SAT scores data and construct a boxplot.
SAT Scores