0% found this document useful (0 votes)
7 views9 pages

AP Statistics Chapter 1: Data Exploration

Chapter 1 of 'The Practice of Statistics' covers the fundamentals of data analysis, including identifying individuals and variables, classifying them as categorical or quantitative, and creating various types of graphs such as bar graphs and histograms. It emphasizes the importance of understanding distributions, measures of center and spread, and how to analyze relationships between categorical variables. The chapter also introduces key statistical concepts and terms necessary for effective data interpretation and comparison.

Uploaded by

nailspa4one
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views9 pages

AP Statistics Chapter 1: Data Exploration

Chapter 1 of 'The Practice of Statistics' covers the fundamentals of data analysis, including identifying individuals and variables, classifying them as categorical or quantitative, and creating various types of graphs such as bar graphs and histograms. It emphasizes the importance of understanding distributions, measures of center and spread, and how to analyze relationships between categorical variables. The chapter also introduces key statistical concepts and terms necessary for effective data interpretation and comparison.

Uploaded by

nailspa4one
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Name ________________________________

Period ____________
Chapter 1: Exploring Data
The Practice of Statistics, 5th Edition by Starnes, Yates, Moore

Section Topics Objectives: Students will be able to… Homework

● Identify the individuals and variables in a set of


data.
● Classify variables as categorical or quantitative.
● Identify units of measurement for a quantitative
variable.
• Chapter 1 Introduction ● Make a bar graph of the distribution of a
• Airline Pilot Simulation Activity Read (+ additional
categorical variable or, in general, to compare
notes, if necessary)
• Bar Graphs and Pie Charts related quantities.
p. 2–20
• Graphs: Good and Bad ● Recognize the appropriate use of a pie chart.
1.1 • Two-Way Tables and Marginal ● Identify what makes some graphs of categorical
Exercises:
Distributions data deceptive.
p. 6–7: #1, 3, 5, 7, 8;
• Relationships Between Categorical ● Form a two-way table of counts, answer
p. 21–24: #11, 13, 15a,
Variables: Conditional Distributions questions involving marginal and conditional
17a, 19, 21, 27–32
• Organizing a Statistical Problem distributions.
● Describe the relationship between two
categorical variables by computing appropriate
conditional distributions.
● Construct bar graphs to display the relationship
between two categorical variables.

● Make and interpret a dotplot or stemplot to Read (+ additional


• Dotplots display small sets of quantitative data. notes, if necessary)
● Describe the overall pattern (shape, center, p. 25–40
• Describing Shape
spread) of a distribution and identify any major
• Comparing Distributions
1.2 departures from the pattern (like outliers). Exercises:
• Stemplots
● Identify the shape of a distribution. p. 41–42: #37, 39, 41,
• Histograms ● Make a histogram with a reasonable choice of 45
• Using Histograms Wisely quantitative data and classes. p. 44–47: #53–59 odd,
● Interpret histograms. 60, 65, 69–74

● Calculate and interpret measures of center


• Measuring Center: Mean and Median (mean, median)
Read (+ additional
● Calculate and interpret measures of spread
• Comparing Mean and Median notes, if necessary)
(range, IQR)
• Measuring Spread: IQR p. 48–68
● Identify outliers using the 1.5  IQR rule.
• Identifying Outliers
● Construct a boxplot
1.3 • Five Number Summary and Boxplots Exercises:
● Calculate and interpret measures of spread
• Measuring Spread: Standard (standard deviation)
p. 69: #79, 81, 83
Deviation p. 70: #89–95 odd
● Choose appropriate measures of center and
• Choosing Measures of Center and spread
p. 71–73: #97, 103,
Spread 105, 107–110
● Use appropriate graphs and numerical summaries
to compare distributions of quantitative variables.

Chapter 1 Review & Practice Test


FRAPPY!!! p. 76–81
(optional, but recommended)

Chapter 1 Test

AP Statistics 2023-2024, Fagyas


Chapter 1: Exploring Data
Introduction. Data Analysis: Making Sense of Data
p. 2–6

What is data analysis?


Data analysis is the process of _______________________ , _______________________ , _______________________, and
_______________________ about data.

What is inference?
Inference is _______________________________ about a ___________________ on the basis of _______________________.

Given a set of data, what are each of the following?


(a) individuals: the _________________________________ described by a set of data.

(b) variables: any ____________________________ of an individual of interest.

(i) categorical variable: places an individual into one of several _______________________.

(ii) quantitative variable: takes _______________________ for which it makes sense to find an _______________.

AP EXAM TIP: If you learn to distinguish categorical variables from quantitative variables now, it will pay off
later. You will be expected to analyze categorical and quantitative variables correctly on the AP exam.

EXAMPLE: Below is information about 10 randomly selected U.S. residents from a recent census. Describe the individuals and
label each variable as either categorical or quantitative.

What is a distribution?
A distribution of a variable tells us _______________________ the variable takes and _______________________ it takes these
values

Other Important Terms:


• Population: the ____________________________ of individuals about which information is desired.

• Census: a study that attempts to collect data from ______________________________________________________.

• Sample: a _____________________________________ selected from the population; who/what we actually collect


data from.

• Frequency: the _______________________ of times a specific outcome occurs; also called a ________________

• Relative Frequency: the _______________________ of times a specific outcome occurs, relative to the whole

2
1.1 Analyzing Categorical Data

The distribution of a categorical variable lists the categories and gives either the ____________ or the ________________ of
individuals who fall within each category.

Bar Graphs and Pie Graphs


p. 7–11
Features of Bar Graphs:
• Used only for _______________________ data
• Bars should be ______________________
• Bars DO NOT ________________
• Comparisons are made by examining/comparing the ___________________ of the bars
• Bar heights can be either ____________________ or ______________________________

Example p.10-11
*Watch the scaling of your bar graphs (vertical axis should always start at 0)!
*Beware of pictographs (bar graphs which use pictures to represent the bars); our eyes react to the areas of the bars, but we should
be focusing on the heights!
Pie Graphs:
• Compares ________________________ visually using _____________________________________
• Should ONLY be used when you want to emphasize each category’s _________________________________________

Two-Way Tables and Marginal Distributions


p. 11–14
What is a two-way table?
A two-way table displays information about ____________________________________________.

Example: I’m Gonna Be Rich! p.12


A survey of 4826 randomly selected young adults (aged 19-25) asked, “What do you think the chances are that you will have
much more than a middle-class income at age 30?” Below is a two-way table summarizing the results.

Young Adults by Gender and Chance of Getting Rich by Age 30


Gender
Opinion Female Male Total
Almost No Chance 96 98
Some Chance, But Probably Not 426 286
50-50 Chance 696 720
Good Chance 663 758
Almost Certain 486 597
Total

What is a marginal distribution?


The marginal distribution of a specific categorical variable in a two-way table of counts is the distribution of values of that
variable among ALL individuals described by the table. Marginal distributions are given as ____________________.

List the marginal distribution of opinions. Make a bar graph to display this marginal distribution and describe what you see.
Opinion Percent (of total)
Almost No Chance
Some Chance, But Probably Not
50-50 Chance
Good Chance
Almost Certain
3
Relationships between Categorical Variables: Conditional Distributions
p. 14–20
Marginal distributions tell us nothing about the relationship between two variables!
To describe a relationship between two categorical variables, we must calculate some well-chosen percents/proportions from the
counts in the body of the table.

Refer to the two-way table from the previous page. We can study the opinions of men alone by looking at the “Male” column in
the two-way table. Consider how we would find the percent of men who think they are almost certain to be rich by age 30.

Using the above method for all five opinions in the “Male” column gives the conditional distribution of opinion among young
men.
Opinion Percent (of men)
Almost No Chance
Some Chance, But Probably Not
50-50 Chance
Good Chance
Almost Certain

So… what is a conditional distribution?


A conditional distribution of a specific variable describes the values of that variable among individuals who have a specific value
of another variable.

Calculate the conditional distribution of opinion among young women.


Opinion Percent (of women)
Almost No Chance
Some Chance, But Probably Not
50-50 Chance
Good Chance
Almost Certain

Question: Based on the survey data, can we conclude that young men and women differ in their opinions about the likelihood of
future wealth? Provide appropriate evidence to support your answer.

Side-By-Side Bar Graphs and Segmented Bar Graphs


A side-by-side bar graph is used to compare the _________________ of a categorical variable in each of __________________.
For each value of the variable, there is a bar corresponding to each group. The height of each bar is determined by the ________
or __________ of individuals in the group with that specific variable value.

A segmented bar graph is an alternative graph used to compare the distribution of a categorical variable in each of several
groups. For each group, there is a single bar made up of __________________ that correspond to the different values of the
categorical variable, with each segment length determined by the _____________ of individuals in the group with that value (the
conditional distribution). Each bar has a total height of ______%.

Below are examples of side-by-side and segmented bar graphs to represent the “I’m Gonna Be Rich” data.

4
What is an association?
We say there is an association between two variables if knowing the value of ________________________ helps _____________
the value of the other.

Example/Discussion: A Titanic Disaster (p. 19)

1.2 Displaying Quantitative Data with Graphs


p. 25–40
Dotplots
A dotplot is a graph where each ____________________ is shown as a ________ above its location on a ___________________.

The following is data on the number of goals scored by the 2012 U.S. women’s soccer team in their 25 matches leading up to the
2012 Olympics:
1 3 1 14 13 4 3 4 2 5 2 0 4 1 3 4 3 4 2 4 3 1 2 4 2

Construct a dotplot that represents this data.

How to Examine the Distribution of a Quantitative Variable:


When examining any graph of quantitative data, you should look at and discuss the overall pattern and any striking departures
from that pattern. For the “overall pattern” we describe the ___________, ___________, and ___________. Any striking
departures from that pattern would include ___________, ___________, ___________, or _____________________________.
To help remember what to look for, you can remember either SOCS (shape, outliers, center, spread) or SUCS (shape, unusual
features, center, spread).

Shape
• Symmetric: the right and left sides of the graph are approximately _____________ images.
• Right-Skewed: the right side of the graph is much ___________ than the left side; there is a _________ on the right side.
• Left-Skewed: the left side of the graph is much ___________ than the right side; there is a _________ on the left side.
*The direction of the skewness is the direction of the “tail”
• Uniform: each outcome has _________________ of observations so there’s a consistent ________ throughout the graph.

Outliers or Unusual Features


• An outlier is a value that lies ______________from the rest of the data.
• Gaps
• Clusters
• Any other unusual occurrences you see

Center
• Discuss where the _______________________________ of the data falls
• Three main measures of center: ___________, ___________, ___________

Spread
• Discuss how _______________________________ the data is
• Three main measures of spread: ____________________________, ___________, ___________

5
Other Important Terms:
• Univariate: data dealing with _____________ variable
• Bivariate: data dealing with _____________ variables
• Unimodal: graphical display of quantitative data that has a _____________ peak
• Bimodal: graphical display of quantitative data that has _____________ peaks

Stemplots and Back-to-Back Stemplots (p.31–33)


• Sometimes called “Stem-and-Leaf Plots”
• Used to display ________________ data.
• MUST include a __________ so that we know how to read the display/numbers (COMMON AP EXAM MISTAKE)

Things to consider before constructing a stemplot:


▪ Stemplots do not work well for ________ data sets
▪ There is no “magic number” for the number of stems to use, but ___ is a good minimum
▪ Can __________ stems if you have a long list of “leaves” for some stems.
▪ Back-to-back stemplots compare _________ sets of data using the same stems with two sets of “leaves” (one set to
the left of the stems, one set to the right)

Example: A random sample of 20 female students were asked, “How many pairs of shoes do you have?” Here is the data:
50 26 26 31 57 19 24 22 23 38
13 50 13 34 23 30 49 13 15 51

Construct a stemplot to display this data.

Histograms (p.33–40)
A histogram is a graphical display for _____________________ that uses bars to display the __________________ or
________________________________ of a variable’s outcomes. In a histogram, the bars should touch if they are directly next to
each other.

There are two types of histograms:


1. Discrete
• Bars are centered over discrete (exact) values.
• Should be used when the variable’s outcomes come from a list of finite values

2. Continuous
• Bars cover a class (interval) of values.
• Should be used when the variable’s outcomes could be one of infinitely many values; Could also be used if you have
a large data set and want to group sets of data points together into classes (intervals of values) to display fewer bars.

Cautions and Tips for Constructing Histograms:


• Don’t confuse histograms and __________________.
• The ______________ axis (x-axis) should be the axis that has a number scale corresponding to your variable outcomes.
• Choose classes that are all the same __________. There is no “magic number” of classes, but _____ is a good minimum.
• Use percents instead of counts on the _____________ axis (y-axis) when comparing distributions with different numbers
of observations.
• Just because a graph looks nice doesn’t make it a ____________________ display of data.

6
Example: The following table presents the average points scored per game (PPG) for the 30 NBA teams in the 2012-2013 regular
season. Construct a histogram for the data.
NBA Scoring Averages
Team PPG Team PPG Team PPG
Atlanta 98.0 Houston 106.0 Oklahoma City 105.7
Boston 96.5 Indiana 94.7 Orlando 94.1
Brooklyn 96.9 LA Clippers 101.1 Philadelphia 93.2
Charlotte 93.4 LA Lakers 102.2 Phoenix 95.2
Chicago 93.2 Memphis 93.4 Portland 97.5
Cleveland 96.5 Miami 102.9 Sacramento 100.2
Dallas 101.1 Milwaukee 98.9 San Antonio 103.0
Denver 106.1 Minnesota 95.7 Toronto 97.2
Detroit 94.9 New Orleans 94.1 Utah 98.0
Golden State 101.2 New York 100.0 Washington 93.2

Comparing Distributions
When asked to compare distributions, you MUST:
1. Always address all four characteristics of the distribution (SOCS or SUCS) in context with approximate values. Use
the problem’s context and information in your response.

2. Use comparative terms in your response (i.e. “larger”, “smaller”, “more”, “less”, “greater than”, “less than”, etc.).
Compare the shapes, compare the outliers or unusual features, compare the centers, compare the spreads. It doesn’t
matter which order you compare these things, but you must COMPARE them. ☺

1.3 Describing Quantitative Data with Numbers

Measuring Center: The Mean (Average)


To find the mean 𝒙̅ (pronounced “x-bar”) of a set of observations, add their values and divide by the number of observations. If
the n observations are 𝑥1 , 𝑥2 , … , 𝑥𝑛 , then their mean is

or, in a more compact notation,

Measuring Center: The Median


The median is the midpoint of a distribution. This is the number such that half of the observations are smaller and about half are
larger than it. To find the median of a distribution:
1. Arrange all observations in order of size, from smallest to largest.
2. If the number of observations n is odd, the median is the exact center observation of the ordered list.
3. If the number of observations n is even, then the median is the average of the two center observations in the ordered list.

Example: Find the mean and median of the US women’s soccer team data.
1 3 1 14 13 4 3 4 2 5 2 0 4 1 3 4 3 4 2 4 3 1 2 4 2

7
Comparing the Mean and Median
• If the distribution is ___________________________, then the mean and median are _______________________.
• If the distribution is ___________________________, then the mean and median are _______________________.
• If the distribution is _______________, then the mean is usually further out in the __________.
o For right-skewed distributions, the mean is to the _________ of the median.
o For left-skewed distributions, the mean is to the _________ of the median.

Measuring Spread: Range and Interquartile Range (IQR)


What is the range of a distribution?
The range is the distance between the __________________ and __________________ values.

How to Calculate the Quartiles Q1 and Q3, and the Interquartile Range:
1. Arrange the observations in increasing order and locate the ___________ of the ordered list of observations.
2. The first quartile Q1 is the median of the observations that are ________ than the median (to the left) in the ordered list.
3. The third quartile Q3 is the median of the observations that are ________ than the median (to the right) in the ordered
list.
4. The interquartile range (IQR) is defined as follows:

Example: Find the range, Q1, Q3, and IQR of the US women’s soccer data.

Identifying Outliers
In addition to serving as a measure of spread, the IQR is used as part of a rule of thumb for identifying outliers by using
The 1.5 x IQR Rule for Outliers: an observation an outlier if it falls more than 1.5 x IQR above the third quartile or 1.5 x IQR
below the first quartile.

Example: Using the 1.5 x IQR Rule for Outliers, determine if the US women’s soccer data contains any outliers.

The Five-Number Summary and Boxplots

What is the five-number summary of a distribution?


1. The smallest observation (_____________________)
2. The first quartile ( )
3. The ________________
4. The third quartile ( )
5. The largest observation (_____________________)

How do we construct a boxplot?


• Find the ______________________________ from the ordered list of data
• Draw a central box spanning from ______ to ______
• Draw a line in the box to mark the ______________
• Draw lines (called whiskers) extending from Q1 and Q3 to the ________________ and ________________ observations
that are not outliers
• Mark any outliers individuals with a special symbol, such as an asterisk *

8
Example: Using the US women’s soccer data, find the five-number summary and construct a boxplot.

Measuring Spread: The Standard Deviation


The standard deviation sx measures the typical (or average) distance of the values in a distribution from the mean. It is calculated
by finding an average of the squared deviations (or differences) of each value from the mean and then taking the square root.

Steps to find the standard deviation:


1. Find the difference (deviation) between each observation and the ________, and then _______ each of these differences.
2. Take the sum of the squared differences and divide this sum by ________
3. Square root your result and you will have the ___________________________.

The average squared deviation is called the _____________. (This would be the value you get after step 2 above, before taking
the square root)

Notes about standard deviation:


• standard deviation sx is always positive. sx = 0 only when there is NO variability.
• sx measures spread about the mean
• sx has the same units of measurement as the original observations (as does the mean)
• sx is NOT resistant; few outliers can greatly increase standard deviation.

Suppose you have the following data: 24 34 26 30 37 16 28

(a) Find the mean

(b) Find the standard deviation

Choosing Measures of Center and Spread


• _____________________ and _________________ are general better to describe _______________ distributions or
those with outliers because they are __________________ measures.
• _____________________ and _________________ tend to tell us more information about a data set because they are
calculated using the specific value of each data point; However, these measures should really only be used to describe
___________________ or _______________ ___________________ distributions because they are
________________________ measures and can be greatly impacted by ________________.
9

You might also like