0% found this document useful (0 votes)
3 views6 pages

Descriptive Statistics Overview

Chapter 2 covers descriptive statistics, including various graphical representations such as stem-and-leaf plots, line graphs, and bar graphs, as well as methods for calculating measures of central tendency and spread, including percentiles, quartiles, and standard deviation. It also discusses the construction of box plots and the interpretation of skewness in data distributions. The chapter provides examples and strategies for analyzing data sets and calculating statistical measures.

Uploaded by

skyrunner799
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views6 pages

Descriptive Statistics Overview

Chapter 2 covers descriptive statistics, including various graphical representations such as stem-and-leaf plots, line graphs, and bar graphs, as well as methods for calculating measures of central tendency and spread, including percentiles, quartiles, and standard deviation. It also discusses the construction of box plots and the interpretation of skewness in data distributions. The chapter provides examples and strategies for analyzing data sets and calculating statistical measures.

Uploaded by

skyrunner799
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Chapter 2 — Descriptive Statistics 1

1 Stem-and-Leaf Graphs, Line Graphs, and Bar Graphs


Stem and Leaf Plot (Stemplot):
Example 1.1. The following data consist of the scores of prior statistic class. Construct a Stem and Leaf plot for the
data.

Stem Leaf

33 42 49 49 53 55 55 61 63 67
68 68 69 69 72 73 74 78 80 83
88 88 88 90 92 94 94 94 94 96

Outlier (Extreme Value):

Line Graph:

Example 1.2. A restaurant server records his weekly earning over a six week period, as shown in the table below (in
hundreds of dollars). Construct a line graph for these data.

Week Earnings ($) 10

1 6
8
2 8
6
3 4
4 3 4

5 8
2
6 10
x
1 2 3 4 5 6

Bar Graph:
Example 1.3. A prior statistics class was surveyed regarding their pets. The survey results are given below. Construct
a bar graph that shows the frequency of each pet.

y
20
Pet Frequency
18
Cat 10 16
Dog 19 14
12
Lizard 3 10
Snake 4 8
6
Fish 8
4
Chicken 2 2
x
Chapter 2 — Descriptive Statistics 2

2 Histograms, Frequency Polygons, and Time Series Graphs


Histogram:
Example 2.1. The histogram below shows the results of Test 2 for 102 statistics students last semester.
y
26
24
22
20
18
16
14
12
10
8
6
4
2
x
9 19 29 39 49 59 69 79 89 100
0− 10− 20− 30− 40− 50− 60− 70− 80− 90−

3 Measures of the Location of the Data


Percentile:

Quartiles:
ˆ First Quartile (Q1 ):
ˆ Second Quartile (Q2 ):
ˆ Third Quartile (Q3 ):
Interquartile Range:

Note: Quartiles and Percentiles may or may not be observations in the data set. That is, sometimes a Quartile or
Percentile falls between two observations.
Strategy for Finding the kth Percentile

ˆ k = the percentile desired (e.g. 70%, 35%, 90%, etc.)

ˆ i = the index (ranking or position of a data value in a ranked data set.)

ˆ n = the total number of observations

1. Rank the data.


k
2. Calculate i = 100 (n + 1)
ˆ If i is an whole number, then the kth percentile is the data value in the ith
position in the ranked data set.
ˆ If i is not a whole number, round i up and round i down to the nearest whole
numbers. The kth percentile is the mean of the values in these two positions of
the ranked data set.

Percentile Rank:
Strategy for Finding Percentile Rank of an Observation
To calculate the percentile rank of an observation a:
1. Rank the data.
2. Let x be the number of values less than a and y be the frequency of a.

3. Percentile Rank of a = x+0.5y



n · 100 %
Chapter 2 — Descriptive Statistics 3

Example 3.1. The data set below is composed of the salaries of all 11 employees at a small company.
33, 000 64, 500 28, 000 54, 000 2, 000 68, 500 69, 000 42, 000 54, 000 120, 000 40, 500
Calculate the following items for these data:
ˆ Q1 , Q2 , and Q3
ˆ Interquartile Range
ˆ The 70th percentile and the 35th percentile
ˆ Identify any outliers
ˆ The percentile rank of 54,000 and the percentile rank of 40,500

4 Box Plots
Box plot (Box-and-Whisker Plot):
ˆ Composed of 5 values:
ˆ Visual Representation of 3 characteristics:
Example 4.1. Create a Box-and-Whisker Plot for the following data set.
75 69 84 112 74 104 81 90 94 144 79 98
1. Rank the data:
2. Calculate: Find the minimum/maximum values and calculate Q1 , Q2 and Q3 .
3. To construct a box plot, use a horizontal or vertical number line and a rectangular box. The smallest and largest
data values label the endpoints of the axis. The first quartile marks one end of the box and the third quartile marks
the other end of the box. Approximately the middle 50% of the data fall inside the box. The “whiskers” extend
from the ends of the box to the smallest and largest data values. The median or second quartile can be between
the first and third quartiles, or it can be one, or the other, or both. The box plot gives a good, quick picture of the
data.

60 70 80 90 100 110 120 130 140 150


Take note of the position, spread, and skew of the data.

5 Measures of the Center of the Data


Median:

Mode:

Mean:

ˆ Sample Mean (ungrouped data): ˆ Sample Mean (grouped data):

ˆ Population Mean (ungrouped data): ˆ Population Mean (grouped data):


Chapter 2 — Descriptive Statistics 4

Example 5.1. The number of books checked out from the library from 25 IRSC students are as follows:
0 0 0 1 2
3 3 4 4 5
5 7 7 7 7
8 8 8 9 10
10 11 11 12 12
Calculate the mean, median, and mode for the data set.

ˆ Mean:

ˆ Median:

ˆ Mode:

6 Skewness and the Mean, Median, and Mode


Uniform Distribution:

Symmetric Distribution:

Skewed Left Distribution:

ˆ Relationship between Mean and Median

Skewed Right Distribution:

ˆ Relationship between Mean and Median

y y
16 Symmetric Uniform
4
14
12
3
10
8 2
6
1
4
2 x
x 1 2 3 4 5
1 2 3 4 5 6 7 8 9 10

y y
9
5 Skewed Left
8
7 Skewed Right 4
6
5 3

4 2
3
2 1
1
x
x 1 2 3 4 5 6
1 2 3 4 5 6 7 8

7 Measures of the Spread of the Data


Range:
Chapter 2 — Descriptive Statistics 5

Variance:

Standard Deviation:
Grouped Data Ungrouped Data

ˆ Sample Variance: ˆ Sample Variance:

ˆ Sample Standard Deviation: ˆ Sample Standard Deviation:

ˆ Population Variance: ˆ Population Variance:

ˆ Population Standard Deviation: ˆ Population Standard Deviation:

Example 7.1. The data below give the salaries (in thousands of dollars) of all six employees at a small company.

88.50 108.40 65.50 52.50 79.80 54.60

Calculate the range, variance, and standard deviation for these data.

Example 7.2. The following data give the average commute time (in minutes) for 25 employees at a large company.

Commute Time Frequency (f ) Midpoint (m) mf m2 f


0 to < 10 4
10 to < 20 9
20 to < 30 6
30 to < 40 4
40 to < 50 2

Calculate the mean, variance, and standard deviation for the data above.
Chapter 2 — Descriptive Statistics 6

Coefficient of Variation:

ˆ Sample:

ˆ Population:

Example 7.3. Use the Coefficient of Variation to compare the standard deviations calculated in Example 7.1 and Example
7.2. Which data set has greater variability?

8 Two Ways to Use Standard Deviation


Using the mean and standard deviation, we can find out approximately what percent of the total observations in a data
set should fall within a given interval about the mean (centered around the mean).

Chebyshev’s Theorem:

ˆ At least of the data fall within k standard deviations of the mean.

ˆ At least of the data fall within 2 standard deviations of the mean.


ˆ At least of the data fall within 3 standard deviations of the mean.

Empirical Rule:

ˆ Approximately of the data fall within 1 standard deviations of the mean.


ˆ Approximately of the data fall within 2 standard deviations of the mean.
ˆ Approximately of the data fall within 3 standard deviations of the mean.

Example 8.1. The scores on a recent statistics test were skewed left with a mean of 75 and a standard deviation of 10.
Determine at least what percentage of students scored between a 55 and a 95.

Example 8.2. If the data in Example 8.1 were known to be normally distributed, approximately what percentage of the
test scores would fall between 55 and 95?

Example 8.3. A data set consisting of 500 Wechsler IQ Test scores is normally distributed with a mean of 100 standard
deviation of 15. Approximately how many scores fall between a Wechsler IQ score of 55 and a Wechsler IQ score of 145?

Example 8.4. Suppose the data set in Example 8.3 were not known to be normally distributed. At least 96% of the test
scores should lie between what two scores?

You might also like