Overview and
1 Descriptive Statistics
Copyright © Cengage Learning. All rights reserved.
1.3 Measures of Location
Copyright © Cengage Learning. All rights reserved.
Measures of Location
Visual summaries of data are excellent tools for obtaining
preliminary impressions and insights. More formal data
analysis often requires the calculation and interpretation of
numerical summary measures.
That is, from the data we try to extract several summarizing
numbers—numbers that might serve to characterize the data
set and convey some of its salient features.
Suppose, then, that our data set is of the form x1, x2,. . ., xn,
where each xi is a number. What features of such a set of
numbers are of most interest and deserve emphasis? One
important characteristic of a set of numbers is its location, and
in particular its center.
3
The Mean
4
The Mean
For a given set of numbers x1, x2,. . ., xn, the most familiar and
useful measure of the center is the mean, or arithmetic
average of the set. Because we will almost always think of the
xi’s as constituting a sample, we will often refer to the arithmetic
average as the sample mean and denote it by x.
5
Example 1.14 cont’d
Here are the 24-hour water-absorption percentages for the
specimens:
Figure 1.14 shows a dotplot of the data; a water-absorption
percentage in the mid-teens appears to be “typical.”
With 229.0, the sample mean is
6
The Mean
Just as x represents the average value of the observations
in a sample, the average of all values in the population can
be calculated. This average is called the population mean
and is denoted by the Greek letter . When there are N
values in the population (a finite population), then
= (sum of the N population values)/N.
When sampling from such a population (a normal or bell-
shaped population being the most important example), the
sample mean will tend to be stable and quite representative
of the sample
7
The Mean
The mean suffers from one deficiency that makes it an
inappropriate measure of center under some circumstances: - -
Its value can be greatly affected by the presence of even a
single outlier (unusually large or small observation).
For example, if a sample of employees contains nine who earn
$50,000 per year and one whose yearly salary is $150,000, the
sample mean salary is $60,000; this value certainly does not
seem representative of the data.
In such situations, it is desirable to employ a measure that is
less sensitive to outlying values than x, and we will momentarily
propose one.
8
The Median
9
The word median is synonymous with “middle,” and the sample
median is indeed the middle value once the observations are
ordered from smallest to largest.
10
Example 1.15 cont’d
An author went to the Web site [Link] and selected
a sample of 12 recordings of Beethoven’s Symphony
#9,yielding the following durations (min) listed in increasing
order:
62.3 62.8 63.6 65.2 65.7 66.4 67.4 68.4 68.8 70.8 75.7
79.0
Here is a dotplot of the data:
Dotplot of the data from Example 14
Figure 1.16 11
Example 1.15 cont’d
Since n = 12 is even, the sample median is the average of
the n/2 = 6th and (n/2 + 1) = 7th values from the ordered list:
If the largest observation 79.0 had not been included in the
sample, the resulting sample median for the n = 11
remaining observations would have been the single middle
value 66.4 (the [n + 1]/2 = 6th ordered value, i.e. the 6th
value in from either end of the ordered list).
12
Example 1.15 cont’d
The sample mean is x = xi = 816.1/12 = 68.01, a bit more
than a full minute larger than the median.
The mean is pulled out a bit relative to the median because
the sample “stretches out” somewhat more on the upper
end than on the lower end.
The mean is quite sensitive to a single outlier, whereas the
median is impervious to many outliers.
13
The Median
The data in Example 1.15 illustrates an important property
of in contrast to x: The sample median is not affected by
outliers. If, for example, we increased the two largest xis
from 75.7 and 79.0 to 85.7 and 89.0, respectively,
would be unaffected.
Thus, in the treatment of outlying data values, x and are
at opposite ends of a spectrum. Both quantities describe
where the data is centered, but they will not in general be
equal because they focus on different aspects of the
sample.
14
The Median
population median, denoted by
The population mean and median will not generally be
identical. If the population distribution is positively or
negatively skewed, as pictured in Figure 1.16, then
(a) Negative skew (b) Symmetric (c) Positive skew
(b) ( Mean < Median) ( Mean = Median) ( Mean > Median)
Figure 1.16
Three different shapes for a population distribution
15
Other Measures of Location:
Quartiles,
Percentiles
16
Other Measures of Location: Quartiles, Percentiles,
The median (population or sample) divides the data set into
two parts of equal size. To obtain finer measures of
location, we could divide the data into more than two such
parts.
Roughly speaking, quartiles divide the data set into four
equal parts, with the observations above the third quartile
constituting the upper quarter of the data set, the second
quartile being identical to the median, and the first quartile
separating the lower quarter from the upper three-quarters.
Similarly, a data set (sample or population) can be even
more finely divided using percentiles; the 99th percentile
separates the highest 1% from the bottom 99%, and so on. 17
62.3 62.8 63.6 65.2 65.7 66.4 67.4 68.4 68.8 70.8
75.7 79.0
What is:
90th percentile
First quartile
18
1.4 Measures of Variability
Copyright © Cengage Learning. All rights reserved.
19
Measures of Variability
Reporting a measure of center gives only partial information
about a data set or distribution. Different samples or populations
may have identical measures of center yet differ from one
another in other important ways.
Figure 1.18 shows dotplots of three samples with the same
mean and median, yet the extent of spread about the center is
different for all three samples. The first sample has the largest
amount of variability, the third has the smallest amount, and the
second is intermediate to the other two in this respect
Samples with identical measures of center but different amounts of variability
Figure 1.18 20
Measures of Variability for
Sample Data
21
Measures of Variability for Sample Data
The simplest measure of variability in a sample is the
range, which is the difference between the largest and
smallest sample values. The value of the range for sample
1 in Figure 1.18 is much larger than it is for sample 3,
reflecting more variability in the first sample than in the
third.
Samples with identical measures of center but different amounts of variability
Figure 1.18
22
Measures of Variability for Sample Data
A defect of the range, though, is that it depends on only the
two most extreme observations and disregards the
positions of the remaining n – 2 values. Samples 1 and 2 in
Figure 1.18 have identical ranges, yet when we take into
account the observations between the two extremes, there
is much less variability or dispersion in the second sample
than in the first.
Our primary measures of variability involve the deviations
from the mean, That is, the
deviations from the mean are obtained by subtracting
from each of the n sample observations.
23
Measures of Variability for Sample Data
A deviation will be positive if the observation is larger than
the mean (to the right of the mean on the measurement
axis) and negative if the observation is smaller than the
mean. If all the deviations are small in magnitude, then all
xis are close to the mean and there is little variability.
Alternatively, if some of the deviations are large in
magnitude, then some xis lie far from suggesting a
greater amount of variability.
A simple way to combine the deviations into a single
quantity is to average them.
24
Measures of Variability for Sample Data
Note that s2 and s are both nonnegative. The unit for s is
the same as the unit for each of the xis.
25
Example 1.17
Consider the following sample of n = 11 efficiencies for the
2009 Ford Focus equipped with an automatic transmission
The data is represented in the blue table below:
26
Example 1.17
Effects of rounding account for the sum of deviations not
being exactly zero. The numerator of s2 is Sxx = 314.106,
from which
The size of a representative deviation from the sample
mean 33.26 is roughly 5.6 mpg.
27
A Computing Formula for s2
28
A Computing Formula for s2
It is best to obtain s2 from statistical software or else use a
calculator that allows you to enter data into memory and
then view s2 with a single keystroke. If your calculator does
not have this capability, there is an alternative formula for
Sxx that avoids calculating the deviations.
The formula involves both summing and then
squaring, and squaring and then summing.
29
Example 1.18
The data below are the compressive strengths in pounds per
square inch (psi) of 13 specimens of a new aluminum-lithium
alloy undergoing evaluation as a possible material for aircraft
structural elements
154 142 137 133 122 126 135 135 108 120 127 134 122
The sum of these 13 sample observations is
and the sum of their squares is
Thus the numerator of the sample variance is
30
Example 1.18
from which
s2 = 1579.0769/12
= 131.59
and
s = 11.47.
31
Boxplots
32
Boxplots
Histograms convey rather general impressions about a
data set, whereas a single summary such as the mean or
standard deviation focuses on just one aspect of the data.
In recent years, a pictorial summary called a boxplot has
been used successfully to describe several of a data set’s
most prominent features.
These features include (1) center, (2) spread, (3) the extent
and nature of any departure from symmetry, and (4)
identification of “outliers,” observations that lie unusually far
from the main body of the data.
33
Boxplots
Because even a single outlier can drastically affect the
values of and s, a boxplot is based on measures that are
“resistant” to the presence of a few outliers—the median
and a measure of variability called the fourth spread or
Interquartile range (IQR).
Definition
34
Boxplots
Roughly speaking, the fourth spread is unaffected by the
positions of those observations in the smallest 25% or the
largest 25% of the data. Hence it is resistant to outliers.
The simplest boxplot is based on the following five-
number summary:
smallest xi lower fourth median upper fourth largest xi
How to create BoxPlot?
1- First, draw a horizontal measurement scale. Then place
a rectangle above this axis; the left edge of the rectangle is
at the lower fourth, and the right edge is at the upper fourth
(so box width = fs).
35
Boxplots
2- Place a vertical line segment or some other symbol
inside the rectangle at the location of the median; the
position of the median symbol relative to the two edges
conveys information about skewness in the middle 50% of
the data.
3- Finally, draw “whiskers” out from either end of the
rectangle to the smallest and largest observations.
36
Example 1.19
The accompanying data consists of observations on the
time until failure (1000s of hours) for a sample of
turbochargers from one type of engine (from “The Beta
Generalized Weibull Distribution: Properties and
Applications,” Reliability Engr. and System Safety, 2012: 5–
15).
From the output in fig. 1.19, find the five-number summary?
37
Example 1.19
Figure 1.19 shows Minitab output from a request to describe
the data. Q1 and Q3 are the lower and upper quartiles,
respectively, and IQR (interquartile range) is the difference
between these quartiles. SE Mean is, the “standard
error of the mean”; it will be important in our subsequent
development of several widely used procedures for making
inferences about the population mean µ.
Five number summary are:
smallest: 1.6 lower fourth: 5.05 median: 6.5 upper fourth:
7.85 largest: 9.0 38
Example 1.19
Figure 1.20 shows both a dotplot of the data and a boxplot.
Both plots indicate that there is a reasonable amount of
symmetry in the middle 50% of the data, but overall values
stretch out more toward the low end than toward the high
end—a negative skew. The box itself is not very narrow,
indicating a fair amount of variability in the middle half of
the data, and the lower whisker is especially long.
39
Boxplots That Show Outliers
40
Boxplots That Show Outliers
A boxplot can be embellished to indicate explicitly the
presence of outliers. Many inferential procedures are based
on the assumption that the population distribution is normal
(a certain type of bell curve). Even a single extreme outlier
in the sample warns the investigator that such procedures
may be unreliable, and the presence of several mild
outliers conveys the same message.
Definition
41
Boxplots That Show Outliers
Let’s now modify our previous construction of a boxplot by
drawing a whisker out from each end of the box to the
smallest and largest observations that are not outliers.
Now represent each mild outlier by a closed circle and
each extreme outlier by an open circle. Some statistical
computer packages do not distinguish between mild and
extreme outliers.
42
Example 1.20
Among the data considered is the following sample of TN (total
nitrogen) loads (kg N/day) from a particular Chesapeake Bay
location, displayed here in increasing order.
Is there any mild or Extreme outliers?
43
Example 1.20
Relevant summary quantities are
Subtracting 1.5fs from the lower 4th gives a negative
number, and none of the observations are negative, so
there are no outliers on the lower end of the data.
However,
upper 4th + 1.5fs = 351.015 upper 4th + 3fs = 534.24
Thus the four largest observations—563.92, 690.11,
826.54, and 1529.35—are extreme outliers, and 352.09,
371.47, 444.68, and 460.86 are mild outliers.
44
Example 20
The whiskers in the boxplot in Figure 1.21 extend out to the
smallest observation, 9.69, on the low end and 312.45, the
largest observation that is not an outlier, on the upper end.
A boxplot of the nitrogen load data showing mild and extreme outliers
Figure 1.21
45
Example 1.21
A comparative or side-by-side boxplot is a very effective way of
revealing similarities and differences between two or more data
sets consisting of observations on the same variable.
Example:
High levels of sodium in food products represent a growing health
concern. The accompanying data consists of values of sodium
content in one serving of cereal for one sample of cereals
manufactured by General Mills, another sample manufactured by
Kellogg, and a third sample produced by Post
46
Example 1.21
Figure 1.22 shows a comparative boxplot of the data from
the software package R. The typical sodium content
(median) is roughly the same for all three companies. But
the distributions differ markedly in other respects.
47
Example 1.21
The General Mills data shows a substantial positive skew
both in the middle 50% and overall, with two outliers at the
upper end.
The Kellogg data exhibits a negative skew in the middle
50% and a positive skew overall, except for the outlier at
the low end (this outlier is not identified by Minitab).
The Post data is negatively skewed both in the middle 50%
and overall with no outliers.
The box G is very narrow, indicating a small amount of
variability in the middle half of the data
48