Contents
Chapter 01
Descriptive Statistics
PHOK Ponna and PHAUK Sokkhey
Department of Applied Mathematics and Statistics
Institute of Technology of Cambodia
October 6, 2024
AMS (ITC) Descriptive Statistics October 6, 2024 1 / 63
Contents
Contents
1 The Nature of Probability and Statistics
2 Frequency Distributions and Graphs
3 Data Description
Measures of Central Tendency
Measures of Variation
Measures of Position
4 Exploratory Data Analysis
AMS (ITC) Descriptive Statistics October 6, 2024 1 / 63
The Nature of Probability and Statistics
Contents
1 The Nature of Probability and Statistics
2 Frequency Distributions and Graphs
3 Data Description
Measures of Central Tendency
Measures of Variation
Measures of Position
4 Exploratory Data Analysis
AMS (ITC) Descriptive Statistics October 6, 2024 2 / 63
The Nature of Probability and Statistics
Definition 1
Statistics is the heart of Data Analytic concern with conducting
studies to collect, organize, summarize, analyze, and draw conclusions
from data.
AMS (ITC) Descriptive Statistics October 6, 2024 3 / 63
The Nature of Probability and Statistics
There are two types of statistics.
Descriptive Statistics is a Statistics that use to Describe datasets. It
consists of the collection, organization, summarization, and
presentation of data.
Inferential statistics that use sample for making Inference. It
concerns with making predictions and drawing conclusions about a
larger population based on a sample of data.
AMS (ITC) Descriptive Statistics October 6, 2024 4 / 63
The Nature of Probability and Statistics
Definition 2
A population consists of all subjects (human or otherwise) that are
being studied.
Definition 3
A sample is a group of subjects selected from a population.
AMS (ITC) Descriptive Statistics October 6, 2024 5 / 63
The Nature of Probability and Statistics
Example 1
Determine whether descriptive or inferential statistics were used.
a. The average jackpot for the top five lottery winners was $367.6
million.
b. A study done by the American Academy of Neurology suggests
that older people who had a high caloric diet more than doubled
their risk of memory loss.
c. Based on a survey of 9317 consumers done by the National Retail
Federation, the average amount that consumers spent on
Valentine’s Day in 2011 was $116.
d. Scientists at the University of Oxford in England found that a
good laugh significantly raises a person’s pain level tolerance.
AMS (ITC) Descriptive Statistics October 6, 2024 6 / 63
The Nature of Probability and Statistics
Solution
a. Descriptive statistics were used because this is an average, and it
is based on data obtained from the top five lottery winners at this
time.
b. Inferential statistics were used since this is a generalization made
from a sample to a population.
c. Descriptive statistics were used since this is an average based on a
sample of 9317 respondents.
d. Inferential statistics were used since an inference is made from a
sample to a population
AMS (ITC) Descriptive Statistics October 6, 2024 7 / 63
The Nature of Probability and Statistics
A variable is any characteristic, number, or quantity that can be
measured or counted. Data refers to a set of values, which are usually
organized by variables
There are four levels of measurement: nominal, ordinal, interval, and
ratio.
Source
AMS (ITC) Descriptive Statistics October 6, 2024 8 / 63
The Nature of Probability and Statistics
AMS (ITC) Descriptive Statistics October 6, 2024 9 / 63
The Nature of Probability and Statistics
Example 2
What level of measurement would be used to measure each variable?
a. The ages of patients in a local hospital
b. The ratings of movies released this month
c. Colors of athletic shirts sold by Oak Park Health Club
d. Temperatures of hot tubs in local health clubs
Solution
a. Ratio
b. Ordinal
c. Nominal
d. Interval
AMS (ITC) Descriptive Statistics October 6, 2024 10 / 63
The Nature of Probability and Statistics
AMS (ITC) Descriptive Statistics October 6, 2024 11 / 63
The Nature of Probability and Statistics
TYPES OF VARIABLES
There are two basic types of variables: (1) qualitative and (2)
quantitative.
AMS (ITC) Descriptive Statistics October 6, 2024 12 / 63
The Nature of Probability and Statistics
Example 3
Classify each variable as a discrete variable or a continuous variable.
a. The highest wind speed of a hurricane
b. The weight of baggage on an airplane
c. The number of pages in a statistics book
d. The amount of money a person spends per year for online
purchases
Solution
a. Continuous, since wind speed must be measured
b. Continuous, since weight is measured
c. Discrete, since the number of pages is countable
d. Discrete, since the smallest value that money can assume is in
cents
AMS (ITC) Descriptive Statistics October 6, 2024 13 / 63
Frequency Distributions and Graphs
Contents
1 The Nature of Probability and Statistics
2 Frequency Distributions and Graphs
3 Data Description
Measures of Central Tendency
Measures of Variation
Measures of Position
4 Exploratory Data Analysis
AMS (ITC) Descriptive Statistics October 6, 2024 14 / 63
Frequency Distributions and Graphs
A frequency distribution is the organization of raw data in table
form, using classes and frequencies.
The categorical frequency distribution is used for data that can
be placed in specific categories, such as nominal- or ordinal-level
data.
When the range of the data is large, the data must be grouped
into classes that are more than one unit in width, in what is called
a grouped frequency distribution.
AMS (ITC) Descriptive Statistics October 6, 2024 15 / 63
Frequency Distributions and Graphs
Example 4 (Distribution of Blood Types)
Twenty-five army inductees were given a blood test to determine their
blood type. The data set is
A B B AB O O O B AB B AB A O
B B O A O A O O O AB B A
Construct a categorical frequency distribution for the data.
AMS (ITC) Descriptive Statistics October 6, 2024 16 / 63
Frequency Distributions and Graphs
Example 5 (Record High Temperatures)
These data represent the record high temperatures in degrees
Fahrenheit ( °F) for each of the 50 states. Construct a grouped
frequency distribution for the data, using 7 classes.
112 100 127 120 134 118 105 110 109 112
110 118 117 116 118 122 114 114 105 109
107 112 114 115 118 117 118 122 106 110
116 108 110 121 113 120 119 111 104 111
120 113 120 117 105 110 118 112 114 114
AMS (ITC) Descriptive Statistics October 6, 2024 17 / 63
Frequency Distributions and Graphs
Definition 4
A cumulative frequency distribution is a distribution that shows the
number of data values less than or equal to a specific value (usually an
upper boundary). The values are found by adding the frequencies of
the classes less than or equal to the upper class boundary of a specific
class. This gives an ascending cumulative frequency.
AMS (ITC) Descriptive Statistics October 6, 2024 18 / 63
Frequency Distributions and Graphs
Procedure to Construct a Grouped Frequency Distribution
1. Determine the classes.
Find the highest and lowest values.
Select the number of classes desired k such that 2k > n.
Find the width of the class i where i ≥ maximum valu−minimum
k
value
and rounding up to the nearest integer.
Select a starting point (usually the lowest value or any convenient
number less than the lowest value); add the width to get the lower
limits.
Find the upper class limits.
Find the boundaries.
2. Tally the data.
3. Find the numerical frequencies from the tallies, and find the
cumulative frequencies.
AMS (ITC) Descriptive Statistics October 6, 2024 19 / 63
Frequency Distributions and Graphs
Example 6 (Record High Temperatures)
These data represent the record high temperatures in degrees
Fahrenheit (o F) for each of the 50 states. Construct a grouped
frequency distribution for the data, using 7 classes.
112 100 127 120 134 118 105 110 109 112
110 118 117 116 118 122 114 114 105 109
107 112 114 115 118 117 118 122 106 110
116 108 110 121 113 120 119 111 104 111
120 113 120 117 105 110 118 112 114 114
AMS (ITC) Descriptive Statistics October 6, 2024 20 / 63
Frequency Distributions and Graphs
Histograms graphical
representation of the
distribution of numerical
data.
Frequency Polygon is a
graph that displays the data
by using lines that connect
points plotted for the
frequencies at the midpoints
of the classes.
Ogive (Cumulative frequency
curve) is a graph that
represents the cumulative
frequencies for the classes in
a frequency distribution.
AMS (ITC) Descriptive Statistics October 6, 2024 21 / 63
Frequency Distributions and Graphs
Bar graphs use size contrast
to compare two or more
values.
Bar charts with horizontal
bars effectively show data
that are ranked, with bars
arranged in ascending or
descending order.
Line graphs are a type of
visualization that can help
your audience understand
shifts or changes in your
data.
AMS (ITC) Descriptive Statistics October 6, 2024 22 / 63
Frequency Distributions and Graphs
Example 7
Construct a histogram, a frequency polygon and ogive to represent the
data shown for the record high temperatures for each of the 50 states
(see Example above).
AMS (ITC) Descriptive Statistics October 6, 2024 23 / 63
Frequency Distributions and Graphs
When the data are qualitative or categorical, bar graphs can be used
to represent the data. A bar graph can be drawn using either
horizontal or vertical bars.
Definition 5
A bar graph represents the data by using vertical or horizontal bars
whose heights or lengths represent the frequencies of the data.
Example 8 (College Spending for First-Year Students)
The table shows the average money spent by first-year college
students. Draw a horizontal and vertical bar graph for the data.
Electronics $728
Dorm decor 344
Clothing 141
Shoes 72
AMS (ITC) Descriptive Statistics October 6, 2024 24 / 63
Frequency Distributions and Graphs
Bar graphs can also be used to compare data for two or more groups.
These types of bar graphs are called compound bar graphs or
multiple bar graphs.
Example 9
Consider the following data for the number (in millions) of never
married adults in the United States.
Year Males Females
1960 15.3 12.3
1980 24.2 20.2
2000 32.3 27.8
2010 40.2 34.0
Construct a multiple bar graphs for this data.
AMS (ITC) Descriptive Statistics October 6, 2024 25 / 63
Frequency Distributions and Graphs
When data are collected over a period of time, they can be
represented by a time series graph.
Definition 6
A time series graph represents data that occur over a specific period
of time.
Example 10
The data show the percentage of U.S. adults who smoke. Draw and
analyze a time series graph for the data.
Year 1970 1980 1990 2000 2010
Percent 37 33 25 23 19
AMS (ITC) Descriptive Statistics October 6, 2024 26 / 63
Frequency Distributions and Graphs
Two or more data sets can be compared on the same graph called a
compound time series graph if two or more lines are used, as shown
below:
This graph shows the percentage of elderly males and females in the
U.S. labor force from 1960 to 2010. It shows that the percentage of
elderly men decreased significantly from 1960 to 1990 and then
increased slightly after that. For the elderly females, the percentage
decreased slightly from 1960 to 1980 and then increased from 1980 to
2010.
AMS (ITC) Descriptive Statistics October 6, 2024 27 / 63
Frequency Distributions and Graphs
The purpose of the pie graph is to show the relationship of the parts to
the whole by visually comparing the sizes of the sections. Percentages
or proportions can be used. The variable is nominal or categorical.
Definition 7
A pie graph is a circle that is divided into sections or wedges according
to the percentage of frequencies in each category of the distribution.
Example 11 (Super Bowl Snack Foods)
This frequency distribution shows the number of pounds of each snack
food eaten during the Super Bowl. Construct a pie graph for the data.
Snack Pounds (frequency)
Potato chips 11.2 million
Tortilla chips 8.2 million
Pretzels 4.3 million
Popcorn 3.8 million
Snack nuts 2.5 million
Total n = 30.0 million
AMS (ITC) Descriptive Statistics October 6, 2024 28 / 63
Data Description
Contents
1 The Nature of Probability and Statistics
2 Frequency Distributions and Graphs
3 Data Description
Measures of Central Tendency
Measures of Variation
Measures of Position
4 Exploratory Data Analysis
AMS (ITC) Descriptive Statistics October 6, 2024 29 / 63
Data Description Measures of Central Tendency
Large amounts of data can be overwhelming. A single number can
summarize information about a dataset, such as the central tendency
or the dispersion of the dataset. A common data summary is the
arithmetic mean or mean , which is the sum of the data values in a
dataset divided by the number of values in the dataset.
Mathematically, the mean of a set of n data values x1 , x2 , . . . , xn is
denoted x̄ and is defined as follows.
n
1X x1 + x2 + · · · + xn
SAMPLE MEAN x̄ = xi =
n n
i=1
Similarly, computation for population but different notation
P
x
POPULATION MEAN µ=
N
where µ represents the population mean and N is the number of
values in the population.
AMS (ITC) Descriptive Statistics October 6, 2024 30 / 63
Data Description Measures of Central Tendency
Definition 8
Any measurable characteristic of a population is called a parameter.
The mean of a population is an example of a parameter.
Definition 9
Any measure based on sample data, is called a statistic. The mean of
a sample is an example of a statistic.
AMS (ITC) Descriptive Statistics October 6, 2024 31 / 63
Data Description Measures of Central Tendency
Example 12
There are 42 exits on I-75 through the state of Kentucky. Listed below
are the distances between exits (in miles).
11 4 10 4 9 3 8 10 3 14 1 10 3 5
2 2 5 6 1 2 2 3 7 1 3 7 8 10
1 4 7 5 2 2 5 1 1 3 3 1 2 1
Why is this information a population? What is the mean number of
miles between exits?
AMS (ITC) Descriptive Statistics October 6, 2024 32 / 63
Data Description Measures of Central Tendency
Example 13
Verizon is studying the number of monthly minutes used by clients in
a particular cell phone rate plan. A random sample of 12 clients
showed the following number of minutes used last month.
90 77 94 89 119 112
91 110 92 100 113 83
What is the arithmetic mean number of minutes used last month?
AMS (ITC) Descriptive Statistics October 6, 2024 33 / 63
Data Description Measures of Central Tendency
Definition 10
The type of mean that considers an additional factor is called the
weighted mean, and it is used when the values are not all equally
represented.
Find the weighted mean of a variable X by multiplying each value by
its corresponding weight and dividing the sum of the products by the
sum of the weights.
P
w1 X1 + w2 X2 + . . . + wn Xn wX
X̄ = = P
w1 + w2 + . . . + wn w
Example 14
A student received an A in English Composition I (3 credits), a C in
Introduction to Psychology (3 credits), a B in Biology I (4 credits),
and a D in Physical Education (2 credits). Assuming A= 4 grade
points, B = 3 grade points, C = 2 grade points, D = 1 grade point,
and F = 0 grade points, find the student’s grade point average.
AMS (ITC) Descriptive Statistics October 6, 2024 34 / 63
Data Description Measures of Central Tendency
SOLUTION
Major Credit (w ) Grade point (xi )
Engish Composition I 3 A (4 appoints)
Introduction to Psychology 3 C (2 points )
Biology I 4 B (3 points
Physical Education 2 D (1 point)
Hence, the GPA of the student is
P
w ·x (3)(4) + (3)(2) + (4)(3) + (2)(1) 8
x̄ = P = = ≈2·7
w 3+3+4+2 3
AMS (ITC) Descriptive Statistics October 6, 2024 35 / 63
Data Description Measures of Central Tendency
Definition 11 (The Median)
MEDIAN is the midpoint of the values after they have been ordered
from the minimum to the maximum values.
Steps in computing the median of a data array
1 Arrange the data in order X , X , . . . , X .
1 2 n
2 Select the middle point.
X n+1 if n is odd
MD = X n 2+X n +1
2 2
2 if n is even
AMS (ITC) Descriptive Statistics October 6, 2024 36 / 63
Data Description Measures of Central Tendency
Properties of the median
1 It is not affected by extremely large or small values. Therefore,
the median is a valuable measure of location when such values do
occur.
2 It can be computed for ordinal-level data or higher. Recall that
ordinal-level data can be ranked from low to high.
Example 15
Facebook is a popular social networking website. Users can add
friends, send them messages, and update their personal profiles to
notify friends about themselves and their activities. A sample of 10
adults revealed they spent the following number of hours last month
using Facebook.
3 5 7 5 9 1 3 9 17 10
Find the median number of hours.
AMS (ITC) Descriptive Statistics October 6, 2024 37 / 63
Data Description Measures of Central Tendency
Definition 12
MODE is the value of the observation that appears most frequently.
A data set that has only one value that occurs with the greatest
frequency is said to be unimodal.
If a data set has two values that occur with the same greatest
frequency, both values are considered to be the mode and the
data set is said to be bimodal.
If a data set has more than two values that occur with the same
greatest frequency, each value is used as the mode, and the data
set is said to be multimodal.
When no data value occurs more than once, the data set is said
to have no mode.
AMS (ITC) Descriptive Statistics October 6, 2024 38 / 63
Data Description Measures of Central Tendency
Example 16
Recall the data regarding the distance in miles between exits on I-75 in
Kentucky. The information is repeated below.
11 4 10 4 9 3 8 10 3 14 1 10 3 5
2 2 5 6 1 2 2 3 7 1 3 7 8 10
1 4 7 5 2 2 5 1 1 3 3 1 2 1
What is the modal distance?
AMS (ITC) Descriptive Statistics October 6, 2024 39 / 63
Data Description Measures of Central Tendency
Properties of Mode
1 The mode is used when the most typical case is desired.
2 The mode is the easiest average to compute.
3 The mode can be used when the data are nominal or categorical,
such as religious preference, gender, or political affiliation.
4 The mode is not always unique. A data set can have more than
one mode, or the mode may not exist for a data set.
AMS (ITC) Descriptive Statistics October 6, 2024 40 / 63
Data Description Measures of Central Tendency
Definition 13
The midrange is defined as the sum of the lowest and highest values in
the data set, divided by 2. The symbol MR is used for the midrange.
lowest value + highest value
MR =
2
Example 17
The number of bank failures for a recent five-year period is shown.
Find the midrange.
3, 30, 148, 157, 71
Properties of MR
1 The midrange is easy to compute.
2 The midrange gives the midpoint.
3 The midrange is affected by extremely high or low values in a
data set.
AMS (ITC) Descriptive Statistics October 6, 2024 41 / 63
Data Description Measures of Variation
For the spread or variability of a data set, three measures are
commonly used: range, variance, and standard deviation.
Definition 14
The range is the highest value minus the lowest value. The symbol R
is used for the range.
R = highest value − lowest value
Definition 15
The variance of the population is denoted by σ 2 defined by
(X − µ)2
P
2
σ =
N
The standard deviation of the population √ denoted by σ is the
square root of the variance, that is, σ = σ 2 .
AMS (ITC) Descriptive Statistics October 6, 2024 42 / 63
Data Description Measures of Variation
Example 18
The number of traffic citations issued last year by month in Beaufort
County, South Carolina, is reported below.
Determine the population variance.
AMS (ITC) Descriptive Statistics October 6, 2024 43 / 63
Data Description Measures of Variation
Example 19
The Philadelphia office of PricewaterhouseCoopers hired five
accounting trainees this year. Their monthly starting salaries were
$3,536; $3,173; $3,448; $3,121; and $3,622.
(a) Compute the population mean.
(b) Compute the population variance.
(c) Compute the population standard deviation.
(d) The Pittsburgh office hired six trainees. Their mean monthly
salary was $3,550, and the standard deviation was $250. Compare
the two groups.
AMS (ITC) Descriptive Statistics October 6, 2024 44 / 63
Data Description Measures of Variation
Definition 16
The variance of the sample (or sample variance) is denoted by
s 2 defined by
(X − X̄ )2 n( X 2 ) − ( X )2
P P P
s2 = =
n−1 n(n − 1)
The standard deviation of the sample
√ denoted by s is the square
root of the variance, that is, s = s 2 .
Example 20
Find the sample variance and standard deviation for the amount of
European auto sales for a sample of 6 years shown. The data are in
millions of dollars.
11.2, 11.9, 12.0, 12.8, 13.4, 14.3
AMS (ITC) Descriptive Statistics October 6, 2024 45 / 63
Data Description Measures of Variation
Theorem 1 (Chebyshev’s theorem)
The proportion of values from a data set that will fall within k
standard deviations of the mean will be at least 1 − k12 , where k is a
number greater than 1 (k is not necessarily an integer).
In summary, Chebyshev’s theorem states
1 At least three-fourths, or 75%, of all data values fall within 2
standard deviations of the mean.
2 At least eight-ninths, or 89%, of all data values fall within 3
standard deviations of the mean.
AMS (ITC) Descriptive Statistics October 6, 2024 46 / 63
Data Description Measures of Variation
Example 21 ( Prices of Homes)
The mean price of houses in a certain neighborhood is $50,000, and
the standard deviation is $10,000. Find the price range for which at
least 75% of the houses will sell.
Example 22 (Travel Allowances)
A survey of local companies found that the mean amount of travel
allowance for couriers was $0.25 per mile. The standard deviation was
$0.02. Using Chebyshev’s theorem, find the minimum percentage of
the data values that will fall between $0.20 and $0.30.
AMS (ITC) Descriptive Statistics October 6, 2024 47 / 63
Data Description Measures of Variation
The Empirical (Normal) Rule
When a distribution is bell-shaped (or what is called normal), the
following statements, which make up the empirical rule, are true.
Approximately 68% of the data values will fall within 1 standard
deviation of the mean.
Approximately 95% of the data values will fall within 2 standard
deviations of the mean.
Approximately 99.7% of the data values will fall within 3 standard
deviations of the mean.
AMS (ITC) Descriptive Statistics October 6, 2024 48 / 63
Data Description Measures of Position
Definition 17
A z score or standard score for a value is obtained by subtracting the
mean from the value and dividing the result by the standard deviation.
The symbol for a standard score is z. The formula is
value − mean
z=
standard deviation
The z score represents the number of standard deviations that a data
value falls above or below the mean.
Example 23 (Test Scores)
A student scored 65 on a calculus test that had a mean of 50 and a
standard deviation of 10; she scored 30 on a history test with a mean
of 25 and a standard deviation of 5. Compare her relative positions on
the two tests
AMS (ITC) Descriptive Statistics October 6, 2024 49 / 63
Data Description Measures of Position
Definition 18
Percentiles are position measures used in educational and
health-related fields to indicate the position of an individual in a
group. Percentiles divide the data set into 100 equal groups. When
the data are arranged in order from lowest to highest, the percentile
corresponding to a given value X is computed by using the following
formula:
(number of values below X ) + 0.5
Percentile = × 100
total number of values
Example 24
A teacher gives a 20-point test to 10 students. The scores are shown
here. Find the percentile rank of a score of 12 and then of 6.
18, 15, 12, 6, 8, 2, 3, 5, 20, 10
AMS (ITC) Descriptive Statistics October 6, 2024 50 / 63
Data Description Measures of Position
Example 25
Using the scores in Example above, find the value corresponding to the
25th percentile and the value that corresponds to the 60th percentile.
AMS (ITC) Descriptive Statistics October 6, 2024 51 / 63
Data Description Measures of Position
Definition 19
Quartiles divide the distribution into four groups, separated by
Q1 , Q2 , Q3 . Note that Q1 is the same as the 25th percentile; Q2 is the
same as the 50th percentile, or the median; Q3 corresponds to the
75th percentile, as shown:
Quartiles can be computed by using the formula given for computing
percentiles. For Q1 use p = 25. For Q2 use p = 50. For Q3 use
p = 75. The interquartile range (IQR) is defined as the difference
between Q1 and Q3 and is the range of the middle 50% of the data.
Example 26
Find Q1 , Q2 , Q3 , and IQR for the data set 15, 13, 6, 5, 12, 50, 22, 18.
AMS (ITC) Descriptive Statistics October 6, 2024 52 / 63
Data Description Measures of Position
A data set should be checked for extremely high or extremely low
values. These values are called outliers.
Definition 20
An outlier is an extremely high or an extremely low data value when
compared with the rest of the data values.
Remark 1
An outlier can strongly affect the mean and standard deviation of a
variable. For example, suppose a researcher mistakenly recorded an
extremely high data value. This value would then make the mean and
standard deviation of the variable much larger than they really were.
Outliers can have an effect on other statistics as well.
AMS (ITC) Descriptive Statistics October 6, 2024 53 / 63
Data Description Measures of Position
Example 27
Check the following data set for outliers.
5, 6, 12, 13, 15, 18, 22, 50
AMS (ITC) Descriptive Statistics October 6, 2024 54 / 63
Exploratory Data Analysis
Contents
1 The Nature of Probability and Statistics
2 Frequency Distributions and Graphs
3 Data Description
Measures of Central Tendency
Measures of Variation
Measures of Position
4 Exploratory Data Analysis
AMS (ITC) Descriptive Statistics October 6, 2024 55 / 63
Exploratory Data Analysis
In exploratory data analysis (EDA), data can be organized using a
stem and leaf [Link] measure of central tendency used in EDA is the
median. The measure of variation used in EDA is the interquartile
range Q3 –Q1 . In EDA the data are represented graphically using a
boxplot (sometimes called a box and whisker plot). The purpose of
exploratory data analysis is to examine data to find out what
information can be discovered about the data, such as the center and
the spread. Exploratory data analysis was developed by John Tukey
and presented in his book Exploratory Data Analysis (Addison-Wesley,
1977).
AMS (ITC) Descriptive Statistics October 6, 2024 56 / 63
Exploratory Data Analysis
The Five-Number Summary and Boxplots
A boxplot can be used to graphically represent the data set. These
plots involve five specific values:
1 The lowest value of the data set (i.e., minimum)
2 Q1
3 The median
4 Q3
5 The highest value of the data set (i.e., maximum)
Definition 21
A boxplot is a graph of a data set obtained by drawing a horizontal
line from the minimum data value to Q1 , drawing a horizontal line
from Q3 to the maximum data value, and drawing a box whose
vertical sides pass through Q1 and Q3 with a vertical line inside the
box passing through the median or Q2 .
AMS (ITC) Descriptive Statistics October 6, 2024 57 / 63
Exploratory Data Analysis
Procedure for constructing a boxplot
1. Find the five-number summary for the data values, that is, the
maximum and minimum data values, Q1 and Q3 , and the median.
2. Draw a horizontal axis with a scale such that it includes the
maximum and minimum data values.
3. Draw a box whose vertical sides go through Q1 and Q3 , and draw
a vertical line though the median.
4. Draw a line from the minimum data value to the left side of the
box and a line from the maximum data value to the right side of
the box.
Example 28
The number of meteorites found in 10 states of the United States is
89, 47, 164, 296, 30, 215, 138, 78, 48, 39. Construct a boxplot for the
data.
AMS (ITC) Descriptive Statistics October 6, 2024 58 / 63
Exploratory Data Analysis
Information Obtained from a Boxplot
a. If the median is near the center of the box, the distribution is
approximately symmetric.
b. If the median falls to the left of the center of the box, the
distribution is positively skewed.
c. If the median falls to the right of the center, the distribution is
negatively skewed.
AMS (ITC) Descriptive Statistics October 6, 2024 59 / 63
Exploratory Data Analysis
SKEWNESS
Another characteristic of a distribution is the shape. There are four
shapes commonly observed: symmetric, positively skewed, negatively
skewed, and bimodal. In a symmetric distribution the mean and
median are equal and the data values are evenly spread around these
values. The shape of the distribution below the mean and median is a
mirror image of distribution above the mean and median. A
distribution of values is skewed to the right or positively skewed if
there is a single peak, but the values extend much farther to the right
of the peak than to the left of the peak. In this case, the mean is
larger than the median. In a negatively skewed distribution there is a
single peak, but the observations extend farther to the left, in the
negative direction, than to the right. In a negatively skewed
distribution, the mean is smaller than the median. Positively skewed
distributions are more common. Salaries often follow this pattern. A
bimodal distribution will have two or more peaks. This is often the
case when the values are from two or more populations.
AMS (ITC) Descriptive Statistics October 6, 2024 60 / 63
Exploratory Data Analysis
Skewness
There are several formulas in the statistical literature used to calculate
skewness. The simplest, developed by Professor Karl Pearson
(1857–1936), is based on the difference between the mean and the
median.
AMS (ITC) Descriptive Statistics October 6, 2024 61 / 63
Exploratory Data Analysis
PEARSON’S COEFFICIENT OF SKEWNESS
3 (x̄ − Median)
sk =
s
Using this relationship, the coefficient of skewness can range from -3
up to 3. A value near -3, such as -2.57, indicates considerable negative
skewness. A value such as 1.63 indicates moderate positive skewness.
A value of 0, which will occur when the mean and median are equal,
indicates the distribution is symmetrical and there is no skewness
present
SOFTWARE COEFFICIENT OF SKEWNESS
" #
n X x − x̄ 3
sk =
(n − 1)(n − 2) s
AMS (ITC) Descriptive Statistics October 6, 2024 62 / 63
Exploratory Data Analysis
Example 29
Following are the earnings per share for a sample of 15 software
companies for the year 2017. The earnings per share are arranged from
smallest to largest.
$0.09 $0.13 $0.41 $0.51 $ 1.12 $ 1.20 $ 1.49 $3.18
3.50 6.36 7.83 8.92 10.13 12.99 16.40
Compute the mean, median, and standard deviation. Find the
coefficient of skewness using Pearson’s estimate and the software
methods. What is your conclusion regarding the shape of the
distribution?
AMS (ITC) Descriptive Statistics October 6, 2024 63 / 63