Week 2
DISC: 203 – PROBABILITY &
STATISTICS
Methods for Describing Sets of Data
1 Lecturer: Muhammad Asim
METHODS FOR DESCRIBING SETS OF
DATA
Describing qualitative data
Describing quantitative data
Numerical measures to describe data
Detecting outliers
2
DESCRIBING QUALITATIVE DATA
Pie Diagram
Bar Chart
3
EXAMPLE
Construct a pie diagram and a bar chart for highest
degrees attained by the CEOs.
CEO Highest Degree Company
A Masters AZ
B None BZ
C Doctorate CZ
D None DZ
E Bachelors EZ
F Masters FZ
G Bachelors GZ
H Bachelors HZ
I Doctorate IZ
J Bachelors JZ
K Doctorate KZ
L Bachelors LZ
M Masters MZ
N Masters NZ
O None OZ
P Doctorate PZ
Q Doctorate QZ
R Doctorate RZ
S Bachelors SZ
T Masters TZ
U Bachelors UZ
V Bachelors VZ 4
W None WZ
X Bachelors XZ
Y Bachelors YZ
FREQUENCY CHART FOR EDUCATION
ATTAINED BY EXECUTIVES IN THE
SAMPLE
Class Class Relative
Class Frequency Frequency
None 4 0.16
Bachelors 10 0.40
Masters 5 0.20
Doctorate 6 0.24
Total 25 1.00
5
PIE DIAGRAM
6
BAR CHART – CLASS FREQUENCY
7
BAR CHART– CLASS RELATIVE
FREQUENCY
8
DESCRIBING QUANTITATIVE DATA
Histogram
Scatterplot
Comparing central tendency and variability
9
HISTOGRAM
A histogram is a graphical depiction of the
class frequency or class relative frequency
table.
10
EXAMPLE
A large company is interested in determining
the length of service of its employees.
Twenty five employees are randomly chosen,
and the length of service (in years) is
recorded for each. The data are as follows:
3.1, 1.8, 6.4, 10.2, 11.2, 15.6, 11.6, 6.8, 1.5,
2.9, 3.4,7.2, 0.5, 7.7, 8.4, 0.7, 3.9, 8.2, 8.0,
5.5, 10.3, 12.1, 3.9, 0.9, 4.3
Construct a frequency histogram for the data.
11
HISTOGRAM
A Class Frequency table for length of service of employees in
the company is:
0.5 0.7 0.9 1.5 1.8 2.9 3.1 3.4
3.9 3.9 4.3 5.5 6.4 6.8 7.2 7.7
8.0 8.2
8.4 10.2 10.3 11.2 11.6 12.1 15.6
Class Class Frequency
0.0- 5.0 11
5.0-10.0 8
10.0-15.0 5
15.0-20.0 1
12
CLASS FREQUENCY HISTOGRAM
13
EFFECT OF SIZE OF DATA SET ON
HISTOGRAM
14
HISTOGRAM
While histograms provide good visual
representation of data sets (particularly of large
data sets)–they do not let us identify individual
measurements.
15
SCATTER PLOT
16
METHODS FOR DESCRIBING SETS OF
DATA
Describing qualitative data
Describing quantitative data
Numerical measures to describe data
Detecting outliers
17
NUMERICAL MEASURES
1. Central 2. Variability 3. Using mean and
Tendency Range standard deviation to
Mean Variance
describe data
Chebyshev’s Rule
Median Standard
The empirical rule
Mode deviation
Skewness
4. Relative
standing
Percentiles and
quartiles
Z-score
18
COMPARING CENTRAL TENDENCY
VERSUS VARIABILITY
Consider a histogram with a trend line:
Distribution
Cente
r
19
Spread/ Variability
EXERCISE
Which car would you buy?
20
MEASURES OF CENTRAL TENDENCY
Mean: The sample mean is defined as:
n
x i
Mean x i 1
n
Symbols for sample and population mean:
x Sample Mean
Population Mean
is often unknown so x is often used as an
approximater or estimator for
21
EXAMPLE: SUMMATION NOTATION
Suppose we have some sample data values
x1=2, x2=3, x3=6, x4=4, x5=6,
• And we want to sum (2+3+6+4+6) or in other words x1+
x2+ x3+ x4 + x5
• We may rewrite this as
5
Where x
i 1
i
x
Or we may rewrite as
i 1
i x1 x2 x3 x4 x5
2 3 6 4 6 21
5 22
x i 1
i 21
MEASURES OF CENTRAL TENDENCY
Median: Middle number when the measurement in
sample are arranged in ascending (or descending)
order.
Calculating a sample median m
Arrange n measurements from smallest to largest.
If n is odd, m is the middle number
If n is even, m is the mean of middle two number
Example: Calculate m for n={2,4,5,5,6,7,20} and 23
n={4,5,5,6,7,20}
MEASURES OF CENTRAL TENDENCY
Mode: Most frequently occurring measurement in the
data set
Modal class: identifies the area in which the data are
most concentrated.
Example: Find mode in previous example
Finding mode in quantitative (continuous) data Find
Modal Class Take mid point of modal class as mode
24
n={2,4,5,5,6,7,20}
25
MEASURES OF VARIABILITY
Range : Largest measurement in data set –
smallest measurement in data set
Average Absolute Deviation: Average
distance of each measurement from the
mean
Sample Variance = s2 : for a sample of n
measurements is equal to the sum of
squared distances from the mean divided by
n
(n-1)
2
(x x)
i
2
Sample Variance s i 1
n 1 26
MEASURES OF VARIABILITY
Sample Standard Deviation = s : is defined
as a positive square root of the sample
variance
s s2
Note that for population N
i
( x ) 2
Variance 2 i 1
N
Standard Deviation 2
27
MEASURES OF VARIABILITY
Example
Sample 1={1,3,5,7,9}
Sample 2={3,4,5,6,7}
Standard deviation is a measure of variability. So while
comparing two sample with equal means, but one sample
having higher standard deviation would imply the higher
variability in the sample with higher sample standard
deviation
28
CHEBYSHEV’S RULE
Applies to any set of data regardless of the
shape of the distribution
No useful information is provided within 1
standard deviation (μ – σ, μ + σ) of the mean.
At least ¾ of the measurements will fall
within 2 standard deviations of the mean (i.e.
μ – 2σ, μ + 2σ).
At least 8/9 of the measurements will fall
within 3 standard deviations of the mean (i.e.
μ – 3σ, μ + 3σ).
Generally, for any number k > 1, at least 1-
1/k2 of the measurements shall be within k 29
deviations of the mean (i.e. μ – kσ, μ + kσ).
EXERCISE
What percentage of values will approximately
fall between 10 and 26 for a data set with
mean of 18 and standard deviation of 2.
30
EXERCISE
What percentage of values will approximately
fall between 10 and 26 for a data set with
mean of 18 and standard deviation of 2.
Answer:
k = 4. So, at least 93.75% (1-1/16) of
the measurements will fall within the
interval [10, 26].
31
THE EMPIRICAL RULE
Applies to any data with symmetric or mount shape
distribution
Approximately 68% of the measurements will fall
within 1 standard deviation (μ – σ, μ + σ) of the
mean.
Approximately 95% of the measurements will fall
within 2 standard deviation (μ – 2σ, μ + 2σ).
Approximately 99.7% of the measurements will
fall within 3 standard deviation (μ – 3σ, μ + 3σ).
32
EMPIRICAL RULE
33
EXAMPLE
A manufacturer of automobile batteries claims that the
average length of life for its grade A battery is 60
months. However, the guarantee on this brand is for
just 36 months. Suppose the standard deviation of the
life length is known to be 10 months, and the frequency
distribution of the life-length data is known to be
mound-shaped.
(a) Approximately what percentage of the manufacturer’s
grade A batteries will last more than 50 months,
assuming that the manufacturer’s claim is true?
(b) Approximately what percentage of the manufacturer’s
batteries will last less than 40 months, assuming
manufacturer’s claim is true?
(c) Suppose your battery lasts 37 months. What could 34
you infer about the manufacturer’s claim?
A manufacturer of automobile batteries claims that the average
length of life for its grade A battery is 60 months. However, the
guarantee on this brand is for just 36 months. Suppose the
standard deviation of the life length is known to be 10 months, and
the frequency distribution of the life-length data is known to be
mound-shaped.
(a) Approximately what percentage of the manufacturer’s grade A
batteries will last more than 50 months, assuming that the
manufacturer’s claim is true? 84%
(b) Approximately what percentage of the manufacturer’s batteries
will last less than 40 months, assuming manufacturer’s claim is
true? 2.5%
(c) Suppose your battery lasts 37 months. What could you infer
about the manufacturer’s claim?
35
MEASURES OF RELATIVE STANDING
zscore: a relative standing measure that
uses mean and standard deviation
x μ x x
Population z score , Sample z score
s
Suppose a collected salary data with n = 200
ranges from $18,000 to $30,000 with mean
of $24,000 and standard deviation of $2,000.
Joe’s salary is $22,000. What is his z score?
z = (22,000-24,000)/2,000 = -1.0
Which tells us that Joe’s salary is 1.0 sample
standard deviation below average salary 36
EMPIRICAL RULE IN TERMS OF Z-SCORE
For mound shaped distributions,
Approximately 68% of the measurements shall have z
scores between -1 and + 1
Approximately 95% of the measurements shall have z
scores between -2 and + 2
Approximately 99.7% of the measurements shall have z
scores between -3 and + 3
34% 34%
2.5% 13.5 13.5 2.5%
% %
37
MEASURES OF RELATIVE STANDING
We may also be interested in relative quantitative
location of a particular data set. (example SAT
scores)
Percentile Rank: For any n measurements
(arranged descending or ascending), pth percentile
of a number is such that p% of the measurements
are same or fall below pth percentile and (100 – p)%
fall above it.
Example: (1,4,5,6,7,11,15,20,22,30}
pth percentile for 22= 90%
38
QUARTILES
Based on quartiles of data set.
Quartiles are the values that partition the data set
into four groups, each containing 25% of the
measurements.
Q is the 25th percentile of the data set. Middle
L
quartile m
is the median and
QU is the 75th
percentile
Interquartile range
(IQR) = QU-QL
39
QL m QU
METHODS FOR DESCRIBING SETS OF
DATA
Describing qualitative data
Describing quantitative data
Numerical measures to describe data
Detecting outliers
40
BOX PLOT
Box Plot: useful for comparing different samples and for
detecting outliers.
Box with hinges at Q and Q ;
L U
A line inside the box for Median ;
Whiskers (one from Q to the smallest measurement
L
between the inner fences and the other from QU to
the largest measurement between the inner fences);
Inner and Outer Fences. Neither visible. Inner at Q -
L
1.5(IQR) and QU + 1.5(IQR) ; and outer at QL - 3(IQR)
and QU + 3(IQR)
mild outliers (*) : values within and including the
inner fences and outer fences;
extreme outliers (o): values outside the outer fences.
EXAMPLE – BOX PLOT
Consider the 20 customer satisfaction
ratings:
1, 3, 5, 5, 7, 8, 8, 8, 8, 8, 8, 9, 9, 9, 9, 9, 10,
10, 10, 10
Construct a box plot for this data, identifying
mild and extreme outliers.
42
BOX PLOT
43
FOR NEXT WEEK
Download R at CRAN (Comprehensive R
Archive Network) website:
[Link]
Then download RStudio:
[Link]
44
RSTUDIO’S LOOK
45
R RESOURCES
R Manual:
[Link]
/[Link]
RStudio's own introductory primers:
[Link]
46
TEXT REFERENCE
Chapter 2
47