(3 ) Chapter
(3 ) Chapter
Hazar Khogeer
School of Mathematical
Umm-Al-Qura University
December 2, 2022
Introduction
Solution:
P
X 4 + 46 + 98 + 115 + 88 + 44 + 73 + 48 + 62
X̄ = = =
n 9
578
≈ 64.2
9
Hence, the mean number of flu cases over the 9-year period is
64.2.
The data show the systemwide sales (in millions) for U.S.
franchises of a well-known donut store for a 5-year period. Find the
mean : $221,$239, $262, $281 ,$318.
Solution:
P
X 221 + 239 + 262 + 281 + 318 1321
X̄ = = = = 264.2
n 5 5
The mean amount of sales for the stores over the 5-year period is
$264.2 million.
weighted mean
n
P
f · Xm
m=1
X̄ =
n P
Note: The symbols f · Xm mean to find the sum of the product of
the frequency (f ) and the midpoint (Xm ) for each class.
How to find a mean for the group data
Step 1: make a table as shown:
A B C D
Class Frequency f Midpoint xm f · xm
Step 2: find the midpoints of each class and place them in column C.
Step 3: multiply the frequency by the midpoint for each class, and place the
product in column D.
Step 4: find the sum of column D and divide the sum obtained in column D by the
sum of frequencies obtained in column B.
Example 3-4:
Mohammed year jogging data, he kept records of how many miles
he run each week. Find the mean of the miles that he run per
week?
172
Solution: x̄ = = 3.308.
52
11 /
129 Hazar Khogeer Introduction to statistics
3-1: Measures of Central Tendency: Median
Median: is the midpoint of the data after arranging the data in
ascending or descending order. The median is often denoted by
MD.
12 /
129 Hazar Khogeer Introduction to statistics
13 /
129 Hazar Khogeer Introduction to statistics
3-1: Measures of Central Tendency: Median
Example 3-4: Tablet Sales The data show the number of tablet
sales in millions of units for a 5-year period. Find the median of the
data: 108.2, 17.6, 159.8, 69.8, 222.6.
Solution:
Step 1: arrange the data in order:
17.6, 69.8, 108.2, 159.8, 222.6.
The median of the number of tablet sales for the 5-year period is
108.2 million
14 /
129 Hazar Khogeer Introduction to statistics
3-1: Measures of Central Tendency: Median
Example 3-4: Tablet Sales The data show the number of tablet
sales in millions of units for a 5-year period. Find the median of the
data: 108.2, 17.6, 159.8, 69.8, 222.6.
Solution:
Step 1: arrange the data in order:
17.6, 69.8, 108.2, 159.8, 222.6.
The median of the number of tablet sales for the 5-year period is
108.2 million
14 /
129 Hazar Khogeer Introduction to statistics
3-1: Measures of Central Tendency: Median
Example 3.6:
The number of tornadoes that have occurred in the United States
over an 8-year period follows. Find the median.
684, 764, 656, 702, 856, 1133, 1132, 1303
Solution:
Step 1: arrange the data in order:
656, 684, 702, 764, 856, 1132,1133,1303.
Step 4: since the middle point falls halfway between 764 and 856,
find the median MD by adding the two values and dividing by 2.
764 + 856
15 /
MD= = 810.
129 2 Khogeer
Hazar Introduction to statistics
3-1: Measures of Central Tendency: Median
Example 3.6:
The number of tornadoes that have occurred in the United States
over an 8-year period follows. Find the median.
684, 764, 656, 702, 856, 1133, 1132, 1303
Solution:
Step 1: arrange the data in order:
656, 684, 702, 764, 856, 1132,1133,1303.
Step 4: since the middle point falls halfway between 764 and 856,
find the median MD by adding the two values and dividing by 2.
764 + 856
15 /
MD= = 810.
129 2 Khogeer
Hazar Introduction to statistics
3-1: Measures of Central Tendency: Median for Grouped
Data
An approximate median can be found for grouped data.
First it is necessary to find the median class.
This is the class that contains the median value.
n/2 − cf
The formula is MD= × (w ) + Lm .
f
Where:
n sum of frequencies.
n/2 − cf
The formula is MD= × (w ) + Lm .
f
17 /
129 Hazar Khogeer Introduction to statistics
3-1 Summarize data, Measures of Central Tendency: Mode
Mode: is the value that occurs most often in a data set.
18 /
129 Hazar Khogeer Introduction to statistics
3-1: Measures of Central Tendency: Mode
Solution:
It is helpful to arrange the data in order, although it is not
necessary.
21,77,77,101, 159, 311, 382.
19 /
129 Hazar Khogeer Introduction to statistics
3-1: Measures of Central Tendency: Mode
Solution:
It is helpful to arrange the data in order, although it is not
necessary.
21,77,77,101, 159, 311, 382.
19 /
129 Hazar Khogeer Introduction to statistics
3-1: Measures of Central Tendency: Mode
EXAMPLE 3-7: Licensed Nuclear Reactors
The data show the number of licensed nuclear reactors in the
United States for a recent 15-year period.
Find the mode :
104 104 104 104 104
107 109 109 109 110
109 111 112 111 109
Solution:
Since the values 104 and 109 both occur 5 times, the modes are
104 and 109. The data set is said to be bimodal.
Solution:
Class Frequency
15.5-20.5 13
20.5-25.5 6
25.5-30.5 4
30.5-35.5 1
35.5-40.5 1
Since the class 15.5-20.5 has the largest frequency, 13, it is the
modal class. Sometimes the midpoint of the class is used. In this
case, it is 18.
21 /
129 Hazar Khogeer Introduction to statistics
3-1: Measures of Central Tendency: Mode
Solution:
Drinks Gallons
Soft drinks 52
Water 34
Milk 26
Coffee 21
Since the category of soft drinks has the largest frequency, 52, we
can say that the mode or most typical drink is a soft drink.
22 /
129 Hazar Khogeer Introduction to statistics
3-1: Measures of Central Tendency: Mean, Median,
Mode
EXAMPLE 3-11 Salaries of Personnel
A small company consists of the owner, the manager, the
salesperson, and two technicians, all of whose annual salaries are
listed here. (Assume that this is the entire population).
Staff Salary
Owner $100,0000
Manager 40,000
Salesperson 24,000
Technician 18,000
Technician 18,000
Find the mean, median, and mode?
Solution:
100, 000 + 40, 000 + 24, 000 + 18, 000 + 18, 000 $200, 000
P
X
µ= = = =
N 5 5
$40,000.
Hence, the mean is $40,000, the median is $24,000, and the mode
23 /
129 is $18,000. Hazar Khogeer Introduction to statistics
3-1 Summarize data, Measures of Central Tendency: Midrange
Midrange
Low value + high value
MR=
2
24 /
129 Hazar Khogeer Introduction to statistics
3-1 Summarize data, Measures of Central Tendency: Midrange
EXAMPLE 3-12 Bank Failures
3 + 157
MR= = 80.
2
The midrange for the number of bank failures is 80.
Example 3.8
The following data represent ages of a sample of 9 students:
18, 17, 19, 22, 20, 24, 19, 21, 20. Find the midrange ?
Solution:
17 + 24
MR= = 20.5.
25 /
2
129 Hazar Khogeer Introduction to statistics
3-1 Summarize data, Measures of Central Tendency: Midrange
EXAMPLE 3-12 Bank Failures
3 + 157
MR= = 80.
2
The midrange for the number of bank failures is 80.
Example 3.8
The following data represent ages of a sample of 9 students:
18, 17, 19, 22, 20, 24, 19, 21, 20. Find the midrange ?
Solution:
17 + 24
MR= = 20.5.
25 /
2
129 Hazar Khogeer Introduction to statistics
3-1 Summarize data, Measures of Central Tendency: Midrange
EXAMPLE 3-12 Bank Failures
3 + 157
MR= = 80.
2
The midrange for the number of bank failures is 80.
Example 3.8
The following data represent ages of a sample of 9 students:
18, 17, 19, 22, 20, 24, 19, 21, 20. Find the midrange ?
Solution:
17 + 24
MR= = 20.5.
25 /
2
129 Hazar Khogeer Introduction to statistics
3-1 Summarize data, Measures of Central Tendency: Weighted Mean
When to use?
Used when values are grouped by frequency or relative
importance.
Weighted Mean
P
w1 X1 + w2 X2 + w3 X3 + · · · + wn Xn w ·X
X¯w = = P
w1 + w2 + w3 + · · · + wn w
26 /
129 Hazar Khogeer Introduction to statistics
3-1 Summarize data, Measures of Central Tendency: Weighted Mean
EXAMPLE 3-14 Grade Point Average
A student received the following grades. A in English Composition
(3 credits), a C in Introduction to Psychology (3 credits), a B in
Biology (4 credits), and a D in Physical Education (2 credits).
Assuming: A = 4 grade points , B = 3 grade points , C = 2 grade
points , D = 1 grade point, and F= 0 grade points.
Find the students grade point average?
Solution:
P
w i xi 3·4+3·2+4·3+2·1 32
X¯w = P = = ≈ 2.7 GPA
wi 3+3+4+2 12
29 /
129 Hazar Khogeer Introduction to statistics
The Median:
1 The median is used to find the center or middle value of a
data set.
2 The median is used when it is necessary to find out whether
the data values fall into the upper half or lower half of the
distribution.
3 The median is used for an open-ended distribution.
4 The median is affected less than the mean by extremely high
or extremely low values.
The Mode:
1 The mode is used when the most typical case is desired.
2 The mode is the easiest average to compute.
3 The mode can be used when the data are nominal or
categorical, such as religious preference, gender, or political
affiliation.
4 The mode is not always unique. A data set can have more
than one mode, or the mode may not exist for a data set.
30 /
129 Hazar Khogeer Introduction to statistics
The Midrange:
1 The midrange is easy to compute.
2 The midrange gives the midpoint.
3 The midrange is affected by extremely high or low values in a
data set.
31 /
129 Hazar Khogeer Introduction to statistics
3-1: Measures of Central Tendency: Distributions
Types of Distributions:
32 /
129 Hazar Khogeer Introduction to statistics
3-2 Measures of Variability: Range
Measures of Variation:
Range,
Variance,
Standard Deviation.
33 /
129 Hazar Khogeer Introduction to statistics
3-2: Measures of Variation
EXAMPLE 3-15 Comparison of Outdoor Paint
A testing lab wishes to test two experimental brands of outdoor
paint to see how long each will last before fading. The testing lab
makes 6 gallons of each paint to test. Since different chemical
agents are added to each group and only six cans are involved,
these two groups constitute two small populations. The results (in
months) are shown. Find the mean of each group:
Solution: P
X 210
The mean for brand A is µ= = = 35 months
N 6
P
X 210
34 / The mean for brand B is µ= = = 35 months
129 Hazar Khogeer Introduction to6
N statistics
3-2 Measures of Variation
Range
R= xmax − xmin
36 /
129 Hazar Khogeer Introduction to statistics
Example 3.16: Comparison of Outdoor Paint
Find the ranges for the paints in Example 3-15.
Solution:
Rang of Brand A: Rang of Brand B:
R = 60 − 10 = 50 months, R = 45 − 25 = 20 months.
Make sure the range is given as a single number. The range for
brand A shows that 50 months separate the largest data value
from the smallest data value. For brand B , 20 months separate
the largest data value from the smallest data value, which is less
than one-half of brand A’s range.
37 /
129 Hazar Khogeer Introduction to statistics
3-2 Measure of Variability
Variance: Measures of variation gives information on the speed of
variability of the data values. It measures the average deviation of
the measurements about their mean.
40 /
129 Hazar Khogeer Introduction to statistics
3-2 Measure of Variability: Variance, and Standard Deviation
Solution:
Step 1: Find the mean for the data
P
X 10 + 60 + 50 + 30 + 40 + 20
µ= = = 35
N 6
Step 2: Subtract the mean from each data value (xi − µ)
Step 3: Square each result (xi − µ)2
Step 4: Find the sum of the squares (xi − µ)2
P
(xi − µ)2
P
Step 5: Divide the sum by N to get the variance : σ2 =
N rP
(xi − µ)2
Step 6: Take the square root to get the standard deviation σ =
N
41 /
129 Hazar Khogeer Introduction to statistics
EXAMPLE 3-18 (continued)
Solution:
A B C D
i xi µ xi − µ (xi − µ)2
1 20 35 -15 225
2 40 35 5 25
3 30 35 -5 25
4 50 35 15 225
5 60 35 25 625
6 10 35 -25 625
(xi − µ) = 0 (xi − µ)2 =1750
P P P
Total xi = 210
N = 6P
(xi − µ)2 1750
σ2 = = ≈ 291.7
rP N 6
(xi − µ)2 √
σ= = 291.7 = 17.1
N
Column A contains the raw data X. Column B contains the mean of X . Column C
differences X − µ obtained in step 2. Column D contains the squares of the
differences obtained in step 3.
42 /
129 Hazar Khogeer Introduction to statistics
EXAMPLE 3-19 Comparison of Outdoor Paint
Find the variance and standard deviation for brand B paint data in
Example 3-15. The months brand B lasted before fading was :
35 45 30 35 40 25
Solution:
A B C D
i Xi µ Xi − µ (Xi − µ)2
1 35 35 0 0
2 45 35 10 100
3 30 35 -5 25
4 35 35 0 0
5 40 35 5 25
6 25 35 -10 100
(xi − µ) = 0 (xi − µ)2 =250
P P P
Total xi = 210
(xi − µ)2
P P
xi 1210 250
µ= = = 35, σ =
2
= = 41.7,
N 6 rP N 6
(xi − µ)2 √
σ= = 41.7 ≈ 6.5.
43 /
N
129 Hence, the standard deviation is 6.5.
Hazar Khogeer Introduction to statistics
3-2 Measure of Variability: Variance, and Standard Deviation
Sample variance
(xi − x̄ )2
P
2
s = .
n−1
44 /
129 Hazar Khogeer Introduction to statistics
3-2 Measure of Variability: Variance, and Standard Deviation
Stander deviation: is the square root of the variance.
Stander deviation
√
s= s2
Comparing Standard Deviations
45 /
129 Hazar Khogeer Introduction to statistics
3-2 Measure of Variability: Variance, and Standard Deviation
EXAMPLE 3-20 Teacher Strikes
The number of public school teacher strikes in Pennsylvania for a
random sample of school years is shown. Find the sample
variance and the sample standard deviation:
9 10 14 7 8 3
Solution:
Step 1: Find the mean of the data values :
P
X 9 + 10 + 14 + 7 + 8 + 3 51
X̄ = = = = 8.5,
n 6 6
Step 2: Find the deviation for each data value(xi − x̄ ),
Step 3: Square each of the deviations (xi − x̄ )2 ,
Step 4: Find the sum of the squares (xi − x̄ )2 ,
P
Step 5: Divide by n − 1 to get the variance,
Step 6: Take the square root of the variance to get the standard
deviation.
46 /
129 Hazar Khogeer Introduction to statistics
3-2 Measure of Variability: Variance, and Standard Deviation
EXAMPLE 3-20 (continued)
Solution:
i xi x̄ xi − x̄ (xi − x̄ )2
1 9 8.5 0.5 0.25
2 10 8.5 1.5 2.25
3 14 8.5 5.5 30.25
4 7 8.5 -1.5 2.25
5 8 8.5 -0.5 0.25
6 3 8.5 -5.5 30.25
(xi − x̄ )2 = 65.5
P P P
Total xi = 51 (xi − x̄ ) = 0
(xi − x̄ )2 65.5
P
2
s = = = 13.1 and
n − 1r 5
√ (xi − x̄ )2 √
P
s = s2 = = 13.1 ≈ 3.6(rounded ). Here the
n−1
sample variance is 13.1, and the sample standard deviation is 3.6.
47 /
129 Hazar Khogeer Introduction to statistics
3-2 Measure of Variability: Variance, and Standard Deviation
48 /
129 Hazar Khogeer Introduction to statistics
3-2 Measure of Variability: Variance, and Standard Deviation
(xi − x̄ )2 n (xi )2 − ( xi )2
P P P
2
Prove: s = =
n−1 n(n − 1)
49 /
129 Hazar Khogeer Introduction to statistics
EXAMPLE 3-21 Teacher Strikes
The number of public school teacher strikes in Pennsylvania for a
random sample of school years is shown. Find the sample
variance and the sample standard deviation:
9 10 14 7 8 3
Solution:
Solution : Step 1: Find the sum of the values :
X = 9 + 10 + 14 + 7 + 8 + 3 = 51.
P
Step 2: Square each value and find the sum :
X = 92 + 102 + 142 + 72 + 82 + 32 = 499.
P 2
50 /
129 Hazar Khogeer Introduction to statistics
3-2 Measure of Variability
xi2 − ( xi )2
P P
2
n
s =
n(n − 1)
6(499) − 512
=
6(6 − 1)
2994 − 2601
=
6(5)
393
= = 13.1
30
The variance is 13.1.
√
s= 13.1 ≈ 3.6 (rounded)
Hence, the sample variance is 13.1, and the sample standard deviation is 3.6.
Notice that these are the same results as the results in Example 3-20.
51 /
129 Hazar Khogeer Introduction to statistics
3-2 Measure of Variability
Example 3.9:
The following are weights of 5 persons: 60 90 80 70 50. Find the
variance and standard deviation?
Solution:
i xi xi2
1 60 3600
2 90 8100
3 80 6400
4 70 4900
5 50 2500
P P 2
Total xi = 350 xi = 25500
53 /
129 Hazar Khogeer Introduction to statistics
3-2 Variance and Standard Deviation for Grouped Data
Find the sample variance and the sample standard deviation for
the frequency distribution of the data shown. The data represent
the number of miles that 20 runners ran during one week:
54 /
129 Hazar Khogeer Introduction to statistics
3-2 Variance and Standard Deviation for Grouped Data
EXAMPLE 3-22 (continued)
Solution :
Step 1: Make a table as shown, and find the midpoint of each class
A B C D E
2
Class bounderies Frequency Midpoint Xm f · xm f · xm
5.5-10.5 1 8
10.5-15.5 2 13
15.5-20.5 3 18
20.5-25.5 5 23
25.5-30.5 4 28
30.5-35.5 3 33
35.5-40.5 2 38
Step 2: Multiply the frequency by the midpoint for each class, and
place the Products in column D :
1·8=8 2 · 13 = 26 ··· 2 · 38 = 76
55 /
129 Hazar Khogeer Introduction to statistics
3-2 Variance and Standard Deviation for Grouped Data
56 /
129 Hazar Khogeer Introduction to statistics
A B C D E
Class bounderies Frequency Midpoint Xm f · xm f · xm2
5.5-10.5 1 8 8 64
10.5-15.5 2 13 26 338
15.5-20.5 3 18 54 972
20.5-25.5 5 23 115 2,645
25.5-30.5 4 28 112 3,136
30.5-35.5 3 33 99 3,267
35.5-40.5 2 38 76 2,888
f · xm2 = 13, 310
P P
n = 20 f · xm = 490
2
Step 5: SubstitutePin the formula
P and 2solve for s to get the
2
n f · xm − ( f · xm )
variance : s 2 =
n(n − 1)
20(13, 310) − 4902
=
20(20 − 1)
266, 200 − 240, 100 26, 100
= = ≈ 68.7
20(19) 380
Step 6: Take the √
square root to get the standard deviation :
57 / s ≈ 68.7 ≈ 8.3
129 Hazar Khogeer Introduction to statistics
3-2 Measure of Variability
58 /
129 Hazar Khogeer Introduction to statistics
59 /
129 Hazar Khogeer Introduction to statistics
3-2 Measure of Variability
60 /
129 Hazar Khogeer Introduction to statistics
3-2 Measure of Variability
Solution:
s 5.59cm
Height: CVar = · 100% = · 100% = 7.5%
x̄ 75cm
s 15.81kg
Weight: CVar = · 100% = · 100% = 22.6%
x̄ 70kg
The weight distribution has a higher variability than the height.
61 /
129 Hazar Khogeer Introduction to statistics
3-2 Measure of Variability
63 /
129 Hazar Khogeer Introduction to statistics
3-2 Measure of Variability
64 /
129 Hazar Khogeer Introduction to statistics
3-2 Chebyshevs Theorem
65 /
129 Hazar Khogeer Introduction to statistics
3-2 Chebyshevs Theorem
EXAMPLE 3-26 Travel Allowances
A survey of local companies found that the mean amount of travel
allowance for couriers was $0.25 per mile. The standard deviation
was $0.02. Using Chebyshevs theorem, find the minimum
percentage of the data values that will fall between $0.20 and
$0.30.
Solution :
Step 1: Subtract the mean from the larger value :
$0.30-$0.25=$0.05.
Step 2 Divide the difference by the standard deviation to get k.
0.05
k= = 2.5.
0.02
Step 3: Use Chebyshevs theorem to find the percentage :
1 1 1
1− 2 =1− =1− = 1 − 0.16 = 0.84
k 2.52 6.25
Hence, at least 84% of the data values will fall between $0.20 and
66 /
129 $0.30. Hazar Khogeer Introduction to statistics
3-3 Measures of Position
67 /
129 Hazar Khogeer Introduction to statistics
3-3 Measures of Position
value − mean
z=
standard deviation
X − X̄
For Samples, the formula is: z= .
s
X −µ
For Population, the formula is: z= .
σ
68 /
129 Hazar Khogeer Introduction to statistics
3-3 Measures of Position
EXAMPLE 3-27 Test Scores
A student scored 85 on an English test while the mean score of all
the students was 76 and the standard deviation was 4. She also
scored 42 on a French test where the class mean was 36 and the
standard deviation was 3. Compare the relative positions on the
two tests.
Solution :
First find the z scores.
For the English test :
X − X̄ 85 − 76
z= = = 2.25.
s 4
For the French test
X − X̄ 42 − 36
z= = = 2.00.
s 3
Since the z score for the English test is higher than the z score for
the French test, her relative position in the English class is higher
than her relative position in the French class.
69 /
129 Hazar Khogeer Introduction to statistics
3-3 Measures of Position
Percentiles: divide the data set into 100 equal groups.
70 /
129 Hazar Khogeer Introduction to statistics
3-3 Measures of Position
EXAMPLE 3-29 Systolic Blood Pressure
The frequency distribution for the systolic blood pressure readings
(in millimeters of mercury, mm Hg) of 200 randomly selected
college students is shown here. Construct a percentile graph :
71 /
129 Hazar Khogeer Introduction to statistics
Solution :
Step 1: Find the cumulative frequencies and place them in column
C.
Step 2: Find the cumulative percentages and place them in
column D. To do this step, use the formula :
cumulative frequency
Cumulative % = · 100
n
24
For the first class, Cumulative % = · 100 = 12%
200
72 /
129 Hazar Khogeer Introduction to statistics
Step 3: Graph the data, using class boundaries for the x axis and
the percentages for the y axis, as shown in Figure.
73 /
129 Hazar Khogeer Introduction to statistics
74 /
129 Hazar Khogeer Introduction to statistics
We used the following rule after ascending order the data:
# of values less than X + 0.5
P= × 100
n
EXAMPLE 3-30 Traffic Violations
The number of traffic violations recorded by a police department
for a 10-day period is shown. Find the percentile rank of 16.
22 19 25 24 18 15 9 12 16 20
Solution :
Arrange the data in order from lowest to highest.
9 12 15 16 18 19 20 22 24 25
Then substitute into the formula.
Step 2: Compute :
n·p
c=
100
where n = total number of values, p = percentile
Thus,
10 · 65
c= = 6.5
100
Since c is not a whole number, round it up to the next whole
number; in this case, it is c = 7. Start at the lowest value and
count over to the 7th value, which is 20. Hence, the value of 20
corresponds to the 65 th percentile.
77 /
129 Hazar Khogeer Introduction to statistics
EXAMPLE 3-33 Traffic Violations
Using the data in Example 3-30, find the data value corresponding
to the 30th percentile.
Solution :
Step 1: Arrange the data in order from lowest to highest.
9 12 15 16 18 19 20 22 24 25
Step 2: Substitute in the formula.
n·p 10 · 30
c= , c= =3
100 100
In this case, it is the 3rd and 4th values.
Step 3: Since c is a whole number, use the value halfway between
the c and c + 1 values when counting up from the lowest. In this
case, it is the third and fourth values.
79 /
129 Hazar Khogeer Introduction to statistics
3-3: Measures of position
Other Location Measures
Quantiles: divides the data set into 4 equal parts
Q1 = P25 , Q2 = P50 = MD , Q3 = P75 .
Solution :
FindQ1 and Q3 . This was done in Example 3-34.
Now Q1 = 15 and Q3 = 22.
84 /
129 Hazar Khogeer Introduction to statistics
3-3 Measures of position
Outlier :
An outlier is an extremely high or low data value when compared
with the rest of the data value.
To detect an outlier:
Obtain Q1 , Q3 and IQR.
Outlier
a data that is smaller than: Q1 − 1.5 × (IQR ), or
85 /
129 Hazar Khogeer Introduction to statistics
3-3 Measures of position
EXAMPLE 3-36 Outliers
50 22 18 15 13 6 12 5
Solution :
The data value 50 is extremely suspect. These are the steps in
checking for an outlier.
6 + 12 18 + 22
Q1 = =9 Q3 = = 20.
2 2
Step 2: Find the interquartile range IQR = Q3 − Q1 :
86 /
IQR = Q3 − Q1 = 20 − 9 = 11
129 Hazar Khogeer Introduction to statistics
3-3 Measures of position
EXAMPLE 3-36 Outliers
50 22 18 15 13 6 12 5
Solution :
The data value 50 is extremely suspect. These are the steps in
checking for an outlier.
6 + 12 18 + 22
Q1 = =9 Q3 = = 20.
2 2
Step 2: Find the interquartile range IQR = Q3 − Q1 :
86 /
IQR = Q3 − Q1 = 20 − 9 = 11
129 Hazar Khogeer Introduction to statistics
Step 3: Multiply this value by 1.5 :
1.5(11)=16.5.
Step 4: Subtract the value obtained in step 3 from Q1 , and add the
value obtained in step 3 to Q3 .
Step 5: Check the data set for any data values that fall outside the
interval from -7.5 to 36.5. The value 50 is outside this interval;
hence, it can be considered an outlier.
87 /
129 Hazar Khogeer Introduction to statistics
3-4: Exploratory data analysis
88 /
129 Hazar Khogeer Introduction to statistics
89 /
129 Hazar Khogeer Introduction to statistics
3-4: Exploratory data analysis
EXAMPLE 3-37 Number of Meteorites Found
The number of meteorites found in 10 states of the United States
is:
89,47,164,296,30,215,138,78,48,39
30,39,47,48,78,89,138,164,215,296
↑
Median
78 + 89
median = = 83.5
2
90 /
129 Hazar Khogeer Introduction to statistics
3-4: Exploratory data analysis
EXAMPLE 3-37 Number of Meteorites Found
The number of meteorites found in 10 states of the United States
is:
89,47,164,296,30,215,138,78,48,39
30,39,47,48,78,89,138,164,215,296
↑
Median
78 + 89
median = = 83.5
2
90 /
129 Hazar Khogeer Introduction to statistics
Find Q1 :
30,39,47,48,78
↑
Q1
Find Q3 :
89,138,164,215,296
↑
Q3
The minimum data value is 30, and the maximum data value is
296.
91 /
129 Hazar Khogeer Introduction to statistics
Step 3: Draw the box above the scale using Q1 and Q3 . Draw a
vertical line through the median, and draw lines from the lowest
data value to the box and from the highest data value to the box.
See Figure 3-7.
92 /
129 Hazar Khogeer Introduction to statistics
Distribution Shape and Box Plot
93 /
129 Hazar Khogeer Introduction to statistics
3-4: Exploratory data analysis
EXAMPLE 3-38 Speeds of Roller Coasters
The data shown are the speeds in miles per hour of a sample of
wooden roller coasters and a sample of steel roller coasters.
Compare the distributions by using boxplots :
Solution:
Step 1: For the wooden coasters :
35 48 50 56 60 67 68 72
↑ ↑ ↑
Q1 MD Q3
48 + 50 56 + 60 67 + 68
Q1 = = 49, MD = = 58, Q3 = = 67.5.
2 2 2
94 /
129 Hazar Khogeer Introduction to statistics
3-4: Exploratory data analysis
EXAMPLE 3-38 Speeds of Roller Coasters
The data shown are the speeds in miles per hour of a sample of
wooden roller coasters and a sample of steel roller coasters.
Compare the distributions by using boxplots :
Solution:
Step 1: For the wooden coasters :
35 48 50 56 60 67 68 72
↑ ↑ ↑
Q1 MD Q3
48 + 50 56 + 60 67 + 68
Q1 = = 49, MD = = 58, Q3 = = 67.5.
2 2 2
94 /
129 Hazar Khogeer Introduction to statistics
Step 2 For the steel coasters : 28 48 55 70 100 102 106 120
↑ ↑ ↑
Q1 MD Q3
48 + 55
Q1 = = 51.5,
2
70 + 100
MD = = 85,
2
102 + 106
Q3 = = 104.
2
95 /
129 Hazar Khogeer Introduction to statistics
Step 2 For the steel coasters : 28 48 55 70 100 102 106 120
↑ ↑ ↑
Q1 MD Q3
48 + 55
Q1 = = 51.5,
2
70 + 100
MD = = 85,
2
102 + 106
Q3 = = 104.
2
95 /
129 Hazar Khogeer Introduction to statistics
Step 3: Draw the boxplots. See Figure 3-8.
The boxplots show that the median of the speeds of the steel
coasters is much higher than the median speeds of the wooden
coasters. The interquartile range (spread) of the steel coasters is
much larger than that of the wooden coasters. Finally, the range of
the speeds of the steel coasters is larger than that of the wooden
coasters.
96 /
129 Hazar Khogeer Introduction to statistics
3-4 Statistical Modeling, Scientific Inspection, and Graphical Diagnostics
Example 3.13:Drew the box-plot for patient stay in the hospital ?
Patient 1 2 3 4 5 6 7 8 9
Solution:
First arrange the data in order: 3,4,6,7,9,12,12,18,55
Patient 1 2 3 4 5 6 7 8 9
Solution:
First arrange the data in order: 3,4,6,7,9,12,12,18,55
98 /
129 Hazar Khogeer Introduction to statistics
3-4 Data Description: Box Plot
Example 3.14:Check the values of data set for any outliers for
patient stay in the hospital in(Example 3-13)?
3,4,6,7,9,12,12,18,55
↑ ↑ ↑ ↑ ↑
Min Q1 MD Q3 Max
Q1 = 5
Q2 = 9
Q3 15
Solution:
IQR = Q3 − Q1 = 15 − 5 = 10
Lower bound=Q1 − 1.5 × (IQR ) = 5 − 1.5 × 10 = −10,
Upper bound= Q3 + 1.5 × (IQR ) = 15 + 1.5 × 10 = 30.
99 /
129 Hazar Khogeer Introduction to statistics
3-4 Data Description: Box Plot
Example 3.14:Check the values of data set for any outliers for
patient stay in the hospital in(Example 3-13)?
3,4,6,7,9,12,12,18,55
↑ ↑ ↑ ↑ ↑
Min Q1 MD Q3 Max
Q1 = 5
Q2 = 9
Q3 15
Solution:
IQR = Q3 − Q1 = 15 − 5 = 10
Lower bound=Q1 − 1.5 × (IQR ) = 5 − 1.5 × 10 = −10,
Upper bound= Q3 + 1.5 × (IQR ) = 15 + 1.5 × 10 = 30.
99 /
129 Hazar Khogeer Introduction to statistics
3-4 Data Description: Box Plot
Example 3.14:Check the values of data set for any outliers for
patient stay in the hospital in(Example 3-13)?
3,4,6,7,9,12,12,18,55
↑ ↑ ↑ ↑ ↑
Min Q1 MD Q3 Max
Q1 = 5
Q2 = 9
Q3 15
Solution:
IQR = Q3 − Q1 = 15 − 5 = 10
Lower bound=Q1 − 1.5 × (IQR ) = 5 − 1.5 × 10 = −10,
Upper bound= Q3 + 1.5 × (IQR ) = 15 + 1.5 × 10 = 30.
99 /
129 Hazar Khogeer Introduction to statistics
3-4: Exploratory data analysis
100 /
129 Hazar Khogeer Introduction to statistics
Further Moments of the Distribution
101 /
129 Hazar Khogeer Introduction to statistics
Moments
102 /
129 Hazar Khogeer Introduction to statistics
Central (or Mean) Moments
In mean moments: the deviations are taken from the mean.
For Ungrouped Data:
103 /
129 Hazar Khogeer Introduction to statistics
Moments
Assignment
104 /
129 Hazar Khogeer Introduction to statistics
Moments(Assigment)
P
f · Xm
x̄ = P = 122.5
P f
f (xm − x̄ )2
P
f (xm − x̄ )
m1 = P = 0, m2 = P = 1216,
f f
f (xm − x̄ )3 f (xm − x̄ )4
P P
m3 = P = 23104, m4 = P = 3717632.
f f
106 /
129 Hazar Khogeer Introduction to statistics
Moments about (arbitrary) Origin
If the deviations are taken from some arbitrary number (’a’ called
origin), then moments are called moments about arbitrary origin ’a’.
For Ungrouped Data:
107 /
129 Hazar Khogeer Introduction to statistics
Moments about zero
108 /
129 Hazar Khogeer Introduction to statistics
Moments
Assignment
Example: Calculate first four moments about zero (origin) for the
following two set of sets of data.
109 /
129 Hazar Khogeer Introduction to statistics
Moments(Assignment)(Continued)
Example: Calculate first four moments about zero (origin) for the
following set of data.
X (X )2 (X )3 (X )4
45 2025 91125 4100625
32 1024 32768 1048576
37 1369 50653 1874161
46 2116 97336 4477456
39 1521 59319 2313441
36 1296 46656 1679616
41 1681 68921 2825761
48 2304 110592 5308416
36 1296 46656 16796
Total 14632 604026 25307668
(xi )2
P P
xi
m1 = = 40, m2 = = 1625.778, n=9
Pn 3 n P
4
(xi ) ( xi )
m3 = = 67114, m4 = = 2811963.1111.
n n
110 /
129 Hazar Khogeer Introduction to statistics
Moments(Assignment)(Continued)
Example: Calculate first four moments about zero (origin) for the
following set of data.
X f xm 2
Xm 3
Xm 4
Xm f · Xm f · (Xm )2 f · (X m )3 f · (X m )4
65-84 9 74.5 5550.25
85-104 10 94.5 8930.25
105-124 17 114.5 13110.24
125-144 10 134.5 18090.25
145-164 5 154.5 23870.25
165-184 4 174.5 30450.25
185-204 5 194.5 37830.25
Total 60 941.5 137831.75 21551170.37 3837824815 7350
2 3
= 138494977.5,
P P
f · Xm = 973335, f · Xm
f · Xm = 20982703863.5,
4
f = 60.
P P
f (xi )2
P P
0 f ( xi ) 0
m1 = P = 122.5, m2 = P = 16222.25,
P f 3 f
0 f ( xi )
m3 = P = 2308249.625,
P f 4
0 f ( xi )
m4 = P = 349711731.0583.
111 /
f
129 Hazar Khogeer Introduction to statistics
Moments
113 /
129 Hazar Khogeer Introduction to statistics
Skewness
114 /
129 Hazar Khogeer Introduction to statistics
Measures of Symmetry (1)
Skewness
Symmetric distribution.
Negatively skewed (Left) distribution.
Positively skewed (Right) distribution.
If Mean = Mode, the skewness is zero.
If Mean < Mode, the skewness is negative.
If Mean > Mode, the skewness is positive.
115 /
129 Hazar Khogeer Introduction to statistics
Measures of Symmetry (2)
They may be tail off to right or to the left and as such said to
be skewed.
116 /
129 Hazar Khogeer Introduction to statistics
Measures of Skewness
(Mean − Mode )
Person’s coefficient of skewness =
Standard deviation
3(Mean − Median)
Person’s coefficient of skewness =
Standard deviation
D9 + D1 − 2Me
Kelley’s coefficient of skewness =
D9 − D1
(Q3 + Q1 ) − 2Me
Bowler’s coefficient of skewness =
Q3 − Q1
µ23
Pn
i =1 (xi − x̄ )3
Moment based coefficient of skewness =
117 /
µ32 ns 3
129 Hazar Khogeer Introduction to statistics
Symmetric Distribution
Symmetric Distribution
118 /
129 Hazar Khogeer Introduction to statistics
Right Skewed Distributions
120 /
2. When the median is greater the the mean.
129 Hazar Khogeer Introduction to statistics
Example(1)
121 /
129 Hazar Khogeer Introduction to statistics
3-4 Measures of Skewness
122 /
129 Hazar Khogeer Introduction to statistics
Example: Skewness using Moment Based Coefficient
Xi f Xi · f (Xi − X̄ ) (X − X̄ )2 f · (X − X̄ )2 (X − X̄ )3 f · (X − X̄ )3
1 1 1 -3.23 10.43 10.44 -33.70 -33.70
2 4 8 -2.23 4.97 19.91 -11.09 -44.36
3 8 24 -1.23 1.51 12.12 -1.86 -14.89
4 4 16 -0.23 0.05 0.21 -0.01 -0.05
5 3 15 0.77 0.59 1.78 0.46 1.37
6 2 12 1.77 3.13 6.26 5.54 11.09
7 1 7 2.77 7.67 7.67 21.25 21.25
8 1 8 3.77 14.21 14.21 53.58 53.58
9 1 9 4.77 22.75 22.75 108.53 108.53
10 1 10 5.77 33.28 33.28 192.10 192.10
sum 26 110 128.62 294.94
Pn
i =1 f × Xi 110
Mean X̄ = Pn = = 4.23
i =1 f s P26
128.62
n 2
r
i =1 f × (Xi − X̄ )
Standard Deviations = Pn = = 2.27
( i =1 f ) − 1 25
294.94
Pn
f × (Xi − X̄ )3
Skewness = i =1 = = 0.97
123 /
ns 3 (26)(2.27)3
129 Hazar Khogeer Introduction to statistics
Assignment
124 /
129 Hazar Khogeer Introduction to statistics
3 Data Description: Kurtosis
Kurtosis is a measure of whether the data are heavy-tailed or light-tailed relative
to a normal distribution.
Leptokurtic
higher peak than the normal curve.
Platykurtic
the curve is more flat-topped than the normal curve.
Mesokurtic
neither too peaked nor too flat-topped.
Kurtosis is based on the size of a distribution’s tails.
Positive kurtosis (leptokurtic)- distributions with relatively long tails.
Negative kurtosis (platykurtic)- distributions with short tails.
125 /
129 Hazar Khogeer Introduction to statistics
3 Data Description: Kurtosis
126 /
129 Hazar Khogeer Introduction to statistics
3 Data Description: Kurtosis
Thus,
If the index
β2 − 3 > 0, the distribution is leptokurtic. (positive kurtosis)
β2 − 3 < 0, the distribution is platykurtic. (negative kurtosis)
β2 − 3 = 0, the distribution is mesokurtic.
127 /
129 Hazar Khogeer Introduction to statistics
3-2 Data Description: Kurtosis
i xi x̄ (xi − x̄ )2 (xi − x̄ )4
1 12 32 400 (−20)4
2 13 32 361 (−19)4
3 54 32 484 (22)4
4 56 32 576 (24)4
5 25 32 49 (−7)4
(xi − x̄ )2 = 1870 (xi − x̄ )4 = 858754
P P P
Total xi = 160
P5
(xi − x̄ )4 858754
β2 − 3 = Pni =1 − 3= − 3 = −2.75.
( i =1 (xi − x̄ )2 )2 (1870)2
The excess kurtosis is negative.
128 /
129 Hazar Khogeer Introduction to statistics
Homework: Chapter 3
Exercises 3-1: page 122-125 (1,8,16).
129 /
129 Hazar Khogeer Introduction to statistics