Module 1 - Complete Notes
Module 1 - Complete Notes
Introduction to Statistics
Measures of Central Tendency:
The following are the five measures of central tendency that are in common use:
i. Arithmetic mean or simple mean
ii. Median
iii. Mode
iv. Geometric mean
v. Harmonic mean
Arithmetic Mean:
𝒏
𝟏 𝟏
̅ = (𝒙𝟏 +𝒙𝟐 + 𝒙𝟑 + ⋯ + 𝒙𝒏 ) = ∑ 𝒙𝒊
𝒙
𝒏 𝒏
𝒊=𝟏
In case of grouped or continuous frequency distribution, x is taken as the as the mid value
of the corresponding class.
It may be noted that if the values of x or f are large, the calculation of mean by formula is
1 𝑛
∑ 𝑓 𝑥 is quite time-consuming and tedious. The arithmetic is reduced to a great extent
𝑁 𝑖=1 𝑖 𝑖
by taking the deviations of the given values from any arbitrary point ‘A’ as explained
below:
Let 𝑑𝑖 = 𝑥𝑖 − 𝐴. Then 𝑓𝑖 𝑑𝑖 = 𝑓𝑖 (𝑥𝑖 − 𝐴) = 𝑓𝑖 𝑥𝑖 − 𝐴𝑓𝑖
Summing both sides over 𝑖 from 1 to n, we get
𝑛 𝑛 𝑛 𝑛
∑ 𝑓𝑖 𝑑𝑖 = ∑ 𝑓𝑖 𝑥𝑖 − 𝐴 ∑ 𝑓𝑖 = ∑ 𝑓𝑖 𝑥𝑖 − 𝐴𝑁
𝑖=1 𝑖=1 𝑖=1 𝑖=1
𝑛 𝑛 𝑛 𝑛
1 1 1 1
∑ 𝑓𝑖 𝑑𝑖 = ∑ 𝑓𝑖 𝑥𝑖 − 𝐴 ∑ 𝑓𝑖 = ∑ 𝑓𝑖 𝑥𝑖 − 𝐴
𝑁 𝑁 𝑁 𝑁
𝑖=1 𝑖=1 𝑖=1 𝑖=1
𝑛
1
∑ 𝑓𝑖 𝑑𝑖 = 𝑥̅ − 𝐴
𝑁
𝑖=1
In case of grouped (or) continuous frequency distribution, the arithmetic is reduced to still
greater extent by taking point h is the common magnitude of class interval. In this case, we
have h𝑑𝑖 = 𝑥𝑖 − 𝐴 and proceeding exactly similarly above, we get
𝑛
h
𝑥̅ = 𝐴 + ∑ 𝑓𝑖 𝑑𝑖
𝑁
𝑖=1
Problem 1:
Solution:
∑ 𝑋 70+120+ 110+ 101+ 88+ 83+ 95+ 98+ 107+ 100 972
𝑋̅ = 𝑛 = 10
= 10 = 97.2
Problem 2:
The following frequency is distribution of the number of telephone calls received in 245
successive one-minute intervals at an exchange:
No. of calls 0 1 2 3 4 5 6 7
frequency 14 21 25 43 51 40 39 12
0 14 0
1 21 21
2 25 50
3 43 129
4 51 204
5 40 200
6 39 234
7 12 84
𝑁 = 245 ∑ 𝑓𝑋 = 922
𝑿 1 2 3 4 5 6 7
𝒇 5 9 12 17 14 10 6
Solution:
𝑋 𝑓 𝑓𝑋
1 5 5
2 9 18
3 12 36
4 17 68
5 14 70
6 10 60
7 7 42
Total 73 299
7
1 299
𝑥̅ = ∑ 𝑓𝑖 𝑥𝑖 = = 4.09
𝑁 73
𝑖=1
Problem 4:
For a certain frequency table which has only been partly reproduced here, the mean was
found to be 1.46.
No. of 0 1 2 3 4 5 Total
accidents
Frequency 46 ? ? 25 10 5 200
0 46 0
1 𝑓1 𝑓1
2 𝑓2 2𝑓2
3 25 75
4 10 40
5 5 25
86 + 𝑓1 + 𝑓2 140
= 200 + 𝑓1
+ 2𝑓2
200 = 86 + 𝑓1 + 𝑓2
𝑓1 + 𝑓2 = 200 − 86 = 114
𝑓1 + 𝑓2 = 114 (1)
1 𝑓1 + 2𝑓2 + 140
𝑋̅ = ∑ 𝑓𝑋 = = 1.46
𝑁 200
𝑓1 + 2𝑓2 + 140 = 1.46 × 200 = 292
𝑓1 + 2𝑓2 = 292 − 140 = 152
𝑓1 + 2𝑓2 = 152 (2)
Solving equations (1) and (2), we get
𝑓1 = 38, 𝑓2 = 76
Problem 5:
Calculate the arithmetic mean of the marks from the following table:
Solution:
0-10 12 5 60
10-20 18 15 270
20-30 27 25 675
30-40 20 35 700
40-50 17 45 765
50-60 6 55 330
Total 100 2800
1 1
𝐴𝑟𝑖𝑡h𝑚𝑒𝑡𝑖𝑐 𝑚𝑒𝑎𝑛 = 𝑥̅ = ∑ 𝑓𝑥 = × 2800 = 28
𝑁 100
Problem 6:
Solution:
Here we take 𝐴 = 28, h = 8
h ∑ 𝑓𝑑
𝑥̅ = 𝐴 +
𝑁
8 × (−25) 20
= 28 + = 28 − = 25.404
77 77
Median:
Median of a distribution is the value of the variable which divides it into two equal parts. It
is the value which exceeds and is exceeded by the same number of observations. That is, it
is the value such that the number of observations above it is equal to the number of
observations below it. The median is thus a positional average.
In case of ungrouped data, if the number of observations is odd then median is the middle
value after the values have been arranged in ascending or descending order of magnitude.
In case of even number of observations, there are two middle terms are median is obtained
by taking the arithmetic mean of the middle terms. For example, the median of the values
8,4,7,6,2, i.e., 2,4,6,7,8 is 6 and the median of 10,15,30,70,40,80, i.e., 10,15,30,40,70,80 is
1
(30 + 40) = 35.
2
𝑁
▪ Find 2 , where 𝑁 = ∑𝑛𝑖=1 𝑓𝑖
𝑁
▪ See the (less than) cumulative frequency (c.f.) just greater than 2 .
▪ The corresponding value of 𝑥 is median.
Example 1:
𝑥 1 2 3 4 5 6 7 8 9
𝑓 8 10 11 16 20 25 15 9 6
Solution:
𝑥 𝑓 𝑐. 𝑓.
1 8 8
2 10 18
3 11 29
4 16 45
5 20 65
6 25 90
7 15 105
8 9 114
9 6 120
𝑁 = 120
𝑁 = 120
𝑁
= 60
2
𝑁
The cumulative frequency (c.f.) 65 just greater than is 60 and the value of 𝑥
2
corresponding to 65 is 5. So, median is 5.
Example 2:
Eight coins were tossed together and the number of heads (𝑥) resulting was noted. The
operation was repeated 256 times and the frequency distribution of the number of heads is
given below:
𝑥 𝑓 𝑐. 𝑓.
0 1 1
1 9 10
2 26 36
3 59 95
4 72 167
5 52 219
6 29 248
7 7 255
8 1 256
𝑁 =256
𝑁 = 256
𝑁
= 128
2
𝑁
The cumulative frequency (c.f.) just greater than is 128 and the value of 𝑥 corresponding
2
to 167 is 5. So, median is 4.
Derivation of Median:
Median for Continuous frequency distribution:
In the case of continuous frequency distribution, the class corresponding to the c.f. just
𝑁
greater than 2 is called the median class and the value of median is obtained by the
following formula:
h 𝑁
𝑀𝑒𝑑𝑖𝑎𝑛 = 𝑙 + ( − 𝑐)
𝑓 2
Where 𝑙 is the lower limit of the median class
𝑓 is the frequency of the median class
h is the magnitude of the median class
𝑐 is the c.f. of the class preceding the median class
𝑁 = ∑𝑓
Example 1:
Solution:
Now we have to write the given distribution into continuous frequency distribution.
𝑁 = 43
𝑁
= 21.5
2
Cumulative frequency is just greater than 21.5 is 8 and the corresponding class is 4000-
5000.
Median class is 4000-5000.
h 𝑁
𝑀𝑒𝑑𝑖𝑎𝑛 = 𝑙 + ( − 𝑐)
𝑓 2
𝑙 = 4000, h = 1000, 𝑓 = 20, 𝑁 = 43, 𝑐 = 8
1000 43
𝑀𝑑 = 4000 + 20
(2 − 8) = 4675 rupees.
Example 2:
Find the frequency distribution of weight in grams of mangoes of a given variety is given
below. Then find the median.
Solution:
Continuous frequency distribution we have to convert the given inclusive class interval
series into exclusive class interval series
𝑁 = 200
𝑁
= 100
2
The c.f. just greater than 100 is 76
So, the corresponding class 439.5-449.5 is the median class.
Now we have to find the median for continuous data
h 𝑁
𝑀𝑒𝑑𝑖𝑎𝑛 = 𝑙 + ( − 𝑐)
𝑓 2
𝑙 = 439.5, h = 10, 𝑓 = 54, 𝑁 = 200, 𝑐 = 76
10 200
= 439.5 + 54 ( 2
− 76) = 443.94 gms.
Example 3:
Find the missing frequency from the following distribution of daily sales of shops, given
that the median sale of shops in Rs. 2,400.
Solution:
Let the frequency be 𝑥.
Given that the median sale of shops is 24 hundred.
0-10 5 5
10-20 25 30
20-30 𝑥 30 + 𝑥
30-40 18 48 + 𝑥
40-50 7 55 + 𝑥
𝑁 = 55 + 𝑥
h 𝑁
𝑀𝑒𝑑𝑖𝑎𝑛 = 𝑙 + ( − 𝑐)
𝑓 2
𝑙 = 20, h = 10, 𝑓 = 𝑥, 𝑁 = 55 + 𝑥, 𝑐 = 30
10 55 + 𝑥
24 = 20 + ( − 30)
𝑥 2
𝑥 = 25
Example 4:
In the frequency distribution of 100 families given below, the number of families
corresponding to expenditure groups 20-40 and 60-80 are missing from table. However, the
median is known to be 50. Find the missing frequencies.
Solution:
Let the missing frequencies for the classes 20-40 and 60-80 be 𝑓1 𝑎𝑛𝑑 𝑓2 respectively.
0-20 14 14
20-40 𝑓1 14 + 𝑓1
40-60 27 41 + 𝑓1
60-80 𝑓2 41 + 𝑓1 + 𝑓2
80-100 15 56 + 𝑓1 + 𝑓2
𝑁 = 56 + 𝑓1 + 𝑓2
Try these:
2. A number of particular articles has been classified according to their weights. After
drying two weeks the same articles have again been weighted and similarly classified. It
is known that the median weight in the first weighing is 20.83 gm, while in the second
weighing it was 17.35 gm. Some frequencies 𝑎 𝑎𝑛𝑑 𝑏 in the first weighing and 𝑥 𝑎𝑛𝑑 𝑦
1 1
in the second are missing. It is known that 𝑎 = 3 𝑥 𝑎𝑛𝑑 𝑏 = 2 𝑦. Find out the values of the
missing frequencies.
Geometric Mean:
The geometric mean, usually abbreviated as G.M. of a set of 𝑛 observations is the 𝑛𝑡h root of
their product. Thus, if 𝑋1 , 𝑋2 , 𝑋3 , … , 𝑋𝑛 are the 𝑛 observations then their G.M. is given by
1
𝐺. 𝑀 = 𝑛√𝑋1 × 𝑋2 × 𝑋3 × … × 𝑋𝑛 = (𝑋1 × 𝑋2 × 𝑋3 × … × 𝑋𝑛 )𝑛 (1)
If 𝑛 number of observations, then the 𝑛𝑡h root is very tedious. In such a case the calculations by
making use of the logarithms.
Now take logarithm both sides of eq. (1), we get
1
log(𝐺. 𝑀) =log(𝑋1 × 𝑋2 × 𝑋3 × … × 𝑋𝑛 )𝑛
1
= 𝑛 log(𝑋1 × 𝑋2 × 𝑋3 × … × 𝑋𝑛 )
1
= (log𝑋1 + 𝑙𝑜𝑔𝑋2 + 𝑙𝑜𝑔𝑋3 + … + 𝑙𝑜𝑔 𝑋𝑛 )
𝑛
1
log(𝐺. 𝑀) = 𝑛 ∑ 𝑙𝑜𝑔𝑋 (2)
i.e., the logarithm of the G.M of a set of observations is the arithmetic mean of their logarithms.
Now, taking Antilog on both sides of eq. (2)
1
𝐺. 𝑀 = 𝐴𝑛𝑡𝑖𝑙𝑜𝑔 ( ∑ 𝑙𝑜𝑔𝑋)
𝑛
In case of frequency distribution (𝑋𝑖 , 𝑓𝑖 ), 𝑖 = 1,2,3, … , 𝑛 , where the total no. of
observations is 𝑁 = ∑ 𝑓.
[(𝑋1 × 𝑋1 × 𝑋1 × … × 𝑓1 𝑡𝑖𝑚𝑒𝑠) × (𝑋2 × 𝑋2 × 𝑋2 × … × 𝑓2 𝑡𝑖𝑚𝑒𝑠) × … ×
1
(𝑋n × 𝑋n × 𝑋n × … × 𝑓𝑛 𝑡𝑖𝑚𝑒𝑠)]𝑛 (3)
1
𝐺. 𝑀 = (𝑋1 × 𝑋2 × 𝑋3 × … × 𝑋𝑛 )𝑁
Taking logarithm on both sides of eq. (3)
1
log(𝐺. 𝑀) = [𝑙𝑜𝑔(𝑋1 𝑓1 × 𝑋2 𝑓2 × … × 𝑋𝑛 𝑓𝑛 )]
𝑛
1
= [𝑙𝑜𝑔𝑋1 𝑓1 + 𝑙𝑜𝑔𝑋2 𝑓2 + ⋯ + 𝑙𝑜𝑔𝑋𝑛 𝑓𝑛 ]
𝑛
1
= [𝑓 𝑙𝑜𝑔𝑋1 + 𝑓2 𝑙𝑜𝑔𝑋2 + ⋯ + 𝑓𝑛 𝑙𝑜𝑔𝑋𝑛 ]
𝑛 1
1
log(𝐺. 𝑀) = ∑ 𝑓 𝑙𝑜𝑔𝑋
𝑛
1
𝐺. 𝑀 = 𝐴𝑛𝑡𝑖𝑙𝑜𝑔 [ ∑ 𝑓 𝑙𝑜𝑔𝑋]
𝑛
Draw the histogram, frequency polygon and frequency curve for the following frequency
distribution
Mid-value of the class 5.5 10.5 15.5 20.5 25.5 30.5 35.5 40.5
interval
Frequency 6 8 12 15 20 22 28 32
Example 1:
Find the geometric mean of 2,4,8,12,16 and 24.
Solution:
𝑋 𝑙𝑜𝑔𝑋
2 0.3010
4 0.6021
8 0.9031
12 1.0792
16 1.2041
24 1.3802
∑ 𝑙𝑜𝑔𝑋
= 5.4697
1
log(𝐺. 𝑀) = ∑ 𝑙𝑜𝑔𝑋
𝑛
1
= (5.4697) = 0.9116
6
= 𝐴𝑛𝑡𝑖𝑙𝑜𝑔(0.9116) = 8.158
𝐺. 𝑀 = 8.158
Example 2:
Find the geometric mean for the following distribution:
Solution:
1
𝐺. 𝑀 = 𝐴𝑛𝑡𝑖𝑙𝑜𝑔 [ ∑ 𝑓 𝑙𝑜𝑔𝑋]
𝑁
1
= 𝐴𝑛𝑡𝑖𝑙𝑜𝑔 [ (84.5243)]
60
= 𝐴𝑛𝑡𝑖𝑙𝑜𝑔(1.4087) = 25.64 marks
Example 3:
The geometric mean of 10 observations on a certain variable was calculated as 16.2. It was later
discovered that one of the observations was wrongly recorded as 12.9; in fact, it was 2.19. Apply
approximate correction and calculate the correct geometric mean.
Solution:
Geometric mean of observations is given by
1
𝐺 = (𝑋1 × 𝑋2 × 𝑋3 × … × 𝑋𝑛 )𝑛 (1)
𝐺 𝑛 = 𝑋1 × 𝑋2 × 𝑋3 × … × 𝑋𝑛
The product of the numbers is given by 𝑋1 × 𝑋2 × 𝑋3 × … × 𝑋𝑛 = 𝐺 𝑛 = (16.20)10 (2)
If the wrong observation 12.9 is replaced by the correct values 21.9, then the corrected value of
the product of 10 numbers is obtained on dividing the expression in (2) by wrong observation
and multiplying by the product of correct observation. Thus, the corrected product
(16.20)10 × 21.9
(𝑋1 × 𝑋2 × 𝑋3 × … × 𝑋𝑛 ) =
12.9
(Or)
Harmonic mean is the value of the variable or the mid-value of the class (in case of grouped or
continuous frequency distribution) and 𝑓 is the corresponding frequency of 𝑋.
In case of frequency distribution, we have
1 1 𝑓1 𝑓2 𝑓𝑛
= ( + +⋯+ )
𝐻 𝑁 𝑋1 𝑋2 𝑋𝑛
1 𝑓
= ∑( )
𝑁 𝑋
𝑁
𝐻= 1 , where, 𝑁 = ∑ 𝑓
∑( )
𝑋
Example 1:
A cyclist pedals from his house to his college at a speed of 10 kmph and back from college to his
house at 15 kmph. Find the average speed.
Solution:
Let the distance from house to college be 𝑥 𝑘𝑚.
𝑥
In going from house to college, the distance 𝑥 𝑘𝑚 is covered by 10 hours, while in coming from
𝑥
college to house, the distance covered in 15 hours.
𝑥 𝑥
Total distance of 2𝑥 𝑘𝑚 is covered in (10 + 15) hours.
𝑡𝑜𝑡𝑎𝑙 𝑑𝑖𝑠𝑡𝑎𝑛𝑐𝑒 𝑐𝑜𝑣𝑒𝑟𝑒𝑑
𝐴𝑣𝑒𝑟𝑎𝑔𝑒 𝑠𝑝𝑒𝑒𝑑 =
𝑡𝑜𝑡𝑎𝑙 𝑡𝑖𝑚𝑒 𝑡𝑎𝑘𝑒𝑛
2𝑥 2
= 𝑥 𝑥 = = 12 𝑘𝑚𝑝h
(10 + ) 1 1
15 (10 + )
15
Example 2:
A vehicle when climbing up a gradient, consumes petrol at the rate of 1 liter per 8 km. While
coming down it gives 12 km per liter. Find its average consumption for to and fro travel between
two places situated at the two ends of a 25 km long gradient. Verify your answer?
Solution:
Since the consumption of petrol is different for upward and downward journeys (at a constant
speed of 25 km), the appropriate average consumption for to and fro journey is given by
harmonic mean of 8km and 12 km.
2 48
𝐴𝑣𝑒𝑟𝑎𝑔𝑒 𝑐𝑜𝑛𝑠𝑢𝑚𝑝𝑡𝑖𝑜𝑛 = = = 9.6 𝑘𝑚/𝑙𝑖𝑡𝑒𝑟
1 1 5
(8 + 12)
Example 3:
In a certain office, a letter is typed by 𝐴 in 4 times. The same letter is typed by 𝐵, 𝐶, 𝑎𝑛𝑑 𝐷 are
5,6,10 minutes respectively. What is the average time taken in completing one letter? How many
letters do you expect to be typed in one day comprising of 8 working hours?
Solution:
The average time taken by each of 𝐴, 𝐵, 𝐶, 𝑎𝑛𝑑 𝐷 in competing one letter is the harmonic mean
of 4,5,6, and 10.
4 4 240
= = = 5.5814 𝑚𝑖𝑛/𝑙𝑒𝑡𝑡𝑒𝑟
1 1 1 1 15 + 12 + 1 + 6 43
+ + + ( )
4 5 6 10 60
43
Expected no. of letters typed by each of 𝐴, 𝐵, 𝐶, 𝑎𝑛𝑑 𝐷 is 240 letters/min.
Mode:
In case of continuous frequency distribution, mode is given by the formula
ℎ(𝑓1 − 𝑓0 )
𝑀𝑜𝑑𝑒 = 𝑙 +
2𝑓1 − 𝑓0 − 𝑓2
Problems:
ii. The median and mode of the following wage distribution are known to be Rs. 3,350 and Rs.
3,400 respectively. Find the values of 𝑓3 , 𝑓4 and𝑓5 .
Dispersion is the variation of the values which helps one to known as how the variates are
closely passed around or widely scattered away from the point of central tendency.
Range:
Range is the difference between the greatest (maximum) and the smallest (minimum)
observation of the distribution.
𝑅𝑎𝑛𝑔𝑒 = 𝑋𝑚𝑎𝑥 − 𝑋𝑚𝑖𝑛
Example 1:
Find the range of the following distribution
Solution:
Solution:
Quartile deviation:
It is a measure of dispersion based on the upper quartile 𝑄3 and the lower quartile 𝑄1.
𝑄3 − 𝑄1
𝑄𝑢𝑎𝑟𝑡𝑖𝑙𝑒 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 (𝑄. 𝐷) =
2
h 𝑁∗𝑖
where 𝑄𝑖 = 𝑙 + 𝑓 ( 4
− 𝑐) , 𝑓𝑜𝑟 𝑖 = 1,2,3
Or frequency distribution:
1
𝑀. 𝐷 = ∑ 𝑓𝑖 |𝑥𝑖 − 𝑥̅ |
𝑁
𝑖
(Or)
If 𝑓𝑖 , 𝑖 = 1,2,3, … , 𝑛 is the frequency distribution corresponding observations 𝑥𝑖 , then mean
deviation from average (usually mean, median and mode) is given by;
1
Mean deviation from average 𝐴 = 𝑁 ∑𝑛𝑖=1 𝑓𝑖 |𝑥𝑖 − 𝐴| , ∑𝑖 𝑓𝑖 = 𝑁, where |𝑥𝑖 − 𝐴| represents
modulus or absolute value of the deviation (𝑥𝑖 − 𝐴), where negative sign ignored.
Example:
Calculate the Mean deviation from the following data:
Solution:
1
Mean deviation from average 𝐴 = 𝑁 ∑𝑛𝑖=1 𝑓𝑖 |𝑥𝑖 − 𝐴|
1
We have to find the mean deviation from mean is 𝑀. 𝐷 = 𝑁 ∑ 𝑓 |𝑥 − 𝑥̅ |
1 659.2
𝑀. 𝐷 = ∑ 𝑓 |𝑥 − 𝑥̅ | = = 13.184
𝑁 50
Standard Deviation:
1 1
𝑉𝑎𝑟𝑖𝑎𝑛𝑐𝑒 = ∑(𝑥𝑖 − 𝑥̅ )2 = ∑ 𝑥𝑖 2 − 𝑥̅ 2
𝑛 𝑛
1 1
𝑆𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝐷𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 = 𝑆. 𝐷 = 𝜎 = √𝑛 ∑(𝑥𝑖 − 𝑥̅ )2 (or) √𝑛 ∑ 𝑥𝑖 2 − 𝑥̅ 2
𝑄3 −𝑄1
Coefficient of dispersion = 𝑄3 +𝑄1
𝜎
Coefficient of variance= 𝐶. 𝑉 = 100 × 𝑥̅
Example 1:
Lives of two modes of refrigerators recorded in a received survey are given in the following
table. Find the average life of each model. Which model shows more uniformity?
Life (No. of years) Model A Model B
0-2 5 2
2-4 16 7
4-6 13 12
8-6 7 19
8-10 5 9
10-12 4 1
Solution:
0-2 1 1 5 2 -1 -3 -5 -6 5 2
2-4 3 9 16 7 0 -2 0 -14 144 63
4-6 5 25 13 12 1 -1 13 -12 325 300
8-6 7 49 7 19 2 0 14 0 343 931
8-10 9 81 5 9 3 1 15 9 405 729
10-12 11 121 4 1 4 2 16 2 484 121
∑ 𝑥2 𝑁 𝑁 ∑ 𝑑𝐴 𝑓𝐴 ∑ 𝑑𝐵 𝑓𝐵 1706 2146
= 50 = 50
= 286 = 53 = − 21
h
𝑥̅𝐴 = 𝐴 + (∑ 𝑑𝐴 𝑓𝐴 )
𝑁
2
=3+ (53) = 5.12
50
h
𝑥̅𝐵 = 𝐵 + (∑ 𝑑𝐵 𝑓𝐵 )
𝑁
2
=7+ (−21) = 6.16
50
1
𝜎𝐴 = √ ∑ 𝑓𝐴 𝑥𝑖 2 − 𝑥̅ 2𝐴
𝑁
1
= √ × 286 − (5.12)2 = 2.8115
50
1
𝜎𝐵 = √ ∑ 𝑓𝐵 𝑥𝑖 2 − 𝑥̅ 2 𝐵
𝑁
1
= √ × 2146 − (6.16)2 = 2.23
50
𝜎𝐴
𝐶𝑉𝐴 = 100 × = 53.95
𝑥̅𝐴
𝜎𝐵
𝐶𝑉𝐵 = 100 × = 36.20
𝑥̅𝐵
𝐶𝑉𝐵 < 𝐶𝑉𝐴
Example 2:
Find the range, all three quartiles, quartile deviation, mean deviation, absolute mean deviation
and standard deviation for the following distribution
Range:
𝑅𝑎𝑛𝑔𝑒 = 𝑋𝑚𝑎𝑥 − 𝑋𝑚𝑖𝑛 = 12 − 0 = 12
Quartiles:
h 𝑁
𝑄1 = 𝑙 + ( − 𝑐)
𝑓 4
𝑙 = 2, h = 2, 𝑓 = 16, 𝑐 = 5, 𝑁 = 50
𝑄1 = 2.9375
h 𝑁
𝑄2 = 𝑙 + ( − 𝑐)
𝑓 2
𝑙 = 4, h = 2, 𝑓 = 13, 𝑐 = 21, 𝑁 = 50
𝑄2 = 4.6153
h 3𝑁
𝑄3 = 𝑙 + ( − 𝑐)
𝑓 4
𝑙 = 6, h = 2, 𝑓 = 7, 𝑐 = 34, 𝑁 = 50
𝑄3 = 7
7 − 2.9375
𝑄. 𝐷 = = 2.0312
2
Mean Deviation:
1
𝑀. 𝐷 = ∑ 𝑓. (𝑥 − 𝑥̅ )
𝑁
Standard Deviation:
1
𝑆. 𝐷 = 𝜎 = √ ∑ 𝑥𝑖 2 − 𝑥̅ 2
𝑛
The 𝑟 𝑡h moment of a variable 𝑥 about the mean 𝑥̅ , usually denoted by 𝜇𝑟 is given by:
1
𝜇𝑟 = 𝑁 ∑𝑖 𝑓𝑖 𝑧𝑖 𝑟 , 𝑤h𝑒𝑟𝑒 𝑧𝑖 = 𝑥𝑖 − 𝑥̅ (3)
1 1
In particular, 𝜇0 = 𝑁 ∑𝑖 𝑓𝑖 (𝑥𝑖 − 𝑥̅ )0 = 1 , 𝜇1 = 𝑁 ∑𝑖 𝑓𝑖 (𝑥𝑖 − 𝑥̅ )1 = 0 ,
1
𝜇2 = 𝑁 ∑𝑖 𝑓𝑖 (𝑥𝑖 − 𝑥̅ )2 = 𝜎 2 , …
1 1 1
𝜇𝑟 ′ = 𝑁 ∑𝑖 𝑓𝑖 (𝑥𝑖 − 𝐴)𝑟 = 𝑁 ∑𝑖 𝑓𝑖 (𝑥𝑖 − 𝑥̅ + 𝑥̅ − 𝐴)𝑟 = 𝑁 ∑𝑖 𝑓𝑖 (𝑧𝑖 + 𝜇1 ′ )𝑟 ,
where 𝑥𝑖 − 𝑥̅ = 𝑧𝑖 𝑎𝑛𝑑 𝑥̅ = 𝐴 + 𝜇1 ′
1 2 3 𝑟
Thus 𝜇𝑟 ′ = 𝑁 ∑𝑖 𝑓𝑖 (𝑧𝑖 𝑟 + 𝑟𝐶1 𝑧𝑖 𝑟−1 𝜇1 ′ + 𝑟𝐶2 𝑧𝑖 𝑟−2 𝜇1 ′ + 𝑟𝐶3 𝑧𝑖 𝑟−3 𝜇1 ′ + ⋯ + 𝜇1 ′ )
2 𝑟
= 𝜇𝑟 + 𝑟𝐶1 𝜇𝑟−1 𝜇1 ′ + 𝑟𝐶2 𝜇𝑟−2 𝜇1 ′ + ⋯ + 𝜇1 ′
Example:
Calculate the first four moments of the following distribution about the mean and hence
find 𝛽1 𝑎𝑛𝑑 𝛽2.
𝑥 0 1 2 3 4 5 6 7 8
𝑓 1 8 28 56 70 56 28 8 1
Solution:
0 1 -4 -4 16 -64 256
1 8 -3 -24 72 -216 648
2 28 -2 -56 112 -224 448
3 56 -1 -56 56 -56 56
4 70 0 0 0 0 0
5 56 1 56 56 56 56
6 28 2 56 112 224 448
7 8 3 24 72 216 648
8 1 4 4 16 64 256
Total 𝑵= 0 ∑ 𝒇𝒅 ∑ 𝒇𝒅𝟐 = 512 ∑ 𝒇𝒅𝟑 = 0 ∑ 𝒇𝒅𝟒 = 2816
𝟐𝟓𝟔
=𝟎
Skewness:
Literally, skewness means lack of symmetry i.e., skewness to have an idea about the shape
of the curve which can draw with help of the given data. A distribution is said to be
skewed, if
▪ Mean(𝑀), median 𝑀𝑑 and mode 𝑀0 fall at different points i.e., 𝑀𝑒𝑎𝑛 ≠ 𝑀𝑒𝑑𝑖𝑎𝑛 ≠ 𝑀𝑜𝑑𝑒
▪ Quartiles are not equidistant from median
▪ The curve drawn with the help of the given data is not symmetrical but stretched more to
one side than to the other.
Measures of Skewness:
Various measures of skewness (𝑆𝑘 ) are:
▪ 𝑆𝑘 = 𝑀 − 𝑀𝑑
▪ 𝑆𝑘 = 𝑀 − 𝑀0
▪ 𝑆𝑘 = (𝑄3 − 𝑀𝑑 ) − (𝑀𝑑 − 𝑄1 )
𝑀−𝑀0
𝑆𝑘 =
𝜎
, 𝜎 is the standard deviation of the distribution.
If mode is ill-defined, then using the empirical relation, 𝑀0 = 3𝑀𝑑 − 2𝑀, for a
moderately asymmetrical distribution, we get
3(𝑀 − 𝑀𝑑 )
𝑆𝑘 =
𝜎
𝑆𝑘 = 0, if 𝑀 = 𝑀0 = 𝑀𝑑 . Hence for a symmetrical distribution all are coincide.
2. Prof. Bowley’s Coefficient of Skewness:
Kurtosis:
▪ Prof. Karl Pearson’s calls as the ‘convexity of the frequency curve’ or Kurtosis.
▪ Kurtosis enables the flatness or peakedness of the frequency curve.
▪ It is measured by the coefficient 𝛽2 or its derivation is given by
𝜇
𝛽2 = 𝜇 42 , 𝛾2 = 𝛽2 − 3
2
C: Leptokurtic Curve (which is more peaked than the normal curve𝛽2 > 3 𝑖. 𝑒. 𝛾2 > 0 )
A: Normal Curve or Mesokurtic Curve (which is neither flat nor peaked 𝛽2 = 3 𝑖. 𝑒. 𝛾2 =
0)
B: Platykurtic Curve (which is flatter than the normal curve 𝛽2 < 3 𝑖. 𝑒. 𝛾2 < 0)
Example 1:
For a distribution, the mean is 10, variance is 16, 𝛾1 is +1 and 𝛽2 is 4. Obtain the first four
moments about the origin, i.e. zero. Comment upon the nature of distribution.
Solution:
3
Therefore , 𝜇3 = 𝜇3 ′ − 3𝜇2 ′ 𝜇1 ′ + 2𝜇1 ′
3
𝜇3 ′ = 𝜇3 + 3𝜇2 𝜇1 ′ + 𝜇1 ′ = 64 + 3(116)(1) − 2(1000) = −1544
𝜇
Now 𝛽2 = 𝜇 42 = 4, then 𝜇4 = 4(256) = 1024
2
2 4
𝜇4 = 𝜇4 ′ − 4𝜇3 ′ 𝜇1 ′ + 6𝜇2 ′ 𝜇1 ′ − 3𝜇1 ′
𝜇4 = 1024 + 4(1,544)(10) − 6(116)(100) + 3(10,000) = 23,184
Scores Frequency
50-60 1
60-70 0
70-80 0
80-90 1
90-100 1
100-110 2
110-120 1
120-130 0
130-140 4
140-150 4
150-160 2
160-170 5
170-180 10
180-190 11
190-200 4
200-210 1
210-220 1
220-230 2
Find also find the correct values of the moments after Sheppard’s corrections are applied.
Also obtain moment coefficient of skewness and kurtosis and comment on the nature of the
distribution.
Solution:
2 4
𝜇4 = (𝜇4 ′ − 4𝜇3 ′ 𝜇1 ′ + 6𝜇2 ′ 𝜇1 ′ − 3𝜇1 ′ ) × h4
= (796.68 − 4 × 91.68 × 3 + 6 × 20.76 × 9 − 3 × 91) × 10000 = 5687091.67
Sheppard’s corrections for moments:
h2 100
𝜇2 = 𝜇2 −
̅̅̅ = 1176 − = 1167.67
12 12
h2
𝜇3 = 𝜇3 −
̅̅̅ = −41160
12
h2 7 4
𝜇4 = 𝜇4 −
̅̅̅ 𝜇2 + h = 5745600 − 58800 + 291.67 = 5687091.67
2 240
Moment coefficient of skewness is given by
𝜇3 −41160
𝛾1 = √𝛽1 = 3/2
= = −1.02
𝜇2 1176√1176
𝜇 5745600
Moment coefficient of kurtosis=𝛽2 = 𝜇 42 = (1176)2
= 4.15
2
Since 𝛾1 = −1.02 is negative, the distribution is moderately negatively skewed, i.e., the
frequency curve of the given distribution has a longer tail towards the left.
Further, since 𝛽2 = 4.15 > 3 is positive the distribution is leptokurtic, i.e., the frequency
curve is more peaked than the normal curve.
Combined Mean:
Given the sample size 𝑛1 𝑛2
Sample mean 𝑥1
̅̅̅ 𝑥2
̅̅̅
𝑛1̅𝑥̅̅1̅+𝑛2 ̅𝑥̅̅2̅
Combined mean 𝑥̅ = 𝑛1 + 𝑛2
𝑑1 = 𝑥1 − 𝑥̅
Example:
A distribution consists of 25 measurements, it was found that the mean and standard
deviation are 36 cm and 12 cm. after these results were calculated, it was noticed that 2
measurements were wrongly recorded as 60 cm and 36 cm, instead of 40 cm and 30 cm.
Find the corrected values of the mean and standard deviation.
Solution:
Given that 𝑛 = 25, 𝑥̅ = 36, 𝜎 = 12
∑ 𝑥𝑖
We know that, 𝑥̅ = 𝑛
∑ 𝑥𝑖 = 𝑛. 𝑥̅ = 25 × 36 = 900
1
𝜎2 = ∑ 𝑥𝑖 2 − 𝑥̅ 2
𝑛