0% found this document useful (0 votes)
11 views41 pages

Module 1 - Complete Notes

The document provides an introduction to statistics, focusing on measures of central tendency including arithmetic mean, median, mode, geometric mean, and harmonic mean. It explains how to calculate the arithmetic mean using both ungrouped and grouped data, along with several examples and problems. Additionally, it covers the concept of median and its calculation in both ungrouped and discrete frequency distributions.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
11 views41 pages

Module 1 - Complete Notes

The document provides an introduction to statistics, focusing on measures of central tendency including arithmetic mean, median, mode, geometric mean, and harmonic mean. It explains how to calculate the arithmetic mean using both ungrouped and grouped data, along with several examples and problems. Additionally, it covers the concept of median and its calculation in both ungrouped and discrete frequency distributions.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Module-1

Introduction to Statistics
Measures of Central Tendency:

The following are the five measures of central tendency that are in common use:
i. Arithmetic mean or simple mean
ii. Median
iii. Mode
iv. Geometric mean
v. Harmonic mean

Arithmetic Mean:

Arithmetic mean of a set of observations is their sum divided by number of observations,


for example, the arithmetic mean 𝑥̅ of n observations 𝑥1, 𝑥2 , 𝑥3 , … , 𝑥𝑛 is given by:

𝒏
𝟏 𝟏
̅ = (𝒙𝟏 +𝒙𝟐 + 𝒙𝟑 + ⋯ + 𝒙𝒏 ) = ∑ 𝒙𝒊
𝒙
𝒏 𝒏
𝒊=𝟏

In case the frequency distribution 𝑓𝑖 , 𝑖 = 1,2,3, … , 𝑛, where 𝑓𝑖 is the frequency


of the variable 𝑥𝑖 ,
𝑓1 𝑥1 +𝑓2 𝑥2 +𝑓3 𝑥3 +⋯+𝑓𝑛 𝑥𝑛 1
𝑥̅ = 𝑓1 +𝑓2 +𝑓2 +⋯+𝑓𝑛
= 𝑁 ∑𝑛𝑖=1 𝑓𝑖 𝑥𝑖 , where 𝑁 = ∑𝑛𝑖=1 𝑓𝑖

In case of grouped or continuous frequency distribution, x is taken as the as the mid value
of the corresponding class.
It may be noted that if the values of x or f are large, the calculation of mean by formula is
1 𝑛
∑ 𝑓 𝑥 is quite time-consuming and tedious. The arithmetic is reduced to a great extent
𝑁 𝑖=1 𝑖 𝑖
by taking the deviations of the given values from any arbitrary point ‘A’ as explained
below:
Let 𝑑𝑖 = 𝑥𝑖 − 𝐴. Then 𝑓𝑖 𝑑𝑖 = 𝑓𝑖 (𝑥𝑖 − 𝐴) = 𝑓𝑖 𝑥𝑖 − 𝐴𝑓𝑖
Summing both sides over 𝑖 from 1 to n, we get
𝑛 𝑛 𝑛 𝑛

∑ 𝑓𝑖 𝑑𝑖 = ∑ 𝑓𝑖 𝑥𝑖 − 𝐴 ∑ 𝑓𝑖 = ∑ 𝑓𝑖 𝑥𝑖 − 𝐴𝑁
𝑖=1 𝑖=1 𝑖=1 𝑖=1
𝑛 𝑛 𝑛 𝑛
1 1 1 1
∑ 𝑓𝑖 𝑑𝑖 = ∑ 𝑓𝑖 𝑥𝑖 − 𝐴 ∑ 𝑓𝑖 = ∑ 𝑓𝑖 𝑥𝑖 − 𝐴
𝑁 𝑁 𝑁 𝑁
𝑖=1 𝑖=1 𝑖=1 𝑖=1
𝑛
1
∑ 𝑓𝑖 𝑑𝑖 = 𝑥̅ − 𝐴
𝑁
𝑖=1

Where 𝑥̅ is the arithmetic mean of the distribution.


𝑛
1
𝑥̅ = 𝐴 + ∑ 𝑓𝑖 𝑑𝑖
𝑁
𝑖=1

In case of grouped (or) continuous frequency distribution, the arithmetic is reduced to still
greater extent by taking point h is the common magnitude of class interval. In this case, we
have h𝑑𝑖 = 𝑥𝑖 − 𝐴 and proceeding exactly similarly above, we get
𝑛
h
𝑥̅ = 𝐴 + ∑ 𝑓𝑖 𝑑𝑖
𝑁
𝑖=1

Problem 1:

The intelligence quotient (IQ’s) of 10 boys in a class are given below:


70, 120, 110, 101, 88, 83, 95, 98, 107, 100. Then find the mean I.Q.

Solution:

Mean I.Q of 10 boys in a class are given below:

∑ 𝑋 70+120+ 110+ 101+ 88+ 83+ 95+ 98+ 107+ 100 972
𝑋̅ = 𝑛 = 10
= 10 = 97.2

Problem 2:
The following frequency is distribution of the number of telephone calls received in 245
successive one-minute intervals at an exchange:

No. of calls 0 1 2 3 4 5 6 7
frequency 14 21 25 43 51 40 39 12

Obtain the mean number of calls per minute.


Solution:

No. of Calls (𝑋) Frequency (𝑓) 𝑓𝑋

0 14 0
1 21 21
2 25 50
3 43 129
4 51 204
5 40 200
6 39 234
7 12 84
𝑁 = 245 ∑ 𝑓𝑋 = 922

Mean number of calls per minute at the at the exchange is given by


∑ 𝑓𝑋 922
𝑋̅ = = = 3.763
𝑁 245
Problem 3:

Find the arithmetic mean of the following frequency distribution:

𝑿 1 2 3 4 5 6 7
𝒇 5 9 12 17 14 10 6

Solution:

𝑋 𝑓 𝑓𝑋

1 5 5
2 9 18
3 12 36
4 17 68
5 14 70
6 10 60
7 7 42
Total 73 299

7
1 299
𝑥̅ = ∑ 𝑓𝑖 𝑥𝑖 = = 4.09
𝑁 73
𝑖=1

Problem 4:

For a certain frequency table which has only been partly reproduced here, the mean was
found to be 1.46.

No. of 0 1 2 3 4 5 Total
accidents
Frequency 46 ? ? 25 10 5 200

Calculate the missing frequencies.


Solution:
Let 𝑋 denote the number of accidents and let the missing frequencies corresponding to
𝑋 = 1 and 𝑋 = 2 be 𝑓1 𝑎𝑛𝑑 𝑓2 respectively.

No. of accidents (𝑋) Frequency (𝑓) 𝑓𝑋

0 46 0
1 𝑓1 𝑓1

2 𝑓2 2𝑓2

3 25 75
4 10 40
5 5 25
86 + 𝑓1 + 𝑓2 140
= 200 + 𝑓1
+ 2𝑓2

200 = 86 + 𝑓1 + 𝑓2

𝑓1 + 𝑓2 = 200 − 86 = 114
𝑓1 + 𝑓2 = 114 (1)
1 𝑓1 + 2𝑓2 + 140
𝑋̅ = ∑ 𝑓𝑋 = = 1.46
𝑁 200
𝑓1 + 2𝑓2 + 140 = 1.46 × 200 = 292
𝑓1 + 2𝑓2 = 292 − 140 = 152
𝑓1 + 2𝑓2 = 152 (2)
Solving equations (1) and (2), we get
𝑓1 = 38, 𝑓2 = 76

Problem 5:
Calculate the arithmetic mean of the marks from the following table:

Marks 0-10 10-20 20-30 30-40 40-50 50-60


No. of Students 12 18 27 20 17 6

Solution:

Marks No. of Students Mid-point (𝑓𝑥)


(𝑓) (𝑋)

0-10 12 5 60
10-20 18 15 270
20-30 27 25 675
30-40 20 35 700
40-50 17 45 765
50-60 6 55 330
Total 100 2800

1 1
𝐴𝑟𝑖𝑡h𝑚𝑒𝑡𝑖𝑐 𝑚𝑒𝑎𝑛 = 𝑥̅ = ∑ 𝑓𝑥 = × 2800 = 28
𝑁 100

Problem 6:

Calculate the mean for the following frequency distribution:

Class interval 0-8 8-16 16-24 24-32 32-40 40-48


Frequency 8 7 16 24 15 7

Solution:
Here we take 𝐴 = 28, h = 8

Class interval Mid-value Frequency 𝑥−𝐴 (𝑓𝑑)


𝑑=
(𝑥) (𝑓) h
0-8 4 8 -3 -24
8-16 12 7 -2 -14
16-24 20 16 -1 -16
24-32 28 24 0 0
32-40 36 15 1 15
40-48 44 7 2 14
Total 77 -25

h ∑ 𝑓𝑑
𝑥̅ = 𝐴 +
𝑁
8 × (−25) 20
= 28 + = 28 − = 25.404
77 77

Problem 7: Try this?

Calculate the mean for the following frequency distribution:

Marks 0-10 10-20 20-30 30-40 40-50 50-60 60-70


Number of students 6 5 8 15 7 6 3

Median:

Median of a distribution is the value of the variable which divides it into two equal parts. It
is the value which exceeds and is exceeded by the same number of observations. That is, it
is the value such that the number of observations above it is equal to the number of
observations below it. The median is thus a positional average.
In case of ungrouped data, if the number of observations is odd then median is the middle
value after the values have been arranged in ascending or descending order of magnitude.
In case of even number of observations, there are two middle terms are median is obtained
by taking the arithmetic mean of the middle terms. For example, the median of the values
8,4,7,6,2, i.e., 2,4,6,7,8 is 6 and the median of 10,15,30,70,40,80, i.e., 10,15,30,40,70,80 is
1
(30 + 40) = 35.
2

In case of discrete frequency distribution median is obtained by considering the cumulative


frequencies. The steps for calculating median are given below:

𝑁
▪ Find 2 , where 𝑁 = ∑𝑛𝑖=1 𝑓𝑖
𝑁
▪ See the (less than) cumulative frequency (c.f.) just greater than 2 .
▪ The corresponding value of 𝑥 is median.

Example 1:

Obtain the median for the following frequency distribution:

𝑥 1 2 3 4 5 6 7 8 9
𝑓 8 10 11 16 20 25 15 9 6

Solution:

𝑥 𝑓 𝑐. 𝑓.

1 8 8
2 10 18
3 11 29
4 16 45
5 20 65
6 25 90
7 15 105
8 9 114
9 6 120
𝑁 = 120

𝑁 = 120

𝑁
= 60
2
𝑁
The cumulative frequency (c.f.) 65 just greater than is 60 and the value of 𝑥
2
corresponding to 65 is 5. So, median is 5.

Example 2:

Eight coins were tossed together and the number of heads (𝑥) resulting was noted. The
operation was repeated 256 times and the frequency distribution of the number of heads is
given below:

No. of heads (𝑥) 0 1 2 3 4 5 6 7 8


Frequency (𝑓) 1 9 26 59 72 52 29 7 1

Find the median.


Solution:

𝑥 𝑓 𝑐. 𝑓.

0 1 1
1 9 10
2 26 36
3 59 95
4 72 167
5 52 219
6 29 248
7 7 255
8 1 256
𝑁 =256

𝑁 = 256
𝑁
= 128
2
𝑁
The cumulative frequency (c.f.) just greater than is 128 and the value of 𝑥 corresponding
2
to 167 is 5. So, median is 4.

Derivation of Median:
Median for Continuous frequency distribution:
In the case of continuous frequency distribution, the class corresponding to the c.f. just
𝑁
greater than 2 is called the median class and the value of median is obtained by the
following formula:
h 𝑁
𝑀𝑒𝑑𝑖𝑎𝑛 = 𝑙 + ( − 𝑐)
𝑓 2
Where 𝑙 is the lower limit of the median class
𝑓 is the frequency of the median class
h is the magnitude of the median class
𝑐 is the c.f. of the class preceding the median class
𝑁 = ∑𝑓

Example 1:

Find the median wage of the following distribution:

Wages No. of workers


(in rupees)
2000-3000 3
3000-4000 5
4000-5000 20
5000-6000 10
6000-7000 5

Solution:
Now we have to write the given distribution into continuous frequency distribution.

Wages No. of workers (𝑓) 𝑐. 𝑓.


(in rupees)
2000-3000 3 3
3000-4000 5 8
4000-5000 20 28
5000-6000 10 38
6000-7000 5 43
𝑁 = 43

𝑁 = 43
𝑁
= 21.5
2
Cumulative frequency is just greater than 21.5 is 8 and the corresponding class is 4000-
5000.
Median class is 4000-5000.
h 𝑁
𝑀𝑒𝑑𝑖𝑎𝑛 = 𝑙 + ( − 𝑐)
𝑓 2
𝑙 = 4000, h = 1000, 𝑓 = 20, 𝑁 = 43, 𝑐 = 8
1000 43
𝑀𝑑 = 4000 + 20
(2 − 8) = 4675 rupees.

The median wage is 4675 rupees.

Example 2:

Find the frequency distribution of weight in grams of mangoes of a given variety is given
below. Then find the median.

Weight in grams 410-419 420-429 430-439 440-449 450-459 460-469 470-


479
No. of mangoes 14 20 42 54 45 18 7

Solution:
Continuous frequency distribution we have to convert the given inclusive class interval
series into exclusive class interval series

Weight in grams No. of mangoes (𝑓) 𝑐. 𝑓. (Less than)


409.5-419.5 14 14
419.5-429.5 20 34
429.5-439.5 42 76
439.5-449.5 54 130
449.5-459.5 45 175
459.5-469.5 18 193
469.5-479.5 7 200
𝑁 = 200

𝑁 = 200
𝑁
= 100
2
The c.f. just greater than 100 is 76
So, the corresponding class 439.5-449.5 is the median class.
Now we have to find the median for continuous data
h 𝑁
𝑀𝑒𝑑𝑖𝑎𝑛 = 𝑙 + ( − 𝑐)
𝑓 2
𝑙 = 439.5, h = 10, 𝑓 = 54, 𝑁 = 200, 𝑐 = 76
10 200
= 439.5 + 54 ( 2
− 76) = 443.94 gms.

Example 3:

Find the missing frequency from the following distribution of daily sales of shops, given
that the median sale of shops in Rs. 2,400.

Sales (in hundred rupees) 0-10 10-20 20-30 30-40 40-50


No. of shops 5 25 ? 18 7

Solution:
Let the frequency be 𝑥.
Given that the median sale of shops is 24 hundred.

Sales (in hundred rupees) No. of shops(𝑓) 𝑐. 𝑓.

0-10 5 5
10-20 25 30
20-30 𝑥 30 + 𝑥

30-40 18 48 + 𝑥

40-50 7 55 + 𝑥

𝑁 = 55 + 𝑥

h 𝑁
𝑀𝑒𝑑𝑖𝑎𝑛 = 𝑙 + ( − 𝑐)
𝑓 2
𝑙 = 20, h = 10, 𝑓 = 𝑥, 𝑁 = 55 + 𝑥, 𝑐 = 30
10 55 + 𝑥
24 = 20 + ( − 30)
𝑥 2
𝑥 = 25

Example 4:

In the frequency distribution of 100 families given below, the number of families
corresponding to expenditure groups 20-40 and 60-80 are missing from table. However, the
median is known to be 50. Find the missing frequencies.

Expenditure 0-20 20-40 40-60 60-80 80-100


No. of families 14 ? 27 ? 15

Solution:
Let the missing frequencies for the classes 20-40 and 60-80 be 𝑓1 𝑎𝑛𝑑 𝑓2 respectively.

Expenditure (in rupees) No. of families (𝑓) 𝑐. 𝑓.

0-20 14 14
20-40 𝑓1 14 + 𝑓1

40-60 27 41 + 𝑓1

60-80 𝑓2 41 + 𝑓1 + 𝑓2

80-100 15 56 + 𝑓1 + 𝑓2

𝑁 = 56 + 𝑓1 + 𝑓2

The no. of families is 100, therefore


𝑁 = 100 = 56 + 𝑓1 + 𝑓2
𝑓1 + 𝑓2 = 44
Given that median is 50, which lies class 40-60.
therefore, 40-60 is the median class.
𝑙 = 40, h = 20, 𝑓 = 27, 𝑁 = 56 + 𝑓1 + 𝑓2 , 𝑐 = 14 + 𝑓1
h 𝑁
𝑀𝑒𝑑𝑖𝑎𝑛 = 𝑙 + ( − 𝑐)
𝑓 2
20 100
50 = 40 + ( − (14 + 𝑓1 ))
27 2
20
10 = (36 − 𝑓1 )
27
𝑓1 = 22.5 ≈ 23
𝑓2 = 21

Try these:

1. Calculate the mean and median from the following distribution


Class 10-20 20-30 30-40 40-50 50-60 60-70 70-80 80-90
Frequency 4 12 40 41 27 13 9 4

2. A number of particular articles has been classified according to their weights. After
drying two weeks the same articles have again been weighted and similarly classified. It
is known that the median weight in the first weighing is 20.83 gm, while in the second
weighing it was 17.35 gm. Some frequencies 𝑎 𝑎𝑛𝑑 𝑏 in the first weighing and 𝑥 𝑎𝑛𝑑 𝑦
1 1
in the second are missing. It is known that 𝑎 = 3 𝑥 𝑎𝑛𝑑 𝑏 = 2 𝑦. Find out the values of the
missing frequencies.

Geometric Mean:
The geometric mean, usually abbreviated as G.M. of a set of 𝑛 observations is the 𝑛𝑡h root of
their product. Thus, if 𝑋1 , 𝑋2 , 𝑋3 , … , 𝑋𝑛 are the 𝑛 observations then their G.M. is given by
1
𝐺. 𝑀 = 𝑛√𝑋1 × 𝑋2 × 𝑋3 × … × 𝑋𝑛 = (𝑋1 × 𝑋2 × 𝑋3 × … × 𝑋𝑛 )𝑛 (1)

If 𝑛 = 2 i.e., if we take two observations, then 𝐺. 𝑀 = √𝑋1 × 𝑋2

If 𝑛 number of observations, then the 𝑛𝑡h root is very tedious. In such a case the calculations by
making use of the logarithms.
Now take logarithm both sides of eq. (1), we get
1
log(𝐺. 𝑀) =log(𝑋1 × 𝑋2 × 𝑋3 × … × 𝑋𝑛 )𝑛
1
= 𝑛 log⁡(𝑋1 × 𝑋2 × 𝑋3 × … × 𝑋𝑛 )
1
= (log𝑋1 + 𝑙𝑜𝑔𝑋2 + 𝑙𝑜𝑔𝑋3 + … + 𝑙𝑜𝑔 𝑋𝑛 )
𝑛

1
log(𝐺. 𝑀) = 𝑛 ∑ 𝑙𝑜𝑔𝑋 (2)

i.e., the logarithm of the G.M of a set of observations is the arithmetic mean of their logarithms.
Now, taking Antilog on both sides of eq. (2)
1
𝐺. 𝑀 = 𝐴𝑛𝑡𝑖𝑙𝑜𝑔 ( ∑ 𝑙𝑜𝑔𝑋)
𝑛
In case of frequency distribution (𝑋𝑖 , 𝑓𝑖 ), 𝑖 = 1,2,3, … , 𝑛 , where the total no. of
observations is 𝑁 = ∑ 𝑓.
[(𝑋1 × 𝑋1 × 𝑋1 × … × 𝑓1 𝑡𝑖𝑚𝑒𝑠) × (𝑋2 × 𝑋2 × 𝑋2 × … × 𝑓2 𝑡𝑖𝑚𝑒𝑠) × … ×
1
(𝑋n × 𝑋n × 𝑋n × … × 𝑓𝑛 𝑡𝑖𝑚𝑒𝑠)]𝑛 (3)
1
𝐺. 𝑀 = (𝑋1 × 𝑋2 × 𝑋3 × … × 𝑋𝑛 )𝑁
Taking logarithm on both sides of eq. (3)
1
log(𝐺. 𝑀) = [𝑙𝑜𝑔(𝑋1 𝑓1 × 𝑋2 𝑓2 × … × 𝑋𝑛 𝑓𝑛 )]
𝑛
1
= [𝑙𝑜𝑔𝑋1 𝑓1 + 𝑙𝑜𝑔𝑋2 𝑓2 + ⋯ + 𝑙𝑜𝑔𝑋𝑛 𝑓𝑛 ]
𝑛
1
= [𝑓 𝑙𝑜𝑔𝑋1 + 𝑓2 𝑙𝑜𝑔𝑋2 + ⋯ + 𝑓𝑛 𝑙𝑜𝑔𝑋𝑛 ]
𝑛 1
1
log(𝐺. 𝑀) = ∑ 𝑓 𝑙𝑜𝑔𝑋
𝑛
1
𝐺. 𝑀 = 𝐴𝑛𝑡𝑖𝑙𝑜𝑔 [ ∑ 𝑓 𝑙𝑜𝑔𝑋]
𝑛

Draw the histogram, frequency polygon and frequency curve for the following frequency
distribution
Mid-value of the class 5.5 10.5 15.5 20.5 25.5 30.5 35.5 40.5
interval
Frequency 6 8 12 15 20 22 28 32

Example 1:
Find the geometric mean of 2,4,8,12,16 and 24.

Solution:

𝑋 𝑙𝑜𝑔𝑋

2 0.3010
4 0.6021
8 0.9031
12 1.0792
16 1.2041
24 1.3802

∑ 𝑙𝑜𝑔𝑋
= 5.4697

1
log(𝐺. 𝑀) = ∑ 𝑙𝑜𝑔𝑋
𝑛
1
= (5.4697) = 0.9116
6
= 𝐴𝑛𝑡𝑖𝑙𝑜𝑔(0.9116) = 8.158
𝐺. 𝑀 = 8.158

Example 2:
Find the geometric mean for the following distribution:

Marks 0-10 10-20 20-30 30-40 40-50


No. of students 5 7 15 25 8

Solution:

Marks Mid-point (𝑋) No. of students(𝑓) 𝑙𝑜𝑔𝑋 𝑓 𝑙𝑜𝑔𝑋

0-10 5 5 0.6990 3.4950


10-20 15 7 1.1761 8.2327
20-30 25 15 1.3979 20.9685
30-40 35 25 1.5441 38.6025
40-50 45 8 1.6532 13.2256
𝑁 = 60 ∑ 𝑓 𝑙𝑜𝑔𝑋
= 84.5243

1
𝐺. 𝑀 = 𝐴𝑛𝑡𝑖𝑙𝑜𝑔 [ ∑ 𝑓 𝑙𝑜𝑔𝑋]
𝑁
1
= 𝐴𝑛𝑡𝑖𝑙𝑜𝑔 [ (84.5243)]
60
= 𝐴𝑛𝑡𝑖𝑙𝑜𝑔(1.4087) = 25.64 marks

Example 3:
The geometric mean of 10 observations on a certain variable was calculated as 16.2. It was later
discovered that one of the observations was wrongly recorded as 12.9; in fact, it was 2.19. Apply
approximate correction and calculate the correct geometric mean.

Solution:
Geometric mean of observations is given by
1
𝐺 = (𝑋1 × 𝑋2 × 𝑋3 × … × 𝑋𝑛 )𝑛 (1)
𝐺 𝑛 = 𝑋1 × 𝑋2 × 𝑋3 × … × 𝑋𝑛
The product of the numbers is given by 𝑋1 × 𝑋2 × 𝑋3 × … × 𝑋𝑛 = 𝐺 𝑛 = (16.20)10 (2)
If the wrong observation 12.9 is replaced by the correct values 21.9, then the corrected value of
the product of 10 numbers is obtained on dividing the expression in (2) by wrong observation
and multiplying by the product of correct observation. Thus, the corrected product
(16.20)10 × 21.9
(𝑋1 × 𝑋2 × 𝑋3 × … × 𝑋𝑛 ) =
12.9

The corrected value of G.M, is say 𝐺 ′


1

(16.20)10 × 21.9 10
𝐺 =[ ]
12.9
1
(16.20)10 × 21.9 10
𝑙𝑜𝑔𝐺 ′ = 𝑙𝑜𝑔 [ ]
12.9
1
= [[𝑙𝑜𝑔(16.20)10 + 𝑙𝑜𝑔(21.9) − 𝑙𝑜𝑔(12.9)]]
10
1
= [10 𝑙𝑜𝑔(16.20) + 𝑙𝑜𝑔(21.9) − 𝑙𝑜𝑔(12.9)]
10
1
𝑙𝑜𝑔𝐺 ′ = [10 (1.2095) + 1.3404 − 1.1106]
10
1
𝑙𝑜𝑔𝐺 ′ = [10 (1.2095) + 1.3404 − 1.1106] = 1.2325
10
𝑙𝑜𝑔𝐺 ′ = 1.2325
𝐺 ′ = 𝐴𝑛𝑡𝑖𝑙𝑜𝑔(1.2325) = 17.08
Harmonic Mean:
If 𝑋1 , 𝑋2 , 𝑋3 , … , 𝑋𝑛 is a given 𝑛 set of observations, then their harmonic mean, abbreviated as
H.M.
1
𝐻=
1 1 1 1 1
𝑛 [𝑋1 + 𝑋2 + 𝑋3 + ⋯ + 𝑋𝑛 ]
1
𝐻=
1 1
∑( )
𝑛 𝑋

(Or)

Harmonic mean is the value of the variable or the mid-value of the class (in case of grouped or
continuous frequency distribution) and 𝑓 is the corresponding frequency of 𝑋.
In case of frequency distribution, we have
1 1 𝑓1 𝑓2 𝑓𝑛
= ( + +⋯+ )
𝐻 𝑁 𝑋1 𝑋2 𝑋𝑛
1 𝑓
= ∑( )
𝑁 𝑋
𝑁
𝐻= 1 , where, 𝑁 = ∑ 𝑓
∑( )
𝑋

𝑋= mid-value of the variable or mid-value of the class


𝑓= frequency of 𝑋

Example 1:

A cyclist pedals from his house to his college at a speed of 10 kmph and back from college to his
house at 15 kmph. Find the average speed.

Solution:
Let the distance from house to college be 𝑥 𝑘𝑚.
𝑥
In going from house to college, the distance 𝑥 𝑘𝑚 is covered by 10 hours, while in coming from
𝑥
college to house, the distance covered in 15 hours.
𝑥 𝑥
Total distance of 2𝑥 𝑘𝑚 is covered in (10 + 15) hours.
𝑡𝑜𝑡𝑎𝑙 𝑑𝑖𝑠𝑡𝑎𝑛𝑐𝑒 𝑐𝑜𝑣𝑒𝑟𝑒𝑑
𝐴𝑣𝑒𝑟𝑎𝑔𝑒 𝑠𝑝𝑒𝑒𝑑 =
𝑡𝑜𝑡𝑎𝑙 𝑡𝑖𝑚𝑒 𝑡𝑎𝑘𝑒𝑛
2𝑥 2
= 𝑥 𝑥 = = 12 𝑘𝑚𝑝h
(10 + ) 1 1
15 (10 + )
15

Example 2:
A vehicle when climbing up a gradient, consumes petrol at the rate of 1 liter per 8 km. While
coming down it gives 12 km per liter. Find its average consumption for to and fro travel between
two places situated at the two ends of a 25 km long gradient. Verify your answer?

Solution:
Since the consumption of petrol is different for upward and downward journeys (at a constant
speed of 25 km), the appropriate average consumption for to and fro journey is given by
harmonic mean of 8km and 12 km.
2 48
𝐴𝑣𝑒𝑟𝑎𝑔𝑒 𝑐𝑜𝑛𝑠𝑢𝑚𝑝𝑡𝑖𝑜𝑛 = = = 9.6 𝑘𝑚/𝑙𝑖𝑡𝑒𝑟
1 1 5
(8 + 12)

Example 3:
In a certain office, a letter is typed by 𝐴 in 4 times. The same letter is typed by 𝐵, 𝐶, 𝑎𝑛𝑑 𝐷 are
5,6,10 minutes respectively. What is the average time taken in completing one letter? How many
letters do you expect to be typed in one day comprising of 8 working hours?

Solution:
The average time taken by each of 𝐴, 𝐵, 𝐶, 𝑎𝑛𝑑 𝐷 in competing one letter is the harmonic mean
of 4,5,6, and 10.
4 4 240
= = = 5.5814 𝑚𝑖𝑛/𝑙𝑒𝑡𝑡𝑒𝑟
1 1 1 1 15 + 12 + 1 + 6 43
+ + + ( )
4 5 6 10 60
43
Expected no. of letters typed by each of 𝐴, 𝐵, 𝐶, 𝑎𝑛𝑑 𝐷 is 240 letters/min.

In a day comprising 8 hours = 8 × 60 minutes.


43
Total letters typed all of them 𝐴, 𝐵, 𝐶, 𝑎𝑛𝑑 𝐷 = 240 × 4 × 8 × 60 = 344.

Mode:
In case of continuous frequency distribution, mode is given by the formula

ℎ(𝑓1 − 𝑓0 )
𝑀𝑜𝑑𝑒 = 𝑙 +
2𝑓1 − 𝑓0 − 𝑓2

Where, ℎ is the magnitude

𝑙 is the lower limit

𝑓1 is the frequency of the modal class

𝑓0 is the frequency of the preceding the modal class

𝑓2 is the frequency of the succeeding the modal class

Problems:

i. Find the mode of the following data:

Class 0-10 10-20 20-30 30-40 40-50 50-60 60-70 70-80


Interval
frequency 5 8 7 12 28 20 10 10

ii. The median and mode of the following wage distribution are known to be Rs. 3,350 and Rs.
3,400 respectively. Find the values of 𝑓3 , 𝑓4 ⁡and⁡𝑓5 .

Wages 0-1000 1000- 2000- 3000- 4000- 5000- 6000- Total


(Rs.) 2000 3000 4000 5000 6000 7000
No. of 4 16 𝑓3 𝑓4 𝑓5 6 4 200
employees
Measures of variability (dispersion):

Dispersion is the variation of the values which helps one to known as how the variates are
closely passed around or widely scattered away from the point of central tendency.

The various measures of dispersion are:


Range
Quartile Deviation
Mean deviation or Absolute Mean deviation
Standard Deviation

Range:
Range is the difference between the greatest (maximum) and the smallest (minimum)
observation of the distribution.
𝑅𝑎𝑛𝑔𝑒 = 𝑋𝑚𝑎𝑥 − 𝑋𝑚𝑖𝑛
Example 1:
Find the range of the following distribution

Class interval 0-2 2-4 4-6 6-8 8-10 10-12


Frequency 5 16 13 7 5 4

Solution:

𝑅𝑎𝑛𝑔𝑒 = 𝑋𝑚𝑎𝑥 − 𝑋𝑚𝑖𝑛


= 12 − 0 = 12
Example 2:
The following table given the age of distribution of a group 50 individuals

Age (in years) 16-20 21-25 26-30 31-36


No. of persons 10 15 17 8

Solution:

Age (in years) No. of persons


15.5-20.5 10
20.5-25.5 15
25.5-30.5 17
30.5-35.5 8

𝑅𝑎𝑛𝑔𝑒 = 𝑋𝑚𝑎𝑥 − 𝑋𝑚𝑖𝑛


= 35.5 − 15.5 = 20

Quartile deviation:
It is a measure of dispersion based on the upper quartile 𝑄3 and the lower quartile 𝑄1.

𝑄3 − 𝑄1
𝑄𝑢𝑎𝑟𝑡𝑖𝑙𝑒 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 (𝑄. 𝐷) =
2

h 𝑁∗𝑖
where 𝑄𝑖 = 𝑙 + 𝑓 ( 4
− 𝑐) , 𝑓𝑜𝑟 𝑖 = 1,2,3

𝑄1 = 𝑓𝑖𝑟𝑠𝑡 𝑞𝑢𝑎𝑟𝑡𝑖𝑙𝑒 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛


𝑄2 = 𝑠𝑒𝑐𝑜𝑛𝑑 𝑞𝑢𝑎𝑟𝑡𝑖𝑙𝑒 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛
𝑄3 = 𝑡h𝑖𝑟𝑑 𝑞𝑢𝑎𝑟𝑡𝑖𝑙𝑒 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛

Mean Deviation (or) absolute Mean Deviation:

For ungrouped data or raw data:


1
𝑀. 𝐷 = ∑|𝑥𝑖 − 𝑥̅ |
𝑛
𝑖

Or frequency distribution:
1
𝑀. 𝐷 = ∑ 𝑓𝑖 |𝑥𝑖 − 𝑥̅ |
𝑁
𝑖
(Or)
If 𝑓𝑖 , 𝑖 = 1,2,3, … , 𝑛 is the frequency distribution corresponding observations 𝑥𝑖 , then mean
deviation from average (usually mean, median and mode) is given by;
1
Mean deviation from average 𝐴 = 𝑁 ∑𝑛𝑖=1 𝑓𝑖 |𝑥𝑖 − 𝐴| , ∑𝑖 𝑓𝑖 = 𝑁, where |𝑥𝑖 − 𝐴| represents
modulus or absolute value of the deviation (𝑥𝑖 − 𝐴), where negative sign ignored.

Example:
Calculate the Mean deviation from the following data:

Marks 0-10 10-20 20-30 30-40 40-50 50-60 60-70


No. of students 6 5 8 35 7 6 3

Solution:
1
Mean deviation from average 𝐴 = 𝑁 ∑𝑛𝑖=1 𝑓𝑖 |𝑥𝑖 − 𝐴|
1
We have to find the mean deviation from mean is 𝑀. 𝐷 = 𝑁 ∑ 𝑓 |𝑥 − 𝑥̅ |

So, first we have to find the mean 𝑥̅ .


h
𝑥̅ = 𝐴 + ∑ 𝑓𝑑
𝑁

Marks Mid-value (𝑥) No. of students (𝑓) 𝑥 − 35 𝑓𝑑


𝑑=
h
0-10 5 6 -3 -18
10-20 15 5 -2 -10
20-30 25 8 -1 -8
30-40 35 15 0 0
40-50 45 7 1 7
50-60 55 6 2 12
60-70 65 3 3 9
𝑁 = 50 ∑ 𝑓𝑑 = −8
h
𝑥̅ = 𝐴 + ∑ 𝑓𝑑
𝑁
10
𝑥̅ = 35 + (−8) = 33.4 𝑚𝑎𝑟𝑘𝑠
50

Marks Mid-value No. of students (𝑓) 𝑥 − 35 𝑓𝑑 |𝑥 − 𝑥̅ | 𝑓|𝑥 − 𝑥̅ |


𝑑=
(𝑥) h = |𝑥 − 33.4|

0-10 5 6 -3 -18 28.4 170.4


10-20 15 5 -2 -10 18.4 92
20-30 25 8 -1 -8 8.4 67.2
30-40 35 15 0 0 1.6 24
40-50 45 7 1 7 11.6 81.2
50-60 55 6 2 12 21.6 129.6
60-70 65 3 3 9 31.6 94.8
𝑁 = 50 ∑ 𝑓 |𝑥 − 𝑥̅ |
= 659.2

1 659.2
𝑀. 𝐷 = ∑ 𝑓 |𝑥 − 𝑥̅ | = = 13.184
𝑁 50

Standard Deviation:
1 1
𝑉𝑎𝑟𝑖𝑎𝑛𝑐𝑒 = ∑(𝑥𝑖 − 𝑥̅ )2 = ∑ 𝑥𝑖 2 − 𝑥̅ 2
𝑛 𝑛

1 1
𝑆𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝐷𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 = 𝑆. 𝐷 = 𝜎 = √𝑛 ∑(𝑥𝑖 − 𝑥̅ )2 (or) √𝑛 ∑ 𝑥𝑖 2 − 𝑥̅ 2

𝑄3 −𝑄1
Coefficient of dispersion = 𝑄3 +𝑄1
𝜎
Coefficient of variance= 𝐶. 𝑉 = 100 × 𝑥̅

Example 1:

Lives of two modes of refrigerators recorded in a received survey are given in the following
table. Find the average life of each model. Which model shows more uniformity?
Life (No. of years) Model A Model B
0-2 5 2
2-4 16 7
4-6 13 12
8-6 7 19
8-10 5 9
10-12 4 1

Solution:

Life Mid- 𝑥2 Model Model 𝑑𝐴 𝑑𝐵 𝑑𝐴 𝑓𝐴 𝑑𝐵 𝑓𝐵 𝑓𝐴 𝑥 2 𝑓𝐵 𝑥 2


(No. of value A B 𝑥 − 𝐴 𝑥−𝐵
= =
years) (𝑥) (𝑓𝐴 ) (𝑓𝐵 ) h h

0-2 1 1 5 2 -1 -3 -5 -6 5 2
2-4 3 9 16 7 0 -2 0 -14 144 63
4-6 5 25 13 12 1 -1 13 -12 325 300
8-6 7 49 7 19 2 0 14 0 343 931
8-10 9 81 5 9 3 1 15 9 405 729
10-12 11 121 4 1 4 2 16 2 484 121

∑ 𝑥2 𝑁 𝑁 ∑ 𝑑𝐴 𝑓𝐴 ∑ 𝑑𝐵 𝑓𝐵 1706 2146
= 50 = 50
= 286 = 53 = − 21

h
𝑥̅𝐴 = 𝐴 + (∑ 𝑑𝐴 𝑓𝐴 )
𝑁
2
=3+ (53) = 5.12
50

h
𝑥̅𝐵 = 𝐵 + (∑ 𝑑𝐵 𝑓𝐵 )
𝑁
2
=7+ (−21) = 6.16
50
1
𝜎𝐴 = √ ∑ 𝑓𝐴 𝑥𝑖 2 − 𝑥̅ 2𝐴
𝑁

1
= √ × 286 − (5.12)2 = 2.8115
50

1
𝜎𝐵 = √ ∑ 𝑓𝐵 𝑥𝑖 2 − 𝑥̅ 2 𝐵
𝑁

1
= √ × 2146 − (6.16)2 = 2.23
50
𝜎𝐴
𝐶𝑉𝐴 = 100 × = 53.95
𝑥̅𝐴
𝜎𝐵
𝐶𝑉𝐵 = 100 × = 36.20
𝑥̅𝐵
𝐶𝑉𝐵 < 𝐶𝑉𝐴

Therefore, Model B has more uniformity.

Example 2:

Find the range, all three quartiles, quartile deviation, mean deviation, absolute mean deviation
and standard deviation for the following distribution

Class Interval frequency


0-2 5
2-4 16
4-6 13
6-8 7
8-10 5
10-12 4
Solution:

Range:
𝑅𝑎𝑛𝑔𝑒 = 𝑋𝑚𝑎𝑥 − 𝑋𝑚𝑖𝑛 = 12 − 0 = 12

Quartiles:

h 𝑁
𝑄1 = 𝑙 + ( − 𝑐)
𝑓 4
𝑙 = 2, h = 2, 𝑓 = 16, 𝑐 = 5, 𝑁 = 50
𝑄1 = 2.9375

h 𝑁
𝑄2 = 𝑙 + ( − 𝑐)
𝑓 2
𝑙 = 4, h = 2, 𝑓 = 13, 𝑐 = 21, 𝑁 = 50
𝑄2 = 4.6153

h 3𝑁
𝑄3 = 𝑙 + ( − 𝑐)
𝑓 4
𝑙 = 6, h = 2, 𝑓 = 7, 𝑐 = 34, 𝑁 = 50
𝑄3 = 7

7 − 2.9375
𝑄. 𝐷 = = 2.0312
2

Mean Deviation:

1
𝑀. 𝐷 = ∑ 𝑓. (𝑥 − 𝑥̅ )
𝑁

Absolute Mean Deviation:


1
𝑀. 𝐷 = ∑ 𝑓𝑖 . |𝑥𝑖 − 𝑥̅ |
𝑁
𝑖

Standard Deviation:

1
𝑆. 𝐷 = 𝜎 = √ ∑ 𝑥𝑖 2 − 𝑥̅ 2
𝑛

Moments, Skewness and Kurtosis:


Moments:
The 𝑟 𝑡h moment of a variable 𝑥 about any point 𝑥 = 𝐴, usually denoted by 𝜇𝑟 ′ is given by:
1
𝜇𝑟 ′ = 𝑁 ∑𝑖 𝑓𝑖 (𝑥𝑖 − 𝐴)𝑟 , ∑𝑖 𝑓𝑖 = 𝑁 (1)
1
𝜇𝑟 ′ = 𝑁 ∑𝑖 𝑓𝑖 𝑑𝑖 𝑟 , 𝑤h𝑒𝑟𝑒 𝑑𝑖 = 𝑥𝑖 − 𝐴 (2)

The 𝑟 𝑡h moment of a variable 𝑥 about the mean 𝑥̅ , usually denoted by 𝜇𝑟 is given by:
1
𝜇𝑟 = 𝑁 ∑𝑖 𝑓𝑖 𝑧𝑖 𝑟 , 𝑤h𝑒𝑟𝑒 𝑧𝑖 = 𝑥𝑖 − 𝑥̅ (3)
1 1
In particular, 𝜇0 = 𝑁 ∑𝑖 𝑓𝑖 (𝑥𝑖 − 𝑥̅ )0 = 1 , 𝜇1 = 𝑁 ∑𝑖 𝑓𝑖 (𝑥𝑖 − 𝑥̅ )1 = 0 ,
1
𝜇2 = 𝑁 ∑𝑖 𝑓𝑖 (𝑥𝑖 − 𝑥̅ )2 = 𝜎 2 , …

We know that if 𝑑𝑖 = 𝑥𝑖 − 𝐴, then


1
𝑥̅ = 𝐴 + 𝑁 ∑𝑖 𝑓𝑖 𝑑𝑖 = 𝐴 + 𝜇1 ′ (4)

Relation between Moments about Mean in terms of Moments


about any point and vice versa:
1 1 1
𝜇𝑟 = 𝑁 ∑𝑖 𝑓𝑖 (𝑥𝑖 − 𝑥̅ )𝑟 = 𝑁 ∑𝑖 𝑓𝑖 (𝑥𝑖 − A + A − 𝑥̅ )𝑟 = ∑ 𝑓 (𝑑
𝑁 𝑖 𝑖 𝑖
+ A − 𝑥̅ )𝑟 , where 𝑑𝑖 = 𝑥𝑖 − 𝐴.

Using eq. (4), we get


1
𝜇𝑟 = ∑ 𝑓𝑖 (𝑑𝑖 − 𝜇1 ′ )𝑟
𝑁
𝑖
1 2 3 𝑟
𝜇𝑟 = 𝑁 ∑𝑖 𝑓𝑖 (𝑑𝑖 𝑟 − 𝑟𝐶1 𝑑𝑖 𝑟−1 𝜇1 ′ + 𝑟𝐶2 𝑑𝑖 𝑟−2 𝜇1 ′ − 𝑟𝐶3 𝑑𝑖 𝑟−3 𝜇1 ′ + ⋯ + (−1)𝑟 𝜇1 ′ ) (5)
2 3 𝑟
𝜇𝑟 = 𝜇𝑟 ′ − 𝑟𝐶1 𝜇𝑟−1 ′ 𝜇1 ′ + 𝑟𝐶2 𝜇𝑟−2 ′ 𝜇1 ′ − 𝑟𝐶3 𝜇𝑟−3 ′ 𝜇1 ′ + ⋯ + (−1)𝑟 𝜇1 ′ (Using eq. (3))

In particular on putting 𝑟 = 2,3 𝑎𝑛𝑑 4 in eq. (5) and simplifying, we get


2
𝜇2 = 𝜇2 ′ − 𝜇1 ′
3
𝜇3 = 𝜇3 ′ − 3𝜇2 ′ 𝜇1 ′ + 2𝜇1 ′
2 4
𝜇4 = 𝜇4 ′ − 4𝜇3 ′ 𝜇1 ′ + 6𝜇2 ′ 𝜇1 ′ − 3𝜇1 ′

1 1 1
𝜇𝑟 ′ = 𝑁 ∑𝑖 𝑓𝑖 (𝑥𝑖 − 𝐴)𝑟 = 𝑁 ∑𝑖 𝑓𝑖 (𝑥𝑖 − 𝑥̅ + 𝑥̅ − 𝐴)𝑟 = 𝑁 ∑𝑖 𝑓𝑖 (𝑧𝑖 + 𝜇1 ′ )𝑟 ,

where 𝑥𝑖 − 𝑥̅ = 𝑧𝑖 𝑎𝑛𝑑 𝑥̅ = 𝐴 + 𝜇1 ′
1 2 3 𝑟
Thus 𝜇𝑟 ′ = 𝑁 ∑𝑖 𝑓𝑖 (𝑧𝑖 𝑟 + 𝑟𝐶1 𝑧𝑖 𝑟−1 𝜇1 ′ + 𝑟𝐶2 𝑧𝑖 𝑟−2 𝜇1 ′ + 𝑟𝐶3 𝑧𝑖 𝑟−3 𝜇1 ′ + ⋯ + 𝜇1 ′ )
2 𝑟
= 𝜇𝑟 + 𝑟𝐶1 𝜇𝑟−1 𝜇1 ′ + 𝑟𝐶2 𝜇𝑟−2 𝜇1 ′ + ⋯ + 𝜇1 ′

In particular on putting 𝑟 = 2,3, 4 𝑎𝑛𝑑 𝑛𝑜𝑡𝑖𝑛𝑔 𝑡h𝑎𝑡 𝜇1 = 𝜇1 ′ , we get


𝜇1 ′ = 𝜇1
2
𝜇2 ′ = 𝜇2 + 𝜇1 ′
3
𝜇3 ′ = 𝜇3 + 3𝜇2 𝜇1 ′ − 2𝜇1 ′
2 4
𝜇4 ′ = 𝜇4 + 4𝜇3 𝜇1 ′ − 6𝜇2 𝜇1 ′ + 3𝜇1 ′
These formulae unable to find the moments about any point, once the mean and the
moments about mean are known.
Pearson’s 𝜷 𝒂𝒏𝒅 𝜸 Coefficients:
Karl Pearson defined the following four coefficients, based upon the first four moments
about mean:
𝜇3 2 𝜇4
𝛽1 = 3 , 𝛾1 = +√𝛽1 𝑎𝑛𝑑 𝛽2 = 2 , 𝛾2 = 𝛽2 − 3
𝜇2 𝜇2
It may be pointed out that these coefficients are pure numbers independent of units of
measurement.

Example:
Calculate the first four moments of the following distribution about the mean and hence
find 𝛽1 𝑎𝑛𝑑 𝛽2.
𝑥 0 1 2 3 4 5 6 7 8
𝑓 1 8 28 56 70 56 28 8 1

Solution:

𝑥 𝑓 𝑑 𝑓𝑑 𝑓𝑑2 𝑓𝑑3 𝑓𝑑4


=𝑥
−4

0 1 -4 -4 16 -64 256
1 8 -3 -24 72 -216 648
2 28 -2 -56 112 -224 448
3 56 -1 -56 56 -56 56
4 70 0 0 0 0 0
5 56 1 56 56 56 56
6 28 2 56 112 224 448
7 8 3 24 72 216 648
8 1 4 4 16 64 256
Total 𝑵= 0 ∑ 𝒇𝒅 ∑ 𝒇𝒅𝟐 = 512 ∑ 𝒇𝒅𝟑 = 0 ∑ 𝒇𝒅𝟒 = 2816
𝟐𝟓𝟔
=𝟎

Moments about the point 𝑥 = 4 are


1 1 512
𝜇1 ′ = ∑ 𝑓𝑑 = 0, 𝜇2 ′ = ∑ 𝑓𝑑2 = = 2,
𝑁 𝑁 256
1 1 2816
𝜇3 ′ = ∑ 𝑓𝑑3 = 0, 𝜇4 ′ = ∑ 𝑓𝑑4 = = 11
𝑁 𝑁 256
Moments about mean are
2 3
𝜇1 = 0, 𝜇2 = 𝜇2 ′ − 𝜇1 ′ = 2, 𝜇3 = 𝜇3 ′ − 3𝜇2 ′ 𝜇1 ′ + 2𝜇1 ′ = 0
2 4
𝜇4 = 𝜇4 ′ − 4𝜇3 ′ 𝜇1 ′ + 6𝜇2 ′ 𝜇1 ′ − 3𝜇1 ′ = 11
𝜇3 2 𝜇4 11
𝛽1 = = 0, 𝛽2 = = = 2.75
𝜇2 3 𝜇2 2 4
Skewness and Kurtosis:

Skewness:
Literally, skewness means lack of symmetry i.e., skewness to have an idea about the shape
of the curve which can draw with help of the given data. A distribution is said to be
skewed, if
▪ Mean(𝑀), median 𝑀𝑑 and mode 𝑀0 fall at different points i.e., 𝑀𝑒𝑎𝑛 ≠ 𝑀𝑒𝑑𝑖𝑎𝑛 ≠ 𝑀𝑜𝑑𝑒
▪ Quartiles are not equidistant from median
▪ The curve drawn with the help of the given data is not symmetrical but stretched more to
one side than to the other.
Measures of Skewness:
Various measures of skewness (𝑆𝑘 ) are:

▪ 𝑆𝑘 = 𝑀 − 𝑀𝑑
▪ 𝑆𝑘 = 𝑀 − 𝑀0
▪ 𝑆𝑘 = (𝑄3 − 𝑀𝑑 ) − (𝑀𝑑 − 𝑄1 )

These are the absolute measures of skewness.


1. Prof. Karl Pearson’s Coefficient of Skewness:

𝑀−𝑀0
𝑆𝑘 =
𝜎
, 𝜎 is the standard deviation of the distribution.
If mode is ill-defined, then using the empirical relation, 𝑀0 = 3𝑀𝑑 − 2𝑀, for a
moderately asymmetrical distribution, we get

3(𝑀 − 𝑀𝑑 )
𝑆𝑘 =
𝜎
𝑆𝑘 = 0, if 𝑀 = 𝑀0 = 𝑀𝑑 . Hence for a symmetrical distribution all are coincide.
2. Prof. Bowley’s Coefficient of Skewness:

(𝑄3 − 𝑀𝑑 ) − (𝑀𝑑 − 𝑄1 ) 𝑄3 + 𝑄1 − 2𝑀𝑑


𝑆𝑘 = =
(𝑄3 − 𝑀𝑑 ) + (𝑀𝑑 − 𝑄1 ) 𝑄3 − 𝑄1
𝑆𝑘 = 0, if 𝑄3 − 𝑀𝑑 = 𝑀𝑑 − 𝑄1 . Hence for a symmetrical distribution. Median is equidistant
from the upper and lower quartiles.

3. Based upon the moments, coefficient of skewness:


√𝛽1 (𝛽2 + 3)
𝑆𝑘 =
2(5𝛽2 − 6𝛽1 − 9)
𝑆𝑘 = 0, if either 𝛽1 = 0 or 𝛽2 = −3

Kurtosis:

▪ Prof. Karl Pearson’s calls as the ‘convexity of the frequency curve’ or Kurtosis.
▪ Kurtosis enables the flatness or peakedness of the frequency curve.
▪ It is measured by the coefficient 𝛽2 or its derivation is given by
𝜇
𝛽2 = 𝜇 42 , 𝛾2 = 𝛽2 − 3
2
C: Leptokurtic Curve (which is more peaked than the normal curve⁡𝛽2 > 3 𝑖. 𝑒. 𝛾2 > 0 )
A: Normal Curve or Mesokurtic Curve (which is neither flat nor peaked 𝛽2 = 3 𝑖. 𝑒. 𝛾2 =
0)
B: Platykurtic Curve (which is flatter than the normal curve 𝛽2 < 3 𝑖. 𝑒. 𝛾2 < 0)

Example 1:
For a distribution, the mean is 10, variance is 16, 𝛾1 is +1 and 𝛽2 is 4. Obtain the first four
moments about the origin, i.e. zero. Comment upon the nature of distribution.

Solution:

Given that Mean=10, 𝜇2 = 16, 𝑡h𝑒𝑛 𝑆. 𝐷 = 4, 𝛾1 = + 1 and 𝛽2 = 4


First four moments about the origin is 𝜇1 ′ , 𝜇2 ′ , 𝜇3 ′ , 𝜇4 ′ .
𝜇1 ′ = 𝑓𝑖𝑟𝑠𝑡 𝑚𝑜𝑚𝑒𝑛𝑡 𝑎𝑏𝑜𝑢𝑡 𝑡h𝑒 𝑜𝑟𝑖𝑔𝑖𝑛 = 𝑀𝑒𝑎𝑛 = 10
2 ′ 2
𝜇2 = 𝜇2 ′ − 𝜇1 ′ 𝑡h𝑒𝑛 𝜇2 = 𝜇2 + 𝜇1 ′ = 16 + 100 = 116
𝜇3 𝜇
We have 𝛾1 = + 1 = 𝛾1 = +√𝛽1 then 𝜇 3/2
= 𝜎33 = 1 𝑜𝑟 𝜇3 = 𝜎 3 = 64
2

3
Therefore , 𝜇3 = 𝜇3 ′ − 3𝜇2 ′ 𝜇1 ′ + 2𝜇1 ′
3
𝜇3 ′ = 𝜇3 + 3𝜇2 𝜇1 ′ + 𝜇1 ′ = 64 + 3(116)(1) − 2(1000) = −1544
𝜇
Now 𝛽2 = 𝜇 42 = 4, then 𝜇4 = 4(256) = 1024
2

2 4
𝜇4 = 𝜇4 ′ − 4𝜇3 ′ 𝜇1 ′ + 6𝜇2 ′ 𝜇1 ′ − 3𝜇1 ′
𝜇4 = 1024 + 4(1,544)(10) − 6(116)(100) + 3(10,000) = 23,184

Comments on nature of the distribution:


Since 𝛾1 = +1, the distribution is moderately positively skewed, i.e., if we draw the curve of
given distribution, it will have longer tail towards the right. Further since 𝛽2 = 4 > 3, the
distribution is leptokurtic, i.e., it will be slightly more peaked than the normal curve.
Example 2:
For the frequency distribution of scores in mathematics of 50 candidates selected at random
from among those appearing at a certain examination, compute the first four moments
about the mean of the distribution.

Scores Frequency
50-60 1
60-70 0
70-80 0
80-90 1
90-100 1
100-110 2
110-120 1
120-130 0
130-140 4
140-150 4
150-160 2
160-170 5
170-180 10
180-190 11
190-200 4
200-210 1
210-220 1
220-230 2
Find also find the correct values of the moments after Sheppard’s corrections are applied.
Also obtain moment coefficient of skewness and kurtosis and comment on the nature of the
distribution.
Solution:

Mid-value (𝑥) frequency 𝑥−4 𝑓𝑑 𝑓𝑑2 𝑓𝑑3 𝑓𝑑4


𝑑=
(𝑓) 10
55 1 -8 -8 64 -512 4096
65 0 -7 0 0 0 0
75 0 -6 0 0 0
85 1 -5 -5 25 -125 625
95 1 -4 -4 16 -64 256
105 2 -3 -6 18 -54 162
115 1 -2 -2 4 -8 16
125 0 -1 0 0 0 0
135 4 0 0 0 0 0
145 4 1 4 4 4 4
155 2 2 4 8 16 32
165 5 3 15 45 135 405
175 10 4 40 160 640 2560
185 11 5 55 275 1375 6875
195 4 6 24 144 864 5184
205 1 7 7 49 343 2401
215 1 8 8 64 512 4096
225 2 9 18 162 1458 13122
Total 50 150 1038 4584 39834

The raw moments of variable 𝑑 (about origin) are computed as


1 150 1 1038
𝜇1 ′ = ∑ 𝑓𝑑 = = 3, 𝜇2 ′ = ∑ 𝑓𝑑2 = = 20.76,
𝑁 50 𝑁 50
1 4584 1 39834
𝜇3 ′ = ∑ 𝑓𝑑3 = = 91.68, 𝜇4 ′ = ∑ 𝑓𝑑4 = = 796.68
𝑁 50 𝑁 50
The central moments of variable 𝑋 are then computed as shown below
2
𝜇2 = (𝜇2 ′ − 𝜇1 ′ ) × h2 = (20.76 − 9) × 100 = 1176
3
𝜇3 = (𝜇3 ′ − 3𝜇2 ′ 𝜇1 ′ + 3𝜇1 ′ ) × h3 = (91.68 − 3 × 20.76 × 3 + 2(27)) × 1000 = −41160

2 4
𝜇4 = (𝜇4 ′ − 4𝜇3 ′ 𝜇1 ′ + 6𝜇2 ′ 𝜇1 ′ − 3𝜇1 ′ ) × h4
= (796.68 − 4 × 91.68 × 3 + 6 × 20.76 × 9 − 3 × 91) × 10000 = 5687091.67
Sheppard’s corrections for moments:
h2 100
𝜇2 = 𝜇2 −
̅̅̅ = 1176 − = 1167.67
12 12
h2
𝜇3 = 𝜇3 −
̅̅̅ = −41160
12
h2 7 4
𝜇4 = 𝜇4 −
̅̅̅ 𝜇2 + h = 5745600 − 58800 + 291.67 = 5687091.67
2 240
Moment coefficient of skewness is given by
𝜇3 −41160
𝛾1 = √𝛽1 = 3/2
= = −1.02
𝜇2 1176√1176
𝜇 5745600
Moment coefficient of kurtosis=𝛽2 = 𝜇 42 = (1176)2
= 4.15
2

Comments on nature of the distribution:

Since 𝛾1 = −1.02 is negative, the distribution is moderately negatively skewed, i.e., the
frequency curve of the given distribution has a longer tail towards the left.
Further, since 𝛽2 = 4.15 > 3 is positive the distribution is leptokurtic, i.e., the frequency
curve is more peaked than the normal curve.

Combined Mean:
Given the sample size 𝑛1 𝑛2
Sample mean 𝑥1
̅̅̅ 𝑥2
̅̅̅
𝑛1̅𝑥̅̅1̅+𝑛2 ̅𝑥̅̅2̅
Combined mean 𝑥̅ = 𝑛1 + 𝑛2

Also, sample variance 𝜎1 2 𝜎2 2


𝑛1 (𝜎1 2 +𝑑1 2 )+𝑛2 (𝜎2 2 +𝑑2 2 )
Combined variance 𝜎2 = 𝑛1 + 𝑛2

𝑑1 = 𝑥1 − 𝑥̅
Example:
A distribution consists of 25 measurements, it was found that the mean and standard
deviation are 36 cm and 12 cm. after these results were calculated, it was noticed that 2
measurements were wrongly recorded as 60 cm and 36 cm, instead of 40 cm and 30 cm.
Find the corrected values of the mean and standard deviation.
Solution:
Given that 𝑛 = 25, 𝑥̅ = 36, 𝜎 = 12
∑ 𝑥𝑖
We know that, 𝑥̅ = 𝑛

∑ 𝑥𝑖 = 𝑛. 𝑥̅ = 25 × 36 = 900

1
𝜎2 = ∑ 𝑥𝑖 2 − 𝑥̅ 2
𝑛

∑ 𝑥𝑖 2 = 𝑛(𝜎 2 + 𝑥̅ 2 ) = 25(144 + 296) = 36000

Corrected ∑ 𝑥𝑖 = 900 − 60 − 36 + 40 + 30 = 874


Corrected ∑ 𝑥𝑖 2 = 36000 − 602 − 362 + 402 + 302 = 33604
𝑐𝑜𝑟𝑟𝑒𝑐𝑡𝑒𝑑 ∑ 𝑥𝑖 874
Corrected mean= = = 34.96
𝑛 25
33604
Corrected variance(𝜎 2 ) = − (34.96)2 = 121.96
25

Corrected standard deviation (𝜎) = √121.96 = 11.04

You might also like