Unit 3
Unit 3
UNIT 3
MEASURES OF CENTRAL
TENDENCY
Structure
3.1 Introduction 3.9 Relationship Between Mean,
Median and Mode
Objectives
3.10 Merits and Demerits of Mode
3.2 Measures of Central
Tendency 3.11 Geometric Mean
Properties for a Suitable For Ungrouped Data
Measure of Central Tendency
For Grouped Data
Different Measures of Central
3.12 Harmonic Mean
Tendency
For Ungrouped Data
3.3 Arithmetic Mean or Mean
For Grouped Data
For Ungrouped Data
3.13 Merits and Demerits of
For Grouped Data
Harmonic Mean
3.4 Properties of Arithmetic
3.14 Relation Between AM, GM
Mean
and HM
3.5 Merits and Demerits of
3.15 Partition Values
Arithmetic Mean
Quartiles
3.6 Median
Deciles
Median for Ungrouped Data
Percentiles
Median for Ungrouped Data
(When Frequencies Are Given) 3.16 Summary
Median for Grouped Data 3.17 Terminal Questions
3.7 Merits and Demerits of 3.18 Answers
Median
3.8 Mode
For Ungrouped Data
For Grouped Data
3.1 INTRODUCTION
As we know, after the classification and tabulation of data, one often needs
more details for the many uses that may be made on the information available.
67
Block 1 Descriptive of Statistics
We, therefore, need further analysis of the tabulated data to draw inferences.
In this unit, we will discuss measures of central tendency. For the purpose of
analysis, a very important and powerful tool is the only average value that
stands for the entire mass of data. In statistics, an average is a one-figure
representation of a distribution. It provides a value that the distribution is
concentrated around it. This is why the measure of central tendency is another
name for the average. For instance, eye surgery typically costs Rs. 15,000 in a
city on an average. Thus, we are able to estimate the cost of eye surgery in
that city. We can compare the average scores on the same test that was
administered to the two classes in order to assess how well they performed.
Therefore, a distribution is reduced to a single value that is meant to represent
the distribution when the average is calculated. This facilitates comparisons
with other distributions as well as individual evaluations of a distribution. In the
previous unit, you have learnt about classification, tabulation, and graphical
representations of data. This unit will address the measures of central
tendency: mean, median, and mode. We will explore these techniques,
highlighting their properties, advantages, and limitations. Additionally, we will
learn how to determine the mean, median, and mode for both grouped and
ungrouped data.
Objectives
After studying this unit, you should be able to:
A measure of central tendency should be calculated for the data with open-
end classes.
2. Median
3. Mode
4. Geometric Mean
5. Harmonic Mean 69
Block 1 Descriptive of Statistics
SAQ 1
a) Discuss in detail: Measures of Central Tendency.
1. Direct method
x
,
x
,
.
.
.
.
.
.
.
.
,
x
calculated mathematically as
x1 + x2 + ..... + xn
X =
n
X =
n
x
i =1 i
n
If f i is the frequency of xi , (i = 1, 2..., n) then the formula for mean would be
f1 x1 + f 2 x2 + ... + f n xn
X=
f1 + f 2 + ... + f n
X=
n
fx
i =1 i i
n
f
i =1 i
2. Short-cut method
The mean can also be calculated by taking differences from a random point
‘A’, therefore the formula will be as
X = A+
n
i =1 di
n
Where, d i = xi − A
70
If f i is the frequency of xi , (i = 1, 2..., n) then the formula for mean would be
Unit 3 Measures of Central Tendency
X = A+
n
i =1 fd i i
n
f
i =1 i
Where, d i = xi − A
Ex.1. Find the mean of the weights of seven students (in kg) observed in a
medical check-up.
Ans: If we denote the weight of students by x then the mean can be calculated
by
x1 + x2 + ..... + xn
X =
n
Therefore,
Ex.2. Find the average weight of the students by using the short-cut method
for the given data in Ex.1.
Ans: To apply the short-cut method, we will use the following formula
X = A+
n
i =1 di
n
Where, d i = xi − A
Let us assume A equals 50 in the given data in Example 1. Then, for the
calculation of di we prepare the following table:
x d = x− A
51 51 50 1
45 45 50 5
46 46 50 4
50 50 50 0
48 48 50 2
61 61 50 11
56 56 50 6
d
i =1
i =7
71
Block 1 Descriptive of Statistics
We have considered A=50; therefore, the mean can be obtained by using the
formula stated above. Then,
X = A + i =1
7
di 7
= 50 + = 50 + 1 = 51
7 7
Ex.3. Calculate the mean from the table where the frequency distribution of 12
students according to their weight is given below.
Frequency 3 7 2
x f fx
40 3 120
50 7 350
60 2 120
3 3
f
i =1
i = 12 fx
i =1
i i = 590
Therefore,
X= 3
f x 590
i =1 i i
= = 49.1
3
f
i =1 i 12
SAQ 2
a) In a water pollution study, a mussels sample was taken, and lead
concentration (milligrams per gram dry weight) was measured from each
one. The following data were obtained:
No. of students 3 7 2
72
Unit 3 Measures of Central Tendency
f1 x1 + f 2 x2 + ... + f n xn
X=
f1 + f 2 + ... + f n
X =
n
i =1 f i x i
=
fx = fx
n
i =1 f i f N
where, N = f1 + f 2 + ... + f n
2. Short-cut method
If f i is the frequency of xi ,(i = 1, 2..., n) where xi is the mid value of the ith class
interval, then the formula for the arithmetic mean will be as follows
X = A+
n
i =1 i ifd
n
i =1 if
where, d i = xi − A
Ex. 4. Calculate the mean from the table below, which shows the frequency
distribution of 25 diabetic patients according to their carbohydrate (in
gm.) intake.
Frequency 3 8 5 4 5
i =1
fi = 25 i =1
f i xi = 1095
Therefore,
X= 5
f x 1095
i =1 i i
= = 43.8 gm.
5
f
i =1 i 25
73
Block 1 Descriptive of Statistics
SAQ
SAQ 3
The following data are on the protein intake (in gm) of 25 players:
Frequency 3 8 5 4 10
Property 2. The sum of squares of deviations taken from the mean is the least
in comparison to the same taken from any other average.
Property 3. The arithmetic mean is affected by both the change of origin and
scale.
• It is rigidly defined;
3.6 MEDIAN
The variable's value that splits the whole distribution in half is called the
median. The data should be arranged in either ascending or descending order
74 of magnitude, it should be noted. The median is the data's middle value when
Unit 3 Measures of Central Tendency
the number of observations is odd. There will be two middle values for an even
number of observations. Therefore, we take these two middle values and take
their arithmetic mean. Number of the observations below and above the
median are the same. The median is not affected by extremely large or
extremely small values (as it corresponds to the middle value), and it is also
not affected by open-end class intervals.
SAQ 4
a) What is a statistical average? Write down the properties of arithmetic
mean.
n +1
Median = th observation, when n is odd.
2
n n
th observation + + 1 th observation
2 2
Median =
2
When n is even.
7, 8, 5, 11, 12, 9, 15
5, 7, 8, 9, 11, 12, 15
(n + 1) (7 + 1)
th observation = = 4th observation
2 2
5, 7, 8, 9, 11, 12
5, 7, 8, 9, 11, 12 75
Block 1 Descriptive of Statistics
Here, the total no. of observations is 6, i.e. n is even. So we will calculate the
median by
n n
th observation + + 1 th observation
2 2
Median =
2
6 6
th observation + + 1 th observation
2 2
=
2
3rd observation + 4th observation
=
2
8 + 9 17
Median = = = 8.5
2 2
SAQ 5
Find the median of the following values:
f N
Median = Value of variable corresponding to th = th cumulative
2 2
frequency.
N
Note: If is not the exact cumulative frequency, then median is the value of
2
the variable that corresponds to the subsequent cumulative frequency.
f
i =1
i = 20
20
Therefore, Median = Value of the variable corresponding to = 10 . Since
2
10 is not among the cumulative frequency, and the next cumulative frequency
is 12, and the value of variable against 12 cumulative frequency is 45. So, the
median is 45 gm.
N
−F
2
Median = l1 + ×c
fm
N= total frequency;
Median class is the class in which the (N/2)th observation falls. If N/2 is not
among any cumulative frequency, then the next class to the N/2 will be
considered the median class.
f
i =1
i = 25
77
Block 1 Descriptive of Statistics
Here
5
N
f
i −1
i = N = 25
2
= 12.5
Since 12.5 is not among the cumulative frequencies, the class with the next
cumulative frequency, i.e. 16. Therefore, our median class will be 40−50.
N= total frequency = 25
N
−F
2
Median = l1 + ×c
fm
12.5 − 11
= 40 + × 10 = 43
5
SAQ 6
a) The following data are on the protein intake (in gm) by 30 players:
Protein intake (in gm) 30 40 50 60 70
Frequency 3 8 5 4 10
Calculate median.
4. It can be calculated even for open-end classes (like "less than 10" or "50
78 and above").
Unit 3 Measures of Central Tendency
Demerits of Median
SAQ 7
Write a short note about median and its merits and demerits.
3.8 MODE
The most frequent observation in the distribution is known as a mode. In other
words, the mode is that observation in a distribution which has the maximum
frequency. For example, when we say that the average blood sugar level
marked in a clinic is 150 mg, and it is the modal value which is observed most
frequently among diabetic patients.
140, 130, 150, 180, 160, 140, 150, 150, 170, 160
The above table shows that 150 has the maximum frequency. Thus, the mode
is 150 mmHg.
The modal class is the class that has the maximum frequency.
In line with the highest frequency, 8, modal class is 30−40, and we have
Therefore,
( f 0 − f −1 )
Mode = l1 + ×c
2 f 0 − f −1 − f1
(8 − 3)
Mode = 30 + × 10 = 36.25 gm.
(2 × 8 − 3 − 5)
SAQ 8
Calculate the mode for the given data
Daily protein intake (in gm.) 0−20 20−40 40−60 60−80 80−100
No. of students 5 15 30 16 14
80
Unit 3 Measures of Central Tendency
SAQ 9
The mean is 41.6 and the mode is 34.4 in an asymmetrical distribution.
Determine the median.
Demerits of Mode
1. It is not rigidly defined. A distribution can have more than one mode;
Ans:
1
GM = ( x1 x2 x3 ...xn ) n
1
= (118 × 120 × 122 × 160) = 128.9 mmHg. 4
SAQ 10
Take a look at the four systolic blood sugar readings in mmHg that follow:
) i
f
GM = ( x1 x2 x3 ...xn
f1 f2 f3 fn
1
( f f
GM = x1 1 x 2 2 x 3 3 ...x n
f fn N
)
Where, N = f1 + f 2 + f 3 + ..... + f n
1
log GM = log ( x1 f1 x2 f2 x3 f3 ...xn fn )
N
1
log GM = (f1 log x1 + f2 log x 2 + f3 log x 3 + ... + f n log x n
N
1 n
f log x
A
n
t
i
l
o
g
GM =
N i i
i =1
Frequency 3 5 7 9 4
82
Unit 3 Measures of Central Tendency
Ans:
1 5
f log x
A
n
t
i
l
o
g
GM =
28 i i
i =1
38.2725
A
n
t
i
l
o
g
GM =
28
A
n
t
i
i
l
o
g
Only in cases like GM, where no observation is zero, is the harmonic mean
defined.
1
HM =
11 1 1
+ + ... +
n x1 x2 xn
n
HM =
1
n
i =1
xi
3 3
HM = =
1 1 1 1+1+ 1
+ + 3 5 10
x1 x2 x3
= 4.73 ml
1
HM =
1 f1 f 2 fn
+ + .... +
N x1 x2 xn
N
HM =
fi
n
i =1
xi
where N = n
f
i =1 i
When equal distances are travelled at different speeds, the average speed is
calculated by the harmonic mean.
Therefore,
N 28
HM = = = 17.956
fi 1.555
5
i =1
xi
SAQ 11
Calculate the HM for the data given in Ex.4.
84
Unit 3 Measures of Central Tendency
1. It is rigidly defined;
1. AM ≥ GM ≥ HM
2. GM = AM × HM
3.15.1 Quartiles
The entire distribution is divided into four equal parts by quartiles. These are
first quartile (Q1), second quartile (Q2), third quartile (Q3). Here, Q1 carries ¼
part of the data, Q2 carries ½ of the data, Q3 carries ¾ part of the data. It may
be noted here that the data should be arranged in ascending or descending
order of magnitude.
iN
−F
4 × c for i = 1, 2, 3.
Qi = l1 +
fq
N = total frequency;
i× N
Here ‘i’ denotes the ith quartile class in which th observation falls in
4
cumulative frequency.
3.15.2 Deciles
Deciles divide the data into ten equal parts. This means that there is a total of
nine deciles. These nine deciles are denoted by D1, D2,…, D9, where D1
stands for the 1st Decile, D2 stands for 2nd Decile, and so on. Here, ith Decile
contains × th part of the data. It may be noted that the data should be
i N
10
arranged in ascending or descending order of magnitude.
If x1 , x2 ,..., xN are the N values of a variable X, then mathematically, then the ith
Decile is defined as
iN
−F
10
Di = l1 + × c for i = 1, 2, 3,…., 9
fd
N = total frequency;
i× N
Here ‘i’ denotes the ith Decile class in which th observation falls in
10
cumulative frequency.
3.15.3 Percentiles
The data is divided into 100 equal parts using percentiles. It indicates that
there are 99 percentiles in total. The notation for these ninety-nine percentiles
is P1, P2,…, P99, where P1 represents the first percentile, P2 the second, and
i× N
so forth. Here ith percentile contains th part of data. The data should
100
be arranged either in ascending or descending order of magnitude, it should
be noted.
(i = 1, 2, 3,…, 99)
iN
−F
Pi = li + × c for i = 1, 2, 3,…., 99
100
fp
N = total frequency;
i× N
Here ‘i’ denotes the ith percentile class where th observation falls in
100
cumulative frequency.
Ex.15. Determine the first and third quartiles for the data in Example 4.
Ans: First, we find the cumulative frequency given in the given cumulative
frequency table:
f
i =1
i = 25
In this case, N/4 = 25/4 = 6.25. The observation thus belongs to class 30–40.
Thus, it is the class in the first quartile. 3N/4 = 75/4 = 18.75 is likewise true. As
a result, the observation belongs to class 50−60. It is, therefore, the third
88 quartile class.
Unit 3 Measures of Central Tendency
For the first quartile, l1 = 30;
N = 25;
F=3
f p = 8 and
c = 10.
(6.25 − 3)
Q1 = 30 + × 10 = 34.06
8
For the third quartile, l1 = 50;
N = 25;
F = 16
f p = 4 and
c = 10.
(18.75 − 16)
Q3 = 50 + × 10 = 56.88
4
SAQ 12
a) What is meant by quartiles?
3.16 SUMMARY
In this unit, we have discussed:
• The geometric mean (GM) is the nth root of nth observations, used for
averaging ratios and proportions. However, it fails to provide the correct
average when an observation is zero or negative.
• Partition values are variables that divide a distribution into equal parts,
typically in ascending or descending order of magnitude.
2. Describe the formula for arithmetic mean for both grouped and
ungrouped data. Also, mention the merits and demerits of arithmetic
mean.
4. What are the quartiles? How would you calculate quartiles for both
grouped and ungrouped data?
3.18 ANSWERS
Self-Assessment Questions
1. a) Professor Bowley described the mean or average as "statistical
constant which enable us to comprehend in a single effort the
significance of the whole." It provides us with an idea of how
concentrated the values are in the middle portion of the
distribution. In simple words, an average or mean of a statistical
series is the value of the variable which acts as a good
representation of the entire distribution.
So,
X =
x
n
112.5 + 140.5 + 162.3 + 170.8 + 181.9 + 201.4
= = 161.57 miligram per gram
6
91
Block 1 Descriptive of Statistics
i =1
fi = 12 fx
i =1
i i = 505
X=
3
fx
i =1 i i
=
505
= 168.33 gm
3
f
i =1 i 3
i =1
f i = 30 f x
i =1
i i = 1705
X =
5
i =1fi x i
=
1705
= 56.83 gm.
5
i =1fi 30
1. Direct method
x
,
x
,
.
.
.
.
.
.
.
.
,
x
x1 + x2 + ..... + xn
X =
92 n
Unit 3 Measures of Central Tendency
The above equation can also be written as
X =
n
x
i =1 i
n
If f i is the frequency of xi , (i = 1, 2..., n) then the formula for mean
would be
f1 x1 + f 2 x2 + ... + f n xn
X=
f1 + f 2 + ... + f n
X=
n
fx
i =1 i i
n
i =1 if
f1 x1 + f 2 x2 + ... + f n xn
X=
f1 + f 2 + ... + f n
X=
n
i =1 f i x i
=
fx = fx
n
i =1 f i f N
where, N = f1 + f 2 + ... + f n
• It is rigidly defined; 93
Block 1 Descriptive of Statistics
2, 3, 6, 8, 10, 12, 15
Therefore
(n + 1) 8
Median = th observation = th = 4th observation
2 2
Therefore, Median = 8
2, 3, 5, 6, 8, 10, 12, 15
Therefore,
n n
th observatio n + + 1 th observatio n
Median = 2
2
2
8 8
th observatio n + + 1 th observatio n
Median =
2 2
2
( + )
4
t
h
o
b
s
e
r
v
a
t
i
o
n
5
t
h
o
b
s
e
r
v
a
t
i
o
n
6
8
+
M
e
d
i
a
n
= = =
2
f
i =1
i = 30
The Median class is 40−60 since the median value is given as 50,
N
−F
2
Median = l1 + ×c
fm
69 + x
− 27
2
50 = 40 + × 20
30
(69 + x − 54)
50 − 40 =
3
30 = 15 + x
x = 30 − 15 = 15
7. The variable's value that splits the whole distribution in half is called the
median. The data should be arranged in either ascending or descending
order of magnitude. The median is the data's middle value when the
number of observations is odd. There will be two middle values for an
even number of observations. Therefore, we take these two middle
values and take their arithmetic mean. A number of the observations
below and above the median are the same. The median is not affected
by extremely large or extremely small values (as it corresponds to the
middle value), and it is also not affected by open-end class intervals.
Merits of Median
• It is rigidly defined;
• It can be calculated even for open-end classes (like "less than 10"
or "50 and above"). 95
Block 1 Descriptive of Statistics
Demerits of Median
Therefore,
f 0 − f −1
Mode = l1 + ×c
2 f 0 − f −1 − f1
(30 − 15)
Mode = 40 + × 20 = 50.34 gm
(60 − 15 − 16)
9. We know that
Mode = 3 Median − 2 Mean
3 Median = 117.6
117.6
Median = = 39.2
3
N 25
HM = = = 36.23 gm.
fi 0.69
5
i =1
xi
12. a) The entire distribution is divided into four equal parts by quartiles.
Those are first quartile (Q1), second quartile (Q2), third quartile
(Q3). Here, Q1 carries ¼ part of the data, Q2 carries ½ of the data,
Q3 carries ¾ part of the data. It may be noted here that the data
should be arranged in ascending or descending order of
magnitude.
iN
−F
4
Qi = l1 + × c for i = 1, 2, 3.
fq
N = total frequency;
F = cumulative frequency of the class prior to ith quartile class and
f
i =1
i = 80
Here, N/4 = 80/4 = 20. Therefore, the observation falls in the class
20-40. So it is the first quartile class. Similarly, 3(80)/4 = 240/4 =
60. Therefore, the observation falls in the class 60−80. So it is the
third quartile class.
N = 80;
F = 5;
f q = 15 and
c = 20
(20 − 5)
Q1 = 20 + × 10 = 30 gm.
15
For the third quartile, l1 = 60;
N = 80;
F = 50
f q = 16 and
c = 20.
(60 − 50)
Q3 = 60 + × 10 = 66.25 gm.
16
98
Unit 3 Measures of Central Tendency
Terminal Questions
1. Refer to Section 3.2 and Subsection 3.2.1.
2. Refer to Sections 3.3 and 3.5 and Subsections 3.3.1 and 3.3.2.
Suggested Readings
Belle G. V., Fisher, L. D., Heagerty P.J. and Lumley, T. (2004).
Biostatistics: A methodology for the Health sciences, Hoboken, New
Jersey. : John Wiley and Sons.
99