0% found this document useful (0 votes)
2 views33 pages

Unit 3

Unit 3 focuses on measures of central tendency, including mean, median, and mode, and their properties, advantages, and limitations. It explains how to calculate these measures for both ungrouped and grouped data, emphasizing the importance of averages in data analysis. The unit also outlines desirable properties for a suitable measure of central tendency and provides examples for better understanding.

Uploaded by

skcyberwork
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views33 pages

Unit 3

Unit 3 focuses on measures of central tendency, including mean, median, and mode, and their properties, advantages, and limitations. It explains how to calculate these measures for both ungrouped and grouped data, emphasizing the importance of averages in data analysis. The unit also outlines desirable properties for a suitable measure of central tendency and provides examples for better understanding.

Uploaded by

skcyberwork
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Unit 3 Measures of Central Tendency

UNIT 3
MEASURES OF CENTRAL
TENDENCY

Structure
3.1 Introduction 3.9 Relationship Between Mean,
Median and Mode
Objectives
3.10 Merits and Demerits of Mode
3.2 Measures of Central
Tendency 3.11 Geometric Mean
Properties for a Suitable For Ungrouped Data
Measure of Central Tendency
For Grouped Data
Different Measures of Central
3.12 Harmonic Mean
Tendency
For Ungrouped Data
3.3 Arithmetic Mean or Mean
For Grouped Data
For Ungrouped Data
3.13 Merits and Demerits of
For Grouped Data
Harmonic Mean
3.4 Properties of Arithmetic
3.14 Relation Between AM, GM
Mean
and HM
3.5 Merits and Demerits of
3.15 Partition Values
Arithmetic Mean
Quartiles
3.6 Median
Deciles
Median for Ungrouped Data
Percentiles
Median for Ungrouped Data
(When Frequencies Are Given) 3.16 Summary
Median for Grouped Data 3.17 Terminal Questions
3.7 Merits and Demerits of 3.18 Answers
Median
3.8 Mode
For Ungrouped Data
For Grouped Data

3.1 INTRODUCTION
As we know, after the classification and tabulation of data, one often needs
more details for the many uses that may be made on the information available.
67
Block 1 Descriptive of Statistics

We, therefore, need further analysis of the tabulated data to draw inferences.
In this unit, we will discuss measures of central tendency. For the purpose of
analysis, a very important and powerful tool is the only average value that
stands for the entire mass of data. In statistics, an average is a one-figure
representation of a distribution. It provides a value that the distribution is
concentrated around it. This is why the measure of central tendency is another
name for the average. For instance, eye surgery typically costs Rs. 15,000 in a
city on an average. Thus, we are able to estimate the cost of eye surgery in
that city. We can compare the average scores on the same test that was
administered to the two classes in order to assess how well they performed.
Therefore, a distribution is reduced to a single value that is meant to represent
the distribution when the average is calculated. This facilitates comparisons
with other distributions as well as individual evaluations of a distribution. In the
previous unit, you have learnt about classification, tabulation, and graphical
representations of data. This unit will address the measures of central
tendency: mean, median, and mode. We will explore these techniques,
highlighting their properties, advantages, and limitations. Additionally, we will
learn how to determine the mean, median, and mode for both grouped and
ungrouped data.

Objectives
After studying this unit, you should be able to:

 explain the concept and importance of central tendency of data;


 describe the various measures of central tendency;
 explain the properties, advantages and limitations of mean, median
and mode;
 calculate the different measures of central tendency for ungrouped and
grouped data; and
 describe the methods of calculation of partition values.

3.2 MEASURES OF CENTRAL TENDENCY


Professor Bowley described the mean or average as "statistical constant
which enables us to comprehend in a single effort the significance of the
whole." It provides us with an idea of how concentrated the values are in the
middle portion of the distribution. In simple words, an average or mean of a
statistical series is the value of the variable which acts as a good
representation of the entire distribution.

3.2.1 Properties for a Suitable Measure of


Central Tendency
Before we start the detailed discussion about measures of central tendency,
we should know the desirable properties of it. The properties are listed below
68 as:
Unit 3 Measures of Central Tendency
1. It needs to be precisely defined.

For a measure of central tendency to have a suitable interpretation, it must be


properly defined. In order for different people to calculate the average from the
same figures and obtain the same result, it should also have an algebraic
formula. It ought to be founded on every observation. It should be readily
comprehensible.

2. It must be readily comprehensible

An average should be easy to understand because we use measures of


central tendency to simplify the complexity of data; otherwise, its use is likely
to be very limited.

3. It is required to be easily calculated.

An average should be easy to understand and simple to compute so that it can


be used as widely as possible.

4. It should not be affected by sampling fluctuations.

A tool with sampling stability should be preferred. To put it another way, we


should anticipate obtaining roughly the same results if we choose ten distinct
groups of observations from the same population and calculate the average
for each group. Because of the sampling fluctuation alone, there might not be
much of a difference.

5. It shouldn’t be affected by extreme values.

Every observation is thought to have an impact on the average's value. The


average cannot be regarded as a good average if it is significantly impacted by
one or two extremely small or extremely large observations, either by
increasing or decreasing its value.

6. It should be possible to calculate even for open-end class


intervals.

A measure of central tendency should be calculated for the data with open-
end classes.

7. It should be readily lend itself to algebraic treatment.

The algebraic manipulations should be attributed to a measure of central


tendency. One can find information about the combined set even if there are
two sets of data and the individual information is available for both sets.

3.2.2 Different Measures of Central Tendency


1. Arithmetic Mean or simply Mean

2. Median

3. Mode

4. Geometric Mean

5. Harmonic Mean 69
Block 1 Descriptive of Statistics

SAQ 1
a) Discuss in detail: Measures of Central Tendency.

b) What are the properties that construct a worthy measure of central


tendency?

3.3 ARITHMETIC MEAN OR MEAN


The formula for calculating the arithmetic mean is defined as sum of all
observations divided by total number of observations. Arithmetic mean (AM)
may be calculated for two types of data given as follows

3.3.1 For Ungrouped Data


For ungrouped data, arithmetic mean may be computed by applying any of the
following methods:

1. Direct method
x
,
x
,
.
.
.
.
.
.
.
.
,
x

If are the n number of observations, then their mean can be


1

calculated mathematically as

x1 + x2 + ..... + xn
X =
n

The above equation can also be written as

X =
 n
x
i =1 i

n
If f i is the frequency of xi , (i = 1, 2..., n) then the formula for mean would be

f1 x1 + f 2 x2 + ... + f n xn
X=
f1 + f 2 + ... + f n

The above equation can also be written as

X=
 n
fx
i =1 i i

 n
f
i =1 i

2. Short-cut method

The mean can also be calculated by taking differences from a random point
‘A’, therefore the formula will be as

X = A+
 n
i =1 di
n
Where, d i = xi − A

70
If f i is the frequency of xi , (i = 1, 2..., n) then the formula for mean would be
Unit 3 Measures of Central Tendency

X = A+ 
n
i =1 fd i i

 n
f
i =1 i

Where, d i = xi − A

We usually use the short-cut method when data are large.

Ex.1. Find the mean of the weights of seven students (in kg) observed in a
medical check-up.

51, 45, 46, 50, 48, 61, 56

Ans: If we denote the weight of students by x then the mean can be calculated
by

x1 + x2 + ..... + xn
X =
n

Therefore,

(51 + 45 + 46 + 50 + 48 + 61 + 56) 357


X = = = 51
7 7

So the average weight of students is 51 Kg.

Ex.2. Find the average weight of the students by using the short-cut method
for the given data in Ex.1.

Ans: To apply the short-cut method, we will use the following formula

X = A+
 n
i =1 di
n
Where, d i = xi − A

Let us assume A equals 50 in the given data in Example 1. Then, for the
calculation of di we prepare the following table:

x d = x− A
51 51 50 1

45 45 50 5

46 46 50 4

50 50 50 0

48 48 50 2

61 61 50 11

56 56 50 6

d
i =1
i =7

71
Block 1 Descriptive of Statistics

We have considered A=50; therefore, the mean can be obtained by using the
formula stated above. Then,

X = A +  i =1
7
di 7
= 50 + = 50 + 1 = 51
7 7

So the mean weight of students is 51 Kg.

Ex.3. Calculate the mean from the table where the frequency distribution of 12
students according to their weight is given below.

Weight (in Kg.) 40 50 60

Frequency 3 7 2

Ans: If we denote the weights of students by x and frequency by f then we


have the following frequency distribution:

x f fx

40 3 120

50 7 350

60 2 120
3 3

f
i =1
i = 12 fx
i =1
i i = 590

Therefore,

X=  3
f x 590
i =1 i i
= = 49.1
 3
f
i =1 i 12

So, the mean weight of students is 49.1 Kg.

SAQ 2
a) In a water pollution study, a mussels sample was taken, and lead
concentration (milligrams per gram dry weight) was measured from each
one. The following data were obtained:

112.5, 140.5, 162.3, 170.8, 181.9, 201.4

Calculate average lead concentration.

b) Calculate the mean level of carbohydrates from the table where


frequency distribution of 12 students according to their carbohydrates (in
gm.) intake is given below.

Carbohydrates (in gm.) 30 45 50

No. of students 3 7 2

72
Unit 3 Measures of Central Tendency

3.3.2 For Grouped Data


1. Direct method
If f i is the frequency of xi ,(i = 1, 2..., n) where xi is the mid value of the ith class
interval, then the formula for the arithmetic mean will be as follows:

f1 x1 + f 2 x2 + ... + f n xn
X=
f1 + f 2 + ... + f n

X =
 n
i =1 f i x i
=
 fx =  fx
 n
i =1 f i f N
where, N = f1 + f 2 + ... + f n

2. Short-cut method
If f i is the frequency of xi ,(i = 1, 2..., n) where xi is the mid value of the ith class
interval, then the formula for the arithmetic mean will be as follows

X = A+ 
n
i =1 i ifd
 n
i =1 if

where, d i = xi − A

Ex. 4. Calculate the mean from the table below, which shows the frequency
distribution of 25 diabetic patients according to their carbohydrate (in
gm.) intake.

Carbohydrates (in gm.) 0−30 30−40 40−50 50−60 60−70

Frequency 3 8 5 4 5

Ans: If we denote the carbohydrates (in gm.) intake of patients by class


interval, frequency by f and mid value by x then we have the following
frequency distribution:

Class interval Mid value x f fx


0−30 15 3 45
30−40 35 8 280
40−50 45 5 225
50−60 55 4 220
60−70 65 5 325
5 5

i =1
fi = 25 i =1
f i xi = 1095

Therefore,

X=  5
f x 1095
i =1 i i
= = 43.8 gm.
 5
f
i =1 i 25
73
Block 1 Descriptive of Statistics

SAQ
SAQ 3
The following data are on the protein intake (in gm) of 25 players:

Protein intake (in gm) 0−40 40−50 50−60 60−70 70−80

Frequency 3 8 5 4 10

Calculate the mean.

3.4 PROPERTIES OF ARITHMETIC MEAN


The arithmetic mean satisfies all properties to become a good average except
the last two. It is particularly useful when dealing with a sample as it is least
affected by sampling fluctuations. It is the most popular average and should
always be our first choice unless there is a lack of suitability.

Some algebraic properties of the arithmetic mean are given below:

Property 1. The algrebraic sum of deviations of observations from their mean


is zero.

Property 2. The sum of squares of deviations taken from the mean is the least
in comparison to the same taken from any other average.

Property 3. The arithmetic mean is affected by both the change of origin and
scale.

3.5 MERITS AND DEMERITS OF ARITHMETIC


MEAN
Merits of Arithmetic Mean

• It utilizes all the observations;

• It is rigidly defined;

• It is simple to comprehend and compute; and

• It can be used for further mathematical treatments.

Demerits of Arithmetic Mean


• It is badly affected by extremely small or extremely large values;

• It can’t be evaluated for open-end class intervals; and

• It is generally not preferred for highly skewed distributions.

3.6 MEDIAN
The variable's value that splits the whole distribution in half is called the
median. The data should be arranged in either ascending or descending order
74 of magnitude, it should be noted. The median is the data's middle value when
Unit 3 Measures of Central Tendency
the number of observations is odd. There will be two middle values for an even
number of observations. Therefore, we take these two middle values and take
their arithmetic mean. Number of the observations below and above the
median are the same. The median is not affected by extremely large or
extremely small values (as it corresponds to the middle value), and it is also
not affected by open-end class intervals.

SAQ 4
a) What is a statistical average? Write down the properties of arithmetic
mean.

b) State the merits and demerits of arithmetic mean.

3.6.1 Median for Ungrouped Data


If x1 , x2 ,......., xn are the n observations, then for obtaining the median, first of
all, we have to arrange these n values either in ascending order or in
descending order. If n is odd, the median is the middle value when the
observations are arranged in either ascending or descending order of
magnitude There will be two intermediate values for an even number of
observations. We, therefore, calculate the arithmetic mean of these two
numbers.

 n +1
Median =   th observation, when n is odd.
 2 

n n 
  th observation +  + 1 th observation
2 2 
Median =  
2

When n is even.

Ex.5. Find the median of the following observations

7, 8, 5, 11, 12, 9, 15

Ans: The provided data is first arranged in ascending order as

5, 7, 8, 9, 11, 12, 15

Here, the total no. of observations is 7, i.e. n is odd.

Therefore, the median would be the

(n + 1) (7 + 1)
th observation = = 4th observation
2 2

So the median would be the middle value, i.e. 9.

Ex.6. Find the median of the following observations

5, 7, 8, 9, 11, 12

Ans: First, we arrange the given data in ascending order as

5, 7, 8, 9, 11, 12 75
Block 1 Descriptive of Statistics

Here, the total no. of observations is 6, i.e. n is even. So we will calculate the
median by

n n 
  th observation +  + 1 th observation
2 2 
Median =  
2

6 6 
  th observation +  + 1 th observation
2 2 
= 
2
3rd observation + 4th observation
=
2

8 + 9 17
Median = = = 8.5
2 2

SAQ 5
Find the median of the following values:

a) 10, 6, 15, 2, 3, 12, 8;

b) 10, 6, 2, 3, 8, 15, 12, 5.

3.6.2 Median for Ungrouped Data (When Frequencies


Are Given)
If xi are the different values of variable with frequencies f i , then we calculate
cumulative frequencies from f i then, the median is defined by

 f  N
Median = Value of variable corresponding to   th =   th cumulative
 2  2
frequency.

N
Note: If is not the exact cumulative frequency, then median is the value of
2
the variable that corresponds to the subsequent cumulative frequency.

Ex.7. Calculate the median from the given frequency distribution.

Carbohydrates (in gm.) 30 45 50 60


Frequency 3 9 2 6

Ans: First, we will calculate the cumulative frequency.

Carbohydrates (in gm.) Frequency Cumulative Frequency


30 3 03
45 9 12
50 2 14
76
Unit 3 Measures of Central Tendency
60 6 20
4

f
i =1
i = 20

 20 
Therefore, Median = Value of the variable corresponding to   = 10 . Since
 2 
10 is not among the cumulative frequency, and the next cumulative frequency
is 12, and the value of variable against 12 cumulative frequency is 45. So, the
median is 45 gm.

3.6.3 Median for Grouped Data


For class interval, first, we find cumulative frequencies from the given
frequencies and the class in which cumulative frequency lies, i.e. N/2, is
termed as the median class. Now, we will use the following formula to
calculate the median:

N 
 −F
2
Median = l1 +   ×c
fm

where l1 = lower boundary of median class;

N= total frequency;

F= cumulative frequency of the class prior to median class;

f m = frequency of median class; and

c = width of median class.

Median class is the class in which the (N/2)th observation falls. If N/2 is not
among any cumulative frequency, then the next class to the N/2 will be
considered the median class.

Ex.8. Calculate the median for the data given in Ex.4.

Ans: We start by creating a cumulative frequency table, which shows the


number of observations below the class's upper limit (officially known as
cumulative frequency of less than type); we also have cumulative frequency of
more than type, which shows the number of observations above or equal to a
class's lower limit:
Class interval frequency Cumulative frequency
(less than type)
0−30 3 3
30−40 8 11
40−50 5 16
50−60 4 20
60−70 5 25
5

f
i =1
i = 25

77
Block 1 Descriptive of Statistics
Here
5
N
f
i −1
i = N = 25 
2
= 12.5

Since 12.5 is not among the cumulative frequencies, the class with the next
cumulative frequency, i.e. 16. Therefore, our median class will be 40−50.

So we have, l1 = lower boundary of median class = 40

N= total frequency = 25

F= cumulative frequency below l1 = 11

f m = frequency of median class = 5

c = width of median class = 10

Now putting all these values in the formula of the Median

N 
 −F
2
Median = l1 +   ×c
fm

12.5 − 11
= 40 + × 10 = 43
5

Therefore, median is 43 gm.

SAQ 6
a) The following data are on the protein intake (in gm) by 30 players:
Protein intake (in gm) 30 40 50 60 70

Frequency 3 8 5 4 10

Calculate median.

b) Find the missing frequency when the median is given as 50 gm.


Daily protein intake 0−20 20−40 40−60 60−80 80−100
(in gm)
No. of students 8 19 30 -- 12

3.7 MERITS AND DEMERITS OF MEDIAN


Merits of Median
1. It is rigidly defined;

2. It is easy to understand and compute;

3. It is not affected by extremely small or extremely large values; and

4. It can be calculated even for open-end classes (like "less than 10" or "50
78 and above").
Unit 3 Measures of Central Tendency
Demerits of Median

1. In the case of even number of observations, we get only an estimate of


the median by taking the mean of the two middle values. We don't get its
exact value;

2. It does not utilize all the observations. The median of 1, 2 and 3 is 2. If


observation 3 is replaced by any number higher than or equal to 2 and if
the number 1 is replaced by any number lower than or equal to 2, the
median value will be unaffected. This means 1 and 3 are not being
utilized;

3. It is not amenable to algebraic treatment; and

4. It is affected by sampling fluctuations

SAQ 7
Write a short note about median and its merits and demerits.

3.8 MODE
The most frequent observation in the distribution is known as a mode. In other
words, the mode is that observation in a distribution which has the maximum
frequency. For example, when we say that the average blood sugar level
marked in a clinic is 150 mg, and it is the modal value which is observed most
frequently among diabetic patients.

3.8.1 For Ungrouped Data


Mathematically, if x1 , x 2 ,..., x n are the n observations, and if some of the
observations are repeated in the data, say x i is repeated the highest times.
Then we can say x i would be the modal value.

Ex.9. The systolic blood pressure (SBP) in mmHg of 10 patients is given


below:

140, 130, 150, 180, 160, 140, 150, 150, 170, 160

Ans: First, we prepare the frequency table as

SBP 130 140 150 160 170 180


No. of 1 2 3 2 1 1
patients
(Frequency)

The above table shows that 150 has the maximum frequency. Thus, the mode
is 150 mmHg.

3.8.2 For Grouped Data


The following mode formula is applied to data with multiple classes. 79
Block 1 Descriptive of Statistics
( f 0 − f −1 )
Mode = l1 + ×c
2 f 0 − f −1 − f1

where, l1 = lower class limit of the modal class;

f 0 = frequency of the modal class;

f1 = frequency of the pre-modal class;

f1 = frequency of the post-modal class and

c = width of the modal class.

The modal class is the class that has the maximum frequency.

Ex.10. For the data given in Ex.4, calculate mode.

Ans: Here, the frequency distribution is

Class interval frequency


0−30 3
30−40 8
40−50 5
50−60 4
60−70 5

In line with the highest frequency, 8, modal class is 30−40, and we have

l1 = lower class limit of the modal class = 30

f −1 = frequency of the pre-modal class = 3

f 0 = frequency of the modal class = 8

f1 = frequency of the post-modal class = 5

c = width of the modal class = 10

Therefore,

( f 0 − f −1 )
Mode = l1 + ×c
2 f 0 − f −1 − f1

(8 − 3)
Mode = 30 + × 10 = 36.25 gm.
(2 × 8 − 3 − 5)

SAQ 8
Calculate the mode for the given data

Daily protein intake (in gm.) 0−20 20−40 40−60 60−80 80−100
No. of students 5 15 30 16 14

80
Unit 3 Measures of Central Tendency

3.9 RELATIONSHIP BETWEEN MEAN, MEDIAN


AND MODE
In a symmetrical distribution, the mode, median, and mean all fall within the
same range. There is an empirical connection between them, though, if the
distribution is somewhat asymmetrical. The partnership is
Mean − Mode = 3 (Mean − Median )

Mode = 3 Median – 2 Mean

Using this formula, we can calculate mean/median/mode if the other two of


them are known.

SAQ 9
The mean is 41.6 and the mode is 34.4 in an asymmetrical distribution.
Determine the median.

3.10 MERITS AND DEMERITS OF MODE


Merits of Mode

1. Mode is the easiest average to understand and also easy to calculate;

2. It is not affected by extreme values;

3. It can be calculated for open end classes;

4. Mode can be calculated even if the classes are of unequal width,


provided the model class, pre-model class and the post model class are
of equal width.

Demerits of Mode

1. It is not rigidly defined. A distribution can have more than one mode;

2. It is not utilizing all the observations;

3. It is not amenable to algebraic treatment; and

4. It is greatly affected by sampling fluctuations.

3.11 GEOMETRIC MEAN


The nth root of the product of the n observations is the geometric mean (GM)
of n observations. It is useful for averaging ratios or proportions. It fails to give
the correct average if an observation is zero or negative.

3.11.1 For Ungrouped Data


If x1 , x 2 ,..., x n are the n observations of a variable X, then geometric mean
(GM) is given as
81
Block 1 Descriptive of Statistics
1
GM = ( x1 x 2 x3 ...x n ) n

Ex.11. Consider the following four observations of systolic blood pressure in


mmHg:

118, 120, 122, 160

Find the GM.

Ans:
1
GM = ( x1 x2 x3 ...xn ) n

1
= (118 × 120 × 122 × 160) = 128.9 mmHg. 4

SAQ 10
Take a look at the four systolic blood sugar readings in mmHg that follow:

141, 220, 182, 170

Find the geometric mean.

3.11.2 For Grouped Data


If x1 , x2 ,..., xn are the n values (or mid values for class intervals) of a variable X
where f i is the frequency of xi ,(i = 1, 2,..., n) then

) i
f
GM = ( x1 x2 x3 ...xn
f1 f2 f3 fn

1
( f f
GM = x1 1 x 2 2 x 3 3 ...x n
f fn N
)
Where, N = f1 + f 2 + f 3 + ..... + f n

Taking log both sides

1
log GM = log ( x1 f1 x2 f2 x3 f3 ...xn fn )
N
1
log GM = (f1 log x1 + f2 log x 2 + f3 log x 3 + ... + f n log x n
N

1 n

 f log x 
A
n
t
i
l
o
g

 GM = 
N i i
 i =1 

Ex.12. Find GM for the data given below

Fasting Serum Insulin 0−10 10−20 20−30 30−40 40−50


(in µU / mL )

Frequency 3 5 7 9 4
82
Unit 3 Measures of Central Tendency
Ans:

Class interval Mid Frequency (f) log x f × log x


value(x)

0−10 5 3 0.6989 2.0970

10−20 15 5 1.1761 5.8805

20−30 25 7 1.3979 9.7853

30−40 35 9 1.5441 13.8969

40−50 45 4 1.6532 6.6128


5 5
N =  fi = 28  f log x
i i = 38.2725
i =1 i =1

By applying the formula

 1 5

 f log x
A
n
t
i
l
o
g

GM =  
 28 i i 
 i =1 

 38.2725 
A
n
t
i
l
o
g

GM =  
 28 
A
n
t
i
i
l
o
g

GM = (1.3669 ) = 23.28 µU / mL.

3.12 HARMONIC MEAN


A set of observations' harmonic mean is equal to the reciprocal of the
arithmetic mean of its reciprocals.

Only in cases like GM, where no observation is zero, is the harmonic mean
defined.

3.12.1 For Ungrouped Data


If x1 , x2 ,..., xn are the n observations of a variable X, then harmonic mean
(HM) is defined as

1
HM =
11 1 1
 + + ... + 
n  x1 x2 xn 

n
HM =
1
 n
i =1
xi

Ex.13. A patient is treated with three different doses of an antibiotic injection.


The doses are 3ml, 5ml and 10ml at different hours of the day. What is
the average dose of injection?

Ans: HM is the appropriate average here. 83


Block 1 Descriptive of Statistics
n
HM =
1
 n
i =1
xi

3 3
HM = =
1 1 1  1+1+ 1
 + +  3 5 10
 x1 x2 x3 
= 4.73 ml

3.12.2 For Grouped Data


If x1 , x2 ,..., xn are the n values (or mid values in case of class intervals) of a
variable X where f i be the frequency of xi ,(i = 1, 2..., n) then

1
HM =
1  f1 f 2 fn 
 + + .... + 
N  x1 x2 xn 

N
HM =
fi
 n
i =1
xi

where N =  n
f
i =1 i

When equal distances are travelled at different speeds, the average speed is
calculated by the harmonic mean.

Ex.14. Calculate the HM for the data given in Ex.12.

Ans: We have the following distribution


Class Mid Value (x) Frequency (f) f/x
0−10 5 3 0.600
10−20 15 5 0.330
20−30 25 7 0.280
30−40 35 9 0.257
40−50 45 4 0.088
5 5
fi
N =  fi = 28
i =1
x
i =1
= 1.555
i

Therefore,

N 28
HM = = = 17.956
fi 1.555
 5
i =1
xi

SAQ 11
Calculate the HM for the data given in Ex.4.
84
Unit 3 Measures of Central Tendency

3.13 MERITS AND DEMERITS OF HARMONIC


MEAN
Merits of Harmonic Mean

1. It is rigidly defined;

2. It utilizes all the observations;

3. It is amenable to algebraic treatment; and

4. It gives greater importance to small items.

Demerits of Harmonic Mean

1. It is a difficult concept to understand and compute.

3.14 RELATION BETWEEN AM, GM AND HM


There are two relations between AM, GM and HM.

1. AM ≥ GM ≥ HM

2. GM = AM × HM

3.15 PARTITION VALUES


Partition values are those values of variable which divide the distribution into a
certain number of equal parts. Here it may be noted that the data should be
arranged in ascending or descending order of magnitude. Commonly used
partition values are quartiles, deciles and percentiles. For example, quartiles
divide the data into four equal parts. Similarly, deciles and percentiles divide
the distribution into ten and hundred equal parts, respectively.

3.15.1 Quartiles
The entire distribution is divided into four equal parts by quartiles. These are
first quartile (Q1), second quartile (Q2), third quartile (Q3). Here, Q1 carries ¼
part of the data, Q2 carries ½ of the data, Q3 carries ¾ part of the data. It may
be noted here that the data should be arranged in ascending or descending
order of magnitude.

For Ungrouped Data

We should arrange the data in either ascending or descending order of


magnitude in order to obtain the quartiles. The first quartile (Q1), second
quartile (Q2), and third quartile (Q3) are then determined by finding the (N/4)th,
(N/2)th, and (3N/4)th positioned items in the organized data. Thus, the first
quartile (Q1) would be the value of the (N/4)th placed item, the second quartile
(Q2) would be the value of the (N/2)th placed item, and the third quartile (Q3)
would be the value of the (3N/4)th placed item.

If x1 , x2 ,..., xN are the N values of a variable X, then mathematically, Q1, Q2 and


Q3 can be defined as 85
Block 1 Descriptive of Statistics
First quartile (Q1) = (N/4)th placed item in the arranged data.

Second quartile (Q2) = (N/2)th placed item in the arranged data.

Third quartile (Q3) = (3N/4)th placed item in the arranged data.

For Grouped Data


If x1 , x2 ,..., xn are the n values (or mid values in case of class intervals) of a
variable X, where f i be the frequency of xi , (i = 1, 2..., n) . Next, we prepare the
cumulative frequency distribution. After that, we will identify the ith quartile
class, which is the same as the case of the median.

Here, the ith quartile is denoted by Qi and is defined as

 iN 
 −F
 4  × c for i = 1, 2, 3.
Qi = l1 +
fq

Where l1 = lower boundary of ith quartile class;

N = total frequency;

F = cumulative frequency of the class prior to ith quartile class and

f q = frequency of ith quartile class;

c = width of ith quartile class.

 i× N 
Here ‘i’ denotes the ith quartile class in which   th observation falls in
 4 
cumulative frequency.

Note: The second quartile, i.e. Q2, is the median.

3.15.2 Deciles
Deciles divide the data into ten equal parts. This means that there is a total of
nine deciles. These nine deciles are denoted by D1, D2,…, D9, where D1
stands for the 1st Decile, D2 stands for 2nd Decile, and so on. Here, ith Decile
contains  ×  th part of the data. It may be noted that the data should be
i N
 10 
arranged in ascending or descending order of magnitude.

For Ungrouped Data


The data must first be arranged in either ascending or descending order of
magnitude. The first decile (D1), second decile (D2), …, and ninth decile (D9)
are then determined by locating the item in the organized data. The value of
the  i × N  th placed item would be the ith Decile. For i = 1, then it would be 1st
 10 
Decile, i = 2, it would be 2nd Decile, and so on.

If x1 , x2 ,..., xN are the N values of a variable X, then mathematically, then the ith
Decile is defined as

ith Decile (Di) =  i × N  th placed item in the arranged data (i = 1, 2, 3…, 9)


86  10 
Unit 3 Measures of Central Tendency
For Grouped Data
If x1 , x2 ,..., xN are the n values (or mid values for class intervals) of a variable
X where f i be the frequency of xi ,(i = 1, 2..., n) . Next, we prepare cumulative
frequency distribution. After that, we will identify the ith Decile class, the same
as in the case of quartiles.

Here, the ith Decile denoted by Di is defined as

 iN 
 −F
10
Di = l1 +   × c for i = 1, 2, 3,…., 9
fd

Where l1 = lower boundary of ith Decile class;

N = total frequency;

F = cumulative frequency of the class prior to ith Decile class and

f d = frequency of ith Decile class;

c = width of ith Decile class.

 i× N 
Here ‘i’ denotes the ith Decile class in which   th observation falls in
 10 
cumulative frequency.

Note: The fifth Decile, i.e. D5, is the median.

3.15.3 Percentiles
The data is divided into 100 equal parts using percentiles. It indicates that
there are 99 percentiles in total. The notation for these ninety-nine percentiles
is P1, P2,…, P99, where P1 represents the first percentile, P2 the second, and
 i× N 
so forth. Here ith percentile contains   th part of data. The data should
 100 
be arranged either in ascending or descending order of magnitude, it should
be noted.

For Ungrouped Data


The data should be arranged either in ascending or descending order of
 N   2N   99 N 
magnitude. Then we find the   th,   th,...,   th placed item in
 100   100   10 
the arrange data to find the 1st percentile (P1), 2nd percentile (P2),…,99th
 i× N 
percentile (P99), respectively. The value of the   th placed item would be
 100 
the ith percentile. For i = 1, then it would be 1st percentile, i = 2, it would be 2nd
 99 N 
percentile, and so on. When i = 99, then the value of   th placed item
 100 
would be the 99th percentile of that data.

If x1 , x2 ,..., xN are the N values of a variable X, then mathematically, the ith


percentile is defined as 87
Block 1 Descriptive of Statistics

ith percentile (Pi) = 


i× N 
 th placed item in the arranged data
 100 

(i = 1, 2, 3,…, 99)

For Grouped Data


If x1 , x2 ,..., xn are the n values (or mid values for class intervals) of a variable X,
where be the frequency of xi ,(i = 1, 2..., n) . Next, we prepare the cumulative
frequency distribution. After that, we will identify the ith percentile class, the
same as in the case of the median.

Here, the ith percentile is denoted by Pi and defined as

 iN 
 −F
Pi = li +   × c for i = 1, 2, 3,…., 99
100
fp

Where l1 = lower boundary of ith percentile class;

N = total frequency;

F = cumulative frequency of the class prior to ith percentile class

f p = frequency of ith percentile class and

c = width of ith percentile class.

 i× N 
Here ‘i’ denotes the ith percentile class where   th observation falls in
 100 
cumulative frequency.

Note: The fiftieth percentile, i.e. P50, is the median.

Ex.15. Determine the first and third quartiles for the data in Example 4.

Ans: First, we find the cumulative frequency given in the given cumulative
frequency table:

Class interval frequency Cumulative frequency


(less than type)
0−30 3 3
30−40 8 11
40−50 5 16
50−60 4 20
60−70 5 25
5

f
i =1
i = 25

In this case, N/4 = 25/4 = 6.25. The observation thus belongs to class 30–40.
Thus, it is the class in the first quartile. 3N/4 = 75/4 = 18.75 is likewise true. As
a result, the observation belongs to class 50−60. It is, therefore, the third
88 quartile class.
Unit 3 Measures of Central Tendency
For the first quartile, l1 = 30;

N = 25;

F=3

f p = 8 and

c = 10.

(6.25 − 3)
Q1 = 30 + × 10 = 34.06
8
For the third quartile, l1 = 50;

N = 25;

F = 16

f p = 4 and

c = 10.

(18.75 − 16)
Q3 = 50 + × 10 = 56.88
4

SAQ 12
a) What is meant by quartiles?

b) Determine the data's first and third quartiles in SAQ 5.

3.16 SUMMARY
In this unit, we have discussed:

• Professor Bowley defined the mean or average as a statistical constant


that enables us to understand the significance of the whole distribution.
It indicates the concentration of values in the middle portion of the
distribution, making it a useful representation of the entire distribution.

• Measures of central tendency should be precisely defined,


comprehensible, easily calculated, and not affected by sampling
fluctuations or extreme values. They should also be able to calculate for
open-end class intervals and lend themselves to algebraic treatment.
This ensures that the combined set of data is available for both sets,
avoiding missing information.

• Among several measures of central tendency are arithmetic mean,


median, mode, geometric mean, and harmonic mean.

• The arithmetic mean (AM) is calculated by dividing the sum of all


observations by the total number of observations and can be applied to
two types of data such as ungrouped and grouped data. 89
Block 1 Descriptive of Statistics

• The median is the middle value in an ordered distribution, representing a


point where half of the values lie below and above it.

• The mode is the value in a distribution that appears most frequently.

• The geometric mean (GM) is the nth root of nth observations, used for
averaging ratios and proportions. However, it fails to provide the correct
average when an observation is zero or negative.

• The harmonic mean is a type of average that is especially useful when


dealing with rates, ratios, or situations in which the average of
reciprocals is relevant. It is defined as the reciprocal of the arithmetic
mean of the reciprocals of a group of numbers.

• Partition values are variables that divide a distribution into equal parts,
typically in ascending or descending order of magnitude.

• Common partition values include quartiles, deciles, and percentiles.


Quartiles divide data into four parts, while deciles and percentiles divide
the distribution into ten and hundred parts, respectively.

3.17 TERMINAL QUESTIONS


1. What are the central tendency measures? Additionally, record the
central tendency properties.

2. Describe the formula for arithmetic mean for both grouped and
ungrouped data. Also, mention the merits and demerits of arithmetic
mean.

3. Describe the geometric and harmonic mean calculations for both


grouped and ungrouped data. Also, mention the merits and demerits of
arithmetic mean.

4. What are the quartiles? How would you calculate quartiles for both
grouped and ungrouped data?

3.18 ANSWERS
Self-Assessment Questions
1. a) Professor Bowley described the mean or average as "statistical
constant which enable us to comprehend in a single effort the
significance of the whole." It provides us with an idea of how
concentrated the values are in the middle portion of the
distribution. In simple words, an average or mean of a statistical
series is the value of the variable which acts as a good
representation of the entire distribution.

b) The following properties are important to be a good measure of


central tendency.

It needs to be precisely defined.

For a measure of central tendency to have a suitable


90 interpretation, it must be properly defined. In order for different
Unit 3 Measures of Central Tendency
people to calculate the average from the same figures and obtain
the same result, it should also have an algebraic formula. It ought
to be founded on every observation. It should be readily
comprehensible.

It must be readily comprehensible

An average should be easy to understand because we use


measures of central tendency to simplify the complexity of data;
otherwise, its use is likely to be very limited.

It is required to be easily calculated.

An average should be easy to understand and simple to compute


so that it can be used as widely as possible.

It should not be affected by sampling fluctuations.

A tool with sampling stability should be preferred. To put it another


way, we should anticipate obtaining roughly the same results if we
choose ten distinct groups of observations from the same
population and calculate the average for each group. Because of
the sampling fluctuation alone, there might not be much of a
difference.

It shouldn’t be affected by extreme values.

Every observation is thought to have an impact on the average's


value. The average cannot be regarded as a good average if it is
significantly impacted by one or two extremely small or extremely
large observations, either by increasing or decreasing its value.

It should be possible to calculate even for open-end


class intervals.

A measure of central tendency should be calculated for the data


with open-end classes.

It should be readily lend itself to algebraic treatment.

The algebraic manipulations should be attributed to a measure of


central tendency. One can find information about the combined set
even if there are two sets of data and the individual information is
available for both sets. In this case, something is missing.

2. a) To calculate the average lead concentration, we will add all the


observations and divide by the total number of observations.

So,

X =
x
n
112.5 + 140.5 + 162.3 + 170.8 + 181.9 + 201.4
= = 161.57 miligram per gram
6
91
Block 1 Descriptive of Statistics

b) We have the distribution as given below


Carbohydrates (in gm) No. of students fx
(x) (f)
30 3 90
45 7 315
50 2 100
3 3

 i =1
fi = 12 fx
i =1
i i = 505

Therefore, by using the formula

X=
 3
fx
i =1 i i
=
505
= 168.33 gm
 3
f
i =1 i 3

3. If we denote the protein intake (in gm.) by class interval, frequency by f


and mid value by x then we have the frequency distribution as:

Class interval Mid value x f fx


0−40 20 3 60
40−50 45 8 360
50−60 55 5 275
60−70 65 4 260
70−80 75 10 750
5 5

i =1
f i = 30 f x
i =1
i i = 1705

Therefore, by using the formula

X =
 5
i =1fi x i
=
1705
= 56.83 gm.
 5
i =1fi 30

4. a) The formula for calculating the arithmetic mean is defined as sum


of all observations divided by total number of observations.
Arithmetic mean (AM) may be calculated for two types of data
given as follows

• For ungrouped data


For ungrouped data, arithmetic mean may be computed by
applying any of the following methods:

1. Direct method
x
,
x
,
.
.
.
.
.
.
.
.
,
x

If are the n number of observations, then their


1

mean can be calculated mathematically as

x1 + x2 + ..... + xn
X =
92 n
Unit 3 Measures of Central Tendency
The above equation can also be written as

X =
 n
x
i =1 i

n
If f i is the frequency of xi , (i = 1, 2..., n) then the formula for mean
would be

f1 x1 + f 2 x2 + ... + f n xn
X=
f1 + f 2 + ... + f n

The above equation can also be written as

X=
 n
fx
i =1 i i

 n
i =1 if

• For grouped data


• Direct method
If f i is the frequency of xi ,(i = 1, 2..., n) where xi is the mid value of
the ith class interval, then the formula for the arithmetic mean will
be as follows:

f1 x1 + f 2 x2 + ... + f n xn
X=
f1 + f 2 + ... + f n

X=
 n
i =1 f i x i
=
 fx =  fx
 n
i =1 f i f N
where, N = f1 + f 2 + ... + f n

• Properties of arithmetic mean


The arithmetic mean satisfies all properties to become a good
average except the last two. It is particularly useful when dealing
with a sample as it is least affected by sampling fluctuations. It is
the most popular average and should always be our first choice
unless there is a lack of suitability.

Some algebraic properties of the arithmetic mean are given below:

Property 1. The algebraic sum of deviations of observations from


their mean is zero.

Property 2. The sum of squares of deviations taken from the


mean is the least in comparison to the same taken from any other
average.

Property 3. The arithmetic mean is affected by both the change of


origin and scale.

b) Merits of Arithmetic Mean

• It utilizes all the observations;

• It is rigidly defined; 93
Block 1 Descriptive of Statistics

• It is simple to comprehend and compute; and

• It can be used for further mathematical treatments.

Demerits of Arithmetic Mean


• It is badly affected by extremely small or extremely large
values;

• It can’t be evaluated for open-end class intervals; and

• It is generally not preferred for highly skewed distributions.

5. a) When we arrange n = 7 in ascending order, we obtain

2, 3, 6, 8, 10, 12, 15

Therefore

(n + 1) 8
Median = th observation = th = 4th observation
2 2
Therefore, Median = 8

b) When we arrange n = 8 in ascending order, we obtain

2, 3, 5, 6, 8, 10, 12, 15

Therefore,

n n 
  th observatio n +  + 1 th observatio n
Median =   2 
2
2

8 8 
  th observatio n +  + 1 th observatio n
Median =    
2 2
2
( + )
4
t
h
o
b
s
e
r
v
a
t
i
o
n
5
t
h
o
b
s
e
r
v
a
t
i
o
n

6
8

+
M
e
d
i
a
n

= = =
2

6. a) First of all, we construct the cumulative frequency distribution

Protein intake (gm) No. of players Cumulative


frequency
30 3 3
40 8 11
50 5 16
60 4 20
70 10 30
5

f
i =1
i = 30

Therefore, Median will be equal to the value of the variable


corresponding to the N th = 30 = 15 th cumulative frequency.
94 2 2
Unit 3 Measures of Central Tendency
Since 15 is not among cumulative frequencies, so, the next
cumulative frequency is 16, and the value of variable against 16 is
50. Therefore median is 50 gm.

b) First of all, we construct the cumulative frequency distribution

Protein intake No. of Cumulative


(gm) students frequency
0−20 8 8
20−40 19 27
40−60 30 57
60−80 x 57+x
80−100 12 69+x

The Median class is 40−60 since the median value is given as 50,

N 
 −F
2
Median = l1 +   ×c
fm

 69 + x 
 − 27 
2
50 = 40 +   × 20
30
(69 + x − 54)
50 − 40 =
3
30 = 15 + x
x = 30 − 15 = 15

Hence, the missing frequency is found to be 15.

7. The variable's value that splits the whole distribution in half is called the
median. The data should be arranged in either ascending or descending
order of magnitude. The median is the data's middle value when the
number of observations is odd. There will be two middle values for an
even number of observations. Therefore, we take these two middle
values and take their arithmetic mean. A number of the observations
below and above the median are the same. The median is not affected
by extremely large or extremely small values (as it corresponds to the
middle value), and it is also not affected by open-end class intervals.

Merits of Median

• It is rigidly defined;

• It is easy to understand and compute;

• It is not affected by extremely small or extremely large values; and

• It can be calculated even for open-end classes (like "less than 10"
or "50 and above"). 95
Block 1 Descriptive of Statistics

Demerits of Median

• In the case of even number of observations, we get only an


estimate of the median by taking the mean of the two middle
values. In this case we don't get its exact value;

• It does not utilize all the observations. The median of 1, 2 and 3 is


2. If observation 3 is replaced by any number higher than or equal
to 2 and if the number 1 is replaced by any number lower than or
equal to 2, the median value will be unaffected. This means 1 and
3 are not being utilized;

• It is not amenable to algebraic treatment; and

• It is affected by sampling fluctuations

8. First, we shall form the frequency distribution

Protein intake (gm) No. of students (f)


0−20 5
20−40 15
40−60 30
60−80 16
80−100 14

Here, the modal class is 40−60, corresponding to the highest frequency


of 30.

Therefore,

f 0 − f −1
Mode = l1 + ×c
2 f 0 − f −1 − f1

(30 − 15)
Mode = 40 + × 20 = 50.34 gm
(60 − 15 − 16)

9. We know that
Mode = 3 Median − 2 Mean

34.4 = 3 Median − 2( 41.6)

3 Median = 117.6

117.6
Median = = 39.2
3

10. To find the geometric mean we use the formula


1
GM = ( x1 x2 x3 ...xn ) n

96 GM = (141× 220 × 182 × 170) 4 = 176.01mmHg.


Unit 3 Measures of Central Tendency
11. We have the following distribution

Class Mid value (x) Frequency f/x


(f)
0−30 15 3 0.200
30−40 35 8 0.229
40−50 45 5 0.111
50−60 55 4 0.073
60−70 65 5 0.077
5 5
fi
f
i =1
i = 25 x
i =1
= 0.69
i

Therefore, by using the formula

N 25
HM = = = 36.23 gm.
fi 0.69
 5
i =1
xi

12. a) The entire distribution is divided into four equal parts by quartiles.
Those are first quartile (Q1), second quartile (Q2), third quartile
(Q3). Here, Q1 carries ¼ part of the data, Q2 carries ½ of the data,
Q3 carries ¾ part of the data. It may be noted here that the data
should be arranged in ascending or descending order of
magnitude.

For Ungrouped Data


We should arrange the data in either ascending or descending
order of magnitude in order to obtain the quartiles. The first quartile
(Q1), second quartile (Q2), and third quartile (Q3), are then
determined by finding the (N/4)th, (N/2)th, and (3N/4)th positioned
items in the organized data. Thus, the first quartile (Q1) would be
the value of the (N/4)th placed item, the second quartile (Q2) would
be the value of the (N/2)th placed item, and the third quartile (Q3)
would be the value of the (3N/4)th placed item.

If x1 , x2 ,..., xN are the N values of a variable X, then mathematically,


Q1, Q2 and Q3 can be defined as

First quartile (Q1) = (N/4)th placed item in the arranged data.

Second quartile (Q2) = (N/2)th placed item in the arranged data.

Third quartile (Q3) = (3N/4)th placed item in the arranged data.

For Grouped Data


If x1 , x2 ,..., xn are the n values (or mid values in case of class
intervals) of a variable X, where f i be the frequency
of xi , (i = 1, 2..., n) . Next, we prepare the cumulative frequency
distribution. After that, we will identify the ith quartile class, which is
the same as the case of the median. 97
Block 1 Descriptive of Statistics
Here, the ith quartile is denoted by Qi and is defined as

 iN 
 −F
4
Qi = l1 +   × c for i = 1, 2, 3.
fq

Where l1 = lower boundary of ith quartile class;

N = total frequency;
F = cumulative frequency of the class prior to ith quartile class and

f q = frequency of ith quartile class;

c = width of ith quartile class.


b) We will construct the cumulative frequency distribution.
Protein intake (gm) No. of students (f) Cumulative frequency
0−20 5 5
20−40 15 20
40−60 30 50
60−80 16 66
80−100 14 80
5

f
i =1
i = 80

Here, N/4 = 80/4 = 20. Therefore, the observation falls in the class
20-40. So it is the first quartile class. Similarly, 3(80)/4 = 240/4 =
60. Therefore, the observation falls in the class 60−80. So it is the
third quartile class.

For the first quartile, l1 = 20;

N = 80;
F = 5;
f q = 15 and

c = 20
(20 − 5)
Q1 = 20 + × 10 = 30 gm.
15
For the third quartile, l1 = 60;

N = 80;
F = 50
f q = 16 and

c = 20.
(60 − 50)
Q3 = 60 + × 10 = 66.25 gm.
16
98
Unit 3 Measures of Central Tendency

Terminal Questions
1. Refer to Section 3.2 and Subsection 3.2.1.

2. Refer to Sections 3.3 and 3.5 and Subsections 3.3.1 and 3.3.2.

3. Refer to Sections 3.11, 3.12, 3.13 and Subsections 3.11.1, 3.11.2,


3.12.1 and 3.12.2.

4. Refer to Section 3.15 and Subsections 3.15.1, [Link], [Link]

Suggested Readings
 Belle G. V., Fisher, L. D., Heagerty P.J. and Lumley, T. (2004).
Biostatistics: A methodology for the Health sciences, Hoboken, New
Jersey. : John Wiley and Sons.

 Forthofer, R. N., Lee E.S. and Hernandez, M. (2007). Biostatistics: A


guide to design, analysis and discovery, Burlington, USA. : Academic
Press.

 Le, Chap T. (2003). Introductory Biostatistics, Hoboken, New Jersey. :


John Wiley and Sons.

99

You might also like