Measures of Dispersion Explained
Measures of Dispersion Explained
Measures of Dispersion
Introduction:
We have studied methods of finding the average or a measure of
central tendency of a frequency distribution and of locating certain
values, such as median, decile and percentiles. However, it is not
enough to compare averages. The extent and nature of the spread of
the values around the measure of central tendency is to be
ascertained. Thus it may be said that an average becomes meaningful
in the real sense only when it is accompanied by a measure of
dispersion. The study of dispersion is useful when-
1. it is necessary to know the structure of the frequency
distributions whose averages are similar.
2. it is necessary to know the structure of the frequency
distributions, whose formations are alike but which have
different means.
3. we want to have a comprehensive idea of the formation of a
series, it is necessary to know its dispersion.
1
Following are the characteristic of an Ideal Measure of Dispersion:
1. It should be rigidly defined.
2. It should be easy to calculate and easy to understand.
3. It should be based on all the observations.
4. It should be amenable to further algebraic treatments.
5. It should be less affected by sampling fluctuation.
6. It should have sampling stability.
Application of Measures or Dispersion:
1. Compare the characteristics between two or more
observations or information.
2. To determine the reliability of an average;
3. To serve as a basic for the control of the variation;
4. To compare two or more series with regard to their variation;
5. To facilitate the use of other statistical measures.
Measures of Dispersion
Following are the measures of dispersion:
a) Absolute measures
(i) Range. (ii) Quartile deviation
(iii) Standard deviation (iv) Mean deviation
b) Relative measures
(i) Co- efficient of Range (ii) Co- efficient of
quartile deviation
(ii) Co- efficient of Variation (iv) Co- efficient of
Mean deviation
Discussion:
RANGE
Range is the difference between the highest and the lowest
observations in a set. If x1, x2, x3.........., xn are the values of n
observations in a sample. Generally it is denoted by R.
Symbolically, Range,
R = max(x1, x2, x3.........., xn) - min (x1, x2, x3.........., xn).
On the other hand the range is define by R = L – S, where L =
Largest value, and S = Smallest value.
A relative measure known as coefficient of range is given as,
2
LS
Co- efficient of Range, C. R = L S 100.
Lesser the range or coefficient of range, better the result.
Properties:
(i) It is the simplest measure and can easily be understood.
(ii) Besides the above merit, it hardly satisfies any property of a good
measure of dispersion e.g. it is based on two extreme values only,
ignoring the orders. It is not liable to further algebraic treatment.
Example-1: The following are the prices of shares of a company
from Monday to Saturday:
Day Price (Taka)
Monday 200
Tuesday 210
Wednesday 208
Thursday 160
Friday 220
Saturday 250
Calculate range and coefficient of range.
Solution-1: We known that
Range, R = L – S
Where, L = Largest value and S = Smallest value
= 250 taka and = 160 taka
R = 250-160 = 90 taka
Co-efficient of range,
LS 250 160
100 100
LS 250 160
90
100 0.219 100 21.90
C. R = 410
This is the required result of our given information.
Limitation of range
1. Range is not based on each and all observation of a
distribution.
2. It is subject to fluctuations of considerable magnitude from
sample to sample.
3. Range cannot be determined in case of open-end
distributions.
4. Range cannot tell us anything about the character of the
distribution within two extreme observations.
Application of the range
There are the different applications of the range. We can describe
the application of range in the following:
1. Range can use in the Quality control;
2. It is used in the share prices;
3. It is used in the Weather forecasts and etc.
Interpretation of the range, R
4
The R is no more than a rough measure of dispersion. It gives a
comprehensive value for the data in the sense that it includes the
limits within which all of the items occurred. The range can be
interpreted as an intensive measure of variability except in very small
samples.
Advantages of QD
1. The QD is based on the middle 50% of a distribution and is
complementary to the median.
2. The QD is easy to compute and easy to understand.
3. QD is not affected by the extreme values.
4. QD is superior to the range as a rough measure of dispersion.
Limitations of QD
1. Quartile deviation ignores 50% items, i.e., the first 25% and
the last 25%.
2. The value of QD does not depend upon every observation it
cannot be regarded as a good method of measuring variation.
3. QD is not capable of mathematical manipulation.
4. Its value is very much affected by sampling fluctuations.
5
5. It is in fact not a measure of dispersion as it really does not
show the scatter around as average but rather a distance on a
scale,.i.e.,
Example 3. The values of Q1, Q2 and Q3 are worked out for a
company are Q1=174.90, Q2 = 190.23 and Q3 =203.83. Find the
quartile deviation and Co-efficient of quartile deviation
.
Mean deviation:
Mean deviation or Average deviation is obtained by calculating the
absolute deviations of each observation from median (or mean), and
then averaging these deviations by taking their arithmetic mean.
Definition:
Mean deviation is the average of the absolute deviations taken from
a central value, generally the mean or median. It is denoted by M.D
MD for raw or ungrouped data:
Let x1, x2, ........,xn are n observations and their arithmetic mean x ,
then the mean deviation is MD is defined by;
x n
Xi X
i 1
M.D ( ) = n , when taken from the mean
n
Xi Me
i 1
M.D (Me) = n , when taken from the median.
Computation of Mean Deviation (ungrouped data):
6
In the deviation method of measuring scatter, the following steps
are to be taken:
1. Compute the mean or median of the observations.i.e.,
X or Me.
2. Find the deviations of each observation from the
mean. i.e., . or
xi x xi Me
1 xi Me
1
xi x
3. Compute MD = n or n .
MD for grouped data:
Let x1,x2,...........,xn denote the mid-values of n classes and f1,f2,.....,fn
are the corresponding frequencies, then the MD is defined by;
x n
fi Xi X
i 1
M.D ( ) = N , when mean deviation from mean.
Where, fi is the class frequency and xi be the mid-value for every
class, N is the total frequency.
n
fi Xi Me
i 1
M.D (Me) = N , when mean deviation from median.
The relative measure corresponding to the MD, called the coefficient
of MD, is obtained, by dividing MD by the particular mean used in
computing MD. Thus, if MD has been computed from median, the
coefficient of MD shall be obtained by dividing MD by the median
or mean.
Coefficient of mean deviation,
MD
100
[Link] = X when, MD from mean.
If median has been used while calculating the value of mean
deviation, in such a case coefficient of mean deviation shall be
obtained by dividing mean deviation by the median. A co-efficient of
mean deviation;
MD
100
[Link] = median when, MD from median.
7
Computation of Mean Deviation (for grouped data)
Following steps are involved in calculating the mean deviation from
mean or median:
1. Median or mean of the series is calculated .i.e., X or Me.
2. Deviations of the items from median or mean are ascertained
ignoring plus and minus signs.
3. Deviations computed above are multiplied by the respective
n n
fi xi x . fi xi Me .
frequencies. i.e., i 1 or i 1
n n
fi xi x . fi xi Me .
4. i 1 or i 1 is divided by the number of items.
The quotient obtained shall be the value of mean deviation.
Properties:
1. Mean deviation removes one main objection of the earlier
measures, that it involves each value of the set.
2. It is not affected much by extreme values.
3. It has no relationship with any of the other measures of
dispersion.
4. Its main drawback is that algebraic negative signs of the
deviations are ignored which mathematically unsound.
5. Mean deviation is minimum when the derivations are taken
from median.
6. Mean deviation is independent in origin but depend on scale.
Interpretation of the mean deviation:
1. The mean deviation may help us to find the percentage of
observations falling in a range of Average mean deviation.
8
3. If the MD is comparatively small, then more than half of the
items in the data fall within a small range around the average.
This concentration would mean compactness of the
distribution.
Application of the MD:
1. The application of the MD is overshadowed to a large extent
by the use of the standard deviation (SD).But the computation
of the MD is less difficult.
2. For purpose of interpreting the significance of a series of
ratios of an item it is the most valuable. Because of its
simplicity in meaning and computation, it is especially
effective in reports presented to the general public or to
groups not familiar with statistical methods.
Advantage of mean deviation:
1. The outstanding advantage of the MD is its relative
simplicity.
2. It is simple to understand and easy to compute.
3. It is based on each and every observation of the data.
Consequently change in the value of any observation would
change the value of average deviation.
4. Mean deviation is less affected by the values of extreme
observation.
5. Deviation are taken from a central value, comparison about
formation of different distributions can easily ve made.
Limitations of mean deviation:
1. The greatest drawback of this method is that algebraic signs
are ignored while taking the deviations of the items.
2. This method may not give us accurate results. The reason is
that mean deviation gives us best results when deviations are
taken from median. But median is not a satisfactory measure
when the degree of variability in a series is very high.
3. It is not capable of further algebraic treatment.
4. It is rarely used in sociological and business studies.
9
It is the ratio of the mean deviation and arithmetic mean product into
100, we can defined as
MD
100
Coefficient of mean deviation, [Link] = X
Example 4.
Calculate mean deviation and coefficient of mean deviation taken the
from mean from the following data:
Sales(in 15-19 19-23 23-27 27-31 31-35 35-39
thousand $)
No. of days 8 59 47 23 6 4
Problem 5.
Calculate mean deviation from mean from the following data:
Sales(in 10 - 20 20 - 30 30 - 40 40 - 50 50 - 60
thousand $)
No. of days 3 6 11 3 2
Also calculate the co-efficient of mean deviation.
Solution: We want to calculate Mean deviation from the mean.
Calculation table of mean deviation
Sales (in Mid- frequency fi xi xi x xi x
thousand value( (fi) fi
$) xi)
10 -20 15 3 45 18 36
20 -30 25 6 150 08 48
30 - 40 35 11 385 02 22
40 - 50 45 3 135 12 36
50 - 60 55 2 110 22 44
Total N = 25 825 186
x n
fi Xi X
i 1
.D ( ) = N , when mean deviation from mean.
Where, fi = class frequency, Xi = Mid-value; X = Mean of X and
N = Total no. of observations.
11
X 825
33
Where, = 25
Mean deviation about mean by the formula is
x fi xi x 186 $7.44
25
M.D ( ) = N =
Coefficient of mean deviation;
M .D 7.44
100 100 0.2255 100 22.55
C.M.D = X = 33
Thus the mean sales are $33 thousand per day and the mean
deviation of sales is $ 22.55 % thousand.
Please read HSC 2nd Paper book for more clearance
VARIANCE: Recommended: Ketab Uddin's Book (HSC 2nd Paper)
Definition: The variance is the average of the squares of the
deviations taken from mean i.e., sum of the square deviation and
dividing by the number of observations is called variance.
1 n n X
n X i
2
= { i 1 Xi2 - ( i 1 )2/n}; where is the population mean.
The sample variance of the set x1,x2, .........,xn of n observations is
given by the formula,
1 n x
S2 = n 1 i 1 (x - )2
i
1 n
1 n
xi 2 xi 2
n 1 i 1 n i 1
= for i = 1, 2, ........., n.
x 1 n
n
xi
where = i 1 .
12
Variance for grouped data:
Let x1,x2,...........,xn denote the mid-values of n classes and f1,f2,.....,fn
are the corresponding frequencies and the variance is defined by
1 n X
2
= N i 1 fi (Xi - )2 for i = 1, 2, .......n. Where, fi is the class
frequency and xi be the mid-value for every class, N is the total
X 1 n n
N f i Xi fi
frequency, = i 1 and N = i 1
1 n 1 n
fiXi 2 fiXi 2
N i 1 N i 1
2
=
If the observation xi occurs fi times for i = 1, 2, ........., n, then the
sample variance,
1 n x
S2 = n 1 i 1 f (x - )2
i i
1 n 1 n
fixi 2 fixi 2
n 1 i 1 n i 1
= for i = 1, 2, ........., n.
x 1 n n
n
fixi fi.
where = i 1 . and n = i 1
Properties of variance:
1. The variance has mostly removed the lacunae which are
present in the measures of dispersion given before it.
2. The main disadvantage of variance is, that its unit is square of
the unit of measurement of variate values. For clarity, say, the
variable X is measured in ms, the unit of variance is m2.
Generally this value is large and makes it difficult to decide
about the magnitude of variation.
3. The variance gives more weightage to the extreme values as
compared to those which are near to mean value, because the
difference is squared in variance.
13
Theorem: Prove that the variance is independent in origin but
depend on scale.
Proof: Let us consider x1, x2, .............,xn are n observations and their
mean is denoted by x . The variance is denoted by Sx2 and defined by
2
1 n
xi x
Sx2 = n i 1 ..............................(i)
xi a
Let us consider ui = c be a new variate, where a is origin and
xi a
c is scale. We have ui = c
xi a cui
xi a cui
x a cu
Putting the value in (i), we have
2
1 n
xi x
Sx2 = n i 1
2
1 n
a cu i a cu
n i 1
2
1 n
cui cu
n i 1
2
ui u
n
2 1
c .
n i 1
c 2 .S u 2
Therefore, the variance is independent in origin but depend on scale.
(Proved)
Example: Find the variance for the first n natural numbers.
Solution: We know that, set of the first n natural numbers
x: {1, 2, 3, ...............,n}.
14
So, the variance of x is defined by
2 2
1 n
n
n
2
xi x xi xi
n i 1 i 1
i 1
n n
Sx2 = =
2
x 2 x 2 2 .......... x n 2 x1 x 2 .......... x n
1
n n
2
12 2 2 ........ n 2 1 2 .......... n
n n
2
n(n 1)(2n 1) n(n 1)
6n 2n
(n 1)(2n 1) (n 1) 2
6 4
(n 1) 2n 1 n 1
2 3 2
(n 1) 4n 2 3n 3
2 6
(n 1) (n 1)
2 6
n2 1
12
So, the variance of the first n natural numbers.
STANDARD DEVIATION:
16
There is another method of summing up deviations from measure of
central tendency and finding a measure of dispersion of the data. It is
standard deviation. Standard deviation considered superior to other
measures of dispersion because of its advantages in mathematically
representing the variability, which is very important for interpreting
statistically data.
Properties
1. Standard deviation is considered to be the best measure of
dispersion and is used widely.
2. There is however one difficulty with it. If the unit of
measurement of variances of two series is not the same, then
their variability can not be compared by comparing the values
of standard deviation.
3. Standard deviation must be a positive quantity.
4. It is based on all the observations and is readily understood.
5. It is more difficult to compute than the other measures of
dispersion but is easy to use mathematically.
6. It has relatively a small sampling error.
7. Its main advantage is that it is amenable to algebraic
treatment and comparatively stable under sampling
fluctuations.
8. Standard deviation is independent of change of origin but not
scale.
17
9. The variance is the minimum of all mean squared deviations
(MSD) and the SD is the minimum of all root mean squared
deviation (RMSD).
10. If X and S denote the mean and standard deviation,
respectively of n non-negative quantities x1, x2,........,xn then,
X n 1 S
11. For a symmetrical distribution, the following area
relationships hold good.
Mean 1 covers 68.27% observations.
Mean 2 covers 95.45% observations.
Mean 3 covers 99.73% observations.
Advantage of Standard deviation
1. The standard deviation is the best measure of the dispersion.
2. It is possible to calculate the combine standard deviation of
two or more groups.
3. For comparing the variability of two or more distributions
coefficient of variation is considered to be most appropriate
and this measure is based on mean and standard deviation.
4. SD is most prominently used in further statistical work.
2
1 n
xi x
n i 1
SD = ..............................(i)
xi a
Let us consider ui = c be a new variate, where a is origin and
xi a
c is scale. We have ui = c
xi a cui
xi a cui
x a cu
Putting the value in (i), we have
2
1 n
a cui a cu
n i 1
SD =
1 n
cui cu
n i 1
2
1 2 n
c ui u
n i 1
2
c
1 n
ui u
n i 1
2
c. SD( x)
Therefore, the standard deviation is independent in origin but depend
on scale.
(Proved)
19
Theorem: If the n term of the positive observation and their
arithmetic mean X and standard deviation S then show that
X n 1 S .
Proof: Let us consider x1, x2, ……..,xn are n term positive
observations and their arithmetic mean X and standard deviation S;
then we have
n n
X xi n x xi
i 1 i 1 and
2
n
n
xi xi
2
i 1
i 1
n n
S=
2
n
n
x i
xi
2
S
2 i 1
i 1 2
n n
2
n
n
xi
nS x i in 1
2 2
i 1 n
Problem5: Find the variance and standard deviation from the weekly
wages of ten workers working in a factory:
Workers Weekly wages (TK.)
A 1320
B 1310
C 1315
D 1322
E 1326
F 1340
G 1325
H 1321
I 1320
20
J 1331
1 n
n
xi x
Solution: We know that the variance, 2
and2
= i 1
Standard deviation, S.D. = Variance = 2 = .
Calculations of Variance and Standard deviation by using the above
information;
Workers Weekly
wages (xi)
(xi- x ) (xi- x )2
A 1320 -3 9
B 1310 -13 169
C 1315 -8 64
D 1322 -1 1
E 1326 +3 9
F 1340 +17 289
G 1325 +2 4
H 1321 -2 4
I 1320 -3 9
J 1331 +8 64
n =10
xi =13230 (xi- x ) (xi- x )2 = 622
We have,
X xi 13230
1323
= n 10 taka.
1 n
n
xi x 2
Variance, 2
= i 1 from table we get,
622
= 10 = 62.20 taka.
Hence, the variance for our given information is 62.20 taka. So, the
22
X fixi 4435 36.96
= N 120 .
1 n
N
fi xi x 2
variance, 2
= i 1 from table we get,
5444.60
= 120 = 45.37.
Hence, the variance for our given information is 45.37. So, the
standard deviation S.D. = Variance = 2 = = 45.37 =6.74.
COEFFICIENT OF VARIATION (C.V):
Coefficient of variation is more useful when the two distributions are
entirely different and the units of measurement are also different.
When the relative variation is stated in terms of the arithmetic mean
and the standard deviation, the resulting percentage is known as the
coefficient of variation or coefficient of variability.
Definition: Coefficient of variation of a series of variate values is
the ratio of the standard deviation to the mean multiplied by 100.
If is the standard deviation and X is the mean of the set of values,
the coefficient of variation is,
100
C.V = X ;
Coefficient of variation is of great practical significance and is the
best measures of comparing the variability of two series or two
groups. It shows whether the items included in a series are consistent
or not as well be evident from the examples solved hereafter. The
series or groups for which the coefficient of variation is greater is
said to be more variable or less consistent. On the other hand, the
23
series for which the coefficient of variation is less is said to be less
variables or more consistent.
[This measure was given by Professor Karl Pearson.] By using the
above Example 6 information, we can calculate the coefficient of
variation. We have,
X
100
C.V = X ; Here, = 6.74 and =36.96, then
6.74
100
= 36.96
=18.24 (Answer).
Example: A sample of six items is taken from the production of a
company. Length and weight of the six items are given below:
Length (inches): 3 8 10 12 14 15
Weight (ounces): 9 11 12 14 16 17
Find out co-efficient of variation for length and weight.
100
Solution: We know that C.V = X ,
1 n n
xi x
n i 1
X
xi
i 1
Where 2
= 2
= n , n = no. of observations and
2 is the variance & is the standard deviation.
Now, we want to construct a calculation table by using given data.
Calculation table:
Length (xi) Weight
(inches) (xi - x ) (xi - x )2 (yi) (yi - y ) (yi - y )2
ounces
3 -7.33 53.729 9 -4.170 17.389
8 -2.33 5.429 11 -2.170 4.709
10 -0.33 0.109 12 -1.170 1.369
24
12 1.67 2.789 14 0.830 0.689
14 3.67 13.469 16 2.830 8.009
15 4.67 21.809 17 3.830 14.669
6 6 x 6 6 y
xi yi
i 1 62 i 1 (xi - )2 i 1 =79 i 1 (yi- )2
X = 97.334 y =13.17 =46.834
=10.33
For length:
X n
xi
i 1
So, = n , n = no. of observations = 6
62
= 6 = 10.33
1 n
xi x 2
n i 1
The variance of length is x
2
= , from the calculation
1 97.334 16.222
table we can get, x
2
= 6 and the standard
deviation, x = 16.222 = 4.028
Therefore the coefficient of variation for length is
x 4.028
100 100
C.V = x = 10.33 = 38.99
For weight:
y n
yi
i 1
So, = n , n = no. of observations = 6
25
79
= 6 = 13.17
1 n
yi y 2
n i 1
The variance of length is y
2
= , from the calculation
1 46.834 7.806
table we can get, y
2
= 6 and the standard
7.806 =2.794
deviation, y =
Therefore the coefficient of variation for length is
y 2.794
100 100
C.V = y = 13.17 = 21.22
Comment: The coefficient of variation for length is greater than the
coefficient of variation for weight. So, the length of variation is less
consistent than the variation of weight.
Shares - X Shares-Y
xi (xi - x ) (xi - x ) 2 yi yi - y (yi - y )2
26
318 -3.38 11.424 2545 -1.5 2.25
322 0.62 0.384 2524 -22.5 506.25
325 3.62 13.104 2544 -2.5 6.25
320 -1.38 1.904 2533 -13.5 182.25
324 2.62 6.864 2565 18.5 342.25
315 -6.38 40.704 2535 -11.5 132.25
318 -3.38 11.424 2567 20.5 420.25
329 7.62 58.064 2559 12.5 156.25
8 8 x 8 8 y
xi 2571 yi
i 1 (xi - )2
i 1 i 1 (yi - )2
i 1
= 143.872 20372 =1748.25
For Shares Y:
Y n
yi
i 1
Mean of Y; = n , n = no. of observations = 8
27
20372
2546.5
= 8
1748.25
1 n
n
yi
y 8
218.531
Variance of X; y = i 1
2 2
=
Standard deviation, y = 218.531 = 14.783
So, the coefficients of variation for shares Y,
y
100
C.V(X) = Y
14.783
100
= 2546.5 = 0.581
Comment: Since it is evident that prices of shares for company Y
are much larger than those for company X, the two standard
deviations are not directly comparable. As a share X has more
coefficient of variation, so, it shows greater variability or shares Y is
more stable. Thus the coefficient of variation can be employed for
comparing the relative consistency of the prices of shares of two or
more companies. This will help a genuine investor in selecting share
the price of which is relatively stable. Thus shares which are more
consistent in the fluctuation of prices will be preferred by him.
X N1X 1 N 2 X 2
= N1 N 2 ;
Then combined variance of the two groups is given by the formula,
28
N 1 12 X 1 X } N 2 22 X 2 X }
2
2
2= N1 N 2
2= N1 N 2
29
Again the combined standard deviation, =
Combine var iance = 455.21 =21.34 (Answer).
2= N1 N 2
1
= 45 65 [45{82 +(85-82.045)2} + 65{132 +(80-82.045)2]
(45 72.732) (65 173.182)
= 110
3272.940 11256.830 14529.770
132.089
= 110 110
30
Example 8: In two factories A and B engaged in the same industry,
the average monthly wages and standard deviations are as follows:
Factory Average Monthly S.D of No. of
wage Wages (TK) Wages (TK) Earners
A 4600 500 100
B 4900 400 80
(i) Which factory A or B pays larger amount as monthly wages?
(ii) Which factory shows greater variability in the distribution of
wages?
(iii) What is the mean and standard deviation of all the workers in
two factories taken together?
Example: The first of two samples has 100 items with mean 15 and
standard deviation 3. If the whole group has 250 items with mean
15.6 and standard deviation 3.666, find the standard deviation of the
second group.
31
Solution: Let us consider X 1 and X 2 be two mean of the first and
second samples respectively,
X N1X 1 N 2 X 2
= N1 N 2
Where, X 1 = 15 and X 2 = ?
N1 = 100 N2 = 150
1 = 3 2 =?
= Combined standard deviation = 3.666
= N1 N 2 ]
(100 9) 100(15 15.6)2 15 2 2 150(16 15.6)2
3.666 = 250
32
13.44 250 = 960 + 150 σ22
150 σ22 = 3359.889 - 960
2400
16
σ2 = 150
2
σ2 = 4 .
Thus standard deviation of the second sample is 4.
33