Dr.
Ajit Patil
VARIABILITY
Meaning of Variability:
Variability means ‘Scatter’ or ‘Spread’.
Thus measures of variability refer to the scatter or spread of scores around
their central tendency.
The measures of variability indicate how the distribution scatter above and
below the central tender.
It shows how far data items lie from each other.
It shows the distance from the center of the distribution.
It also provides a descriptive analysis of the picture.
For example 1 Suppose, there are two groups. In one group there are 50 boys
and in another group 50 girls. A test is administered to both these groups. The
mean score of boys and is 54.4 and girls is we compare the mean score of both
the groups, we find that that there is no difference in the performance of the two
groups. But suppose the boys’ scores are found to range from 20 to 80 and the
girls’ scores range from 40 to 60.
This difference in range shows that the boys are more variable, because they cover
more territory than the girls. If the group contains individuals of widely differing
capacities, scores will be scattered from high to low, the range will be relatively
wide and variability becomes large.
Some definitions are:
“The scatter or variability of the observations of a distribution about some
measure of central tendency.”
Dr. Ajit Patil
“Dispersion is the spread of a distribution”
“Dispersion is the measure of the variation of the items.”
“Dispersion or spread is the degree of the scatter or variation of the
variables about a central value.”
Need of Variability:
Helps to as-certain the measures of deviation:
It helps to compare different group:
It is useful to supplement the information provided by the measures of
central tendency.
It is useful to calculate further advance statistics based on the measures of
dispersion.
There are two types of measures of variability
1. Absolute measures
2. Relatives measures
Absolute Measures of Variability
1. The Range
2. The Quartile Deviation
3. The Average Deviation
4. The Standard Deviation
Relative Measures of Variability
1. Coefficient of Range
2. Coefficient of Quartile Deviation
Dr. Ajit Patil
3. Coefficient of Average Deviation
4. Coefficient of Standard Deviation
Difference between Absolute and Relative Measures
Basis Absolute Relative
Definition When Dispersion is When it is the ratio of a
measured in original measure of absolute
units. dispersion.
Source Calculated from Raw Calculated from
Data absolute measure of
dispersion.
Ration & Percentages These cannot be These are generally
expressed in the Ratios expressed in terms of
& Percentage. ratios and percentage.
Uses This cannot be used for In those cases, relative
comparison when two dispersion is used for
data sets are expressed comparison.
in different units.
1. The Range:
Range is the difference between in a series.
It is the most general measure of spread or scatter.
It is a measure of variability of the varieties or observation among
themselves and does not given an idea about the spread of the observations
around some central value.
Dr. Ajit Patil
Range = H—L
Where, H = Highest score
L = Lowest score
Properties:
If the range is higher than the group indicates more heterogeneity.
if the range is lower than the group indicates more homogeneity.
Thus range provides us an instant and rough indication of the variability of
a distribution.
Merits of Range:
Range is easily calculated and readily understood.
It provides a quick estimate of the measure of variability.
It is the simplest measure of variability.
Demerits of Range:
Range is greatly affected by fluctuation of scores.
In case of open ended distributions range cannot be used.
It is affected greatly by fluctuations in sampling.
It is not based on all the observations of the series. It only takes the highest
and the lowest scores in to account.
The series is not truly represented by range. A symmetrical and A
symmetrical distribution may have same range but not the same dispersion.
It is affected greatly by extreme scores.
Uses of Range:
Range is used as a measure of dispersion when variations in the value of
the variable are not much.
Dr. Ajit Patil
Range is used when the knowledge of extreme score or total spread is
wanted.
Range is the best measure of variability when the data are too scattered or
too scant.
When a quick estimate of variability is required then, range is used.
Coefficient of Range
The ratio of difference between the highest and lowest value of frequency to the
sum of highest and lowest value of frequency.
Mathematically, = H – L / H+ L
Where, L = Lowest Score
H- Highest Score
Example 2:
In a class, 20 students have secured the marks as following:
16, 48, 47, 60, 45, 25, 17, 35, 35, 56, 51, 70, 25, 40, 22, 18, 35, 44, 52, 19
Here—The Highest score is 70
The Lowest score is 15
Range = H — L = 70 –15 = 55
Coefficient of Range = H – L / H + L
70-15 / 70 +15
= 55/ 85
0.647
Dr. Ajit Patil
2. The Quartile Deviation (Q):
Quartile deviation is second important measure of variability.
It is based upon the interval containing the middle fifty percent of cases in
a given distribution.
One quarter means 1/4th of something, when a scale is divided in to four
equal parts. “The quartile deviation or Q is the one-half the scale
distance between the 75th percentile and 25th percentiles in a frequency
distribution.”
From the figure we see that the 1st quartile or Q1 is position in a
distribution below which 25% cases, and above which 75% cases lie.
The 2nd quartile or Q2 is a position below and above which 50% cases lie.
It is the Median of the distribution.
The 3rd quartile or Q3 is the 75th percentile, below which 75% cases and
above which 25% cases lie.
So the quartile deviation (Q) is one half the scale distances between the 3rd
quartile (Q3) and the 1st quartile (Q1).
It is also known as the Semi-Interquartile Range i.e.
(Q3 – Q1) / 2
While, Quartile Range = Q3 – Q1
Dr. Ajit Patil
Symbolically:
Therefore, order to compute quartile deviation first of all we have to compute 1st
quartile (Q1) and the 3rd quartile (Q3)
Where = L = Lower limit of the 1st quartile class,
The 1st quartile class is that class; whose cumulative frequency is greater than the
value of N/4 when it is calculated from lower end.
N/4 = One fourth of the total number of cases.
F = Cumulative frequency of the class interval below the 1st quartile class.
Fq = The frequency of the Q1 class
W = Size of the class interval 3N
Where: L = Lower limit of the 3rd quartile class
The 3rd quartile class is that class whose cumulative frequency (Cf) is greater
than the value of 3N/4 i.e. Cf > 3N/4, when Cf is calculated from lower end.
3N/4= ¾ th of N or 75% of the total number of cases.
F = Cumulative frequency of the class below the class.
Dr. Ajit Patil
fq= The frequency of the Q3 class.
i = Size of the class interval.
Merits of quartile deviation:
It is simple to calculate and easy to understand.
It is trust worthy and more representative than range.
It is used in studying measures of dispersion, In case of open ended class
intervals
It is a good index of score density at the middle of the distribution.
When we take Median as the measure of central tendency at that time Q is
preferred as the measure of dispersion.
Like range it is not affected by extreme scores.
Demerits of quartile deviation:
It ignores the first 25% and the last 25% of the scores. Thus, It is not based
on all the observations of data.
Further algebraic treatment is not possible in case of Q. It is only a
positional average. It does not study variation of the values of a variable
from any average. It merely indicates a distance on a scale.
Its value is affected in any case, by a change in the value of a single score.
It is affected by fluctuation of scores.
Q is not a suitable measure of dispersion, when in a series there is a
considerable variation in the values of various scores.
Uses of quartile deviation:
When Median is the measure of central tendency at that time Q is used is
used as the measure of dispersion.
Dr. Ajit Patil
When our primary interest is to know the concentration around the median-
the middle 50% of cases, at that time Q is used.
When extreme scores affect S.D. or the scores are scattered at that time Q
is used as measure of variability.
When the class intervals are open ended, Q is used as measure of
dispersion.
Coefficient of Quartile Deviation
A relative measure of dispersion based on the quartile deviation is called
the coefficient of quartile deviation.
It is pure number free of any units of measurement.
It can be used for comparing the dispersion in two or more than two sets of
data.
The Formula is:
= (Q3 – Q1) / (Q3 + Q1)
Example 3: Consider a data set of following numbers: 22, 12, 14, 7, 18, 16, 11,
15, 12. You are required to calculate the Quartile Deviation.
Solution:
First, we need to arrange data in ascending order to find Q3 and Q1 and avoid
any duplicates.
7, 11, 12, 13, 14, 15, 16, 18, 22
Calculation of Q1 can be done as follows,
Q1 = ¼ (9 + 1)
=¼ (10)
Dr. Ajit Patil
Q1=2.5 Term
Calculation of Q3 can be done as follows,
Q3=¾ (9 + 1)
=¾ (10)
Q3= 7.5 Term
Calculation of quartile deviation can be done as follows,
Q1 is an average of 2nd and 3rd term so, (11 + 12) / 2 = 11.50.
Q3 is an average of 7th and 8rth term so, (16+18)/ 2 = 17.
Q.D. = Q3 – Q1 / 2
Using the quartile deviation formula, we have (17-11.50) / 2
=5.5/2
Q.D.=2.75.
Coefficient of Quartile Deviation = (Q3 – Q1) + (Q3 + Q1)
= 17-11.50 / 17+11.50
= 5.5 / 28.5
= 19.29
Interquartile Range = (Q3 – Q1)
= 17 - 11.50
= 5.5
Dr. Ajit Patil
3. The Average Deviation / Mean Deviation (A.D.):
Range and quartile deviation have been used in the initial parts. But none
of these dispersion indicate about the composition of the distribution. It is
because both the dispersions do not take all individual scores into account.
The serious shortcomings of range and quartile deviation can be overcome
by using another dispersion called Average deviation.
Also known as Mean Deviation.
“Average deviation is the arithmetic mean of all the deviations of
different scores from the mean value of the scores without the regard
for sign of the deviation.”
Thus average deviation is arithmetic mean of the deviations of a series
computed from any measure of central tendency.
So average deviation is the mean of the deviations taken from their mean
(Sometimes from Median and Mode.)
Definitions:
Collins Dictionary of Statistics:
“Average deviation is the mean of the absolute values of the differences between
the values of a variable and the mean of its distribution.”
Dictionary of Education, C.V. Good:
“A measure expressing the average amount by which the individual items in a
distribution deviate from a measure of central tendency such as mean of median.”
H.E. Garrett:
“The average deviation or AD is the mean of the deviations of all the separate
scores in a series taken from their mean (occasionally from the Median or
Mode).”
Thus it can be said that average deviation or Mean deviation as it is called
is the mean of the deviations of all the scores.
Dr. Ajit Patil
No account is taken of signs and all deviations whether +ve or —ve treated
as positive.
1) Individual Series: The formula to find the Mean Deviation for an individual
series is:
=
∑ = Summation
X = Observation / Values
µ = Mean
N = Number of observations
2) Discrete Series: The formula of Mean Deviation from Mean for a discrete
series is:
∑ = Summation
X = Observation / Values
µ = Mean
f = frequency of observations
3) Continuous Series: The formula to find the Mean Deviation for a discrete
series is:
∑ = Summation
X = Mid-Value of the class
µ = Mean
Dr. Ajit Patil
f = frequency of observations
Mean Deviation from Median
1) Individual Series: The formula to find the Mean Deviation for an individual
series is:
MD = ∑∣X−M∣ / N
Where,
∑ = Summation
X = Mid-Value of the class
M = median
N = Number of observations
2) Discrete Series: The formula to find the Mean Deviation for a discrete series
is:
MD = ∑f∣X−M∣ / ∑f
Where,
∑ = Summation
X = Observation / Values
M = median
f = frequency of observations
3) Continuous Series: The formula to find the Mean Deviation for a continuous
series is:
MD= ∑f∣X−M∣ /∑f
Where,
Dr. Ajit Patil
∑ = Summation
X = Mid-Value of the class
M = Median
f = frequency of observations
Mean Deviation from Mode
1) Individual Series: The formula to find the Mean Deviation from Mode for an
individual series is:
MD= ∑∣X−Mode∣ / N
Where,
∑ = Summation
X = Observation / Values
M = Mode
N = Number of observations
2) Discrete Series: The formula to find the Mean Deviation from Mode for a
discrete series is:
MD=∑f∣X−Mode∣ / ∑f
Where,
∑ = Summation
X = Observation / Values
M = Mode
f = frequency of observations
Dr. Ajit Patil
3) Continuous Series: The formula to find the Mean Deviation from Mode for a
continuous series is:
MD= ∑f ∣X−Mode∣ / ∑f
∑ = Summation
X = Mid-Value of the class
M = Mode
f = frequency of observations
Merits of A.D.:
Average deviation is rigidly defined and its value is precise and
definite.
It is easy to calculate.
It is based on all the observations.
It is easy to understand. Because it is the average of the deviations
from a measure of central tendency.
It is less affected by the value of extreme scores.
Demerits of A.D.:
The most serious drawback with average deviation is that it ignores
the algebraic signs of the deviations which is against the
fundamental rules of mathematics.
It is very rarely used. Because of standard deviation is generally used
as a measure of dispersion.
Further algebraic treatment is not possible in case of AD.
When calculated from mode AD does not give accurate measure of
dispersion.
Dr. Ajit Patil
Uses of Average Deviation:
Average deviation is used when it is desired to weight all the
deviations from the mean according to their size.
AD is used when we want to know the extent to which the measures
are spread out either side of the mean.
When extreme scores influence standard deviation at that time AD
is the best measure of dispersion.
Coefficient of the Mean Deviation
A relative measure of dispersion based on the mean deviation is called
the coefficient of the mean deviation or the coefficient of dispersion.
It is defined as the ratio of the mean deviation of the average used in
the calculation of the mean deviation. Thus:
Coefficient of M.D (about mean) = Mean Deviation from Mean / Mean
Coefficient of M.D (about median) = Mean Deviation from Median /
Median
Coefficient of M.D (about mode) = Mean Deviation from Mode / Mode
Example 4: Calculate the Mean Deviation about the Mean using the following
Data
6, 7, 10, 12, 13, 4, 8, 12.
Solution: First we have to find the Mean of the Data that we are provided with
Mean of the given data=
Sum of all the terms / total number of terms
=
Dr. Ajit Patil
(6+7+10+12+13+4+8+12) / 8
72 / 8
=9
Next, we have to find the Mean Deviation
Xi xi−x¯ |xi−x¯|
6 6-9=-3 |-3|=3
7 7-9=-2 |-2|=2
10 10-9=1 |1|=1
12 12-9=3 |3|=3
13 13-9=4 |4|=4
4 4-8=-5 |-5|=5
8 8-9=-1 |-1|=1
12 12-9=3 |3|=3
=22
Mean deviation about mean=
∑∣Xi−Mean∣ / 8
22/8
=2.75
4. The Standard Deviation (SD):
We have discussed three measures of variability namely, Range, Quartile
Deviation and Average Deviation. We also found that all of them suffer
from serious drawbacks.
Dr. Ajit Patil
The range just taken in to account only the highest score and the lowest
score. The quartile deviation takes into account only the middle 50% of
scores and in case of average deviation we ignore the signs.
Therefore, in order to overcome all these difficulties we use another
measure of dispersion called Standard Deviation. It is commonly used in
experimental research as it is the most stable index of variability.
Its symbol is σ (Greek small letter sigma).
Definitions:
“Standard deviation is a measure of spread or dispersion. It is root
mean squared deviation.”
“A widely used measure of variability, consisting of the square root of the
mean of the squared deviations of scores from the mean of the
distribution.”
Standard deviation is the square root of the average value of the
squared deviations of the scores from their arithmetical mean.
The SD is computed by summing the squared deviation of each measure
from the mean, divided by the number of cases and extracting the square
root.
To be more clear, we should note here that in computing the SD we square
all the deviations separately, find their sum, divide the sum by the total
number of scores and then find the square root of the mean of the squared
deviation.
So that it is also called the ‘root mean square deviation’.
The square of standard deviation is called as Variance (σ2). It is referred to
as the mean square deviation.
It is also called as the second moment dispersion
Dr. Ajit Patil
Formula of Ungrouped data:
In grouped data SD can be calculated in two methods:
1. Direct method or Long method
2. Short method or Assumed Mean method
Properties of Standard Deviation:
If a constant value is added to each score or subtracted from each score,
the value of S.D. remains unchanged.
If a constant value is multiplied or divided to the original scores, the value
of S.D. is also multiplied or divided by the same number
Merits of S.D:
Standard deviation is rigidly defined and its value is always definite.
It is based on all the observations of data.
It is capable of further algebraic treatment and possesses many
mathematical properties.
Unlike Q and AD it is less affected by fluctuations of scores.
Dr. Ajit Patil
Unlike AD, it does not ignore the negative signs. By squaring of deviations
it overcomes these difficulties.
It is the reliable and most accurate measure of variability. It always goes
with the mean which is the most stable measure of central tendency.
S.D. gives a measure that is comparable meaning from one test to other.
Above all the normal curve units are expressed in a unit.
Demerits of S.D:
S.D. is difficult to understand and not easy to calculate.
S.D. gives more weight to extreme scores and loss to those which are nearer
to the mean. It is because the squares of the deviations, which are big in
size, would be proportionately greater than the squares of those deviations
which are comparatively small.
Uses of S.D:
S.D. is used when our thrust is to measure the variability having greatest
stability.
When extreme deviations might affect the variability at that time S.D. is
used.
S.D. is used for calculating the further statistics like coefficient of
correlation, standard scores, standard errors, Analysis of Variance,
Analysis of Co-variance etc.
When the interpretation of scores is made in terms of the NPC, S.D is used.
When we want to determine the reliability and validity of test scores, S.D.
is used.
Coefficient of standard deviation
The standard deviation is the absolute measure of dispersion.
Dr. Ajit Patil
Its relative measure is called the standard coefficient of dispersion or
coefficient of standard deviation. It is defined as:
Coefficient of Standard Deviation = Standard Deviation / Mean
The most important of all the relative measures of dispersion is the coefficient
of variation. This word is variation not variance. The coefficient of
variation (C.V) is defined as:
Coefficient of Variation = (Standard Deviation / Mean) x 100
Variance = Square of standard Deviation
Example 5:
Find out the SD of the following distribution:
Solution:
Step-1:
Find out the midpoint of each class interval.
Dr. Ajit Patil
Step-2:
Find out the mean of the distribution:
Here M= ∑fx/N = 3540/50
= 70.80
Step -3:
Find out the deviation (x) by deducting the mean from points.
Step -4:
Find out the fx by multiplying the f (col-2) with x (col-5)
Step-5:
Find out the fx by multiplying fx (col- 2) with x (col-5)
Step-6:
Compute ∑ fx by adding the values in col-7.
Step-7:
Put the values in formula.
Dr. Ajit Patil
2. Short Method or Assumed Mean Method:
In short method, calculation of SD is easy and less time consuming.
If the mid-points of the class intervals are decimal numbers then it
becomes more complicated to compute S.D in long method.
This method consists essentially in ‘guessing’ or assuming a mean and
later applying a correction to give actual mean. So that it is called as
assumed mean method.
Example 6 : Compute the S.D, of the following distribution:
Dr. Ajit Patil
Solution:
Step-1:
Assume the mid-point of any class interval as ‘Assumed Mean’. But it is better
to assume the mid-point of the class interval in the middle having highest
frequency as assumed mean. Here let is assume = 72 as Assumed mean.
Step-2:
Find out x (deviation of the scores from the assumed mean) as shown in col-3.
x’ = X – M/i
Step-3:
Compute fx’, by multiplying x’ with f (col-4).
Step-4:
Compute fx2 by multiplying x’ (col-3) with fx (col-5).
Step-5:
Find out ∑ fx’ and ∑fx’2 it’ by adding the values in col-4 and col-5 respectively.
Step-6:
Put the values in formula:
Dr. Ajit Patil
Example 7: Calculate the coefficient of standard deviation and coefficient of
variation for the following sample data: 2, 4, 8, 6, 10, and 12.
Solution:
X (X– Mean )2
22 (2–7)2=25(2–7)2=25
44 (4–7)2=9(4–7)2=9
88 (8–7)2=1(8–7)2=1
66 (6–7)2=1(6–7)2=1
1010 (10–7)2=9(10–7)2=9
1212 (12–7)2=25(12–7)2=25
∑X=42 ∑(X– Mean) 2=70
Mean =∑X / n
=42 / 6
=7
Dr. Ajit Patil
S = √ (∑(X– Mean)2 / n)
S= √ (70 / 6)
S = √ (35 / 3)
S =3.42
Coefficient of Standard Deviation = S / Mean
=3.42 / 7
=0.49
Coefficient of Variation
(C.V)= (S / Mean) x 100
=3.42 / 7×100
=48.86%
Formulas at a glance