0% found this document useful (0 votes)
2 views14 pages

Chapter Four

Chapter Four discusses measures of dispersion, which quantify the spread of data around an average value. It outlines the objectives of measuring variation, types of measures including absolute and relative measures, and specific calculations for range, mean deviation, variance, and standard deviation. The chapter emphasizes the importance of these measures in statistical analysis and comparison of data distributions.

Uploaded by

wondebelete92
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views14 pages

Chapter Four

Chapter Four discusses measures of dispersion, which quantify the spread of data around an average value. It outlines the objectives of measuring variation, types of measures including absolute and relative measures, and specific calculations for range, mean deviation, variance, and standard deviation. The chapter emphasizes the importance of these measures in statistical analysis and comparison of data distributions.

Uploaded by

wondebelete92
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Chapter Four

Measures of Dispersion (Variation) and Shape

1.1 Introduction and Objectives of Measuring Variation

-The scatter or spread of items of a distribution is known as dispersion or variation. In other


words the degree to which numerical data tend to spread about an average value is called
dispersion or variation of the data.
-Measures of dispersions are statistical measures which provide ways of measuring the extent in
which data are dispersed or spread out.
Objectives of measuring Variation:
 To judge the reliability of measures of central tendency
 To control variability itself.
 To compare two or more groups of numbers in terms of their variability.
 To make further statistical analysis.

Properties of a good measure of dispersion


A good measure of dispersion should:
- be rigidly defined by a mathematical formula,
- be simple to understand and easy to calculate,
- be unique,
- be fundamental of all observations in the series,
- not be affected by some extreme values existing in the series,
- have sampling stability property, and
- be capable of further algebraic treatment as well as further statistical analysis.

1.2 Absolute and Relative Measures of Dispersion

Absolute and Relative Measures of Dispersion

Absolute measures of dispersion are expressed in terms of the original unit of a series.
Such measures are not suitable for comparing the variability of two distributions which
are expressed in different units of measurement and different average size.
Example: Range, standard deviation, mean deviation
Relative measures of dispersions are a ratio or percentage of a measure of absolute
dispersion to an appropriate measure of central tendency and are thus pure numbers
independent of the units of measurement. For comparing the variability of two
distributions (even if they are measured in the same unit), we compute the relative
measure of dispersion instead of absolute measures of dispersion.
Example: Relative Range, coefficient of variation, coefficient of mean deviation

1.3 Types of Measures of Dispersion


1.3.1 The Range and Relative Range

Range : The range is the largest score minus the smallest score. It is a quick and dirty measure of
variability, although when a test is given back to students they very often wish to know the range
of scores. Because the range is greatly affected by extreme scores, it may give a distorted picture
of the scores. The following two distributions have the same range, 13, yet appear to differ
greatly in the amount of variability.

Distribution 1: 32 35 36 36 37 38 40 42 42 43 43 45
Distribution 2: 32 32 33 33 33 34 34 34 34 34 35 45

For this reason, among others, the range is not the most important measure of variability.

R  LS , L  l arg est observation


S  smallest observation
Range for grouped data:
If data are given in the shape of continuous frequency distribution, the range is computed as:
R  UCLk  LCL1 , UCLk is upperclass lim it of the last class.
UCL1 is lower class lim it of the first class.
This is sometimes expressed as:
R  X k  X1 , X k is class mark of the last class.
X 1 is classmark of the first class.
Relative Range (RR) for ungrouped data -it is also sometimes called coefficient of range and
given by:
LS R
RR  
LS LS
Relative Range (RR) for grouped data
UCLk  LCL1
RR  , UCLk is upperclass lim it of the last class.
UCLk  LCL1
UCL1 is lower class lim it of the first class.
Properties of Range and Relative Range
- Range and relative range are easy to calculate and simple to understand.
- Both cannot be computed for grouped data with open ended classes.
- They do not tell us anything about the distribution of values in the series.

Example 1: Find the range and relative range for the monthly salary of ten workers in a certain paint
factory given below.

462 480 534 624 498 552 606 588 516 570

Solution:
xmax  624 birr xmin  462 birr
R  xmax  xmin  624 birr  462 birr  162 birr
xmax  xmin 624 birr  462 birr 162 birr
RR     0.149
xmax  xmin 624 birr  462 birr 1086 birr

Example 2: Find the values of the range and relative range for the following frequency distribution: which
shows the distribution of the maximum loads supported by a certain number of cables.
Maximum load Number
(in kilo-Newton) of cables
93 – 97 2
98 – 102 5
103 – 107 12
108 – 112 17
113 – 117 14
118 – 122 6
123 – 127 3
128 – 132 1
Solution:
M first  95 kN M last  130 kN
R  M last  M first  130 kN  95 kN  35 kN
M last  M first 130 kN  95 kN 35 kN
RR     0.156
M last  M first 130 kN  95 kN 225 kN
Example 3: If the range and relative range of a series are 4 and 0.25 respectively. Then what is the
value of:
a) Smallest observation
b) Largest observation
Solutions :( 2)
R  4  L  S  4 _________________(1)
RR  0.25  L  S  16 _____________(2)
Solving (1) and ( 2) at the same time, one can obtainthe following value
L  10 and S  6

1.3.2 The Mean Deviation and Coefficient of Mean Deviation

The mean deviation (MD) measures the average deviation of a set of observations about their central
value, generally the mean or the median, ignoring the plus/minus sign of the deviations.

The mean deviation of a sample of n observations x1 , x2 , ... , xn is given as

MD 
x i A
Where A is a central measure (the mean or the median)
n
In case of grouped data, the formula for MD becomes
MD 
f i xi  A
Where x i is the class mark of the i th class, f i is the frequency of the i th
n
class and n  f i .

 The mean deviation about the arithmetic mean is, therefore, given by

MD 
x i x
.... for ungrouped data
n

MD 
f i xi  x
.... for grouped frequency distribution; where x i is the class mark of the i th
n
class, f i is the frequency of the i th class and n  f i

 The mean deviation about the median is also given by

MD 
x i ~
x
.... for ungrouped data
n

MD 
 f i xi  ~x .... for grouped frequency distribution; where x is the class mark of the i th
i
n
class, f i is the frequency of the i th class and n   f i .

The coefficient of mean deviation (CMD) is the ratio of the mean deviation of the observations to their
appropriate measure of central tendency: the arithmetic mean or the median.
MD
In general, CMD  where A is a measure of central tendency: the arithmetic mean or the median.
A
MD
That is, CMD about the arithmetic mean is given by CMD  where MD is the mean deviation
x
calculated about the arithmetic mean. On the other hand CMD about the median is given by
MD
CMD  ~ in which case MD is calculated about the median of the observations.
x

Properties of Mean Deviation and coefficient of mean deviation


- It is easy to understand and compute
- It is based on all observations
- It is not affected very much by the of extreme value(s).
- It is not capable of further mathematical treatments and it is not a very accurate measure of
dispersion.

1.3.3 The Variance, the Standard Deviation and Coefficient of Variation

The Variance
Variance is the arithmetic mean of the square of the deviation of observations from their arithmetic mean.
Population Variance
If we divide the variation by the number of values in the population, we get something called the
population variance. This variance is the "average squared deviation from the mean".
1
Population Varince   2   ( X i   ) 2 , i  1,2,..... N
N

For the case of grouped frequency distribution it is expressed as:

1
Population Varince   2   f i ( X i   ) 2 , i  1,2,.....k
N

Sample Variance

One would expect the sample variance to simply be the population variance with the population mean
replaced by the sample mean. However, one of the major uses of statistics is to estimate the
corresponding parameter. This formula has the problem that the estimated value isn't the same as the
parameter. To counteract this, the sum of the squares of the deviations is divided by one less than the
sample size.

1
Sample Varince  S 2   ( X i  X ) 2 , i  1,2,....., n
n 1

For the case of grouped frequency distribution it is expressed as:

1
Sample Varince  S 2   f i ( X i  X ) 2 , i  1,2,.....k
n 1

We usually use the following short cut formula.

n
 X i  nX 2
2

S2  i 1
, for raw data.
n 1

k
 f i X i  nX 2
2

S2  i 1
, for frequency distribution.
n 1

Standard Deviation

There is a problem with variances. Recall that the deviations were squared. That means that the units were
also squared. To get the units back the same as the original data values, the square root must be taken.

Population s tan dard deviation     2


Samples tan dard deviation  s  S 2

 The following steps are used to calculate the sample variance:

1. Find the arithmetic mean.


2. Find the difference between each observation and the mean.
3. Square these differences.
4. Sum the squared differences.
5. Since the data is a sample, divide the number (from step 4 above) by the number of
observations minus one, i.e., n-1 (where n is equal to the number of observations in the data set).

Examples:
1. Find the variance and standard deviation of the following sample data 5, 17, 12, 10.
2. The data is given in the form of frequency distribution.

Class Frequency
40-44 7
45-49 10
50-54 22
55-59 15
60-64 12
65-69 6
70-74 3

Solutions:
1. X  11

Xi 5 10 12 17 Total
(Xi- X )2 36 1 1 36 74
n
 ( X i  X )2 74
 S2  i 1
  24.67.
n 1 3
S  S2  24.67  4.97.

2. X  55
Xi(C.M) 42 47 52 57 62 67 72 Total
(Xi- X )2 169 64 4 4 49 144 289

fi(Xi- X ) 2 1183 640 198 60 588 864 867 4400

n
 fi ( X i  X )2 4400
 S2  i 1
  59.46.
n 1 74
S  S 2
 59.46  7.71.

Coefficient of Variation
 The standard deviation is an absolute measure of dispersion. The corresponding relative measure
is known as the coefficient of variation (CV).
 Coefficient of variation is used in such problems where we want to compare the variability of two
or more than two different series.
 Coefficient of variation is the ratio of the standard deviation to the arithmetic mean, usually
expressed in percent.
S
CV   100 . Where S is the standard deviation of the observations.
x
A distribution having less coefficient of variation is said to be less variable or more consistent or more
uniform or more homogeneous.

Example: Last semester, the students of Biology and Chemistry Departments took Stat 273 course. At the
end of the semester, the following information was recorded.

Department Biology Chemistry


Mean score 79 64
Standard deviation 23 11

Compare the relative dispersions of the two departments’ scores using the appropriate way.

Solution:
Biology Department Chemistry Department
S S
CV   100 CV   100
x x
23 11
  100  29.11%   100  17.19%
79 64

Interpretation: Since the CV of Biology Department students is greater than that of Chemistry Department
students, we can say that there is more dispersion relative to the mean in the distribution of Biology
students’ scores compared with that of Chemistry students.

Properties of the Variance and the Standard Deviation


Variance
– It removes most of the demerits or drawbacks of the measures of dispersion discussed so far.
– Its unit is the square of the unit of measurement of values. For example, if the variable is measured in
kg, the unit of variance is kg2.
– It is calculated based on all the observations/data in the series.
– It gives more weight to extreme values and less to those which are near to the mean.

Standard Deviation
– It is considered to be the best measure of dispersion.
– [Demerits] If the values of two series have different unit of measurement, then we can not compare
their variability just by comparing the values of their respective standard deviations.
– It is calculated based on all the observations/data in the series. Standard deviation is capable of further
algebraic treatment.
– Standard deviation is as such neither easy to calculate nor to understand.
– Similar to the variance, standard deviation gives more weight to extreme values and less to those which
are near to the mean.

The Standard Scores (Z-Scores)


 A standard score is a measure that describes the relative position of a single score in the entire
distribution of scores in terms of the mean and standard deviation.
 It also gives us the number of standard deviations a particular observation lie above or below the
mean.
x
Population standard score: Z  where x is the value of the observation,  and  are the mean

and standard deviation of the population respectively.
xx
Sample standard score: Z  where x is the value of the observation, x and S are the mean and
S
standard deviation of the sample respectively.

Interpretation:

Example:1. Two sections were given an exam in a course. The average score was 72 with standard
deviation of 6 for section 1 and 85 with standard deviation of 5 for section 2. Student A from section 1
scored 84 and student B from section 2 scored 90. Who performed better relative to his/her group?
Solution: Section 1: x = 72, S = 6 and score of student A from Section 1; x A = 84
Section 2: x = 85, S = 5 and score of student B from Section 2; x B = 90

x A  x1 84  72
Z-score of student A: Z    2.00
S1 6
x B  x 2 90  85
Z-score of student B: Z    1.00
S2 5
From these two standard scores, we can conclude that student A has performed better relative to his/her
section students because his/her score is two standard deviations above the mean score of selection 1
while the score of student B is only one standard deviation above the mean score of section 2 students.

Example: 2. Two groups of people were trained to perform a certain task and tested to find out
which group is faster to learn the task. For the two groups the following information was given:

Value Group one Group two

Mean 10.4 min 11.9 min

[Link]. 1.2 min 1.3 min

Relatively speaking:
a) Which group is more consistent in its performance
b) Suppose a person A from group one take 9.2 minutes while person B
from Group two take 9.3 minutes, who was faster in performing the
task? Why?

Solutions:
a) Use coefficient of variation.
S 1.2
C.V1  1 *100  *100  11.54%
X1 10.4
S2 1.3
C.V2  *100  *100  10.92%
X2 11.9
Since C.V2 < C.V1, group 2 is more consistent.
b) Calculate the standard score of A and B

X A  X1 9.2  10.4
ZA    1
S1 1.2
XB  X2 9.3  11.9
ZB    2
S2 1.3
Child B is faster because the time taken by child B is two standard deviation shorter than the
average time taken by group 2 while, the time taken by child A is only one standard deviation
shorter than the average time taken by group 1.

Measures of Shape
• Measure of shape describes the pattern or distribution of data within dataset.
• It indicates how variations is distributed about the locations, that is, it tells about if
variation is symmetric about certain centeral value or it is skewed or if possibly
multimodal.
• The most common types of measures of shape are:
i. Moments
ii. Skewness and
iii. Kurtosis

Moments
- It measures the distribution of data around the origin, mean or any other value A.
moment about origin for n observations is given by;
X  X 2  ...  X n
k k k
m  1
k

n
n

X
k
i
 i 1
n

- For the case of frequency distribution this is expressed as:


n

f
k
i Xi
m  k i 1

n
- If r  1 , it is the simple arithmetic mean, this is called the first moment.
2. The kth moment about the mean (the kth central moment) for observation is
given by;

- Denoted by Mr and defined as:


n

(X  X ) i
k

m 
k i 1
n
- For the case of frequency distribution this is expressed as:
n

f i ( X i  X )k
mk  i 1
n

[Link] rth moment about any number A is defined as:


n

(X i  A) k
mk  i 1

n
- For the case of frequency distribution this is expressed as:
n

f i ( X i  A) k
mk  i 1

Example:
1. Find the first two moments for the following set of numbers 2, 3, 7
2. Find the first three central moments of the numbers in problem 1
3. Find the third moment about the number 3 of the numbers in problem 1.
Solutions:

1. Use the rth moment formula.


n

X
k
i
M k
 i 1

n
2  3 7
 M1  4 X
3
2 2  32  7 2
M2   20 .67
3
2. Use the rth central moment formula.
n

 fi ( X i  X )k
M k
 i 1
n
( 2  4)  (3  4)  (7  4)
 M 1
 0
3
( 2  4) 2  (3  4) 2  (7  4) 2
M 2
  4.67
3
( 2  4) 3  (3  4) 3  (7  4) 3
M3  6
3
3. Use the rth moment about A.
n

(X i  A) k
M k
 i 1

n
( 2  3) 3  (3  3) 3  (7  3) 3
M2   21
3

Skewness
- Skewness is the degree of asymmetry or departure from symmetry of a distribution.
- A skewed frequency distribution is one that is not symmetrical.
- Skewness is concerned with the shape of the curve not size.
- If the frequency curve (smoothed frequency polygon) of a distribution has a longer tail to the
right of the central maximum than to the left, the distribution is said to be skewed to the right
or said to have positive skewness.
- If it has a longer tail to the left of the central maximum than to the right, it is said to be skewed
to the left or said to have negative skewness.
- For moderately skewed distribution, the following relation holds among the three commonly
used measures of central tendency.
Mean  Mode  3 * ( Mean  Median)
Measures of Skewness
- Denoted by  3
- There are various measures of skewness.
1. The Pearsonian coefficient of skewness
Mean  Mode X  Xˆ
3  
S tan dard deviation S
2. The Bowley’s coefficient of skewness ( coefficient of skewness
based on quartiles)
(Q3  Q2 )  (Q2  Q1 ) Q3  Q1  2Q2
3  
Q3  Q1 Q3  Q1
3. The moment coefficient of skewness
M3 M M
3   2 33 2  33 , Where  is the populations tan dard deviation.
M2
32
( ) 
The shape of the curve is determined by the value of  3
 If  3  0 then the distribution is positively skewed .
 If  3  0 then the distribution is symmetric.
 If  3  0 then the distribution is negatively skewed .
Remark:
o In a positively skewed distribution, smaller observations are more frequent
than larger observations. i.e. the majority of the observations have a value
below an average.
o In a negatively skewed distribution, smaller observations are less frequent
than larger observations. i.e. the majority of the observations have a value
above an average.

Examples:
1. Suppose the mean, the mode, and the standard deviation of a certain distribution are 32,
30.5 and 10 respectively. What is the shape of the curve representing the distribution?
Solutions:
Use the Pearsonian coefficient of skewness
Mean  Mode 32  30.5
3    0.15
S tan dard deviation 10
 3  0  The distribution is positively skewed .
2. In a frequency distribution, the coefficient of skewness based on the quartiles is given to
be 0.5. If the sum of the upper and lower quartile is 28 and the median is 11, find the
values of the upper and lower quartiles.
Solutions:
Given: ~
 3  0.5, X  Q2  11
Q1  Q3  28 .......... .......... .......(*)
Required: Q1 , Q3

(Q3  Q2 )  (Q2  Q1 ) Q3  Q1  2Q2


3    0.5
Q3  Q1 Q3  Q1
Substituting the given values , one can obtainthe following
Q3  Q1  12.......... .......... .......... .....(**)
Solving (*) and (**) at the same time we obtain the following values
Q1  8 and Q3  20

3. Some characteristics of annually family income distribution (in Birr) in two regions is as
follows:
Region Mean Median Standard Deviation
A 6250 5100 960
B 6980 5500 940
a) Calculate coefficient of skewness for each region
b) For which region is, the income distribution more skewed. Give your interpretation for
this Region
c) For which region is the income more consistent?

Solutions: (exercise)

4. For a moderately skewed frequency distribution, the mean is 10 and the median is 8.5. If
the coefficient of variation is 20%, find the Pearsonian coefficient of skewness and the
probable mode of the distribution. (exercise)
5. The sum of fifteen observations, whose mode is 8, was found to be 150 with coefficient
of variation of 20%
(a) Calculate the pearsonian coefficient of skewness and give appropriate conclusion.
(b) Are smaller values more or less frequent than bigger values for this distribution?
(c) If a constant k was added on each observation, what will be the new pearsonian
coefficient of skewness? Show your steps. What do you conclude from this?
Solutions: (exercise)

Kurtosis

 Kurtosis is the degree of peakdness of a distribution, usally taken relative to a normal


distribution.
 A distribution having relatively high peak is called leptokurtic.
 If a curve representing a distribution is flat topped, it is called platykurtic.
 The normal distribution which is not very high peaked or flat topped is called mesokurtic.
Measures of kurtosis
The moment coefficient of kurtosis:
 Denoted by  4 and given by
M4 M
4   44
M2
2

Where : M 4 is the fourth moment about the mean.
M 2 is the sec ond moment about the mean.
 is the populations tan dard deviation.
The peakdness depends on the value of  4 .
 If  4  3 then the curve is leptokurti c.(morepeaked )
 If  4  3 then the curve is mesokurtic .( Normalcurv e)
 If  4  3 then the curve is platykurti c(lesspeaked )
Examples:
1. If the first four central moments of a distribution are:
M 1  0, M 2  16, M 3  60, M 4  162
a) Compute a measure of skewness
b) Compute a measure of kurtosis and give your interpretation.

Solutions:
M3  60
3    0.94  0
a) M2
32
16 3 2
 The distribution is negatively skewed .
M4 162
b) 4  2
  0.6  3
M2 16 2
 The curve is platykurtic.
2. The median and the mode of a mesokurtic distribution are 32 and 34 respectively. The 4th
moment about the mean is 243. Compute the Pearsonian coefficient of skewness and identify the
type of skewness. Assume (n-1 = n).

3. If the standard deviation of a symmetric distribution is 10, what should be the value of the fourth
moment so that the distribution is mesokurtic?
Solutions (exercise).

Exercises
1. Consider the marks of 20 students out of 20% in biology test as follows

Marks of Students’ 0-5 5-10 10-15 15-20 Total

Number of students 2 6 8 4 20
Find
i. Range
ii. Quartile deviation
iii. Mean and median deviation
iv. Variance and standard deviation
2. The final exam of a course consists of two exams: mathematics and History. If a student scored
66 in Mathematics and 80 in History. However, all students’ average score is 51 with a standard
deviation of 12 in mathematics and 72 with the standard deviation 16 in history.
a. In which subject a student had better performance?
b. In which subject all students have similar (consistent) results?

You might also like