STA111 Lecture Note Updated
STA111 Lecture Note Updated
STA111
Descriptive Statistics
(3 UNITS)
Department of
Statistics
University of Abuja
1
Course Aim and Objectives
This course teaches Descriptive Statistics required by 100L Students. On
successful completion of the course, you are expected to have adequate
knowledge of the course contents listed below:
COURSE Contents:
● Definitions and classifications of Statistics
● Data, methods of data collection, Classification and Tabulation
● Frequency distributions (grouped and ungrouped)
● Diagrammatic and Graphic Representations of data
● Averages (Measures of Central Tendency), Measures of Dispersion,
and Partition
● Skewness and Kurtosis
● Errors and Approximations
● Rates and Ratio
● Index Numbers
Texts
This course has detailed lecture notes; it should not be necessary to buy
a book for this course. However, you may consult any elementary books
in Statistics.
Lecture notes:
The course attempts to convey a large amount of information in a short space
of time. Some of the material is of a technical nature and may not†be covered
2
explicitly in the lectures and classes. You are expected to read the lecture
notes thoroughly. The syllabus is defined by the contents of the lecture notes
Additional reading is suggested at the end of each week’s classes. Despite my
best efforts, there will be mistakes in the notes. If you spot something that
looks wrong, please let me know. Solutions to some problems have been
provided while some have been intentionally omitted. All problems will be
solved together in the lecture rooms.
Definitions of Statistics
Statistics has been defined differently by different writers from time to time so
much so that scholarly articles have collected together hundreds of
definitions, emphasizing precisely the meaning, scope and limitations of the
subject. The reasons for such definitions may be broadly classified as follows:
i. The field of utility of statistics has been increasing steadily and thus
different people defined it differently according to the development of
the subject. In old days statistics was regarded as the “science of state
craft” but today it embraces almost every sphere of natural and human
activity. Accordingly, the old definitions which were confined to and very
limited and narrow field of enquiry were replaced by the new definitions
which are more exhaustive and elaborate in approach.
ii. The word statistics has been used to convey different meanings in
singular and plural sense. When use as plural statics means numerical
set of data and when used singular (statistic) sense it means the
science of statistical methods embodying the theory and techniques
use for collecting analyzing and drawing inferences from numerical
data.
3
1. “Statistics are the classified facts representing the conditions of the
people in a state. Specially those facts which can be stated in numbers
or in tables of numbers or in any tabular or classified arrangement.” -
Webster
2. “Statistics are numerical statements of facts in any departments of
enquiry place in relation to each other.” – Bowley.
3. By statistics we mean quantitative data affected to a marked extent by
multiplication of causes.” – Yule and Kendall.
4. Statistics may be defined as the aggregate of facts affected for a
marked extent by multiplicity of causes, numerically expressed
enumerated or estimated according to a reasonable standard of
accuracy, collected in a systematic manner, for a predetermined
purpose and placed in relation to each other”. – Prof. Horace Secrist.
Classification of Statistics
It has become accepted in today’s world that in order to learn about
something, you must first collect data. Statistics is the art of learning data. It is
concerned with the collection of data, its subsequent descriptions, and its
analysis, which often leads to the drawing of conclusion. At the end, the data
should be described for instance the scores of two groups of teaching
methods should be presented. In addition, summary measures such as the
average scores of members of each of the groups should be presented. This
part of statics, concerned with the description and summarization of data, is
called descriptive statistics. After the preceding experiment is completed and
the data are described and summarized, we hope to be able to draw
conclusion about which teaching method is superior. This part of statistics.
Concerned with the drawing of conclusion, is called inferential statistics.
Importance of Statistics
To a very striking degree our culture has become a statistical culture. Even a
person who may never have heard of an index number is affected by those
index numbers, which describe the cost of living. It is impossible to understand
4
psychology, sociology, economics, finance or natural and physical sciences
without some general idea of the meaning of an average, of variation, of
concomitance, of sampling, of how to interpret charts and tables. Statistics is
important in all area of human endeavors. For example, we have: statistics in
planning, statistics in state, statistics in mathematics, statistics in physics,
statistics in chemistry, in Biology, in Economics, in industry, in insurance in
astronomy, in psychology, in education, in war, in medical science, among
others, the list of area where statistics is important is endless.
Collection of Data
Statistics are set of numerical data. In fact only numerical data constitute
statistics. This means that the phenomenon under study must be capable of
quantitative measurement. Thus, the raw material of statistics always
originates from the operation of counting (enumeration), or measurements.
For any statistical enquiry, whether it is in business, economic, or Sciences, the
basic problem is to collect facts and figures relating to particular phenomenon
under study. The items in which the measurements are taken are called
statistical units. On the face of it, it might appear that the collection of data is
the first step for any statistical investigation. But in a scientifically prepared
(efficient and well-plane) statistical enquiry, the collection of data is by no
means the first step. Before we embark upon the collection of data for any
given statistics enquiry, it is imperative to examine carefully the following
points which may be termed as preliminaries to data collection: objectives and
scope of the enquiry, statistical units to be used, sources of information (data),
method of data collection, degree of accuracy aimed at in the final result, type
of enquiry. A good data however, possessed the following characteristics: it
should be unambiguous, it should be specific, it should be uniform, it should be
stable, and it should be appropriate.
Methods of Data Collection
Primary and Secondary Data
The data which are originally collected by an investigators or agency for the
first time for any statistical investigation and used by them in the statistical
analysis are termed as primary data. On the other hand, the data (published or
unpublished), which have already been collected and processed by some
5
agency or person and taken over from there and used by any other agency or
person for their statistical work are termed as secondary data as far as
second agency is concerned. The second agency if and when it publishes and
files such data becomes the secondary source to anyone who latter uses
these data. In other words, secondary source is the agency who publishes or
releases for use by others the data which was not originally collected and
processed by it. It may be observed that the distinction between primary and
secondary data is a matter of degree or relativity only. The same set of data
may be secondary in the hands of one and primary in the hands of others. In
general, the data are primary to the source who collects and processes them
for the first time and are secondary for all sources who latter use such data.
The methods commonly use for the collection of primary data are as follows:
direct personal investigation. Indirect oral interviews. Information received
through local agencies. Mailed questionnaire method. Schedules sent through
enumerators. The chief source of secondary data may be broadly classified
into the following two groups: published sources and unpublished sources.
6
● Sex
● Age
● The state to which they belong
● Religion
● Different faculties
● Heights or weights and so on
The functions of classification may be briefly summarized as follows:
i. It condenses the data
ii. It facilitates comparisons statistical treatment of the data
iii. it helps to study the relationships between groups of data
iv. it facilitates the statistical treatments of the data
The basis or the criteria with respect to (w.r.t) which the data are classified
primarily depends on the objectives and the purpose of the enquiry. Generally,
the data can be classified on the following four bases;
i. Geographical i.e., Area-wise or regional
ii. Chronological i.e., w.r.t occurrence of time
iii. Qualitative i.e., w.r.t some character or attribute
iv. Quantitative i.e., w.r.t numerical values or magnitudes
Frequency Distribution
A frequency table is a table that displays all the observed values of a variable
under study and shows many times each values occurs. The distribution of
the total number of observations among the various categories or classes of
the variable is called a frequency distribution.
The organization of the data pertaining to a quantitative phenomenon involves
the following four stages;
i. The set or series of individual observations - unorganized (raw) or
organized (arrayed data)
7
ii. Discrete or Ungrouped distribution
iii. Continuous frequency distribution
iv. Grouped frequency distribution
We shall explain the various stages by means of a numerical illustration.
Example: the table below is a frequency table distributing 100 female students
in a class according to their marital status (nominal data).
Relative frequency
This would be very useful if we need to compare tow data set of different
sizes.
8
fewer classes by grouping them. A number of rules of the thumb have been
proposed for calculating the proper number of classes. However, an elegant,
though approximate formula seems to be one given by Prof. Sturges known as
Sturges’srule, According to which K = 1+3.322 log10 N
Which K is the number of class intervals (classes) and N is the total frequency
i.e, total number of observations in the data. The value obtained is rounded to
the next higher integer.
Accordingly, the Sturges formula very ingeniously restricts the number of
classes between 4 and 20, which is fairly reasonable number from practical
point of view. The rule, however, fails if the number of observations is very
large or very small.
One has also to decide on the width of the class intervals (or size of class
intervals). In general, one should aim at class of equal width. Each class
intervals width, w, is obtained by w= , where R is the range, and K is the
number of class intervals. Another rule of the thumb for determining the rule
of the class interval should not be greater than th of the estimated population
standard deviation.
of marks Monthly
scholarshi
p
60-65 2,500
9
65-70 3,000
70-75 3,500
75-80 4,000
80-85 4,500
30-50 5 5
50-55 5 60+5=65
55-60 5 65+5=70
Marks F
10
Less than 30 0
30-35 5 65+5=70
50-55 5 5+5=10
55-60 5 5
11
Less than 50 10
Tabulation
By tabulation we mean the systematic presentation of the information
contained in the data, in rows and columns in accordance with some salient
features or characteristics. Rows are horizontal arrangements and columns
and vertical arrangements. In the word of A. M. Tuttle; “A statistical table is the
logical listing of related quantitative data in rows and columns with sufficient
explanatory and qualifying words, phrases and statements in the form of
titles, heading and notes to make clear the full meaning of data and their
origin.” Professor Bowley in his manual of statistics refers to tabulation as “the
intermediate process between the accumulation of data in whatever form
they are obtained, and the final reasoned account of the result shown by the
statistics.”
The various parts of a table include: the table number, title, head notes or
prefatory notes, captions and stubs, body of the table, foot-note, and source
note.
12
etc. we shall consider the most commonly used once, they include: line
graphs, Bar charts, Pie charts, Histogram, Frequency polygon, and Ogive.
1. Line Diagram: This is the simplest of all the diagrams. It consists in
drawing vertical line being equal to the frequency. Line graph facilitate
comparisons.
Example: The following data shows the number of accidents sustained
by 314 drivers of a public utility company over a period of five years.
Number of accidents:
0 1 2 3 4 5 6 7 8 9 10 11
Number of drivers:
82 44 68 41 25 20 13 7 5 4 3 2
Represent the data by a line diagram.
2. Bar charts: Bar diagrams are one of the easiest and the most
commonly used devices of presenting most of the business and
economic data. These are especially satisfactory for categorical data or
series. The height (or length) of each bar indicating the size of the figure
represented, where the length of a single bar is proportional to the
magnitude of each part of the data. The sizes of the bars are the same.
The following are the various types of bar diagram in common use:
I. Simple bar chart
II. Sub- divided or component bar chart
III. Percentage bar chart
IV. Multiple bar diagram
V. Deviation or bilateral bar diagram
13
workers employed in the factories tabulated below.
Factory A B C D
Pacific 70.8
Atlantic 41.2
Indian 28.5
Antarctic 7.6
14
Arctic 4.8
15
with the bases (sections) equal to the width of the corresponding class
intervals and heights are so taken that the areas of the rectangles are equal to
the frequencies of the corresponding classes. The values are taken along the
x-axis and the frequencies along the y-axis. This however, involves two cases:
case (i) Histogram with equal classes, case (ii) Histogram with unequal
classes.
Example: represent the adjoining distribution of marks of 100 students in the
examination by a histogram
16
Weekly 10-15 15-20 20-25 25-30 30-40 40-60 60-80
wages
Weekly 20-2 25-2 30-3 35-3 40-4 45-4 50-5 55- 60- Total
wages 4 9 4 9 4 9 4 59 64
( ’00 N
17
workers
X = mean
ii. Moderately Asymmetrical (Skewed) Frequency Curves.
18
A frequency curve is
said to be skewed (asymmetrical) if it is not symmetrical. Such curves are
stretched more to one side than to the other. If the curve is stretched more to
the right (i.e, it has a longer tail towards the right), it is said to be positively
skewed and if it is stretched more to the left (i.e. has a longer tail towards
theleft), it is said to be negatively skewed.
19
iv. U-Curve.
The frequency distribution in which the maximum frequency occurs at the
extremes (i.e, both ends) of the range and the frequency keeps on falling
symmetrically (about the middle), the minimum frequency being attained at
the center give rise to a U-Shaped curve.
U-shaped Curve
Bi
-
m
o
v.
d
Mixed Curves
Sometimes, though very rarely, we come across certain distributions in which
al
maximum frequency is attained at two or more points in an irregular manner.
Such curves are obtained in a distribution where as the value of the variable
C
increase, the frequencies increase and decrease, then again increase and
decrease twice or thrice as shown in the diagram or even more than that.
ur Tri-modal
ve 20
f
e
q
r
u
frequency
Variable
Variabl
The table below give the marks obtained by 70 candidates in STA 101
21
examination
Marks No. of Candidates (f) Less than C.F More than C.F
0-10 2 2 70
plot both less than and more than ogiv for the data.
22
AVERAGES
One of the important objectives of statistical analysis is to determine various
numerical measures which describe the inherent characteristics of a
frequency distribution. The first of such measures is average. The averages
are the measures which condense a huge unwieldy set of numerical data into
single numerical value which are representative of the entire distribution.
Averages are the typical values around which other items of the distribution
congregate. They are the values which lie between the two extreme
observation, (i.e., the smallest and the largest observations), of the
distribution and give us an idea about the concentration of the values in the
central part of the distribution. Accordingly they are also sometimes referred
to as the measures of central tendency. Averages are very much useful:
i) For describing the distribution in concise manner
ii) For comparative study of different distributions
iii) For computing various other statistical measures such as dispersion,
skewness, kurtosis and various other basic characteristics of a mass
data. Averages are also sometimes referred to as measures of
location since they enable us to locate the position or place of the
distribution in question.
“Statistical analysis seeks to develop concise4 summary figures which
describe a large body of quantitative data. One of the most widely used set of
summary figures is known as measures of location, which are often referred
to as averages, measures of central tendency or central location. The purpose
for computing an average value for a set of observation is to obtain a single
value which is representative of all the items and which the mind can grasp
simply and quickly. The single value is the point or location around cluster”
_______ Lawrence J. Kaplan
The following are the five measures of central tendency or measure of
location which are commonly used in practice
(i) Arithmetic mean or simply mean
23
(ii) Geometric mean
(iii)Harmonic mean
(iv) Median
(v) Mode
We shall discuss them in detail one by one.
Arithmetic mean (AR)
Arithmetic mean of a given set of observations is their sum divided by the
number of observations. For example, the AR of 5, 8, 10, 15, 24 and 28 is
In general, if X1, X2, ---, Xn are the given n observations, then their arithmetic
mean, usually denoted by is given by:
X=
Example: the following table gives the daily income of ten operators in a
machine tool factory. Find the mean.
Name of A B C D E F G H I J
operators
Income 12 15 18 20 25 30 22 35 37 26
Solution
If income is represented by x, then
=
In case of a discrete frequency distribution:
X - - -
- - -
24
The arithmetic mean is given by:
=
Illustration: The following is the frequency distribution of the number of
telephone calls received in 245 successive one-minute intervals at an
exchange:
No of Calls 0 1 2 3 4 5 6 7
One Minute 14 21 25 43 51 40 39 12
intervals
No of 0 1 2 3 4 5 6 7
Calls (x)
Frequenc 14 21 25 43 51 40 39 12 ∑f=245
y (f)
= =3.8
ARITHMETIC MEAN FOR GROUPED DATA
In calculating the mean of a grouped data set it is customary to assume that
all values falling in a particular class interval are located at the class mark or
midpoint of that interval.
To calculate the mean of such a data set; we multiply each class mark by the
25
corresponding class frequencies, sum these products over all the class
intervals, and divide the results by the total frequency. Thus, the mean is
calculated as = i=1, 2, - - -, k, where xi represent the class mark
Illustration: The data below showing the frequency distribution of a lifetime of
720 television tubes:
∑fi=720 ∑fixi=
26
167090.
0
=
Mathematical Properties of Arithmetic Mean
Arithmetic Mean possesses some very interesting and important
mathematical properties as given below:
Property 1: The algebraic sum of the deviations of the given set of
observations from their arithmetic mean is zero.
Mathematically, ∑ (x - ) = 0
Or for a frequency distribution: ∑f (x - ) = 0
Proof: ∑f (x - ) = ∑(fx - f) = ∑fx - ∑f
= ∑ f x - ∑f (: is a constant)
= ∑ f x - N (: ∑f = N)
But = ∑ fx = ∑ fx = N
: ∑ f (x - = N - N = 0
Property 2: The sum of squares of deviations of the given set of observations
is minimum when taken from the arithmetic mean compare to the sum of
squares of deviation of other value of the distribution.
Mathematically, for a given frequency distribution, the sum
2
S = ∑ f (x – A)
Which represent the sum of the squares of deviations of given observations
from any arbitrary value ‘A’ is minimum when A =
2
S1 = sum of squared deviations from mean = ∑ (x - ) , and
2
S = sum of squared deviations from any arbitrary point A = ∑ (X – A) : A
27
Then S1 is always less than S i.e, S1 <S
Property 3: The product of the arithmetic mean and number of values on
which the mean is based is equal to the sum of all given values. that is,
= fx
Property 4: Mean of the combined series
The mean of all the sum (or, differences) of corresponding observations in two
series, number of observations being equal in the two, is equal to the sum (or,
difference) of the means of the two series. If n1 and n2 are the sizes and 1 and 2
are the respective means of two groups then the mean of the combined
group of sizes n1 + n2 is given by:
= or 1,2 =
In general, if 1, 2, - - - k are the arithmetic means of k groups with n1, n2, - - -, nk
observations respectively, then
=
Illustration: The mean of marks in statistics of 100 students in a class was 72.
The mean of marks of boys was 75, while their number was 70. Find out the
mean marks of girls in the class.
Solution: In the usual notations we are given:
n1 = 70, 1 = 75; n1 + n2 = 100, = 72 : n2 = 100 – 70 = 30, we want 2
= = 72 =
2 = = = 65
Hence the mean of marks of girls in the class is 65.
Step Deviation Method for Computing Arithmetic Mean
It may be pointed out that the formula can be used conveniently if the values
of x or/and f are small. However, if the values of x or/and f are large, the
calculations of mean by is quite tedious and time consuming. In such a case
the calculation can be reduced to a great extent by using the step deviation (or
28
assumed mean) method which consists in taking the deviation (differences)
of the given observation from any arbitrary value A.
Let d = X – A, then fd = f (X – A) = Fx – A.F
Taking the sum over various values of x, we get
∑fd = ∑fx - A∑f
∑fd = ∑fx – A.N (:∑f = N)
: Dividing both sides by N, we get
=A+
In case of grouped or continuous frequency distribution, with class intervals of
equal magnitude, the calculations are further simplified by taking:
d = , where X is the mid-value of the class and h is the common magnitude of
the class intervals.
From d = , we get hd = (X – A)
Multiplying both sides by f, we get hfd = f (X – A) = fx – FA
Summing both sides over the values of X, we get:
h∑fd = ∑fx - A∑f = ∑fx – N.A
Dividing both sides by N, we get
h
No. of 6 5 8 15 7 6 3
Students
29
(i) By the direct formula (ii) By the step deviation method
Solution
(i) = =
(ii) A = 35, h = 10
: =A+
Geometric Mean
The geometric mean usually abbreviated as (G.M) of a set of n observation is
th
the n root of their product. Thus if x1, x2, - - - xn are the given n observation then
30
their G.M is given by
1/n
G.M = = (X1. X2. - - -. Xn)
For example, the G.M of 4, 8, 16 is
G.M = = 8
th
But if n, the number of observations is large, then the computation of the n
root is very tedious. In such a case the calculations are facilitated by making
use of the logarithms. Taking logarithm of both sides
Solution
G.M. = Antilog
31
No. of 5 7 15 25 8
Students
Solution
∑f = 60 ∑ f log x =
84.5243
HARMONIC MEAN
Another important mean is the harmonic mean which is used for averaging
the rates. If X1, X2, - - -, Xn is a given set of n observations, then their harmonic
mean (H.M) or simply H is given by:
H=
In other words, H.M. is the reciprocal of the arithmetic mean of the reciprocal
of the given observations. In case of frequency distribution, we have:
32
=H=
Example: The following table gives the weights of 31 persons on a sample
enquiry. Calculate the mean weight using harmonic mean
Weights 13 13 14 14 14 14 14 15 15
(lbs) 0 5 0 5 6 8 9 0 7
No. of 3 4 6 6 3 5 2 2 1
persons
Solution
130 3 0.0231
135 4 0.0296
140 6 0.0429
145 6 0.0414
146 3 0.0205
148 5 0.0338
149 2 0.0134
150 1 0.0067
33
157 1 0.0064
∑f = 31 ∑ = 0.2178
H.M. =
-1
Example: A cyclist pedals from his house to his college at a speed of 10 km h
-1
and back from the college to his house at 15km h . Find the average speed.
Solution
Let the distance from the house to the college be xkms. In going from house
to college, the distance (x kms) is covered in hours, while in coming from
college to house, the distance is covered in hours. Thus a total distance of 2x
kms is covered in hours.
Hence, average speed =
-1
H = 12 km h
MEDIAN
In the words of L.R. Connor:
“The median is that value of the variable which divides the group into two
equal parts, one part comprising all the values greater and the other, all the
values less than median”. Thus Median of a distribution may be define as that
value of the variable which exceeds and is exceeded by the same number of
observations i.e, it is the value such that the number of observations above it
is equal to the number of observations below it. Thus the median is a
positional average i.e, its value depends on the position occupied by a value in
the frequency distribution.
Calculation of Median
Case (1): Ungrouped Data: if the number of observations is odd, then the
median is the middle value after the observations have been arranged in
34
ascending or descending order of magnitude. For example, the median of 5
observations 32, 12, 40, 8, 60 i.e, 8, 12, 35, 40, 60 is 35.
In case of even number of observations median is obtained as the arithmetic
mean of the two middle observations after they are arranged in ascending or
descending order of magnitude.
For example 8, 12, 35, 40, 50, 60
The median =
Case (II): Frequency Distribution (Discrete type): In case of frequency
distribution where the variable takes the values X1, X2, - - -, Xn with respective
frequencies f1, f2, - - -, fn with , total frequency, median is the size of the item or
observation.
In this case the use of cumulative frequency distribution facilitates the
calculations.
Example: eight coins were tossed together and the number of heads (x)
resulting was noted. The operation was repeated 256 times and the frequency
distribution of the number of heads is given below:
No of heads 0 1 2 3 4 5 6 7 8
(x)
X F Less than
c.f.
0 1 1
1 9 9
2 26 36
35
3 59 95
4 72 167
5 52 219
6 29 248
7 7 255
8 1 256
∑f = 256
Here, ∑f = 256, = = 128. the c.f. just greater than 128 is 167 and the value of X
corresponding to 167 is 4. hence, median number of heads is 4.
36
Example: The following table shows the frequency distribution of weight in
grams of mangoes of a given variety. Calculate the median.
Solution
= = 100 The c.f. just greater than 100 is 130. Hence the corresponding class
439.5 – 449.5 is the median class.
Median = l +
= 439.5 +
= 443.94 grams
37
We can also get the median by plotting the ogive of the distribution, median is
the value below which item lie.
MODE
Mode is the value which occurs most frequently in a set of observations and
around which the other items of the set cluster densely. In other words, mode
is the value of a series which is predominant in it. In the words of Coxton and
Cowden, “The mode of a distribution is value at the point around which the
items tend to be most heavily concentrated. It may be regarded as the most
typical of a series of values”. According to A.M. Tuttle. ‘Mode is the value
which has the greatest frequency density in its immediate neighborhood’.
Illustration: The wheat yield in a particular region over the past 12 years (in
millions of tons) are: 1.5, 1.3, 1.2, 1.0, 1.3, 1.4, 1.6, 1.7, 1.5, 1.3, 1.2 and 1.4.
The mode is 1.3 (million tons).\
In case of a frequency distribution, mode is the value of the variable
corresponding to the maximum frequency. This method can be applied with
ease and simplicity if the distribution is ‘unimodal’. For example, in the
following distribution:
X 1 2 3 4 5 6 7 8 9
Fre
qu
enc
y
38
0 Mode x
Mode = l + , where
Assumptions:
Example: find the value of mode from the data given below:
Weight (in 93-97 98-102 103-10 108-11 113-11 118-12 123-12 128-13
kg) 7 2 7 2 7 2
39
No. of 3 5 12 17 14 6 3 1
students
Solution
Class boundaries f
92.5 – 97.5 3
97.5 – 102.5 5
117.5 – 122.5 6
122.5 – 127.5 3
127.5 – 132.5 1
Here maximum frequency is 17. The corresponding class 107.5 – 112.5 is the
modal class.
Mode = l + =
40
EMPIRICAL RELATION BETWEEN MEAN (M), MEDIAN (MD), MODE (MO)
Centre of gravity
MO MD M
41
Quartiles: The values which divide the given data into four equal parts are
known as quartiles. Obviously there will be three such points Q1, Q2 and Q3 such
that Q1 ≤ Q2 ≤ Q3, termed as the three quartiles. Q1, known as the lower or first
quartile is the value which has 25% of the items of the distribution below it and
consequently 75% of the items are greater than it. Incidentally Q2, the second
quartile, coincide with the median and has an equal number of observations
above it and below it. Q3, known as the upper or third quartile, has 75% of the
observations below it and consequently 25% of the observations above it. The
working principle for computing the quartiles is basically the same as that of
computing the median.
To computer Q1, find , where N = ∑f, see the (less than) c.f. just greater than ,
the corresponding value of x gives the value of Q1. In case of continuous
frequency distribution, the corresponding class containing Q1 and the value of
Q1 is obtained by the interpolation formula:
Q1 = l + , where all symbols has this usual meaning.
Similarly to compute Q3, see the (less than) c.f., just greater than , and for
continuous distribution, Q3 = l +
Deciles
Deciles are the values which divide the series into ten equal parts. Obviously
there are Nine deciles, D1, D2, D3, - - -, D9 (Say), such that D1 ≤ D2 ≤ D3 ≤ - - - ≤ D9.
Incidentally D5 coincides with the median. The method of computing the
deciles Di (i = 1, 2, 3, - - -, 9) is the same as discussed for Q1 and Q3. To compute
the ith decile = Di (i = 1, 2, 3, - - -, 9) see the c.f. just greater than . the
corresponding value of X is Di . In case of continuous frequency distribution the
corresponding class contains Di and its value is obtained by the interpolation
formula:
Di = l +
Percentiles:
Percentiles are the values which divide the series into 100 equal parts.
Obviously, there are 99 percentiles, P1, P2, P3, - - -, P99 such that P1 ≤ P2 ≤ P3 ≤ - - -≤
P99. The ith percentile Pi (i = 1, 2, 3, - - -, 99) is the value of X corresponding to
42
c.f. just greater than . In case of continuous frequency distribution, the
corresponding class contains Pi and its value is obtained by the interpolation
formula:
Pi = l +
In particular, we shall have:
P25 = Q1, P50 = D5 = Q2, P75 = Q3, D9 = P90, D1 = P10, D2 = P20
D3 = P30
The various partition values quartiles, deciles and percentiles can be easily
located graphically with the help of Ogive.
Example: The following data gives the distribution of marks of 100 students.
th th
Obtain the values of quartiles, 6 decile and 70 percentile
∑f =
100
43
Quartiles: Q1 = =
The c.f. just greater than is 32. Hence the corresponding class is 29.5 – 39.5
Q1 = 29.5 +
The c.f. just greater than is 80. Hence, the class is 49.5 – 59.5 is the Q3 class
Q3 = 49.5 +
th
6 Deciles D6 =
D6 = 49.5 +
th
70 Percentile =
P70 = 49.5 +
DISPERSION
44
The first two measures, range and quartile deviation are termed a position
measures since they depend upon the values of the variables of particular
position of the distribution. The last measure, Lorenz curve is a graphical
method of studying variability.
1. Range: Range is the difference between the greatest (maximum) and
the smallest (minimum) observation of the distribution. Thus
Range = Xmax - Xmin
In case of a grouped frequency distribution (for discrete values) or the
continuous frequency distribution, range is defined as the difference
between the upper limit of the highest class and the lower limit of the
smallest class.
Illustration: Calculate the range and the coefficient of range of A’s monthly
earnings for a year.
Earning 139 150 151 151 157 158 160 161 162 162 173 175
(N1000)
Solution
L = 175000, 5 = 139000
Range = L – S = 175000 – 139000 = 36000
Coefficient of range =
45
Age (in 16 – 20 21 – 25 26 – 30 31 – 35
years)
Solution
Convert into continuous classes. The first class will then become 15.5 –
20.5 and the last class will become 30.5 – 35.5
L = 35.5, 5 = 15.5
Range = L – S = 35.5 – 15.5 = 20 years
Coefficient of range =
46
Coefficient of Q.D =
Percentile Range
This is a measure of dispersion based on the difference between certain
percentiles. If Pi is the ith percentile and Pj is the jth percentile then the so-
called i-j percentile range is given by i-j percentile Range = Pj – Pi (i < j).
Thus i – j semi-percentile Range is given by:
(Pj – Pi)/2, (i < j)
th
The commonly used percentile range is the one which corresponds to the 10
th
and 90 percentile. Thus,
10 – 90 percentile Range = P90 – P10 and
10 – 90 semi-percentile Range = (P90 – P10)/2.
The above measures are absolute measures only. The relative measure of
variability based on percentile is given by:
Coefficient of 10 – 90 percentile =
47
is often called the mean deviation”,
If X1, X2, - - -, Xn are n given observations then the mean deviation (M.D)
about an average A, say, is given by:
M.D =
Where =
Steps:
(i) Calculate the average A of the distribution by the usual method
(ii) Take the deviation d = X – A of each observation from the Average A.
(iii)Ignore the negative signs of deviation, taking all the deviation to be
positive to obtain the absolute deviation, = .
(iv) Obtain the sum of the absolute deviations obtained in step (iii)
(v) Divide the total obtained in step (iv) by n, the number of observation.
The result gives the value of the mean deviation about the average A. In case
of frequency distribution or grouped or continuous frequency distribution,
mean deviation about an average A is given by:
M.D. = , where x is the value of variable or it is the mid-value of the class
interval.
Relative Measures of Mean Deviation. The measures of mean deviation as
defined above are absolute measure depending on the units of measurement.
The relative measure of dispersion called the coefficient of mean deviation is
given by:
Coefficient of M.D. =
coefficient of M.D. about mean =
And coefficient of M.D. about median =
The coefficients of mean deviation defined above are pure numbers
independent of the units of measurement and are useful for comparing the
variability of different distribution.
48
Example: Calculate the mean deviation from the following data given marks
obtained by 11 students in a class test 14, 15, 23, 20, 10, 30, 19, 18, 16, 25, 12
49
Solution
M.D. =
Example: Calculate mean deviation from median of the following distribution.
∑ F = 5250
Here
Median = 1 +
M.D. = about median =
50
Coefficient of M.D. =
4. Standard Deviation
Standard deviation, usual denoted by the letter (small sigma) of the Greek
alphabet was first suggested by Karl Pearson as a measure of dispersion in
1893. It is defined as the positive square root of the arithmetic mean of the
squares of the deviations of the given observations from the arithmetic mean.
Thus if X1, X2, - - -, Xn is a set of n observations then its standard deviation is
given by:
2
= , where
, is the arithmetic mean of the given values.
Steps:
(i) Compute the arithmetic mean
(ii) Compute the deviation (X - ) of each observation from arithmetic mean,
i.e., obtain X1 - , X2 - , - - -, Xn - .
2
(iii)Square each of the deviations obtained in step (ii) i.e., compute (X1 - ) ,
2 2
(X2 - ) , - - -, (Xn - .
(iv) Find the sum of the squared deviations in step (iii) and divide by n given
by:
2 2 2 2
∑f (X - ) /n = (X1 - ) + (X2 - ) + - - - + (Xn - .
(v) Take the positive square root of the value obtained in step (v)
(vi) The resulting value gives the standard deviation of the distribution.
In case of frequency distribution, the standard deviation is given by:
2
= , N = ∑f, X is the value of the variable or the mid-value of class (in case of
grouped or continuous frequency distribution); f is the corresponding
frequency of the value x.
Thus the value of will be greater if the values of X are scattered widely away
51
from the mean. Thus a small value of will imply that the distribution is
homogeneous and a large value of will imply that it is heterogeneous.
Variance and Mean Square Deviation
According to William I. Greenwald the variance is the mean of the squared
deviations about the mean of a series. Thus, variance is the square of the
2
standard deviation and is denoted by . For a frequency distribution variance is
given by:
2 2
= .
2
The mean square deviation, usually denoted by S is defined as
2 2
S = , where A is any arbitrary number.
The square root of the mean square deviation is called root mean square
2
deviation and given by: S =
2 2
Relation between and S . We have
2 2
S =
2
=
2 2
= +( +2
2 2
= + .
being constant is taken outside the summation sign.
∑f
2 2 2
S = +
2 2 2
so, S = + [
(, being the square of a real quantity is always non-negative.
2 2
Thus S = + (A non-negative quantity)
22
S
In other words, mean square deviation is not less than the variance or the root
52
mean square deviation is not less than the square deviation.
2 2
S = iff
2
( =O
so, ,
2
Thus, S will be least when = A. Hence, mean square deviation or equivalently
root mean square deviation is least when deviations are taken from the
arithmetic mean and variance (standard deviation) is the minimum value of
mean square deviation (root mean square deviation).
Different Formula:
2 2
● x = , ∑f = N
● 2 22 2 2
x = = -
If d = X – A, when A is an arbitrary constant, then
● 2 2 2 2
x = d = -
If we change the origin and scale in X i.e., if we take
d = ; h > O, then
2 22 2 2 2
● x =h d =h
solution
53
2
X X– (X -
240.16 0.00 0
2
∑X = ∑ (X – ∑ (X - = 0.0106
2401.60
Example: Calculate the mean and standard deviation from the following:
54
Value 90 – 99 80 – 89 70 – 79 60 – 69 50 – 59 40 – 49 30 – 39
Solution
2
Class Mid-value (X) f d= Fd fd
2
∑f = 75 ∑fd = ∑fd = 127
27
= 68.1
55
= h. = 12.505
It has been pointed out that we need statistical measures which will reveal
clearly the salient features of a frequency distribution. The measures of
central tendency tells us about the concentration of the observations about
the middle of the distribution and the measure of dispersion gives us an idea
about the spread or scatter of the observations about some measure of
central tendency. We may come across frequency distributions which differ
widely in their nature and composition and yet may have the same central
tendency and dispersion, but yet may give histograms which differ very widely
in shape and size.
Skewness
(ii) The values of mean, median and mode fall at different point, i.e., they
do not coincide.
56
(iii)Quartiles Q1 and Q3 are not equidistant from the median Q3 – md ≠ md
– Q1
(iv) The corresponding pairs of deciles and percentiles are not equidistant
from the median i.e.,
D5 – D5 – i ≠ D5+1 – D5 (i = 1, 2, 3, 4)
(v) The sum of the positive deviations from the median is not equal to the
sum of the negative deviation from the median.
Skewness =
But quite often, mode is ill-defined and is thus quite difficult to locate. In
such a situation, we use the following empirical relationship between the
mean, median and mode for a moderately asymmetrical (skewed)
distribution.
MO = 3md – 2m
= skewness =
Skewness =
57
from the median value. The refinement was suggested by Kelly. Kelly’s
percentile (or decile) measure of skewness is given by:
Sk (Kelly) =
KURTOSIS
A – Lepto - Kurtic
B – Meso - Kurtic
C – Platy - Kurtic
58
Kurtosis is concerned with flatness or peakedness of the frequency curve.
Curve of type B which is neither flat nor peaked is known as normal curve and
shape of its hump (middle part) is accepted as a standard one. Curve with
humps of the form of a normal curve are said to have normal kurtosis and are
termed as meso-kurtic. The curve of type A, which is more peaked than the
normal curve are known as lepto-kurtic and are said to lack Kurtosis or to have
negative kurtosis. On the other hand, curve of type C, which are flatter than
the normal curve are called platy-kurtic and they are said to possess kurtosis
in excess or have positive kurtosis.
59
measurements will consistently show weights that are 1% higher than their
actual weights.
Absolute Error =
-Relative Error: The absolute error expressed as a fraction of the true value,
often expressed as a percentage.
Approximations
In statistics, approximations are used to simplify complex calculations.
Common methods include rounding, truncating, and using linear
approximations.
1. Rounding
Rounding involves reducing the number of digits to make numbers easier to
work with.
Example: Round 3.567 to two decimal places:
3.567 ≈ 3.57
60
2. Truncation
Truncation eliminates digits beyond a certain point without rounding.
Example: Truncate 3.567 to two decimal places:
at x=2: f(2)=4
61
= +
Result z = 15±1.1
Understanding errors and approximations is essential for accurate data
analysis in statistics. By recognizing the types of errors and how to propagate
them, statisticians can provide more reliable insights and improve the quality
of their conclusions. Always consider both types of errors when reporting
results to convey a complete picture of uncertainty.
Additional Practice Problems
1. If the measured length of a metal rod is 100.5 cm, and the true length is
100 cm, calculate the absolute and relative errors.
2. Use linear approximation to estimate f(x) = sin(x) at x = and x = +0.1
3. A car travels 150 km at a speed of 50 km/h with a possible error of ±2
km/h. Find the propagated error in the travel time.
62
fields.
Definitions
1. Ratio: A ratio is a quantitative relationship between two numbers,
showing how many times one value contains or is contained within the
other. It can be expressed in the form of a fraction, using a colon, or as a
decimal.
Ratio = , where (A\) and (B\) are the two quantities compared.
Example: If there are 20 boys and 30 girls in a classroom, the ratio of boys to
girls can be calculated as: = . This means that for every 2 boys, there are 3
girls in the classroom.
2. Rate: A rate is a specific kind of ratio that compares two quantities of
different units, usually expressing one quantity per unit of another. Rates often
provide context and can indicate a frequency, intensity, or prevalence.
Rate =
Example: If a car travels 150 miles in 3 hours, the rate of speed can be
calculated as:
Rate of speed = .=50 miles per hour
63
2. Health
Example: Incidence rate is a measure used in epidemiology to describe the
frequency of new cases of a disease in a particular population over a specific
period.
If 50 new cases of a disease are reported in a population of 10,000 over one
year, the incidence rate is calculated as:
Incidence Rate = x100 = 0.5%
This indicates that 0.5% of the population developed the disease during that
year.
3. Demographics
Example: The birth rate is an example of a demographic rate, calculated as
the number of live births in a year per 1,000 people in the population.
If a country has 10,000 live births in a year and a population of 1,000,000, the
birth rate is:
Birth Rate = x1,000= 10 births per 1,000 people.
This means there are 10 births for every 1,000 individuals in the population.
Properties of Ratios and Rates
- Ratios can be simplified to their lowest terms, which can help facilitate
comparison.
- Rates are crucial in comparing different scales and sizes, making them useful
for standardized measurements.
- Understanding the context of the rates and ratios is essential, as they can
mislead if the underlying data changes.
Rates and ratios are fundamental statistical tools that help interpret and
summarize data efficiently. By comparing different quantities, they provide
valuable insights across various fields, including finance, health, and
demographics. Mastery of calculating and interpreting these measures
permits better decision-making and more profound understanding of the data
64
trends and patterns.
Practice Problems:
1. Calculate the unemployment rate if 300 out of 10,000 people in the
workforce are unemployed. Provide your answer per 1,000 people.
2. You are given that in a town of 5,000 individuals, 1,250 are children, and the
rest are adults. Compute the ratio of children to adults.
INDEX NUMBERS
Definition:
Index numbers are statistical devices designed to measure the relative
changes in the level of a phenomenon (variable or a group of variables) with
respect to time, geographical location or other characteristics such as income,
profession, etc. in other words, Index numbers are specialized types of rates,
ratios, percentages which give the general level of magnitude of a group of
distinct but related variables in two or more situations. Index numbers is a
statistical device which enables us to arrive at a single representative figure
which gives the general level of the price of the commodities in an extensive
group.
65
USES OF INDEX NUMBERS
The first index number was constructed by an Italian, Mr. Carli, in 1764 to
compare the changes in price for the year 1750 (current year) with the price
level of in 1500 (base year) in order to study the effect of discovery of America
on the price level in Italy. Though originally designed to study the general level
of prices or accordingly purchasing power of money, today index numbers are
extensively used for a variety of purposes in economics, business,
management, etc. and for quantitative data relating to production,
construction, consumption, profits, personnel and financial matters, etc. for
comparing changes in the level of phenomenon for two periods, places, etc.,
there is hardly any field of quantitative measurements where index numbers
are not constructed, they are used in almost all sciences. The main uses of
Index numbers can be summarized as follows:
1. Index numbers serve as economic barometer
2. Index numbers help in studying trends and tendencies
3. Index numbers help in formulating decisions and policies.
4. Price index measure the purchasing power of money.
5. Index numbers are used for deflation
66
b. Retail Price Index Numbers: These indices reflect the general
changes in the retail prices of various commodities such as
consumption of goods, stocks and shares, bank deposits,
government bonds, etc.
2. Quantity Index Numbers: this study the changes in the volume of goods
produced (manufactured), consumed or distributed, like the indices of
agricultural production, industrial production, imports and exports, etc.
they are extremely helpful in studying the level of physical output in an
economy.
3. Value Index Numbers: these are intended to study the change in the
total value (Price multiplied by quantity) of production such as indices of
retail sales or profits or inventories. However, these indices are not as
common as price and quantity indices.
67
ii. The base period should not be too distant from the given period.
NOTATION AND TERMINOLOGY
Base Year: The year selected for comparison i.e. the year with respect to
which comparison are made. It is denoted by the suffix zero ‘0’.
Current Year: The year for which comparisons are sought or required. It is
denoted by the suffix ‘1’.
P0: Price of commodity in the base year.
P1: Price of commodity in the current year.
q0: Quantity of a commodity consumed or purchased during the base year.
q1: Quantity of a commodity consumed or purchased in the current year.
w: Weight assigned to a commodity according to its relative importance in the
group.
I: simple index number or price relative obtained on expressing current year
price as a percentage of the base year price and is given as
I = Price Relative =
P01: Price Index Number for the current year with respect to the base year.
P10: Price Index Number for the base year with respect to the current year.
Q01: Quantity Index Number for the current year with respect to the base year.
Q10: Quantity Index Number for the base year with respect to the current year.
V01: Value Index for the current year with respect to the base year.
METHODS OF CONSTRUCTING INDEX NUMBERS
We shall now discuss the various techniques or methods used for the
construction of index numbers. The price indices is the most important of all
the indices, we shall discuss their construction in detail. The quantity indices
can be obtained from price indices by interchanging the price (p) and quantity
(q) in the final formula.
1. Simple (Unweighted) Aggregate Method
68
This is the simplest of all the methods, it consists in expressing the total price
in the current year as a percentage of the aggregate of prices in the base year.
Thus
P01 =
The quantity index is given by
Q01 =
Example:
From the following data calculate Index Number by simple Aggregate method.
Commodity A B C D
Price in 1980 162 256 25 132
7
Price in 1981 171 164 18 145
9
(P0) = 807
(P1) = 669
69
By using different systems of weighting we get a number of formulae, some of
the important are given below
i. Laspeyre’s Price Index or Base Year Method: Taking base year quantities as
weights i.e. w = q0 in the above formula (*), we get the Laspeyre’s price Index
as:
La
P01 = 100
This formula was devised by French Economist Laspeyre in 1817.
ii. Paasche’s Price Index: if we take current year quantities as weights in (*),
we obtain paasche’s price index which is given by:
Pa
P01 = 100
This formula was given by German Statistician Paasche in 1874.
iii. Dorbish – Bowley Price Index: This index is given by the arithmetic mean of
Laspeyre’s and Paasche’s price index numbers and we have:
DB
P01 = 100
iv. Fisher’s Price Index: Irving Fisher advocated the geometric cross of La and
Pa price index numbers and is given by:
F La Pa
P01 = [ P01 P0 ] = 100
Fisher’s index is termed as an ideal index since it satisfies time reversal and
factor reversal tests for the consistency of index numbers.
v. Marshall-Edgeworth Price Index: taking the arithmetic cross of the
quantities in the base year and the current year as weights i.e. w = , we obtain
the Marshal-Edgeworth (M.E.) formula given by:
ME
P01 =
=
vi. Walsch Price Index: instead of taking the arithmetic of base year and
current year quantities as weights, if we take their geometric mean, i.e., w = ,
70
then we obtain Welsch Index given by:
Wa
P01 =
vii. Kelly’s Price Index or Fixed Weight Index: this formula, named after Truman
L. Kelly, requires the weights to be fixed for all periods and is also sometimes
known as aggregative index with fixed weights and is given by the formula:
K
P01 =
Where the weights are the quantities (q) which may refer to some period (not
necessarily the base year or the current year) and are kept constant for all
periods.
VALUE INDICES
Value Index numbers are obtained on expressing the total value (or
expenditure) in any given year as a percentage of the same in the base year.
Symbolically, we write
V01 =
V01 =
71
EXAMPLE
a. From the following data, calculate price index numbers for 1980 with
1970 as base year by
i. Laspeyre’s method
ii. Paasches’s method
iii. Marshall-Edgeworth method, and
iv. Fisher’s Ideal method
b. It is stated that Marshall-Edgeworth index number is a good
approximation to Fisher’s Ideal index number. Verify this for the data.
A 20 8 40 6
72
SIMPLE AVERAGE OF PRICE RELATIVES
In this method, first of all we obtain the price relatives for each commodity.
The price relatives are obtained by expressing the price of the commodity in
the current year as a percentage of its price in the base year. i.e.
P = Price relative for a commodity =
Price-relatives are the simplest form of the index numbers for each
commodity. The price index for the composite group is obtained on averaging
these price-relatives by using some suitable measures of central tendency,
usually arithmetic mean (A.M.) or geometric mean (G.M.). Price Index using
simply arithmetic mean of relatives is given by:
P01 (A.M.) =
Where n is the number of commodities in the group.
Using simple geometric mean of the price relatives, the price index is given by:
P01 (G.M.) = =
Where Π denotes the product of the price-relatives for the n commodities. To
evaluate this, we use logarithms. Taking logarithms of both sides, we get
LogP01 (G.M.) =
P01 (G.M.) = Antilog
73
influence the index unduly. It gives equal importance to all observations.
The drawback of this method is that it gives equal weights to all the
commodities and thus neglects their relative importance in the group. This
drawback is removed by taking the weighted average of the price-relative.
74
Example: Construct Index number for each year from the following average
wholesale prices of cotton with 1993 as base
Year Price Year Price
Example: The following are the prices of commodities in 1995 and 2000.
Calculate a price index based on price-relative using the arithmetic mean as
well as geometric mean.
Year Commodity
A B C D E F
75
WEIGHTED AVERAGE OF PRICE RELATIVE
The shortcoming of Simple Average of Relatives method which assumes that
all the relatives are equally important is overcome in this method which
consists in assigning appropriate weights to the relatives according to the
relative importance of the different commodities in the group. Thus, the index
for the whole group is obtained on taking the weighted average, usually A.M.
or G.M. of the price relatives. Thus, based on weighted A.M., the price index is
given by:
P01 (A.M) = =
Where W is the weight attached to the price-relative P.
Steps:
1. Find the price-relative (p) for each commodity, i.e., compute P =
2. Multiply the price relatives in step 1 by the corresponding weights (W)
assigned to get the product WP.
3. Obtain the sum of products obtained in step 2 for all the commodities
to get .
4. Divide the sum in step 3 by, the total of the weight assigned.
The resulting figure gives the price index based on the weighted average of
price-relatives.
The price index based on the weighted geometric mean of price relatives is
given by
P01 (weighted G.M.) =
Taking logarithms of both sides, we get
log [P01 (weighted G.M.)] =
P01 (weighted G.M.) = Antilog
Steps:
76
1. Compute the price-relatives P = , for each commodity.
2. Find the logarithms of all the price relatives, log P.
3. Multiply log P values for each commodity by the corresponding weights
(W) assigned. This will give ([Link] P) values.
4. Find the sum of the values in step 3 over all the commodities to get .
5. Divide the sum obtained in step 4 by , the sum of weights.
6. Antilog of the value obtained in step 5 gives required price-index.
Example: the following table gives the prices of some food items in the base
year and current year and the quantities sold in the base year. calculate the
weighted index number by using the weighted average of price relatives.
Items Base year Quantities Base Year Price Current year Price
A 7 18.00 21.60
B 6 3.00 4.65
77