100% found this document useful (1 vote)
11 views77 pages

STA111 Lecture Note Updated

The lecture notes for STA111 on Descriptive Statistics at the University of Abuja cover essential topics such as definitions and classifications of statistics, data collection methods, frequency distributions, and measures of central tendency and dispersion. The course aims to equip 100L students with foundational knowledge in statistics, emphasizing the importance of data collection and analysis in various fields. Students are encouraged to thoroughly read the provided lecture notes and engage with additional readings to enhance their understanding.

Uploaded by

jahdielkarma
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
100% found this document useful (1 vote)
11 views77 pages

STA111 Lecture Note Updated

The lecture notes for STA111 on Descriptive Statistics at the University of Abuja cover essential topics such as definitions and classifications of statistics, data collection methods, frequency distributions, and measures of central tendency and dispersion. The course aims to equip 100L students with foundational knowledge in statistics, emphasizing the importance of data collection and analysis in various fields. Students are encouraged to thoroughly read the provided lecture notes and engage with additional readings to enhance their understanding.

Uploaded by

jahdielkarma
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Lecture Notes for

STA111
Descriptive Statistics
(3 UNITS)
Department of
Statistics
University of Abuja

This lecture notes is not for sale.

1
Course Aim and Objectives
This course teaches Descriptive Statistics required by 100L Students. On
successful completion of the course, you are expected to have adequate
knowledge of the course contents listed below:
COURSE Contents:
● Definitions and classifications of Statistics
● Data, methods of data collection, Classification and Tabulation
● Frequency distributions (grouped and ungrouped)
● Diagrammatic and Graphic Representations of data
● Averages (Measures of Central Tendency), Measures of Dispersion,
and Partition
● Skewness and Kurtosis
● Errors and Approximations
● Rates and Ratio
● Index Numbers

Texts
This course has detailed lecture notes; it should not be necessary to buy
a book for this course. However, you may consult any elementary books
in Statistics.
Lecture notes:
The course attempts to convey a large amount of information in a short space
of time. Some of the material is of a technical nature and may not†be covered

2
explicitly in the lectures and classes. You are expected to read the lecture
notes thoroughly. The syllabus is defined by the contents of the lecture notes
Additional reading is suggested at the end of each week’s classes. Despite my
best efforts, there will be mistakes in the notes. If you spot something that
looks wrong, please let me know. Solutions to some problems have been
provided while some have been intentionally omitted. All problems will be
solved together in the lecture rooms.

Definitions of Statistics
Statistics has been defined differently by different writers from time to time so
much so that scholarly articles have collected together hundreds of
definitions, emphasizing precisely the meaning, scope and limitations of the
subject. The reasons for such definitions may be broadly classified as follows:
i. The field of utility of statistics has been increasing steadily and thus
different people defined it differently according to the development of
the subject. In old days statistics was regarded as the “science of state
craft” but today it embraces almost every sphere of natural and human
activity. Accordingly, the old definitions which were confined to and very
limited and narrow field of enquiry were replaced by the new definitions
which are more exhaustive and elaborate in approach.
ii. The word statistics has been used to convey different meanings in
singular and plural sense. When use as plural statics means numerical
set of data and when used singular (statistic) sense it means the
science of statistical methods embodying the theory and techniques
use for collecting analyzing and drawing inferences from numerical
data.

We give below some selected definitions of statistics.

3
1. “Statistics are the classified facts representing the conditions of the
people in a state. Specially those facts which can be stated in numbers
or in tables of numbers or in any tabular or classified arrangement.” -
Webster
2. “Statistics are numerical statements of facts in any departments of
enquiry place in relation to each other.” – Bowley.
3. By statistics we mean quantitative data affected to a marked extent by
multiplication of causes.” – Yule and Kendall.
4. Statistics may be defined as the aggregate of facts affected for a
marked extent by multiplicity of causes, numerically expressed
enumerated or estimated according to a reasonable standard of
accuracy, collected in a systematic manner, for a predetermined
purpose and placed in relation to each other”. – Prof. Horace Secrist.
Classification of Statistics
It has become accepted in today’s world that in order to learn about
something, you must first collect data. Statistics is the art of learning data. It is
concerned with the collection of data, its subsequent descriptions, and its
analysis, which often leads to the drawing of conclusion. At the end, the data
should be described for instance the scores of two groups of teaching
methods should be presented. In addition, summary measures such as the
average scores of members of each of the groups should be presented. This
part of statics, concerned with the description and summarization of data, is
called descriptive statistics. After the preceding experiment is completed and
the data are described and summarized, we hope to be able to draw
conclusion about which teaching method is superior. This part of statistics.
Concerned with the drawing of conclusion, is called inferential statistics.
Importance of Statistics
To a very striking degree our culture has become a statistical culture. Even a
person who may never have heard of an index number is affected by those
index numbers, which describe the cost of living. It is impossible to understand

4
psychology, sociology, economics, finance or natural and physical sciences
without some general idea of the meaning of an average, of variation, of
concomitance, of sampling, of how to interpret charts and tables. Statistics is
important in all area of human endeavors. For example, we have: statistics in
planning, statistics in state, statistics in mathematics, statistics in physics,
statistics in chemistry, in Biology, in Economics, in industry, in insurance in
astronomy, in psychology, in education, in war, in medical science, among
others, the list of area where statistics is important is endless.
Collection of Data
Statistics are set of numerical data. In fact only numerical data constitute
statistics. This means that the phenomenon under study must be capable of
quantitative measurement. Thus, the raw material of statistics always
originates from the operation of counting (enumeration), or measurements.
For any statistical enquiry, whether it is in business, economic, or Sciences, the
basic problem is to collect facts and figures relating to particular phenomenon
under study. The items in which the measurements are taken are called
statistical units. On the face of it, it might appear that the collection of data is
the first step for any statistical investigation. But in a scientifically prepared
(efficient and well-plane) statistical enquiry, the collection of data is by no
means the first step. Before we embark upon the collection of data for any
given statistics enquiry, it is imperative to examine carefully the following
points which may be termed as preliminaries to data collection: objectives and
scope of the enquiry, statistical units to be used, sources of information (data),
method of data collection, degree of accuracy aimed at in the final result, type
of enquiry. A good data however, possessed the following characteristics: it
should be unambiguous, it should be specific, it should be uniform, it should be
stable, and it should be appropriate.
Methods of Data Collection
Primary and Secondary Data
The data which are originally collected by an investigators or agency for the
first time for any statistical investigation and used by them in the statistical
analysis are termed as primary data. On the other hand, the data (published or
unpublished), which have already been collected and processed by some

5
agency or person and taken over from there and used by any other agency or
person for their statistical work are termed as secondary data as far as
second agency is concerned. The second agency if and when it publishes and
files such data becomes the secondary source to anyone who latter uses
these data. In other words, secondary source is the agency who publishes or
releases for use by others the data which was not originally collected and
processed by it. It may be observed that the distinction between primary and
secondary data is a matter of degree or relativity only. The same set of data
may be secondary in the hands of one and primary in the hands of others. In
general, the data are primary to the source who collects and processes them
for the first time and are secondary for all sources who latter use such data.
The methods commonly use for the collection of primary data are as follows:
direct personal investigation. Indirect oral interviews. Information received
through local agencies. Mailed questionnaire method. Schedules sent through
enumerators. The chief source of secondary data may be broadly classified
into the following two groups: published sources and unpublished sources.

Classification and Tabulation


Classification: it is of interest to give below the following definitions of
classification:
i. Classification is the process of arranging data into sequences and
groups according to their common characteristics or separating them
into different but related parts. – securest
ii. A classification is a scheme for breaking a category into a set of parts,
called classes, according to some precisely defined differing
characteristics possessed by all the elements of the category – Tuttle A.
M
Thus classification impresses upon the arrangement of the data into different
classes, which are to be determined depending upon the nature, objectives
and scope of the enquiry. For instance the number of students registered in
university of Abuja may be classified on the basis of any of the following
criterion:

6
● Sex
● Age
● The state to which they belong
● Religion
● Different faculties
● Heights or weights and so on
The functions of classification may be briefly summarized as follows:
i. It condenses the data
ii. It facilitates comparisons statistical treatment of the data
iii. it helps to study the relationships between groups of data
iv. it facilitates the statistical treatments of the data
The basis or the criteria with respect to (w.r.t) which the data are classified
primarily depends on the objectives and the purpose of the enquiry. Generally,
the data can be classified on the following four bases;
i. Geographical i.e., Area-wise or regional
ii. Chronological i.e., w.r.t occurrence of time
iii. Qualitative i.e., w.r.t some character or attribute
iv. Quantitative i.e., w.r.t numerical values or magnitudes
Frequency Distribution
A frequency table is a table that displays all the observed values of a variable
under study and shows many times each values occurs. The distribution of
the total number of observations among the various categories or classes of
the variable is called a frequency distribution.
The organization of the data pertaining to a quantitative phenomenon involves
the following four stages;
i. The set or series of individual observations - unorganized (raw) or
organized (arrayed data)

7
ii. Discrete or Ungrouped distribution
iii. Continuous frequency distribution
iv. Grouped frequency distribution
We shall explain the various stages by means of a numerical illustration.

Example: the table below is a frequency table distributing 100 female students
in a class according to their marital status (nominal data).

S/N Marital Status Frequency (number in


each class)

1. Married 28

2. Divorced 11

3. Separated 17

4. Single 44

Relative frequency

This would be very useful if we need to compare tow data set of different
sizes.

Grouped Frequency Distribution


Frequently one has a collection of sets of data whose tabulation would result
in lengthy frequency table. To summarize such data sets and make them
more comprehensible in a frequency table we often collapse the data into

8
fewer classes by grouping them. A number of rules of the thumb have been
proposed for calculating the proper number of classes. However, an elegant,
though approximate formula seems to be one given by Prof. Sturges known as
Sturges’srule, According to which K = 1+3.322 log10 N
Which K is the number of class intervals (classes) and N is the total frequency
i.e, total number of observations in the data. The value obtained is rounded to
the next higher integer.
Accordingly, the Sturges formula very ingeniously restricts the number of
classes between 4 and 20, which is fairly reasonable number from practical
point of view. The rule, however, fails if the number of observations is very
large or very small.
One has also to decide on the width of the class intervals (or size of class
intervals). In general, one should aim at class of equal width. Each class
intervals width, w, is obtained by w= , where R is the range, and K is the
number of class intervals. Another rule of the thumb for determining the rule
of the class interval should not be greater than th of the estimated population
standard deviation.

Examples: form a grouped frequency distribution from the following data by


inclusive method, using Sturges rule.
10, 17, 15, 11, 16, 19, 24, 29, 18, 25, 26, 32, 14, 22, 17, 20, 23, 27, 30, 12, 15, 18,
24, 36, 18, 15, 21, 28, 33, 38, 34, 13, 10, 16, 20, 22, 29, 19, 23, 31.
Example: a college management wanted to give scholarship to students
securing 60 and above marks in the following manner: 74, 62, 84, 72, 61, 83,
72, 81, 64, 71, 63, 61, 60, 67, 74, 66, 64, 79, 73, 75, 76, 69, 68, 78, 67. Calculate
the monthly scholarship paid to the students.

of marks Monthly
scholarshi
p

60-65 2,500

9
65-70 3,000

70-75 3,500

75-80 4,000

80-85 4,500

Cumulative Frequency Distribution:


A frequency distribution simply tells us how frequently a particular value of the
variable (class) is occurring. However, if we want to know the total number of
observations getting a value “less than” or “more than: a particular value of the
variable, this frequency table fails to furnish the information as such. This
information can be obtained very conveniently from the cumulative frequency
distribution.
Examples:
Less than c.f. of marks of 70 students less than c.f. distribution

Marks f Less than c.f

30-50 5 5

35-40 10 5+10=15

40-45 15 15+15=30

45-50 30 30+30=60

50-55 5 60+5=65

55-60 5 65+5=70

Marks F

10
Less than 30 0

Less than 45 30

Less than 50 60

Less than 55 65

Less than 60 70

More than c.f

Marks f More than thanc.f

30-35 5 65+5=70

35-40 10 55+10=65

40-45 15 40+15=55

45-50 30 10+30=40

50-55 5 5+5=10

55-60 5 5

more than frequency distribution


Marks F

Less than 30 70

Less than 35 65

Less than 40 55

Less than 45 40

11
Less than 50 10

Less than 60 0

Tabulation
By tabulation we mean the systematic presentation of the information
contained in the data, in rows and columns in accordance with some salient
features or characteristics. Rows are horizontal arrangements and columns
and vertical arrangements. In the word of A. M. Tuttle; “A statistical table is the
logical listing of related quantitative data in rows and columns with sufficient
explanatory and qualifying words, phrases and statements in the form of
titles, heading and notes to make clear the full meaning of data and their
origin.” Professor Bowley in his manual of statistics refers to tabulation as “the
intermediate process between the accumulation of data in whatever form
they are obtained, and the final reasoned account of the result shown by the
statistics.”
The various parts of a table include: the table number, title, head notes or
prefatory notes, captions and stubs, body of the table, foot-note, and source
note.

Diagrammatic and Graphic Representation


Another important, convincing, appealing and easily understood method of
presenting the statistical data is the use of diagrams and graphs. They are
nothing but geometrical figures like points, lines, bars, squares, circles, cubes,

12
etc. we shall consider the most commonly used once, they include: line
graphs, Bar charts, Pie charts, Histogram, Frequency polygon, and Ogive.
1. Line Diagram: This is the simplest of all the diagrams. It consists in
drawing vertical line being equal to the frequency. Line graph facilitate
comparisons.
Example: The following data shows the number of accidents sustained
by 314 drivers of a public utility company over a period of five years.
Number of accidents:
0 1 2 3 4 5 6 7 8 9 10 11
Number of drivers:
82 44 68 41 25 20 13 7 5 4 3 2
Represent the data by a line diagram.

2. Bar charts: Bar diagrams are one of the easiest and the most
commonly used devices of presenting most of the business and
economic data. These are especially satisfactory for categorical data or
series. The height (or length) of each bar indicating the size of the figure
represented, where the length of a single bar is proportional to the
magnitude of each part of the data. The sizes of the bars are the same.
The following are the various types of bar diagram in common use:
I. Simple bar chart
II. Sub- divided or component bar chart
III. Percentage bar chart
IV. Multiple bar diagram
V. Deviation or bilateral bar diagram

Example: use a simple bar chart to illustrate the number of

13
workers employed in the factories tabulated below.

Factory A B C D

No. of Employee 120 300 250 150

3. Angular or Pie Diagram: just as sub- divided and percentage bars or


rectangles are used to represent the total magnitude and its various
components, the circle (representing the total) may be divided into
various sections or segments viz., sectors representing certain
proportion or percentage of the various components parts to the total.
Such a sub- divided circle diagram is known as an angular or pie chart,
named so because the various segments resemble slices cut from a
[Link] degree represented by the various component parts of a given
magnitude can be obtained directly without computing their percentage
to the total value as follows:
Degree of any component part = 360

Examples: the following tables shows the area in millions of [Link]. of


ocean of the world:

Ocean Area (million [Link])

Pacific 70.8

Atlantic 41.2

Indian 28.5

Antarctic 7.6

14
Arctic 4.8

Draw a pie diagram to represent the data.

GRAPHIC REPRESENTATION OF DATA.


The difference between the diagrams and graphs is that; diagrams are useful
for visual presentation of categorical and geographical data while the data
relating to time series and frequency distribution is best represented through
graphs. Diagrams are primarily used for comparative studies and can’t be
used to study the relationship, (not necessary functional) between the
variables under study. This is done through graphs. The most commonly used
graphs for charting a frequency distribution for the general understanding of
the details of the data are:
i. Histogram
ii. Frequency polygon
iii. Frequency curve
iv. Ogive or Cumulative frequency Curve
Histogram: It is one of the most popular and commonly used devices for
charting continuous frequency distribution. It consist in erecting a series of
adjacent vertical rectangles on the sections of the horizontal axis ( x-axis),

15
with the bases (sections) equal to the width of the corresponding class
intervals and heights are so taken that the areas of the rectangles are equal to
the frequencies of the corresponding classes. The values are taken along the
x-axis and the frequencies along the y-axis. This however, involves two cases:
case (i) Histogram with equal classes, case (ii) Histogram with unequal
classes.
Example: represent the adjoining distribution of marks of 100 students in the
examination by a histogram

Marks obtained No. of students (c.f)

Less than 10 4

Less than 20 6

Less than 30 24

Less than 40 46

Less than 50 67

Less than 60 86

Less than 70 96

Less than 80 99

Less than 90 100

Example: represent the following data by means of a histogram

16
Weekly 10-15 15-20 20-25 25-30 30-40 40-60 60-80
wages

Frequency Polygon: frequency polygon is another device of graphic


presentation of a frequency distribution (continuous, grouped or discrete).
In case of discrete frequency distribution, frequency polygon is obtained on
plotting the frequencies on the vertical axis (y-axis) against the corresponding
values of the variables on the horizontal axis (x-axis) and joining the points so
obtained by straight lines.
Examples: the following data show the number of accidents sustained by 313
drivers of a public utility company over a period of 5 years.

No. of 0 1 2 3 4 5 6 7 8 9 10 11


accidents

No. of drivers 8 44 6 41 25 20 13 7 5 4 3 2


0 8

Draw the frequency polygon.


In case of grouped or continuous frequency distribution, frequency polygon
may be drawn in two ways. Case (1) from histogram, case (2) without
constructing histogram.
Example. The following table gives give the frequency distribution of the
weekly wages
(in ’00 N) of 100 workers in a factory.

Weekly 20-2 25-2 30-3 35-3 40-4 45-4 50-5 55- 60- Total
wages 4 9 4 9 4 9 4 59 64
( ’00 N

No of 4 5 12 23 31 10 8 5 2 100

17
workers

Draw the histogram and frequency polygon of the distribution.


Frequency Curve. A frequency curve is a smooth free hand curve drawn
through the vertices of a frequency polygon. The object of smoothing of the
frequency polygon is to eliminate, as far possible the random or erratic
fluctuations that might be present in the data. The area enclosed by the
frequency curve is same as that of the histogram or frequency polygon but its
shape is smooth one and not with sharp edges. Frequency curve may be
regarded as a limiting form of the frequency polygon as the number of
observations (total frequency) becomes very large and the class intervals are
made smaller and smaller.
Example Draw a frequency curve for the following distribution

Age (year) 17-19 19-21 21-23 23-25 25-27 27-29 29-31

No. of students 7 13 24 30 22 15 6

However, frequency curves are of different types. Some of the important


curves which in general, describe most of the data observed in practice are:
i. Curves of symmetrical distribution. In a symmetrical distribution, the
class frequencies first rise steadily, reach a maximum and then diminish
in the same identical manner. The most commonly and widely used
symmetrical curve in statistics is the normal frequency curve.
Normal Probability Curve

X = mean
ii. Moderately Asymmetrical (Skewed) Frequency Curves.

18
A frequency curve is
said to be skewed (asymmetrical) if it is not symmetrical. Such curves are
stretched more to one side than to the other. If the curve is stretched more to
the right (i.e, it has a longer tail towards the right), it is said to be positively
skewed and if it is stretched more to the left (i.e. has a longer tail towards
theleft), it is said to be negatively skewed.

iii. Extremely Asymmetrical or J-Shaped Curve.


The distribution in which the value of the variable correspondingly to the
maximum frequency is at one of the ranges, give rise to highly skewed curves.
When plotted, they give a J-shaped or inverted J-shaped curve and
accordingly such curves are also called J-Shaped curves.

J – Shaped Curve Inverted J – Shaped Curve

19
iv. U-Curve.
The frequency distribution in which the maximum frequency occurs at the
extremes (i.e, both ends) of the range and the frequency keeps on falling
symmetrically (about the middle), the minimum frequency being attained at
the center give rise to a U-Shaped curve.

U-shaped Curve
Bi
-
m
o
v.
d
Mixed Curves
Sometimes, though very rarely, we come across certain distributions in which
al
maximum frequency is attained at two or more points in an irregular manner.
Such curves are obtained in a distribution where as the value of the variable

C
increase, the frequencies increase and decrease, then again increase and
decrease twice or thrice as shown in the diagram or even more than that.

ur Tri-modal

ve 20
f

e
q
r

u
frequency

Variable
Variabl

OGIVE: Ogive is a graphic presentation of the cumulative frequency (C.F)


distribution of continuous variables. It consists in plotting the C.F (along the y-
axis) against the class boundaries (along x-axis). since there are two types of
cumulative frequency distribution viz, ‘less than’ C.F and ‘more than’ C.F. we
have accordingly two types of ogives, viz.,
i. ‘Less Than’ Ogive,
ii. ‘More Than’ Ogive
‘Less Than’ Ogive; this consists in plotting the ‘less than’ cumulative
frequencies against the upper-class boundaries of the respective classes. The
point so obtained are joined by the smooth freehand curve to give ‘Less Than’
Ogive. Obviously, ‘Less than’ Ogive is an increasing curve, sloping upwards
form left to right and has the shape of an elongated S.
‘More Than’ Ogive; Similarly, in ‘More Than’ Ogive, the ‘More Than’ cumulative
frequencies are plotted against the lower class boundaries of the respective
classes. The point so obtained are joined by a smooth freehand curve to give
‘More Than’ Ogive. ‘More Than’ Ogive is a decreasing curve and slopes
downwards from left to right and has the shape of an elongated S, upside
down
Remarks: we may draw both the ‘Less Than’ Ogive and ‘More Than’ Ogive on
the same graph. If done so, they intersect at a pint. The foot of the
perpendicular from their point of intersection on the x-axis gives the value of
median.
Example

The table below give the marks obtained by 70 candidates in STA 101

21
examination

Marks No. of Candidates (f) Less than C.F More than C.F

0-10 2 2 70

10-20 3 2+3 =5 70-2 =68

20-30 6 5+6 =11 68-3 = 65

30-40 11 11+11 =22 65-6 = 59

40-50 12 22+12 =34 59-11 = 48

50-60 15 34+15 =49 48-12 = 36

60-70 10 49+10 =59 36-15 = 21

70-80 7 59+7 = 66 21-10 = 11

80-90 4 66+4 =70 11-7 = 4

plot both less than and more than ogiv for the data.

22
AVERAGES
One of the important objectives of statistical analysis is to determine various
numerical measures which describe the inherent characteristics of a
frequency distribution. The first of such measures is average. The averages
are the measures which condense a huge unwieldy set of numerical data into
single numerical value which are representative of the entire distribution.
Averages are the typical values around which other items of the distribution
congregate. They are the values which lie between the two extreme
observation, (i.e., the smallest and the largest observations), of the
distribution and give us an idea about the concentration of the values in the
central part of the distribution. Accordingly they are also sometimes referred
to as the measures of central tendency. Averages are very much useful:
i) For describing the distribution in concise manner
ii) For comparative study of different distributions
iii) For computing various other statistical measures such as dispersion,
skewness, kurtosis and various other basic characteristics of a mass
data. Averages are also sometimes referred to as measures of
location since they enable us to locate the position or place of the
distribution in question.
“Statistical analysis seeks to develop concise4 summary figures which
describe a large body of quantitative data. One of the most widely used set of
summary figures is known as measures of location, which are often referred
to as averages, measures of central tendency or central location. The purpose
for computing an average value for a set of observation is to obtain a single
value which is representative of all the items and which the mind can grasp
simply and quickly. The single value is the point or location around cluster”
_______ Lawrence J. Kaplan
The following are the five measures of central tendency or measure of
location which are commonly used in practice
(i) Arithmetic mean or simply mean

23
(ii) Geometric mean
(iii)Harmonic mean
(iv) Median
(v) Mode
We shall discuss them in detail one by one.
Arithmetic mean (AR)
Arithmetic mean of a given set of observations is their sum divided by the
number of observations. For example, the AR of 5, 8, 10, 15, 24 and 28 is

In general, if X1, X2, ---, Xn are the given n observations, then their arithmetic
mean, usually denoted by is given by:
X=
Example: the following table gives the daily income of ten operators in a
machine tool factory. Find the mean.

Name of A B C D E F G H I J
operators

Income 12 15 18 20 25 30 22 35 37 26

Solution
If income is represented by x, then
=
In case of a discrete frequency distribution:

X - - -

- - -

24
The arithmetic mean is given by:
=
Illustration: The following is the frequency distribution of the number of
telephone calls received in 245 successive one-minute intervals at an
exchange:

No of Calls 0 1 2 3 4 5 6 7

One Minute 14 21 25 43 51 40 39 12
intervals

Obtain the mean number of calls per minute


Solution
Let the variable x denote the number of calls received per minute at the
exchange

No of 0 1 2 3 4 5 6 7
Calls (x)

Frequenc 14 21 25 43 51 40 39 12 ∑f=245
y (f)

Fx 0 21 50 129 204 200 234 84 ∑fx=245

= =3.8
ARITHMETIC MEAN FOR GROUPED DATA
In calculating the mean of a grouped data set it is customary to assume that
all values falling in a particular class interval are located at the class mark or
midpoint of that interval.
To calculate the mean of such a data set; we multiply each class mark by the

25
corresponding class frequencies, sum these products over all the class
intervals, and divide the results by the total frequency. Thus, the mean is
calculated as = i=1, 2, - - -, k, where xi represent the class mark
Illustration: The data below showing the frequency distribution of a lifetime of
720 television tubes:

Lifetimes in 0- 50 – 100 - 150 - 200 - 250 - 300 - 350 -


hours 49 99 149 199 249 299 349 399

No. of Tubes 5 16 29 128 206 314 17 5

Obtain the arithmetic mean.


Solution:

S/ Class Class Mark Fi Fi xi


No Interval (x)

1 0 – 49 24.5 5 122.5

2 50 – 99 74.5 16 1192.0

3 100 – 149 124.5 29 3610.5

4 150 – 199 174.5 128 22336.0

5 200 – 249 224.5 206 46247.0

6 250 – 299 274.5 314 86193.0

7 300 – 349 324.5 17 5516.5

8 350 – 399 374.5 5 1872.5

∑fi=720 ∑fixi=

26
167090.
0

=
Mathematical Properties of Arithmetic Mean
Arithmetic Mean possesses some very interesting and important
mathematical properties as given below:
Property 1: The algebraic sum of the deviations of the given set of
observations from their arithmetic mean is zero.
Mathematically, ∑ (x - ) = 0
Or for a frequency distribution: ∑f (x - ) = 0
Proof: ∑f (x - ) = ∑(fx - f) = ∑fx - ∑f
= ∑ f x - ∑f (: is a constant)
= ∑ f x - N (: ∑f = N)
But = ∑ fx = ∑ fx = N
: ∑ f (x - = N - N = 0
Property 2: The sum of squares of deviations of the given set of observations
is minimum when taken from the arithmetic mean compare to the sum of
squares of deviation of other value of the distribution.
Mathematically, for a given frequency distribution, the sum
2
S = ∑ f (x – A)
Which represent the sum of the squares of deviations of given observations
from any arbitrary value ‘A’ is minimum when A =
2
S1 = sum of squared deviations from mean = ∑ (x - ) , and
2
S = sum of squared deviations from any arbitrary point A = ∑ (X – A) : A

27
Then S1 is always less than S i.e, S1 <S
Property 3: The product of the arithmetic mean and number of values on
which the mean is based is equal to the sum of all given values. that is,
= fx
Property 4: Mean of the combined series
The mean of all the sum (or, differences) of corresponding observations in two
series, number of observations being equal in the two, is equal to the sum (or,
difference) of the means of the two series. If n1 and n2 are the sizes and 1 and 2
are the respective means of two groups then the mean of the combined
group of sizes n1 + n2 is given by:
= or 1,2 =
In general, if 1, 2, - - - k are the arithmetic means of k groups with n1, n2, - - -, nk
observations respectively, then
=
Illustration: The mean of marks in statistics of 100 students in a class was 72.
The mean of marks of boys was 75, while their number was 70. Find out the
mean marks of girls in the class.
Solution: In the usual notations we are given:
n1 = 70, 1 = 75; n1 + n2 = 100, = 72 : n2 = 100 – 70 = 30, we want 2
= = 72 =

2 = = = 65
Hence the mean of marks of girls in the class is 65.
Step Deviation Method for Computing Arithmetic Mean
It may be pointed out that the formula can be used conveniently if the values
of x or/and f are small. However, if the values of x or/and f are large, the
calculations of mean by is quite tedious and time consuming. In such a case
the calculation can be reduced to a great extent by using the step deviation (or

28
assumed mean) method which consists in taking the deviation (differences)
of the given observation from any arbitrary value A.
Let d = X – A, then fd = f (X – A) = Fx – A.F
Taking the sum over various values of x, we get
∑fd = ∑fx - A∑f
∑fd = ∑fx – A.N (:∑f = N)
: Dividing both sides by N, we get

=A+
In case of grouped or continuous frequency distribution, with class intervals of
equal magnitude, the calculations are further simplified by taking:
d = , where X is the mid-value of the class and h is the common magnitude of
the class intervals.
From d = , we get hd = (X – A)
Multiplying both sides by f, we get hfd = f (X – A) = fx – FA
Summing both sides over the values of X, we get:
h∑fd = ∑fx - A∑f = ∑fx – N.A
Dividing both sides by N, we get
h

Illustration: calculate the mean for the following frequency distribution:

Marks 0 - 10 10 – 20 20 – 30 30 – 40 40 – 50 50 – 60 60 – 70

No. of 6 5 8 15 7 6 3
Students

29
(i) By the direct formula (ii) By the step deviation method
Solution

S/N Marks Mid- F Fx d= fd


value (x)

1 0 - 10 5 6 30 -3 -18

2 10 – 20 15 5 75 -2 -10

3 20 – 30 25 8 200 -1 -8

4 30 – 40 35 15 525 0 0

5 40 – 50 45 7 315 1 7

6 50 – 60 55 6 330 2 12

7 60 – 70 65 3 195 3 9

∑f = 50 ∑fx = ∑fd = -8


1670

(i) = =
(ii) A = 35, h = 10
: =A+

Geometric Mean
The geometric mean usually abbreviated as (G.M) of a set of n observation is
th
the n root of their product. Thus if x1, x2, - - - xn are the given n observation then

30
their G.M is given by
1/n
G.M = = (X1. X2. - - -. Xn)
For example, the G.M of 4, 8, 16 is
G.M = = 8
th
But if n, the number of observations is large, then the computation of the n
root is very tedious. In such a case the calculations are facilitated by making
use of the logarithms. Taking logarithm of both sides

Taking anti log of both sides


G.M = Antilog
In case of frequency distribution
G.M. Antilog
Example: find the G.M. of 2, 4, 8, 12, 16, 24

Solution

X 2 4 8 12 16 24 Total

Log x 0.3010 0.6021 0.9031 1.0791 1.2041 1.3802 5.4697

G.M. = Antilog

Example: Find the G.M. for the following distribution.

Marks 0-10 10-20 20-30 30-40 40-50

31
No. of 5 7 15 25 8
Students

Solution

Marks Mid-point (x) F Log x F log x

0-10 5 5 0.6990 3.4950

10-20 15 7 1.1761 8.2327

20-30 25 15 1.3979 20.9685

30-40 35 25 1.5441 38.6025

40-50 45 8 1.6532 13.2256

∑f = 60 ∑ f log x =
84.5243

G.M. Antilog = 25.64 marks

HARMONIC MEAN
Another important mean is the harmonic mean which is used for averaging
the rates. If X1, X2, - - -, Xn is a given set of n observations, then their harmonic
mean (H.M) or simply H is given by:
H=
In other words, H.M. is the reciprocal of the arithmetic mean of the reciprocal
of the given observations. In case of frequency distribution, we have:

32
=H=
Example: The following table gives the weights of 31 persons on a sample
enquiry. Calculate the mean weight using harmonic mean

Weights 13 13 14 14 14 14 14 15 15
(lbs) 0 5 0 5 6 8 9 0 7

No. of 3 4 6 6 3 5 2 2 1
persons

Solution

Weights (lbs) (x) F Fx

130 3 0.0231

135 4 0.0296

140 6 0.0429

145 6 0.0414

146 3 0.0205

148 5 0.0338

149 2 0.0134

150 1 0.0067

33
157 1 0.0064

∑f = 31 ∑ = 0.2178

H.M. =
-1
Example: A cyclist pedals from his house to his college at a speed of 10 km h
-1
and back from the college to his house at 15km h . Find the average speed.
Solution
Let the distance from the house to the college be xkms. In going from house
to college, the distance (x kms) is covered in hours, while in coming from
college to house, the distance is covered in hours. Thus a total distance of 2x
kms is covered in hours.
Hence, average speed =
-1
H = 12 km h

MEDIAN
In the words of L.R. Connor:
“The median is that value of the variable which divides the group into two
equal parts, one part comprising all the values greater and the other, all the
values less than median”. Thus Median of a distribution may be define as that
value of the variable which exceeds and is exceeded by the same number of
observations i.e, it is the value such that the number of observations above it
is equal to the number of observations below it. Thus the median is a
positional average i.e, its value depends on the position occupied by a value in
the frequency distribution.
Calculation of Median
Case (1): Ungrouped Data: if the number of observations is odd, then the
median is the middle value after the observations have been arranged in

34
ascending or descending order of magnitude. For example, the median of 5
observations 32, 12, 40, 8, 60 i.e, 8, 12, 35, 40, 60 is 35.
In case of even number of observations median is obtained as the arithmetic
mean of the two middle observations after they are arranged in ascending or
descending order of magnitude.
For example 8, 12, 35, 40, 50, 60
The median =
Case (II): Frequency Distribution (Discrete type): In case of frequency
distribution where the variable takes the values X1, X2, - - -, Xn with respective
frequencies f1, f2, - - -, fn with , total frequency, median is the size of the item or
observation.
In this case the use of cumulative frequency distribution facilitates the
calculations.
Example: eight coins were tossed together and the number of heads (x)
resulting was noted. The operation was repeated 256 times and the frequency
distribution of the number of heads is given below:

No of heads 0 1 2 3 4 5 6 7 8
(x)

Frequency 1 9 26 59 72 52 29 7 1

Calculate the median


Solution

X F Less than
c.f.

0 1 1

1 9 9

2 26 36

35
3 59 95

4 72 167

5 52 219

6 29 248

7 7 255

8 1 256

∑f = 256

Here, ∑f = 256, = = 128. the c.f. just greater than 128 is 167 and the value of X
corresponding to 167 is 4. hence, median number of heads is 4.

Case (III): Continuous Frequency Distribution


The value of median is now obtained by using the interpolation formula:
median = l + , where
l is the lower limit of the median class, h is the magnitude or width of the
median class, N = , is the total frequency, f is the frequency of the median
class, c is the cumulative frequency of the class preceding the median class.
The interpolation formula is based on the following assumptions:
(i) The distribution of the variable under consideration is continuous with
exclusive type classes without any gaps.
(ii) There is an orderly and even distribution of observations with each
class.
However, if the data are given as a grouped frequency distribution where
classes are not continuous, then it must be converted into a continuous
frequency distribution before applying the formula. This adjustment will affect
only the value of l

36
Example: The following table shows the frequency distribution of weight in
grams of mangoes of a given variety. Calculate the median.

Weight in 410-419 420-429 430-439 440-449 450-459 460-469 470-47


grams 9

No. of 14 20 42 54 45 18 7


Mangoes

Solution

Class F Less than


boundaries c.f.

409.5 - 419.5 14 14

419.5 – 429.5 20 34

429.5 – 439.5 42 76

439.5 – 449.5 54 130

449.5 – 459.5 45 175

459.5 – 469.5 18 193

469.5 – 479.5 7 200

= = 100 The c.f. just greater than 100 is 130. Hence the corresponding class
439.5 – 449.5 is the median class.
Median = l +
= 439.5 +
= 443.94 grams

37
We can also get the median by plotting the ogive of the distribution, median is
the value below which item lie.
MODE
Mode is the value which occurs most frequently in a set of observations and
around which the other items of the set cluster densely. In other words, mode
is the value of a series which is predominant in it. In the words of Coxton and
Cowden, “The mode of a distribution is value at the point around which the
items tend to be most heavily concentrated. It may be regarded as the most
typical of a series of values”. According to A.M. Tuttle. ‘Mode is the value
which has the greatest frequency density in its immediate neighborhood’.
Illustration: The wheat yield in a particular region over the past 12 years (in
millions of tons) are: 1.5, 1.3, 1.2, 1.0, 1.3, 1.4, 1.6, 1.7, 1.5, 1.3, 1.2 and 1.4.
The mode is 1.3 (million tons).\
In case of a frequency distribution, mode is the value of the variable
corresponding to the maximum frequency. This method can be applied with
ease and simplicity if the distribution is ‘unimodal’. For example, in the
following distribution:

X 1 2 3 4 5 6 7 8 9

F 3 1 18 25 40 30 22 10 6

The maximum frequency is 40 and therefore, the corresponding value 5 gives


the value of mode. In case of a frequency curve mode corresponds to the
peak of the curve.

Fre
qu
enc
y

38
0 Mode x

In the case of continuous distribution, the class corresponding to the


maximum frequency is called the modal class and the value of mode is
obtained by the interpolation formula:

Mode = l + , where

l is the lower limit of the modal class,

f1 is the frequency of the modal class,

f0 is the frequency of the class preceding the modal class

f2 is the frequency of the class succeeding the modal class

h is the magnitude of the modal class.

Assumptions:

(i) The frequency distribution must be continuous with exclusive type


classes without any gaps.

(ii) The class intervals must be uniform throughout

Example: find the value of mode from the data given below:

Weight (in 93-97 98-102 103-10 108-11 113-11 118-12 123-12 128-13
kg) 7 2 7 2 7 2

39
No. of 3 5 12 17 14 6 3 1
students

Solution

Class boundaries f

92.5 – 97.5 3

97.5 – 102.5 5

102.5 – 107.5 12

107.5 – 112.5 17

112.5 – 117.5 14

117.5 – 122.5 6

122.5 – 127.5 3

127.5 – 132.5 1

Here maximum frequency is 17. The corresponding class 107.5 – 112.5 is the
modal class.

Mode = l + =

= 107.5 + = 110.625 kgs

40
EMPIRICAL RELATION BETWEEN MEAN (M), MEDIAN (MD), MODE (MO)

In case of a symmetrical distribution mean, median and mode coincide i.e.


mean = median = mode. However, for a moderately asymmetrical (non-
symmetrical or skewed) distribution, mean and mode usually lie on the two
ends and median lies in between them and they obey the following important
empirical relationship given by Prof. Karl Pearson.

Mode = mean – 3 (mean – median)

mean – mode = 3 (mean – median)

mean – median = (mean – mode)

The above relation can be exhibited diagrammatically as follows:

Divides area in halves

Under peak of curve

Centre of gravity

MO MD M

QUARTILES, DECILES AND PERCENTILES


We have defined the median as the value of items which is located at the
centre of the array, we can define other measures which are located at other
specified points.

41
Quartiles: The values which divide the given data into four equal parts are
known as quartiles. Obviously there will be three such points Q1, Q2 and Q3 such
that Q1 ≤ Q2 ≤ Q3, termed as the three quartiles. Q1, known as the lower or first
quartile is the value which has 25% of the items of the distribution below it and
consequently 75% of the items are greater than it. Incidentally Q2, the second
quartile, coincide with the median and has an equal number of observations
above it and below it. Q3, known as the upper or third quartile, has 75% of the
observations below it and consequently 25% of the observations above it. The
working principle for computing the quartiles is basically the same as that of
computing the median.
To computer Q1, find , where N = ∑f, see the (less than) c.f. just greater than ,
the corresponding value of x gives the value of Q1. In case of continuous
frequency distribution, the corresponding class containing Q1 and the value of
Q1 is obtained by the interpolation formula:
Q1 = l + , where all symbols has this usual meaning.
Similarly to compute Q3, see the (less than) c.f., just greater than , and for
continuous distribution, Q3 = l +
Deciles
Deciles are the values which divide the series into ten equal parts. Obviously
there are Nine deciles, D1, D2, D3, - - -, D9 (Say), such that D1 ≤ D2 ≤ D3 ≤ - - - ≤ D9.
Incidentally D5 coincides with the median. The method of computing the
deciles Di (i = 1, 2, 3, - - -, 9) is the same as discussed for Q1 and Q3. To compute
the ith decile = Di (i = 1, 2, 3, - - -, 9) see the c.f. just greater than . the
corresponding value of X is Di . In case of continuous frequency distribution the
corresponding class contains Di and its value is obtained by the interpolation
formula:
Di = l +
Percentiles:
Percentiles are the values which divide the series into 100 equal parts.
Obviously, there are 99 percentiles, P1, P2, P3, - - -, P99 such that P1 ≤ P2 ≤ P3 ≤ - - -≤
P99. The ith percentile Pi (i = 1, 2, 3, - - -, 99) is the value of X corresponding to

42
c.f. just greater than . In case of continuous frequency distribution, the
corresponding class contains Pi and its value is obtained by the interpolation
formula:
Pi = l +
In particular, we shall have:
P25 = Q1, P50 = D5 = Q2, P75 = Q3, D9 = P90, D1 = P10, D2 = P20
D3 = P30
The various partition values quartiles, deciles and percentiles can be easily
located graphically with the help of Ogive.
Example: The following data gives the distribution of marks of 100 students.
th th
Obtain the values of quartiles, 6 decile and 70 percentile

Class F Less than Class


c.f. boundaries

Less than 10 5 5 Below 9.5

10 – 19 8 13 9.5 - 19.5

20 – 29 7 20 19.5 - 29.5

30 – 39 12 32 29.5 - 39.5

40 – 49 28 60 39.5 - 49.5

50 – 59 20 80 49.5 - 59.5

60 – 69 10 90 59.5 - 69.5

70 – 79 10 100 69.5 - 79.5

∑f =
100

43
Quartiles: Q1 = =

The c.f. just greater than is 32. Hence the corresponding class is 29.5 – 39.5

Q1 = 29.5 +

The c.f. just greater than is 80. Hence, the class is 49.5 – 59.5 is the Q3 class

Q3 = 49.5 +
th
6 Deciles D6 =

D6 = 49.5 +
th
70 Percentile =

P70 = 49.5 +

DISPERSION

Averages or the measures of central tendency give us an idea of the


concentration of the observations about the central party of the distribution. In
spite of their great utility in statistical analysis, they have their own limitations.
If we are given only the average of a series of observations, we cannot form
complete idea about the distribution since there may exist a number of
distribution where averages are same but which may differ widely from each
other in a number of ways. Thus, the measures of central tendency must be
supported and supplemented by some other measures, one such measure is
‘Dispersion’ literal meaning of dispersion is “scatteredness”. We study
dispersion to have an idea of the homogeneity (compactness) or
heterogeneity (scatter) of the distribution. Dispersion is the measure of the
variation of the items.
The various measures of dispersion are:
(i) Range (ii) Quartile deviation or semi-interquartile range (iii)mean
deviation (iv) Standard deviation (v) Lorenz curve.

44
The first two measures, range and quartile deviation are termed a position
measures since they depend upon the values of the variables of particular
position of the distribution. The last measure, Lorenz curve is a graphical
method of studying variability.
1. Range: Range is the difference between the greatest (maximum) and
the smallest (minimum) observation of the distribution. Thus
Range = Xmax - Xmin
In case of a grouped frequency distribution (for discrete values) or the
continuous frequency distribution, range is defined as the difference
between the upper limit of the highest class and the lower limit of the
smallest class.

Coefficient of Range (Relative measure of range) =

Illustration: Calculate the range and the coefficient of range of A’s monthly
earnings for a year.

Month 1 2 3 4 5 6 7 8 9 10 11 12

Earning 139 150 151 151 157 158 160 161 162 162 173 175
(N1000)

Solution
L = 175000, 5 = 139000
Range = L – S = 175000 – 139000 = 36000
Coefficient of range =

Illustration: The following table gives the age distribution of a group of 50


individuals.

45
Age (in 16 – 20 21 – 25 26 – 30 31 – 35
years)

No of 10 15 17 8


Persons

Calculate range and the coefficient of range.

Solution
Convert into continuous classes. The first class will then become 15.5 –
20.5 and the last class will become 30.5 – 35.5
L = 35.5, 5 = 15.5
Range = L – S = 35.5 – 15.5 = 20 years
Coefficient of range =

2. Quartile Deviation or Semi Inter-Quartile Range


It is a measure of dispersion based on the upper quartile Q3 and the
lower quartile Q1.
Inter-quartile range = Q3 – Q1
Quartile Deviation (Q.D) =

Q.D as defined above is only an absolute measure of dispersion for


comparative studies of variability of two distributions we need a relative
measure which is known as efficient for Quartile Deviation and is given
by:

46
Coefficient of Q.D =

Percentile Range
This is a measure of dispersion based on the difference between certain
percentiles. If Pi is the ith percentile and Pj is the jth percentile then the so-
called i-j percentile range is given by i-j percentile Range = Pj – Pi (i < j).
Thus i – j semi-percentile Range is given by:
(Pj – Pi)/2, (i < j)
th
The commonly used percentile range is the one which corresponds to the 10
th
and 90 percentile. Thus,
10 – 90 percentile Range = P90 – P10 and
10 – 90 semi-percentile Range = (P90 – P10)/2.
The above measures are absolute measures only. The relative measure of
variability based on percentile is given by:
Coefficient of 10 – 90 percentile =

3. Mean Deviation or Average Deviation


As already pointed out, the two measures of dispersion discussed so far,
range and Q.D are not based on all the observations also they do not exhibit
any scatter of the observations. Also they do not exhibit any scatter of the
observation from an average and thus completely ignore the composition
of the series. Average Deviation overcomes both these drawbacks.
According to Clark and Schkade: “Average” deviation is the average amount
of scatter of the items in a distribution from either the mean or the median,
ignoring the signs of the deviations. The average that is taken of the
scatter is an arithmetic mean, which account for the fact that this measure

47
is often called the mean deviation”,
If X1, X2, - - -, Xn are n given observations then the mean deviation (M.D)
about an average A, say, is given by:
M.D =
Where =
Steps:
(i) Calculate the average A of the distribution by the usual method
(ii) Take the deviation d = X – A of each observation from the Average A.
(iii)Ignore the negative signs of deviation, taking all the deviation to be
positive to obtain the absolute deviation, = .
(iv) Obtain the sum of the absolute deviations obtained in step (iii)
(v) Divide the total obtained in step (iv) by n, the number of observation.
The result gives the value of the mean deviation about the average A. In case
of frequency distribution or grouped or continuous frequency distribution,
mean deviation about an average A is given by:
M.D. = , where x is the value of variable or it is the mid-value of the class
interval.
Relative Measures of Mean Deviation. The measures of mean deviation as
defined above are absolute measure depending on the units of measurement.
The relative measure of dispersion called the coefficient of mean deviation is
given by:
Coefficient of M.D. =
coefficient of M.D. about mean =
And coefficient of M.D. about median =
The coefficients of mean deviation defined above are pure numbers
independent of the units of measurement and are useful for comparing the
variability of different distribution.

48
Example: Calculate the mean deviation from the following data given marks
obtained by 11 students in a class test 14, 15, 23, 20, 10, 30, 19, 18, 16, 25, 12

49
Solution

M.D. =
Example: Calculate mean deviation from median of the following distribution.

Class 50 – 100 100 – 150 – 200 – 250 – 300 –


Interval 150 200 250 300 350

f 7 18 25 31 15 4

Also calculate the coefficient of mean deviation from median.


Solution

Less than C.F. Mid-value (X) = F

7 75 125 875

25 125 75 1350

50 175 25 625

81 225 25 775

96 275 75 1125

100 325 125 500

∑ F = 5250

Here
Median = 1 +
M.D. = about median =

50
Coefficient of M.D. =

4. Standard Deviation
Standard deviation, usual denoted by the letter (small sigma) of the Greek
alphabet was first suggested by Karl Pearson as a measure of dispersion in
1893. It is defined as the positive square root of the arithmetic mean of the
squares of the deviations of the given observations from the arithmetic mean.
Thus if X1, X2, - - -, Xn is a set of n observations then its standard deviation is
given by:
2
= , where
, is the arithmetic mean of the given values.
Steps:
(i) Compute the arithmetic mean
(ii) Compute the deviation (X - ) of each observation from arithmetic mean,
i.e., obtain X1 - , X2 - , - - -, Xn - .
2
(iii)Square each of the deviations obtained in step (ii) i.e., compute (X1 - ) ,
2 2
(X2 - ) , - - -, (Xn - .
(iv) Find the sum of the squared deviations in step (iii) and divide by n given
by:
2 2 2 2
∑f (X - ) /n = (X1 - ) + (X2 - ) + - - - + (Xn - .
(v) Take the positive square root of the value obtained in step (v)
(vi) The resulting value gives the standard deviation of the distribution.
In case of frequency distribution, the standard deviation is given by:
2
= , N = ∑f, X is the value of the variable or the mid-value of class (in case of
grouped or continuous frequency distribution); f is the corresponding
frequency of the value x.
Thus the value of will be greater if the values of X are scattered widely away

51
from the mean. Thus a small value of will imply that the distribution is
homogeneous and a large value of will imply that it is heterogeneous.
Variance and Mean Square Deviation
According to William I. Greenwald the variance is the mean of the squared
deviations about the mean of a series. Thus, variance is the square of the
2
standard deviation and is denoted by . For a frequency distribution variance is
given by:
2 2
= .
2
The mean square deviation, usually denoted by S is defined as
2 2
S = , where A is any arbitrary number.
The square root of the mean square deviation is called root mean square
2
deviation and given by: S =
2 2
Relation between and S . We have
2 2
S =
2
=
2 2
= +( +2
2 2
= + .
being constant is taken outside the summation sign.
∑f
2 2 2
S = +
2 2 2
so, S = + [
(, being the square of a real quantity is always non-negative.
2 2
Thus S = + (A non-negative quantity)
22
S
In other words, mean square deviation is not less than the variance or the root

52
mean square deviation is not less than the square deviation.
2 2
S = iff
2
( =O
so, ,
2
Thus, S will be least when = A. Hence, mean square deviation or equivalently
root mean square deviation is least when deviations are taken from the
arithmetic mean and variance (standard deviation) is the minimum value of
mean square deviation (root mean square deviation).
Different Formula:
2 2
● x = , ∑f = N
● 2 22 2 2
x = = -
If d = X – A, when A is an arbitrary constant, then
● 2 2 2 2
x = d = -
If we change the origin and scale in X i.e., if we take
d = ; h > O, then
2 22 2 2 2
● x =h d =h

● Coefficient of standard deviation = , coefficient of variation = x 100

Example, calculate the standard deviation of the frequency observation on a


certain variable:

240.12, 240.13, 240.15, 240.12, 240.17,

240.15, 240.17, 240.16, 240.22, 240.21.

solution

53
2
X X– (X -

240.12 - 0.04 0.0016

240.13 - 0.03 0.0009

240.15 - 0.01 0.0001

240.12 - 0.04 0.0016

240.17 0.01 0.0001

240.15 - 0.01 0.0001

240.17 0.01 0.0001

240.16 0.00 0

240.22 0.06 0.0036

240.21 0.05 0.0025

2
∑X = ∑ (X – ∑ (X - = 0.0106
2401.60

=240.16, variance=0.00106, sd=0.03256

Example: Calculate the mean and standard deviation from the following:

54
Value 90 – 99 80 – 89 70 – 79 60 – 69 50 – 59 40 – 49 30 – 39

Solution

2
Class Mid-value (X) f d= Fd fd

90 – 99 94.5 2 3 6 18

80 – 89 84.5 12 2 24 48

70 – 79 74.5 22 1 22 22

60 – 69 64.5 20 0 0 0

50 – 59 54.5 14 -1 -14 14

40 – 49 44.5 4 -2 -8 16

30 – 39 34.5 1 -3 -3 9

2
∑f = 75 ∑fd = ∑fd = 127
27

= 68.1

55
= h. = 12.505

coefficient of variation = x 100 = 18.36%

Skewness and Kurtosis

It has been pointed out that we need statistical measures which will reveal
clearly the salient features of a frequency distribution. The measures of
central tendency tells us about the concentration of the observations about
the middle of the distribution and the measure of dispersion gives us an idea
about the spread or scatter of the observations about some measure of
central tendency. We may come across frequency distributions which differ
widely in their nature and composition and yet may have the same central
tendency and dispersion, but yet may give histograms which differ very widely
in shape and size.

Thus the measures of central tendency and dispersion are inadequate to


characterize a distribution completely and they must be supported and
supplemented by two more measures; ‘skewness’ and ‘kurtosis’.

Skewness helps us to study the shape i.e., symmetry or asymmetry of the


distribution while kurtosis refers to the flatness or peakedness of the curve
which can be drawn with the help of the given data.

Skewness

Literal meaning of skewness is ‘lack of symmetry’. We study skewness to


have an idea about the shape of the curve which we can draw with the help of
the given frequency distribution. It helps us to determine the nature and
extent of the concentration of the observation towards the higher or lower
values of the variable. A distribution is said to be skewed if:

(i) The frequency curve of the distribution is not a symmetric bell-shaped


curve but it is stretched more to one side than to the other.

(ii) The values of mean, median and mode fall at different point, i.e., they
do not coincide.

56
(iii)Quartiles Q1 and Q3 are not equidistant from the median Q3 – md ≠ md
– Q1

(iv) The corresponding pairs of deciles and percentiles are not equidistant
from the median i.e.,

D5 – D5 – i ≠ D5+1 – D5 (i = 1, 2, 3, 4)

P50 – P50-i ≠ P50 + 1 – P50 (i = 1, 2, - - - 49)

(v) The sum of the positive deviations from the median is not equal to the
sum of the negative deviation from the median.

the following are coefficient of skewness which are commonly used:

1. Karl Pearson’s coefficient of skewness. This is given by the formula:

Skewness =

But quite often, mode is ill-defined and is thus quite difficult to locate. In
such a situation, we use the following empirical relationship between the
mean, median and mode for a moderately asymmetrical (skewed)
distribution.

MO = 3md – 2m

= skewness =

2. Bowley’s coefficient of skewness. Prof. A. L. Bowley’s coefficient of


skewness is based on the quartiles and is given by :

Skewness =

This is also known as quartile coefficient of skewness and is especially


useful in situation where quartiles and median are used.

3. Kelly’s measure of skewness. The drawbacks of Bowley’s coefficient of


skewness (that it ignores the 50% of the data towards the extremes),
can be partially removed by taking two deciles or percentiles equidistant

57
from the median value. The refinement was suggested by Kelly. Kelly’s
percentile (or decile) measure of skewness is given by:

Skewness = (P90 – P50) – (P50 – P10) = P90 + P10 – 2P50

But P50 = D5, P90 = D9 and P10 = D1. Hence

Skewness = (D9 – D5) – (D5 – D1) = D9 + D1 – 2D5

The kelly’s coefficient of skewness is given by:

Sk (Kelly) =

4. Coefficient of skewness based on moments. This coefficient is based


nd rd
on the 2 and 3 moment about mean.

KURTOSIS

So far we have studied three measures; central tendency, dispersion and


skewness to describe the characteristics of a frequency distribution. However,
even if we know all these three measures we are not in a position to
characterize a distribution completely. The following diagram will clarify the
point.

A – Lepto - Kurtic

B – Meso - Kurtic

C – Platy - Kurtic

58
Kurtosis is concerned with flatness or peakedness of the frequency curve.
Curve of type B which is neither flat nor peaked is known as normal curve and
shape of its hump (middle part) is accepted as a standard one. Curve with
humps of the form of a normal curve are said to have normal kurtosis and are
termed as meso-kurtic. The curve of type A, which is more peaked than the
normal curve are known as lepto-kurtic and are said to lack Kurtosis or to have
negative kurtosis. On the other hand, curve of type C, which are flatter than
the normal curve are called platy-kurtic and they are said to possess kurtosis
in excess or have positive kurtosis.

Errors and Approximations


In statistics, errors and approximations are critical for understanding the
reliability and accuracy of statistical estimates and predictions. Errors can
occur in various stages of data collection, analysis, and interpretation, while
approximations are often necessary to simplify complex calculations.
Types of Errors
Errors in statistics can be broadly classified into two categories: systematic
errors and random errors.
1. Systematic Errors: These are consistent, repeatable errors that occur due to
a flaw in the measurement system. They can arise from equipment
calibration, environmental conditions, or biases in data collection methods.
Example: If a weighing scale is incorrectly calibrated to weigh 1% heavy, all

59
measurements will consistently show weights that are 1% higher than their
actual weights.

2. Random Errors: These arise from unpredictable fluctuations in the


measurement process. They can be caused by human error, variations in the
environment, or other unforeseen factors.
Example: A person's weight may fluctuate slightly due to daily changes in
water intake, food consumption, or clothing.
Measurement of Errors
Errors can often be quantified using estimates such as:
- Absolute Error: The difference between the measured value and the true
value.

Absolute Error =

-Relative Error: The absolute error expressed as a fraction of the true value,
often expressed as a percentage.

Relative Error = ×100%

Approximations
In statistics, approximations are used to simplify complex calculations.
Common methods include rounding, truncating, and using linear
approximations.

1. Rounding
Rounding involves reducing the number of digits to make numbers easier to
work with.
Example: Round 3.567 to two decimal places:

3.567 ≈ 3.57

60
2. Truncation
Truncation eliminates digits beyond a certain point without rounding.
Example: Truncate 3.567 to two decimal places:

3.567 truncated to two decimal places ≈ 3.56.


3. Linear Approximation
Linear approximation uses the concept of tangent lines to approximate the
value of a function near a known point.
Example: To approximate f(x) = at x = 2 using linear approximation near x = 2:
1. Calculate the derivative:
f(x) = 2x

at x=2: f(2)=4

2. Use the linear approximation formula:

f(x) ≈ f(2) + f)2).(x-2(‫ﹶﹶ‬

[Link] x=2.1: f(2.1) ≈ 4+4.(0.1) = 4+0.4 = 4


2
The actual value f(2.1) = (2.1) = 4.41. The approximation is close, with a small
error.
Error Propagation
When combining measurements, it's important to understand how errors
propagate. The propagation of uncertainty can be calculated using the
following rules:
[Link] and Subtraction: When adding or subtracting values, the absolute
errors are summed.
z = x + y
2. Multiplication and Division: When multiplying or dividing values, the relative
errors are summed.

61
= +

Example of Error Propagation

Given 5±0.2 and y=3±0.1, find the error in z=x.y

Calculate the product: z=5.3=15,

Calculate the relative errors: = = 0.04, = ≈ 0.0333

Total relative error: = 0.04+0.0333 = 0.0733

Absolute error: z = z. = 15.0.0733=1.0995≈1.1

Result z = 15±1.1
Understanding errors and approximations is essential for accurate data
analysis in statistics. By recognizing the types of errors and how to propagate
them, statisticians can provide more reliable insights and improve the quality
of their conclusions. Always consider both types of errors when reporting
results to convey a complete picture of uncertainty.
Additional Practice Problems
1. If the measured length of a metal rod is 100.5 cm, and the true length is
100 cm, calculate the absolute and relative errors.
2. Use linear approximation to estimate f(x) = sin(x) at x = and x = +0.1
3. A car travels 150 km at a speed of 50 km/h with a possible error of ±2
km/h. Find the propagated error in the travel time.

Rates and Ratios in Descriptive Statistics


In the realm of descriptive statistics, rates and ratios play a pivotal role in
summarizing data and providing insights into relationships between different
quantities. Understanding how to calculate and interpret these statistical
measures is essential for analyzing data effectively. This lecture note defines
rates and ratios, provide examples, and discuss their applications in various

62
fields.
Definitions
1. Ratio: A ratio is a quantitative relationship between two numbers,
showing how many times one value contains or is contained within the
other. It can be expressed in the form of a fraction, using a colon, or as a
decimal.
Ratio = , where (A\) and (B\) are the two quantities compared.
Example: If there are 20 boys and 30 girls in a classroom, the ratio of boys to
girls can be calculated as: = . This means that for every 2 boys, there are 3
girls in the classroom.
2. Rate: A rate is a specific kind of ratio that compares two quantities of
different units, usually expressing one quantity per unit of another. Rates often
provide context and can indicate a frequency, intensity, or prevalence.
Rate =
Example: If a car travels 150 miles in 3 hours, the rate of speed can be
calculated as:
Rate of speed = .=50 miles per hour

Applications of Ratios and Rates


1. Finance
Example: In finance, the debt-to-equity ratio is a common measure that
indicates the relative proportion of shareholders' equity and debt used to
finance a company's assets.
Debt-to-Equity Ratio = .
If a company has $200,000 in debt and $100,000 in equity, the ratio is:
Debt-to-Equity Ratio = .=2.
This means the company has $2 of debt for every $1 of equity.

63
2. Health
Example: Incidence rate is a measure used in epidemiology to describe the
frequency of new cases of a disease in a particular population over a specific
period.
If 50 new cases of a disease are reported in a population of 10,000 over one
year, the incidence rate is calculated as:
Incidence Rate = x100 = 0.5%
This indicates that 0.5% of the population developed the disease during that
year.
3. Demographics
Example: The birth rate is an example of a demographic rate, calculated as
the number of live births in a year per 1,000 people in the population.
If a country has 10,000 live births in a year and a population of 1,000,000, the
birth rate is:
Birth Rate = x1,000= 10 births per 1,000 people.
This means there are 10 births for every 1,000 individuals in the population.
Properties of Ratios and Rates
- Ratios can be simplified to their lowest terms, which can help facilitate
comparison.
- Rates are crucial in comparing different scales and sizes, making them useful
for standardized measurements.
- Understanding the context of the rates and ratios is essential, as they can
mislead if the underlying data changes.
Rates and ratios are fundamental statistical tools that help interpret and
summarize data efficiently. By comparing different quantities, they provide
valuable insights across various fields, including finance, health, and
demographics. Mastery of calculating and interpreting these measures
permits better decision-making and more profound understanding of the data

64
trends and patterns.
Practice Problems:
1. Calculate the unemployment rate if 300 out of 10,000 people in the
workforce are unemployed. Provide your answer per 1,000 people.
2. You are given that in a town of 5,000 individuals, 1,250 are children, and the
rest are adults. Compute the ratio of children to adults.

INDEX NUMBERS
Definition:
Index numbers are statistical devices designed to measure the relative
changes in the level of a phenomenon (variable or a group of variables) with
respect to time, geographical location or other characteristics such as income,
profession, etc. in other words, Index numbers are specialized types of rates,
ratios, percentages which give the general level of magnitude of a group of
distinct but related variables in two or more situations. Index numbers is a
statistical device which enables us to arrive at a single representative figure
which gives the general level of the price of the commodities in an extensive
group.

65
USES OF INDEX NUMBERS
The first index number was constructed by an Italian, Mr. Carli, in 1764 to
compare the changes in price for the year 1750 (current year) with the price
level of in 1500 (base year) in order to study the effect of discovery of America
on the price level in Italy. Though originally designed to study the general level
of prices or accordingly purchasing power of money, today index numbers are
extensively used for a variety of purposes in economics, business,
management, etc. and for quantitative data relating to production,
construction, consumption, profits, personnel and financial matters, etc. for
comparing changes in the level of phenomenon for two periods, places, etc.,
there is hardly any field of quantitative measurements where index numbers
are not constructed, they are used in almost all sciences. The main uses of
Index numbers can be summarized as follows:
1. Index numbers serve as economic barometer
2. Index numbers help in studying trends and tendencies
3. Index numbers help in formulating decisions and policies.
4. Price index measure the purchasing power of money.
5. Index numbers are used for deflation

TYPES OF INDEX NUMBERS


Index numbers may be broadly classified into various categories depending
upon the type of the phenomenon or variable in which the relative changes are
to be studied. It can be broadly classified into the following three categories:
1. Price Index Number: this measures the general changes in the prices.
They are further sub-divided into the following classes
a. Wholesale Price Index Numbers: this reflects the changes in the
general price level of a country.

66
b. Retail Price Index Numbers: These indices reflect the general
changes in the retail prices of various commodities such as
consumption of goods, stocks and shares, bank deposits,
government bonds, etc.
2. Quantity Index Numbers: this study the changes in the volume of goods
produced (manufactured), consumed or distributed, like the indices of
agricultural production, industrial production, imports and exports, etc.
they are extremely helpful in studying the level of physical output in an
economy.
3. Value Index Numbers: these are intended to study the change in the
total value (Price multiplied by quantity) of production such as indices of
retail sales or profits or inventories. However, these indices are not as
common as price and quantity indices.

Consumer Price Index


Consumer Price Index (CPI), commonly known as the cost of living index is a
specialized kind of retail price index and enables us to study the effect of
changes in the prices of a basket of goods or commodities on the purchasing
power or cost of living of a particular class or section of the people.

Selection of base Period


Base period is the period selected for comparisons of the relative changes in
the level of a phenomenon from time to time. The index for base period is
always taken as 100. The following points in conforming with the objectives of
the index should serve as guideline for selecting a base period.
i. Base period should be a period of normal and stable economic
conditions.

67
ii. The base period should not be too distant from the given period.
NOTATION AND TERMINOLOGY
Base Year: The year selected for comparison i.e. the year with respect to
which comparison are made. It is denoted by the suffix zero ‘0’.
Current Year: The year for which comparisons are sought or required. It is
denoted by the suffix ‘1’.
P0: Price of commodity in the base year.
P1: Price of commodity in the current year.
q0: Quantity of a commodity consumed or purchased during the base year.
q1: Quantity of a commodity consumed or purchased in the current year.
w: Weight assigned to a commodity according to its relative importance in the
group.
I: simple index number or price relative obtained on expressing current year
price as a percentage of the base year price and is given as
I = Price Relative =
P01: Price Index Number for the current year with respect to the base year.
P10: Price Index Number for the base year with respect to the current year.
Q01: Quantity Index Number for the current year with respect to the base year.
Q10: Quantity Index Number for the base year with respect to the current year.
V01: Value Index for the current year with respect to the base year.
METHODS OF CONSTRUCTING INDEX NUMBERS
We shall now discuss the various techniques or methods used for the
construction of index numbers. The price indices is the most important of all
the indices, we shall discuss their construction in detail. The quantity indices
can be obtained from price indices by interchanging the price (p) and quantity
(q) in the final formula.
1. Simple (Unweighted) Aggregate Method

68
This is the simplest of all the methods, it consists in expressing the total price
in the current year as a percentage of the aggregate of prices in the base year.
Thus
P01 =
The quantity index is given by
Q01 =
Example:
From the following data calculate Index Number by simple Aggregate method.
Commodity A B C D
Price in 1980 162 256 25 132
7
Price in 1981 171 164 18 145
9

(P0) = 807
(P1) = 669

2. Weighted Aggregate Method


In this method, appropriate weights are assigned to various commodities to
reflect their relative importance in the group. For the construction of the price
index numbers, quantity weights are used, i.e. the amount of quantity
consumed, purchased or marketed. If W is the weight attached to a
commodity, then the price index is given by
P01 = *

69
By using different systems of weighting we get a number of formulae, some of
the important are given below
i. Laspeyre’s Price Index or Base Year Method: Taking base year quantities as
weights i.e. w = q0 in the above formula (*), we get the Laspeyre’s price Index
as:
La
P01 = 100
This formula was devised by French Economist Laspeyre in 1817.

ii. Paasche’s Price Index: if we take current year quantities as weights in (*),
we obtain paasche’s price index which is given by:
Pa
P01 = 100
This formula was given by German Statistician Paasche in 1874.
iii. Dorbish – Bowley Price Index: This index is given by the arithmetic mean of
Laspeyre’s and Paasche’s price index numbers and we have:
DB
P01 = 100
iv. Fisher’s Price Index: Irving Fisher advocated the geometric cross of La and
Pa price index numbers and is given by:
F La Pa
P01 = [ P01 P0 ] = 100
Fisher’s index is termed as an ideal index since it satisfies time reversal and
factor reversal tests for the consistency of index numbers.
v. Marshall-Edgeworth Price Index: taking the arithmetic cross of the
quantities in the base year and the current year as weights i.e. w = , we obtain
the Marshal-Edgeworth (M.E.) formula given by:
ME
P01 =
=
vi. Walsch Price Index: instead of taking the arithmetic of base year and
current year quantities as weights, if we take their geometric mean, i.e., w = ,

70
then we obtain Welsch Index given by:
Wa
P01 =
vii. Kelly’s Price Index or Fixed Weight Index: this formula, named after Truman
L. Kelly, requires the weights to be fixed for all periods and is also sometimes
known as aggregative index with fixed weights and is given by the formula:
K
P01 =
Where the weights are the quantities (q) which may refer to some period (not
necessarily the base year or the current year) and are kept constant for all
periods.

QUANTITY INDICES (Q)


As already pointed out, Q numbers reflect the relative changes in the quantity
or volume of goods produced, consumed, marketed or distributed in any given
year with respect to some base year.
La
Q01 =
F
Q01 = 100
Pa
Q01 =
ME
Q01 =

VALUE INDICES
Value Index numbers are obtained on expressing the total value (or
expenditure) in any given year as a percentage of the same in the base year.
Symbolically, we write
V01 =
V01 =

71
EXAMPLE
a. From the following data, calculate price index numbers for 1980 with
1970 as base year by
i. Laspeyre’s method
ii. Paasches’s method
iii. Marshall-Edgeworth method, and
iv. Fisher’s Ideal method
b. It is stated that Marshall-Edgeworth index number is a good
approximation to Fisher’s Ideal index number. Verify this for the data.

Commodities 1970 1980

Price Quantity Price Quantity

A 20 8 40 6

B 50 10 60 5

C 40 15 50 15

D 20 20 20 25

La = 124.699, Pa = 121.77, ME = 123.32, F = 123.23


ME F
Since P01 = 123.32 and P01 = 123.32, are approximately equal, ME index
number is a good approximation to Fisher’s index number.

72
SIMPLE AVERAGE OF PRICE RELATIVES
In this method, first of all we obtain the price relatives for each commodity.
The price relatives are obtained by expressing the price of the commodity in
the current year as a percentage of its price in the base year. i.e.
P = Price relative for a commodity =
Price-relatives are the simplest form of the index numbers for each
commodity. The price index for the composite group is obtained on averaging
these price-relatives by using some suitable measures of central tendency,
usually arithmetic mean (A.M.) or geometric mean (G.M.). Price Index using
simply arithmetic mean of relatives is given by:
P01 (A.M.) =
Where n is the number of commodities in the group.

Using simple geometric mean of the price relatives, the price index is given by:
P01 (G.M.) = =
Where Π denotes the product of the price-relatives for the n commodities. To
evaluate this, we use logarithms. Taking logarithms of both sides, we get
LogP01 (G.M.) =
P01 (G.M.) = Antilog

MERITS AND DEMERITS


The index number based on the single average of the price-relatives
overcomes some of the drawbacks of the ‘simple aggregate method, viz.,
i. Price-relative are pure numbers independent of the units of
measurement and hence the index number based on their average is
not affected by the units in which the prices are quoted.
ii. The extreme observations (large and small quotations) do not

73
influence the index unduly. It gives equal importance to all observations.
The drawback of this method is that it gives equal weights to all the
commodities and thus neglects their relative importance in the group. This
drawback is removed by taking the weighted average of the price-relative.

Remark: the distribution of the price-relative is found to be positively skewed


and the skewness increases as the base is shifted more and more away from
the given year.

74
Example: Construct Index number for each year from the following average
wholesale prices of cotton with 1993 as base
Year Price Year Price

1993 75 1998 70

1994 50 1999 69

1995 65 2000 75

1996 60 2001 84

1997 72 2002 80

Example: The following are the prices of commodities in 1995 and 2000.
Calculate a price index based on price-relative using the arithmetic mean as
well as geometric mean.
Year Commodity

A B C D E F

1995 45 60 20 50 85 120

2000 55 70 30 75 90 130

75
WEIGHTED AVERAGE OF PRICE RELATIVE
The shortcoming of Simple Average of Relatives method which assumes that
all the relatives are equally important is overcome in this method which
consists in assigning appropriate weights to the relatives according to the
relative importance of the different commodities in the group. Thus, the index
for the whole group is obtained on taking the weighted average, usually A.M.
or G.M. of the price relatives. Thus, based on weighted A.M., the price index is
given by:
P01 (A.M) = =
Where W is the weight attached to the price-relative P.
Steps:
1. Find the price-relative (p) for each commodity, i.e., compute P =
2. Multiply the price relatives in step 1 by the corresponding weights (W)
assigned to get the product WP.
3. Obtain the sum of products obtained in step 2 for all the commodities
to get .
4. Divide the sum in step 3 by, the total of the weight assigned.
The resulting figure gives the price index based on the weighted average of
price-relatives.
The price index based on the weighted geometric mean of price relatives is
given by
P01 (weighted G.M.) =
Taking logarithms of both sides, we get
log [P01 (weighted G.M.)] =
P01 (weighted G.M.) = Antilog

Steps:

76
1. Compute the price-relatives P = , for each commodity.
2. Find the logarithms of all the price relatives, log P.
3. Multiply log P values for each commodity by the corresponding weights
(W) assigned. This will give ([Link] P) values.
4. Find the sum of the values in step 3 over all the commodities to get .
5. Divide the sum obtained in step 4 by , the sum of weights.
6. Antilog of the value obtained in step 5 gives required price-index.
Example: the following table gives the prices of some food items in the base
year and current year and the quantities sold in the base year. calculate the
weighted index number by using the weighted average of price relatives.
Items Base year Quantities Base Year Price Current year Price

A 7 18.00 21.60

B 6 3.00 4.65

C 16 7.50 9.00

D 21 2.50 2.25

This lecture notes is not for sale.

77

You might also like