Quantitative Methods
Quantitative Methods
2
3
1: INTRODUCTION TO STATISTICS
Today our culture has become a statistical culture because every citizen finds statistics in
newspaper, magazines, advertisement in radios and televisions etc.
The figures relating to the various aspects of his life social political and economic
represent and support the absolute facts and situations. The reader analysis the figures and
arrives at certain conclusion.
EXAMPLE:
A prospective investor is interested in the profitability of the business enterprise and his
analysis the income statement and balance sheet of the enterprise for a number of years. Such
analysis helps him to decide whether to invest or not.
DEFINITION OF STATISTICS.
(b)Plural sense.
DEFENITION 2
By Prof Horace Secrist: Statistics may be defined as the aggregate of facts affected to a
marked extent by multiplicity of causes, numerically expressed enumerated or estimated
according to a reasonable standard of accuracy, collected in a systematic manner, for a
predetermined purpose and placed in relation to each other.
DEFENITION 3
COLLECTION:
It is the first step in statistical investigation. The collection part is the backbone of the enquiry.
If the of data is not in proper form in that case conclusion drawn can never be reliable.
4
The source of collection may be primary or secondary data.
Primary data: If the data are collected originally by the investigator for the given enquiry it
is termed as primary data.
Secondary data: if he makes use of the data which had been earlier collected by someone
else, it is termed as secondary data.
PRESENTATION: After collection and organizing the data should be presented. If the
data are presented systematically, the statistical analysis gets facilitated. As far as the
presentation of data is concerned either through diagrams or graphs. The classified data is to be
presented in such a way that it becomes easily understandable.
ANALYSIS :
Once the data is collected, presented the next step is that of analysis is to prepare data in such a
fashion so as to arrive at certain definite conclusion.
2. Measure of dispersion .
INTERPRETATION:
The last stage of statistical investigation is to drive the results and give comments on
the enquiry in question. Interpretation is to draw conclusion from the data collected and
analyzed. If the analyzed data is not properly interpretated the whole object of the enquiry may
be erroneous.
USES OF STATISTICS:
Every field of human activity makes use of statistics. Business houses, Government,
departments public sector undertaking, social reformers, economist etc, all use statistics.
The important uses of statistics to different sectors of the society are given below.
1. Statistics is the arithmetic of human welfare because the problems relating to unemployment,
food storage, poverty etc, cannot be analyzed without the use of statistics.
2. Statistics are great use to business man. A business man can estimate demand for his
product in different parts of the year correctly with the help of statistics.
2. It is useful to insurance companies because the insurance companies need exact information
of vital statistics to prepare mortality table and premium plans.
3. Statistics is very useful to bankers and banking industry in deciding the policies regarding
deposits advances etc.
5
4. Statistics analysis of railway working is very useful in determining the efficiency of services
their requirements and their expansion programmes.
5. Statistics is very useful to finance ministers who prepare annual budget and outlines the
physical policy on the basis of statistics collected in the various departments.
6. The educationist makes use of statistics for evaluating the performance of students and
teacher.
7. Statistics helps the administrator to know the social and economic ambitions of the society.
8. It is very much useful to the economists who prepares economic planes and frame economic
policies for the state.
9. It is useful to social reformer who has to carry on his activities based on statistics. The
incident of certain diseases, seriousness of beggar problem, child marriage etc, all need
statistics for their solution.
OBJECTS OF STATISTICS:
1. To throw light on the complex man data and make sense from them.
2. To take action on the data processed.
3. To draw conclusion on the analyzed data.
4. To estimate and forecast the trends in future from the past data.
5. To ascertain the reliability of statements.
6. To assist the research studies in gaining knowledge.
7. To prove unknown from the known data.
8. To examine change in the economic activities.
FUNCTIONS OF STATISTICS:
LIMITATION OF STATISTICS:
Science of statistics suffers from various limitations. If the limitations are not kept in mind the
conclusions reached will not be dependable. The following are the main limitations of statistics.
6
2. Quantitative aspect ignored:
Statistical methods cannot study the nature of phenomenon which cannot be
expressed in quantitative form. The phenomenon which cannot be a part of the study of
statistics. These characteristics include health intelligence, beauty, honesty etc. there is no
doubt that the data which cannot be quantitatively expressed needs conversion of qualitative
data into quantitative data then we can easily apply the statistics techniques on this converted
data.
DISTRUST OF STATISTICS:
The improper use of statistical tools by unscrupulous people with an improper statistical bend
of mind has led to the public distrust in statistics. By this we mean that public loses its belief,
faith and confidence in the science of statistics and starts condemning it. Such irresponsible,
inexperienced and dishonest persons who use statistical data and statistical techniques to fulfill
their selfish motives have discredited the science of statistics with some very interesting
comments, some of which are stated below:
7
a) Figures are innocent and believable and the facts base on them are psychologically more
convincing. By it is a pity that figures do not have the label of quality on their face.
b) Arguments are put forward to establish certain result which are not true by making use of
in-accurate figures or by using incomplete data. This destroying the truth.
c) Though accurate, the figures might be molded and manipulated by dishonest and
unscrupulous persons to conceal the truth and present a working and distorted picture of the
facts to the public for personal and selfish motives.
Hence if statistics and its tools are misused the fault does not lies
with the science of statistics rather, it is the people who misuse it, are to be blamed.
8
COLLECTION AND TABULATION OF DATA.
Sources of Data
1) Primary source.
2) Secondary source.
Primary Source:
The data which are originally collected for the first time are called Primary
Data.
“Census” or “sample” technique is adopted in collecting the primary data. The collected raw
data is edited and scrutinized by detecting the errors and omitting the irrelevant and
inconsistent information.
iv. Questionnaires:
A questionnaire is a list of logically arranged questions relating to the field of
enquiry and providing space for the answer to be filled in by the respondents.
v. Schedules:
9
Secondary sources are sources of data which consists of readily
available compendia and already compiled information. They consist of published records,
reports and unpublished but already used records.
2) Unpublished Sources:
The data which are not available in printed form. There are various data
available which are unwritten in any format.
Census Technique:
When each and every unit of population under study is considered and
collected for a statistical investigation.
Sample Technique:
The terminology “sampling” means the process of selecting a part of the
population under study with a view to obtaining information about the whole population.
Classification of Data.
Classification is the process of arranging the data into sequences and groups according to their
common characteristics them in to different but related parts.
-by- Horace Secrist.
10
Qualitative Classification: Data are classified according to attribute or characteristics or
qualities. Generally the qualitative phenomena are not measurable. However, they can be
studied with reference to their presence or absence like educated or uneducated, married and
unmarried, boys and girls so on.
Simple classification - It involves only 2 fold (two-fold).
Arbitrary-> It refers to the attribute, which are not clearly defined and which no demarcation
can be made. Ex:- tall persons and short persons . we should give the correct height above
which the persons are considered tall and below which they are short.
Quantitative Classification.
Data are classified according to quantities that are measurable such as
age, weight, marks, price and so on.
11
Tabulation
A statistical table is the logical listing of related quantitative data is vertical columns and
horizontal rows of number with sufficient explanatory and qualifying words, phrases and
statements in the form of titles, heading and footnotes to make clear the full meaning of data
and their origin.
-by Tuttle.
Types of Tabulation.
1. Simple Tabulation (one-way):-
Below 30 15
30-40 18
40-50 9
50 and 3
Above
12
2. Double Table:-
30-40 10 8 18
40-50 2 7 9
50 and Above - 3 3
Total 21 24 45
3. Complex Table:- Table showing number of employees in Canara Bank according to age, sex
and marital status.
Below 30 3 6 9 2 4 6 15
30-40 2 8 10 4 4 8 18
40-50 1 1 2 7 - 7 9
50 and - - - 3 - 3 3
Above
Total 6 15 21 16 8 24 45
Ex:-
In a sample study about coffee habit in 2 town the following
information was received.
13
Town-A : Females were 40%, Total coffee drinkers were 45%
and Males non-coffee drinkers were 20%.
Town-B : Males were 55% Males non-coffee drinkers were
30% and females coffee drinkers were 15%.
Represent the above in tabular form.
Town A B
Coffee drinker 40 5 45 25 15 40
Non coffee 20 35 55 30 30 60
drinker
DIAGRAMS
Diagrams are visual aids of presenting the data in pictures, geometric figures
and curves. They present huge mass of quantitative data in a condensed form
attractively.
OBJECTS:
1) TITLE: Every diagram and graph must be given a suitable title. The title
should communicate and details of data presented in that diagram.
2) SCALE: As a guiding principle, the scale should be selected consistent with
the size of the paper and nature of the data to be displayed, so that the diagram obtained
is either too small or too big.
3) WIDTH AND HEIGHT: A proper proportion between the width and height
of the diagram should be maintained to make it look attractive and understandable.
4) INDEX: An index explaining different types of shades colours used should be
given.
5) NEAT AND SIMPLE: To be effective, the diagrams should be prepared
neatly and they should be simple to understand.
USES:
14
LIMITATIONS:
DIAGRAMS GRAPHS
They can be drawn on only paper- They can only be drawn on graph
plane (or) graph. papers.
Lines, bars, rectangles, circles, cubes, Dots, dashes, dot-dashes and curves
pictures are used in diagrams. are used in graphs.
They furnish information They furnish information in detail.
approximately. They depict time series and
They depict categorical and frequently distribution.
geographical data. They are easy to draw.
They are not easy to draw.
15
2: MEASURES OF CENTRAL TENDENCY
REQUISTIES OF A GOOD AVERAGE:
ARITHMETIC MEAN
It is the quantity obtained by dividing the sum of the values of the items in a variable
by the number of items.
PROBLEM:
data .
No of workers: 8 3 11 14 5 7 2
72-80 14 76 1064 0 0 0 0
80-88 5 84 420 8 40 1 5
96- 2 200 24 48 3 6
104 100
n=∑f=50 3672 -128 -16
16
1 .DIRECT METHOD:
2. SHORTCUT METHOD:
MERITS:
DEMERITS:
1) IT is very much affected by extreme value
2) Not useful in case of open-end classes.
3) It is not always reliable.
4) It may give wrong impression, if proper weights are not given.
MEDIAN:
Median is the value which divides the given series into two equal parts.
INDIVIDUAL OBSERVATION:
25 36 45 68 75
𝒏+𝟏 𝒕𝒉
Median is the size of( ) item
𝟐
5+1
n=5 = 3 item , so
2
Median is 45
DISCRETE SERIES
Problem:
Calculate median from the following data.
Value: 10 20 30 40 50 60.
Frequency: 28 36 24 32 40 16
17
Values Frequency(f) Cumulative
frequency(c.f)
10 28 28
20 36 64
30 24 88
40 32 120
50 40 160
60 16 176
176
𝒏+𝟏 𝑡ℎ
Median is the size of( ) item
𝟐
CONTINUOUS SERIES
Problem:
Calculate median from the following data.
Value: 10-20 20-30 30-40 40-50 50-60 60-70.
Frequency: 28 36 24 32 40 16
Me= l + i/f(n/2 - m)
= 40+ 10/32(88-88)
𝑀𝑒 is 40
18
MERITS OF MEDIAN:
DEMERITS OF MEDIAN:
MODE:
Mode is the value which occurs more frequently in a set of observation and around which
the other items of the set of observations clustered densely. It is denoted by „Z‟.
INDIVIDUAL OBSERVATION:
The mode is 64, as it occurs the highest number of times in the series. So the mode is clearly
defined as it is a uni - modal case.
DISCRETE SERIES:
A Frequency
SERIES
X F
1 2
2 8
3 11
4 18
5 9
6 7
Highest frequency is 18.
19
So, the Mode is 4.
CONTINUOUS SERIES
Problem:
Calculate mode.
f: 1 3 10 6 11 9 1
behind 𝑓1 is 𝑓0 i.e.,
𝑓0 = 6
next to 𝑓1 is 𝑓2 i.e.,
𝑓2 = 9
𝑓1− 𝑓2
Z= l+ ×𝑖
2𝑓1− 𝑓2 −𝑓0
Z= 20+ 11−6 ×5
2(11)−6−9
Z=23.5
f: 1 3 10 6 10 9 1
X 𝑓1 𝑓2 𝑓3 𝑓4 𝑓5 𝑓6
0-5 1 4 13
5-10 3 13 19
10-15 10 16 26
15-20 6 16 25
20-25 10 19 20
25-30 9 10
30-35 1
20
ANALYSIS TABLE:
Z= 20+ 10−6 ×5
2(10)−6−9
Z=24
Problem
From the following data find out the missing frequency if the median is
50.
Frequency: 2 8 6 - 15 10
Sol:
C.I f cf
10-20 2 2
20-30 8 10
30-40 6 16
40-50 A 16+A
50-60 15 31+A
21
60-70 10 41+A
41+A
Me= 50
CI is 50-60
Me= l+i/f[n/2-m]
50-50=0.667[ 41+𝐴−2(16+𝐴)
2
]
0= 5[9-A]
A=9
Problem
Class intervals: 10-20 20-30 30-40 40-50 50-60 60-70 70-80 total:229
Frequency: 12 30 - 65 - 25 18
Sol:
CI f cf
10-20 12 12
20-30 30 42
30-40 A 42+A
22
40-50 65 107+A
60-70 25 211
70-80 18 229
Total 229
12+30+A+65+B+25+18=229
A+B+150=229
A+B=229-150
A+B=79
B=79-A
Me=46
Me= l+i/f[n/2+m]
Me= 40+10/65[299/2+(42+A)]
A=33.5
B=79-33.5
B=45.5
Problem
From the following data find out the missing frequency when its
arithmetic mean is 25.
Frequency: 5 - 15 - 5 45
Sol:
CI x f fx
23
0-10 5 5 15
10-20 15 A 15A
20-30 25 15 375
40-50 45 5 225
1325−20A
𝟐𝟓 =
𝟒𝟓
A=10 B=20-A=20-10
B=10
24
3: MEASURES OF DISPERSION
Measures of Dispersion are the averages of second order. They are based on averages of
deviations of the values obtained from central tendencies. The average of second order is an
average of difference of all items of the series from an average of those items.
1. Range
2. Quartile Deviation
3. Mean Deviations
4. Standard Deviation.
RANGE:
Range represents the difference between the values of the extremes, the largest value and
the smallest value.
25
R=L-S
R = Range
L = largest value
S = smallest value
Merits of Range:
Demerits of Range:
Problem:
Compute the range and co-efficient of range of the series and state which one is more
dispersed and which one is more uniform.
1 13 14 15 16 17 (mean=15)
2 9 12 15 18 21 (mean=15)
“Central tendency is same but formation differs”.
SERIES 1
R=L–S
R = 17 - 13
R=4
= 17 −13
17 +13
= 0.133
26
SERIES 2
R=L–S
R = 21 – 9
R = 12
= 21 −9
21 +9
= 0.4
Conclusion :
QUARTILE DEVIATION
Quartile deviation is the measures of dispersion based on the upper quartile 𝑄3 and lower
quartile 𝑄1.
1. It ignores the portions below the lower quartile and upper quartile.
2. It is not capable of further mathematical treatment.
3. It is greatly affected by the fluctuations in the sampling
Problem:
Compute Q.D and its Co-efficient for the following data.
Value: 10 20 30 40 50 60.
Frequency: 28 36 24 32 40 16
27
Values Frequency(f) Cumulative
frequency(c.f)
10 28 28
20 36 64
30 24 88
40 32 120
50 40 160
60 16 176
176
𝒏+𝟏 𝒕𝒉
𝑸𝟏 𝒊s the size of( ) item.
𝟒
n=176 → 176+1
4
=44.25.
So, 𝑄1= 20
𝟑(𝒏+𝟏) 𝒕𝒉
𝑸𝟑 is the size of( ) item.
𝟒
3(176+1) 𝑡ℎ
n = 176. → ( ) item= 132.75.
4
So, 𝑄3 = 50.
Q.D = 𝑄3−2 𝑄 1
= 50 −20
2
Q.D = 15.
Co-efficient of Q.D = 𝑄3 − 𝑄 1.
𝑄3 + 𝑄1
MEAN DEVIATION
It is the average amount of difference of the items in a distribution from either the mean or
median or mode ignoring the signs of the deviation.
INDIVIDUAL OBSERVATION:
δ= ∑d
𝑛
Coefficient of δ = 𝛿
or 𝛿 or 𝛿
𝑀𝑒 z 𝑋̅
28
DISCRETE AND CONTINUOUS SERIES
δ = ∑𝑓𝑑
𝑛
Coefficient of δ = 𝛿
or 𝛿 or 𝛿
𝑀𝑒 z 𝑋̅
Question: From the following variable find the mean deviation and coefficient of mean
deviation from the mean:
X: 68 49 32 21 54 38 59 66 41
Solution: Arrange the data in an ascending order to have the short cut method applicable:
𝑥=∑x/n=428/9=47.56
X (x-𝑥)
21 26.56
32 15.56
38 9.56
41 6.56
49 1.44
54 6.44
59 11.44
66 18.44
69 20.44
∑x=428 ∑d=116.44
∑𝑑
𝛿=
𝑛
=116.44/9= 12.937
Coefficient of δ=δ/𝑥
=12.9378/47.56
=0.272
Question: Following are runs scored by the batsmen in different innings of cricket tests.
Runs: 20 40 60 80 100 120 140 160 180
No of batsmen 6 19 40 23 65 83 55 20 9
Compute mean deviation from mode and its coefficient
29
Solution: Computation of mean deviation:
X f (x-z)=d fd
20 6 100 600
40 19 80 1520
60 40 60 2400
80 23 40 920
100 65 20 1300
120 83 0 0
140 55 20 1100
160 20 40 800
180 9 60 540
n=320 ∑fd=9180
Problem: Calculate mean deviation from mean from the following data.
f : 8 3 11 14 5 7 2
(f) Mid
point
48-56 8 52 21.44 171.52 416
30
δ = ∑𝑓𝑑
𝑛
= 543.36/50 = 10.8672
Coefficient of δ = 𝛿
𝑋̅
= 10.8672/73.44
= 0.14797
problem:
Calculate mean deviation from median and also find its co-efficient.
Value: 10 20 30 40 50 60.
Frequency: 28 36 24 32 40 16
X f Cumulative d = X- fx fd
frequency(c.f) 𝑀𝑒
10 28 28 30 280 840
20 36 64 20 720 720
30 24 88 10 720 240
40 32 120 0 1280 0
𝑛+1 𝑡ℎ
𝑀 is the size of( ) item.
𝑒 2
n=176 = 176+1 =88.5
2
Therefore, 𝑀𝑒 is 40
δ = ∑𝑓𝑑
𝑛
31
= 2520/176
δ = 14.3181
Coefficient of δ = 𝛿 .
𝑀𝑒
= 14.3181/40
= 0.35795
Problem:
X F d=X-Z Fd
20 6 100 600
40 19 80 1520
60 40 60 2400
δ = ∑𝑓𝑑
140 55 20 1100 𝑛
= 9180/320
160 20 40 800
δ = 28.687.
180 9 60 540 Coefficient of δ = 𝛿 .
Z
∑fd=9180 = 28.687/120
= 0.2390
32
Merits of Mean Deviation:
1. It is rigidly defined and easy to compute and understand.
2. It takes all items into consideration and gives weight to deviation according to
their size.
3. It is less affected by extreme values of variable.
4. It removes all the irregularities by obtaining deviations and provides a correct
measure.
33
σ =√∑fd2/n where d= X- 𝑿
Coefficient of Variation σ/x *100
Problem:
Calculate S.D for the following deviation and also find co-efficient of variation.
Tests: 0 1 2 3 4
Marks: 17 9 6 5 3
X f Fx A=2 d2 fd fd2
d =X-A
0 17 0 -2 4 -34 68
1 9 9 -1 1 -9 9
2 6 12 0 0 0 0
3 5 15 1 1 1 5
4 3 12 2 4 6 12
40 ∑𝑓𝑑 =- ∑fd2=94
32
σ = √∑fd2/𝑛– (∑𝑓𝑑/𝑛)2
σ = √94/40– (−32/40)2
= √1.71
= 1.30
S.D = √𝒗𝒂𝒓𝒊𝒂𝒏𝒄𝒆.
Problem:
The numbers of employees, wages per employees for two factories are given below.
In which factory there is greater variance in the distribution of wages per employee and
which factory pays more wages.
Factory X Factory Y
No of employees 50 100
Average wages 60 42
per employee per
month
Variance wages 16 64
per employee
per month.
34
FACTORY X FACTORY Y
n = 50 n = 100
𝑋̅ =60 𝑋̅ =42
Variance = 16 Variance = 64
S.D= √𝒗𝒂𝒓𝒊𝒂𝒏𝒄𝒆. S.D = √𝒗𝒂𝒓𝒊𝒂𝒏𝒄𝒆.
= √𝟏𝟔 = √𝟔𝟒
= 4 = 8
To find total number of To find total number of
employees employees
∑𝑓𝑥= n ×𝑋̅ ∑𝑓𝑥= n ×𝑋̅
= 50×60 = 100×42
= 3000 = 4200
𝜎 𝜎
C.V = ( × 100 ) C.V = ( × 100 )
𝑋̅ 𝑋̅
= (4/60 × 100 ) = 8/6(42 × 100 )
= 6.66 = 19.04
Conclusion:
Question: Find which of the Batsman is more consistent in scoring. Would you also
accept him as a better run getter? Why?
Batsman
A: 5 7 16 27 39 53 56 61 80 101 105
B: 0 4 16 21 41 43 57 78 83 90 95
35
Solution:
x-𝒙 𝒅𝟐 y y-𝒚 𝒅2
x Batsman A (d) Batsman B(d)
= 33.5368/50*100 =33.3657/48*100
= 67.0736% = 69.5119 %
36
4. CORRELATION AND REGRESSION ANALYSIS
Correlation
Regression
“Regression” means returning or stepping back to the average value. With the help of values
of one variable (independent) we can establish most likely values of other variable (dependent).
On the basis of two available correlated variables, we can forecast the future data or events or
values.
In statistics the term “Regression” means simply the “Average Relationship”. We can predict
or estimate the values of dependent variable from the given related values of independent
variable with the help of a Regression Technique.
Regression Analysis
“Regression Analysis” refers to the methods by which estimates are made about the values of
a dependent variable from the values of an independent variable. It is a technique of predicting
the unknown values on the basis of the “average relationship”.
Correlation Regression
37
4. It is merely a tool of ascertaining 4. It is also a tool of studying cause
the degree of relationship. and effect of relationship.
Types of Correlation
Correlation is classified, on the basis of its nature, into the following ways:
If the values of the two variables deviate in the same direction, correlation is said to be „positive‟
or „direct‟. It means, the increase in the values of one variable results, on an average, in a
corresponding increase in the values of other variable and vice versa.
If the values of the two variables deviate in the opposite direction, correlation is said to be
„negative‟ or „indirect‟. It means, the increase in the values of one variable results, on an average,
in a corresponding decrease in the values of other variable and vice versa.
When one variable is independent and the other variable is dependent on the former, it is the case
of a „partial correlation‟.
When only two variables are studied, it is called a „simple correlation‟. It means the study
involves only two variables, which are changing either in the same or opposite direction.
38
Probable Error
“Probable Error” is a difference resulting due to taking samples from the mass of population.
𝟏−𝒓𝟐
P.E. = Probable factor × Standard error = 0.6745
√𝒏
𝐧 𝐧
1. Following are the results of [Link]. Examination in a college. Compute coefficient of correlation
between age and success in the examination and interpret the result.
Age of Candidates: 20-21 21-22 22-23 23-24 24-25 25-26
Candidates Appeared: 120 100 70 40 10 5
Successful Candidates: 72 55 35 18 4 1
Solution: Let us obtain the Mid-value of age group and convert the successful candidates into
percentages.
22-23 22.5 -1 1 50 0 0 0
23-24 23.5 0 0 45 -5 25 0
39
′
(𝐝 𝐱)(𝐝 𝐲) ′
′ ′
𝐝 𝐱𝐝 𝐲 − 𝐧
𝐫=
√𝐝′𝐱𝟐 − (𝐝′𝐱) √𝐝′𝐲𝟐 − (𝐝′𝐲)
𝟐 𝟐
𝐧 𝐧
(−𝟑) (−𝟔)
−𝟐𝟐 −
= 𝟔
√𝟏𝟗 − (−𝟑) √𝟒𝟔 − (−𝟔)
𝟐 𝟐
𝟔 𝟔
−𝟐𝟓 −𝟐. 𝟓
= =
𝟒. 𝟏𝟖𝟑𝟑 × 𝟔. 𝟑𝟐𝟒𝟓 𝟐𝟔. 𝟒𝟓𝟕
r= -0.94492
X: 06 08 12 15 18 20 24 28 31
Y: 10 12 15 15 18 25 22 26 28
Sol:
12 15 -6 36 -4 16 +24
15 15 -3 09 -4 16 +12
18 18 0 00 -1 01 00
20 25 +2 04 +6 36 +12
24 22 +6 36 +3 09 +18
40
𝐝𝐱𝐝𝐲
𝐫=
√𝐝𝐱𝟐 √𝐝𝐲𝟐
𝟒𝟑𝟏
𝐫=
√𝟓𝟗𝟖 √𝟑𝟑𝟖
𝟒𝟑𝟏
𝐫=
(𝟐𝟒. 𝟒𝟓) (𝟏𝟖. 𝟑𝟖)
𝟒𝟑𝟏
𝐫=
𝟒𝟒𝟗. 𝟑𝟗
R= 0.959
X y Rx Ry d d2
60 75 3 1 +2 4
34 32 11 10 +1 1
40 35 10 8 +2 4
50 40 4 6 -2 4
45 45 6 4 +2 4
41 33 9 9 0 0
22 12 12 12 0 0
43 30 7 11 -4 16
42 36 8 7 +1 1
66 72 1 2 -1 1
64 41 2 5 -3 9
46 57 5 3 +2 4
d=0 d2=48
41
6d2
rs = 1-
n3−n
6×48
= 1- 123−1
6×48
= 1-
123−12
288
= 1- 1716
= 1 - 0.16783
rs = +0.83217
X Rx Y Ry d= Rx-Ry d2
60 2 10 8 -6 36
15 7 40 3 4 16
20 5.5 30 5 0.5 0.25
28 4 50 2 2 4
12 8 30 5 3 9
40 3 20 7 -4 16
80 1 60 1 0 0
20 5.5 30 5 0.5 0.25
d2 = 81.5
In ‘X’ series 5th and 6th ranks are tied(2 ranks equal m=2)
So 5+6/2 = 5.5
In ‘Y’ series 4th, 5th and 6th ranks are tied(3 ranks equal m=3)
So 4+5+6/2 = 5
1 1
𝟔[d 2 + ( m3 −m )+ (m 3−m)]
12 12
rs = 𝟏 − n3 −n
where m= no. of ranks equal
1 1
𝟔[81.5+ (23−2 )+ (33−3)]
rs = 𝟏 − 12
83 −8
12
42
1 1
𝟔[𝟖𝟏.𝟓+ ( 6) + ( 24) ]
rs = 𝟏 − 12 12
504
1
𝟔[81.5+ +2]
rs = 𝟏 − 2
504
𝟔[81.5+2.5]
rs = 𝟏 − 504
𝟓𝟎𝟒
rs = 𝟏 −
504
rs = 𝟎
28 26 -8 -3 64 9
24
30 29 -6 0 36 0
0
32 30 -4 +1 16 1
-4
35 25 -1 -4 1 16
+4
36 18 0 -11 0 121
0
38 26 2 -3 4 9
-6
39 35 3 +6 9 36
+18
42 35 6 +6 36 36
+36
43
∑𝑥 360 ∑𝑦 290
𝑋̅ = = = 36 𝑌= = = 29
𝑛 10 𝑛 10
y on x x on y
∑ 𝑥𝑦 ∑ 𝑥𝑦
(𝑌 − 𝑌) = (𝑋̅ − 𝑋̅) (𝑋̅ − 𝑋̅) = (𝑌 − 𝑌)
∑ 𝑥2 ∑ 𝑥2
494 494
(Y−29) = (X−36) (X−36) = (Y−29)
648 598
(Y−29) = 0.7624 (X−36) (Y−36) = 0.8261 (Y−29)
Y= 0.7624X−27.4464+29 X= 0.8261Y−23.9569+36
Y= 0.7624X +1.5536 X= 0.8261Y+12.0431
4. You are given below the information about advertising and sales of a company:
Advertising expenses(x) Sales(y)
(Rs in lakhs) (Rs in lakhs)
Mean 10 90
Variance 9 144
Correlation Coefficient 0.8
a) Calculate the two regression lines.
b) Find the likely sales when advertising expenses are Rs 15 and Rs 20.
c) What should be the advertising expenses when the company wants to attain a sales target of Rs 140 lakhs
and Rs 190 lakhs.
Solution:
a) Calculation of Two Regression Lines
𝜎𝑥 = √𝑣𝑎𝑟𝑖𝑎𝑛𝑐𝑒 = √9 = 3, 𝜎𝑦 = √𝑣𝑎𝑟𝑖𝑎𝑛𝑐𝑒 = √144 = 12
Y on x x on y
y=bx x=by
(𝑌 − 𝑌) = 𝑏𝑦𝑥 (𝑋̅ − 𝑋̅) (𝑋̅ − 𝑋̅) = 𝑏𝑥𝑦 (𝑌 − 𝑌)
𝜎𝑦 𝜎𝑥
(𝑌 − 90) = 𝑟 (𝑋̅ − 10) (𝑋̅ − 10) = 𝑟 (𝑌 − 90)
𝜎𝑥 𝜎𝑦
12 3
(𝑌 − 90) = 0.8 (𝑋̅ − 10) (𝑋̅ − 10) = 0.8 (𝑌 − 90)
3 12
44
Y – 90 = 3.2 (X − 10) X – 10 = 0.2 (Y − 90)
Y −90 = 3.2X −32 X −10 = 0.2Y −18
Y=3.2X −32+90 X=0.2Y −18+10
Y=3.2X+58 X=0.2Y−8
45
5: INDEX NUMBER
According to A.M Tuttle „Index number‟ is a single ratio which measures the
combine change of several variables between two different times, places, (or) situations.
It can also be defined as a quantity which by reference to a base period shows variations or
changes in the magnitude over a period of time.
Purpose:
i. To measure and compare changes in a variable or set of variables with
some base year.
ii. To measure purchasing power of money which help in finding out real
wages of people
iii. To provide guidelines to policymaking in business field.
Importance:
Index numbers are economic barometers as they indicate the pulses of economy and
tendencies in price fluctuations .They also measure the pressure of economic activities of the
country.
Base year:
Base year is a year with which index number are to be compared (or) to which they have to
be refered. It should be a period which is having normal (or) stable economic activities.
Current year
It refers to the year for which comparison is done.
METHODS OF CONSTRUCTING INDEX NUMBER.
INDEX NUMBER
UNWEIGHTED WEIGHTED
(PRICE AND
(PRICE)
QUALITY)
NOTATIONS:
𝑃0= Base year prices.
46
𝑃1= Current year prices.
𝑃01= Price index number for current year with reference to the base year.
𝑞01= Quantity index number for the current year with reference to the base year.
𝑉01= Value index number for the current year with reference to the base year.
Where;
Items A B C D E
Price in 6 10 2 12 5
2001(𝑝0)
Price in 8 15 4 8 5
2003(𝑝1)
Solution:
= 40×100
35
= 114.28.
47
Conclusion:
This means during the year 2003 the prices have rise up by 14.28% on an average
compared to the prices in the year 2001.
Where;
I = 𝑝1×100
𝑝0
I= Price relative.
n = number of commodities.
Problem:
Calculate index number by average of price relatives method.
Items A B C D E
Price in 6 10 2 12 5
2001(𝑝0)
Price in 8 15 4 8 5
2003(𝑝1)
Solution:
A 6 8 133.3
B 10 15 150
C 2 4 200
D 12 8 66.67
E 5 5 100
∑𝐼= 650
= 650/5= 130
48
WEIGHTED INDEX NUMBER:
1. Laspeyre‟s method:
In this method the base year quantity are taken as weights.
∑𝑝 1𝑞 0
𝑃01= × 100
∑𝑝 0𝑞 0
2 Paasche‟s method:
4 Fisher‟s method:
∑𝑝1 𝑞1
𝑃 = √∑𝑝1𝑞0 × ×100
01 ∑𝑝 0𝑞 0 ∑𝑝0𝑞1
∑𝑝 𝑞 ∑𝑝 𝑞
𝑃01 = √∑𝑝 1𝑞0 × ∑𝑝1 𝑞1
0 0 0 1
∑𝑝 𝑞 ∑𝑝0𝑞0
𝑃01 = √∑𝑝0𝑞1 ×
1 1 ∑𝑝1𝑞0
𝑃01 × 𝑃01 = 1.
49
Factor reversal test:
The formula should true value ratio, if the price index number multiplied by quantity index
number.
∑𝑝1𝑞1
𝑃 = √∑𝑝1𝑞0 ×
01 ∑𝑝0𝑞0 ∑𝑝0𝑞1
∑𝑞1𝑝1
𝑄01 = √∑𝑞 1𝑝 0 ×
∑𝑞 0𝑝 0 ∑𝑞0 𝑝1
∑𝑝1𝑞1
𝑉01 =
∑𝑝0 𝑞0
Construct the Laspeyre‟s, Paasche‟s, Marshall- Edge worth and Index number for the
following data and prove that it satisfies both TRT and FRT.
Solution:
50
1. Fisher‟s method:
∑𝑝1𝑞1
𝑃 = √∑𝑝1𝑞0 × ×100
01 ∑𝑝0𝑞0 ∑𝑝0𝑞1
5060 5490
𝑃01 = √ 5000 × 5470 × 100
= 100.78
2. Laspeyre‟s method:
∑𝑝 1𝑞 0
𝑃01 = × 100
∑𝑝 0𝑞 0
= 5060×100
5000
= 101.2
3. Paasche‟s method
𝑃01 = ∑𝑝 1 𝑞 1 × 100
∑𝑝 0𝑞 1
5490
×100
5470
= 100.36
4. Marshall- Edge worth :
=5060 +5490
5000 +5470
× 100
=1.0076
5. Time reversal test:
𝑃01 × 𝑃01 = 1.
∑𝑝1 𝑞0 ∑𝑝1 𝑞1 ∑𝑝0 𝑞1 ∑𝑝0 𝑞0
√ × ×√ ×
∑𝑝0 𝑞0 ∑𝑝0 𝑞1 ∑𝑝1 𝑞1 ∑𝑝1 𝑞0
√ 5060
5000
× 5490
5470
× 5470
5490
× 5000
5060
= 1
51
6. Factor reversal test:
√∑𝑝 1𝑞 0 × ∑𝑝 1𝑞 1 ×
∑𝑝 0𝑞 0 ∑𝑝 0𝑞 1
√∑𝑞 1𝑝 0 × ∑𝑞 1𝑝 1
∑𝑞 0𝑝 0 ∑𝑞 0𝑝 1
= 5490/5000 = ∑𝑝1𝑞1 = 𝑉
∑𝑝0 𝑞0 01
Methods of computation:
∑𝑝 1𝑞 0
Aggregate Expenditure Method….𝑃01 = × 100
∑𝑝 0𝑞 0
Question: A textile worker earns Rs. 350 per month . The cost of living indeex
for that particular month is known as 136. Using the following data finid the
amounts spend by him on house rent and clothing.
Group Food Clothing Rent Fuel Miscellanous
Expenditure 140 ? ? 56 63
Group Index 180 150 100 110 80
52
Solution: Assume the amounts spend on clothing and house rent as A & B
respectively. The total wages spent will also be assumed as his earnings. Thus we can
have following equations:
140+A+B+56+63=350
A+B=91.....................................................................(i)
Computation of cost of living Index
Group Expenditure(W) I IW
Food 140 180 25,200
Clothing A 150 (150A)
House rent B 100 (100B)
Fuel 56 110 6160
Miscellaneous 63 80 5040
∑W=350 ∑IW=36400+150A+100B
C.L.I= ∑𝐈𝐖
∑𝐖
× 100
136=36400+150A+100B/350
150A+100B=11,200 ........................................... (ii)
Solving (i) & (ii) we get the value of A & B AS
A=42 & B=49 ; so the worker has spent Rs 42 on clothing and Rs 49 on House rent
53
6. TIME SERIES
Time series is a series of numerical data which have been recorded at different
intervals of time. It is a record of changes in variables over a period of time.
54
2. Using the 4 yearly moving average determine the trend value and plot the graph for the
actual and trend values.
Production: 75 62 76 78 94 84 96
1994 75
62
1995 76
291 291/4=72.75
78
1996 75.125 96
94
310 310/4=77.5
1997 84 80.25 97
96 332 332/4=83
1998 85.5 98
352 352/4=88
1999
2000
55
2) Using 5 yearly moving average determine the trend values and plot the graph.
Year: 1986 1987 1988 1989 1990 1991 1992 1993 1994 1995 1996
Sales: 12 14 18 24 22 20 16 25 26 34
31
a =∑y/n
b = ∑xy/∑𝒙𝟐
𝒀𝒄=𝒂+𝒃𝒙
Sales: 20 23 22 25 26 29 30
56
a =∑y/n b = ∑xy/∑𝒙𝟐 𝒀𝒄=𝒂+𝒃𝒙
57