0% found this document useful (0 votes)
2 views68 pages

Statistics

The document provides an introduction to statistics, defining it as the collection and analysis of numerical data, and differentiating between quantitative and qualitative data. It discusses the characteristics and limitations of statistics, the importance of data collection methods, and the concepts of population and sample in statistical studies. Additionally, it covers classification, variables, and types of frequency distributions, along with methods for presenting data effectively.

Uploaded by

suneetatiwari8
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views68 pages

Statistics

The document provides an introduction to statistics, defining it as the collection and analysis of numerical data, and differentiating between quantitative and qualitative data. It discusses the characteristics and limitations of statistics, the importance of data collection methods, and the concepts of population and sample in statistical studies. Additionally, it covers classification, variables, and types of frequency distributions, along with methods for presenting data effectively.

Uploaded by

suneetatiwari8
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Statistics Ch-1 Introduction

1. Statistics Meaning and Definition

 Statistics may be defined as the collection, organization, presentation, analysis


and interpretation of numerical data.

 Statistics refers to information in terms of numbers or numerical data.

 Statistics means data or quantitative information capable of meaningful


conclusion.

 Quantitative Data – This type of data can be expressed in numerical terms.


Example - Income, expenditure, investment.

 Qualitative data – This type of data cannot be expressed in numerical terms.


Example – beauty, intelligence. These can be rated or ranked like good, very good,
excellent or 1, 2, 3.

 All Statistics are data, but all data are not statistics.
Example – ‘Ram has a pocket allowance of 500’ is not statistics because it is neither
an aggregate nor an average. ‘Average pocket allowance of students of class 11 in a
school is 500’ is a statistics.

2. Features/Characteristics of Statistics

 Aggregate of facts – A single number does not constitute statistics.


Example – ‘There are 1000 students in a college’ is not statistics because no
conclusion can be drawn from it. ‘There are 300 students in Arts, 400 in Commerce
and 300 in Science in a college’ is Statistics.

 Numerically expressed – Statistics are expressed in terms of numbers.


Example – Irfan Pathan is taller than Sachin Tendulkar is not statistics. ‘Irfan Pathan is
6 ft tall and Sachin Tendulkar is 5 ft tall’ is statistics.

 Multiplicity of causes – Statistics are affected by multiple factors.


Example – 30% rise in prices could be due to several factors.
 Reasonable Accuracy – A reasonable degree of accuracy must be kept in view
while collecting statistical data.

 Mutually related and comparable


Example – ‘Ram is 40 years old, Mohan is 5 ft tall, and Sohan has 60 kg of weight’ is
not Statistics as these numbers are not mutually related and cannot be compared.

 Pre-determined Objective – Statistics are collected with some pre determined


objectives.

3. Limitations of Statistics

 Study numerical facts only – Statistics does not study qualitative phenomenon
such as honesty, friendship, knowledge.

 Study aggregates only

 Results are true only on average


Example – If per capita income of India is 30,000 rupees per month, then it does not
mean than every Indian earn 30,000 rupees per month.

 Without Reference, result may be wrong or misleading


Example – In following example, average profit of firm 2 is higher but its profit is
decreasing over time.
Firm 1 Firm 2
Profit Year 1 1,00,000 6,00,000
Profit Year 2 2,00,000 4,00,000
Profit Year 3 3,00,000 2,00,000
Average 2,00,000 4,00,000

 Statistics is prone to misuse


There are three kind of lies – lies, damned lies and Statistics

 Can only be used by experts


4. Function/Importance of Statistics

 Quantitative expression of economic problem – First task of economists is to


understand magnitude of economic problem like unemployment, inflation by quantifying
them.

 Intertemporal and Intersectoral comparison – Intertemporal comparison means


comparison over time (unemployment rate over different months or years) and
Intersectoral comparison means comparison across space (rural versus urban poverty)

 Working out cause and effect relationship – to diagnose the problem (finding
the cause of unemployment) and then offering suggestions (like increasing expenditure
on education).

 Economic forecasting – to predict the future course of a variable (likely


unemployment in coming years)

 Formulation of policies
Statistics Ch-2 Collection of Data

5. Primary and Secondary data

A. Primary Data

Primary data is the first hand data collected by an investigator for its own purpose and
for the first time. It is the data that is collected by the investigator from its source of
origin.

B. Secondary data

Secondary data is the second hand data which is already in existence and which has
been collected for some other purpose. It implies collection of data from some agency
or institution that has already collected the data.

6. Method of Data Collection

 Personal interview

a. Direct personal investigation


Data is personally collected by the investigator from the respondent.
It is suitable for limited coverage, has high degree of accuracy but it is costly.

b. Indirect oral investigation


In this method, information is obtained not from the person regarding whom the
information is needed but from the other persons who are expected to possess
the necessary information
Example – Collection of workers data from employer rather than workers.
It has wide coverage and less expensive, but is less accurate and biased.

 Questionnaires

a. Mailing (questionnaire) surveys


It is suitable when area of study is large and respondents are educated. It is
economical.
It lacks flexibility, and of limited use (questionnaires can only be answered by
educated people)
b. Enumerator’s method
Enumerator himself approaches the respondents with the questionnaires.
Questionnaire is filled by enumerator.
It has high degree of accuracy and has personal contact but is expensive and
time consuming.

 Telephonic Interview

This method has high degree of accuracy, reliability and is cheaper.


It has limited access (some respondents may be poor and don’t have telephone)

7. Types of questions in a questionnaire

a. Simple Alternative questions – Yes/No


b. Multiple Choice Questions (MCQs)
c. Specific information questions – How much is your salary?
d. Open Questions – These questions ask for the opinion. Example – How can we
solve the issue of global warming?

8. Sources of secondary data

 Published sources
a. Government publications – Annual survey of industries, RBI bulletin
b. Reports of committee and commission – Finance commission, NITI Aayog
c. Publication of research institutes – Indian Statistical Institute, NCERT
d. Journals and newspaper
e. International publication – IMF, World Bank, ILO

 Two important sources of secondary data

a. Census of India
It is a decennial publication related to population size and its related parameters.
It is published by Registrar General and Census Commissioner, India. It
includes information on
 Population size, growth and distribution
 Density of population
 Sex composition
 State of literacy
Information on above parameter relates to country as a whole as well as different
states and UTs. It covers each and every household of country.
b. National Sample Survey Office(NSSO)
It is a government organization under MoSPI (Ministry of Statistics and Program
Implementation). It conducts regular sample surveys to collect a variety of
information like
 Land & Livestock Holding
 Employment and Unemployment
 Consumer expenditure
Statistics Ch-3 Census and Sample Method

1. Concepts of ‘Population’ and ‘Sample’

 Population in Statistics means the aggregate of all items to be studied for an


investigation.

 Sample is a part or subset of population.

 Example - If an investigation relates to all 2000 students in a college then 2000 would
be taken as population. But if investigation covers only some students, then it is a
sample inquiry.

 Characteristics of sample are supposed to represent characteristics of entire


population.

2. Census and sample method

Census Method Sample Method

Coverage In census method data is collected In sample method data is collected only
covering every item of the from a part or subset of population.
population

Suitability This method is suitable when area This method is suitable when area of
of investigation is relatively small investigation is relatively large and
and population is diverse. population is homogeneous

Accuracy This method has greater accuracy This method has less accuracy and
because every item of population is reliability because it studies only few
studied items of populations.

Cost and It is more time consuming and It is less time consuming and less
time costly expensive
3. Method of sampling

A. Random sampling

 It is a method of sampling in which every item of population has equal chance of


being selected.
 This method is free from personal bias of the investigator.
 It is particularly used when population is homogeneous.
 Methods of random sampling – Lottery method, Table of Random numbers
(Tippet’s table).

B. Non- random Sampling

It is a method of sampling in which all item of population do not have equal chance of
being selected.

Method of Non- random Sampling –

a. Purposive sampling
 In this method, investigator himself makes the choice of the sample item.
 It is suitable when some items of population are special and should be
included in sample.
 Example – Inclusion of Tata Iron and Steel Company in the study of Iron
and Steel industry in India.
 There is possibility of personal bias and sample could be tuned to serve
the purpose of investigator.

b. Stratified sampling
 In this method, population is divided into different strata/groups having
different characteristics and then some items are selected proportionally
from each strata/group.
 This method covers diverse characteristics of the population.
 Example – Division of the population into low, middle, and high-income
groups and then a random selection from each group proportionally.

c. Systematic sampling
 According to this method, population is numerically arranged and then
every nth item is selected.
 It is a short cut method of random sampling in which only first item is
selected randomly.
 Example – A population consist of 1000 items and a sample of 100 is
required. One item from the first 10 items is selected randomly. If the 5 th
item is selected then subsequent items selected would be 15 th, 25th, 35th
and so on.

d. Quota sampling
 In this method population is divided into groups and the enumerator is told
to select a fixed number of items from each group.
 Example – a sample consisting of 250 males and 250 females irrespective
of the proportion of male and female in the population.

e. Convenience sampling
 Sampling is done is such as manner that suits the investigator
convenience.
 Example – surveying people in a mall, on the street, and in other crowded
locations.
Statistics Ch-4 Organisation of Data

1. Classification

 Classification is the process of arranging things into groups or classes on the basis of
their characteristics.

 Basis of classification –

a. Geographical/Spatial classification –
This type of classification is based on geographical or locational differences.
Example –
State Rice production (mt)
Punjab 200
Haryana 300
UP 500

b. Chronological classification –
This type of classification is based on time.
Example – Annual sales of a firm
Year Sales(₹)
2020 2,000
2021 3,000
2022 5,000

c. Qualitative classification –
 This type of classification is based on Qualities or Attributes of the data,
 Example – Classification of data on the basis of gender, occupation or
religion of the population.
 It is of two types – Simple and Manifold
 Simple Classification – It is called classification according to dichotomy
(existence or absence of quality). Example – Male/Female
Educated/Uneducated etc.
 Manifold classification – This type of classification involve more than one
characteristic. Example – Following classification of factory workers
Worker
s

Male Female

Literate Illiterate Literate Illiterate

d. Quantitative classification –
This type of classification is based on numerical value of the facts.
Example – Annual profit of small scale firms in UP in 2022
Annual Profit (₹) Number of firms
Below 1,00,000 500
1,00,000 – 2,00,000 150
2,00,000 – 3,00,000 2000
3,00,000 – 4,00,000 40
Above 4,00,000 30

2. Concept of Variable

 A characteristics or phenomenon which can be measured and changes overtime is


called a variable.
 Example – Weight, Temperature, Number of students passing a test etc.
 Types of variables – Discrete and Continuous

Discrete Variable Continuous Variable

These variables are expressed in These variables assume a range of values


complete numbers like 1,2,5,7,200 including fractions.

These variables increases in jumps These variables increase continuously.

Example – Number of students passing a Example – Height and weight students


test.
3. Statistical series

 Series refers to the data presented in some order and sequence


 Types of Statistical series –

a. Individual series
 It is a series without frequency in which all items are listed singly.
 Example – marks obtained by students in a test
Roll No Marks(out of 50)
1 35
2 40
3 47
4 13
5 25

b. Discrete series/Frequency Array


 In this type of series, every item is measured exactly along with its
frequency.
 Frequency refers to the number of times an item repeats itself in a
series.
 Example –
Marks Frequency
35 4
40 3
47 1
13 4
25 5

c. Frequency Distribution
 In this series, data is classified into different ranges or limits called class
intervals and corresponding to each class interval there is some frequency.
 Every class interval has a lower limit as well as upper limit.

Class size OR magnitude of class interval = Upper limit – Lower limit

Marks Frequency
0-10 4
10-20 3
20-30 1
30-40 4
40-50 3
4. Types of Frequency Distribution

a. Exclusive series

Exclusive series is that series in which every class interval excludes items
corresponding to its upper limit.

Example - 10 will be included in interval 10-20 and not in 0-10

Marks Frequency
0-10 4
10-20 3
20-30 1
30-40 4
40-50 3

b. Inclusive series

Inclusive series is that series in which every class interval includes items corresponding
to its upper limit.

Marks Frequency
0-9 4
10-19 3
20-29 1
30-39 4
40-49 3

c. Open end series

In open end series, lower limit of the first class interval and upper limit of last class
interval is missing.

Marks Frequency
Below 10 4
10-20 3
20-30 1
30-40 4
40 and Above 3
d. Cumulative series

It is that series in which frequencies are continuously added corresponding to each


class interval.

Marks Frequency
Less than 10 4
Less than 20 7
Less than 30 8
Less than 40 12
Less than 50 15

e. Mid-values series

In this type of series, we have only mid-values of the class interval and the
corresponding frequencies.

Mid-value = (Upper limit + Lower limit)/2

Mid-value Frequency
5 4
15 3
25 1
35 4
45 3
Statistics Ch-5 Presentation of Data – Textual and
Tabular

1. Type of data presentation

a. Textual
b. Tabular
c. Diagrammatic –

1. Geometric – Bar Diagrams, Pie Diagrams


2. Frequency Diagrams – Histogram, Polygon, Ogives
3. Arithmetic Line - Graphs OR Time Series Graphs

2. Textual Presentation

 This is also called descriptive presentation of data and is most common

 Example – Survey conduction by an NGO, in the state of Punjab, revealed that area
under pulses cultivation tended to shrink by 40% and area under rice and wheat
cultivation has tended to expand by 20%, between the years 2020-21.

3. Tabular Presentation

 A statistical table is a systematic organization of data in rows and columns.

 Components of a table –
Table Number -
Title -
Captions (Column Headings)
Stubs (Row Heading) Column Heading 1 Column Heading 2
Row Heading 1 Cell Cell
Row Heading 2 Cell Cell Body of
Row Heading 3 Cell Cell the Table
Row Heading 4 Cell Cell

Footnotes: Body of the Table


Source:
 Simple or one way table – shows only one characteristic of the data.

Number of Students in a College

Class Number of Students


B.A. (I) 100
B.A. (II) 80
B.A. (III) 40
Total 220

 Complex table – shows more than one characteristics of the data.

Types –
a. Two way or double – shows two characteristics
b. Three way or Treble – shows three characteristics
c. Manifold – shows more than three characteristics

Example of a two way table -

Number of Students in a College

Number of Students
Class Total
Boys Girls
B.A. (I) 40 60 100
B.A. (II) 60 20 80
B.A. (III) 20 20 40
Total 120 100 220
Statistics Ch-6 Bar Diagrams and Pie Diagrams

1. Types of Bar diagrams

a. Simple bar diagrams


Simple bar diagrams are based on a single set of numerical data.
Two types – Vertical and horizontal

Year Birth Rate


1951-60 45
1961-70 35
1971-80 30
1981-90 28
1991-00 24

Birth rate in India (Vertical)


50 45
45
40 35
35
BIRTH RATE

30
30 28
24
25
20
15
10
5
0
1951-60 1961-70 1971-80 1981-90 1991-00
YEAR
Birth rate in India (Horizontal)

1991-00 24

YEAR 1981-90 28

1971-80 30

1961-70 35

1951-60 45

0 10 20 30 40 50
BIRTH RATE

b. Multiple bar diagrams


Multiple bar diagrams are those diagrams which show two or more sets of data
simultaneously.
These bar diagrams are used to make comparison.

Birth Death
Year
rate rate
1951-60 40 27
1961-70 42 23
1971-80 41 19
1981-90 37 15
1991-00 32.5 11.5

Birth and Death rate in India


45
40
BIRTH AND DEATH RATE

35
30
25
20 BIRTH RATE
15 DEATH RATE
10
5
0
1951-60 1961-70 1971-80 1981-90 1991-00
YEAR
c. Component bar diagrams
These diagrams simultaneously present total value as well as part values of the
data.

Year Hydro Thermal Total


2017-18 46 64 110
2018-19 49 71 120
2019-20 48 82 130
2020-21 51 89 140

Electricity Production in India


160
140
140 130
120
120 110
PRODUCTION

100 89
71 82
80 64
Thermal
60
Hydro
40
20 46 49 48 51

0
2017-18 2018-19 2019-20 2020-21
YEAR

d. Percentage bar diagrams


These diagrams simultaneously show different parts of the values of a set of data
in terms of percentages.

Family A Family B
Expenditure Expenditure Expenditure Expenditure
Item
(₹) (%) (₹) (%)
Food 200 (200/500)*100 = 40 300 (300/1,000)*100 = 30
Clothing 100 (100/500)*100 = 20 200 (200/1,000)*100 = 20
Rent 150 (150/500)*100 = 30 100 (100/1,000)*100 = 10
Fuel 50 (50/500)*100 = 10 400 (400/1,000)*100 = 40
Total 500 100 1,000 100
Y

100 100
100 -
10
90 -
Food
80 - 40
30
70 - Clothing
Expenditure 60 -
10 10 Rent
50 - 20
40 - 20 Fuel
30 -
20 - 40
30
10 -
0
Family A Family B X

e. Deviation bar diagrams


These diagrams are used to used to compare net deviation of related variables with
respect to time or location
Bars representing positive and negative values are shown above and below base line.

Year Savings/Deficit
2016 30
2017 -20
2018 10
2019 15
2020 -25
Saving/Deficit
40
30
30
Saving/Deficit
20 15
10
10
0
2016 2017 2018 2019 2020
-10
-20
-20
-30 -25
YEAR

2. Pie diagrams
Pie diagram is a circle divided into various segments showing different values of a series.

Food Type Number of People Component in 360°

150
North Indian 150 × 360° = 108°
500

100
South Indian 100 × 360° = 72°
500

125
Chinese 125 × 360° = 90°
500

75
Italian 75 × 360° = 54°
500

50
Mexican 50 × 360° = 36°
500

Total 500 360


Mexican
36°

North Indian,
Italian, 54° 108°

Chinese, 90°
South Indian,
72°
Statistics Ch-7 Frequency Diagrams – Histogram,
Polygon and Ogive

1. Histogram
A histogram is a graphical representation of a frequency distribution of a continuous series.
Histograms of frequency distribution are of two types –

a. Histogram of equal class intervals

Marks Number of students


0-10 5
10-20 10
20-30 15
30-40 20
40-50 12
50-60 8
60-70 4

Y
Marks in Statistics
30

25
Number of 20
20
Students 15
15 12
10
10 8
5 4
5

X
O 10 20 30 40 50 60 70

Marks
b. Histogram of Unequal class intervals

With unequal class interval, frequencies are to be adjusted first before presenting data.

𝐂𝐥𝐚𝐬𝐬 𝐬𝐢𝐳𝐞 𝐨𝐟 𝐭𝐡𝐞 𝐜𝐨𝐧𝐜𝐞𝐫𝐧𝐞𝐝 𝐜𝐥𝐚𝐬𝐬 𝐢𝐧𝐭𝐞𝐫𝐯𝐚𝐥


Adjustment Factor =
𝐂𝐥𝐚𝐬𝐬 𝐬𝐢𝐳𝐞 𝐨𝐟 𝐭𝐡𝐞 𝐬𝐦𝐚𝐥𝐥𝐞𝐬𝐭 𝐜𝐥𝐚𝐬𝐬 𝐢𝐧𝐭𝐞𝐫𝐯𝐚𝐥

Weekly wages Number of workers Adjustment factor Adjusted frequency


10-15 7 5/5 = 1 7/1 = 7
15-20 10 5/5 = 1 10/1 = 10
20-25 27 5/5 = 1 27/1 = 27
25-30 15 5/5 = 1 15/1 = 15
30-40 12 10/5 = 2 12/2 = 6
40-60 12 20/5 = 4 12/4 = 3
60-80 8 20/5 = 4 8/4 = 2

*Smallest class interval size = 5

Weekly Wages of Workers


Y

30
27
Number 25
of
Workers 20
15
15 12
10
10 8
7 6
5 3
2
X
O 5 5 10 10 1515 20
20 25
25 30
30 3535 40 4045 4550 50
55 60
55 65
60 7065 75 7080 75 80

Weekly Wages
2. Frequency polygon

It can be formed by joining mid points of the tops of all rectangles in a histogram.

Marks Number of students


0-10 5
10-20 10
20-30 15
30-40 20
40-50 12
50-60 8

Y
Marks in Statistics
30

25
Number of 20
20
Students 15
15 12
10
10 8
5
5

X
-10 0 10 20 30 40 50 60 70

Marks

3. Ogives or Cumulative frequency curves

Two ways of constructing ogives –


a. Less than method
b. More than method
Number of Cumulative Cumulative
Marks Marks Marks
students Frequency Frequency
0-5 4 Less than 5 4 More than 0 100
5-10 6 Less than 10 4+6 = 10 More than 5 100 – 4 = 96
10-15 10 Less than 15 10+10 = 20 More than 10 96 – 6 = 90
15-20 10 Less than 20 20+10 = 30 More than 15 90 – 10 = 80
20-25 25 Less than 25 30+25 = 55 More than 20 80 -10 = 70
25-30 22 Less than 30 55+22 = 77 More than 25 70 – 25 = 45
30-35 18 Less than 35 77+18 = 95 More than 30 45 – 22 = 23
35-40 5 Less than 40 95+5 = 100 More than 35 23 –18 = 5
More than 40 5–5=0

Marks in Statistics(Less than Ogive)


120
100
95
Number of students

100
77
80
55
60
40 30
20
20 10
4
0
0 5 10 15 20 25 30 35 40 45
Marks

Marks in Statistics(More than Ogive)


120
100 96
Number of students

100 90
80
80 70

60 45
40
23
20 5
0
0
0 5 10 15 20 25 30 35 40 45
Marks
Statistics Ch-8 Arithmetic line graphs OR Time series
graphs

1. One variable graph

Year Production
2015 5
2016 8
2017 13
2018 16
2019 20
2020 17

Production of rice in India


25
20
20 17
16
Production

15 13

10 8
5
5

0
2014 2015 2016 2017 2018 2019 2020 2021
Year
2. Two or more variables graph

Year Exports Imports


2015 300 420
2016 350 460
2017 400 600
2018 380 480
2019 450 550
2020 280 450

Exports and Imports of India


700
600
600 550
480
Exports/Imports

500 460 450


420
400 450
400 380 Exports
300 350
300 280 Imports
200
100
0
2014 2015 2016 2017 2018 2019 2020 2021
Year
Statistics Ch-9 Measures of Central Tendency –
Arithmetic Mean

1. Central Tendency and Averages

 A Central Tendency refers to an average or a central value of a statistical series.

 An Average is a single value within the range of data that is used to represent all
values in the series.

 An Average is a figure that represents the whole group.

 Essentials of a good average –

a. Stable definition
b. Representative – should be based on all observation.
c. Capable of further algebraic treatment.
d. Least effect of a change in sample.

 Types of Statictical Averages –

a. Mathematical Averages – Arithmetic Mean, Geometric Mean, Harmonic Mean


b. Positional Averages – Median, Partition Value, Mode
2. Arithmetic Mean OR simply Mean

A. Individual series
∑𝐗
̅=
a. Direct Method: 𝐗
𝐍

Example – 1, Pocket Allowances of 10 students

Pocket
Allowances (X)
15
20
30
22
25
18
40
50
55
65
∑ 𝐗 = 340

̅ = ∑ 𝐗 = 𝟑𝟒𝟎 = 34
𝐗
𝐍 𝟏𝟎

∑𝐝
̅=𝐀+
b. Assumed Mean Method: 𝐗 ,
𝐍
Where
‘A’ is assumed mean
‘d’ is deviation from arithmetic mean, d = X - A

Example – 2, Pocket Allowances of 10 students


Pocket d=X–A
Allowances (X) (A = 40)
15 15 – 40 = - 25
20 20 – 40 = - 20
30 30 – 40 = - 10
22 22 – 40 = - 18
25 25 – 40 = - 15
18 18 – 40 = - 22
40 40 – 40 = 0
50 50 – 40 = 10
55 55 – 40 = 15
65 65 – 40 = 25
∑ 𝐝 = - 60

̅ = 𝐀 + ∑ 𝐝 = 40 +
𝐗
−𝟔𝟎
= 40 – 6 = 34
𝐍 𝟏𝟎
∑ 𝐝′
̅= 𝐀+
c. Step-Deviation Method: 𝐗 ×C
𝐍
Where
d’ = d/C and C is the common factor

Example – 3, Monthly Expenditure of 10 Households

Monthly Expenditure d=X–A d' = d/C


Household
(X) (A = 50,000) (C = 10,000)
1 50,000 0 0
2 60,000 10,000 1
3 40,000 -10,000 -1
4 70,000 20,000 2
5 80,000 30,000 3
6 20,000 -30,000 -3
7 10,000 -40,000 -4
8 30,000 -20,000 -2
9 90,000 40,000 4
10 1,00,000 50,000 5
∑ 𝐝′ = 5

̅ = 𝐀 + ∑ 𝐝′×C = 𝟓𝟎, 𝟎𝟎𝟎 + 𝟓 ×10,000 = 55,000


𝐗
𝐍 𝟏𝟎

B. Discrete series
∑𝐟 𝐗
̅
a. Direct Method : 𝐗 = ∑𝐟
Where ‘ f ’ is frequency
Example – 4, Weekly wage earning of 19 workers

Wages(X) Number of workers (f) fX


10 4 40
20 5 100
30 3 90
40 2 80
50 5 250
∑ 𝐟 = 19 ∑ 𝐟 𝐗 = 560

̅ = ∑ 𝐟 𝐗 = 𝟓𝟔𝟎 = 29.47
𝐗 ∑ 𝐟 𝟏𝟗
̅ = 𝐀 + ∑𝐟 𝐝
b. Assumed Mean Method : 𝐗
∑𝐟
Example – 5, Weekly wage earning of 19 workers

Wages(X) Number of workers (f) d = X - 30 fd


10 4 -20 -80
20 5 -10 -50
30 3 0 0
40 2 10 20
50 5 20 100
∑ 𝐟 = 19 ∑ 𝐟 𝐝 = -10

̅ = 𝐀 + ∑ 𝐟 𝐝 = 𝟑𝟎 + −𝟏𝟎 = 29.47
𝐗 ∑𝐟 𝟏𝟗

∑ 𝐟 𝐝′
̅= 𝐀+
c. Step-Deviation Method : 𝐗 ×C
∑𝐟

Example – 6, Weekly wage earning of 19 workers

Wages(X) Number of workers (f) d = X - 30 d' = d/10 f d'


10 4 -20 -2 -8
20 5 -10 -1 -5
30 3 0 0 0
40 2 10 1 2
50 5 20 2 10
∑ 𝐟 = 19 ∑ 𝐟 𝐝 ′ = -1

̅ = 𝐀 + ∑ 𝐟 𝐝′ × 𝐂 = 𝟑𝟎 +
𝐗
−𝟏
× 𝟏𝟎 = 29.47
∑𝐟 𝟏𝟗
C. Continuous series

∑𝒇 𝒎
̅
a. Direct Method : 𝑿 = ∑𝒇
Where ‘m’ is mid value
𝑳+𝑼
m= , L = lower limit, U = upper limit
𝟐
Example – 7, Marks secured by Class X students in English

𝑳+𝑼
Marks (X) Number of students (f) m= fm
𝟐
0 - 10 20 5 100
10 - 20 24 15 360
20 - 30 40 25 1,000
30 - 40 36 35 1,260
40 - 50 20 45 900
∑ 𝒇 = 140 ∑ 𝒇 𝒎 = 3,620

̅ = ∑ 𝐟 𝐦 = 𝟑,𝟔𝟐𝟎 = 25.86
𝐗 ∑ 𝐟 𝟏𝟒𝟎

∑𝒇 𝒅
̅ =𝑨+
b. Assumed Mean Method : 𝑿
∑𝒇
Where d = m – A

Example – 8, Marks secured by Class X students in English

𝑳+𝑼
Marks (X) Number of students (f) m= d = m - 25 fd
𝟐
0 - 10 20 5 -20 -400
10 - 20 24 15 -10 -240
20 - 30 40 25 0 0
30 - 40 36 35 10 360
40 - 50 20 45 20 400
∑ 𝒇 = 140 ∑ 𝒇 𝒅 = 120

̅ = 𝐀 + ∑ 𝐟 𝐝 = 𝟐𝟓 +
𝐗
𝟏𝟐𝟎
= 25.86
∑ 𝐟 𝟏𝟒𝟎
∑ 𝒇 𝒅′
̅ = 𝑨+
c. Step-Deviation Method : 𝑿 ×C
∑𝒇
𝒎−𝑨
Where d’ =
𝑪

Example – 9, Marks secured by Class X students in English

𝑳+𝑼 𝒅
Marks (X) Number of students (f) m= d = m - 25 d’ = f d'
𝟐 𝟏𝟎
0 - 10 20 5 -20 -2 -40
10 - 20 24 15 -10 -1 -24
20 - 30 40 25 0 0 0
30 - 40 36 35 10 1 36
40 - 50 20 45 20 2 40
∑ 𝐟 = 140 ∑ 𝐟 𝐝 ′ = 12

̅ = 𝐀 + ∑ 𝐟 𝐝′ × 𝐂 = 𝟐𝟓 +
𝐗
𝟏𝟐
× 𝟏𝟎 = 25.86
∑ 𝐟 𝟏𝟒𝟎

3. Weighted Arithmetic Mean

 Weighted arithmetic mean is the mean of the weighted items.


 When all items of a series are not equally important, then we use weighted mean.
 Example – In household expenditure, food is more significant than entertainment.

∑ 𝐖𝐗
 Weighted Arithmetic Mean: XW =
∑𝐖

 Example – 10

Wages (X) Weight (W) WX


10 4 40
20 5 100
30 3 90
40 2 80
50 5 250
∑ 𝐖 = 19 ∑ 𝐖𝐗 = 560

∑ 𝐖𝐗 ∑ 𝟓𝟔𝟎
XW = = = 29.47
∑𝐖 ∑ 𝟏𝟗
4. Combined Arithmetic Mean

 Combined Arithmetic Mean is the mean of the series formed by combining two or more
series.

 Combined Arithmetic Mean: X1,2 = (X1N1 + X2N2) / (N1 + N2)


Where,
X1 is the mean of the first series
X2 is the mean of the second series
N1 is the number of observation/items in first series
N2 is the number of observation/items in second series

 Example – 11
60 students of Section A of Class XI obtained 40 mean marks in Statistics. 40 students
of Section B obtained 35 mean marks in Statistics. Find mean marks in Statistics for
Class XI as a whole.

Given: N1 = 60, X1 = 40, N2 = 40, X2 = 35

X1,2 = (X1N1 + X2N2) / (N1 + N2)

X1,2 = (40×60 + 35×40) / (60 + 40) = 3,800/100 = 38


5. Properties of the arithmetic mean

a. Sum of deviations taken from arithmetic mean is always zero, ∑ 𝐝 = 𝟎, where


d = X- X
Example – 12, Marks of 5 students

Marks (X) d = X - 30
10 -20
20 -10
30 0
40 10
50 20
∑ 𝐗 = 150 ∑𝐝 = 0

̅ = ∑ 𝐗 = 𝟏𝟓𝟎 = 30
𝐗
𝐍 𝟓

b. If a specific value is added to or subtracted from all items in a series, the mean
value if the items would increase or decrease by same value. Likewise, if all
items of a series are multiplied or divided by any specific value, the mean value
of the items would be multiplied or divided by the same specific value.

Example – 13, Marks of 5 students

Marks (X) L=X+2 M=X-2 N=X×2 O=X÷2


10 12 8 20 5
20 22 18 40 10
30 32 28 60 15
40 42 38 80 20
50 52 48 100 25
∑ 𝐗 =150 ∑ 𝐋 =160 ∑ 𝐌 =140 ∑ 𝐍 = 300 ∑ 𝐎 =75
̅
𝐗 = 30 ̅ ̅+2
𝐋 = 32 = 𝐗 ̅ ̅−2
𝐌 = 28 = 𝐗 𝐍 = 60 = 𝐗 × 2 𝐎 = 15 = ̅𝐗 ÷ 2
̅ ̅ ̅
6. Merit and Demerits of Arithmetic Mean

Merits –
 Simple and Certain
 Based on all items therefore it is stable (least affected by sampling fluctuations)
 Capable of algebraic treatment.

Demerit –

 Affected by Extreme Values :


Example – Consider income of 5 individuals – 100, 80, 70, 50, 2000
Mean income is (100+80+70+50+2000)/5 = 2,300/5 = 460
Clearly 460 is not an accurate mean due to the presence of extreme value of
2000

 Mean value may not figure in the series at all :


Example – Mean of 2,3 and 7 is 4, which is not there in the series.

 Arithmetic mean is not a suitable average in case of ratios and proportions.

 Misleading conclusions :
Example – In following situation, average profit of firm 1 and firm 2 are same but
the profit of firm 1 is increasing overtime and that of firm 2 is decreasing overtime.

Firm 1 Firm 2
Profit Year 1 1,00,000 3,00,000
Profit Year 2 2,00,000 2,00,000
Profit Year 3 3,00,000 1,00,000
Average 2,00,000 3,00,000
∑𝐗
Direct Method
𝐍

Individual Assumed Mean ∑𝐝


𝐀+ d=X–A
Series Method 𝐍

∑ 𝐟 𝐝′ 𝐗−𝐀
Step-Deviation Method 𝐀+ ×C d' =
𝐍 𝐂

∑𝐟𝐗
Direct Method
∑𝐟

Simple Discrete Assumed Mean ∑𝐟 𝐝


Mean Series 𝐀+ d=X–A
Method ∑𝐟

∑ 𝐟 𝐝′ 𝐗−𝐀
Step-Deviation Method 𝐀+ ∑𝐟
×C d' =
𝐂

∑𝐟𝐦 𝐋+𝐔
Direct Method m=
∑𝐟 𝟐

Continuous Assumed Mean ∑𝐟 𝐝


Series 𝐀+ d=m-A
Method ∑𝐟

∑ 𝐟 𝐝′ 𝐦−𝐀
Step-Deviation Method 𝐀+ ∑𝐟
×C d' =
𝐂

Weighted ∑ 𝐖𝐗
XW = ∑𝐖
Mean
Combined
X1,2 = (X1N1 + X2N2) / (N1 + N2)
Mean
Arithmetic Mean (Exercises)

1. Mean of 100 observations is found to be 40. If at the time of computations two


items are wrongly taken as 30 and 27 instead of 3 and 72, find the correct
mean.

Solution:
ΣX
̅
X=
N
ΣX
40 =
100
ΣX = 4,000
Now ΣX (wrong) = 4,000
ΣX (correct) = 4,000 – 30 – 27 + 3 + 72 = 4,018
̅ = ΣX(correct) = ΣX= 4,018 = 40.18
Correct X
N N 100

2. In a class of 50 students 10 have failed and their average of their marks is 2.5.
The total marks secured by the entire class were 281. Find the average marks
of those who have passed.

Solution:

Total marks of the 10 students who have failed = 2.5×10 = 25


Total marks of the 40 students who have passed = 281 – 25 = 256
256
Average marks obtained by those who have passed = = 6.4
40
3. Find the missing information in the following table

A B C Combined
Number (N) 10 8 x 24
Mean (X) 20 y 6 15

Solution:
x = 24 – 10 – 8 = 6
Now,
XA,B,C = (XANA + XBNB +XCNC) / (NA + NB + NC)

15 = [(20×10) + (8×y) + (6×6)] / (10 + 8 + 6)


15×24 = 200 + 8y + 36
360 = 236 + 8y
8y = 124
y = 124/8 = 15.5

4. Calculate arithmetic mean of the following

Solution :

L+U d
Class Interval (X) Frequency (f) m = d = m - 45 d’ = f d'
2 10
0-10 12 5 -40 -4 -48
10-20 16 15 -30 -3 -48
20-30 32 25 -20 -2 -64
30-40 52 35 -10 -1 -52
40-50 42 45 0 0 0
50-60 32 55 10 1 32
60-70 18 65 20 2 36
70-80 12 75 30 3 36
∑ f = 216 ∑ f d ′ = -108

∑ f d′ −108
̅
X = A + ∑ × C = 45 + × 10 = 45 – 5 = 40
f 216
5. In the following frequency distribution, the frequency of the class interval 40-50
is not known. Find it if the mean is 52.

Solution:

L+U
Wages(X) Number of workers (f) m= fm
2
10-20 5 15 75
20-30 3 25 75
30-40 4 35 140
40-50 f 45 45 f
50-60 2 55 110
60-70 6 65 390
70-80 13 75 975
∑ f = 33 + f ∑ f m = 1,765 + 45 f

∑f X
̅
X= ∑
f
1765+45f
52 =
33+f
52(33 + f) = 1,765 + 45f
1,716 + 52f = 1,765 + 45f
7f = 49
f=7
Statistics Ch-10 Measures of Central Tendency –
Median and Mode

1. Median

 Median is that value of a variable which divides the group into to equal
parts.
 Median is not affected by extreme observation and can be located graphically.
 But Median is not based on all observation and is less stable than Mean.

Calculating Median

a. Individual Series

 Arrange the data into increasing or decreasing order.


 Let N be the total number of items.
𝑵+𝟏
 Median = size of the( )th item, if N is odd
𝟐
𝑵 𝑵
𝐬𝐢𝐳𝐞 𝐨𝐟 𝐭𝐡𝐞 ( )𝐭𝐡 𝐢𝐭𝐞𝐦 + 𝐬𝐢𝐳𝐞 𝐨𝐟 𝐭𝐡𝐞 ( + 𝟏)𝐭𝐡 𝐢𝐭𝐞𝐦
𝟐 𝟐
 Median = ( ),
𝟐
if N is even

 Example – 1, Weight of eight students in kg – 71, 72, 64, 68, 70, 76, 73, 75

Solution
Weight in increasing order – 64, 68, 70, 71, 72, 73, 75, 76
N = 8(even), therefore
𝑁 𝑁
size of the ( 2 )th item + size of the ( 2 + 1)th item
Median = ( )
2
size of the 4th item + size of the 5th item
Median = ( )
2
71+72
Median = ( )
2
Median = 71.5
b. Discrete Series

 Arrange the data into increasing or decreasing order.


 Convert simple frequencies into cumulative frequencies.
 N = ∑𝐟
𝑵+𝟏
 Median = Size of the ( )th item
𝟐

 Example – 2,

Cumulative
Size Frequency (f)
frequency (cf)
2 2 2
3 3 5
4 8 13
5 10 23
6 12 35
7 16 51
8 10 61
9 8 69
10 6 75
N = ∑ 𝐟 = 75

𝑁+1
Median = Size of the ( )th item
2
Median = Size of the 38th item

38th item will first appear in 51st cf, therefore

Median = 7
c. Continuous Series

 Arrange the data into increasing or decreasing order of their class interval.
 Convert simple frequencies into cumulative frequencies.
 Median class is the class interval corresponding to that cumulative frequency
𝑁
which includes ( )th item.
2

 Formula

𝐍 − 𝐜𝐟
M=L+ 𝟐 ×i
𝐟
Where,
L = lower limit of the median class
cf = cumulative frequency of the class preceding median class
f = frequency of the median class
i = size of the median class interval

 Example – 3,

Number of Cumulative
Wage Rate
Workers (f) Frequency (cf)
0 - 10 22 22
10 - 20 38 60 (cf)
(L) 20 - 30 46 (f) 106
30 - 40 35 141
40 - 50 20 161
N = ∑ 𝐟 = 161

N
( )th item = 80.5th item, which will lie in class interval 20 – 30 (Median class)
2
Now,
𝐍 − 𝐜𝐟
M=L+ 𝟐 ×i
𝐟
80.5 − 60 × 10
M = 20 +
46
20.5 × 10
M = 20 +
46
M = 20 + 4.46
M = 24.46
Graphical Estimation of Median

Graphically Median is given by the corresponding value on X-axis where Less than Ogive
and More than Ogive cut each other.

Number of Cumulative Cumulative


Marks Marks Marks
Students Frequency (cf) Frequency (cf)
0-5 4 Less than 5 4 More than 0 100
5 - 10 6 Less than 10 10 More than 5 96
10 - 15 10 Less than 15 20 More than 10 90
15 - 20 10 Less than 20 30 More than 15 80
20 - 25 25 Less than 25 55 More than 20 70
25 - 30 22 Less than 30 77 More than 25 45
30 - 35 18 Less than 35 95 More than 30 23
35 - 40 5 Less than 40 100 More than 35 5
More than 40 0

Less than Ogive More than Ogive

120
100 100
100 96 95
90
Number of Students

80 77
80
70

60 55

45
40
20 23
10 30
20
0 4 5
0
0
0 5 10 15 20 25 30 35 40 45
Marks
Median = 24
2. Partition Values : Quartile, Decile and Percentile

A. Quartiles

 Quartiles divide data into 4 equal parts. There are 3 quartiles,

1. Q1 (Lower Quartile)
2. Q2 (Median)
3. Q3 (Upper Quartile)

 25% of the data lies below Q1 it and 75% of data above Q1.
 50% of the data lies below Q2 it and 50% of data above Q2.
 75% of the data lies below Q3 it and 25% of data above Q3.

B. Deciles

 Deciles divide data into 10 equal parts. There are 9 deciles,


 D5 = Median

C. Percentile

 Percentiles divide data into 100 equal parts. There are 99 percentiles.
 P50 = Median
3. Mode

 Mode is the most frequently occurring value of a variable.


 Mode can be located graphically using histogram of the corresponding data.

Calculating Mode

a. Individual Series

 If the number of observations is small, then mode can be located by inspection


method. Mode is the value that occurs most frequently.
 If the number of observations is large, using inspection method is difficult.
 Then, first convert individual series into a discrete series.
 Mode is the value corresponding to highest frequency.

b. Discrete Series

 Mode is the value corresponding to highest frequency.


 If more than one value commands the highest frequency, then we can use
grouping method.

 Grouping method:

1. It involves construction of two tables – Grouping Table and Analysis


Table

2. Grouping Table

a. There are 6 columns in the grouping table.


b. In Column I, frequencies are noted against each item as it is.
c. In Column II, frequencies are grouped into twos beginning with the first
item.
d. In Column III, frequencies are grouped into twos beginning with the
second item.
e. In Column IV, frequencies are grouped into threes beginning with the
first item.
f. In Column V, frequencies are grouped into threes beginning with the
second item.
g. In Column VI, frequencies are grouped into threes beginning with the
third item.
h. Highest value in all columns is encircled.
3. Analysis Table
It shows those items corresponding to which there are highest frequencies
in different column.

c. Continuous Series

 First identify modal class which is the class interval with the highest frequency. If
there is more than one class interval with highest frequency, then we can use
grouping method to identify modal class.

 After identifying modal class, use the following formula to find exact value of
mode

𝐟𝟏 −𝐟𝟎
Z=L+ ×i
𝟐𝐟𝟏 −𝐟𝟎−𝐟𝟐

Where,
L = Lower limit of the modal class
f1 = frequency of the modal class
f0 = frequency of pre modal class (class before modal class)
f2 = frequency of the post modal class (class after modal class)

 Example – 4,

Size Frequency
0-5 1
5 - 10 2
10 - 15 10
15 - 20 4
20 - 25 10
25 - 30 9
30 - 35 2
Grouping Table
Column Column Column Column Column Column
(I) (II) (III) (IV) (V) (VI)
Size (1) (1+2) (2+3) (1+2+3) (2+3+4) (3+4+5)
0-5 1 1+2=
5 - 10 3 1 + 2 + 10 =
2 2 + 10 = 13 2 + 10 + 4 =
10 - 15 10 10 + 4 = 12
16 10 + 4 + 10 =
15 - 20 4 14 4 + 10 =
4 + 10 + 9 = 24
20 - 25 10 10 + 9 = 14
23 10 + 9 + 2 =
25 - 30 9 19 9+2= 21
30 - 35 2 11

Analysis Table
Column 0 - 5 5 - 10 10 - 15 15 - 20 20 - 25 25 - 30 30 – 35
I ✔ ✔
II ✔ ✔
III ✔ ✔
IV ✔ ✔ ✔
V ✔ ✔ ✔
VI ✔ ✔ ✔
Total 0 0 2 3 6 3 1

Class Interval 20 – 25 has the largest number of ticks. Therefore modal class is
20 – 25

Size Frequency
0-5 1
5 - 10 2
10 - 15 10
15 - 20 4 ( 𝐟𝟎 )
20 - 25 10 ( 𝐟𝟏 )
25 - 30 9 ( 𝐟𝟐 )
30 - 35 2

𝐟𝟏 −𝐟𝟎 10−4 × 5 = 20 + 4.29 = 24.29


Z=L+ × i = 20 +
𝟐𝐟𝟏 −𝐟𝟎−𝐟𝟐 20−4−9
4. Relationship between Mean, Median and Mode

 Mode = 3 Median – 2 mean


 In symmetric distribution, Mean = Median = mode
 In positively skewed distribution, Mean > Median > Mode
 In negatively skewed distribution, Mean < Median < Mode

Y Positively skewed distribution

O X
Mode Median Mean

Y Negatively skewed distribution

O X
Mean Median Mode
Y
Symmetric distribution

O X
Mean = Median = Mode
Median and Mode Exercises

1. Calculate Median of the following data

Marks Number of students


Less than 10 4
Less than 20 16
Less than 30 40
Less than 40 76
Less than 50 96
Less than 60 112
Less than 70 120
Less than 80 125
Solution:

Number of Number of
Marks
students (cf) students (f)
0 - 10 4 4
10 - 20 16 16 – 4 = 12
20 - 30 40 (cf) 40 – 16 = 24
(L) 30 - 40 76 76 – 40 = 36 (f)
40 - 50 96 96 – 76 = 20
50 - 60 112 112 – 96 = 16
60 - 70 120 120 – 112 = 8
70 - 80 125 125 – 120 = 5
N = ∑ 𝐟 = 125

N
( )th item = 62.5th item, which will lie in class interval 30 – 40 (Median class)
2
Now,
𝐍 − 𝐜𝐟
M=L+ 𝟐 ×i
𝐟
M = 30 +
62.5 − 40 × 10
36
22.5 × 10
M = 30 +
36
M = 30 + 6.25
M = 36.25
2. Find the missing frequency of the class interval 20-30 when the Median is 28

C . I. f
0 - 10 5
10 - 20 8
20 - 30 x
30 - 40 16
40 - 50 6

Solution:

C . I. f cf
0 - 10 5 5
10 - 20 8 13 (cf)
(L) 20 - 30 x (f) 13 + x
30 - 40 16 29 + x
40 - 50 6 35 + x

N = 𝐟 = 35 + x

Median is 28, therefore Median Class is 20 – 30


Now,
𝐍 − 𝐜𝐟
M=L+ 𝟐 ×i
𝐟
35+ x − 13
28 = 20 + 2 x × 10
35+ x
8x = ( – 13 ) × 10
2
8x = 5 (35 + x) - 130
8x = 175 + 5x - 130
3x = 45
x = 15
3. Calculate mode of the following data:

C. I. f
20-24 3
25-29 5
30-34 10
35-39 20
40-44 12
45-49 6
50-54 3
55-59 1
Solution:

C. I. f
19.5 - 24.5 3
24.5 - 29.5 5
29.5 - 34.5 10 (f0)
(L) 34.5 - 39.5 20 (f1)
39.5 - 44.5 12 (f2)
44.5 - 49.5 6
49.5 - 54.5 3
54.5 - 59.5 1

𝐟𝟏 −𝐟𝟎
Z=L+ ×i
𝟐𝐟𝟏 −𝐟𝟎−𝐟𝟐
20 −10
Z = 34.5 + ×5
40−10−12
Z=L+
10 × 5
18
Z = 34.5 + 2.78
Z = 37.28

4. The mode and mean are 26.6 and 28.1 respectively in an asymmetric
distribution. Find out the value of the median.

Solution:

Mode = 3 Median – 2 Mean


26.6 = 3 Median – 2 (28.1)
3 Median = 26.6 + 56.2
82.8
Median = = 27.6
3
Statistics Ch-11 Correlation

1. Concept of Correlation

 Correlation is a measure of degree of relationship between two or more variables.

 Positive correlation –
a. Positive correlation exists when two variables moves in the same direction.
b. When one variable increases, other variable also increases and when one
variable decreases, other variable also decreases.
c. Example – Price and supply.

X 1 2 3 4 5
Y 10 20 30 40 50

 Negative correlation –
a. Negative correlation exists when two variables moves in the opposite direction.
b. When one variable increases, other variable decreases and when one variable
decreases, other variable increases.
c. Example – Price and demand

X 1 2 3 4 5
Y 50 40 30 20 10

 Linear correlation –
a. Linear correlation exists when two variables change in constant proportion.
b. It implies a straight line relationship.

X 2 4 6 8 10
Y 5 10 15 20 25

 Non-Linear correlation –
a. Non-Linear correlation exists when two variables change in variable proportion.
b. It implies a curvy relationship.

X 2 4 6 8
Y 3 8 18 20
 Simple correlation –
It implies the study of relationship between two variables.

 Multiple correlation –
It studies relationship among three or more variables simultaneously.
Example – Effect of rainfall and fertilizer on agricultural production.

2. Degree of Correlation

Degree Positive Negative


Perfect +1 -1
High (+ 0.75, + 1) (- 0.75, - 1)
Moderate (+ 0.25, + 0.75) (- 0.25, - 0.75)
Low (0, + 0.25) (0, - 0.25)
Zero 0 0

3. Methods of estimating Correlation

I. Scatter Diagram –
It is a very simple method but it does not measure the precise degree of correlation.
Y Perfect Y Y
No Correlation (0) Perfect
Negative (-1)
Positive (+1)

O X O X O X
Y Y
Negative (0, -1) Positive (0, +1)

O X O X
II. Karl Pearson’s Coefficient of correlation –

 Properties of correlation coefficient –


1) It has no units. It is a pure number.
2) Value of correlation coefficient lies between -1 and +1.
3) When r = 0, there is no correlation.
4) When r = 1, there is perfect positive correlation.
5) When r = -1, there is perfect negative correlation.

∑ 𝒙𝒚
a. Direct Method r= ̅,y=Y-𝐘
Here, x = X - 𝐗 ̅
√∑ 𝒙𝟐 ∑ 𝒚𝟐
Example – 1

∑ X 80
̅
X= = = 16
N 5
∑ Y 100
̅
Y= = = 34
N 5

X Y x=X-𝐗 ̅ x2 y=Y-𝐘 ̅ y2 xy
10 14 -6 36 -6 36 36
12 17 -4 16 -3 9 12
15 23 -1 1 3 9 -3
23 25 7 49 5 25 35
20 21 4 16 1 1 4
∑ 𝐗 = 80 ∑ 𝐗 = 100 ∑ 𝒙 = 0 ∑ 𝒙𝟐= 118 ∑ 𝒚 = 0 ∑ 𝒚𝟐= 80 ∑ 𝒙𝒚 = 84

∑ 𝒙𝒚
r=
√∑ 𝒙𝟐 ∑ 𝒚𝟐

𝟖𝟒
r=
√𝟏𝟏𝟖×𝟖𝟎
r = 0.864
(∑ 𝒅𝒙)×(∑ 𝒅𝒚)
∑ 𝒅𝒙𝒅𝒚 –
𝑵
b. Short cut Method r=
𝟐 𝟐
√∑ 𝒅𝒙𝟐 − (∑ 𝒅𝒙) ×√∑ 𝒅𝒚𝟐 − (∑ 𝒅𝒚)
𝑵 𝑵

Here,
dx = X - AX, where AX is assumed mean of series X
dy = Y - AY, where AY is assumed mean of series Y

This method is useful when mean value is not a whole number

Example – 2

X Y dx = X - AX dx2 dy = Y – AY dy2 dxdy


4 10 -4 16 - 10 100 40
6 15 -2 4 -5 25 10
8 (AX) 20 (AY) 0 0 0 0 0
15 25 7 49 5 25 35
20 30 12 144 10 100 120
∑ 𝐗 = 53 ∑ 𝐘 = 100 ∑ 𝒅𝒙 = 13 ∑ 𝒅𝒙𝟐= 213 ∑ 𝒅𝒚 = 0 ∑ 𝒅𝒚𝟐 = 250 ∑ 𝒅𝒙𝒅𝒚 = 205

∑ X 53
̅
X= = = 10.6, not a whole number
N 5
∑ Y 100
̅
Y= = = 20
N 5
(∑ 𝒅𝒙)×(∑ 𝒅𝒚)
∑ 𝒅𝒙𝒅𝒚 –
𝑵
r=
𝟐 𝟐
√∑ 𝒅𝒙𝟐 − (∑ 𝒅𝒙) ×√∑ 𝒅𝒚𝟐 − (∑ 𝒅𝒚)
𝑵 𝑵

(𝟏𝟑)×(𝟎)
𝟐𝟎𝟓 – 𝟓
r=
(𝟏𝟑)𝟐 (𝟎) 𝟐
√𝟐𝟏𝟑− ×√𝟐𝟓𝟎− 𝟓
𝟓

𝟐𝟎𝟓
r=
√𝟏𝟕𝟗.𝟐×√𝟐𝟓𝟎
𝟐𝟎𝟓
r=
𝟐𝟏𝟏.𝟕
r = 0.97
(∑ 𝐝𝐱′)×(∑ 𝐝𝐲′)
∑ 𝐝𝐱′𝐝𝐲′ –
𝐍
c. Step-Deviation Method r=
𝟐 𝟐
√∑(𝐝𝐱 ′ )𝟐 − (∑ 𝐝𝐱′) ×√∑(𝐝𝐲 ′ )𝟐 − (∑ 𝐝𝐲′)
𝐍 𝐍

Here,
𝑿 − 𝑨𝑿
dx = , where AX is assumed mean of X and CX is common factor of X.
𝑪𝑿
𝒀 − 𝑨𝒚
dy = , where AY is assumed mean of X and CY is common factor of Y.
𝑪𝒀
Example – 3

dx = dx' = dy = dy' =
X Y (dx’)2 (dy’)2 dx'dy’
X - AX dx/5 Y – AY dy/5
5 40 - 10 -2 4 10 2 4 -4
10 35 -5 -1 1 5 1 1 -1
15(AX) 30(AY) 0 0 0 0 0 0 0
20 25 5 1 1 -5 -1 1 -1
25 20 10 2 4 - 10 -2 4 -4
∑ 𝒅𝒙′ ∑(𝒅𝒙′)𝟐 ∑ 𝒅𝒚′ ∑(𝒅𝒚′)𝟐 ∑ 𝒅𝒙′𝒅𝒚′
=0 = 10 =0 = 10 = - 10

(∑ 𝐝𝐱′)×(∑ 𝐝𝐲 ′ )
∑ 𝐝𝐱 ′ 𝐝𝐲 ′ –
𝐍
r=
𝟐 𝟐
√∑(𝐝𝐱 ′ ) – (∑ 𝐝𝐱′) ×√∑(𝐝𝐲 ′ )𝟐 – (∑ 𝐝𝐲′)
𝟐
𝐍 𝐍

𝟎
− 𝟏𝟎 – 𝟓
r=
𝟎 𝟎
√𝟏𝟎− ×√𝟏𝟎−
𝟓 𝟓

−𝟏𝟎
r=
√𝟏𝟎×√𝟏𝟎

r = -1
III. Spearman’s Rank Correlation

It helps in calculating correlation of qualitative variables such as beauty,


bravery, ability, etc.
𝟔 ∑ 𝐃𝟐
Formula: rK = 1 –
𝐍𝟑− 𝐍
D = rank difference, N = number of pairs

Example – 4, Score of 2 judges in a beauty contest


X R1 Y R2 D = R1 – R2 D2
15 7 15 10 -3 9
17 9 12 9 0 0
14 6 4 3 3 9
13 5 6 5 0 0
11 3 7 6 -3 9
12 4 9 7 -3 9
16 8 3 2 6 36
18 10 10 8 2 4
10 2 2 1 1 1
9 1 5 4 -3 9
𝟐
∑ 𝐃 = 86

𝟔 ∑ 𝐃𝟐
rK = 1 –
𝐍𝟑− 𝐍
𝟔×𝟖𝟔
rK = 1 –
𝟏𝟎𝟑 − 𝟏𝟎
𝟓𝟏𝟔
rK = 1 –
𝟗𝟗𝟎
rK = 1 – 0.52
rK = 0.48
Statistics Ch-12 Index Numbers

1. Concept of Index Numbers

 Index numbers are devices for measuring change in a group of variables with
respect to time.

 Features of index number –

a. Index number measures relative or percentage change.

b. Index numbers shows change in terms of average.


Example – If inflation is 7%, it does not mean that price of every item has
increased by 7%.

 Difficulties in the construction of index numbers –

a. Purpose of index number – Before constructing an index number, one must


define an objective.

b. Selection of Base Year – Base Year is the reference year. It is the year with
which comparison is made. It should be a normal year.

c. Selection of variables – For example, to construct price index, one must decide
how many and which all goods and services are to be included in the index.

d. Selection of data source and collection of data – For example, if we have to


construct Price index number, then we have to decide whether to use retail price
or wholesale price.

e. Choice of average – a choice is to be made regarding which type of average to


use such as simple or weighted mean, arithmetic mean or geometric mean.

f. Selection of weights

g. Selection of formula
 Advantages of Index Numbers –

a. Price Index numbers can be used to measure the value of money over time.
Example - Consumer Price Index (CPI)

b. Price index is also used to adjust salaries and allowances.

c. Quantity index is used to obtain information regarding production.


Example – Index of industrial Production (IIP)

d. Index number gives information regarding economic condition can be used by


politicians to formulate economic policies. With the help of index number,
government determines its monetary and fiscal policy

2. Methods of Constructing Index Numbers

A. Simple Index Numbers – All items are given equal weights

a. Simple Aggregative Method


∑ 𝐏𝟏
Formula: P01 = × 100
∑ 𝐏𝟎
0 stands for base year
1 stands for current year
P1 is the price in current year
P0 is the price in base year

Example – 1

Commodity 2011 Price 2021 Price


A 50 80
B 40 60
C 10 20
D 5 10
E 2 6
∑ 𝐏𝟎 = 107 ∑ 𝐏𝟏 = 176

∑ 𝐏𝟏 𝟏𝟕𝟔
P01 = × 100 = × 100 = 164.49
∑ 𝐏𝟎 𝟏𝟎𝟕
b. Simple Average of Price Relative Method
𝐏𝟏
Price Relatives, R = × 100
𝐏𝟎
𝐏
∑𝐑 ∑( 𝟏 × 𝟏𝟎𝟎)
𝐏𝟎
Formula: P01 = =
𝐍 𝐍
Example – 2

𝐏𝟏
Commodity 2011 Price 2021 Price R= × 100
𝐏𝟎
200
Wheat 100 200 × 100 = 200
100
40
Ghee 8 40 × 100 = 500
8
16
Milk 2 16 × 100 = 800
2
800
Rice 200 800 × 100 = 400
200
6
Sugar 1 6 × 100 = 600
1
∑ 𝐑 = 2,500

∑𝐑 𝟐,𝟓𝟎𝟎
P01 = = = 50
𝐍 𝟓

B. Weighted Index Numbers – items are wighted according to their relative


importance.

a. Weighted Average of Price Relatives Method


𝐏𝟏
Price Relatives, R = × 100
𝐏𝟎
∑ 𝐑𝐖
Formula: P01 =
∑𝐖
Example – 3

2011 2021 𝐏𝟏
Commodity W R= × 100 RW
Price Price 𝐏𝟎
200
Wheat 40 100 200 × 100 = 200 8,000
100
40
Ghee 10 8 40 × 100 = 500 5,000
8
16
Milk 15 2 16 × 100 = 800 12,000
2
800
Rice 30 200 800 × 100 = 400 12,000
200
6
Sugar 5 1 6 × 100 = 600 3,000
1
∑ 𝐖 = 100 ∑ 𝐑 𝐖 = 40,000

∑ 𝐑𝐖 𝟒𝟎,𝟎𝟎𝟎
P01 = = = 400
∑𝐖 𝟏𝟎𝟎

b. Weighted Aggregative Methods – Goods are weighted according to the


quantity purchased

1) Laspeyre’s Method – uses base year quanitites as weights.


𝐋
∑ 𝐩𝟏 𝐪𝟎
Formula: 𝐏𝟎𝟏 = × 100
∑ 𝐩𝟎 𝐪𝟎

2) Paasche’s Method – uses current year quanitites as weights.


𝐏
∑ 𝐩𝟏 𝐪𝟏
Formula: 𝐏𝟎𝟏 = × 100
∑ 𝐩𝟎 𝐪𝟏

3) Fisher’s Method – combines both Laspeyre’s and Paasche’s method.


It is also considered as an ideal index because it satisfies both time
reversal and factor reversal Test.
∑𝐩 𝐪 ∑𝐩 𝐪 𝐋 𝐏
𝐅
Formula: 𝐏𝟎𝟏 = √[∑ 𝐩𝟏𝐪𝟎 × 𝟏𝟎𝟎] × [∑ 𝐩𝟏𝐪𝟏 × 𝟏𝟎𝟎] = √[𝐏𝟎𝟏 ] × [𝐏𝟎𝟏 ]
𝟎 𝟎 𝟎 𝟏
Example – 4

2011(Base Year) 2021 (Current Year)


Items
Price Quantity Price Quantity
A 10 10 20 25
B 35 3 40 10
C 30 5 20 15
D 10 20 8 20
E 40 2 40 5

2011 2021
(Base Year) (Current Year)
Items p0q0 p0q1 p1q0 p1q1
Price Quantity Price Quantity
(p0) (q0) (p1) (q1)
A 10 10 20 25 100 250 200 500
B 35 3 40 10 105 350 120 400
C 30 5 20 15 150 450 100 300
D 10 20 8 20 200 200 160 160
E 40 2 40 5 80 200 80 200
∑ 𝐩𝟎 𝐪𝟎 ∑ 𝐩𝟎 𝐪𝟏 ∑ 𝐩𝟏 𝐪𝟎 ∑ 𝐩𝟏 𝐪𝟏
= 635 = 1,450 = 660 = 1,560

𝐋
∑ 𝐩𝟏 𝐪𝟎 𝟔𝟔𝟎
𝐏𝟎𝟏 = × 100 = × 100 = 103.94
∑ 𝐩𝟎 𝐪𝟎 𝟔𝟑𝟓

𝐏
∑ 𝐩𝟏 𝐪𝟏 𝟏,𝟓𝟔𝟎
𝐏𝟎𝟏 = × 100 = × 100 = 107.59
∑ 𝐩𝟎 𝐪𝟏 𝟏,𝟒𝟓𝟎

∑𝐩 𝐪 ∑𝐩 𝐪 𝟔𝟔𝟎 𝟏,𝟓𝟔𝟎
𝐅
𝐏𝟎𝟏 = √[∑ 𝐩𝟏 𝐪𝟎 × 𝟏𝟎𝟎] × [∑ 𝐩𝟏 𝐪𝟏 × 𝟏𝟎𝟎] = √[ ] [ ]
𝟎 𝟎 𝟎 𝟏 𝟔𝟑𝟓 × 𝟏𝟎𝟎 × 𝟏,𝟒𝟓𝟎 × 𝟏𝟎𝟎 = 105
3. Consumer Price Index[CPI] or Cost of Living Index Number

 CPI is an index that measures the change in prices paid by a specific class of
consumers for goods and services consumed by them.

 In India CPI is constructed for following consumers groups –


a) Industrial Workers (IW)
b) Urban-Non Manual Employees (UNME)
c) Agricultural Labourers (AL)

 Base Year – 2011-12

 It is used by government for formulation of price policy and minimum wage


adjustment and calculation of dearness allowance (DA).

It is also used for measurement of real value of money.

 Methods of constructing CPI

a Aggregative Expenditure Method – It is similar to Laspeyre’s method.

∑ 𝐩𝟏 𝐪𝟎
Consumer Price Index (CPI) = × 100
∑ 𝐩𝟎 𝐪𝟎

b Family Budget Method – It is weighted average of price relative.

∑ 𝐑𝐖
Consumer Price Index (CPI) =
∑𝐖
, Where

𝐏𝟏
R= × 100
𝐏𝟎

W = p0q0 (expenditure on good in base year)

𝐂𝐏𝐈𝟏 −𝐂𝐏𝐈𝟎
 Rate of inflation from time period 0 to 1 = ( ) × 100
𝐂𝐏𝐈𝟎
Example – 5

Quantity Price Price


Items
Base year Base Year Current Year
Rice 5 24 30
Wheat 1 16 20
Pulses 2 12 18
Ghee 4 5 6.25
Oil 5 4 5
Clothing 40 1 1.5
Firewood 10 2 2.5
House Rent 1 20 25

a Aggregative Expenditure Method

Quantity Price Price Aggregate Aggregate


Base Base Current Expenditure in Expenditure in
Items
year Year Year Base Year Base Year
(q0) (p0) (p1) (p0q0) (p0q0)
Rice 5 24 30 120 150
Wheat 1 16 20 16 20
Pulses 2 12 18 24 36
Ghee 4 5 6.25 20 25
Oil 5 4 5 20 25
Clothing 40 1 1.5 40 60
Firewood 10 2 2.5 20 25
House Rent 1 20 25 20 25
∑ 𝐩𝟎 𝐪𝟎 = 280 ∑ 𝐩𝟏 𝐪𝟎 = 366

∑ 𝐩𝟏 𝐪𝟎 𝟑𝟔𝟔
Consumer Price Index (CPI) = × 100 = × 100 = 130.71
∑ 𝐩𝟎 𝐪𝟎 𝟐𝟖𝟎
b Family Budget Method

Quantity Price Price


Base Base Current 𝐏𝟏
Items R= × 100 W = p0q0 RW
Year Year Year 𝐏𝟎
(q0) (p0) (p1)
Rice 5 24 30 125 120 15,000
Wheat 1 16 20 125 16 2,000
Pulses 2 12 18 150 24 3,600
Ghee 4 5 6.25 125 20 2,500
Oil 5 4 5 125 20 2,500
Clothing 40 1 1.5 150 40 6,000
Firewood 10 2 2.5 125 20 2,500
House Rent 1 20 25 125 20 2,500
∑ 𝐑𝐖 =
∑ 𝐖 =280
36,600

∑ 𝐑𝐖 𝟑𝟔,𝟔𝟎𝟎
Consumer Price Index (CPI) = = × 100 = 130.71
∑𝐖 𝟐𝟖𝟎
4. Important Index Numbers

A. Wholesale Price Index (WPI) –

 Base Year – 2011-12

B. Index of Industrial Production –

 It measures the change in level of industrial production.


 Base Year – 2011-12
 Different Groups of items – Mining, Manufacturing, Electricity
∑ 𝐑𝐖
 Formula: IIP =
∑𝐖
𝐪𝟏
R= × 100
𝐪𝟎

C. SENSEX –

 It measures the change in stock prices of 30 major Indian Companies.

D. Human Development Index (HDI) –

 It an index developed by UN to measure level of social and economic


development of different countries.

 It focuses on three basic indicators – Life Expectancy, Education, Standard of


Living.

You might also like