Statistics
Statistics
All Statistics are data, but all data are not statistics.
Example – ‘Ram has a pocket allowance of 500’ is not statistics because it is neither
an aggregate nor an average. ‘Average pocket allowance of students of class 11 in a
school is 500’ is a statistics.
2. Features/Characteristics of Statistics
3. Limitations of Statistics
Study numerical facts only – Statistics does not study qualitative phenomenon
such as honesty, friendship, knowledge.
Working out cause and effect relationship – to diagnose the problem (finding
the cause of unemployment) and then offering suggestions (like increasing expenditure
on education).
Formulation of policies
Statistics Ch-2 Collection of Data
A. Primary Data
Primary data is the first hand data collected by an investigator for its own purpose and
for the first time. It is the data that is collected by the investigator from its source of
origin.
B. Secondary data
Secondary data is the second hand data which is already in existence and which has
been collected for some other purpose. It implies collection of data from some agency
or institution that has already collected the data.
Personal interview
Questionnaires
Telephonic Interview
Published sources
a. Government publications – Annual survey of industries, RBI bulletin
b. Reports of committee and commission – Finance commission, NITI Aayog
c. Publication of research institutes – Indian Statistical Institute, NCERT
d. Journals and newspaper
e. International publication – IMF, World Bank, ILO
a. Census of India
It is a decennial publication related to population size and its related parameters.
It is published by Registrar General and Census Commissioner, India. It
includes information on
Population size, growth and distribution
Density of population
Sex composition
State of literacy
Information on above parameter relates to country as a whole as well as different
states and UTs. It covers each and every household of country.
b. National Sample Survey Office(NSSO)
It is a government organization under MoSPI (Ministry of Statistics and Program
Implementation). It conducts regular sample surveys to collect a variety of
information like
Land & Livestock Holding
Employment and Unemployment
Consumer expenditure
Statistics Ch-3 Census and Sample Method
Example - If an investigation relates to all 2000 students in a college then 2000 would
be taken as population. But if investigation covers only some students, then it is a
sample inquiry.
Coverage In census method data is collected In sample method data is collected only
covering every item of the from a part or subset of population.
population
Suitability This method is suitable when area This method is suitable when area of
of investigation is relatively small investigation is relatively large and
and population is diverse. population is homogeneous
Accuracy This method has greater accuracy This method has less accuracy and
because every item of population is reliability because it studies only few
studied items of populations.
Cost and It is more time consuming and It is less time consuming and less
time costly expensive
3. Method of sampling
A. Random sampling
It is a method of sampling in which all item of population do not have equal chance of
being selected.
a. Purposive sampling
In this method, investigator himself makes the choice of the sample item.
It is suitable when some items of population are special and should be
included in sample.
Example – Inclusion of Tata Iron and Steel Company in the study of Iron
and Steel industry in India.
There is possibility of personal bias and sample could be tuned to serve
the purpose of investigator.
b. Stratified sampling
In this method, population is divided into different strata/groups having
different characteristics and then some items are selected proportionally
from each strata/group.
This method covers diverse characteristics of the population.
Example – Division of the population into low, middle, and high-income
groups and then a random selection from each group proportionally.
c. Systematic sampling
According to this method, population is numerically arranged and then
every nth item is selected.
It is a short cut method of random sampling in which only first item is
selected randomly.
Example – A population consist of 1000 items and a sample of 100 is
required. One item from the first 10 items is selected randomly. If the 5 th
item is selected then subsequent items selected would be 15 th, 25th, 35th
and so on.
d. Quota sampling
In this method population is divided into groups and the enumerator is told
to select a fixed number of items from each group.
Example – a sample consisting of 250 males and 250 females irrespective
of the proportion of male and female in the population.
e. Convenience sampling
Sampling is done is such as manner that suits the investigator
convenience.
Example – surveying people in a mall, on the street, and in other crowded
locations.
Statistics Ch-4 Organisation of Data
1. Classification
Classification is the process of arranging things into groups or classes on the basis of
their characteristics.
Basis of classification –
a. Geographical/Spatial classification –
This type of classification is based on geographical or locational differences.
Example –
State Rice production (mt)
Punjab 200
Haryana 300
UP 500
b. Chronological classification –
This type of classification is based on time.
Example – Annual sales of a firm
Year Sales(₹)
2020 2,000
2021 3,000
2022 5,000
c. Qualitative classification –
This type of classification is based on Qualities or Attributes of the data,
Example – Classification of data on the basis of gender, occupation or
religion of the population.
It is of two types – Simple and Manifold
Simple Classification – It is called classification according to dichotomy
(existence or absence of quality). Example – Male/Female
Educated/Uneducated etc.
Manifold classification – This type of classification involve more than one
characteristic. Example – Following classification of factory workers
Worker
s
Male Female
d. Quantitative classification –
This type of classification is based on numerical value of the facts.
Example – Annual profit of small scale firms in UP in 2022
Annual Profit (₹) Number of firms
Below 1,00,000 500
1,00,000 – 2,00,000 150
2,00,000 – 3,00,000 2000
3,00,000 – 4,00,000 40
Above 4,00,000 30
2. Concept of Variable
a. Individual series
It is a series without frequency in which all items are listed singly.
Example – marks obtained by students in a test
Roll No Marks(out of 50)
1 35
2 40
3 47
4 13
5 25
c. Frequency Distribution
In this series, data is classified into different ranges or limits called class
intervals and corresponding to each class interval there is some frequency.
Every class interval has a lower limit as well as upper limit.
Marks Frequency
0-10 4
10-20 3
20-30 1
30-40 4
40-50 3
4. Types of Frequency Distribution
a. Exclusive series
Exclusive series is that series in which every class interval excludes items
corresponding to its upper limit.
Marks Frequency
0-10 4
10-20 3
20-30 1
30-40 4
40-50 3
b. Inclusive series
Inclusive series is that series in which every class interval includes items corresponding
to its upper limit.
Marks Frequency
0-9 4
10-19 3
20-29 1
30-39 4
40-49 3
In open end series, lower limit of the first class interval and upper limit of last class
interval is missing.
Marks Frequency
Below 10 4
10-20 3
20-30 1
30-40 4
40 and Above 3
d. Cumulative series
Marks Frequency
Less than 10 4
Less than 20 7
Less than 30 8
Less than 40 12
Less than 50 15
e. Mid-values series
In this type of series, we have only mid-values of the class interval and the
corresponding frequencies.
Mid-value Frequency
5 4
15 3
25 1
35 4
45 3
Statistics Ch-5 Presentation of Data – Textual and
Tabular
a. Textual
b. Tabular
c. Diagrammatic –
2. Textual Presentation
Example – Survey conduction by an NGO, in the state of Punjab, revealed that area
under pulses cultivation tended to shrink by 40% and area under rice and wheat
cultivation has tended to expand by 20%, between the years 2020-21.
3. Tabular Presentation
Components of a table –
Table Number -
Title -
Captions (Column Headings)
Stubs (Row Heading) Column Heading 1 Column Heading 2
Row Heading 1 Cell Cell
Row Heading 2 Cell Cell Body of
Row Heading 3 Cell Cell the Table
Row Heading 4 Cell Cell
Types –
a. Two way or double – shows two characteristics
b. Three way or Treble – shows three characteristics
c. Manifold – shows more than three characteristics
Number of Students
Class Total
Boys Girls
B.A. (I) 40 60 100
B.A. (II) 60 20 80
B.A. (III) 20 20 40
Total 120 100 220
Statistics Ch-6 Bar Diagrams and Pie Diagrams
30
30 28
24
25
20
15
10
5
0
1951-60 1961-70 1971-80 1981-90 1991-00
YEAR
Birth rate in India (Horizontal)
1991-00 24
YEAR 1981-90 28
1971-80 30
1961-70 35
1951-60 45
0 10 20 30 40 50
BIRTH RATE
Birth Death
Year
rate rate
1951-60 40 27
1961-70 42 23
1971-80 41 19
1981-90 37 15
1991-00 32.5 11.5
35
30
25
20 BIRTH RATE
15 DEATH RATE
10
5
0
1951-60 1961-70 1971-80 1981-90 1991-00
YEAR
c. Component bar diagrams
These diagrams simultaneously present total value as well as part values of the
data.
100 89
71 82
80 64
Thermal
60
Hydro
40
20 46 49 48 51
0
2017-18 2018-19 2019-20 2020-21
YEAR
Family A Family B
Expenditure Expenditure Expenditure Expenditure
Item
(₹) (%) (₹) (%)
Food 200 (200/500)*100 = 40 300 (300/1,000)*100 = 30
Clothing 100 (100/500)*100 = 20 200 (200/1,000)*100 = 20
Rent 150 (150/500)*100 = 30 100 (100/1,000)*100 = 10
Fuel 50 (50/500)*100 = 10 400 (400/1,000)*100 = 40
Total 500 100 1,000 100
Y
100 100
100 -
10
90 -
Food
80 - 40
30
70 - Clothing
Expenditure 60 -
10 10 Rent
50 - 20
40 - 20 Fuel
30 -
20 - 40
30
10 -
0
Family A Family B X
Year Savings/Deficit
2016 30
2017 -20
2018 10
2019 15
2020 -25
Saving/Deficit
40
30
30
Saving/Deficit
20 15
10
10
0
2016 2017 2018 2019 2020
-10
-20
-20
-30 -25
YEAR
2. Pie diagrams
Pie diagram is a circle divided into various segments showing different values of a series.
150
North Indian 150 × 360° = 108°
500
100
South Indian 100 × 360° = 72°
500
125
Chinese 125 × 360° = 90°
500
75
Italian 75 × 360° = 54°
500
50
Mexican 50 × 360° = 36°
500
North Indian,
Italian, 54° 108°
Chinese, 90°
South Indian,
72°
Statistics Ch-7 Frequency Diagrams – Histogram,
Polygon and Ogive
1. Histogram
A histogram is a graphical representation of a frequency distribution of a continuous series.
Histograms of frequency distribution are of two types –
Y
Marks in Statistics
30
25
Number of 20
20
Students 15
15 12
10
10 8
5 4
5
X
O 10 20 30 40 50 60 70
Marks
b. Histogram of Unequal class intervals
With unequal class interval, frequencies are to be adjusted first before presenting data.
30
27
Number 25
of
Workers 20
15
15 12
10
10 8
7 6
5 3
2
X
O 5 5 10 10 1515 20
20 25
25 30
30 3535 40 4045 4550 50
55 60
55 65
60 7065 75 7080 75 80
Weekly Wages
2. Frequency polygon
It can be formed by joining mid points of the tops of all rectangles in a histogram.
Y
Marks in Statistics
30
25
Number of 20
20
Students 15
15 12
10
10 8
5
5
X
-10 0 10 20 30 40 50 60 70
Marks
100
77
80
55
60
40 30
20
20 10
4
0
0 5 10 15 20 25 30 35 40 45
Marks
100 90
80
80 70
60 45
40
23
20 5
0
0
0 5 10 15 20 25 30 35 40 45
Marks
Statistics Ch-8 Arithmetic line graphs OR Time series
graphs
Year Production
2015 5
2016 8
2017 13
2018 16
2019 20
2020 17
15 13
10 8
5
5
0
2014 2015 2016 2017 2018 2019 2020 2021
Year
2. Two or more variables graph
An Average is a single value within the range of data that is used to represent all
values in the series.
a. Stable definition
b. Representative – should be based on all observation.
c. Capable of further algebraic treatment.
d. Least effect of a change in sample.
A. Individual series
∑𝐗
̅=
a. Direct Method: 𝐗
𝐍
Pocket
Allowances (X)
15
20
30
22
25
18
40
50
55
65
∑ 𝐗 = 340
̅ = ∑ 𝐗 = 𝟑𝟒𝟎 = 34
𝐗
𝐍 𝟏𝟎
∑𝐝
̅=𝐀+
b. Assumed Mean Method: 𝐗 ,
𝐍
Where
‘A’ is assumed mean
‘d’ is deviation from arithmetic mean, d = X - A
̅ = 𝐀 + ∑ 𝐝 = 40 +
𝐗
−𝟔𝟎
= 40 – 6 = 34
𝐍 𝟏𝟎
∑ 𝐝′
̅= 𝐀+
c. Step-Deviation Method: 𝐗 ×C
𝐍
Where
d’ = d/C and C is the common factor
B. Discrete series
∑𝐟 𝐗
̅
a. Direct Method : 𝐗 = ∑𝐟
Where ‘ f ’ is frequency
Example – 4, Weekly wage earning of 19 workers
̅ = ∑ 𝐟 𝐗 = 𝟓𝟔𝟎 = 29.47
𝐗 ∑ 𝐟 𝟏𝟗
̅ = 𝐀 + ∑𝐟 𝐝
b. Assumed Mean Method : 𝐗
∑𝐟
Example – 5, Weekly wage earning of 19 workers
̅ = 𝐀 + ∑ 𝐟 𝐝 = 𝟑𝟎 + −𝟏𝟎 = 29.47
𝐗 ∑𝐟 𝟏𝟗
∑ 𝐟 𝐝′
̅= 𝐀+
c. Step-Deviation Method : 𝐗 ×C
∑𝐟
̅ = 𝐀 + ∑ 𝐟 𝐝′ × 𝐂 = 𝟑𝟎 +
𝐗
−𝟏
× 𝟏𝟎 = 29.47
∑𝐟 𝟏𝟗
C. Continuous series
∑𝒇 𝒎
̅
a. Direct Method : 𝑿 = ∑𝒇
Where ‘m’ is mid value
𝑳+𝑼
m= , L = lower limit, U = upper limit
𝟐
Example – 7, Marks secured by Class X students in English
𝑳+𝑼
Marks (X) Number of students (f) m= fm
𝟐
0 - 10 20 5 100
10 - 20 24 15 360
20 - 30 40 25 1,000
30 - 40 36 35 1,260
40 - 50 20 45 900
∑ 𝒇 = 140 ∑ 𝒇 𝒎 = 3,620
̅ = ∑ 𝐟 𝐦 = 𝟑,𝟔𝟐𝟎 = 25.86
𝐗 ∑ 𝐟 𝟏𝟒𝟎
∑𝒇 𝒅
̅ =𝑨+
b. Assumed Mean Method : 𝑿
∑𝒇
Where d = m – A
𝑳+𝑼
Marks (X) Number of students (f) m= d = m - 25 fd
𝟐
0 - 10 20 5 -20 -400
10 - 20 24 15 -10 -240
20 - 30 40 25 0 0
30 - 40 36 35 10 360
40 - 50 20 45 20 400
∑ 𝒇 = 140 ∑ 𝒇 𝒅 = 120
̅ = 𝐀 + ∑ 𝐟 𝐝 = 𝟐𝟓 +
𝐗
𝟏𝟐𝟎
= 25.86
∑ 𝐟 𝟏𝟒𝟎
∑ 𝒇 𝒅′
̅ = 𝑨+
c. Step-Deviation Method : 𝑿 ×C
∑𝒇
𝒎−𝑨
Where d’ =
𝑪
𝑳+𝑼 𝒅
Marks (X) Number of students (f) m= d = m - 25 d’ = f d'
𝟐 𝟏𝟎
0 - 10 20 5 -20 -2 -40
10 - 20 24 15 -10 -1 -24
20 - 30 40 25 0 0 0
30 - 40 36 35 10 1 36
40 - 50 20 45 20 2 40
∑ 𝐟 = 140 ∑ 𝐟 𝐝 ′ = 12
̅ = 𝐀 + ∑ 𝐟 𝐝′ × 𝐂 = 𝟐𝟓 +
𝐗
𝟏𝟐
× 𝟏𝟎 = 25.86
∑ 𝐟 𝟏𝟒𝟎
∑ 𝐖𝐗
Weighted Arithmetic Mean: XW =
∑𝐖
Example – 10
∑ 𝐖𝐗 ∑ 𝟓𝟔𝟎
XW = = = 29.47
∑𝐖 ∑ 𝟏𝟗
4. Combined Arithmetic Mean
Combined Arithmetic Mean is the mean of the series formed by combining two or more
series.
Example – 11
60 students of Section A of Class XI obtained 40 mean marks in Statistics. 40 students
of Section B obtained 35 mean marks in Statistics. Find mean marks in Statistics for
Class XI as a whole.
Marks (X) d = X - 30
10 -20
20 -10
30 0
40 10
50 20
∑ 𝐗 = 150 ∑𝐝 = 0
̅ = ∑ 𝐗 = 𝟏𝟓𝟎 = 30
𝐗
𝐍 𝟓
b. If a specific value is added to or subtracted from all items in a series, the mean
value if the items would increase or decrease by same value. Likewise, if all
items of a series are multiplied or divided by any specific value, the mean value
of the items would be multiplied or divided by the same specific value.
Merits –
Simple and Certain
Based on all items therefore it is stable (least affected by sampling fluctuations)
Capable of algebraic treatment.
Demerit –
Misleading conclusions :
Example – In following situation, average profit of firm 1 and firm 2 are same but
the profit of firm 1 is increasing overtime and that of firm 2 is decreasing overtime.
Firm 1 Firm 2
Profit Year 1 1,00,000 3,00,000
Profit Year 2 2,00,000 2,00,000
Profit Year 3 3,00,000 1,00,000
Average 2,00,000 3,00,000
∑𝐗
Direct Method
𝐍
∑ 𝐟 𝐝′ 𝐗−𝐀
Step-Deviation Method 𝐀+ ×C d' =
𝐍 𝐂
∑𝐟𝐗
Direct Method
∑𝐟
∑ 𝐟 𝐝′ 𝐗−𝐀
Step-Deviation Method 𝐀+ ∑𝐟
×C d' =
𝐂
∑𝐟𝐦 𝐋+𝐔
Direct Method m=
∑𝐟 𝟐
∑ 𝐟 𝐝′ 𝐦−𝐀
Step-Deviation Method 𝐀+ ∑𝐟
×C d' =
𝐂
Weighted ∑ 𝐖𝐗
XW = ∑𝐖
Mean
Combined
X1,2 = (X1N1 + X2N2) / (N1 + N2)
Mean
Arithmetic Mean (Exercises)
Solution:
ΣX
̅
X=
N
ΣX
40 =
100
ΣX = 4,000
Now ΣX (wrong) = 4,000
ΣX (correct) = 4,000 – 30 – 27 + 3 + 72 = 4,018
̅ = ΣX(correct) = ΣX= 4,018 = 40.18
Correct X
N N 100
2. In a class of 50 students 10 have failed and their average of their marks is 2.5.
The total marks secured by the entire class were 281. Find the average marks
of those who have passed.
Solution:
A B C Combined
Number (N) 10 8 x 24
Mean (X) 20 y 6 15
Solution:
x = 24 – 10 – 8 = 6
Now,
XA,B,C = (XANA + XBNB +XCNC) / (NA + NB + NC)
Solution :
L+U d
Class Interval (X) Frequency (f) m = d = m - 45 d’ = f d'
2 10
0-10 12 5 -40 -4 -48
10-20 16 15 -30 -3 -48
20-30 32 25 -20 -2 -64
30-40 52 35 -10 -1 -52
40-50 42 45 0 0 0
50-60 32 55 10 1 32
60-70 18 65 20 2 36
70-80 12 75 30 3 36
∑ f = 216 ∑ f d ′ = -108
∑ f d′ −108
̅
X = A + ∑ × C = 45 + × 10 = 45 – 5 = 40
f 216
5. In the following frequency distribution, the frequency of the class interval 40-50
is not known. Find it if the mean is 52.
Solution:
L+U
Wages(X) Number of workers (f) m= fm
2
10-20 5 15 75
20-30 3 25 75
30-40 4 35 140
40-50 f 45 45 f
50-60 2 55 110
60-70 6 65 390
70-80 13 75 975
∑ f = 33 + f ∑ f m = 1,765 + 45 f
∑f X
̅
X= ∑
f
1765+45f
52 =
33+f
52(33 + f) = 1,765 + 45f
1,716 + 52f = 1,765 + 45f
7f = 49
f=7
Statistics Ch-10 Measures of Central Tendency –
Median and Mode
1. Median
Median is that value of a variable which divides the group into to equal
parts.
Median is not affected by extreme observation and can be located graphically.
But Median is not based on all observation and is less stable than Mean.
Calculating Median
a. Individual Series
Example – 1, Weight of eight students in kg – 71, 72, 64, 68, 70, 76, 73, 75
Solution
Weight in increasing order – 64, 68, 70, 71, 72, 73, 75, 76
N = 8(even), therefore
𝑁 𝑁
size of the ( 2 )th item + size of the ( 2 + 1)th item
Median = ( )
2
size of the 4th item + size of the 5th item
Median = ( )
2
71+72
Median = ( )
2
Median = 71.5
b. Discrete Series
Example – 2,
Cumulative
Size Frequency (f)
frequency (cf)
2 2 2
3 3 5
4 8 13
5 10 23
6 12 35
7 16 51
8 10 61
9 8 69
10 6 75
N = ∑ 𝐟 = 75
𝑁+1
Median = Size of the ( )th item
2
Median = Size of the 38th item
Median = 7
c. Continuous Series
Arrange the data into increasing or decreasing order of their class interval.
Convert simple frequencies into cumulative frequencies.
Median class is the class interval corresponding to that cumulative frequency
𝑁
which includes ( )th item.
2
Formula
𝐍 − 𝐜𝐟
M=L+ 𝟐 ×i
𝐟
Where,
L = lower limit of the median class
cf = cumulative frequency of the class preceding median class
f = frequency of the median class
i = size of the median class interval
Example – 3,
Number of Cumulative
Wage Rate
Workers (f) Frequency (cf)
0 - 10 22 22
10 - 20 38 60 (cf)
(L) 20 - 30 46 (f) 106
30 - 40 35 141
40 - 50 20 161
N = ∑ 𝐟 = 161
N
( )th item = 80.5th item, which will lie in class interval 20 – 30 (Median class)
2
Now,
𝐍 − 𝐜𝐟
M=L+ 𝟐 ×i
𝐟
80.5 − 60 × 10
M = 20 +
46
20.5 × 10
M = 20 +
46
M = 20 + 4.46
M = 24.46
Graphical Estimation of Median
Graphically Median is given by the corresponding value on X-axis where Less than Ogive
and More than Ogive cut each other.
120
100 100
100 96 95
90
Number of Students
80 77
80
70
60 55
45
40
20 23
10 30
20
0 4 5
0
0
0 5 10 15 20 25 30 35 40 45
Marks
Median = 24
2. Partition Values : Quartile, Decile and Percentile
A. Quartiles
1. Q1 (Lower Quartile)
2. Q2 (Median)
3. Q3 (Upper Quartile)
25% of the data lies below Q1 it and 75% of data above Q1.
50% of the data lies below Q2 it and 50% of data above Q2.
75% of the data lies below Q3 it and 25% of data above Q3.
B. Deciles
C. Percentile
Percentiles divide data into 100 equal parts. There are 99 percentiles.
P50 = Median
3. Mode
Calculating Mode
a. Individual Series
b. Discrete Series
Grouping method:
2. Grouping Table
c. Continuous Series
First identify modal class which is the class interval with the highest frequency. If
there is more than one class interval with highest frequency, then we can use
grouping method to identify modal class.
After identifying modal class, use the following formula to find exact value of
mode
𝐟𝟏 −𝐟𝟎
Z=L+ ×i
𝟐𝐟𝟏 −𝐟𝟎−𝐟𝟐
Where,
L = Lower limit of the modal class
f1 = frequency of the modal class
f0 = frequency of pre modal class (class before modal class)
f2 = frequency of the post modal class (class after modal class)
Example – 4,
Size Frequency
0-5 1
5 - 10 2
10 - 15 10
15 - 20 4
20 - 25 10
25 - 30 9
30 - 35 2
Grouping Table
Column Column Column Column Column Column
(I) (II) (III) (IV) (V) (VI)
Size (1) (1+2) (2+3) (1+2+3) (2+3+4) (3+4+5)
0-5 1 1+2=
5 - 10 3 1 + 2 + 10 =
2 2 + 10 = 13 2 + 10 + 4 =
10 - 15 10 10 + 4 = 12
16 10 + 4 + 10 =
15 - 20 4 14 4 + 10 =
4 + 10 + 9 = 24
20 - 25 10 10 + 9 = 14
23 10 + 9 + 2 =
25 - 30 9 19 9+2= 21
30 - 35 2 11
Analysis Table
Column 0 - 5 5 - 10 10 - 15 15 - 20 20 - 25 25 - 30 30 – 35
I ✔ ✔
II ✔ ✔
III ✔ ✔
IV ✔ ✔ ✔
V ✔ ✔ ✔
VI ✔ ✔ ✔
Total 0 0 2 3 6 3 1
Class Interval 20 – 25 has the largest number of ticks. Therefore modal class is
20 – 25
Size Frequency
0-5 1
5 - 10 2
10 - 15 10
15 - 20 4 ( 𝐟𝟎 )
20 - 25 10 ( 𝐟𝟏 )
25 - 30 9 ( 𝐟𝟐 )
30 - 35 2
O X
Mode Median Mean
O X
Mean Median Mode
Y
Symmetric distribution
O X
Mean = Median = Mode
Median and Mode Exercises
Number of Number of
Marks
students (cf) students (f)
0 - 10 4 4
10 - 20 16 16 – 4 = 12
20 - 30 40 (cf) 40 – 16 = 24
(L) 30 - 40 76 76 – 40 = 36 (f)
40 - 50 96 96 – 76 = 20
50 - 60 112 112 – 96 = 16
60 - 70 120 120 – 112 = 8
70 - 80 125 125 – 120 = 5
N = ∑ 𝐟 = 125
N
( )th item = 62.5th item, which will lie in class interval 30 – 40 (Median class)
2
Now,
𝐍 − 𝐜𝐟
M=L+ 𝟐 ×i
𝐟
M = 30 +
62.5 − 40 × 10
36
22.5 × 10
M = 30 +
36
M = 30 + 6.25
M = 36.25
2. Find the missing frequency of the class interval 20-30 when the Median is 28
C . I. f
0 - 10 5
10 - 20 8
20 - 30 x
30 - 40 16
40 - 50 6
Solution:
C . I. f cf
0 - 10 5 5
10 - 20 8 13 (cf)
(L) 20 - 30 x (f) 13 + x
30 - 40 16 29 + x
40 - 50 6 35 + x
∑
N = 𝐟 = 35 + x
C. I. f
20-24 3
25-29 5
30-34 10
35-39 20
40-44 12
45-49 6
50-54 3
55-59 1
Solution:
C. I. f
19.5 - 24.5 3
24.5 - 29.5 5
29.5 - 34.5 10 (f0)
(L) 34.5 - 39.5 20 (f1)
39.5 - 44.5 12 (f2)
44.5 - 49.5 6
49.5 - 54.5 3
54.5 - 59.5 1
𝐟𝟏 −𝐟𝟎
Z=L+ ×i
𝟐𝐟𝟏 −𝐟𝟎−𝐟𝟐
20 −10
Z = 34.5 + ×5
40−10−12
Z=L+
10 × 5
18
Z = 34.5 + 2.78
Z = 37.28
4. The mode and mean are 26.6 and 28.1 respectively in an asymmetric
distribution. Find out the value of the median.
Solution:
1. Concept of Correlation
Positive correlation –
a. Positive correlation exists when two variables moves in the same direction.
b. When one variable increases, other variable also increases and when one
variable decreases, other variable also decreases.
c. Example – Price and supply.
X 1 2 3 4 5
Y 10 20 30 40 50
Negative correlation –
a. Negative correlation exists when two variables moves in the opposite direction.
b. When one variable increases, other variable decreases and when one variable
decreases, other variable increases.
c. Example – Price and demand
X 1 2 3 4 5
Y 50 40 30 20 10
Linear correlation –
a. Linear correlation exists when two variables change in constant proportion.
b. It implies a straight line relationship.
X 2 4 6 8 10
Y 5 10 15 20 25
Non-Linear correlation –
a. Non-Linear correlation exists when two variables change in variable proportion.
b. It implies a curvy relationship.
X 2 4 6 8
Y 3 8 18 20
Simple correlation –
It implies the study of relationship between two variables.
Multiple correlation –
It studies relationship among three or more variables simultaneously.
Example – Effect of rainfall and fertilizer on agricultural production.
2. Degree of Correlation
I. Scatter Diagram –
It is a very simple method but it does not measure the precise degree of correlation.
Y Perfect Y Y
No Correlation (0) Perfect
Negative (-1)
Positive (+1)
O X O X O X
Y Y
Negative (0, -1) Positive (0, +1)
O X O X
II. Karl Pearson’s Coefficient of correlation –
∑ 𝒙𝒚
a. Direct Method r= ̅,y=Y-𝐘
Here, x = X - 𝐗 ̅
√∑ 𝒙𝟐 ∑ 𝒚𝟐
Example – 1
∑ X 80
̅
X= = = 16
N 5
∑ Y 100
̅
Y= = = 34
N 5
X Y x=X-𝐗 ̅ x2 y=Y-𝐘 ̅ y2 xy
10 14 -6 36 -6 36 36
12 17 -4 16 -3 9 12
15 23 -1 1 3 9 -3
23 25 7 49 5 25 35
20 21 4 16 1 1 4
∑ 𝐗 = 80 ∑ 𝐗 = 100 ∑ 𝒙 = 0 ∑ 𝒙𝟐= 118 ∑ 𝒚 = 0 ∑ 𝒚𝟐= 80 ∑ 𝒙𝒚 = 84
∑ 𝒙𝒚
r=
√∑ 𝒙𝟐 ∑ 𝒚𝟐
𝟖𝟒
r=
√𝟏𝟏𝟖×𝟖𝟎
r = 0.864
(∑ 𝒅𝒙)×(∑ 𝒅𝒚)
∑ 𝒅𝒙𝒅𝒚 –
𝑵
b. Short cut Method r=
𝟐 𝟐
√∑ 𝒅𝒙𝟐 − (∑ 𝒅𝒙) ×√∑ 𝒅𝒚𝟐 − (∑ 𝒅𝒚)
𝑵 𝑵
Here,
dx = X - AX, where AX is assumed mean of series X
dy = Y - AY, where AY is assumed mean of series Y
Example – 2
∑ X 53
̅
X= = = 10.6, not a whole number
N 5
∑ Y 100
̅
Y= = = 20
N 5
(∑ 𝒅𝒙)×(∑ 𝒅𝒚)
∑ 𝒅𝒙𝒅𝒚 –
𝑵
r=
𝟐 𝟐
√∑ 𝒅𝒙𝟐 − (∑ 𝒅𝒙) ×√∑ 𝒅𝒚𝟐 − (∑ 𝒅𝒚)
𝑵 𝑵
(𝟏𝟑)×(𝟎)
𝟐𝟎𝟓 – 𝟓
r=
(𝟏𝟑)𝟐 (𝟎) 𝟐
√𝟐𝟏𝟑− ×√𝟐𝟓𝟎− 𝟓
𝟓
𝟐𝟎𝟓
r=
√𝟏𝟕𝟗.𝟐×√𝟐𝟓𝟎
𝟐𝟎𝟓
r=
𝟐𝟏𝟏.𝟕
r = 0.97
(∑ 𝐝𝐱′)×(∑ 𝐝𝐲′)
∑ 𝐝𝐱′𝐝𝐲′ –
𝐍
c. Step-Deviation Method r=
𝟐 𝟐
√∑(𝐝𝐱 ′ )𝟐 − (∑ 𝐝𝐱′) ×√∑(𝐝𝐲 ′ )𝟐 − (∑ 𝐝𝐲′)
𝐍 𝐍
Here,
𝑿 − 𝑨𝑿
dx = , where AX is assumed mean of X and CX is common factor of X.
𝑪𝑿
𝒀 − 𝑨𝒚
dy = , where AY is assumed mean of X and CY is common factor of Y.
𝑪𝒀
Example – 3
dx = dx' = dy = dy' =
X Y (dx’)2 (dy’)2 dx'dy’
X - AX dx/5 Y – AY dy/5
5 40 - 10 -2 4 10 2 4 -4
10 35 -5 -1 1 5 1 1 -1
15(AX) 30(AY) 0 0 0 0 0 0 0
20 25 5 1 1 -5 -1 1 -1
25 20 10 2 4 - 10 -2 4 -4
∑ 𝒅𝒙′ ∑(𝒅𝒙′)𝟐 ∑ 𝒅𝒚′ ∑(𝒅𝒚′)𝟐 ∑ 𝒅𝒙′𝒅𝒚′
=0 = 10 =0 = 10 = - 10
(∑ 𝐝𝐱′)×(∑ 𝐝𝐲 ′ )
∑ 𝐝𝐱 ′ 𝐝𝐲 ′ –
𝐍
r=
𝟐 𝟐
√∑(𝐝𝐱 ′ ) – (∑ 𝐝𝐱′) ×√∑(𝐝𝐲 ′ )𝟐 – (∑ 𝐝𝐲′)
𝟐
𝐍 𝐍
𝟎
− 𝟏𝟎 – 𝟓
r=
𝟎 𝟎
√𝟏𝟎− ×√𝟏𝟎−
𝟓 𝟓
−𝟏𝟎
r=
√𝟏𝟎×√𝟏𝟎
r = -1
III. Spearman’s Rank Correlation
𝟔 ∑ 𝐃𝟐
rK = 1 –
𝐍𝟑− 𝐍
𝟔×𝟖𝟔
rK = 1 –
𝟏𝟎𝟑 − 𝟏𝟎
𝟓𝟏𝟔
rK = 1 –
𝟗𝟗𝟎
rK = 1 – 0.52
rK = 0.48
Statistics Ch-12 Index Numbers
Index numbers are devices for measuring change in a group of variables with
respect to time.
b. Selection of Base Year – Base Year is the reference year. It is the year with
which comparison is made. It should be a normal year.
c. Selection of variables – For example, to construct price index, one must decide
how many and which all goods and services are to be included in the index.
f. Selection of weights
g. Selection of formula
Advantages of Index Numbers –
a. Price Index numbers can be used to measure the value of money over time.
Example - Consumer Price Index (CPI)
Example – 1
∑ 𝐏𝟏 𝟏𝟕𝟔
P01 = × 100 = × 100 = 164.49
∑ 𝐏𝟎 𝟏𝟎𝟕
b. Simple Average of Price Relative Method
𝐏𝟏
Price Relatives, R = × 100
𝐏𝟎
𝐏
∑𝐑 ∑( 𝟏 × 𝟏𝟎𝟎)
𝐏𝟎
Formula: P01 = =
𝐍 𝐍
Example – 2
𝐏𝟏
Commodity 2011 Price 2021 Price R= × 100
𝐏𝟎
200
Wheat 100 200 × 100 = 200
100
40
Ghee 8 40 × 100 = 500
8
16
Milk 2 16 × 100 = 800
2
800
Rice 200 800 × 100 = 400
200
6
Sugar 1 6 × 100 = 600
1
∑ 𝐑 = 2,500
∑𝐑 𝟐,𝟓𝟎𝟎
P01 = = = 50
𝐍 𝟓
2011 2021 𝐏𝟏
Commodity W R= × 100 RW
Price Price 𝐏𝟎
200
Wheat 40 100 200 × 100 = 200 8,000
100
40
Ghee 10 8 40 × 100 = 500 5,000
8
16
Milk 15 2 16 × 100 = 800 12,000
2
800
Rice 30 200 800 × 100 = 400 12,000
200
6
Sugar 5 1 6 × 100 = 600 3,000
1
∑ 𝐖 = 100 ∑ 𝐑 𝐖 = 40,000
∑ 𝐑𝐖 𝟒𝟎,𝟎𝟎𝟎
P01 = = = 400
∑𝐖 𝟏𝟎𝟎
2011 2021
(Base Year) (Current Year)
Items p0q0 p0q1 p1q0 p1q1
Price Quantity Price Quantity
(p0) (q0) (p1) (q1)
A 10 10 20 25 100 250 200 500
B 35 3 40 10 105 350 120 400
C 30 5 20 15 150 450 100 300
D 10 20 8 20 200 200 160 160
E 40 2 40 5 80 200 80 200
∑ 𝐩𝟎 𝐪𝟎 ∑ 𝐩𝟎 𝐪𝟏 ∑ 𝐩𝟏 𝐪𝟎 ∑ 𝐩𝟏 𝐪𝟏
= 635 = 1,450 = 660 = 1,560
𝐋
∑ 𝐩𝟏 𝐪𝟎 𝟔𝟔𝟎
𝐏𝟎𝟏 = × 100 = × 100 = 103.94
∑ 𝐩𝟎 𝐪𝟎 𝟔𝟑𝟓
𝐏
∑ 𝐩𝟏 𝐪𝟏 𝟏,𝟓𝟔𝟎
𝐏𝟎𝟏 = × 100 = × 100 = 107.59
∑ 𝐩𝟎 𝐪𝟏 𝟏,𝟒𝟓𝟎
∑𝐩 𝐪 ∑𝐩 𝐪 𝟔𝟔𝟎 𝟏,𝟓𝟔𝟎
𝐅
𝐏𝟎𝟏 = √[∑ 𝐩𝟏 𝐪𝟎 × 𝟏𝟎𝟎] × [∑ 𝐩𝟏 𝐪𝟏 × 𝟏𝟎𝟎] = √[ ] [ ]
𝟎 𝟎 𝟎 𝟏 𝟔𝟑𝟓 × 𝟏𝟎𝟎 × 𝟏,𝟒𝟓𝟎 × 𝟏𝟎𝟎 = 105
3. Consumer Price Index[CPI] or Cost of Living Index Number
CPI is an index that measures the change in prices paid by a specific class of
consumers for goods and services consumed by them.
∑ 𝐩𝟏 𝐪𝟎
Consumer Price Index (CPI) = × 100
∑ 𝐩𝟎 𝐪𝟎
∑ 𝐑𝐖
Consumer Price Index (CPI) =
∑𝐖
, Where
𝐏𝟏
R= × 100
𝐏𝟎
𝐂𝐏𝐈𝟏 −𝐂𝐏𝐈𝟎
Rate of inflation from time period 0 to 1 = ( ) × 100
𝐂𝐏𝐈𝟎
Example – 5
∑ 𝐩𝟏 𝐪𝟎 𝟑𝟔𝟔
Consumer Price Index (CPI) = × 100 = × 100 = 130.71
∑ 𝐩𝟎 𝐪𝟎 𝟐𝟖𝟎
b Family Budget Method
∑ 𝐑𝐖 𝟑𝟔,𝟔𝟎𝟎
Consumer Price Index (CPI) = = × 100 = 130.71
∑𝐖 𝟐𝟖𝟎
4. Important Index Numbers
C. SENSEX –