Statistics
Statistics
0
Statistics (Advanced Level Combined Mathematics)
Statistics:
The Science of obtaining and analyzing quantitative data with a view to make inferences and decisions, is called
‘Statistics’.
E.g.: Mean
Areas of Statistics:
There are two areas to be considered.
1. Descriptive statistics.
2. Inferential statistics.
Descriptive statistics:
This consists of methods for organizing, displaying and describing data by using tables, graphs and summery
measures.
Inferential statistics:
This consists of methods that use sample results to help make decisions or predictions about a population.
Information:
The manipulated and processed form of data is called information.
Experiment:
An activity to obtain data is called an experiment.
E.g.
a. Drawing lots.
b. Taking a count.
c. Measure a value or a size.
Types of Data:
There are two types of data.
1. Discrete data
2. Continuous data.
1
Discrete data:
A variable whose values are countable or assume only certain values with no intermediate values is called a
discrete data.
E.g.
a. Shoe size.
b. Number of heads.
Continuous data:
A variable that can take ay numerical value over a certain interval is called a continuous data.
E.g.
a. Length.
b. Duration of time.
Methods of Classification:
1. Array.
2. Frequency Distribution.
3. Stem and Leaf Diagram.
Array:
A systematic arrangement of members (Data) usually in rows and columns, enabling to derive other measures is
called an array.
Frequency Distribution:
A tabular presentation of a number of observations arranged according to their magnitudes either individually (in
discrete data) or in a grouped form as classes (in both discrete and continuous data) with their frequencies, is
called a’ frequency distribution’.
2
E.g.
Variable Frequency
x1 f1
x2 f2
x3 f3
. .
. .
Individual data . .
xn fn
Variable Frequency
x1-y1 f1
x2-y2 f2
x3-y3 f3
. .
. .
Grouped data . .
xn-yn fn
Stem-and-Leaf Diagram :
A stem and leaf diagram or plot is a special table where each data value is split into a ‘stem’ (the first digit or
digits) and a ‘leaf’ (usually the last digit)/
E.g. 15, 16, 21, 23, 23, 26, 26, 30, 32, 41
Stem Leaf
1 56
2 13366
3 02
4 1
Usually we order the numbers and present with a key and frequency as follows.
Stem Leaf
1 5 6
2 1 3 3 6 6
3 0 2
4 4 1
3
Tabulation of data and Information
Three tabulation techniques are used.
E.g.
The small groups of data that breaks the range (difference between the maximum value of data and the minimum
value of data) into smaller intervals, are called class intervals. (The desired number of classes is usually between 5
and 20).
Class Frequency
Class Limits
The two terminal values of a class interval are called class limits. The upper terminal is called the upper class
limit and the lower is called the lower class limit.
Class Boundaries
Class limits calculated by taking the average of the upper class limit of a certain class and the lower class limit of
the next class, are called class boundaries.
4
Class Mark
The average of upper class limit (boundary) and lower class limit (boundary) is called class mark or class mid-
point.
Class Size
The difference between the upper class boundary and the lower class boundary is called the class size (width).
Usually this is an odd value. All classes must be in the same size (the exception here is the first or last class).
Cumulative Frequency
A table showing cumulative frequencies of all classes is called a cumulative frequency distribution.
Note:
A graph showing cumulative frequencies less than upper class boundary of each class is known as a less than
cumulative frequency graph or an Ogive.
E.g.
5
Solution:
Bar Chart:
A graph made of bars whose height represents the frequency of respective category, is called a bar chart.
E.g.:
6
Pie Chart:
A circle divided into sectors that represents the relative frequencies or percentages of the categories they
represent.
E.g.:
Histogram:
Histogram is a bar chart without gaps in which the area of the bar is proportional to the frequency of the particular
class.
E.g.:
Line Graph:
Line graphs consists of vertical lines, the height of a line represents the frequency of an ungrouped discrete data.
E.g.:
7
Box Plot:
A box plot that shows three quartiles and whiskers extends from the box to the minimum and maximum values.
The box represents the central 50% of the data.
E.g.:
8
Measures of Central Tendency
A set of data given in terms of numerical measures can be interpreted in terms of numerical a measure which tells us
something about the centre of the set of data. These measures are called measures of central tendency. Three of such
measures are,
1. Mean
2. Median
3. Mode
Mean
The mean x� of a set of data 𝑥1 , 𝑥2 , … , 𝑥𝑛 is defined by
𝑥1 ,+ 𝑥2+,… ,+ 𝑥𝑛
𝑥̅ = 𝑛
or
𝑛
�𝑖=1 𝑥𝑖
𝑥̅ =
𝑛
Let 𝑥1 , 𝑥2 , … , 𝑥𝑛 be a set of data with frequencies 𝑓𝑓1 , 𝑓𝑓2 , … , 𝑓𝑓𝑛 respectively. Mean (Arithmetic Mean) of an
ungrouped data is defined as
or
𝑛
�𝑖=1 𝑓𝑖 𝑥𝑖
𝑥̅ = 𝑛
�𝑖=1 𝑓𝑖
Note:
For grouped data 𝑥𝑖 denotes the mid-point of the ith class interval.
9
Coding method
𝑥𝑖 −𝐴
Here we transform 𝑥𝑖 to 𝑢𝑖 using , 𝑈𝑖 =
𝑐𝑐
𝑛
�𝑖=1 𝑓𝑖 𝑥𝑖
Since, 𝑥̅ = 𝑛
�𝑖=1 𝑓𝑖
𝑛
�𝑖=1 𝑓𝑖 (𝐴+𝑐𝑐𝑢𝑖 )
𝑥̅ = 𝑛
�𝑖=1 𝑓𝑖
𝑛 𝑛
�𝑖=1 𝑓𝑖 𝐴+ �𝑖=1 𝑓𝑖 𝑢𝑖𝑐𝑐
𝑥̅ = 𝑛
�𝑖=1 𝑓𝑖
𝑛
�𝑖=1 𝑓𝑖 𝑢𝑖
= A+c 𝑛
�𝑖=1 𝑓𝑖
= A + c 𝑢�
𝑛
�𝑖=1 𝑓𝑖 𝑑𝑖
𝑥̅ = 𝑛 +A
�𝑖=1 𝑓𝑖
x� = A + 𝑑̅
10
Rule : Sum of the deviations from the mean is zero.
∑𝑛𝑖=1(𝑥𝑖 − 𝑥̅ ) = 0 .
Proof:
∑𝑛𝑖=1(𝑥𝑖 − 𝑥̅ ) = ∑𝑛𝑖=1 𝑥𝑖 - ∑𝑛𝑖=1 𝑥̅
= ∑𝑛𝑖=1 𝑥𝑖 – n x� .
𝑛
�𝑖=1 𝑥𝑖
= ∑𝑛𝑖=1 𝑥𝑖 –n .
𝑛
= ∑𝑛𝑖=1 𝑥𝑖 - ∑𝑛𝑖=1 𝑥𝑖
= 0.
Rule : If mean of 𝑓𝑓1 numbers is 𝑚1 , 𝑓𝑓2 numbers is 𝑚2 , . . . 𝑓𝑓𝑘 numbers is 𝑚𝑘 ,then the mean of 𝑓𝑓1 , + 𝑓𝑓2 +, … , + 𝑓𝑓𝑘
numbers is given by,
𝑓1 𝑚𝑚1 ,+𝑓2 𝑚𝑚2+,… ,+𝑓𝑘 𝑚𝑚𝑘
𝑥̅ = 𝑓1 ,+ 𝑓2 +,… ,+ 𝑓𝑛
Weighted Mean
If the values 𝑥1 , 𝑥2 , … , 𝑥𝑛 are weighted with weights 𝑤1 , 𝑤2 , … , 𝑤𝑛 then the weighted mean of
values 𝑥1 , 𝑥2 , … , 𝑥𝑛 is given by ,
𝑛
�𝑖=1 𝑤𝑖 𝑥𝑖
𝑥
����
𝑤= 𝑛 , where 𝑤𝑖 is the weight of 𝑥𝑖 .
�𝑖=1 𝑤𝑖
Mode
The most frequent value or the value with greatest frequency in a data set is called mode. In some data sets this
may have more than one value.
For a grouped frequency distribution mode is given by,
∆1
Mode = 𝐿𝑚𝑚𝑚𝑚 + c � �
∆1 + ∆2
11
Proof of formula :
∆2
∆1
Mode
L𝑚𝑚𝑚𝑚 c
x c−x
=
∆𝟏 ∆𝟐
∆1
∴ x = c �∆ �
1 + ∆2
∵ Mode = 𝐿𝑚𝑚𝑚𝑚 + x
∆1
∴ Mode = 𝐿𝑚𝑚𝑚𝑚 + c �∆ �
1 + ∆2
Median
Median is the middle value of an ordered set of data.
𝑛+1 𝑡ℎ
Let 𝑥1 , 𝑥2 , … , 𝑥𝑛 be the ordered set of n data. Then median is the � � value of the ordered set.
2
𝑁
� − 𝑓𝑐�𝑐𝑐
2
Median = b + 𝑓
,
12
Proof of formula :
C
M
f
𝑓𝑓𝑐𝑐
N
2
𝑀−𝑏 𝑐𝑐
𝑁 =𝑓
− 𝑓𝑐
2
𝑁
� − 𝑓𝑐�𝑐𝑐
2
Median ( M) = b + 𝑓
Quartiles
𝑛+1 𝑡ℎ
𝑄1 is the � � value of the data arranged in the ascending order.
4
𝑛+1 𝑡ℎ
𝑄2 is the � � value of the data arranged in the ascending order.
2
13
For Grouped data
𝑁
� − 𝑓𝑐 �𝑐𝑐
4
𝑄1= b + 𝑓
3𝑁
� − 𝑓𝑐 �𝑐𝑐
4
𝑄3 = b + 𝑓
Percentiles
𝑝𝑛 𝑡ℎ
pth percentile of the data set arranged in the ascending order is given by � � value of the data set
100
arranged in the ascending order .
14
Measures of Dispersion
Dispersion indicates the spread of data, about a central measure. This is used to represent the spread within
data.
I.Q.R. = 𝑄3 - 𝑄1
4. Mean Deviation
∑𝑛
𝑖=1|𝑥𝑖 − 𝑥̅ |
Mean Deviation =
𝑛
∑𝑛
𝑖=1 𝑓𝑖 |𝑥𝑖 − 𝑥̅ |
Mean Deviation = ∑𝑛
( For a grouped frequency distribution 𝑥𝑖 is the mid value of the 𝑖𝑡ℎ class.)
𝑖=1 𝑓𝑖
5. Variation (𝝈𝟐 ) :
For a set of data 𝑥1 , 𝑥2 , … , 𝑥𝑛 ,
𝑛
�𝑖=1(𝑥𝑖 − 𝑥̅ )2
Variance =
𝑛
Note :
𝑛
�𝑖=1 𝑥𝑖 2
Variance = - 𝑥̅ 2
𝑛
15
Proof :
1 𝑛
Var(𝑥) = �𝑖=1(𝑥𝑖 − 𝑥̅ )2
𝑛
𝑛 𝑛
�𝑖=1(𝑥𝑖 − 𝑥̅ )2 = �𝑖=1(𝑥𝑖 2 − 2𝑥𝑖 𝑥� + 𝑥̅ 2 )
𝑛
= �𝑖=1 𝑥𝑖 2 - 2𝑥̅ ∑𝑛𝑖=1 𝑥𝑖 + n𝑥̅ 2
𝑛
= �𝑖=1 𝑥𝑖 2 - 2n𝑥̅ 2 + n𝑥̅ 2
𝑛
= �𝑖=1 𝑥𝑖 2 - n𝑥̅ 2
1 𝑛
∴ Var(𝑥) = ��𝑖=1 𝑥𝑖 2 − n𝑥̅ 2 �
𝑛
𝑛
�𝑖=1 𝑥𝑖 2
= - 𝑥̅ 2
𝑛
𝑛 𝑛
�𝑖=1 𝑓𝑓𝑖 (𝑥𝑖 − 𝑥̅ )2 = �𝑖=1 𝑓𝑓𝑖 (𝑥𝑖 2 − 2𝑥𝑖 𝑥� + 𝑥̅ 2 )
𝑛
= �𝑖=1(𝑓𝑓𝑖 𝑥𝑖 2 − 2𝑓𝑓𝑖 𝑥𝑖 𝑥� + 𝑓𝑓𝑖 𝑥̅ 2 )
𝑛
=� 𝑓𝑓𝑖 𝑥𝑖 2 - 2𝑥̅ ∑𝑛𝑖=1 𝑓𝑓𝑖 𝑥𝑖 + 𝑥̅ 2 ∑𝑛𝑖=1 𝑓𝑓𝑖
𝑖=1
𝑛
= � 𝑓𝑓𝑥𝑖 2 - 2n𝑥̅ 2 + n𝑥̅ 2
𝑖=1
𝑛
= � 𝑓𝑓𝑥𝑖 2 - n𝑥̅ 2
𝑖=1
𝑛
1
∴ Var(𝑥) = ∑𝑛 �� 𝑓𝑓𝑖 𝑥𝑖 2 − n𝑥̅ 2 �
𝑖=1 𝑓𝑖 𝑖=1
𝑛
2
� 𝑓𝑖 𝑥𝑖
= 𝑖=1
𝑛
∑𝑖=1 𝑓𝑖
- 𝑥̅ 2
16
6). Standard deviation (𝝈)
i) . 𝑦� = a𝑥̅ + b
Proof (i) :
𝑦𝑖 = a𝑥𝑖 + b
∑𝑛
𝑖=1 𝑦𝑖
𝑦� =
𝑛
∑𝑛𝑖=1 𝑦𝑖 = ∑𝑛𝑖=1(a𝑥𝑖 + b)
= a∑𝑛𝑖=1 𝑥𝑖 + nb
∑𝑛
𝑖=1 𝑦𝑖 ∑𝑛
𝑖=1 𝑥𝑖
÷ n, =a +b
𝑛 𝑛
∴ � = a𝒙
𝒚 �+b
Proof (ii) :
𝑦𝑖 = a𝑥𝑖 + b
𝑛
�𝑖=1 𝑦𝑖 2
𝜎𝜎𝑦 2 = - 𝑦� 2
𝑛
𝑛 𝑛
�𝑖=1 𝑦𝑖 2 = �𝑖=1(a𝑥𝑖 + b)2
𝑛
= �𝑖=1(𝑎2 𝑥𝑖 2 + 2𝑎𝑏𝑥𝑖 + 𝑏 2 )
𝑛
= 𝑎2 �𝑖=1 𝑥𝑖 2 + 2ab ∑𝑛𝑖=1 𝑥𝑖 + n 𝑏 2
𝑛
�𝑖=1 𝑦𝑖 2 𝑎2 𝑛
∴ ÷ n, = �𝑖=1 𝑥𝑖 2 + 2ab x� + 𝑏 2 ……..………. ( 1 )
𝑛 𝑛
17
∑𝑛
𝑖=1 𝑦𝑖
𝑦� =
𝑛
∑𝑛
𝑖=1(a𝑥𝑖 + b)
=
𝑛
1
= ∑𝑛𝑖=1(a𝑥𝑖 + b)
𝑛
𝑎 1
= ∑𝑛𝑖=1 𝑥𝑖 + x nb
𝑛 𝑛
𝑎
= ∑𝑛𝑖=1 𝑥𝑖 + b
𝑛
𝑎 2
𝑦� 2 = � ∑𝑛𝑖=1 𝑥𝑖 + b�
𝑛
𝑎2 2𝑎𝑏
= (∑𝑛𝑖=1 𝑥𝑖 )2 + ∑𝑛𝑖=1 𝑥𝑖 + 𝑏 2
𝑛2 𝑛
2
∑𝑛
𝑖=1 𝑥𝑖 ∑𝑛
𝑖=1 𝑥𝑖
= 𝑎2 � � + 2ab + 𝑏2
𝑛 𝑛
= 𝑎2 𝑥̅ 2 + 2ab 𝑥̅ + 𝑏 2 …………………( 2 )
From ( 1 ) ad ( 2 ),
𝑛
�𝑖=1 𝑦𝑖 2
𝜎𝜎𝑦 2 = - 𝑦� 2
𝑛
𝑎2
= (∑𝑛𝑖=1 𝑥𝑖 )2 + 2𝑎𝑏 x� + 𝑏 2 - 𝑎2 𝑥̅ 2 - 2ab 𝑥̅ - 𝑏 2
𝑛
𝑎2 𝑛
= �𝑖=1 𝑥𝑖 2 - 𝑎2 𝑥̅ 2
𝑛2
𝑛
�𝑖=1 𝑥𝑖 2
= 𝑎2 � − 𝑥̅ 2 �
𝑛
= 𝑎2 𝜎𝜎𝑥 2
∴ 𝝈𝒚 = |𝒂| 𝝈𝒙
7). Z – Score:
Let 𝑥̅ be the mean and σ𝑥 be the standard deviation for a set of data 𝑥1 , 𝑥2 , … , 𝑥𝑛 .
𝑥𝑖 − 𝑥̅
For each 𝑥𝑖 , 𝑧𝑖 is defined as , 𝑧𝑖 = , where 𝑧𝑖 is called Z – Score of 𝑥𝑖 .
𝜎𝑥
18
Proof i) :
𝑥𝑖 − 𝑥̅
𝑧𝑖 =
𝜎𝑥
∑𝑛
𝑖=1 𝑧𝑖
𝑧̅ =
𝑛
𝑛 �
𝑥 −𝑥
� � 𝑖 �
𝑖=1 𝜎𝑥
=
𝑛
∑𝑛
𝑖=1 𝑥𝑖 ∑𝑛
𝑖=1 𝑥̅
= -
𝑛𝜎𝑥 𝑛𝜎𝑥
𝑥̅ 𝑥̅
= -
𝜎𝑥 𝜎𝑥
= 0.
Proof ii) :
𝑛
�𝑖=1 𝑧𝑖 2
𝜎𝜎𝑧 2 = - 𝑧̅ 2
𝑛
𝑛
�
𝑥 −𝑥 2
� � 𝑖 �
𝜎𝑥
= 𝑖=1
- 𝑧̅ 2
𝑛
𝑛
�𝑖=1�𝑥𝑖 2 −2𝑥𝑖𝑥̅ +𝑥̅ 2 �
= - 𝑧̅ 2
𝑛𝜎2 𝑥
𝑛 𝑛
� 𝑥𝑖 2 � 𝑥𝑖
𝑖=1
𝑛
− 2x� 𝑖=1
𝑛
+ 𝑥̅ 2
= –0
𝜎2 𝑥
𝑛
� 𝑥𝑖 2
𝑖=1
𝑛
− 𝑥̅ 2
=
𝜎𝑥 2
𝜎𝑥 2
=
𝜎𝑥 2
=1
Let x� 1 𝑎𝑛𝑑 x� 2 be the means of two sets of data with sizes 𝑛1 and 𝑛2 respectively. Then the pooled mean 𝑥̅ is given by
𝑛1 𝑥̅ 1 +𝑛2 𝑥̅2
𝑥̅ = .
𝑛1 + 𝑛2
Proof:
Let the elements of a set of data with size 𝑛1 be 𝑥11 , 𝑥12 , … , 𝑥1𝑛1 and the elements of a set of data with size 𝑛2 be
∑𝑛
𝑖=1 𝑥1𝑖 ∑𝑛
𝑖=1 𝑥2𝑖
∴ ���
𝑥1 = and ���
𝑥2 =
𝑛1 𝑛2
19
𝑛 𝑛
∑𝑖=1
1
𝑥1𝑖 + 𝑛 ∑𝑖=1 𝑥2𝑖
2
𝑥̅ = 𝑛1 + 𝑛2
𝑛1 𝑥̅ 1 +𝑛2 𝑥̅ 2
𝑥̅ = 𝑛1 + 𝑛2
Let 𝜎𝜎1 2 and 𝜎𝜎2 2 be the variances of sets of data with sizes 𝑛1 , 𝑛2 respectively. Then,
1 𝑛1 𝑛2
𝜎𝜎 2 = {𝑛1 𝜎𝜎1 2 + 𝑛2 𝜎𝜎2 2 } +
(𝑛 2
(𝑥̅1 − 𝑥̅2 )2
𝑛1 + 𝑛2 1 + 𝑛2 )
Proof:
Where,
𝑛 𝑛
�𝑖=1(𝑥1𝑖 − 𝑥̅ )2 = �𝑖=1{(𝑥1𝑖 − 𝑥̅1 ) + (𝑥̅1 − 𝑥̅ )}2
𝑛 𝑛
= �𝑖=1(𝑥1𝑖 − 𝑥̅1 )2 + 2(𝑥̅1 − 𝑥̅ ) ∑𝑛𝑖=1(𝑥1𝑖 − 𝑥̅1 ) + �𝑖=1(𝑥̅1 − 𝑥̅ )2
= 𝑛1 𝜎𝜎1 2 + 0 + 𝑛1 (𝑥̅1 − 𝑥̅ )2
= 𝑛1 𝜎𝜎1 2 + 𝑛1 (𝑥̅1 − 𝑥̅ )2
Similarly,
𝑛
�𝑖=1(𝑥2𝑖 − 𝑥̅ )2 = 𝑛2 𝜎𝜎2 2 + 𝑛2 (𝑥̅2 − 𝑥̅ )2
𝑛1 𝜎1 2 + 𝑛1 (𝑥̅ 1 − 𝑥̅ )2 + 𝑛2 𝜎2 2 + 𝑛2 (𝑥̅ 2 − 𝑥̅ )2
∴ 𝜎𝜎 2 =
𝑛1 + 𝑛2
1
= (𝑛 {𝑛1 𝜎𝜎1 2 + 𝑛2 𝜎𝜎2 2 + 𝑛1 (𝑥̅1 − 𝑥̅ )2 + 𝑛2 (𝑥̅2 − 𝑥̅ )2 }
1 𝑛2 )
+
𝑛1 𝑥̅1 +𝑛2 𝑥̅ 2
but, 𝑥̅1 − 𝑥̅ = 𝑥̅1 - � �
𝑛1 + 𝑛2
𝑛2 𝑥̅1 − 𝑛2 𝑥̅ 12
=
𝑛1 + 𝑛2
𝑛2 (𝑥̅ 1 − 𝑥̅ 2 )
=
𝑛1 + 𝑛2
20
Similarly,
𝑛1 (𝑥̅2 − 𝑥̅ 1 )
𝑥̅2 − 𝑥̅ =
𝑛1 + 𝑛2
1 𝑛2 (𝑥̅1 − 𝑥̅ 2 ) 2 𝑛1 (𝑥̅2 − 𝑥̅ 1 ) 2
∴ 𝜎𝜎 2 = = (𝑛 �𝑛1 𝜎𝜎1 2 + 𝑛2 𝜎𝜎2 2 + 𝑛1 � � + 𝑛2 � � �
1 + 𝑛2 ) 𝑛1 + 𝑛2 𝑛1 + 𝑛2
1 𝑛1 𝑛2
𝜎𝜎 2 = = (𝑛 {𝑛1 𝜎𝜎1 2 + 𝑛2 𝜎𝜎2 2 } + 3
(𝑥̅1 − 𝑥̅2 )2 (𝑛1 + 𝑛2 )
1 + 𝑛2 ) (𝑛 1 + 𝑛2 )
𝟏 𝒏𝟏 𝒏𝟐
𝝈𝟐 = �𝒏𝟏 𝝈𝟏 𝟐 + 𝒏𝟐 𝝈𝟐 𝟐 � + (𝒏 𝟐
(𝒙 �𝟐 )𝟐
�𝟏 − 𝒙
𝒏𝟏 + 𝒏𝟐 𝟏 + 𝒏𝟐 )
Skewness
This determines the shape of the distribution or this tells us something about the symmetry of the frequency
distribution. The amount of asymmetry is measured using skewness. Basically there are three shapes.
Symmetric Distribution
21
Mean – Mode = 0
Negatively skewed distribution
Measuring Skewness
Mean – Mode , measures skewness but the unit of this measure depends o the distribution concerned. So that this
measure cannot be used in comparing two different distributions. Therefore, for the comparison of distributions, there are
two other measures known as Pearson’s coefficients of skewness, defined such that,
𝑀𝑒𝑎𝑛−𝑀𝑚𝑚𝑑𝑒
i). k = 𝑆𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑚𝑚𝑛
𝑥̅ − 𝑀𝑜
k= 𝜎
3(𝑀𝑒𝑎𝑛−𝑀𝑒𝑑𝑖𝑎𝑛)
ii). k = 𝑆𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑚𝑚𝑛
3(𝑥̅ − 𝑀)
k= 𝜎
Note :
This is an empirical formula derived using observed results. We cannot prove this mathematically.
22
Recent Past Papers
01. The mean and the standard deviation of a set of observations {x1,x2,…, xn} are 𝑥̅ and Sx respectively. Suppose that a linear
transformation yi = a + bxi, where a and b are constants, transforms the original data set {x1,x2,…, xn} to the set {y1,y2,…, yn}.
Show that 𝑦� = a + b𝑥̅ and 𝑆𝑦2 = 𝑏2 𝑆𝑥2 , where 𝑦� and 𝑆𝑦 are the mean and standard deviation of the set {y1,y2,…, yn}.
(i) Find the mean and the standard deviation of the set of observations { 1, 2, 3, 4, 5, 6, 7 }.
Hence, find
(a) The mean and the standard deviation of the set of observations { 2.01, 3.02, 4.03, 5.04, 6.05, 7.06, 8.07}
(b) Seven values whose mean is 5 and the standard deviation is 6.
(ii) Salt is packed in the bags which the manufacturer claims contain 25kg each. The following information is given for 80 such
bags whose actual weights are not known:
∑80 80 2
𝑖=1(𝑥𝑖 − 25) = 27.2 and ∑𝑖=1(𝑥𝑖 − 25) = 85.1
where xi (I = 1, 2, …, 80) denotes the actual weight of the ith bag, Using an appropriate linear transformation or otherwise, find
the mean and the variance of the actual weights of the eighty of the eighty bags. (2011)
02. Students in a certain class were given a question paper in Statistics. The marks obtained by these students are given in the
following grouped frequency table:
03. A frequency distribution of diameters of a set containing of 50 small metal balls is given in the following table:
Find the mean and variance of the diameters of the combined set of 150 metal balls.
It is subsequently discovered that the measuring instrument used for the second set of 100 metal balls was faulty and the diameter
of each ball has been underestimated by 0.015 cm. Find the true mean and true standard deviation of the diameters of these 100
metal balls.(2013 – AL)
23
04 Let the mean and the variance of the set of data {𝑥1 , 𝑥2 , … , 𝑥𝑛 } be 𝑥̅ and 𝜎𝜎𝑥2 respectively.
1 𝑛
(i). Show that 𝜎𝜎𝑥2 = �𝑖=1 𝑥𝑖 2 − 𝑥̅ 2
𝑛
𝑛
(ii). Let α and β be real constants. Show that �𝑖=1(a𝑥𝑖 + b)2 = n𝛼 2 𝜎𝜎𝑥2 + n(𝛼𝑥̅ + 𝛽)2
Let 𝑦𝑖 = α𝑥𝑖 + β for I = 1, 2, … , n. Show that 𝑦� = 𝛼𝑥̅ + 𝛽, and using (i) and (ii) above, deduce that 𝜎𝜎𝑦2 = 𝛼 2 σ2𝑥 ,
where 𝑦� and 𝜎𝜎𝑦2 are the mean and the variance of the set �𝑦1 , 𝑦2 , … , 𝑦𝑛 � respectively.
The mean of the marks obtained by candidates in a certain examination is 45. These marks are to be scaled linearly to give a mean
of 50 and a standard deviation 15. It is given that the scaled mark 68 corresponds to the original mark 60. Calculate the standard
deviation of the original marks.
It is given further that the original mark m obtained by a candidate is not lowered by the above scaling. Show that m ≥ 20. (2014)
05. A group of 100 technical college students measured the length of a certain portion of a main road, and their measurements are
given in the following frequency table.
𝑥− 𝑥̅ 𝑎
By means of transformation y = , for an assumed mean 𝑥̅𝑎 = 100.1 and d =0.1 , extend the above table to include
𝑑
the corresponding values of y and 𝑦 2 . Find the mean of y ad hence show that the mean of x is 100.123.
Taking that √1.917 ≈ 1.385 , calculate, approximately the standard deviation of the frequency distribution, correct to three
decimal places. (2015)
06. The mean and standard deviation of n umbers 𝑥1 , 𝑥2 , … , 𝑥𝑛 are 𝜇𝜇1 and 𝜎𝜎1 respectively, and the mean and standard
deviation of n umbers 𝑦1 , 𝑦2 , … , 𝑦𝑛 are 𝜇𝜇2 and 𝜎𝜎2 respectively. Let mean and standard deviation of these n + m numbers be
𝜇𝜇3 and 𝜎𝜎3 respectively.
𝑛𝜇1 +𝑚𝑚𝜇2
Show that 𝜇𝜇3 = .
𝑛+𝑚𝑚
𝑛
Let 𝑑1 = 𝜇𝜇3 - 𝜇𝜇1 . Show that �𝑖=1(𝑥𝑖 − 𝑥̅ )2 = n (𝜎𝜎12 + 𝑑12 ).
𝑛
By taking 𝑑2 = 𝜇𝜇3 - 𝜇𝜇2 , write down a similar expression for �𝑖=1(𝑦𝑖 − 𝑦�)2 .
2 2
�𝑛𝜎𝜎1 2 +m𝜎𝜎2 2�+�𝑛𝑑1 +m𝑑2 �
Deduce that 𝜎𝜎32 = .
𝑛 + 𝑚𝑚
The number of copies sold per day, during the first 100 days after publishing a new book, had mean 2.3 and variance 0.8. During
the next 100 days, the number of copies sold per day had mean 1.7 and variance 0.5. Fid the mean and the variance of the number
of copies sold per day during the first 200 days. ( 2016 )
07. The following table gives the distribution of marks obtained by a group of 10 students for their answers to a Statistics question.
25