Module 5
Module 5
STATISTICAL METHODS
AND PROBABILITY
ISC XI- Applied Mathematics
5
(i) Measures of Dispersion
SCOPE
Measures of dispersion: range, Quartile Deviation, mean deviation, variance and standard deviation of
ungrouped/grouped data.
• Range
• Quartiles, Interquartile range, Quartile deviation, Coefficient of Quartile deviation
• Mean deviation about mean and median, coefficient of mean deviation
• Mean deviation about median.
• Standard deviation - by direct method, short cut method and step deviation method, Variance and
coefficient of variance
• Combined mean and standard deviation
• Mode of grouped and ungrouped data.
• Differentiate between range, quartile deviation, mean deviation and standard deviation.
• Choose appropriate measure of dispersion to calculate spread of data.
LEARNING OUTCOMES
PRE-REQUISITES
1
ISC XI- Applied Mathematics
Introduction
Statistics is about collecting, presenting, analyzing and interpreting, and data such that we can draw
meaningful conclusions from datasets.
The study of statistics can be categorized into two main types:
• Descriptive Statistics – Describing the main features of the dataset.
• Inferential Statistics – Making inferences from a data.
Descriptive Statistics
Descriptive statistics involves methods for summarizing and describing the main features of a dataset. These
methods include calculating measures of central tendency (mean, median, mode) and measures of variability
or dispersion (standard deviation, variance, range). Descriptive statistics are used to present data in a clear and
understandable way.
Dispersion
Suppose there are three series of 5 items each as follows:
Series A 40 40 40 40 40
Series B 38 39 40 41 42
Series C 6 21 40 57 76
2
ISC XI- Applied Mathematics
In all the three series, the mean is 40. But in Series A all the items are identical, in series B the items are not
identical but are not scattered a lot. In contrast, in Series C the non-identical items are widely scattered.
The term “dispersion” refers to how scattered or dispersed a set of data is. The measure of dispersion is always
a non-negative real number that starts at zero when all the items in the data are the same and rise as the data
get more varied. These measures help determine the homogeneity or heterogeneity of the data distribution,
providing insight into how much the values differ from each other.
Measures of Dispersion
Measures of dispersion describe how spread out or varied the data in a dataset is.
3
ISC XI- Applied Mathematics
Range
Range is the simplest possible measure of dispersion or the extent of scatter. It is mathematically expressed as
‘The difference between the largest and the smallest values in a data set’.
When we buy things, they are always sold within a price range. Take the example of your favourite pair
of jeans.
The store from where you made the purchase probably had a range of colours, a range of fits, a range of
sizes, and a range of prices.
A wide range indicates high variability meaning the data is spread out, and a small range specifies low
variability in the distribution.
4
ISC XI- Applied Mathematics
Example 1.
(i) The profits of a company for the last 6 years are given below. Calculate the range.
Solution:
Range = largest value – smallest value = 120 – 30 = 90
Classes 115 - 125 125 - 135 135 - 145 145 - 155 155 - 165 165 - 175
Frequency 1 4 6 1 3 5
Solution:
Range = upper limit of highest class – lower limit of the lowest class = 175 – 115 = 60
Limitations of Range
• Range is greatly affected by fluctuations of sampling.
• It is not based on all the observations of the series.
• Only the values of the variables are taken into account and the frequencies are completely ignored.
5
ISC XI- Applied Mathematics
Quartiles
Just as the median is a positional average that divides a data set into two equal parts, there are other positional
values that split a series into multiple sections. One of the most commonly used among them is quartiles.
Quartiles divide a data set into four equal parts, helping to analyze the spread and distribution of the data.
The three quartiles are:
1. First Quartile (Q₁) or Lower Quartile – This is the value that separates the lowest one-fourth (25%) of
the data from the rest.
2. Second Quartile (Q₂) or Median – This is the middle value, dividing the data set into two equal halves.
3. Third Quartile (Q₃) or Upper Quartile – This is the value that separates the lowest three-fourths (75%)
of the data from the highest one-fourth.
𝑁𝑁+1
For ungrouped data: 𝑄𝑄𝑖𝑖 = the value of the 𝑖𝑖 � � th item
4
For grouped data: Identify the 𝑄𝑄𝑖𝑖 class that has the
𝑁𝑁
cumulative frequency as 𝑖𝑖 � �. Name it as 𝑄𝑄𝑖𝑖 class
4
Then,
𝑁𝑁
𝑁𝑁 𝑖𝑖� �−𝑐𝑐𝑐𝑐
𝑄𝑄𝑖𝑖 = the value of the 𝑖𝑖 � � th item = 𝐿𝐿 + 4
× 𝐼𝐼
4 𝑓𝑓
where i = 1,2,3.
L = lower boundary of the 𝑄𝑄𝑖𝑖 class
cf = cumulative frequency of the class preceding the 𝑄𝑄𝑖𝑖
class
f = frequency of the 𝑄𝑄𝑖𝑖 class
I = size of class (assuming all classes have equal size)
Example 2. Calculate the lower and upper quartiles from the following data.
No of babies 2 4 5 7 3 2 1
Solutions:
N = 24
Q1 = �
N+1 th
� item = �
24+1 th
� item x f c.f. less than
4 4
1th 4 2 2
=6 item
4
5 4 6
1
= 6th item +
4
�7𝑡𝑡ℎ − 6𝑡𝑡ℎ � item
6 5 11
1
= 5 + (6 − 5) = 5 ⋅ 25 𝑙𝑙𝑙𝑙𝑙𝑙 8 7 18
4
Q3 = 3 �
N+1 th
� item = 3 �
24+1 th
� item 11 3 21
4 4
3th 13 2 23
= 18 item
4
3
14 1 24
= 18th item +
4
�19𝑡𝑡ℎ − 18𝑡𝑡ℎ � item
6
ISC XI- Applied Mathematics
3
= 8 + (11 − 8) = 10 ⋅ 25 𝑙𝑙𝑙𝑙𝑙𝑙
4
Example 3. Calculate the lower and upper quartiles from the following data.
CI 10 - 19 20 - 29 30 - 39 40 - 49 50 - 59 60 - 69 70 - 79 80 - 89
Frequency 5 9 14 20 25 15 8 4
Solution:
Cumulative
Class interval Exclusive CI Frequency (f) frequency
(c.f less than)
10 – 19 9.5 – 19.5 5 5
20 – 29 19.5 – 29.5 9 14
30 – 39 29.5 – 39.5 14 28
40 – 49 39.5 – 49.5 20 48
50 – 59 49.5 – 59.5 25 73
60 - 69 59.5 – 69.5 15 88
70 - 79 69.5 – 79.5 8 96
N = 100
𝑁𝑁 100 th
𝑸𝑸𝟏𝟏 =� �th item =� � item = 25th item
4 4
𝑁𝑁 100 th
𝑸𝑸𝟐𝟐 =2 � �th item =2 � � item = 50th item
4 4
𝑁𝑁 100 th
𝑸𝑸𝟑𝟑 =3 � �th item =3 � � item =75th item
4 4
7
ISC XI- Applied Mathematics
Quartile Deviation
Quartile deviation is also called the Semi-inter quartile range and is calculated as
𝑸𝑸𝟑𝟑 −𝑸𝑸𝟏𝟏
Quartile deviation = 𝟐𝟐
The coefficient of quartile deviation, also known as the quartile coefficient of dispersion, is a statistical
measure that quantifies the spread or dispersion of a dataset, specifically focusing on the central 50% of the
data.
𝑸𝑸 −𝑸𝑸
Coefficient of Quartile Deviation = 𝑸𝑸𝟑𝟑+𝑸𝑸𝟏𝟏
𝟑𝟑 𝟏𝟏
It's a relative measure, making it suitable for comparing data sets with different scales or units. Since it relies
on quartiles (Q1 and Q3) rather than the mean and standard deviation, it's less sensitive to the influence of
extreme values (outliers) in the data.
Example 4. Find the quartile deviation for the following given data.
23, 8, 5, 16, 33, 7, 24, 5, 30, 33, 37, 30, 9, 11, 26, 32
Solution:
Let us arrange this data in the following ascending order.
5, 5, 7, 8, 9, 11, 16, 23, 24, 26, 30, 30, 32, 33, 33, 37
N =16
N+1 th 16+1 th 1th 1
Q1 = �
4
� item = �
4
� item = 4
4
item =4th item +
4
�5𝑡𝑡ℎ − 4𝑡𝑡ℎ � item
1
= 8 + (9 − 8) = 8 ⋅ 25
4
N+1 th 16+1 th 3th 3
Q3 = 3 �
4
� item = 3 �
4
� item = 12
4
item =12th item +
4
�13𝑡𝑡ℎ − 12𝑡𝑡ℎ � item
3
= 30 + (32 − 30) = 31.5
4
𝑄𝑄3 −𝑄𝑄1 31⋅5−8.25
Quartile Deviation = 2
= = 11.625
2
Example 5. Calculate Quartile deviation from the following and hence calculate the coefficient of quartile
deviation.
Class 0-2 2-4 4-6 6-8 8 - 10 10 - 12
Frequency 5 16 13 7 5 4
8
ISC XI- Applied Mathematics
Solution:
Here, N =50 Class Frequency c.f. (less than)
𝑁𝑁 50
𝑄𝑄1 =� 4 �th item =� 4 �th item =12.5th item 0-2 5 5
2-4 16 21
∴ 𝑄𝑄1 lies in the class = 2 - 4
𝑁𝑁
4-6 13 34
� �−𝑐𝑐𝑐𝑐 12.5−5
𝑄𝑄1 = 𝐿𝐿 + 4
× 𝐼𝐼 = 2 + × 2 = 2.93 6-8 7 41
𝑓𝑓 16
𝑁𝑁 50
𝑄𝑄3 =3 � �th item =3 � �th item =37.5th item 8 - 10 5 46
4 4
10 - 12 4 50
∴ 𝑄𝑄3 lies in the class = 6 - 8
𝑛𝑛
3 � �−𝑐𝑐𝑐𝑐 37.5−34
𝑄𝑄3 = 𝐿𝐿 + 4
× 𝐼𝐼 = 6 + ×2=7
𝑓𝑓 7
Scenario: An HR team is determining salary ranges for new hires with varying experience and
qualifications.
Application: They use quartile deviation to understand the spread of salaries within different experience
levels. A high quartile deviation might indicate a wide range of salaries for a particular role, while a low
deviation suggests a more uniform salary structure.
Interpretation:
High QD: Suggests a significant difference between the salaries of the most and least experienced
employees within a certain range, potentially indicating a wide range of skills and responsibilities.
Low QD: Indicates that salaries are more closely clustered around the median, suggesting a more
standardized compensation structure.
9
ISC XI- Applied Mathematics
−10+5+20
Mean = =5
3
Now a deviation from the mean for the different values is,
• (-10 -5) = -15
• (5 – 5) = 0
• (20 – 5) = 15
Now adding the deviations, we see that there is zero deviation from the mean which is incorrect. Thus, to
counter this problem only the absolute values of the difference are taken while calculating the mean
deviation.
10
ISC XI- Applied Mathematics
The Mean Deviation (MD) about the mean (𝑥𝑥̅ ) is the mean of the absolute values of the deviations of the
items from the mean, and is given by
𝜮𝜮𝒇𝒇𝒊𝒊 |𝒙𝒙𝒊𝒊 − 𝒙𝒙
�|
𝑴𝑴𝑴𝑴 =
𝜮𝜮𝒇𝒇𝒊𝒊
The coefficient of mean deviation from the mean is a measure of dispersion that indicates the variability of
a dataset around its average. It’s calculated by dividing the mean deviation (average of absolute differences
from the mean) by the mean itself.
𝑴𝑴𝑴𝑴𝑴𝑴𝑴𝑴 𝒅𝒅𝒅𝒅𝒅𝒅𝒅𝒅𝒅𝒅𝒅𝒅𝒅𝒅𝒅𝒅𝒅𝒅
Coefficient of Mean Deviation from mean =
�
𝒙𝒙
This ratio helps compare the variability of different datasets, regardless of the scale of the data.
Example 7. Find the mean deviation about mean for the following and hence calculate the coefficient of
mean deviation.
Marks 5 15 25 35 45
No of students 5 8 15 16 6
Solution:
1350
1. Calculate the mean: 𝑥𝑥̅ = = 27
50
472
2. Compute the Mean Deviation about mean: 𝑀𝑀𝑀𝑀 = 50
= 9.44
9.44
3. Coefficient of Mean Deviation = = 0.3496
27
11
ISC XI- Applied Mathematics
No of students 6 5 8 15 7 6 3
Solution:
1670
1. Calculate the mean: 𝑥𝑥̅ = = 33.4
50
659.2
2. Compute the Mean Deviation about mean: 𝑀𝑀𝑀𝑀 = 50
= 13.18
13.18
3. Coefficient of Mean Deviation = = 0.3496
33.4
Example 9. Find the mean deviation about mean using step deviation method for the following:
Solution:
Height (in cm) No of students 𝒙𝒙𝒊𝒊 𝐱𝐱𝐢𝐢 − 𝟏𝟏𝟏𝟏𝟏𝟏 𝒇𝒇𝒊𝒊 𝒖𝒖𝒊𝒊 |𝒙𝒙𝒊𝒊 − 𝒙𝒙
�| 𝒇𝒇𝒊𝒊 |𝒙𝒙𝒊𝒊 − 𝒙𝒙
�|
𝐮𝐮𝐢𝐢 =
(𝒇𝒇𝒊𝒊 ) 𝟏𝟏𝟏𝟏
95-105 9 100 -2 -18 25.3 227.7
105-115 13 110 -1 -13 15.3 198.9
115-125 26 120 0 0 5.3 137.8
125-135 30 130 1 30 4.7 141
135-145 12 140 2 24 14.7 176.4
145-155 10 150 3 30 24.7 247
100 53 1128.8
12
ISC XI- Applied Mathematics
Median
Median is the middle item of a distribution, when items are arranged in order of magnitude. Unlike mean, the
median does not take into account the values of all items in a series.
𝑛𝑛+1 th
For ungrouped data: Median = the value of the � � item
2
(ii) We first arrange in ascending order 32, 33, 48, 68, 70, 80, 86, 94
𝑛𝑛+1 th 8+1 th
Median = � � item = � � item = 4.5th item
2 2
1 1
= 2 [4𝑡𝑡ℎ 𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡 + 5𝑡𝑡ℎ 𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡] = 2 [68 + 70]= 69.
No of persons 14 26 40 53 50 37 25
13
ISC XI- Applied Mathematics
Solution:
36 – 40 35.5 – 40.5 14 14
41 – 45 40.5 – 45.5 26 40
46 – 50 45.5 – 50.5 40 80
51 – 55 50.5– 55.5 53 133
56 – 60 55.5– 60.5 50 183
61 – 65 60.5– 65.5 37 220
66 – 70 65.5– 70.5 25 245
245
N = 245
𝑁𝑁 245 th
Median = � �th item =� � item =122.5th item
2 2
Frequency 5 a 20 15 b 5 60
Solution:
14
ISC XI- Applied Mathematics
𝑁𝑁
−𝑐𝑐𝑐𝑐
∴ Median = 𝐿𝐿 + 2
× 𝐼𝐼
𝑓𝑓
30 –(5+𝑎𝑎)
⇒ 28.5 = 20 + × 10
20
30 –(5+𝑎𝑎)
⇒ 28.5-20 = 20
× 10
25−𝑎𝑎
⇒ 8.5 = 2
⇒ 17 = 25 – 𝑎𝑎
⇒ 𝑎𝑎 = 8.
Given N = 60
⇒ 45 + 𝑎𝑎 + 𝑏𝑏 =60
⇒ 45 + 8 + b = 60
⇒ 53 + b = 60
⇒ b =7
�|
𝛴𝛴𝑓𝑓𝑖𝑖 |𝑥𝑥𝑖𝑖 −𝑀𝑀
𝑀𝑀𝑀𝑀 = , where n is the number of observations.
𝑛𝑛
The coefficient of mean deviation from the median is a statistical measure that indicates the average degree
of dispersion or variability in a dataset, relative to the median. It’s calculated by dividing the mean deviation
from the median by the median itself.
𝑀𝑀𝑀𝑀𝑀𝑀𝑀𝑀𝑀𝑀𝑀𝑀 𝑑𝑑𝑑𝑑𝑑𝑑𝑑𝑑𝑑𝑑𝑑𝑑𝑑𝑑𝑑𝑑𝑑𝑑
Coefficient of Mean Deviation from median = �
𝑀𝑀
This helps in comparing the variability of different datasets, regardless of their central tendency.
Marks 5 7 9 10 12 15
No of students 8 6 2 2 2 6
15
ISC XI- Applied Mathematics
Solution:
1. N = 26
𝑁𝑁+1 26+1
2. Calculate the median: Median = � � th item = � � th item = 13.5 item
2 2
1 1
= [13𝑡𝑡ℎ 𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡 + 14𝑡𝑡ℎ 𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡] = [7 + 7]= 7.
2 2
84
3. Compute the Mean Deviation about Median: 𝑀𝑀𝑀𝑀 = = 3.23
26
Scores 0 -10 10 – 20 20 – 30 30 – 40 40 – 50 50 – 60
No of students 8 10 10 16 4 2
Solution:
1. N = 50
𝑁𝑁 50
2. Median = � �th item =� �th item =25th item
2 2
∴ Median class = 20 – 30
𝑁𝑁
−𝑐𝑐𝑐𝑐
25−18
3. Calculate the median: Median = 𝐿𝐿 + 2 𝑓𝑓 × 𝐼𝐼 = 20 + × 10 = 27
10
572
4. Compute the Mean Deviation about Median: 𝑀𝑀𝑀𝑀 = = 11.44
50
𝑀𝑀𝑀𝑀𝑀𝑀𝑀𝑀𝑀𝑀𝑀𝑀 𝑑𝑑𝑑𝑑𝑑𝑑𝑑𝑑𝑑𝑑𝑑𝑑𝑑𝑑𝑑𝑑𝑑𝑑 11.44
5. Coefficient of Mean Deviation from median = �
= = 0.4237
𝑀𝑀 27
16
ISC XI- Applied Mathematics
Variance
Recall that the absolute values were taken while calculating mean deviation, otherwise the deviations may
cancel among themselves. Another way to overcome this difficulty due to signs of deviation is to take squares
of all deviations. The mean of squared deviations is called variance. It is denoted as σ2. This leads to a
limitation for variance as its units are different from the units of the variable ‘x’ making it difficult to compare
datasets that use different units of measurement.
Therefore, it leads us to define standard deviation which is also called the root mean squared deviation and is
considered the square root of variance.
Standard Deviation
Standard deviation is a key statistical measure that quantifies the amount of variation or dispersion in a set of
data values. It tells us how much individual data points deviate from the mean (average) of the dataset.
Standard Deviation is the most reliable measure of dispersion because it is based on all observations in the
dataset and allows further mathematical treatment due to squaring of deviations.
A low standard deviation indicates that the data points are closely clustered around the mean, whereas a high
standard deviation suggests that the data points are spread out over a wider range.
For example, there are five people of different heights
standing against a wall. The average (mean) height of
these five individuals is 106.5 units. The standard
deviation, which measures how much individual It is Root Mean Square Deviation from Mean.
heights deviate from the mean, is 32.8. Since this value As per the name take actions in the reverse order
is relatively high, it indicates that some people in the i.e.
group are significantly taller or shorter than the
average height. 1. Take deviations of all the entries from mean.
2. Take square of the deviations.
Standard Deviation (denoted as σ) is defined as the
3. Get the mean of the squared deviations.
positive square root of the arithmetic mean of the
squares of the deviations of all the observations from 4. Finally take positive square root of this mean.
their arithmetic mean.
1
For ungrouped data: 𝜎𝜎 = � ∑𝑛𝑛
𝑖𝑖=1(𝑥𝑥𝑖𝑖 − 𝑥𝑥̅ )
2
𝑛𝑛
1 𝑛𝑛
For grouped data: 𝜎𝜎 = � ∑𝑛𝑛
𝑖𝑖=1 𝑓𝑓𝑖𝑖 (𝑥𝑥𝑖𝑖 − 𝑥𝑥̅ ) where 𝑁𝑁 = ∑𝑖𝑖=1 𝑓𝑓𝑖𝑖
2
𝑁𝑁
17
ISC XI- Applied Mathematics
𝛴𝛴𝑓𝑓𝑓𝑓2 𝛴𝛴𝛴𝛴𝛴𝛴 2
ii. Short-cut method 𝜎𝜎 = � −� � , where 𝑢𝑢 = 𝑥𝑥 − 𝐴𝐴.
𝑁𝑁 𝑁𝑁
𝛴𝛴𝑓𝑓𝑓𝑓2 𝛴𝛴𝛴𝛴𝛴𝛴 2 𝑥𝑥−𝐴𝐴
iii. Step deviation method 𝜎𝜎 = � 𝑁𝑁
−� 𝑁𝑁
� × 𝐼𝐼, where 𝑑𝑑 = 𝐼𝐼
.
The coefficient of variation (CV) is the standardised measure of the dispersion of data points around the
mean. It is also called relative standard deviation (RSD). It’s calculated as the ratio of the standard deviation
to the mean and is often expressed as a percentage, making it easy to compare the variability of datasets with
different units or scales.
𝜎𝜎
𝐶𝐶𝐶𝐶 = × 100
𝑥𝑥̅
A lower CV indicates less variability and higher stability, while a higher CV suggests greater variability.
x x2
1 1
2 4
3 9
4 16
5 25
6 36
7 49
8 64
9 81
10 100
55 385
x 10 20 30 40 50 60
f 9 18 25 27 14 7
18
ISC XI- Applied Mathematics
Solution:
x f fx fx2
10 9 90 900
20 18 360 7200
30 25 750 22500
40 27 1080 43200
50 14 700 35000
60 7 420 25200
100 3400 134000
Example 17. Find the standard deviation using the short-cut method for the following:
No of articles produced 18 19 20 21 22 23 24 25 26 27
No of workers 3 7 11 14 18 17 13 8 5 4
Solution:
x f 𝒖𝒖 = 𝒙𝒙 − 𝑨𝑨 fu fu2
18 3 -4 -12 48
19 7 -3 -21 63
20 11 -2 -22 44
21 14 -1 -14 14
22 18 0 0 0
23 17 1 17 17
24 13 2 26 52
25 8 3 24 72
26 5 4 20 80
27 4 5 20 100
100 38 490
Assumed mean A = 22
Example 18. Find the standard deviation using the step deviation method for the following:
19
ISC XI- Applied Mathematics
Solution:
CI x f 𝐱𝐱 − 𝐀𝐀 fd fd2
𝐝𝐝 =
𝐈𝐈
Example 19. The mean and standard deviation of 25 observations is 60 and 3 respectively. Later, it was
decided to omit an observation which was incorrectly recorded as 50. Calculate the mean and standard
deviation of the remaining 24 observations.
Solution:
∑𝑥𝑥
𝑥𝑥̅ = 60 ⇒ = 60 ⇒ ∑𝑥𝑥 = 25 × 60 = 1500
25
∑𝑥𝑥 2 ∑𝑥𝑥 2 ∑𝑥𝑥 2
𝜎𝜎 = 3 ⇒ −� � =9⇒ − (60)2 = 9
𝑛𝑛 𝑛𝑛 25
∑𝑥𝑥 2
= 3609 ⇒ ∑𝑥𝑥 2 = 90225
25
1450
∴ New mean = = 60 ⋅ 42
24
87725
New Standard deviation = � − (60 ⋅ 42)2 = 2 ⋅ 15
24
20
ISC XI- Applied Mathematics
Mean Deviation does not take into account the Standard Deviation takes into account the algebraic
algebraic signs in its calculations. signs in its calculations.
Mean Deviation can be computed from Mean, Standard Deviation is always computed from the
Median or Mode. Mean.
It is less frequently used. It is one of the most commonly used measures of
variability and frequently used.
When there are a greater number of outliers in the When there are fewer outliers in the data, the
data, mean absolute deviation is used. standard deviation is employed.
Standard deviation can be used to compare the variability of temperatures between two cities. A city
with a higher standard deviation in daily temperatures will have more extreme temperature swings than
a city with a lower standard deviation.
Imagine comparing the average daily maximum temperatures for Kolkata (a coastal city) and Hyderabad
(an inland city). Kolkata might have a lower standard deviation, meaning the daily temperatures are
more consistent, while Hyderabad might have a higher standard deviation, indicating more significant
temperature variations.
21
ISC XI- Applied Mathematics
Example 20. A company has two production units. The details of the number of workers and their average
monthly wages (in ₹) are given below:
Production Unit Number of Workers Mean Salary (₹) Standard Deviation (₹)
A 40 25,000 3,000
B 60 30,000 4,000
(a) Find the combined mean salary of all the workers.
(b) Find the combined standard deviation of the salaries.
Solution:
𝑛𝑛1 = 40, 𝑛𝑛2 = 60
𝑥𝑥̅1 = 25,000, 𝑥𝑥̅2 = 30,000
𝜎𝜎1 = 3000 𝜎𝜎2 = 4000
Combined mean
𝑛𝑛1 𝑥𝑥1 + 𝑛𝑛2 𝑥𝑥2 40 × 25000 + 60 × 30000 2800000
𝑥𝑥̅ = = = = ₹ 28,000
𝑛𝑛1 + 𝑛𝑛2 40 + 60 100
= 4381.78
Example 21. The following table gives the result of three centres of a public examination.
Calculate the mean and standard deviation of the combined scores of the three centres.
22
ISC XI- Applied Mathematics
Solution:
𝑛𝑛1 (𝜎𝜎12 + 𝑑𝑑12 ) + 𝑛𝑛2 (𝜎𝜎22 + 𝑑𝑑22 ) + 𝑛𝑛3 (𝜎𝜎32 + 𝑑𝑑32 ) 200(9 + 81) + 250(16 + 36) + 300(25 + 1)
𝜎𝜎 = � =�
𝑛𝑛1 + 𝑛𝑛2 + 𝑛𝑛3 200 + 250 + 300
38800
= � = 7.19
750
Mode
Mode is the value of the item which occurs the maximum number of times i.e. which has the highest frequency.
For ungrouped data: the mode can be obtained by simple observation.
For grouped data: Identify the class interval having the maximum frequency. Call it as Modal class. From
the data set find the following corresponding to the modal class.
L = lower boundary of the modal class
𝑓𝑓𝑚𝑚−1 = frequency of the class preceding the modal class
𝑓𝑓𝑚𝑚 = frequency of the modal class
𝑓𝑓𝑚𝑚+1 = frequency of the class following the modal class
I = size of class (assuming that all classes have equal size).
N = total number of items
𝑓𝑓𝑚𝑚 −𝑓𝑓𝑚𝑚−1
Then, Mode = 𝐿𝐿 + × 𝐼𝐼
2𝑓𝑓𝑚𝑚 −𝑓𝑓𝑚𝑚−1 −𝑓𝑓𝑚𝑚+1
NOTE:
• If a data set has only 1 value that occurs most often, the set is called unimodal.
• If a data set has 2 values that occur with the greatest frequency, the set is called bimodal.
• If a set of data has more than 2 values that occur with the same greatest frequency, the set is called
multimodal.
23
ISC XI- Applied Mathematics
Solution:
(i) 4 occurs three times – which is the maximum number of occurrences.
∴ Mode = 4
(ii) Both 3 and 4 occurs twice – which is the maximum number of occurrences.
∴ Mode = 3 and 4
(iii) All numbers occur the same number of times. So, there is no mode.
Example 23. The following table represents the number of minutes that students spend studying for a math
test:
What is the mode of the amount of time spent studying for the math test?
Solution: Changing the given data set into continuous class intervals as shown below:
Since, the highest frequency is 10, therefore the modal class is 10.5 - 20.5.
= 17.17 minutes
Solution:
Marks No of students
0 - 20 10
20 – 40 x
40 -–60 30
60 – 80 y
80 -–100 14
24
ISC XI- Applied Mathematics
𝑓𝑓𝑚𝑚 −𝑓𝑓𝑚𝑚−1
Mode = 𝐿𝐿 + 2𝑓𝑓 × 𝐼𝐼
𝑚𝑚 −𝑓𝑓𝑚𝑚−1 −𝑓𝑓𝑚𝑚+1
30−𝑥𝑥
⇒ 54 = 40 + × 20
2(30)−𝑥𝑥−𝑦𝑦
30−𝑥𝑥
⇒ 54 - 40 = × 20
60−𝑥𝑥−𝑦𝑦
For Venue A:
𝜎𝜎 8000
𝐶𝐶𝐶𝐶 = × 100 = × 100 = 10%
𝑥𝑥̅ 80000
For Venue B:
𝜎𝜎 12000
𝐶𝐶𝐶𝐶 = × 100 = × 100 = 20%
𝑥𝑥̅ 60000
Since Venue A has a lower CV, its rental charges are more consistent, making it a better investment from a
stability standpoint.
25
ISC XI- Applied Mathematics
26
ISC XI- Applied Mathematics
Exercises
1. Choose the correct option:
(i) The mean of 100 observations is 50 and their standard deviation is 10. If 5 is added to each
observation, then new mean and new standard deviation respectively will be
(a) 50, 10 (c) 55, 10
(b) 55, 15 (d) 60, 10
(ii) The mean of 5 observations is 3 and their variance is 2. If three of the observations are 1, 3 and 5
then the other two observations are
(a) 2 and 4 (c) 7 and 2
(b) 2 and 6 (d) 7 and 9
(iii) The median of a set of 9 distinct observations is 20.5. If each of the largest 4 observations of the set
is increased by 2, then the median of the new set
(a) is increased by 2
(b) is decreased by 2
(c) is two times the original median
(d) remains the same as that of the original set
(v) Assertion (A): If the standard deviation of a dataset is zero, then all the values in the dataset are
equal.
Reason (R): Standard deviation is the square root of the arithmetic mean of the squares of deviations
from the actual mean.
(a) Both A and R individually true and R is the correct explanation of A.
(b) Both A and R individually true but R is not the correct explanation of A.
(c) A is true but R is false.
(d) A is false but R is true.
2. Given below are the prices of shares of a company for the last 10 days.
172,164,188,214,190,237,200,195,208,230
Find the quartile deviation.
27
ISC XI- Applied Mathematics
3. Calculate the range, quartile deviation and coefficient of quartile deviation of wages.
Wages (Rs) 30 - 32 32 - 34 34 - 36 36 - 38 38 - 40 40 - 42 42 - 44
No of members 12 18 16 14 12 8 6
4. For a group of 50 male workers the mean and S.D of their daily wages are Rs 630 and Rs 90 respectively
for a group of 40 female workers these are Rs 540 and Rs 60 respectively. Find the mean and S.D of these
90 workers.
7. Find the mean deviation and coefficient of mean deviation about median for the following data:
Heights (in cm) 95-105 105-115 115-125 125-135 135-145 145-155
Number of Girls 9 15 23 30 13 10
8. A gardener buys 10 packets of seeds from two different companies. Each pack contains 20 seeds, and he
records the number of plants which grow from each pack.
Company A 20 5 20 20 20 6 20 20 20 8
Company B 17 18 15 16 18 18 17 15 17 18
(i) Find the median and mode for each company’s seeds.
(ii) Which company does the mode suggest is best?
(iii) Find the range for each company’s seeds.
9. The diameter of circles (in mm) drawn in a design are given below:
Diameter (in mm) 33 – 36 37 – 40 41 – 44 45 – 48 49 – 52
No. of circles 15 17 21 22 25
Calculate the standard deviation and mean diameter of the circles.
28
ISC XI- Applied Mathematics
11. A catering company is comparing the weekly transportation costs of two delivery vehicles - a diesel van
and an electric van - over a few months. The data shows:
• Diesel Van: Mean weekly cost = ₹5,000; Standard Deviation = ₹750
• Electric Van: Mean weekly cost = ₹4,200; Standard Deviation = ₹1,050
(i) Calculate the coefficient of variation (CV) for both vehicles.
(ii) Which vehicle shows more consistent costs and would be better for budgeting?
Answers
1. (i) (c) 55, 10
(v) (a) Both A and R individually true and R is the correct explanation of A.
2. 17
3. range = 14, quartile deviation = 2.85, coefficient of quartile deviation = 0.0792
4. mean = 590, S.D = 90
5. (i) Case 1. Variance = 9.2, standard deviation = 3.03
Case 2. Variance = 9.2, standard deviation = 3.03
(ii) There is no change in variance and standard deviation of the given data if the origin is changed i.e., if a
constant is added to each observation.
6. 5.003 and 0.0769
7. 11.54 and 0.0916
8. (a) Company A: Median = 20, Mode = 20 & Company B: Median = 17, Mode = 18
(b) The mode for Company A is 20 and the mode for Company B is 18. Since a larger number of seeds
grew from the packs in Company A, the mode suggests that Company A is better.
(c) Company A's range: 15, Company B's range: 3
9, Mean = 43.5 Standard deviation = 5.55
10. (i) 12.575 (ii) 12.74
11. (i) CV of Diesel Van = 15%
CV of Electric Van = 25%
(ii) Since the Diesel Van has a lower CV, its weekly costs are more consistent. Therefore, it would be a
better choice for budgeting and planning, despite slightly higher average costs.
29
ISC XI- Applied Mathematics
Project Work
"Can Measures of Dispersion Be Manipulated to Misrepresent Data?"
Case Study:
A leading consumer electronics company, TechTrend, recently launched a new smartphone. To promote their
product, they conducted a survey comparing customer satisfaction scores between their phone and a
competitor’s model.
The company published the following statistical summary in their advertisement:
TechTrend claimed that their smartphone had a much higher customer satisfaction rating. However, a
group of independent analysts questioned their use of dispersion measures.
Discussion Questions:
1. What role does Standard Deviation play in understanding customer satisfaction in this case?
2. Does TechTrend’s claim seem misleading? Why or why not?
3. How could a company selectively use measures of dispersion to create a biased representation of data?
4. What other statistical measures could be considered before making a judgment?
5. How can businesses ensure ethical use of statistics in advertising?
30
ISC XI- Applied Mathematics
HISTORICAL NOTE
The concept of dispersion in statistics has evolved over centuries, with contributions from various
mathematicians and statisticians:
• 17th Century – Beginnings of Probability Theory:
− Blaise Pascal and Pierre de Fermat laid the foundation for probability theory, which later
influenced statistical measures.
• 18th Century – Early Statistical Concepts:
− Abraham de Moivre introduced the normal distribution in his book The Doctrine of Chances
(1733), hinting at the importance of variability in data.
• 19th Century – Formal Development of Dispersion Measures:
− Carl Friedrich Gauss (1809) developed the method of least squares, closely linked to variance
and standard deviation.
− Francis Galton (late 1800s) introduced the concept of statistical correlation and spread,
influencing modern dispersion measures.
• Early 20th Century – Standard Deviation and Variance:
− Karl Pearson formally introduced standard deviation as a measure of spread in 1894.
− Ronald Fisher (1920s) developed analysis of variance (ANOVA), further expanding
applications of dispersion in research.
• Modern Era – Computational Advancements:
− The widespread use of computers in the 20th century enabled large-scale data analysis, making
dispersion measures essential in fields like finance, healthcare, and machine learning.
31
ISC XI- Applied Mathematics
SCOPE
LEARNING OUTCOMES
PRE-REQUISITES
32
ISC XI- Applied Mathematics
Introduction to Moments
In Statistics, moments are used to describe the characteristic of a distribution. They represent a convenient and
unifying method for summarizing many of the most commonly used statistical measures.
Generally, in any frequency distribution, four moments are obtained which are known as first, second, third
and fourth moments. These four moments describe the information about mean, variance, skewness, and
kurtosis of a frequency distribution.
Calculation of moments gives some features of a distribution which are of statistical importance. Moments
can be classified as raw and central moment. Raw moments are measured about any arbitrary point A (say). If
A is taken to be zero, then raw moments are called moments about origin. When A is taken to be arithmetic
mean we get central moments. The first raw moment about origin is the mean whereas the first central moment
is zero. The second central moment is variance and the third and fourth moments are useful in measuring
skewness and kurtosis.
Central Moments
If the arbitrary origin of moments of variable X is taken as arithmetic mean i.e. A = 𝑥𝑥̅ the moments are called
central moments.
For ungrouped data, the central moment of variable X is given by
𝑛𝑛
1
𝜇𝜇𝑟𝑟 = �(𝑥𝑥𝑖𝑖 − 𝑥𝑥̅ )𝑟𝑟 ; 𝑟𝑟 = 0,1,2, …
𝑛𝑛
𝑖𝑖=1
When r = 0,
𝑛𝑛
1 𝛴𝛴𝑓𝑓𝑖𝑖
𝜇𝜇0 = � 𝑓𝑓𝑖𝑖 (𝑥𝑥𝑖𝑖 − 𝑥𝑥̅ )0 = =1
𝑛𝑛 𝑛𝑛
𝑖𝑖=1
33
ISC XI- Applied Mathematics
When r = 1,
𝑛𝑛
1 𝛴𝛴𝑓𝑓𝑖𝑖 𝑥𝑥𝑖𝑖 𝛴𝛴𝛴𝛴
𝜇𝜇1 = � 𝑓𝑓𝑖𝑖 (𝑥𝑥𝑖𝑖 − 𝑥𝑥̅ )1 = − 𝑥𝑥̅ = 𝑥𝑥̅ − 𝑥𝑥̅ = 0
𝑛𝑛 𝑛𝑛 𝑛𝑛
𝑖𝑖=1
When r = 2,
𝑛𝑛
1
𝜇𝜇2 = � 𝑓𝑓𝑖𝑖 (𝑥𝑥𝑖𝑖 − 𝑥𝑥̅ )2 = 𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉
𝑛𝑛
𝑖𝑖=1
Moments can be extended to higher powers but in actual practice the first four moments are enough to describe
a series.
Introduction to Skewness
In statistics, when we collect and organize data, we often create a graph to see how the data is distributed.
A normal curve, also known as a Gaussian or bell curve, is a symmetrical, bell-shaped probability
distribution [Link] normal curve is symmetrical around its mean. In a normal distribution, the mean,
median, and mode are all equaland located at the centre of the curve. The curve has a characteristic bell shape,
with the highest point at the mean and the tails tapering off towards both ends.
But real-life data is not always symmetrical. Sometimes one side of the graph stretches out more than the
other. This is called skewness.
What is Skewness?
A distribution is said to be 'skewed' when the mean and the median fall at different points in the distribution,
and the balance (or centre of gravity) is shifted to one side or the other-to left or right. Measures of skewness
tell us the direction and the degree of asymmetry.
In symmetrical distribution the mean, median and mode are identical. The more the mean moves away from
the mode, the larger the asymmetry or skewness.
Types of Skewness
1. Symmetrical Distribution
34
ISC XI- Applied Mathematics
Measures of Skewness
To study the extent of asymmetry and direction in a series, various measures of skewness are employed. These
measures of skewness can be both absolute or relative.
35
ISC XI- Applied Mathematics
The formula shows if (Q3−Md) > (Md−Q1), skewness is positive otherwise it is negative. This is true
because (Q3−Md) > (Md−Q1) implies that there is a longer tail on the right hand side i.e., skewness is
right handed or positive.
𝑀𝑀𝑀𝑀𝑀𝑀𝑀𝑀 − 𝑀𝑀𝑀𝑀𝑀𝑀𝑀𝑀
𝑠𝑠𝑘𝑘 =
𝑆𝑆𝑆𝑆𝑆𝑆𝑆𝑆𝑆𝑆𝑆𝑆𝑆𝑆𝑆𝑆 𝐷𝐷𝐷𝐷𝐷𝐷𝐷𝐷𝐷𝐷𝐷𝐷𝐷𝐷𝐷𝐷𝐷𝐷
36
ISC XI- Applied Mathematics
This method is particularly useful in case of open end distributions and where extreme values are present or
when class-intervals are unequal. Skewness should be measured by Bowley's method also when positional
measures are called for.
But the main drawback of this measure is that it is based on central 50% of the data and it ignores the
remaiaing 50% of the data i.e., 25% of the data below Q1 and 25% of the data above Q3.
Then the sign of skewness would depend upon whether the value of µ3 and is positive or negative. That is
why γ1 is preffered to measure skewness.
Example 1. The table below shows the weekly pocket money (₹) received by 103 students:
Pocket Money 100 200 300 400 500 600 700 800 900 1000
No of students 3 5 8 10 12 15 20 15 10 5
From the graph we can infer that for the curve the tail is longer on the left. So the data is negatively skewed.
Example 2. From the marks secured by 120 students in Sections A and B of a class, the following measures
are obtained:
Section A : 𝑥𝑥̅ = 46.83, 𝜎𝜎 = 14.8, Mode = 51.67
Section B : 𝑥𝑥̅ = 47.83, 𝜎𝜎 = 14.8, Mode = 47.07
Find the skewness of marks in both sections. Which section has more skewed marks and in which direction?
Solution:
𝑥𝑥̅ −𝑀𝑀𝑀𝑀𝑀𝑀𝑀𝑀 46.83−51.67
For Section A: 𝑆𝑆𝑘𝑘 = = = −0.327
𝜎𝜎 14.8
38
ISC XI- Applied Mathematics
Example 3. For a given data, Q1 = 58, Md = 59 and Q3 = 61, find the coefficient of skewness.
Solution . We know
𝑄𝑄3 + 𝑄𝑄1 − 2𝑀𝑀𝑑𝑑 61 + 58 − 2 × 59
𝑆𝑆𝑘𝑘 = = = 0.33
𝑄𝑄3 − 𝑄𝑄1 61 − 58
Example 4. In a frequency distribution, the coefficient of skewness based upon quartiles is 0.6. If the sum of
the upper and lower quartiles is 100 and the median is 38, find the value of the upper quartile.
Solution. We know,
𝑄𝑄3 + 𝑄𝑄1 − 2𝑀𝑀𝑑𝑑
𝑆𝑆𝑘𝑘 =
𝑄𝑄3 − 𝑄𝑄1
Substituting the given values,
100−2×38 100−2×38
0.6 = 𝑄𝑄3 −𝑄𝑄1
⇒ 𝑄𝑄3 − 𝑄𝑄1 = 0.6
= 40
Now,
(𝑄𝑄3 + 𝑄𝑄1 ) + (𝑄𝑄3 − 𝑄𝑄1 ) = 100 + 40 ⇒ 2𝑄𝑄3 = 140 ⇒ 𝑄𝑄3 = 70
Example 5. Assume that a firm has selected a random sample of 100 employees from its production line
representing their monthly working hours and has obtained the data shown in table below.
Monthly 130 - 134 135 - 139 140 - 144 145 - 149 150 - 154 155 - 159 160-164
working
hours
No of 3 12 21 28 19 2 5
employees
Compute:
(i) the arithmetic mean
(ii) the standard deviation
(iii) Karl Pearson’s Coefficient of Skewness.
Solution:From the given data we obtained
CI x f d fd fd2
160-164 162 5 3 9 45
100 4 208
39
ISC XI- Applied Mathematics
A = 147
4 208 4 2
𝑥𝑥̅ = 147 + � × 5� = 147 ⋅ 2 𝝈𝝈 = � −� � × 5 = 7 ⋅ 21
100 100 100
Example 6. Calculate the Bowley’s Coefficient of Skewness from the data given below:
CI 30 - 40 40 - 50 50 - 60 60 - 70 70 - 80 80 - 90 90 - 100
Frequency 1 3 11 21 43 32 9
Solution:From the given data we obtained,
CI f Cf
30 - 40 1 1
40 - 50 3 4
50 - 60 11 15
60 - 70 21 36
70 - 80 43 79
80 - 90 32 111
90 - 100 9 120
n = 120
𝑛𝑛 120
𝑄𝑄1 =�4�th item =� 4 �th item =30th item
∴ 𝑄𝑄1 lies in the class = 60 – 70
𝑛𝑛
� �−𝑐𝑐𝑐𝑐 30−15
𝑄𝑄1 = 𝐿𝐿 + 4
× 𝐼𝐼 = 60 + × 10 = 67.14
𝑓𝑓 21
𝑛𝑛 120
Median =2 �4�th item =2 � 4 �th item = 60th item
∴𝑀𝑀𝑀𝑀𝑀𝑀𝑀𝑀𝑀𝑀𝑀𝑀 lies in the class = 70 – 80
𝑛𝑛
2 � �−𝑐𝑐𝑐𝑐 60−36
𝑄𝑄2 = 𝐿𝐿 + 4
× 𝐼𝐼 = 70 + × 10 = 75.5
𝑓𝑓 83
𝑛𝑛 120 th
𝑄𝑄3 =3 �4�th item =3 � 4
� item =75th item
40
ISC XI- Applied Mathematics
Example 7. Find out if there is skewness in the distribution and whether the skewness is positive or negative
given that 𝜇𝜇1 = 0, 𝜇𝜇2 = 40, 𝜇𝜇3 = −100, 𝜇𝜇4 = 200.
Solution:We know
𝜇𝜇32 (−100)2 1
𝛽𝛽1 = 3 = 3
=
𝜇𝜇2 (40) 64
1
Since, 𝜇𝜇3 is negative we can infer that skewness is also negative and of the value
64
.
Example 8. You are given the following information before and after the settlement of an industrial dispute:
3(𝑀𝑀𝑀𝑀𝑀𝑀𝑀𝑀−𝑀𝑀𝑀𝑀𝑀𝑀𝑀𝑀𝑀𝑀𝑀𝑀) 3(325−300)
After settlement 𝑠𝑠𝑘𝑘 = = = +3.75
𝑆𝑆𝑆𝑆𝑆𝑆𝑆𝑆𝑆𝑆𝑆𝑆𝑆𝑆𝑆𝑆 𝐷𝐷𝐷𝐷𝐷𝐷𝐷𝐷𝐷𝐷𝐷𝐷𝐷𝐷𝐷𝐷𝐷𝐷 20
Thus, from negative skewness before the settlement the skewness has become positive so there is no longer a
tail towards the right of the curve. This indicates that the number of workers getting more wages has gone
down.
4. If the sums of positive and negative deviations obtained from mean, median or mode are equal then there
is no asymmetry;
5. If a graph of a data become a normal curve and when it is folded at middle and one part overlap fully on
the other one then there is no asymmetry.
Introduction to Kurtosis
In statistics, kurtosis is a measure that describes the "peakedness" or flatness of a distribution curve
compared to a normal distribution.
While skewness tells us how stretched the curve is to the left or right, kurtosis tells us how tall and sharp
or flat and wide the peak of the curve is, which further indicates the presence and extent of outliers by
assessing the tails and peaks.
What is Kurtosis?
Kurtosis is another measure of the shape of a frequency curve.
It is a Greek word, which means bulginess. While skewness
signifies the extent of asymmetry, kurtosis measures the degree
of peakedness of a frequency distribution. Karl Pearson
classified curves into three types on the basis of the shape of
their peaks. These are Mesokurtic, leptokurtic and platykurtic.
These three types of curves are shown in figure below:
Types of Kurtosis
42
ISC XI- Applied Mathematics
William Gosset, an English statistician, in his 1927 paper, used the visual analogy of a platypus for
platykurtic and two leaping kangaroos for leptokurtic distributions to aid in remembering the shapes.
(Source: ‘Student’, 1972, Errors of routine [Link],19,[Link] p.160)
43
ISC XI- Applied Mathematics
Karl Pearson defined the following β and γ coefficients of kurtosis, based upon the second and fourth central
moments:
𝜇𝜇4
𝛽𝛽2 =
𝜇𝜇22
Since the curve is taller and sharper than normal, it is a leptokurtic curve.
44
ISC XI- Applied Mathematics
Example 10. The first four moments of a distribution are 0, 2.5, 0.7 and 18.75. Comment on the Kurtosis of
the distribution.
Solution:
𝜇𝜇4 18.75
𝛽𝛽2 = = =3
𝜇𝜇22 (2.5)²
The curve is Mesokurtic – Normal Kurtosis.
Example 11. Calculate Skewness and Kurtosis from the following grouped data.
Class Frequency
0-2 5
2-4 16
4-6 13
6-8 7
8 - 10 5
10 - 12 4
x f fx �
𝒙𝒙 − 𝒙𝒙 �)𝟐𝟐
𝒇𝒇(𝒙𝒙 − 𝒙𝒙 �)𝟑𝟑
𝒇𝒇(𝒙𝒙 − 𝒙𝒙 �)𝟒𝟒
𝒇𝒇(𝒙𝒙 − 𝒙𝒙
Class
0-2 1 5 5 -4.12 84.872 -349.6726 1440.6513
45
ISC XI- Applied Mathematics
Now,
∑𝒇𝒇𝒇𝒇 256
x� = = = 5 ⋅ 12,
Σf 50
1 𝑛𝑛 395.28
𝜇𝜇2 = �𝑖𝑖=1 𝑓𝑓𝑖𝑖 (𝑥𝑥𝑖𝑖 − 𝑥𝑥̅ )2 = = 7.9056,
𝑛𝑛 50
𝑛𝑛
1 649.6128
𝜇𝜇3 = � 𝑓𝑓𝑖𝑖 (𝑥𝑥𝑖𝑖 − 𝑥𝑥̅ )3 = = 12.99,
𝑛𝑛 50
𝑖𝑖=1
1 𝑛𝑛 7766.0233
𝜇𝜇4 = 𝑛𝑛 �𝑖𝑖=1 𝑓𝑓𝑖𝑖 (𝑥𝑥𝑖𝑖 − 𝑥𝑥̅ )3 = 50
= 155.32,
𝜇𝜇2 (12.99)2
𝛽𝛽1 = 𝜇𝜇33 = (7.9056)3
= 0.342,
2
Since, 𝜇𝜇3 is positive we can infer that skewness is also positive and of the value 0.342.
𝝁𝝁𝟒𝟒 155.32
𝜷𝜷𝟐𝟐 = 𝟐𝟐
= = 2.4852
𝝁𝝁𝟐𝟐 (7.9056)²
The curve is Platykurtic – Low Kurtosis.
46
ISC XI- Applied Mathematics
Exercises
1. Choose the correct option:
(i) In Uni-model distribution, if the mode = 50 and mean = 60, then the distribution will be
(a) Negatively skewed (c) Positively skewed
(b) Symmetrical (d) Not possible to draw inference
(ii) If the third moment about mean is zero, then the distribution is:
(a) Negatively skewed (c) Positively skewed
(b) Symmetrical (d) Not possible to draw inference
47
ISC XI- Applied Mathematics
2. For 40 observations, it is given that ∑𝑓𝑓𝑓𝑓 = 322, 𝛴𝛴𝛴𝛴𝑥𝑥 2 = 2780 and mode = 8. Calculate the Karl
Pearson’s coefficient of skewness with the given data and comment on the degree of skewness.
3. The standard deviation of symmetrical distribution is 3. What must be the value of the fourth moment
about the mean in order that the distribution be mesokurtic?
4. In a frequency distribution the coefficient of skewness based on quartiles is 0.6. If the sum of the upper
and the lower quartiles is 100 and the median is 38, find the value of the upper quartile.
5. Calculate the Karl Pearson’s coefficient of skewness for the following data:
6. Calculate Bowley’s coefficient of skewness for the following distribution of weekly wage of workers.
Wages Below 300 300 - 400 400 - 500 500 - 600 600 - 700 Above 700
5 8 18 35 27 7
No of workers
7. The first four central moments of distribution are 0, 2.5, 0.7 and 18.75. Comment on the skewness and
kurtosis of the distribution.
Marks 10 - 20 20 - 30 30 - 40 40 - 50 50 - 60
No. of marks 2 4 7 6 1
48
ISC XI- Applied Mathematics
10. Compute the moment coefficient of kurtosis from the following data.
x 5 10 15 20 25 30 35
f 8 15 18 29 23 17 5
11.
(i) Calculate the Pearson's coefficient of skewness
for Curve B.
(ii) Identify which curve is symmetrical, positively
skewed, and leptokurtic, respectively. Justify
your answer.
(iii) Based on the given standard deviations, which
curve is most dispersed?
12.
(i) Which of the three curves shows no skewness?
(ii) Which curve indicates positive skewness?
(iii) Which curve has highest kurtosis?
Answers
1. (i) (c) Positively skewed
(iii) (d) 25
(v) (a) Both A and R individually true and R is the correct explanation of A.
2. 0.023(fairly symmetrical)
3. 243
4. Upper quartile is 70
5. – 0.122
49
ISC XI- Applied Mathematics
Project Work
"Analysing Health and Lifestyle Patterns Using Skewness and Kurtosis"
Objective:
To collect real-life data, calculate skewness and kurtosis, and interpret what they tell us about health and
lifestyle trends in a selected group.
Steps to Follow:
1. Choose a Dataset
Pick any one of the following and try to collect data from 20–30 individuals for better analysis.
50
ISC XI- Applied Mathematics
6. Conclusion
• Are Class XII students healthy?
• Is the lifestyle consistent or highly varied?
• Are there outliers?
• What could improve?
HISTORICAL NOTE
• Skewness and kurtosis, measures of distribution shape, were introduced by Karl Pearson, an English
biostatistician and mathematician, in the late 19th and early 20th centuries. Karl Pearson introduced
the concept of kurtosis in 1905 and developed a system of probability distributions that varied in
skewness, measured relative to the mode.
• Ronald Fisher, British mathematical statistician and biologist, with his 1925 text "Statistical
Methods for Research Workers," helped introduce skewness and kurtosis to many generations of
statisticians.
• While Thiele, a Danish astronomer, actuary, and mathematician, did not use the terms "skewness"
and "kurtosis”, but his work laid the groundwork for these concepts.
51
ISC XI- Applied Mathematics
SCOPE
∑ ( x - x )( y - y ) ∑𝑑𝑑𝑥𝑥 𝑑𝑑𝑦𝑦
r= = ∑𝑑𝑑 2
×∑𝑑𝑑𝑥𝑥 2
∑ (x - x ) ∑(y - y)
2 2 𝑥𝑥
1
∑ xy − ∑ x∑ y
r= N
1
2
∑x − (∑ x )2 ∑ y 2 − 1 (∑ y )2
N N
1
( ∑ u )( ∑ v )
∑ uv -
r= N
2 1 2 2 1 2
∑ u − (∑ u) ∑ v − (∑ v)
N N
LEARNING OUTCOMES
PRE-REQUISITES
52
ISC XI- Applied Mathematics
Introduction
In statistics, earlier methods such as measures of central tendency and dispersion are primarily used to analyse
a single variable at a time. This kind of analysis, where only one variable is involved, is known as Univariate
Distribution.
However, in many real-world situations, data involves two variables simultaneously. For example, we may
study relationships between income and expenditure, price and demand, or height and weight. When two
variables are considered together, the data is referred to as a Bivariate Distribution.
In such cases, it becomes essential to examine whether and how the variables are related. This leads us to the
concept of Correlation, which is a statistical technique used to measure the strength and direction of a
relationship between two variables. Understanding correlation helps in drawing meaningful conclusions and
making informed decisions based on patterns observed in the data.
Before understanding the idea of correlation let us first learn about the covariance.
Covariance
Covariance is a statistical representation of the degree to which the two given variables vary together.
Basically, Covariance is a number reflecting the degree to which the two variables vary together.
• If both variables increase or decrease together, the covariance is positive.
• If one variable increase while the other decreases, the covariance is negative.
• If there is no consistent pattern, the covariance is close to zero.
The symbol of Covariance of two variables (say X and Y) is denoted by Cov (X, Y).
Note: Covariance is independent of the choice of the origin, but it depends on the scale of measurement.
We can also say,
1 1 1
Cov(𝑋𝑋, 𝑌𝑌) = �𝛴𝛴𝑥𝑥𝑖𝑖 𝑦𝑦𝑖𝑖 − Σ𝑥𝑥𝑖𝑖 ∑𝑦𝑦𝑖𝑖 �= 2 [𝑛𝑛𝑛𝑛𝑥𝑥𝑖𝑖 𝑦𝑦𝑖𝑖 − Σ𝑥𝑥𝑖𝑖 ∑𝑦𝑦𝑖𝑖 ]
𝑛𝑛 𝑛𝑛 𝑛𝑛
Even though, the value of covariance indicates whether statistical relationship exists between two variables or
not, it is not often used as its value depends on the dimensions(units) used for variables. The value of
covariance may be very large if some particular units of measurement for X and Y are used, while it may be
small if some other units of measurements are used.
Therefore, we define a new quantity which does not depend on dimensions of variables X and Y, that is we
define a non-dimensional quantity to study correlation. This measure of correlation is called Coefficient of
Correlation.
Correlation
Correlation is a statistical tool used to measure and analyze the relationship between two variables. It helps us
understand how one variable behaves in relation to another, which is particularly useful in studying economic
and social behavior. By examining correlation, we can identify whether changes in one variable are associated
with changes in another. For example, suppose we want to study the relationship between hours studied and
marks obtained by students. If students who study more hours generally score higher marks, then there is a
positive correlation between study hours and marks.
53
ISC XI- Applied Mathematics
Does this mean that an increased number of high school graduates is leading to more pizza consumption in
India? Not quite!
The more likely explanation is that India’s population has been increasing over time, which means that the
number of people receiving a high school degree and the total pizza being consumed are both increasing as
population increases.
Although these two variables are correlated, one does not cause the other.
• Negative Correlation:
When two variables move in opposite directions, i.e., when one increases the other decreases, and vice-
versa, then such a relation is called a Negative Correlation.
For example,
− An increase in the price of a commodity often leads to a decrease in the quantity demanded.
− A decrease in temperature immediately leads to an increased sale of woollen garments.
• Perfect Correlation:
If the relationship between the two variables is in such a way that it varies in equal proportion (increase
or decrease) it is said to be perfectly correlated. This can be of two types:
− Positive Correlation: When the proportional change in two variables is in the same direction, it is said
to be positively correlated. In this case, the Coefficient of Correlation is shown as +1.
− Negative Correlation: When the proportional change in two variables is in the opposite direction, it is
said to be negatively correlated. In this case, the Coefficient of Correlation is shown as -1.
• Zero Correlation:
55
ISC XI- Applied Mathematics
If there is no relation between two series or variables, it is said to have zero or no correlation. It means
that if one variable varies and it does not have any impact on the other variable, then there is a lack of
correlation between them. In such cases, the Coefficient of Correlation will be 0.
Degree of Correlation
Degree of Correlation Positive Correlation Negative Correlation
Perfect +1 -1
Zero 0 0
Scatter Diagram
A simple and attractive method of measuring correlation by diagrammatically representing bivariate
distribution for determination of the nature of the correlation between the variables is known as the Scatter
Diagram Method. This method gives the investigator/analyst a visual idea of the nature of the association
between the two variables. The values of the two variables are drawn in a cartesian plane by points/dots. Then
by the scatter or closeness of the points the degree of the correlation between the variables can be discussed.
56
ISC XI- Applied Mathematics
Positive Correlation
Negative Correlation
57
ISC XI- Applied Mathematics
No Correlation
The formula for calculating Karl Pearson's Coefficient of Correlation can be transformed into a working
formula:
𝛴𝛴𝑑𝑑𝑥𝑥 𝑑𝑑𝑦𝑦
⇒ 𝑟𝑟 =
2 2
𝑑𝑑 𝑑𝑑𝑦𝑦
𝑛𝑛 × � 𝑛𝑛𝑥𝑥 × � 𝑛𝑛
𝛴𝛴𝑑𝑑𝑥𝑥 𝑑𝑑𝑦𝑦
∴ 𝑟𝑟 =
∑𝑑𝑑𝑥𝑥 2 × ∑𝑑𝑑𝑦𝑦 2
58
ISC XI- Applied Mathematics
This method of determining Coefficient of Correlation should be applied only when the deviations of items
are taken from actual means and not from assumed means.
𝑟𝑟 2 , the square of the coefficient of correlation is called coefficient of determination. So, 0 ≤ 𝑟𝑟 2 ≤ 1.
Variation between X and Y is indicated by 𝑟𝑟 2 and not r.
For example, imagine you are studying the relationship between hours a student studies and marks scored in
an exam. You collect data and find that 𝑟𝑟 = 0 ⋅ 8 and 𝑟𝑟 2 = 0 ⋅ 64. As 𝑟𝑟 = 0 ⋅ 8 there is a strong positive
relationship between study time and exam [Link] more importantly, 𝑟𝑟 2 = 0 ⋅ 64 means that 64% of the
variation in exam marks is explained by the variation in study time. The remaining 36% variation is due to
other factors (like sleep, health, test anxiety, etc.).
Example 2. If the Covariance between two variables X and Y is 9.4 and the variance of Series X and Y are
10.6 and 12.5, respectively, then calculate the coefficient of correlation.
Solution:
𝛴𝛴𝑑𝑑𝑥𝑥 𝑑𝑑𝑦𝑦
Cov(𝑋𝑋, 𝑌𝑌) = = 9.4
𝑛𝑛
Variance of X = 𝜎𝜎𝑥𝑥 2 = 10.6 ⇒ 𝜎𝜎𝑥𝑥 = 3 ⋅ 25
Variance of Y = 𝜎𝜎𝑦𝑦 2= 12.5 ⇒ 𝜎𝜎𝑦𝑦 = 3 ⋅ 53
∑𝑑𝑑𝑥𝑥 𝑑𝑑𝑦𝑦 ∑𝑑𝑑𝑥𝑥 𝑑𝑑𝑦𝑦 1 1 1 1
𝑟𝑟 = = × × = 9.4 × × = 0.816
𝑛𝑛 × 𝜎𝜎𝑥𝑥 × 𝜎𝜎𝑦𝑦 𝑛𝑛 𝜎𝜎𝑥𝑥 𝜎𝜎𝑦𝑦 3.25 3.53
Coefficient of Correlation = 0.816 (positive, fairly high correlation)
Example 3. If ∑𝑥𝑥𝑖𝑖 = 15, ∑𝑦𝑦𝑖𝑖 = 40, ∑(𝑥𝑥𝑖𝑖 − 2)(𝑦𝑦𝑖𝑖 − 3) = 20 and n = 10, find Cov(x,y).
Solution:
∑(𝑥𝑥𝑖𝑖 − 2)(𝑦𝑦𝑖𝑖 − 3) = 20
⇒ ∑xi yi − 3∑xi − 2Σyi + 6n = 20
⇒ ∑xi yi − 3 × 15 − 2 × 40 + 6 × 10 = 20
⇒ ∑xi yi = 85
1 1 1 1
Cov(𝑥𝑥, 𝑦𝑦) = �𝛴𝛴𝛴𝛴𝛴𝛴 − Σx∑y� = �85 − (15)(40)� = 2.5
𝑛𝑛 𝑛𝑛 10 10
59
ISC XI- Applied Mathematics
∑ ( x - x )( y - y ) ∑𝑑𝑑𝑥𝑥 𝑑𝑑𝑦𝑦
r= =
∑𝑑𝑑𝑥𝑥 2 ×∑𝑑𝑑𝑥𝑥 2
∑ (x - x ) ∑(y - y)
2 2
Example 4. Using Actual Mean Method, determine the coefficient of correlation for the following data:
X 12 16 20 24 28 32 36
Y 6 9 12 15 18 21 24
Solution:
X �
𝑑𝑑𝑥𝑥 = 𝑿𝑿 − 𝑿𝑿 𝑑𝑑𝑥𝑥 2 Y 𝑑𝑑𝑦𝑦 � 𝑑𝑑𝑦𝑦 2
= 𝒀𝒀 − 𝒀𝒀 𝑑𝑑𝑥𝑥 𝑑𝑑𝑦𝑦
12 -12 144 6 -9 81 108
16 -8 64 9 -6 36 48
20 -4 16 12 -3 9 12
24 0 0 15 0 0 0
28 4 16 18 3 9 12
32 8 64 21 6 36 48
36 12 144 24 9 81 108
168 448 105 252 336
� = ∑X = 168 = 24, Y
X � = ∑Y = 105 = 15
n 7 n 7
Direct Method
Though this method can be used for all situations, it is mainly used when the values of X and Y are small.
1
∑XY − n ΣX∑Y 𝑛𝑛∑XY − ΣX∑Y
𝒓𝒓 = =
�ΣX 2 − 1 (∑X)2 �∑Y 2 − 1 (∑Y)2 �nΣX 2 − (∑X)2 �n∑Y 2 − (∑Y)2
n n
60
ISC XI- Applied Mathematics
Example 5. Using Direct Method, determine the coefficient of correlation for the following data:
X 5 4 3 2 1
Y 4 2 10 8 6
Solution:
X X2 Y Y2 XY
5 25 4 16 20
4 16 2 4 8
3 9 10 100 30
2 4 8 64 16
1 1 6 36 6
15 55 30 220 80
1
∑uv − n Σu∑v
𝒓𝒓 =
�Σ𝑢𝑢2 − 1 (∑u)2 �∑𝑣𝑣 2 − 1 (∑v)2
n n
Any value can be considered as assumed mean but for ease of calculations the integer value closest to the
average of the highest and the lowest value for a particular variable is taken as the assumed mean.
Example 6. Using Short-Cut Method, determine the coefficient of correlation for the following data:
X 78 89 96 69 59 79 68 62
Y 125 137 156 112 107 136 123 104
Solution:
X u = 𝑿𝑿 − 𝟖𝟖𝟖𝟖 u2 Y v = 𝒀𝒀 − 𝟏𝟏𝟏𝟏𝟏𝟏 v2 uv
78 -2 4 125 -5 25 10
89 9 81 137 7 49 63
96 16 256 156 26 676 416
69 -11 121 112 -18 324 198
59 -21 441 107 -23 529 483
79 -1 1 136 6 36 -6
68 -12 144 123 -7 49 84
62 -18 324 104 -26 676 468
-40 1372 -40 2364 1716
61
ISC XI- Applied Mathematics
1 1
∑uv − n Σu∑v 1716 − 8 (−40)(−40)
𝒓𝒓 = = = 0.95
�Σ𝑢𝑢2 1 1
− n (∑u)2 �∑𝑣𝑣 2 − n (∑v)2 �1372 − 1 (−40)2 �2364 − 1 (−40)2
8 8
Example 7. Calculate the value of the correlation coefficient for the following data:
(1, 13), (2, 23), (3, 33), (4, 43), (5, 53), (6, 63), (7, 73), (8, 83), (9, 93), (10, 103), (11, 10.5), (12, 11), (13,
11.5), (14, 12), (15, 12.5), (16, 13), (17, 13.5), (18, 14), (19, 14.5), (20, 15)
Solution:
We take 10 and 13 as the assumed mean for X and Y respectively.
X 𝒖𝒖 = 𝑿𝑿 − 𝟏𝟏𝟏𝟏 u2 Y 𝒗𝒗 = 𝒀𝒀 − 𝟏𝟏𝟏𝟏 v2 uv
1 -9 81 13 0 0 0
2 -8 64 23 10 100 -80
3 -7 49 33 20 400 -140
4 -6 36 43 30 900 -180
5 -5 25 53 40 1600 -200
6 -4 16 63 50 2500 -200
7 -3 9 73 60 3600 -180
8 -2 4 83 70 4900 -140
9 -1 1 93 80 6400 -80
10 0 0 103 90 8100 0
11 1 1 10.5 -2.5 6.25 -2.5
12 2 4 11 -2 4 -4
13 3 9 11.5 -1.5 2.25 -4.5
14 4 16 12 -1 1 -4
15 5 25 12.5 -0.5 0.25 -2.5
16 6 36 13 0 0 0
17 7 49 13.5 0.5 0.25 3.5
18 8 64 14 1 1 8
19 9 81 14.5 1.5 2.25 13.5
20 10 100 15 2 4 20
10 447.5 28521.25 -1172.5
1 1
∑uv − n Σu∑v −1172.5 − 20 (10)(447.5)
𝒓𝒓 = = = −0.39798
�Σ𝑢𝑢2 1 1
− n (∑u)2 �∑𝑣𝑣 2 − n (∑v)2 �670 − 1 (10)2 �28521.25 − 1 (447.5)2
20 20
What is Ranking?
A ranking is a relationship between a set of items such that, for any two items, the first is either 'ranked
higher than', 'ranked lower than' or 'ranked equal to' the second.
In mathematics, this is known as a weak order or total pre order of objects. It is not necessarily a total order
of objects because two different objects can have the same ranking. The rankings themselves are totally
ordered.
6∑𝐷𝐷 2
𝑅𝑅 = 1 −
𝑛𝑛(𝑛𝑛2 − 1)
Example 8. 10 competitors in a competition were ranked by two judges X and Y in the following manner:
Ranked by X 1 6 5 10 3 2 4 9 7 8
Ranked by Y 6 4 9 8 1 2 3 10 5 7
63
ISC XI- Applied Mathematics
Solution:
6∑𝐷𝐷 2 6×60
𝑅𝑅 = 1 − 𝑛𝑛(𝑛𝑛2−1) = 1 − 10(102−1) = 0.64 (positive, moderate correlation)
Example 9. The marks obtained by nine students in Accounts and Commerce are given below:
Accounts marks 35 23 47 17 10 43 9 6 28
Commerce marks 30 33 45 23 8 49 12 4 31
6∑𝐷𝐷 2 6×12
𝑅𝑅 = 1 − 𝑛𝑛(𝑛𝑛2−1) = 1 − 9(92−1) = 0.9 (positive, very high correlation)
This implies that a student scoring high marks in accounts is expected to score high marks in commerce as
well.
64
ISC XI- Applied Mathematics
1
Correction factor = CF = ∑ (𝑚𝑚3 − 𝑚𝑚) … …
12
6[∑𝐷𝐷2 + 𝐶𝐶𝐶𝐶]
𝑅𝑅 = 1 −
𝑛𝑛(𝑛𝑛2 − 1)
Example 10. Calculate the Rank Coefficient from the sales (in thousand units) and expenses (in thousand
Rupees) of 10 firms as given below:
Sales 50 50 55 60 65 65 65 60 60 50
Expenses 11 13 14 16 16 15 15 14 13 13
Solution:
Now,
1 1 1 1 1 1
CF = 12 (33 − 3) + 12
(33 − 3) +
12
(33 − 3) +
12
(23 − 2) +
12
(23 − 2) +
12
(23 − 2) +
1
(33 − 3) = 9.5
12
65
ISC XI- Applied Mathematics
Example 11. The coefficient of rank correlation of the marks obtained by 10 students in Mathematics and
Statistics was found to be 0.5. It was then detected that the difference in ranks in the two subjects for one
particular student was wrongly taken to be 3 in place of 7. What should be the correct rank correlation
coefficient?
Solution:
n = 10, R = 0.5
wrong value of D = 3
correct value of D = 7
6∑𝐷𝐷2 6∑𝐷𝐷2
𝑅𝑅 = 1 − ⇒ 0.5 = 1 − ⇒ ∑𝐷𝐷2 = 82.5
𝑛𝑛(𝑛𝑛2 − 1) 10(102 − 1)
66
ISC XI- Applied Mathematics
Exercises
1. Match List-I with List-II:
List I List-II
(A) If you water your plants then they grow (I) Correlation but not Causation
(B) Shoe size of an employee and his salary (II) Causation
(C) The people who smoke and alcoholism. (III) Neither correlation nor Causation
Choose the correct answer from the options given below:
(a) (A) - (I), (B) - (II), (C) - (III)
(b) (A) - (II), (B) - (III), (C) - (I)
(c) (A) - (II), (B) - (I), (C) - (III)
(d) (A) - (III), (B) - (II), (C) - (I)
67
ISC XI- Applied Mathematics
4. For each of the following statements, state whether it describes a correlation, causation, or neither:
(i) When people increase the number of hours they exercise, they tend to lose weight.
(ii) A survey finds that students who eat breakfast regularly score higher on math tests.
(iii) Turning up the volume on your speaker makes the sound louder.
(iv) People who own more books tend to have higher income.
(v) A new medicine is found to directly lower blood pressure in patients.
(vi) It is observed that students who wear red shirts tend to perform better in exams.
(vii) As ice cream sales increase, the number of drowning incidents also increases.
(viii) Scientists discover that smoking damages lung tissues and leads to cancer.
5. Each situation describes a correlation. Suggest whether causation is likely or not and explain briefly.
(i) Students who spend more time studying tend to get better grades.
(ii) People who sleep less than 5 hours a night often reports higher levels of stress.
(iii) Cities with more police officers also have higher crime rates.
(iv) People who own luxury cars are more likely to take vacations abroad.
(ii) If ∑𝑢𝑢𝑢𝑢 = 50 and n = 15 where u and v are deviations of X and Y series from their respective mean,
then Cov(X, Y) is
(a) 2.43 (c) 3.33
(b) 3.24 (d) 3.63
(iii) If the values of two variables move in opposite directions, their correlation is said to be
(a) linear (c) positive
(b) non-linear (d) negative
68
ISC XI- Applied Mathematics
(v) Assertion: A perfect positive correlation means that all data points lie exactly on a straight line
sloping upward.
Reason: When r = 1, the values of the two variables increase in exact proportion to each other.
(a) Both A and R individually true and R is the correct explanation of A.
(b) Both A and R individually true but R is not the correct explanation of A.
(c) A is true but R is false.
(d) A is false but R is true.
7. Can you directly conclude that drinking more coffee causes better productivity just because they are
correlated? What else would you need to check?
8. Scientists found a correlation between kids who play a lot of video games and better hand-eye
coordination.
(i) Does this necessarily prove that playing video games causes better hand-eye coordination?
(ii) Suggest a way to find out if causation is true.
9. A study shows a correlation between eating more green vegetables and a lower risk of heart disease. How
could researchers test if this is truly causation?
12. Calculate the coefficient of correlation by Karl Pearson’s method from the following data relating to
overhead expenses and cost of production.
Overhead in ’000 Rs 80 90 100 110 120 130 140 150 160
Cost in ’000 Rs 15 15 16 19 17 18 16 11 19
13. Calculate the Pearson coefficient of correlation between advertisement cost and sales as per the data
given below:
Advertisement cost 39 65 62 90 82 75 25 98 36 78
in ’000 Rs
Sales in lakhs Rs 47 53 58 86 62 68 60 91 51 84
14. A computer while calculating the correlation coefficient between two variables x and y obtained the
following results:
𝑛𝑛 = 30, ∑𝑥𝑥 = 120, ∑𝑦𝑦 = 90, ∑𝑥𝑥 2 = 600, ∑𝑦𝑦 2 = 250, ∑𝑥𝑥𝑥𝑥 = 356
It was, however, later discovered that two pairs of observations (8,10) and (12,7) had been wrongly
entered. The correct values were (8,12) and (10,8). Obtain the correct value of the correlation coefficient
between x and y.
15. Calculate the rank correlation coefficient for the following data:
x 92 89 87 86 83 77 71 63 53 50
y 86 83 91 77 68 85 52 82 37 57
69
ISC XI- Applied Mathematics
16. In a contest, two judges ranked eight candidates in order of their preference as given below. Find the
rank correlation coefficient.
Judge 1 5 2 8 1 4 6 3 7
Judge 2 4 5 7 3 2 8 1 6
18. The coefficient of rank correlation between the marks in Economics and Statistics obtained by a group
9
of students is and the sum of squares of the difference in rank is 30. Find the number of students in
11
the group.
Answers
1. (b) (A) - (II), (B) - (III), (C) - (I)
2. (c)Statement 1 only correlation and Statement 2 is only causation
3. (c)Statement 1 is correlation, also causation and Statement 2 is neither correlation nor causation.
4. (i) Correlation — Exercise and weight are related, but other factors may influence
weight loss.
(ii) Correlation — Breakfast and better scores are linked, but no proof breakfast causes better scores.
(iii) Causation — Direct cause-and-effect: turning up volume causes louder sound.
(iv) Correlation — Book ownership and income are related, but one doesn't
necessarily cause the other.
(v) Causation — Medicine directly causes a drop in blood pressure (assuming a properly done clinical
trial).
(vi) Correlation — Wearing red shirts and exam performance are just observed together; no proven
cause.
(vii) Correlation — Ice cream and drowning rise together (due to summer), but one does not cause the
other.
(viii) Causation — Smoking has been proven to cause lung cancer scientifically.
5. (i) Causation is likely — More studying probably helps grades, but need controlled
experiments to be sure.
(ii) Causation is possible — Lack of sleep may cause stress, but stress might also cause poor sleep;
more study needed.
(iii) Correlation only — High crime may lead to hiring more police; not that police cause crime.
(iv) Correlation only — Wealth might cause both luxury car ownership and more
vacations — a third factor (wealth).
6. (i) (c) change of scale and origin both
(ii) (c) 3.33
(iii) (d) negative
(iv) (b) only 2
(v) (a) Both A and R individually true and R is the correct explanation of A
70
ISC XI- Applied Mathematics
7. No, you cannot directly conclude causation. You would need to check for:
• Other factors (like overall health, motivation, job type).
• Experimental studies or control groups.
8. (a) No, correlation does not prove causation.
(b) Conduct an experiment — randomly assign some kids to play video games and compare hand-eye
coordination over time.
9. Researchers could do a long-term controlled study, tracking people who eat more vs. fewer green
vegetables, while controlling other lifestyle factors like exercise, smoking, genetics, etc.
(Or conduct a randomized controlled trial, if ethical.)
10. (i) -0.5 (ii) 0.8
11. 5
12. 0.055
13. 0.78
14. 0.05
15. 0.73
16. 0.67
17. - 0.024
18. 10
Classroom Discussion
Give one example of your own where two things are correlated but not causally related. Explain the hidden
third factor that could cause both.
Project Work
Project Title: "Does Study Time Affect Academic Performance?"
Objective:
To investigate whether there is a statistical relationship between the number of hours students spend studying
and their academic performance (measured by test scores).
Data Collection:
• Participants: 30 students from the same grade or class.
• Variables:
− Independent Variable: Number of hours studied per week.
− Dependent Variable: Scores obtained in the most recent mathematics test.
• Tool: Use a questionnaire or Google Form to collect data.
Steps:
1. Collect Data:
Ask each student:
− How many hours they studied in the last week before the test.
− Their score on the test (out of 100).
71
ISC XI- Applied Mathematics
HISTORICAL NOTE
• Auguste Bravais's work on the bivariate normal distribution in the mid-19th century laid a foundation
for understanding correlation.
• Francis Galton is credited with formally defining the concept of correlation in 1885. He defined
correlation as the relationship between two variables, where a change in one is associated with a
change in the other.
• Karl Pearson developed the Pearson correlation coefficient, a widely used statistical measure that
quantifies the strength and direction of a linear relationship between two variables.
• Charles Spearman introduced his rank correlation coefficient in [Link] method was particularly
useful when data did not meet the assumptions of parametric tests.
72
ISC XI- Applied Mathematics
SCOPE
LEARNING OUTCOMES
PRE-REQUISITES
73
ISC XI- Applied Mathematics
Introduction
The dictionary meaning of the word regression is ‘stepping or going back’. The use of this word dates to the
times of early studies made by Francis Galton in the later half of the nineteenth century. He studied the
relationship between the heights of fathers and sons and arrived at the following conclusions:
(i) Tall fathers have tall sons and short fathers have short sons.
(ii) If fathers are tall, mean height of the sons < mean height of fathers.
(iii) If fathers are short, mean height of the sons > mean height of fathers.
(iv) Deviations in the mean height of sons from the mean height of the race < Deviations in the mean
height of fathers from the mean height of the race.
So, he concluded that when the fathers move above the mean or below the mean, the sons tend to go back or
regress towards the mean.
Galton studied the average relationship between two variables graphically and called the line describing the
relationship the line of regression.
What is Regression?
When two or more variables are correlated, a
statistical method to formulate an algebraic
relationship between two or more variables in the
form of equation to estimate the value of one
variable when the values of all other variables are
known or specified, is called Regression.
The variable whose value is predicted is called
response or dependent variable. The variables
used for prediction the response variable or
dependent variable is called predictors or
independent variables.
A regression shows the changes in a dependent
variable on the y-axis to the changes in the
independent variable on the x-axis. Basically,
regression is the task of predicting a continuous
quantity.
74
ISC XI- Applied Mathematics
Correlation Regression
Regression is used to numerically describe how a
Correlation is used to determine whether variables
dependent variable changes with a change in an
are related or not.
independent variable
Correlation tries to establish a linear relationship It finds the best-fitted regression line to estimate an
between variables. unknown variable on the basis of the known variable.
In our scope of study, we will be dealing only with Linear Regression. The main purpose of curve fitting is
to predict or estimate the value of the dependent variable from the value of the independent variable. The
regression line which is used to predict the values of y for the given values of x is called the regression line of
y on x and the regression line which is used to predict the values of x for the given values of y is called the
regression line of x on y.
75
ISC XI- Applied Mathematics
Regression lines can be drawn by free hand curve method but it is a very cumbersome and difficult method.
It is also dependent on individual judgement and different observers will obtain different curves and equations.
If the line of regression is so chosen that the sum of squares of deviations parallel to the x-axis is minimised,
it is called the regression line of x on y. Its equation is 𝑥𝑥 = 𝑐𝑐 + 𝑑𝑑𝑑𝑑 and b is called the regression coefficient
of x on y and is represented as 𝑏𝑏𝑥𝑥𝑥𝑥 . (Fig. B)
Fig . A Fig. B
For y on x, the equation of least square line is given by 𝑦𝑦 = 𝑎𝑎 + 𝑏𝑏𝑏𝑏. The corresponding normal equations
are:
∑𝑦𝑦 = 𝑛𝑛𝑛𝑛 + 𝑏𝑏∑𝑥𝑥
∑𝑥𝑥𝑥𝑥 = 𝑎𝑎∑𝑥𝑥 + 𝑏𝑏∑𝑥𝑥 2
Solving these two normal equations we can get the required trend line equation.
76
ISC XI- Applied Mathematics
Thus, we can get the line of best fit with formula 𝑦𝑦 = 𝑎𝑎 + 𝑏𝑏𝑏𝑏.
For x on y, the equation of least square line is given by 𝑥𝑥 = 𝑐𝑐 + 𝑑𝑑𝑑𝑑. The corresponding normal equations
are:
∑𝑥𝑥 = 𝑛𝑛𝑛𝑛 + 𝑑𝑑∑𝑦𝑦
∑𝑥𝑥𝑥𝑥 = 𝑐𝑐∑𝑦𝑦 + 𝑑𝑑∑𝑦𝑦 2
Solving these two normal equations we can get the required trend line equation.
Thus, we can get the line of best fit with formula 𝑥𝑥 = 𝑐𝑐 + 𝑑𝑑𝑑𝑑.
Example 1. Fit the regression lines of y on x and x on y for the following data.
x 2 4 6 8 10
y 5 7 9 8 11
Solution:
x y x2 y2 xy
2 5 4 25 10
4 7 16 49 28
6 9 36 81 54
8 8 64 64 64
10 11 100 121 110
∑𝑥𝑥 =30 ∑𝑦𝑦 =40 ∑𝑥𝑥 2 =220 ∑𝑦𝑦 2 =340 ∑𝑥𝑥𝑥𝑥 = 266
77
ISC XI- Applied Mathematics
Regression Coefficients
Solving the normal equations,
Since, 𝑏𝑏𝑦𝑦𝑦𝑦 = 𝑏𝑏 the regression equation of y on x can be written as 𝑦𝑦 − 𝑦𝑦� = 𝑏𝑏𝑦𝑦𝑦𝑦 (𝑥𝑥 − 𝑥𝑥̅ ).
Since, 𝑏𝑏𝑥𝑥𝑥𝑥 = 𝑏𝑏 the regression equation of x on y can be written as 𝑥𝑥 − 𝑥𝑥̅ = 𝑏𝑏𝑥𝑥𝑥𝑥 (𝑦𝑦 − 𝑦𝑦�)
Let 𝑑𝑑𝑥𝑥 = 𝑥𝑥 − 𝑥𝑥̅ and 𝑑𝑑𝑦𝑦 = 𝑦𝑦 − 𝑦𝑦� be the deviations of the variables x and y from the actual mean, then
∑𝑑𝑑𝑥𝑥 𝑑𝑑𝑦𝑦 = ∑(𝑥𝑥 − 𝑥𝑥̅ )(𝑦𝑦 − 𝑦𝑦�) = ∑𝑥𝑥𝑥𝑥 − 𝑥𝑥̅ (𝑦𝑦1 +𝑦𝑦2 + ⋯ + 𝑦𝑦𝑛𝑛 ) − 𝑦𝑦�(𝑥𝑥1 + 𝑥𝑥2 + ⋯ + 𝑥𝑥𝑛𝑛 ) + 𝑛𝑛𝑥𝑥̅ 𝑦𝑦�
= (∑𝑥𝑥𝑥𝑥 + 𝑛𝑛𝑥𝑥̅ 𝑦𝑦�) − 𝑥𝑥.
� 𝑛𝑛𝑦𝑦� − 𝑦𝑦.
� 𝑛𝑛𝑥𝑥̅
= ∑𝑥𝑥𝑥𝑥 + 𝑛𝑛𝑥𝑥̅ 𝑦𝑦� − 2𝑛𝑛𝑥𝑥̅ 𝑦𝑦� = ∑𝑥𝑥𝑥𝑥 − 𝑛𝑛𝑥𝑥̅ 𝑦𝑦�
∑𝑥𝑥 ∑𝑦𝑦 ∑𝑥𝑥∑𝑦𝑦
= ∑𝑥𝑥𝑥𝑥 − 𝑛𝑛. ⋅ = ∑𝑥𝑥𝑥𝑥 −
𝑛𝑛 𝑛𝑛 𝑛𝑛
(∑𝑥𝑥)2 (∑𝑦𝑦)2
Also, 𝑛𝑛𝜎𝜎𝑥𝑥2 = ∑𝑥𝑥 2 − and 𝑛𝑛𝜎𝜎𝑦𝑦2 = ∑𝑦𝑦 2 −
𝑛𝑛 𝑛𝑛
78
ISC XI- Applied Mathematics
3. If two lines of regression, when plotted on a graph paper, coincide, there is a perfect correlation between
the two series.
When 𝒓𝒓 = ±𝟏𝟏, 𝒕𝒕𝒕𝒕𝒕𝒕 𝜽𝜽 = 𝟎𝟎. This implies θ is either 0 or π. Hence, the lines of regression coincide and
there is perfect correlation between the variables x and y.
4. Greater the angle between the regression lines, lesser is the correlation between the two series.
79
ISC XI- Applied Mathematics
5. If the two regression lines cut each other at right angles there is no
correlation between them.
𝝅𝝅
When 𝒓𝒓 = 𝟎𝟎, 𝜽𝜽 = . Hence, the estimated value of y will be the same for
𝟐𝟐
all values of x and vice- versa.
𝑟𝑟 = �𝑏𝑏𝑦𝑦𝑦𝑦 . 𝑏𝑏𝑥𝑥𝑥𝑥
2. If one of the regression coefficients is greater than unity, then the other must be less than unity numerically.
0 ≤ 𝑏𝑏𝑦𝑦𝑦𝑦 ⋅ 𝑏𝑏𝑥𝑥𝑥𝑥 ≤ 1
3. Arithmetic mean of regression coefficients is greater than the correlation coefficient.
𝑏𝑏𝑦𝑦𝑦𝑦 + 𝑏𝑏𝑥𝑥𝑥𝑥 > 2𝑟𝑟
4. The correlation coefficient and the two regression coefficients have the same sign.
5. Regression coefficients are independent of the change in origin but are dependent on the change in scale.
80
ISC XI- Applied Mathematics
Example 2. From the following data find the two regression lines and predict the value of y when x= 2.5
x 1 2 3 4 5
y 2 3 5 4 6
Solution:
Method 1:
x y x2 y2 xy
1 2 1 4 2
2 3 4 9 6
3 5 9 25 15
4 4 16 16 16
5 6 25 36 30
15 20 55 90 69
∑𝑥𝑥 15 ∑𝑦𝑦 20
𝑥𝑥̅ = = = 3, 𝑦𝑦� = = =4
𝑛𝑛 5 𝑛𝑛 5
1 1
∑𝑥𝑥𝑥𝑥 − 𝑛𝑛 ∑𝑥𝑥∑𝑦𝑦 69 − 5 (15)(20)
𝑏𝑏𝑦𝑦𝑦𝑦 = = = 0.9
1 1
∑𝑥𝑥 − (∑𝑥𝑥)
2 2 55 − (15) 2
𝑛𝑛 5
1 1
∑𝑥𝑥𝑥𝑥 − 𝑛𝑛 ∑𝑥𝑥∑𝑦𝑦 69 − 5 (15)(20)
𝑏𝑏𝑥𝑥𝑥𝑥 = = = 0.9
1 1
∑𝑦𝑦 2 − 𝑛𝑛 (∑𝑦𝑦)2 90 − (20)2
5
Method 2:
∑𝑥𝑥 15 ∑𝑦𝑦 20
𝑥𝑥̅ = = = 3, 𝑦𝑦� = = =4
𝑛𝑛 5 𝑛𝑛 5
∑ 𝑑𝑑𝑥𝑥 𝑑𝑑𝑦𝑦 ∑ 𝑑𝑑𝑥𝑥 𝑑𝑑𝑦𝑦
The actual mean are integer values, hence we use 𝑏𝑏𝑦𝑦𝑦𝑦 = and 𝑏𝑏𝑥𝑥𝑥𝑥 = .
∑ 𝑑𝑑𝑥𝑥2 2
∑ 𝑑𝑑𝑦𝑦
∑ 𝑑𝑑𝑥𝑥 𝑑𝑑𝑦𝑦 9
𝑏𝑏𝑦𝑦𝑦𝑦 = == 0.9
∑ 𝑑𝑑𝑥𝑥2 10
∑ 𝑑𝑑𝑥𝑥 𝑑𝑑𝑦𝑦 9
𝑏𝑏𝑥𝑥𝑥𝑥 = 2 = 10 = 0.9
∑ 𝑑𝑑𝑦𝑦
The regression equation of y on x is 𝑦𝑦 − 𝑦𝑦� = 𝑏𝑏𝑦𝑦𝑦𝑦 (𝑥𝑥 − 𝑥𝑥̅ ) ⇒ 𝑦𝑦 − 4 = 0.9(𝑥𝑥 − 3)
⇒ 𝑦𝑦 = 0.9𝑥𝑥 + 1.3
The regression equation of x on y is 𝑥𝑥 − 𝑥𝑥̅ = 𝑏𝑏𝑥𝑥𝑥𝑥 (𝑦𝑦 − 𝑦𝑦�) ⇒ 𝑥𝑥 − 3 = 0.9(𝑦𝑦 − 4)
⇒ 𝑥𝑥 = 0.9𝑦𝑦 − 0.6
When x = 2.5, from 𝑦𝑦 = 0.9𝑥𝑥 + 1.3 we get y = 3.55
81
ISC XI- Applied Mathematics
Example 3. The grade of 9 students for the half yearly and the annual examination in Mathematics are as
follows:
x 77 50 71 72 81 94 96 99 67
y 82 66 78 84 47 85 99 99 68
Find a suitable linear regression equation for the data and then estimate the grade in the annual examination
for a student who received 85 in half yearly but was sick at the time of annual examinations and hence was
absent.
Solution:
x y u = x -72 v = y - 78 𝒖𝒖𝟐𝟐 uv
77 82 5 4 25 20
50 66 - 22 -12 484 264
71 78 -1 0 1 0
72 84 0 6 0 0
81 47 9 - 31 81 -279
94 85 22 7 484 154
96 99 24 21 576 504
99 99 27 21 729 567
67 68 -5 -10 25 50
707 708 59 6 2405 1280
∑x 707 ∑y 708
x� = n
= = 78.56, y� = n = 9 = 78.67
9
1 1
∑𝑢𝑢𝑢𝑢 − 𝑛𝑛 ∑𝑢𝑢∑𝑣𝑣 1280 − 9 (59)(6)
𝑏𝑏𝑦𝑦𝑦𝑦 = = = 0.615
1 1
∑𝑢𝑢2 − 𝑛𝑛 (∑𝑢𝑢)2 2405 − 𝑛𝑛 (59)2
The regression equation of y on x is 𝑦𝑦 − 𝑦𝑦� = 𝑏𝑏𝑦𝑦𝑦𝑦 (𝑥𝑥 − 𝑥𝑥̅ ) ⇒ 𝑦𝑦 − 78.67 = 0.615(𝑥𝑥 − 78.56)
⇒ 𝑦𝑦 = 0.615𝑥𝑥 + 30.36
When x = 85, we get y = 82.635≈ 83.
His Mathematics grade in the annual examination will be 83(approx.)
Example 4. The coefficient of correlation between the ages of husbands and wives in a community was found
to be 0.8; the mean of the husband’s age was 25 years and that of the wives 22 years. Their standard deviations
were 4 and 5 years respectively. Find the two lines of regression and hence obtain:
(a) Husband’s age when wife’s age is 22.
(b) Wife’s age when husband’s age is 33.
Solution: r = 0 ⋅ 8, x� = 25, y
� = 22, σx = 4, σy = 5
𝜎𝜎𝑦𝑦 5
𝑏𝑏𝑦𝑦𝑦𝑦 = 𝑟𝑟 = 0.8 × = 1
𝜎𝜎𝑥𝑥 4
𝜎𝜎𝑥𝑥 4
𝑏𝑏𝑥𝑥𝑥𝑥 = 𝑟𝑟 = 0.8 × = 0.64
𝜎𝜎𝑦𝑦 5
82
ISC XI- Applied Mathematics
⇒ 𝑥𝑥 = 0.64𝑦𝑦 + 10.92
When y = 32, we get x = 25 years.
(a) When wife’s age is 22, husband’s age is 25 years.
(b) When husband’s age is 33, wife’s age is 30 years.
Example 5. The following results were obtained from the record of age(x) and blood pressure(y) of a group
of 10 women.
x y
Mean 53 142
Variance 130 165
Given 𝛴𝛴(𝑥𝑥 − 𝑥𝑥̅ )(𝑦𝑦 − 𝑦𝑦�) = 1220, find the regression equation of y on x and use it to estimate the blood
pressure of a 47-year-old woman.
Solution:
1 1
∑(𝑥𝑥 − 𝑥𝑥̅ )(𝑦𝑦 − 𝑦𝑦�) 𝑛𝑛 ∑(𝑥𝑥 − 𝑥𝑥̅ )(𝑦𝑦 − 𝑦𝑦�) 10 × 1220
𝑏𝑏𝑦𝑦𝑦𝑦 = = = = 0 ⋅ 938
∑(𝑥𝑥 − 𝑥𝑥̅ )2 1 130
) 2
𝑛𝑛 ∑(𝑥𝑥 − 𝑥𝑥̅
Example 6. If two regression lines of a bivariate distribution are 4𝑥𝑥 − 5𝑦𝑦 + 33 = 0 and 20𝑥𝑥 −
9𝑦𝑦 − 107 = 0, calculate
(i) 𝑥𝑥̅ and 𝑦𝑦� , the arithmetic means of x and y respectively.
(ii) Estimate the value of x when y = 7.
(iii) Find the variance of y when 𝜎𝜎𝑥𝑥 = 3.
Solution:
4 9 3
𝑟𝑟 = +�𝑏𝑏𝑦𝑦𝑦𝑦 ⋅ 𝑏𝑏𝑥𝑥𝑥𝑥 = +� × =+
5 20 5
3 3 3
𝑟𝑟 = −�𝑏𝑏𝑦𝑦𝑦𝑦 ⋅ 𝑏𝑏𝑦𝑦𝑦𝑦 = −��− � × �− � = −
4 4 4
84
ISC XI- Applied Mathematics
Exercises
1. Choose the correct answer.
(ii) Which of the below statements concerning two regression coefficients’ arithmetic mean is correct?
(a) It is greater than the correlation (c) It is equal to or greater than the
coefficient. correlation coefficient.
(b) It is less than that of the correlation (d) It’s equal to the correlation coefficient.
coefficient.
85
ISC XI- Applied Mathematics
(iii) What is the sum of squares of error if all the actual and estimated values of Y on the regression line
are the same?
(a) Maximum (c) Zero
(b) Minimum (d) Unknown
(iv) Which of the following statement(s) is/are correct with respect to regression coefficients?
Statement 1. It measures the degree of linear relationship between two variables.
Statement 2. It gives the value by which one variable changes for a unit change in the other variable.
(a) Only 1 (c) Both 1 and 2
(b) Only 2 (d) Neither 1 nor 2
2. The two lines of regressions are x + 2y - 5 = 0 and 2x + 3y - 8 = 0 and the variance of x is 12. Find the
variance of y and the coefficient of correlation.
3. The following table shows the sales and advertisement expenditure of a firm.
The correlation coefficient is 0.9. Estimate the likely sales for a proposed expenditure of Rs 10 crores.
x 1 2 3 4 5
y 7 6 5 4 3
5. The following observations are given: (1,4), (2,8), (3,2), (4,12), (5,10), (6,14), (7,16), (8,6), (9,18).
Estimate the values of
(i) y when the value of x is 10.
(ii) x when the value of y is 5.
6. The following data represents the number of hours 12 different students used mobile phones during the
weekend and the scores of each student who took a test the following Monday.
(i) Find the equation of the suitable regression line.
(ii) Use the equation to find the expected test score for a student who used the mobile phone for 9
hours.
Hours 0 1 2 3 3 5 5 5 6 7 7 10
Test score 96 85 82 74 95 68 76 84 58 65 75 50
86
ISC XI- Applied Mathematics
7. Calculate the two regression equations of x on y and y on x from the data given below, taking deviations
from actual means of x and y.
Price (₹) 10 12 13 12 16 15
Demand(kg) 40 38 43 45 37 43
x 40 50 38 60 65 50 35
y 38 60 55 70 60 48 30
9. Find the means of x and y and the coefficient of correlation between them from the following two
regression equations:
2y – x – 50 = 0
3y – 2x – 10 = 0.
x 8 3 2 10 11 3 6 5 6 8
y 4 12 1 12 9 4 9 6 1 14
Use the least square method to determine the equation of line of best fit (y on x) for the data.
Answers
1 (i) (c) Regression coefficient of y on x
(ii) (a) It is greater than the correlation coefficient.
(iii) (c) Zero
(iv) (b) Only 2
(v) (a) Both A and R individually true and R is the correct explanation of A.
2. Variance of y = 4, r = - 0.866
3. ₹ 64 crores.
4. y+ x = 8, y = 2
50 10
5. y = , x=
3 3
87
ISC XI- Applied Mathematics
Project Work
Project Title: "Predicting Monthly Electricity Bill Based on Number of Units Consumed"
Objective:
To apply regression analysis to predict the monthly electricity bill of a household based on the number of
electricity units consumed, using real data.
Concepts Applied:
• Regression line (linear regression)
• Scatter diagram
• Correlation coefficient
• Regression coefficients
• Equation of regression line
Real-Life Context:
Electricity bills usually depend on the number of units used. Understanding this relationship helps families
predict bills, budget wisely, or adopt energy-saving strategies.
Data Collection:
Collect data from your or a few households over 6–12 months. Record:
88
ISC XI- Applied Mathematics
HISTORICAL NOTE
• Early Roots: Isaac Newton's work in 1700 on equinoxes is considered an early form of regression
analysis, including averaging data and forcing the regression line through the average point.
• Method of Least Squares: Legendre (1805) and Gauss (1809) independently developed the method
of least squares, which is fundamental to modern regression analysis. They applied it to astronomical
observations, like determining orbits of celestial bodies.
• Francis Galton's Contribution: Galton, a cousin of Charles Darwin, studied heredity, particularly
the inheritance of height. He noticed that tall parents tended to have somewhat shorter offspring, and
vice versa. He coined the term "regression" to describe this "reversion" towards the mean.
• Evolution of Regression: Galton's work laid the groundwork for understanding the relationship
between variables and predicting one variable based on another. Since then, numerous researchers
have expanded on these concepts.
89
ISC XI- Applied Mathematics
(v) Probability
SCOPE
Random experiments; outcomes, sample spaces (set representation). Events; occurrence of events, 'not',
'and' and 'or' events, exhaustive events, mutually exclusive events, Axiomatic (set theoretic) probability,
connections with other theories studied in earlier classes. Probability of an event, probability of' not',
'and' and 'or' events.
• Random experiments and their outcomes.
• Events: sure events, impossible events, mutually exclusive and exhaustive events.
- Definition of probability of an event
- Laws of probability addition theorem.
LEARNING OUTCOMES
PRE-REQUISITES
Introduction
In common parlance the term Probability refers to the chance of happening or not happening of an event. The
use of the word chance in any statement indicates that there is an element of uncertainty about that statement.
For example, consider the statements,
(a) The probability that it will rain tomorrow is low.
(b) The chance that the prices of Company A will go up in the next financial year is very high.
There is an element of uncertainty associated with each one of the above statements as there is no numerical
value assigned to these statements.
Probability provides a numerical measure of the element of uncertainty. It enables us to take decisions under
conditions of uncertainty with a calculated risk.
90
ISC XI- Applied Mathematics
Random Experiment
It is an experiment which if conducted repeatedly under homogeneous conditions does not give the same
result. It has more than one possible outcome and it is not possible to predict the outcome in advance. The
result maybe any one of the various possible outcomes. For example, if an unbiased die is thrown any one of
the six numbers can appear as an outcome.
Sample space
This is a set of all possible outcomes of a random
experiment. It is also called event space or
probability set. It is denoted by S. Each outcome is
called an element or a sample point of the sample
space.
Events
Events could be either simple/elementary or compound/composite. An event is called simple if it corresponds
to a single possible outcome. Thus, in throwing a die, the chance of getting a 3 is a simple event. However,
the chance of getting an odd number is a compound event.
A compound event can be further decomposed into simple events. So, for the above example for compound
event, the decomposed simple events are getting either 1 or 3 or 5 on throwing the die.
Exhaustive Cases: All possible outcomes of an event are known as exhaustive cases. So, in the case of a
throwing a die there are 6 exhaustive cases but if two dice are thrown there are 36 exhaustive cases to consider.
Favourable Cases: The number of outcomes which result in the occurrence of a desired event are called
favourable cases. In drawing a card from a pack, the cases favourable to getting a spade are 13.
Mutually Exclusive Events: Two or more events are said to be mutually exclusive if the occurrence of one
of them excludes the occurrence of the others. So, when tossing a coin getting a head or a tail are mutually
exclusive events.
Equally likely Events: Two or more events are said to be equally likely if their chances of happening are
equal. For the tossing of a coin, it is equally likely to get a head or a tail as a result.
91
ISC XI- Applied Mathematics
Not Equally likely Events: Not equally likely events are those where each possible
outcome of an experiment does not have the same chances of occurring. For example, a
spinner with unequal sections will have different probabilities for landing on each
colour.
Sure Event: Since the set S is a subset of itself, it represents an event. Since all the
outcomes of a random experiment belong to its sample space S, the event S represents
the sure event.
Impossible Event: Since the null set is also a subset of S, therefore, it also represents an event. As no outcomes
of the experiment belongs to 𝜙𝜙, the event 𝜙𝜙 can never occur when the experiment is performed. For this
reason, the event 𝜙𝜙 is called the impossible event.
Complementary Event: Let S be the sample space associated with a random experiment and let E be an
event. The event of non-occurrence of E is called the complementary event of E. It is represented by 𝐸𝐸 ′ or 𝐸𝐸� .
We also call this event as not E.
Notations
Sample space S
Event A has occurred A
92
ISC XI- Applied Mathematics
What is Probability?
The probability of an event denotes the likelihood of its occurrence. The value of probability ranges from 0 to
1. If an event is certain to happen its probability would be 1 and it is certain that it will not take place its
probability would be 0.
𝑛𝑛(𝐴𝐴)
For an event A we can write this as, 𝑃𝑃(𝐴𝐴) = ,
𝑛𝑛(𝑆𝑆)
where,
𝑃𝑃(𝐴𝐴) is the probability of an event A,
𝑛𝑛(𝐴𝐴) = the number of favourable outcomes for event A,
𝑛𝑛(𝑆𝑆) is the total number of outcomes in the sample space.
93
ISC XI- Applied Mathematics
Do you know?
Artificial intelligence (AI) systems rely heavily on probability. When we type a search query in Google,
it predicts what we might type next. The next word is not equally likely. Rather it depends on past
searches and user behaviour. Classical probability cannot be used because all words are not equally
likely. AI systems use probability models to calculate the likelihood of different words.
𝑘𝑘
𝑖𝑖 𝑘𝑘 𝑁𝑁𝑁𝑁𝑁𝑁𝑁𝑁𝑁𝑁𝑁𝑁 𝑜𝑜𝑜𝑜 𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠 𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝 𝑖𝑖𝑖𝑖 𝐴𝐴 𝑁𝑁𝑁𝑁𝑁𝑁𝑁𝑁𝑁𝑁𝑁𝑁 𝑜𝑜𝑜𝑜 𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐 𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓 𝑡𝑡𝑡𝑡 𝐴𝐴 𝑛𝑛(𝐴𝐴)
𝑃𝑃(𝐴𝐴) = � = = = =
𝑁𝑁 𝑁𝑁 𝑁𝑁𝑁𝑁𝑁𝑁𝑁𝑁𝑁𝑁𝑁𝑁 𝑜𝑜𝑜𝑜 𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠 𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝 𝑖𝑖𝑖𝑖 𝑆𝑆 Number of sample points in S 𝑛𝑛(𝑆𝑆)
𝑖𝑖=1
Probability Vs Odds
Odds represent the likelihood of an event happening compared to it not happening, while probability is the
chance of an event occurring out of all possible outcomes. Odds are expressed as a ratio, while probability is
usually a fraction or percentage.
Given the probability of an event ‘E’, i.e., P(E),
𝑃𝑃(𝐸𝐸)
Odds for 𝐸𝐸 =
1−𝑃𝑃(𝐸𝐸)
1−𝑃𝑃(𝐸𝐸)
Odds against 𝐸𝐸 = 𝑃𝑃(𝐸𝐸)
94
ISC XI- Applied Mathematics
Usage Used in statistics and probability Often used in betting, gambling, and
theory. risk assessments.
Interpretation Represents how likely an event will Compares the occurrence of an event to
happen based on a ratio to all possible the non-occurrence.
outcomes.
Example 1. A spinner with three equal sectors coloured red, green and blue is spun
three times.
(i) Represent the sample space on a tree diagram.
Use the tree diagram to work out the probability of getting:
(ii) the same colour on all three spins.
(iii) a green and two reds in any order.
(iv) a different colour on all three spins.
Solution:
(i) Total number of outcomes = 27
95
ISC XI- Applied Mathematics
Example 2. A fair six-sided die is rolled once. Find the probability of getting:
(i) a multiple of 3
(ii) a number greater than 4
(iii) an odd number
Solution:
The sample space 𝑆𝑆 = {1,2,3,4,5,6}
Total number of outcomes = 6
(i) Multiples of 3 = {3,6} ∴ 2 𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓 𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜
2 1
P (Multiple of 3) = =
6 3
(ii) Numbers greater than 4 = {5,6} ∴ 2 𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓 𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜
2 1
P (Number > 4) = =
6 3
(iii) Odd numbers = {1,3,5} ∴ 3 𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓 𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜𝑜
3 1
P (Odd number) = =
6 2
96
ISC XI- Applied Mathematics
Example 3. A card is drawn randomly from a standard deck of 52 playing cards. Find the probability that
the card is:
(i) A king
(ii) A red card
(iii) A face card (King, Queen, Jack)
Solution:
(i) There are 4 Kings in a deck.
4 1
P(King)= = 13
52
(ii) There are 26 red cards (13 hearts + 13 diamonds).
26 1
P (Red card) = =2
52
(iii) Face cards: 12 total (4 Kings, 4 Queens, 4 Jacks)
12 3
P (Face card) = = 13
52
Example 4. Sahil makes a game involving tossing three coins and noting the sequence of heads and tails. If
more heads than tails appear, he wins. If more tails than heads appear, you win. Is this game fair?
Solution:
To see if the game is fair, we calculate Sahil's chance of winning and your chance of winning. For
the experiment of tossing three coins, the sample space is
{𝐻𝐻𝐻𝐻𝐻𝐻, 𝐻𝐻𝐻𝐻𝐻𝐻, 𝐻𝐻𝐻𝐻𝐻𝐻, 𝑇𝑇𝑇𝑇𝑇𝑇, 𝐻𝐻𝐻𝐻𝐻𝐻, 𝑇𝑇𝑇𝑇𝑇𝑇, 𝑇𝑇𝑇𝑇𝑇𝑇, 𝑇𝑇𝑇𝑇𝑇𝑇}.
4 1
P (Sahil winning) = =2
8
4 1
P (you winning) = =2
8
This is a fair game.
Example 5. Consider the experiment of rolling a pair of dice. Let A be the event that the numbers on the dice
have a sum of 8. Let B be the event that the two numbers on the dice are a 3 and a 5.
(i) What is P(A)?
(ii) What is P(B)?
Solution:
The sample space for this experiment has 36 outcomes.
𝑆𝑆 = {(1,1), (1,2), (1,3), (1,4), (1,5), (1,6), (2,1), (2,2), (2,3), (2,4), (2,5), (2,6),
(3,1), (3,2), (3,3), (3,4), (3,5), (3,6), (4,1), (4,2), (4,3), (4,4), (4,5), (4,6),
(5,1), (5,2), (5,3), (5,4), (5,5), (5,6), (6,1), (6,2), (6,3), (6,4), (6,5), (6,6)}
5
(i) There are 5 pairs of numbers that have a sum of 8, so 𝑃𝑃(𝐴𝐴) =
36
2 1
(ii) There are 2 pairs of numbers that are a 3 and a 5, so 𝑃𝑃(𝐵𝐵 ) = =
36 18
97
ISC XI- Applied Mathematics
Example 6. Find the probability that, in a random arrangement of the letters of "SCHOOL", the two O’s
appear together.
Solution:
Total letters = 6
6! 720
Since there are two O’s, the total permutations = = = 360
2! 2
Treating "OO" as a single unit, remaining letters = 5
Total favourable arrangements = 5! = 120
120 1
P (O’s together) = =3
360
Example 7. Two cards are drawn at random from a well-shuffled deck of 52 cards. Find the probability that
both are aces.
Solution:
52! 52×51
Total ways to choose 2 cards from 52 = 𝐶𝐶(52,2) = = = 1326
[2!(52−2)!] 2
4! 4×3
Ways to choose 2 aces from 4 aces = 𝐶𝐶(4,2) = = =6
2!(4−2)! 2
6 1
P (Both Aces) = = 221
1326
Example 8. Two dice are thrown together. Find the probability of getting a sum of at least 10.
Solution:
Total outcomes = 6 × 6 = 36
Cases when sum ≥ 10 = {(4,6), (5,5), (6,4), (5,6), (6,5), (6,6)} ∴ 6 favourable cases
6 1
P (Sum ≥ 10) = =
36 6
Example 9. A committee of 3 is to be formed from a group of 4 men and 5 women. Find the probability that
the committee has at least one woman.
Solution:
9!
Total ways to select 3 members from 9 = 𝐶𝐶(9,3) = = 84
3!(9−3)!
4!
Ways to select only men (no women) = 𝐶𝐶(4,3) = =4
3!(4−3)!
4 80 20
P (At least 1 Woman) = 1 − 𝑃𝑃 (𝑂𝑂𝑂𝑂𝑂𝑂𝑂𝑂 𝑀𝑀𝑀𝑀𝑀𝑀) = 1 − = =
84 84 21
Example 10. A 4-digit PIN is created using the numbers 0 to 9. Find the probability that all four digits are
different.
Solution:
Total possible PINs =10×10×10×10=10000
Ways to select 4 different digits:
• First digit: 10 choices
• Second digit: 9 choices
• Third digit: 8 choices
• Fourth digit: 7 choices
Total ways to select 4 different digits = 10 × 9 × 8 × 7 = 5040
P (All digits different) = 5040/10000 = 0.504
98
ISC XI- Applied Mathematics
Example 11. A box contains five blue balls, two green balls, and six yellow balls. What are the
(i) odds of drawing a blue ball?
(ii) odds against drawing a blue ball from the box?
Solution:
Let P(B) represent the event that a blue ball is drawn from the box.
5
Therefore, 𝑃𝑃(𝐵𝐵) =
13
5 5
5 13 5
(i) Odds of drawing a blue ball from the box = 13
5 = 13
8 =( )×( )=
�1− � � � 13 8 8
13 13
5 8
�1− � 8 13 8
(ii) Odds against drawing a blue ball from the box= 5
13
= 13
5 = � �� � = 5
� � 13 5
13 13
Suppose in a class of 30 students, 8 students are in the band, 15 students play a sport, and 5 students are both
in the band and play a sport. Let A be the event that a student is in the band and let B be the event that a student
plays a sport.
P(A∪B) is the probability that a student is in the band or plays a sport or both.
With the help of the Venn diagram,
3 + 10 + 5 18 3
𝑃𝑃(𝐴𝐴 ∪ 𝐵𝐵) = = =
30 30 5
You could also compute this probability using the Addition Rule:
Note that by using the Addition Rule, you avoid having to determine that 3 people
are in a band and don't play a sport and 10 people who play a sport but are not in
the band. The Addition Rule is easier when you have not created a Venn diagram.
Addition Theorem of Probability for Three Events
If A, B, C are three events, then
𝑃𝑃(𝐴𝐴 ∪ 𝐵𝐵 ∪ 𝐶𝐶) = 𝑃𝑃(𝐴𝐴) + 𝑃𝑃(𝐵𝐵) + 𝑃𝑃(𝐶𝐶) − 𝑃𝑃(𝐴𝐴 ∩ 𝐵𝐵) − 𝑃𝑃(𝐵𝐵 ∩ 𝐶𝐶) − 𝑃𝑃(𝐴𝐴 ∩ 𝐶𝐶) + 𝑃𝑃(𝐴𝐴 ∩ 𝐵𝐵 ∩ 𝐶𝐶)
This theorem extends the two-event addition rule and ensures that each outcome is counted exactly once.
Corollary:
If A, B, C are mutually exclusive 𝑃𝑃(𝐴𝐴 ∩ 𝐵𝐵) = 𝑃𝑃(𝐵𝐵 ∩ 𝐶𝐶) = 𝑃𝑃(𝐴𝐴 ∩ 𝐶𝐶) = 𝑃𝑃(𝐴𝐴 ∩ 𝐵𝐵 ∩ 𝐶𝐶) = 0. Therefore,
𝑃𝑃(𝐴𝐴 ∪ 𝐵𝐵 ∪ 𝐶𝐶) = 𝑃𝑃(𝐴𝐴) + 𝑃𝑃(𝐵𝐵) + 𝑃𝑃(𝐶𝐶)
99
ISC XI- Applied Mathematics
Special Cases of the Addition Theorem - Probability of any one of two events occurring exactly
For two events A and B, the probability that only one of them occurs but not both is given by
This formula ensures that we count only the cases in which either A or B occurs individually and not the
instances in which they occur together.
P(A) and P(B) include all cases where either event A or B happens.
𝑃𝑃(𝐴𝐴 ∩ 𝐵𝐵) is the probability of both events occurring simultaneously. If we simply subtract 𝑃𝑃(𝐴𝐴 ∩ 𝐵𝐵) once (as
in the standard addition theorem), we still count cases where both events occur. To ensure that these cases are
completely excluded, we subtract 𝑃𝑃(𝐴𝐴 ∩ 𝐵𝐵) twice, ensuring that the overlap is fully removed.
Example 13. Nick enjoys dessert after dinner and on any given night the probability that Nick has a
cookie (A) is 10%, the probability that Nick has ice cream (B) is 50% and the probability that Nick has both
cookie and ice cream (A∩B) is 5%. Calculate the probability that Nick has exactly one of the two desserts.
Solution:
Using the formula, 𝑃𝑃((𝐴𝐴 ∩ 𝐵𝐵�) ∪ (𝐴𝐴̅ ∩ 𝐵𝐵)) = 𝑃𝑃(𝐴𝐴) + 𝑃𝑃(𝐵𝐵) − 2𝑃𝑃(𝐴𝐴 ∩ 𝐵𝐵)
= 0.10 + 0.50 − 2(0.05) = 0.10 + 0.50 − 0.10 = 0.50
There is a 50% chance that Nick has either only a cookie or only ice cream but not
both.
Remark 1:
𝑃𝑃(𝐴𝐴 ∩ 𝐵𝐵�) or P (A – B) is the probability of occurrence of A only. This means that
among the outcomes in event A, we exclude those that also belong to event B.
100
ISC XI- Applied Mathematics
Remark 2:
Remark 3:
𝑃𝑃((𝐴𝐴 ∩ 𝐵𝐵�) ∪ (𝐴𝐴̅ ∩ 𝐵𝐵)) is the probability of exactly one of two events occurring.
This means either event A occurs without event B or event B occurs without
event A.
Remark 4:
If A and B are events such that A⊂B, then B splits into A and B. Therefore,
𝑃𝑃(𝐴𝐴) ≤ 𝑃𝑃(𝐵𝐵). This states that if A is completely contained within B, then the
probability of A occurring can never exceed the probability of B.
Suppose the probability of rain today is 0.6 and the probability of high winds is 0.3. The probability of both
rain and high winds occurring together is 0.2.
𝑃𝑃(𝑟𝑟𝑟𝑟𝑟𝑟𝑟𝑟 ∩ 𝑤𝑤𝑤𝑤𝑤𝑤𝑤𝑤) = 0.2 ≤ 𝑃𝑃(𝑟𝑟𝑟𝑟𝑟𝑟𝑟𝑟) = 0.6 ≤ 𝑃𝑃(𝑟𝑟𝑟𝑟𝑟𝑟𝑟𝑟 ∪ 𝑤𝑤𝑤𝑤𝑤𝑤𝑤𝑤) ≤ 0.6 + 0.3 = 0.9
Thus, the probability of either rain or high winds (or both) must be between 0.6 and 0.9.
3 1
Example 14. Given 𝑃𝑃(𝐴𝐴) = and 𝑃𝑃(𝐵𝐵) = , find P(A∪B) if A and B are mutually exclusive events.
5 5
Solution:
Since A and B are mutually exclusive, they do not overlap, meaning 𝑃𝑃(𝐴𝐴 ∩ 𝐵𝐵) = 0.
The addition theorem states 𝑃𝑃(𝐴𝐴 ∪ 𝐵𝐵) = 𝑃𝑃(𝐴𝐴) + 𝑃𝑃(𝐵𝐵)
3 1 4
Substituting the given values, 𝑃𝑃(𝐴𝐴 ∪ 𝐵𝐵) = + =
5 5 5
Example 15. Two mutually exclusive events A and B are given in an experiment. If 𝑃𝑃(𝐴𝐴̅) = 0.65 and
𝑃𝑃(𝐴𝐴 ∪ 𝐵𝐵) = 0.65, find P(B).
Solution:
𝑃𝑃(𝐴𝐴) = 1 − 𝑃𝑃(𝐴𝐴̅) = 1 − 0.65 = 0.35
Since A and B are mutually exclusive,
101
ISC XI- Applied Mathematics
1 2 1
Example 16. For two non-mutually exclusive events,𝑃𝑃(𝐴𝐴) = , 𝑃𝑃(𝐵𝐵) = and 𝑃𝑃(𝐴𝐴 ∪ 𝐵𝐵) = . Find the values
4 5 2
of 𝑃𝑃(𝐴𝐴 ∩ 𝐵𝐵) and 𝑃𝑃(𝐴𝐴 ∩ 𝐵𝐵�).
Solution:
1 1 2
𝑃𝑃(𝐴𝐴 ∪ 𝐵𝐵) = 𝑃𝑃(𝐴𝐴) + 𝑃𝑃(𝐵𝐵) − 𝑃𝑃(𝐴𝐴 ∩ 𝐵𝐵) ⇒ = + − 𝑃𝑃(𝐴𝐴 ∩ 𝐵𝐵)
2 4 5
1 5 8 1 13 3
2
= 20 + 20 − 𝑃𝑃(𝐴𝐴 ∩ 𝐵𝐵) ⇒ 2 = 20 − 𝑃𝑃(𝐴𝐴 ∩ 𝐵𝐵) ⇒ 𝑃𝑃(𝐴𝐴 ∩ 𝐵𝐵) = 20
1 3 1
� ) = 𝑃𝑃(𝐴𝐴) − 𝑃𝑃(𝐴𝐴 ∩ 𝐵𝐵) = −
Now, using 𝑃𝑃(𝐴𝐴 ∩ 𝐵𝐵 = 10
4 20
1 1 1
Example 17. If 𝑃𝑃(𝐸𝐸) = , 𝑃𝑃(𝐹𝐹) = and 𝑃𝑃(𝐸𝐸 ∩ 𝐹𝐹) = , find
4 2 8
(i) 𝑃𝑃(𝐸𝐸 ∪ 𝐹𝐹)
(ii) 𝑃𝑃(𝐸𝐸� ∩ 𝐹𝐹� ) (Probability that neither event occurs).
Solution:
1 1 1 5
(i) 𝑃𝑃(𝐸𝐸 ∪ 𝐹𝐹) = 𝑃𝑃(𝐸𝐸) + 𝑃𝑃(𝐹𝐹) − 𝑃𝑃(𝐸𝐸 ∩ 𝐹𝐹) = 4 + 2 − 8 = 8
5 3
(ii) 𝑃𝑃(𝐸𝐸� ∩ 𝐹𝐹� ) = 1 − 𝑃𝑃(𝐸𝐸 ∪ 𝐹𝐹) = 1 − 8 = 8
Example 18. The probability that at least one of the events A and B occurs is 0.6. If 𝑃𝑃(𝐴𝐴 ∩ 𝐵𝐵) =
0.2, find 𝑃𝑃(𝐴𝐴) + 𝑃𝑃(𝐵𝐵).
Solution:
𝑃𝑃(𝐴𝐴 ∪ 𝐵𝐵) = 𝑃𝑃(𝐴𝐴) + 𝑃𝑃(𝐵𝐵) − 𝑃𝑃(𝐴𝐴 ∩ 𝐵𝐵) ⇒ 0.6 = 𝑃𝑃(𝐴𝐴) + 𝑃𝑃(𝐵𝐵) − 0.2
∴ 𝑃𝑃(𝐴𝐴) + 𝑃𝑃(𝐵𝐵) = 0.6 + 0.2 = 0.8
Example 19. A number is chosen at random from 1 to 200. What is the probability that it is divisible by 4 or
6?
Solution:
Let A be the event that the number is divisible by 4 and B be the event that the number is divisible by 6.
For A, 200 = 4 + (𝑛𝑛 − 1)4 ⇒ 4(𝑛𝑛 − 1) = 196 ⇒ 𝑛𝑛 − 1 = 49 ⇒ 𝑛𝑛 = 50.
For B, 198 = 6 + (𝑛𝑛 − 1)6 ⇒ 6(𝑛𝑛 − 1) = 192 ⇒ 𝑛𝑛 − 1 = 32 ⇒ 𝑛𝑛 = 33.
For 𝐴𝐴 ∩ 𝐵𝐵, 192 = 12 + (𝑛𝑛 − 1)12 ⇒ 12(𝑛𝑛 − 1) = 180 ⇒ 𝑛𝑛 − 1 = 15 ⇒ 𝑛𝑛 = 16.
50 33 16
𝑃𝑃(𝐴𝐴) = , 𝑃𝑃(𝐵𝐵) = , 𝑃𝑃(𝐴𝐴 ∩ 𝐵𝐵) =
200 200 200 Recall, Arithmetic
50 33 16 67 progression to find the
𝑃𝑃(𝐴𝐴 ∪ 𝐵𝐵) = 𝑃𝑃(𝐴𝐴) + 𝑃𝑃(𝐵𝐵) − 𝑃𝑃(𝐴𝐴 ∩ 𝐵𝐵) = + − = number of terms.
200 200 200 200
102
ISC XI- Applied Mathematics
Solution:
𝑃𝑃(𝐹𝐹 ∪ 𝐼𝐼) = 𝑃𝑃(𝐹𝐹) + 𝑃𝑃(𝐼𝐼) − 𝑃𝑃(𝐹𝐹 ∩ 𝐼𝐼) = 0.7 + 0.5 − 0.3 = 0.9
The probability that a random person uses either Facebook, Instagram, or both is 0.9 (90%).
Example 21. The Union Budget is going to be announced by the government this week. The probability that
it will be announced on a day is given below:
Find the probability of the budget getting announced between Monday to Wednesday.
Solution:
1 3 1 5
P(Monday to Wednesday) = P(Monday) + P(Tuesday) + P(Wednesday) = + + =
7 7 7 7
Example 22. A and B are two candidates seeking admission to IIT. The probability that A getting selected is
0.5 and the probability that both A and B getting selected is 0.3. Prove that the probability of B being selected
is at most 0.8.
Solution
𝑃𝑃 (𝐴𝐴) = 0.5, 𝑃𝑃 (𝐴𝐴 ∩ 𝐵𝐵) = 0.3
We have P (A U B) ≤ 1 ⇒ P (A) + P(B) − P (A ∩ B) ≤ 1
∴ 0.5 + P (B) − 0.3 ≤ 1 ⇒ P (B) ≤ 1 − 0.2 ⇒ P (B) ≤ 0.8
Therefore, probability of B getting selected is at most 0.8.
103
ISC XI- Applied Mathematics
• Elections and Polling - Predicting election results using random samples (exit polls).
• Sports - Calculating winning chances, player performance, and match outcomes.
• Genetics and Biology - Determining the probability of inheriting certain traits or diseases.
• Computer Algorithms (e.g., AI and Machine Learning) - Making predictions or decisions based on
data probabilities.
• Project Management - Assessing the probability of meeting deadlines and budget constraints.
• Cryptography and Cybersecurity - Estimating chances of data breaches or decoding encrypted data.
Exercises
1. Choose the correct answer.
1 1
(i) If A, B, C are mutually exclusive and exhaustive events, find P(B) if 𝑃𝑃(𝐶𝐶) = 2 𝑃𝑃(𝐴𝐴) = 𝑃𝑃(𝐵𝐵).
3
1 1
(a) (c)
6 3
1 1
(b) (d)
4 2
(ii) There are four hotels in a hill station. If three tourists check into hotels in a day, what is the
probability that each check into a different hotel?
1 3
(a) (c)
4 8
1 3
(b) (d)
3 4
(iii) One ticket is drawn from a bag containing 30 tickets numbered 1 to 30. Find the probability that it is
a multiple of 3 or 5.
1 7
(a) (c)
15 15
1 8
(b) (d)
3 15
(iv) Let A and B be two events such that 𝑃𝑃(𝐴𝐴 ∪ 𝐵𝐵) = 𝑃𝑃(𝐴𝐴 ∩ 𝐵𝐵). Then
�) − P(A
Statement 1. P(A ∩ B � ∩ B) = 0.
Statement 2. P(A) + P(B) = 1.
(a) Only 1 (c) Both 1 and 2
(b) Only 2 (d) Neither 1 nor 2
(v) Assertion: The probability that a person visiting a zoo will see tigers is 0.72, the probability that he
will see the elephants is 0.84 and the probability that he will see both is 0.52.
Reason: 0 ≤ 𝑃𝑃(𝐴𝐴) ≤ 1
(a) Both A and R individually true and R is the correct explanation of A.
(b) Both A and R individually true but R is not the correct explanation of A.
(c) A is true but R is false.
(d) A is false but R is true.
104
ISC XI- Applied Mathematics
2. An ice cream shop has vanilla, chocolate or strawberry ice cream and you can have either chocolate or
caramel sauce on top.
(i) Represent this information on a tree diagram.
Use the tree diagram to answer the following questions: What is the probability that a customer
chosen at random will order?
(ii) a chocolate ice cream?
(iii) an ice cream with caramel sauce?
(iv) a strawberry ice cream with caramel sauce?
3. You are buying a jumper from a shop. You have to choose a size, colour and style. You can choose
between small, medium or large for size; blue, pink, red or brown for colour; and round-neck or V-neck
for style.
(i) Draw a tree diagram to represent the sample space.
Use the tree diagram to answer the following questions:
(ii) How many different options do you have?
(iii) What is the probability that you pick a brown jumper?
(iv) What is the probability that you pick a jumper that has a V-neck?
(v) What is the probability that you pick a large, red, round-neck jumper?
4. A bag contains 4 white and 2 black balls. Two balls are drawn at random. Find the probability that they
are white.
5. A, B and C are bidding for a contract. It is believed that A has exactly half the chance of B and B in turn
4
is th. as likely as C to win the contract. What is the probability for each to win the contract?
5
6. Three dice are tossed. What is the probability that the numbers shown will all be different?
8. From a pack of cards if one card is drawn at random what is the probability that it is either a King or a
Queen?
10. Four-digit numbers are formed by using the digits 1, 2, 3, 4 and 5 without repetition. Find the probability
that a number, chosen at random, is an odd number.
11. The letters of the word ‘SOCIETY’ are placed at random in a row. What is the probability that three
vowels come together?
12. An urn contains 9 balls, 2 of which are white, 3 blue and 4 black. 3 balls are drawn at random from the
urn. What is the chance that (i) the three balls will be of different colours (ii) 2 balls will be of the same
colour and the third a different colour (iii) three balls will be of the same colour?
105
ISC XI- Applied Mathematics
13. There are 10 persons who are to be seated around a circular table. Find the probability that two particular
persons will always sit together.
14. Eight people in a start-up of 25 people are graduates. If 3 are picked out at random, what is the probability
that
(i) they are all graduates.
(ii) there is at least one graduate.
15. Given are the ages of six people in an old age home:62,90,78, 85,79, 68
If two people are selected as representatives, what is the probability that at least one will have age lower
than the average.
16. Two dice are thrown together. Find the probability that the total is
(i) five
(ii) two
(iii) two or five
(iv) at least five
17. In the Annual sports meet, among the 260 students in XI standard in the school, 90 participated in Kabaddi,
120 participated in Hockey, and 50 participated in Kabaddi and Hockey. A Student is selected at random.
Find the probability that the student participated in
(i) Either Kabaddi or Hockey,
(ii) Neither of the two tournaments,
(iii) Hockey only,
(iv) Kabaddi only,
(v) Exactly one of the tournaments.
18. Twenty-six people were surveyed about their choice of mobile phones. The survey finds that 14 people
have Apple iPhones, 10 have Samsungs and five have Motorola. Four have Apple iPhones and Samsungs,
three have Apple iPhones and Motorola and one has a Samsung and a Motorola. No one has all three types
of phones. Represent this information on a Venn diagram and use it to calculate the probability that a
person chosen at random has:
(i) An iphone
(ii) A Samsung and a Motorola
(iii) Two phones
(iv) No phone
(v) No iphone
19. Two candidates, Riya and Megha, recently sat for a national-level entrance examination.
Riya has been juggling her preparation with a part-time internship, limiting the number of hours she could
dedicate to study. Megha, although studying full-time, has had difficulty grasping some critical topics.
Based on their preparation patterns and previous mock test performances:
• The probability that Riya qualifies the exam is 0.05,
• The probability that Megha qualifies is 0.10, and
• The probability that both Riya and Megha qualify is 0.02.
Answer the following:
(i) What is the probability that at least one of them will qualify the examination?
(ii) What is the probability that at least one of them will not qualify?
106
ISC XI- Applied Mathematics
(iii) What is the probability that neither Riya nor Megha qualifies the exam?
(iv) What is the probability that only one of them qualifies the exam?
20. Out of four team members — Arjun, Nisha, Faizan, and Leena — only one person will be selected for a
leadership training program next month. Therefore, the sample space is:
S = {Arjun selected, Nisha selected, Faizan selected, Leena selected}
You are given the following information:
• The chances of Arjun being selected are equal to Leena's.
• Nisha's chances are twice that of Arjun.
• Faizan's chances are four times that of Arjun.
Based on this, answer the following:
(i) What is the probability that Arjun gets selected?
(ii) What is the probability that Nisha gets selected?
(iii) What is the probability that Faizan gets selected?
(iv) What is the probability that Leena gets selected?
21. A sample space consists of 9 elementary outcomes E1, E2, ..., E9 whose probabilities are
P(E1) = P(E2) = 0.08,
P (E3) = P (E4) = 0.1
P (E6) = P (E7) = 0.2
P (E8) = P (E9) = 0.07
Suppose A = {E1, E5, E8}, B = {E2, E5, E8, E9}
(i) Calculate P (A), P (B), and P (A ∩ B).
(ii) Calculate P (A ∪ B).
(iii) Calculate P (𝐵𝐵�).
Answers
1
1 (i) (a)
6
3
(ii) (c)
8
7
(iii) (c)
15
107
ISC XI- Applied Mathematics
1
(ii) P (Chocolate) =
3
1
(iii) (iii) P (Caramel) =
2
1
(iv) (iv) P (Strawberry with Caramel) =
6
3. (i)
(ii) 24
1
(iii) P(Brown) =
4
1
(iv) P (V neck) =
2
1
(v) P (large, red, round-neck) =
24
2
4
5
2 4 5
5. P(A) = , P(B) = , P(C) =
11 11 11
5
6
9
7. (i) 3:2 (ii) 7:3
2
8.
13
9. (i) 0.3 (ii) 0.6 (iii) 0.4
10. 0.6
1
11.
7
2 55 5
12. (i) (ii) 84 (iii) 84
7
2
13.
9
14 81
14. (i) (ii) 115
575
3
15.
5
1 1 5 5
16. (i) (ii) 36 (iii) 36 (iv) 6
9
8 5 7 2 11
17. (i) (ii) 13 (iii) 26 (iv) 13 (v) 26
13
108
ISC XI- Applied Mathematics
7 1 4 5 6
18. (i) (ii) 26 (iii) 13 (iv) 26 (v) 13
13
19. (i) 0.13 (ii) 0.98 (iii) 0.87 (iv) 0.11
1 1 1 1
20. (i) (ii) (iii)2 (iv)8
8 4
21. P(A) = 0.25, P(B) = 0.32, P (𝐴𝐴 ∩ 𝐵𝐵) = 0.17, P( 𝐴𝐴 ∪ 𝐵𝐵) = 0.4, 𝑃𝑃(𝐵𝐵�) = 0.68
HISTORICAL NOTE
• Early Beginnings: The earliest motivation for probability came from gambling and games of
chance. People wanted to understand how likely they were to win or lose in games involving dice,
cards, or coins.
• A physician and mathematician, Gerolamo Cardano, wrote a book titled "Liber de Ludo Aleae"
(Book on Games of Chance), which included rules of probability based on dice throws. He was one
of the first to define probability as a ratio of favourable to total outcomes.
• The modern theory of probability began with a famous correspondence between Blaise Pascal and
Pierre de Fermat in [Link] solved the "Problem of Points", which dealt with how to fairly divide
winnings in an unfinished game — laying the foundation of expected value and fair division. This
is considered the birth of modern probability theory.
• Christiaan Huygens published the first formal textbook on probability, "De Ratiociniis in Ludo
Aleae" (On Reasoning in Games of Chance), in 1657. He used mathematical methods to calculate
probabilities in dice games and introduced terminology like "expectation."
• Mathematicians like Jacob Bernoulli, Abraham de Moivre, and Pierre-Simon Laplace expanded the
subject by developing laws of large numbers, the normal distribution, and more complex probability
models. Set theory and formal logic, developed later, helped formalize the ideas of events, sample
spaces, and operations on events like union and intersection.
• With the development of set theory in the 19th century by Georg Cantor and others, events in
probability could be represented as sets, allowing for more systematic calculation.
109