P&S Module-2 Notes
P&S Module-2 Notes
Contents
2 Probability 4
2.1 Classical probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4
2.2 Frequency interpretation of probability . . . . . . . . . . . . . . . . . . . . . 6
2.3 Subjective probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
1
Prof. V. K. Narla Module 2: Introduction to Probability
6 Bayes’ Theorem 35
6.1 Partition of the Sample Space . . . . . . . . . . . . . . . . . . . . . . . . . . 35
6.2 Statement and Derivation of Bayes’ Theorem . . . . . . . . . . . . . . . . . . 36
6.3 Interpretation of the Formula . . . . . . . . . . . . . . . . . . . . . . . . . . 36
6.4 Procedure for Applying Bayes’ Theorem . . . . . . . . . . . . . . . . . . . . 37
6.5 Examples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38
6.6 Important Observations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42
6.7 Formula Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42
Page 2
Prof. V. K. Narla Module 2: Introduction to Probability
Page 3
Prof. V. K. Narla Module 2: Introduction to Probability
S = {0, 1, 2, . . .}
1.5 Events
Any subset of a sample space is called an event. An event may contain one outcome, several
outcomes, the entire sample space, or no outcomes.
The set containing no outcomes is called the empty set and is denoted by
∅.
2. Probability
After identifying what is possible in an experiment, the next step is to describe what is
probable or improbable. Three common interpretations of probability are:
1. the classical interpretation;
2. the frequency interpretation;
3. the subjective interpretation.
Page 4
Prof. V. K. Narla Module 2: Introduction to Probability
If there are m equally likely possibilities, one of which must occur, and s of them are
favourable to an event, then the probability of that event is
s
P (A) = .
m
Equivalently,
number of favourable outcomes
P (A) = .
total number of equally likely outcomes
The terms favourable and success are used in a technical sense. A favourable outcome
does not necessarily represent something desirable.
Example 1: Well-shuffled cards are equally likely to be selected
What is the probability of drawing an ace from a well-shuffled deck of 52 playing cards?
Solution. There are 4 aces among the 52 cards. Hence,
m = 52, s = 4.
Therefore,
s 4 1
P (ace) = = = .
m 52 13
The assumption of equally likely outcomes is commonly used in games of chance and also
in carefully designed random-selection procedures.
Example 2: Random selection in the equally likely case
Ten miniature electric motors are available. Tests show that 8 motors operate satisfac-
torily and 2 do not. Two motors are selected at random, with every pair having the same
chance of being selected.
Find the probability that:
(a) both motors operate satisfactorily;
(b) one motor operates satisfactorily and the other does not.
Solution. The total number of possible pairs is
(︃ )︃
10
m= = 45.
2
Page 5
Prof. V. K. Narla Module 2: Introduction to Probability
The number of ways of selecting one satisfactory motor and one unsatisfactory motor is
(︃ )︃(︃ )︃
8 2
s= = 8(2) = 16.
1 1
Hence,
16
P (one satisfactory and one unsatisfactory) = ≈ 0.356.
45
Page 6
Prof. V. K. Narla Module 2: Introduction to Probability
P (S) = 1.
Since one of the possible outcomes must occur, the probability of the certain event is 1.
P (A ∪ B) = P (A) + P (B).
Page 7
Prof. V. K. Narla Module 2: Introduction to Probability
Page 8
Prof. V. K. Narla Module 2: Introduction to Probability
Explanation
The third axiom gives the result for two mutually exclusive events:
Repeating this argument gives the result for any finite number n of mutually exclusive
events.
Example 5: Probabilities add for mutually exclusive events
A consumer testing service assigns the following probabilities to the ratings of a new
antipollution device for cars:
By Theorem 3.4,
P (A) = 0.07 + 0.12 + 0.17 + 0.32
= 0.19 + 0.17 + 0.32
= 0.36 + 0.32
= 0.68.
Thus,
P (very poor, poor, fair, or good) = 0.68.
(b) Good, very good, or excellent
Let B denote this event. Then
Page 9
Prof. V. K. Narla Module 2: Introduction to Probability
Therefore,
P (B) = 0.32 + 0.21 + 0.11
= 0.53 + 0.11
= 0.64.
Hence,
P (good, very good, or excellent) = 0.64.
Detailed justification
Only one individual outcome can occur in a single performance of the experiment. Hence,
Ei ∩ Ej = ∅ for i ̸= j.
The event A occurs precisely when one of the outcomes E1 , E2 , . . . , En occurs. Therefore,
P (A) = P (E1 ∪ E2 ∪ · · · ∪ En ),
Page 10
Prof. V. K. Narla Module 2: Introduction to Probability
Derivation
The union A ∪ B can be divided into the following three mutually exclusive parts:
A ∩ B, A ∩ B, A ∩ B.
Thus,
A ∪ B = (A ∩ B) ∪ (A ∩ B) ∪ (A ∩ B).
By Theorem 3.4,
P (A ∪ B) = P (A ∩ B) + P (A ∩ B) + P (A ∩ B).
Also,
P (A) = P (A ∩ B) + P (A ∩ B),
and
P (B) = P (A ∩ B) + P (A ∩ B).
Adding the last two equations gives
Therefore,
P (A ∪ B) = P (A) + P (B) − P (A ∩ B).
The intersection probability is subtracted because outcomes in A ∩ B are counted once
in P (A) and once again in P (B).
If A and B are mutually exclusive, then
P (A ∩ B) = 0,
and the general addition rule reduces to the special addition rule
P (A ∪ B) = P (A) + P (B).
and
C3 = {the car is expensive to operate}.
Suppose
P (M1 ) = 0.20, P (C3 ) = 0.40, P (M1 ∩ C3 ) = 0.08.
Page 11
Prof. V. K. Narla Module 2: Introduction to Probability
Find the probability that a car has low mileage or is expensive to operate; that is, find
P (M1 ∪ C3 ).
Solution.
The word “or” indicates the union of the two events. Since a car may have both low
mileage and high operating cost, the events are not necessarily mutually exclusive. Hence,
Theorem 3.6 must be used:
P (M1 ∪ C3 ) = P (M1 ) + P (C3 ) − P (M1 ∩ C3 ).
Substituting the given values,
P (M1 ∪ C3 ) = 0.20 + 0.40 − 0.08
= 0.60 − 0.08
= 0.52.
Therefore,
P (low mileage or expensive to operate) = 0.52.
Page 12
Prof. V. K. Narla Module 2: Introduction to Probability
Proof
The event A and its complement A are mutually exclusive:
A ∩ A = ∅.
A ∪ A = S.
P (A) + P (A) = 1.
Therefore,
P (A) = 1 − P (A).
As a special case,
P (∅) = 1 − P (S) = 1 − 1 = 0.
find:
(a) the probability that a used car does not have low mileage;
(b) the probability that a used car either does not have low mileage or is not expensive
to operate.
Solution.
(a) The car does not have low mileage
The event “does not have low mileage” is M1 . By the complement rule,
P (M1 ) = 1 − P (M1 )
= 1 − 0.20
= 0.80.
Page 13
Prof. V. K. Narla Module 2: Introduction to Probability
Thus,
P (M1 ) = 0.80.
(b) The car either does not have low mileage or is not expensive to operate
The required event is
M1 ∪ C3 .
By De Morgan’s law,
M1 ∪ C3 = M1 ∩ C3 .
Therefore, by the complement rule,
(︁ )︁
P (M1 ∪ C3 ) = P M1 ∩ C3
= 1 − P (M1 ∩ C3 )
= 1 − 0.08
= 0.92.
Hence,
P (M1 ∪ C3 ) = 0.92.
n
∑︂
P (A1 ∪ A2 ∪ · · · ∪ An ) = P (Ai ), if the events are mutually exclusive,
i=1
∑︂
P (A) = P (Ei ), for a finite sample space,
Ei ∈A
Page 14
Prof. V. K. Narla Module 2: Introduction to Probability
Answer.
When two balanced dice are rolled, each die has 6 possible outcomes. Therefore, the
total number of equally likely ordered outcomes is
6 × 6 = 36.
(a) A sum of 7
The favourable outcomes are
(1, 6), (2, 5), (3, 4), (4, 3), (5, 2), (6, 1).
There are 6 favourable outcomes. Hence,
6 1
P (sum 7) = = .
36 6
(b) A sum of 11
The favourable outcomes are
(5, 6), (6, 5).
Thus,
2 1
P (sum 11) = = .
36 18
(c) A sum of 7 or 11
The events “sum 7” and “sum 11” are mutually exclusive. Therefore,
P (sum 7 or 11) = P (sum 7) + P (sum 11)
6 2
= +
36 36
8
=
36
2
= .
9
(d) A sum of 3
The favourable outcomes are
(1, 2), (2, 1).
Therefore,
2 1
P (sum 3) = = .
36 18
(e) A sum of 2 or 12
The sum 2 occurs only for (1, 1), and the sum 12 occurs only for (6, 6). Hence,
1+1 2 1
P (sum 2 or 12) = = = .
36 36 18
(f) A sum of 2, 3, or 12
The numbers of favourable outcomes are 1, 2, and 1, respectively. Thus,
1+2+1 4 1
P (sum 2, 3, or 12) = = = .
36 36 9
Page 15
Prof. V. K. Narla Module 2: Introduction to Probability
Page 16
Prof. V. K. Narla Module 2: Introduction to Probability
0.359 .
0.033 ,
or about 3.3%.
Page 17
Prof. V. K. Narla Module 2: Introduction to Probability
Let
S = {students enrolled in advanced statistics}
and
O = {students enrolled in operations research}.
Using the addition rule,
n(S ∪ O) = 92 + 63 − 40
= 155 − 40
= 115.
Hence,
45 students
are enrolled in neither course.
(b)
P (A) = 0.27, P (B) = 0.30, P (C) = 0.28, P (D) = 0.16.
(c)
P (A) = 0.32, P (B) = 0.27, P (C) = −0.06, P (D) = 0.47.
(d)
1 1 1 1
P (A) = , P (B) = , P (C) = , P (D) = .
2 4 8 16
(e)
5 1 1 2
P (A) = , P (B) = , P (C) = , P (D) = .
18 6 3 9
Page 18
Prof. V. K. Narla Module 2: Introduction to Probability
Answer.
A permissible probability assignment must satisfy
and
P (A) + P (B) + P (C) + P (D) = 1.
(a) All probabilities are nonnegative, and
Page 19
Prof. V. K. Narla Module 2: Introduction to Probability
Therefore,
0.262 + 0.314 + 0.242 + 0.182 = 1.
Page 20
Prof. V. K. Narla Module 2: Introduction to Probability
For event B,
B = {(0, 0), (0, 1), (0, 2), (0, 3)}.
Hence,
P (B) = 0.080 + 0.032 + 0.086 + 0.064
= 0.262.
C = {(0, 1), (0, 2), (0, 3), (1, 2), (1, 3), (2, 3)}.
Therefore,
P (C) = 0.032 + 0.086 + 0.064 + 0.065 + 0.091 + 0.075
= 0.413.
Hence,
P (A) = 0.240, P (B) = 0.262, P (C) = 0.413.
Thus,
Page 21
Prof. V. K. Narla Module 2: Introduction to Probability
Find:
(a) P (A);
(b) P (A ∪ B);
(c) P (A ∩ B);
(d) P (A ∩ B).
Answer.
Since A and B are mutually exclusive,
A ∩ B = ∅ and P (A ∩ B) = 0.
P (A ∪ B) = P (A) + P (B)
= 0.45 + 0.30
= 0.75.
(c) Since A and B cannot occur together, every outcome in A belongs to B. Therefore,
A ∩ B = A,
and
P (A ∩ B) = P (A) = 0.45.
Therefore,
Page 22
Prof. V. K. Narla Module 2: Introduction to Probability
0.01, 0.03, 0.07, 0.15, 0.19, 0.18, 0.14, 0.12, 0.09, 0.02.
X = 0, 1, 2, 3, or 4.
Therefore,
P (X ≤ 4) = 0.01 + 0.03 + 0.07 + 0.15 + 0.19
= 0.45.
X = 6, 7, 8, or at least 9.
Thus,
P (X ≥ 6) = 0.14 + 0.12 + 0.09 + 0.02
= 0.37.
X = 5, 6, 7, or 8.
Therefore,
P (5 ≤ X ≤ 8) = 0.18 + 0.14 + 0.12 + 0.09
= 0.53.
Hence,
Page 23
Prof. V. K. Narla Module 2: Introduction to Probability
P (A ∪ B) = P (A) + P (B) − P (A ∩ B)
= 0.30 + 0.62 − 0.12
= 0.80.
P (B) = P (A ∩ B) + P (A ∩ B).
Hence,
P (A ∩ B) = P (B) − P (A ∩ B)
= 0.62 − 0.12
= 0.50.
P (A ∩ B) = P (A) − P (A ∩ B)
= 0.30 − 0.12
= 0.18.
Page 24
Prof. V. K. Narla Module 2: Introduction to Probability
P (B) > 0.
P (A ∩ B)
P (A | B) = .
P (B)
Here:
P (A | B) means the probability that A occurs, given that B has occurred;
A ∩ B represents the event that both A and B occur;
P (B) is the probability of the conditioning event.
Once it is known that B has occurred, the sample space is effectively restricted to B.
Among the outcomes in B, only those belonging to A ∩ B are favourable to event A. This
explains the ratio
P (A ∩ B)
.
P (B)
Page 25
Prof. V. K. Narla Module 2: Introduction to Probability
Page 26
Prof. V. K. Narla Module 2: Introduction to Probability
P (A | B) = P (A).
P (B | A) = P (B).
Thus, for independent events, knowing that one event has occurred provides no additional
information about the occurrence of the other event.
In Example 2, the conditional probability was
P (M1 | C3 ) = 0.20.
P (M1 ) = 0.20,
then
P (M1 | C3 ) = P (M1 ),
so the events M1 and C3 are independent.
or equivalently,
P (A ∩ B) = P (B) P (A | B), P (B) > 0.
Derivation
From the definition of conditional probability,
P (A ∩ B)
P (A | B) = .
P (B)
P (A ∩ B) = P (B)P (A | B).
Page 27
Prof. V. K. Narla Module 2: Introduction to Probability
Similarly,
P (A ∩ B)
P (B | A) = ,
P (A)
and multiplication by P (A) gives
P (A ∩ B) = P (A)P (B | A).
The multiplication rule is particularly useful for sequential experiments, especially when
the result of the first selection changes the probability of the second selection.
Example 3: Using the General Multiplication Rule
A supervisor has a group of 20 construction workers. Of these, 12 workers favour new
safety regulations and 8 are against them. The supervisor randomly selects two workers
without replacement.
Find the probability that both selected workers are against the new safety regulations.
Solution.
Let
A = {the first selected worker is against the regulations},
and
B = {the second selected worker is against the regulations}.
For the first selection, 8 of the 20 workers are against the regulations. Therefore,
8
P (A) = .
20
Since the first selected worker is not replaced, if the first worker is against the regulations,
then 7 such workers remain among the remaining 19 workers. Hence,
7
P (B | A) = .
19
Using the general multiplication rule,
P (A ∩ B) = P (A)P (B | A).
Therefore,
8 7
P (A ∩ B) = ·
20 19
56
=
380
14
= .
95
Hence,
14
P (both workers are against the regulations) = ≈ 0.1474.
95
The two selections are dependent because the first worker is not replaced.
Page 28
Prof. V. K. Narla Module 2: Introduction to Probability
P (A ∩ B) = P (A)P (B).
Explanation
For independent events,
P (A | B) = P (A).
Substituting this relation into the general multiplication rule,
P (A ∩ B) = P (B)P (A | B),
gives
P (A ∩ B) = P (B)P (A).
Thus, when two events are independent, the probability that both events occur is the
product of their individual probabilities.
Conversely, if
P (A ∩ B) = P (A)P (B),
then
P (A ∩ B) P (A)P (B)
P (A | B) = = = P (A),
P (B) P (B)
provided P (B) > 0. Hence, A and B are independent.
Example 4: Two Heads in Two Tosses of a Balanced Coin
Find the probability of obtaining two heads in two tosses of a balanced coin.
Solution.
Let
H1 = {a head occurs on the first toss},
and
H2 = {a head occurs on the second toss}.
For a balanced coin,
1
P (H1 ) =
2
and
1
P (H2 ) = .
2
The two tosses are independent because the result of the first toss does not affect the
result of the second toss.
Page 29
Prof. V. K. Narla Module 2: Introduction to Probability
Therefore,
P (H1 ∩ H2 ) = P (H1 )P (H2 )
1 1
= ·
2 2
1
= .
4
Hence,
1
P (two heads) = .
4
Since the card is replaced, the deck again contains 52 cards with 4 aces. Therefore,
the probability that the second card is also an ace is
4
.
52
Thus,
1
P (two aces with replacement) = .
169
Page 30
Prof. V. K. Narla Module 2: Introduction to Probability
If an ace is drawn first and not replaced, then only 3 aces remain among 51 cards.
Therefore,
3
P (second ace | first ace) = .
51
Using the general multiplication rule,
4 3
P (two aces without replacement) = ·
52 51
12
=
2652
1
= .
221
Hence,
1
P (two aces without replacement) = .
221
Notice that
1 4 4
̸= · .
221 52 52
Therefore, the two selections are not independent when sampling is performed without
replacement.
Example 6: Checking Whether Two Events Are Independent
Suppose
P (C) = 0.65, P (D) = 0.40,
and
P (C ∩ D) = 0.24.
Determine whether the events C and D are independent.
Solution.
If C and D are independent, then they must satisfy
P (C ∩ D) = P (C)P (D).
P (C ∩ D) = 0.24.
Page 31
Prof. V. K. Narla Module 2: Introduction to Probability
Since
0.24 ̸= 0.26,
we conclude that
C and D are not independent.
P (A ∩ B) = P (A)P (B).
P (A ∩ B) = (0.8)(0.7)
= 0.56.
Therefore,
P (A ∩ B) = 0.56.
Thus, the probability that the raw material is available and the machining time is less
than one hour is 0.56.
Page 32
Prof. V. K. Narla Module 2: Introduction to Probability
For repeated independent trials, the probability that the same type of event occurs on
every trial is the product of the corresponding trial probabilities.
Example 8: Extended Special Product Rule
A balanced die is rolled four times. Find the probability of not obtaining a 6 on any of
the four rolls.
Solution.
For one roll of a balanced die,
5
P (not obtaining a 6) = .
6
The four rolls are independent. Therefore,
(︃ )︃4
5
P (no 6 in four rolls) =
6
4
5
= 4
6
625
= .
1296
Hence,
625
P (no 6 in four rolls) = ≈ 0.4823.
1296
P (A ∩ B)
P (A | B) = , P (B) > 0,
P (B)
P (A ∩ B) = P (A)P (B | A), P (A) > 0,
P (A ∩ B) = P (B)P (A | B), P (B) > 0,
P (A ∩ B) = P (A)P (B), if A and B are independent,
n
∏︂
P (A1 ∩ · · · ∩ An ) = P (Ai ), for mutually independent events.
i=1
Page 33
Prof. V. K. Narla Module 2: Introduction to Probability
use the special product rule when the events are independent;
use conditional probability when probability is required under a stated condition.
Page 34
Prof. V. K. Narla Module 2: Introduction to Probability
6. Bayes’ Theorem
Bayes’ theorem is used to revise the probability of a possible cause after observing an out-
come. It combines:
the probability of each possible cause before the new evidence is observed;
the probability of observing the evidence under each possible cause;
the total probability of observing the evidence.
It is therefore an important rule for inverse probability: we begin with an observed
effect and calculate the probability of the cause that produced it.
Page 35
Prof. V. K. Narla Module 2: Introduction to Probability
Derivation
By the definition of conditional probability,
P (Br ∩ A)
P (Br | A) = .
P (A)
Using the multiplication rule in the numerator,
P (Br ∩ A) = P (Br )P (A | Br ).
Using the law of total probability in the denominator,
n
∑︂
P (A) = P (Bi )P (A | Bi ).
i=1
Page 36
Prof. V. K. Narla Module 2: Introduction to Probability
P (A | Bi ).
P (Bi )P (A | Bi ).
8. Divide the route probability associated with the required cause by the total probability
of the evidence.
Page 37
Prof. V. K. Narla Module 2: Introduction to Probability
6.5 Examples
For the next breakdown, the production-line diagnosis indicates that the initial repair
was incomplete. Find the probability that Janet made the initial repair.
Solution.
Let
A = {the initial repair was incomplete}.
Define the technician events:
P (A | B1 ) = 0.05,
P (A | B2 ) = 0.10,
P (A | B3 ) = 0.10,
and
P (A | B4 ) = 0.05.
Page 38
Prof. V. K. Narla Module 2: Introduction to Probability
P (B1 )P (A | B1 )
P (B1 | A) =
P (A)
(0.20)(0.05)
=
(0.20)(0.05) + (0.60)(0.10) + (0.15)(0.10) + (0.05)(0.05)
0.0100
=
0.0875
= 0.1142857.
Hence,
P (Janet made the repair | repair was incomplete) ≈ 0.114.
Thus, approximately 11.4% of all incomplete repairs are attributable to Janet.
Page 39
Prof. V. K. Narla Module 2: Introduction to Probability
Interpretation
Janet services 20% of the breakdowns, but her incomplete-repair rate is only 5%. Tom
services a much larger fraction of the breakdowns and has a higher incomplete-repair rate.
Consequently, after observing an incomplete repair, the probability that Janet was respon-
sible decreases from the prior value 0.20 to the posterior value 0.114.
Example 2: Identifying Spam Using Bayes’ Theorem
A spam-filtering system uses a list of words that occur more frequently in spam messages
than in normal messages.
A database contains 5000 messages, of which:
Page 40
Prof. V. K. Narla Module 2: Introduction to Probability
The probability that a normal message contains words from the list is
297
P (A | B2 ) =
3300
= 0.09.
Step 3: Find the total probability that a message contains words from the list
Using the law of total probability,
P (A) = P (B1 )P (A | B1 ) + P (B2 )P (A | B2 )
= (0.34)(0.79) + (0.66)(0.09)
= 0.2686 + 0.0594
= 0.3280.
Interpretation
Before examining the message, the probability that it is spam is
P (B1 ) = 0.34.
After observing that the message contains words from the list, the probability increases
to
P (B1 | A) = 0.819.
Thus, the observed words provide strong evidence in favour of the message being spam.
However, the probability is not 1, because some normal messages also contain words from
the list.
Page 41
Prof. V. K. Narla Module 2: Introduction to Probability
P (Br )P (A | Br )
P (Br | A) = n
∑︂
P (Bi )P (A | Bi )
i=1
Prior × Likelihood
Posterior =
Total probability of the evidence
Page 42
Prof. V. K. Narla Module 2: Introduction to Probability
P (X ∩ Y ) = 0.30.
However,
0.30 ̸= 0.2475.
Therefore, the condition
P (X ∩ Y ) = P (X)P (Y )
is not satisfied.
Hence,
X and Y are not independent.
P (X | Y ).
P (X ∩ Y )
P (X | Y ) = .
P (Y )
Since
P (X | Y ) = 0.40
but
P (X) = 0.33,
we have
P (X | Y ) ̸= P (X).
This again confirms that X and Y are dependent events.
Page 43
Prof. V. K. Narla Module 2: Introduction to Probability
P (T | C) = 0.7.
The test falsely indicates corrosion when corrosion is absent with probability 0.2. Hence,
P (T | C) = 0.2.
P (T | C) = 1 − P (T | C) = 1 − 0.7 = 0.3,
and
P (T | C) = 1 − P (T | C) = 1 − 0.2 = 0.8.
Page 44
Prof. V. K. Narla Module 2: Introduction to Probability
P (T ) = (0.10)(0.70) + (0.90)(0.20)
= 0.07 + 0.18
= 0.25.
Hence,
P (T ) = 0.25.
Thus, the test indicates corrosion for 25% of the pipe sections.
Therefore,
P (C | T ) = 0.28.
Thus, when the test indicates that corrosion is present, the probability that the pipe
section actually has internal corrosion is
28%.
Page 45
Prof. V. K. Narla Module 2: Introduction to Probability
Interpretation
Although the test correctly detects corrosion 70% of the time when corrosion is present,
the actual prevalence of corrosion is only 10%. Furthermore, the false-positive probability
is 20%. Since non-corroded pipe sections are much more common than corroded sections,
false-positive results form a substantial proportion of all positive test results.
To see this clearly, consider 1000 pipe sections:
Page 46
Prof. V. K. Narla Module 2: Introduction to Probability
Illustration
Suppose a balanced die is rolled. The sample space is
S = {1, 2, 3, 4, 5, 6}.
If X denotes the number appearing on the upper face, then
X(1) = 1, X(2) = 2, ..., X(6) = 6.
Here, the random variable converts each outcome into its corresponding numerical value.
Page 47
Prof. V. K. Narla Module 2: Introduction to Probability
The first condition states that a probability cannot be negative. The second condition
states that one of the possible values of the random variable must occur.
Page 48
Prof. V. K. Narla Module 2: Introduction to Probability
Page 49
Prof. V. K. Narla Module 2: Introduction to Probability
Therefore,
x−2
f (x) = cannot serve as a probability distribution.
2
(b) For
x2
h(x) = , x = 0, 1, 2, 3, 4,
25
all the function values are nonnegative. However, their sum must also equal 1.
We calculate
4 4
∑︂ ∑︂ x2
h(x) =
x=0 x=0
25
02 + 12 + 22 + 32 + 42
=
25
0 + 1 + 4 + 9 + 16
=
25
30
=
25
6
= .
5
Since
6
̸= 1,
5
the total-probability condition is violated.
Therefore,
x2
h(x) = cannot serve as a probability distribution.
25
Thus, F (x) accumulates all probabilities assigned to values of the random variable that
are less than or equal to x.
Page 50
Prof. V. K. Narla Module 2: Introduction to Probability
4.
lim F (x) = 1.
x→∞
For a discrete random variable, the probability at a particular value may be recovered
from the jumps of the cumulative distribution function.
(b)
f (1) = 0.24, f (2) = 0.24, f (3) = 0.24, f (4) = 0.24.
(c)
f (1) = 0.35, f (2) = 0.33, f (3) = 0.34, f (4) = −0.02.
Answer:
For a function f (x) to be a probability distribution of a discrete random variable, it must
satisfy both conditions:
and ∑︂
f (x) = 1.
x
Page 51
Prof. V. K. Narla Module 2: Introduction to Probability
Their sum is
4
∑︂
f (x) = 0.19 + 0.27 + 0.27 + 0.27
x=1
= 0.46 + 0.27 + 0.27
= 0.73 + 0.27
= 1.
Since
0.96 ̸= 1,
the total-probability condition is not satisfied. Therefore,
Page 52
Prof. V. K. Narla Module 2: Introduction to Probability
f (x) ≥ 0
There are four allowed values: 10, 11, 12, and 13. Hence,
∑︂ 1 1 1 1
f (x) = + + +
x
4 4 4 4
(︃ )︃
1
=4
4
= 1.
Therefore,
1
f (x) = , x = 10, 11, 12, 13, is a valid probability distribution.
4
Page 53
Prof. V. K. Narla Module 2: Introduction to Probability
Since
6 ̸= 1,
the function does not satisfy the total-probability condition. Also, for example,
6
f (3) = > 1,
5
which cannot be a probability.
Therefore,
2x
f (x) = is not a valid probability distribution.
5
For example,
8 − 15 7
f (8) = = − < 0,
20 20
and
12 − 15 3
f (12) = = − < 0.
20 20
Thus, the function assigns negative values to the possible outcomes and violates the
non-negativity condition. Therefore,
x − 15
f (x) = is not a valid probability distribution.
20
Page 54
Prof. V. K. Narla Module 2: Introduction to Probability
Since
02 + 12 + 22 + 32 + 42 + 52 = 0 + 1 + 4 + 9 + 16 + 25 = 55,
we obtain
5
∑︂ 6 + 55
f (x) =
x=0
61
61
=
61
= 1.
1 + x2
f (x) = , x = 0, 1, 2, 3, 4, 5, is a valid probability distribution.
61
f (x) = P (X = x).
The probability distribution describes the possible values of X, but a numerical summary
is often needed to describe:
the central or average value of the distribution; and
the amount of variation or spread around that average.
The mean measures the centre of the distribution, whereas the variance and standard
deviation measure its dispersion.
Page 55
Prof. V. K. Narla Module 2: Introduction to Probability
Thus, each possible value x is multiplied by its probability f (x), and the resulting prod-
ucts are added.
The mean need not be one of the possible values of the random variable. It represents
the long-run average value obtained when the experiment is repeated many times.
The variance is expressed in squared units, while the standard deviation is expressed in
the same units as the random variable.
Since ∑︂
E(X 2 ) = x2 f (x),
all x
8.5 Examples
Example 1: Number of Heads in Three Tosses of a Fair Coin
A fair coin is tossed three times. Let X denote the number of heads obtained. Find the
mean, variance, and standard deviation of X.
Page 56
Prof. V. K. Narla Module 2: Introduction to Probability
Solution.
The possible values of X are
0, 1, 2, 3.
The probability distribution is
x 0 1 2 3
1 3 3 1
f (x)
8 8 8 8
we obtain (︃ )︃ (︃ )︃ (︃ )︃ (︃ )︃
1 3 3 1
µ=0 +1 +2 +3
8 8 8 8
3 6 3
=0+ + +
8 8 8
12
=
8
3
= .
2
Therefore,
3
µ= = 1.5.
2
Page 57
Prof. V. K. Narla Module 2: Introduction to Probability
Thus,
3
σ2 = = 0.75.
4
Hence,
µ = 1.5, σ 2 = 0.75, σ ≈ 0.866.
x 0 1 2 3
f (x) 0.18 0.50 0.29 0.03
Find the mean, variance, and standard deviation of X.
Solution.
Therefore,
µ = 1.17.
Page 58
Prof. V. K. Narla Module 2: Introduction to Probability
σ 2 = E(X 2 ) − µ2
= 1.93 − (1.17)2
= 1.93 − 1.3689
= 0.5611.
Therefore,
σ 2 = 0.5611.
Hence,
µ = 1.17, σ 2 = 0.5611, σ ≈ 0.7491.
Page 59
Prof. V. K. Narla Module 2: Introduction to Probability
Summary of Formulas
∑︂
µ = E(X) = x f (x)
all x
∑︂
E(X 2 ) = x2 f (x)
all x
∑︂
σ2 = (x − µ)2 f (x) = E(X 2 ) − µ2
all x
√
σ= σ2
Page 60
Prof. V. K. Narla Module 2: Introduction to Probability
P (X = x) = 0
Therefore, in the continuous case, the distinction between < and ≤ does not affect an
interval probability.
Page 61
Prof. V. K. Narla Module 2: Introduction to Probability
as well.
The function f (x) itself is not a probability. The probability is the area under f (x) over
an interval.
Condition 1: Non-negativity
This condition states that the total probability of all possible values of the random
variable is 1.
Thus, the two defining conditions are
f (x) ≥ 0
and ∫︂ ∞
f (x) dx = 1.
−∞
F (x) = P (X ≤ x).
Thus, F (x) is the total area under the density curve to the left of x.
Page 62
Prof. V. K. Narla Module 2: Introduction to Probability
P (a ≤ X ≤ b) = F (b) − F (a).
dF (x)
= f (x).
dx
Hence, the cumulative distribution function is obtained by integrating the density, while
the density is obtained by differentiating the cumulative distribution function.
2. F (x) is nondecreasing;
3.
lim F (x) = 0;
x→−∞
4.
lim F (x) = 1.
x→∞
9.6 Examples
Example 1: Calculating Probabilities from a Probability Density Function
Suppose a continuous random variable X has the probability density function
{︄ −2x
2e , x > 0,
f (x) =
0, x ≤ 0.
Find:
(a) P (1 ≤ X ≤ 3);
(b) P (X > 0.5).
Page 63
Prof. V. K. Narla Module 2: Introduction to Probability
Solution.
(a) Probability that X lies between 1 and 3
Since f (x) = 2e−2x for x > 0,
∫︂ 3
P (1 ≤ X ≤ 3) = 2e−2x dx.
1
Because ∫︂
2e−2x dx = −e−2x ,
we obtain ]︁3
P (1 ≤ X ≤ 3) = −e−2x 1
[︁
= −e−6 + e−2
= e−2 − e−6
≈ 0.1353 − 0.0025
≈ 0.1329.
Therefore,
P (1 ≤ X ≤ 3) ≈ 0.133.
(b) Probability that X > 0.5
∫︂ ∞
P (X > 0.5) = 2e−2x dx.
0.5
Hence, ]︁∞
P (X > 0.5) = −e−2x 0.5
[︁
= 0 − (−e−1 )
= e−1
≈ 0.3679.
Therefore,
P (X > 0.5) ≈ 0.368.
P (X ≤ 1).
Solution.
Page 64
Prof. V. K. Narla Module 2: Introduction to Probability
By definition, ∫︂ x
F (x) = f (t) dt.
−∞
F (x) = 0.
Case 2: x > 0
∫︂ x
F (x) = 2e−2t dt
[︁ 0 −2t ]︁x
= −e 0
= −e−2x + 1
= 1 − e−2x .
Therefore,
{︄
0, x ≤ 0,
F (x) =
1 − e−2x , x > 0.
Now,
P (X ≤ 1) = F (1)
= 1 − e−2
≈ 1 − 0.1353
≈ 0.8647.
Hence,
P (X ≤ 1) ≈ 0.865.
Page 65
Prof. V. K. Narla Module 2: Introduction to Probability
Let
u = 4x2 .
Then
du
du = 8x dx =⇒ x dx = .
8
Therefore,
∞
k ∞ −u
∫︂ ∫︂
−4x2
kxe dx = e du
0 8 0
k [︁ −u ]︁∞
= −e 0
8
k
= (1)
8
k
= .
8
The total area must equal 1, so
k
= 1.
8
Hence,
k = 8.
Thus, the required density is
{︄
0, x ≤ 0,
f (x) = 2
8xe−4x , x > 0.
Page 66
Prof. V. K. Narla Module 2: Introduction to Probability
In particular,
µ′1 = E(X) = µ
and ∫︂ ∞
µ′2 = E(X ) = 2
x2 f (x) dx.
−∞
The kth moment about the mean is
∫︂ ∞
k
µk = E[(X − µ) ] = (x − µ)k f (x) dx.
−∞
Therefore,
σ 2 = E(X 2 ) − [E(X)]2 .
The standard deviation is √
σ= σ2.
Page 67
Prof. V. K. Narla Module 2: Introduction to Probability
u = x, dv = 2e−2x dx.
Then
du = dx, v = −e−2x .
Therefore, ∫︂ ∞
−2x ∞
e−2x dx
[︁ ]︁
µ = −xe 0
+
[︃ ]︃∞0
1 −2x
=0+ − e
2 0
1
= .
2
Hence,
1
µ= .
2
Step 2: Calculate E(X 2 )
∫︂ ∞
E(X ) = 2
2x2 e−2x dx.
0
Using the standard integral
∫︂ ∞
n!
xn e−ax dx = , a > 0,
0 an+1
with n = 2 and a = 2, ∫︂ ∞
2! 1
x2 e−2x dx = 3
= .
0 2 4
Therefore, (︃ )︃
2 1 1
E(X ) = 2 = .
4 2
Step 3: Calculate the variance
σ 2 = E(X 2 ) − µ2
(︃ )︃2
1 1
= −
2 2
1 1
= −
2 4
1
= .
4
Thus,
1
σ2 = .
4
Page 68
Prof. V. K. Narla Module 2: Introduction to Probability
P (a ≤ X ≤ b) = F (b) − F (a).
dF (x)
= f (x).
dx
∫︂ ∞
µ = E(X) = xf (x) dx.
−∞
∫︂ ∞
E(X 2 ) = x2 f (x) dx.
−∞
∫︂ ∞
σ2 = (x − µ)2 f (x) dx = E(X 2 ) − µ2 .
−∞
√
σ= σ2.
Page 69
Prof. V. K. Narla Module 2: Introduction to Probability
3
(a) greater than ;
4
1 2
(b) between and .
3 3
Answer:
For f (x) to be a probability density function, it must satisfy
∫︂ ∞
f (x) dx = 1.
−∞
Therefore,
1 ]︃1
x4
∫︂ [︃
3
(k + 2x ) dx = kx +
0 2 0
1
=k+ .
2
Hence,
1
k+ = 1,
2
so that
1
k= .
2
Thus, the density is ⎧
⎨ 1 + 2x3 , 0 < x < 1,
f (x) = 2
⎩0, elsewhere.
3
(a) Probability that X >
4
(︃ )︃ ∫︂ 1 (︃ )︃
3 1
P X> = + 2x3 dx.
4 3/4 2
Now,
x x4
∫︂ (︃ )︃
1
+ 2x3 dx = + .
2 2 2
Page 70
Prof. V. K. Narla Module 2: Introduction to Probability
Therefore,
]︃1
x x4
(︃ )︃ [︃
3
P X> = +
4 2 2 3/4
(︃ )︃ (︄ (︃ )︃4 )︄
1 1 3 1 3
= + − +
2 2 8 2 4
(︃ )︃
3 81
=1− +
8 512
273
=1−
512
239
= .
512
Hence,
(︃ )︃
3 239
P X> = ≈ 0.4668.
4 512
1 2
(b) Probability that <X<
3 3
(︃ )︃ ∫︂ 2/3 (︃ )︃
1 2 1 3
P <X< = + 2x dx.
3 3 1/3 2
Thus,
]︃2/3
x x4
(︃ )︃ [︃
1 2
P <X< = +
3 3 2 2 1/3
(︃ )︃ (︃ )︃
1 8 1 1
= + − +
3 81 6 162
35 14
= −
81 81
21
=
81
7
= .
27
Therefore,
(︃ )︃
1 2 7
P <X< = ≈ 0.2593.
3 3 27
Page 71
Prof. V. K. Narla Module 2: Introduction to Probability
Question 2
If the probability density of a random variable is given by
⎧
⎪
⎪ x, 0 < x < 1,
⎨
f (x) = 2 − x, 1 ≤ x < 2,
⎪
⎪
0, elsewhere,
⎩
find the probabilities that a random variable having this probability density will take on a
value
(a) between 0.2 and 0.8;
(b) between 0.6 and 1.2.
Answer:
The density changes its formula at x = 1. Therefore, an interval that crosses x = 1 must
be divided into two integrals.
f (x) = x.
Hence, ∫︂ 0.8
P (0.2 < X < 0.8) = x dx
0.2
[︃ 2 ]︃0.8
x
=
2 0.2
0.82 − 0.22
=
2
0.64 − 0.04
=
2
0.60
=
2
= 0.30.
Therefore,
P (0.2 < X < 0.8) = 0.30.
Page 72
Prof. V. K. Narla Module 2: Introduction to Probability
Therefore,
P (0.6 < X < 1.2) = 0.32 + 0.18
= 0.50.
Hence,
P (0.6 < X < 1.2) = 0.50.
Question 3
Given the probability density
k
f (x) = , −∞ < x < ∞,
1 + x2
find k.
Answer:
A probability density must satisfy
∫︂ ∞
f (x) dx = 1.
−∞
Therefore, ∫︂ ∞
k
dx = 1.
−∞ 1 + x2
Taking k outside the integral,
∫︂ ∞
1
k dx = 1.
−∞ 1 + x2
Page 73
Prof. V. K. Narla Module 2: Introduction to Probability
Since ∫︂
1
2
dx = tan−1 x,
1+x
we obtain ∫︂ ∞
1 [︁ −1 ]︁∞
k dx = k tan x −∞
−∞ 1 + x2
(︂ π (︂ π )︂)︂
=k − −
2 2
= kπ.
Thus,
kπ = 1,
and hence
1
k= .
π
Therefore, the normalized density is
1
f (x) = , −∞ < x < ∞.
π(1 + x2 )
Question 4
If the distribution function of a random variable is given by
⎨1 − 4 , x > 2,
⎧
F (x) = x2
0, x ≤ 2,
⎩
find the probabilities that this random variable will take on a value
(a) less than 3;
(b) between 4 and 5.
Answer:
For a continuous random variable,
F (x) = P (X ≤ x).
Also,
P (a < X < b) = F (b) − F (a).
Page 74
Prof. V. K. Narla Module 2: Introduction to Probability
P (X < 3) = P (X ≤ 3) = F (3).
Therefore,
4
P (X < 3) = 1 −
32
4
=1−
9
5
= .
9
Hence,
5
P (X < 3) = ≈ 0.5556.
9
Question 5
Let the phase error in a tracking device have probability density
f (x) = 2
⎩0, elsewhere.
Page 75
Prof. V. K. Narla Module 2: Introduction to Probability
π
(b) greater than .
3
Answer:
First, observe that
∫︂ π/2
π/2
cos x dx = [sin x]0 = 1,
0
so the given function is a valid probability density.
π
(a) Probability that 0 < X <
4
∫︂ π/4
(︂π )︂
P 0<X< = cos x dx
4 0
π/4
= [sin x]0
π
= sin − sin 0
√ 4
2
= .
2
Hence,
√
(︂ π )︂ 2
P 0<X< = ≈ 0.7071.
4 2
π
(b) Probability that X >
3
∫︂ π/2
π )︂
(︂
P X> = cos x dx
3 π/3
π/2
= [sin x]π/3
√
3
=1− .
2
Therefore,
√
π )︂
(︂ 3
P X> =1− ≈ 0.1340.
3 2
Question 6
The length of satisfactory service, in years, provided by a certain model of laptop computer
is a random variable having the probability density
⎨ 1 e−x/4.5 , x > 0,
⎧
f (x) = 4.5
0, x ≤ 0.
⎩
Find the probabilities that one of these laptops will provide satisfactory service for
Page 76
Prof. V. K. Narla Module 2: Introduction to Probability
P (X > x) = e−x/4.5 .
= e−4/4.5 − e−6/4.5
= e−8/9 − e−4/3
≈ 0.4111 − 0.2636
≈ 0.1475.
Hence,
P (4 ≤ X ≤ 6) ≈ 0.1475.
Page 77
Prof. V. K. Narla Module 2: Introduction to Probability
P (X ≥ 6.75) = e−6.75/4.5
= e−1.5
≈ 0.2231.
Therefore,
P (X ≥ 6.75) ≈ 0.2231.
Module 2
END OF THE UNIT
INTRODUCTION TO PROBABILITY
Review the key concepts, formulas, examples, and practice problems before proceeding to the next
unit.
Page 78