0% found this document useful (0 votes)
3 views78 pages

P&S Module-2 Notes

Module 2 of 'Probability and Statistics for Engineers' covers the fundamentals of probability and random variables, including sample spaces, events, and various interpretations of probability. Key concepts such as the axioms of probability, conditional probability, and Bayes' theorem are introduced, along with examples and applications. The module also discusses random variables, probability distributions, and provides practice questions to reinforce understanding.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views78 pages

P&S Module-2 Notes

Module 2 of 'Probability and Statistics for Engineers' covers the fundamentals of probability and random variables, including sample spaces, events, and various interpretations of probability. Key concepts such as the axioms of probability, conditional probability, and Bayes' theorem are introduced, along with examples and applications. The module also discusses random variables, probability distributions, and provides practice questions to reinforce understanding.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

MODULE 2

Introduction to Probability and Random


Variables
Prof. V. K. Narla
MILLER & FREUND’S
PROBABILITY AND STATISTICS FOR ENGINEERS

Contents

1 Sample Spaces and Events 3


1.1 Experiment, outcome, and sample space . . . . . . . . . . . . . . . . . . . . 3
1.2 Illustration: awarding a construction contract . . . . . . . . . . . . . . . . . 3
1.3 Illustration: locating research facilities . . . . . . . . . . . . . . . . . . . . . 3
1.4 Discrete and continuous sample spaces . . . . . . . . . . . . . . . . . . . . . 4
1.5 Events . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4

2 Probability 4
2.1 Classical probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4
2.2 Frequency interpretation of probability . . . . . . . . . . . . . . . . . . . . . 6
2.3 Subjective probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7

3 The Axioms of Probability 7


3.1 Axiom 1: Non-negativity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
3.2 Axiom 2: Probability of the sample space . . . . . . . . . . . . . . . . . . . . 7
3.3 Axiom 3: Addition for mutually exclusive events . . . . . . . . . . . . . . . . 7

4 Basic Theorems of Probability 8


4.1 Generalization of the third axiom . . . . . . . . . . . . . . . . . . . . . . . . 8
4.2 Probability of an event in a finite sample space . . . . . . . . . . . . . . . . . 10
4.3 General addition rule . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10
4.4 Probability rule of the complement . . . . . . . . . . . . . . . . . . . . . . . 13

1
Prof. V. K. Narla Module 2: Introduction to Probability

5 Conditional Probability, Multiplication Rules, and Independence 25


5.1 Conditional Probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
5.2 Independent Events . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
5.3 General Multiplication Rule of Probability . . . . . . . . . . . . . . . . . . . 27
5.4 Special Product Rule for Independent Events . . . . . . . . . . . . . . . . . . 29
5.5 Assigning Probabilities by the Special Product Rule . . . . . . . . . . . . . . 32
5.6 Extended Product Rule for Several Independent Events . . . . . . . . . . . . 33

6 Bayes’ Theorem 35
6.1 Partition of the Sample Space . . . . . . . . . . . . . . . . . . . . . . . . . . 35
6.2 Statement and Derivation of Bayes’ Theorem . . . . . . . . . . . . . . . . . . 36
6.3 Interpretation of the Formula . . . . . . . . . . . . . . . . . . . . . . . . . . 36
6.4 Procedure for Applying Bayes’ Theorem . . . . . . . . . . . . . . . . . . . . 37
6.5 Examples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38
6.6 Important Observations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42
6.7 Formula Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42

7 Random Variables and Probability Distributions 47


7.1 Random Variables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 47
7.2 Classification of Random Variables . . . . . . . . . . . . . . . . . . . . . . . 47
7.3 Probability Distribution of a Discrete Random Variable . . . . . . . . . . . . 48
7.4 Cumulative Distribution Function . . . . . . . . . . . . . . . . . . . . . . . . 50

8 Practice Questions and Answers 51


8.1 The Mean and Variance of a Discrete Distribution . . . . . . . . . . . . . . . 55
8.2 Mean or Expected Value . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 56
8.3 Variance and Standard Deviation . . . . . . . . . . . . . . . . . . . . . . . . 56
8.4 Alternative Computing Formula for Variance . . . . . . . . . . . . . . . . . . 56
8.5 Examples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 56

9 Continuous Random Variables and Probability Density Functions 61


9.1 Probability at an Individual Point . . . . . . . . . . . . . . . . . . . . . . . . 61
9.2 Probability Density Function . . . . . . . . . . . . . . . . . . . . . . . . . . . 61
9.3 Conditions for a Probability Density Function . . . . . . . . . . . . . . . . . 62
9.4 Cumulative Distribution Function . . . . . . . . . . . . . . . . . . . . . . . . 62
9.5 Properties of the Cumulative Distribution Function . . . . . . . . . . . . . . 63
9.6 Examples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 63
9.7 Mean of a Continuous Probability Distribution . . . . . . . . . . . . . . . . . 66
9.8 Moments of a Continuous Distribution . . . . . . . . . . . . . . . . . . . . . 67
9.9 Variance and Standard Deviation . . . . . . . . . . . . . . . . . . . . . . . . 67

Page 2
Prof. V. K. Narla Module 2: Introduction to Probability

1. Sample Spaces and Events


Probability provides a numerical description of the variability associated with the outcome
of an experiment whose exact result cannot be predicted with certainty. Before probability
can be introduced, it is necessary to identify the possible outcomes of the experiment and
the events to which probabilities will be assigned.

1.1 Experiment, outcome, and sample space


In statistics, the terms experiment and outcome are used in a broad sense. An experiment
may be as simple as observing whether a switch is on or off, or it may involve a measurement
such as the time required for a car to accelerate to a specified speed. It may also involve a
complicated scientific measurement.
The result obtained from an experiment is called an outcome. The complete set of all
possible outcomes is called the sample space and is denoted by S.
Thus,
S = {all possible outcomes of the experiment}.

1.2 Illustration: awarding a construction contract


Suppose four contractors submit bids for a highway construction project. Let
a, b, c, d
denote the events that the contract is awarded to Mr. Adam, Mrs. Brown, Mr. Clark, and
Ms. Dean, respectively. The sample space is
S = {a, b, c, d}.
Each element of the sample space represents one possible outcome of the experiment.

1.3 Illustration: locating research facilities


Suppose a government agency must decide where to locate two new research facilities. The
possible locations are Texas and California. Let the first coordinate denote the number of
facilities located in Texas and the second coordinate denote the number located in California.
The sample space is
S = {(0, 0), (1, 0), (0, 1), (2, 0), (1, 1), (0, 2)}.
For example:
(2, 0)
means that two facilities are located in Texas and none in California, while
(1, 1)
means that one facility is located in each state.
Geometrically, the outcomes may be represented as points in the coordinate plane.

Page 3
Prof. V. K. Narla Module 2: Introduction to Probability

1.4 Discrete and continuous sample spaces


A sample space is called a discrete sample space when it contains either finitely many
outcomes or a countably infinite number of outcomes.
Examples include:
S = {1, 2, 3, 4, 5, 6}
for the outcome of a die roll, and

S = {0, 1, 2, . . .}

for the number of occurrences of a particular event.


A sample space is called a continuous sample space when its outcomes form a con-
tinuum. Examples include:
ˆ all points on a line;
ˆ all points on a line segment;
ˆ all points in a plane;
ˆ all possible values of a measured physical quantity.

1.5 Events
Any subset of a sample space is called an event. An event may contain one outcome, several
outcomes, the entire sample space, or no outcomes.
The set containing no outcomes is called the empty set and is denoted by

∅.

Thus, if A ⊆ S, then A is an event associated with the sample space S.

2. Probability
After identifying what is possible in an experiment, the next step is to describe what is
probable or improbable. Three common interpretations of probability are:
1. the classical interpretation;
2. the frequency interpretation;
3. the subjective interpretation.

2.1 Classical probability


The classical probability concept applies when all possible outcomes are equally likely.
Classical probability concept

Page 4
Prof. V. K. Narla Module 2: Introduction to Probability

If there are m equally likely possibilities, one of which must occur, and s of them are
favourable to an event, then the probability of that event is
s
P (A) = .
m
Equivalently,
number of favourable outcomes
P (A) = .
total number of equally likely outcomes
The terms favourable and success are used in a technical sense. A favourable outcome
does not necessarily represent something desirable.
Example 1: Well-shuffled cards are equally likely to be selected
What is the probability of drawing an ace from a well-shuffled deck of 52 playing cards?
Solution. There are 4 aces among the 52 cards. Hence,

m = 52, s = 4.

Therefore,
s 4 1
P (ace) = = = .
m 52 13
The assumption of equally likely outcomes is commonly used in games of chance and also
in carefully designed random-selection procedures.
Example 2: Random selection in the equally likely case
Ten miniature electric motors are available. Tests show that 8 motors operate satisfac-
torily and 2 do not. Two motors are selected at random, with every pair having the same
chance of being selected.
Find the probability that:
(a) both motors operate satisfactorily;
(b) one motor operates satisfactorily and the other does not.
Solution. The total number of possible pairs is
(︃ )︃
10
m= = 45.
2

(a) Both motors operate satisfactorily


The number of ways of selecting two satisfactory motors from the eight satisfactory
motors is (︃ )︃
8
s= = 28.
2
Therefore,
28
≈ 0.622.
P (both satisfactory) =
45
(b) One satisfactory and one unsatisfactory motor

Page 5
Prof. V. K. Narla Module 2: Introduction to Probability

The number of ways of selecting one satisfactory motor and one unsatisfactory motor is
(︃ )︃(︃ )︃
8 2
s= = 8(2) = 16.
1 1

Hence,
16
P (one satisfactory and one unsatisfactory) = ≈ 0.356.
45

2.2 Frequency interpretation of probability


The classical approach has limited applicability because many practical situations do not
have equally likely outcomes. In such situations, probability may be interpreted through
repeated experimentation.
Frequency interpretation of probability
The probability of an event is interpreted as the long-run proportion of times that the
event occurs in repeated trials.
If an event A occurs nA times in N repetitions of an experiment, its relative frequency is
nA
rN = .
N
Equivalently,
number of times A occurs in N trials
rN = .
N
As the number of trials becomes large, the relative frequency tends to stabilize near the
probability of the event.
For example, if a weather service states that the probability of rain is 0.40, this means
that under similar weather conditions, rain is expected to occur approximately 40% of the
time.
Example 3: Long-run relative frequency approximation to proba-
bility
Records show that 294 of 300 ceramic insulators tested were able to withstand a specified
thermal shock. Estimate the probability that an untested insulator will withstand the same
thermal shock.
Solution. The observed relative frequency is
294
= 0.98.
300
Therefore, the estimated probability that an untested insulator will withstand the thermal
shock is
P (withstanding the shock) ≈ 0.98.

Page 6
Prof. V. K. Narla Module 2: Introduction to Probability

2.3 Subjective probability


A third interpretation treats probability as a personal or subjective evaluation.
Subjective probabilities express the strength of a person’s belief concerning uncertain
outcomes. They are especially useful when there is little or no direct experimental evidence.
Such probabilities may be based on:
ˆ indirect evidence;
ˆ educated judgement;
ˆ experience;
ˆ intuition;
ˆ other relevant information.
Subjective probability is frequently used in risk analysis, forecasting, and decision-making
under uncertainty.

3. The Axioms of Probability


Probability may be defined mathematically as a set function that assigns a number P (A) to
each event A in a sample space S.
For a finite sample space, the probability assignment must satisfy the following three
axioms.

3.1 Axiom 1: Non-negativity


For every event A in S,
0 ≤ P (A) ≤ 1.
Thus, a probability cannot be negative and cannot exceed 1.

3.2 Axiom 2: Probability of the sample space


The probability of the entire sample space is

P (S) = 1.

Since one of the possible outcomes must occur, the probability of the certain event is 1.

3.3 Axiom 3: Addition for mutually exclusive events


If A and B are mutually exclusive events in S, then

P (A ∪ B) = P (A) + P (B).

Mutually exclusive events cannot occur simultaneously.

Page 7
Prof. V. K. Narla Module 2: Introduction to Probability

Example 4: Checking possible assignments of probability


An experiment has three possible and mutually exclusive outcomes A, B, and C. Deter-
mine whether each of the following probability assignments is permissible.
(a)
1 1 1
P (A) = , P (B) = , P (C) = .
3 3 3
(b)
P (A) = 0.64, P (B) = 0.38, P (C) = −0.02.
(c)
P (A) = 0.35, P (B) = 0.52, P (C) = 0.26.
(d)
P (A) = 0.57, P (B) = 0.24, P (C) = 0.19.
Solution.
(a) The assignment is permissible because all values lie between 0 and 1, and
1 1 1
+ + = 1.
3 3 3
(b) The assignment is not permissible because
P (C) = −0.02 < 0.
This violates the non-negativity axiom.
(c) The assignment is not permissible because
0.35 + 0.52 + 0.26 = 1.13 > 1.
The total probability must equal 1.
(d) The assignment is permissible because all probabilities lie between 0 and 1, and
0.57 + 0.24 + 0.19 = 1.

4. Basic Theorems of Probability


The three axioms lead to several rules that are used repeatedly in probability calculations.
The following results are direct extensions and consequences of the axioms.

4.1 Generalization of the third axiom


Theorem 3.4. Addition of probabilities for mutually exclusive events
If A1 , A2 , . . . , An are mutually exclusive events in a sample space S, then
P (A1 ∪ A2 ∪ · · · ∪ An ) = P (A1 ) + P (A2 ) + · · · + P (An ).
Here, mutually exclusive means that no two of the events can occur simultaneously:
Ai ∩ Aj = ∅, i ̸= j.

Page 8
Prof. V. K. Narla Module 2: Introduction to Probability

Explanation
The third axiom gives the result for two mutually exclusive events:

P (A1 ∪ A2 ) = P (A1 ) + P (A2 ).

Because A1 ∪ A2 is also mutually exclusive of A3 ,

P (A1 ∪ A2 ∪ A3 ) = P (A1 ∪ A2 ) + P (A3 )


= P (A1 ) + P (A2 ) + P (A3 ).

Repeating this argument gives the result for any finite number n of mutually exclusive
events.
Example 5: Probabilities add for mutually exclusive events
A consumer testing service assigns the following probabilities to the ratings of a new
antipollution device for cars:

Rating Very poor Poor Fair Good Very good Excellent


Probability 0.07 0.12 0.17 0.32 0.21 0.11

Find the probability that the service rates the device:


(a) very poor, poor, fair, or good;
(b) good, very good, or excellent.
Solution.
Each device receives exactly one rating. Therefore, the rating events are mutually exclu-
sive, and their probabilities may be added directly.
(a) Very poor, poor, fair, or good
Let A denote this event. Then

A = {very poor} ∪ {poor} ∪ {fair} ∪ {good}.

By Theorem 3.4,
P (A) = 0.07 + 0.12 + 0.17 + 0.32
= 0.19 + 0.17 + 0.32
= 0.36 + 0.32
= 0.68.
Thus,
P (very poor, poor, fair, or good) = 0.68.
(b) Good, very good, or excellent
Let B denote this event. Then

B = {good} ∪ {very good} ∪ {excellent}.

Page 9
Prof. V. K. Narla Module 2: Introduction to Probability

Therefore,
P (B) = 0.32 + 0.21 + 0.11
= 0.53 + 0.11
= 0.64.
Hence,
P (good, very good, or excellent) = 0.64.

4.2 Probability of an event in a finite sample space

Theorem 3.5. Rule for calculating the probability of an event


If A is an event in a finite sample space S, then P (A) equals the sum of the probabilities
of the individual outcomes that constitute A.
Suppose
A = {E1 , E2 , . . . , En },
where E1 , E2 , . . . , En are the individual outcomes in A. Since distinct individual outcomes
are mutually exclusive,
A = E1 ∪ E2 ∪ · · · ∪ En .
Using Theorem 3.4,

P (A) = P (E1 ) + P (E2 ) + · · · + P (En ).

Detailed justification
Only one individual outcome can occur in a single performance of the experiment. Hence,

Ei ∩ Ej = ∅ for i ̸= j.

The event A occurs precisely when one of the outcomes E1 , E2 , . . . , En occurs. Therefore,

P (A) = P (E1 ∪ E2 ∪ · · · ∪ En ),

and the generalized addition rule gives


n
∑︂
P (A) = P (Ei ).
i=1

4.3 General addition rule

Theorem 3.6. General addition rule for probability


If A and B are any two events in S, then

P (A ∪ B) = P (A) + P (B) − P (A ∩ B).

Page 10
Prof. V. K. Narla Module 2: Introduction to Probability

Derivation
The union A ∪ B can be divided into the following three mutually exclusive parts:

A ∩ B, A ∩ B, A ∩ B.

Thus,
A ∪ B = (A ∩ B) ∪ (A ∩ B) ∪ (A ∩ B).
By Theorem 3.4,

P (A ∪ B) = P (A ∩ B) + P (A ∩ B) + P (A ∩ B).

Also,
P (A) = P (A ∩ B) + P (A ∩ B),
and
P (B) = P (A ∩ B) + P (A ∩ B).
Adding the last two equations gives

P (A) + P (B) = 2P (A ∩ B) + P (A ∩ B) + P (A ∩ B).

Comparing this with the expression for P (A ∪ B), we obtain

P (A) + P (B) = P (A ∪ B) + P (A ∩ B).

Therefore,
P (A ∪ B) = P (A) + P (B) − P (A ∩ B).
The intersection probability is subtracted because outcomes in A ∩ B are counted once
in P (A) and once again in P (B).
If A and B are mutually exclusive, then

P (A ∩ B) = 0,

and the general addition rule reduces to the special addition rule

P (A ∪ B) = P (A) + P (B).

Example 6: Using the general addition rule for probability


In a used-car classification, let:

M1 = {the car has low mileage},

and
C3 = {the car is expensive to operate}.
Suppose
P (M1 ) = 0.20, P (C3 ) = 0.40, P (M1 ∩ C3 ) = 0.08.

Page 11
Prof. V. K. Narla Module 2: Introduction to Probability

Find the probability that a car has low mileage or is expensive to operate; that is, find
P (M1 ∪ C3 ).
Solution.
The word “or” indicates the union of the two events. Since a car may have both low
mileage and high operating cost, the events are not necessarily mutually exclusive. Hence,
Theorem 3.6 must be used:
P (M1 ∪ C3 ) = P (M1 ) + P (C3 ) − P (M1 ∩ C3 ).
Substituting the given values,
P (M1 ∪ C3 ) = 0.20 + 0.40 − 0.08
= 0.60 − 0.08
= 0.52.
Therefore,
P (low mileage or expensive to operate) = 0.52.

Example 7: Probability of requiring repair under warranty


For a new car under warranty, suppose:
P (E) = 0.87,
where E is the event that engine repairs are required,
P (D) = 0.36,
where D is the event that drive-train repairs are required, and
P (E ∩ D) = 0.29.
Find the probability that the car requires repairs to the engine, the drive train, or both
during the warranty period.
Solution.
The phrase “engine, drive train, or both” describes the union
E ∪ D.
Because a car can require both types of repairs, the events overlap. Therefore,
P (E ∪ D) = P (E) + P (D) − P (E ∩ D).
Substituting the data,
P (E ∪ D) = 0.87 + 0.36 − 0.29
= 1.23 − 0.29
= 0.94.
Hence,
P (engine or drive-train repairs, or both) = 0.94.
The subtraction of 0.29 prevents cars requiring both types of repair from being counted
twice.

Page 12
Prof. V. K. Narla Module 2: Introduction to Probability

4.4 Probability rule of the complement

Theorem 3.7. Probability of the complement of an event


If A is any event in S, then
P (A) = 1 − P (A).

Proof
The event A and its complement A are mutually exclusive:

A ∩ A = ∅.

Together, they contain every outcome of the sample space:

A ∪ A = S.

By the addition axiom,


P (A ∪ A) = P (A) + P (A).
Since A ∪ A = S and P (S) = 1,

P (A) + P (A) = 1.

Therefore,
P (A) = 1 − P (A).
As a special case,
P (∅) = 1 − P (S) = 1 − 1 = 0.

Example 8: Using the probability rule of the complement


Using the used-car events from Example 17, where

P (M1 ) = 0.20 and P (M1 ∩ C3 ) = 0.08,

find:
(a) the probability that a used car does not have low mileage;
(b) the probability that a used car either does not have low mileage or is not expensive
to operate.
Solution.
(a) The car does not have low mileage
The event “does not have low mileage” is M1 . By the complement rule,

P (M1 ) = 1 − P (M1 )
= 1 − 0.20
= 0.80.

Page 13
Prof. V. K. Narla Module 2: Introduction to Probability

Thus,
P (M1 ) = 0.80.
(b) The car either does not have low mileage or is not expensive to operate
The required event is
M1 ∪ C3 .
By De Morgan’s law,
M1 ∪ C3 = M1 ∩ C3 .
Therefore, by the complement rule,
(︁ )︁
P (M1 ∪ C3 ) = P M1 ∩ C3
= 1 − P (M1 ∩ C3 )
= 1 − 0.08
= 0.92.

Hence,
P (M1 ∪ C3 ) = 0.92.

Summary of the probability rules

n
∑︂
P (A1 ∪ A2 ∪ · · · ∪ An ) = P (Ai ), if the events are mutually exclusive,
i=1
∑︂
P (A) = P (Ei ), for a finite sample space,
Ei ∈A

P (A ∪ B) = P (A) + P (B) − P (A ∩ B), for any two events,


P (A) = 1 − P (A), for any event A.

Practice Questions and Answers


Question 1: Probabilities for the Sum of Two Balanced Dice
A pair of balanced dice is rolled. Find the probability of obtaining:
(a) a sum of 7;
(b) a sum of 11;
(c) a sum of 7 or 11;
(d) a sum of 3;
(e) a sum of 2 or 12;
(f) a sum of 2, 3, or 12.

Page 14
Prof. V. K. Narla Module 2: Introduction to Probability

Answer.
When two balanced dice are rolled, each die has 6 possible outcomes. Therefore, the
total number of equally likely ordered outcomes is
6 × 6 = 36.
(a) A sum of 7
The favourable outcomes are
(1, 6), (2, 5), (3, 4), (4, 3), (5, 2), (6, 1).
There are 6 favourable outcomes. Hence,
6 1
P (sum 7) = = .
36 6
(b) A sum of 11
The favourable outcomes are
(5, 6), (6, 5).
Thus,
2 1
P (sum 11) = = .
36 18
(c) A sum of 7 or 11
The events “sum 7” and “sum 11” are mutually exclusive. Therefore,
P (sum 7 or 11) = P (sum 7) + P (sum 11)
6 2
= +
36 36
8
=
36
2
= .
9
(d) A sum of 3
The favourable outcomes are
(1, 2), (2, 1).
Therefore,
2 1
P (sum 3) = = .
36 18
(e) A sum of 2 or 12
The sum 2 occurs only for (1, 1), and the sum 12 occurs only for (6, 6). Hence,
1+1 2 1
P (sum 2 or 12) = = = .
36 36 18
(f) A sum of 2, 3, or 12
The numbers of favourable outcomes are 1, 2, and 1, respectively. Thus,
1+2+1 4 1
P (sum 2, 3, or 12) = = = .
36 36 9

Page 15
Prof. V. K. Narla Module 2: Introduction to Probability

Question 2: Registration Number Divisible by 40


The registration numbers for the candidates of an entrance test are numbered from 000001
to 200000. Find the probability that a candidate receives a registration number divisible by
40.
Answer.
The total number of possible registration numbers is
200000.
The registration numbers divisible by 40 are
40, 80, 120, . . . , 200000.
The number of multiples of 40 in this range is
200000
= 5000.
40
Therefore,
5000
P (registration number divisible by 40) =
200000
1
=
40
= 0.025.
Hence, the required probability is
0.025 .

Question 3: Selection of Compact and Intermediate-Size Cars


A car-rental agency has 19 compact cars and 12 intermediate-size cars. Four cars are selected
at random for a safety inspection. Find the probability that exactly two compact cars and
two intermediate-size cars are selected.
Answer.
The total number of cars is
19 + 12 = 31.
The total number of ways of selecting 4 cars from 31 cars is
(︃ )︃
31
.
4
The number of ways of selecting exactly two compact cars is
(︃ )︃
19
,
2
and the number of ways of selecting exactly two intermediate-size cars is
(︃ )︃
12
.
2

Page 16
Prof. V. K. Narla Module 2: Introduction to Probability

Therefore, the number of favourable selections is


(︃ )︃(︃ )︃
19 12
.
2 2
Hence, (︁19)︁(︁12)︁
2
P (two compact and two intermediate) = (︁31)︁2
4
171 × 66
=
31465
11286
=
31465
≈ 0.3587.
Thus, the required probability is approximately

0.359 .

Question 4: Relative-Frequency Estimate for Temperature


During the previous year, the maximum daily temperature at a plant exceeded 68◦ F on 12
days. Estimate the probability that the maximum temperature will exceed 68◦ F tomorrow.
Answer.
Using the relative-frequency interpretation of probability,
number of days on which the temperature exceeded 68◦ F
Pˆ︁ = .
total number of days observed
Assuming that the previous year had 365 days,
12
Pˆ︁ = ≈ 0.03288.
365
Therefore,
P (tomorrow’s maximum exceeds 68◦ F) ≈ 0.033.
Thus, the estimated probability is approximately

0.033 ,

or about 3.3%.

Question 5: Students Enrolled in Neither Course


Among 160 graduate engineering students, 92 are enrolled in an advanced statistics course,
63 are enrolled in an operations-research course, and 40 are enrolled in both courses. How
many students are enrolled in neither course?
Answer.

Page 17
Prof. V. K. Narla Module 2: Introduction to Probability

Let
S = {students enrolled in advanced statistics}
and
O = {students enrolled in operations research}.
Using the addition rule,

n(S ∪ O) = n(S) + n(O) − n(S ∩ O).

Substituting the given values,

n(S ∪ O) = 92 + 63 − 40
= 155 − 40
= 115.

Therefore, the number of students enrolled in neither course is

160 − n(S ∪ O) = 160 − 115


= 45.

Hence,
45 students
are enrolled in neither course.

Question 6: Checking Assignments of Probability


An experiment has four possible mutually exclusive and exhaustive outcomes A, B, C, and
D. Determine whether each assignment is permissible.
(a)
P (A) = 0.38, P (B) = 0.16, P (C) = 0.11, P (D) = 0.35.

(b)
P (A) = 0.27, P (B) = 0.30, P (C) = 0.28, P (D) = 0.16.

(c)
P (A) = 0.32, P (B) = 0.27, P (C) = −0.06, P (D) = 0.47.

(d)
1 1 1 1
P (A) = , P (B) = , P (C) = , P (D) = .
2 4 8 16
(e)
5 1 1 2
P (A) = , P (B) = , P (C) = , P (D) = .
18 6 3 9

Page 18
Prof. V. K. Narla Module 2: Introduction to Probability

Answer.
A permissible probability assignment must satisfy

P (A), P (B), P (C), P (D) ≥ 0

and
P (A) + P (B) + P (C) + P (D) = 1.
(a) All probabilities are nonnegative, and

0.38 + 0.16 + 0.11 + 0.35 = 1.

Therefore, the assignment is permissible.

(b) All probabilities are nonnegative, but

0.27 + 0.30 + 0.28 + 0.16 = 1.01.

Since the total is not 1, the assignment is not permissible.

(c) The probability


P (C) = −0.06
is negative. Therefore, the assignment violates the non-negativity axiom and is not
permissible.

(d) All probabilities are nonnegative, but


1 1 1 1 8 4 2 1
+ + + = + + +
2 4 8 16 16 16 16 16
15
=
16
̸= 1.

Therefore, the assignment is not permissible.

(e) All probabilities are nonnegative, and


5 1 1 2 5 3 6 4
+ + + = + + +
18 6 3 9 18 18 18 18
18
=
18
= 1.

Therefore, the assignment is permissible.

Page 19
Prof. V. K. Narla Module 2: Introduction to Probability

Question 7: Probabilities on a Coordinate Sample Space


A civil engineer tests three units of white cement and three units of black cement for adul-
teration. An outcome (x, y) means that x white-cement units and y black-cement units are
adulterated.
The outcomes and their probabilities are:

(x, y) (0, 0) (0, 1) (0, 2) (0, 3)


P (x, y) 0.080 0.032 0.086 0.064

(x, y) (1, 0) (1, 1) (1, 2) (1, 3)


P (x, y) 0.085 0.073 0.065 0.091

(x, y) (2, 0) (2, 1) (2, 2) (2, 3)


P (x, y) 0.071 0.050 0.046 0.075

(x, y) (3, 0) (3, 1) (3, 2) (3, 3)


P (x, y) 0.040 0.021 0.080 0.041
Let
A = {equal numbers of white and black units are adulterated},
B = {no white-cement unit is adulterated},
and
C = {fewer white-cement units than black-cement units are adulterated}.
(a) Verify that the probability assignment is permissible.
(b) Find P (A), P (B), and P (C).
(c) Find the probabilities that exactly one, exactly two, or exactly three white-cement
units are adulterated.
Answer.
(a) Verification of the probability assignment
Every listed probability is nonnegative. The row totals are

0.080 + 0.032 + 0.086 + 0.064 = 0.262,

0.085 + 0.073 + 0.065 + 0.091 = 0.314,


0.071 + 0.050 + 0.046 + 0.075 = 0.242,
and
0.040 + 0.021 + 0.080 + 0.041 = 0.182.

Therefore,
0.262 + 0.314 + 0.242 + 0.182 = 1.

Hence, the probability assignment is permissible.

Page 20
Prof. V. K. Narla Module 2: Introduction to Probability

(b) Probabilities of A, B, and C


For event A,
A = {(0, 0), (1, 1), (2, 2), (3, 3)}.
Therefore,
P (A) = 0.080 + 0.073 + 0.046 + 0.041
= 0.240.

For event B,
B = {(0, 0), (0, 1), (0, 2), (0, 3)}.
Hence,
P (B) = 0.080 + 0.032 + 0.086 + 0.064
= 0.262.

For event C, we require x < y. Thus,

C = {(0, 1), (0, 2), (0, 3), (1, 2), (1, 3), (2, 3)}.

Therefore,
P (C) = 0.032 + 0.086 + 0.064 + 0.065 + 0.091 + 0.075
= 0.413.

Hence,
P (A) = 0.240, P (B) = 0.262, P (C) = 0.413.

(c) Number of adulterated white-cement units


Exactly one white-cement unit is adulterated when x = 1:

P (x = 1) = 0.085 + 0.073 + 0.065 + 0.091


= 0.314.

Exactly two white-cement units are adulterated when x = 2:

P (x = 2) = 0.071 + 0.050 + 0.046 + 0.075


= 0.242.

Exactly three white-cement units are adulterated when x = 3:

P (x = 3) = 0.040 + 0.021 + 0.080 + 0.041


= 0.182.

Thus,

P (x = 1) = 0.314, P (x = 2) = 0.242, P (x = 3) = 0.182.

Page 21
Prof. V. K. Narla Module 2: Introduction to Probability

Question 8: Mutually Exclusive Events and Complements


Suppose A and B are mutually exclusive events with

P (A) = 0.45, P (B) = 0.30.

Find:
(a) P (A);
(b) P (A ∪ B);
(c) P (A ∩ B);
(d) P (A ∩ B).
Answer.
Since A and B are mutually exclusive,

A ∩ B = ∅ and P (A ∩ B) = 0.

(a) By the complement rule,


P (A) = 1 − P (A)
= 1 − 0.45
= 0.55.

(b) Since A and B are mutually exclusive,

P (A ∪ B) = P (A) + P (B)
= 0.45 + 0.30
= 0.75.

(c) Since A and B cannot occur together, every outcome in A belongs to B. Therefore,

A ∩ B = A,

and
P (A ∩ B) = P (A) = 0.45.

(d) By De Morgan’s law,


A ∩ B = A ∪ B.
Hence,
P (A ∩ B) = 1 − P (A ∪ B)
= 1 − 0.75
= 0.25.

Therefore,

P (A) = 0.55, P (A ∪ B) = 0.75, P (A ∩ B) = 0.45, P (A ∩ B) = 0.25.

Page 22
Prof. V. K. Narla Module 2: Introduction to Probability

Question 9: Number of Complaints Received by a Television Sta-


tion
The probabilities that a television station receives 0, 1, 2, . . . , 8, or at least 9 complaints after
showing a controversial programme are, respectively,

0.01, 0.03, 0.07, 0.15, 0.19, 0.18, 0.14, 0.12, 0.09, 0.02.

Find the probability that the station receives:


(a) at most 4 complaints;
(b) at least 6 complaints;
(c) from 5 to 8 complaints, inclusive.
Answer.
Let X denote the number of complaints.
(a) At most 4 complaints means

X = 0, 1, 2, 3, or 4.

Therefore,
P (X ≤ 4) = 0.01 + 0.03 + 0.07 + 0.15 + 0.19
= 0.45.

(b) At least 6 complaints means

X = 6, 7, 8, or at least 9.

Thus,
P (X ≥ 6) = 0.14 + 0.12 + 0.09 + 0.02
= 0.37.

(c) From 5 to 8 complaints, inclusive, means

X = 5, 6, 7, or 8.

Therefore,
P (5 ≤ X ≤ 8) = 0.18 + 0.14 + 0.12 + 0.09
= 0.53.

Hence,

P (X ≤ 4) = 0.45, P (X ≥ 6) = 0.37, P (5 ≤ X ≤ 8) = 0.53.

Page 23
Prof. V. K. Narla Module 2: Introduction to Probability

Question 10: General Addition and Complement Rules


Given
P (A) = 0.30, P (B) = 0.62, P (A ∩ B) = 0.12,
find:
(a) P (A ∪ B);
(b) P (A ∩ B);
(c) P (A ∩ B);
(d) P (A ∪ B).
Answer.
(a) Using the general addition rule,

P (A ∪ B) = P (A) + P (B) − P (A ∩ B)
= 0.30 + 0.62 − 0.12
= 0.80.

(b) Event B is the disjoint union of A ∩ B and A ∩ B. Therefore,

P (B) = P (A ∩ B) + P (A ∩ B).

Hence,
P (A ∩ B) = P (B) − P (A ∩ B)
= 0.62 − 0.12
= 0.50.

(c) Event A is the disjoint union of A ∩ B and A ∩ B. Thus,

P (A ∩ B) = P (A) − P (A ∩ B)
= 0.30 − 0.12
= 0.18.

(d) By De Morgan’s law,


A ∪ B = A ∩ B.
Therefore,
P (A ∪ B) = 1 − P (A ∩ B)
= 1 − 0.12
= 0.88.
Hence,

P (A ∪ B) = 0.80, P (A ∩ B) = 0.50, P (A ∩ B) = 0.18, P (A ∪ B) = 0.88.

Page 24
Prof. V. K. Narla Module 2: Introduction to Probability

5. Conditional Probability, Multiplication Rules, and


Independence
Conditional probability is used when the probability of an event is required under the con-
dition that another event has already occurred. The additional information changes the
effective sample space and may therefore change the probability of the event under consid-
eration.

5.1 Conditional Probability


Let A and B be events in a sample space S, with

P (B) > 0.

The conditional probability of A given B is defined by

P (A ∩ B)
P (A | B) = .
P (B)

Here:
ˆ P (A | B) means the probability that A occurs, given that B has occurred;
ˆ A ∩ B represents the event that both A and B occur;
ˆ P (B) is the probability of the conditioning event.
Once it is known that B has occurred, the sample space is effectively restricted to B.
Among the outcomes in B, only those belonging to A ∩ B are favourable to event A. This
explains the ratio
P (A ∩ B)
.
P (B)

Example 1: Calculating a Conditional Probability


The probability that a communication system has high fidelity is 0.81. The probability
that it has both high fidelity and high selectivity is 0.18. Find the probability that a system
with high fidelity also has high selectivity.
Solution.
Let
A = {the communication system has high selectivity},
and
B = {the communication system has high fidelity}.
The given probabilities are
P (B) = 0.81
and
P (A ∩ B) = 0.18.

Page 25
Prof. V. K. Narla Module 2: Introduction to Probability

The required probability is the conditional probability of A given B:


P (A ∩ B)
P (A | B) = .
P (B)
Substituting the given values,
0.18
P (A | B) =
0.81
18
=
81
2
= .
9
Therefore,
2
P (high selectivity | high fidelity) =
≈ 0.2222.
9
Thus, among systems having high fidelity, approximately 22.22% also have high selectiv-
ity.
Example 2: Conditional Probability for a Used Car
In a used-car classification, let
M1 = {the car has low mileage}
and
C3 = {the car is expensive to operate}.
Suppose
P (M1 ∩ C3 ) = 0.08
and
P (C3 ) = 0.40.
Find the conditional probability that a used car has low mileage, given that it is expensive
to operate.
Solution.
The required probability is
P (M1 | C3 ).
Using the definition of conditional probability,
P (M1 ∩ C3 )
P (M1 | C3 ) = .
P (C3 )
Substituting the given probabilities,
0.08
P (M1 | C3 ) =
0.40
= 0.20.
Hence,
P (M1 | C3 ) = 0.20.
Therefore, among used cars that are expensive to operate, 20% have low mileage.

Page 26
Prof. V. K. Narla Module 2: Introduction to Probability

5.2 Independent Events


Two events are independent when the occurrence of one event does not change the probability
of the other.
If A and B are events and P (B) > 0, then A is independent of B when

P (A | B) = P (A).

Similarly, if P (A) > 0, then B is independent of A when

P (B | A) = P (B).

Thus, for independent events, knowing that one event has occurred provides no additional
information about the occurrence of the other event.
In Example 2, the conditional probability was

P (M1 | C3 ) = 0.20.

If the unconditional probability is also

P (M1 ) = 0.20,

then
P (M1 | C3 ) = P (M1 ),
so the events M1 and C3 are independent.

5.3 General Multiplication Rule of Probability

Theorem 1. General Multiplication Rule


If A and B are events in a sample space S, then

P (A ∩ B) = P (A) P (B | A), P (A) > 0,

or equivalently,
P (A ∩ B) = P (B) P (A | B), P (B) > 0.

Derivation
From the definition of conditional probability,

P (A ∩ B)
P (A | B) = .
P (B)

Multiplying both sides by P (B), we obtain

P (A ∩ B) = P (B)P (A | B).

Page 27
Prof. V. K. Narla Module 2: Introduction to Probability

Similarly,
P (A ∩ B)
P (B | A) = ,
P (A)
and multiplication by P (A) gives

P (A ∩ B) = P (A)P (B | A).

The multiplication rule is particularly useful for sequential experiments, especially when
the result of the first selection changes the probability of the second selection.
Example 3: Using the General Multiplication Rule
A supervisor has a group of 20 construction workers. Of these, 12 workers favour new
safety regulations and 8 are against them. The supervisor randomly selects two workers
without replacement.
Find the probability that both selected workers are against the new safety regulations.
Solution.
Let
A = {the first selected worker is against the regulations},
and
B = {the second selected worker is against the regulations}.
For the first selection, 8 of the 20 workers are against the regulations. Therefore,
8
P (A) = .
20
Since the first selected worker is not replaced, if the first worker is against the regulations,
then 7 such workers remain among the remaining 19 workers. Hence,
7
P (B | A) = .
19
Using the general multiplication rule,

P (A ∩ B) = P (A)P (B | A).

Therefore,
8 7
P (A ∩ B) = ·
20 19
56
=
380
14
= .
95
Hence,
14
P (both workers are against the regulations) = ≈ 0.1474.
95
The two selections are dependent because the first worker is not replaced.

Page 28
Prof. V. K. Narla Module 2: Introduction to Probability

5.4 Special Product Rule for Independent Events

Theorem 2. Special Product Rule


Two events A and B are independent if and only if

P (A ∩ B) = P (A)P (B).

Explanation
For independent events,
P (A | B) = P (A).
Substituting this relation into the general multiplication rule,

P (A ∩ B) = P (B)P (A | B),

gives
P (A ∩ B) = P (B)P (A).
Thus, when two events are independent, the probability that both events occur is the
product of their individual probabilities.
Conversely, if
P (A ∩ B) = P (A)P (B),
then
P (A ∩ B) P (A)P (B)
P (A | B) = = = P (A),
P (B) P (B)
provided P (B) > 0. Hence, A and B are independent.
Example 4: Two Heads in Two Tosses of a Balanced Coin
Find the probability of obtaining two heads in two tosses of a balanced coin.
Solution.
Let
H1 = {a head occurs on the first toss},
and
H2 = {a head occurs on the second toss}.
For a balanced coin,
1
P (H1 ) =
2
and
1
P (H2 ) = .
2
The two tosses are independent because the result of the first toss does not affect the
result of the second toss.

Page 29
Prof. V. K. Narla Module 2: Introduction to Probability

Therefore,
P (H1 ∩ H2 ) = P (H1 )P (H2 )
1 1
= ·
2 2
1
= .
4
Hence,
1
P (two heads) = .
4

Example 5: Selection With and Without Replacement


Two cards are drawn at random from an ordinary deck of 52 playing cards. Find the
probability of drawing two aces when:
(a) the first card is replaced before the second card is drawn;
(b) the first card is not replaced before the second card is drawn.
Solution.
There are 4 aces in a deck of 52 cards.
(a) With replacement
The probability that the first card is an ace is
4
.
52

Since the card is replaced, the deck again contains 52 cards with 4 aces. Therefore,
the probability that the second card is also an ace is
4
.
52

The two selections are independent. Hence,


4 4
P (two aces with replacement) = ·
52 52
1 1
= ·
13 13
1
= .
169

Thus,
1
P (two aces with replacement) = .
169

Page 30
Prof. V. K. Narla Module 2: Introduction to Probability

(b) Without replacement


The probability that the first card is an ace is
4
.
52

If an ace is drawn first and not replaced, then only 3 aces remain among 51 cards.
Therefore,
3
P (second ace | first ace) = .
51
Using the general multiplication rule,
4 3
P (two aces without replacement) = ·
52 51
12
=
2652
1
= .
221

Hence,
1
P (two aces without replacement) = .
221
Notice that
1 4 4
̸= · .
221 52 52
Therefore, the two selections are not independent when sampling is performed without
replacement.
Example 6: Checking Whether Two Events Are Independent
Suppose
P (C) = 0.65, P (D) = 0.40,
and
P (C ∩ D) = 0.24.
Determine whether the events C and D are independent.
Solution.
If C and D are independent, then they must satisfy

P (C ∩ D) = P (C)P (D).

Compute the product:


P (C)P (D) = (0.65)(0.40)
= 0.26.
However, the given intersection probability is

P (C ∩ D) = 0.24.

Page 31
Prof. V. K. Narla Module 2: Introduction to Probability

Since
0.24 ̸= 0.26,
we conclude that
C and D are not independent.

5.5 Assigning Probabilities by the Special Product Rule


When two events describe unrelated stages of an experiment or manufacturing process, they
may be modelled as independent. Their joint probability can then be obtained by multiplying
their individual probabilities.
Example 7: Assigning Probability by the Special Product Rule
Let
A = {raw material is available when required}
and
B = {machining time is less than one hour}.
Suppose
P (A) = 0.8
and
P (B) = 0.7.
Assuming that the two events concern unrelated steps of the manufacturing process, find
P (A ∩ B).
Solution.
Because the events are treated as independent,

P (A ∩ B) = P (A)P (B).

Substituting the given probabilities,

P (A ∩ B) = (0.8)(0.7)
= 0.56.

Therefore,
P (A ∩ B) = 0.56.
Thus, the probability that the raw material is available and the machining time is less
than one hour is 0.56.

Page 32
Prof. V. K. Narla Module 2: Introduction to Probability

5.6 Extended Product Rule for Several Independent Events


If A1 , A2 , . . . , An are mutually independent events, then

P (A1 ∩ A2 ∩ · · · ∩ An ) = P (A1 )P (A2 ) · · · P (An ).

For repeated independent trials, the probability that the same type of event occurs on
every trial is the product of the corresponding trial probabilities.
Example 8: Extended Special Product Rule
A balanced die is rolled four times. Find the probability of not obtaining a 6 on any of
the four rolls.
Solution.
For one roll of a balanced die,
5
P (not obtaining a 6) = .
6
The four rolls are independent. Therefore,
(︃ )︃4
5
P (no 6 in four rolls) =
6
4
5
= 4
6
625
= .
1296
Hence,
625
P (no 6 in four rolls) = ≈ 0.4823.
1296

Summary of the Main Results

P (A ∩ B)
P (A | B) = , P (B) > 0,
P (B)
P (A ∩ B) = P (A)P (B | A), P (A) > 0,
P (A ∩ B) = P (B)P (A | B), P (B) > 0,
P (A ∩ B) = P (A)P (B), if A and B are independent,
n
∏︂
P (A1 ∩ · · · ∩ An ) = P (Ai ), for mutually independent events.
i=1

The essential distinction is:


ˆ use the general multiplication rule when one event changes the probability of
another;

Page 33
Prof. V. K. Narla Module 2: Introduction to Probability

ˆ use the special product rule when the events are independent;
ˆ use conditional probability when probability is required under a stated condition.

Page 34
Prof. V. K. Narla Module 2: Introduction to Probability

6. Bayes’ Theorem
Bayes’ theorem is used to revise the probability of a possible cause after observing an out-
come. It combines:
ˆ the probability of each possible cause before the new evidence is observed;
ˆ the probability of observing the evidence under each possible cause;
ˆ the total probability of observing the evidence.
It is therefore an important rule for inverse probability: we begin with an observed
effect and calculate the probability of the cause that produced it.

6.1 Partition of the Sample Space


Let
B1 , B2 , . . . , Bn
be events that form a partition of the sample space S. This means that:
1. the events are mutually exclusive:
Bi ∩ Bj = ∅, i ̸= j;

2. one of them must occur:


B1 ∪ B2 ∪ · · · ∪ Bn = S;
3. each event has positive probability:
P (Bi ) > 0.

Suppose A is an observed event with


P (A) > 0.
Because B1 , B2 , . . . , Bn form a partition, the event A may be written as
A = (A ∩ B1 ) ∪ (A ∩ B2 ) ∪ · · · ∪ (A ∩ Bn ).
The events on the right-hand side are mutually exclusive. Therefore,
n
∑︂
P (A) = P (A ∩ Bi ).
i=1

Using the multiplication rule,


P (A ∩ Bi ) = P (Bi )P (A | Bi ).
Hence,
n
∑︂
P (A) = P (Bi )P (A | Bi ).
i=1

This expression is called the law of total probability.

Page 35
Prof. V. K. Narla Module 2: Introduction to Probability

6.2 Statement and Derivation of Bayes’ Theorem


Theorem 1. Bayes’ Theorem
If
B1 , B2 , . . . , Bn
are mutually exclusive and exhaustive events, one of which must occur, and A is an event
with P (A) > 0, then for
r = 1, 2, . . . , n,
P (Br )P (A | Br )
P (Br | A) = n .
∑︂
P (Bi )P (A | Bi )
i=1

Derivation
By the definition of conditional probability,
P (Br ∩ A)
P (Br | A) = .
P (A)
Using the multiplication rule in the numerator,
P (Br ∩ A) = P (Br )P (A | Br ).
Using the law of total probability in the denominator,
n
∑︂
P (A) = P (Bi )P (A | Bi ).
i=1

Substituting these two expressions gives


P (Br )P (A | Br )
P (Br | A) = n .
∑︂
P (Bi )P (A | Bi )
i=1

6.3 Interpretation of the Formula


The numerator
P (Br )P (A | Br )
is the probability of reaching the observed event A through the particular cause Br .
The denominator n
∑︂
P (Bi )P (A | Bi )
i=1
is the total probability of reaching A through all possible causes.
Therefore, Bayes’ theorem may be interpreted as
probability of the selected route to the evidence
posterior probability = .
probability of all routes to the evidence

Page 36
Prof. V. K. Narla Module 2: Introduction to Probability

Prior and Posterior Probabilities


The probabilities
P (B1 ), P (B2 ), . . . , P (Bn )
are called prior probabilities. They represent the probabilities assigned before the event
A is observed.
The probabilities
P (B1 | A), P (B2 | A), . . . , P (Bn | A)
are called posterior probabilities. They represent the revised probabilities after the evi-
dence A has been observed.
The quantities
P (A | Bi )
are the probabilities of observing the evidence under the different possible causes. They are
often called likelihoods.

6.4 Procedure for Applying Bayes’ Theorem


A Bayes’ theorem problem can be solved systematically as follows:
1. Define the possible causes
B1 , B2 , . . . , Bn .

2. Verify that the causes are mutually exclusive and exhaustive.


3. Define the observed event A.
4. Write the prior probabilities
P (Bi ).

5. Write the conditional probabilities

P (A | Bi ).

6. Calculate the joint-route probabilities

P (Bi )P (A | Bi ).

7. Add the joint-route probabilities to obtain


n
∑︂
P (A) = P (Bi )P (A | Bi ).
i=1

8. Divide the route probability associated with the required cause by the total probability
of the evidence.

Page 37
Prof. V. K. Narla Module 2: Introduction to Probability

6.5 Examples

Example 1: Using Bayes’ Theorem for Repair Responsibility


Four technicians regularly make repairs when breakdowns occur on an automated pro-
duction line.

Technician Fraction of breakdowns serviced Probability of incomplete repair

Janet 0.20 1/20 = 0.05


Tom 0.60 1/10 = 0.10
Georgia 0.15 1/10 = 0.10
Peter 0.05 1/20 = 0.05

For the next breakdown, the production-line diagnosis indicates that the initial repair
was incomplete. Find the probability that Janet made the initial repair.
Solution.
Let
A = {the initial repair was incomplete}.
Define the technician events:

B1 = {Janet made the repair},

B2 = {Tom made the repair},


B3 = {Georgia made the repair},
and
B4 = {Peter made the repair}.
The prior probabilities are

P (B1 ) = 0.20, P (B2 ) = 0.60, P (B3 ) = 0.15, P (B4 ) = 0.05.

The conditional probabilities of an incomplete repair are

P (A | B1 ) = 0.05,

P (A | B2 ) = 0.10,
P (A | B3 ) = 0.10,
and
P (A | B4 ) = 0.05.

Page 38
Prof. V. K. Narla Module 2: Introduction to Probability

Step 1: Compute the probability of each route to an incomplete repair


For Janet,
P (B1 )P (A | B1 ) = (0.20)(0.05)
= 0.0100.
For Tom,
P (B2 )P (A | B2 ) = (0.60)(0.10)
= 0.0600.
For Georgia,
P (B3 )P (A | B3 ) = (0.15)(0.10)
= 0.0150.
For Peter,
P (B4 )P (A | B4 ) = (0.05)(0.05)
= 0.0025.

Step 2: Find the total probability of an incomplete repair


Using the law of total probability,
4
∑︂
P (A) = P (Bi )P (A | Bi )
i=1
= 0.0100 + 0.0600 + 0.0150 + 0.0025
= 0.0875.

Step 3: Apply Bayes’ theorem


The required probability is
P (B1 | A).
Therefore,

P (B1 )P (A | B1 )
P (B1 | A) =
P (A)
(0.20)(0.05)
=
(0.20)(0.05) + (0.60)(0.10) + (0.15)(0.10) + (0.05)(0.05)
0.0100
=
0.0875
= 0.1142857.

Hence,
P (Janet made the repair | repair was incomplete) ≈ 0.114.
Thus, approximately 11.4% of all incomplete repairs are attributable to Janet.

Page 39
Prof. V. K. Narla Module 2: Introduction to Probability

Interpretation
Janet services 20% of the breakdowns, but her incomplete-repair rate is only 5%. Tom
services a much larger fraction of the breakdowns and has a higher incomplete-repair rate.
Consequently, after observing an incomplete repair, the probability that Janet was respon-
sible decreases from the prior value 0.20 to the posterior value 0.114.
Example 2: Identifying Spam Using Bayes’ Theorem
A spam-filtering system uses a list of words that occur more frequently in spam messages
than in normal messages.
A database contains 5000 messages, of which:

1700 are spam


and
3300 are normal.
Among the spam messages, 1343 contain words from the specified list. Among the normal
messages, 297 contain words from the list.
Find the probability that a message is spam, given that it contains words from the list.
Solution.
Let
A = {the message contains words from the list},
B1 = {the message is spam},
and
B2 = {the message is normal}.

Step 1: Calculate the prior probabilities


The prior probability that a message is spam is
1700
P (B1 ) =
5000
= 0.34.

The prior probability that a message is normal is


3300
P (B2 ) =
5000
= 0.66.

Step 2: Calculate the conditional probabilities


The probability that a spam message contains words from the list is
1343
P (A | B1 ) =
1700
= 0.79.

Page 40
Prof. V. K. Narla Module 2: Introduction to Probability

The probability that a normal message contains words from the list is
297
P (A | B2 ) =
3300
= 0.09.

Step 3: Find the total probability that a message contains words from the list
Using the law of total probability,
P (A) = P (B1 )P (A | B1 ) + P (B2 )P (A | B2 )
= (0.34)(0.79) + (0.66)(0.09)
= 0.2686 + 0.0594
= 0.3280.

Step 4: Apply Bayes’ theorem


The required probability is
P (B1 | A).
Therefore,
P (B1 )P (A | B1 )
P (B1 | A) =
P (B1 )P (A | B1 ) + P (B2 )P (A | B2 )
(0.34)(0.79)
=
(0.34)(0.79) + (0.66)(0.09)
0.2686
=
0.3280
= 0.8189024.
Hence,
P (spam | message contains listed words) ≈ 0.819.
Therefore, a message containing words from the specified list has an approximately 81.9%
probability of being spam.

Interpretation
Before examining the message, the probability that it is spam is

P (B1 ) = 0.34.

After observing that the message contains words from the list, the probability increases
to
P (B1 | A) = 0.819.
Thus, the observed words provide strong evidence in favour of the message being spam.
However, the probability is not 1, because some normal messages also contain words from
the list.

Page 41
Prof. V. K. Narla Module 2: Introduction to Probability

6.6 Important Observations


ˆ Bayes’ theorem reverses the direction of a conditional probability. It obtains P (Bi | A)
from information about P (A | Bi ).
ˆ The prior probability measures belief before observing the evidence.
ˆ The posterior probability measures the revised belief after observing the evidence.
ˆ A large value of P (A | Bi ) does not by itself guarantee a large posterior probability
P (Bi | A). The prior probability P (Bi ) must also be considered.
ˆ The denominator in Bayes’ theorem is the total probability of the observed evidence.
ˆ Bayes’ theorem is widely used in reliability analysis, quality control, medical diagnosis,
classification, machine learning, and spam filtering.

6.7 Formula Summary


n
∑︂
P (A) = P (Bi )P (A | Bi )
i=1

P (Br )P (A | Br )
P (Br | A) = n
∑︂
P (Bi )P (A | Bi )
i=1

Prior × Likelihood
Posterior =
Total probability of the evidence

Practice Questions and Answers


Question 1: Testing Whether Two Events Are Independent
Suppose
P (X) = 0.33, P (Y ) = 0.75, P (X ∩ Y ) = 0.30.
Determine whether the events X and Y are independent.
Answer.
Two events X and Y are independent if and only if
P (X ∩ Y ) = P (X)P (Y ).

Step 1: Calculate P (X)P (Y )


Using the given probabilities,
P (X)P (Y ) = (0.33)(0.75)
= 0.2475.

Page 42
Prof. V. K. Narla Module 2: Introduction to Probability

Step 2: Compare with the given intersection probability


The given intersection probability is

P (X ∩ Y ) = 0.30.

However,
0.30 ̸= 0.2475.
Therefore, the condition
P (X ∩ Y ) = P (X)P (Y )
is not satisfied.
Hence,
X and Y are not independent.

Verification Using Conditional Probability


We may also check independence by calculating

P (X | Y ).

By the definition of conditional probability,

P (X ∩ Y )
P (X | Y ) = .
P (Y )

Substituting the given values,


0.30
P (X | Y ) =
0.75
= 0.40.

Since
P (X | Y ) = 0.40
but
P (X) = 0.33,
we have
P (X | Y ) ̸= P (X).
This again confirms that X and Y are dependent events.

Page 43
Prof. V. K. Narla Module 2: Introduction to Probability

Question 2: Detection of Internal Corrosion in a Pipe


Engineers in charge of maintaining our nuclear fleet must continually check for corrosion
inside the pipes that are part of the cooling systems. The inside condition of the pipes
cannot be observed directly but a nondestructive test can give an indication of possible
corrosion. This test is not infallible. The test has probability 0.7 of detecting corrosion
when it is present but it also has probability 0.2 of falsely indicating internal corrosion.
Suppose the probability that any section of pipe has internal corrosion is 0.1.
(a) Determine the probability that a section of pipe has internal corrosion, given that the
test indicates its presence.
(b) Determine the probability that a section of pipe has internal corrosion, given that the
test is negative.
Answer:
Definition of events
Let
C = {the pipe section has internal corrosion},
and
C = {the pipe section does not have internal corrosion}.
Let
T = {the test indicates that corrosion is present},
and
T = {the test is negative}.
From the question,
P (C) = 0.1.
Therefore,
P (C) = 1 − P (C) = 1 − 0.1 = 0.9.
The test detects corrosion when it is actually present with probability 0.7. Hence,

P (T | C) = 0.7.

The test falsely indicates corrosion when corrosion is absent with probability 0.2. Hence,

P (T | C) = 0.2.

The corresponding negative-test probabilities are

P (T | C) = 1 − P (T | C) = 1 − 0.7 = 0.3,

and
P (T | C) = 1 − P (T | C) = 1 − 0.2 = 0.8.

Page 44
Prof. V. K. Narla Module 2: Introduction to Probability

Probability That the Test Indicates Corrosion


The test can indicate corrosion through either of two mutually exclusive routes:
1. corrosion is present and the test correctly detects it;
2. corrosion is absent and the test gives a false positive.
Therefore, by the law of total probability,

P (T ) = P (C)P (T | C) + P (C)P (T | C).

Substituting the given values,

P (T ) = (0.10)(0.70) + (0.90)(0.20)
= 0.07 + 0.18
= 0.25.

Hence,
P (T ) = 0.25.
Thus, the test indicates corrosion for 25% of the pipe sections.

(a) Probability of corrosion given that the test is positive


The required probability is
P (C | T ).
By Bayes’ theorem,
P (C)P (T | C)
P (C | T ) = .
P (C)P (T | C) + P (C)P (T | C)
Substituting the given values,
(0.1)(0.7)
P (C | T ) =
(0.1)(0.7) + (0.9)(0.2)
0.07
=
0.07 + 0.18
0.07
=
0.25
= 0.28.

Therefore,
P (C | T ) = 0.28.
Thus, when the test indicates that corrosion is present, the probability that the pipe
section actually has internal corrosion is

28%.

Page 45
Prof. V. K. Narla Module 2: Introduction to Probability

(b) Probability of corrosion given that the test is negative


The required probability is
P (C | T ).
By Bayes’ theorem,
P (C)P (T | C)
P (C | T ) = .
P (C)P (T | C) + P (C)P (T | C)
Substituting the probabilities,
(0.1)(0.3)
P (C | T ) =
(0.1)(0.3) + (0.9)(0.8)
0.03
=
0.03 + 0.72
0.03
=
0.75
= 0.04.
Therefore,
P (C | T ) = 0.04.
Thus, even when the test is negative, there remains a
4%
probability that the pipe section actually has internal corrosion.
Final answers
P (C | T ) = 0.28, P (C | T ) = 0.04.

Interpretation
Although the test correctly detects corrosion 70% of the time when corrosion is present,
the actual prevalence of corrosion is only 10%. Furthermore, the false-positive probability
is 20%. Since non-corroded pipe sections are much more common than corroded sections,
false-positive results form a substantial proportion of all positive test results.
To see this clearly, consider 1000 pipe sections:

Condition Number of sections Positive-test rate Expected positive tests

Corrosion present 100 0.70 70


Corrosion absent 900 0.20 180

Total 1000 — 250


Of the 250 expected positive results, only 70 correspond to actual corrosion. Therefore,
70
= 0.28,
250
which agrees with the Bayes’ theorem calculation.

Page 46
Prof. V. K. Narla Module 2: Introduction to Probability

7. Random Variables and Probability Distributions


A random experiment may have outcomes that are descriptive rather than numerical. To
analyse such outcomes mathematically, we assign a numerical value to each possible outcome.

7.1 Random Variables


A random variable is a function that assigns a numerical value to each possible outcome
of a random experiment.
If S is the sample space, then a random variable X may be written as
X : S −→ R.
Thus, for every outcome ω ∈ S, the random variable assigns a real number
X(ω).
The numerical value assigned by the random variable should represent an important
characteristic of the outcome.

Illustration
Suppose a balanced die is rolled. The sample space is
S = {1, 2, 3, 4, 5, 6}.
If X denotes the number appearing on the upper face, then
X(1) = 1, X(2) = 2, ..., X(6) = 6.
Here, the random variable converts each outcome into its corresponding numerical value.

7.2 Classification of Random Variables


Random variables are commonly classified according to the number of values that they can
assume.

Discrete Random Variable


A random variable is called a discrete random variable if it can assume:
ˆ only a finite number of values; or
ˆ a countably infinite number of values.
Examples include:
ˆ the number shown on a die;
ˆ the number of defective items in a sample;
ˆ the number of customers arriving during a fixed time interval.

Page 47
Prof. V. K. Narla Module 2: Introduction to Probability

Continuous Random Variable


A random variable is called a continuous random variable if it can assume any value in
an interval or in a collection of intervals.
Examples include:
ˆ temperature;
ˆ time;
ˆ length;
ˆ mass;
ˆ pressure.
In this section, the discussion is restricted to discrete random variables.

7.3 Probability Distribution of a Discrete Random Variable


Let X be a discrete random variable. The probability that X assumes the value x is denoted
by
f (x) = P (X = x).
The function f (x) is called the probability distribution or probability mass func-
tion of X.
The probability distribution of a discrete random variable is a list of all possible values
of X, together with their corresponding probabilities.
Thus,
f (x) = P (X = x).

Conditions for a Probability Distribution


A function f (x) can serve as the probability distribution of a discrete random variable only
if it satisfies the following two conditions:

f (x) ≥ 0 for every possible value of x


and
∑︂
f (x) = 1.
all x

The first condition states that a probability cannot be negative. The second condition
states that one of the possible values of the random variable must occur.

Page 48
Prof. V. K. Narla Module 2: Introduction to Probability

Example 1: Distribution of the Outcome of a Balanced Die


Let X denote the number obtained when a balanced die is rolled. Then the possible values
of X are
1, 2, 3, 4, 5, 6.
Since all six outcomes are equally likely,
1
f (x) = , x = 1, 2, 3, 4, 5, 6.
6
Equivalently, ⎧
⎨ 1 , x = 1, 2, 3, 4, 5, 6,
f (x) = 6
⎩0, otherwise.

The non-negativity condition is satisfied because


1
f (x) = > 0.
6
Also,
6 (︃ )︃
∑︂ 1
f (x) = 6 = 1.
x=1
6
Hence, f (x) is a valid probability distribution.

Example 2: Checking the Conditions for a Probability Distribution


Determine whether each of the following functions can serve as a probability distribution.
(a)
x−2
f (x) = , x = 1, 2, 3, 4.
2
(b)
x2
h(x) = , x = 0, 1, 2, 3, 4.
25
Solution.
(a) For
x−2
f (x) = ,
2
evaluate the function at x = 1:
1−2 1
f (1) = =− .
2 2
Since
f (1) < 0,
the non-negativity condition is violated.

Page 49
Prof. V. K. Narla Module 2: Introduction to Probability

Therefore,

x−2
f (x) = cannot serve as a probability distribution.
2

(b) For
x2
h(x) = , x = 0, 1, 2, 3, 4,
25
all the function values are nonnegative. However, their sum must also equal 1.
We calculate
4 4
∑︂ ∑︂ x2
h(x) =
x=0 x=0
25
02 + 12 + 22 + 32 + 42
=
25
0 + 1 + 4 + 9 + 16
=
25
30
=
25
6
= .
5
Since
6
̸= 1,
5
the total-probability condition is violated.
Therefore,
x2
h(x) = cannot serve as a probability distribution.
25

7.4 Cumulative Distribution Function


In addition to the probability that a random variable assumes a specific value, it is often
useful to determine the probability that the random variable is less than or equal to a
specified value.
The cumulative distribution function of a random variable X is defined by

F (x) = P (X ≤ x), −∞ < x < ∞.

For a discrete random variable with probability mass function f (x),


∑︂
F (x) = f (t).
t≤x

Thus, F (x) accumulates all probabilities assigned to values of the random variable that
are less than or equal to x.

Page 50
Prof. V. K. Narla Module 2: Introduction to Probability

Basic Properties of the Cumulative Distribution Function


The cumulative distribution function satisfies:
1.
0 ≤ F (x) ≤ 1;

2. F (x) is a nondecreasing function;


3.
lim F (x) = 0;
x→−∞

4.
lim F (x) = 1.
x→∞

For a discrete random variable, the probability at a particular value may be recovered
from the jumps of the cumulative distribution function.

8. Practice Questions and Answers


Question 1: Checking Assigned Probabilities
Determine whether the following can be probability distributions of a random variable that
can take only the values 1, 2, 3, and 4.
(a)
f (1) = 0.19, f (2) = 0.27, f (3) = 0.27, f (4) = 0.27.

(b)
f (1) = 0.24, f (2) = 0.24, f (3) = 0.24, f (4) = 0.24.

(c)
f (1) = 0.35, f (2) = 0.33, f (3) = 0.34, f (4) = −0.02.

Answer:
For a function f (x) to be a probability distribution of a discrete random variable, it must
satisfy both conditions:

f (x) ≥ 0 for every possible value of x,

and ∑︂
f (x) = 1.
x

Page 51
Prof. V. K. Narla Module 2: Introduction to Probability

(a) All the assigned probabilities are nonnegative:

0.19, 0.27, 0.27, 0.27 ≥ 0.

Their sum is
4
∑︂
f (x) = 0.19 + 0.27 + 0.27 + 0.27
x=1
= 0.46 + 0.27 + 0.27
= 0.73 + 0.27
= 1.

Both probability-distribution conditions are satisfied. Therefore,

The assignment in part (a) is a valid probability distribution.

(b) All four assigned values are nonnegative. However,


4
∑︂
f (x) = 0.24 + 0.24 + 0.24 + 0.24
x=1
= 4(0.24)
= 0.96.

Since
0.96 ̸= 1,
the total-probability condition is not satisfied. Therefore,

The assignment in part (b) is not a valid probability distribution.

(c) First, calculate the sum:


4
∑︂
f (x) = 0.35 + 0.33 + 0.34 − 0.02
x=1
= 0.68 + 0.34 − 0.02
= 1.02 − 0.02
= 1.

Although the probabilities sum to 1, one assigned value is negative:

f (4) = −0.02 < 0.

This violates the non-negativity condition. Therefore,

The assignment in part (c) is not a valid probability distribution.

Page 52
Prof. V. K. Narla Module 2: Introduction to Probability

Question 2: Checking Probability Functions


Check whether each of the following functions can define a probability distribution. Explain
your answer.
(a)
1
f (x) = , x = 10, 11, 12, 13.
4
(b)
2x
f (x) = , x = 0, 1, 2, 3, 4, 5.
5
(c)
x − 15
f (x) = , x = 8, 9, 10, 11, 12.
20
(d)
1 + x2
f (x) = , x = 0, 1, 2, 3, 4, 5.
61
Answer:
Again, a valid probability distribution must satisfy

f (x) ≥ 0

for every allowed value of x, and ∑︂


f (x) = 1.
x

(a) For each of the four possible values,


1
f (x) = > 0.
4

There are four allowed values: 10, 11, 12, and 13. Hence,
∑︂ 1 1 1 1
f (x) = + + +
x
4 4 4 4
(︃ )︃
1
=4
4
= 1.

Therefore,

1
f (x) = , x = 10, 11, 12, 13, is a valid probability distribution.
4

Page 53
Prof. V. K. Narla Module 2: Introduction to Probability

(b) The function is


2x
f (x) = , x = 0, 1, 2, 3, 4, 5.
5
All values are nonnegative because x ≥ 0. Now calculate the total:
5 5
∑︂ ∑︂ 2x
f (x) =
x=0 x=0
5
5
2 ∑︂
= x
5 x=0
2
= (0 + 1 + 2 + 3 + 4 + 5)
5
2
= (15)
5
= 6.

Since
6 ̸= 1,
the function does not satisfy the total-probability condition. Also, for example,
6
f (3) = > 1,
5
which cannot be a probability.
Therefore,
2x
f (x) = is not a valid probability distribution.
5

(c) The function is


x − 15
f (x) = , x = 8, 9, 10, 11, 12.
20
Since every allowed value of x is less than 15,
x − 15 < 0.

For example,
8 − 15 7
f (8) = = − < 0,
20 20
and
12 − 15 3
f (12) = = − < 0.
20 20
Thus, the function assigns negative values to the possible outcomes and violates the
non-negativity condition. Therefore,
x − 15
f (x) = is not a valid probability distribution.
20

Page 54
Prof. V. K. Narla Module 2: Introduction to Probability

(d) The function is


1 + x2
f (x) = , x = 0, 1, 2, 3, 4, 5.
61
Since
1 + x2 > 0,
all function values are positive. Now calculate the total probability:
5 5
∑︂ ∑︂ 1 + x2
f (x) =
x=0 x=0
61
5
1 ∑︂
= (1 + x2 )
61 x=0
1 [︁
6 + 02 + 12 + 22 + 32 + 42 + 52 .
(︁ )︁]︁
=
61

Since
02 + 12 + 22 + 32 + 42 + 52 = 0 + 1 + 4 + 9 + 16 + 25 = 55,
we obtain
5
∑︂ 6 + 55
f (x) =
x=0
61
61
=
61
= 1.

Both conditions are satisfied. Therefore,

1 + x2
f (x) = , x = 0, 1, 2, 3, 4, 5, is a valid probability distribution.
61

8.1 The Mean and Variance of a Discrete Distribution


Let X be a discrete random variable with probability distribution

f (x) = P (X = x).

The probability distribution describes the possible values of X, but a numerical summary
is often needed to describe:
ˆ the central or average value of the distribution; and
ˆ the amount of variation or spread around that average.
The mean measures the centre of the distribution, whereas the variance and standard
deviation measure its dispersion.

Page 55
Prof. V. K. Narla Module 2: Introduction to Probability

8.2 Mean or Expected Value


The mean or expected value of a discrete random variable X is
∑︂
µ = E(X) = x f (x).
all x

Thus, each possible value x is multiplied by its probability f (x), and the resulting prod-
ucts are added.
The mean need not be one of the possible values of the random variable. It represents
the long-run average value obtained when the experiment is repeated many times.

8.3 Variance and Standard Deviation


The variance of X is the expected squared distance of X from its mean:
∑︂
σ 2 = Var(X) = (x − µ)2 f (x).
all x

The standard deviation is the positive square root of the variance:



σ = σ2.

The variance is expressed in squared units, while the standard deviation is expressed in
the same units as the random variable.

8.4 Alternative Computing Formula for Variance


The variance can also be calculated from
∑︂
σ2 = x2 f (x) − µ2 .
all x

Since ∑︂
E(X 2 ) = x2 f (x),
all x

the alternative formula is


σ 2 = E(X 2 ) − [E(X)]2 .
This form is often more convenient because it avoids calculating each individual squared
deviation (x − µ)2 .

8.5 Examples
Example 1: Number of Heads in Three Tosses of a Fair Coin
A fair coin is tossed three times. Let X denote the number of heads obtained. Find the
mean, variance, and standard deviation of X.

Page 56
Prof. V. K. Narla Module 2: Introduction to Probability

Solution.
The possible values of X are
0, 1, 2, 3.
The probability distribution is

x 0 1 2 3
1 3 3 1
f (x)
8 8 8 8

Step 1: Calculate the mean. Using


∑︂
µ= xf (x),

we obtain (︃ )︃ (︃ )︃ (︃ )︃ (︃ )︃
1 3 3 1
µ=0 +1 +2 +3
8 8 8 8
3 6 3
=0+ + +
8 8 8
12
=
8
3
= .
2
Therefore,
3
µ= = 1.5.
2

Step 2: Calculate E(X 2 ).


∑︂
E(X 2 ) = x2 f (x)
(︃ )︃ (︃ )︃ (︃ )︃ (︃ )︃
2 1 2 3 2 3 2 1
=0 +1 +2 +3
8 8 8 8
3 12 9
=0+ + +
8 8 8
24
=
8
= 3.

Step 3: Calculate the variance.


σ 2 = E(X 2 ) − µ2
(︃ )︃2
3
=3−
2
9
=3−
4
3
= .
4

Page 57
Prof. V. K. Narla Module 2: Introduction to Probability

Thus,
3
σ2 = = 0.75.
4

Step 4: Calculate the standard deviation.


√︃
3
σ=
4

3
=
2
≈ 0.866.

Hence,
µ = 1.5, σ 2 = 0.75, σ ≈ 0.866.

Example 2: Number of Preferred Attributes of a Used Car


Let X denote the number of preferred attributes possessed by a used car. The probability
distribution is

x 0 1 2 3
f (x) 0.18 0.50 0.29 0.03
Find the mean, variance, and standard deviation of X.
Solution.

Step 1: Calculate the mean.


∑︂
µ= xf (x)
= 0(0.18) + 1(0.50) + 2(0.29) + 3(0.03)
= 0 + 0.50 + 0.58 + 0.09
= 1.17.

Therefore,
µ = 1.17.

Step 2: Calculate E(X 2 ).


∑︂
E(X 2 ) = x2 f (x)
= 02 (0.18) + 12 (0.50) + 22 (0.29) + 32 (0.03)
= 0 + 0.50 + 1.16 + 0.27
= 1.93.

Page 58
Prof. V. K. Narla Module 2: Introduction to Probability

Step 3: Calculate the variance.

σ 2 = E(X 2 ) − µ2
= 1.93 − (1.17)2
= 1.93 − 1.3689
= 0.5611.

Therefore,
σ 2 = 0.5611.

Step 4: Calculate the standard deviation.



σ = 0.5611
≈ 0.7491.

Hence,
µ = 1.17, σ 2 = 0.5611, σ ≈ 0.7491.

Example 3: Number Obtained in a Roll of a Balanced Die


Let X denote the number obtained when a balanced die is rolled. Find the mean, variance,
and standard deviation of X.
Solution.
For a balanced die,
1
f (x) = , x = 1, 2, 3, 4, 5, 6.
6

Step 1: Calculate the mean.


6
∑︂
µ= xf (x)
x=1
(︃ )︃ (︃ )︃ (︃ )︃ (︃ )︃ (︃ )︃ (︃ )︃
1 1 1 1 1 1
=1 +2 +3 +4 +5 +6
6 6 6 6 6 6
1+2+3+4+5+6
=
6
21
=
6
7
= .
2
Therefore,
7
µ= = 3.5.
2

Page 59
Prof. V. K. Narla Module 2: Introduction to Probability

Step 2: Calculate E(X 2 ).


6
∑︂
E(X 2 ) = x2 f (x)
x=1
(︃ )︃ (︃ )︃ (︃ )︃ (︃ )︃ (︃ )︃ (︃ )︃
2 1 2 1 2 1 2 1 2 1 2 1
=1 +2 +3 +4 +5 +6
6 6 6 6 6 6
1 + 4 + 9 + 16 + 25 + 36
=
6
91
= .
6

Step 3: Calculate the variance.


σ 2 = E(X 2 ) − µ2
(︃ )︃2
91 7
= −
6 2
91 49
= −
6 4
182 − 147
=
12
35
= .
12
Therefore,
35
σ2 = ≈ 2.9167.
12

Step 4: Calculate the standard deviation.


√︃
35
σ=
12
≈ 1.7078.
Hence,
µ = 3.5, σ 2 ≈ 2.9167, σ ≈ 1.7078.

Summary of Formulas
∑︂
µ = E(X) = x f (x)
all x
∑︂
E(X 2 ) = x2 f (x)
all x
∑︂
σ2 = (x − µ)2 f (x) = E(X 2 ) − µ2
all x

σ= σ2

Page 60
Prof. V. K. Narla Module 2: Introduction to Probability

9. Continuous Random Variables and Probability Den-


sity Functions
A random variable X is called a continuous random variable if it can assume any value
in an interval or in a collection of intervals of the real line.
Examples of continuous random variables include:
ˆ the lifetime of a component;
ˆ the temperature of a system;
ˆ the mass of a manufactured product;
ˆ the pressure in a pipe;
ˆ the time required to complete a task.
Unlike a discrete random variable, a continuous random variable is described by a prob-
ability density function rather than by probabilities assigned to individual points.

9.1 Probability at an Individual Point


For a continuous random variable,

P (X = x) = 0

for every real number x.


This does not mean that the value x is impossible. It means that a single point has zero
area under a probability density curve.
Consequently, the probability associated with an interval is unchanged by including or
excluding either endpoint. Thus,

P (a ≤ X ≤ b) = P (a ≤ X < b) = P (a < X ≤ b) = P (a < X < b).

Therefore, in the continuous case, the distinction between < and ≤ does not affect an
interval probability.

9.2 Probability Density Function


A function f (x) is called a probability density function, or simply a density function,
of a continuous random variable X if probabilities are obtained from areas under the graph
of f .
For any real numbers a < b,
∫︂ b
P (a ≤ X ≤ b) = f (x) dx.
a

Page 61
Prof. V. K. Narla Module 2: Introduction to Probability

Since endpoint probabilities are zero,


∫︂ b
P (a < X < b) = f (x) dx
a

as well.
The function f (x) itself is not a probability. The probability is the area under f (x) over
an interval.

9.3 Conditions for a Probability Density Function


A real-valued function f (x) can serve as a probability density function only if it satisfies the
following two conditions.

Condition 1: Non-negativity

f (x) ≥ 0 for all x.


A density cannot be negative because probabilities cannot be negative.

Condition 2: Total Area Equal to One


∫︂ ∞
f (x) dx = 1.
−∞

This condition states that the total probability of all possible values of the random
variable is 1.
Thus, the two defining conditions are

f (x) ≥ 0

and ∫︂ ∞
f (x) dx = 1.
−∞

9.4 Cumulative Distribution Function


The cumulative distribution function of a continuous random variable X is defined by

F (x) = P (X ≤ x).

If f (x) is the probability density function of X, then


∫︂ x
F (x) = f (t) dt.
−∞

Thus, F (x) is the total area under the density curve to the left of x.

Page 62
Prof. V. K. Narla Module 2: Introduction to Probability

For any a < b,


P (a < X ≤ b) = F (b) − F (a).
Since endpoint probabilities are zero,

P (a ≤ X ≤ b) = F (b) − F (a).

Whenever the derivative exists,

dF (x)
= f (x).
dx

Hence, the cumulative distribution function is obtained by integrating the density, while
the density is obtained by differentiating the cumulative distribution function.

9.5 Properties of the Cumulative Distribution Function


Every cumulative distribution function satisfies:
1.
0 ≤ F (x) ≤ 1;

2. F (x) is nondecreasing;
3.
lim F (x) = 0;
x→−∞

4.
lim F (x) = 1.
x→∞

9.6 Examples
Example 1: Calculating Probabilities from a Probability Density Function
Suppose a continuous random variable X has the probability density function
{︄ −2x
2e , x > 0,
f (x) =
0, x ≤ 0.

Find:
(a) P (1 ≤ X ≤ 3);
(b) P (X > 0.5).

Page 63
Prof. V. K. Narla Module 2: Introduction to Probability

Solution.
(a) Probability that X lies between 1 and 3
Since f (x) = 2e−2x for x > 0,
∫︂ 3
P (1 ≤ X ≤ 3) = 2e−2x dx.
1

Because ∫︂
2e−2x dx = −e−2x ,

we obtain ]︁3
P (1 ≤ X ≤ 3) = −e−2x 1
[︁

= −e−6 + e−2
= e−2 − e−6
≈ 0.1353 − 0.0025
≈ 0.1329.
Therefore,
P (1 ≤ X ≤ 3) ≈ 0.133.
(b) Probability that X > 0.5
∫︂ ∞
P (X > 0.5) = 2e−2x dx.
0.5
Hence, ]︁∞
P (X > 0.5) = −e−2x 0.5
[︁

= 0 − (−e−1 )
= e−1
≈ 0.3679.
Therefore,
P (X > 0.5) ≈ 0.368.

Example 2: Determining a Distribution Function from a Density Function


For the density {︄ −2x
2e , x > 0,
f (x) =
0, x ≤ 0,
determine the cumulative distribution function F (x). Then find

P (X ≤ 1).

Solution.

Page 64
Prof. V. K. Narla Module 2: Introduction to Probability

By definition, ∫︂ x
F (x) = f (t) dt.
−∞

We consider two cases.


Case 1: x ≤ 0
Since the density is zero for t ≤ 0,

F (x) = 0.

Case 2: x > 0
∫︂ x
F (x) = 2e−2t dt
[︁ 0 −2t ]︁x
= −e 0
= −e−2x + 1
= 1 − e−2x .
Therefore,
{︄
0, x ≤ 0,
F (x) =
1 − e−2x , x > 0.

Now,
P (X ≤ 1) = F (1)
= 1 − e−2
≈ 1 − 0.1353
≈ 0.8647.
Hence,
P (X ≤ 1) ≈ 0.865.

Example 3: Finding a Constant so that a Function Is a Probability Density


Find the value of k such that
{︄
0, x ≤ 0,
f (x) =
−4x2
kxe , x > 0,

is a probability density function.


Solution.
Since
2
x > 0, e−4x > 0,
the function is nonnegative whenever
k ≥ 0.

Page 65
Prof. V. K. Narla Module 2: Introduction to Probability

To satisfy the total-area condition,


∫︂ ∞
f (x) dx = 1.
−∞

Because f (x) = 0 for x ≤ 0,


∫︂ ∞
2
kxe−4x dx = 1.
0

Let
u = 4x2 .
Then
du
du = 8x dx =⇒ x dx = .
8
Therefore,

k ∞ −u
∫︂ ∫︂
−4x2
kxe dx = e du
0 8 0
k [︁ −u ]︁∞
= −e 0
8
k
= (1)
8
k
= .
8
The total area must equal 1, so
k
= 1.
8
Hence,
k = 8.
Thus, the required density is
{︄
0, x ≤ 0,
f (x) = 2
8xe−4x , x > 0.

9.7 Mean of a Continuous Probability Distribution


Let X be a continuous random variable with density f (x). The mean or expected value
of X is ∫︂ ∞
µ = E(X) = xf (x) dx,
−∞

provided that the integral exists.


The mean gives the long-run average value of the random variable.

Page 66
Prof. V. K. Narla Module 2: Introduction to Probability

9.8 Moments of a Continuous Distribution


The kth moment about the origin is
∫︂ ∞
µ′k = E(X ) = k
xk f (x) dx.
−∞

In particular,
µ′1 = E(X) = µ
and ∫︂ ∞
µ′2 = E(X ) = 2
x2 f (x) dx.
−∞
The kth moment about the mean is
∫︂ ∞
k
µk = E[(X − µ) ] = (x − µ)k f (x) dx.
−∞

9.9 Variance and Standard Deviation


The second moment about the mean is called the variance:
∫︂ ∞
2
σ = Var(X) = (x − µ)2 f (x) dx.
−∞

An equivalent computing formula is


∫︂ ∞
2
σ = x2 f (x) dx − µ2 .
−∞

Therefore,
σ 2 = E(X 2 ) − [E(X)]2 .
The standard deviation is √
σ= σ2.

Example 4: Mean, Variance, and Standard Deviation from a Density Function


For the density {︄ −2x
2e , x > 0,
f (x) =
0, x ≤ 0,
find the mean, variance, and standard deviation.
Solution.
Step 1: Calculate the mean
∫︂ ∞ ∫︂ ∞
µ= xf (x) dx = 2xe−2x dx.
−∞ 0

Page 67
Prof. V. K. Narla Module 2: Introduction to Probability

Using integration by parts, let

u = x, dv = 2e−2x dx.

Then
du = dx, v = −e−2x .
Therefore, ∫︂ ∞
−2x ∞
e−2x dx
[︁ ]︁
µ = −xe 0
+
[︃ ]︃∞0
1 −2x
=0+ − e
2 0
1
= .
2
Hence,
1
µ= .
2
Step 2: Calculate E(X 2 )
∫︂ ∞
E(X ) = 2
2x2 e−2x dx.
0
Using the standard integral
∫︂ ∞
n!
xn e−ax dx = , a > 0,
0 an+1
with n = 2 and a = 2, ∫︂ ∞
2! 1
x2 e−2x dx = 3
= .
0 2 4
Therefore, (︃ )︃
2 1 1
E(X ) = 2 = .
4 2
Step 3: Calculate the variance

σ 2 = E(X 2 ) − µ2
(︃ )︃2
1 1
= −
2 2
1 1
= −
2 4
1
= .
4
Thus,
1
σ2 = .
4

Page 68
Prof. V. K. Narla Module 2: Introduction to Probability

Step 4: Calculate the standard deviation


√︃
1
σ=
4
1
= .
2
Therefore,
1 1 1
µ= , σ2 = , σ= .
2 4 2

Summary of Main Formulas


∫︂ b
P (a ≤ X ≤ b) = f (x) dx.
a
∫︂ ∞
f (x) ≥ 0, f (x) dx = 1.
−∞
∫︂ x
F (x) = P (X ≤ x) = f (t) dt.
−∞

P (a ≤ X ≤ b) = F (b) − F (a).

dF (x)
= f (x).
dx
∫︂ ∞
µ = E(X) = xf (x) dx.
−∞
∫︂ ∞
E(X 2 ) = x2 f (x) dx.
−∞
∫︂ ∞
σ2 = (x − µ)2 f (x) dx = E(X 2 ) − µ2 .
−∞

σ= σ2.

Practice Questions and Answers


Question 1
If the probability density of a random variable is given by
{︄
k + 2x3 , 0 < x < 1,
f (x) =
0, elsewhere,
find the value of k and the probability that the random variable takes on a value

Page 69
Prof. V. K. Narla Module 2: Introduction to Probability

3
(a) greater than ;
4
1 2
(b) between and .
3 3
Answer:
For f (x) to be a probability density function, it must satisfy
∫︂ ∞
f (x) dx = 1.
−∞

Since f (x) = 0 outside (0, 1),


∫︂ 1
(k + 2x3 ) dx = 1.
0

Therefore,
1 ]︃1
x4
∫︂ [︃
3
(k + 2x ) dx = kx +
0 2 0
1
=k+ .
2
Hence,
1
k+ = 1,
2
so that
1
k= .
2
Thus, the density is ⎧
⎨ 1 + 2x3 , 0 < x < 1,
f (x) = 2
⎩0, elsewhere.

3
(a) Probability that X >
4
(︃ )︃ ∫︂ 1 (︃ )︃
3 1
P X> = + 2x3 dx.
4 3/4 2
Now,
x x4
∫︂ (︃ )︃
1
+ 2x3 dx = + .
2 2 2

Page 70
Prof. V. K. Narla Module 2: Introduction to Probability

Therefore,
]︃1
x x4
(︃ )︃ [︃
3
P X> = +
4 2 2 3/4
(︃ )︃ (︄ (︃ )︃4 )︄
1 1 3 1 3
= + − +
2 2 8 2 4
(︃ )︃
3 81
=1− +
8 512
273
=1−
512
239
= .
512
Hence,
(︃ )︃
3 239
P X> = ≈ 0.4668.
4 512

1 2
(b) Probability that <X<
3 3
(︃ )︃ ∫︂ 2/3 (︃ )︃
1 2 1 3
P <X< = + 2x dx.
3 3 1/3 2
Thus,
]︃2/3
x x4
(︃ )︃ [︃
1 2
P <X< = +
3 3 2 2 1/3
(︃ )︃ (︃ )︃
1 8 1 1
= + − +
3 81 6 162
35 14
= −
81 81
21
=
81
7
= .
27
Therefore,
(︃ )︃
1 2 7
P <X< = ≈ 0.2593.
3 3 27

Page 71
Prof. V. K. Narla Module 2: Introduction to Probability

Question 2
If the probability density of a random variable is given by


⎪ x, 0 < x < 1,

f (x) = 2 − x, 1 ≤ x < 2,


0, elsewhere,

find the probabilities that a random variable having this probability density will take on a
value
(a) between 0.2 and 0.8;
(b) between 0.6 and 1.2.
Answer:
The density changes its formula at x = 1. Therefore, an interval that crosses x = 1 must
be divided into two integrals.

(a) Probability that 0.2 < X < 0.8


The entire interval lies in 0 < x < 1, where

f (x) = x.

Hence, ∫︂ 0.8
P (0.2 < X < 0.8) = x dx
0.2
[︃ 2 ]︃0.8
x
=
2 0.2
0.82 − 0.22
=
2
0.64 − 0.04
=
2
0.60
=
2
= 0.30.
Therefore,
P (0.2 < X < 0.8) = 0.30.

(b) Probability that 0.6 < X < 1.2


Since the interval crosses x = 1,
∫︂ 1 ∫︂ 1.2
P (0.6 < X < 1.2) = x dx + (2 − x) dx.
0.6 1

Page 72
Prof. V. K. Narla Module 2: Introduction to Probability

For the first part,


1 ]︃1
x2
∫︂ [︃
x dx =
0.6 2 0.6
1 − 0.36
=
2
= 0.32.
For the second part,
1.2 ]︃1.2
x2
∫︂ [︃
(2 − x) dx = 2x −
1 2 1
(︃ )︃ (︃ )︃
1.44 1
= 2.4 − − 2−
2 2
= (2.4 − 0.72) − 1.5
= 1.68 − 1.5
= 0.18.

Therefore,
P (0.6 < X < 1.2) = 0.32 + 0.18
= 0.50.
Hence,
P (0.6 < X < 1.2) = 0.50.

Question 3
Given the probability density
k
f (x) = , −∞ < x < ∞,
1 + x2
find k.
Answer:
A probability density must satisfy
∫︂ ∞
f (x) dx = 1.
−∞

Therefore, ∫︂ ∞
k
dx = 1.
−∞ 1 + x2
Taking k outside the integral,
∫︂ ∞
1
k dx = 1.
−∞ 1 + x2

Page 73
Prof. V. K. Narla Module 2: Introduction to Probability

Since ∫︂
1
2
dx = tan−1 x,
1+x
we obtain ∫︂ ∞
1 [︁ −1 ]︁∞
k dx = k tan x −∞
−∞ 1 + x2
(︂ π (︂ π )︂)︂
=k − −
2 2
= kπ.
Thus,
kπ = 1,
and hence
1
k= .
π
Therefore, the normalized density is
1
f (x) = , −∞ < x < ∞.
π(1 + x2 )

Question 4
If the distribution function of a random variable is given by

⎨1 − 4 , x > 2,

F (x) = x2
0, x ≤ 2,

find the probabilities that this random variable will take on a value
(a) less than 3;
(b) between 4 and 5.
Answer:
For a continuous random variable,

F (x) = P (X ≤ x).

Also,
P (a < X < b) = F (b) − F (a).

Page 74
Prof. V. K. Narla Module 2: Introduction to Probability

(a) Probability that X < 3


Since the variable is continuous,

P (X < 3) = P (X ≤ 3) = F (3).

Therefore,
4
P (X < 3) = 1 −
32
4
=1−
9
5
= .
9
Hence,
5
P (X < 3) = ≈ 0.5556.
9

(b) Probability that 4 < X < 5


P (4 < X < 5) = F (5) − F (4).
Now,
4 21
F (5) = 1 − = ,
25 25
and
4 1 3
F (4) = 1 − =1− = .
16 4 4
Therefore,
21 3
P (4 < X < 5) = −
25 4
84 − 75
=
100
9
= .
100
Thus,
P (4 < X < 5) = 0.09.

Question 5
Let the phase error in a tracking device have probability density

⎨cos x, 0 < x < π ,


f (x) = 2
⎩0, elsewhere.

Find the probability that the phase error is


π
(a) between 0 and ;
4

Page 75
Prof. V. K. Narla Module 2: Introduction to Probability

π
(b) greater than .
3
Answer:
First, observe that
∫︂ π/2
π/2
cos x dx = [sin x]0 = 1,
0
so the given function is a valid probability density.

π
(a) Probability that 0 < X <
4
∫︂ π/4
(︂π )︂
P 0<X< = cos x dx
4 0
π/4
= [sin x]0
π
= sin − sin 0
√ 4
2
= .
2
Hence,

(︂ π )︂ 2
P 0<X< = ≈ 0.7071.
4 2

π
(b) Probability that X >
3
∫︂ π/2
π )︂
(︂
P X> = cos x dx
3 π/3
π/2
= [sin x]π/3

3
=1− .
2
Therefore,

π )︂
(︂ 3
P X> =1− ≈ 0.1340.
3 2

Question 6
The length of satisfactory service, in years, provided by a certain model of laptop computer
is a random variable having the probability density

⎨ 1 e−x/4.5 , x > 0,

f (x) = 4.5
0, x ≤ 0.

Find the probabilities that one of these laptops will provide satisfactory service for

Page 76
Prof. V. K. Narla Module 2: Introduction to Probability

(a) at most 2.5 years;


(b) anywhere from 4 to 6 years;
(c) at least 6.75 years.
Answer:
The given density is exponential with mean 4.5. For x > 0, its cumulative distribution
function is ∫︂ x
1 −t/4.5
F (x) = e dt.
0 4.5
Since ∫︂
1 −t/4.5
e dt = −e−t/4.5 ,
4.5
we obtain
F (x) = 1 − e−x/4.5 , x > 0.
Also, the survival probability is

P (X > x) = e−x/4.5 .

(a) At most 2.5 years


P (X ≤ 2.5) = F (2.5)
= 1 − e−2.5/4.5
= 1 − e−5/9
≈ 0.4262.
Therefore,
P (X ≤ 2.5) ≈ 0.4262.

(b) From 4 to 6 years


P (4 ≤ X ≤ 6) = F (6) − F (4).
Therefore,
P (4 ≤ X ≤ 6) = 1 − e−6/4.5 − 1 − e−4/4.5
(︁ )︁ (︁ )︁

= e−4/4.5 − e−6/4.5
= e−8/9 − e−4/3
≈ 0.4111 − 0.2636
≈ 0.1475.
Hence,
P (4 ≤ X ≤ 6) ≈ 0.1475.

Page 77
Prof. V. K. Narla Module 2: Introduction to Probability

(c) At least 6.75 years

P (X ≥ 6.75) = e−6.75/4.5
= e−1.5
≈ 0.2231.
Therefore,
P (X ≥ 6.75) ≈ 0.2231.

Module 2
END OF THE UNIT
INTRODUCTION TO PROBABILITY

Review the key concepts, formulas, examples, and practice problems before proceeding to the next
unit.

Page 78

You might also like