0% found this document useful (0 votes)
12 views135 pages

Probability 1 Lecture Notes Overview

The Probability 1 lecture notes by Clare Wallace cover fundamental concepts in probability, including axioms, counting principles, conditional probability, and random variables. The document is structured into eight main sections, each detailing various aspects of probability theory and its applications. It also includes historical context and interpretations of probability, making it a comprehensive resource for understanding the subject.

Uploaded by

khoshkenarian
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
12 views135 pages

Probability 1 Lecture Notes Overview

The Probability 1 lecture notes by Clare Wallace cover fundamental concepts in probability, including axioms, counting principles, conditional probability, and random variables. The document is structured into eight main sections, each detailing various aspects of probability theory and its applications. It also includes historical context and interpretations of probability, making it a comprehensive resource for understanding the subject.

Uploaded by

khoshkenarian
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Probability 1 lecture notes

Clare Wallace

2025-10-02
Table of contents

Welcome to Probability 1 3
How to use these notes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3

1 Axioms of probability 5
1.1 Sets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
1.2 Sample space and events . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
1.3 Event calculus . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
1.4 Sigma-algebras . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
1.5 The axioms of probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13
1.6 Consequences of the axioms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
1.7 Historical context . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18

2 Equally likely outcomes and counting principles 20


2.1 Classical probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
2.2 Counting principles . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21
2.2.1 The multiplication principle . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22
2.2.2 Order matters; objects are distinct . . . . . . . . . . . . . . . . . . . . . . . . . . . 23
2.2.3 Order doesn’t matter; objects are distinct . . . . . . . . . . . . . . . . . . . . . . . 24
2.2.4 Separating objects into groups . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
2.3 Historical context . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28

3 Conditional probability and independence 29


3.1 Conditional probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29
3.2 Properties of conditional probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31
3.3 Independence of events . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 36
3.4 Historical context . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 39

4 Interpretations of probability 41
4.1 Relative frequency interpretation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41
4.2 Betting interpretation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42
4.3 Interpretation and the axioms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42

5 Some applications of probability 44


5.1 Reliability of networks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44
5.2 Genetics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 46
5.3 Hardy-Weinberg equilibrium . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50
5.4 Historical context . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 51

6 Random variables 54
6.1 Definition and notation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 55
6.2 Discrete random variables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 57
6.3 The binomial and geometric distributions . . . . . . . . . . . . . . . . . . . . . . . . . . . 59
6.4 The Poisson distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61

2
6.5 Continuous random variables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 64
6.6 The uniform distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 67
6.7 The exponential distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 67
6.8 The normal distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 68
6.9 Cumulative distribution functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69
6.10 Standard normal tables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 72
6.11 Functions of random variables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 73
6.12 Historical context . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 75

7 Multiple random variables 77


7.1 Joint probability distributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 77
7.2 Jointly distributed discrete random variables . . . . . . . . . . . . . . . . . . . . . . . . . 79
7.3 Jointly continuously distributed random variables . . . . . . . . . . . . . . . . . . . . . . . 85
7.4 Functions of multiple random variables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 91

8 Expectation 94
8.1 Definition and interpretation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 94
8.2 Expectation of functions of random variables . . . . . . . . . . . . . . . . . . . . . . . . . 97
8.3 Linearity of expectation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 100
8.4 Variance and covariance . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 102
8.5 Conditional expectation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 111
8.6 Independence: multiplication rule for expectation . . . . . . . . . . . . . . . . . . . . . . . 115
8.7 Expectation and probability inequalities . . . . . . . . . . . . . . . . . . . . . . . . . . . . 118
8.8 Historical context . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 119

9 Limit theorems 121


9.1 The weak law of large numbers . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 121
9.2 The central limit theorem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 123
9.3 Moment generating functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 128
9.4 Historical context . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 133

References 134

3
Welcome to Probability 1

Welcome to Probability 1! These lecture notes contain all the mathematical content you’ll need to know
to succeed in Probability this year.
If you have questions about any of the content here, try one of the following:

• ask a friend!
• ask me! I like to answer emails, and I am often in my office (MCS3060): you can (and should) pop
by to see if I’m around. You can do this during my official office hours (Mondays, 10-12) for a
guaranteed speedy response, but you definitely shouldn’t wait until then, especially if it’s a short or
quick question.
• Google it, or try a textbook. There are some good ones on the reading list (see below).

These notes have been developed over the years by several members of the Statistics and Probability
groups, including (most recently) Debleena Thacker and Andrew Wade.

Warning

There could still be typos. If you find one, let me know about it and you can have a free bag of
Skittles.

How to use these notes

The notes contain all the mathematical content for the course. In lectures, we will start at the beginning
and work our way through the whole document, until we reach the end (hopefully, this will happen exactly
at the end of term).
Throughout the notes, there are boxes like this one:

� Try it out

You can do the “Introductory” exercises on the problem sheet already.

These contain examples you can work through to check your understanding. Wherever possible, I’ve also
worked examples into the text, but there are some places where I want to give you an extra example.
These come in purple boxes.
Content that’s particularly important for the course is highlighted in red:

� Key idea

Probability is cooler than statistics

while advanced material is highlighted in blue:

4
Advanced content
Probability is almost surely cooler than statistics.

You’ll also find textbook recommendations, with the relevant sections:

� Textbook references

If you want more help with this section, check out:

• Appendix A.1 in (Blitzstein and Hwang 2019);


• Appendix B in (Anderson, Seppäläinen, and Valkó 2018);
• or the Appendix to Chapter 0 in (Stirzaker 2003).

The library has lots of good books on introductory probability, and there are even more available online/to
buy. The following four textbooks are a good starting point:

• (Blitzstein and Hwang 2019) covers the material in depth and uses simulation code to illustrate the
theory.
• (Anderson, Seppäläinen, and Valkó 2018) covers just about everything in the course at about the
right level of detail.
• (Stirzaker 2003) is concise and the most mathematically advanced, and will be useful for students
taking 2H probability.
• (DeGroot and Schervish 2013) has a statistical perspective, covering this course as well as a lot of
Statistics.

5
1 Axioms of probability

� Goals

1. Understand elementary set theory and how to use it to formulate probabilistic scenarios and to
describe the calculus of events.

2. Be familiar with the axioms of probability and their consequences, and how these properties
may be deduced from the axioms.

In this chapter, we lay the foundations of probability calculus, and establish the main techniques for
practical calculations with probabilities. The mathematical theory of probability is based on axioms, like
Euclidean geometry. In classical geometry, the fundamental objects posited by the axioms are points and
lines; in probability, they are events and their probabilities. The language and apparatus of set theory is
used to express these concepts and to work with them.
There is a lot of ambiguity inherent in probability, because we are often using mathematical approaches to
describe real-world scenarios. In some cases, there are several different ways to represent the real-world
scenario as a probabilistic model, and the choices we make could affect our conclusions. In others, an
unambiguous mathematical setup could have different real-world interpretations, depending on how we
view it. Either way, once we have a probabilistic model, the axioms help us to ensure that the maths
remains the same.
The axioms and properties of probability we develop in this chapter lay the foundations for all the rest of
the theory we will build later in the course.

1.1 Sets

One of the key tools we need in this chapter is a good understanding of set theory. You’ll see all of this
much more formally in Analysis, but in this section we give a quick rundown of the essentials we need for
Probability.
In essence, a set is an unordered collection of distinguishable objects; these objects can be numbers,
functions, other sets, and so on—any mathematical object can belong to a set.
The formal notation for a set is an opening curly bracket, followed by a list of elements that belong to
the set, followed by a closing curly bracket. For instance, the set containing the elements 2, 4, and 5 is
denoted by
{2, 4, 5}.
Because the ordering of the elements is irrelevant, {2, 4, 5} and {4, 5, 2} denote the same set.

6
Definition: empty set

The set with no outcomes is called the empty set, and is denoted by ∅:

∅ ∶= {}.

A set is often denoted by a capital letter such as 𝐴, 𝐵, 𝐶, and so on.

Definition: subset

For two sets 𝐴 and 𝐵, we say that 𝐴 is a subset of 𝐵, and we write 𝐴 ⊆ 𝐵 (or 𝐵 ⊇ 𝐴), whenever
every element that belongs to 𝐴 also belongs to 𝐵, that is, for all 𝑥 ∈ 𝐴 we have 𝑥 ∈ 𝐵.

For instance, {2, 4, 5} ⊆ {1, 2, 3, 4, 5}. Note that for every set 𝐴, we have 𝐴 ⊆ 𝐴 and ∅ ⊆ 𝐴. We can also
use strict subsets, when the subset is not equal to the larger set: {2, 4, 5} ⊂ {1, 2, 3, 4, 5}.

Definition: power set

The set consisting of all subsets of a set 𝐴 is called the power set of 𝐴, and is denoted as 2𝐴 :

2𝐴 ∶= {𝐵 ∶ 𝐵 ⊆ 𝐴}.

For example, the power set of the set 𝐴 = {1, 2, 3} is

2𝐴 = {∅, {1}, {2}, {3}, {1, 2}, {1, 3}, {2, 3}, {1, 2, 3}}.

The notation 2𝐴 alludes to the size of the power set. When 𝐴 is a finite set, its power set contains 2|𝐴|
subsets. This can be proved by constructing a bijection from 2𝐴 to ordered |𝐴|-tuples of 0s and 1s, where
a 1 indicates that the corresponding element of 𝐴 is in the subset.

� Textbook references

If you want more help with this section, check out:

• Appendix A.1 in (Blitzstein and Hwang 2019);


• Appendix B in (Anderson, Seppäläinen, and Valkó 2018);
• or the Appendix to Chapter 0 in (Stirzaker 2003).

1.2 Sample space and events

Definition: scenarios, outcomes and sample space

Whenever we do some probability, it is based on a scenario in which there are various outcomes. We
assume that we know the (set of all) possible outcomes, but we are unsure about which outcome will
occur.
A sample space is a set of outcomes for this scenario with the property that one (and only one) of
these outcomes must occur.
In this course, we will usually denote the sample space by Ω, and a generic outcome by 𝜔 ∈ Ω.

For instance, suppose we roll a standard six-sided die.

7
The most obvious sample space is Ω = {1, 2, 3, 4, 5, 6}, but if one was interested only in whether the die
was odd or even, or a six or not, one could use Ω = {odd, even}, or Ω = {not a 6, 6}.
Often, like in the above example, we may enumerate the elements of the sample space Ω in a finite or
infinite list Ω = {𝜔1 , 𝜔2 , …}, in which case we say the set Ω is countable or discrete.
A set is said to be countable when its elements can be enumerated in a (possibly infinite) sequence. Every
finite set is countable, and so is the set of natural numbers ℕ ∶= {1, 2, 3, …}. The set of integers ℤ is
countable as well. The set of real numbers ℝ is not countable, and neither is any interval [𝑎, 𝑏] when
𝑎 < 𝑏.

Definition: countable

A set 𝐴 is countable if either:

• 𝐴 is finite, or
• there is a bijection (one-to-one and onto mapping) between 𝐴 and the set of natural numbers
ℕ.

One can prove that the set of rational numbers ℚ is countable.


When we perform an experiment we are interested in the occurence, or otherwise, of events. An event is
just a collection of possible outcomes, i.e., a subset of Ω.

� Key idea: Definition: events

Associated to our sample space Ω is a collection ℱ of events:

𝐴 ⊆ Ω for every 𝐴 ∈ ℱ.

We say that an event 𝐴 occurs when the outcome that occurs at the end of the scenario is in the set
𝐴.

If Ω is discrete, we can always take ℱ = 2Ω , so that every subset of Ω is an event. If Ω is not discrete, we
need to be a little more careful: see Section 1.4 below.
The empty set ∅ represents the impossible event, i.e., it will never occur. The sample space Ω represents
the certain event, i.e., it will always occur. Most interesting events are somewhere in between.
The representation of an event as a set obviously depends on the choice of sample space Ω for the specific
scenario under study, as shown by the following two examples.

Examples

1. For rolling a standard cubic die (with Ω = {1, 2, 3, 4, 5, 6}), the event “throw an odd number”
is the subset 𝐴 = {1, 3, 5} consisting of three outcomes. If we roll the die and it comes up a 3,
then event 𝐴 has occurred.

2. For the same scenario, but with Ω = {odd, even}, the event ‘throw an odd number’ is the
subset 𝐴 = {odd} consisting of just one outcome.

8
� Try it out

Suppose we roll three standard six-sided dice and record the outcome of the experiment by an ordered
triple (𝑖, 𝑗, 𝑘) where 𝑖, 𝑗, 𝑘 ∈ {1, 2, … , 6}. What is Ω? How big could ℱ be?
Answer:
In this case Ω = {1, 2, … , 6}3 has 63 = 216 individual outcomes.
The number of events (2216 ) is enormous! One is the event that the scores on the three dice are the
same: 𝐴 = {(1, 1, 1), (2, 2, 2), … , (6, 6, 6)}.

� Textbook references

If you want more help with this section, check out:

• Section 1.2 in (Blitzstein and Hwang 2019);


• Section 1.1 in (Anderson, Seppäläinen, and Valkó 2018);
• Sections 1.1 and 1.2 in (Stirzaker 2003);
• or Section 1.4 in (DeGroot and Schervish 2013).

1.3 Event calculus

Once we’ve defined our sample space and the set of all possible events, we need to be able to refer to
combinations of events. To do so, we use standard notation from set theory.

Definition: complements

For an event 𝐴 ∈ ℱ, we define its complement, denoted 𝐴c (or sometimes 𝐴)̄ and read “not 𝐴”, to be

𝐴c ∶= Ω\𝐴 = {𝜔 ∈ Ω ∶ 𝜔 ∉ 𝐴}.

Notice that:

• the complement of 𝐴c is 𝐴: (𝐴c )c = 𝐴;


• there are no outcomes in both 𝐴 and 𝐴c : 𝐴 ∩ 𝐴c = ∅;
• and every outcome is in one or the other: 𝐴 ∪ 𝐴c = Ω.

� Key idea: event calculus

Given any two events 𝐴 and 𝐵 that are associated with the same sample space (i.e. 𝐴 ⊆ Ω and
𝐵 ⊆ Ω for the same Ω), here are some of the other events we can define, along with how we would
read them out:

Notation We say (as sets) We say (as events) Meaning (as events)
𝐴∪𝐵 𝐴 union 𝐵 𝐴 or 𝐵 𝐴 occurs or 𝐵 occurs or both 𝐴
and 𝐵 occur
𝐴∩𝐵 𝐴 intersect 𝐵 𝐴 and 𝐵 𝐴 occurs and 𝐵 occurs
c
𝐴 ∶= Ω\𝐴 𝐴 complement not 𝐴 𝐴 does not occur
𝐴\𝐵 𝐴 minus 𝐵 𝐴 but not 𝐵 𝐴 occurs but 𝐵 does not
𝐴⊆𝐵 𝐴 is a subset of 𝐵 𝐴 implies 𝐵 if 𝐴 occurs, then 𝐵 must occur

9
(In the final row, “𝐴 ⊆ 𝐵” is not an event but rather a statement about how two events relate to
each other. I still wanted to include it because I think it’s helpful)

� Try it out

Prove that 𝐴\𝐵 = 𝐴 ∩ 𝐵c .


Answer:
We can do this by working with the events as sets. We have

𝐴\𝐵 = {𝜔 ∈ Ω ∶ 𝜔 ∈ 𝐴, 𝜔 ∉ 𝐵} = {𝜔 ∈ Ω ∶ 𝜔 ∈ 𝐴} ∩ {𝜔 ∈ Ω ∶ 𝜔 ∉ 𝐵} = 𝐴 ∩ 𝐵c .

� Key idea: Definition: disjoint

We say that events 𝐴 and 𝐵 are disjoint, mutually exclusive, or incompatible if 𝐴 ∩ 𝐵 = ∅, i.e., it is
impossible for 𝐴 and 𝐵 both to occur.

� Try it out

Consider the sample space Ω ∶= {1, 2, 3, 4, 5, 6}, and the events

𝐴 ∶= {2, 4, 6},
𝐵 ∶= {1, 3, 5},
𝐶 ∶= {1, 2, 3}.

In other words, 𝐴 is the event “throw an even number”, 𝐵 is the event “throw an odd number”, and
𝐶 is the event “throw at most three”. Use some of the ideas from the table above to combine events
𝐴, 𝐵, and 𝐶 in different ways. Are any of your new events disjoint?
Answer:
Some combinations:

𝐴∪𝐵=Ω (even or odd)


𝐴∩𝐵=∅ (even and odd)
𝐴c = 𝐵 (not even)
𝐶\𝐴 = {1, 3} (at most 3 but not even)
𝐴 ∪ 𝐶 = {1, 2, 3, 4, 6} (even or at most 3)
𝐴 ∩ 𝐶 = {2} (even and at most 3).

The events 𝐴 and 𝐵 are disjoint as 𝐴 ∩ 𝐵 = ∅. We have also created two disjoint events: 𝐶\𝐴 and
𝐴 ∩ 𝐶. Think about why these two events would always be disjoint, however we define 𝐴 and 𝐶.

10
� Try it out

Toss a coin twice and denote the sample space by Ω = {HH, HT, TH, TT}. Consider the events

𝐴 ∶= {HH, HT} (first toss H)


𝐵 ∶= {HT, TT} (second toss T)
𝐶 ∶= {HH} (both H).

How do these events relate to each other?


Answer:
Some things you might notice:

• 𝐶 ⊆ 𝐴, i.e., if 𝐶 occurs then 𝐴 must occur;


• 𝐴 ∪ 𝐵 = {HH, HT, TT} is the event that either the first toss is H, the second toss is T, or both;
• 𝐴 ∩ 𝐵 = {HT};
• 𝐴c = {TH, TT};
• 𝐵 ∩ 𝐶 = ∅.

� Try it out

Draw a card from a standard deck of 52 playing cards. Take Ω to consist of each of the 52 possible
draws: Ω = {A♣, A♢, … , K♡, K♠}. Events in ℱ = 2Ω include

𝐸 = {eight} = {8♠, 8♡, 8♢, 8♣},


𝑆 = {spade} = {𝐴♠, 2♠, … , 𝐾♠},

and we can combine them to form other events, such as

𝐸 ∩ 𝑆 = {8♠},
𝐸\𝑆 = {8♡, 8♢, 8♣},
𝑆\𝐸 = {𝐴♠, 2♠, … , 7♠, 9♠, … , 𝐾♠}.

As with sums (∑) and products (Π) of multiple numbers, we also have shorthands for unions and
intersections of multiple sets:
𝑛
⋃ 𝐴𝑖 ∶= 𝐴1 ∪ 𝐴2 ∪ ⋯ ∪ 𝐴𝑛
𝑖=1
is the event that at least one of 𝐴1 , 𝐴2 , … 𝐴𝑛 occurs (or the set of all 𝜔 ∈ Ω which are contained in at
least one of the 𝐴𝑖 s), and
𝑛
⋂ 𝐴𝑖 ∶= 𝐴1 ∩ 𝐴2 ∩ ⋯ ∩ 𝐴𝑛
𝑖=1
is the event that all of 𝐴1 , 𝐴2 , … 𝐴𝑛 occur (or the set of all 𝜔 ∈ Ω which are in every 𝐴𝑖 ).
Occasionally, we will also need to take infinite unions and intersections over sequences of sets:

⋃ 𝐴𝑖 ∶= 𝐴1 ∪ 𝐴2 ∪ 𝐴3 ∪ …
𝑖=1

⋂ 𝐴𝑖 ∶= 𝐴1 ∩ 𝐴2 ∩ 𝐴3 ∩ … .
𝑖=1

We will also sometimes need De Morgan’s Laws: for a (possibly infinite) collection of events 𝐴𝑖 ,

11
c
a. The complement of the union is the intersection of the complements: (⋃𝑖 𝐴𝑖 ) = ⋂𝑖 𝐴c𝑖 , and
c
b. The complement of the intersection is the union of the complements: (⋂𝑖 𝐴𝑖 ) = ⋃𝑖 𝐴c𝑖 .

These could be more intuitive than they appear: the negation of “some of these things happened” is “none
of these things happened”, and the negation of “all of these things happened” is “some of these things did
not happen”.
It is often useful to visualize the sample space in a Venn diagram. Then events such as 𝐴 are subsets of
the sample space. It is a helpful analogy to imagine the probability of an event as the area in the Venn
diagram.

Figure 1.1: Venn diagram

Advanced content

This analogy is more apt than it first appears, since the mathematical foundations of rigorous
probability theory are built on measure theory, which is the same theory that gives rigorous foundation
to the concepts of length, area, and volume.

� Textbook references

If you want more help with this section, check out:

• Section 1.2 in (Blitzstein and Hwang 2019);


• Section 1.2 in (Stirzaker 2003);
• or Section 1.4 in (DeGroot and Schervish 2013).

1.4 Sigma-algebras

In the last section we described some of the ways in which events can be combined. Now we can set out
the rules for our collection of events, ℱ, to ensure that it’s possible to use these different combinations.
We said that in the case where Ω is discrete, one can take ℱ = 2Ω .

12
In general, if Ω is uncountable, it is too much to demand that probabilities should be defined on all subsets
of Ω. The reason why this is a problem goes beyond the scope of this course (see the Bibliographical notes
at the end of this chapter for references), but the essence is that for uncountable sample spaces, such as
Ω = [0, 1], there exist subsets of Ω that cannot be assigned a probability in a way that is consistent. The
construction of such non-measurable sets is also the basis of the famous Banach–Tarski paradox.
Uncountable Ω are unavoidable: we will see an infinite coin-tossing space at the end of section Section 1.6,
and other examples occur whenever we have an experiment whose outcome is modelled by a continuous
distribution such as the normal distribution (more on this later).
The upshot of all this is that we can, in general, only demand that probabilities are defined for all events
in some collection ℱ of subsets of Ω (i.e., for some ℱ ⊆ 2Ω ). What properties should the collection
ℱ of events possess? Consideration of the set operations in the previous section suggests the following
definition.

Definition: �-algebra

A collection ℱ of subsets of Ω is called a 𝜎-algebra over Ω if it satisfies the following properties.


(S1) Ω ∈ ℱ;
(S2) 𝐴 ∈ ℱ implies that 𝐴c ∈ ℱ;
(S3) if 𝐴1 , 𝐴2 , … ∈ ℱ then ⋃𝑖=1 𝐴𝑖 ∈ ℱ.

Property S2 says that ℱ is closed under complementation, while S3 says that ℱ is closed under countable
unions.
We can combine S1 and S2 to see that we must have ∅ ∈ ℱ. Also note that, we can get to a finite-union
version of S3 by taking 𝐴𝑛+1 = 𝐴𝑛+2 = ⋯ = ∅: so ℱ is also closed under finite unions.

Examples

1. The power set 2Ω is a 𝜎-algebra over Ω, and in fact it is the biggest possible 𝜎-algebra over Ω.
As described above, for uncountable Ω the set 2Ω may be too unwieldy, in which case we would
consider a smaller 𝜎-algebra.

2. The trivial 𝜎-algebra {∅, Ω} is the smallest possible 𝜎-algebra over Ω. It’s very nicely behaved
(just two elements!) but it carries no information about the outcome of the experiment.

� Try it out

Consider the sample space Ω = {1, 2, 3, 4, 5, 6} for the experiment of rolling a fair die. The choice of
𝜎-algebra determines the resolution at which we observe the experiment, and may depend on exactly
what we are interested in:

• ℱ0 = {∅, Ω} (carries no information);


• ℱ1 = {∅, {1, 3, 5}, {2, 4, 6}, Ω} (if we only care whether the score is odd or even);
• ℱ2 = 2Ω (if we are interested in the exact score).

Note the inclusions ℱ0 ⊂ ℱ1 ⊂ ℱ2 .


Let us check that ℱ1 is indeed a 𝜎-algebra.
S1 is immediate: we can see it from the definition of ℱ1 .
For S2, we need that every 𝐴 ∈ ℱ1 is accompanied by its complement 𝐴c ; we see that this is the
case.

13
Since ℱ1 is a finite set it suffices to check S3 for finite unions. In other words, it is enough to check
that if 𝐴, 𝐵 ∈ ℱ1 , then 𝐴 ∪ 𝐵 ∈ ℱ1 too. Since there are only two sets, this is also quick: we see that
it is the case.

� Textbook references

If you want more help with this section, check out:

• Section 1.2 in (Stirzaker 2003).

1.5 The axioms of probability

� Key idea: Definition: probability

A probability ℙ on a sample space Ω with collection ℱ of events is a function mapping every event
𝐴 ∈ ℱ to a real number ℙ(𝐴), obeying the following axioms:
(A1) ℙ(𝐴) ≥ 0 for every 𝐴 ∈ ℱ;
(A2) ℙ(Ω) = 1; and
(A3) if 𝐴 and 𝐵 are disjoint events (i.e. if 𝐴, 𝐵 ∈ ℱ have 𝐴 ∩ 𝐵 = ∅) then

ℙ(𝐴 ∪ 𝐵) = ℙ(𝐴) + ℙ(𝐵).

We call the number ℙ(𝐴) the probability of 𝐴.

We will see shortly that a consequence of these axioms is that the probabilities ℙ(𝐴) must lie between 0
and 1: 0 ≤ ℙ(𝐴) ≤ 1.
We can upgrade (A3) to a slightly more technical version:
(A4) For any infinite sequence 𝐴1 , 𝐴2 , … of pairwise disjoint events (so 𝐴𝑖 ∩ 𝐴𝑗 = ∅ for all 𝑖 ≠ 𝑗),
∞ ∞
ℙ ( ⋃ 𝐴𝑖 ) = ∑ ℙ(𝐴𝑖 ).
𝑖=1 𝑖=1

� Key idea: A small request

If you only take one thing away from this course, please let it be this:
Probabilities are numbers and events are sets.
We can add up numbers (but not sets) and we can take unions and intersections of sets (but not
numbers).

For the axioms to make sense, we can’t just use any old event set ℱ. For one thing, we need Ω ∈ ℱ; in
fact all the events in (A1-4) need to be in ℱ. Our definition of a 𝜎-algebra from the previous section
gives us exactly the event set we need.

14
� Key idea: Definition: probability space

If Ω is a set and ℱ is a 𝜎-algebra of subsets of Ω, and if ℙ satisfies (A1–4) for events in ℱ, then the
triple (Ω, ℱ, ℙ) is called a probability space.

� Try it out

Consider a finite sample space Ω = {𝜔1 , … , 𝜔𝑚 } of size |Ω| = 𝑚. Then we can define a valid
probability ℙ by taking any numbers 𝑝1 , … , 𝑝𝑚 with 𝑝𝑖 ≥ 0 for all 𝑖 and ∑𝑖=1 𝑝𝑖 = 1 and declaring
𝑚

that for any event 𝐴,


ℙ(𝐴) = ∑ 𝑝𝑖 .
𝑖∶𝜔𝑖 ∈𝐴

This satisfies the axioms (A1–4) (don’t just believe me - check them for yourself).
By considering the event 𝐴 = {𝜔𝑖 }, we see that 𝑝𝑖 = ℙ(𝜔𝑖 ) is the probability of the elementary
outcome 𝜔𝑖 .
In the simplest setting, we might assume that all the outcomes are equally likely, that is, 𝑝𝑖 = 1/𝑚
for all 𝑖. Note that in this case probability reduces to counting, since

1 |𝐴|
ℙ(𝐴) = ∑ = .
𝑖∶𝜔𝑖 ∈𝐴
𝑚 |Ω|

As a concrete example, for tossing a fair die we would have Ω = {1, 2, … , 6}, and ℙ(𝐴) = |𝐴|/6 so, for
example,
3 1
ℙ(score is odd) = ℙ({1, 3, 5}) = = .
6 2
We examine this setting in detail in Chapter 2.

� Try it out

Consider a countably infinite sample space Ω = {𝜔1 , 𝜔2 , …}. Then we can define a valid probability
ℙ by taking any numbers 𝑝1 , 𝑝2 , … with 𝑝𝑖 ≥ 0 for all 𝑖 and ∑𝑖=1 𝑝𝑖 = 1 and declaring that for any

event 𝐴,
ℙ(𝐴) = ∑ 𝑝𝑖 .
𝑖∶𝜔𝑖 ∈𝐴

This definition of a probability satisfies all of the axioms (A1-A4).

For this course, we will usually assume that the probability distribution is given (and satisfies the axioms),
without worrying too much about how the important practical task of finding the probabilities was carried
out.

� Textbook references

If you want more help with this section, check out:

• Section 1.6 in (Blitzstein and Hwang 2019);


• Section 1.1 in (Anderson, Seppäläinen, and Valkó 2018);
• or Section 1.3 in (Stirzaker 2003).

15
1.6 Consequences of the axioms

A host of useful results can be derived from A1–4.

� Key idea: Consequences of the axioms

(C1) For any two events 𝐴 and 𝐵,

ℙ(𝐵\𝐴) = ℙ(𝐵) − ℙ(𝐴 ∩ 𝐵).

(C2) For any event 𝐴, ℙ(𝐴c ) = 1 − ℙ(𝐴).


(C3) The probability of ∅ is ℙ(∅) = 0.
(C4) For any event 𝐴, ℙ(𝐴) ≤ 1.
(C5) If 𝐴 ⊆ 𝐵 then ℙ(𝐴) ≤ ℙ(𝐵) (“monotonicity”).
(C6) For any two events 𝐴 and 𝐵,

ℙ(𝐴 ∪ 𝐵) = ℙ(𝐴) + ℙ(𝐵) − ℙ(𝐴 ∩ 𝐵).

(C7) If 𝐴1 , 𝐴2 , … , 𝐴𝑘 are pairwise disjoint (so 𝐴𝑖 ∩ 𝐴𝑗 = ∅ if 𝑖 ≠ 𝑗) then

𝑘 𝑘
ℙ ( ⋃ 𝐴𝑖 ) = ∑ ℙ(𝐴𝑖 ).
𝑖=1 𝑖=1

(This property is called “finite additivity” in textbooks.)


(C8) For any events 𝐴1 , 𝐴2 , …, (these need not be pairwise disjoint),
∞ ∞
ℙ( ⋃ 𝐴𝑖 ) ≤ ∑ ℙ(𝐴𝑖 ).
𝑖=1 𝑖=1

(This one is sometimes referred to as “Boole’s inequality.”)


(C9) If 𝐴1 ⊆ 𝐴2 ⊆ ⋯ is an increasing sequence of events, then

ℙ ( ⋃ 𝐴𝑛 ) = lim ℙ(𝐴𝑛 ).
𝑛→∞
𝑛=1

If 𝐴1 ⊇ 𝐴2 ⊇ ⋯ is a decreasing sequence of events, then



ℙ ( ⋂ 𝐴𝑛 ) = lim ℙ(𝐴𝑛 ).
𝑛→∞
𝑛=1

(This property is a bit more sophisticated than the previous ones. It establishes the “continuity of
probability along monotone limits:” we can take limits, as long as the events in question form a
monotone sequence. It will be really important in Probability II.)

Just one more consequence to go! Before we get there, we need the following simple but extremely useful
idea: partitions.

� Key idea: Definition: partition

We say that the events 𝐸1 , 𝐸2 , … , 𝐸𝑘 ∈ ℱ form a (finite) partition of the sample space Ω if:

i. they all have positive probability, i.e., ℙ(𝐸𝑖 ) > 0 for all 𝑖;

16
ii. they are pairwise disjoint, i.e., 𝐸𝑖 ∩ 𝐸𝑗 = ∅ whenever 𝑖 ≠ 𝑗; and
iii. their union is the whole sample space: ∪𝑘𝑖=1 𝐸𝑖 = Ω.

The definition extends to countably infinite partitions. We say that 𝐸1 , 𝐸2 , … ∈ ℱ form an infinite
partition of Ω if:

i. ℙ(𝐸𝑖 ) > 0 for all 𝑖;


ii. 𝐸𝑖 ∩ 𝐸𝑗 = ∅ whenever 𝑖 ≠ 𝑗; and
iii. ∪∞
𝑖=1 𝐸𝑖 = Ω.

For example, consider the sample space Ω = {1, 2, 3, 4, 5, 6}. Some partitions are:

{1}, {2}, {3}, {4}, {5}, {6}


{1, 2}, {3, 4}, {5, 6}
{1, 2, 3}, {4, 5, 6}
{1}, {2, 3}, {4, 5, 6}
{1, 2, 3, 4, 5, 6}

and so on.

� Key idea

(C10) If 𝐸1 , 𝐸2 , …, 𝐸𝑘 form a partition then


𝑘
∑ ℙ(𝐸𝑖 ) = 1.
𝑖=1

These consequences have an enormous effect on the way we work with probability. In particular, it turns
out that we can solve most problems without ever having to explicitly write down the outcomes in our
sample space, as in the next example. In fact, some people do probability without even defining a sample
space.

� Try it out

Jimmy’s die has the numbers 2,2,2,2,5,5. Your die has numbers 1,1,4,4,4,4. You both throw and the
highest number wins. Assuming all outcomes are equally likely, what is the probability that Jimmy
wins?
Answer:
The event, 𝐽, that Jimmy wins happens if either Jimmy throws a 5 (call this event 𝐹), or if you
throw a 1 (call this event 𝐴). Therefore 𝐽 = 𝐴 ∪ 𝐹 and by C6,

ℙ(𝐽 ) = ℙ(𝐴) + ℙ(𝐹 ) − ℙ(𝐴 ∩ 𝐹 ).

As ℙ(𝐹 ) = 1/3, ℙ(𝐴) = 1/3 and ℙ(𝐴 ∩ 𝐹 ) = 4/36 = 1/9 (by counting equally likely outcomes) we
have
ℙ(𝐽 ) = 1/3 + 1/3 − 1/9 = 5/9.

17
Finite sample spaces are a great way to build up our intuition for probability calculations. However, it is
surprisingly easy to end up in a situation where things start to get complicated.

� Try it out

What is the probability that, in an indefinitely long sequence of tosses of a fair coin, we will eventually
see heads?
Answer:
The sample space Ω is infinite and consists of all sequences 𝜔 = (𝜔1 , 𝜔2 , …) with 𝜔𝑖 ∈ {H, T}.
What is ℙ? Well, it certainly would be desirable that if we restrict to just a finite sequence of 𝑛
tosses, then each of the 2𝑛 possible outcomes (sequences) should be equally likely. It is a special case
of a general theorem that such a ℙ exists and is unique.
Now, let 𝐴 = {H occurs}. Then the only way 𝐴 can not occur is if there are no heads, i.e.,
𝐴c = {(TTT⋯)}. This is a single sequence, out of infinitely many, and it is intuitively clear that it
should have probability 0. To prove this, it is enough to observe that 𝐴c ⊆ {first 𝑛 tosses T}, so by
monotonicity (C5),
ℙ(𝐴c ) ≤ ℙ({first 𝑛 tosses are T}) ≤ 2−𝑛 .
But this is true for any 𝑛, so we must have ℙ(𝐴c ) = 0.
Another way to see this is as follows. Consider events defined for 𝑛 = 1, 2, … by

𝐴𝑛 = {first H occurs on toss 𝑛} = {𝜔 ∶ 𝜔𝑘 = T, for all 𝑘 < 𝑛, 𝜔𝑛 = H}.

This means that 𝐴1 consists of sequences H⋯, 𝐴2 consists of sequences TH⋯, and so on.
𝑛=1 𝐴𝑛 . So, by (A4),
Now the event we are interested in is 𝐴 = ∪∞
∞ ∞
ℙ(𝐴) = ∑ ℙ(𝐴𝑛 ) = ∑ 2−𝑛 = 1.
𝑛=1 𝑛=1

Note that a similar argument works if the coin is biased with probability 𝑝 ∈ (0, 1) of heads.

Advanced content

In fact, the sequence space Ω in the previous example is not even countable! To see this, a sequence
(𝑑1 , 𝑑2 , …) with each 𝑑𝑖 ∈ {0, 1} is called a dyadic expansion of 𝑥 ∈ [0, 1] if 𝑥 = ∑𝑖=1 2−𝑖 𝑑𝑖 . For

example, (1, 0, 0, …) is a dyadic expansion of 1/2, (1, 1, 0, 0, …) is 3/4, and so on. The map between
𝑥 and (𝑑1 , 𝑑2 , …) is almost a bijection. It is not a bijection because of possible non-uniqueness of the
dyadic expansion: e.g. (0, 1, 1, 1, …) is another expansion of 1/2. It turns out that this problem only
occurs for rational 𝑥, and can be circumvented. Thus we have (essentially) a bijection between [0, 1]
and the space of infinite sequences of 0s and 1s, which is another name for our coin tossing space Ω.
This shows that Ω is uncountable.
It is remarkable that the probability ℙ on infinite sequences of coin tosses turns out to correspond
(under the bijection by dyadic expansion) to nothing other than the uniform distribution on [0, 1],
that is the measure defined by lengths of intervals. This is the famous Lebesgue measure.

� Textbook references

If you want more help with this section, check out:

• Section 1.6 in (Blitzstein and Hwang 2019);


• Section 1.4 in (Anderson, Seppäläinen, and Valkó 2018);

18
• or Sections 1.4 and 1.5 in (Stirzaker 2003).

1.7 Historical context

Sets are important not only for probability theory, but for all of mathematics. In fact, all of standard
mathematics can be formulated in terms of set theory, under the assumption that sets satisfy the ZFC
axioms; see for instance this Wikipedia page.
The foundations of probability have a long and interesting history (Hacking 2006; Todhunter 2014). The
classical theory owes much to Pierre-Simon Laplace (1749–1827): see (Laplace 1825). However, a rigorous
mathematical foundation for the theory was lacking, and was posed as part of one of David Hilbert’s
(1862–1943) famous list of problems in 1900 (the 5th problem). After important work by Henri Lebesgue
(1875–1941) and 'Emile Borel (1871–1956), it was Andrey Kolmogorov who succeeded in 1933 in providing
the axioms that we use today (see the 1950 edition of his book (Kolmogorov 1950)). This approach declares
that probabilities are measures.
A measure 𝜇 can be defined on any set Ω with a 𝜎-algebra of subsets ℱ, and the defining axioms are
versions of A1 and A4. The special property of a probability measure is just that 𝜇(Ω) = 1. Measure
theory is the theory that gives mathematical foundation to the concepts of length, area, and volume. For
example, on ℝ the unique measure that has 𝜇(𝑎, 𝑏) = 𝑏 − 𝑎 for intervals (𝑎, 𝑏) is the Lebesgue measure.

(b) Boole (c) Venn


(a) Laplace (d) Kolmogorov

Figure 1.2: Laplace, Boole, Venn, and Kolmogorov

George Boole (1815–1864) and John Venn (1834–1923) both wrote books concerned with probability
theory (Boole 1854), (Venn 1888); both were working before the formulation of Kolmogorov’s axioms.
As mentioned in Section 1.4, it is necessary in the general theory of probability to restricting events to
some 𝜎-algebra. The reason for this is that in standard ZFC set theory, when Ω is uncountable (such as
Ω = [0, 1] the unit interval), it follows from an argument by Vitali (1905) that many natural probability
assessments, such as the continuous uniform distribution, cannot be modelled by a probability defined on
all subsets of Ω satisfying A1–4: see for instance Chapter 1 of (Rosenthal 2007). In the case where Ω is
countable, one can always define ℙ on the whole of 2Ω . In the case where Ω is uncountable, we usually do
not explicitly mention Ω at all (when we work with continuous random variables, for example).

19
The formulation of the infinite coin-tossing experiment in Section 1.6 leads to the connection between coin
tossing and the Lebesgue measure, as first described by Hugo Steinhaus in a 1923 paper.
An alternative approach to probability theory is to do away with axiom A4, in which case some of these
technical issues can be avoided, at the expense of certain pathologies; however, in the standard approach
to modern probability, based on measure theory, A4 is a central part of the theory.

20
2 Equally likely outcomes and counting principles

� Goals

1. Understand the equally likely outcomes model of classical probability.


2. Know counting principles, and when and how to apply them on specific problems.

In Chapter 1 we have seen the abstract formulation of probability theory; next we turn to the question of
how the probabilities themselves may be assigned.
The most basic scenario occurs when our experiment has a finite number of possible outcomes which we
deem to be equally likely.
Such situations rarely—not to say never—occur in practice, but serve as good models in extremely
controlled environments such as in gambling or games of chance. However, this situation (which will
essentially come down to counting) gives us a good initial setting in which to obtain some very useful
insights into the nature and calculus of probability.

2.1 Classical probability

Suppose that we have a finite sample space Ω. Since Ω is finite, we can list it as a collection of 𝑚 = |Ω|
possible outcomes:
Ω = {𝜔1 , … , 𝜔𝑚 }.
In the equally likely outcomes model (also sometimes known as classical probability) we suppose that each
outcome has the same probability:
1
ℙ(𝜔) = for each 𝜔 ∈ Ω,
|Ω|

or, in the notation above, ℙ(𝜔𝑖 ) = 1/𝑚 for each 𝑖.


The axioms of probability then allow us to determine the probability of any event 𝐴 ⊆ Ω: by C7,

|𝐴|
ℙ(𝐴) = ∑ ℙ(𝜔) = for any event 𝐴 ⊆ Ω.
𝜔∈𝐴
|Ω|

This is a particular case of the discrete sample space discussed in Chapter 1.

Definition: Equally likely outcomes

Consider a scenario with 𝑚 equally likely outcomes enumerated as Ω = {𝜔1 , … , 𝜔𝑚 }. In the equally
likely outcomes model, the probability of an event 𝐴 ⊆ Ω is declared to be

|𝐴|
ℙ(𝐴) ∶= .
|Ω|

21
Using this definition, we meet all of the axioms (A1–A4) (checking each of them comes down to what we
know about counting). Remember that in the case of a finite state space, we always have the option to
take ℱ = 2Ω as our 𝜎-algebra.

Examples

1. Draw a card at random from a well-shuffled pack, so that each of the 52 cards is equally likely
to be chosen. Typical events are that the card is a spade (a set of 13 outcomes), the card is a
queen (a set of 4 outcomes), the card is the queen of spades (a set of a single outcome). In the
equally likely outcomes model, the probability of drawing the queen of spades (or any other
specified card) is 1/52, the probability of drawing a spade is 13/52, and the probability of
drawing a queen is 4/52.

2. Flip a coin and see whether it falls heads or tails, each assumed equally likely; then ‘heads’ or
‘tails’ each has probability 1/2.

3. Roll a fair cubic die to get a number from 1 to 6. Here the word ‘fair’ is used to mean
each outcome is equally likely. Then Ω = {1, … , 6} and ℙ(𝐴) = |𝐴|/6. For example, if
𝐴1 = {2} (the score is 2) we get ℙ(𝐴1 ) = 1/6, while if 𝐴2 = {1, 3, 5} (the score is odd) we get
ℙ(𝐴2 ) = 3/6 = 1/2.

4. If we roll a pair of fair dice then outcomes are pairs (𝑖, 𝑗) so there are 36 possible outcomes. If
we assume that the outcomes are equally likely, then the probability of getting a pair of 6’s is
1/36, for example.

The classical interpretation of probability is the most straightforward approach we can take, just as
counting can be seen as “basic” mathematics. It is a good place to start and there are many important
situations where intuitively it seems reasonable to say that each outcome of a particular collection is
equally likely.
To extend the theory or apply it in practice we have to address situations where there are no candidates
for equally likely outcomes or where there are infinitely many possible outcomes and work out how to find
probabilities to put into calculations that give useful predictions. We will come back to some of these
issues later; but bear in mind that however we come up with our probability model, the same system of
axioms that we saw in Chapter 1 applies.

� Textbook references

If you want more help with this section, check out:

• Section 1.3 in (Blitzstein and Hwang 2019);


• or Section 1.2 in (Anderson, Seppäläinen, and Valkó 2018).

2.2 Counting principles

Given a finite sample space and assuming that outcomes are equally likely, to determine probabilities of
certain events comes down to counting.
For example, in drawing a poker hand of five cards from a well-shuffled deck of 52 cards, the probability of
having a ‘full house’ (meaning two cards of one denomination and three of another, e.g., two Kings and

22
three 7s) is given by the number of hands that are full houses divided by the total number of hands (each
hand being equally likely).
These counting problems need careful thought, and we will describe some counting principles for some of
the most common situations. There is some common ground with the Discrete Maths course; here we have
a slightly different emphasis.

2.2.1 The multiplication principle

Counting principle: Multiplication

Suppose that we must make 𝑘 choices in succession where there are:

• 𝑚1 possibilities for the first choice,

• 𝑚2 possibilities for the second choice,

• ⋮

• 𝑚𝑘 possibilities for the 𝑘th choice,

and the number of possibilities at each stage does not depend on the outcomes of any previous
choices. The total number of distinct possible selections is
𝑘
𝑚1 × 𝑚2 × 𝑚3 × ⋯ × 𝑚𝑘 = ∏ 𝑚𝑖 .
𝑖=1

For instance, in a standard deck of playing cards, each card has a denomination and a suit. There are 13
possible denominations: A(ce), 2, 3, …, 10, J(ack), Q(ueen), K(ing). There are 4 possible suits: ♡ (heart),
♢ (diamond), ♣ (club), ♠ (spade). Because all combinations of denomination and suit are allowed, the
multiplication principle applies: there are 13 × 4 = 52 cards in a standard deck.
We will see many applications of counting to dealing cards from a well-shuffled deck. Counting the
possibilities is affected by (i) whether the order of dealing is important, and (ii) how we distinguish the
cards: e.g. we may only be interested in their colour (so all red cards are the same) or their suit or their
denomination.

Examples

1. A hotel serves 3 choices of breakfast, 4 choices of lunch and 5 choices of dinner so a guest selects
from 3 × 4 × 5 different combinations of the three meals (assuming we opt to have all three).

2. A coffee bar has 5 different papers to choose from, 19 types of coffee and 7 different snacks.
This means there are 6 × 20 × 8 = 960 distinct selections of coffee, snack and paper. Of these 5
involve no coffee or snack (which the staff may object to) plus one has no coffee, snack or paper!

3. PINs are made up of 4 digits (0–9) with the exceptions that (i) they cannot be four repetitions
of a single digit; (ii) they cannot form increasing or decreasing consecutive sequences, e.g. 3456
and 8765 are excluded. How many possible four-digit PINs are there?
Ignoring restrictions there are 104 = 10, 000 distinct PINs. There are 10 PINs with the same
digit repeated, namely 0000, 1111, …, 9999. Increasing sequences start with 0, 1, 2, ..., 6 and

23
decreasing sequences start with 9, 8, ..., 3, so there are seven options for each. This leaves
10, 000 − 24 = 9, 976 permitted PINs.

All of the following counting principles are effectively consequences of the multiplication principle.

2.2.2 Order matters; objects are distinct

First, we look at ordered choices of distinct objects. In this case, we distinguish between selection with
replacement, where the same object can be selected multiple times, and selection without replacement,
where each object can only be selected at most once.

Counting principle: Selection with replacement for ordered choices

Suppose that we have a collection of 𝑚 distinct objects and we select 𝑟 of them with replacement.
The number of different ordered lists (ordered 𝑟-tuples) is

𝑚⏟
⏟ ×⏟
⋯×
⏟⏟𝑚 = 𝑚𝑟 .
𝑟 times

Counting principle: Selection without replacement for ordered choices

Suppose that we have a collection of 𝑚 distinct objects and we select 𝑟 ≤ 𝑚 of them without
replacement. The number of different ordered lists (ordered 𝑟-tuples) is

𝑚!
(𝑚)𝑟 ∶= 𝑚 × (𝑚 − 1) × (𝑚 − 2) × ⋯ × (𝑚 − 𝑟 + 1) =
⏟⏟⏟⏟⏟⏟⏟⏟⏟⏟⏟⏟⏟⏟⏟⏟⏟⏟⏟ .
(𝑚 − 𝑟)!
𝑟 terms

The falling factorial notation (𝑚)𝑟 (sometimes also denoted 𝑚𝑟 ) is simply a convenient way to write (𝑚−𝑟)!
𝑚!
.
In the special case where 𝑟 = 𝑚 we set 0! = 1 and then (𝑚)𝑚 = 𝑚! is the number of permutations of the
𝑚 objects. If 𝑚 is large, and 𝑟 is much smaller than 𝑚, then (𝑚)𝑟 ≈ 𝑚𝑟 .

Example

The number of ways we can deal out four cards in order from a pack of cards is (52)4 and the number
of ways we can arrange the four aces in order is 4! so the probability of finding the four aces on top
of a well-shuffled deck is
4! 4×3×2×1
= .
(52)4 52 × 51 × 50 × 49
This probability is approximately 3.7 × 10−6 or about 1 in 270, 000.

� Try it out

There are 𝑛 < 365 people in a room. Let 𝐵 be the event that (at least) two of them have the same
birthday. (We ignore leap years.)
What is ℙ(𝐵)? How big must 𝑛 be so that ℙ(𝐵) > 1/2?
Answer:

24
Here the equally likely outcomes are the ordered length-𝑛 lists of possible birthdays:

(person 1’s birthday, person 2’s birthday, … , person 𝑛’s birthday).

The number of possible outcomes is

365 × 365 × ⋯ × 365 = 365𝑛 .

This is the denominator in our probability.


For the numerator, we must work out how many outcomes are in 𝐵. In fact, it is easier to count
outcomes in 𝐵c , where everyone has a different birthday. There are

365 × 364 × ⋯ × (365 − 𝑛 + 1) = (365)𝑛

of these. So
(365)𝑛
ℙ(𝐵) = 1 − ℙ(𝐵c ) = 1 − .
365𝑛
It turns out that ℙ(𝐵) ≈ 1/2 for 𝑛 = 23.

Advanced content
Here’s one way to get this. Note that

(365)𝑛 1 2 𝑛−1
= 1 × (1 − ) × (1 − ) × ⋯ × (1 − ).
365𝑛 365 365 365
Now 1 − 𝑥 ≤ 𝑒−𝑥 , and in fact the inequality is very close to equality for 𝑥 = 1/365, being close
to zero. In any case,
(365)𝑛
≤ 𝑒−𝑥 𝑒−2𝑥 𝑒−3𝑥 ⋯ 𝑒−(𝑛−1)𝑥
365𝑛
= exp {−(1 + 2 + ⋯ + 𝑛 − 1)𝑥}
(𝑛 − 1)𝑛
= exp {− }.
2 × 365

2.2.3 Order doesn’t matter; objects are distinct

In this section, we move on to think about the scenario where the order in which objects are selected
doesn’t matter. This can arise in situations such as dealing a hand of cards, or separating a class into two
teams.

Counting principle: Selection without replacement for unordered choices

Suppose that we have a collection of 𝑚 distinct objects and we select a subset of 𝑟 ≤ 𝑚 of them
without replacement. The number of distinct subsets of size 𝑟 is

𝑚 (𝑚)𝑟 𝑚!
( ) ∶= = .
𝑟 𝑟! 𝑟! (𝑚 − 𝑟)!

25
To see this, first count the number of distinct ordered lists of 𝑟 objects—this is (𝑚)𝑟 . Each unordered
subset has been counted (𝑟)𝑟 = 𝑟! times as this is the number of distinct ways of arranging 𝑟 different
objects. Therefore the (𝑚)𝑟 ordered selections can be grouped into collections of size 𝑟!, each representing
a particular subset, and the result follows by dividing.
The expression (𝑚
𝑟 ) is the binomial coefficient for choosing 𝑟 objects from 𝑚 and is often called ‘𝑚-choose-𝑟’.
Note that
𝑚 𝑚
( )=( )
𝑟 𝑚−𝑟
as we can choose to take 𝑟 objects from 𝑚 in exactly the same number of ways that we can choose to leave
behind 𝑟 objects i.e., take 𝑚 − 𝑟 objects.

� Try it out

What is the probability of finding no aces in a four-card hand dealt from a well-shuffled deck?
Answer:
Let’s answer this by treating hands as unordered selections. Then there are

52 52 × 51 × 50 × 49
( )= = 270, 725
4 4×3×2×1

distinct unordered hands of four cards. The number of these with no aces is

48 48 × 47 × 46 × 45
( )= = 194, 580,
4 4×3×2×1

and so the probability of finding no aces in a four card hand is

48 52 48 × 47 × 46 × 45
( )/( ) = ≈ 0.7187.
4 4 52 × 51 × 50 × 49

Alternatively, we could answer this by treating the hands as ordered selections (the order corresponding
to the order of the deal, say). Of course, this will give different numerator and denominator in our
calculation, but the final answer must be the same! As ordered selections, there are

(52)4 = 52 × 51 × 50 × 49

distinct hands. The number of these with no ace is

(48)4 = 48 × 47 × 46 × 45.

Our probability is then (48)4 /(52)4 which is the same as before.

In this simple example, either method is relatively straightforward, but in many examples, it is much more
natural to treat the selections as ordered/undordered. For hands of cards, treating them as unordered
selections usually works best. For something like rolling dice, it usually makes sense to treat them as
ordered selections.

26
� Try it out

You are dealt five cards from a well-shuffled deck. Let 𝐴 be the event that exactly four cards are of
the same suit. What is ℙ(𝐴)?
Answer:
There are (52
5 ) different unordered selections for the hand, and all are equally likely. How many of
these unordered selections are in 𝐴? We need to describe a subset of 5 elements such that exactly 4
have the same suit. We build this up sequentially:

• We first choose the suit that we are going to use for the four cards: 4 possibilities.

• Then we choose the four denominations (unordered) for those cards: (13
4 ) possibilities.

• All that remains is to choose the last card, which must be of a different suit than the four
already chosen: 3 × 13 = 39 possibilities.

So the answer is
4 × (13
4 ) × 39
ℙ(𝐴) = ≈ 0.0429.
(52
5)

� Try it out

In ‘Lotto Extra’ you have to select 6 numbers from 1 to 49. You win the big prize if 6 randomly
drawn numbers match your selection. Let 𝑊 be the event that you win. Let 𝑀4 be the event that
you match exactly 4 out of 6 numbers. Find the probabilities of 𝑊 and 𝑀4 .
Answer:
We model the outcomes of the Lotto draw as unordered selections, so there are (49
6 ) = 13, 983, 816
outcomes in total. The event 𝑊 contains only one of them (your entry)! So ℙ(𝑊 ) = 1/13, 983, 816.
Now 𝑀4 uses any 4 of your numbers plus any 2 of the remaining 49 − 6 = 43 numbers. So the
number of outcomes in 𝑀4 is

6 43 6 × 5 43 × 42
( )×( )= × = 15 × 43 × 21 = 13, 545.
4 2 2×1 2×1

Then ℙ(𝑀4 ) = 13, 545/13, 983, 816 ≈ 0.001.

Advanced content
The same counting arguments can be used when we need to divide 𝑚 objects into 𝑘 > 2 groups:
arranging 𝑚 distinguishable objects into 𝑘 groups with sizes 𝑟1 , … , 𝑟𝑘 where 𝑟1 + ⋯ + 𝑟𝑘 = 𝑚 can
be done in
𝑚 𝑚!
( ) ∶=
𝑟1 , … , 𝑟 𝑘 𝑟1 ! ⋯ 𝑟 𝑘 !

ways. The expression (𝑟 , 𝑚 ) is called the multinomial coefficient (Anderson, Seppäläinen, and
1 … ,𝑟𝑘
Valkó 2018 Example 6.7).

27
2.2.4 Separating objects into groups

In the final section of this chapter, we look into how we can group objects: either by combining different
types of object into one big group, or by separating a big group into smaller ones.

Counting principle: Two types of object

Suppose that we have 𝑚 objects, 𝑟 of type 1 and 𝑚 − 𝑟 of type 2, where objects are indistinguishable
from others of their type. The number of distinct, ordered choices of the 𝑚 objects is

𝑚
( ).
𝑟

For example, suppose we have four red tokens, and three black ones. Then there are 7!/(4! 3!) = 35
different ways to lay them out in a row. The probability that they will be alternately red and black is
1/35 as there is only one such ordering.
To see why, note that each distinct order for laying out all of the token in a row is precisely the same as
choosing 4 of the 7 positions for red ones. In other words, it is an unordered choice of 4 positions from the
7 distinct positions.

� Try it out

A coin is tossed 7 times. Let 𝐸 be the event that a total of 3 heads is obtained. What is ℙ(𝐸)?
Answer:
Consider ordered sequences of H and T: then there are 27 = 128 possible sequences, e.g. HTHTHTT.
How many of them are in 𝐸? We choose the 3 places where H occurs: (73) = 35 ways to do this. The
other places are taken by Ts. So the answer is
35
ℙ(𝐸) = .
128

Advanced content

More generally, using the positions argument again, the multinomial coefficient is the number of
ordered choices of objects with 𝑘 types, 𝑟𝑖 of type 𝑖, which are indistinguishable within each type.

Counting principle: Separating into groups

The number of ways to divide 𝑚 indistinguishable objects into 𝑘 distinct groups is

𝑚+𝑘−1 𝑚+𝑘−1
( )=( ).
𝑚 𝑘−1

This counting principle lets us work out how many different ways there are to divide one group into smaller
groups. My favourite example is a packet of Skittles: if there are 16 or 17 of them in a bag, how many
different combinations of the five different flavours could we have?
We can count the number of choices with the ‘sheep-and-fences’ method. Placing all the objects in a line,
separated into their groups, there are 𝑘 − 1 “fenceposts” between the 𝑘 groups of sheep (or Skittles).

28
For example, with 6 objects in 4 groups, we could represent “three in group A, one in group B, none in
group C and two in group D” with the drawing ∗ ∗ ∗ ∣ ∗ ∣∣ ∗∗.
We draw 𝑚 + 𝑘 − 1 ‘things’ in total (stars and fences). This means that the number of groupings of the
objects is the same as the number of choices for the locations of the 𝑘 − 1 fences among the 𝑚 + 𝑘 − 1
‘things’, or (𝑚+𝑘−1 𝑚+𝑘−1
𝑘−1 ) = ( 𝑚 ).

� Textbook references

If you want more help with this section, check out:

• Section 1.4 in (Blitzstein and Hwang 2019);


• Appendix C in (Anderson, Seppäläinen, and Valkó 2018);
• or Chapter 3 in (Stirzaker 2003).

2.3 Historical context

Classical probability theory originated in calculation of odds for games of chance; as well as contributions
by Pierre de Fermat (1601–1665) and Blaise Pascal (1623–1662), comprehensive approaches were given
by Abraham de Moivre (1667–1754) (Moivre 1756), Laplace (1749–1827) (Laplace 1825), and Sim'eon
Poisson (1781–1840). A collection of these classical methods made just before the advent of the modern
axiomatic theory can be found in (Whitworth 1901).

29
3 Conditional probability and independence

� Goals

1. Know the definition of conditional probability and its properties


2. Have a solid knowledge of the partition theorem and Bayes’ theorem, recognizing situations
where one can apply them.
3. Understand the concept of independence.

3.1 Conditional probability

Definition: conditional probability

For events 𝐴, 𝐵 ⊆ Ω, the conditional probability of 𝐴 given 𝐵 is

ℙ(𝐴 ∩ 𝐵)
ℙ(𝐴 ∣ 𝐵) ∶= whenever ℙ(𝐵) > 0.
ℙ(𝐵)

In this course, when ℙ(𝐵) = 0, ℙ(𝐴 ∣ 𝐵) is undefined. The usual interpretation is that ℙ(𝐴 ∣ 𝐵) represents
our probability for 𝐴 after we have observed 𝐵. Conditional probability is therefore very important for
statistical reasoning, for example:

• In legal trials. How can we use DNA (or other) evidence to determine the chance that an accused
person is guilty?
• Medical screening. How can we make best use of the information from large scale cancer screening
programs?

Unfortunately, conditional probability is not always well understood. There are several well-known legal
cases that have involved a serious error in probabilistic reasoning: see e.g. Example 2.4.5 of (Anderson,
Seppäläinen, and Valkó 2018).
For example, if we roll a fair six-sided die, the conditional probability that the score is odd, given that the
score is at most 3, is
ℙ({1, 3}) 2/6 2
ℙ(odd ∣ at most 3) = = = .
ℙ({1, 2, 3}) 3/6 3

� Try it out

Throw three fair coins. What is the conditional probability of at least one head (event A) given at
least one tail (event B)?
Answer:
Let 𝐻 be the event ‘all heads’, 𝑇 the event ‘all tails’. Then ℙ(𝐵) = 1 − ℙ(𝐻) = 7/8 and ℙ(𝐴 ∩ 𝐵) =

30
1 − ℙ(𝐻) − ℙ(𝑇 ) = 6/8 so that

ℙ(𝐴 ∩ 𝐵) 6/8 6
ℙ(𝐴 ∣ 𝐵) = = = .
ℙ(𝐵) 7/8 7

� Try it out

Consider a family with two children, whose sex we do not know. The possible sexes are listed by the
sample space Ω = {BB, BG, GB, GG}, with the eldest first. Assume that all outcomes are equally
likely. Consider the events

𝐴1 = {GG} = {both girls},


𝐴2 = {GB, BG, GG} = {at least one girl},
𝐴3 = {GB, GG} = {first child is a girl}.

Find ℙ(𝐴1 ∣ 𝐴2 ), ℙ(𝐴2 ∣ 𝐴1 ), and ℙ(𝐴1 ∣ 𝐴3 ).


Answer:
We compute
ℙ(𝐴1 ∩ 𝐴2 ) ℙ({GG})
ℙ(𝐴1 ∣ 𝐴2 ) = =
ℙ(𝐴2 ) ℙ({GB, BG, GG})
1/4 1
= = .
3/4 3
Similarly,
ℙ(𝐴1 ∩ 𝐴2 ) ℙ({GG})
ℙ(𝐴2 ∣ 𝐴1 ) = = = 1,
ℙ(𝐴1 ) ℙ({GG})
and
ℙ(𝐴1 ∩ 𝐴3 ) ℙ({GG}) 1/4 1
ℙ(𝐴1 ∣ 𝐴3 ) = = = = .
ℙ(𝐴3 ) ℙ({GB, GG}) 2/4 2

� Try it out

Consider throwing two standard dice. Consider the events 𝐹 = first die shows 6, and 𝑇 = total is 10.
Calculate ℙ(𝐹 ) and ℙ(𝐹 ∣ 𝑇 ). Before doing any calculation, do you expect ℙ(𝐹 ∣ 𝑇 ) to be higher or
lower than ℙ(𝐹 )? (Hint: 10 is a high total. We’ll see later that the ‘average’ total score on two dice
is 7.)
Answer:
The possible outcomes are ordered pairs of the numbers 1 to 6, so |Ω| = 62 = 36. In 𝐹 are all
outcomes of the form (6, ?). There are 6 of those, so ℙ(𝐹 ) = 6/36 = 1/6.
Now 𝑇 = {(6, 4), (5, 5), (4, 6)} so 𝐹 ∩ 𝑇 = {(6, 4)}, and ℙ(𝐹 ∣ 𝑇 ) = (1/36)/(3/36) = 1/3 > 1/6.
Similarly, if the total had been 5 we would know that 𝐹 was impossible!

� Textbook references

If you want more help with this section, check out:

• Section 2.2 in (Blitzstein and Hwang 2019);


• Section 2.1 in (Anderson, Seppäläinen, and Valkó 2018);

31
• or Section 2.1 in (Stirzaker 2003).

3.2 Properties of conditional probability

In this section, we’ll meet five key properties of conditional probability.

� Key idea: properties of conditional probability

(P1) For any event 𝐵 ⊆ Ω for which ℙ(𝐵) > 0, ℙ( ⋅ ∣ 𝐵) satisfies axioms A1–A4 (i.e., is a probability
on Ω) and therefore also satisfies C1–C10.

For example, C6 for conditional probabilities says that, if ℙ(𝐶) > 0,

ℙ(𝐴 ∪ 𝐵 ∣ 𝐶) = ℙ(𝐴 ∣ 𝐶) + ℙ(𝐵 ∣ 𝐶) − ℙ(𝐴 ∩ 𝐵 ∣ 𝐶).

� Key idea: properties of conditional probability: multiplication

(P2) For any events 𝐴 and 𝐵 with ℙ(𝐴) > 0 and ℙ(𝐵) > 0,

ℙ(𝐴 ∩ 𝐵) = ℙ(𝐵) ℙ(𝐴 ∣ 𝐵) = ℙ(𝐴) ℙ(𝐵 ∣ 𝐴).

More generally, for any 𝐴, 𝐵, and 𝐶,

ℙ(𝐴 ∩ 𝐵 ∣ 𝐶) = ℙ(𝐵 ∣ 𝐶) ℙ(𝐴 ∣ 𝐵 ∩ 𝐶), if ℙ(𝐵 ∩ 𝐶) > 0. (3.1)

Some people refer to P2 as the multiplication rule for probabilities.


Both P1 and P2 can be deduced from the definition of probability. For example, Equation 3.1 follows
from the fact that
ℙ(𝐵 ∩ 𝐶) ℙ(𝐴 ∩ 𝐵 ∩ 𝐶) ℙ(𝐴 ∩ 𝐵 ∩ 𝐶)
ℙ(𝐵 ∣ 𝐶) ℙ(𝐴 ∣ 𝐵 ∩ 𝐶) = ⋅ = = ℙ(𝐴 ∩ 𝐵 ∣ 𝐶).
ℙ(𝐶) ℙ(𝐵 ∩ 𝐶) ℙ(𝐶)

� Try it out

Derek is playing Dungarees & Dragons. He rolls an octahedral die to generate the occupant of the
room he has just entered. He knows that with probability 3/8 it will be a Goblin, otherwise it will
be a Hobbit. A Goblin has a 1 in 4 chance of being equipped with a spiky club. What is the chance
that he encounters a Goblin with a spiky club?
Answer:
Let 𝐺 be the event that the occupant is a Goblin, and let 𝐶 be the event that the occupant has
a spiky club. We are told that ℙ(𝐺) = 3/8 and ℙ(𝐶 ∣ 𝐺) = 1/4, so ℙ(𝐺 ∩ 𝐶) = ℙ(𝐺)ℙ(𝐶 ∣ 𝐺) =
(3/8) × (1/4) = 3/32.

Our next property is a more general version of the multiplication rule.

32
� Key idea: properties of conditional probability: multiplication (again)

(P3): For any events 𝐴0 , 𝐴1 , … , 𝐴𝑘 with ℙ (∩𝑘−1


𝑖=0 𝐴𝑖 ) > 0,

𝑘 𝑘−2 𝑘−1
ℙ ( ⋂ 𝐴𝑖 ∣ 𝐴0 ) = ℙ(𝐴1 ∣ 𝐴0 ) × ℙ(𝐴2 ∣ 𝐴1 ∩ 𝐴0 ) × ⋯ × ℙ (𝐴𝑘−1 ∣ ⋂ 𝐴𝑖 ) × ℙ (𝐴𝑘 ∣ ⋂ 𝐴𝑖 ) .
𝑖=1 𝑖=0 𝑖=0

When 𝑘 = 2, we get P2; for 𝑘 = 3, this becomes

ℙ(𝐴 ∩ 𝐵 ∩ 𝐶) = ℙ(𝐴) ℙ(𝐵 ∣ 𝐴) ℙ(𝐶 ∣ 𝐴 ∩ 𝐵).

We can prove this by repeatedly applying Equation 3.1 (in this case, we use it twice).

� Try it out

If Derek encounters a Goblin armed with a spiky club, the Goblin will attack, causing a wound with
probability 1/2. A Goblin without a spiky club will flee. If Derek encounters a Hobbit, the Hobbit
will offer him a cup of tea. What is the probability that Derek is wounded by this encounter?
Answer: Let 𝑊 be the event that Derek is wounded. Then
3 1 1 3
ℙ𝑊 = ℙ(𝐺 ∩ 𝐶 ∩ 𝑊 ) = ℙ(𝐺)ℙ(𝐶 ∣ 𝐺)ℙ(𝑊 ∣ 𝐶 ∩ 𝐺) = ⋅ ⋅ = .
8 4 2 64

� Key idea: properties of conditional probability: partitions

(P4) If 𝐸1 , 𝐸2 , … , 𝐸𝑘 form a partition then, for any event 𝐴, we have


𝑘
ℙ(𝐴) = ∑ 𝑃 (𝐸𝑖 ) 𝑃 (𝐴 ∣ 𝐸𝑖 ) . (3.2)
𝑖=1

More generally, if ℙ(𝐵) > 0,


𝑘
ℙ(𝐴 ∣ 𝐵) = ∑ 𝑃 (𝐸𝑖 ∣ 𝐵) 𝑃 (𝐴 ∣ 𝐸𝑖 ∩ 𝐵) .
𝑖=1

This result is often called the partition theorem, or the law of total probability. (If you’ve forgotten what a
partition is, head back to Section 1.6.)
To prove P4 is true, we first use P2 on the right-hand side of Equation 3.2 to get
𝑘 𝑘
∑ 𝑃 (𝐸𝑖 ) 𝑃 (𝐴 ∣ 𝐸𝑖 ) = ∑ 𝑃 (𝐴 ∩ 𝐸𝑖 ) .
𝑖=1 𝑖=1

But since the 𝐸𝑖 form a partition, they are pairwise disjoint, and hence so are the 𝐴 ∩ 𝐸𝑖 , so by C7
𝑘
∑ 𝑃 (𝐴 ∩ 𝐸𝑖 ) = 𝑃 (∪𝑘𝑖=1 (𝐴 ∩ 𝐸𝑖 )) = 𝑃 (𝐴 ∩ (∪𝑘𝑖=1 𝐸𝑖 )) ,
𝑖=1

but since the 𝐸𝑖 form a partition, ∪𝑘𝑖=1 𝐸𝑖 = Ω, giving the result. You should check that P4 remains true
(with 𝑘 = ∞) for infinite partitions.

33
� Try it out

Back in the land of Dungarees and Dragons, suppose that the Goblin player has a special token that
he will play so that even unarmed Goblins will attack, rather than flee. An unarmed Goblin causes a
wound with probability 1/6. Now what is the chance that Derek is wounded?
Answer:
The partition we use is 𝐺c , 𝐺∩𝐶, and 𝐺∩𝐶 c . We know from our previous examples that 𝑃 (𝐺c ) = 5/8,
ℙ(𝐺 ∩ 𝐶) = 3/32, and 𝑃 (𝐺 ∩ 𝐶 c ) = ℙ(𝐺)𝑃 (𝐶 c ∣ 𝐺) = 9/32.
We also know that 𝑃 (𝑊 ∣ 𝐺c ) = 0, 𝑃 (𝑊 ∣ 𝐺 ∩ 𝐶) = 1/2, and, now, 𝑃 (𝑊 ∣ 𝐺 ∩ 𝐶 c ) = 1/6. So

ℙ𝑊 = 𝑃 (𝐺c ) 𝑃 (𝑊 ∣ 𝐺c ) + ℙ(𝐺 ∩ 𝐶)𝑃 (𝑊 ∣ 𝐺 ∩ 𝐶) + 𝑃 (𝐺 ∩ 𝐶 c ) 𝑃 (𝑊 ∣ 𝐺 ∩ 𝐶 c )


3 1 9 1 3
=0+ ⋅ + ⋅ = .
32 2 32 6 32

� Try it out

Three machines, A, B and C, produce components. 10% of components from A are faulty, 20% of
components from B are faulty and 30% of components from C are faulty. Equal numbers from each
machine are collected in a packet. One component is selected at random from the packet. What is
the probability that it is faulty?
Answer:
Let 𝐹 be the event that the component is faulty. Let 𝑀𝐴 , 𝑀𝐵 , 𝑀𝐶 be the events that the component
is from machines A, B, C respectively. Then 𝑀𝐴 , 𝑀𝐵 , 𝑀𝐶 form a partition so

𝑃 (𝐹) = 𝑃 (𝑀𝐴 ) 𝑃 (𝐹 ∣ 𝑀𝐴 ) + 𝑃 (𝑀𝐵 ) 𝑃 (𝐹 ∣ 𝑀𝐵 ) + 𝑃 (𝑀𝐶 ) 𝑃 (𝐹 ∣ 𝑀𝐶 )


1 1 1
= 0.1 × + 0.2 × + 0.3 × = 0.2.
3 3 3

The most important result in conditional probability is Bayes’ theorem. It allows us to express the
conditional probability of an event 𝐴 given 𝐵 in terms of the “inverse” conditional probability of 𝐵 given
𝐴.

� Key idea: properties of conditional probability: Bayes theorem

(P5) For any events 𝐴 and 𝐵 with ℙ(𝐴) > 0 and ℙ(𝐵) > 0,

ℙ(𝐴)ℙ(𝐵 ∣ 𝐴)
ℙ(𝐴 ∣ 𝐵) = .
ℙ(𝐵)

More generally, if ℙ(𝐴 ∣ 𝐶) > 0 and ℙ(𝐵 ∣ 𝐶) > 0, then

ℙ(𝐴 ∣ 𝐶)𝑃 (𝐵 ∣ 𝐴 ∩ 𝐶)
𝑃 (𝐴 ∣ 𝐵 ∩ 𝐶) = .
ℙ(𝐵 ∣ 𝐶)

� Try it out

Suppose that in the previous example, the component was indeed faulty. What is the probability
that it came from machine A?
Answer:

34
We have, by Bayes’ theorem (P5),

𝑃 (𝐹 ∣ 𝑀𝐴 ) 𝑃 (𝑀𝐴 ) (1/10)(1/3) 1
𝑃 (𝑀𝐴 ∣ 𝐹) = = = .
𝑃 (𝐹) 1/5 6

� Try it out

Pat ends up in the pub of an evening with probability 3/10. If she goes to the pub, she will get
drunk with probability 1/2. If she stays in, she will get drunk with probability 1/5. What is the
probability that she gets drunk? Given that she does get drunk, what is the probability that she
went to the pub?
Answer:
Let 𝑃 be the event that she goes to the pub, and 𝐷 be the event that she gets drunk. Then we are
told that 𝑃 (𝑃) = 3/10, 𝑃 (𝐷 ∣ 𝑃) = 1/2 and 𝑃 (𝐷 ∣ 𝑃 c ) = 1/5. Using the partition 𝑃 , 𝑃 c we get

3 1 7 1 29
ℙ(𝐷) = 𝑃 (𝑃) 𝑃 (𝐷 ∣ 𝑃) + 𝑃 (𝑃 c ) 𝑃 (𝐷 ∣ 𝑃 c ) = ⋅ + ⋅ = .
10 2 10 5 100
Then, by Bayes’ theorem (P5),
3 1
𝑃 (𝐷 ∣ 𝑃) 𝑃 (𝑃) 10 ⋅ 2 15
𝑃 (𝑃 ∣ 𝐷) = = 29
= .
ℙ(𝐷) 100
29

� Try it out

There are three regions (A,B,C) in a country with populations in relative proportions 5 ∶ 3 ∶ 2. In
region A, 5% of people own a rabbit. In region B, it is 10%, and in region C, it is 15%.
[Link] proportion of people nationally own rabbits? ii. What proportion of rabbit-owners come
from region A?
Answer:
Let 𝐴, 𝐵, 𝐶 be the events that a randomly-chosen individual comes from regions A, B, C respectively.
Let 𝑅 be the event that the individual is a rabbit owner. Then
5 1 3 1 2 3 17
𝑃 (𝑅) = ℙ(𝐴)𝑃 (𝑅 ∣ 𝐴) + ℙ(𝐵)𝑃 (𝑅 ∣ 𝐵) + ℙ(𝐶)𝑃 (𝑅 ∣ 𝐶) = ⋅ + ⋅ + ⋅ = .
10 20 10 10 10 20 200
And, by Bayes’ theorem,
5 1
𝑃 (𝑅 ∣ 𝐴) ℙ(𝐴) 10 ⋅ 20 5
𝑃 (𝐴 ∣ 𝑅) = = 17
= .
𝑃 (𝑅) 200
17

� Try it out

One of a set of 𝑛 people committed a crime. A suspect has been arrested, and DNA evidence is
a match. Consider the events 𝐺 = suspect is guilty, and 𝐸 = DNA evidence is a match. Suppose
that we initially believe that ℙ(𝐺) = 𝛼/𝑛. The probability of a ‘false positive’ DNA match is
𝑃 (𝐸 ∣ 𝐺c ) = 𝑝.
What is our new probability that the suspect is guilty, given the DNA evidence?
Answer:

35
We use the partition 𝐺, 𝐺c . Then, by P4,

𝑃 (𝐸) = 𝑃 (𝐸 ∣ 𝐺) ℙ(𝐺) + 𝑃 (𝐸 ∣ 𝐺c ) 𝑃 (𝐺c )


𝛼 𝛼
= 1 × + 𝑝 × (1 − )
𝑛 𝑛
𝛼 + (𝑛 − 𝛼)𝑝
= .
𝑛
Then by Bayes’ theorem (P5),

𝑃 (𝐸 ∣ 𝐺) ℙ(𝐺) 𝛼/𝑛 𝛼
𝑃 (𝐺 ∣ 𝐸) = = = .
𝑃 (𝐸) (𝛼 + (𝑛 − 𝛼)𝑝)/𝑛 𝛼 + (𝑛 − 𝛼)𝑝

Typically: 𝛼 ≈ 1 and 𝑛 is very large, and fairly easy to asses. On the other hand 𝑝 is very small, and
is difficult to assess as it requires a lot of information about the genetic make up of a (potentially
large) group of people. When 𝑛 is small it may be possible to test all of the group. A great variety of
mistakes have been made in using complex evidence of this type in courts. The famous ‘prosecutor’s
fallacy’ is pretending that 𝑃 (𝐺c ∣ 𝐸) = 𝑃 (𝐸 ∣ 𝐺c ) (of course this is wrong).

We can also combine properties P4 and P5 to make a mega-property of conditional expectation: Bayes’
theorem for partitions.

Theorem: properties of conditional probability: Bayes theorem for partitions

For any partition 𝐴1 , …, 𝐴𝑘 and any 𝐵 with ℙ(𝐵) > 0,

𝑃 (𝐴𝑖 ) 𝑃 (𝐵 ∣ 𝐴𝑖 )
𝑃 (𝐴𝑖 ∣ 𝐵) = 𝑘
.
∑𝑗=1 𝑃 (𝐴𝑗 ) 𝑃 (𝐵 ∣ 𝐴𝑗 )

More generally, if ℙ(𝐵 ∣ 𝐶) > 0,

𝑃 (𝐴𝑖 ∣ 𝐶) 𝑃 (𝐵 ∣ 𝐴𝑖 ∩ 𝐶)
𝑃 (𝐴𝑖 ∣ 𝐵 ∩ 𝐶) = 𝑘
.
∑𝑗=1 𝑃 (𝐴𝑗 ∣ 𝐶) 𝑃 (𝐵 ∣ 𝐴𝑗 ∩ 𝐶)

� Try it out

On any given day, it rains with probability 1/2. If it rains, Charlie the cat will go outside with
probability 1/10; if it is dry, the probability is 3/5. If Charlie goes outside, what is the conditional
probability that it has rained?
Answer:
Let 𝑅 = it rains, 𝐶 = Charlie goes outside. Then 𝑅, 𝑅c form a partition with 𝑃 (𝑅) = 𝑃 (𝑅c ) = 1/2.
Also, 𝑃 (𝐶 ∣ 𝑅) = 1/10 and 𝑃 (𝐶 ∣ 𝑅c ) = 3/5. So, by Bayes’ theorem (P6),

𝑃 (𝐶 ∣ 𝑅) 𝑃 (𝑅)
𝑃 (𝑅 ∣ 𝐶) =
𝑃 (𝐶 ∣ 𝑅) 𝑃 (𝑅) + 𝑃 (𝐶 ∣ 𝑅c ) 𝑃 (𝑅c )
1
⋅ 12 1
= 1 10 1 3 1
= .
10 ⋅ 2 + 5 ⋅ 2
7

36
� Textbook references

If you want more help with this section, check out:

• Sections 2.3 and 2.4 in (Blitzstein and Hwang 2019);


• Sections 2.1 and 2.2 in (Anderson, Seppäläinen, and Valkó 2018);
• or Section 2.1 in (Stirzaker 2003).

3.3 Independence of events

Tied to the idea of conditional probability is the idea of independence: the property that two events are
unrelated, or have no bearing on each other’s likelihood.

� Key idea: Independence of two events

We say that two events 𝐴 and 𝐵 are independent whenever

𝑃 (𝐴 ∩ 𝐵) = ℙ(𝐴)ℙ(𝐵).

We say that two events 𝐴 and 𝐵 are conditionally independent given a third event 𝐶 with ℙ(𝐶) > 0
whenever
𝑃 (𝐴 ∩ 𝐵 ∣ 𝐶) = ℙ(𝐴 ∣ 𝐶)ℙ(𝐵 ∣ 𝐶).

For example, if we pick a card from a well-shuffled deck, the events “the card is red” (𝑅) and “the card is
an Ace” (𝐴) are independent.
By counting, we have that 𝑃 (𝑅) = 26 52 = 2 and ℙ(𝐴)
1
= 524 1
= 13 . Now, 𝐴 ∩ 𝑅 = {𝐴♢, 𝐴♡} so
2
𝑃 (𝐴 ∩ 𝑅) = 52 1
= 26 , We check that 26
1
= 12 ⋅ 13 , so 𝑅 and
1
𝐴 are indeed independent.

� Try it out

Roll two standard dice. Let 𝐸 be the event that we have an even outcome on the first die. Let 𝐹 be
the event that we have a 4 or 5 on the second die. Are 𝐸 and 𝐹 independent?
Answer:
We will verify using a counting argument.
There are 36 equally likely outcomes, namely:

Ω = {(𝑖, 𝑗) ∶ 𝑖 ∈ {1, … , 6} and 𝑗 ∈ {1, … , 6}}

Of those, 3 × 6 are in 𝐸, and 6 × 2 are in 𝐹, so 𝑃 (𝐸) = 18/36 = 1/2 and 𝑃 (𝐹) = 12/36 = 1/3.
Moreover, 3 × 2 of these outcomes belong to both 𝐸 and 𝐹, so 𝑃 (𝐸 ∩ 𝐹) = 6/36 = 1/6. Indeed,

𝑃 (𝐸 ∩ 𝐹) = 1/6 = 1/2 × 1/3 = 𝑃 (𝐸) 𝑃 (𝐹) .

So 𝐸 and 𝐹 are independent.

37
� Try it out

Roll a fair die. Consider the events

𝐴1 = {2, 4, 6}, 𝐴2 = {3, 6}, 𝐴3 = {4, 5, 6}, and 𝐴4 = {1, 2}.

Which pairs of events are independent?


Answer:
Note that 𝐴1 ∩ 𝐴2 = {6} so 𝑃 (𝐴1 ∩ 𝐴2 ) = 16 , while 𝑃 (𝐴1 ) 𝑃 (𝐴2 ) = 36 ⋅ 26 = 16 too. So 𝐴1 and 𝐴2
are independent.
On the other hand, 𝐴1 ∩ 𝐴3 = {4, 6} so 𝑃 (𝐴1 ∩ 𝐴3 ) = 26 = 13 , but 𝑃 (𝐴1 ) 𝑃 (𝐴3 ) = 36 ⋅ 36 = 14 . So 𝐴1
and 𝐴3 are not independent.

Never confuse disjoint events with independent events! For independent events, we have that 𝑃 (𝐴 ∩ 𝐵) =
ℙ(𝐴)ℙ(𝐵), but for disjoint events, 𝑃 (𝐴 ∩ 𝐵) = 0 because 𝐴 ∩ 𝐵 = ∅.
Disjointness is a property of the sets only (it can be seen from the Venn diagram). Independence is a
property of probabilities (it cannot be seen from the Venn diagram).

� Try it out

In the context of the previous example, 𝐴3 ∩ 𝐴4 = ∅, so 𝐴3 and 𝐴4 are disjoint. They are certainly
not independent, since 𝑃 (𝐴3 ∩ 𝐴4 ) = 0 but 𝑃 (𝐴3 ) 𝑃 (𝐴4 ) = 12 ⋅ 13 = 16 ≠ 0.

The next theorem explains why independence is called independence:

Theorem: equivalent forms for independence

Consider any two events 𝐴 and 𝐵 with ℙ(𝐴) > 0 and ℙ(𝐵) > 0. The following statements are
equivalent.

(i) 𝑃 (𝐴 ∩ 𝐵) = ℙ(𝐴)ℙ(𝐵).

(ii) ℙ(𝐴 ∣ 𝐵) = ℙ(𝐴).

(iii) ℙ(𝐵 ∣ 𝐴) = ℙ(𝐵).

In other words, learning about 𝐵 will not tell us anything new about 𝐴, and similarly, learning about 𝐴
will not tell us anything new about 𝐵.
For conditional independence, we have a similar result.

Theorem: equivalent forms for conditional independence

Consider any three events 𝐴, 𝐵, and 𝐶, with 𝑃 (𝐴 ∩ 𝐵 ∩ 𝐶) > 0. The following statements are
equivalent.

(i) 𝑃 (𝐴 ∩ 𝐵 ∣ 𝐶) = ℙ(𝐴 ∣ 𝐶)ℙ(𝐵 ∣ 𝐶).

(ii) 𝑃 (𝐴 ∣ 𝐵 ∩ 𝐶) = ℙ(𝐴 ∣ 𝐶).

(iii) 𝑃 (𝐵 ∣ 𝐴 ∩ 𝐶) = ℙ(𝐵 ∣ 𝐶).

38
In other words, if we know 𝐶 then learning about 𝐵 will not tell us anything new about 𝐴, and similarly,
if we know 𝐶 then learning about 𝐴 will not tell us anything new about 𝐵.
Consider the card-shuffling example again. The probability that our card is an Ace is ℙ(𝐴) = 1/13 and
the probabilitiy that it is an Ace, given it is red, is
𝑃 (𝐴 ∩ 𝑅)
𝑃 (𝐴 ∣ 𝑅) = = ℙ(𝐴),
𝑃 (𝑅)

by independence. The ‘reason’ for the independence is that the proportion of aces in the deck (4/52) is
the same as that of aces among the red cards (2/26).

� Key idea

It is possible for two events to be conditionally independent on particular events, but not to be
(unconditionally) independent. We will see an example of this when we discuss genetics, in Section 5.2.

It can be extremely useful to recognize situations where (conditional) independence can be applied.
Of course, it is equally important not to assume (conditional) independence where there really are
dependencies.

Definition: independence for multiple events

A (possibly infinite) collection of events 𝒜 ⊆ ℱ are mutually independent if for every finite non-empty
𝒞 ⊆ 𝒜 (that is, 𝒞 is a finite subcollection of the events in question),

𝑃 ( ⋂ 𝐴) = ∏ ℙ(𝐴).
𝐴∈𝒞 𝐴∈𝒞

A collection of events 𝒜 ⊆ ℱ are mutually conditionally independent given another event 𝐵 if for
every finite non-empty subcollection 𝒞 ⊆ 𝒜,

𝑃 ( ⋂ 𝐴 ∣ 𝐵) = ∏ ℙ(𝐴 ∣ 𝐵).
𝐴∈𝒞 𝐴∈𝒞

The smallest case here is to consider three events. We say that the events 𝐴, 𝐵, and 𝐶 are mutually
independent if all of the following equalities are satisfied:

𝑃 (𝐴 ∩ 𝐵 ∩ 𝐶) = ℙ(𝐴)ℙ(𝐵)ℙ(𝐶),
𝑃 (𝐴 ∩ 𝐵) = ℙ(𝐴)ℙ(𝐵),
𝑃 (𝐵 ∩ 𝐶) = ℙ(𝐵)ℙ(𝐶),
𝑃 (𝐶 ∩ 𝐴) = ℙ(𝐶)ℙ(𝐴).

Suppose we roll 4 dice and their values are independent.


To find the probability that we throw no sixes let 𝐴𝑖 be the event ‘the 𝑖th throw is not a 6’. By assumption
𝐴1 , …, 𝐴4 are independent so
4 4
5 4
𝑃 (no sixes on 4 dice) = 𝑃 ( ⋂ 𝐴𝑖 ) = ∏ 𝑃 (𝐴𝑖 ) = ( ) .
𝑖=1 𝑖=1
6

The same result is obtained from the classical model, by selection with replacement.

39
It is possible for events to be pairwise independent without being mutually independent, as the next
example demonstrates.

Examples: Example

Toss two fair coins. The sample space is Ω = {𝐻𝐻, 𝐻𝑇 , 𝑇 𝐻, 𝑇 𝑇 } and each outcome has probability
1/4.
Let 𝐴 = {𝐻𝐻, 𝐻𝑇 } be the event that the first coin comes up ‘heads’, 𝐵 = {𝐻𝐻, 𝑇 𝐻} the event that
the second coin comes up ‘heads’, and 𝐶 = {𝐻𝐻, 𝑇 𝑇 } the event that the coins come up the same.
Then since ℙ(𝐴) = ℙ(𝐵) = ℙ(𝐶) = 1/2 and each pairwise intersection has probability 1/4, it is easy
to see that the events are pairwise independent. However, 𝑃 (𝐴 ∩ 𝐵 ∩ 𝐶) = 𝑃 (𝐻𝐻) = 1/4 which is
not the same as ℙ(𝐴)ℙ(𝐵)ℙ(𝐶) = 1/8, so the three events are not mutually independent.
To interpret this in words, if we consider any two of the events, the occurrence of one tells us nothing
about the occurrence of the other. As soon as we consider statements involving all three events,
however, we see the dependence. For example,

𝑃 (𝐶 ∣ 𝐴 ∩ 𝐵) = 1,

since 𝐴 ∩ 𝐵 = {𝐻𝐻} and {𝐻𝐻} ⊆ 𝐶, compared to the unconditional probability ℙ(𝐶) = 1/2.

� Textbook references

If you want more help with this section, check out:

• Section 2.5 in (Blitzstein and Hwang 2019);


• Section 2.3 in (Anderson, Seppäläinen, and Valkó 2018);
• or Section 2.2 in (Stirzaker 2003).

3.4 Historical context

Bayes’ theorem is named after the Reverend Thomas Bayes (1701–1761); it was published after his death,
in 1763. In our modern approach to probability, the theorem is a very simple consequence of our definitions;
however, the result may be interpreted more widely, and is one of the most important results regarding
statistical reasoning.

40
Figure 3.1: Thomas Bayes

41
4 Interpretations of probability

� Goals

1. Understand that there are different ways to interpret probability values.

This chapter covers some different ways in which we can interpret probabilities in the real world. The
axioms in Chapter 1 are helpful to determine a framework for mathematical probability, but they leave us
lots of room to choose a model within that framework.
We have already discussed one approach in Chapter 2: the “classical” approach, in which each outcome in
the sample space is assigned the same probability. This interpretation has some obvious limitations in
practice. Often we cannot find a set of outcomes that it is reasonable to think of as a priori equally likely.
Therefore, it is essential to have more widely applicable models to deal with uncertainty.
In this chapter, we discuss two more approaches to determining probabilities in real-world applications:
the relative frequencies approach and the subjective probability (or betting) approach. Either can be used,
depending on the context, to help us to assign probabilities to events.
The goal of this chapter is to help you to develop more intuition and probabilistic thinking. For the rest
of the course, we will assume that we “know” the probability of each event, without worrying too much
about how it was determined.

4.1 Relative frequency interpretation

This interpretation applies to trials giving chance outcomes of an experiment that can be repeated
indefinitely under essentially unchanged conditions and which exhibits long term regularity.
Suppose that we run 𝑛 trials of an experiment with a known list of possible outcomes and the number of
trials on which event 𝐴 occurs is 𝑛𝐴 (𝐴 is again a set of possible outcomes). The relative frequency of
occurrence of 𝐴 is 𝑛𝐴 /𝑛.
For example, if we toss a coin 1000 times and observe 490 heads, then the relative frequency of heads is
490/1000.
For some experiments, it may be reasonable to suppose that relative frequencies are stable for very large
𝑛.
If we toss a fair coin one billion times, we might expect that the relative frequency of heads after the first
few thousand throws would remain very close to 1/2.
As a mathematical idealization, we suppose that there is a unique, empirical limiting value for 𝑛𝐴 /𝑛, as 𝑛
tends to infinity, which we call the relative frequency probability of 𝐴.
For our coin, the statement 𝑃 (heads) = 1/2 means ‘if we tossed the coin an extremely large number of
times, then the proportion of heads would be arbitrarily close to 1/2’.

42
This interpretation is widely used, especially in physics, where experiments are designed for repeatability
and we can expect future trials to behave like those in the past. In this view, probability is a property of
the experimental setup and may be “objectively’ ’ discovered by sufficient repetitions of the experiment.
Amongst the problems with this interpretation are:

• it is often impossible to decide what “essentially unchanged conditions’ ’ are;


• we often have no way of knowing when limiting frequencies become stable (how many trials should
we do to test this?);
• we can only use it in situations that are repeatable.

4.2 Betting interpretation

A very different way of interpreting probability goes by considering probability as a quantification of


someone’s (yours, mine, your neighbour, …) belief that an event will occur. There are various different
ways in which we can measure this belief numerically. Here is one of the simplest.
Your subjective probability that 𝐴 will occur is measured by the amount £𝑝𝐴 that you would consider
to be a fair price for the following gamble:

• if 𝐴 occurs, you receive £1;


• if 𝐴 does not occur, you receive nothing.

In this interpretation, there are no “true’ ’ probabilities. Different individuals will have different information
relevant to a problem and so may validly make different probability assessments.
For instance, if you say your probability that ‘Your Team’ wins its next match is 1/2 this means that you
view £1/2 as a fair price for the gamble winning you £1 if Your Team wins but otherwise nothing. Others
may disagree with you.
Subjective probability ideas are often used by decision makers who have to consider problems concerning
unique, non-repeatable events, based on their informed but subjective judgements. The advantages of this
interpretation are that probability measures the belief of a subject, and is no longer seen as a property of
the experimental setup. Potential issues are that the highest ‘buying price’ may differ from lowest ‘selling
price’; a subject may have reason to misrepresent their fair price; placing the bet itself might affect the
experiment.

4.3 Interpretation and the axioms

We claimed that the axioms of probability are the same regardless of the interpretation of the probabilities
that we are using. A1 and A2 are clearly very sensible in any interpretation. The justification of A3
(and, by extension, A4) needs some more thought.
A3 feels intuitive for the classical model of probability by its relation to counting: in the classical
model if 𝐴 contains 𝑚𝐴 outcomes and 𝐵 contains 𝑚𝐵 outcomes, with none in common with 𝐴, then
𝑚𝐴∪𝐵 = 𝑚𝐴 + 𝑚𝐵 . The argument is very similar for the relative frequency model and only slightly more
subtle for the betting model.

43
� Textbook references

For more information on these ideas, check out:

• Section 1.2 in (DeGroot and Schervish 2013);


• Chapter 0 in (Stirzaker 2003);
• or (Hájek 2012).

44
5 Some applications of probability

� Goals

1. Understand the meaning of a reliability network.


2. Know how to evaluate the reliability of networks.
3. Understand the probabilistic nature of genetics.
4. Know how to derive the probability of genotypes of children given the genotypes of parents.
5. Know how and when to exploit the general rules of (conditional) probability to solve complex
reliability networks and problems in genetics.

5.1 Reliability of networks

Reliability theory concerns mathematical models of systems that are made up of individual components
which may be faulty.
If components fail randomly, a key objective of the theory is to determine the probability that the system
as a whole works. This will depend on the structure of the system (how the components are organized).
This is an important problem in industrial (or other) applications, such as electronic systems, mechanical
systems, or networks of roads, railways, telephone lines, and so on.
Once we know how to work out (or estimate) failure probabilities of these systems, we can start to
ask more sophisticated questions, such as: How should the system be designed to minimize the failure
probability, given certain practical constraints? What is a good inspection, servicing and maintenance
policy to maximize the life of the system for a minimal cost?
In this course, to demonstrate an application of the probabilistic ideas we have covered so far, we address
the basic question: Given a system made up of finitely many components, what is the probability that the
system works? Whether the system functions depends on whether the components function, and on the
configuration of those components.

Example

The figure below shows (a) two components in series, (b) three in parallel, (c) a four component
system. In each case assume that the system works if it is possible to get from the left end to the
right through functioning components.

Figure 5.1: a) Two components in series

45
Figure 5.2: b) Three components in parallel

Figure 5.3: c) A system with four components

The system in (a) works if and only if both components 1 and 2 work.
The system in (b) works if any of 1, 2, 3 work.
The system in (c) works if either both 1 and 2 work, or both 3 and 4 work (or all work).

Definition: reliability network

A reliability network is a diagram of nodes and arcs. The nodes represent components of a multi-
component system, where each node is either working or is broken, and where the entire system works
if it is possible to get from the left end to the right of the diagram through working components only.
Suppose the 𝑖th component functions with probability 𝑝𝑖 , 𝑖 ∈ {1, 2, … , 𝑘}, and different components
are independent. The probability that the system works is then a function of the probabilities 𝑝1 , …,
𝑝𝑘 . We denote this function by 𝑟(𝑝1 , 𝑝2 , … , 𝑝𝑘 ), and call it the reliability function. It is determined
by the layout of the reliability network.

Looking at reliability networks, and determining their reliability functions, is a source of lots of good
examples to practice working with the axioms of probability (in particular A3, C1, and C6), as well
as building up some more intuition about independence. Let’s have one more example (and there are a
couple on the problem sheet, too).

� Try it out

Consider the three networks in the previous example. We consider the events
𝑊𝑖 = component 𝑖 works, 𝑆 = system works.
Calculate 𝑃 (𝑆) for each of the networks.
Suppose that the system in (c) works. What is the conditional probability that component 1 works?
Answer:
The first step is to represent 𝑆 in terms of the 𝑊𝑖 and the operations of set theory. Then we can
compute 𝑃 (𝑆) using our rules for probabilities.
In (a),
𝑆 = 𝑊 1 ∩ 𝑊2 .

46
By independence, 𝑃 (𝑆) = 𝑃 (𝑊1 ) 𝑃 (𝑊2 ) = 𝑝1 𝑝2 .
For (b), we have
𝑆 = 𝑊 1 ∪ 𝑊2 ∪ 𝑊3 .
It is easiest to compute

𝑃 (𝑆 c ) = 𝑃 (𝑊1c ∩ 𝑊2c ∩ 𝑊3c ) = 𝑃 (𝑊1c ) 𝑃 (𝑊2c ) 𝑃 (𝑊3c ) = (1 − 𝑝1 )(1 − 𝑝2 )(1 − 𝑝3 ),

by independence. So
𝑃 (𝑆) = 1 − (1 − 𝑝1 )(1 − 𝑝2 )(1 − 𝑝3 ).
For (c), we have
𝑆 = (𝑊1 ∩ 𝑊2 ) ∪ (𝑊3 ∩ 𝑊4 ).
Then, by C6,
𝑃 (𝑆) = 𝑃 (𝑊1 ∩ 𝑊2 ) + 𝑃 (𝑊3 ∩ 𝑊4 ) − 𝑃 (𝑊1 ∩ 𝑊2 ∩ 𝑊3 ∩ 𝑊4 )
= 𝑝1 𝑝2 + 𝑝3 𝑝4 − 𝑝1 𝑝2 𝑝3 𝑝4 .
To find the conditional probability that component 1 works, given that system (c) works, we go back
to the definition of conditional probability:

𝑃 (𝑊1 ∩ 𝑆)
𝑃 (𝑊1 ∣ 𝑆) =
𝑃 (𝑆)
𝑃 ((𝑊1 ∩ 𝑊2 ) ∪ (𝑊1 ∩ 𝑊3 ∩ 𝑊4 ))
=
𝑃 (𝑆)
𝑝 𝑝 + 𝑝1 𝑝3 𝑝4 − 𝑝1 𝑝2 𝑝3 𝑝4
= 1 2 .
𝑝1 𝑝2 + 𝑝3 𝑝4 − 𝑝1 𝑝2 𝑝3 𝑝4

� Textbook references

If you want more help with this section, check out:

• Sections 4.1–4.4 in (Billinton and Allan 1996);


• or Chapter 9 in (Ross 2010).

5.2 Genetics

Inherited characteristics are determined by genes. The mechanism governing inheritance is random and so
the laws of probability are crucial to understanding genetics.
Your cells contain 23 pairs of chromosomes, each containing many genes (while 23 pairs is specific to
humans the idea is similar for all animals and plants). The genes take different forms called alleles and
this is one reason why people differ (there are also environmental factors). Of the 23 pairs of chromosomes,
22 pairs are homologous (each of the pair has an allele for any gene located on this pair). People with
different alleles are grouped by visible characteristics into phenotypes; often one allele, 𝐴 say, is dominant
and another, 𝑎, is recessive in which case 𝐴𝐴 and 𝐴𝑎 are of the same phenotype while 𝑎𝑎 is distinct.
Sometimes, the recessive gene is rare and the corresponding phenotype is harmful, for example haemophilia
or sickle-cell anaemia.

47
For instance, in certain types of mice, the gene for coat colour (a phenotype) has alleles 𝐵 (black) or 𝑏
(brown). 𝐵 is dominant, so 𝐵𝐵 or 𝐵𝑏 mice are black, while 𝑏𝑏 mice are brown with no difference between
𝐵𝑏 and 𝑏𝐵.
With sickle-cell anaemia, allele 𝐴 produces normal red blood cells but 𝑎 produces deformed cells. Genotype
𝑎𝑎 is fatal but 𝐴𝑎 provides protection against malarial infection (which is often fatal) and so allele 𝑎 is
common in some areas of high malaria risk.
To apply probability theory to the study of genetics, we use the basic principle of genetics: For each
gene on a homologous chromosome, a child receives one allele from each parent, where each allele received
is chosen independently and at random from each parent’s two alleles for that gene.

Examples: Example

For example, in a certain type of flowering pea, flower colour is determined by a gene with alleles 𝑅
and 𝑊, with phenotypes 𝑅𝑅 (red), 𝑅𝑊 (pink) and 𝑊 𝑊 (white). The table of offspring genotype
probabilities given parental genotypes is

Parental genotype RR RR RR RW RR WW RW RW RW WW WW WW
offspring RR 1 1/2 0 1/4 0 0
genotype RW 0 1/2 1 1/2 1/2 0
WW 0 0 0 1/4 1/2 1

How to read the table: for example, parents RR and RW produce RR offspring with chance 1/2 (the
RR parent must supply an R but the RW parent supplies either R or W, each with probability 1/2).
When we cross red and white peas, all the offspring will be pink but when we cross red and pink
peas, about half of the peas will be red, half pink. Mendel carried out experiments like these to
establish the genetic basis of inheritance.

Advanced content

Similar but larger tables are relevant when there are more than two alleles.

It is extremely important to note that genotypes of siblings are dependent unless we condition
on parental genotypes. For example, if two black mice (which may each be BB or Bb) have 100 black
offspring, you may conclude that the next offspring is overwhelmingly likely to also be black, because it is
very likely that at least one parent is BB.
Genes can also affect reproductive fitness, as we see in the next example.

� Try it out

A gene has alleles 𝐴 and 𝑎 but 𝑎 is recessive and harmful, so genotype 𝑎𝑎 does not reproduce while
𝐴𝐴, 𝐴𝑎 are indistinguishable. With proportions 1 − 𝜆, 𝜆 of 𝐴𝐴, 𝐴𝑎 in the healthy population, show
that the probability of an 𝑎𝑎 offspring is 𝜆2 /4.
Answer: To show this we can use the partition 𝐹𝐴𝐴 , 𝐹𝐴𝑎 (father 𝐴𝐴, 𝐴𝑎 respectively) to calculate
𝑃 (𝐹 𝑎) = 𝑃 (𝐹𝐴𝐴 ) 𝑃 (𝐹 𝑎 ∣ 𝐹𝐴𝐴 ) + 𝑃 (𝐹𝐴𝑎 ) 𝑃 (𝐹 𝑎 ∣ 𝐹𝐴𝑎 ) = 0 + 𝜆 × (1/2) = 𝜆/2
for the event 𝐹 𝑎 that the father provides allele 𝑎.
By symmetry, the mother also supplies allele 𝑎 with probability 𝜆/2 and by independence (random
mating) the probability that they both supply allele 𝑎 is 𝜆2 /4 e.g. when 𝜆 ≈ 1/2, about 6% of

48
offspring will be 𝑎𝑎. Over time the proportion of allele 𝑎 will decrease unless 𝐴𝑎 has a reproductive
advantage over 𝐴𝐴.

Things are slightly different for genes on the X or Y chromosomes (sex-linked genes).
These are the final chromosome pair, known as the sex chromosomes. Each may be X, a long chromosome,
or Y, a short chromosome. Most of the genes on X do not occur on Y. Most people have sex determined
as XX (female) or XY (male); YY is not possible.1

� Try it out

A gene carried on the 𝑋 chromosome has alleles 𝐴 and 𝑎 (so men have only one allele, while women
have two).

• 𝑎𝑎 women are unhealthy;


• 𝑎 men are unhealthy;
• otherwise the person is healthy.

A male child inherits his gene on the 𝑋 chromosome from his mother (as he must get his 𝑌 from
his father) with equal chance of the two alleles that the mother carries. A female child inherits her
father’s single allele as well as one of her mother’s two alleles.
Jane is healthy. Her maternal aunt has an unhealthy son (Jane’s cousin). Jane’s maternal grandparents
and her father are all healthy.

i. What is the probability that Jane is genotype 𝐴𝑎?

Now suppose that Jane has two healthy brothers.

ii. What now is the probability that Jane is genotype 𝐴𝑎?

Answer: We start with part i. From the information given, we can add some information to the
genetic tree. The healthy men are 𝐴. A male child receives his single (X-carried) allele as a random
selection from his mother’s two alleles (the genotype of his father has no bearing). Thus Jane’s Aunt
must carry an 𝑎. She cannot have inherited this from the healthy grandfather, so the grandmother
must also carry an 𝑎. This gives us the picture below.

Figure 5.4: Jane’s family tree, pt1

1
At least when restricted to the Riemann integral.

49
Consider the events 𝐽 = {Jane is 𝐴𝑎}, 𝑀1 = {mother is 𝐴𝐴}, and 𝑀2 = {mother is 𝐴𝑎}. From the
tree above, we have 𝑀1 occurs if and only if Jane’s mother inherited an 𝐴 from her mother, i.e.,
𝑃 (𝑀1 ) = 1/2 and 𝑃 (𝑀2 ) = 1/2 too. Given the mother’s genotype, we can work out the probabilities
for Jane’s inheritance. Thus, by the partition theorem,

𝑃 (𝐽) = 𝑃 (𝑀1 ) 𝑃 (𝐽 ∣ 𝑀1 ) + 𝑃 (𝑀2 ) 𝑃 (𝐽 ∣ 𝑀2 )


1 1 1 1
= ⋅0+ ⋅ = .
2 2 2 4
Now we move on to part ii. The tree is now augmented by the additional information about Jane’s
siblings:

Figure 5.5: Jane’s family tree, pt2

Let 𝐵 = {Two brothers are 𝐴}. We want 𝑃 (𝐽 ∣ 𝐵). Note that our knowledge of 𝐵 changes our
beliefs about the genotype of Jane’s mother. To see this more clearly, imagine that Jane had 100
brothers, all of whom were of type 𝐴. Then we would be very nearly sure that Jane’s mother was of
type 𝐴𝐴, and so Jane would be almost certainly of type 𝐴𝐴 too.
For the calculation, we use the partition theorem for conditional probabilities:

𝑃 (𝐽 ∣ 𝐵) = 𝑃 (𝑀1 ∣ 𝐵) 𝑃 (𝐽 ∣ 𝑀1 ∩ 𝐵) + 𝑃 (𝑀2 ∣ 𝐵) 𝑃 (𝐽 ∣ 𝑀2 ∩ 𝐵) .

But given 𝑀𝑖 , 𝐽 is independent of 𝐵 so

𝑃 (𝐽 ∣ 𝐵) = 𝑃 (𝑀1 ∣ 𝐵) 𝑃 (𝐽 ∣ 𝑀1 ) + 𝑃 (𝑀2 ∣ 𝐵) 𝑃 (𝐽 ∣ 𝑀2 ) .

As above, we have 𝑃 (𝐽 ∣ 𝑀1 ) = 0 and 𝑃 (𝐽 ∣ 𝑀2 ) = 1/2. By Bayes’s theorem,

𝑃 (𝐵 ∣ 𝑀2 ) 𝑃 (𝑀2 )
𝑃 (𝑀2 ∣ 𝐵) =
𝑃 (𝐵 ∣ 𝑀1 ) 𝑃 (𝑀1 ) + 𝑃 (𝐵 ∣ 𝑀2 ) 𝑃 (𝑀2 )
(1/2)2 ⋅ (1/2) 1
= 2 = .
1 (1/2) + (1/2)2 ⋅ (1/2) 5

So
1 1 1
𝑃 (𝐽 ∣ 𝐵) = 0 + ⋅ = .
2 5 10
So seeing that Jane has two healthy brothers significantly reduces the chance that Jane is carrying
an 𝑎.

50
� Textbook references

If you want more help with this section, check out

• Section V.5 in (Feller 1968);


• or Section 5.6 in (Chung and AitSahlia 2003).

5.3 Hardy-Weinberg equilibrium

Consider a population of a large number of individuals evolving over successive generations. Consider
a gene (on a homologous chromosome) with two alleles 𝐴 and 𝑎 and genotypes {𝐴𝐴, 𝐴𝑎, 𝑎𝑎}. Suppose
the genotype proportions in the population (uniformly for males and females) at generation 𝑛 = 0, 1, 2, …
are

𝐴𝐴 𝐴𝑎 𝑎𝑎
𝑢𝑛 2𝑣𝑛 𝑤𝑛

where we have 𝑢𝑛 + 2𝑣𝑛 + 𝑤𝑛 = 1. Suppose also that the proportions of the alleles in the population are

𝐴 𝑎
𝑝 𝑞

where 𝑝𝑛 + 𝑞𝑛 = 1. We see that

2𝑢𝑛 + 2𝑣𝑛
𝑝𝑛 = = 𝑢 𝑛 + 𝑣𝑛
2𝑢𝑛 + 4𝑣𝑛 + 2𝑤𝑛

and, similarly, 𝑞𝑛 = 𝑣𝑛 + 𝑤𝑛 .
Suppose that

• the gene is neutral, meaning that different genotypes have equal reproductive success;
• there is random mating with respect to this gene, meaning that each individual in generation 𝑛 + 1
draws randomly two parents whose genotypes are independently in the proportions 𝑢𝑛 , 2𝑣𝑛 , 𝑤𝑛 .

How do the genotype proportions evolve over successive generations?


Consider the offspring of generation 0. Let 𝐹 𝐴 = event that child gets allele 𝐴 from father, 𝑀 𝐴 = event
that child gets allele 𝐴 from mother, 𝐹𝐴𝐴 = event that father is 𝐴𝐴, 𝐹𝐴𝑎 = event that father is 𝐴𝑎, 𝐹𝑎𝑎
= event that father is 𝑎𝑎. Then
𝑃 (𝐹 𝐴) = 𝑃 (𝐹𝐴𝐴 ) 𝑃 (𝐹 𝐴 ∣ 𝐹𝐴𝐴 ) + 𝑃 (𝐹𝐴𝑎 ) 𝑃 (𝐹 𝐴 ∣ 𝐹𝐴𝑎 ) + 𝑃 (𝐹𝑎𝑎 ) 𝑃 (𝐹 𝐴 ∣ 𝐹𝑎𝑎 )
1
= 1 ⋅ 𝑢0 + ⋅ 2𝑣0 + 0 ⋅ 𝑤0 = 𝑢0 + 𝑣0 = 𝑝0 .
2
Similarly, 𝑃 (𝑀 𝐴) = 𝑝0 . In particular, since parents contribute alleles independently, the probability
distribution of the genotype of an individual in generation 1 is

51
𝐴𝐴 𝐴𝑎 𝑎𝑎
𝑝02 2𝑝0 (1 − 𝑝0 ) (1 − 𝑝0 )2

Provided that the population is large enough (see the law of large numbers in Chapter 9) these will also
be the generation 1 proportions of 𝐴𝐴, 𝐴𝑎, 𝑎𝑎, i.e.,

𝑢1 = 𝑝02 , 𝑣1 = 𝑝0 (1 − 𝑝0 ), 𝑤1 = (1 − 𝑝0 )2 .

Now let 𝑝1 = 𝑢1 + 𝑣1 be the proportion of 𝐴 in the gene pool at generation 1. Substituting the values of
𝑢1 , 𝑣1 we find that
𝑝1 = 𝑢1 + 𝑣1 = 𝑝02 + 𝑝0 (1 − 𝑝0 ) = 𝑝0 ,
i.e. the proportions of 𝐴 and 𝑎 in the gene pool are constant.
The same argument applies for later generations, so that 𝑝𝑛 = 𝑝0 for all 𝑛, i.e., the proportions of the two
alleles in the gene pool remain constant. This means that, for 𝑛 ≥ 1,

𝑢𝑛 = 𝑝02 , 𝑣𝑛 = 𝑝0 (1 − 𝑝0 ), 𝑤𝑛 = (1 − 𝑝0 )2 ,

so that the proportions of the three genotypes in the population remain constant in every generation after
the first. This is called the Hardy–Weinberg equilibrium.

5.4 Historical context

Reliability for systems of infinitely many components is related to percolation.


On the infinite square lattice ℤ2 , declare each vertex to be open, independently, with probability 𝑝 ∈ [0, 1],
else it is closed. Consider the open cluster containing the origin, that is, the set of vertices that can be
reached by nearest-neighbour steps from the origin using only open vertices. Percolation asks the question:
for which values of 𝑝 is the open cluster containing the origin infinite with positive probability? It turns
out that for this model, the answer is: for all 𝑝 > 𝑝c where 𝑝c ≈ 0.593.
The picture shows part of a percolation configuration, with open sites indicated by black dots and edges
between open sites indicated by unbroken lines.
Percolation is an important example of a probability model that displays a phase transition. You may see
more about it if you do later probability courses.

52
Figure 5.6: A lattice, with some edges missing

The laws governing the statistical nature of inheritance were first observed and formulated by monk Gregor
Johann Mendel (1822–1884).
Biologist William Bateson (1861–1926)) coined the terms “genetics” and “allele”.
The Hardy of the Hardy–Weinberg law is G.H. Hardy (1877–1947), the famous mathematical analyst, who
published it in 1908.
The statistician R.A. Fisher (1890–1962) made significant contributions to genetics, and much early work
in statistics was concerned with genetical problems. A lot of this work contributed to a legacy of eugenics,
which was used as a justification for racial discrimination.
The Wright–Fisher model formulates a random model for the evolution of genes in a population with
mutation as an urn model (Mahmoud 2009, chap. 9).
The deep influence of probability theory on genetics has continued in recent times, with significant
developments including the coalescent of J.F.C. Kingman.

53
(b) Hardy
(a) Mendel (c) Fisher

Figure 5.7: Mendel, Hardy, and Fisher.

54
6 Random variables

� Goals

1. Understand the definition of a random variable as a function on the sample space.

2. Master the notation for events and probabilities of events relating to random variables.

3. Know how to recognize a discrete random variable, how to identify its probability mass function,
and how to derive probabilities of associated events.

4. Know how to recognize a continuous real-valued random variable, how to identify its probability
density function, and how to derive probabilities of associated events.

5. Know the following distributions. This includes identifying the scenarios in which they hold,
the assumptions behind them, and how to identify their parameters.

• the binomial distribution


• the geometric distributions
• the Poisson distribution, including how and when it can be use to approximate a binomial
distribution.
• the uniform distribution.
• the exponential distribution, and how it arises from the Poisson distribution.
• the normal distribution, and how we can derive probabilities of events using its cumulative
distribution function and standard normal tables.

6. Work with functions of random variables.

In many experiments we are often interested in a numerical value rather than the elementary event 𝜔 ∈ Ω
per se: for example, in the financial industry we may not care about the behaviour of the stock price
throughout a given period, only whether it reached a certain level or not; in weather forecasting we may
not be interested in the detailed variation of atmospheric pressure and temperature, only in how much
rain is going to fall, and so on. These uncertain quantities associated with random scenarios have as their
mathematical idealization the concept of random variable.
Put simply, a random variable is a function or mapping of the sample space: for each 𝜔 ∈ Ω, the random
variable 𝑋 gives the output 𝑋(𝜔). In this chapter, we study both discrete and continuous univariate random
variables, and discuss some important examples: binomial, geometric, Poisson, uniform, exponential, and
normal distributions. To enable practical calculations, we also discuss cumulative distribution functions,
standard tables, and how probabilities behave under transformations.

55
6.1 Definition and notation

� Key idea: Definition: random variable

A random variable on Ω is a mapping from the sample space Ω to some set of possible values
𝑋(Ω) ∶= {𝑋(𝜔) ∶ 𝜔 ∈ Ω}:
𝑋 ∶ Ω → 𝑋(Ω) given by 𝜔 ↦ 𝑋(𝜔).

Typically, we find ourselves in one of the following situations:

• 𝑋(Ω) ⊆ ℝ, in which case we say that 𝑋 is a real-valued random variable or a univariate random
variable; or
• 𝑋(Ω) ⊆ ℝ𝑑 , in which case we say that 𝑋 is a vector-valued random variable, or a multivariate random
variable; in this case we can identify 𝑋 with a vector (𝑋1 , … , 𝑋𝑑 ) of real-valued random variables,
where 𝑋𝑖 (𝜔) ∶= [𝑋(𝜔)]𝑖 , that is, 𝑋𝑖 (𝜔) is the 𝑖th component of 𝑋(𝜔).

For any 𝐵 ⊆ 𝑋(Ω), we write ‘𝑋 ∈ 𝐵’ to denote the event {𝜔 ∈ Ω ∶ 𝑋(𝜔) ∈ 𝐵}. For any 𝑥 ∈ 𝑋(Ω), we
write ‘𝑋 = 𝑥’ to mean the event 𝑋 ∈ {𝑥}, that is, {𝜔 ∈ Ω ∶ 𝑋(𝜔) = 𝑥}. We sometimes also write {𝑋 = 𝑥}
and {𝑋 ∈ 𝐵} to emphasize that these are sets.
For example, if we consider throwing two standard dice, with sample space

Ω = {(𝑖, 𝑗) ∶ 𝑖 ∈ {1, 2, 3, 4, 5, 6}, 𝑗 ∈ {1, 2, 3, 4, 5, 6}},

then the sum of the numbers that show on the dice corresponds to a real-valued random variable 𝑋 defined
by:
𝑋(𝑖, 𝑗) ∶= 𝑖 + 𝑗, for all (𝑖, 𝑗) ∈ Ω.

Then the notation 𝑋 = 10 denotes the event {(4, 6), (5, 5), (6, 4)}, and 𝑋 ∈ [0, 4] denotes the event

{(1, 1), (1, 2), (2, 1), (1, 3), (2, 2), (3, 1)}.

� Try it out

Toss 3 fair coins, and let 𝑋 denote the total number of heads.

1. Describe the function 𝑋 by tabulating its values:

𝜔 HHH HHT HTH THH HTT THT TTH TTT


𝑋(𝜔) - - - - - - - -

56
2. Consider the events

𝐴1 = {𝑋 = 2}, 𝐴2 = {𝑋 ∈ [0, 1.5]}, and 𝐴3 = {𝑋 ∈ [10, 20]}.

To which subsets of Ω do these events correspond? What are their probabilities?

Answer:

1. The table of values of 𝑋 looks like this:

𝜔 HHH HHT HTH THH HTT THT TTH TTT


𝑋(𝜔) 3 2 2 2 1 1 1 0

2. We have that
𝐴1 = {𝜔 ∈ Ω ∶ 𝑋(𝜔) = 2} = {HHT, HTH, THH},
so, assuming all 8 outcomes are equally likely, 𝑃 (𝑋 = 2) = 3/8.

Similarly,
𝐴2 = {𝜔 ∈ Ω ∶ 0 ≤ 𝑋(𝜔) ≤ 1.5} = {TTT, TTH, THT, HTT},
so 𝑃 (𝑋 ∈ [0, 1.5]) = 4/8 = 1/2, and 𝐴3 = ∅ so 𝑃 (𝑋 ∈ [10, 20]) = 0.

A simple but important class of random variable is formed by the indicator random variables, denoted for
an event 𝐴 ∈ ℱ by 𝟙𝐴 and given by the mapping

1 if 𝜔 ∈ 𝐴,
𝟙𝐴 (𝜔) = {
0 otherwise.

Note that 𝑃 (𝟙𝐴 = 1) = 𝑃 ({𝜔 ∈ Ω ∶ 𝜔 ∈ 𝐴}) = 𝑃 (𝐴) and 𝑃 (𝟙𝐴 = 0) = 1 − 𝑃 (𝐴).


Recall the definition of probability from Section 1.5. Given a probability distribution 𝑃 () on the sample
space Ω, a random variable 𝑋 ∶ Ω → 𝑋(Ω) induces a probability distribution ℙ𝑋 (⋅) on the sample space
𝑋(Ω) as follows.

� Key idea: Theorem: defining probailities

The function ℙ𝑋 (⋅), mapping sets 𝐵 ⊆ 𝑋(Ω) to a real number ℙ𝑋 (𝐵), defined by

ℙ𝑋 (𝐵) ∶= 𝑃 (𝑋 ∈ 𝐵) = 𝑃 ({𝜔 ∈ Ω ∶ 𝑋(𝜔) ∈ 𝐵}) ,

is a probability on 𝑋(Ω), that is, ℙ𝑋 (⋅) satisfies the probability axioms (A1–A4).

Proof

The proof of this theorem is an exercise! It’s 6.17 on the problem sheet.

57
Advanced content

As mentioned in Chapter 1, in general sample spaces Ω one cannot expect to be able to assign
probabilities to every 𝐴 ∈ 2Ω , and one must restrict to some ℱ ⊂ 2Ω of ‘nice events’.
This has consequences for the definition of random variable we gave above, as can be gleaned by
examination of the Theorem: we must have that {𝜔 ∈ Ω ∶ 𝑋(𝜔) ∈ 𝐵} is a genuine event in ℱ, at
least for a large enough class of 𝐵 ⊆ 𝑋(Ω) to tell us everything we would like to know about the
random variable 𝑋.
For real-valued random variables, the relevant 𝐵 are the Borel sets: these are the members of the
smallest 𝜎-algebra on ℝ that contains all open intervals. Those of you who do Probability II (and
later courses) will see that the relevant concept for a more complete definition of random variable is
that of a measurable function.

� Textbook references

If you want more help with this section, check out:

• Section 3.1 in (Blitzstein and Hwang 2019);


• Section 1.5 in (Anderson, Seppäläinen, and Valkó 2018);
• or Section 4.1 in (Stirzaker 2003).

6.2 Discrete random variables

The function ℙ𝑋 (⋅) tells us everything we might need to know about the random variable 𝑋: it is called
the distribution of 𝑋. In general it is a large and unwieldy object, as we need a value of ℙ𝑋 (𝐵) for
every 𝐵 ⊆ 𝑋(Ω). However, there are two special cases where we can give an efficient description of the
distribution. The first is in the case of discrete random variables (the second, which we will see a little
later, is the case of continuous random variables).
Recall that a set is countable if there exists a bijection between that set and a subset of the natural
numbers ℕ.

� Key idea: definition: discrete random variable and probability mass function

A random variable 𝑋 ∶ Ω → 𝑋(Ω) is said to be discrete when there is a finite or countable set of
values 𝒳 ⊆ 𝑋(Ω) such that 𝑃 (𝑋 ∈ 𝒳) = 1. The function 𝑝() ∶ 𝒳 → [0, 1] defined by

𝑝(𝑥) = 𝑃 (𝑋 = 𝑥) , for all 𝑥 ∈ 𝒳,

is called the probability mass function of 𝑋.

The fact that 𝑝(𝑥) ∈ [0, 1] is an immediate consequence of A1 and C4. Here are some further important
properties of the probability mass function.

� Key idea: probability mass functions for discrete random variables

Suppose that 𝑋 is a discrete random variable and 𝑝() ∶ 𝒳 → [0, 1] is its probability mass function.

58
Then
𝑃 (𝑋 ∈ 𝐵) = ∑ 𝑝(𝑥), for all 𝐵 ⊆ 𝒳,
𝑥∈𝐵

and
∑ 𝑝(𝑥) = 1.
𝑥∈𝒳

Proof

Any 𝐴 ⊆ 𝒳 is finite or countable (because it is a subset of a finite or countable set) and so

{𝑋 ∈ 𝐵} = ⋃ {𝑋 = 𝑥},
𝑥∈𝐵

where the union runs over a countable number of events. Thus we may apply A4 to get

𝑃 (𝑋 ∈ 𝐵) = ∑ 𝑃 (𝑋 = 𝑥) = ∑ 𝑝(𝑥).
𝑥∈𝐵 𝑥∈𝐵

In particular, taking 𝐵 = 𝒳 we get ∑𝑥∈𝒳 𝑝(𝑥) = 𝑃 (𝑋 ∈ 𝒳) = 1.

The probability mass function of a discrete random variable summarizes all information we have about 𝑋.
Specifically, it allows us to calculate the probability of every event of the form {𝑋 ∈ 𝐵}.
In simple cases, the set 𝒳 is often just 𝑋(Ω), but this definition is necessary to cover all cases. If the
random variable under consideration is not clear from the context, we may write 𝑝𝑋 () for the probability
mass function of 𝑋.

Theorem: alternative characterisation of discrete random variables

A random variable 𝑋 ∶ Ω → 𝑋(Ω) is discrete whenever

(i) 𝑋(Ω) is finite or countable, or


(ii) Ω is finite or countable.

Proof

Note that (ii) implies (i).


Then, if (i) holds, the statement is immediate from the definitition of a discrete random variable,
since one can simply take 𝒳 = 𝑋(Ω).

� Try it out

Continuing our previous example, toss three fair coins and let 𝑋 be the total number of heads
obtained. This example is discrete since the possible values are 0, 1, 2, 3. Examining the table, and
grouping the 8 outcomes by the value of 𝑋, we find that the probability mass function of 𝑋 is

𝑥 0 1 2 3
1 3 3 1
𝑝(𝑥) 8 8 8 8

59
For example,
|{HHT, THH, HTH}| 3
𝑝(2) = = .
|Ω| 8
A quick way to get this is to observe that the number of ways of getting 𝑥 heads is (𝑥3) so 𝑃 (𝑋 = 𝑥) =
(𝑥3) 18 for 𝑥 ∈ 𝑋(Ω) = {0, 1, 2, 3}. We will see shortly that this is an example of the binomial
distribution.

� Textbook references

If you want more help with this section, check out:

• Section 3.2 in (Blitzstein and Hwang 2019);


• Section 3.1 in (Anderson, Seppäläinen, and Valkó 2018);
• or Section 4.2 in (Stirzaker 2003).

6.3 The binomial and geometric distributions

Consider the following random experiment, called a binomial scenario:

• A sequence of 𝑛 trials will be carried out, where 𝑛 is known in advance of the experiment.
• Trials are independent.
• Each trial has only two outcomes, usually denoted ‘success’ or ‘failure’.
• Each trial succeeds independently with the same probability 𝑝.

Consider the random variable 𝑋, the total number of successes in the 𝑛 trials.
The usual sample space Ω for the binomial scenario is the set of all the possible length-𝑛 sequences of
successes and failures; if we represent a success by 1 and a failure by 0, then each 𝜔 ∈ Ω is a string
𝜔 = 𝜔1 𝜔2 ⋯ 𝜔𝑛 with each 𝜔𝑖 ∈ {0, 1}. The random variable 𝑋 takes values in 𝑋(Ω) ∶= {0, 1, … , 𝑛}, which
is finite, so 𝑋 is a discrete random variable. As a mapping on the sample space, we have 𝑋(𝜔) = ∑𝑖=1 𝜔𝑖 ,
𝑛

the total numbers of 1s in the string.


For each 𝑥 ∈ {0, 1, … , 𝑛}:

• because trials are independent (see the definition of independence of multiple events in Section 3.3),
every sequence 𝜔 ∈ Ω with exactly 𝑥 successes and 𝑛−𝑥 failures has probability 𝑃 ({𝜔}) = 𝑝𝑥 (1−𝑝)𝑛−𝑥 ;
and
• there are (𝑛𝑥) sequences with exactly 𝑥 successes.

By (C7), we can sum the probabilities of the outcomes in the event

{𝑋 = 𝑥} = {𝜔 ∈ Ω ∶ ∑ 𝜔𝑖 = 𝑥}
𝑖

to obtain
𝑝(𝑥) = 𝑃 (𝑋 = 𝑥) = ∑ 𝑃 (𝜔) .
𝜔∈Ω∶∑𝑖 𝜔𝑖 =𝑥

60
Putting everything together, the probability mass function of 𝑋 is given by the following, which we take
as a definition.

� Key idea: Definition: binomial distribution

We say that a discrete random variable 𝑋 is binomially distributed with parameters 𝑛 ∈ ℕ and
𝑝 ∈ [0, 1], and we write 𝑋 ∼ Bin(𝑛, 𝑝), when 𝒳 = {0, 1, … , 𝑛} and

𝑛
𝑝(𝑥) = ( )𝑝𝑥 (1 − 𝑝)𝑛−𝑥 for all 𝑥 ∈ {0, 1, 2, … , 𝑛}.
𝑥

In the case of just a single trial (𝑛 = 1), the binomial scenario is often referred to as a Bernoulli trial, and
𝑋 ∼ Bin(1, 𝑝) is often referred to as a Bernoulli random variable with parameter 𝑝.

Example

If we roll 4 fair cubic dice and let 𝑋 be the number of 6s then 𝑃 (𝑋 = 𝑥) = (𝑥4) ( 16 )𝑥 ( 56 )4−𝑥 so the
probability mass function of 𝑋 is
625 500 150 20 1
𝑝(0) = , 𝑝(1) = , 𝑝(2) = , 𝑝(3) = , 𝑝(4) = .
1296 1296 1296 1296 1296
Hopefully you can see that these probabilities sum to 1.

� Try it out

105 people bought tickets for a flight. Each person independently has chance 0.04 of missing the
flight. Find
(a) the probability that nobody misses the flight;
(b) the probability that three or more people miss the flight.
Answer:
Let 𝑋 be the number of people that miss the flight. Then 𝑋 ∼ Bin(105, 0.04) so

105
𝑝(𝑥) = ( ) ⋅ 0.04𝑥 ⋅ 0.96105−𝑥 .
𝑥

a. We find that 𝑃 (𝑋 = 0) = 𝑝(0) = (105 0


0 ) ⋅ 0.04 ⋅ 0.96
105
= 0.96105 ≈ 0.014.
b. The trick here is to apply (C2), so that
𝑃 (𝑋 ≥ 3) = 1 − 𝑃 (𝑋 < 3) = 1 − 𝑝(0) − 𝑝(1) − 𝑝(2),
where
𝑝(0) ≈ 0.014 as before,
105
𝑝(1) = ( ) ⋅ 0.041 ⋅ 0.96104 ≈ 0.060,
1
105
𝑝(2) = ( ) ⋅ 0.042 ⋅ 0.96103 ≈ 0.130,
2
so 𝑃 (𝑋 ≥ 3) ≈ 0.796.

61
Suppose that we extend the binomial scenario indefinitely, to an unlimited number of trials, and we repeat
the trials until we obtain the first success. The (random) number of trials up to and including the first
success is called the geometric distribution. Note that

𝑃 (first success occurs on trial 𝑛) = 𝑃 (first 𝑛 − 1 trials are failure, then trail 𝑛 is a success)
(1 − 𝑝)𝑛−1 𝑝,

by independence of the trials and the definition of “independence of multiple events”. Note that, provided
𝑝 > 0, this is a probability distribution on {1, 2, 3, …} because, by the geometric series formula,

1
∑(1 − 𝑝)𝑛−1 𝑝 = 𝑝 ⋅ = 1.
𝑛=1
1 − (1 − 𝑝)

Thus we are led to the following definition.

� Key idea: Definition: geometric distribution

We say that a discrete random variable 𝑋 is geometrically distributed with parameter 𝑝 ∈ (0, 1], and
we write 𝑋 ∼ Geo(𝑝), when 𝒳 = ℕ ∶= {1, 2, 3 …} and

𝑝(𝑥) = (1 − 𝑝)𝑥−1 𝑝, for all 𝑥 ∈ {1, 2, 3, …}.

� Textbook references

If you want more help with this section, check out:

• Section 3.3 in (Blitzstein and Hwang 2019);


• Section 2.4 in (Anderson, Seppäläinen, and Valkó 2018);
• or Section 4.2 in (Stirzaker 2003).

6.4 The Poisson distribution

� Key idea: Definition: Poisson distribution

We say that a discrete random variable 𝑋 is Poisson distributed with parameter 𝜆, and we write
𝑋 ∼ Po(𝜆), when 𝒳 = ℤ+ ∶= {0, 1, 2, …} and

𝑒−𝜆 𝜆𝑥
𝑝(𝑥) = for 𝑥 ∈ ℤ+ .
𝑥!

The Poisson distribution is used to model counts of events which occur randomly in time at a constant
average rate 𝑟 per unit time, under some natural assumptions. Specifically,

𝑃 (event occurs in (𝑥, 𝑥 + ℎ)) ≈ 𝑟ℎ, for small ℎ.

If 𝑋 is the count of the number of events over a period of length 𝑡 then it has distribution Po(𝑟𝑡). Thus,
the interpretation of the parameter 𝜆 in Po(𝜆) is as the average number of events: we will return to this
more formally later. Typical applications are

• calls to a telephone exchange,

62
• radioactive decay events,
• jobs at a printer queue,
• accidents at busy traffic intersections,
• earthquakes at a tectonic boundary,
• fish biting at an angler’s line.

� Try it out

From a particular fleet of aircraft, there have been 32 crashes over a 25-year period. Let 𝑊 , 𝑀, and
𝑌 denote the number of crashes in the next week, month, and year, respectively. Assume that a
year has 365 days and that a month has 30 days. Suppose that crashes occur at random so that the
number of crashes in a particular period can be modelled by a Poisson distribution.

(a) How are 𝑊 , 𝑀, and 𝑌 distributed?

(b) Find 𝑃 (no crashes in the next week).

(c) Find 𝑃 (no crashes in the next month).

(d) Find 𝑃 (no crashes in the next year).

Answer:
For part (a), we compute the daily rate of crashes. 25 years is 9125 days. So the daily rate of crashes
is 𝑟 = 9125
32
≈ 0.0035. In a week the average number of crashes is 7 ⋅ 9125
32
, so 𝑊 ∼ Po( 9125
7⋅32
). Similarly,
𝑀 ∼ Po( 9125 ) and 𝑌 ∼ Po( 9125 ).
30⋅32 365⋅32

For part (b), the probability that a Po(𝜆) random variable takes value 0 is 𝑒−𝜆 . So 𝑃 (𝑊 = 0) =
7⋅32
𝑒− 9125 ≈ 0.976.
Similarly, for part (c) we get 𝑃 (𝑀 = 0) = 𝑒− 9125 ≈ 0.900, and for part (d) we get 𝑃 (𝑌 = 0) =
30⋅32

365⋅32
𝑒− 9125 ≈ 0.278.

Another situation where the Poisson distribution arises is as an approximation to the binomial distribution
when 𝑝 is small and 𝑛 is large, i.e., events are rare. More precisely, we have the following result.

� Key idea: Theorem: Poisson approximation for Binomial distributions

Consider any 𝜆 > 0. Let 𝑋𝑛 ∼ Bin(𝑛, 𝑝𝑛 ) where lim𝑛→∞ 𝑛𝑝𝑛 = 𝜆, and let 𝑌 ∼ Po(𝜆). Then for all
𝑥 ∈ ℤ+ ,
lim 𝑝𝑋𝑛 (𝑥) = 𝑝𝑌 (𝑥).
𝑛→∞
We describe this by saying that 𝑋𝑛 converges in distribution to 𝑌.

Proof

Note that since 𝑛𝑝𝑛 → 𝜆, we have 𝑝𝑛 → 0. For fixed 𝑥 we have that, for 𝑛 ≥ 𝑥,

𝑛 𝑛!
𝑃 (𝑋𝑛 = 𝑥) = ( )(𝑝𝑛 )𝑥 (1 − 𝑝𝑛 )𝑛−𝑥 = 𝑛−𝑥 (𝑛𝑝𝑛 )𝑥 (1 − 𝑝𝑛 )𝑛−𝑥 ,
𝑥 (𝑛 − 𝑥)!𝑥!

63
where we observe that, by some calculus of limits,

lim (𝑛𝑝𝑛 )𝑥 = 𝜆𝑥 ,
𝑛→∞
𝑛! 𝑛 𝑛−1 𝑛−𝑥+1
lim 𝑛−𝑥 = lim ⋅ ⋯ = 1,
𝑛→∞ (𝑛 − 𝑥)! 𝑛→∞ 𝑛 𝑛 𝑛
𝑛𝑝 𝑛
lim (1 − 𝑝𝑛 )𝑛 = lim (1 − 𝑛 ) = 𝑒−𝜆 ,
𝑛→∞ 𝑛→∞ 𝑛
lim (1 − 𝑝𝑛 ) = 1.
−𝑥
𝑛→∞

Collecting up terms gives


𝑒−𝜆 𝜆𝑥
lim 𝑃 (𝑋𝑛 = 𝑥) = ,
𝑛→∞ 𝑥!
as claimed.

This means that if 𝑋 ∼ Bin(𝑛, 𝑝), where 𝑛 is large and 𝑝 is small, then approximately 𝑋 ∼ Po(𝑛𝑝). As
a rule of thumb, with 𝑝 ≤ 0.05, we find 𝑛 = 20 gives a reasonable approximation, while 𝑛 = 100 gives a
good approximation. The approximation is useful when it allows us to sidestep calculating (𝑛𝑥) for large
values of 𝑛 and 𝑥.

Example

A typist produces a page of 1000 symbols but has probability 0.001 of mistyping any single symbol
and such errors are independent (note: neither assumption is particularly realistic). The probability
that a page contains more than two mistakes is 𝑃 (𝑋 > 2) where 𝑋 is the number of mistakes on the
page.
We can use a Binomial distribution: 𝑋 ∼ Bin(1000, 0.001). As 𝑛 is large and 𝑝 is small, we
approximately have that 𝑋 ∼ Po(1000 × 0.001) = Po(1). Therefore,

10 11 12
𝑃 (𝑋 > 2) = 1 − 𝑃 (𝑋 ≤ 2) ≈ 1 − 𝑒−1 ( + + ) ≈ 0.0803.
0! 1! 2!

� Try it out

From 1979 to 1981, 1103 Bristol postmen reported 245 dog-biting incidents (the dogs biting the
postmen, that is). In all 191 postmen were bitten, 145 of them just once. Are these numbers
consistent with dogs attacking postmen at random?
Answer:
Let 𝑋 = number of incidents suffered by a particular postman. Supposing that dogs attack at random,
then each of the 245 incidents is a random trial, where our postman has chance 1/1103 to be involved.
So 𝑋 ∼ Bin(245, 1/1103). This is suitable for a Poisson approximation with 𝜆 = 𝑛𝑝 = 1103245
≈ 0.222.
Approximately, 𝑋 ∼ Po(0.222), and then

𝑃 (𝑋 = 0) ≈ 𝑒−0.222 ≈ 0.80,
𝑃 (𝑋 = 1) ≈ 0.222𝑒−0.222 ≈ 0.18,
𝑃 (𝑋 ≥ 2) ≈ 1 − 0.80 − 0.18 = 0.02.

64
Compare this to the observed data for the proportion of postmen who were
1103 − 191
not attacked: ≈ 0.83,
1103
145
once attacked: ≈ 0.13,
1103
191 − 145
more than once attacked: ≈ 0.04.
1103
This looks like a reasonably good fit. One can test this using a goodness of fit test that you may
have seen in Statistics courses.

� Textbook references

If you want more help with this section, check out:

• Section 4.7 in (Blitzstein and Hwang 2019);


• Section 4.4 in (Anderson, Seppäläinen, and Valkó 2018);
• or Section 4.2 in (Stirzaker 2003).

6.5 Continuous random variables

� Key idea: Definition: continuous random variables

Consider a real-valued random variable 𝑋 ∶ Ω → ℝ. We say that 𝑋 is a continuous random variable,


or that 𝑋 has a continuous probability distribution, or that 𝑋 is continuously distributed, when there
is a non-negative function 𝑓 ∶ ℝ → ℝ such that
𝑏
𝑃 (𝑋 ∈ [𝑎, 𝑏]) = ∫ 𝑓(𝑡) d𝑡, (6.1)
𝑎

for all [𝑎, 𝑏] ⊆ ℝ. In this case, 𝑓 is called the probability density function of 𝑋.

If 𝑋 is continuously distributed, then taking 𝑎 = 𝑏 = 𝑥 in Equation 6.1 we see that 𝑃 (𝑋 = 𝑥) = 0 for all
𝑥 ∈ ℝ, and so for any 𝑎 < 𝑏 we have

𝑃 (𝑋 ∈ [𝑎, 𝑏]) = 𝑃 (𝑋 ∈ [𝑎, 𝑏)) = 𝑃 (𝑋 ∈ (𝑎, 𝑏]) = 𝑃 (𝑋 ∈ (𝑎, 𝑏)) .

Roughly speaking, 𝑓(𝑥) has the interpretation

𝑃 (𝑋 ∈ [𝑥, 𝑥 + d𝑥]) = 𝑓(𝑥) d𝑥, for all 𝑥 at which 𝑓 is continuous. (6.2)

All of the continuous random variables that we will see in this course have a probability density function
that is piecewise continuous.
If the random variable under consideration is not clear from the context, we may write 𝑓𝑋 (⋅) for the
probability density function of 𝑋.

65
Advanced content

A probability density function satisfying Equation 6.1 is not quite unique. In fact, we can change the
value of 𝑓(𝑥) at countably many 𝑥 and not affect Equation 6.1. For example, the two probability
density functions

1 if 0 ≤ 𝑥 ≤ 1, 1 if 0 < 𝑥 < 1,
𝑓1 (𝑥) = { , and 𝑓2 (𝑥) = { ,
0 otherwise, 0 otherwise,

both integrate to 1 and both give rise to the same distribution, i.e., the same collection of 𝑃 (𝑋 ∈ 𝐵)
over sensible sets 𝐵.

The probability density function of a random variable determines its probability distribution for all events
in a reasonable collection of subsets of ℝ (but not all subsets of ℝ). The relevant concept here is again
that of a 𝜎-algebra. We will not give the exact definition of this 𝜎-algebra here, but events of the following
type can be assigned probabilities.

� Key idea: Theorem: density functions determine probabilities

If 𝑋 is continuously distributed with probability density function 𝑓(⋅), then for any 𝐵 ⊆ ℝ that is a
finite union of intervals,
𝑃 (𝑋 ∈ 𝐵) = ∫ 𝑓(𝑥) d𝑥.
𝐵
The probability density function of a continuous random variable summarizes practically all informa-
tion we have about 𝑋. Specifically, it allows us to calculate the probability of every event of the
form {𝑋 ∈ 𝐵} where 𝐵 is a finite union of intervals.

Advanced content

The conclusion of this theorem can be extended to all 𝐵 in the Borel 𝜎-algebra, which is the smallest
𝜎-algebra on ℝ that contains all intervals. Pretty much every subset of ℝ that you will ever encounter
is Borel.

Note that we can also evaluate probabilities of unbounded intervals:


𝑏
𝑃 (−∞ < 𝑋 ≤ 𝑏) = ∫ 𝑓(𝑥) d𝑥;
−∞

𝑃 (𝑎 ≤ 𝑋 < ∞) = ∫ 𝑓(𝑥) d𝑥;
𝑎

𝑃 (−∞ < 𝑋 < ∞) = ∫ 𝑓(𝑥) d𝑥 = 1.
−∞

To see this, we can for example use (C9) (continuity along monotone limits) to get

𝑃 (𝑎 ≤ 𝑋 < ∞) = 𝑃 (∪∞
𝑛=1 {𝑎 ≤ 𝑋 ≤ 𝑎 + 𝑛})
= lim 𝑃 (𝑎 ≤ 𝑋 ≤ 𝑎 + 𝑛)
𝑛→∞
𝑎+𝑛
= lim ∫ 𝑓(𝑥) d𝑥
𝑛→∞
𝑎

=∫ 𝑓(𝑥) d𝑥,
𝑎

66
as claimed. In particular, we recover a version of (C10) for probability density functions:

Theorem: Corollary: densities integrate to one

Let 𝑋 be a continuous random variable. Then its probability density function 𝑓(⋅) integrates to one:

∫ 𝑓(𝑥) d𝑥 = 1.
−∞

� Try it out

Suppose that a continuous random variable 𝑋 has probability density given by 𝑓(𝑥) = 𝑘𝑥 for 𝑥 ∈ [0, 2],
where 𝑘 is some constant, and 𝑓(𝑥) = 0 for 𝑥 ∉ [0, 2]. Find 𝑘 and then compute 𝑃 (𝑋 ∈ [0, 1]).
Answer:
To find 𝑘, we use the fact that the density integrates to one:
∞ 2
1=∫ 𝑓(𝑥) d𝑥 = ∫ 𝑘𝑥 d𝑥 = 2𝑘,
−∞ 0

and therefore 𝑘 = 1/2. With 𝐵 = [0, 1], we have 𝑃 (𝑋 ∈ 𝐵) = ∫ 𝑥/2 d𝑥 = 1/4.


1
0

� Try it out

Suppose that a continuous random variable 𝑋 has probability density given by

⎧𝑘(1 + 𝑥) if − 1 ≤ 𝑥 < 0,
{
𝑓(𝑥) = 𝑘(2 − 𝑥)
⎨ if 0 ≤ 𝑥 ≤ 2,
{0 elsewhere,

where 𝑘 is some constant. Find 𝑘 and then compute 𝑃 (𝑋 ∈ [0, 1]).


Answer:
To find 𝑘, note that

1=∫ 𝑓(𝑥) d𝑥
−∞
0 2
= ∫ 𝑘(1 + 𝑥) d𝑥 + ∫ 𝑘(2 − 𝑥) d𝑥
−1 0
2 0 2
𝑥 𝑥2
= 𝑘 [𝑥 + ] + 𝑘 [2𝑥 − ]
2 −1 2 0
𝑘 5𝑘
= + 2𝑘 = ,
2 2
so 𝑘 = 2/5. Then
1 1
2 2 𝑥2 3
𝑃 (𝑋 ∈ [0, 1]) = ∫ (2 − 𝑥) d𝑥 = [2𝑥 − ] = .
0
5 5 2 0 5

There are random variables that are neither discrete nor continuous, but they do not arise in practice very
often. Can you think how to construct one? We will return to this briefly in Section 6.9.

67
� Textbook references

If you want more help with this section, check out:

• Section 5.1 in (Blitzstein and Hwang 2019);


• Section 3.1 in (Anderson, Seppäläinen, and Valkó 2018);
• or Section 7.1 in (Stirzaker 2003).

6.6 The uniform distribution

� Key idea: Definition: uniform distribution

Let 𝑎 and 𝑏 be real numbers with 𝑎 < 𝑏. We say a continuous random variable 𝑋 is uniformly
distributed on [𝑎, 𝑏], and we write 𝑋 ∼ U(𝑎, 𝑏), when

1/(𝑏 − 𝑎) for all 𝑥 ∈ [𝑎, 𝑏],


𝑓(𝑥) = {
0 elsewhere.

When 𝑋 ∼ U(𝑎, 𝑏) then 𝑋 can take any value in the continuous range of values from 𝑎 to 𝑏 and the
probability of finding 𝑋 in any interval [𝑥, 𝑥 + ℎ] ⊆ [𝑎, 𝑏] does not depend on 𝑥.

� Try it out

Suppose that 𝑋 ∼ U(0, 3). What is 𝑃 (𝑋 ≤ 1)?


Answer:
We compute
1 1
1 1
𝑃 (𝑋 ≤ 1) = ∫ 𝑓(𝑥) d𝑥 = ∫ d𝑥 = .
−∞ 0
3 3

� Textbook references

If you want more help with this section, check out:

• Section 5.2 in (Blitzstein and Hwang 2019);


• Section 3.1 in (Anderson, Seppäläinen, and Valkó 2018);
• or Section 7.1 in (Stirzaker 2003).

6.7 The exponential distribution

Suppose a bell chimes randomly in time at rate 𝛽 > 0. Let 𝑇 > 0 denote the time of the first chime (a
random variable). The events {no chimes in [0, 𝜏 ]} and {𝑇 > 𝜏 } are the same. We know that the number
of chimes in the interval [0, 𝜏 ] is Po(𝛽𝜏 ) and so the probability of no chimes is 𝑒−𝛽𝜏 . Hence
𝑃 (𝑇 > 𝜏) = 𝑒−𝛽𝜏 or 𝑃 (𝑇 ≤ 𝜏) = 1 − 𝑒−𝛽𝜏 for all 𝜏 ≥ 0
As 1 − 𝑒−𝛽𝜏 = ∫ 𝛽𝑒−𝛽𝑡 d𝑡 we see that 𝑇 is a continuous random variable with probability density function
𝜏
0
𝑓(𝑡) = 𝛽𝑒−𝛽𝑡 for all 𝑡 ≥ 0.

68
� Key idea: Definition: exponential distribution

Let 𝛽 > 0. We say a continuous random variable 𝑋 is exponentially distributed with parameter 𝛽,
and we write 𝑋 ∼ ℰ(𝛽), when

𝛽𝑒−𝛽𝑥 for all 𝑥 ≥ 0,


𝑓(𝑥) = {
0 elsewhere.

� Textbook references

If you want more help with this section, check out:

• Section 5.5 in (Blitzstein and Hwang 2019);


• Section 4.5 in (Anderson, Seppäläinen, and Valkó 2018);
• or Section 7.1 in (Stirzaker 2003).

6.8 The normal distribution

The normal distribution is one of the most important probability distributions for several reasons, some of
which we will see later in this course. As fate would have it, its density is also a little more complicated.

� Key idea: Definition: normal distribution

Let 𝜇, 𝜎 be real numbers with 𝜎 > 0. We say a continuous random variable 𝑋 is normally distributed
with parameters 𝜇 and 𝜎2 , and we write 𝑋 ∼ 𝒩(𝜇, 𝜎2 ), when

1 1 𝑥−𝜇 2
𝑓(𝑥) = √ 𝑒− 2 ( 𝜎 ) for all 𝑥 ∈ ℝ.
𝜎 2𝜋

The normal distribution is also sometimes called the Gaussian distribution. The probability density
function of the normal distribution is bell shaped. Roughly speaking, the parameter 𝜇 determines the
location of the bell, and 𝜎 determines the spread of the bell: for smaller 𝜎, the bell is narrower. The
relevant formal concepts are expectation and standard deviation, which will be introduced later on.
Unfortunately, there is no closed analytical form for 𝑃 (𝑋 ∈ [𝑎, 𝑏]) = ∫ 𝑓(𝑥) d𝑥 when 𝑋 is normally
𝑏
𝑎
distributed.
We can compute 𝑃 (𝑋 ∈ [𝑎, 𝑏]) by numerical integration, but we would like a quick reference to these
computations. This looks unwieldy, since there are four parameters upon which 𝑃 (𝑋 ∈ [𝑎, 𝑏]) depends: 𝑎,
𝑏, 𝜇, and 𝜎. We show by a series of steps how to reduce everything to a single parameter, so that values
can be easily tabulated. The relevant concepts that we will need to carry out this simplification are the
cumulative distribution function, and transformations of random variables. Both of these concepts are
important for many other applications too, so we spend a little time on them.

� Textbook references

If you want more help with this section, check out:

• Section 5.4 in (Blitzstein and Hwang 2019);


• Section 3.5 in (Anderson, Seppäläinen, and Valkó 2018);

69
• or Section 7.1 in (Stirzaker 2003).

6.9 Cumulative distribution functions

� Key idea: Definition: cumulative distribution function

For any real-valued random variable 𝑋, the function 𝐹 ∶ ℝ → [0, 1] defined by

𝐹 (𝑥) ∶= 𝑃 (𝑋 ≤ 𝑥) for 𝑥 ∈ ℝ

is called the cumulative distribution function of 𝑋.

If the random variable under consideration is not clear from the context, we may write 𝐹𝑋 for the
cumulative distribution function of 𝑋.
By (C1), it follows that
𝑃 (𝑋 ∈ (𝑎, 𝑏]) = 𝑃 (𝑋 ≤ 𝑏) − 𝑃 (𝑋 ≤ 𝑎) = 𝐹 (𝑏) − 𝐹 (𝑎).
Now for a continuous random variable 𝑋 this is the same as 𝑃 (𝑋 ∈ (𝑎, 𝑏)), 𝑃 (𝑋 ∈ [𝑎, 𝑏]), and so on, so
knowledge of the cumulative distribution function is sufficient to calculate the probability of virtually any
event of practical interest, that is, any finite union of intervals. This is also true for the discrete case, but
is a little more complicated as it requires a limiting procedure. The continuous case is thus somewhat
simpler, so we start with that.

� Key idea: Theorem: cumulative distribution functions and probability density functions

Suppose that 𝑋 is a continuously distributed random variable on ℝ with probability density function
𝑓(⋅). Then 𝐹 (⋅) is a continuous function and, for all 𝑥 ∈ ℝ,
𝑥
d𝐹
𝐹 (𝑥) = ∫ 𝑓(𝑡) d𝑡, 𝑓(𝑥) = (𝑥) when 𝑓(⋅) is continuous at 𝑥. (6.3)
−∞
d𝑥

Proof

The first equality follows from the definition of continuous random variables, and it implies that 𝐹
is continuous. Then the second equality is a consequence of the fundamental theorem of calculus
(correspondence between derivative and integral).

� Try it out

Recall from the final example in Section 6.5 that we have a continuous random variable 𝑋 with a
piecewise continuous density given by
⎧ 25 (1 + 𝑥) if − 1 ≤ 𝑥 ≤ 0,
{
𝑓(𝑥) = ⎨ 25 (2 − 𝑥) if 0 < 𝑥 ≤ 2,
{0 elsewhere.

Find 𝐹 (𝑥) for all 𝑥 ∈ ℝ, and use it to calculate 𝑃 (0 ≤ 𝑋 ≤ 1).

70
Answer:
Since 𝑓(⋅) is defined piecewise, we must compute 𝐹 piecewise too, always using 𝐹 (𝑥) = ∫ 𝑓(𝑡) d𝑡.
𝑥
−∞
The sensible way to do this is to work ‘left to right’, since then we can use the fact that for 𝑥 > 𝑥0 ,
𝑥0 𝑥 𝑥
𝐹 (𝑥) = ∫ 𝑓(𝑡) d𝑡 + ∫ 𝑓(𝑡) d𝑡 = 𝐹 (𝑥0 ) + ∫ 𝑓(𝑡) d𝑡,
−∞ 𝑥0 𝑥0

where we choose 𝑥0 to correspond to the piecewise definition of 𝑓(⋅). To start with, for 𝑥 ≤ −1,
𝐹 (𝑥) = ∫ 0 d𝑡 = 0. Next suppose that −1 ≤ 𝑥 ≤ 0. Then
𝑥
−∞

𝑥
2
𝐹 (𝑥) = 𝐹 (−1) + ∫ (1 + 𝑡) d𝑡
−1
5
2 𝑥
2 𝑡 1
=0+ [𝑡 + ] = (𝑥 + 1)2 .
5 2 −1 5

Now suppose that 0 ≤ 𝑥 ≤ 2. Then


𝑥
2
𝐹 (𝑥) = 𝐹 (0) + ∫ (2 − 𝑡) d𝑡
0
5
𝑥
1 2 𝑡2
= + [2𝑡 − ]
5 5 2 0
(𝑥 − 2)2
=1− .
5
Finally, if 𝑥 ≥ 2,
𝑥
𝐹 (𝑥) = 𝐹 (2) + ∫ 0 d𝑡 = 𝐹 (2) = 1.
2
So we conclude that
⎧0 if 𝑥 ≤ −1,
{ (𝑥+1)2
{
5 if − 1 ≤ 𝑥 ≤ 0,
𝐹 (𝑥) =
⎨1 − (𝑥−2)2 if 0 ≤ 𝑥 ≤ 2,
{ 5
{1 if 𝑥 ≥ 2.

Note that, as expected, 𝐹 is continuous everywhere and differentiable at all but finitely many points.
Now, since 𝑋 is continuous,

𝑃 (0 ≤ 𝑋 ≤ 1) = 𝑃 (0 < 𝑋 ≤ 1)
= 𝑃 (𝑋 ≤ 1) − 𝑃 (𝑋 ≤ 0)
= 𝐹 (1) − 𝐹 (0)
4 1 3
= − = .
5 5 5

� Try it out

Recall from the first example in Section Section 6.5 that the continuous random variable 𝑋 has
𝑓(𝑥) = 𝑥/2 for 𝑥 ∈ [0, 2] and 𝑓(𝑥) = 0 elsewhere. By the same method as the previous example, the

71
cumulative distribution function is

⎧0 if 𝑥 < 0,
𝑥 {
𝐹 (𝑥) = ∫ 𝑓(𝑡) d𝑡 = 𝑥 /4 if 𝑥 ∈ [0, 2],
2

−∞ {1 if 𝑥 > 2.

Now let us move on to the case of a discrete random variable. Let us start with an example.

Example

Suppose 𝑋 is a discrete random variable taking values in {0, 1, 4} with 𝑝(0) = 12 , 𝑝(1) = 𝑝(4) = 14 .
Now 𝐹 (𝑥) = 𝑃 (𝑋 ≤ 𝑥) so that 𝐹 (𝑥) = 0 for 𝑥 < 0. At 𝑥 = 0, we have 𝐹 (0) = 𝑃 (𝑋 ≤ 0) = 𝑝(0) = 12 ,
so the cumulative distribution function jumps and the magnitude of the jump is the value of the
probability mass function at that point. This is the general picture. Then 𝐹 (𝑥) = 12 for all 𝑥 ∈ [0, 1),
until 𝐹 (1) = 𝑝(0) + 𝑝(1) = 34 . Continuing, we get

⎧0 if 𝑥 < 0,
{1
{ if 0 ≤ 𝑥 < 1,
𝐹 (𝑥) = ⎨ 23
{4 if 1 ≤ 𝑥 < 4,
{1 if 𝑥 ≥ 4.

� Key idea: Theorem: properties of discrete cdfs

If 𝑋 is a discrete real-valued random variable with probability mass function 𝑝(), then 𝐹 is piecewise
constant, and
𝐹 (𝑥) = ∑ 𝑝(𝑡) 𝑝(𝑥) = 𝐹 (𝑥) − 𝐹 (𝑥− ).
𝑡 ∶ 𝑡≤𝑥

Here 𝐹 (𝑥 ) means the limit from the left 𝐹 (𝑥− ) = lim𝑦↑𝑥 𝐹 (𝑦).

Proof

The first equality follows from the definition of discrete random variables.. Now

𝑝(𝑥) = 𝑃 (𝑋 = 𝑥) = 𝑃 ({𝑋 ≤ 𝑥}\{𝑋 < 𝑥}) = 𝐹 (𝑥) − 𝑃 (𝑋 < 𝑥) .

Let 𝑥𝑛 be an increasing sequence with 𝑥𝑛 < 𝑥 and 𝑥𝑛 → 𝑥. Then to prove the theorem it is enough
to show that
𝑃 (𝑋 < 𝑥) = 𝐹 (𝑥− ) = lim 𝐹 (𝑥𝑛 ).
𝑛→∞

Note that since 𝑥𝑛+1 > 𝑥𝑛 , 𝐹 (𝑥𝑛 ) ≤ 𝐹 (𝑥𝑛+1 ) ≤ 1, so 𝐹 (𝑥𝑛 ) is bounded and increasing, so the limit
does exist. Moreover, 𝑋 < 𝑥 if and only if 𝑋 ≤ 𝑥𝑛 for some 𝑛 in ℕ, i.e.,

𝑛=1 {𝑋 ≤ 𝑥𝑛 }) = lim 𝑃 (𝑋 ≤ 𝑥𝑛 ) ,
𝑃 (𝑋 < 𝑥) = 𝑃 (∪∞
𝑛→∞

by C9 (continuity along monotone limits), because the events 𝐴𝑛 = {𝑋 ≤ 𝑥𝑛 } are increasing in 𝑛.


But 𝑃 (𝑋 ≤ 𝑥𝑛 ) = 𝐹 (𝑥𝑛 ) and we are done.

We state a result on the properties of the cumulative distribution function in general, that are already

72
apparent from our examples in the special cases of discrete and continuous random variables.

Theorem: properties of continuous cdfs

Let 𝐹 be the cumulative distribution function of a real-valued random variable. Then 𝐹 has the
following properties.
F1: lim𝑡→−∞ 𝐹 (𝑡) = 0 and lim𝑡→+∞ 𝐹 (𝑡) = 1.
F2: Monotonicity. For any 𝑠 ≤ 𝑡, 𝐹 (𝑠) ≤ 𝐹 (𝑡).
F3: Right-continuity. For any 𝑡 ∈ ℝ, 𝐹 (𝑡) = 𝐹 (𝑡+ ) where 𝐹 (𝑡+ ) is the limit from the right
𝐹 (𝑡+ ) = lim𝑠↓𝑡 𝐹 (𝑠).

We do not give the proof here, but you can try to prove this or consult the recommended text books.
We have already seen, in the discrete and continuous cases, that we can recover the probability mass
function and probability density function, respectively, from the cumulative distribution function. In other
words, the cumulative distribution function determines the distribution. This is true in general.

Theorem: cdfs determine distributions

The cumulative distribution function 𝐹 of a real-valued random variable 𝑋 completely determines


the distribution of 𝑋.

Advanced content

In general this means that 𝐹 determines 𝑃 (𝑋 ∈ 𝐵) for all Borel sets 𝐵.

Now we can see that there are some cumulative distribution functions that correspond to random variables
that are neither discrete nor continuous. For example, we might have some jumps but also some continuously
increasing parts. There are even examples where the cumulative distribution function is continuous but no
probability density function exists: these singular distributions are quite pathological and rarely occur in
practice.

� Textbook references

If you want more help with this section, check out:

• Section 3.6 in (Blitzstein and Hwang 2019);


• Section 3.2 in (Anderson, Seppäläinen, and Valkó 2018);
• or Section 7.2 in (Stirzaker 2003).

6.10 Standard normal tables

Definition: standard normal distribution

A continuous random variable 𝑍 is standard normally distributed when 𝑍 ∼ 𝒩(0, 1).


In other words, a standard normal is normal with 𝜇 = 0 and 𝜎 = 1.

Because the standard normal distribution plays such a central role in many practical probability calculations,

73
we use a special symbol to denote its probability density function and cumulative distribution function:1
𝑧
1
𝜙(𝑧) ∶= 𝑓(𝑧) = √ 𝑒−𝑧 /2 , Φ(𝑧) ∶= 𝐹 (𝑧) = ∫ 𝜙(𝑡) d𝑡.
2

2𝜋 −∞

As already mentioned, Φ has no closed analytical form, however, it can be tabulated. Such tabulation is
called a standard normal table. Because 𝜙(𝑧) = 𝜙(−𝑧), it follows that Φ(𝑧) = 1 − Φ(−𝑧), and so we only
need to tabulate Φ for non-negative values of 𝑧. Some values that are often useful are:

𝑧 0 1.28 1.64 1.96 2.58


Φ(𝑧) 0.5 0.9 0.95 0.975 0.995

� Try it out

Suppose 𝑍 ∼ 𝒩(0, 1). Calculate 𝑃 (−1.28 ≤ 𝑍 ≤ 1.64).


Answer:
We compute, making good use of symmetry,

𝑃 (−1.28 ≤ 𝑍 ≤ 1.64) = 𝑃 (𝑍 ≤ 1.64) − 𝑃 (𝑍 ≤ −1.28)


= 𝑃 (𝑍 ≤ 1.64) − 𝑃 (𝑍 ≥ 1.28)
= 𝑃 (𝑍 ≤ 1.64) − (1 − 𝑃 (𝑍 ≤ 1.28))
= Φ(1.64) − (1 − Φ(1.28))
= 0.95 − (1 − 0.9) = 0.85.

We can use normal tables for Φ to also calculate 𝑃 (𝑋 ∈ [𝑎, 𝑏]) for any 𝑋 ∼ 𝒩(𝜇, 𝜎2 ). To explain this, we
need to look at transformations of random variables, introduced next.

� Textbook references

If you want more help with this section, check out:

• Section 5.4 in (Blitzstein and Hwang 2019);


• or Section 3.5 in (Anderson, Seppäläinen, and Valkó 2018).

6.11 Functions of random variables

Suppose 𝑋 ∶ Ω → 𝑋(Ω) is a random variable, and 𝑔 ∶ 𝑋(Ω) → 𝒮 is some function. Then 𝑔(𝑋) is also a
random variable, namely the outcome to a ‘new experiment’ obtained by running the ‘old experiment’ to
produce a value 𝑥 for 𝑋, and then evaluating 𝑔(𝑥). Formally, as a function of 𝜔 ∈ Ω, 𝑔(𝑋) ∶= 𝑔 ∘ 𝑋, i.e.,

𝑔(𝑋)(𝜔) ∶= 𝑔(𝑋(𝜔)) for all 𝜔 ∈ Ω.

For example,
𝑃 (𝑔(𝑋) ∈ 𝐵) = 𝑃 ({𝜔 ∈ Ω ∶ 𝑔(𝑋(𝜔)) ∈ 𝐵}) for all 𝐵 ⊆ 𝒮.

1
At least when restricted to the Riemann integral.

74
Examples

1. For any random variable 𝑋, we can consider sin(𝑋), 𝑒3𝑋 , 𝑋 3 , and so on, which are all again
random variables.

2. Let 𝑋 be the score when you roll a fair die and let 𝑌 = (𝑋 − 3)2 . Then 𝑌 is a discrete random
variable with probability mass function

𝑦 0 1 4 9
𝑝(𝑦) 1/6 1/3 1/3 1/6

and zero elsewhere. To see this, note that {𝑌 = 4} = {𝑋 ∈ {1, 5}}, and so on.

3. If 𝑋 is Bin(𝑛, 𝑝), then 𝑛 − 𝑋 is Bin(𝑛, 1 − 𝑝), as shown in Exercise 6.6.

4. Let 𝑋 ∼ U(0, 1). For any constants 𝑎 and 𝑏 > 0, define 𝑌 ∶= 𝑎 + 𝑏𝑋. Then 𝑌 ∼ U(𝑎, 𝑎 + 𝑏),
because for any 𝑥 ∈ [0, 1],

{𝑌 ≤ 𝑎 + 𝑏𝑥} = {𝑎 + 𝑏𝑋 ≤ 𝑎 + 𝑏𝑥} = {𝑋 ≤ 𝑥}

so that 𝑃 (𝑌 ≤ 𝑎 + 𝑏𝑥) = 𝑃 (𝑋 ≤ 𝑥) = 𝑥, and consequently 𝐹 (𝑦) = 𝑃 (𝑌 ≤ 𝑦) = 𝑦−𝑎 𝑏 whenever


𝑦 ∈ [𝑎, 𝑎 + 𝑏]. Therefore, by Equation 6.3, indeed, 𝑓(𝑦) = 1/𝑏 for 𝑦 ∈ [𝑎, 𝑎 + 𝑏] (and zero
elsewhere), so 𝑌 is uniformly distributed on [𝑎, 𝑎+𝑏] by the definition of the Uniform distribution.

A similar result holds when 𝑏 < 0.

Example

If 𝑈 ∼ U(0, 1) and 𝑋 is a continuous random variable with a cumulative distribution function 𝐹𝑋


which is strictly increasing on [𝑎, 𝑏], with 𝐹 + 𝑋(𝑎) = 0 and 𝐹𝑋 (𝑏) = 1, then 𝐹𝑋
−1
exists and 𝐹𝑋
−1
(𝑈 )
has the same distribution as 𝑋 because for any 𝑥 ∈ [𝑎, 𝑏],
−1
𝑃 (𝐹𝑋 (𝑈 ) ≤ 𝑥) = 𝑃 (𝑈 ≤ 𝐹𝑋 (𝑥)) = 𝐹𝑋 (𝑥) = 𝑃 (𝑋 ≤ 𝑥) .

This is a special case of the probability integral transform and is very useful for generating random
samples of 𝑋 with computer generated ‘uniform random numbers’. For example, − 𝛽1 log(1 − 𝑈 ) is
ℰ(𝛽) and so is 𝛽1 log 1/𝑈 (as 𝑈 and 1 − 𝑈 are both U(0, 1)).

There is one particularly important function which enables us to get the cumulative distribution function of
any normally distributed random variable, using just the the standard normal tables (i.e. Φ, the cumulative
distribution function of the standard normal).

� Key idea: Theorem: standardizing the normal distribution

Suppose 𝜇 ∈ ℝ and 𝜎 > 0. If 𝑋 ∼ 𝒩(𝜇, 𝜎2 ) and 𝑍 ∼ 𝒩(0, 1), then

𝑋−𝜇
∼ 𝒩(0, 1), and 𝜎𝑍 + 𝜇 ∼ 𝒩(𝜇, 𝜎2 ).
𝜎

75
Proof

We can prove this via a change of variable in the integral for the cumulative distribution function:
see Exercise 6.15. A shorter proof goes via the moment generating function, which will be introduced
later.

Corollary

If 𝑋 ∼ 𝒩(𝜇, 𝜎2 ) then
𝑥−𝜇
𝐹 (𝑥) = Φ ( ).
𝜎

� Try it out

Suppose 𝑋 ∼ 𝒩(2, 4). Find 𝑃 (𝑋 ≥ 5.28).


Answer:
We compute
𝑋−2 5.28 − 2
𝑃 (𝑋 ≥ 5.28) = 𝑃 ( ≥ )
2 2
= 𝑃 (𝑍 ≥ 1.64) ,
where 𝑍 ∼ 𝒩(0, 1). Hence

𝑃 (𝑋 ≥ 5.28) = 1 − Φ(1.64) = 1 − 0.95 = 0.05.

� Textbook references
For more help with this section, check out:

• Section 3.7 in (Blitzstein and Hwang 2019);


• Section 5.2 in (Anderson, Seppäläinen, and Valkó 2018);
• or Section 7.2 in (Stirzaker 2003).

6.12 Historical context

The normal distribution appeared already in work of de Moivre, and is sometimes known as the Gaussian
distribution after Carl Friedrich Gauss (1777–1855). The name ‘normal distribution’ was applied by
eugenicist and biometrician Sir Francis Galton (1822–1911) and statistician Karl Pearson (1857–1936) to
mark the distribution’s ubiquity in biometric data.
There is a great deal of subtle and interesting mathematics on the subject of what functions are integrable
over what sets. You may see some of this in the third year probability course. The Riemann integral that
we use here is sufficient for integrating piecewise continuous functions over finite unions of intervals. Here,
we will only consider continuous random variables which have a piecewise continuous probability density
function. Other approaches to integration are required to deal with more general functions.
For instance, for infinite countable unions of intervals, we would need the Lebesgue integral (see for
instance (Rosenthal 2007)). More precisely, 𝑓(⋅) still determines the value of 𝑃 (𝑋 ∈ 𝐵) when 𝐵 is an
infinite countable union of intervals, but that value is not necessarily given by the Riemann integral.
The treatment of discrete and continuous random variables separately is a little irksome. A general

76
(b) Gauss

(c) Galton
(a) de Moivre
(d) Pearson

Figure 6.1: (left to right) de Moivre, Gauss, Galton, and Pearson.

treatment of random variables, which covers both cases, as well as cases that are neither discrete nor
continuous, in a unified setting, requires the mathematical framework of measure theory; you will see some
of this if you take later probability courses.

77
7 Multiple random variables

� Goals

1. Understand a multivariate random variable as a function from the sample space to a higher
dimensional space.

2. Understand jointly distributed discrete random variables:

• their joint probability mass function, marginal probability mass functions, and conditional
probability mass functions.
• the partition theorem for discrete random variables.
• independence, and how to apply it.
• the properties of and links between the different probability mass functions, and how those
properties and links arise from the axioms of probability distributions.

3. Understand continuously distributed random variables:

• their joint probability density function, marginal probability density functions, and conditional
probability density functions.
• the partition theorem for jointly continuously distributed random variables.
• independence, and how to apply it.
• the properties of and links among the different probability density functions.

4. Understand how to work with functions of multiple random variables.

7.1 Joint probability distributions

It is essential for most useful applications of probability to have a theory which can handle many random
variables simultaneously. To start, we consider having two random variables. The theory for more than
two random variables is an obvious extension of the bivariate case covered below.
Remember that, formally, random variables are simply mappings from Ω into some set. A bivariate random
variable is a mapping from Ω into a Cartesian product of two sets, i.e., a random variable whose values are
ordered pairs of the form (𝑥, 𝑦). Of course, a bivariate random variable is a random variable according to
our original definition, just with a special kind of set of possible values. However, the concept of bivariate
random variable is a useful one if the individual components of the bivariate variable have their own
meaning or interest.

Definition: bivariate random variable

Consider random variables 𝑋 and 𝑌 defined on the same sample space Ω, 𝑋 ∶ Ω → 𝑋(Ω) and

78
𝑌 ∶ Ω → 𝑌 (Ω). The mapping (𝑋, 𝑌 ) ∶ Ω → (𝑋, 𝑌 )(Ω) defined by

(𝑋, 𝑌 )(𝜔) ∶= (𝑋(𝜔), 𝑌 (𝜔))

is then a bivariate random variable.

Here is a picture:
Note that the set of possible values (𝑋, 𝑌 )(Ω) = {(𝑋(𝜔), 𝑌 (𝜔)) ∶ 𝜔 ∈ Ω} is a subset of the Cartesian
product 𝑋(Ω) × 𝑌 (Ω).

Example

On sample space Ω = {1, 2, 3, 4, 5, 6}, define random variables 𝑋 and 𝑌 by

𝜔 1 2 3 4 5 6
𝑋(𝜔) 0 0 0 1 1 1
𝑌 (𝜔) 0 1 0 2 0 3

Then the bivariate random variable (𝑋, 𝑌 ) is given by

𝜔 1 2 3 4 5 6
(𝑋, 𝑌 )(𝜔) (0,0) (0,1) (0,0) (1,2) (1,0) (1,3)

Note that 𝑋(Ω) = {0, 1}, 𝑌 (Ω) = {0, 1, 2, 3}, and (𝑋, 𝑌 )(Ω) = {(0, 0), (0, 1), (1, 0), (1, 2), (1, 3)}
which is a strict subset of {0, 1} × {0, 1, 2, 3} (the outcome (0, 3) does not appear, for example).

Similarly to before, for any 𝐴 ⊆ 𝑋(Ω) × 𝑌 (Ω), we write ‘(𝑋, 𝑌 ) ∈ 𝐴’ to denote the event

{𝜔 ∈ Ω ∶ (𝑋(𝜔), 𝑌 (𝜔)) ∈ 𝐴}.

For any 𝑥 ∈ 𝑋(Ω) and 𝑦 ∈ 𝑌 (Ω), we write ‘𝑋 = 𝑥, 𝑌 = 𝑦’ to mean the event (𝑋, 𝑌 ) ∈ {(𝑥, 𝑦)}. We

79
sometimes also write {(𝑋, 𝑌 ) = (𝑥, 𝑦)} and {(𝑋, 𝑌 ) ∈ 𝐴} to emphasize that these are sets:

{𝑋 = 𝑥, 𝑌 = 𝑦} ∶= {(𝑋, 𝑌 ) ∈ {(𝑥, 𝑦)}}


= {𝑋 = 𝑥} ∩ {𝑌 = 𝑦}
= {𝜔 ∈ Ω ∶ 𝑋(𝜔) = 𝑥 and 𝑌 (𝜔) = 𝑦}.

We may also write more complex expressions like:

{0 ≤ 𝑋 ≤ 𝑌 2 ≤ 1} = {𝜔 ∈ Ω ∶ 0 ≤ 𝑋(𝜔) ≤ 𝑌 (𝜔)2 ≤ 1}.

� Key idea: Definition: Independence of two random variables

Two random variables 𝑋 and 𝑌 on the same sample space Ω are independent if

𝑃 (𝑋 ∈ 𝐴, 𝑌 ∈ 𝐵) = 𝑃 (𝑋 ∈ 𝐴) 𝑃 (𝑌 ∈ 𝐵) for all 𝐴 ⊆ 𝑋(Ω) and 𝐵 ⊆ 𝑌 (Ω).

In other words, 𝑋 and 𝑌 are independent (as random variables) if and only if {𝑋 ∈ 𝐴} and {𝑌 ∈ 𝐵}
are independent as events for all sets 𝐴 and 𝐵.

Three random variables, 𝑋, 𝑌, and 𝑍, say, are independent if the events {𝑋 ∈ 𝐴}, {𝑌 ∈ 𝐵}, {𝑍 ∈ 𝐶}
are mutually independent. Similarly for any finite collection of random variables. This definition is a bit
unwieldy: it reduces to simpler statements in the cases of pairs of discrete or continuous random variables,
which we look at next.

Advanced content

Later (when we talk about limit theorems such as the law of large numbers) we will need to talk
about infinite sequences of independent random variables. A (possibly infinite) collection of random
variables 𝑋𝑖 , 𝑖 ∈ ℐ, is independent if every finite nonempty sub-collection 𝒥 ⊆ ℐ is independent. In
other words, 𝑋𝑖 , 𝑖 ∈ ℐ, are independent if for every finite 𝒥 ⊆ ℐ,

𝑃 ( ⋂ {𝑋𝑗 ∈ 𝐴𝑗 }) = ∏ 𝑃 (𝑋𝑗 ∈ 𝐴𝑗 ) .
𝑗∈𝒥 𝑗∈𝒥

7.2 Jointly distributed discrete random variables

� Key idea: Key definition: joint probability mass function

Let (𝑋, 𝑌 ) be a bivariate discrete random variable with 𝑃 ((𝑋, 𝑌 ) ∈ 𝒵) = 1 for a finite or countable
𝒵 ⊆ (𝑋, 𝑌 )(Ω). The joint probability mass function 𝑝(⋅) of 𝑋 and 𝑌 is defined by

𝑝(𝑥, 𝑦) ∶= 𝑃 (𝑋 = 𝑥, 𝑌 = 𝑦) for all (𝑥, 𝑦) ∈ 𝒵.

Note there is nothing really new here, other than the terminology: this is just the definition of a discrete
random variable written for the special case of a random variable (𝑋, 𝑌 ) whose values are ordered pairs of
the form (𝑥, 𝑦).
To avoid ambiguity, we sometimes write if 𝑝𝑋,𝑌 (⋅) if it is not clear from the context which random variables
the two arguments refer to; note that 𝑝𝑋,𝑌 (⋅) is not the same as 𝑝𝑌 ,𝑋 (⋅).

80
It is the case that (𝑋, 𝑌 ) is discrete if and only if the individual random variables 𝑋 and 𝑌 are discrete.
Proving this is the purpose of the following theorem, which and also explains how the marginal probability
mass functions 𝑝(𝑥) and 𝑝(𝑦) of the individual random variables 𝑋 and 𝑌 are connected to the joint
probability mass function 𝑝(𝑥, 𝑦). Note that it does no harm to take 𝒵 = 𝒳 × 𝒴, extending the definition
of 𝑝(𝑥, 𝑦) with extra 0 values if necessary.

� Key idea: Theorem: Joint and single discrete random variables

Let 𝑋 and 𝑌 be two random variables on the same sample space Ω. Then the bivariate random
variable (𝑋, 𝑌 ) is discrete if and only if 𝑋 and 𝑌 are both discrete. Moreover, if 𝑋 and 𝑌 are
both discrete with 𝑃 (𝑋 ∈ 𝒳) = 𝑃 (𝑌 ∈ 𝒴) = 1 for finite or countable 𝒳 and 𝒴, then their marginal
probability mass functions are given in terms of the joint probability mass function by

𝑝𝑋 (𝑥) = ∑ 𝑝(𝑥, 𝑦) for all 𝑥 ∈ 𝒳, 𝑝𝑌 (𝑦) = ∑ 𝑝(𝑥, 𝑦) for all 𝑦 ∈ 𝒴.


𝑦∈𝒴 𝑥∈𝒳

Proof

First suppose that 𝑋 and 𝑌 are discrete. Then there exist finite or countable sets of values 𝒳 and 𝒴
such that 𝑃 (𝑋 ∈ 𝒳) = 𝑃 (𝑌 ∈ 𝒴) = 1. Hence 𝑃 (𝑋 ∈ 𝒳, 𝑌 ∈ 𝒴) = 1. The bivariate random variable
(𝑋, 𝑌 ) thus has 𝑃 ((𝑋, 𝑌 ) ∈ 𝒳 × 𝒴) = 1. Since 𝒳 and 𝒴 are finite or countable, the Cartesian
product 𝒳 × 𝒴 is also finite or countable. Hence (𝑋, 𝑌 ) is discrete.
On the other hand, suppose that (𝑋, 𝑌 ) is discrete. The possible values in (𝑋, 𝑌 )(Ω) are ordered pairs
of the form (𝑥, 𝑦), and there is a finite or countable set 𝒵 ⊆ (𝑋, 𝑌 )(Ω) such that 𝑃 ((𝑋, 𝑌 ) ∈ 𝒵) = 1.
But if we set 𝒳 = {𝑥 ∶ (𝑥, 𝑦) ∈ 𝒵} we have 𝑃 (𝑋 ∈ 𝒳) = 𝑃 ((𝑋, 𝑌 ) ∈ 𝒵) = 1, and 𝒳 is finite or
countable (check this!), so 𝑋 is discrete. Similarly for 𝑌. It remains to notice that, for example, with
the same definition of 𝒳,
𝑃 (𝑌 = 𝑦) = 𝑃 (𝑋 ∈ 𝒳, 𝑌 = 𝑦)
= 𝑃 (∪𝑥∈𝒳 {𝑋 = 𝑥, 𝑌 = 𝑦})
= ∑ 𝑝(𝑥, 𝑦),
𝑥∈𝒳

where we have used A4, the fact that 𝒳 is countable, and 𝑃 (𝑋 ∈ 𝒳) = 1.

� Try it out

Roll two fair six-sided dice. Let 𝑋 be the number of 6s rolled, and let 𝑌 be the number of 1s and 2s.
Find the joint probability mass function 𝑝(𝑥, 𝑦).
Answer:
The best way to present this is in a table:

𝑝(𝑥, 𝑦) 𝑥=0 𝑥=1 𝑥=2


𝑦=0 9/36 6/36 1/36
𝑦=1 12/36 4/36 0
𝑦=2 4/36 0 0

81
For example,
9
𝑝(0, 0) = 𝑃 (both scores in {3, 4, 5}) = ,
36
𝑝(0, 1) = 𝑃 (first is 1, 2 and second 3, 4, 5) + 𝑃 (first is 3, 4, 5 and second 1, 2)
2⋅3 3⋅2 12
= + = ,
36 36 36
and so on. Note that 𝑝(𝑥, 𝑦) sums to 1!

Again, by C7, the joint probability mass function determines the joint probability distribution of 𝑋 and 𝑌,
and the values of the joint probability mass function sum to one:

Theorem: probability mass functions determine distributions for multiple random variables

Let 𝑋 and 𝑌 be discrete random variables with 𝑃 ((𝑋, 𝑌 ) ∈ 𝒵) = 1 for a finite or countable 𝒵. Then
we have that
𝑃 ((𝑋, 𝑌 ) ∈ 𝐴) = ∑ 𝑝(𝑥, 𝑦) for all 𝐴 ⊆ 𝒵. (7.1)
(𝑥,𝑦)∈𝐴

In particular,
∑ 𝑝(𝑥, 𝑦) = 1.
(𝑥,𝑦)∈𝒵

This isn’t really a new idea: it is just the result about probability mass functions from Section 6.2 rewritten
for the case of a discrete random variable whose possible values are ordered pairs (𝑥, 𝑦).

� Try it out

Continuing with the “two-dice” example, use 𝑝(𝑥, 𝑦) to find 𝑃 (𝑋 ≥ 1, 𝑌 ≤ 1). Also, compute the
marginal probability mass functions 𝑝𝑋 (𝑥) and 𝑝𝑌 (𝑦).
Answer:
We have that
𝑃 (𝑋 ≥ 1, 𝑌 ≤ 1) = 𝑃 ((𝑋, 𝑌 ) ∈ {(1, 0), (1, 1), (2, 0), (2, 1)})
= 𝑝(1, 0) + 𝑝(1, 1) + 𝑝(2, 0) + 𝑝(2, 1)
6+4+1+0 11
= = .
36 36
For the marginal distributions, we sum down columns and along rows in the table:

𝑝(𝑥, 𝑦) 𝑥=0 𝑥=1 𝑥=2 𝑝𝑌 (𝑦)


𝑦=0 9/36 6/36 1/36 16/36
𝑦=1 12/36 4/36 0 16/36
𝑦=2 4/36 0 0 4/36
𝑝𝑋 (𝑥) 25/36 10/36 1/36 -

82
Note that this gives the same result as the binomial distribution, since 𝑋 ∼ Bin(2, 1/6) and
𝑌 ∼ Bin(2, 1/3), so, for example,

2 10
𝑃 (𝑋 = 1) = ( ) ⋅ (1/6)1 ⋅ (5/6)1 = .
1 36

Examples

1. In a card game played with a standard 52 card deck, hearts are worth 1, the queen of spades is
worth 13 and all other cards worth 0. Let 𝑋, 𝑌 denote the values of the first and second cards
dealt (without replacement, as usual). The possible outcomes of the bivariate random variable
(𝑋, 𝑌 ) are
{(0, 0), (1, 0), (0, 1), (1, 1), (0, 13), (13, 0), (1, 13), (13, 1)}
(not (13, 13)). With a well shuffled deck, the event 𝐴 that the pair does not include the queen
of spades, is
{𝑋 ≤ 1, 𝑌 ≤ 1} = {(𝑋, 𝑌 ) ∈ 𝐴}
where
𝐴 = {(0, 0), (1, 0), (0, 1), (1, 1)}.
So, by Equation 7.1, and a few counting arguments,
38 37 1 38 38 13 1 12 51 × 50 50
𝑃 ((𝑋, 𝑌 ) ∈ 𝐴) = + + + = = .
52 51 4 51 52 51 4 51 52 × 51 52

2. Discrete random variables 𝑋 and 𝑌 are such that 𝑋 takes possible values 0, 1, 2 while 𝑌 takes
values 1, 2, 3, 4, and their joint distribution is given by

𝑝(𝑥, 𝑦) 𝑦=1 𝑦=2 𝑦=3 𝑦=4


𝑥=0 0 0 0 1/4
𝑥=1 0 1/4 1/4 0
𝑥=2 1/4 0 0 0

From this table we can calculate 𝑃 (𝑋 = 𝑥) = 1/4, 1/2, 1/4 for 𝑥 = 0, 1, 2 respectively, i.e.,
𝑋 ∼ Bin(2, 1/2). Similarly 𝑃 (𝑌 = 𝑦) = 1/4 for 𝑦 = 1, 2, 3, 4 so 𝑌 is uniformly distributed on
{1, 2, 3, 4}.

Often we want to know the distribution of one random variable conditional on the value of another random
variable.

Definition: Conditional probability mass function

Let 𝑋 and 𝑌 be discrete random variables. For 𝑦 ∈ 𝒴, the conditional probability mass function of 𝑋
given 𝑌 = 𝑦 is defined by
𝑝𝑋,𝑌 (𝑥, 𝑦)
𝑝𝑋|𝑌 (𝑥, 𝑦) ∶= 𝑃 (𝑋 = 𝑥 ∣ 𝑌 = 𝑦) =
𝑝𝑌 (𝑦)

83
for all 𝑥 ∈ 𝒳 and 𝑦 ∈ 𝒴 such that 𝑝𝑌 (𝑦) > 0 and similarly, the conditional probability mass function
of 𝑌 given 𝑋 = 𝑥 is
𝑝𝑋,𝑌 (𝑥, 𝑦)
𝑝𝑌 |𝑋 (𝑦, 𝑥) ∶= 𝑃 (𝑌 = 𝑦 ∣ 𝑋 = 𝑥) =
𝑝𝑋 (𝑥)
for all 𝑦 ∈ 𝒴 and 𝑥 ∈ 𝒳 such that 𝑝𝑋 (𝑥) > 0.

There’s nothing new here, apart from the notation: this is just a particular case of the usual definition of
conditional probability of one event given another, e.g.,

𝑃 (𝑋 = 𝑥, 𝑌 = 𝑦) 𝑝𝑋,𝑌 (𝑥, 𝑦)
𝑃 (𝑋 = 𝑥 ∣ 𝑌 = 𝑦) = = .
𝑃 (𝑌 = 𝑦) 𝑝𝑌 (𝑦)

� Try it out

Continuing from the “two dice” examples above, find 𝑝(𝑥|𝑦) and calculate the conditional probability
𝑃 (𝑋 ≥ 1 ∣ 𝑌 = 0).
Answer:
Again, this is best presented in a table:

𝑝𝑋|𝑌 (𝑥, 𝑦) 𝑥=0 𝑥=1 𝑥=2


𝑦=0 9/16 6/16 1/16
𝑦=1 12/16 4/16 0
𝑦=2 4/4 0 0

For example,
𝑃 (𝑋 = 0, 𝑌 = 2)
𝑝𝑋|𝑌 (0, 2) = 𝑃 (𝑋 = 0 ∣ 𝑌 = 2) =
𝑃 (𝑌 = 2)
𝑝𝑋,𝑌 (0, 2)
=
𝑝𝑌 (2)
4/36
= = 1.
4/36
Hence
6 1 7
𝑃 (𝑋 ≥ 1 ∣ 𝑌 = 0) = 𝑝𝑋|𝑌 (1, 0) + 𝑝𝑋|𝑌 (2, 0) = + = .
16 16 16

There is also a version of P4 (partition theorem or, law of total probability) for discrete random variables;
again, only the notation is new here.

Theorem: partition theorem for discrete random variables

Let 𝑋 and 𝑌 be discrete random variables. Then,

𝑝𝑋 (𝑥) = ∑ 𝑝𝑋|𝑌 (𝑥, 𝑦)𝑝𝑌 (𝑦).


𝑦∈𝒴

Moreover, if 𝑋 is real-valued, then for any functions 𝑔 ∶ 𝒴 → ℝ and ℎ ∶ 𝒴 → ℝ with 𝑔 ≤ ℎ, we can

84
write
𝑃 (𝑔(𝑌 ) ≤ 𝑋 ≤ ℎ(𝑌 )) = ∑ 𝑃 (𝑔(𝑦) ≤ 𝑋 ≤ ℎ(𝑦) ∣ 𝑌 = 𝑦) 𝑃 (𝑌 = 𝑦)
𝑦∈𝒴

=∑ ∑ 𝑝𝑋|𝑌 (𝑥, 𝑦)𝑝𝑌 (𝑦).


𝑦∈𝒴 𝑥∈[𝑔(𝑦),ℎ(𝑦)]∩𝒳

A version of the above result also holds with 𝑋 and 𝑌 swapped.


Recall from the definition of independence of random variables that 𝑋 and 𝑌 are independent random
variables if events {𝑋 ∈ 𝐴} and {𝑌 ∈ 𝐵} are independent for all 𝐴, 𝐵. If 𝑋 and 𝑌 are discrete, the
following result gives a simpler characterization of independence.

� Key idea: Lemma: Independence of discrete random variables

Two discrete random variables 𝑋 and 𝑌 on the same sample space Ω are independent if and only if

𝑝𝑋,𝑌 (𝑥, 𝑦) = 𝑝𝑋 (𝑥)𝑝𝑌 (𝑦) for all 𝑥 ∈ 𝒳 and 𝑦 ∈ 𝒴.

It follows that for independent discrete random variables 𝑋 and 𝑌, 𝑝𝑋|𝑌 (𝑥, 𝑦) = 𝑝𝑋 (𝑥) whenever 𝑝𝑌 (𝑦) > 0,
and 𝑝𝑌 |𝑋 (𝑦, 𝑥) = 𝑝𝑌 (𝑦) whenever 𝑝𝑋 (𝑥) > 0.

Proof

By the definition of independence, we have that

𝑃 (𝑋 = 𝑥, 𝑌 = 𝑦) = 𝑃 (𝑋 = 𝑥) 𝑃 (𝑌 = 𝑦) ,

as required. On the other hand, suppose that 𝑝𝑋,𝑌 (𝑥, 𝑦) = 𝑝𝑋 (𝑥)𝑝𝑌 (𝑦). Then by the “probability
mass functions determine distributions for multiple random variables” theorem, for any sets 𝐴 and 𝐵,

𝑃 (𝑋 ∈ 𝐴, 𝑌 ∈ 𝐵) = 𝑃 ((𝑋, 𝑌 ) ∈ 𝐴 × 𝐵) = ∑ ∑ 𝑝𝑋,𝑌 (𝑥, 𝑦)


𝑥∈𝐴 𝑦∈𝐵

= ∑ 𝑝𝑋 (𝑥) ∑ 𝑝𝑌 (𝑦)
𝑥∈𝐴 𝑦∈𝐵
= 𝑃 (𝑋 ∈ 𝐴) 𝑃 (𝑌 ∈ 𝐵) ,

so 𝑋 and 𝑌 are independent.

� Try it out

Continuing the dice example, we saw that 𝑃 (𝑋 = 2, 𝑌 = 2) = 0 but 𝑃 (𝑋 = 2) 𝑃 (𝑌 = 2) = 361 4


⋅ 36 ≠ 0.
Hence 𝑋 and 𝑌 are not independent. To show dependence it is enough to find a single pair (𝑥, 𝑦)
where the joint probability mass function does not factorize. Conversely, to show independence it is
necessary to consider all pairs (𝑥, 𝑦).

� Textbook references

If you want more help with this section, check out:

• Section 7.1 in (Blitzstein and Hwang 2019);

85
• or Chapter 6 in (Anderson, Seppäläinen, and Valkó 2018).

7.3 Jointly continuously distributed random variables

Definition: jointly continuous random variables

Consider two real-valued random variables 𝑋 ∶ Ω → ℝ and 𝑌 ∶ Ω → ℝ. We say that 𝑋 and 𝑌


are jointly continuously distributed when there is a non-negative piecewise continuous function
𝑓(⋅) ∶ ℝ2 → ℝ, called the joint probability density function, such that
𝑏 𝑑
𝑃 (𝑋 ∈ [𝑎, 𝑏], 𝑌 ∈ [𝑐, 𝑑]) = ∫ (∫ 𝑓(𝑥, 𝑦) 𝑑𝑦) 𝑑𝑥
𝑎 𝑐

for all [𝑎, 𝑏] × [𝑐, 𝑑] ⊆ ℝ2 .

To avoid ambiguity we sometimes write 𝑓𝑋,𝑌 (𝑥, 𝑦) for the joint probability density of 𝑋 and 𝑌. The
interpretation of 𝑓(𝑥, 𝑦) is that

𝑃 (𝑋 ∈ [𝑥, 𝑥 + 𝑑𝑥], 𝑌 ∈ [𝑦, 𝑦 + 𝑑𝑦]) = 𝑓(𝑥, 𝑦) 𝑑𝑥 𝑑𝑦 (7.2)

for 𝑥, 𝑦 at which 𝑓(𝑥, 𝑦) is continuous.


As before, the joint probability density function 𝑓(𝑥, 𝑦) determines the joint probability 𝑃 ((𝑋, 𝑌 ) ∈ 𝐴) for
most events 𝐴. More precisely, remember the definition of type I and II regions in the theory of multiple
integration: a type I region is of the form

𝐷 = {𝑎0 ≤ 𝑥 ≤ 𝑎1 , 𝜙1 (𝑥) ≤ 𝑦 ≤ 𝜙2 (𝑥)}

for some continuous functions 𝜙1 and 𝜙2 , and a type II region is of the form

𝐷 = {𝜓1 (𝑦) ≤ 𝑥 ≤ 𝜓2 (𝑦), 𝑏0 ≤ 𝑦 ≤ 𝑏1 }

for some continuous functions 𝜓1 and 𝜓2 .

86
Figure 7.1: Regions of types 1 and 2

As we know from calculus, unions of regions of these types are precisely the regions over which we can
integrate1 . So, we have:

Theorem: probabilities as integrals for joint distributions

If 𝑋 and 𝑌 are jointly continuously distributed, then for any 𝐴 ⊆ ℝ2 that is a finite union of type I
and type II regions:
𝑃 ((𝑋, 𝑌 ) ∈ 𝐴) = ∬ 𝑓(𝑥, 𝑦) 𝑑𝑥 𝑑𝑦.
𝐴

Again as before, 𝑓(𝑥, 𝑦) integrates to one:

Corollary: joint densities integrate to 1

Let 𝑋 and 𝑌 be jointly continuously distributed random variables. Then their joint probability
density function integrates to one:
∬ 𝑓(𝑥, 𝑦) 𝑑𝑥 𝑑𝑦 = 1.
ℝ2

� Try it out

Suppose that 𝑋 and 𝑌 have joint probability density function

𝑓(𝑥, 𝑦) = 𝑐(𝑥2 + 𝑦), for − 1 ≤ 𝑥 ≤ 1, 0 ≤ 𝑦 ≤ 1 − 𝑥2 ,

1
At least when restricted to the Riemann integral.

87
with 𝑓(𝑥, 𝑦) = 0 otherwise. What is the value of 𝑐?
Answer: This is an exercise in multiple integration. We have from the Corollary that

1 = ∬ 𝑓(𝑥, 𝑦) 𝑑𝑥 𝑑𝑦
ℝ2
1 1−𝑥2
= ∫ 𝑑𝑥 ∫ 𝑐(𝑥2 + 𝑦) 𝑑𝑦
−1 0
1 1−𝑥2
𝑦2
= 𝑐 ∫ 𝑑𝑥 [𝑥2 𝑦 + ]
−1
2 0
𝑐 1
= ∫ (1 − 𝑥4 ) 𝑑𝑥
2 −1
1
𝑐 𝑥5
= [𝑥 − ]
2 5 −1
4
= 𝑐.
5
So we find that 𝑐 = 5/4.

� Try it out

Consider random variables 𝑋 and 𝑌 with joint probability density function

𝑥+𝑦 if (𝑥, 𝑦) ∈ [0, 1]2


𝑓(𝑥, 𝑦) = {
0 otherwise.

(a) Calculate 𝑃 (1/4 < 𝑋 < 3/4, 0 < 𝑌 < 1/2)

(b) Calculate 𝑃 (𝑋 2 < 𝑌 < 𝑋)

Answer:
Again, this is an exercise in multiple integration. For part (a) we have
3/4 1/2
𝑃 (1/4 < 𝑋 < 3/4, 0 < 𝑌 < 1/2) = ∫ 𝑑𝑥 ∫ (𝑥 + 𝑦) 𝑑𝑦
1/4 0
3/4 1/2
𝑦2
=∫ 𝑑𝑥 [𝑥𝑦 + ]
1/4
2 0
3/4
𝑥 1
=∫ ( + ) 𝑑𝑥
1/4
2 8
3
= .
16

88
For part (b), we have
1 𝑥 1 𝑥
𝑦2
𝑃 (𝑋 2 < 𝑌 < 𝑋) = ∫ 𝑑𝑥 ∫ (𝑥 + 𝑦) 𝑑𝑦 = ∫ 𝑑𝑥 [𝑥𝑦 + ]
0 𝑥2 0
2 𝑥2
1 2 4
3𝑥 𝑥 3
=∫ ( − 𝑥3 − ) 𝑑𝑥 = .
0
2 4 20

Note that, if 𝑋 and 𝑌 are jointly continuously distributed, then for any interval [𝑎, 𝑏],
𝑏 ∞
𝑃 (𝑋 ∈ [𝑎, 𝑏]) = 𝑃 (𝑋 ∈ [𝑎, 𝑏], 𝑌 ∈ ℝ) = ∫ (∫ 𝑓(𝑥, 𝑦) 𝑑𝑦) 𝑑𝑥,
𝑎 −∞

so, by the definition of a continuous random variable, 𝑋 is continuously distributed as well, as is 𝑌 by a


similar argument. We have shown the following:

Corollary

Let 𝑋 and 𝑌 be jointly continuously distributed random variables. Then 𝑋 and 𝑌 are (each separately)
continuously distributed, with
∞ ∞
𝑓𝑋 (𝑥) = ∫ 𝑓(𝑥, 𝑦) 𝑑𝑦, for all 𝑥 ∈ ℝ, 𝑓𝑌 (𝑦) = ∫ 𝑓(𝑥, 𝑦) 𝑑𝑥, for all 𝑦 ∈ ℝ.
−∞ −∞

In a multivariate context, the probability density functions 𝑓𝑋 (𝑥) and 𝑓𝑌 (𝑦) are also called marginal
probability density functions.

� Try it out

Continuing from the previous example, we have that for 𝑥 ∈ [0, 1],
1 1
𝑦2 1
𝑓𝑋 (𝑥) = ∫ (𝑥 + 𝑦) 𝑑𝑦 = [𝑥𝑦 + ] =𝑥+ ,
0
2 0 2

so
𝑥+ 1
2 if 0 ≤ 𝑥 ≤ 1,
𝑓𝑋 (𝑥) = {
0 otherwise.

� Try it out

As in a previous example, suppose that 𝑋 and 𝑌 have joint probability density function

𝑓(𝑥, 𝑦) = 𝑐(𝑥2 + 𝑦), for − 1 ≤ 𝑥 ≤ 1, 0 ≤ 𝑦 ≤ 1 − 𝑥2 ,

with 𝑓(𝑥, 𝑦) = 0 otherwise. we find


2
5 1−𝑥 2 5
𝑓( 𝑥) = ∫ (𝑥 + 𝑦) 𝑑𝑦 = (1 − 𝑥4 )
4 0 8

89
for −1 ≤ 𝑥 ≤ 1 and 0 otherwise;

1−𝑦
5 5
𝑓( 𝑦) = ∫ (𝑥2 + 𝑦) 𝑑𝑥 = (1 + 2𝑦)√1 − 𝑦
4 −√1−𝑦 6

for 0 ≤ 𝑦 ≤ 1 and 0 otherwise.

Advanced content

So, if 𝑋 and 𝑌 are jointly continuously distributed, then both 𝑋 and 𝑌 are also continuously distributed
separately. Unlike the discrete case (the “joint and single discrete random variables” theorem), the
converse, however, is not true in general, as the following example shows.

Example

Let 𝑋 be a continuously distributed random variable, and let 𝑌 ∶= 2𝑋. Then both 𝑋 and 𝑌 are
continuously distributed separately, however 𝑋 and 𝑌 are not jointly continuously distributed.
Indeed, suppose that 𝑋 and 𝑌 were jointly continuously distributed with density function
𝑓(𝑥, 𝑦). Then, with 𝐴 = {(𝑥, 𝑦) ∶ 2𝑥 = 𝑦},
+∞ 2𝑥
∬ 𝑓(𝑥, 𝑦) 𝑑𝑥 𝑑𝑦 = ∫ (∫ 𝑓(𝑥, 𝑦) 𝑑𝑦) 𝑑𝑥 = 0.
−∞ 2𝑥
𝐴

This implies that 𝑃 (2𝑋 = 𝑌) = 0 for every two jointly continuously distributed random variables
𝑋 and 𝑌. But, because for our choice of 𝑋 and 𝑌, obviously 𝑃 (2𝑋 = 𝑌) = 1; so, by contradiction,
𝑋 and 𝑌 cannot be jointly continuously distributed.

Conditional probability density function

Let 𝑋 and 𝑌 be jointly continuously distributed random variables. The conditional probability density
function of 𝑋 at 𝑌 = 𝑦 is defined by

𝑓𝑋,𝑌 (𝑥, 𝑦)
𝑓𝑋|𝑌 (𝑥|𝑦) ∶= for all 𝑥, 𝑦 ∈ ℝ such that 𝑓𝑌 (𝑦) > 0.
𝑓𝑌 (𝑦)

and similarly, the conditional probability density function of 𝑌 at 𝑋 = 𝑥 is

𝑓𝑋,𝑌 (𝑥, 𝑦)
𝑓𝑌 |𝑋 (𝑦|𝑥) ∶= for all 𝑥, 𝑦 ∈ ℝ such that 𝑓𝑋 (𝑥) > 0.
𝑓𝑋 (𝑥)

Advanced content

Roughly speaking, 𝑓𝑋|𝑌 (𝑥|𝑦) should be thought of as the probability density function of 𝑋 conditional
on 𝑌 = 𝑦. Because the event 𝑌 = 𝑦 has probability 0, this interpretation needs a bit of work to

90
realize rigorously. A formal manipulation using Equation 6.2 and Equation 7.2 goes as follows:

𝑃 (𝑋 ∈ [𝑥, 𝑥 + 𝑑𝑥], 𝑌 ∈ [𝑦, 𝑦 + 𝑑𝑦])


𝑃 (𝑋 ∈ [𝑥, 𝑥 + 𝑑𝑥] ∣ 𝑌 ∈ [𝑦, 𝑦 + 𝑑𝑦]) =
𝑃 (𝑌 ∈ [𝑦, 𝑦 + 𝑑𝑦])
𝑓𝑋,𝑌 (𝑥, 𝑦) 𝑑𝑥 𝑑𝑦
= = 𝑓𝑋|𝑌 (𝑥|𝑦) 𝑑𝑥,
𝑓𝑌 (𝑦) 𝑑𝑦

so 𝑓𝑋|𝑌 (𝑥|𝑦) is the probability density of 𝑋 conditional on 𝑌 ∈ [𝑦, 𝑦 + 𝑑𝑦]. This relies on 𝑓𝑋,𝑌 and
𝑓𝑌 being continuous so that Equation 6.2 and Equation 7.2 are valid.

Theorem: Partition theorem for jointly continuous random variables

Let 𝑋 and 𝑌 be jointly continuously distributed random variables. Then,


+∞
𝑓𝑋 (𝑥) = ∫ 𝑓𝑋|𝑌 (𝑥|𝑦)𝑓𝑌 (𝑦) 𝑑𝑦.
−∞

Moreover, for any piecewise continuous functions 𝑔 ∶ ℝ → ℝ and ℎ ∶ ℝ → ℝ with 𝑔 ≤ ℎ, we can write,
+∞ ℎ(𝑦)
𝑃 (𝑔(𝑌 ) ≤ 𝑋 ≤ ℎ(𝑌 )) = ∫ (∫ 𝑓𝑋|𝑌 (𝑥|𝑦) 𝑑𝑥) 𝑓𝑌 (𝑦) 𝑑𝑦. (7.3)
−∞ 𝑔(𝑦)

Advanced content

We can formulate this theorem in a way that is more similar to the discrete case. For any event 𝐴,
and any continuously distributed random variable 𝑌, we define

𝑃 (𝐴 ∣ 𝑌 = 𝑦) ∶= lim 𝑃 (𝐴 ∣ 𝑌 ∈ [𝑦, 𝑦 + ℎ])


ℎ→0

whenever 𝑃 (𝐴 ∣ 𝑌 ∈ [𝑦, 𝑦 + ℎ]) > 0 for all ℎ sufficiently small—this happens when 𝑓( 𝑦) is continuous at
𝑦 and 𝑓( 𝑦) > 0. Our usual definition of conditional probability does not apply, because 𝑃 (𝑌 = 𝑦) = 0,
so 𝑃 (𝐴 ∣ 𝑌 = 𝑦) is not really a conditional probability. One should be warned that the notation
𝑃 (𝐴 ∣ 𝑌 = 𝑦) for continuously distributed 𝑌 can lead to extremely confusing issues, such as Borel’s
paradox. Anyway, ignoring potential pitfalls, with this notation
ℎ(𝑦)
𝑃 (𝑔(𝑦) ≤ 𝑋 ≤ ℎ(𝑦) ∣ 𝑌 = 𝑦) = ∫ 𝑓(𝑥|𝑦) 𝑑𝑥,
𝑔(𝑦)

and consequently, we can rewrite Equation 7.3 as


+∞
𝑃 (𝑔(𝑌 ) ≤ 𝑋 ≤ ℎ(𝑌 )) = ∫ 𝑃 (𝑔(𝑦) ≤ 𝑋 ≤ ℎ(𝑦) ∣ 𝑌 = 𝑦) 𝑓(𝑦) 𝑑𝑦.
−∞

Similarly to the discrete case, independence can be characterized as a factorization property of the joint
probability density function.

91
� Key idea: Lemma: independence of jointly continuous random variables

Two jointly continuously distributed random variables 𝑋 and 𝑌 are independent if and only if

𝑓𝑋,𝑌 (𝑥, 𝑦) = 𝑓𝑋 (𝑥)𝑓𝑌 (𝑦) for all 𝑥 and 𝑦 ∈ ℝ.

For independent jointly continuously distributed 𝑋 and 𝑌, 𝑓𝑋|𝑌 (𝑥|𝑦) = 𝑓𝑋 (𝑥) whenever 𝑓𝑌 (𝑦) > 0, and
𝑓𝑌 |𝑋 (𝑦|𝑥) = 𝑓𝑌 (𝑦) whenever 𝑓𝑋 (𝑥) > 0.

Example

Suppose 𝑋 and 𝑌 have joint probability density function

3𝑒−(𝑥+3𝑦) if 𝑥 ≥ 0 and 𝑦 ≥ 0,
𝑓(𝑥, 𝑦) = {
0 otherwise.

Because we can write 𝑓(𝑥, 𝑦) as 𝑒−𝑥 ⋅ 3𝑒−3𝑦 (for 𝑥 ≥ 0 and 𝑦 ≥ 0), it follows that 𝑋 ∼ ℰ(1) and
𝑌 ∼ ℰ(3), and they are independent. We can calculate things like, for 𝑎 > 0,
∞ ∞
𝑃 (𝑎𝑌 < 𝑋) = ∫ (∫ 𝑓(𝑥|𝑦) 𝑑𝑥) 𝑓( 𝑦) 𝑑𝑦
0 𝑎𝑦
∞ ∞ ∞
=∫ (∫ 𝑒−𝑥 𝑑𝑥) 3𝑒−3𝑦 𝑑𝑦 = ∫ 𝑒−𝑎𝑦 ⋅ 3𝑒−3𝑦 𝑑𝑦 = 3/(3 + 𝑎).
0 𝑎𝑦 0

� Textbook references

If you want more help on this section, check out:

• Section 7.1 in (Blitzstein and Hwang 2019);


• Section 6.2 in (Anderson, Seppäläinen, and Valkó 2018);
• or Sections 8.1 and 8.3 in (Stirzaker 2003).

7.4 Functions of multiple random variables

Suppose 𝑋 ∶ Ω → 𝑋(Ω) and 𝑌 ∶ Ω → 𝑌 (Ω) are (discrete or continuous) random variables, and 𝑔 ∶
𝑋(Ω) × 𝑌 (Ω) → 𝒮 is some function assigning a value 𝑔(𝑥, 𝑦) ∈ 𝒮 to each point (𝑥, 𝑦). Then 𝑔(𝑋, 𝑌 ) is also
a random variable, namely the outcome to a ‘new experiment’ obtained by running the ‘old experiments’
to produce values 𝑥 for 𝑋 and 𝑦 for 𝑌, and then evaluating 𝑔(𝑥, 𝑦).

92
Figure 7.2: 𝑔(𝑋, 𝑌 ) as a random variable

Formally, 𝑔(𝑋, 𝑌 ) ∶= 𝑔 ∘ (𝑋, 𝑌 ), or in more specific terms, the random variable 𝑔(𝑋, 𝑌 ) ∶ Ω → 𝒮 is defined
by:
𝑔(𝑋, 𝑌 )(𝜔) ∶= 𝑔(𝑋(𝜔), 𝑌 (𝜔)) for all 𝜔 ∈ Ω.
For example:
𝑃 (𝑔(𝑋, 𝑌 ) ∈ 𝐴) = 𝑃 ({𝜔 ∈ Ω ∶ 𝑔(𝑋(𝜔), 𝑌 (𝜔)) ∈ 𝐴}) for all 𝐴 ⊆ 𝒮.

For any random variables 𝑋 and 𝑌, 𝑋 + 𝑌, 𝑋 − 𝑌, 𝑋𝑌, min(𝑋, 𝑌 ), 𝑒𝑡(𝑋+𝑌 ) , and so on, are all random
variables as well.

� Try it out

Consider the jointly continuous random variables 𝑋 and 𝑌 from . Define a new random variable
𝑆 = 𝑋 + 𝑌.

(a) Calculate 𝐹𝑆 (⋅) and hence deduce that 𝑆 is a continuous random variable.
(b) Identify 𝑓𝑆 .

Answer:
For part (a), we note that 𝑃 (0 ≤ 𝑆 ≤ 2) = 1 so 𝐹𝑆 (𝑠) = 0 for 𝑠 < 0 and 𝐹𝑆 (𝑠) = 1 for 𝑠 ≥ 2. Suppose
that 0 ≤ 𝑠 ≤ 1. Then
𝑠 𝑠−𝑦
𝐹𝑆 (𝑠) = 𝑃 (𝑋 + 𝑌 ≤ 𝑠) = ∫ (∫ (𝑥 + 𝑦) 𝑑𝑥) 𝑑𝑦
0 0
𝑠
𝑥=𝑠−𝑦
= ∫ [𝑥2 /2 + 𝑥𝑦]𝑥=0 𝑑𝑦
0
1 𝑠 2
= ∫ (𝑠 − 𝑦2 ) 𝑑𝑦
2 0
𝑠3
= .
3

93
Next, for 1 < 𝑠 ≤ 2,

𝐹𝑆 (𝑠) = 𝑃 (𝑋 + 𝑌 ≤ 𝑠)
𝑠−1 1 1 𝑠−𝑦
=∫ (∫ (𝑥 + 𝑦) 𝑑𝑥) 𝑑𝑦 + ∫ (∫ (𝑥 + 𝑦) 𝑑𝑥) 𝑑𝑦
0 0 𝑠−1 0
𝑠−1
1 1 2
=∫ (𝑦 + 1/2) 𝑑𝑦 + ∫ (𝑠 − 𝑦2 ) 𝑑𝑦
0
2 𝑠−1
1 + 𝑠3
= 𝑠2 − .
3
We see that 𝐹𝑆 (⋅) is continuous, and piecewise differentiable. Hence there is a density which is
obtained by differentiation:

⎧𝑠2 if 0 ≤ 𝑠 ≤ 1,
𝑑𝐹𝑆 (𝑠) {
𝑓𝑆 (𝑠) = = 2𝑠 − 𝑠2
⎨ if 1 < 𝑠 ≤ 2,
𝑑𝑠 {0
⎩ elsewhere,

and in fact 𝑓𝑆 (𝑠) is continuous everywhere.

� Textbook references

If you want more help on this section, check out:

• Section 3.9 in (DeGroot and Schervish 2013).

94
8 Expectation

� Goals

1. Have an intuitive as well as mathematical understanding of expectation, variance, and covariance.


Know how expectation, variance, and covariance, behave under linear transformations and
sums.

2. Know how to evaluate expectation of functions of random variables.

3. Know the properties of expectation, variance, and covariance, and the relations between them.

4. Understand the difference between variance and standard deviation.

5. Know conditional expectation, the partition theorem for conditional expectation and the special
notation associated with it.

6. Know how expectation, variance, and standard deviation behave under independence, and for
sums of independent random variables in particular.

7. Know the Markov and Chebyshev inequalities, where they come from, and how to apply them.

8.1 Definition and interpretation

In a relative frequency interpretation (discussed earlier in Section 4.1), suppose that we run 𝑛 trials on an
experiment where we observe the outcome of some real-valued random variable 𝑋 ∶ Ω → ℝ in each trial.
Let 𝑥𝑖 denote the observed value of 𝑋 in the 𝑖th trial; the sequence of observations 𝑥1 , 𝑥2 , …, 𝑥𝑛 is called
a sample. The sample mean is then simply 𝑛1 ∑𝑖=1 𝑥𝑖 . As a mathematical idealization, we may suppose
𝑛

that there is a unique, empirical limiting value for the sample mean, as 𝑛 tends to infinity, which we call
the expectation of 𝑋.
In a betting interpretation (discussed earlier in Section 4.2), you can simply consider your ‘fair price’ for a
bet which pays 𝑋; that price, we call your expectation of 𝑋.
The idea of expectation is very interesting mathematically, and also provides ways to use probability
in a host of practical applications. Regardless of interpretation, the expectation of 𝑋 can be connected
to the probability mass function 𝑝(𝑥) (if 𝑋 is discrete) or the probability density function 𝑓(𝑥) (if 𝑋 is
continuously distributed) in the following way:

� Key idea: Definition: expectation

For any real-valued random variable 𝑋, the expectation (also called expected value or mean) of 𝑋,
denoted as 𝔼[𝑋], is defined as:
𝔼[𝑋] ∶= ∑ 𝑥 𝑝(𝑥) (8.1)
𝑥∈𝒳

95
if 𝑋 is discrete, and

𝔼[𝑋] ∶= ∫ 𝑥 𝑓(𝑥) 𝑑𝑥 (8.2)
−∞
if 𝑋 is continuously distributed, provided that the sum or integral exists.

Examples

1. Suppose that 𝑋 is discrete with probability mass function

𝑥 1 2
1 1
𝑝(𝑥) 2 2

Then 𝔼[𝑋] = 1
2 ⋅1+ 1
2 ⋅ 2 = 1.5.

2. Consider the following ‘game’. You pay Jimmy a pound and then you both throw a fair die.
If you get the higher number you get back the difference in pounds, otherwise you lose your
pound. Call the return from a game 𝑋, with possible values 0, 1, 2, 3, 4 and 5. By counting
outcomes,

𝑥 0 1 2 3 4 5
21 5 4 3 2 1
𝑝(𝑥) 36 36 36 36 36 36

so that
5
𝔼[𝑋] = ∑ 𝑥 𝑝(𝑥) = (0 × 21 + 1 × 5 + 2 × 4 + 3 × 3 + 4 × 2 + 5 × 1)/36 = 35/36.
𝑥=0

Since it costs £1 to play, this means that the expected profit is −£1/36. We can interpret this value
as meaning that over a long series of games you will get back £35 for every £36 paid out.

3. To find the expectation of a discrete random variable 𝑋 where 𝑝(𝑥) = 1/𝑛 for 𝑥 ∈ {1, 2, … , 𝑛},
we compute
𝑛
𝔼[𝑋] = ∑ 𝑥𝑝(𝑥)
𝑥=1
1 𝑛+1
= (1 + 2 + ⋯ + 𝑛) = .
𝑛 2
4. If 𝑋 ∼ U(𝑎, 𝑏) then
𝑏 𝑏
𝑥 𝑥2 /2 𝑎+𝑏
𝔼[𝑋] = ∫ 𝑑𝑥 = [ ] = .
𝑎
𝑏−𝑎 𝑏−𝑎 𝑎 2
5. If 𝑍 ∼ 𝒩(0, 1) then
+∞
𝔼[𝑍] = ∫ 𝑧𝜙(𝑧) d𝑧 = 0,
−∞

since the integrand is an odd function (𝜙(𝑧) = 𝜙(−𝑧)).

96
� Try it out

Find the expectation of a continuous random variable 𝑋 where

𝑥/2 if 𝑥 ∈ [0, 2],


𝑓(𝑥) = {
0 elsewhere.

Answer: We compute
2 2
𝑥 𝑥3 4
𝔼[𝑋] = ∫ 𝑥 ⋅ d𝑥 = [ ] = .
0
2 6 0 3

Advanced content

If the range of possible values for a random variable 𝑋 is unbounded, then the sum or integral in may
fail to exist. In this case, the preceding formulas may still be used to assign a meaningful expectation
in some cases, provided we interpret them with care.
For example, if 𝑋 is discrete with probability mass function 𝑝(𝑥), consider

𝔼[𝑋] = ∑ 𝑥𝑝(𝑥) = ∑ 𝑥𝑝(𝑥) + ∑ 𝑥𝑝(𝑥);


𝑥∈𝑋(Ω) ⏟⏟ ⏟⏟⏟⏟⏟
𝑥∈𝑋(Ω)∶𝑥≥0 ⏟⏟ ⏟⏟⏟⏟⏟
𝑥∈𝑋(Ω)∶𝑥≤0
𝑆+ 𝑆−

now the individual sums 𝑆+ and 𝑆− always exist, but may be equal to +∞.
In fact, if we write 𝑋 + = max(𝑋, 0) and 𝑋 − = max(−𝑋, 0), then 𝑋 = 𝑋 + − 𝑋 − and 𝔼[𝑋 + ] = 𝑆+
and 𝔼[𝑋 − ] = 𝑆− .
To see this, note for example that 𝑋 + is a random variable with 𝑝𝑋+ 𝑥 = 𝑝𝑋 (𝑥) for 𝑥 > 0 and
𝑝𝑋+ (0) = 𝑃 (𝑋 ≤ 0), but only the positive terms contribute to 𝔼[𝑋 + ].
It makes sense to say that 𝔼[𝑋] = 𝔼[𝑋 + ] − 𝔼[𝑋 − ] (being possibly −∞ or +∞) as long as at most
one of 𝑆+ and 𝑆− are infinite, using the rules ∞ − 𝑥 = ∞ and 𝑥 − ∞ = −∞ for finite 𝑥. (There is no
sensible interpretation of ∞ − ∞.) A similar argument applies in the continuous case, with integrals
instead of sums. This is summarized in the following table, which shows the values of 𝔼[𝑋] in each
case.

𝔼[𝑋 + ] < ∞ 𝔼[𝑋 + ] = ∞


𝔼[𝑋 − ] < ∞ 𝔼[𝑋 + ] − 𝔼[𝑋 − ] +∞
𝔼[𝑋 − ] = ∞ −∞ undefined

Examples

Suppose that 𝑋 is discrete with probability mass function 𝑝(𝑥) = 𝑐𝛼 𝑥−𝛼 for 𝑥 ∈ {1, 2, …}. This
is only a proper probability mass function if 𝜁(𝛼) ∶= ∑∞𝑥=1
𝑥−𝛼 < ∞, so we need 𝛼 > 1. Then
the normalizing constant must be 𝑐𝛼 = 1/𝜁(𝛼). But 𝔼[𝑋] = 𝑐𝛼 ∑𝑥=1 𝑥1−𝛼 . If 𝛼 ∈ (1, 2], this

sum diverges, so 𝔼[𝑋] = +∞. This is the case if, for instance, 𝑝(𝑥) = (6/𝜋2 )𝑥−2 .

� Textbook references

If you want more help with this section, check out:

• Sections 4.1 and 5.1 in (Blitzstein and Hwang 2019);

97
• Section 3.3 in (Anderson, Seppäläinen, and Valkó 2018);
• or Sections 4.3 and 7.4 in (Stirzaker 2003).

8.2 Expectation of functions of random variables

Let 𝑋 be a discrete random variable with 𝑃 (𝑋 ∈ 𝒳) = 1 for a finite or countable set 𝒳, and let 𝑔 ∶ 𝒳 → ℝ
be a real-valued function. As seen in Section 6.11, 𝑔(𝑋) ∶= 𝑔 ∘ 𝑋 is again a random variable. Indeed, 𝑔(𝑋)
is discrete, since 𝑃 (𝑔(𝑋) ∈ 𝑔(𝒳)) = 1 where 𝑔(𝒳) ∶= {𝑔(𝑥) ∶ 𝑥 ∈ 𝒳} is finite or countable, and 𝑔(𝑋) is a
real-valued random variable, so we can define its expectation.
To find the expectation of 𝑔(𝑋), by Equation 8.1, according to the definition we need to find the probability
mass function 𝑝𝑔(𝑋) () first. It turns out however that we can express 𝔼[𝑔(𝑋)] directly in terms of 𝑝𝑋 (),
saving us the effort of having to calculate 𝑝𝑔(𝑋) () from 𝑝𝑋 ().
For any 𝑦 ∈ 𝑔(𝒳),

𝑝𝑔(𝑋) (𝑦) = 𝑃 (𝑔(𝑋) = 𝑦) = ∑ 𝑃 (𝑔(𝑋) = 𝑦 ∣ 𝑋 = 𝑥) 𝑃 (𝑋 = 𝑥) = ∑ 𝑝(𝑥), (8.3)


𝑥∈𝒳 𝑥∈𝒳∶𝑔(𝑥)=𝑦

since
1 if 𝑦 = 𝑔(𝑥),
𝑃 (𝑔(𝑋) = 𝑦 ∣ 𝑋 = 𝑥) = {
0 otherwise.
It follows that
𝔼[𝑔(𝑋)] = ∑ 𝑦 𝑝𝑔(𝑋) (𝑦) = ∑ 𝑦 ( ∑ 𝑝(𝑥))
𝑦∈𝑔(𝒳) 𝑦∈𝑔(𝒳) 𝑥∈𝒳∶𝑔(𝑥)=𝑦

= ∑ ( ∑ 𝑦𝑝(𝑥)) = ∑ ( ∑ 𝑦𝑝(𝑥))
𝑦∈𝑔(𝒳) 𝑥∈𝒳∶𝑔(𝑥)=𝑦 𝑥∈𝒳 𝑦∈𝑔(𝒳)∶𝑦=𝑔(𝑥)

= ∑( ∑ 𝑦) 𝑝(𝑥) = ∑ 𝑔(𝑥)𝑝(𝑥),
𝑥∈𝒳 𝑦∈𝑔(𝒳)∶𝑦=𝑔(𝑥) 𝑥∈𝒳

where we applied the definition of expectation, Equation 8.3, distributivity, change of order of summation,
and distributivity again. A similar result can be proven when 𝑋 is continuously distributed. Concluding,
we have the following result, which is sometimes known as the Law of the Unconscious Statistician:

� Key idea: Theorem: expectation of a function of a random variable

For any discrete random variable 𝑋 taking values in 𝒳, and any function 𝑔 ∶ 𝒳 → ℝ,

𝔼[𝑔(𝑋)] = ∑ 𝑔(𝑥)𝑝(𝑥), (8.4)


𝑥∈𝒳

provided that the sum exists. Similarly, for any continuous random variable 𝑋 and any function
𝑔 ∶ ℝ → ℝ,

𝔼[𝑔(𝑋)] = ∫ 𝑔(𝑥)𝑓(𝑥) d𝑥, (8.5)
−∞
provided that the integral exists.

98
Examples

1. Suppose that 𝑋 takes values 0, 1, 2, 3, 4 each with probability 1/5. Then


4
𝔼[(𝑋 − 3)2 ] = ∑(𝑥 − 3)2 𝑝(𝑥)
𝑥=0
1
= ((0 − 3)2 + (1 − 3)2 + (2 − 3)2 + (3 − 3)2 + (4 − 3)2 )
5
1
= (9 + 4 + 1 + 0 + 1) = 3.
5

2. Suppose 𝑋 takes values −2, −1, 0, 1, 2, 3 each with probability 1/6. Then

1 19
𝔼[𝑋 2 ] = ((−2)2 + (−1)2 + 0 + 1 + 22 + 32 ) = ;
6 6
1 √ √ √ 1
𝔼[sin(𝜋𝑋/4)] = (−1 − 1/ 2 + 0 + 1/ 2 + 1 + 1/ 2) = √ ;
6 6 2
and so on.

3. If 𝑋 ∼ U(−1, 1) then 𝑓(𝑥) = 1


2 for 𝑥 ∈ [−1, 1], and zero elsewhere, so

𝔼[𝑋 2 ] = ∫ 𝑥2 𝑓(𝑥) d𝑥
−∞
1
1
= ∫ 𝑥2 ⋅ d𝑥
−1
2
3 1
𝑥 1
=[ ] = .
6 −1 3

4. Note that although 𝑔(𝑋) is discrete if 𝑋 is discrete, if 𝑋 is continuous then 𝑔(𝑋) need not
be continuous: for example, if 𝑋 ∼ U(0, 2) and 𝑔(𝑥) = 1 if 𝑥 ∈ (0, 1) and 𝑔(𝑥) = 0 otherwise,
we have that 𝑔(𝑋) is discrete with 𝑃 (𝑔(𝑋) = 1) = 1/2 and 𝑃 (𝑔(𝑋) = 0) = 1/2. In this case
𝑓(𝑥) = 1/2 for 𝑥 ∈ (0, 2), and says that

1 2 1
𝔼[𝑔(𝑋)] = ∫ 𝑔(𝑥) d𝑥 = ,
2 0 2

as we would get from a direct calculation for the discrete random variable 𝑔(𝑋) as 𝔼[𝑔(𝑋)] =
2 ⋅ 0 + 2 ⋅ 1 = 2.
1 1 1

Advanced content

Similar comments apply here about extensions of 𝔼[𝑔(𝑋)] to include +∞ or −∞ as at the end of the
previous section.

99
Example

Suppose 𝑋 ∼ U(−1, 1) i.e., 𝑋 is uniformly distributed on the interval (−1, 1), and we set
𝑔(𝑥) = 1/𝑥 for 𝑥 ≠ 0 and 𝑔(0) = 0. Then 𝔼[𝑔(𝑋)] is not defined because 𝔼[𝑔(𝑋)+ ] =
𝑑𝑥 = ∞ and similarly 𝔼[𝑔(𝑋)− ] = ∞.
0 1 1
∫ 0 12 𝑑𝑥 + ∫ 2𝑥
−1 0

For multiple random variables, the Law of the Unconscious Statistician reads as follows:

� Key idea: Theorem: Expectation of a function of a multivariate random variable

For any discrete random variables 𝑋 and 𝑌 taking values in 𝒳 and 𝒴, and any function 𝑔 ∶ 𝒳 × 𝒴 → ℝ,

𝔼[𝑔(𝑋, 𝑌 )] = ∑ ∑ 𝑔(𝑥, 𝑦)𝑝(𝑥, 𝑦),


𝑥∈𝒳 𝑦∈𝒴

provided that the sum exists. Similarly, for any jointly continuously distributed random variables 𝑋
and 𝑌, and any function 𝑔 ∶ ℝ2 → ℝ,

𝔼[𝑔(𝑋, 𝑌 )] = ∬ 𝑔(𝑥, 𝑦)𝑓(𝑥, 𝑦) d𝑥 d𝑦,


ℝ2

provided that the integral exists.

Examples

1. Consider discrete random variables 𝑋 and 𝑌 with joint probability mass function:

𝑝(𝑥, 𝑦) 𝑥=1 𝑥=2 𝑥=3


𝑦=1 1/2 0 1/8
𝑦=2 0 1/4 1/8

100
Then
𝔼[(𝑋 − 2)𝑌 ] = ∑ ∑(𝑥 − 2)𝑦𝑝(𝑥, 𝑦)
𝑥 𝑦
1 1 1 1
= (1 − 2) ⋅ 1 ⋅ + (2 − 2) ⋅ 2 ⋅ + (3 − 2) ⋅ 1 ⋅ + (3 − 2) ⋅ 2 ⋅
2 4 8 8
1
=− .
8
2. Consider discrete random variables 𝑋 and 𝑌 with joint probability mass function:

𝑝(𝑥, 𝑦) 𝑥 = −1 𝑥=0 𝑥=1


𝑦=0 1/4 0 1/4
𝑦=1 0 1/4 1/4

Then 𝔼[𝑋𝑌 ] = 14 ((−1) × 0 + 0 × 1 + 1 × 0 + 1 × 1) = 1/4.

� Try it out

Let 𝑋 and 𝑌 be jointly continuously distributed random variables, with

1 if (𝑥, 𝑦) ∈ [0, 1]2 ,


𝑓(𝑥, 𝑦) = {
0 otherwise.

Find 𝔼[𝑋𝑌 ].
Answer:
Using the theorem, 𝔼[𝑋𝑌 ] = ∫ ∫ 𝑥𝑦 d𝑥 d𝑦 = (∫ 𝑥 d𝑥)(∫ 𝑦 d𝑦) = (1/2)2 = 1/4.
1 1 1 1
0 0 0 0

� Textbook references

If you want more help with this section, check out:

• Sections 4.5 and 5.1 in (Blitzstein and Hwang 2019);


• Section 3.3 in (Anderson, Seppäläinen, and Valkó 2018);
• or Sections 4.5, 5.3, 7.4, and 8.5 in (Stirzaker 2003).

8.3 Linearity of expectation

Remember that summation and integration are linear operators, i.e.,


∑ 𝛼𝑓(𝑥𝑖 ) + 𝛽𝑔(𝑥𝑖 ) = 𝛼 ∑ 𝑓(𝑥𝑖 ) + 𝛽 ∑ 𝑔(𝑥𝑖 ),
𝑖 𝑖 𝑖

and
∫ (𝛼𝑓(𝑥) + 𝛽𝑔(𝑥)) d𝑥 = 𝛼 ∫ 𝑓(𝑥) d𝑥 + 𝛽 ∫ 𝑔(𝑥) d𝑥.
𝐴 𝐴 𝐴
Consequently,

101
� Key idea: Theorem: linearity of expectation 1

For any real-valued random variable 𝑋, and any constants 𝛼 and 𝛽 ∈ ℝ,

𝔼[𝛼𝑋 + 𝛽] = 𝛼𝔼[𝑋] + 𝛽.

A similar, but deeper, result is the following.

� Key idea: Theorem: linearity of expectation 2

For any two real-valued random variables 𝑋 and 𝑌 on the same sample space Ω,

𝔼[𝑋 + 𝑌 ] = 𝔼[𝑋] + 𝔼[𝑌 ].

More generally, for any real-valued random variables 𝑋1 , 𝑋2 , …, 𝑋𝑛 ,


𝑛 𝑛
𝔼 [∑ 𝑋𝑖 ] = ∑ 𝔼[𝑋𝑖 ].
𝑖=1 𝑖=1

Proof

We give the proof in the case where 𝑋 and 𝑌 are discrete. Consider the multiple random variable
(𝑋, 𝑌 ) and the function 𝑔(𝑥, 𝑦) = 𝑥 + 𝑦. By the Law of the Unconscious Statistician we get

𝔼[𝑔(𝑋, 𝑌 )] = ∑ ∑(𝑥 + 𝑦)𝑝(𝑥, 𝑦) = ∑ 𝑥 ∑ 𝑝(𝑥, 𝑦) + ∑ 𝑦 ∑ 𝑝(𝑥, 𝑦)


𝑥∈𝒳 𝑦∈𝒴 𝑥∈𝒳 𝑦∈𝒴 𝑦∈𝒴 𝑥∈𝒳

= ∑ 𝑥𝑝𝑋 (𝑥) + ∑ 𝑦𝑝𝑌 (𝑦) = 𝔼[𝑋] + 𝔼[𝑌 ],


𝑥∈𝒳 𝑦∈𝒴

as claimed. A similar calculation applies in the jointly continuous case, and the extension to more
than two random variables follows by induction.

� Try it out

Suppose that 𝑋 ∼ Bin(𝑛, 𝑝). What is 𝔼[𝑋]?


Answer: We could use the probability mass function and compute
𝑛 𝑛
𝑛
𝔼[𝑋] = ∑ 𝑥𝑝(𝑥) = ∑ ( )𝑥𝑝𝑥 (1 − 𝑝)𝑛−𝑥 ,
𝑥=0 𝑥=0
𝑥

but now some work is needed to evaluate this (exercise!).


Here is a neater way that will also be useful later on. Recall that 𝑋 counts the number of success
on 𝑛 independent trails. If we let 𝑌𝑖 = 1 if trial 𝑖 is a success and 𝑌𝑖 = 0 if trial 𝑖 is a failure, then
in the binomial scenario 𝑌1 , … , 𝑌𝑛 are independent with 𝑃 (𝑌𝑖 = 1) = 𝑝 and 𝑃 (𝑌𝑖 = 0) = 1 − 𝑝. In
other words, we may write
𝑋 = 𝑌 1 + 𝑌2 + ⋯ + 𝑌 𝑛 ,
where 𝑌𝑖 ∼ Bin(1, 𝑝) are independent Bernoulli random variables. Then 𝔼[𝑌𝑖 ] = 𝑝 and so 𝔼[𝑋] =
𝔼[𝑌1 + ⋯ + 𝑌𝑛 ] = 𝑛𝑝. Note that the independence of the trials is not necessary for this result.

102
Advanced content

We have proved linearity of expectation for discrete and jointly continuous random variables, but
linearity of expectation is true for all random variables, and follows from the general measure-theoretic
approach to probability theory, in which expectation is defined as a Lebesgue integral with respect to
a probability measure; this includes the discrete and continuous settings we study as special cases.
Students who take later probability courses will see some of this general approach. Note that without
measure theory it is hard to prove the theorem in the case where say 𝑋 is discrete but 𝑌 is continuous.
If 𝑋 and 𝑌 are independent, and at least one of them is continuous, then the sum 𝑋 + 𝑌 is also
continuous: this is a theorem (see, e.g. (Moran 1968 Theorem 5.9, p.230)) and an example can be
seen in Exercise 6.31, but this may fail without independence.

� Textbook references

If you want more help with this section, check out:

• Section 4.2 in (Blitzstein and Hwang 2019);


• Sections 4.2 and 4.6 in (DeGroot and Schervish 2013);
• or Section 5.3 in (Stirzaker 2003).

8.4 Variance and covariance

As mentioned earlier, we can interpret the expectation of 𝑋 as a long-run average of a sample from
distribution 𝑋. A popular and mathematically convenient way to measure the variability of 𝑋—i.e. to
measure how much 𝑋 varies from 𝔼[𝑋] in the long run—goes via the expectation of the random variable
(𝑋 − 𝔼[𝑋])2 .

� Key idea: Definition: Variance

Let 𝑋 be any real-valued random variable. The variance of 𝑋 is defined as


2
Var (𝑋) ∶= 𝔼 [(𝑋 − 𝔼[𝑋]) ] ,

and the standard deviation of 𝑋 is defined as

𝜎 (𝑋) ∶= √Var (𝑋).

Note that both Var (𝑋) and 𝜎 (𝑋) are non-negative numbers.
Using LOTUS, we can immediately derive the following expressions for the variance:

Var (𝑋) = ∑ (𝑥 − 𝔼[𝑋])2 𝑝(𝑥) if 𝑋 is discrete, and


𝑥∈𝒳

Var (𝑋) = ∫ (𝑥 − 𝔼[𝑋])2 𝑓(𝑥) d𝑥 if 𝑋 is continuously distributed,
−∞

provided that the sum or integral exists.

103
� Try it out

As in a previous example, suppose that 𝑋 takes values 0, 1, 2, 3, 4 each with probability 1/5. What is
Var (𝑋)?
Answer:
First we need to compute 𝔼[𝑋], so
4
0+1+2+3+4 10
𝔼[𝑋] = ∑ 𝑥𝑝(𝑥) = = = 2.
𝑥=0
5 5

Then
4
Var (𝑋) = ∑(𝑥 − 𝔼[𝑋])2 𝑝(𝑥)
𝑥=0
(−2)2 + (−1)2 + 02 + 12 + 22
= = 2.
5

Examples

1. If 𝑋 takes values 0, 10, 20 each with probability 1/3, then

1 1 1
𝔼[𝑋] = 0 × + 10 × + 20 × = 10.
3 3 3
Consequently, using this value,
1 200
Var (𝑋) = × ((0 − 10)2 + (10 − 10)2 + (20 − 10)2 ) = ,
3 3

and so 𝜎 (𝑋) = √ 200


3 ≈ 8.16.

2. Let 𝑍 ∼ 𝒩(0, 1). We know from an earlier example that 𝔼[𝑍] = 0 and by Exercise 8.10,
𝔼[𝑍 2 ] = 1. Consequently,

Var (𝑍) = 𝔼[(𝑍 − 𝐸(𝑍))2 ] = 𝔼[𝑍 2 ] = 1,

and 𝜎 (𝑍) = √Var (𝑍) = 1.

For two real-valued random variables, we can ask ourselves how they vary jointly.

� Key idea: Definition: covariance

Let 𝑋 and 𝑌 be two real-valued random variables on the same sample space. The covariance of 𝑋
and 𝑌 is defined as
Cov (𝑋, 𝑌) ∶= 𝔼[(𝑋 − 𝔼[𝑋])(𝑌 − 𝔼[𝑌 ])].

We also use the following qualitative terminology.

• If Cov (𝑋, 𝑌) > 0 it means that 𝑋 − 𝔼[𝑋] and 𝑌 − 𝔼[𝑌 ] tend to have the same sign. That is, if
𝑋 > 𝔼[𝑋] then it tends to be the case that 𝑌 > 𝔼[𝑌 ] (or, conversely, if 𝑋 < 𝔼[𝑋] then it tends to
be the case that 𝑌 < 𝔼[𝑌 ] too). In this case we say that 𝑋 and 𝑌 are positively correlated.
• If Cov (𝑋, 𝑌) < 0 we say that 𝑋 and 𝑌 are negatively correlated. Now 𝑋 − 𝔼[𝑋] and 𝑌 − 𝔼[𝑌 ] tend

104
to have opposite signs.
• If Cov (𝑋, 𝑌) = 0 we say that 𝑋 and 𝑌 are uncorrelated.

Note: uncorrelated is not the same as independent (more on this later). A quantification of the correlation
is provided by the correlation coefficient, given by

Cov (𝑋, 𝑌)
𝜌(𝑋, 𝑌 ) ∶= .
√Var (𝑋) Var (𝑌)

It can be proved (see Exercises 8.x and 8.x) that

−1 ≤ 𝜌(𝑋, 𝑌 ) ≤ 1.

We will see an example below where the correlation coefficient is 1.


Using LOTUS for multiple random variables, we immediately derive the following expressions for the
covariance:
Cov (𝑋, 𝑌) = ∑ ∑(𝑥 − 𝔼[𝑋])(𝑦 − 𝔼[𝑌 ])𝑝(𝑥, 𝑦)
𝑥∈𝒳 𝑦∈𝒴

if 𝑋 and 𝑌 are discrete, and

Cov (𝑋, 𝑌) = ∬(𝑥 − 𝔼[𝑋])(𝑦 − 𝔼[𝑌 ])𝑓(𝑥, 𝑦) d𝑥 d𝑦,


ℝ2

if 𝑋 and 𝑌 are jointly continuously distributed, provided that the double sum or double integral exists.

� Try it out

Consider discrete random variables 𝑋 and 𝑌 with distribution given by

𝑝(𝑥, 𝑦) 𝑥=1 𝑥=2 𝑝𝑌 (𝑦)


𝑦=1 1/4 0 1/4
𝑦=4 0 3/4 3/4
𝑝𝑋 (𝑥) 1/4 3/4

Find their expectations and their covariance.


Answer:
Then 𝔼[𝑋] = 1 ⋅ 14 + 2 ⋅ 34 = 74 and 𝔼[𝑌 ] = 1 ⋅ 1
4 +4⋅ 3
4 = 4 .
13
We compute, using the formula for
expectation of a function (LOTUS again),

7 13
Cov (𝑋, 𝑌) = 𝔼[(𝑋 − ) (𝑌 − )]
4 4
7 13
= ∑ ∑ (𝑥 − ) (𝑦 − ) 𝑝(𝑥, 𝑦)
𝑥 𝑦
4 4
1 7 13 3 7 13
= (1 − ) (1 − ) + (2 − ) (4 − )
4 4 4 4 4 4
9
= .
16
This means that 𝑋 and 𝑌 are positively correlated, which makes sense from the shape of the table.

105
We note two simple but important properties.

Proposition: Variance as covariance

For any real-valued random variable 𝑋,

Var (𝑋) = Cov (𝑋, 𝑋) .

Proposition: Symmetry of covariance

For any real-valued random variables 𝑋 and 𝑌,

Cov (𝑋, 𝑌) = Cov (𝑌 , 𝑋) .

As immediate consequences of linearity of expectation (see Section 8.3), we obtain the formulæ:

Corollary: Variance and covariance of linear combinations

For any real-valued random variable 𝑋, and any constants 𝛼 and 𝛽 ∈ ℝ,

Var (𝛼 + 𝛽𝑋) = 𝛽 2 Var (𝑋) .


For any real-valued random variables 𝑋, 𝑌, and 𝑍, and any constants 𝛼, 𝛽, 𝛾, and 𝛿 ∈ ℝ,

Cov (𝛼 + 𝛽𝑋, 𝛾 + 𝛿𝑌) = 𝛽𝛿Cov (𝑋, 𝑌) ; (8.6)

Cov (𝑋 + 𝑌 , 𝑍) = Cov (𝑋, 𝑍) + Cov (𝑌 , 𝑍) ; (8.7)


and
Cov (𝑋, 𝑌 + 𝑍) = Cov (𝑋, 𝑌) + Cov (𝑋, 𝑍) . (8.8)

Note: Equation 8.6 through Equation 8.8 mean that Cov () is a bilinear operator.

Proof

We give an example of the type of calculation:

Var (𝛼 + 𝛽𝑋) = 𝔼[(𝛼 + 𝛽𝑋 − 𝔼[𝛼 + 𝛽𝑋]) ]


2

2
= 𝔼[(𝛼 + 𝛽𝑋 − 𝛼 − 𝛽𝔼[𝑋]) ]
= 𝔼[(𝛽𝑋 − 𝛽𝔼[𝑋])2 ]
= 𝛽 2 𝔼[(𝑋 − 𝔼[𝑋])2 ]
= 𝛽 2 Var (𝑋) .

The other statements are similar.

� Try it out

Let 𝑋 ∼ 𝒩(𝜇, 𝜎2 ). Then, 𝑍 ∶= 𝑋−𝜇


𝜎 ∼ 𝒩(0, 1), and as we saw earlier, 𝔼[𝑍] = 0 and Var (𝑍) = 1.
Consequently,
𝔼[𝑋] = 𝔼 [𝜇 + 𝜎𝑍] = 𝜇 + 𝜎𝔼[𝑍] = 𝜇,
Var (𝑋) = 𝜎2 Var (𝑍) = 𝜎2 .

106
In other words, the parameters 𝜇 and 𝜎2 of a normal distribution correspond to the expectation and
variance, respectively.

We also obtain a slightly different way of calculating the variance:

Corollary: Variance and expectation

For any real-valued random variable 𝑋,


2
Var (𝑋) = 𝔼 [𝑋 2 ] − (𝔼[𝑋]) .

Proof

Observe that, by linearity of expectation,

Var (𝑋) = 𝔼 [(𝑋 − 𝔼[𝑋]) ]


2

= 𝔼 [𝑋 2 − 2𝑋𝔼[𝑋] + (𝔼[𝑋])2 ]
= 𝔼 [𝑋 2 ] − 2𝔼[𝑋]𝔼[𝑋] + (𝔼[𝑋])2
= 𝔼 [𝑋 2 ] − (𝔼[𝑋])2 ,

as required.

Example

If 𝑋 takes values 0, 10, 20 each with probability 1/3, then

1 1 1
𝔼[𝑋] = 0 × + 10 × + 20 × = 10,
3 3 3
1 1 1
𝔼 [𝑋 2 ] = 02 × + 102 × + 202 × = 500/3.
3 3 3
Consequently,
Var (𝑋) = 𝔼 [𝑋 2 ] − (𝔼[𝑋])2 = 100 − 500/3 = 200/3,
which agrees with the value that we found earlier.

The next example shows how our various formulae can be put to good use.

� Try it out

Suppose that 𝑋 has probability mass function

𝑥 0 3 10
𝑝(𝑥) 1/4 1/2 1/4

107
Define 𝑌 = 2𝑋 − 6. Find

(a) 𝔼[𝑋], 𝔼 [𝑋 2 ], and Var (𝑋);

(b) 𝔼[𝑌 ] and Var (𝑌);

(c) Cov (𝑋, 𝑌).

Answer:
For (a) we calculate that 𝔼[𝑋] = 3 ⋅ 12 + 10 ⋅ 1
4 = 4 and 𝔼 [𝑋 2 ] = 32 ⋅ 1
2 + 102 ⋅ 1
4 = 2 .
59
So
Var (𝑋) = 𝔼 [𝑋 2 ] − (𝔼[𝑋])2 = 59
2 −4 = 2 .
2 27

For (b), we compute swiftly that

𝔼[𝑌 ] = 𝔼 [2𝑋 − 6] = 2𝔼[𝑋] − 6 = 2,

and
Var (𝑌) = Var (2𝑋 − 6) = 4Var (𝑋) = 54.
Finally, for (c),

Cov (𝑋, 𝑌) = Cov (𝑋, 2𝑋 − 6) = 2Cov (𝑋, 𝑋) = 2Var (𝑋) = 27.

Note that
Cov (𝑋, 𝑌)
𝜌(𝑋, 𝑌 ) = = 1,
√Var (𝑋) Var (𝑌)
which makes sense since 𝑋 and 𝑌 are perfectly positively correlated.

We also obtain a slightly different way of calculating the covariance:

Corollary: Covariance via expectations

For any real-valued random variables 𝑋 and 𝑌,

Cov (𝑋, 𝑌) = 𝔼[𝑋𝑌 ] − 𝔼[𝑋]𝔼[𝑌 ].

Proof

The calculation should now be familiar:


Cov (𝑋, 𝑌) = 𝔼 [(𝑋 − 𝔼[𝑋])(𝑌 − 𝔼[𝑌 ])]
= 𝔼 [𝑋𝑌 − 𝑌 𝔼[𝑋] − 𝑋𝔼[𝑌 ] + 𝔼[𝑋]𝔼[𝑌 ]]
= 𝔼 [𝑋𝑌] − 𝔼[𝑌 ]𝔼[𝑋] − 𝔼[𝑋]𝔼[𝑌 ] + 𝔼[𝑋]𝔼[𝑌 ]
= 𝔼[𝑋𝑌 ] − 𝔼[𝑋]𝔼[𝑌 ],

as required.

108
Examples

We return to the previous example, with joint pmf given by

𝑝(𝑥, 𝑦) 𝑥=1 𝑥=2 𝑝𝑌 (𝑦)


𝑦=1 1/4 0 1/4
𝑦=4 0 3/4 3/4
𝑝𝑋 (𝑥) 1/4 3/4

We already saw that 𝔼[𝑋] = 7/4 and 𝔼[𝑌 ] = 13/4. Now we can compute 𝔼[𝑋𝑌 ] = 1
4 ⋅1+ 3
4 ⋅8= 4 ,
25

so that Cov (𝑋, 𝑌) = 25


4 − 4 ⋅ 4 = 16 , as we obtained before.
7 13 9

� Try it out

Suppose 𝑋 takes values 0, 1, 2 with probabilities 1/4, 1/2, 1/4. Let 𝑌 ∶= 𝑋 2 . What is Var (𝑋),
Var (𝑌), and Cov (𝑋, 𝑌)?
Answer:
First, 𝔼[𝑋] = 1.
Next, 𝑌 = 𝑋 2 takes values 0, 1, 4 with probabilities 1/4, 1/2, 1/4, so 𝔼 [𝑋 2 ] = 𝔼[𝑌 ] = 3/2. Similarly,
𝑌 2 = 𝑋 4 takes values 0, 1, 16 with probabilities 1/4, 1/2, 1/4, so 𝔼 [𝑌 2 ] = 9/2. Finally, 𝑋𝑌 = 𝑋 3
takes values 0, 1, 8 with probabilities 1/4, 1/2, 1/4, so 𝔼[𝑋𝑌 ] = 5/2. Concluding,

Var (𝑋) = 𝔼 [𝑋 2 ] − (𝔼[𝑋])2 = 3/2 − 1 = 1/2;


Var (𝑌) = 𝔼 [𝑌 2 ] − (𝔼[𝑌 ])2 = 9/2 − 9/4 = 9/4;
Cov (𝑋, 𝑌) = 𝔼[𝑋𝑌 ] − 𝔼[𝑋]𝔼[𝑌 ] = 5/2 − 3/2 = 1.

Finally, we can now also say something about the variance of sums of random variables:

Theorem: Variance of a sum

For any real-valued random variables 𝑋 and 𝑌 on the same sample space,

Var (𝑋 + 𝑌) = Var (𝑋) + Var (𝑌) + 2Cov (𝑋, 𝑌) .

More generally, for any real-valued random variables 𝑋1 , 𝑋2 , …, 𝑋𝑛 ,


𝑛 𝑛 𝑛−1 𝑛
Var (∑ 𝑋𝑖 ) = ∑ Var (𝑋𝑖 ) + 2 ∑ ∑ Cov (𝑋𝑖 , 𝑋𝑗 ) .
𝑖=1 𝑖=1 𝑖=1 𝑗=𝑖+1

Proof

For the first statement,

Var (𝑋 + 𝑌) = Cov (𝑋 + 𝑌 , 𝑋 + 𝑌)
= Cov (𝑋, 𝑋 + 𝑌) + Cov (𝑌 , 𝑋 + 𝑌)
= Cov (𝑋, 𝑋) + Cov (𝑋, 𝑌) + Cov (𝑌 , 𝑋) + Cov (𝑌 , 𝑌) ,

109
which gives the result. In general,
𝑛 𝑛 𝑛
Var (∑ 𝑋𝑖 ) = Cov (∑ 𝑋𝑖 , ∑ 𝑋𝑗 )
𝑖=1 𝑖=1 𝑗=1
𝑛 𝑛
= ∑ Cov (𝑋𝑖 , ∑ 𝑋𝑗 )
𝑖=1 𝑗=1
𝑛 𝑛
= ∑ ∑ Cov (𝑋𝑖 , 𝑋𝑗 )
𝑖=1 𝑗=1
𝑛 𝑛
= ∑ Cov (𝑋𝑖 , 𝑋𝑖 ) + ∑ ∑ Cov (𝑋𝑖 , 𝑋𝑗 )
𝑖=1 𝑖=1 𝑗≠𝑖
𝑛 𝑛 𝑖−1 𝑛 𝑛
= ∑ Var (𝑋𝑖 ) + ∑ ∑ Cov (𝑋𝑖 , 𝑋𝑗 ) + ∑ ∑ Cov (𝑋𝑖 , 𝑋𝑗 ) ,
𝑖=1 𝑖=1 𝑗=1 𝑖=1 𝑗=𝑖+1

with the convention that an empty sum is zero.

So, in general, the variance of a sum is not equal to the sum of the variances, unless all covariances are
zero (zero covariance occurs under independence, covered later).
For example, for three real-valued random variables 𝑋, 𝑌, and 𝑍,

Var (𝑋 + 𝑌 + 𝑍) = Var (𝑋) + Var (𝑌) + Var (𝑍) + 2 (Cov (𝑋, 𝑌) + Cov (𝑋, 𝑍) + Cov (𝑌 , 𝑍)) .

� Try it out

Suppose 𝑋 takes values 0, 1, 2 with probabilities 1/4, 1/2, 1/4. Let 𝑌 ∶= 𝑋 2 .


We already found earlier that Var (𝑋) = 1/2, Var (𝑌) = 9/4, and Cov (𝑋, 𝑌) = 1. Now also find
Var (𝑋 + 𝑌) and Var (𝑋 − 𝑌).
Answer:
By the above, Var (𝑋 + 𝑌) = Var (𝑋) + Var (𝑌) + 2Cov (𝑋, 𝑌) = 1/2 + 9/4 + 2 × 1 = 19/4.
(Note that in this very simple example we can easily calculate Var (𝑋 + 𝑌) directly, as 𝑋 + 𝑌 takes
possible values 0, 2, 6 with probabilities 1/4, 1/2, 1/4 so we can confirm by direct calculation that
Var (𝑋 + 𝑌) = 19/4.)
Next, note that Var (𝑋 − 𝑌) = Var (𝑋 + 𝑍) , where 𝑍 = −𝑌. As Var (𝑍) = (−1)2 Var (𝑌) = Var (𝑌)
and Cov (𝑋, 𝑍) = Cov (𝑋, −𝑌) = −Cov (𝑋, 𝑌) we have

Var (𝑋 − 𝑌) = Var (𝑋) + Var (𝑌) − 2Cov (𝑋, 𝑌)


= 1/2 + 9/4 − 2 × 1 = 3/4.

� Try it out

Suppose that 𝑛 people throw their hats high in the air, and each catches a hat which is equally likely
to be any of the 𝑛 hats. Let 𝐻 be the number of people that catch their own hat. Find 𝔼 [𝐻] and
Var (𝐻).
Answer:

110
The distribution of 𝐻 is quite hard to find, but 𝔼 [𝐻] and Var (𝐻) are fairly straightforward. Define

1 if person 𝑖 catches their own hat,


𝑋𝑖 ∶= {
0 otherwise.

Then 𝐻 = ∑𝑖=1 𝑋𝑖 . Here


𝑛

1 𝑛−1 1
𝔼 [𝑋𝑖 ] = ×1+ ×0= ,
𝑛 𝑛 𝑛
and so
𝑛 𝑛
𝔼 [𝐻] = 𝔼 [∑ 𝑋𝑖 ] = ∑ 𝔼 [𝑋𝑖 ] = 1.
𝑖=1 𝑖=1

Similarly, because 𝑋𝑖2 = 𝑋𝑖 ,

1 1 𝑛−1
Var (𝑋𝑖 ) = 𝔼 [𝑋𝑖2 ] − (𝔼 [𝑋𝑖 ])2 = − 2 = .
𝑛 𝑛 𝑛2
We also need Cov (𝑋𝑖 , 𝑋𝑗 ). We compute

𝔼 [𝑋𝑖 𝑋𝑗 ] = 𝑃 (𝑋𝑖 = 1, 𝑋𝑗 = 1)
= 𝑃 (𝑋𝑖 = 1) 𝑃 (𝑋𝑗 = 1 ∣ 𝑋𝑖 = 1)
1 1
= ⋅ ,
𝑛 𝑛−1
so
Cov (𝑋𝑖 , 𝑋𝑗 ) = 𝔼 [𝑋𝑖 𝑋𝑗 ] − 𝔼 [𝑋𝑖 ] 𝔼 [𝑋𝑗 ]
1 1 1
= − 2 = 2 .
𝑛(𝑛 − 1) 𝑛 𝑛 (𝑛 − 1)
Hence
𝑛 𝑛−1 𝑛
Var (𝐻) = ∑ Var (𝑋𝑖 ) + 2 ∑ ∑ Cov (𝑋𝑖 , 𝑋𝑗 )
𝑖=1 𝑖=1 𝑗=𝑖+1
= 𝑛Var (𝑋1 ) + 𝑛(𝑛 − 1)Cov (𝑋1 , 𝑋2 )
𝑛−1 1
=𝑛× + 𝑛(𝑛 − 1) × 2
𝑛2 𝑛 (𝑛 − 1)
𝑛−1+1
= = 1.
𝑛

� Textbook references

If you want more help with this section, check out:

• Sections 4.6 and 7.3 in (Blitzstein and Hwang 2019);


• Section 3.4 in (Anderson, Seppäläinen, and Valkó 2018);
• or Sections 5.3, 5.3, and 8.5 in (Stirzaker 2003).

111
8.5 Conditional expectation

We now turn to the expectation of a random variable given an event (such as the value of another random
variable). Recall that the indicator random variable of an event 𝐴 is given by

1 if 𝜔 ∈ 𝐴,
𝟙{𝐴}(𝜔) = {
0 otherwise.

Then 𝔼 [𝟙{𝐴}] = 1 ⋅ 𝑃 (𝐴) + 0 ⋅ 𝑃 (𝐴c ) = 𝑃 (𝐴).

� Key idea: Definition: conditional expectation

Let 𝑋 be a real-valued random variable, and let 𝐴 ⊆ Ω be an event. The conditional expectation of
𝑋 given 𝐴 is:
𝔼 [𝑋𝟙{𝐴}]
𝔼 [𝑋|𝐴] ∶= whenever 𝑃 (𝐴) > 0.
𝑃 (𝐴)

Conditional expectation generalizes the concept of conditional probability. For example, if 𝐴 and 𝐵
are any events such that 𝑃 (𝐵) > 0 then because 𝟙{𝐴}𝟙{𝐵} = 𝟙{𝐴 ∩ 𝐵}, it follows that 𝔼 [𝟙{𝐴} ∣ 𝐵] =
𝔼 [𝟙{𝐴}𝟙{𝐵}] /𝑃 (𝐵) = 𝑃 (𝐴 ∩ 𝐵) /𝑃 (𝐵) = 𝑃 (𝐴 ∣ 𝐵).
One may also view 𝔼 [ ⋅ |𝐴] as expectation with respect to the conditional probability 𝑃 ( ⋅ ∣ 𝐴):

� Key idea: Theorem: conditional expectation and probabilities

For any discrete random variable 𝑋 and any event 𝐴 ⊆ Ω,

𝔼 [𝑋|𝐴] = ∑ 𝑥𝑃 (𝑋 = 𝑥 ∣ 𝐴) .
𝑥∈𝒳

Proof

This is an exercise in tracking definitions. Indeed, if 𝑌 = 𝑋𝟙{𝐴}, then for any 𝑥 ≠ 0,

𝑃 (𝑌 = 𝑥) = 𝑃 ({𝑋 = 𝑥} ∩ 𝐴) = 𝑃 (𝑋 = 𝑥 ∣ 𝐴) 𝑃 (𝐴) .

Hence
𝔼[𝑌 ] = ∑ 𝑥𝑃 (𝑌 = 𝑥)
𝑥∈𝒳
= ∑ 𝑥𝑃 (𝑌 = 𝑥)
𝑥∈𝒳, 𝑥≠0

= ∑ 𝑥𝑃 (𝑋 = 𝑥 ∣ 𝐴) 𝑃 (𝐴) ,
𝑥∈𝒳

and so
𝔼[𝑌 ]
𝔼 [𝑋|𝐴] = = ∑ 𝑥𝑃 (𝑋 = 𝑥 ∣ 𝐴) ,
𝑃 (𝐴) 𝑥∈𝒳
as claimed.

112
� Try it out

Suppose that 𝑋 and 𝑌 are discrete random variables, 𝑔 ∶ 𝒳 → ℝ, and 𝐵 ⊆ 𝒴. Then choosing the
event 𝐴 = {𝑌 ∈ 𝐵}, we see that whenever 𝑃 (𝑌 ∈ 𝐵) = ∑𝑦∈𝐵 𝑝𝑌 (𝑦) > 0,

∑𝑥∈𝒳 𝑔(𝑥) ∑𝑦∈𝐵 𝑝𝑋,𝑌 (𝑥, 𝑦)


𝔼 [𝑔(𝑋)|𝑌 ∈ 𝐵] = .
∑𝑦∈𝐵 𝑝𝑌 (𝑦)

As a special case, we have


𝔼 [𝑔(𝑋)|𝑌 = 𝑦] = ∑ 𝑔(𝑥)𝑝𝑋|𝑌 (𝑥|𝑦).
𝑥∈𝒳

Similarly, if 𝑋 and 𝑌 are jointly continuously distributed, then

∫ 𝑔(𝑥) (∫ 𝑓𝑋,𝑌 (𝑥, 𝑦) d𝑦) d𝑥


ℝ 𝐵
𝔼 [𝑔(𝑋)|𝑌 ∈ 𝐵] = ,
∫ 𝑓𝑌 (𝑦) d𝑦
𝐵

provided 𝑃 (𝑌 ∈ 𝐵) = ∫ 𝑓𝑌 (𝑦) d𝑦 > 0.


𝐵
To summarize, conditional expectation is just like ordinary expectation but with probabilities replaced
by conditional probabilities.

� Try it out

In a raffle there is one £500 prize and five £100 prizes. We have one of the 2000 raffle tickets. Let 𝑋
be our winnings and let 𝐴 be the event that we have the top prize.
(a) Calculate 𝔼 [𝑋 ∣ 𝐴] and 𝔼 [𝑋 ∣ 𝐴c ].
(b) Compute 𝔼 [𝑋|𝐴] 𝑃 (𝐴) + 𝔼 [𝑋|𝐴c ] 𝑃 (𝐴c ).
(c) Compute 𝔼[𝑋].
Answer:
By counting we see that 𝑃 (𝑋 = 500 ∣ 𝐴) = 1 and
𝑃 (𝑋 = 500 ∣ 𝐴c ) = 0,
5
𝑃 (𝑋 = 100 ∣ 𝐴c ) = ,
1999
1994
𝑃 (𝑋 = 0 ∣ 𝐴c ) = .
1999
So we get 𝔼 [𝑋|𝐴] = 500 and
5 500
𝔼 [𝑋|𝐴c ] = 100 ⋅ +0= .
1999 1999
Now for(b), we note that 𝑃 (𝐴) = 1/2000 and 𝑃 (𝐴c ) = 1999/2000, so
1 500 1999 1000 1
𝔼 [𝑋|𝐴] 𝑃 (𝐴) + 𝔼 [𝑋|𝐴c ] 𝑃 (𝐴c ) = 500 ⋅ + ⋅ = = .
2000 1999 2000 2000 2
For (c), we compute directly that
1 5 1
𝔼[𝑋] = 500 ⋅ + 100 ⋅ +0= .
2000 2000 2
So the answers to (b) and (c) are the same.

113
It was no coincidence that the answers to parts (b) and (c) of were the same. This is an example of the
following important theorem.

� Key idea: Theorem: partition theorem for expectation

Let 𝑋 be any real-valued random variable. Let 𝐸1 , 𝐸2 , …, 𝐸𝑘 be any events that form a partition.
Then,
𝑘
𝔼[𝑋] = ∑ 𝔼 [𝑋|𝐸𝑖 ] 𝑃 (𝐸𝑖 ) .
𝑖=1
Similarly, if 𝐸1 , 𝐸2 , ... form an infinite partition

𝔼[𝑋] = ∑ 𝔼 [𝑋|𝐸𝑖 ] 𝑃 (𝐸𝑖 ) .
𝑖=1

Proof

Because 𝐸1 , 𝐸2 , … constitute a partition, ∪𝑖 𝐸𝑖 = Ω and the 𝐸𝑖 are pairwise disjoint, so that

1 = 𝟙{Ω} = 𝟙{∪𝑖 𝐸𝑖 } = ∑ 𝟙{𝐸𝑖 },


𝑖

and hence, by linearity of expectation,

𝔼[𝑋] = 𝔼 [𝑋 ∑ 𝟙{𝐸𝑖 }] = ∑ 𝔼 [𝑋𝟙{𝐸𝑖 }] ,


𝑖 𝑖

which gives the result.

� Try it out

Ann and Bob play a sequence of independent games, each with outcomes 𝐴 = {Ann wins}, 𝐵 =
{Bob wins}, 𝐷 = {game drawn} which have probabilities 𝑃 (𝐴) = 𝑝, 𝑃 (𝐵) = 𝑞 and 𝑃 (𝐷) = 𝑟 > 0
with 𝑝 + 𝑞 + 𝑟 = 1. Let 𝑁 denote the number of games played until the first win by either Ann
or Bob. The results of the first game 𝐴, 𝐵, 𝐷 form a partition and 𝔼 [𝑁 |𝐴] = 𝔼 [𝑁 |𝐵] = 1 (as
𝑃 (𝑁 = 1 ∣ 𝐴) = 𝑃 (𝑁 = 1 ∣ 𝐵) = 1) while 𝔼 [𝑁 |𝐷] = 1 + 𝔼 [𝑁] (the future after a drawn game looks
the same as at the start). Hence 𝔼 [𝑁] = 1 × 𝑝 + 1 × 𝑞 + (1 + 𝔼 [𝑁]) × 𝑟 or (1 − 𝑟)𝔼 [𝑁] = 𝑝 + 𝑞 + 𝑟 = 1
i.e. 𝔼 [𝑁] = 1/(1 − 𝑟).

A more demanding concept is the following:

� Key idea: Definition: conditional expectation with respect to a random variable

Let 𝑋 be a real-valued random variable, and let 𝑌 be another random variable. Define the function
𝑔 ∶ 𝑌 (Ω) → ℝ by 𝑔(𝑦) ∶= 𝔼 [𝑋|𝑌 = 𝑦]. Then the conditional expectation of 𝑋 given 𝑌 is the random
variable denoted by 𝔼 [𝑋|𝑌] given by
𝔼 [𝑋|𝑌] ∶= 𝑔(𝑌 ).

You may see this written more compactly in books as

𝔼 [𝑋|𝑌] (𝜔) ∶= 𝔼 [𝑋|𝑌 = 𝑌 (𝜔)] , for all 𝜔 ∈ Ω,

but our definition above is a little easier to digest. In the case where 𝑌 is discrete, 𝔼 [𝑋|𝑌] is a random

114
variable that takes values 𝔼 [𝑋|𝑌 = 𝑦] with probabilities 𝑃 (𝑌 = 𝑦). We concentrate on the discrete case
here.
The next result is of considerable importance.

� Key idea: Theorem: partition theorem for expectation II

For any real-valued random variable 𝑋 and any random variable 𝑌,

𝔼[𝑋] = 𝔼 [𝔼 [𝑋 ∣ 𝑌]] .

Proof

We give a proof in the case when 𝑌 is discrete, which shows why we call this a ‘Partition theorem’.
In this case {𝑌 = 𝑦}, 𝑦 ∈ 𝒴, forms a partition, and so gives

𝔼[𝑋] = ∑ 𝔼 [𝑋|𝑌 = 𝑦] 𝑃 (𝑌 = 𝑦) ,
𝑦∈𝒴

but this last expression is the expectation of the discrete random variable 𝔼 [𝑋|𝑌] which takes values
𝑔(𝑦) = 𝔼 [𝑋|𝑌 = 𝑦] with probabilities 𝑃 (𝑌 = 𝑦):

𝔼 [𝔼 [𝑋 ∣ 𝑌]] = 𝔼 [𝑔(𝑌 )] = ∑ 𝑔(𝑦)𝑃 (𝑌 = 𝑦) ,


𝑦∈𝒴

by the law of the unconscious statistician.

This theorem is sometimes called the law of iterated expectation. We can use this result to calculate 𝔼[𝑋]
in some tricky cases.

� Try it out

Toss three fair coins and let 𝐻 be the total number of heads. The roll a fair die 𝐻 times, and let 𝑇
be the total score on the dice rolls. What is 𝔼 [𝑇]?
Answer:
This looks tricky, but we use conditioning on 𝐻 to take advantage of the structure of the problem.
Note that 𝐻 ∼ Bin(3, 1/2), so and 𝔼 [𝐻] = 3/2. Now, the expected score on a single die is

1 7
(1 + 2 + 3 + 4 + 5 + 6) = .
6 2
Thus, given ℎ = ℎ, the expected total on ℎ rolls of a fair die the expected total score is 72 ℎ. In other
words, 𝔼 [𝑇 |𝐻 = ℎ] = 72 ℎ. Thus
7
𝔼 [𝑇 |𝐻] = 𝐻.
2
Thus, by ,
7 7 3 21
𝔼 [𝑇] = 𝔼 [𝔼 [𝑇 ∣ 𝐻]] = 𝔼 [𝐻] = × = .
2 2 2 4
Note that while this looks very slick, we have actually hidden something here. In fact, we have
implicitly used the fact that the scores rolled on the die are independent of 𝐻, when we claimed that
𝔼 [𝑇 |𝐻 = ℎ] = 72 ℎ. To see where this is being used, let 𝑆1 , 𝑆2 , … be the scores on the die rolls, so

115
𝑇 = ∑𝐻 𝑆 . Then
𝑖=1 𝑖

𝐻 ℎ
𝔼 [𝑇 |𝐻 = ℎ] = 𝔼 [∑ 𝑆𝑖 |𝐻 = ℎ] = 𝔼 [∑ 𝑆𝑖 |𝐻 = ℎ] ,
𝑖=1 𝑖=1

and we can drop the condition here if 𝐻 is independent of the 𝑆𝑖 . Then


ℎ ℎ
7
𝔼 [𝑇 ∣ 𝐻 = ℎ] = 𝔼 [∑ 𝑆𝑖 ] = ∑ 𝔼 [𝑆𝑖 ] = ℎ,
𝑖=1 𝑖=1
2

as claimed. The next example is of the same type.

� Try it out

A shop has 𝑁 customers a day where 𝔼 [𝑁] = 800. Each customer spends £𝑋𝑖 where 𝔼 [𝑋𝑖 ] = 25.
Let 𝑇 = ∑𝑁𝑖=1
𝑋𝑖 be the total takings on a particular day. What is 𝔼 [𝑇]?
Answer:
We condition on the value of 𝑁. Then
𝑛
𝔼 [𝑇 |𝑁 = 𝑛] = 𝔼 [∑ 𝑋𝑖 |𝑁 = 𝑛] .
𝑖=1

This is similar to the last example, and we must here use the fact that the 𝑋𝑖 are independent of 𝑁
to get
𝑛 𝑛 𝑛
𝔼 [∑ 𝑋𝑖 |𝑁 = 𝑛] = 𝔼 [∑ 𝑋𝑖 ] = ∑ 𝔼 [𝑋𝑖 ] = 25𝑛.
𝑖=1 𝑖=1 𝑖=1

Thus 𝔼 [𝑇 |𝑁] = 25𝑁 and

𝔼 [𝑇] = 𝔼 [𝔼 [𝑇 ∣ 𝑁]] = 𝔼 [25𝑁] = 25𝔼 [𝑁] = 25 × 800 = 20, 000.

In this example the assumption of independence between 𝑁 and the 𝑋𝑖 is open to question: if 𝑁 is
very big, perhaps the 𝑋𝑖 might be smaller than usual, since the shop runs low on stock, for example.

� Textbook references

If you want more help with this section, check out:

• Sections 9.1 and 9.2 in (Blitzstein and Hwang 2019);


• Section 10.3 in (Anderson, Seppäläinen, and Valkó 2018);
• or Sections 4.4 and 5.5 in (Stirzaker 2003).

8.6 Independence: multiplication rule for expectation

Remember that two discrete random variables are independent when their joint probability mass function
factorises, i.e. when 𝑝(𝑥, 𝑦) = 𝑝(𝑥)𝑝(𝑦). Similarly, two jointly continuously distributed random variables
are independent when their joint probability density function factorizes, i.e. when 𝑓(𝑥, 𝑦) = 𝑓(𝑥)𝑓(𝑦). It

116
turns out that in these cases, expectation factorizes as well:

� Key idea: Theorem: Independence means multiply

If 𝑋 and 𝑌 are independent real-valued random variables then

𝔼[𝑋𝑌 ] = 𝔼[𝑋]𝔼[𝑌 ].

Moreover, if 𝑔, ℎ ∶ ℝ → ℝ,
𝔼 [𝑔(𝑋)ℎ(𝑌 )] = 𝔼[𝑔(𝑋)]𝔼 [ℎ(𝑌 )] .
More generally, for any mutually independent real-valued random variables 𝑋1 , 𝑋2 , …, 𝑋𝑛 ,
𝑛 𝑛
𝔼 [∏ 𝑋𝑖 ] = ∏ 𝔼 [𝑋𝑖 ] .
𝑖=1 𝑖=1

Proof

We give a proof only in the discrete case. Suppose that 𝑋 and 𝑌 are independent discrete random
variables with joint probability mass function 𝑝(𝑥, 𝑦) = 𝑝𝑋 (𝑥)𝑝𝑌 (𝑦). Then by the Law of the
Unconscious Statistician,

𝔼 [𝑔(𝑋)ℎ(𝑌 )] = ∑ ∑ 𝑔(𝑥)ℎ(𝑦)𝑝(𝑥, 𝑦) = ∑ 𝑔(𝑥)𝑝𝑋 (𝑥) ∑ ℎ(𝑦)𝑝𝑌 (𝑦),


𝑥∈𝒳 𝑦∈𝒴 𝑥∈𝒳 𝑦∈𝒴

which is 𝔼[𝑔(𝑋)]𝔼 [ℎ(𝑌 )].

Corollary: Independence means zero covariance

If 𝑋 and 𝑌 are independent random variables, then Cov (𝑋, 𝑌) = 0.

Proof

In the independent case, Cov (𝑋, 𝑌) = 𝔼[𝑋𝑌 ] − 𝔼[𝑋]𝔼[𝑌 ] = 0.

The converse of the last property is not true: 𝑋 and 𝑌 may be dependent but uncorrelated.

Example

Suppose that (𝑋, 𝑌 ) are jointly distributed taking values (−1, 0), (+1, 0), (0, −1), and (0, +1) with
probability 1/4 of each.
Then 𝑋 and 𝑌 are not independent, because 𝑃 (𝑋 = 1, 𝑌 = 1) = 0 is not the same as
𝑃 (𝑋 = 1) 𝑃 (𝑌 = 1) = 1/16, for instance.
However, 𝑋 and 𝑌 are uncorrelated, because 𝔼[𝑋𝑌 ] = 0 and 𝔼[𝑋] = 𝔼[𝑌 ] = 0 so Cov (𝑋, 𝑌) = 0. (In
fact, in this case 𝑋𝑌 = 0 with probability 1.)

An important consequence of this Corollary is a simplification of the formula for the variance of a sum for
pairwise independent random variables:

117
Corollary: Variance of a sum of independent variables

Consider random variables 𝑋1 , 𝑋2 , …, 𝑋𝑛 . If these random variables are pairwise independent, then
𝑛 𝑛
Var (∑ 𝑋𝑖 ) = ∑ Var (𝑋𝑖 ) .
𝑖=1 𝑖=1

Example

Remember from an earlier example that we can write any 𝑋 ∼ Bin(𝑛, 𝑝) as 𝑋 = ∑1 𝑌𝑖 where 𝑌1 , …,
𝑛

𝑌𝑛 are independent and each 𝑌𝑖 ∼ Bin(1, 𝑝). In , you showed that Var (𝑌𝑖 ) = 𝑝(1 − 𝑝). Consequently,
as the 𝑌𝑖 are independent,
𝑛
Var (𝑋) = ∑ Var (𝑌𝑖 ) = 𝑛𝑝(1 − 𝑝).
𝑖=1

� Try it out

Let 𝑆 be the total score and 𝑇 be the product of the scores from throwing four dice.
To find 𝔼 [𝑆], Var (𝑆) and 𝔼 [𝑇] let the individual scores be 𝑋𝑘 , 𝑘 = 1, 2, 3, 4 so that 𝑆 = ∑𝑘=1 𝑋𝑘 ,
4

𝑇 = ∏𝑘=1 𝑋𝑘 .
4

We readily calculate 𝔼 [𝑋𝑘 ] = ∑𝑖=1 𝑖 × 1/6 = 7/2 and Var (𝑋𝑘 ) = ∑𝑖=1 𝑖2 × 1/6 − (7/2)2 = 35/12.
6 6

Therefore
4 4
𝔼 [𝑆] = 𝔼 [∑ 𝑋𝑘 ] = ∑ 𝔼 [𝑋𝑘 ] = 4 × 7/2 = 14,
𝑘=1 𝑘=1

and further, as the 𝑋𝑘 are independent,


4 4
Var (𝑆) = Var (∑ 𝑋𝑘 ) = ∑ Var (𝑋𝑘 ) = 4 × 35/12 = 35/3.
𝑘=1 𝑘=1

Again as the 𝑋𝑘 are independent,


4 4
7 4
𝔼 [𝑇] = 𝔼 [ ∏ 𝑋𝑘 ] = ∏ 𝔼 [𝑋𝑘 ] = ( ) = 2401/16.
𝑘=1 𝑘=1
2

Note that all of these calculations are possible without having to deal with the joint probability
distribution of the 𝑋𝑘 (which is uniform on the 1296 possible outcomes).

� Textbook references

If you want more help with this section, check out:

• Section7.3 in (Blitzstein and Hwang 2019);


• Section 8.2 in (Anderson, Seppäläinen, and Valkó 2018);
• or Section 5.3 in (Stirzaker 2003).

118
8.7 Expectation and probability inequalities

By the monotonicity properties of summation and integration, namely that if 𝑓(𝑥) ≥ 𝑔(𝑥) then

∑ 𝑓(𝑥𝑖 ) ≥ ∑ 𝑔(𝑥𝑖 ), and ∫ 𝑓(𝑥) d𝑥 ≥ ∫ 𝑔(𝑥) d𝑥,


𝑖 𝑖 𝐴 𝐴

we immediately get the following.

Theorem: Monotonicity of expectation

For any random variable 𝑋, and any 𝑎 ∈ ℝ, if 𝑃 (𝑋 ≥ 𝑎) = 1 then 𝔼[𝑋] ≥ 𝑎.

For instance, suppose that 𝑋 and 𝑌 have 𝑃 (𝑋 ≤ 𝑌) = 1. Then 𝑃 (𝑌 − 𝑋 ≥ 0) = 1 so 𝔼 [𝑌 − 𝑋] ≥ 0 and


hence 𝔼[𝑋] ≤ 𝔼[𝑌 ].
This simple property has various interesting consequences:

Corollary: Variances are positive

For any random variable 𝑋, Var (𝑋) ≥ 0.

Proof

We have Var (𝑋) = 𝔼 [(𝑋 − 𝔼[𝑋])2 ] and the random variable (𝑋 − 𝔼[𝑋])2 is non-negative.

� Key idea: Corollary: Markov’s inequality

If 𝑋 ≥ 0 then, for any 𝑎 > 0,


𝔼[𝑋]
𝑃 (𝑋 ≥ 𝑎) ≤ .
𝑎

Proof

Note that 𝑋 ≥ 𝑎𝟙{𝑋 ≥ 𝑎}, and consequently 0 ≤ 𝔼 [𝑋 − 𝑎𝟙{𝑋 ≥ 𝑎}] = 𝔼[𝑋] − 𝑎𝑃 (𝑋 ≥ 𝑎).

Example

If 𝑋 equals 𝑠 with chance 𝑝 but otherwise equals 0 then 𝔼[𝑋] = 𝑠𝑝. For 𝑎 > 𝑠 we have 0 =
𝑃 (𝑋 ≥ 𝑎) ≤ 𝑝𝑠/𝑎 while for 𝑎 ≤ 𝑠 Markov’s inequality says 𝑝 = 𝑃 (𝑋 ≥ 𝑎) ≤ 𝑝 × 𝑠/𝑎 which is exact
at 𝑎 = 𝑠 so this bound is as strong as possible.

Corollary: Chebyshev’s inequality

For any random variable 𝑋 and any 𝑎 > 0, we have

Var (𝑋)
𝑃 (|𝑋 − 𝔼[𝑋]| ≥ 𝑎) ≤ .
𝑎2

119
Proof

𝑃 ((𝑋 − 𝔼[𝑋])2 ≥ 𝑎2 ) ≤ 𝔼 [(𝑋 − 𝔼[𝑋])2 ] /𝑎2 by Markov’s inequality applied to (𝑋 − 𝔼[𝑋])2 at 𝑎2 .


Now observe that {(𝑋 − 𝔼[𝑋])2 ≥ 𝑎2 } = {|𝑋 − 𝔼[𝑋]| ≥ 𝑎} and of course 𝔼[(𝑋 − 𝔼[𝑋])2 ] = Var (𝑋).

Markov and Chebyshev bounds are often too generous when distributional information is available, as seen
in the next examples. Nevertheless, their generality and simplicity make these inequalities very valuable
for complex probability calculations.

� Try it out

Suppose that 𝑋 ∼ Bin(10, 0.1). Give an upper bound on 𝑃 (𝑋 ≥ 6) using (a) Markov’s inequality,
and (b) Chebyshev’s inequality.
Answer:
For part (a), we get
𝔼[𝑋] 1
𝑃 (𝑋 ≥ 6) ≤ = .
6 6
For (b), we get
𝑃 (𝑋 ≥ 6) ≤ 𝑃 (|𝑋 − 1| ≥ 5)
= 𝑃 (|𝑋 − 𝔼[𝑋]| ≥ 5)
Var (𝑋) 0.9
≤ = = 0.036.
25 25
The exact probability can be calculated and is 𝑃 (𝑋 ≥ 6) ≈ 0.00015.

� Try it out

If 𝑍 ∼ 𝒩(0, 1) then 𝑃 (|𝑍 − 𝔼[𝑍]| ≥ 2) = 𝑃 (|𝑍| ≥ 2) = 1 − 𝑃 (−2 < 𝑍 < 2) = 0.046, while the
Chebyshev bound on this probability is Var (𝑍) /22 = 0.25.

� Textbook references

If you want more help with this section, check out:

• Section 10.1 in (Blitzstein and Hwang 2019);


• Section 4. in (DeGroot and Schervish 2013);
• or Section 4.6 in (Stirzaker 2003).

8.8 Historical context

There are approaches to probability theory that start out from expectation of random variables directly,
rather than starting out from probability of events as we have done here; see e.g. (Whittle 1992).
Pafnuty Chebyshev (1821–1894) and his student Andrei Markov (1856–1922) made several important
contributions to early probability theory. What we call Markov’s inequality was actually published by
Chebyshev, as was what we call Chebyshev’s inequality; our nomenclature is standard, and at least has
the benefit of distinguishing the two. A version of the inequality was first formulated by Irénée-Jules
Bienaymé (1796–1878).

120
(b) Markov
(a) Chebyshev

Figure 8.1: (left to right) Chebyshev and Markov

121
9 Limit theorems

� Goals

1. Understand and know how to prove the weak law of large numbers, for proportions as well as
for general random variables, and know under what conditions the weak law applies.

2. Understand and know how to prove (by means of moment generating functions) the central
limit theorem.

3. Know under what conditions the central limit theorem applies.

4. Know how to exploit the central limit theorem to approximate the binomial distribution, and
under what circumstances.

5. Know the definition and properties of moment generating functions.

6. Know how to derive the moment generating function of a given distribution.

9.1 The weak law of large numbers

The results of this section describe limiting properties of distributions of sums of random variables using
only some assumptions about means and variances.
Toss a coin 𝑛 times, where the probability of heads is 𝑝, independently on each toss. Let 𝑋 be the number
of heads and let 𝐵𝑛 ∶= 𝑋/𝑛 be the proportion of heads in the 𝑛 tosses. Then, because 𝑋 ∼ Binom(𝑛, 𝑝)
has expectation 𝑛𝑝 and variance 𝑛𝑝(1 − 𝑝),

𝔼 [𝐵𝑛 ] = 𝔼 [𝑋] /𝑛 = 𝑝, Var (𝐵𝑛 ) = Var (𝑋) /𝑛2 = 𝑝(1 − 𝑝)/𝑛.

So, by Chebyshev’s inequality, for any 𝜖 > 0,

𝑝(1 − 𝑝)
𝑃 (|𝐵𝑛 − 𝑝| ≥ 𝜖) ≤ .
𝑛𝜖2
Whence,
𝑃 (|𝐵𝑛 − 𝑝| ≥ 𝜖) → 0 as 𝑛 → ∞,
no matter how small 𝜖 is. In other words, as 𝑛 → ∞, the sample proportion is with very high probability
within any tiny interval centred on 𝑝.
The same argument applies more generally.

122
� Key idea: Theorem: the weak law of large numbers

Suppose we have an infinite sequence 𝑋1 , 𝑋2 , … of independent random variables with the same
mean and variance:
𝔼 [𝑋𝑖 ] = 𝜇 and Var (𝑋𝑖 ) = 𝜎2 for all 𝑖.
Consider the sample average 𝑋̄ 𝑛 ∶= ∑𝑖=1 𝑋𝑖 . Then, for any 𝜖 > 0,
1 𝑛
𝑛

lim 𝑃 (|𝑋̄ 𝑛 − 𝜇| > 𝜖) = 0. (9.1)


𝑛→∞

In other words, the sample average has a very high probability of being very near the expected value 𝜇
when 𝑛 is large. The type of convergence in Equation 9.1 is called convergence in probability: the weak
law of large numbers says that “𝑋̄ 𝑛 converges in probability to 𝜇”.

Proof

We use Chebyshev’s inequality to bound the probability that we are trying to show is small.
First note that, by linearity of expectation, 𝔼 [𝑋̄ 𝑛 ] = 𝑛1 ∑𝑖=1 𝔼 [𝑋𝑖 ] = 𝜇 and, by independence,
Var (𝑋̄ 𝑛 ) = 12 ∑ Var (𝑋𝑖 ) = 𝜎 . So, by Chebyshev’s inequality,
𝑛 2
𝑛 𝑖=1 𝑛

𝜎2
𝑃(|𝑋̄ 𝑛 − 𝜇| ≥ 𝜖) ≤ 2 ,
𝑛𝜖
which indeed converges to zero as 𝑛 tends to infinity.

Advanced content

The assumption of finite variances in this theorem is not necessary. The weak law of large numbers
holds assuming only that 𝔼 [𝑋𝑖 ] = 𝜇 for independent 𝑋𝑖 : for this and other more advanced results,
see (Feller 1968, chap. 10).

Examples

1. Measure the heights 𝐻𝑖 of 𝑛 randomly selected people from a very large population, where 𝜇,
𝜎2 are the average and the variance of heights over the whole population.
Then 𝔼 [𝐻𝑖 ] = 𝜇, Var (𝐻𝑖 ) = 𝜎2 and so, as long as the collection of heights is not too asymmetric,
the chance of the average height 𝑋̄ 𝑛 being more than a small amount from 𝜇 is very small;
i.e. almost all large samples have average near 𝜇.

2. Here is an application to repeated sampling. Let 𝑋 be a random variable whose distribution


we want to study by ‘sampling’, i.e., observing a number of independent random variables
𝑋1 , 𝑋2 , … , 𝑋𝑛 which all have the same distribution as 𝑋. We could observe 𝑋1 , 𝑋2 , … , 𝑋𝑛
and consider
1 𝑛
𝜋𝑛 (𝑎, 𝑏) = ∑ 𝟙{𝑋𝑖 ∈ [𝑎, 𝑏)},
𝑛 𝑖=1
the proportion of observations whose value falls in the interval [𝑎, 𝑏). Since the 𝑋𝑖 are
independent, so are the indicator random variables, and we know that 𝔼 [𝟙{𝑋𝑖 ∈ [𝑎, 𝑏)}] =
𝑃 (𝑋𝑖 ∈ [𝑎, 𝑏)). So the law of large numbers says that 𝜋𝑛 (𝑎, 𝑏) will approach 𝑃 (𝑋𝑖 ∈ [𝑎, 𝑏)) for
large 𝑛 with high probability. Another way to see this is to observe that ∑𝑖=1 𝟙{𝑋𝑖 ∈ [𝑎, 𝑏)} is
𝑛

a binomial random variable.

123
If we want to look at the distribution of 𝑋, we would construct a histogram using the proportions
𝜋𝑛 (𝑎𝑖 , 𝑏𝑖 ) over a collection of ‘bins’ [𝑎𝑖 , 𝑏𝑖 ). If there are only finitely many bins, then it follows
that all the 𝜋𝑛 ’s tend to be close to their respective probabilities. For example, the picture in
Figure 9.1 shows a histogram produced by 104 simulations of a 𝑈 (0, 1) random variable: the
fact that the histogram is a good approximation to the probability density function can, in this
instance, be seen as a consequence of the law of large numbers.

Figure 9.1: Histogram generated from 104 simulations of a U(0, 1) random variable. The vertical axis is
the frequency.

� Textbook references

If you want more help with this section, check out:

• Section 10.2 in (Blitzstein and Hwang 2019);


• Section 9.2 in (Anderson, Seppäläinen, and Valkó 2018);
• or Section 5.8 in (Stirzaker 2003).

9.2 The central limit theorem

If the law of large numbers is a ‘first order’ result, a ‘second order result’ is the famous central limit theorem,
which describes fluctuations around the law of large numbers, and explains, in part, why the normal
distribution has a central role in statistics. A sequence of random variables 𝑋1 , 𝑋2 , … are independent and
identically distributed (i.i.d. for short) if they are mutually independent and all have the same (marginal)
distribution.

� Key idea: Theorem: the Central Limit Theorem

Suppose we have a sequence 𝑋1 , 𝑋2 , …of i.i.d. random variables. Let

𝜇 ∶= 𝔼 [𝑋𝑖 ] and 𝜎2 ∶= Var (𝑋𝑖 ) ,

124
with 𝜎 > 0. Let 𝑆𝑛 ∶= ∑𝑖=1 𝑋𝑖 , 𝑋̄ 𝑛 ∶= 𝑆𝑛 /𝑛, and
𝑛

𝑆𝑛 − 𝑛𝜇 𝑋̄ − 𝜇
𝑍𝑛 ∶= √ = 𝑛√
𝜎 𝑛 𝜎/ 𝑛

Then, for any 𝑧 ∈ ℝ,


lim 𝐹𝑍𝑛 (𝑧) = lim 𝑃 (𝑍𝑛 ≤ 𝑧) = Φ(𝑧).
𝑛→∞ 𝑛→∞
We say that 𝑍𝑛 converges in distribution to the standard normal distribution.

In other words, for large 𝑛, we have that approximately 𝑍𝑛 ≈ 𝒩(0, 1). Note that the definition of 𝑍𝑛
is such that 𝔼 [𝑍𝑛 ] = 0 and Var (𝑍𝑛 ) = 1. The content of the central limit theorem is that it should be
approximately normal. Consequently, if we also invoke , we approximately have that 𝑆𝑛 ≈ 𝒩(𝑛𝜇, 𝑛𝜎2 )

and 𝑋̄ 𝑛 ≈ 𝒩(𝜇, 𝜎2 /𝑛), i.e., typical values of 𝑆𝑛 are of order 𝜎 𝑛 from 𝑛𝜇 while typical values of 𝑋̄ 𝑛 are

of order 𝜎/ 𝑛 from 𝜇. In other words, for large enough 𝑛, by ,

𝑥 − 𝑛𝜇
𝑃 (𝑆𝑛 ≤ 𝑥) ≈ Φ ( √ ),
𝜎 𝑛
𝑥−𝜇
𝑃 (𝑋̄ 𝑛 ≤ 𝑥) ≈ Φ ( √ ) .
𝜎/ 𝑛

Example

Suppose 𝑋1 , 𝑋2 , … are i.i.d. exponential random variables with parameter 1, i.e., they have probability
density function 𝑓(𝑥) = 𝑒−𝑥 for 𝑥 > 0 and 𝑓(𝑥) = 0 otherwise.
Consider 𝑆𝑛 = ∑𝑖=1 𝑋𝑖 . The central limit theorem says that 𝑆𝑛 will be approximately normal
𝑛

for large 𝑛, and since we know 𝔼 [𝑋𝑖 ] = 1 and Var (𝑋𝑖 ) = 1 in this case, 𝑆𝑛 will be approximately
𝒩(𝑛, 𝑛) for large 𝑛.
How “large” should 𝑛 be? Well, in this case it turns out we can compute the distribution of 𝑆𝑛
exactly. It is an example of a gamma distribution, and for 𝑛 ≥ 1, 𝑆𝑛 has probability density function

⎧ 𝑒−𝑥 𝑥𝑛−1
{ if 𝑥 > 0,
𝑓𝑛 (𝑥) = (𝑛 − 1)!

{0
⎩ elsewhere.

See Figure 9.2 for some plots of this for various values of 𝑛.

Advanced content

The conditions of the central limit theorem can be substantially weakened.


The assumption that the 𝑋𝑖 are identically distributed can be dropped and replaced with the
Lindeberg condition: see e.g. (Feller 1971, chap. VIII).
In fact, the random variables 𝑋1 , 𝑋2 ,… need be neither identical nor independent: it is sufficient
that we can turn them into a normalized martingale. Without going into too much detail, it suffices
to assume that
𝔼 [𝑋𝑛+1 |𝑋1 = 𝑥1 … 𝑋𝑛 = 𝑥𝑛 ] = 𝔼 [𝑋𝑛+1 ]
Var (𝑋𝑛+1 |𝑋1 = 𝑥1 … 𝑋𝑛 = 𝑥𝑛 ) = Var (𝑋𝑛+1 ) > 0

125
(a) 𝑛 = 1 (b) 𝑛 = 2

(c) 𝑛 = 10 (d) 𝑛 = 100

Figure 9.2: Plots of 𝑓𝑛 for 𝑛 = 1, 2 (top row) and 𝑛 = 10, 100 (bottom row).

126
for all 𝑛 and all possible values for 𝑥1 , …,𝑥𝑛 . In this case, the sequence of random variables

∑𝑛𝑘=1 (𝑋𝑘 − 𝔼 [𝑋𝑘 ])


𝑍𝑛 ∶=
√∑𝑘=1 Var (𝑋𝑘 )
𝑛

converges to a 𝒩(0, 1) random variable whenever the martingale version of the Lindeberg condition
is satisfied. For further details, see for instance (Nelson 1987, chap. 14 & 18).

� Try it out

Potatoes with an average weight of 100g and standard deviation of 40g are packed into bags to
contain at least 2500g. What is the chance that more than 30 potatoes will be needed to fill a given
bag?
Answer:
Let 𝑁 be the number needed to exceed 2500 and 𝑊 be the total weight of 30 potatoes, 𝑊 = ∑𝑖=1 𝑋𝑖 .
30

Then 𝔼 [𝑊] = 30 ⋅ 100 = 3000 and Var (𝑊) = 30 ⋅ 402 = 48000, and the central limit theorem says
that
𝑊 − 3000
√ ≈ 𝒩(0, 1).
48000
Also, {𝑁 > 30} = {𝑊 < 2500} and so

𝑃 (𝑁 > 30) = 𝑃 (𝑊 < 2500)


𝑊 − 3000 2500 − 3000
= 𝑃( √ < √ )
48000 48000
≈ 𝑃 (𝑍 < −2.282) ,

where 𝑍 ∼ 𝒩(0, 1). From the tables, this probability is

𝑃 (𝑍 < −2.282) = 𝑃 (𝑍 > 2.282) = 1 − Φ(2.282) ≈ 0.011.

� Try it out

Measurements from a particular experiment have mean 𝜇 (unknown) and known standard deviation
𝜎 = 2.5. We perform 20 repetitions of the experiment and use 𝑋̄ 20 to estimate 𝜇. What is
𝑃 (|𝑋̄ 20 − 𝜇| < 1)?
Answer:
The central limit theorem says that

𝑋̄ 𝑛 − 𝜇
≈ 𝒩(0, 1).
√2.52 /𝑛

So
𝑃 (|𝑋̄ 20 − 𝜇| < 1) = 𝑃 (−1 ≤ 𝑋̄ 20 − 𝜇 ≤ 1)
1 𝑋̄ 20 − 𝜇 1
= 𝑃 (− ≤ ≤ )
√2.52 /20 2
√2.5 /20 √2.52 /20
≈ 𝑃 (−1.789 ≤ 𝑍 ≤ 1.789) ,

127
where 𝑍 ∼ 𝒩(0, 1). Thus

𝑃 (|𝑋̄ 20 − 𝜇| < 1) ≈ Φ(1.789) − Φ(−1.789)


= 2Φ(1.789) − 1
≈ 0.92,

using normal tables.


Note that Chebyshev’s inequality gives the weaker result

𝑃 (|𝑋̄ 20 − 𝜇| < 1) = 1 − 𝑃 (|𝑋̄ 20 − 𝜇| ≥ 1)


Var (𝑋̄ 20 )
≥1−
12
2
2.5
=1− ≈ 0.69.
20

Remember that we can write any binomially distributed random variable as a sum of independent Bernoulli
random variables. Consequently:

Corollary: Normal approximation to the Binomial

Let 𝑋 ∼ Bin(𝑛, 𝑝) with 0 < 𝑝 < 1, 𝐵𝑛 ∶= 𝑋/𝑛, and

𝑋 − 𝑛𝑝 𝐵𝑛 − 𝑝
𝑍𝑛 ∶= =
√𝑛𝑝(1 − 𝑝) √𝑝(1 − 𝑝)/𝑛

Then
lim 𝐹𝑍𝑛 (𝑧) = lim 𝑃 (𝑍𝑛 ≤ 𝑧) = Φ(𝑧).
𝑛→∞ 𝑛→∞

This means that, when 𝑋 ∼ Bin(𝑛, 𝑝) with 0 < 𝑝 < 1 and large enough 𝑛, then for any 𝑥 ∈ ℝ, we
approximately have that
𝑥 − 𝑛𝑝
𝑃 (𝑋 ≤ 𝑥) ≈ Φ ( )
√𝑛𝑝(1 − 𝑝)
For moderate 𝑛, a continuity correction improves the approximation; e.g. for 𝑘 ∈ ℕ:
𝑘 + 0.5 − 𝜇
𝑃 (𝑋 ≤ 𝑘) = 𝑃 (𝑋 ≤ 𝑘 + 0.5) ≈ Φ ( )
𝜎
𝑘 − 0.5 − 𝜇
𝑃 (𝑘 ≤ 𝑋) = 𝑃 (𝑘 − 0.5 ≤ 𝑋) ≈ 1 − Φ ( )
𝜎

� Try it out

Consider a multiple choice test with 50 questions, one mark for each correct answer. Independently
for each question, a particular student has chance 1/2 of answering correctly. Find, approximately,
the probability of the student scoring at least 30 marks.
Answer
Let 𝑋 = student’s score, so 𝑋 ∼ Bin(50, 1/2). Then 𝔼 [𝑋] = 50/2 = 25 and Var (𝑋) = 50/4 = 25/2.

128
The normal approximation says that 𝑋 ≈ 𝒩(25, 25/2), so

𝑃 (𝑋 ≥ 30) = 1 − 𝑃 (𝑋 < 30)


30 − 25
≈ 1 − Φ( ) = 1 − Φ(1.41) ≈ 0.1.
3.54

� Try it out

A plane has 110 seats and 𝑛 business people book seats. They show up independently for their flight
with chance 𝑞 = 0.85. Let 𝑋𝑛 be the number that arrive to take the flight (the others just take a
different flight) and find 𝑃 (𝑋𝑛 > 110) for 𝑛 = 110, 120, 130, 140.
Answer:

Using the CLT for binomials, 𝔼 [𝑋𝑛 ] = 0.85𝑛 and 𝜎 (𝑋𝑛 ) = √𝑛𝑞(1 − 𝑞) = 0.3571 𝑛 and so

𝑃 (𝑋𝑛 > 110) ≈ 1 − Φ((110.5 − 0.85𝑛)/0.3571 𝑛) which takes values 0, 0.015, 0.500 and 0.978 for
𝑛 = 110, 120, 130 and 140.

The proof of the central limit theorem requires an important new tool: the moment generating function.

� Textbook references

If you want more help with this section, check out:

• Section 10.3 in (Blitzstein and Hwang 2019);


• Section 9.3 in (Anderson, Seppäläinen, and Valkó 2018);
• or Section 8.9 in (Stirzaker 2003).

9.3 Moment generating functions

� Key idea: Definition: moment generating function

For any real-valued random variable 𝑋, the function 𝑀𝑋 ∶ ℝ → [0, +∞] given by

𝑀𝑋 (𝑡) ∶= 𝔼 [𝑒𝑡𝑋 ]

is called the moment generating function of 𝑋.

Because 𝑒𝑡𝑋 ≥ 0, we have that 𝑀𝑋 (𝑡) ≥ 0 by monotonicity of expectations.


Using the Law of the Unconscious Statistician, we can derive the following expressions for 𝑀𝑋 (𝑡):

𝑀𝑋 (𝑡) = ∑ 𝑒𝑡𝑥 𝑝(𝑥) if 𝑋 is discrete, and


𝑥∈𝒳

𝑀𝑋 (𝑡) = ∫ 𝑒𝑡𝑥 𝑓(𝑥)d𝑥 if 𝑋 is continuously distributed.
−∞

The above sum and integral always exist, but can be +∞.

129
Examples

1. If 𝑋 ∼ Bin(1, 𝑝) then 𝑀𝑋 (𝑡) = 𝑝𝑒𝑡 + (1 − 𝑝).

2. If 𝑌 ∼ Po(𝜆) then

𝑀𝑌 (𝑡) = ∑ 𝑒𝑡𝑥 𝑝(𝑥)
𝑥=0

𝜆𝑥 𝑡𝑥
= ∑ 𝑒−𝜆 𝑒
𝑥=0
𝑥!

(𝜆𝑒𝑡 )𝑥
= ∑ 𝑒−𝜆
𝑥=0
𝑥!
= exp(𝜆(𝑒𝑡 − 1)).

3. If 𝑈 ∼ U(𝑎, 𝑏) then
𝑒𝑏𝑡 − 𝑒𝑎𝑡
𝑀𝑈 (𝑡) =
(𝑏 − 𝑎)𝑡
for 𝑡 ≠ 0, and 𝑀𝑈 (0) = 1.

4. If 𝑍 ∼ 𝒩(0, 1) then

𝑀𝑍 (𝑡) = ∫ 𝑓𝑍 (𝑧)𝑒𝑡𝑧 d𝑧
−∞

1
√ 𝑒−𝑧 /2 𝑒𝑡𝑧 d𝑧
2
=∫
−∞ 2𝜋

1
√ 𝑒−(𝑧−𝑡) /2 𝑒𝑡 /2 d𝑧,
2 2
=∫
−∞ 2𝜋
as we see by completing the square in the exponential. Now put 𝑦 = 𝑧 − 𝑡 to get

1
√ 𝑒−𝑦 /2 𝑒𝑡 /2 d𝑦 = 𝑒𝑡 /2 ,
2 2 2
𝑀𝑍 (𝑡) = ∫
−∞ 2𝜋

because the 𝑦-dependent part is just 𝑓𝑍 (𝑦), which integrates to 1.

The moment generating function has several useful properties. The property that gives the name is
revealed by considering the Taylor series for 𝑒𝑡𝑋 : formally,

𝑡2 𝑋 2 𝑡3 𝑋 3
𝑀𝑋 (𝑡) = 𝔼 [𝑒𝑡𝑋 ] = 𝔼 [1 + 𝑡𝑋 + + + ⋯]
2! 3!
𝑡2 𝑡3
= 1 + 𝑡𝔼 [𝑋] + 𝔼 [𝑋 2 ] + 𝔼 [𝑋 3 ] + ⋯ ,
2! 3!
at least if 𝑡 ≈ 0. (Some work is needed to justify this last step, which we omit.) This gives the first of our
properties.

� Key idea: Properties of moment generating functions

M1: (Moment generating functions generate moments.)


For every 𝑘 ∈ ℕ,
𝑑 𝑘 𝑀𝑋
𝔼 [𝑋 𝑘 ] = (0).
𝑑𝑡𝑘

130
M2: (Moment generating function determines distribution.)
Consider any two random variables 𝑋 and 𝑌. If there is an ℎ > 0 such that

𝑀𝑋 (𝑡) = 𝑀𝑌 (𝑡) < +∞ for all 𝑡 ∈ (−ℎ, ℎ),

then
𝐹𝑋 (𝑥) = 𝐹𝑌 (𝑥) for all 𝑥 ∈ ℝ.
Conversely, if 𝐹𝑋 (𝑥) = 𝐹𝑌 (𝑥) for all 𝑥 ∈ ℝ then 𝑀𝑋 (𝑡) = 𝑀𝑌 (𝑡) for all 𝑡 ∈ ℝ.
M3: (Scaling.)
For any random variable 𝑋 and any constants 𝑎, 𝑏 ∈ ℝ,

𝑀𝑎𝑋+𝑏 (𝑡) = 𝑒𝑏𝑡 𝑀𝑋 (𝑎𝑡).

M4: (Product.)
Suppose that 𝑋1 , …, 𝑋𝑛 are independent random variables and let 𝑌 = ∑𝑖=1 𝑋𝑖 . Then
𝑛

𝑛
𝑀𝑌 (𝑡) = ∏ 𝑀𝑋𝑖 (𝑡).
𝑖=1

M5: (Convergence.)
Suppose that 𝑋1 , 𝑋2 , …is an infinite sequence of random variables, and that 𝑋 is a further random
variable. If there is an ℎ > 0 such that

lim 𝑀𝑋𝑛 (𝑡) = 𝑀𝑋 (𝑡) < +∞ for all 𝑡 ∈ (−ℎ, ℎ),


𝑛→∞

then
lim 𝐹𝑋𝑛 (𝑥) = 𝐹𝑋 (𝑥) for all 𝑥 ∈ ℝ where 𝐹𝑋 is continuous,
𝑛→∞
i.e., 𝑋𝑛 converges in distribution to 𝑋.

Proof

The proof of M3 just uses linearity of expectation.


M1, M2, and M5 use some deeper analysis, which we omit.
For M4, we use the fact that “independence means multiply” when working with expectations:

𝑀𝑆𝑛 (𝑡) = 𝔼 [𝑒𝑡(𝑋1 +𝑋2 +⋯+𝑋𝑛 ) ]


= 𝔼 [𝑒𝑡𝑋1 𝑒𝑡𝑋2 ⋯ 𝑒𝑡𝑋𝑛 ]
= 𝔼 [𝑒𝑡𝑋1 ] 𝔼 [𝑒𝑡𝑋2 ] ⋯ 𝔼 [𝑒𝑡𝑋𝑛 ]
𝑛
= ∏ 𝑀𝑋𝑖 (𝑡).
𝑖=1

Advanced content

Regarding M5 (convergence), in the case of the central limit theorem, the limit has 𝐹𝑋 (𝑥) = Φ(𝑥)
which is continuous for all 𝑥 ∈ ℝ, so in that case, 𝐹𝑋𝑛 converges to 𝐹𝑋 everywhere. As we saw earlier
(), 𝐹𝑋 is in fact continuous whenever 𝑋 is continuously distributed, so in that case convergence to

131
𝐹𝑋 (𝑥) takes place for all 𝑥.
In general, one can show that the cumulative distribution function 𝐹 of any random variable 𝑋 has
at most a countable number of points where 𝐹 is not continuous.
To see this, let
1
𝐷𝑛 = {𝑥 ∈ ℝ ∶ 𝐹 (𝑥) − 𝐹 (𝑥−) > } ,
𝑛
the points 𝑥 at which 𝐹 (𝑥) has a jump of size at least 1/𝑛. Then the points at which 𝐹 is discontinuous
can be expressed as
𝐷 = ⋃ 𝐷𝑛 .
𝑛∈ℕ

But 𝐷𝑛 is a finite set since 𝐹 is non-decreasing; indeed, 𝐷𝑛 is at most of size 𝑛, or else the jumps
would add up to more than 1, which is impossible. So 𝐷 is a countable union of finite sets, and is
hence countable.

Examples

1. Suppose that 𝑍 ∼ 𝒩(0, 1). We saw earlier that 𝑀𝑍 (𝑡) = 𝑒𝑡 .


2
/2

So, by M1,
2
𝔼 [𝑍] = 𝑀𝑍′ (0) = 𝑡𝑒𝑡 /2
∣ = 0,
𝑡=0
and
2
𝔼 [𝑍 2 ] = 𝑀𝑍″ (0) = (1 + 𝑡2 )𝑒𝑡 /2
∣ = 1,
𝑡=0

so Var (𝑍) = 𝔼 [𝑍 2 ] − 𝔼 [𝑍] = 1 − 0 = 1.


2

2. Suppose 𝑋 ∼ Po(𝜆). Then, by M1,



𝔼 [𝑋] = 𝑀𝑋 (0) = 𝜆𝑒0 exp(𝜆(𝑒0 − 1)) = 𝜆.

3. If 𝑍 ∼ 𝒩(0, 1), then by M3,

1
𝑀𝜇+𝜎𝑍 (𝑡) = 𝑒𝜇𝑡 𝑀𝑍 (𝜎𝑡) = exp (𝜇𝑡 + 𝜎2 𝑡2 ) .
2

We already know that 𝜇 + 𝜎𝑍 ∼ 𝒩(𝜇, 𝜎2 ) (see the “standardising the normal distribution”
theorem from Chapter 6), so the above expression gives us the moment generating function of
the normal distribution with mean 𝜇 and variance 𝜎2 .
Additionally, by M2, it follows that 𝑋 ∼ 𝒩(𝜇, 𝜎2 ) if and only if 𝑀𝑋 (𝑡) = exp (𝜇𝑡 + 12 𝜎2 𝑡2 ).

4. Suppose that 𝑋1 , …, 𝑋𝑘 are independent random variables with 𝑋𝑖 ∼ 𝒩(𝜇𝑖 , 𝜎𝑖2 ).


If 𝑌 = ∑𝑘𝑖=1 𝑋𝑖 , then, by M4, the moment generating function of 𝑌 is

𝑘
1 1
𝑀𝑌 (𝑡) = ∏ exp (𝜇𝑖 𝑡 + 𝜎𝑖2 𝑡2 ) = exp (𝜇𝑡 + 𝜎2 𝑡2 )
𝑖=1
2 2

where 𝜇 = ∑𝑘𝑖=1 𝜇𝑖 and 𝜎2 = ∑𝑘𝑖=1 𝜎𝑖2 . Thus, by M2, it must be that 𝑌 ∼ 𝒩(𝜇, 𝜎2 ).

132
� Try it out

Let 𝑋1 , … , 𝑋𝑛 be independent with 𝑋𝑖 ∼ Po(𝜆𝑖 ). Identify the distribution of 𝑌𝑛 = ∑𝑛𝑖=1 𝑋𝑖 .


Answer:
By M4 and our earlier calculation of the mgf of a Possion random variable,
𝑛
𝑀𝑌 (𝑡) = ∏ 𝑀𝑋𝑖 (𝑡)
𝑖=1
𝑛
= ∏ exp (𝜆𝑖 (𝑒𝑡 − 1))
𝑖=1
𝑛
= exp ((∑ 𝜆𝑖 ) (𝑒𝑡 − 1)) .
𝑖=1

By uniqueness (M2) this is the moment generating function of Po(∑𝑛𝑖=1 𝜆𝑖 ). So a sum of independent
Poissons is also Poisson!

Proof: Proof of the Central Limit Theorem

Recall from calculus: if 𝑎𝑛 → 𝑎 then


𝑎𝑛 𝑛
(1 + ) → 𝑒𝑎 .
𝑛
We want to show that for 𝑛
𝑋𝑖 − 𝜇
𝑍𝑛 = ∑ √
𝑖=1
𝜎 𝑛

we have 𝑀𝑍𝑛 (𝑡) → 𝑒𝑡 /2 for all 𝑡 in an open interval containing 0. This will give the central limit
2

theorem by M5.

Let 𝑌𝑖 = (𝑋𝑖 − 𝜇)/𝜎 and denote the moment generating function of the 𝑌𝑖 by 𝑚(𝑡). By M3, 𝑌𝑖 / 𝑛
√ √
has moment generating function 𝑚(𝑡/ 𝑛). Then by M4, 𝑍𝑛 = ∑𝑛𝑖=1 𝑌𝑖 / 𝑛 has moment generating
function √ 𝑛
𝑀𝑍𝑛 (𝑡) = (𝑚(𝑡/ 𝑛)) .
Next, by M1,
𝑚(0) = 𝔼 [𝑌𝑖0 ] = 1, 𝑚′ (0) = 𝔼 [𝑌𝑖 ] = 0, 𝑚″ (0) = 𝔼 [𝑌𝑖2 ] = 1.
So by Taylor’s theorem around 0 there is a function ℎ with ℎ(𝑢) → 0 such that

𝑢2
𝑚(𝑢) = 1 + + 𝑢2 ℎ(𝑢).
2
Hence 𝑛
𝑡2 𝑡2 √ 2
𝑀𝑍𝑛 (𝑡) = (1 + + ℎ(𝑡/ 𝑛)) → 𝑒𝑡 /2 ,
2𝑛 𝑛
for any 𝑡 ∈ ℝ.

� Textbook references
For more help with this section, check out:

• Section 6.4 in(Blitzstein and Hwang 2019);

133
• Section 5.1 in (Anderson, Seppäläinen, and Valkó 2018);
• or Sections 6.4 and 7.5 in (Stirzaker 2003).

9.4 Historical context

The law of large numbers and the central limit theorem have long and interesting histories. The weak law
of large numbers for binomial (i.e. sums of Bernoulli) variables was first established by Jacob Bernoulli
(1654–1705) and published in 1713 (Bernoulli 1713). The name ‘law of large numbers’ was given by Poisson.
The modern version is due to Aleksandr Khinchin (1894–1959), and our Central Limit Theorem is only a
special case—the assumption on variances is unnecessary.

(b) Khinchin (c) Lyapunov (d) Polya


(a) Bernoulli

It was apparent to mathematicians in the mid 1700s that a more refined result than Bernoulli’s law of
large numbers could be obtained. A special case of the central limit theorem for binomial (i.e. sums of
Bernoulli) variables was first established by de Moivre in 1733, and extended by Laplace; hence the normal
approximation to the binomial is sometimes known as the de Moivre–Laplace theorem. The name ‘central
limit theorem’ was given by George Pólya (1887–1985) in 1920.
The first modern proof of the central limit theorem was given by Aleksandr Lyapunov (1857–1918)
around 1901 (“Lyapunov Theorem,” n.d.). Lyapunov’s assumptions were relaxed by Jarl Waldemar
Lindeberg (1876–1932) in 1922 (Lindeberg 1922). Many different versions of the central limit theorem
were subsequently proved. The subject of Alan Turing’s (1912–1954) Cambridge University Fellowship
Dissertation of 1934 was a version of the central limit theorem similar to Lindeberg’s; Turing was unaware
of the latter’s work.

134
References

Anderson, D. F., T. Seppäläinen, and B. Valkó. 2018. Introduction to Probability. Cambridge University
Press.
Bernoulli, J. 1713. Ars Conjectandi. Basileæ: Thurnisiorum.
Billinton, R., and R. N. Allan. 1996. Reliability Evaluation of Power Systems. 2nd ed. Plenum Press.
Blitzstein, J. K., and J. Hwang. 2019. Introduction to Probability. Texts in Statistical Science Series. CRC
Press.
Boole, G. 1854. An Investigation of the Laws of Thought: On Which Are Founded the Mathematical
Theories of Logic and Probabilities. London: Walton; Maberly.
Chung, K. L., and F. AitSahlia. 2003. Elementary Probability Theory. 4th ed. Undergraduate Texts in
Mathematics. Springer-Verlag, New York.
DeGroot, M. H., and M. J. Schervish. 2013. Probability and Statistics. 4th ed. Harlow, England: Pearson.
Feller, W. 1968. An Introduction to Probability Theory and Its Applications, Vol. 1. 3rd ed. Wiley, New
York.
———. 1971. An Introduction to Probability Theory and Its Applications, Vol. 2. 2nd ed. Wiley, New
York.
Hacking, I. 2006. The Emergence of Probability. Cambridge University Press.
Hájek, A. 2012. “Interpretations of Probability.” In The Stanford Encyclopedia of Philosophy, edited by
Edward N. Zalta.
Kolmogorov, A. N. 1950. Foundations of the Theory of Probability. New York: Chelsea Publishing
Company.
Laplace, P. S. 1825. Essai Philosophique Sur Les Probabilitiés. Paris: Bachelier.
Lindeberg, J. W. 1922. “Eine Neue Herleitung Des Exponentialgesetzes in Der Wahrscheinlichkeitsrech-
nung.” Mathematische Zeitschrift 15: 211–25.
“Lyapunov Theorem.” n.d. Encyclopedia of Mathematics.
Mahmoud, H. M. 2009. Pólya Urn Models. CRC Press, Boca Raton, FL.
Moivre, A. de. 1756. The Doctrine of Chances: Or, a Method for Calculating the Probabilities of Events
in Play. Third. London: A. Millar.
Moran, P. A. P. 1968. An Introduction to Probability Theory. Clarendon Press, Oxford.
Nelson, E. 1987. Radically Elementary Probability Theory. Princeton University Press.
Rosenthal, J. 2007. A First Look at Rigorous Probability Theory. Second. New York: World Scientific.
Ross, S. M. 2010. Introduction to Probability Models. 10th ed. Academic Press, Amsterdam.
Stirzaker, D. 2003. Elementary Probability. Second. Cambridge University Press.
Todhunter, I. 2014. A History of the Mathematical Theory of Probability. Cambridge University Press.
Venn, J. 1888. The Logic of Chance: An Essay on the Foundations and Province of the Theory of
Probability, with Especial Reference to Its Application to Moral and Social Science. Third. London:
Macmillan.
Whittle, P. 1992. Probability via Expectation. Third. New York: Springer.
Whitworth, W. A. 1901. Choice and Chance. Third. Cambridge: Deighton Bell.

135

You might also like