0% found this document useful (0 votes)
10 views391 pages

Main

The document is a compilation of course lectures covering topics from probability to stochastic processes, aimed at MBA students from 2025-2027. It includes detailed sections on probability basics, combinatorics, random variables, and various probability theorems and examples. The lectures are structured to provide a comprehensive understanding of the subject matter through theoretical concepts and practical applications.

Uploaded by

Atul Mithul
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
10 views391 pages

Main

The document is a compilation of course lectures covering topics from probability to stochastic processes, aimed at MBA students from 2025-2027. It includes detailed sections on probability basics, combinatorics, random variables, and various probability theorems and examples. The lectures are structured to provide a comprehensive understanding of the subject matter through theoretical concepts and practical applications.

Uploaded by

Atul Mithul
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Course Lectures: From Probability to Stochastic Processes

Compiled by [MBA 2025-2027]

April 15, 2026


ii
Contents

I Probability 1

1 Lecture 1 3
1.1 The Basics of Probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
1.2 Sample Space and Events . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
1.3 Axioms of Probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4
1.4 Example . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
1.5 Set Theory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
1.5.1 Definition of a Set . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
1.5.2 Types of Sets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
1.6 Example (Set Theory) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
1.6.1 Venn Diagram . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
1.6.2 Set Containment . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
1.6.3 Probability Calculations . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
1.6.4 For both uniform and non-uniform sample space . . . . . . . . . . . . . . 7

2 Lecture 2 9
2.1 Combinatorics: The Art of Counting . . . . . . . . . . . . . . . . . . . . . . . . . 9
2.1.1 Permutations: When Order Matters . . . . . . . . . . . . . . . . . . . . . 9
2.1.2 Combinations: When Order Doesn’t Matter . . . . . . . . . . . . . . . . . 9
2.2 Solved Examples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10
2.2.1 Problem 1: Distributing Teachers to Schools . . . . . . . . . . . . . . . . 10
2.2.2 Problem 2: Distributing Indistinguishable Objects (Stars and Bars) . . . 12
2.3 The Multinomial Coefficient . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13

3 Lecture 3 15
3.1 Stars and Bars: Basic Idea . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
3.1.1 Distributions into Non-Empty Boxes . . . . . . . . . . . . . . . . . . . . . 16
3.1.2 Distribution Allowing Empty Boxes . . . . . . . . . . . . . . . . . . . . . 17
3.1.3 Equation-Based Counting . . . . . . . . . . . . . . . . . . . . . . . . . . . 17

4 Lecture 4 19
4.1 Sample Space and Events . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19
4.2 Probability Function . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19
4.3 Axioms of Probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19
4.4 Theorem I . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
4.5 Theorem 2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21
4.6 Theorem 3 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23
4.7 Sum of First (n − 1) Natural Numbers . . . . . . . . . . . . . . . . . . . . . . . . 25
4.8 General Formula for Binomial Coefficient . . . . . . . . . . . . . . . . . . . . . . 26

iii
iv CONTENTS

5 Lecture 5 27
5.1 Example . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
5.2 Conditional Probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
5.3 Sample Space in a collection of all basic outcomes ω ∈ Ω of some experiment. . . 28
5.4 Example . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29
5.5 Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30

6 Lecture 7 33
6.1 Mutual Independence, Pairwise Independence, Conditional Independence . . . . 33
6.2 Example: Tossing a Fair Die . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 33
6.3 Conditional Probability and Intersection . . . . . . . . . . . . . . . . . . . . . . . 34
6.4 Axioms of Conditional Probability . . . . . . . . . . . . . . . . . . . . . . . . . . 34
6.5 Sample Space and Events . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35
6.6 Independence of Events . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 36
6.7 Example: Drawing Balls without Replacement . . . . . . . . . . . . . . . . . . . 36
6.8 Conditional Probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37
6.9 Law of Total Probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37
6.10 Bayes’ Theorem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37
6.11 Example: Box and Balls Problem . . . . . . . . . . . . . . . . . . . . . . . . . . . 37
6.12 Disjoint Events . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38
6.13 Unions of Events . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38
6.14 Partitions of the Sample Space . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38
6.15 Example: Tossing a Fair Coin Twice . . . . . . . . . . . . . . . . . . . . . . . . . 39
6.16 Example: Mutually Independent Events . . . . . . . . . . . . . . . . . . . . . . . 39

7 Lecture 8 41
7.1 Random Variables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41
7.1.1 Example . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41
7.1.2 Probability Mass Function (PMF) . . . . . . . . . . . . . . . . . . . . . . 41
7.1.3 Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41
7.1.4 Probability Mass Function (PMF) . . . . . . . . . . . . . . . . . . . . . . 42
7.2 Experiment: Waiting Time for First Head . . . . . . . . . . . . . . . . . . . . . . 42
7.2.1 Probability Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42
7.3 X is a Random Variable . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42
7.3.1 Example . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 43
7.3.2 Properties of CDF . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 43

8 Lecture 9 45
8.1 Probability Example: Uniform over (0, 1) . . . . . . . . . . . . . . . . . . . . . . 45
8.1.1 Facts about the uniform distribution on (0, 1) . . . . . . . . . . . . . . . . 45
8.1.2 Probability of a single point . . . . . . . . . . . . . . . . . . . . . . . . . . 45
8.2 Continuous Random Variable . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 46
8.3 Discrete Random Variable . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 46
8.3.1 Expectation of a Discrete Random Variable . . . . . . . . . . . . . . . . . 46
8.4 Example: Tossing Coins . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 46
8.4.1 Tossing 2 Coins: Number of Heads X . . . . . . . . . . . . . . . . . . . . 46
8.4.2 Tossing 10 Coins: Observed Sample . . . . . . . . . . . . . . . . . . . . . 47
8.5 Indicator Random Variable Example . . . . . . . . . . . . . . . . . . . . . . . . . 47
8.6 Expectation Examples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 47
8.6.1 Expectation of a Fair Die . . . . . . . . . . . . . . . . . . . . . . . . . . . 47
8.6.2 Expectation of a Biased Coin (Two Tosses) . . . . . . . . . . . . . . . . . 47
CONTENTS v

9 Lecture 10 49
9.1 Random Variable in Function . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 51
9.2 Variance of RV X . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 55
9.2.1 For function Y=[x+a] — a ∈ R . . . . . . . . . . . . . . . . . . . . . . . . 55

10 Lecture 11 57
10.0.1 Linearity of Expectation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 57
10.1 Variance Properties . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 57
10.2 Expectation and Variance of Linear Combinations . . . . . . . . . . . . . . . . . 58
10.2.1 Example: Toss a fair coin twice . . . . . . . . . . . . . . . . . . . . . . . . 59
10.3 Example: Picking Balls . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59
10.4 Example: Toss a Coin Once . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59
10.5 Bernoulli Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 60
10.6 Example 2: Toss one biased coin n times . . . . . . . . . . . . . . . . . . . . . . . 60
10.6.1 Case n = 2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 60

11 Lecture 12 61
11.1 Families of Random Variables (RVs) — Bernoulli RV . . . . . . . . . . . . . . . . 61
11.1.1 1. Xi ∼ Bern(p) [One Trial, One Parameter] . . . . . . . . . . . . . . . . . 61
11.1.2 2. n sets of identical & mutually independent trials . . . . . . . . . . . . . 62
11.1.3 Binomial RV and PMF . . . . . . . . . . . . . . . . . . . . . . . . . . . . 62
11.1.4 3. Outcomes of n trials . . . . . . . . . . . . . . . . . . . . . . . . . . . . 63
11.1.5 4. Cumulative Distribution Function (CDF) for Binomial RV . . . . . . . 63
11.1.6 5. Probability Calculation for Different Paths . . . . . . . . . . . . . . . . 64
11.2 Multiplication of probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 64
11.2.1 7. Further Expansion and Multiplication of Probabilities . . . . . . . . . . 64

12 Lecture 13 67
12.1 The Binomial Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 67
12.1.1 Probability Mass Function (PMF) . . . . . . . . . . . . . . . . . . . . . . 67
12.1.2 The Bernoulli Distribution (n = 1) . . . . . . . . . . . . . . . . . . . . . . 67
12.1.3 Expected Value of a Bernoulli Random Variable . . . . . . . . . . . . . . 68
12.2 Binomial Probability Mass Function (PMF) . . . . . . . . . . . . . . . . . . . . . 68
12.2.1 Path Interpretation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 68
12.3 Tree Diagram for n = 2 Trials . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 68
12.4 Probability of a Single Path (Sequence) . . . . . . . . . . . . . . . . . . . . . . . 69
12.5 Total Probability of y Successes . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69
12.5.1 Definition of the Event En,y . . . . . . . . . . . . . . . . . . . . . . . . . . 69
12.5.2 Calculating P (En,y ) using Disjoint Union . . . . . . . . . . . . . . . . . . 69
12.5.3 Final PMF Formula . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 70
12.6 Example: Coin Toss 3 Times (n = 3) . . . . . . . . . . . . . . . . . . . . . . . . . 70
12.6.1 Outcomes and Probabilities for n = 3 . . . . . . . . . . . . . . . . . . . . 70
12.7 Expected Value of a Bernoulli Trial . . . . . . . . . . . . . . . . . . . . . . . . . . 70
12.8 Conditions for Binomial Distribution . . . . . . . . . . . . . . . . . . . . . . . . . 70
12.8.1 Independence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71
12.8.2 Identically Distributed . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71
12.8.3 General Formulation and Paths . . . . . . . . . . . . . . . . . . . . . . . . 71
12.9 Scenario: Events are NOT i.i.d. . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71

13 Lecture 14 73
vi CONTENTS

14 Lecture 15 77
14.1 Probability Distributions and Memoryless Property . . . . . . . . . . . . . . . . . 77
14.1.1 Calculation of P [G > t] . . . . . . . . . . . . . . . . . . . . . . . . . . . . 78
14.1.2 Image 1: Memoryless Property - Conditional Probability . . . . . . . . . . 78
14.1.3 Memoryless Property . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 79
14.1.4 Expected Value and Variance of Bernoulli RV . . . . . . . . . . . . . . . . 80
14.1.5 Random Variables and Expected Value Formula . . . . . . . . . . . . . . 80
14.1.6 Image 3: Law of Expectation (LOE) and Law of Variance (LOV) . . . . . 81
14.1.7 Covariance and Independence . . . . . . . . . . . . . . . . . . . . . . . . . 81
14.1.8 Image 0: Conditional Probability of Geometric RV (Shift) . . . . . . . . . 82

15 Lecture 16 83
15.1 Probability Measure and Axioms . . . . . . . . . . . . . . . . . . . . . . . . . . . 83
15.1.1 Experiment and Sample Space . . . . . . . . . . . . . . . . . . . . . . . . 83
15.1.2 Probability Measure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83
15.1.3 Axioms of Probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83
15.1.4 Countable Collection of Sets . . . . . . . . . . . . . . . . . . . . . . . . . 84
15.2 Concept of Random Variable . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 84
15.2.1 Discrete Random Variable . . . . . . . . . . . . . . . . . . . . . . . . . . . 84
15.2.2 Continuous Random Variable . . . . . . . . . . . . . . . . . . . . . . . . . 84
15.3 Cumulative Distribution Function (CDF) . . . . . . . . . . . . . . . . . . . . . . 85
15.4 Joint Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 85
15.4.1 Example: Two Fair Dice . . . . . . . . . . . . . . . . . . . . . . . . . . . . 85
15.5 Joint Distribution Table . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 86
15.6 Independent Random Variables . . . . . . . . . . . . . . . . . . . . . . . . . . . . 86
15.7 Example . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 86
15.8 i.i.d Random Variables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 87
15.9 Independent Random Variables . . . . . . . . . . . . . . . . . . . . . . . . . . . . 87
15.10Expectation and Covariance . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 87
15.11Continuous Random Variable X . . . . . . . . . . . . . . . . . . . . . . . . . . . 87

16 Lecture 17 89
16.1 Derivation of Expected Value E[Y ] . . . . . . . . . . . . . . . . . . . . . . . . . . 89
16.1.1 Expansion of the Summation . . . . . . . . . . . . . . . . . . . . . . . . . 89
16.1.2 Formal Manipulation of the Term . . . . . . . . . . . . . . . . . . . . . . . 89
16.2 Completing the Derivation of E[Y ] . . . . . . . . . . . . . . . . . . . . . . . . . . 90
16.2.1 Algebraic Manipulation (from i = 1) . . . . . . . . . . . . . . . . . . . . . 90
16.2.2 Factoring out np . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 90
16.2.3 Change of Index and Variable Substitution . . . . . . . . . . . . . . . . . 90
16.2.4 Recognizing the Binomial Theorem . . . . . . . . . . . . . . . . . . . . . . 91
16.2.5 Final Result . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 91
16.3 Derivation of E[Y (Y − 1)] . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 91
16.3.1 Setup and Initial Summation . . . . . . . . . . . . . . . . . . . . . . . . . 91
16.3.2 Algebraic Manipulation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 92
16.3.3 Factoring out n(n − 1)p2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . 92
16.3.4 Change of Index and Variable Substitution . . . . . . . . . . . . . . . . . 92
16.3.5 Recognizing the Binomial Theorem . . . . . . . . . . . . . . . . . . . . . . 92
16.3.6 Final Result . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 92
16.4 Variance of the Binomial Distribution Y ∼ Bin(n, p) . . . . . . . . . . . . . . . . 93
16.4.1 Substitution of Moments . . . . . . . . . . . . . . . . . . . . . . . . . . . . 93
16.5 Expected Value of the Poisson Distribution . . . . . . . . . . . . . . . . . . . . . 93
16.5.1 Derivation of E[X] . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 93
CONTENTS vii

16.5.2 Simplification and Series Recognition . . . . . . . . . . . . . . . . . . . . . 94


16.5.3 Final Result . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 94
16.6 Variance of the Poisson Distribution . . . . . . . . . . . . . . . . . . . . . . . . . 94
16.6.1 Derivation of E[X(X − 1)] . . . . . . . . . . . . . . . . . . . . . . . . . . . 95
16.6.2 Final Variance Calculation . . . . . . . . . . . . . . . . . . . . . . . . . . 95
16.7 Geometric Distribution Derivation X ∼ Geom(p) . . . . . . . . . . . . . . . . . . 96
16.7.1 Probability Mass Function (PMF) . . . . . . . . . . . . . . . . . . . . . . 96
16.7.2 Expected Value E[X] (Derivation) . . . . . . . . . . . . . . . . . . . . . . 96
16.8 Note on Variance Calculation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 97
16.9 Variance of the Geometric Distribution X ∼ Geom(p) . . . . . . . . . . . . . . . 97
16.9.1 Using the Identity Var[X] = E[X 2 ] − (E[X])2 . . . . . . . . . . . . . . . . 97
16.10Negative Binomial Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . 98
16.10.1 Probability Mass Function (PMF) of NB(r, p) . . . . . . . . . . . . . . . . 98

II Statistical Inference 99

17 Lecture 2: The Geometric and Exponential Distributions 101


17.1 The Geometric Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 101
17.1.1 Definition and Context . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 101
17.1.2 Probability Mass Function (PMF) . . . . . . . . . . . . . . . . . . . . . . 101
17.1.3 Tail Probability and Cumulative Distribution Function (CDF) . . . . . . 101
17.1.4 Properties of the Geometric Distribution . . . . . . . . . . . . . . . . . . . 102
17.2 The Exponential Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 102
17.2.1 Definition and Probability Density Function (PDF) . . . . . . . . . . . . . 102
17.2.2 Cumulative Distribution Function (CDF) . . . . . . . . . . . . . . . . . . 103
17.2.3 Tail Probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 103
17.3 Properties of the Exponential Distribution . . . . . . . . . . . . . . . . . . . . . . 103
17.3.1 The Memoryless Property . . . . . . . . . . . . . . . . . . . . . . . . . . . 103
17.3.2 The Shifting Property and Visual Interpretation . . . . . . . . . . . . . . 104
17.3.3 Conditional PDFs and CDFs . . . . . . . . . . . . . . . . . . . . . . . . . 104

18 Lecture 3: Exponential Distribution Derivations 105


18.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 105
18.2 Derivation of E[X] for an Exponential Distribution . . . . . . . . . . . . . . . . . 105
18.3 Derivations of E[X] and E[X 2 ] for an Exponential Distribution . . . . . . . . . . 106
18.3.1 Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 106
18.3.2 1. Compute E[X] . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 106
18.3.3 2. Compute E[X 2 ] . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 107
18.3.4 3. Variance (optional) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 108
18.4 Variance of an Exponential Random Variable . . . . . . . . . . . . . . . . . . . . 108
18.4.1 1. Mean of X . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 108
18.4.2 2. Second Moment E[X 2 ] . . . . . . . . . . . . . . . . . . . . . . . . . . . 108
18.4.3 3. Variance Computation . . . . . . . . . . . . . . . . . . . . . . . . . . . 108
18.4.4 4. Relation to Poisson Distribution . . . . . . . . . . . . . . . . . . . . . . 108
18.4.5 5. Shape of the Exponential Density . . . . . . . . . . . . . . . . . . . . . 109

19 Lecture 4: Normalization and Sample Statistics 111


19.1 Normalization of Random Variables . . . . . . . . . . . . . . . . . . . . . . . . . 111
19.1.1 General Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 111
19.1.2 Example: {1, 2, 3} . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 111
viii CONTENTS
 
19.1.3 Calculations for Set: − √1 , √0 , √+1 . . . . . . . . . . . . . . . . . 112
2/3 2/3 2/3
 
19.1.4 Calculations for Set: √1 , √2 , √3 . . . . . . . . . . . . . . . . . . 112
2/3 2/3 2/3
19.2 Normalization and Sample Statistics . . . . . . . . . . . . . . . . . . . . . . . . . 112
19.2.1 Shifted Exponential Distribution . . . . . . . . . . . . . . . . . . . . . . . 112
19.2.2 Normalization of a Random Variable . . . . . . . . . . . . . . . . . . . . . 113
19.2.3 Proof of Mean for Z . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 113
19.2.4 Proof of Variance for Z . . . . . . . . . . . . . . . . . . . . . . . . . . . . 113
19.3 Sample Statistics for n i.i.d. Variables . . . . . . . . . . . . . . . . . . . . . . . . 113
19.3.1 Sample Sum (Sn ) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 114
19.3.2 Sample Average (X̄n ) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 114
19.4 Statistical Properties of Sample Metrics . . . . . . . . . . . . . . . . . . . . . . . 114
19.4.1 Standardization and Verification . . . . . . . . . . . . . . . . . . . . . . . 114
19.4.2 Numerical Example: Standardizing a Discrete Set . . . . . . . . . . . . . 115
19.5 Expectation and Variance of Sample Sums and Averages . . . . . . . . . . . . . . 115
19.5.1 The Sample Sum (Sn ) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 115
19.5.2 The Sample Average (X̄n ) . . . . . . . . . . . . . . . . . . . . . . . . . . . 115
19.5.3 Standardizing the Sample Mean . . . . . . . . . . . . . . . . . . . . . . . 116
19.6 Fundamentals of Statistics and Calculus . . . . . . . . . . . . . . . . . . . . . . . 116
19.6.1 Normalization of Random Variables . . . . . . . . . . . . . . . . . . . . . 116
19.7 Review of Calculus Fundamentals . . . . . . . . . . . . . . . . . . . . . . . . . . . 116
19.7.1 Differential Calculus: The Slope . . . . . . . . . . . . . . . . . . . . . . . 116
19.7.2 Integral Calculus: Area Under the Curve . . . . . . . . . . . . . . . . . . 117

20 Lectures 5 and 6: Moment Generating Functions and Limit Theorems 119


20.1 Moment Generating Function (MGF): Introduction . . . . . . . . . . . . . . . . . 119
20.2 Power Series Expansion of a Function . . . . . . . . . . . . . . . . . . . . . . . . 119
20.2.1 Step 1: Function value at x = 0 . . . . . . . . . . . . . . . . . . . . . . . . 120
20.2.2 Step 2: First derivative . . . . . . . . . . . . . . . . . . . . . . . . . . . . 120
20.2.3 Step 3: Second derivative . . . . . . . . . . . . . . . . . . . . . . . . . . . 120
20.2.4 Step 4: General nth derivative . . . . . . . . . . . . . . . . . . . . . . . . 120
20.2.5 Step 5: Power series representation . . . . . . . . . . . . . . . . . . . . . . 120
20.2.6 Final Result: Maclaurin Series . . . . . . . . . . . . . . . . . . . . . . . . 121
20.3 Moment Generating Function (MGF) Definitions . . . . . . . . . . . . . . . . . . 121
20.3.1 Moments Using MGF . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 121
20.3.2 MGF for Different Types of Random Variables . . . . . . . . . . . . . . . 121
20.4 Using Taylor Expansion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 121
20.4.1 Substitution into the MGF . . . . . . . . . . . . . . . . . . . . . . . . . . 122
20.4.2 Using Linearity of Expectation . . . . . . . . . . . . . . . . . . . . . . . . 122
20.4.3 Differentiation of the MGF . . . . . . . . . . . . . . . . . . . . . . . . . . 122
20.4.4 Evaluation at t = 0 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 122
20.4.5 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 122
20.5 Moment Generating Function of the Binomial Distribution . . . . . . . . . . . . . 122
20.5.1 Using MGF to Find Moments in Detail . . . . . . . . . . . . . . . . . . . 123
20.6 Uniqueness Property of MGF . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 124
20.6.1 Example: Tossing a Coin Twice . . . . . . . . . . . . . . . . . . . . . . . . 124
20.7 Purpose of the MGF in the Standard Normal Case . . . . . . . . . . . . . . . . . 125
20.8 MGF for Normal Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 125
20.8.1 Standard Normal Case . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 125
20.8.2 MGF of a General Normal Distribution . . . . . . . . . . . . . . . . . . . 126
20.8.3 Substitution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 126
CONTENTS ix

20.8.4 Final MGF of Normal . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 127


20.9 Properties of the MGF . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 127
20.9.1 (i) Scaling Property . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 127
20.9.2 (ii) Sum of Independent Random Variables . . . . . . . . . . . . . . . . . 127
20.9.3 Shift by a Constant . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 127
20.9.4 Normal Distribution Revisited . . . . . . . . . . . . . . . . . . . . . . . . 127
20.10Expectation via MGF . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 128
20.10.1 Second Moment . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 128
20.10.2 Variance . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 128
20.11Example: Sum of Two Independent Normal Random Variables . . . . . . . . . . 128
20.11.1 Using MGF . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 128
20.11.2 Final Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 129
20.12Limit Theorems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 129
20.12.1 Law of Large Numbers (LLN) . . . . . . . . . . . . . . . . . . . . . . . . . 129
20.13Population and Sample . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 130
20.14Central Limit Theorem (CLT) . . . . . . . . . . . . . . . . . . . . . . . . . . . . 130
20.15Purpose of the Central Limit Theorem (CLT) . . . . . . . . . . . . . . . . . . . . 130

21 Lecture 7: Central Limit Theorem 131


21.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 131
21.2 Theoretical Background . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 131
21.3 Standardisation of the Sample Mean . . . . . . . . . . . . . . . . . . . . . . . . . 132
21.3.1 CDF of ZX n . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 132
21.4 Sample Sum Form and MGF Proof Setup . . . . . . . . . . . . . . . . . . . . . . 132

22 Lecture 8: Asymptotic Bounds and Inequalities 133


22.1 MGF Proof of the Central Limit Theorem . . . . . . . . . . . . . . . . . . . . . . 133
22.1.1 Applying L’Hospital’s Rule . . . . . . . . . . . . . . . . . . . . . . . . . . 133
22.2 Asymptotic Bounds . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 134
22.2.1 Markov’s Inequality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 134
22.2.2 Chebyshev’s Inequality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 135

23 Lecture 11: Transformations of Random Variables 137


23.1 Fundamental Theorem of Calculus (FTC) . . . . . . . . . . . . . . . . . . . . . . 137
23.1.1 Example 1: Square Root of an Exponential Variable . . . . . . . . . . . . 137
23.1.2 Example 2: Transformation of a Uniform Variable . . . . . . . . . . . . . 138
23.1.3 Example 3: Standardizing a Normal Distribution . . . . . . . . . . . . . . 138

24 Lecture 12: Jointly Distributed Random Variables 141


24.1 1D vs 2D Probability Representation . . . . . . . . . . . . . . . . . . . . . . . . . 141
24.2 Jointly Continuous Random Variables . . . . . . . . . . . . . . . . . . . . . . . . 142
24.2.1 Properties of Joint PDF . . . . . . . . . . . . . . . . . . . . . . . . . . . . 142
24.3 Joint CDF and its Relationship to PDF . . . . . . . . . . . . . . . . . . . . . . . 142
24.4 Evaluating Joint PDFs and Probabilities . . . . . . . . . . . . . . . . . . . . . . . 143
24.4.1 Checking if a PDF is Valid . . . . . . . . . . . . . . . . . . . . . . . . . . 143
24.4.2 Finding Regional Probabilities . . . . . . . . . . . . . . . . . . . . . . . . 144
24.4.3 Finding Conditional Probabilities . . . . . . . . . . . . . . . . . . . . . . . 144
x CONTENTS

25 Lectures 9 and 10: Joint Distributions and Conditional Expectation 145


25.1 Probability Bounds: Chebyshev’s Inequality . . . . . . . . . . . . . . . . . . . . . 145
25.1.1 Applying Chebyshev’s Inequality . . . . . . . . . . . . . . . . . . . . . . . 145
25.2 Jointly Discrete Random Variables . . . . . . . . . . . . . . . . . . . . . . . . . . 146
25.3 Jointly Continuous Random Variables . . . . . . . . . . . . . . . . . . . . . . . . 146
25.4 Linearity of Expectation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 147
25.4.1 Discrete Case Proof . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 147
25.4.2 Continuous Case Proof . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 147
25.5 Covariance and Independence Example . . . . . . . . . . . . . . . . . . . . . . . . 147
25.6 Variance of Linear Combinations . . . . . . . . . . . . . . . . . . . . . . . . . . . 148
25.7 Conditional Expectation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 149
25.8 Law of Total Expectation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 149
25.8.1 Proof (Discrete Case) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 149
25.8.2 Example 1: Tabular Data . . . . . . . . . . . . . . . . . . . . . . . . . . . 150
25.9 Example: Geometric Distribution Expectation . . . . . . . . . . . . . . . . . . . 150
25.10Example: Elevator Stops . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 151

26 Lecture 11: Transformations of Random Variables 153


26.1 Fundamental Theorem of Calculus (FTC) . . . . . . . . . . . . . . . . . . . . . . 153
26.1.1 Example 1: Square Root of an Exponential Variable . . . . . . . . . . . . 153
26.1.2 Example 2: Transformation of a Uniform Variable . . . . . . . . . . . . . 154
26.1.3 Example 3: Standardizing a Normal Distribution . . . . . . . . . . . . . . 154

27 Lecture 12: Jointly Distributed Random Variables 157


27.1 1D vs 2D Probability Representation . . . . . . . . . . . . . . . . . . . . . . . . . 157
27.2 Jointly Continuous Random Variables . . . . . . . . . . . . . . . . . . . . . . . . 158
27.2.1 Properties of Joint PDF . . . . . . . . . . . . . . . . . . . . . . . . . . . . 158
27.3 Joint CDF and its Relationship to PDF . . . . . . . . . . . . . . . . . . . . . . . 158
27.4 Evaluating Joint PDFs and Probabilities . . . . . . . . . . . . . . . . . . . . . . . 159
27.4.1 Checking if a PDF is Valid . . . . . . . . . . . . . . . . . . . . . . . . . . 159
27.4.2 Finding Regional Probabilities . . . . . . . . . . . . . . . . . . . . . . . . 160
27.4.3 Finding Conditional Probabilities . . . . . . . . . . . . . . . . . . . . . . . 160

28 Lecture 13: Conditional Probability (Continuous Case) 161


28.1 Joint PDF Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 161
28.2 Calculating Conditional Probabilities . . . . . . . . . . . . . . . . . . . . . . . . . 161
28.2.1 Graphical Representation of Events . . . . . . . . . . . . . . . . . . . . . 161
28.3 Marginal Distributions and Expectation . . . . . . . . . . . . . . . . . . . . . . . 162
28.3.1 Marginal PDF of X . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 162
28.3.2 Expectation of X . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 162

29 Lecture 14: Independence and Sums of Variables 163


29.1 Independence of Random Variables . . . . . . . . . . . . . . . . . . . . . . . . . . 163
29.2 Example: Sum of Independent Poisson RVs . . . . . . . . . . . . . . . . . . . . . 163
29.2.1 Derivation via Convolution (PMF Method) . . . . . . . . . . . . . . . . . 163
29.2.2 Derivation via Moment Generating Functions (MGF) . . . . . . . . . . . 164
29.3 Sum of Two Uniform Random Variables . . . . . . . . . . . . . . . . . . . . . . . 164
29.3.1 Problem Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 164
29.3.2 Case 1: a < 0 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 164
29.3.3 Case 2: 0 ≤ a ≤ 1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 164
29.3.4 Case 3: 1 < a ≤ 2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 164
29.3.5 Final CDF and PDF . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 165
CONTENTS xi

30 Lectures 15 and 16: Maximum Likelihood Estimation 167


30.1 Introduction to Estimation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 167
30.1.1 Estimation as an Inverse Problem . . . . . . . . . . . . . . . . . . . . . . 167
30.2 Parameters of Distributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 167
30.2.1 Bernoulli Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 167
30.2.2 Gaussian Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 167
30.3 Likelihood Function . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 168
30.3.1 Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 168
30.3.2 Log-Likelihood . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 168
30.4 Examples of Log-Likelihood . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 168
30.4.1 Gaussian Case . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 168
30.4.2 Bernoulli Case . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 168
30.5 Maximum Likelihood Estimation . . . . . . . . . . . . . . . . . . . . . . . . . . . 168
30.5.1 Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 168
30.5.2 MLE for Bernoulli Distribution . . . . . . . . . . . . . . . . . . . . . . . . 168
30.5.3 MLE for Gaussian Mean . . . . . . . . . . . . . . . . . . . . . . . . . . . . 169
30.6 Application: Social Network Estimation . . . . . . . . . . . . . . . . . . . . . . . 169
30.7 Application: Single-Photon Imaging . . . . . . . . . . . . . . . . . . . . . . . . . 169
30.7.1 Log-Likelihood . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 169
30.8 Estimation as an Inverse Problem . . . . . . . . . . . . . . . . . . . . . . . . . . . 169
30.9 Likelihood and Log-Likelihood . . . . . . . . . . . . . . . . . . . . . . . . . . . . 170
30.10Sample Sum 1: Bernoulli MLE (Fully Worked) . . . . . . . . . . . . . . . . . . . 170
30.11Bernoulli Log-Likelihood Shape (TikZ Plot) . . . . . . . . . . . . . . . . . . . . . 171
30.12Sample Sum 2: Gaussian Mean Estimation . . . . . . . . . . . . . . . . . . . . . 171
30.13Likelihood Matching Intuition (Diagram) . . . . . . . . . . . . . . . . . . . . . . 172
30.14Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 172

31 Lectures 17 and 18: Limit Theorems 173


31.1 Introduction and Motivation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 173
31.2 Probability Setup and Assumptions . . . . . . . . . . . . . . . . . . . . . . . . . . 173
31.3 Definition of the Sample Mean . . . . . . . . . . . . . . . . . . . . . . . . . . . . 173
31.4 Expectation of the Sample Mean . . . . . . . . . . . . . . . . . . . . . . . . . . . 174
31.5 Variance of the Sample Mean . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 174
31.6 Weak Law of Large Numbers (WLLN) . . . . . . . . . . . . . . . . . . . . . . . . 175
31.7 Proof of WLLN Using Chebyshev’s Inequality . . . . . . . . . . . . . . . . . . . . 175
31.8 Event-Based Interpretation of WLLN . . . . . . . . . . . . . . . . . . . . . . . . 175
31.9 Strong Law of Large Numbers (SLLN) . . . . . . . . . . . . . . . . . . . . . . . . 176
31.10Bernoulli Example . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 176
31.11WLLN via Moment Generating Functions . . . . . . . . . . . . . . . . . . . . . . 176
31.12Central Limit Theorem (CLT) . . . . . . . . . . . . . . . . . . . . . . . . . . . . 177
31.13Final Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 177

32 Lectures 19 and 20: Sampling Distributions (χ2 , t, and F) 179


32.1 Sampling Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 179
32.1.1 Population and Samples . . . . . . . . . . . . . . . . . . . . . . . . . . . . 179
32.1.2 After collecting a sample . . . . . . . . . . . . . . . . . . . . . . . . . . . 179
32.1.3 Sample Mean . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 179
32.1.4 Sample Variance . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 179
32.1.5 Useful identity (DOF explanation) . . . . . . . . . . . . . . . . . . . . . . 179
32.2 Distribution from Normal . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 180
32.2.1 Standard Normal . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 180
32.2.2 Sampling distribution of X . . . . . . . . . . . . . . . . . . . . . . . . . . 180
xii CONTENTS

32.3 Chi-square Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 180


32.3.1 Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 180
32.3.2 Chi-square from sample variance . . . . . . . . . . . . . . . . . . . . . . . 180
32.3.3 Right tail critical value . . . . . . . . . . . . . . . . . . . . . . . . . . . . 180
32.3.4 Sketch (right tail) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 180
32.3.5 Additive property . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 181
32.3.6 Expansion idea . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 181
32.3.7 Example: distance in 3D . . . . . . . . . . . . . . . . . . . . . . . . . . . . 181
32.4 Student t Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 181
32.4.1 Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 181
32.4.2 From sample mean when σ unknown . . . . . . . . . . . . . . . . . . . . . 181
32.4.3 Note: t density vs Normal . . . . . . . . . . . . . . . . . . . . . . . . . . . 181
32.4.4 Symmetry property . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 182
32.4.5 Critical values and symmetry relation . . . . . . . . . . . . . . . . . . . . 182
32.4.6 Sketch . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 182
32.4.7 As n → ∞ . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 182
32.5 F Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 182
32.5.1 Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 182
32.5.2 F critical value (upper tail) . . . . . . . . . . . . . . . . . . . . . . . . . . 182
32.5.3 Key reciprocal property derivation . . . . . . . . . . . . . . . . . . . . . . 183
32.5.4 Sketch (right tail) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 183
32.6 Two-Population Variance Ratio . . . . . . . . . . . . . . . . . . . . . . . . . . . . 183
32.6.1 Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 183
32.6.2 Chi-square forms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 183
32.6.3 F ratio . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 183

III Introduction to Machine Learning 185

33 Lectures 1 and 2: Fundamentals of Machine Learning and Hypothesis Testing187

34 Lectures 3 and 4: Hypothesis Testing and Statistical Decisions 193


34.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 193
34.2 Model Formulation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 193
34.2.1 Objective . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 193
34.3 Construction of Critical Region . . . . . . . . . . . . . . . . . . . . . . . . . . . . 193
34.4 Derivation of Critical Value . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 194
34.5 Critical Values for Common α . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 194
34.5.1 α = 0.01 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 194
34.5.2 α = 0.05 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 194
34.5.3 α = 0.10 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 194
34.5.4 Observation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 194
34.6 Diagram: 5% Rejection Region . . . . . . . . . . . . . . . . . . . . . . . . . . . . 195
34.7 Example: Observed Value x = 15 . . . . . . . . . . . . . . . . . . . . . . . . . . . 195
34.8 p-value Derivation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 195
34.8.1 Interpretation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 195
34.9 Diagram: p-value . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 196
34.10Power of the Test . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 196
34.11Diagram: Power of the Test . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 196
34.12Conceptual Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 197
34.13Statistical Decision Framework . . . . . . . . . . . . . . . . . . . . . . . . . . . . 197
34.14Decision Rule and Rejection Region . . . . . . . . . . . . . . . . . . . . . . . . . 197
CONTENTS xiii

34.15Type I and Type II Errors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 197


34.15.1 Type I Error . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 197
34.15.2 Type II Error . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 198
34.15.3 Power . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 198
34.16Simulation Interpretation of Type I Error . . . . . . . . . . . . . . . . . . . . . . 198
34.17Type II Error Derivation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 198
34.17.1 Type II Error when α = 0.01 . . . . . . . . . . . . . . . . . . . . . . . . . 199
34.17.2 Power of the Test Calculation . . . . . . . . . . . . . . . . . . . . . . . . . 199
34.18Graphical Illustration of β and Power . . . . . . . . . . . . . . . . . . . . . . . . 199
34.19One-Sided Test . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 200
34.20Two-Sided Test . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 200
34.21Comparison: One-Sided vs Two-Sided . . . . . . . . . . . . . . . . . . . . . . . . 200
34.22Summary of the Complete Testing Procedure . . . . . . . . . . . . . . . . . . . . 200
34.23Two-Sided Hypothesis Test for Population Mean . . . . . . . . . . . . . . . . . . 201
34.23.1 Problem Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 201
34.23.2 Population Assumption . . . . . . . . . . . . . . . . . . . . . . . . . . . . 201
34.23.3 Test Statistic . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 201
34.23.4 Significance Level and Critical Region . . . . . . . . . . . . . . . . . . . . 201
34.23.5 Graphical Representation . . . . . . . . . . . . . . . . . . . . . . . . . . . 202
34.23.6 Theoretical Interpretation . . . . . . . . . . . . . . . . . . . . . . . . . . . 202
34.23.7 One Realization of the Random Variable . . . . . . . . . . . . . . . . . . . 202
34.23.8 Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 202
34.24A Note on Hypothesis ”Acceptance” . . . . . . . . . . . . . . . . . . . . . . . . . 203
34.24.1 Significance Levels . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 203
34.24.2 Null Hypothesis (H0 ) about θ . . . . . . . . . . . . . . . . . . . . . . . . . 203
34.25Decision Rules and Rejection Regions . . . . . . . . . . . . . . . . . . . . . . . . 203
34.25.1 Rejection Region (Critical Region) . . . . . . . . . . . . . . . . . . . . . . 203
34.25.2 Example: Testing H0 : θ = 1 . . . . . . . . . . . . . . . . . . . . . . . . . 203
34.25.3 Visualizing the Acceptance Region . . . . . . . . . . . . . . . . . . . . . . 204

35 Lectures 5 and 6: Hypothesis Testing for Population Mean 205


35.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 205
35.2 Statistical Hypotheses . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 205
35.3 Errors in Hypothesis Testing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 205
35.4 Testing the Mean of a Normal Population . . . . . . . . . . . . . . . . . . . . . . 206
35.5 Construction of the Test . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 206
35.6 Rejection Region . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 206
35.7 Acceptance Region . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 207
35.8 P-Value . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 207
35.9 Examples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 207
35.10Type II Error and OC Curve . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 207
35.11Statistical Hypothesis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 208
35.12Testing Equality of Two Population Means . . . . . . . . . . . . . . . . . . . . . 208
35.13Test Statistic . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 208
35.14Acceptance and Rejection Regions . . . . . . . . . . . . . . . . . . . . . . . . . . 208
35.15Types of Errors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 209
35.16Decision Table . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 209
35.17Power of a Test . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 209
35.18Likelihood Ratio Test . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 209
35.19Large Sample Z Test . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 209
35.20Right-Tailed Test . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 210
xiv CONTENTS

35.21Left-Tailed Test . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 210


35.22Two-Tailed Test . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 210
35.23Introduction to Hypothesis Testing . . . . . . . . . . . . . . . . . . . . . . . . . . 210
35.23.1 Null vs. Alternative Hypotheses . . . . . . . . . . . . . . . . . . . . . . . 210
35.24The Decision Framework . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 210
35.24.1 Errors in Decision Making . . . . . . . . . . . . . . . . . . . . . . . . . . . 211
35.25Likelihood Ratio Test (LRT) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 211

36 Lectures 7 and 8: Likelihood Ratio Test and Large Sample Z Test 213
36.1 Simple and Composite Alternative Hypotheses . . . . . . . . . . . . . . . . . . . 213
36.2 Likelihood Ratio Test Statistic . . . . . . . . . . . . . . . . . . . . . . . . . . . . 213
36.3 Level of Significance . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 214
36.4 Further Remarks on Likelihood Ratio Test . . . . . . . . . . . . . . . . . . . . . . 214
36.5 Example Setup (i.i.d Sample) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 215
36.6 Likelihood Ratio in this Case . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 215
36.7 Likelihood Function and MLE (Detailed) . . . . . . . . . . . . . . . . . . . . . . 216
36.8 Derivation of Likelihood Ratio (Final Form) . . . . . . . . . . . . . . . . . . . . . 218
36.9 Likelihood Ratio and Rejection Region (Detailed) . . . . . . . . . . . . . . . . . . 219
36.10Recall . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 220
36.11Final Result . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 221

37 Lectures 9 and 10: Two-Population Inference and Error Probabilities 231

38 Lectures 11 and 12: Linear Algebra Foundation - Subspaces and Projections243


38.1 Motivating Example: Why Do Some Systems Have Infinite Solutions? . . . . . . 243
38.1.1 Matrix Form . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 243
38.1.2 Geometric Interpretation . . . . . . . . . . . . . . . . . . . . . . . . . . . 244
38.2 The Null Space . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 244
38.3 The Four Fundamental Subspaces . . . . . . . . . . . . . . . . . . . . . . . . . . . 244
38.3.1 Worked Example . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 245
38.3.2 Rank . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 245
38.4 Orthogonality of the Fundamental Subspaces . . . . . . . . . . . . . . . . . . . . 245
38.4.1 Key Orthogonality Result . . . . . . . . . . . . . . . . . . . . . . . . . . . 245
38.5 Existence and Uniqueness of Solutions to A⃗x = ⃗b . . . . . . . . . . . . . . . . . . 246
38.6 Projections . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 246
38.6.1 2D Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 246
38.6.2 Numerical Example . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 247
38.7 Properties of the Projection Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . 247
38.7.1 Eigenvalues of P . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 248
38.8 Projection onto a Column Space . . . . . . . . . . . . . . . . . . . . . . . . . . . 248
38.8.1 Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 248
38.8.2 Derivation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 248
38.8.3 Special Case: A is Square and Invertible . . . . . . . . . . . . . . . . . . . 248
38.9 Linear Regression (Finally!) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 249
38.9.1 Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 249
38.9.2 Matrix Formulation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 249
38.10Quick Reference: Independence . . . . . . . . . . . . . . . . . . . . . . . . . . . . 249

39 Lectures 13 and 14: Projection and Regression 251


CONTENTS xv

40 Lectures 15 and 16: Fisher Information, MLE, and Regression 261


40.1 Score Function and Fisher Information . . . . . . . . . . . . . . . . . . . . . . . . 261
40.1.1 Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 261
40.1.2 Properties of the Score Function . . . . . . . . . . . . . . . . . . . . . . . 261
40.1.3 Alternative Form of Fisher Information . . . . . . . . . . . . . . . . . . . 262
40.1.4 Fisher Information for a Sample of Size n . . . . . . . . . . . . . . . . . . 262
40.2 Cramér–Rao Lower Bound (CRLB) . . . . . . . . . . . . . . . . . . . . . . . . . . 262
40.2.1 Statement of the CRLB . . . . . . . . . . . . . . . . . . . . . . . . . . . . 262
40.2.2 Proof of the CRLB . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 263
40.2.3 Correlation Coefficient Form . . . . . . . . . . . . . . . . . . . . . . . . . 263
40.2.4 Special Case: Estimating θ Itself . . . . . . . . . . . . . . . . . . . . . . . 263
40.3 Unbiased Estimators and Efficiency . . . . . . . . . . . . . . . . . . . . . . . . . . 263
40.3.1 Unbiased Estimator of µ for Normal Distribution . . . . . . . . . . . . . . 263
40.3.2 CRLB for the Normal Mean . . . . . . . . . . . . . . . . . . . . . . . . . . 264
40.3.3 Efficiency of the Sample Mean . . . . . . . . . . . . . . . . . . . . . . . . 264
40.4 Maximum Likelihood Estimation (MLE) . . . . . . . . . . . . . . . . . . . . . . . 264
40.4.1 Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 264
40.4.2 Asymptotic Property of MLE . . . . . . . . . . . . . . . . . . . . . . . . . 265
40.5 Probabilistic Interpretation of Linear Regression . . . . . . . . . . . . . . . . . . 265
40.5.1 Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 265
40.5.2 Error Model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 265
40.5.3 Distribution of the Model . . . . . . . . . . . . . . . . . . . . . . . . . . . 265
40.5.4 PDF of the Model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 265
40.5.5 MLE for Linear Regression . . . . . . . . . . . . . . . . . . . . . . . . . . 266
40.6 Least Squares and the Normal Equation . . . . . . . . . . . . . . . . . . . . . . . 266
40.6.1 The Cost Function . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 266
40.6.2 Gradient of the Cost . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 266
40.6.3 Normal Equation (Closed-Form Solution) . . . . . . . . . . . . . . . . . . 267
40.6.4 Numerical Example . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 267
40.7 Batch Gradient Descent for Linear Regression . . . . . . . . . . . . . . . . . . . . 267
40.7.1 Gradient Descent Update Rule . . . . . . . . . . . . . . . . . . . . . . . . 267
40.7.2 Expanded Batch Gradient Descent . . . . . . . . . . . . . . . . . . . . . . 267
40.7.3 Stochastic Gradient Descent (SGD) . . . . . . . . . . . . . . . . . . . . . 267
40.8 Logistic Regression — Binary Classification . . . . . . . . . . . . . . . . . . . . . 268
40.8.1 Problem Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 268
40.8.2 Sigmoid (Logistic) Function . . . . . . . . . . . . . . . . . . . . . . . . . . 268
40.8.3 Probabilistic Interpretation . . . . . . . . . . . . . . . . . . . . . . . . . . 269
40.8.4 Likelihood Function for Logistic Regression . . . . . . . . . . . . . . . . . 269
40.8.5 Gradient of the Log-Likelihood . . . . . . . . . . . . . . . . . . . . . . . . 269
40.8.6 Gradient Ascent for Logistic Regression . . . . . . . . . . . . . . . . . . . 270
40.8.7 Stochastic and Batch Gradient Descent for Logistic Regression . . . . . . 270
40.8.8 Stochastic Gradient Descent Algorithm . . . . . . . . . . . . . . . . . . . 271
40.9 Summary of Key Formulae . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 271

41 Lectures 17 and 18: Markov Chains and Naive Bayes Classification 273
41.1 Introduction to Stochastic Processes . . . . . . . . . . . . . . . . . . . . . . . . . 273
41.2 Transition Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 273
41.2.1 Weather Prediction Model . . . . . . . . . . . . . . . . . . . . . . . . . . . 273
41.2.2 Non-Markovian Counter-example . . . . . . . . . . . . . . . . . . . . . . . 274
41.2.3 Two-Step Transition Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . 274
41.3 Gambler’s Ruin Problem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 274
xvi CONTENTS

41.3.1 Problem Statement . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 274


41.3.2 Transition Probabilities . . . . . . . . . . . . . . . . . . . . . . . . . . . . 275

42 Naive Bayes Classifier 277


42.1 Introduction and Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 277
42.1.1 Notation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 277
42.2 The Classification Goal . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 277
42.2.1 Bayes’ Theorem for Classification . . . . . . . . . . . . . . . . . . . . . . . 277
42.3 The Generative Model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 278
42.3.1 Distribution of X⃗ and Tractability . . . . . . . . . . . . . . . . . . . . . . 278
42.4 Worked Example: Spam Classifier . . . . . . . . . . . . . . . . . . . . . . . . . . 278
42.4.1 Dataset Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 278
42.4.2 Training Data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 278
42.4.3 Class-Conditional Likelihoods . . . . . . . . . . . . . . . . . . . . . . . . . 278
42.5 Quick Reference Notation Table . . . . . . . . . . . . . . . . . . . . . . . . . . . . 279

43 Lectures 19 and 20: Asymptotic Dynamics of Markov Chains 281


43.1 Introduction to Finite State Markov Chains . . . . . . . . . . . . . . . . . . . . . 281
43.2 Accessibility and Communication of States . . . . . . . . . . . . . . . . . . . . . 281
43.3 Recurrent and Transient States . . . . . . . . . . . . . . . . . . . . . . . . . . . . 281
43.3.1 Recurrent States . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 282
43.3.2 Transient States . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 282
43.4 A Finer Classification: Positive and Null Recurrence . . . . . . . . . . . . . . . . 282
43.4.1 Positive Recurrent States . . . . . . . . . . . . . . . . . . . . . . . . . . . 282
43.4.2 Null Recurrent States . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 283
43.5 Periodicity and Aperiodicity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 283
43.5.1 Aperiodic States . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 283
43.6 Ergodic States and Ergodic Markov Chains . . . . . . . . . . . . . . . . . . . . . 283
43.7 Irreducibility . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 283
43.8 Limiting Probabilities and the Limiting Distribution . . . . . . . . . . . . . . . . 284
43.9 Stationary Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 284
43.10Mean Recurrence Time . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 285
43.11A Worked Example: Three-State Ergodic Chain . . . . . . . . . . . . . . . . . . 285
43.12The Two-State Markov Chain . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 285
43.13Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 286

44 Lectures 1 and 2: Introduction to Stochastic Processes 289


44.1 Lecture 1: Foundations of Stochastic Processes . . . . . . . . . . . . . . . . . . . 289
44.1.1 Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 289
44.1.2 Random Variable (RV) . . . . . . . . . . . . . . . . . . . . . . . . . . . . 289
44.1.3 Randomness . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 289
44.1.4 Example: Toss Coin Twice . . . . . . . . . . . . . . . . . . . . . . . . . . 289
44.1.5 Defining Different RVs on the Same Sample Space . . . . . . . . . . . . . 290
44.1.6 Classification of Stochastic Processes . . . . . . . . . . . . . . . . . . . . . 290
44.1.7 Sample Paths and Trajectories . . . . . . . . . . . . . . . . . . . . . . . . 290
44.1.8 Filtrations and Paths . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 291
44.1.9 The Power Set and Realizations . . . . . . . . . . . . . . . . . . . . . . . 292
44.2 Lecture 2: Probability and State Spaces . . . . . . . . . . . . . . . . . . . . . . . 292
44.2.1 Stochastic Process Formal Definition . . . . . . . . . . . . . . . . . . . . . 292
44.2.2 Probability Axioms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 292
44.2.3 Example 1: Discrete Sample Space . . . . . . . . . . . . . . . . . . . . . . 293
44.2.4 Example 2: Continuous Random Numbers . . . . . . . . . . . . . . . . . . 293
CONTENTS xvii

44.2.5 Conditional Probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . 293

45 Lectures 3 and 4: Bernoulli Processes and Arrival Times 295


45.1 Introduction to the Bernoulli Process . . . . . . . . . . . . . . . . . . . . . . . . . 295
45.2 System Modeling and Discretization . . . . . . . . . . . . . . . . . . . . . . . . . 295
45.2.1 Parameters . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 295
45.2.2 Assumptions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 295
45.3 The Random Variable and Distribution . . . . . . . . . . . . . . . . . . . . . . . 295
45.4 Mathematical Formulations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 296
45.4.1 Probability Mass Function (PMF) . . . . . . . . . . . . . . . . . . . . . . 296
45.4.2 Statistical Moments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 296
45.5 Summary of Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 296
45.6 Waiting Time for a Job . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 296
45.6.1 Defining the Random Variable T1 . . . . . . . . . . . . . . . . . . . . . . . 296
45.7 The Geometric Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 297
45.7.1 Probability Mass Function (PMF) . . . . . . . . . . . . . . . . . . . . . . 297
45.7.2 Visualizing the Process . . . . . . . . . . . . . . . . . . . . . . . . . . . . 297
45.8 Statistical Properties . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 297
45.9 Example: Lottery Ticket String of Losses . . . . . . . . . . . . . . . . . . . . . . 297
45.9.1 Scenario Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 297
45.9.2 The Geometric Constraint . . . . . . . . . . . . . . . . . . . . . . . . . . . 298
45.10Introduction to Total Waiting Time . . . . . . . . . . . . . . . . . . . . . . . . . 298
45.11The Concept of the Losing Streak . . . . . . . . . . . . . . . . . . . . . . . . . . . 298
45.12Derivation of the Negative Binomial Distribution . . . . . . . . . . . . . . . . . . 298
45.12.1 Logical Breakdown . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 298
45.12.2 Probability Mass Function (PMF) . . . . . . . . . . . . . . . . . . . . . . 298
45.13Waiting Time as a Sum of Geometric Variables . . . . . . . . . . . . . . . . . . . 299
45.14Total Waiting Time for k Arrivals . . . . . . . . . . . . . . . . . . . . . . . . . . 299
45.14.1 The Random Variable Wk . . . . . . . . . . . . . . . . . . . . . . . . . . . 299
45.14.2 Expectation and Variance . . . . . . . . . . . . . . . . . . . . . . . . . . . 299
45.15Splitting a Bernoulli Process . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 299
45.15.1 The Mechanism . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 300
45.15.2 Visual Representation of Splitting . . . . . . . . . . . . . . . . . . . . . . 300
45.15.3 Sub-Process Probabilities . . . . . . . . . . . . . . . . . . . . . . . . . . . 300
45.16Merging of Bernoulli Processes . . . . . . . . . . . . . . . . . . . . . . . . . . . . 300
45.17Tree Diagram Approach: Conditional Probabilities . . . . . . . . . . . . . . . . . 300
45.18Merging Independent Streams . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 301
45.18.1 Individual Process Definitions . . . . . . . . . . . . . . . . . . . . . . . . . 301
45.18.2 The Merged Probability (pM W ) . . . . . . . . . . . . . . . . . . . . . . . . 301
45.18.3 Simplified Merged Equation . . . . . . . . . . . . . . . . . . . . . . . . . . 301
45.19Visual Interpretation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 301
45.20Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 301
45.21Alternative Derivation of Merged Probability . . . . . . . . . . . . . . . . . . . . 301
45.21.1 Algebraic Proof . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 302
45.22Continuous Time: The Exponential Distribution . . . . . . . . . . . . . . . . . . 302
45.22.1 Probability Density Function (PDF) . . . . . . . . . . . . . . . . . . . . . 302
45.22.2 Cumulative Distribution Function (CDF) . . . . . . . . . . . . . . . . . . 302
45.23Summary of Waiting Times . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 302
xviii CONTENTS

46 Lectures 5 and 6: Geometric and Exponential Distributions 303


46.1 Lecture 5: Sums of Geometric Random Variables . . . . . . . . . . . . . . . . . . 303
46.1.1 Mean and Variance . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 303
46.2 Bernoulli Processes: Splitting and Merging . . . . . . . . . . . . . . . . . . . . . 303
46.2.1 1. Splitting a Bernoulli Process . . . . . . . . . . . . . . . . . . . . . . . . 303
46.2.2 2. Merging Bernoulli Processes . . . . . . . . . . . . . . . . . . . . . . . . 304
46.3 The Exponential Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 304
46.3.1 1. Probability Density Function (PDF) . . . . . . . . . . . . . . . . . . . 304
46.3.2 2. Cumulative Distribution Function (CDF) . . . . . . . . . . . . . . . . . 304
46.3.3 3. Expectation (Mean) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 305
46.4 Lecture 6: Further Properties of the Exponential Distribution . . . . . . . . . . . 305
46.4.1 Moments and Variance . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 305
46.4.2 The Memoryless Property . . . . . . . . . . . . . . . . . . . . . . . . . . . 305
46.5 Competing Risks: Time to First Event . . . . . . . . . . . . . . . . . . . . . . . . 306
46.5.1 (i) Distribution of the Minimum: T = min(T1 , T2 ) . . . . . . . . . . . . . 306
46.5.2 (ii) Distribution of the Maximum: T = max(T1 , T2 ) . . . . . . . . . . . . . 306
46.5.3 (iii) Probability of T1 occurring before T2 . . . . . . . . . . . . . . . . . . 306
46.6 Small Interval Approximations . . . . . . . . . . . . . . . . . . . . . . . . . . . . 307
46.6.1 Probability of Two Events in a Small Interval h . . . . . . . . . . . . . . . 307

47 Lecture 9: Poisson Distribution Derivation 315


47.1 Poisson Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 315
47.1.1 Probability Mass Function (PMF) . . . . . . . . . . . . . . . . . . . . . . 315
47.1.2 Key Properties . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 315
47.1.3 Assumptions of Poisson Distribution . . . . . . . . . . . . . . . . . . . . . 315
47.1.4 Relation with Binomial Distribution . . . . . . . . . . . . . . . . . . . . . 316
47.1.5 Cumulative Probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . 316
47.1.6 Applications . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 316

48 Lecture 11 and 12 329

49 Lecture 13 and 14: Poisson Process Properties 339


49.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 339

50 Discrete Time Markov Chains (DTMC) 347


50.1 Introduction to DTMC . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 353
50.2 State Space . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 354
50.3 Transition Probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 354
50.4 Transition Probability Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 354
50.5 Initial Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 355
50.6 State Distribution at Time k . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 355
50.7 Two-Step Transition Probability . . . . . . . . . . . . . . . . . . . . . . . . . . . 355
50.8 n-Step Transition Probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 356
50.9 State Evolution Formula . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 356
50.10Chapman-Kolmogorov Equation . . . . . . . . . . . . . . . . . . . . . . . . . . . 356
50.11Classification of States . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 357
50.12Final Intuition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 357
CONTENTS xix

51 Lecture 18: Markov Chains in matrix form 359


51.0.1 Introduction and Notation . . . . . . . . . . . . . . . . . . . . . . . . . . . 364
51.0.2 Transition Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 365
51.0.3 Transition Probabilities . . . . . . . . . . . . . . . . . . . . . . . . . . . . 365
51.0.4 Transition Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 366
51.0.5 Non-Markovian Counter-example from Notes . . . . . . . . . . . . . . . . 366
51.0.6 Two-Step Transition Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . 366
51.1 Numerical 2 — Gambler’s Ruin Problem . . . . . . . . . . . . . . . . . . . . . . . 367
51.1.1 Problem Statement . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 367
51.1.2 Transition Probabilities . . . . . . . . . . . . . . . . . . . . . . . . . . . . 367
51.1.3 Transition Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 367
51.2 Naive Bayes Classifier . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 367
51.2.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 367
51.2.2 Setup and Notation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 367
51.2.3 The Classification Goal . . . . . . . . . . . . . . . . . . . . . . . . . . . . 368
51.2.4 Bayes’ Theorem (Derivation) . . . . . . . . . . . . . . . . . . . . . . . . . 368
51.2.5 Generative Model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 369
⃗ . . . . . . . . . . . . . . . .
51.2.6 Distribution of X . . . . . . . . . . . . . . . . 369
51.2.7 Worked Example — Spam Classifier . . . . . . . . . . . . . . . . . . . . . 369
51.2.8 Parameter Counting — Tractability Analysis . . . . . . . . . . . . . . . . 371
51.3 Quick Reference — Notation Table . . . . . . . . . . . . . . . . . . . . . . . . . . 371
xx CONTENTS
Part I

Probability

1
Chapter 1

Lecture 1

1.1 The Basics of Probability


Probability is simply the mathematics of chance. It’s a way to measure how likely it is that
something will happen. To talk about probability, we need to know two key things: the Sample
Space and an Event.

1.2 Sample Space and Events


Think of it like ordering from a menu.

ˆ Sample Space (S): This is the list of all possible outcomes that can happen. It’s like
the entire menu at a restaurant. We use curly braces {} to list the outcomes. The total
number of outcomes is called the size of the sample space, written as |S|.

ˆ Event (E): An event is a specific outcome or a set of outcomes that we are interested in.
It’s a subset of the sample space. This is like choosing what you want to eat from the
menu.

Example: Tossing Two Coins If we toss two different coins, say a penny and a nickel,
what are all the possible results? Let ’H’ stand for Heads and ’T’ for Tails.
The Sample Space is all possible combinations:

S = {HH, HT, T H, T T } (1)

There are 4 possible outcomes, so |S| = 4.


Now, let’s define an event. Let’s say we are interested in the event where we get at least
one Head. We’ll call this Event E.
The outcomes from our sample space that fit this description are:

E = {HH, HT, T H} (2)

There are 3 outcomes in this event, so |E| = 3.


The probability of Event E happening is calculated as:

Number of outcomes in the event |E| 3


P (E) = = = (3)
Total number of outcomes in the sample space |S| 4

So, there’s a 75% chance of getting at least one head when you toss two coins.

3
4 CHAPTER 1. LECTURE 1

⇒ RANDOM experiment
RISK


SEVERAL OUTCOMES
W1 , W2 , . . . , Wn , . . .
Set of all possible outcomes


Sample Space of the Random experiment:
Ω = {W1 , W2 , . . . , Wn }


Events: Subset ⊆ Ω in the sample space
E1 ⊆ Ω, E2 ⊆ Ω

Definition: A random experiment is one whose outcome cannot be predicted with


certainty (e.g., tossing a coin, rolling a die). - Each possible result of the experiment is called
an outcome, denoted by W1 , W2 , . . .. - The collection of all outcomes is the sample space Ω.
- An event is a subset of the sample space, which groups together certain outcomes based on
a property of interest.
Define Events: Ei ⊆ Ω that satisfy certain property

Example:
ˆ Outcome has at least one heads:
EH = {H1 H2 , H1 T2 , T1 H2 } ⊆ Ω

ˆ Outcome has head on the first toss:


EH1 = {H1 H2 , H1 T2 } ⊆ Ω
Events: An event is a subset of the sample space Ω consisting of outcomes that satisfy a
particular property. - For example: - EH includes all outcomes where at least one coin shows
H. - EH1 includes all outcomes where the first toss results in H. - In general, any subset of Ω is
an event. - Special cases: - The empty set ∅ is an event (represents the “impossible event”).
- The entire sample space Ω is also an event (represents the “certain event”). - Thus, events
help us represent conditions or situations in terms of outcomes of the experiment.

Coin Toss model is an important model in the Binomial model which we’ll be discussing
in the future.
When there is more than 1 outcome, then you’ll generate “Risk”.

A Valid Probability P defined in Events E ⊆ Ω:

0 ≤ P (E) ≤ 1 (4)

1.3 Axioms of Probability


The axioms of probability are the fundamental rules that any probability measure P must satisfy.
They were first formalized by Andrey Kolmogorov and serve as the foundation of probability
theory.
1.4. EXAMPLE 5

1. Non-negativity: For any event E ⊆ Ω,

0 ≤ P (E) ≤ 1 (5)

This means that probabilities can never be negative and cannot exceed 1. - Example:
When tossing a coin, P (Head) = 0.5, which lies between 0 and 1.

2. Normalization: The probability of the entire sample space is always equal to 1:

P (Ω) = 1 (6)

This reflects the fact that something in the sample space must occur. - Example: For a
dice roll, Ω = {1, 2, 3, 4, 5, 6}, so

P ({1, 2, 3, 4, 5, 6}) = 1 (7)

3. Additivity (Countable Additivity): If E1 , E2 , . . . , En , . . . ⊆ Ω are pairwise disjoint


events (i.e., no two of them overlap), then:
∞ ∞
!
[ X
P Ei = P (Ei ) (8)
i=1 i=1


X
⇒ P (E1 ∪ E2 ∪ . . . ∪ En ∪ . . .) = P (Ei ) (9)
i=1

Where:
Ek ∩ Ej = ∅ for k ̸= j (10)
Ek ⊆ Ω, Ej ⊆ Ω (11)

In words, the probability of the union of mutually exclusive events is the sum of their
probabilities.
- Example: In a dice roll, let E1 = {1} and E2 = {2}. Since E1 ∩ E2 = ∅, we have:
1 1 2
P (E1 ∪ E2 ) = P (E1 ) + P (E2 ) = 6 + 6 = 6 (12)

These three axioms form the basis of probability theory. All other results, such as conditional
probability, Bayes’ theorem, and distributions, can be derived from these axioms.

1.4 Example
Ω = {1, 2, 3, 4, 5, 6, 7, 8, 9} (13)
A = {1, 2, 3}, B = {3, 4, 5}, C = {7, 8} (14)
A ⊂ Ω, B ⊂ Ω, C⊂Ω (15)
|A| = 3, |B| = 3, |C| = 2 (16)
Complement of A:
Ac = {w ∈ Ω | w ∈
/ A} = {4, 5, 6, 7, 8, 9} (17)

Venn Diagram:
6 CHAPTER 1. LECTURE 1

B
1,2 A 3 4,5 C 7,8

1.5 Set Theory


1.5.1 Definition of a Set
A set is a well-defined collection of distinct objects, considered as an object in its own right.
The objects inside a set are called elements or members.

1.5.2 Types of Sets


ˆ Finite Set: A set with a finite number of elements. Example: {a, b, c}.
ˆ Infinite Set: A set with infinitely many elements. Example: N = {1, 2, 3, . . . }.
ˆ Empty Set (Null Set): A set with no elements. Example: ∅ = {}.
ˆ Singleton Set: A set with exactly one element. Example: {5}.
ˆ Universal Set: The set containing all objects under consideration, usually denoted by
Ω or U .

1.6 Example (Set Theory)


Ω = {1, 2, 3, 4, 5, 6, 7, 8, 9} (18)
A = {1, 2, 3}, B = {3, 4, 5}, C = {7, 8} (19)
|A| = 3, |B| = 3, |C| = 2 (20)
Complement of A:
Ac = {4, 5, 6, 7, 8, 9} (21)
Union:
A ∪ B = {1, 2, 3, 4, 5} (22)
Intersection:
A ∩ B = {3} (23)
Difference:
A − B = {1, 2} (24)
Intersection of A and C:
A∩C =∅ (25)
1.6. EXAMPLE (SET THEORY) 7

1.6.1 Venn Diagram


A Venn diagram is a pictorial way to represent sets and their relationships. In probability, the
sample space Ω contains all possible outcomes. Subsets of Ω (such as A and B) represent events.
In the diagram below: - A − B represents the outcomes that belong to A but not to B. -
B − A represents the outcomes that belong to B but not to A. - A ∩ B represents the common
outcomes of both A and B (the intersection).

A B

A−B A∩B B−A

1.6.2 Set Containment


If one event F is completely inside another event E, we say that F is contained in E and write:

F ⊆E (26)

This means that every outcome in F is also an outcome in E.


In set-builder notation:
E = {w ∈ Ω | w ∈ F ⇒ w ∈ E} (27)

1.6.3 Probability Calculations


The probability of an event E is related to its complement E c (the set of outcomes not in E).
Since either E or E c must occur:
P (E) = 1 − P (E c ) (28)

1.6.4 For both uniform and non-uniform sample space


In a uniform sample space, all outcomes are equally likely. If E ⊆ Ω:

|E|
P (E) = (29)
|Ω|

Here, |E| is the number of favorable outcomes, and |Ω| is the total number of possible outcomes.
In a non-uniform sample space, outcomes may not be equally likely. In that case:
X
P (E) = P ({w}) (30)
w∈E

That is, we add up the probabilities of each outcome inside E.


If Ω = {ω1 , ω2 , . . . , ωn }, then each probability is written as:

P ({ωi }) = P (w) (31)

This allows us to calculate probabilities in both uniform (equal likelihood) and non-uniform
(unequal likelihood) cases.
8 CHAPTER 1. LECTURE 1
Chapter 2

Lecture 2

2.1 Combinatorics: The Art of Counting


Before we can find probabilities, we often need to count the number of ways things can happen.
This is where Permutations and Combinations (P&C) come in. The most important question
in P&C is: Does the order matter?

2.1.1 Permutations: When Order Matters


A permutation is an arrangement of items where the order is important.

Analogy: The Race


Imagine 10 people run a race. The first three finishers win Gold, Silver, and Bronze medals.
Does the order they finish in matter? Yes! Alice finishing 1st, Bob 2nd, and Charlie 3rd is very
different from Charlie 1st, Alice 2nd, and Bob 3rd.
The formula for calculating the number of permutations of selecting k items from a set of n
items is:
n!
n Pk = (32)
(n − k)!
The ‘¡ symbol is called a factorial. For example, 5! = 5 × 4 × 3 × 2 × 1.

Derivation
Let’s say we want to pick and arrange 3 people (k = 3) from a group of 10 (n = 10).

ˆ For the 1st position (Gold), we have 10 choices.


ˆ For the 2nd position (Silver), we have 9 remaining choices.
ˆ For the 3rd position (Bronze), we have 8 remaining choices.
Total arrangements = 10 × 9 × 8 = 720.
This is what the formula does:
(
(
10! 10! 10 × 9 × 8 × (
7× (6 (×(·(· ·(
×1
10 P3 = = = ((( = 10 × 9 × 8 = 720 (33)
(10 − 3)! 7! 7
(( × 6
( ×
( ·
( · · × 1

2.1.2 Combinations: When Order Doesn’t Matter


A combination is a selection of items where the order is NOT important.

9
10 CHAPTER 2. LECTURE 2

Analogy: The Pizza Toppings You’re ordering a pizza and can choose 3 toppings from
a list of 10. Does it matter if you choose ”pepperoni, mushrooms, and olives” versus ”olives,
pepperoni, and mushrooms”? No! The final pizza is the same.
The formula for calculating the number of combinations of selecting k items from a set of n
items is:  
n n!
n Ck = = (34)
k k!(n − k)!

Derivation
The key idea is that a combination is just a permutation where we remove the overcounting
caused by order. For our 3 pizza toppings, there are 3! = 3 × 2 × 1 = 6 ways to arrange them.
A permutation counts all 6 of these as different, but a combination counts them as one.
So, to get the number of combinations, we take the number of permutations and divide by
the number of ways to arrange the selected items (k!).

n Pk n!/(n − k)! n!
n Ck = = = (35)
k! k! k!(n − k)!

For the pizza example:


 
10 10! 10! 10 × 9 × 8
10 C3 = = = = = 120 (36)
3 3!(10 − 3)! 3!7! 3×2×1

There are 120 different 3-topping pizzas you can make.

2.2 Solved Examples


2.2.1 Problem 1: Distributing Teachers to Schools
Question: 8 distinct teachers are to be divided among 4 distinct schools.

Part (i): How many ways can this be done with no conditions?
This is a problem about choices. Let’s think from the perspective of each teacher.

ˆ Teacher 1: Has 4 schools they can be assigned to. (4 choices)


ˆ Teacher 2: Also has 4 schools they can be assigned to. (4 choices)
ˆ ...and so on, up to Teacher 8.
Since each teacher’s assignment is an independent choice, we multiply the number of choices
together.
Total Ways = 4 × 4 × 4 × 4 × 4 × 4 × 4 × 4 = 48 = 65, 536 (37)
There are 65,536 ways to distribute the 8 teachers among the 4 schools.

Part (ii): If each school must get at least 1 teacher.


This is a more complex problem. It’s easier to calculate the opposite of what we want and
subtract it from the total. The opposite of ”at least one teacher per school” is ”one or more
schools are empty”.
We use the Principle of Inclusion-Exclusion (PIE) for this. The basic idea is: Total
Ways = (Total) - (Ways where at least 1 school is empty) + (Ways where at least 2 are empty)
- (Ways where at least 3 are empty) + ...
2.2. SOLVED EXAMPLES 11

1. Total Ways: We already calculated this as 48 .

2. Subtract ways where at least ONE school is empty:

ˆFirst, choose which school is empty: 41 ways.




ˆ Then, distribute the 8 teachers among the remaining 3 schools: 3 8 ways.


ˆ Total to subtract: × 3 = 4 × 6561 = 26244.
4
1
8

3. Add back ways where at least TWO schools are empty: (We subtracted these
cases twice in the last step, so we add them back once).

ˆ Choose which 2 schools are empty: 42 ways.




ˆ Distribute the 8teachers among the remaining 2 schools: 2 8 ways.


ˆ Total to add: × 2 = 6 × 256 = 1536.
4
2
8

4. Subtract ways where at least THREE schools are empty:

ˆChoose which 3 schools are empty: 43 ways.




ˆ Distribute the 8 teachers among the remaining 1 school: 1 8 ways.


ˆ Total to subtract: × 1 = 4 × 1 = 4.
4

3
8

(We can’t have 4 empty schools, as we have teachers to distribute).


The final calculation is:
       
4 8 4 8 4 8 4 8
Ways = 4 − 3 + 2 − 1 (38)
0 1 2 3

Ways = 1 × 65536 − 4 × 6561 + 6 × 256 − 4 × 1 (39)


Ways = 65536 − 26244 + 1536 − 4 = 40824 (40)

Part (iii): How many ways to assign exactly 2 teachers to each school?
This question is about partitioning. We are splitting 8 distinct teachers into 4 distinct groups
(the schools), with each group having exactly 2 members. Here’s how we can count this:

ˆ For the first


 school, we need to choose 2 teachers from the 8 available. The number of
8
ways is 2 .

ˆ For the second school, we


 now have 6 teachers left. We need to choose 2 from these 6.
6
The number of ways is 2 .

ˆ For the third school, we have 4 teachers left. We choose 2. The number of ways is 4

2 .

ˆ For the fourth school, we have 2 teachers left. We choose 2. The number of ways is 2

2 .

To get the total number of ways, we multiply the possibilities for each step:
       
8 6 4 2
Total Ways = × × × (41)
2 2 2 2

Let’s expand this using the formula nk = k!(n−k)!


n!

:

8! 6! 4! 2!
= × × × (42)
2!6! 2!4! 2!2! 2!0!
12 CHAPTER 2. LECTURE 2

Notice that many terms cancel out (like the 6! and 4!):
8! 6! 4! 2!
= × × × (43)
2!6! 
  2!4! 2!2! 2!0!
We are left with (since 0! = 1):
8! 40320 40320
Total Ways = = = = 2520 (44)
2! · 2! · 2! · 2! 2·2·2·2 16

2.2.2 Problem 2: Distributing Indistinguishable Objects (Stars and Bars)


Question: How many ways can you distribute n indistinguishable objects among k distinct
groups such that each group has at least one object?
This is a classic problem that can be solved with a beautiful method called ”Stars and Bars”.

The Analogy Imagine the n objects are ”stars” (∗). To divide them into k groups, you need
k − 1 ”bars” (|).
n = 6 objects (stars) and k = 3 groups. We need k − 1 = 2 bars to divide them.
Consider this arrangement:
∗| ∗ ∗| ∗ ∗∗ (45)
This represents: Group 1 gets 1 star, Group 2 gets 2 stars, and Group 3 gets 3 stars. (1+2+3 =
6)

The Logic Since every group must have at least one object, we can imagine our 6 stars laid
out in a row with spaces between them:
∗ ∗ ∗ ∗ ∗ ∗ (46)
There are n − 1 = 5 possible spaces where we can place our bars. To create 3 groups, we need
to choose 2 of these spaces to place our bars. The number of ways to do this is:
     
n−1 6−1 5
= = (47)
k−1 3−1 2
 
5 5! 5×4
= = = 10 (48)
2 2!(5 − 2)! 2×1
There are 10 ways to do this.

General Formula The problem is equivalent to finding the number of positive integer solu-
tions to the equation:
x1 + x2 + · · · + xk = n, where xi ≥ 1 (49)
Let’s make a substitution to simplify the condition. Let yi = xi − 1. Since xi ≥ 1, we know
that yi ≥ 0. Substituting xi = yi + 1 into the equation:
(y1 + 1) + (y2 + 1) + · · · + (yk + 1) = n (50)
y1 + y2 + · · · + yk + k = n (51)
y1 + y2 + · · · + yk = n − k (52)
Now we are distributing n − k items into k groups where the groups are allowed to be empty
(yi ≥ 0). This is a standard stars and bars problem with n − k ”stars” and k − 1 ”bars”. The
total number of arrangements is:
   
(n − k) + (k − 1) n−1
= (53)
k−1 k−1
2.3. THE MULTINOMIAL COEFFICIENT 13

2.3 The Multinomial Coefficient


Suppose
Pr that n distinct objects are to be divided into r groups of sizes n1 , n2 , . . . , nr with
i=1 ni = n. The number of ways of doing this is given by the multinomial coefficient:
 
n n!
= . (54)
n1 , n2 , . . . , nr n1 !n2 ! · · · nr !

Example 2.3.1. Distribute 8 teachers among 4 schools with each school receiving 2 teachers.
The number of possible assignments is
8!
= 2520. (55)
2! 2! 2! 2!
14 CHAPTER 2. LECTURE 2
Chapter 3

Lecture 3

3.1 Stars and Bars: Basic Idea


Consider the problem of finding the number of nonnegative integer solutions to

x1 + x2 + · · · + xr = n. (56)

This is equivalent to distributing n indistinguishable objects into r distinct boxes.


The solution is given by the formula
 
n+r−1
. (57)
r−1

Stars and Bars Representation


For instance, the distribution (3, 1, 2) is represented as:

⋆ ⋆ ⋆| ⋆ | ⋆ ⋆ (58)

Stars represent the items to be distributed. Bars act as dividers to separate the items into
distinct groups or bins. The number of ways to distribute n identical items into k distinct bins
is given by the formula:  
n+k−1
(59)
k−1
This comes from choosing k − 1 positions for the bars among n + k − 1 total positions (stars
plus bars).

Example 3.1.1. How many nonnegative integer solutions are there to

x1 + x2 + x3 = 6? (60)

Solution: Here n = 6 and r = 3. The number of solutions is


   
6+3−1 8
= = 28. (61)
3−1 2

We can represent this combinatorial situation by arranging 6 stars (representing objects)


and 2 bars (dividers). Each arrangement corresponds to one solution.

⋆ ⋆ ⋆| ⋆ | ⋆ ⋆ (62)

⋆ ⋆ ⋆ — ⋆ — ⋆ ⋆

15
16 CHAPTER 3. LECTURE 3

This arrangement corresponds to (x1 , x2 , x3 ) = (3, 1, 2).

Positive Solutions. If we require xi > 0 for all i, then setting yi = xi − 1 ≥ 0, we obtain

y1 + y2 + · · · + yr = n − r. (63)

The number of positive solutions is therefore


 
n−1
. (64)
r−1

Example 3.1.2. How many positive integer solutions are there to

x1 + x2 + x3 = 6? (65)

Solution: Here n = 6, r = 3. The number of solutions is


   
6−1 5
= = 10. (66)
3−1 2

Example 3.1.3. Find the number of solutions in nonnegative integers to

x1 + x2 + x3 = 6. (67)

Here n = 6, k = 3, hence the answer is


   
6+3−1 8
= = 28. (68)
3−1 2

3.1.1 Distributions into Non-Empty Boxes


To account for the non-empty condition:

1. Place one item in each of the k boxes to satisfy the requirement that no box is empty.
This uses up k items.

2. Now, distribute the remaining n − k items into the k boxes with no restrictions (boxes
can receive zero or more items).

3. The number of ways to distribute n − k identical items into k distinct boxes is given by
the standard stars and bars formula:
   
(n − k) + k − 1 n−1
= (69)
k−1 k−1

This counts the number of non-negative integer solutions to y1 +y2 +· · ·+yk = n−k, where
yi represents the additional items in box i, and the total items in box i is xi = yi + 1 ≥ 1.

Conditions

ˆ This formula assumes n ≥ k, because you need at least k items to put one in each box.
ˆ If n < k, it’s impossible to have at least one item in each box, so the number of ways is 0.
3.1. STARS AND BARS: BASIC IDEA 17

Example 3.1.4. Distribute n = 6 indistinguishable objects into k = 3 non-empty boxes.

If each box must contain at least one object, let yi = xi − 1 ≥ 0. Then

y1 + y2 + · · · + yk = n − k. (70)

Thus the number of ways is  


n−1
. (71)
k−1
   
6−1 5
= = 10. (72)
3−1 2

⋆ ⋆ ⋆ — ⋆ — ⋆ ⋆

3.1.2 Distribution Allowing Empty Boxes


To distribute n identical items into k distinct boxes, allowing boxes to be empty, the number
of ways is given by the stars and bars formula:
 
n+k−1
(73)
k−1
This counts the number of non-negative integer solutions to x1 + x2 + · · · + xk = n, where xi ≥ 0
represents the number of items in box i. The formula arises from representing the n items as
stars and using k − 1 bars to separate them into k groups, choosing k − 1 positions for the bars
among n + k − 1 total positions
Example 3.1.5. Distribute n = 4 indistinguishable objects into k = 3 boxes (empty allowed).
Then    
4+3−1 6
= = 15. (74)
3−1 2

⋆ ⋆ — ⋆ — ⋆

This corresponds to the distribution (2, 1, 1).

3.1.3 Equation-Based Counting


Example 3.1.6. Find the number of natural number solutions to

x + y = 5, x, y ∈ N. (75)

The solutions are (1, 4), (2, 3), (3, 2), (4, 1), so there are 4 solutions.
Example 3.1.7. Find the number of solutions to

x + y + z = 6, x, y, z ∈ N. (76)

Substitute x′ = x − 1, y ′ = y − 1, z ′ = z − 1. Then

x′ + y ′ + z ′ = 3. (77)

Thus the number of solutions is


   
3+3−1 5
= = 10. (78)
3−1 2
18 CHAPTER 3. LECTURE 3
Chapter 4

Lecture 4

4.1 Sample Space and Events


Definition 4.1.1 (Sample Space). The sample space, denoted by Ω, is the set of all possible
outcomes of a random experiment. Each outcome w belongs to Ω, i.e., w ∈ Ω.
Definition 4.1.2 (Event). An event E is any subset of the sample space Ω, i.e., E ⊆ Ω. An
event occurs when the actual outcome of the experiment belongs to E.
Example 4.1.1 (Tossing a Coin). If a fair coin is tossed once, the sample space is Ω = {H, T }.
The event E = {H} corresponds to obtaining a Head.

Sample Space:
Set of all possible outcomes. It is denoted by Ω
ˆ w ∈ Ω (w is an element of Ω)
ˆ E is a subset of a sample space.
E⊆Ω

4.2 Probability Function


A probability measure assigns a number between 0 and 1 to each event.
Definition 4.2.1 (Probability Function). A probability function is a mapping:

P : F → [0, 1] (79)

where F is the collection of events. For finite experiments, F can be taken as the power set 2Ω .
Remark 4.2.1. If Ω is finite and all outcomes are equally likely, then for any E ⊆ Ω,
|E|
P (E) = . (80)
|Ω|

4.3 Axioms of Probability


1. 0 ≤ P(E) ≤ 1
P(E) = P(E|Ω)
E⊆Ω

2. P(Ω) = 1

19
20 CHAPTER 4. LECTURE 4

3. If E1 , E2 , . . . , Ei , . . . , En , . . . , Ej . . . are mutually exclusive events:

ˆ E ∩E
i j = ∅ i ̸= j
E1 ∩ E2 = ∅
E1 ∩ E3 = ∅
E2 ∩ E4 = ∅

then
∞ ∞
!
[ X
P Ei = P(Ei ) (81)
i=1 i=1

( ∞
S
i=1 Ei ) : It is a disjoint union. (A disjoint union is when you combine multiple sets,
and none of them share any elements. That means each item in the union comes from
exactly one set — no duplicates, no overlap.)

ˆ Set operation: S P
ˆ Arithmetic operation:
4.4 Theorem I
A⊆Ω (82)
c
P (A ) = 1 − P (A) (83)

Proof

Ac

A⊆Ω (84)
Ac ⊆ Ω (85)
c
A∩A =∅ (86)
A ∪ Ac = Ω (87)

P (A ∪ Ac ) = P (Ω) (From Axiom (2)) (88)


P (Ω) = 1 (89)

P (A) + P (Ac ) = 1 (From Axiom (3)) (90)


c
⇒ P (A ) = 1 − P (A) (91)

Note:
4.5. THEOREM 2 21

ˆ P (A) may be difficult to calculate directly.


ˆ P (A ) might be easier to compute.
c

ˆ Therefore, use: P (A) = 1 − P (A ) c

4.5 Theorem 2
Let A ⊆ Ω, B ⊆ Ω.

P (A ∪ B) = P (A) + P (B) − P (A ∩ B) (92)

Proof

A A∩B B

A = (A ∩ B c ) ∪ (A ∩ B) (93)

These two sets are disjoint.


By Axiom (3),
P (A) = P (A ∩ B c ) + P (A ∩ B) (94)

Also,
P (A ∩ B c ) = P (A) − P (A ∩ B) (95)

Note that:
A ∪ B = (A ∩ B c ) ∪ B (96)

These sets are disjoint.


Applying Axiom (3) again,

P (A ∪ B) = P (A ∩ B c ) + P (B) (97)

Rearranging,
P (A ∩ B c ) = P (A ∪ B) − P (B) (98)

Equating both equations:

P (A) − P (A ∩ B) = P (A ∪ B) − P (B) (99)

Solving for P (A ∪ B):

P (A ∪ B) = P (A) + P (B) − P (A ∩ B) (100)


22 CHAPTER 4. LECTURE 4

Example: Toss the fair coin once


(By fair coin, we mean P (H) = P (T ) = 12 )
Let
Ω1 = {H1 , T1 } (101)

Each outcome:
ω1 = H1 ∈ Ω1 , ω2 = T1 ∈ Ω1 (102)

Let the probability function be:

P : 2Ω1 → [0, 1] (103)

F(Ω1 ) = 2Ω1 = {ϕ, {H1 }, {T1 }, {H1 , T1 }} (104)

P (ϕ) = P (Eϕ ) = 0 (105)

1
P ({H1 }) = P (EH1 ) = p = (106)
2
1
P ({T1 }) = P (ET1 ) = q = (107)
2

P ({H1 , T1 }) = P (EΩ1 ) = P (Ω1 ) = 1 (108)

Eϕ = ϕ ⊆ Ω1 (109)

EH1 = {H1 } ⊆ Ω1 (110)

ET1 = {T1 } ⊆ Ω1 (111)

EΩ1 = {H1 , T1 } ⊆ Ω1 (112)

Ω1 = {H1 , T1 } = {H1 } ∪ {T1 } (disjoint) (113)

So,
P ({H1 , T1 }) = P ({H1 }) + P ({T1 }) (114)

E = {ω1 , ω2 , ω3 } = {{ω1 }, {ω2 }, {ω3 }} (They are mutually disjoint) (115)

Then by Axiom (3),

P (E) = P ({ω1 }) + P ({ω2 }) + P ({ω3 }) (116)

X
P (E) = P ({ω}) (117)
ω∈E
4.6. THEOREM 3 23

Example: Tossing a fair Coin Twice


All outcomes of Ω2 are equally likely (uniform sample space)

Ω2 = {H1 H2 , H1 T2 , T1 H2 , T1 T2 } (118)

1
P ({H1 H2 }) = P ({H1 T2 }) = P ({T1 H2 }) = P ({T1 T2 }) = (119)
4
Let:
A = {H1 T2 , T1 H2 } (120)
Then:
1 1 1
P (A) = P ({H1 T2 }) + P ({T1 H2 }) = + = (121)
4 4 2
Also:
|A| = 2, |Ω2 | = 4 (122)
|A| 2 1
P (A) = = = (123)
|Ω2 | 4 2
Note: This only holds true for a uniform sample space where all the outcomes are equally
likely.

4.6 Theorem 3
Let E ⊆ Ω, then:
|E|
P (E) = (124)
|Ω|
For a uniform sample space.
In general (for any sample space):
X
P (E) = P ({ω}) [General] (125)
ω∈E

(For non-uniform and uniform sample spaces)

Proof: Probability in a Uniform Sample Space


Let E ⊆ Ω be an event in a uniform sample space. Then:
X
P (E) = P ({ω}) (126)
ω∈E

Since the sample space Ω is uniform, all outcomes ω ∈ Ω are equally likely:
1
P ({ωj }) = for all ωj ∈ Ω (127)
|Ω|
Let:
Ω = {ω1 , ω2 , . . . , ω|Ω| } (128)
Then:
1 = P (Ω) = P ({ω1 }) + P ({ω2 }) + . . . + P ({ω|Ω| }) (129)
Since all probabilities are equal:

P ({ωj }) = y for all ωj ∈ Ω (130)


24 CHAPTER 4. LECTURE 4

1
1 = y + y + . . . + y = y · |Ω| ⇒ y = (131)
|Ω|
Thus:
1
P ({ωj }) = for all ωj ∈ Ω (132)
|Ω|
Let the event E = {ω1 , ω2 , . . . , ω|E| }. Then:

X X 1 1 1 1
P (E) = P ({ω}) = = + + ... + (133)
|Ω| |Ω| |Ω| |Ω|
ω∈E ω∈E

(repeated |E| times) (134)

|E|
P (E) = (135)
|Ω|

Example: MBA Class Selection


Suppose we have an MBA class of 50 students:

ˆ 20 students from the North


ˆ 30 students from the South
Three students are chosen at random for the placement cell.
Let the event E be:

E = At least 1 student is from the North in the placement cell (136)

We are asked to find:


P (E) =? (137)

Instead of directly calculating P (E) (which is hard), we find the complement:

E c = No student from North in placement cell (i.e., all 3 are from South) (138)

So,
   
c 30 50
|E | = , |Ω| = (139)
3 3

30

c |E c | 3
P (E ) = = 50 (140)
|Ω|

3

Hence,
30

c 3
P (E) = 1 − P (E ) = 1 − 50
 (141)
3

This method (using the complement) is easier.

50
− 30
 
3 3
P (E) = 50
 (142)
3
4.7. SUM OF FIRST (n − 1) NATURAL NUMBERS 25

Alternate Method for Event E


Let:
E = E1 ∪ E2 ∪ E3 (143)
Where:

E1 = Choose 1 from North and 2 from South (144)


E2 = Choose 2 from North and 1 from South (145)
E3 = Choose 3 from North and 0 from South (146)

Then:
|E1 | |E2 | |E3 |
P (E) = P (E1 ) + P (E2 ) + P (E3 ) = + + (147)
|Ω| |Ω| |Ω|
Where:
  
20 30
|E1 | = (148)
1 2
  
20 30
|E2 | = (149)
2 1
 
20
|E3 | = (150)
3
 
50
|Ω| = (151)
3

Example: Handshakes
Problem: An MBA class has 50 students. How many handshakes occur if each student shakes
hands with every other student?
Let total students be n = 50.
   
n 50
Total handshakes = = (152)
2 2
 
n n! n(n − 1)(n − 2)! n(n − 1)
= = = (153)
2 2!(n − 2)! 2(n − 2)! 2
Note: This is not the same as the sum of the first n natural numbers:
 
n
Sn = 1 + 2 + 3 + . . . + (n − 1) ̸= (154)
2

But in fact,
n(n − 1)
Sn−1 = (155)
2

4.7 Sum of First (n − 1) Natural Numbers


Sn−1 = 0 + 1 + 2 + 3 + · · · + (n − 2) + (n − 1) (156)
This is an arithmetic series where: - First term a = 0 - Last term l = n − 1 - Number of
terms n
So,
n n(n − 1)
Sn−1 = (0 + (n − 1)) = (157)
2 2
26 CHAPTER 4. LECTURE 4

50

Graphical Insight into 2
 
50
= 49 + 48 + 47 + · · · + 2 + 1 + 0 (158)
2
This shows that choosing 2 people out of 50 can be visualized as summing all unique pairs:
- Person 1 shakes hands with 49 people - Person 2 shakes hands with 48 new people (excluding
the one already counted), and so on.
Hence,
  X 49
50 50 · 49
= k= (159)
2 2
k=0

4.8 General Formula for Binomial Coefficient


For any n, r ∈ N with r ≤ n,  
n n!
= (160)
r r!(n − r)!
Chapter 5

Lecture 5

5.1 Example
Roll a pair of fair dice. Find the probability that the second die shows a higher value than the
first die.

|Ω| = 62 = 36, Ω = {(x1 , x2 ) | x1 , x2 ∈ {1, 2, 3, 4, 5, 6}}. (161)

Direct counting

ˆ If first die gives 1, second can be 2, 3, 4, 5, 6 ⇒ 5 outcomes.


ˆ If first die gives 2, second can be 3, 4, 5, 6 ⇒ 4 outcomes.
ˆ If first die gives 3, second can be 4, 5, 6 ⇒ 3 outcomes.
ˆ If first die gives 4, second can be 5, 6 ⇒ 2 outcomes.
ˆ If first die gives 5, second can be 6 ⇒ 1 outcome.
ˆ If first die gives 6, second can be none ⇒ 0 outcomes.
Total favourable outcomes: 5 + 4 + 3 + 2 + 1 + 0 = 15.
15 5
= (162)
36 12

x
P (E2>1 ) = =? (163)
|Ω|
x
= (164)
36
6 2x
1= + (165)
36 36
2x 6
=1− (166)
36 36
2x 36 − 6
= (167)
36 36
2x 30
= (168)
36 36
30
2x = × 36 (169)
36

27
28 CHAPTER 5. LECTURE 5

2x = 30 (170)
30
x= (171)
2
x = 15 (172)

5.2 Conditional Probability


ˆ P : F −→ [0, 1]
ˆ F being set of events.
ˆ E ∈ F =⇒ 0 ≤ P (E) ≤ 1
P(E) = P(E |Ω|)

P(E) is the probability of E given Sample space Ω

P (E|F ) = Conditional Probability of event E (E ⊆ Ω) given event F(F ⊆ Ω) (173)

Suppose F does not give any info about E ⊆ Ω


then P (E|F )?


F ⊆ Ω and E ⊆ Ω
E F P (E ∩ F ) = ∅

Disjointevent ̸= Independentevent (174)

P (E|F ) = P (E|Ω) = P (E) (175)


where, F is independent of E

5.3 Sample Space in a collection of all basic outcomes ω ∈ Ω of


some experiment.
Let E ⊆ Ω and F ⊆ Ω be two events.

P (E ∩ F )
P (E|F ) = conditional probability of E given F = (176)
P (F )
where P (F ) ̸= 0 and 0 < P (F ) ≤ 1 (177)
P(E) = P(E|Ω)
5.4. EXAMPLE 29

P(E ∩ F)

EE∩F F

P (E|F ) ≥ P (E|Ω) (178)


P (E|F ) ≥ P (E) (179)
P (E ∩ F )
P (F |E) = (180)
P (E)

5.4 Example
Toss a Fair Die
Ω = {1, 2, 3, 4, 5, 6} (181)
E = {4} (182)
F = {4, 5, 6} (183)

E 4 F
4 4,5,6

P (E ∩ F )
P (E|F ) = (184)
P (F )

E ∩ F = {4} ∩ {4, 5, 6} = {4} (185)


The number of outcomes in E ∩ F is 1

1
P (E ∩ F ) = (186)
6
30 CHAPTER 5. LECTURE 5

The number of outcomes in F is 3


The number of outcome in Ω is 6

3 1
P (F ) = = (187)
6 2

1 6 1
P (E|F ) = × = (188)
6 3 3

1 2 3 time
E1 occurs E2 E3

P (E ∩ F ) P (F ∩ E) P (E ∩ F )
P (E|F ) = and P (F |E) = = (189)
P (F ) P (E) P (E)

P (E ∩ F ) = P (E|F )P (F ) = P (F |E)P (E) (190)

P (E1 ∩ E2 ∩ E3 ) = P (E1 )P (E2 |E1 )P (E3 |E1 ∩ E2 ) (191)


LHS RHS (192)

P (E1 ∩ E2 ) P (E3 ∩ F )
P (E
 1)
 (193)
P (E
 1)
 P (F )
(194)
P (E3 ∩ F )
P (E1 ∩ E2 ) (195)
(( (((
P (E1 ∩ E2 )
(
( (( (((
(196)
P (E3 ∩ F ) (197)
(198)
P (E1 ∩ E2 ∩ E3 ) (199)

5.5 Definition
Event E ⊆ Ω and F ⊆ Ω are independent events if:

P (E|F ) = P (E|Ω) = P (E) (200)

P (E|F ) = P (E) (201)


P (E ∩ F )
= P (E) (202)
P (F )
P (E ∩ F ) = P (E)P (F ) (203)
where, E and F are independent
5.5. DEFINITION 31

0 < P (F ) ≤ 1
E F
0 < P (E) ≤ 1

E ∩ F = ∅ =⇒ P (E ∩ F ) = P (∅) = 0 (204)
LHS = P (E ∩ F ) = 0 (205)
RHS = P (E)P (F ) ̸= 0 (206)
LHS ̸= RHS (207)

different events are not independent.

P (E) = P (E|Ω) = PΩ (E), Ω, E ⊆ Ω (208)


3 axioms:

1. 0 ≤ P (E) ≤ 1

2. P (Ω) = 1

3. If E1 , E2 , . . . , Ei , . . . , En , . . . , Ej , . . . are mutually exclusive events,


Ei ∩ Ej = ∅ i ̸= j

PΩ (E ∩ F )
PF (E) = P (E|F ) = (209)
PΩ (F )

0 < PΩ (F ) ≤ 1 (210)
32 CHAPTER 5. LECTURE 5
Chapter 6

Lecture 7

6.1 Mutual Independence, Pairwise Independence, Conditional


Independence
ˆ PI ̸⇒ CI
ˆ CI ̸⇒ PI
6.2 Example: Tossing a Fair Die
Sample Space
Ω = {1, 2, 3, 4, 5, 6} (211)

Events

E1 = {1, 2, 3, 4} (212)
E2 = {4, 3, 6} (213)
E3 = {1, 2, 6} (214)

Assume the events are independent if:

P (Ei ∩ Ej ) = P (Ei )P (Ej ) for all i ̸= j (215)

Example Calculation
Let us compute the probability of event {4}:
1
P ({4}) = (216)
6

Verifying Pairwise Independence

4 3 2
P (E1 ) = , P (E2 ) = , P (E1 ∩ E2 ) = (217)
6 6 6
4 3 12 1
P (E1 ) · P (E2 ) = · = = (218)
6 6 36 3
2 1
P (E1 ∩ E2 ) = = (219)
6 3
Here, P (E1 ∩ E2 ) = P (E1 ) · P (E2 ).

33
34 CHAPTER 6. LECTURE 7

Now consider:
1 4 3 3 36 1
P (E1 ∩ E2 ∩ E3 ) = , P (E1 )P (E2 )P (E3 ) = · · = = (220)
6 6 6 6 216 6
So the events are mutually independent.
But:
1 1
P ({4}) = = ̸ (221)
6 3
So some conditional independence relations may not hold.

6.3 Conditional Probability and Intersection


Definition
Let E, F ⊆ Ω and P (F ) > 0, then:

P (E ∩ F )
PF (E) = P (E | F ) = (222)
P (F )

Properties
1. 0 ≤ PF (E) ≤ 1

2. PF (Ω) = 1

3. If E1 , E2 are disjoint, then:

PF (E1 ∪ E2 ) = PF (E1 ) + PF (E2 ) (223)

Venn Diagram Interpretation


Let E, F ⊆ Ω with E ∩ F ̸= ∅:

E1 E2

6.4 Axioms of Conditional Probability


Let F ⊆ Ω with P (F ) > 0. Define the conditional probability of event E given F as:

P (E ∩ F )
PF (E) = (224)
P (F )

Axiom 1: Boundedness
For any event E ⊆ Ω,
0 ≤ PF (E) ≤ 1 (225)
Justification:
Since E ∩ F ⊆ F , we have:
0 ≤ P (E ∩ F ) ≤ P (F ) (226)
6.5. SAMPLE SPACE AND EVENTS 35

Dividing all sides by P (F ) (which is > 0) gives:


P (E ∩ F )
0≤ ≤1 (227)
P (F )
Therefore:
0 ≤ PF (E) ≤ 1 (228)

Axiom 2: Normalization
The conditional probability of the entire sample space Ω, given F , is 1:
P (Ω ∩ F ) P (F )
PF (Ω) = = =1 (229)
P (F ) P (F )
Conclusion: The conditional probability function PF satisfies all properties of a probability
measure on the reduced sample space F .

E1 F E2

(E1 ∩ F ) (E2 ∩ F )

Axiom 3: Additivity
If E1 ∩ E2 = ∅, then:
PF (E1 ∪ E2 ) = PF (E1 ) + PF (E2 ) (230)
Proof:
By the definition of conditional probability:
P ((E1 ∪ E2 ) ∩ F )
PF (E1 ∪ E2 ) = (231)
P (F )
Since E1 ∩ E2 = ∅, their intersections with F are also disjoint:

(E1 ∪ E2 ) ∩ F = (E1 ∩ F ) ∪ (E2 ∩ F ) (232)

So:
P ((E1 ∪ E2 ) ∩ F ) = P (E1 ∩ F ) + P (E2 ∩ F ) (233)
Hence:
P (E1 ∩ F ) P (E2 ∩ F )
PF (E1 ∪ E2 ) = + = PF (E1 ) + PF (E2 ) (234)
P (F ) P (F )
Conclusion: The conditional probability measure PF satisfies all three axioms of a proba-
bility function on the sample space restricted to F .

6.5 Sample Space and Events


A sample space S is the set of all possible outcomes of an experiment. An event is a subset
of the sample space.
36 CHAPTER 6. LECTURE 7

Figure 6.1: Sample space S and an event E.

6.6 Independence of Events


Definition 6.6.1. Two events E and F are said to be independent if

P (E ∩ F ) = P (E) · P (F ). (235)

Definition 6.6.2 (Pairwise Independence). Events E1 , E2 , . . . , En are pairwise independent if

P (Ei ∩ Ej ) = P (Ei ) · P (Ej ), ∀i ̸= j. (236)

Definition 6.6.3 (Mutual Independence). Events E1 , E2 , . . . , En are mutually independent if

P (E1 ∩ E2 ∩ · · · ∩ En ) = P (E1 ) · P (E2 ) · · · P (En ). (237)

Remark 6.6.1. Mutual independence =⇒ pairwise independence. However, pairwise inde-


pendence does not necessarily imply mutual independence.

E E∩F F

Figure 6.2: Venn diagram of two independent events.

6.7 Example: Drawing Balls without Replacement


An urn contains 5 white balls and 5 black balls. One ball is drawn at random.

5 5
P (White) = = 12 , P (Black) = = 12 . (238)
10 10
If a ball is drawn and not replaced, the probability for the second draw changes. For example,
if the first draw is white, then
4 5
P (Second White) = , P (Second Black) = . (239)
9 9
Thus, the two draws are not independent.
6.8. CONDITIONAL PROBABILITY 37

5W

5B

Figure 6.3: Urn containing 5 white and 5 black balls.

6.8 Conditional Probability


Definition 6.8.1. The conditional probability of an event E given event F is defined as

P (E ∩ F )
P (E|F ) = , P (F ) > 0. (240)
P (F )

F E∩F E

Figure 6.4: Conditional probability focuses on the overlap relative to F .

6.9 Law of Total Probability


Theorem 6.9.1 (Law of Total Probability). If F1 , F2 , . . . , Fn are mutually exclusive and ex-
haustive events, then for any event E,
n
X
P (E) = P (E|Fi ) P (Fi ). (241)
i=1

6.10 Bayes’ Theorem


Theorem 6.10.1 (Bayes’ Theorem). For events F1 , F2 , . . . , Fn forming a partition of the sample
space, and an event E with P (E) > 0,

P (E|Fj ) P (Fj )
P (Fj |E) = Pn . (242)
i=1 P (E|Fi ) P (Fi )

Remark 6.10.1. Bayes’ theorem updates prior probabilities P (Fi ) into posterior probabilities
P (Fi |E) after observing evidence E.

6.11 Example: Box and Balls Problem


Suppose there are 3 boxes:

ˆ Box 1: 4 red, 2 blue balls


ˆ Box 2: 3 red, 3 blue balls
38 CHAPTER 6. LECTURE 7

ˆ Box 3: 2 red, 4 blue balls


An experiment consists of: 1. Selecting a box at random. 2. Selecting a ball from that box.
If a red ball is drawn, what is the probability it came from Box 1?

P (Red|Box 1)P (Box 1)


P (Box 1|Red) = P3 (243)
i=1 P (Red|Box i)P (Box i)
4 1 4
6 · 3 18 4
= 4 1 3 1 2 1 = 9 = . (244)
6 · 3 + 6 · 3 + 6 · 3 18
9

Choose Box

1 1 1
3 3 3

Box 1 Box 2 Box 3

4 2 3 3 2 4
6 6 6 6 6 6

Red Blue Red Blue Red Blue

Figure 6.5: Tree diagram for Bayes’ theorem box problem.

6.12 Disjoint Events


Definition 6.12.1. Events E1 , E2 , . . . , En are said to be mutually disjoint (mutually exclu-
sive) if
Ei ∩ Ej = ∅, ∀i ̸= j. (245)

E1 E2

Figure 6.6: Two disjoint events E1 and E2 .

6.13 Unions of Events


For two events E and F ,
P (E ∪ F ) = P (E) + P (F ) − P (E ∩ F ). (246)

6.14 Partitions of the Sample Space


A collection of events {F1 , F2 , . . . , Fn } forms a partition of S if:
n
[
Fi ∩ Fj = ∅ (i ̸= j), Fi = S. (247)
i=1
6.15. EXAMPLE: TOSSING A FAIR COIN TWICE 39

E E∩F F

Figure 6.7: Union of two events.

F1 F2 F3

Figure 6.8: Partition of sample space into disjoint events.

6.15 Example: Tossing a Fair Coin Twice


Sample space: S = {HH, HT, T H, T T }.
1 1
P (H) = , P (T ) = . (248)
2 2
If we define events:
A = {first toss is H}, B = {second toss is H}, (249)
then
1
P (A ∩ B) = P ({HH}) = . (250)
4
1 1
Since P (A) · P (B) = 2 · 2 = 41 , events A and B are independent.

Start

H T

HH HT TH TT

Figure 6.9: Tree diagram of tossing a coin twice.

6.16 Example: Mutually Independent Events


Consider rolling a fair die. Define
A1 = {even outcome}, A2 = {multiple of 3}, A3 = {outcome > 4}. (251)
One can verify pairwise independence, but not mutual independence.
40 CHAPTER 6. LECTURE 7
Chapter 7

Lecture 8

7.1 Random Variables


A random variable (RV) is a function:

X:Ω→R (252)

where Ω is the sample space.

ˆ If S X is finite ⇒ X is a discrete RV.

ˆ If S X is continuous ⇒ X is a continuous RV.

7.1.1 Example
Pick students at random from MBA class and measure their height.

S = {all MBA students}, X = height of students (253)


If h ∈ [4, 7] (feet), then height is a continuous set.

H = [4, 7], h∈H (254)

7.1.2 Probability Mass Function (PMF)


For a discrete RV X:
p(x) = P [X = x], x ∈ SX (255)
That is,
P [X = xi ] = P (Ei ) (256)

7.1.3 Definition
A random variable (RV) is a function

X:Ω→R (257)

that assigns a real number to each outcome in the sample space Ω.

ˆ If the set of possible values S X is finite or countable, then X is a discrete random


variable.

ˆ If the set of possible values S X is an interval or continuous set, then X is a continuous


random variable.

41
42 CHAPTER 7. LECTURE 8

7.1.4 Probability Mass Function (PMF)


If X is a discrete random variable, the probability that X takes the value x is given by the
probability mass function (PMF):

p(x) = P [X = x], x ∈ SX (258)

That is:
P [X = xi ] = P (Ei ) (259)

Properties of PMF
1. p(x) ≥ 0 ∀x ∈ SX
P
2. x∈SX p(x) = 1

7.2 Experiment: Waiting Time for First Head


Experiment: Toss a coin until you see a Head. Each toss is identical and independent. Stop
the experiment when you get a Head.
Let N be the waiting time to see the first Head.
Define the event:

Ek = [N = k] = On the k-th toss we see the first Head. (260)

This means the first k − 1 tosses are Tails, and the k-th toss is a Head:

Ek = {T1 , T2 , . . . , Tk−1 , Hk } (261)

Let:
P (Tails) = q = (1 − p), P (Heads) = p (262)
Therefore:
P (Ek ) = q · q · · · · · q · p = q k−1 p, k ∈ {1, 2, 3, . . . } (263)
For a fair coin, p = q = 21 .

7.2.1 Probability Distribution


q0p + q1p + q2p + q3p + · · · = p 1 + q + q2 + q3 + . . .
 
(264)
Using the geometric series:
1
1 + q + q2 + q3 + · · · = (265)
1−q
So:
1 1
p· =p· =1 (266)
1−q p

7.3 X is a Random Variable


Definition: If X is discrete or continuous R.V., then the function

FX (x) = P (X ≤ x) (267)

is called the Cumulative Distribution Function (CDF) of X.


Note 1: PMF of X =⇒ PX (x) = P (X = x) for discrete R.V.
Note 2: CDF of X is a right-continuous function.
7.3. X IS A RANDOM VARIABLE 43

7.3.1 Example
Toss a fair coin twice:
S = {H1 H2 , H1 T2 , T1 H2 , T1 T2 } (268)
Let X = Number of Heads.

X ∈ SX = {0, 1, 2} (269)

P (X = 0) = 41 , P (X = 1) = 12 , P (X = 2) = 1
4 (270)
So, the PMF of X is:
PX (x) = P (X = x) (271)

7.3.2 Properties of CDF


FX (x) = P (X ≤ x) (272)

1. FX (x) is right-continuous.

2. limx→−∞ FX (x) = 0 [No probability]

3. limx→+∞ FX (x) = 1

4. FX (x) is non-decreasing (increasing or constant only).

x ≤ y =⇒ FX (x) ≤ FX (y) (273)


44 CHAPTER 7. LECTURE 8
Chapter 8

Lecture 9

8.1 Probability Example: Uniform over (0, 1)


Choose a number at random from the interval (0, 1). Let X denote the chosen number. Compute
P (X = 1/2) and P (X = a) for a ∈ [0, 1]. (274)

8.1.1 Facts about the uniform distribution on (0, 1)


If X is chosen uniformly on (0, 1), then for any 0 ≤ u ≤ v ≤ 1:
P (u ≤ X ≤ v) = v − u (275)
Explanation: The probability of choosing a number in an interval equals the length of that
interval.
Examples using (275):
P (0 ≤ X ≤ 1) = 1 − 0 = 1, (276)
 
P 12 ≤ X ≤ 1 = 1 − 12 = 12 , (277)
 
P 0 ≤ X ≤ 21 = 12 − 0 = 12 . (278)

8.1.2 Probability of a single point


A single point {a} has length zero. Therefore:
P (X = a) = 0 for every a ∈ [0, 1]. (279)
Proof via shrinking intervals: For ε > 0 such that [a − ε, a + ε] ⊂ (0, 1):
P (a − ε ≤ X ≤ a + ε) = 2ε. (280)
Since {a} ⊂ [a − ε, a + ε], we have
0 ≤ P (X = a) ≤ 2ε. (281)
Letting ε → 0 gives P (X = a) = 0.
Alternate proof via discrete approximation: Consider the grid
Gn = 0, n1 , . . . , 1 .

(282)
If one chooses a point uniformly from Gn , each point has probability 1/(n + 1). If 1/2 is a grid
point (even n), then
1
Pn (X = 1/2) = → 0 as n → ∞. (283)
n+1
Conclusion: Singleton sets have zero probability under a continuous uniform distribution:
P (X = 1/2) = 0, P (X = a) = 0 for all a ∈ [0, 1]. (284)

45
46 CHAPTER 8. LECTURE 9

8.2 Continuous Random Variable


For a continuous random variable X:

P (X = a) = 0 for all a ∈ (0, 1). (285)

Explanation: Continuous random variables do not have probabilities for single points. Prob-
abilities are assigned only to intervals. There is no probability mass function (PMF), only a
probability density function (PDF).

8.3 Discrete Random Variable


If X is a discrete random variable with support set SX :

PMF of X : P (X = x), x ∈ SX , (286)


CDF of X : FX (x) = P (X ≤ x). (287)

8.3.1 Expectation of a Discrete Random Variable


Let X : Ω → SX be discrete. The expectation is defined as
X
E P [X] = x · P (X = x), (288)
x∈SX

and under another measure Q:


X
E Q [X] = x · Q(X = x). (289)
x∈SX

Interpretation: Expectation is the weighted average of all possible values, weighted by their
probabilities.

8.4 Example: Tossing Coins


8.4.1 Tossing 2 Coins: Number of Heads X
Case 1: Fair Coin (P )

1 1 1
P (X = 0) = , P (X = 1) = , P (X = 2) = . (290)
4 2 4
Expectation:
1 1 1
E P [X] = 0 · +1· +2· =1 (291)
4 2 4
Case 2: Biased Coin (Q)

1 1 1
Q(X = 0) = , Q(X = 1) = , Q(X = 2) = . (292)
3 3 3
Expectation:
1 1 1
E Q [X] = 0 · +1· +2· =1 (293)
3 3 3
Observation: Different distributions can yield the same expected value.
8.5. INDICATOR RANDOM VARIABLE EXAMPLE 47

8.4.2 Tossing 10 Coins: Observed Sample


Sample: 1, 1, 0, 0, 0, 1, 1, 0, 0, 1 (Heads=1, Tails=0)
Average number of heads:
1+1+0+0+0+1+1+0+0+1 5
Avg # of H’s = = = 0.5 (294)
10 10
Expected value for a fair coin (Bernoulli X):
1 1
E[X] = 1 · + 0 · = 0.5 (295)
2 2
Expected value of X 2 :
1 1 1
E[X 2 ] = 12 · + 02 · = (296)
2 2 2
Reason: For a Bernoulli variable, X 2 = X.

8.5 Indicator Random Variable Example


Consider X as the outcome of a die roll, and define
(
1 if X = 2,
I := 1{X=2} = (297)
0 otherwise.
Sample average: If X = 2 occurs k times in n observations:
k
I= (298)
n
Expectation:
E[I] = 1 · P (X = 2) + 0 · P (X ̸= 2) = P (X = 2) (299)
Observation: The expectation of the indicator equals the probability of the event it rep-
resents. Empirical frequency k/n estimates P (X = 2).

8.6 Expectation Examples


8.6.1 Expectation of a Fair Die
Let D represent the outcome of a fair six-sided die:
6
X 1 1+2+3+4+5+6
E[D] = d· = = 3.5 (300)
6 6
d=1

8.6.2 Expectation of a Biased Coin (Two Tosses)


Coin bias: Q(H) = 0.6, Q(T ) = 0.4. Let X be number of heads in two tosses.
Distribution of X:
P (X = 0) = 0.42 = 0.16, (301)
P (X = 1) = 2 · 0.6 · 0.4 = 0.48, (302)
2
P (X = 2) = 0.6 = 0.36 (303)
Expectation:
E[X] = 0 · 0.16 + 1 · 0.48 + 2 · 0.36 = 1.20 (304)
Interpretation: On average, slightly more than 1 head is expected in 2 tosses due to the
bias.
48 CHAPTER 8. LECTURE 9
Chapter 9

Lecture 10

Let X be a discrete random variable, and x ∈ SX


Random Variable: X = Ω → SX

Probability Mass function of RV X:

PX (x) = P (X = x) (305)

Where,x ∈ SX

Cummulative Distributive Function of RV X:

PX (x) = P (X ≤ x) (306)

Expectation of a Discrete Random Variable based on outcome:

ω ∈ Ω, where ω is the outcome.


X
E P [X] = P (ω) × ω (307)
ω∈Ω

Expectation of Discrete Random Variable based on Events:

X
E P [X] = xP [X = x] (308)
x∈SX

EX = [X = x] (309)
X
= xP [EX ] (310)
x∈SX

49
50 CHAPTER 9. LECTURE 10

Example: Toss a fair coin twice

Let X =#of Heads [# - Number of]

Sample Space = Ω2

Ω2 = {H1 H2 = ω1 , H1 T2 = ω2 , T1 H2 = ω3 , T1 T2 = ω4 } (311)

X = 2, E2 = {H1 H2 , } (312)
X = 1, E1 = {H1 T2 , T1 H2 } (313)
X = 0, E0 = {T1 T2 } (314)

X
[Link][X] = P (ω) × ω (315)
ω∈Ω2

(Outcome Based)

X Partition into Ω2 Into disjoint events

In a fair coin toss every trial is mutually independent of each other and are statistically
identical.

X
E P [X] = P (ω) × ω = P (ω1 ) × (ω1 ) + P (ω2 ) × (ω2 ) + P (ω3 ) × (ω3 ) + P (ω4 ) × (ω4 ) (316)
ω∈Ω2

1 1 1 1
= ×2+ ×1+ ×1+ ×0 (317)
4 4 4 4
2 1 1
= + + =1 (318)
4 4 4

X X
2.E P [X] = xP (EX ) = xP (X = x) (319)
x∈SX x∈SX

( Event Based)

= 2P (E2 ) + 1P (E2 ) + 0P (E0 ) (320)

1
P (E2 ) = P {H1H 2} = (321)
4
2
P (E2 ) = P {H1 T2 , T1 H2 } = (322)
4
1
P (E0 ) = P {T1 T2 } = (323)
4

P 1 2 1
EX = 2( ) + 1( ) + 0( ) = 1 (324)
4 4 4
9.1. RANDOM VARIABLE IN FUNCTION 51

[ [
P (Ω2 ) = P (E0 ) P (E1 ) P (E2 ) (325)
Example:2

DAY 1 DAY 2 DAY 3 DAY 4 DAY 5 DAY 6 DAY 7 DAY 8 DAY 9 DAY 10
10 1 5 2 1 20 10 2 1 15
A person deposits coins of different sum with a banker for 10 days. The deposits made are
listed in the above [Link] the expectation of the event

1. Outcome Based:

X
P
EX = P (ω) × (ω) (326)
x∈§X

Ω = {10 = ω1 , 1 = ω2 , 5 = ω3 , 2 = ω4 , 1 = ω5 , 20 = ω6 , 10 = ω7 , 2 = ω8 , 1 = ω9 , 5 = ω1 0} (327)

1 1 1 1 1 1 1 1 1
= (10) + (1) + (5) + (2) + (1) + (20) + (5) + (10) + (2) + (1) (328)
10 10 10 10 10 1 10 10 10 10

57
= = 5.7 (329)
10
[Link] Based:

X
P
EX = xP (EX ) (330)
x∈SX

1 1 1 1 1 1 1 1 1 1
1( + + ) + 2( + ) + 5( + ) + 10( + ) + 20( ) (331)
10 10 10 10 10 10 1 10 10 10
3 + 4 + 10 + 20 + 20 57
= = = 5.7 (332)
10 10

9.1 Random Variable in Function


For Random Variable X

Function: y = g(x)

Y (ω) = g(X(ω)) (333)

Y = g(X) (334)

X
E(Y ) = E[g(x)] = g(x)P (X = x) (335)
x∈SX
X X X
yP (Y = y) = P (ω)y(ω) = g(x)P [X = x] (336)
y∈SY ω∈Ω x∈SX
52 CHAPTER 9. LECTURE 10

X X
P (ω)y(ω) = P (ω)g[X(x)] (337)
ω∈Ω ω∈ω

Ey = [Y = y] = {ω ∈ Ω|Y (ω) = y} = {ω ∈ Ω|g(X(ω)) = g(x)} (338)

Ex = [X = x] = {ω ∈ Ω|X(ω) = x} (339)

Ex = P (Ey ) = P (Ex ) (340)


Event Based :
X
EP [g(x)] = g(x)P (X = x) (341)
x∈SX

Outcome Based:
X
EP [g(x)] = P (ω)g(X(ω) (342)
ω∈Ω

If g(x) = x = X
EP [X] = xP(X = x) (343)
x∈SX
X
EP [X] = P (ω) × ω (344)
ω∈Ω

Example:
PMF of RV X is given by :

SX = {5, 7} (345)
1
P (X = 5) = (346)
3
2
P (X = 7) = (347)
3

Probability Mass Function of X


1

2/3
P (X = x)

1/3

0
5 7
x
9.1. RANDOM VARIABLE IN FUNCTION 53

Scaling of Random Variable X:


1.y = 6x‘

1
P (X = 5) = P (6X = 30) = (348)
3

2
P (X = 7) = P (7X = 42) = (349)
3

SY = {30, 42} (350)

Probability Mass Function of X


1

2/3
P (y = 6x)

1/3

0
30 42
x
Probability (Height) of events does not change in scaling.

P (Y = y) = P (g(X)) = g(x) (351)

P (Y = y) = P (6X = 6x) (352)

P (Y = y) = P (X = x) (353)

[Link] of RV X :
Y = X+9
PMF of X + 9: Y = h(x) = (x + 9)

1
P (X = 5) = P (X + 9 = 14) = (354)
3

2
P (X = 7) = P (X + 9 = 16) = (355)
3

SH = {14, 16} (356)


54 CHAPTER 9. LECTURE 10

Probability Mass Function of X


1
P (H = x + 9)

2/3

1/3

0
14 16
x
1)
EP [S] = P [X = 3] = 1 (357)

ˆ Expectation of any constant is 1.


2.
Ep [ax + b], (a, b ∈ R) (358)
Solution:

y = ax + b (359)

Ep [ax + b] = E↷ P = P (y) = P (ax + b) (360)

Y = aX + b (361)

X
EP [Y ] = P (ω)y(ω) (362)
ω∈Ω
X
= P (ω)[a × (ω) + b] (363)
ω∈Ω
X X
a[ P (ω) × ω] + b( P (ω)) (364)
ω∈Ω ω∈Ω
X
= a × EP [X] + b(1), [ P (ω) = 1] (365)
ω∈Ω

Linearity of Expectation:
[EP [aX + b] = aEP [X] + b] (366)
Example:

SX = {−1, 1} (367)
−1
p(X = −1) = (368)
2
1
P (X = 1) = (369)
2
9.2. VARIANCE OF RV X 55

X
EP = xp(x) (370)
ω∈Ω
1 1
= −1( ) + 1( ) (371)
2 2
=0 (372)

9.2 Variance of RV X
Variance is given by: X
2
σX = varX = (x − µx )2 P (X = x) (373)
x∈SX

µ − expected mean (374)


2
σX − Variance (375)
X
µX = E(X) = P (X = x) (376)
x∈SX
2
σX = VarX = EP [(x − µx )2 ] (377)

EP [x2 + µ2x − 2µx x] (378)


By Linearity of Expectation:
σx2 = E(x2 ) + E(µ2x ) + E(−2µx x) (379)
where E - is the 2nd moment of random variable.
E(x2 ) + µ2x ) − 2µx E(x) (380)

= E(X 2 ) − µ2x (381)


= E(x2 ) − [E(X)]2 (382)
Without squaring: X
V ∗ (X) = (x − µx )P (X = x) (383)
x∈SX
X X
= xP (X = x) − µx P (X = x) (384)
x∈SX x∈§X

V = µx − µx (1) = o (385)

[E − E = 0] (386)

9.2.1 For function Y=[x+a] — a ∈ R


σY2 = E(Y − µy )2 (387)
µy = E[Y ] = E[x + a] = E[X] + a = (µx + a) (388)

= E[(x + a − (µx + a)2 ] (389)


= E[(x + a − µx − a)2 ] (390)

σy2 = σx+a
2
= σx2 = E[(x − µx )2 ] = σx2 (391)
56 CHAPTER 9. LECTURE 10
Chapter 10

Lecture 11

Let X : Ω → S ⊆ R be a discrete random variable (RV).


X X
E[X] = x P (X = x) = X(ω)P (ω) (392)
x∈SX ω∈Ω

where ω ∈ Ω denotes an outcome.

10.0.1 Linearity of Expectation


n n
" #
X X
E ci Xi = ci E[Xi ] (393)
i=1 i=1

where ci ∈ R and Xi are random variables.

10.1 Variance Properties


X
2
σX = V (X) = Var(X) = (x − µX )2 P (X = x) (394)
x∈SX

= E (X − µX )2 = E[X 2 ] − µ2X
 
(395)

X
E[g(X)] = g(x)P (X = x) (2nd moment of RV X) (396)
x∈SX

Var(X + b) = Var(X), b∈R (397)

p q
σX = SD(X) = Var(X) = E[X 2 ] − µ2X (398)

Note:
The variance is invariant under shifting.

Var(aX + b) = Var(aX), a, b ∈ R (399)

µY = E[Y ] = E[aX + b] = aµX + b, Y = aX + b (400)

57
58 CHAPTER 10. LECTURE 11

V (Y ) = E (Y − µY )2 = E (aX + b − (aµX + b))2


   
(401)

= E (a(X − µX ))2 = a2 E[(X − µX )2 ]


 
(402)

=⇒ V (Y ) = a2 V (X) (403)

∴ Var(aX + b) = a2 Var(X) (404)

10.2 Expectation and Variance of Linear Combinations


Let
T = aX + bY + c, a, b, c ∈ R, X, Y are random variables. (405)

E[T ] = E[aX + bY + c] = aE[X] + bE[Y ] + c = aµX + bµY + c (406)

V (T ) = V (aX + bY + c) (407)

Let
D = aX + bY (408)

Var(D + c) = Var(D) (409)

Var(aX + bY ) = E[(aX + bY )2 ] − (E[aX + bY ])2 (410)

µD = E[aX + bY ] = aµX + bµY (411)

E[(aX + bY )2 ] = a2 E[X 2 ] + b2 E[Y 2 ] + 2ab E[XY ] (412)

D2 = (aX + bY )2 = a2 X 2 + b2 Y 2 + 2abXY (413)

Var(D) = a2 E[X 2 ] + b2 E[Y 2 ] + 2abE[XY ] − (aµX + bµY )2 (414)

= a2 (E[X 2 ] − µ2X ) + b2 (E[Y 2 ] − µ2Y ) + 2ab(E[XY ] − µX µY ) (415)

=⇒ Var(D) = a2 Var(X) + b2 Var(Y ) + 2ab Cov(X, Y ) (416)


Cov(X, Y ) = E[XY ] − E[X]E[Y ] (417)



10.3. EXAMPLE: PICKING BALLS 59

10.2.1 Example: Toss a fair coin twice


Let X = number of heads.

PX (x) = P (X = x), x ∈ {0, 1, 2} (418)

P (X = 0) = 14 , P (X = 1) = 12 , P (X = 2) = 1
4 (419)

SX = {0, 1, 2} (420)

10.3 Example: Picking Balls


We have an urn with 100 white balls and 100 red balls. Pick 2 balls consecutively at random
without replacement.
Let
Y = number of red balls selected. (421)

The probability mass function (PMF) of Y is:

PY (y) = P (Y = y), y ∈ SY = {0, 1, 2}. (422)

100

2 100 · 99 99
P (Y = 0) = 200 = = . (423)
200 · 199

2
398

100

2 100 · 99 99
P (Y = 2) = 200 = = . (424)
200 · 199

2
398

99 99 200
P (Y = 1) = 1 − − = . (425)
398 398 398
Thus,

P (Y = 0) = 41 , P (Y = 1) = 12 , P (Y = 2) = 14 . (426)

10.4 Example: Toss a Coin Once


Let X = number of heads in one toss.

Ω = {H, T }, SX = {0, 1} (427)

(
1, with probability p = P (H),
X= (428)
0, with probability q = 1 − p = P (T ).
60 CHAPTER 10. LECTURE 11

10.5 Bernoulli Distribution


Let (
1, {H} ⊆ Ω
X= (429)
0, {T} ⊆ Ω
Then the PMF of X is

pX (x) = P (X = x) = px (1 − p)1−x , x ∈ {0, 1}, q = 1 − p. (430)

Thus,
X ∼ Bernoulli(p). (431)

10.6 Example 2: Toss one biased coin n times


Each toss is statistically identical and independent of the others.

p = P (H), q = 1 − p = P (T ). (432)
Each toss is mutually independent:

Each toss does not statistically influence the other tosses. (433)

10.6.1 Case n = 2
Ω2 = {H1 H2 , H1 T2 , T1 H2 , T1 T2 } (434)

|Ω2 | = 22 = 4 (435)
In general, for n tosses:

Ωn = {H1 H2 . . . Hn , . . . , T1 T2 . . . Tn } (436)

|Ωn | = 2n (437)
Chapter 11

Lecture 12

11.1 Families of Random Variables (RVs) — Bernoulli RV


11.1.1 1. Xi ∼ Bern(p) [One Trial, One Parameter]
Let p = P (E) be the probability of an event of interest E ⊆ Ω.
Then, q = 1 − p = P (E c ).

Indicator Random Variable (RV):


(
1 with probability p = P (E), if E has occurred on the ith trial
Xi = (438)
0 with probability q = 1 − p = P (E c ), if E has not occurred on the ith trial

This indicates presence or absence of the event E.

Probability Mass Function (PMF):

P (X = x) = px q 1−x , x ∈ {0, 1} (439)

Probability
p
q

x
0 1

p+q =1 (440)

ˆ X : Ω → {0, 1}
i

ˆ {H, T } = {E, E }c

ˆ X is a dummy variable in econometrics.


i

ˆ X is also called a counting RV: {0, 1}, indicating occurrence or absence.


i

ˆ p+q =1
ˆ P (X = x) = p qx 1−x

Used to break down logical problems into sub-problems.

61
62 CHAPTER 11. LECTURE 12

11.1.2 2. n sets of identical & mutually independent trials


iid = independent and identically distributed
Each k th trial is independent of the j th trial.

p = P (E) is constant and independent of the trial (441)

Trial j : pj = pi = P (E) Trial k : pk = P (E) (442)

If identical: pj = pk = p (443)

Example 11.1.1. Toss a coin on the j th trial and count # of H’s:


(
1 if H occurs
Xji = (444)
0 otherwise

Let n iid trials, and E is the event of interest. Then:

Yn = # of events (E) in n iid trials (445)

Sample space:
SYn = {0, 1, 2, . . . , n} (446)

Yn : Ωn → SYn (447)

Example 11.1.2. Toss a coin n = 2 times and count # of H’s:

Ω2 = {H1 H2 , H1 T2 , T1 H2 , T1 T2 } |Ω2 | = 22 = 4 (448)


General case: |Ωn | = 2n

Yn ∼ Binomial(n, p) (449)

Every trial has two outcomes:


H→E or T → Ec (450)

11.1.3 Binomial RV and PMF


E is called the success event, E c is the failure event.

Yn = # of events (successes) in n iid trials (451)


Each trial has two outcomes: E and E c

p = P (E), q = P (E c ), p+q =1 (452)

Xj ∼ Bern(p) ⇒ Bin(1, p) (453)

P (X = x) = px q 1−x , x ∈ {0, 1} (454)

 
n y n−y
P (Yn = y) = p q , y ∈ {0, 1, . . . , n} (455)
y
11.1. FAMILIES OF RANDOM VARIABLES (RVS) — BERNOULLI RV 63

Example 11.1.3. Toss a coin n = 2 times.

Start

H1 T1

H2 T2 H2 T2

H1 H2 → 2 ⇒ p2 (456)
H1 T2 → 1 ⇒ pq (457)
T1 H2 → 1 ⇒ qp (458)
T1 T2 → 0 ⇒ q 2 (459)

All paths are equally likely. Ω2 = 4 paths.


 
2 y 2−y
P (Yn = y) = p q (460)
y
 
2 1 1
Example: P (Y2 = 1) = p q = 2pq (461)
1
 
3 1 2
Another example: p q = 3pq 2 (462)
1

11.1.4 3. Outcomes of n trials


We represent outcomes of n trials as a tree structure. For example, a 3-trial experiment where
each trial has 2 possible outcomes (heads or tails):

H1 H2 H3 = p3
H1 H2 T3 = p2 q
H1 T2 H3 = p2 q
H1 T2 T3 = pq 2
(463)
T1 H2 H3 = p2 q
T1 H2 T3 = pq 2
T1 T2 H3 = q 2 p
T1 T2 T3 = q 3
This tree shows different possible sequences of trials, each with an associated probability.

11.1.5 4. Cumulative Distribution Function (CDF) for Binomial RV


The CDF can be used to compute the probability that a binomial random variable Yn takes a
value less than or equal to k:

k
X
FY (k) = P (Yn ≤ k) = P (Yn = y) (464)
y=0
64 CHAPTER 11. LECTURE 12

11.1.6 5. Probability Calculation for Different Paths


We can calculate the probability of different paths in the tree, and the general formula is given
by:

P (ω k ) = pq q n−q (465)
The number of paths that lead to a particular outcome is given by:
 
n
k= (466)
q
The probability of each path is the same, as shown by:

P (E1 ∩ E2 ∩ E3 ) = P (E1 ) × P (E2 ) × P (E3 ) (467)


The total probability of all possible outcomes in n trials is the sum of the probabilities of each
path:

P (E) = kpq q n−q (468)

11.2 Multiplication of probability


n

Let the number of paths be k = q .

P (ω n,q ) = pq q n−q (469)


n−q n−q q
P (ω )=p q (470)
k
X
P (ω k ) = P (ω i ) = k × pq × q n−q (471)
i=1

Each path has an identical probability. Now consider the following:

E n,q = [yn = q] in n trials we see q successes. (472)


Probability Mass Function (PMF) of yn = q:

P (Yn = q) = P (exactly q successes in n trials) (473)


The total probability of q successes is then:

P (E1 ∩ E2 ∩ E3 ) = P (E1 )P (E2 )P (E3 ) (474)

11.2.1 7. Further Expansion and Multiplication of Probabilities


Now, we extend the problem and multiply probabilities as follows:

P (ω q,n ) = pq q n−q k = #of paths (475)

P (E) = k · pq q n−q (476)


Finally, as shown previously:

P (E1 E2 E3 ) = P (E1 )P (E2 )P (E3 ) (477)


11.2. MULTIPLICATION OF PROBABILITY 65

The result leads to a binomial expansion as follows:


 
n q n−q
P (E) = p q (478)
q
66 CHAPTER 11. LECTURE 12
Chapter 12

Lecture 13

12.1 The Binomial Distribution


The random variable Yn , representing the number of successes in n independent Bernoulli
trials (where each trial has only two outcomes, Success E or Failure Ec ), follows a Binomial
Distribution.
Let p be the probability of success on any single trial. We denote this as:
Yn ∼ Bin(n, p) (479)

12.1.1 Probability Mass Function (PMF)


The Probability Mass Function (PMF) of Yn gives the probability of observing exactly y
successes in n trials. The question asked is:
P (Yn = y) =? (480)
The PMF for a Binomial random variable Yn ∼ Bin(n, p) is:
 
n y
P (Yn = y) = p (1 − p)n−y (481)
y
for y ∈ {0, 1, 2, . . . , n}.
Where ny is the binomial coefficient, calculated as:

 
n n!
= (482)
y y!(n − y)!
and it represents the number of ways to choose y successes from n trials. py is the probability
of getting y successes. (1 − p)n−y is the probability of getting n − y failures.

12.1.2 The Bernoulli Distribution (n = 1)


For a single trial (n = 1), the random variable X1 follows a Bernoulli Distribution, denoted
Bern(p).
X1 ∼ Bin(1, p) (483)
The probability mass function (PMF) is:
P (X1 = x) = px q 1−x , x ∈ SX1 = {0, 1} (484)
ˆ If x = 1 (Success): P (X = 1) = p q = p.
1
1 0

ˆ If x = 0 (Failure): P (X = 0) = p q = q.
1
0 1

67
68 CHAPTER 12. LECTURE 13

12.1.3 Expected Value of a Bernoulli Random Variable


The expected value of X1 is:

E[X1 ] = (1) · P (X1 = 1) + (0) · P (X1 = 0) = 1 · p + 0 · q = p (485)

12.2 Binomial Probability Mass Function (PMF)


The PMF for the number of successes y in n trials is:
 
n y n−y
P (Yn = y) = p q , y ∈ SYn = {0, 1, 2, . . . , n} (486)
y

Interpretation:

ˆ : The number of distinct sequences (paths) with exactly y successes.


n
y

ˆ p : The probability of y successes.


y

ˆ q : The probability of n − y failures.


n−y

12.2.1 Path Interpretation


Consider a specific sequence of n trials with y successes (H) and n − y failures (T).

H1 , H2 , . . . , Hy Ty+1 , Ty+2 , . . . , Tn (487)


| {z } | {z }
Successes (y) Failures (n − y)

The probability of this specific path is:

P (Path) = p · p · . . . · p · q · q · . . . · q = py q n−y (488)


| {z } | {z }
y times n−y times

n

The term accounts for the fact that the final result of the path matters, not the
y

paths themselves. ny aggregates all paths that result in y successes.




12.3 Tree Diagram for n = 2 Trials


The diagram illustrates the sample space and probabilities for n = 2 coin tosses:

p H2 ) ⇒ P = p2
(H1 , H2=
H1
q
T2 (H1 , T2 )=⇒ P = pq
p

Start
p H2 (T1 , H2 )=⇒ P = qp
q

T1
q
T2 (T1 , T2 )=⇒ P = q 2
12.4. PROBABILITY OF A SINGLE PATH (SEQUENCE) 69

12.4 Probability of a Single Path (Sequence)


Consider an experiment with n independent trials, where the probability of success is p and
the probability of failure is q = 1 − p.
Let ω be a specific sequence of n outcomes, resulting in exactly y successes and n − y failures.
ˆ P (Success) = p
ˆ P (Failure) = q
The probability of any single specific path ω with y successes is the product of the
probabilities of the individual trials (due to independence):

P (ω with y successes) = p · p · . . . · p · q · q · . . . · q (489)


| {z } | {z }
y times n−y times

P (ω) = py q n−y (490)

12.5 Total Probability of y Successes


12.5.1 Definition of the Event En,y
Let En,y be the event that there are exactly y successes in n trials. This event is composed of
a collection of distinct, mutually exclusive sequences (paths).

ˆ Let ω n,y n,y n,y


1 , ω2 , . . . , ωk be the distinct paths (sequences) that result in exactly y
successes.

ˆ The total number of such distinct paths is k = n


y

. This is the number of ways to choose
the y positions for success out of n trials.

En,y = {ω1n,y , ω2n,y , . . . , ωkn,y } (491)


The number of distinct paths is:  
n
k= (492)
y

12.5.2 Calculating P (En,y ) using Disjoint Union


Since the paths ωpn,y are mutually exclusive (disjoint), the probability of the event En,y is the
sum of the probabilities of these individual paths:

P (En,y ) = P (ω1n,y ) + P (ω2n,y ) + . . . + P (ωkn,y ) (493)

This can be written using summation notation:


k
X
P (En,y ) = P (ωpn,y ) (494)
p=1

Since the probability of *every* single path ωpn,y with y successes is the same,
P (ωpn,y ) = py q n−y , we have:
Xk
 y n−y 
P (En,y ) = p q (495)
p=1

P (En,y ) = py q n−y + py q n−y + . . . + py q n−y


     
(496)
| {z }
k times
70 CHAPTER 12. LECTURE 13

12.5.3 Final PMF Formula


By factoring out the common probability term py q n−y , we get:

P (En,y ) = k · py q n−y (497)


n

Substituting the value of k = y :
 
n y n−y
P (En,y ) = p q (498)
y

This is the Probability Mass Function (PMF) for the Binomial Random Variable Yn :
 
n y
P (Yn = y) = p (1 − p)n−y (499)
y

12.6 Example: Coin Toss 3 Times (n = 3)


The possible outcomes for 3 trials (n = 3), assuming p is the probability of a Head (H) and
q = 1 − p is the probability of a Tail (T). Let Y be the random variable for the number of
Heads.

12.6.1 Outcomes and Probabilities for n = 3


Outcome Probability Value of Y (y)
H1 H2 H3 p3 P (Y = 3)
H1 H2 T3 , H1 T2 H3 , T1 H2 H3 3p2 q P (Y = 2)
H1 T2 T3 , T1 H2 T3 , T1 T2 H3 3pq 2 P (Y = 1)
T1 T2 T3 q3 P (Y = 0)

12.7 Expected Value of a Bernoulli Trial


Let X be a Bernoulli random variable for a single trial: X = 1 for success (H) and X = 0 for
failure (T).

ˆ Expected value for success:


E[X = 1] = 1 · P (X = 1) = p (500)

ˆ Expected value for failure (complement of success):


E[X = 0] = 0 · P (X = 0) = 0 (501)

The expected value of the Bernoulli random variable X is:

E[X] = 1 · p + 0 · q = p (502)

12.8 Conditions for Binomial Distribution


The key assumptions for the Binomial distribution rely on two conditions:
12.9. SCENARIO: EVENTS ARE NOT I.I.D. 71

12.8.1 Independence
The trials must be Independent and Identically Distributed (i.i.d.).

ˆ Independence:The outcome of one trial does not affect the outcome of any other trial.
ˆ Incorrect Summation (Example of what NOT to do in Independent Events):
P (E1 ∩ E2 ∩ E3 ) ̸= P (E1 ) + P (E2 ) + P (E3 ) + . . . (503)

ˆ Correct Probability for Independent Events: For n independent trials, the


probability of a specific sequence of outcomes (e.g., H1 , H2 , H3 ) is the product of their
individual probabilities:

P ({H1 , H2 , H3 }) = P (H1 ) × P (H2 ) × P (H3 ) (504)

12.8.2 Identically Distributed


ˆ Constant Success Probability: The probability of success p is the same for every trial.
p1 = p2 = p3 = . . . = pk = p (505)

ˆ p = P (E) is independent of the trial.


12.8.3 General Formulation and Paths
ˆ The total number of possible paths (sequences of outcomes) is 2 . n

ˆ Each path gives a value of y that is possible (i.e., y is the number of successes in that
path).

12.9 Scenario: Events are NOT i.i.d.


If the trials are not independent, we must use the multiplication rule for dependent events (or
conditional probability). The probability of a sequence is calculated by conditioning on the
previous outcomes.

If events are NOT i.i.d. (Dependent):

P ({H1 , H2 , H3 }) = P (H1 ) × P (H2 |H1 ) × P (H3 |H1 , H2 ) (506)


72 CHAPTER 12. LECTURE 13
Chapter 13

Lecture 14
 x n
lim 1+ → ex (507)
n→∞ n
 x n
1− → e−x (508)
n
 x n
ln = 1 − (509)
n
 x
loge ln = n loge 1 − (510)
n
 
−x
= ∞ loge 1 + = ∞ loge (1) = ∞(0) (511)
n
 x
= lim n loge 1 − (512)
n→∞ n

loge (1) 0
= = (513)
(1/n) 0

loge (1 − nλ )
⇒ lim loge ln = lim (514)
n→∞ n→∞ (1/n)
Apply L’Hopital’s Rule:

d λ . d
⇒ loge (1 − ) (1/n) (515)
dn n dn
λ
λ
n2 (1− n ) −λ
= lim 1 = lim (516)
n→∞ − 2 n→∞ (1 − λ )
n n

⇒ lim loge ln = −λ (517)


n→∞

λ n
 
∴ lim 1 − = e−λ (518)
n→∞ n

(4) Let p = P (E)

X ∼ Geom(p) (519)

73
74 CHAPTER 13. LECTURE 14

X = Number of trials required until 1st success event occurs. Once the 1st success occurs, you stop.
(520)

Trials are i.i.d. ∼ Bern(p) (521)

P [X = n] = q n−1 p, n = 1, 2, 3, . . . (522)

Xj ∼ Bern(p) (523)

P [Xj = x] = px q (1−x) , x = 0, 1, . . . (524)


(2) Y ∼ Bin(n, p) : i.i.d Bernoulli(p) trials
n  
X n k n−k
Y = Xi , P [Y = k] = p q , k = 0, 1, 2, . . . , n (525)
k
i=1

p→0, np→λ
(3) Bin(n, p) −−−−−−−→ Pois(λ)

e−λ λk
Z ∼ Pois(λ), P [Z = k] = , λ ≥ 0, k = 0, 1, 2, . . . (526)
k!

(5) GR ∼ NB(R, p) (Negative Binomial) ⇒ (Pascal Distribution)

GR = Number of trials required until the Rth success event occurs. (527)

q = (R + 1), (R + 2), . . . (528)

 
q − 1 R (q−R)
P [GR = q] = p q (529)
R−1

P [GR = q] = P [At (q−1)th trial we see (R−1) success events, and at the q th trial we see Rth trial]
(530)

= P [At (q − 1)th trial we see (R − 1) success events] × p = P (E) (531)

X ∼ Bin(n, p) (532)
 
n k n−k
P [X = k] = p q , k = 0, 1, 2, . . . , n (533)
k

Recall: 
a
b > 1 ⇒ a > b

a
a > 0, b > 0 ⇒ b > 1 if and only if a > b (534)

a
b <1⇒a<b
75

n − (R − 1) p
⇒ P [X = R] = P [X = R − 1] × × (535)
R q
(R − 1)!(n − R + 1)! p
= × (536)
R!(n − R)! q
P [X = R] (n − R + 1) p
= × (537)
P [X = R − 1] R q

If R < (n + 1)p ⇒ P (X = R) > P (X = R − 1) (538)


If R > (n + 1)p ⇒ P (X = R) < P (X = R − 1) (539)
If R = (n + 1)p ⇒ P (X = R) = P (X = R − 1) (540)
76 CHAPTER 13. LECTURE 14
Chapter 14

Lecture 15

14.1 Probability Distributions and Memoryless Property

Memoryless Property (MP): Only geometric distribution has Ageless property / Anti-aging
property (in discrete case).

G ∼ Geom(p) (541)

p = P (E) (Prob. of Success), q = (1 − p) = P (E c ) (Prob. of Failure). G is the No. of Trials


for getting E first time.

P [G = k] = q (k−1) p (542)

[k = 1, 2, 3, . . .] (543)

E1 stop E2 stop E3 stop

q p q p q p

0 E1c E2c E3c


After 1st Trial

The coin does not give you preference [Ageless property of a coin - coin does not age].
Event [G > t]:

[G > t] = [G = t + 1] ∪ [G = t + 2] ∪ [G = t + 3] ∪ [G = t + 4] ∪ . . . (544)

Disjoint union. (545)

A
P [G > t] =3 P [G = t + 1] + P [G = t + 2] + P [G = t + 3] + . . . (546)


X
= P [G = t + k] (547)
k=1

77
78 CHAPTER 14. LECTURE 15

14.1.1 Calculation of P [G > t]


Using the geometric series sum:

X
P [G > t] = q (k−1) p (548)
k=t+1
X∞
t
=q p q (k−t−1) (549)
k=t+1
X∞
t j
=q p q (Let j = k − t − 1) (550)
j=0

= q p q + q1 + q2 + . . .
t 0
 
(551)
 

X 1 1
 qj = =  (552)
1−q p
j=0
1
= qtp · (553)
p
= qt (554)

P [G > t] = q t (555)
[Till time t the event E has not occurred].

P [{E1c E2c · · · Etc }] = q · q · q · · · q (t times) = q t (556)

It is also called a tail event.

Body
Tail

14.1.2 Image 1: Memoryless Property - Conditional Probability


Geometric Probability Mass Function (PMF)
Pk = P [G = k] = q k−1 p (557)
The distribution is geometrically decreasing (by a factor of q).
p
p
qp
q2p
q3p
q4p

k
1 2 33
t= 4 5

Conditional Probability P [G = k | G > t]


P (G = k ∩ G > t)
P [G = k | G > t] = (558)
P (G > t)
14.1. PROBABILITY DISTRIBUTIONS AND MEMORYLESS PROPERTY 79

ˆ A = [G = k], B = [G > t]. This is conditional probability.


ˆ A ∩ B = [G = k] ∩ [G > t] = [G = k] (Since [G = k] ⊂ [G > t] for k > t).
P (G = k) q (k−1) p
P [G = k | G > t] = = (559)
P (G > t) qt

= q k−1−t p = q [(k−t)−1] p (560)

14.1.3 Memoryless Property


Memoryless Property (MP) in brief:

P [G > t + ∆ | G > t] = P [G > ∆] (561)

ˆ If does not matter how much time have you waited. The coin won’t give you preference.
ˆ Waited extra ∆ units of time, but not arrived.
Using the general conditional probability formula:

P [G > t + ∆ ∩ G > t] P [G > t + ∆]


P [G > t + ∆ | G > t] = = (562)
P [G > t] P [G > t]

q t+∆
= = q∆ (563)
qt

= P [G > ∆] (564)

t=0
P [G > (t + ∆) | G > t] =⇒ P [G > (0 + ∆) | G > 0]

Diagrams of Time Intervals:


t waited ∆

0 1 2 (t − 1) t (t +(∆
1) − 1)∆ Priya

Past history
Waited t units of time, but not arrived.
t=0 ∆−1 ∆ Gowathi

Sample space Ω = {G > 0}.
Distribution says probability is the same.
Application → Fusing of a bulb, Markov chain (Google search engine).

80 CHAPTER 14. LECTURE 15

14.1.4 Expected Value and Variance of Bernoulli RV


Bernoulli Distribution Bern(p):
Ω = {ω1 , ω2 } (565)

P [X = x] = px q (1−x) , x ∈ {0, 1} (566)


(
1 w.p. p (ω1 )
Xj ∼ Bern(p) =⇒ Xj = (567)
0 w.p. q = (1 − p) (ω2 )

Expectation E(Xj ):

µj = E(Xj ) = X(ω1 )P (ω1 ) + X(ω2 )P (ω2 ) = 1 · p + 0 · q = p (568)

1
X
= xP [Xj = x] (Event based) (569)
x=0

=0·q+1·p=p (570)

E(Xj ) = p (571)

Variance Var(Xj ):
Var(X) = E(X 2 ) − µ2 = E[(X − µ)2 ] (572)
(
12 = 1 w.p. p
Xj2 = (573)
02 = 0 w.p. q

E(Xj2 ) = 1 · p + 0 · q = p (574)

Var(Xj ) = p − p2 = p(1 − p) = pq (575)

p
q

0 1
µj = p

14.1.5 Random Variables and Expected Value Formula


Random Variable (RV): X : Ω → R (outcome based).
Expected Value E[X]:
X
E[X] = P (ω)X(ω) (outcome based) (576)
ω∈Ω

X
= xP [X = x] (Event based) (577)
x∈SX


14.1. PROBABILITY DISTRIBUTIONS AND MEMORYLESS PROPERTY 81

14.1.6 Image 3: Law of Expectation (LOE) and Law of Variance (LOV)


Binomial Distribution (Sup. Notes)
Y ∼ Bin(n, p).  
n k (n−k)
P [Y = k] = p q , k ∈ SY = {0, 1, 2, . . . , n} (578)
k
Y can be expressed as Y = nj=1 Xj , where Xj ∼ Bern(p).
P

Expectation E(Y ):
n  
X X n k (n−k)
µY = E(Y ) = yP (Y = y) = k p q (579)
k
y∈SY k=0

LOE Pn Pn
Using LOE: E(Y ) = j=1 E(Xj ) = j=1 p

E(Y ) = p + p + · · · + p = np (580)
| {z }
n times

Variance Var(Y ):
n
LOV
X
Var(Y ) = Var(Xj ) only when they are Independent (581)
j=1

n
X
Var(Y ) = pq = npq (582)
j=1

→ Express it in some Bernoulli (583)

14.1.7 Covariance and Independence


Z = (X + Y ) (584)

Var(X + Y ) = Var(X) + Var(Y ) + 2cov(X, Y ) (585)

cov(X, Y ) = E(XY ) − E(X)E(Y ) (586)

ˆ E(XY ) = P xyP (x = x, y = y) [Joint Distribution]


x,y

ˆ If X ⊥ Y (Independent RV’s), P (x = x, y = y) = P (x = x)P (y = y).


" #" #
X X
E(XY ) = xP (x = x) yP (y = y) = E(X)E(Y ) (587)
x y

Strong Property Weak Property


If X ⊥ Y independent RV’s =⇒ cov(X, Y ) = 0
If cov(X, Y ) = 0 ⇏ X ⊥ Y are independent
Note: LOE always holds regardless of whether the RV’s are independent or not. But LOV
only holds if RV’s are independent or uncorrelated (cov(X, Y ) = 0).

82 CHAPTER 14. LECTURE 15

14.1.8 Image 0: Conditional Probability of Geometric RV (Shift)


Formula:
P [G = k | G > t] = q [(k−t)−1] p (588)

q/p
Waited t units of time for event of occur and it has not occurred.

k starts from (t + 1)
t+1 t+2
Distribution remain the same. It only shifts.
Chapter 15

Lecture 16

15.1 Probability Measure and Axioms


15.1.1 Experiment and Sample Space
Experiment −→ Ω −→ ωi ∈ Ω (589)
Each ωi is an outcome. The set of all possible outcomes is the sample space Ω.

F ⊆ P(Ω) (590)
where F is the set of events (measurable sets).

(Ω, F) = measurable space (591)

15.1.2 Probability Measure


A probability measure is a function

P : F → [0, 1] (592)

that assigns a probability value to each event in F.

15.1.3 Axioms of Probability


1. Non-negativity and Normalization:

P (Ω) = 1 (593)

i.e., the probability of the sample space is 1.

2. For any event E ∈ F,


0 ≤ P (E) ≤ 1 (594)

3. Countable Additivity: For a countable collection of disjoint sets {Ai } ⊆ F such that

Ai ∩ Aj = ∅, i ̸= j, (595)

we have !
[ X
P Ai = P (Ai ) (596)
i i

83
84 CHAPTER 15. LECTURE 16

15.1.4 Countable Collection of Sets


A countable collection of sets can be written as:

A1 , A2 , A3 , . . . , AN , . . . (597)

and can be either:

ˆ Finite: A , A , . . . , A for some finite N


1 2 N

ˆ Infinite: A , A , A , . . . , A , A , . . .
1 2 3 N N +1

15.2 Concept of Random Variable


A random variable (R.V) is a function

X : Ω −→ R (598)

15.2.1 Discrete Random Variable


A random variable X is called discrete if the set of possible outcomes SX is finite or
countable.

PMF: PX (x) = P (X = x) (599)

The PMF (Probability Mass Function) maps the random variable to the interval [0, 1]:

P : X(Ω) −→ [0, 1] (600)

The PMF completely specifies the discrete random variable.


X
P (X = x) = 1 (601)
x∈SX

15.2.2 Continuous Random Variable


A continuous random variable is one that can take real values.
Examples: Temperature in Chennai, weight of a person, etc.
Definition: A continuous random variable has a Probability Density Function (PDF)
p(x) such that
p : R → [0, ∞) (602)

and Z ∞
p(x) dx = 1 (603)
−∞

The PDF cannot take negative values.


Definition: A random variable X is continuous if there exists a continuous PDF such that for
every a < b,
Z b
P [a ≤ X ≤ b] = p(x) dx (604)
a
15.3. CUMULATIVE DISTRIBUTION FUNCTION (CDF) 85

15.3 Cumulative Distribution Function (CDF)


Definition: A Cumulative Distribution Function (CDF) of a continuous random
variable X is defined as: Z x
FX (x) = P (X ≤ x) = fX (t) dt (605)
−∞

where fX (x) is the Probability Density Function (PDF) of X.

fX (x)

x
P (X ≤ x)

The CDF satisfies the following properties:

(a)
P (X > x) = 1 − FX (x), since [X > x] = [X ≤ x]c (606)

(b) For a < b,


Z b
P (a ≤ X ≤ b) = FX (b) − FX (a) = fX (t) dt (607)
a

(c)
lim FX (x) = 1 (608)
x→+∞

(d)
lim FX (x) = 0 (609)
x→−∞

15.4 Joint Distribution


If X and Y are discrete random variables, then they can be completely described by the joint
PMF.

P : R2 −→ [0, 1] (610)
such that
P (X = x, Y = y) = P (X = x ∩ Y = y) = P (x, y) (611)
The joint PMF satisfies: XX
P (x, y) = 1 (612)
x y

15.4.1 Example: Two Fair Dice


Two fair dice are rolled ⇒ (n1 , n2 )

X = sum of numbers = n1 + n2 , (613)


Y = larger of the two numbers = max(n1 , n2 ) (614)
86 CHAPTER 15. LECTURE 16

15.5 Joint Distribution Table


X\Y 1 2 3 4 5 6 Row Sum
2 (2,1) (2,2) (2,3) (2,4) (2,5) (2,6) 1/36
3 (3,1) (3,2) (3,3) (3,4) (3,5) (3,6) 2/36
4 (4,1) (4,2) (4,3) (4,4) (4,5) (4,6) 3/36
5 (5,1) (5,2) (5,3) (5,4) (5,5) (5,6) 4/36
6 (6,1) (6,2) (6,3) (6,4) (6,5) (6,6) 5/36
7 (7,1) (7,2) (7,3) (7,4) (7,5) (7,6) 6/36
8 (8,1) (8,2) (8,3) (8,4) (8,5) (8,6) 5/36
9 (9,1) (9,2) (9,3) (9,4) (9,5) (9,6) 4/36
10 (10,1) (10,2) (10,3) (10,4) (10,5) (10,6) 3/36
11 (11,1) (11,2) (11,3) (11,4) (11,5) (11,6) 2/36
12 (12,1) (12,2) (12,3) (12,4) (12,5) (12,6) 1/36
Column Sum 6/36 6/36 6/36 6/36 6/36 6/36 36/36 = 1

Marginal PMF of X:
X
PX (X = xi ) = PX,Y (xi , y) ⇒ Row Sum (615)
y∈SY

Marginal PMF of Y:
X
PY (Y = yj ) = PX,Y (x, yj ) ⇒ Column Sum (616)
x∈SX

15.6 Independent Random Variables


X and Y are independent random variables if

PX,Y (xi , yj ) = PX (xi ) · PY (yj ) ∀xi , yj (617)

Example:
1
PX,Y (2, 1) = (618)
36
1 1
PX (2) = , PY (1) = (619)
36 36

⇒ X and Y are not independent random variables. (620)

15.7 Example
Pick 2 balls with replacement from an urn containing 4 white and 6 black balls.
( (
1, if the first ball is white 1, if the second ball is white
X1 = X2 = (621)
0, if the first ball is black 0, if the second ball is black

X1 , X2 ∈ {0, 1} (622)
Joint Probability Table:
15.8. I.I.D RANDOM VARIABLES 87

Y=0 Y=1 Row Sum


6 6 36 6 4 24 60
X=0 P (0, 0) = 10 · 10 = 100 P (0, 1) = 10 · 10 = 100 P (X = 0) = 100
4 6 24 4 4 16 40
X=1 P (1, 0) = 10 · 10 = 100 P (1, 1) = 10 · 10 = 100 P (X = 1) = 100
Column Sum P (Y = 0) = 36+24 60
100 = 100 P (Y = 1) = 24+16 40
100 = 100 =1

15.8 i.i.d Random Variables


Let X1 , X2 , . . . , Xn be random variables.
They are said to be i.i.d. (independent and identically distributed) if:

fXi (x) = f (x) ∀i (623)

and
n
Y
fX1 ,X2 ,...,Xn (x1 , x2 , . . . , xn ) = fXi (xi ) (624)
i=1

15.9 Independent Random Variables


X and Y are independent random variables if:

PX,Y (x, y) = PX (x) PY (y) ∀x, y (625)

Expectation property for independent variables:


" #" #
XX X X
E(XY ) = P (x, y) xy = xPX (x) yPY (y) = E(X) · E(Y ) (626)
x y x y

15.10 Expectation and Covariance


E(XY ) = E(X) · E(Y ) (627)

Cov(X, Y ) = E(XY ) − E(X)E(Y ) = 0 (628)

If X and Y are uncorrelated ̸⇒ X and Y are independent. (629)

If X and Y are independent ⇒ X and Y are uncorrelated. (630)

15.11 Continuous Random Variable X


X ∼ N (µ, σ 2 ) (631)

µ = E(X) (632)

σ 2 = V ar(X) (633)

p
σ = SD(X) = + V ar(X) (634)
88 CHAPTER 15. LECTURE 16

µ
Tail Tail
Bell
−σ +σ
x

"  #
1 x−µ 2

1
PDF of X ⇒ fX (x) = √ exp − (635)
σ 2π 2 σ

0 < fX (x) < ∞ (636)


Chapter 16

Lecture 17

16.1 Derivation of Expected Value E[Y ]


The expected value E[Y ] of a discrete random variable Y is defined as the sum of each
possible value i multiplied by its probability P (Y = i).

n
X
E[Y ] = i · P (Y = i) (637)
i=0
n  
X n i n−i
= i· pq (638)
i
i=0

16.1.1 Expansion of the Summation


Expanding the first few terms of the summation:

E[Y ] = 0 · P (Y = 0) + 1 · P (Y = 1) + 2 · P (Y = 2) + . . . + n · P (Y = n) (639)
     
n 0 n−0 n 1 n−1 n 2 n−2
=0· p q +1· p q +2· p q + ... (640)
0 1 2
n(n − 1) 2 n−2
= 0 + 1 · npq n−1 + 2 · p q + ... (641)
2
= 0 + npqn−1 + n(n − 1)p2 qn−2 + . . . (642)

16.1.2 Formal Manipulation of the Term


n

To simplify the summation, we can substitute the definition of the binomial coefficient i :

n
X n!
E[Y ] = i· pi q n−i (643)
i!(n − i)!
i=0

Note on Simplification: Since the term for i = 0 is 0, the sum can start from i = 1. For
n·(n−1)!
i ≥ 1, we use the identity i · i!1 = (i−1)!
1
and n!
i! = i·(i−1)! .

89
90 CHAPTER 16. LECTURE 17

n
X n!
E[Y ] = i· pi q n−i (644)
i!(n − i)!
i=1
n
X n!
= pi q n−i (645)
(i − 1)!(n − i)!
i=1
n
X (n − 1)!
= np pi−1 q n−i (646)
(i − 1)!(n − i)!
i=1

Let k = i − 1 and m = n − 1. The summation simplifies to:


m
X m!
np pk q m−k (647)
k!(m − k)!
k=0

The summation term is the sum of the PMF of a Bin(n − 1, p) distribution, which must equal
1.
n−1
X n − 1
pk q (n−1)−k = (p + q)n−1 = 1n−1 = 1 (648)
k
k=0

Therefore, the final result is:


E[Y ] = np · 1 = np (649)

16.2 Completing the Derivation of E[Y ]


The derivation for the expected value of a Binomial Random Variable Y ∼ Bin(n, p) starts
from the summation form:
n  
X n i n−i
E[Y ] = i· pq (650)
i
i=1

16.2.1 Algebraic Manipulation (from i = 1)


n! n!
Using the identity i · i!(n−i)! = (i−1)!(n−i)! :

n
X n!
E[Y ] = pi q n−i (651)
(i − 1)!(n − i)!
i=1

16.2.2 Factoring out np


We factor out np from the expression, rewriting n! = n · (n − 1)! and pi = p · pi−1 :
n
X (n − 1)!
E[Y ] = np pi−1 q n−i (652)
(i − 1)!(n − i)!
i=1

16.2.3 Change of Index and Variable Substitution


Introduce a new index k = i − 1. When i = 1, k = 0. When i = n, k = n − 1. The term
(n − i)! becomes [n − (k + 1)]! = [(n − 1) − k]!.
The summation becomes:
n−1
X (n − 1)!
E[Y ] = np pk q (n−1)−k (653)
k![(n − 1) − k]!
k=0
16.3. DERIVATION OF E[Y (Y − 1)] 91

Now, let’s introduce a new variable for simplicity: m = n − 1.

m
X m!
E[Y ] = np pk q m−k (654)
k![m − k]!
k=0

16.2.4 Recognizing the Binomial Theorem


m
pk q m−k .

The term inside the summation is the PMF of a Bin(m, p) distribution, which is k

m  
X m
E[Y ] = np pk q m−k (655)
k
k=0

The sum of the probabilities for *any* distribution must equal 1. Specifically, by the Binomial
Theorem, the summation is the expansion of (p + q)m :

m  
X m k m−k
p q = (p + q)m (656)
k
k=0

Since p + q = 1 (the sum of success and failure probabilities):

(p + q)m = (1)m = 1 (657)

16.2.5 Final Result


Substituting 1 back into the E[Y ] expression:

E[Y ] = np · (p + q)m (658)

E[Y ] = np · (1)m (659)

E[Y ] = np (660)

16.3 Derivation of E[Y (Y − 1)]


To find the variance Var[Y ], we first calculate the second factorial moment E[Y (Y − 1)].

16.3.1 Setup and Initial Summation


The formula is:
n
X
E[Y (Y − 1)] = i(i − 1) · P (Y = i) (661)
i=0

n
pi q n−i :

Substituting the PMF P (Y = i) = i

n  
X n i n−i
E[Y (Y − 1)] = i(i − 1) pq (662)
i
i=0
92 CHAPTER 16. LECTURE 17

16.3.2 Algebraic Manipulation


The terms for i = 0 and i = 1 are zero (since i(i − 1) = 0 in both cases), so we start the
summation from i = 2. Substituting ni = i!(n−i)!
n!
:

n
X n!
E[Y (Y − 1)] = i(i − 1) · pi q n−i (663)
i!(n − i)!
i=2

1 1
We use the identity i(i − 1) · i! = (i−2)! and i(i − 1) · n!/i! = n(n − 1)(n − 2)!/(i − 2)!:

n
X n!
E[Y (Y − 1)] = pi q n−i (664)
(i − 2)!(n − i)!
i=2

16.3.3 Factoring out n(n − 1)p2


We factor out n(n − 1)p2 from the expression, rewriting n! = n(n − 1)(n − 2)! and
pi = p2 · pi−2 :
n
X (n − 2)!
E[Y (Y − 1)] = n(n − 1)p2 pi−2 q n−i (665)
(i − 2)!(n − i)!
i=2

16.3.4 Change of Index and Variable Substitution


Introduce a new index k = i − 2. When i = 2, k = 0. When i = n, k = n − 2. The term
(n − i)! becomes [n − (k + 2)]! = [(n − 2) − k]!.
The summation becomes:
n−2
2
X (n − 2)!
E[Y (Y − 1)] = n(n − 1)p pk q (n−2)−k (666)
k![(n − 2) − k]!
k=0

Now, let m = n − 2.
m
X m!
E[Y (Y − 1)] = n(n − 1)p2 pk q m−k (667)
k![m − k]!
k=0

16.3.5 Recognizing the Binomial Theorem


The term inside the summation is the PMF of a Bin(m, p) distribution, which sums to
(p + q)m :
m  
X m k m−k
p q = (p + q)m (668)
k
k=0

Since p + q = 1:
E[Y (Y − 1)] = n(n − 1)p2 (p + q)m (669)

E[Y (Y − 1)] = n(n − 1)p2 (1)m (670)

16.3.6 Final Result


E[Y (Y − 1)] = n(n − 1)p2 (671)
16.4. VARIANCE OF THE BINOMIAL DISTRIBUTION Y ∼ Bin(n, p) 93

16.4 Variance of the Binomial Distribution Y ∼ Bin(n, p)


The variance is calculated using the formula:

Var[Y ] = E[Y 2 ] − (E[Y ])2 (672)

A common method for discrete distributions is to use the factorial moment E[Y (Y − 1)]:

E[Y 2 ] = E[Y (Y − 1)] + E[Y ] (673)

Substituting this into the variance formula:

Var[Y ] = E[Y (Y − 1)] + E[Y ] − (E[Y ])2 (674)

16.4.1 Substitution of Moments


From the previous derivations, we have:

ˆ E[Y ] = np
ˆ E[Y (Y − 1)] = n(n − 1)p 2

Substituting these into the variance equation:

Var[Y ] = n(n − 1)p2 + [np] − [np]2


 
(675)
2 2 2 2 2
= (n p − np ) + np − n p (676)
22 2 22
n p − np + np − 
=  n p  (677)
= np − np2 (678)
= np(1 − p) (679)

Since 1 − p = q, the variance is:


Var[Y ] = npq (680)

16.5 Expected Value of the Poisson Distribution


Let X be a Poisson Random Variable with parameter λ.

X ∼ Pois(λ) (681)

The Probability Mass Function (PMF) is:

e−λ λi
P [X = i] = , for i = 0, 1, 2, . . . , ∞ (682)
i!

16.5.1 Derivation of E[X]


The expected value E[X] is defined as:

X
E[X] = i · P (X = i) (683)
i=0

Substituting the PMF:



X e−λ λi
E[X] = i· (684)
i!
i=0
94 CHAPTER 16. LECTURE 17

16.5.2 Simplification and Series Recognition


The term for i = 0 is 0, so we start the summation from i = 1:


X e−λ λi
= i· (685)
i!
i=1

Factor out the constant term e−λ and simplify the factorial term i/i! = 1/(i − 1)!:


X λi
= e−λ (686)
(i − 1)!
i=1

Factor out one λ from λi = λ · λi−1 :



−λ
X λi−1
= λe (687)
(i − 1)!
i=1

Introduce a new index k = i − 1. When i = 1, k = 0. When i = ∞, k = ∞.


−λ
X λk
= λe (688)
k!
k=0

16.5.3 Final Result


P∞ λk
The summation k=0 k! is the Taylor series expansion of eλ .


X λk
= eλ (689)
k!
k=0

Substituting this back:


E[X] = λe−λ · eλ (690)

E[X] = λe−λ+λ (691)

E[X] = λe0 (692)

The expected value is:


E[X] = λ (693)

16.6 Variance of the Poisson Distribution


The variance is defined as:
Var[X] = E[X 2 ] − (E[X])2 (694)

To simplify the calculation of E[X 2 ], we first calculate the second factorial moment
E[X(X − 1)]:
Var[X] = E[X(X − 1)] + E[X] − (E[X])2 (1) (695)
16.6. VARIANCE OF THE POISSON DISTRIBUTION 95

16.6.1 Derivation of E[X(X − 1)]


The factorial moment is:

X
E[X(X − 1)] = i(i − 1) · P (X = i) (696)
i=0

e−λ λi
Substituting the PMF P (X = i) = i! :

X e−λ λi
E[X(X − 1)] = i(i − 1) · (697)
i!
i=0

The terms for i = 0 and i = 1 are zero, so we start the summation from i = 2.

X e−λ λi
= i(i − 1) · (698)
i!
i=2

Simplification and Series Recognition


Factor out the constant e−λ and simplify i(i − 1) · 1
i! = 1
(i−2)! :

X λi
= e−λ (699)
(i − 2)!
i=2

Factor out λ2 from λi = λ2 · λi−2 :



X λi−2
= λ2 e−λ (700)
(i − 2)!
i=2

Introduce a new index k = i − 2.



2 −λ
X λk
=λ e (701)
k!
k=0
P∞ λk
The summation k=0 k! is the Taylor series expansion of eλ .

X λk
= eλ (702)
k!
k=0

Thus, the factorial moment is:


E[X(X − 1)] = λ2 e−λ · eλ = λ2 (703)

16.6.2 Final Variance Calculation


Substitute the derived moments into equation (1):
ˆ E[X(X − 1)] = λ 2

ˆ E[X] = λ
Var[X] = E[X(X − 1)] + E[X] − (E[X])2 (704)
= λ2 + λ − (λ)2 (705)
= λ2 + λ − λ2 (706)
The variance of a Poisson random variable is:
Var[X] = λ (707)
96 CHAPTER 16. LECTURE 17

16.7 Geometric Distribution Derivation X ∼ Geom(p)


The Geometric distribution models the number of **Bernoulli trials** required to get the first
success. Let p be the probability of success on any single trial, and q = 1 − p be the probability
of failure. The random variable X is the trial number where the first success occurs.

16.7.1 Probability Mass Function (PMF)


The PMF is the probability of n − 1 failures followed by the first success on the n-th trial:

P (X = n) = q n−1 p, for n = 1, 2, 3, . . . (708)

16.7.2 Expected Value E[X] (Derivation)


The expected value E[X] is defined as:

X
E[X] = n · P (X = n) (709)
n=1

Substituting the PMF:



X
E[X] = n · q n−1 p (710)
n=1

Factorization
Factor out the constant term p:

X
E[X] = p nq n−1 (711)
n=1

Recognizing the Geometric Series Derivative


The summation ∞
P n−1 is the derivative of the standard infinite geometric series
P∞ n
n=1 nq n=0 q
with respect to q. The standard geometric series is:

X 1
qn = 1 + q + q2 + q3 + . . . = , for |q| < 1 (712)
1−q
n=0

The derivative with respect to q gives:


∞ ∞
!
d X
n
X
q = nq n−1 (713)
dq
n=0 n=1

The value of the derivative is:


 
d 1 d 1
= (1 − q)−1 = −1(1 − q)−2 (−1) = (714)
dq 1 − q dq (1 − q)2

Final Result
Substitute the result of the summation back into the E[X] equation:
"∞ #
X
E[X] = p · nq n−1 (715)
n=1
16.8. NOTE ON VARIANCE CALCULATION 97

1
E[X] = p · (716)
(1 − q)2
Since 1 − q = p:
1
E[X] = p · (717)
p2

1
E[X] = (718)
p

16.8 Note on Variance Calculation


The variance σ 2 can be calculated using the expected value of X(X − 1) or by using the
relationship:
σ 2 = E[X 2 ] − (E[X])2 (719)

Using the auxiliary moment:

σ 2 = E[X(X − 1)] + E[X] − (E[X])2 (720)

16.9 Variance of the Geometric Distribution X ∼ Geom(p)


The derivation concludes with the final algebraic steps to find the variance Var[X].

16.9.1 Using the Identity Var[X] = E[X 2 ] − (E[X])2


The notes implicitly use the results from the derivation of the second moment, E[X 2 ], which is
related to the derivative of the geometric series twice. The key intermediate results are:

ˆ Expected Value: E[X] = 1


p

ˆ Second Moment (Final Form after substitution): E[X ] = 2 2−p


p2

Substituting the moments into the variance formula Var[X] = E[X 2 ] − (E[X])2 :
 2
2−p 1
Var[X] = 2
− (721)
p p
2−p 1
= − 2 (722)
p2 p
(2 − p) − 1
= (723)
p2
1−p
= (724)
p2

Since 1 − p = q (the probability of failure):

q
Var[X] = (725)
p2
98 CHAPTER 16. LECTURE 17

16.10 Negative Binomial Distribution


The notes introduce the **Negative Binomial Distribution** (NB), which generalizes the
Geometric distribution. While the Geometric distribution counts the number of trials until the
**first** success, the Negative Binomial Distribution counts the number of trials until the
**r-th** success.
Let Xn be the number of trials required to observe **r successes**.

Xn ∼ NB(r, p) (726)

16.10.1 Probability Mass Function (PMF) of NB(r, p)


The event Xn = n means that the r-th success occurred exactly on the n-th trial. This implies:

ˆ There were exactly r − 1 successes in the first n − 1 trials.


ˆ The n-th trial was a success.
The probability of r − 1 successes in n − 1 trials is given by the Binomial PMF:
n−1 r−1 (n−1)−(r−1)
r−1 p q . The probability of the final success is p.
Multiplying these gives the PMF:
 
n − 1 r−1 n−r
P (Xn = n) = p q ·p (727)
r−1
 
n − 1 r n−r
P (Xn = n) = p q (728)
r−1
for n = r, r + 1, r + 2, . . .
Part II

Statistical Inference

99
Chapter 17

Lecture 2: The Geometric and


Exponential Distributions

17.1 The Geometric Distribution


17.1.1 Definition and Context
Following our discussion on Discrete Random Variables (DRVs), we introduce the Geometric
Distribution. Consider a sequence of independent Bernoulli trials, where each trial has two
possible outcomes: ”Success” or ”Failure.”
Let T be a Discrete Random Variable representing the number of trials required to
observe the first success. Alternatively, this can be viewed as the ”waiting time” (in
discrete steps) for the event E to occur.
If we define the probability of success as p = P (E), then the probability of failure is given by
the complement:
q = 1 − p = P (E c ) (729)
A random variable T following this structure is denoted as:

T ∼ Geom(p) (730)

17.1.2 Probability Mass Function (PMF)


The Probability Mass Function (PMF) quantifies the likelihood that the first success occurs
exactly at the i-th trial. For this to happen, we must observe i − 1 failures followed
immediately by 1 success.

P (T = i) = q (i−1) p, for i = 1, 2, 3, . . . (731)

17.1.3 Tail Probability and Cumulative Distribution Function (CDF)


In reliability and survival analysis, we are often interested in the probability that the event
has not occurred by time i. This is known as the Tail Probability.
The probability that more than i trials are needed (i.e., the first i trials were all failures) is:

P (T > i) = q × q × · · · × q = q i (732)
| {z }
i times

From this, we derive the Cumulative Distribution Function (CDF), FT (i), which represents
the probability that the event occurs on or before trial i:

FT (i) = P (T ≤ i) = 1 − P (T > i) = 1 − q i (733)

101
102CHAPTER 17. LECTURE 2: THE GEOMETRIC AND EXPONENTIAL DISTRIBUTIONS

17.1.4 Properties of the Geometric Distribution


The Memoryless Property
The Geometric distribution is the unique discrete distribution that possesses the memoryless
property.
Definition 17.1.1. The memoryless property states that the probability of needing more than
t + ∆ trials, given that we have already surpassed t trials, is equivalent to the probability of
needing more than ∆ trials from the start. The past history of failures does not influence the
future waiting time.
Formally:
P (T > t + ∆ | T > t) = P (T > ∆) (734)

The Shifting Property


This property is a direct consequence of the memoryless nature. It allows us to ”shift” the
distribution’s origin. Using the definition of Conditional Probability:
P [(T = i) ∩ (T > j)]
P (T = i | T > j) = (735)
P (T > j)
Assuming i > j (since we condition on passing j), the intersection implies T = i. Substituting
the PMF and Tail Probability:
P (T = i) q (i−1) p
= = q (i−1−j) p = q (i−j)−1 p (736)
qj qj

Shifting Property
j=1
i→i−j

P Pq P Pq

i
1 2 3 4 4
j= 5 6

i has gone to 5

17.2 The Exponential Distribution


17.2.1 Definition and Probability Density Function (PDF)
We now transition to Continuous Random Variables (CRV). The Exponential Distribution is
often described as the continuous limit of the Geometric Distribution.
Let X (or J) be a continuous random variable representing the waiting time until an event
occurs in a Poisson process. We write X ∼ Exp(λ), where λ > 0 is the rate parameter.
1
λ∼ (Dimensions of Frequency) (737)
time
The Probability Density Function (PDF) is defined as:
(
λe−λx x ≥ 0
fX (x) = (738)
0 x<0
17.3. PROPERTIES OF THE EXPONENTIAL DISTRIBUTION 103

Important Note: As discussed in Lecture 1, fX (x) is not a probability. It represents the


density of probability per unit time. Probabilities are calculated as the area under the
curve.

Figure 2.1: PDF of Exponential Distribution

λe−λx
fX (x) 1

0.5

Probability Area
0
0 1 2 3 4 5
x (time)

Figure 17.1: The probability P (a ≤ X ≤ b) is the integral (area) of the PDF.

17.2.2 Cumulative Distribution Function (CDF)


The CDF, FX (x), gives the probability that the event occurs by time x. It is found by
integrating the PDF:
Z x
FX (x) = λe−λt dt (739)
0
 −λt x
e
=λ (740)
−λ 0
= −[e−λx − 1] = 1 − e−λx (741)

17.2.3 Tail Probability


The probability that the event occurs after time x (survival function) is:

P (X > x) = 1 − FX (x) = e−λx (742)

17.3 Properties of the Exponential Distribution


17.3.1 The Memoryless Property
The Exponential Distribution is the only continuous distribution that possesses the
memoryless property.

Definition 17.3.1. Similar to the discrete case, the future lifetime of the event depends only
on the current duration, not on how much time has already elapsed.

Proof Sketch: Using the definition of conditional probability for continuous variables:

P ((J > t + ∆) ∩ (J > t))


P (J > t + ∆ | J > t) = (743)
P (J > t)
104CHAPTER 17. LECTURE 2: THE GEOMETRIC AND EXPONENTIAL DISTRIBUTIONS

Since survival past t + ∆ implies survival past t, the intersection simplifies to P (J > t + ∆).
Substituting the Tail Probability formula (e−λx ):

P (J > t + ∆) e−λ(t+∆) e−λt · e−λ∆


= = = e−λ∆ (744)
P (J > t) e−λt e−λt

Notice that the result e−λ∆ is exactly equal to P (J > ∆). This confirms that the probability
depends only on the interval ∆, independent of the starting time t.

17.3.2 The Shifting Property and Visual Interpretation


The shifting property implies that if we condition on the variable J surviving past time t, the
distribution of the remaining time is identical to the original Exponential distribution starting
at 0.

Area Ratio Visual


We can visualize this property using the areas under the PDF curve. Let:
ˆA total = P (J > t) (The total ”Tail” area beyond t)

ˆA sub = P (J > t + ∆) (The smaller area beyond t + ∆)


The conditional probability is the ratio of these areas. Mathematically, this ratio scales
perfectly to match the area of the original distribution starting from 0 up to ∆.

fJ (t)

Conditional Area

t
t t+∆

Figure 17.2: Visualizing the Shifting Property: The ratio of the blue area to the gray area is
constant and depends only on ∆.

17.3.3 Conditional PDFs and CDFs


Using the properties above, we can rigorously define the conditional functions.
1. Conditional Tail Probability:
e−λt
P (J > t | J > j) = = e−λ(t−j) for t > j (745)
e−λj
2. Conditional CDF: The probability that the event occurs by time t, given it has survived
past j:
FJ|J>j (t) = 1 − e−λ(t−j) (746)
3. Conditional PDF: Differentiating the Conditional CDF with respect to t yields the
probability density for the remaining time:
d  
fJ|J>j (t) = 1 − e−λ(t−j) = λe−λ(t−j) for t ≥ j (747)
dt
Chapter 18

Lecture 3: Exponential Distribution


Derivations

18.1 Introduction
The exponential distribution is one of the most important continuous probability models,
commonly used to describe the waiting time between independent events that occur at a
constant average rate. If X is an exponentially distributed random variable with rate
parameter λ > 0, then its probability density function is

fX (x) = λe−λx , x > 0. (748)

In this derivation, we compute the mean, second moment, variance, and standard deviation of
the exponential distribution step by step. These results are important because they highlight
the memoryless property of the exponential distribution and also connect it to the Poisson
process, where λ represents the rate of event occurrence.

18.2 Derivation of E[X] for an Exponential Distribution


Given an exponential random variable X with parameter λ, its probability density function is

f (x) = λe−λx , x ≥ 0. (749)

We compute the expectation: Z ∞


E[X] = xλe−λx dx. (750)
0
Using integration by parts, let
u=x ⇒ du = dx, (751)
dv = λe−λx dx ⇒ v = −e−λx . (752)
Applying integration by parts:
∞ Z ∞
E[X] = uv − v du. (753)
0 0

Substitute the values: h i∞ Z ∞


−λx
E[X] = x(−e ) − (−e−λx ) dx. (754)
0 0

Evaluate boundary terms: - As x → ∞, xe−λx → 0. - At x = 0: xe−λx = 0.


Thus boundary part = 0 − 0 = 0.

105
106 CHAPTER 18. LECTURE 3: EXPONENTIAL DISTRIBUTION DERIVATIONS

Now compute remaining integral:


Z ∞
E[X] = e−λx dx. (755)
0

Z ∞  ∞  
1 1 1
e−λx dx = − e−λx =0− − = . (756)
0 λ 0 λ λ

1
E[X] = . (757)
λ
This matches the standard mean of an exponential distribution.

18.3 Derivations of E[X] and E[X 2 ] for an Exponential


Distribution
18.3.1 Setup
We consider a random variable X that follows an exponential distribution with rate parameter
λ > 0. Its probability density function (pdf) is

fX (x) = λe−λx , x ≥ 0. (758)

We will compute the expectation E[X] and the second moment E[X 2 ] with step-by-step
explanations.

18.3.2 1. Compute E[X]


The expectation is defined by the integral
Z ∞ Z ∞
E[X] = x fX (x) dx = x λe−λx dx. (759)
0 0

Step 1: take constants outside the integral. Since λ is a constant (with respect to x),
we write Z ∞
E[X] = λ xe−λx dx. (760)
0

R∞
Step 2: use integration by parts. To evaluate 0 xe−λx dx we apply integration by parts:
choose u = x (so du = dx) and dv = e−λx dx (so v = − λ1 e−λx ). Integration by parts gives
Z ∞   ∞ Z ∞ 
−λx 1 −λx 1
xe dx = x − e − − e−λx dx. (761)
0 λ 0 0 λ

Step 3: evaluate the boundary term. The boundary term is


 x   0 
−λx
lim − e − − e0 . (762)
x→∞ λ λ

Because the exponential e−λx decays faster than the polynomial x grows, we have
limx→∞ xe−λx = 0. Therefore the boundary term equals 0 − 0 = 0.
18.3. DERIVATIONS OF E[X] AND E[X 2 ] FOR AN EXPONENTIAL DISTRIBUTION107

Step 4: evaluate the remaining integral. We are left with


Z ∞
1 ∞ −λx
Z
−λx
xe dx = e dx. (763)
0 λ 0

The integral of the exponential is


Z ∞ ∞  
1 1 1
e−λx dx = − e−λx =0− − = . (764)
0 λ 0 λ λ

Hence Z ∞
1 1 1
xe−λx dx = · = 2. (765)
0 λ λ λ

Step 5: multiply back the constant λ. Recall E[X] = λ times that integral, so

1 1
E[X] = λ · = . (766)
λ2 λ
1
Result: E[X] = .
λ

18.3.3 2. Compute E[X 2 ]


The second moment is
Z ∞ Z ∞ Z ∞
E[X 2 ] = x2 fX (x) dx = x2 λe−λx dx = λ x2 e−λx dx. (767)
0 0 0

R∞
Step 1: integration by parts (first application). To evaluate 0 x2 e−λx dx use
integration by parts with u = x2 (so du = 2x dx) and dv = e−λx dx (so v = − λ1 e−λx ). This
yields
Z ∞ ∞ Z ∞
x2

1
x2 e−λx dx = − e−λx − − 2xe−λx dx. (768)
0 λ 0 0 λ

Step 2: boundary term vanishes. As before, limx→∞ x2 e−λx = 0 and at x = 0 the term
is 0. So the boundary term is 0. Thus
Z ∞
2 ∞ −λx
Z
2 −λx
x e dx = xe dx. (769)
0 λ 0

Step 3:
Z ∞use the previously computed integral. From the computation of E[X] we
1
found xe−λx dx = 2 . Therefore
0 λ
Z ∞
2 1 2
x2 e−λx dx = · 2 = 3 . (770)
0 λ λ λ

Step 4: multiply back the factor λ. Recall E[X 2 ] = λ times the integral, hence

2 2
E[X 2 ] = λ · 3
= 2. (771)
λ λ
2
Result: E[X 2 ] = .
λ2
108 CHAPTER 18. LECTURE 3: EXPONENTIAL DISTRIBUTION DERIVATIONS

18.3.4 3. Variance (optional)


Using the results above, the variance is
 2
2 1 2 1 1
Var(X) = E[X 2 ] − (E[X])2 = − = 2 − 2 = 2. (772)
λ2 λ λ λ λ

1
Result: Var(X) = .
λ2
Remarks: In the integrations above we repeatedly used the fact that for any positive integer
n, lim xn e−λx = 0 (exponential decay dominates any polynomial growth). This justifies
x→∞
discarding boundary terms that involve such expressions.

18.4 Variance of an Exponential Random Variable


Let X ∼ Exp(λ) with probability density function

fX (x) = λe−λx , x > 0. (773)

18.4.1 1. Mean of X
The expectation of an exponential random variable is
1
E[X] = . (774)
λ

18.4.2 2. Second Moment E[X 2 ]


For the exponential distribution,
2
E[X 2 ] = . (775)
λ2

18.4.3 3. Variance Computation


The variance is defined as
σ 2 = Var(X) = E[X 2 ] − (E[X])2 . (776)
Substitute values:  2
2 1 2 1 1
σ2 = − = 2 − 2 = 2. (777)
λ2 λ λ λ λ
Thus,
1 1
σ2 = and σ= . (778)
λ2 λ

18.4.4 4. Relation to Poisson Distribution


If
Y ∼ Poisson(λ), (779)
then
E[Y ] = λ, Var(Y ) = λ. (780)
This shows the same parameter λ appears in both the Poisson distribution (as the mean rate
of occurrences) and the exponential distribution (as the rate of arrival time), but the
distributions are different.
18.4. VARIANCE OF AN EXPONENTIAL RANDOM VARIABLE 109

18.4.5 5. Shape of the Exponential Density


The PDF of X ∼ Exp(λ) is
fX (x) = λe−λx , x > 0. (781)
It is a decreasing curve starting at λ when x = 0 and approaches 0 as x → ∞.
fX (x) = λe−λx x≥0
λ

K x
110 CHAPTER 18. LECTURE 3: EXPONENTIAL DISTRIBUTION DERIVATIONS
Chapter 19

Lecture 4: Normalization and


Sample Statistics

19.1 Normalization of Random Variables


fX (x) = λe−λ(x−k) , x>k (782)

“waited k units of time and event has not occurred.”

19.1.1 General Definition


Let X ∼ D(µ, σ 2 ), where:

ˆ RV X is distributed with distribution D having:


ˆ µ = E(X)
ˆ σ = var(X)
2

The normalized variable Z is defined as:

Z ∼ D(0, 1) (783)

Z ←− X (784)

19.1.2 Example: {1, 2, 3}


1+2+3 6
µ= = =2 (785)
3 3
Centering the data:
(1 − 2, 2 − 2, 3 − 2) −→ (−1, 0, +1) −→ 0 (786)

Calculating Variance σ 2 :

σ 2 = E(X 2 ) − µ2 (787)
2
= E(X ) − 4 (788)
1+4+9
= −4 (789)
3
14 2
= −4= (790)
3 3

111
112 CHAPTER 19. LECTURE 4: NORMALIZATION AND SAMPLE STATISTICS
 
19.1.3 Calculations for Set: − √ , √ , √
1 0 +1
2/3 2/3 2/3

− √1 + √0 + √1
2/3 2/3 2/3
µ= (791)
3
r  
3 1
=0· =0 (792)
2 3

Calculating Variance σ 2 :

σ 2 = E(X 2 ) − µ2 (793)
 2  2  2
√−1 + √0
+ √1
2/3 2/3 2/3
= − 02 (794)
3
1 1
2/3 + 0 + 2/3
= (795)
3
2×3
= =1 (796)
2×3

 
19.1.4 Calculations for Set: √1 , √2 , √3
2/3 2/3 2/3

σ 2 = E(X 2 ) − µ2 (797)
 2  2  2
√ 1 2 3
+ √ + √ !2
2/3 2/3 2/3 2
= − p (798)
3 2/3
 
1 4 9
2/3 + 2/3 + 2/3 4
= − (799)
3 2/3
 
1+4+9
2/3 4×3
= − (800)
3 2
14 × 3
= −6 (801)
2×3×3
14
= −6=7−6=1 (802)
2

19.2 Normalization and Sample Statistics


19.2.1 Shifted Exponential Distribution
We begin by defining a shifted exponential Probability Density Function (PDF). This often
represents a scenario where an event cannot occur before time k.

fX (x) = λe−λ(x−k) , x>k (803)

In this context, the parameter k signifies that we have waited k units of time and the event
has not yet occurred.
19.3. SAMPLE STATISTICS FOR n I.I.D. VARIABLES 113

19.2.2 Normalization of a Random Variable


Normalization (or standardization) is the process of transforming a random variable X with
mean µ and variance σ 2 into a new variable Z with a mean of 0 and a variance of 1.

X −µ
X ∼ D(µ, σ 2 ) =⇒ Z = ∼ D(0, 1) (804)
σ
Where the parameters are defined as:

ˆ µ = E(X)
ˆ σ = Var(X)
2

19.2.3 Proof of Mean for Z


To verify the normalization, we calculate the expected value of Z:
 
X −µ
E(Z) = E
σ
1
= [E(X) − E(µ)] (805)
σ
1
= [µ − µ] = 0 (806)
σ

In step (805), we use the linearity of expectation. Since E(X) = µ, the numerator becomes
zero in (806).

19.2.4 Proof of Variance for Z


Next, we derive the variance to show it equals 1. We use the identity σ 2 = E(X 2 ) − [E(X)]2 :

Var(Z) = E(Z 2 ) − [E(Z)]2


"  #
X − µx 2
=E − 02 (807)
σx
1
= E[X 2 + µ2x − 2Xµx ]
σx2
1
= 2 [E(X 2 ) + E(µ2x ) − 2µx E(X)] (808)
σx
1
= 2 [E(X 2 ) + µ2x − 2µ2x ]
σx
E(X 2 ) − µ2x σx2
= = =1 (809)
σx2 σx2

1
Equations (807) through (809) confirm that the scaling factor σ correctly yields a unit
variance.

19.3 Sample Statistics for n i.i.d. Variables


Consider n Independent and Identically Distributed (i.i.d.) random variables X1 , X2 , . . . , Xn ,
each following distribution D(µ, σ 2 ).
114 CHAPTER 19. LECTURE 4: NORMALIZATION AND SAMPLE STATISTICS

19.3.1 Sample Sum (Sn )


The sum of these variables is defined as:
n
X
Sn = Xi = (X1 + X2 + · · · + Xn ) (810)
i=1

Applying the Linearity of Expectation (LOE):


n
X n
X
E(Sn ) = E(Xi ) = µ = nµ (811)
i=1 i=1

For independent variables, the variance of the sum is the sum of the variances:
n
X n
X
Var(Sn ) = Var(Xi ) = σ 2 = nσ 2 (812)
i=1 i=1

19.3.2 Sample Average (X̄n )


The sample average is the sum divided by the number of observations:
Sn X1 + · · · + Xn
X̄n = = (813)
n n
The expected value remains the population mean:
 
Sn 1 nµ
E(X̄n ) = E = E(Sn ) = =µ (814)
n n n
However, the variance of the average decreases as the sample size n increases:

nσ 2 σ2
 
Sn 1
Var(X̄n ) = Var = 2 Var(Sn ) = 2 = (815)
n n n n
This result in (815) is a fundamental principle in statistics, showing that larger samples
provide more precise estimates of the mean.

19.4 Statistical Properties of Sample Metrics


19.4.1 Standardization and Verification
Normalization ensures that any random variable can be compared on a standard scale. Given
Z = X−µ
σ , we verify that the resulting mean is 0 and the variance is 1.

Proof of Zero Mean


The expected value of the standardized variable Z is calculated as follows:
 
X −µ
E(Z) = E
σ
1
= [E(X) − E(µ)] (816)
σ
1
= [µ − µ] = 0 (817)
σ
In Equation (816), we apply the linearity of expectation, and since E(X) = µ, the result in
(817) confirms the centering property.
19.5. EXPECTATION AND VARIANCE OF SAMPLE SUMS AND AVERAGES 115

Proof of Unit Variance


To show the variance is normalized to 1, we use the property Var(aX) = a2 Var(X):
 
X −µ
Var(Z) = Var
σ
1
= 2 Var(X − µ) (818)
σ
1 σ2
= 2 Var(X) = 2 = 1 (819)
σ σ
Equation (819) completes the proof that Z ∼ D(0, 1).

19.4.2 Numerical Example: Standardizing a Discrete Set


Consider the sample set {1, 2, 3}. We first find the population parameters.
1+2+3
µ= =2 (820)
3
The variance σ 2 is calculated as E(X 2 ) − µ2 :
12 + 2 2 + 3 2 14 2
σ2 = − 22 = −4= (821)
3 3 3

Standardizing the Elements

pmean µ = 2 gives the centered set {−1, 0, 1}. To finish standardization, we


Subtracting the
divide by σ = 2/3, resulting in the set:
!
−1 0 1
p ,p ,p (822)
2/3 2/3 2/3

19.5 Expectation and Variance of Sample Sums and Averages


Let X1 , X2 , . . . , Xn be i.i.d. random variables where Xi ∼ D(µ, σ 2 ).

19.5.1 The Sample Sum (Sn )


Pn
The sum is defined as Sn = i=1 Xi . Its expectation is:
n
X
E(Sn ) = E(Xi ) = nµ (823)
i=1
Because the variables are independent, the variance is:
Xn
Var(Sn ) = Var(Xi ) = nσ 2 (824)
i=1

19.5.2 The Sample Average (X̄n )


Sn
The average is defined as X̄n = n .
The variance of the sample mean is:
 
Sn
Var(X̄n ) = Var
n
1 nσ 2 σ2
= 2 Var(Sn ) = 2 = (825)
n n n
Equation (825) represents the ”Standard Error,” showing that variance decreases as sample
size n increases.
116 CHAPTER 19. LECTURE 4: NORMALIZATION AND SAMPLE STATISTICS

19.5.3 Standardizing the Sample Mean


The standardized version of the sample mean, ZX̄n , is used extensively in the Central Limit
Theorem:
X̄n − E(X̄n ) X̄n − µ
ZX̄n = = √ (826)
SD(X̄n ) σ/ n
Rearranging this gives:

 
X̄n − µ
ZX̄n = n ∼ D(0, 1) (827)
σ

19.6 Fundamentals of Statistics and Calculus


19.6.1 Normalization of Random Variables
A random variable X with a distribution D characterized by mean µ and variance σ 2 can be
transformed into a standard form Z.

Theoretical Definition
Given X ∼ D(µ, σ 2 ), we define the normalized variable Z as:
X − E(X) X −µ
Z= = (828)
SD(X) σ
By definition, Z will have a distribution D(0, 1).

Proof of Standard Properties


We verify that E(Z) = 0 using the linearity of expectation:
 
X −µ 1
E(Z) = E = [E(X) − E(µ)]
σ σ
1
= [µ − µ] = 0 (829)
σ
Next, we prove that Var(Z) = 1 using the identity σ 2 = E(X 2 ) − µ2 :
"  #
2 2 X − µx 2
Var(Z) = E(Z ) − (E(Z)) = E −0
σx
1
= E[X 2 + µ2x − 2Xµx ]
σx2
1 σ2
= 2 [E(X 2 ) + µ2x − 2µ2x ] = x2 = 1 (830)
σx σx
These steps confirm that the transformation in (828) successfully centers and scales the
variable.

19.7 Review of Calculus Fundamentals


19.7.1 Differential Calculus: The Slope
The derivative represents the instantaneous rate of change. For a function f (x), the slope of
the tangent line at point A is defined as the limit of the secant line AB as h → 0:
f (x + h) − f (x)
f ′ (x) = lim (831)
h→0 h
19.7. REVIEW OF CALCULUS FUNDAMENTALS 117

Functions changing rapidly require a small window for analysis, whereas gradual changes allow
for a larger window.

19.7.2 Integral Calculus: Area Under the Curve


To find the area under a curve y = f (x) on [a, b], we sum the areas of n small rectangles with
width ∆j x = (aj − aj−1 ) and height f (xj ):
n Z b
X ∆j x→0
f (xj )(aj − aj−1 ) −−−−→ f (x) dx (832)
n→∞ a
j=1

This lead
R xus to the Fundamental Theorem of Calculus (FTC), stating that if
F (x) = a f (t) dt, then F ′ (x) = f (x).
118 CHAPTER 19. LECTURE 4: NORMALIZATION AND SAMPLE STATISTICS
Chapter 20

Lectures 5 and 6: Moment


Generating Functions and Limit
Theorems

20.1 Moment Generating Function (MGF): Introduction


In probability theory, the Moment Generating Function (MGF) is an important tool used to
study the distributional properties of a random variable. The main purpose of the MGF is to
generate the moments of a random variable in a systematic and convenient way.
For a random variable X, the moment generating function is defined as

MX (t) = E etX ,
 
for those values of t for which the expectation exists. (833)

The term “moment generating” arises from the fact that the derivatives of the MGF evaluated
at t = 0 produce the moments of the random variable. In particular, the nth moment about
the origin is given by
(n)
E[X n ] = MX (0), (834)

provided the derivative exists.


In addition to generating moments, the MGF uniquely characterizes the probability
distribution of a random variable. That is, if two random variables have the same MGF in an
open interval containing zero, then they have the same probability distribution.
Furthermore, the MGF is especially useful in studying sums of independent random variables.
If X1 , X2 , . . . , Xn are independent random variables, then the MGF of their sum is equal to
the product of their individual MGFs. This property plays a crucial role in many theoretical
results, including the Central Limit Theorem.
Hence, the moment generating function provides a powerful and unified framework for
analyzing moments, distributions, and sums of random variables.

20.2 Power Series Expansion of a Function


Let f (x) be a “nice” function such that all derivatives exist at x = 0. Assume that f (x) can
be expressed as a power series:

f (x) = a0 + a1 x + a2 x2 + a3 x3 + · · · + an xn + · · · (835)

119
120CHAPTER 20. LECTURES 5 AND 6: MOMENT GENERATING FUNCTIONS AND LIMIT THEOREM

20.2.1 Step 1: Function value at x = 0


Substituting x = 0 in the above series, we get

f (0) = a0 (836)

Hence,
a0 = f (0) (837)

20.2.2 Step 2: First derivative


Differentiating term by term,

f ′ (x) = a1 + 2a2 x + 3a3 x2 + · · · + nan xn−1 + · · · (838)

Evaluating at x = 0,
f ′ (0) = a1 (839)

Thus,
a1 = f ′ (0) (840)

20.2.3 Step 3: Second derivative


Differentiating again,

f ′′ (x) = 2a2 + 6a3 x + · · · + n(n − 1)an xn−2 + · · · (841)

Evaluating at x = 0,
f ′′ (0) = 2a2 (842)

Hence,
f ′′ (0)
a2 = (843)
2!

20.2.4 Step 4: General nth derivative


Differentiating n times,
f (n) (x) = n!an + terms involving x (844)

Evaluating at x = 0,
f (n) (0) = n!an (845)

Therefore,
f (n) (0)
an = (846)
n!

20.2.5 Step 5: Power series representation


Substituting the coefficients back into the series, we obtain

f ′′ (0) 2 f (n) (0) n


f (x) = f (0) + f ′ (0)x + x + ··· + x + ··· (847)
2! n!
20.3. MOMENT GENERATING FUNCTION (MGF) DEFINITIONS 121

20.2.6 Final Result: Maclaurin Series


Thus, the power series expansion of f (x) about x = 0 (Maclaurin series) is


X f (n) (0)
f (x) = xn (848)
n!
n=0

f ′ (0) f ′′ (0) 2 f (3) (0) 3


f (x) = f (0) + x+ x + x + ··· (849)
1! 2! 3!
Conclusion: The power series expansion of f (x) is unique.

20.3 Moment Generating Function (MGF) Definitions


Definition. Let X be a random variable (discrete or continuous). The moment generating
function (MGF) of X is defined as

MX (t) = E(etX ), t ∈ R. (850)

20.3.1 Moments Using MGF


1st moment of X = E(X) (851)

2nd moment of X = E(X 2 ) (852)

k-th moment of X = E(X k ) (853)


The moment generating function (MGF) of a random variable X is defined as

MX (t) = E etX

(854)

20.3.2 MGF for Different Types of Random Variables


(i) Discrete random variable
If X is a discrete random variable with probability mass function fX (x), then
X
MX (t) = etx fX (x) (855)
x

(ii) Continuous random variable


If X is a continuous random variable with probability density function fX (x), then
Z ∞
MX (t) = etx fX (x) dx (856)
−∞

Thus, in both cases,


MX (t) = E etX

(857)

20.4 Using Taylor Expansion


From Taylor’s theorem, the exponential function can be expanded as

tx t2 x2 t3 x3
etx = 1 + + + + ··· (858)
1! 2! 3!
122CHAPTER 20. LECTURES 5 AND 6: MOMENT GENERATING FUNCTIONS AND LIMIT THEOREM

20.4.1 Substitution into the MGF


By definition,
MX (t) = E(etX ) (859)
Substituting the Taylor expansion of etX , we get

t2 X 2 t3 X 3
 
tX
MX (t) = E 1 + + + + ··· (860)
1! 2! 3!

20.4.2 Using Linearity of Expectation


Since expectation is linear,

t t2 t3
MX (t) = 1 + E(X) + E(X 2 ) + E(X 3 ) + · · · (861)
1! 2! 3!
This can be written in summation form as
∞ n
X t
MX (t) = E(X n ) (862)
n!
n=0

20.4.3 Differentiation of the MGF


Differentiating term by term with respect to t,

X t n−1
d
MX (t) = E(X n ) (863)
dt (n − 1)!
n=1

20.4.4 Evaluation at t = 0
Setting t = 0, we obtain

MX (0) = E(X) (864)
More generally, the nth derivative of the MGF evaluated at t = 0 gives

(n)
MX (0) = E(X n ), n = 1, 2, 3, . . . (865)

20.4.5 Conclusion
Hence, the nth moment of a random variable X about the origin can be obtained by
differentiating its moment generating function n times and evaluating at t = 0.

20.5 Moment Generating Function of the Binomial


Distribution
Let X be a binomial random variable with parameters n and p, denoted by

X ∼ Binomial(n, p), (866)

where n is the number of independent Bernoulli trials and p is the probability of success in
each trial.
The moment generating function (MGF) of X is defined as

MX (t) = E etX .
 
(867)
20.5. MOMENT GENERATING FUNCTION OF THE BINOMIAL DISTRIBUTION 123

Using the probability mass function of the binomial distribution,


 
n x
P (X = x) = p (1 − p)n−x , x = 0, 1, 2, . . . , n, (868)
x

the MGF is obtained as


n  
X n x
tx
MX (t) = e p (1 − p)n−x
x
x=0
n  
X n (869)
= (pet )x (1 − p)n−x
x
x=0
n
= (1 − p) + pet .


The MGF of the binomial distribution is therefore


n
MX (t) = 1 − p + pet . (870)

The MGF is useful for obtaining the moments of the binomial distribution. Differentiating the
MGF and evaluating at t = 0, we obtain the mean and variance:

E[X] = MX (0) = np, (871)

′′ ′
2
Var(X) = MX (0) − MX (0) = np(1 − p). (872)
Thus, the moment generating function provides a convenient method for deriving the
moments and studying the distributional properties of the binomial random variable.

20.5.1 Using MGF to Find Moments in Detail


Let X ∼ Bin(n, p).
n  
tX
X
tk n k
MX (t) = E(e )= e p (1 − p)n−k (873)
k
k=0

Rewrite using the binomial theorem:


n  
X n
MX (t) = (pet )k (1 − p)n−k = (pet + 1 − p)n (874)
k
k=0

First moment:
Given
MX (t) = (pet + 1 − p)n (875)
Differentiate:

MX (t) = n(pet + 1 − p)n−1 (pet ) (876)
Evaluate at t = 0:

E(X) = MX (0) = n(p + 1 − p)n−1 p = np (877)
Thus:
E(X) = np (878)

We have
MX (t) = (pet + 1 − p)n (879)
Differentiate again to compute E(X 2 ):
124CHAPTER 20. LECTURES 5 AND 6: MOMENT GENERATING FUNCTIONS AND LIMIT THEOREM


MX (t) = n(pet + 1 − p)n−1 (pet ) (880)

′′
(t) = n (n − 1)(pet + 1 − p)n−2 (pet )(pet ) + (pet + 1 − p)n−1 (pet )
 
MX (881)
Factor out (pet + 1 − p)n−2 :

′′
(t) = n(pet + 1 − p)n−2 (n − 1)p2 e2t + (pet + 1 − p)(pet )
 
MX (882)
Evaluate at t = 0:
′′
(0) = n(1) n−2 (n − 1)p2 + (1)(p)
 
MX (883)
Hence:
′′
E(X 2 ) = MX (0) = n(n − 1)p2 + np (884)

We know:
Var(X) = E(X 2 ) − [E(X)]2 (885)
Substituting:

Var(X) = n(n − 1)p2 + np − (np)2


 
(886)
Simplify:

Var(X) = n(n − 1)p2 + np − n2 p2 (887)

Var(X) = np(1 − p) (888)


Thus, for X ∼ Bin(n, p):

E(X) = np, Var(X) = np(1 − p) (889)

20.6 Uniqueness Property of MGF


Theorem. If X and Y are random variables (discrete or continuous) such that

MX (t) = MY (t) for all t ∈ R, (890)

then X and Y have the same probability distribution.

20.6.1 Example: Tossing a Coin Twice


Let
P (H) = p, P (T ) = q = 1 − p. (891)
Define the random variable X as:

0, if TT,

X = 1, if HT or TH, (892)

2, if HH.

Then the PMF becomes:

P (X = 0) = q 2 , P (X = 1) = 2pq, P (X = 2) = p2 . (893)

Thus the MGF is:


MX (t) = E(etX ) = q 2 e0 + 2pqet + p2 e2t . (894)
20.7. PURPOSE OF THE MGF IN THE STANDARD NORMAL CASE 125

20.7 Purpose of the MGF in the Standard Normal Case


Let Z be a standard normal random variable, i.e.,
Z ∼ N (0, 1). (895)
The moment generating function (MGF) of Z is defined as
MZ (t) = E etZ .
 
(896)
The purpose of studying the MGF in the standard normal case is multi-fold. First, the MGF
provides a direct method to compute the moments of the standard normal distribution. By
differentiating the MGF and evaluating at t = 0, we can obtain the mean, variance, and
higher-order moments of Z without evaluating complicated integrals involving the normal
density function.
Second, the MGF of the standard normal distribution has a simple closed-form expression,
 2
t
MZ (t) = exp , (897)
2
which clearly shows that all moments of the standard normal distribution exist. From this
expression, it immediately follows that
E[Z] = 0 and Var(Z) = 1. (898)
Third, the MGF uniquely characterizes the standard normal distribution. This means that if a
random variable has the same MGF as Z in a neighborhood of t = 0, then it must follow a
standard normal distribution. This property is frequently used in theoretical proofs.
The MGF of the standard normal distribution plays a crucial role in studying sums of
independent normal random variables. Since the MGF of a sum of independent variables is
the product of their MGFs, the normality of sums and the derivation of results such as the
Central Limit Theorem become mathematically tractable.
Thus, in the standard normal case, the MGF serves as a powerful tool for generating
moments, characterizing the distribution, and analyzing sums of random variables.

20.8 MGF for Normal Distribution


20.8.1 Standard Normal Case
Let X ∼ N (0, 1). Then
Z ∞
1 2
MX (t) = E(e tX
)= etx √ e−x /2 dx. (899)
−∞ 2π
Combine exponents:
x2 1 1
= − x2 − 2tx = − x2 − 2tx + t2 − t2
 
tx − (900)
2 2 2
1 t2
= − (x − t)2 + . (901)
2 2
Thus: Z ∞
1 2 2
MX (t) = √ et /2 e−(x−t) /2 dx. (902)
2π −∞
Since the integral equals 1:
2 /2
MX (t) = et . (903)
Therefore, for X ∼ N (0, 1):
2 /2
MX (t) = et . (904)
126CHAPTER 20. LECTURES 5 AND 6: MOMENT GENERATING FUNCTIONS AND LIMIT THEOREM

20.8.2 MGF of a General Normal Distribution


Let X ∼ N (µ, σ 2 ).
We want to compute:
MX (t) = E(etX ). (905)

The pdf of X is:


 2 !
1 1 x−µ
fX (x) = √ exp − . (906)
σ 2π 2 σ

Thus,
∞  2 !
x−µ
Z
1 tx1
MX (t) = e √ exp − dx. (907)
−∞ σ 2π 2 σ

20.8.3 Substitution
Let
x−µ
Z= ⇒ x = µ + σZ, dx = σ dZ. (908)
σ
Then:

MX (t) = E(etX ) = E et(µ+σZ)



(909)

= E etµ etσZ = etµ E etσZ ,


 
(910)

where Z ∼ N (0, 1), with density


1 2
fZ (z) = √ e−z /2 . (911)

Thus,
Z ∞
1 2
E(etσZ ) = etσz √ e−z /2 dz. (912)
−∞ 2π
Rewrite exponent:

z2 1
= − z 2 − 2tσz

tσz − (913)
2 2

1 2
z − 2tσz + t2 σ 2 − t2 σ 2

=− (914)
2

1 t2 σ 2
= − (z − tσ)2 + . (915)
2 2
Thus,
Z ∞
2 σ 2 /2 1 2
E(etσZ ) = et √ e−(z−tσ) /2 dz. (916)
−∞ 2π
Since this integral equals 1,

2 σ 2 /2
E(etσZ ) = et . (917)
20.9. PROPERTIES OF THE MGF 127

20.8.4 Final MGF of Normal


t2 σ 2
 
2 σ 2 /2
MX (t) = etµ · et = exp tµ + . (918)
2
Thus, for X ∼ N (µ, σ 2 ):

t2 σ 2
 
MX (t) = exp tµ + (919)
2

20.9 Properties of the MGF


Recall:
MX (t) = E(etX ), t ∈ R. (920)

20.9.1 (i) Scaling Property


For any constant a:
   
MaX (t) = E et(aX) = E e(at)X = MX (at). (921)

20.9.2 (ii) Sum of Independent Random Variables


If X and Y are independent, then
 
MX+Y (t) = E et(X+Y ) = E etX etY .

(922)

Since X and Y are independent, etX and etY are also independent, so

E(etX etY ) = E(etX )E(etY ) = MX (t) MY (t). (923)


Thus,

MX+Y (t) = MX (t) MY (t) (924)

Additional Note: If X and Y are independent, then

E(XY ) = E(X)E(Y ), ⇒ Cov(X, Y ) = 0. (925)

20.9.3 Shift by a Constant


Let c be a constant. Then:
 
MX+c (t) = E et(X+c) = E etX etc = etc E(etX )

(926)
Thus:

MX+c (t) = etc MX (t) (927)

20.9.4 Normal Distribution Revisited


For X ∼ N (µ, σ 2 ):
 
1 2 2
MX (t) = exp tµ + σ t . (928)
2
128CHAPTER 20. LECTURES 5 AND 6: MOMENT GENERATING FUNCTIONS AND LIMIT THEOREM

20.10 Expectation via MGF



E[X] = MX (0). (929)
Differentiate:

′ d
MX (t) = MX (t) = MX (t)(σ 2 t + µ). (930)
dt
Thus,


E[X] = MX (0) = MX (0) · µ = µ. (931)

20.10.1 Second Moment


′′
E[X 2 ] = MX (0). (932)
Compute:

′′ ′
MX (t) = MX (t)(σ 2 ) + (σ 2 t + µ)MX (t). (933)
At t = 0:

′′ ′
MX (0) = MX (0)σ 2 + µ MX (0) = σ 2 + µ2 . (934)
Thus,

E[X 2 ] = σ 2 + µ2 . (935)

20.10.2 Variance
Var(X) = E[X 2 ] − (E[X])2 = (σ 2 + µ2 ) − µ2 = σ 2 . (936)

20.11 Example: Sum of Two Independent Normal Random


Variables
Given:

2
X ∼ N (µX , σX ), Y ∼ N (µY , σY2 ), (937)
and X and Y are independent random variables.
We want to determine the distribution of:

X + Y. (938)

20.11.1 Using MGF


We know that for independent variables,

MX+Y (t) = MX (t) MY (t). (939)


For a normal distribution:
 
1 2 2
MX (t) = exp µX t + σX t , (940)
2
20.12. LIMIT THEOREMS 129

 
1 2 2
MY (t) = exp µY t + σY t . (941)
2

Thus:
   
1 2 2 1 2 2
MX+Y (t) = exp µX t + σX t exp µY t + σY t . (942)
2 2

Combine exponents:
 
1 2
MX+Y (t) = exp (µX + µY )t + (σX + σY2 )t2 . (943)
2

20.11.2 Final Distribution


This is the MGF of a normal distribution with

µ = µX + µY , σ 2 = σX
2
+ σY2 . (944)

Therefore:

2
X + Y ∼ N (µX + µY , σX + σY2 ) (945)

20.12 Limit Theorems

Xn −−−→ X, X̄n −−−→ µ. (946)


n→∞ n→∞

20.12.1 Law of Large Numbers (LLN)


Let

n
1X
X̄n = Xi . (947)
n
i=1

Strong Law of Large Numbers (SLLN):

X̄n −−→ µ. (948)


a.s.

Weak Law of Large Numbers (WLLN):

X̄n −
→ µ. (949)
P

Consistency of estimator:

X̄n → µ (strong or weak). (950)


130CHAPTER 20. LECTURES 5 AND 6: MOMENT GENERATING FUNCTIONS AND LIMIT THEOREM

20.13 Population and Sample


Let X ∼ D.
From the population, we take a random sample:

X1 , X2 , . . . , Xn iid RVs. (951)


Sample mean:
n
1X
X̄n = Xi . (952)
n
i=1

20.14 Central Limit Theorem (CLT)


We study the distribution of the sample mean X̄n as n → ∞.
Recall:
σ
E(X̄n ) = µ, SD(X̄n ) = √ . (953)
n
Define the standardized variable:

X̄n − E(X̄n ) X̄n − µ


ZX̄n = = √ . (954)
SD(X̄n ) σ/ n

20.15 Purpose of the Central Limit Theorem (CLT)


The Central Limit Theorem (CLT) is one of the most fundamental results in probability and
statistics. The primary purpose of the CLT is to explain the probabilistic behavior of the sum
or average of a large number of independent random variables.
The theorem states that, under mild conditions, the standardized sum (or sample mean) of
independent and identically distributed random variables with finite mean and finite variance
converges in distribution to a normal distribution as the sample size increases.
The key purpose of the CLT is to provide a justification for using the normal distribution as
an approximation for complex and unknown distributions. Even when the original population
distribution is not normal, the CLT ensures that the distribution of the sample mean becomes
approximately normal for sufficiently large sample sizes.
Another important purpose of the CLT is to simplify statistical inference. It forms the
theoretical foundation for many classical statistical procedures, including confidence intervals,
hypothesis testing, and large-sample approximations. Because of the CLT, practitioners can
apply normal-based methods in a wide range of practical situations.
Furthermore, the CLT explains why the normal distribution appears so frequently in natural,
social, and economic phenomena. Many observed quantities are influenced by the cumulative
effect of several small, independent random factors, and the CLT provides a mathematical
explanation for this universality.
Hence, the Central Limit Theorem serves as a bridge between arbitrary probability
distributions and the normal distribution, making it an indispensable tool in probability
theory and statistical analysis.
Central Limit Theorem Formula:
d
ZX̄n −−−→ N (0, 1). (955)
n→∞

That is, the standardized sample mean converges in distribution to a standard normal random
variable.
Chapter 21

Lecture 7: Central Limit Theorem

21.1 Introduction
In probability and statistics, a central objective is to understand the behavior of random
variables and functions of random variables. In real-world applications, data are collected in
the form of samples from a population, and statistical inference is used to draw conclusions
about unknown population parameters such as the mean and variance.
One of the most important theoretical results in statistics is the Central Limit Theorem
(CLT). The CLT explains why the normal distribution arises naturally in many statistical
problems, even when the underlying population distribution is not normal. It provides the
theoretical foundation for approximation methods, hypothesis testing, and confidence interval
estimation.
In this session, we study the Central Limit Theorem and analyze the asymptotic behavior of
the sample sum and the sample mean of independent and identically distributed random
variables.

21.2 Theoretical Background


Let X be a random variable with finite mean µ = E(X) and finite variance σ 2 = Var(X). We
assume that X follows an arbitrary distribution D(µ, σ 2 ), which is not necessarily normal.

We take n independent samples:


x1 , x2 , x3 , . . . , xi , . . . , xn (956)
From the same population:
i.i.d.
Xi ∼ D(µ, σ 2 ), i = 1, 2, . . . , n (957)
The sample sum is defined as
n
X
Sn = Xi , (958)
i=1

131
132 CHAPTER 21. LECTURE 7: CENTRAL LIMIT THEOREM

and the sample mean (average) is


n
Sn 1X
Xn = = Xi . (959)
n n
i=1
By linearity of expectation,
E(X n ) = µ, (960)
and since the observations are independent,
σ2
Var(X n ) = . (961)
n

21.3 Standardisation of the Sample Mean


The standardised sample mean is given by:
X n − E(X n ) Xn − µ
ZX n = = √ (962)
SD(X n ) σ/ n
By the Central Limit Theorem, as n → ∞, the standardized sample mean converges in
distribution to a standard normal distribution:
n→∞
ZX n −−−→ Z, Z ∼ N (0, 1). (963)
CLT

21.3.1 CDF of ZX n
FZX (z) = P (ZX n ≤ z) −−−→ P (Z ≤ z) = FZ (z) (964)
n n→∞
For all z ∈ R:
⇒ P (a < ZX n < b) −−−→ P (a < Z < b) (965)
n→∞
For the standard normal distribution, the probability is calculated as the area under the curve:
Z z
1 2
P (Z ≤ z) = √ e−x /2 dx, Z ∼ N (0, 1). (966)
2π −∞
Z ∼ N (0, 1)

z
-1 0 1

21.4 Sample Sum Form and MGF Proof Setup


Without loss of generality, we can assume:
µ = 0, σ 2 = 1 ⇒ σ = 1. (967)
Since
Xi − µ
Zi = ∼ D(0, 1), (968)
σ
we can map D(µ, σ 2 ) −→ D(0, 1).
Then the standardized sample mean simplifies to:
n n
Xn √ √ 1X 1 X
ZX n = √ = n Xn = n Xi = √ Xi . (969)
σ/ n n n
i=1 i=1
CD
By CLT, ZX n −−→ Z ∼ N (0, 1). To prove this using Moment Generating Functions (MGF),
we want to show:
2
MZX (t) −−−→ et /2 . (970)
n n→∞
Chapter 22

Lecture 8: Asymptotic Bounds and


Inequalities

22.1 MGF Proof of the Central Limit Theorem


We continue the proof from Lecture 7. Let the sample mean be represented as identical

random variables divided by n:

X1 X2 Xn
ZX̄n = √ + √ + · · · + √ (971)
n n n
The MGF of ZX n is the product of the individual MGFs:

MZX (t) = MX1 /√n (t) · MX2 /√n (t) · · · MXn /√n (t). (972)
n
 
Since Xi are i.i.d., MXi /√n (t) = MX √t , yielding:
n
  n
t
MZX (t) = MX √ (973)
n n
2
We want to show that this converges to et /2 as n → ∞. Taking the natural logarithm on both
sides:
t2
  
t
lim n log MX √ = (974)
n→∞ n 2

Let L(t) = loge MX (t) . The limit can be rewritten as:
 
L √tn
lim (975)
n→∞ 1/n

Since MX (0) = 1, we have L(0) = loge (1) = 0. Thus, as n → ∞, the limit approaches 00 ,
requiring L’Hospital’s Rule.

22.1.1 Applying L’Hospital’s Rule


We evaluate the derivatives of L(t) at t = 0:
′ (t)
MX µ
L′ (t) = =⇒ L′ (0) = = 0 (since we assumed µ = 0) (976)
MX (t) 1
′′ ′
2
′′ MX (t)MX (t) − MX (t) ′′ (1)(1) − (0)2
L (t) = 2 (t) =⇒ L (0) = =1 (977)
MX 12

133
134 CHAPTER 22. LECTURE 8: ASYMPTOTIC BOUNDS AND INEQUALITIES

Now, applying L’Hospital’s rule to our limit:


   
L √tn L’H L′ √tn − 2t n−3/2

lim = lim (978)
n→∞ 1/n n→∞ −1/n2
 
t L′ √tn
= lim √ (979)
2 n→∞ 1/ n
L′ (0)
This limit evaluates to 0 = 00 . We apply L’Hospital’s rule a second time:
   
t L′′ √t
n
d
t dn √1
n t2

t

′′
L’H =⇒ lim   = lim L √ (980)
2 n→∞ d √1 2 n→∞ n
dn n
t2 ′′
= L (0) (981)
2
t2
Since L′′ (0) = 1, the final limit is exactly 2. Thus:

t2
 
t
n log MX √ −−−→ (982)
n n→∞ 2
2 /2
Which confirms the MGF converges to et , proving the Central Limit Theorem.

22.2 Asymptotic Bounds


Given a random variable X where we do not fully know its distribution, how do we bound its
probability behavior?

P(X > a)
−b a
x

22.2.1 Markov’s Inequality


Given µ = E(X) and X ≥ 0 (a non-negative random variable, either discrete or continuous).
Markov’s inequality gives an upper bound on the tail probability:

E(X) µ
P(X > a) ≤ = , ∀ a > 0. (983)
a a
(Note: this may not be the tightest possible upper bound).
Proof via Indicator Variables:
Let I be an indicator variable representing the event of interest [X ≥ a]:

1, if X > a with probability P (X > a) = p
I= (984)
0, if X ≤ a with probability [1 − P (X > a)] = q

For any X ≥ 0, we can establish the inequality X ≥ aI.


22.2. ASYMPTOTIC BOUNDS 135

ˆ Case 1: X > a ⇒ I = 1 ⇒ X ≥ a (True)

ˆ Case 2: X ≤ a ⇒ I = 0 ⇒ X ≥ 0 (True)
Taking the expectation on both sides:

E(X) ≥ E[aI] (985)


E(X) ≥ aE[I] = aP (X > a) (986)
E(X)
P (X > a) ≤ (987)
a

22.2.2 Chebyshev’s Inequality


Chebyshev’s Inequality provides a tighter upper bound by utilizing information from both the
1st and 2nd moments (mean and variance).
For any random variable X with mean µ = E(X) and variance σ 2 = Var[X] = E(X 2 ) − µ2 :
  σ2
P |X − µ| ≥ k ≤ 2 (988)
k
Proof:
The event [ |X − µ| ≥ k ] is strictly equivalent to the event [ (X − µ)2 ≥ k 2 ].

P |X − µ| ≥ k = P (X − µ)2 ≥ k 2
   
(989)

Let Y = (X − µ)2 . Note that Y ≥ 0. We can apply Markov’s inequality to Y , letting the
threshold a = k 2 > 0:
E(Y )
P Y ≥ k2 ≤
 
(990)
k2
Since E(Y ) = E[(X − µ)2 ] = σ 2 , we arrive directly at Chebyshev’s Inequality:
  σ2
P |X − µ| ≥ k ≤ 2 (991)
k
Visualizing the Bounds:

x ≤ (µ − k) x ≥ (µ + k)

x
µ−k k µ k µ+k

Absolute Value Expansions:


(
(X − µ) ≥ k if X ≥ µ
|X − µ| ≥ k =⇒ (992)
−(X − µ) ≥ k if X < µ

Which expands into the two tail events:

X ≥µ+k and X ≤ µ − k (993)

The total probability is the union of these disjoint tail events:



P |X − µ| ≥ k = P (X ≤ µ − k) + P (X ≥ µ + k) (994)
136 CHAPTER 22. LECTURE 8: ASYMPTOTIC BOUNDS AND INEQUALITIES
Chapter 23

Lecture 11: Transformations of


Random Variables

23.1 Fundamental Theorem of Calculus (FTC)


It states that if f (x) is a continuous function on a closed interval [a, b], the FTC provides the
relationship between the Cumulative Distribution Function (CDF) and the Probability
Density Function (PDF) of the function.

1. The CDF is the integral of the PDF:


Z x
F (x) = f (t) dt (995)
a

This calculates the total probability of a random variable X being less than or equal to
x by adding up all probability densities:
Z x
FX (x) = P (X ≤ x) = f (t) dt (996)
a

2. The PDF is the derivative of the CDF:

fX (x) = F ′ (x) (997)

23.1.1 Example 1: Square Root of an Exponential Variable


Let X be√a random variable of exponential distribution: X ∼ exp(λ), (λ > 0). Find the
PDF of X.

CDF of X:

F√X (x) = P ( X ≤ x) (998)
= P (X ≤ x2 ) (999)
Z x2
2
= FX (x ) = fX (t) dt (1000)
−∞


PDF of X: By FTC, we can find the PDF of X by differentiating the CDF:

d
f√X (x) = FX (x2 ) = fX (x2 ) · (2x) (1001)
dx

137
138 CHAPTER 23. LECTURE 11: TRANSFORMATIONS OF RANDOM VARIABLES

For an exponential random variable:


(
λe−λx x≥0
fX (x) = (1002)
0 otherwise

Substituting into the chain rule result:


2
fX (x2 )(2x) = 2xλe−λx , x>0 (1003)
√ 2
The PDF of X is: f√X (x) = 2λxe−λx .

23.1.2 Example 2: Transformation of a Uniform Variable


Let X be a continuous random variable with uniform distribution between closed interval
[0, 1]. Find the PDF of 2X.
X ∼ Uniform[0, 1] (1004)

CDF of 2X:

F2X (x) = P (2X ≤ x) = P (X ≤ x/2) (1005)


Z x/2
FX (x/2) = P (X ≤ x/2) = fX (t) dt (1006)
0

PDF of 2X:
d d
f2X (x) = F2X (x) = FX (x/2) (1007)
dx dx
Using the chain rule:
d 1
f2X (x) = fX (x/2) · (x/2) = fX (x/2) (1008)
dx 2
Since fX (t) = 1 for 0 < t < 1, then f2X (x) = 1/2 for 0 < x < 2.

The PDF of 2X = 1/2. (1009)

23.1.3 Example 3: Standardizing a Normal Distribution


Let X be a continuous random variable with normal distribution (µ, σ 2 ).

X ∼ N (µ, σ 2 ) (1010)
X−µ
Find the PDF of Z = σ , where µ is expectation, σ 2 is variance, and σ is standard deviation.

CDF of Z:

FZ (z) = P (Z ≤ z) (1011)
 
X −µ
=P ≤z (1012)
σ
= P [X ≤ µ + σz] (1013)
Z µ+σz
= fX (t) dt (1014)
−∞

Where fX (t) is the PDF of X ∼ N (µ, σ 2 ):


Z µ+σz
1 1 t−µ 2
FZ (z) = √ e− 2 ( σ ) dt (1015)
−∞ σ 2π
23.1. FUNDAMENTAL THEOREM OF CALCULUS (FTC) 139

PDF of Z: Let g(z) = µ + σz. By FTC and the chain rule:

d
fZ (z) = FZ (z) = fX (g(z)) · g ′ (z) (1016)
dz
d
g ′ (z) = (µ + σz) = σ (1017)
dz
Substituting g(z) into fX :
h i2
1 −1
(µ+σz)−µ
fX (g(z)) = √ e 2 σ
(1018)
σ 2π
1 1 2
= √ e− 2 z (1019)
σ 2π
Then, fZ (z) = σ · fX (g(z)):
1 1 2 1 1 2
fZ (z) = σ · √ e− 2 z = √ e− 2 z (1020)
σ 2π 2π
The result is the PDF of the standard normal distribution:

Z ∼ N (0, 1) (1021)
140 CHAPTER 23. LECTURE 11: TRANSFORMATIONS OF RANDOM VARIABLES
Chapter 24

Lecture 12: Jointly Distributed


Random Variables

24.1 1D vs 2D Probability Representation


In 1D, for a function y = f (x), the probability is represented as the area under the curve:
Z b
P [a ≤ X ≤ b] = f (x) dx (1022)
a

y
y = f (x) curve
Z t=b
= f (t)dt
a
(x, y)
y area under curve x=b
Z x=b
= f (x)dx
a

x
a bx
interval on x-axis

In 3D space, where z = f (x, y), we look at a region in the xy-plane. The input is 2D, and the
output exists in 3D space (x, y, z).

f (x, y)

S : z = f (x, y)

c
d y
a
x R
b
(x, y)

141
142 CHAPTER 24. LECTURE 12: JOINTLY DISTRIBUTED RANDOM VARIABLES

24.2 Jointly Continuous Random Variables


Definition 24.2.1. (X, Y ) is a jointly continuous RV if there exists a non-negative
function f (x, y) such that:
ZZ
P [(x, y) ∈ B] = f (x, y) dx dy (1023)
B

for some set B ⊆ R2 in the (x, y) plane.

24.2.1 Properties of Joint PDF


The function f (x, y) is a valid Probability Density Function (PDF) of (X, Y ) if it satisfies:

1. f (x, y) ≥ 0 ∀(x, y) ∈ R2
R∞ R∞
2. −∞ −∞ f (x, y) dx dy = 1

24.3 Joint CDF and its Relationship to PDF


The Joint Cumulative Distribution Function (CDF) is defined as:
Z b Z a
FX,Y (a, b) = P [X ≤ a, Y ≤ b] = f (u, v) du dv (1024)
−∞ −∞

To find the PDF from the CDF, we use partial differentiation:

∂2F
 
∂ ∂F
f (x, y) = = (1025)
∂x∂y ∂x ∂y

Y
c d
a
re
ct
an
g le

b
B = {a ≤ x < b, c ≤ y ≤ d}
X

When (u, v) is contained in B, along the x-axis, [u, v] ⊆ [a, b]:


Z x=v
P [X ∈ (u, v)] = fX (x) dx ←− marginal PDF of X (1026)
x=u
Z y=d
= f(x,y) (x, y) dy ←− (Keep x fixed as inner integral) (1027)
y=c
24.4. EVALUATING JOINT PDFS AND PROBABILITIES 143

Fact: The PDF of a continuous RV is unique. If X is a continuous RV, and it has a PDF
fX (x), this means:
Z
P [X ∈ A] = fX (x) dx for almost all sets A ⊆ R (1028)
A

For example:
Z b
P [X ∈ (a, b)] = fX (x) dx (1029)
a
(Note: If X is a discrete RV, then the PMF of X is unique.)

24.4 Evaluating Joint PDFs and Probabilities


24.4.1 Checking if a PDF is Valid
Suppose the joint PDF of (X, Y ) is defined as:
6 h 2 xy i
f (x, y) = x + for 0 < x < 1, 0<y<2 (1030)
7 2
To verify this is a valid PDF, we must check two conditions: 1. f (x, y) ≥ 0 within the bounds.
2. The double integral over the bounds equals 1.
Condition 1 (Non-negativity): For 0 < x < 1 and 0 < y < 2:
6 2 6
f (x, y) = x + xy ≥ 0 (1031)
7
|{z} |14{z }
≥0 ≥0

Thus, the first condition is satisfied.


Condition 2 (Total Volume equals 1): We evaluate the double integral:
Z y=2 Z x=1
f (x, y) dx dy (1032)
y=0 x=0

First, evaluate the inner integral with respect to x:


Z x=1 h 1
6 x3 x2 y

6 2 xy i
x + dx = + (1033)
x=0 7 2 7 3 4 0
 
6 1 y
= + (1034)
7 3 4
 
6 4 + 3y
= (1035)
7 12
4 + 3y y
= = (Simplified) (1036)
14 2
Next, evaluate the outer integral with respect to y:
Z y=2
1 2
Z
y
dy = y dy (1037)
y=0 2 2 0
 2
1 y2
= (1038)
2 2 0
 
1 4
= −0 =1 (1039)
2 2
Since the total volume is 1, the joint PDF is valid.
144 CHAPTER 24. LECTURE 12: JOINTLY DISTRIBUTED RANDOM VARIABLES

24.4.2 Finding Regional Probabilities


Find Probability P [X > Y ]
Let Exy be the event where X > Y . First, we find the region in the XY plane:
Exy = [X > Y ]. This is the area below the line x = y, bounded by 1 > x > y.
To find the probability (volume of the region), we set up the double integral:
Z x=1 Z y=x
P (X > Y ) = f (x, y) dy dx (1040)
x=0 y=0

Evaluate the inner integral with respect to y:


Z y=x  x
xy 2

6 2 xy  6 2
x + dy = x y+ (1041)
y=0 7 2 7 4 0
3 6x3 6x3
 
6 3 x
= x + = + (1042)
7 4 7 28

Evaluate the outer integral with respect to x:


Z x=1  3
6x3 6 1 x3
 Z  
6x 3
+ dx = x + dx (1043)
x=0 7 28 7 0 4
1
6 x4 x4

= + (1044)
7 4 16 0
 
6 1 1 6 6
= + = + (1045)
7 4 16 28 112
3 3 12 + 3 15
= + = = (1046)
14 56 56 56

24.4.3 Finding Conditional Probabilities


Find Probability P [Y > 1/2 | X < 1/2]
By definition of conditional probability:

P (Y > 1/2 ∩ X < 1/2)


P (A | B) = (1047)
P (X < 1/2)

1. Numerator: Z 1/2 Z 2
6  2 xy  69
x + dy dx = (1048)
0 1/2 7 2 448

2. Denominator: Z 1/2 Z 2
6  2 xy  5
x + dy dx = (1049)
0 0 7 2 28

Final Result:
69/448 69
= (1050)
5/28 80
Chapter 25

Lectures 9 and 10: Joint


Distributions and Conditional
Expectation

25.1 Probability Bounds: Chebyshev’s Inequality


In this example, we consider a random variable whose distribution is unknown, but whose first
two moments are given. Our objective is to obtain a useful probability bound using only this
limited information.
Let X be a random variable where the mean and variance are both equal to 20.
2
µX = 20, σX = 20 (1051)
We want to find a bound for the probability that X lies within the interval [0, 40], namely:
P (0 ≤ X ≤ 40) (1052)
Without additional information about the distribution of X, we can immediately write the
trivial bound which holds for any event:
0 ≤ P (0 ≤ X ≤ 40) ≤ 1 (1053)
Formally, if the probability density function fX (x) were known, then this probability could be
computed as: Z 40
P (0 ≤ X ≤ 40) = fX (x) dx (1054)
0
However, since the density function of X is not specified, an exact evaluation is impossible.
Instead, we make use of an inequality that depends only on the mean and variance of X,
namely Chebyshev’s inequality.

25.1.1 Applying Chebyshev’s Inequality


Chebyshev’s inequality provides an upper bound on the probability that a random variable
deviates from its mean by a given amount k:
2
σX
P (|X − µX | ≥ k) ≤ (1055)
k2
We begin by defining the event of interest as E = {0 ≤ X ≤ 40}. The complement of this
event consists of values of X that lie outside the interval [0, 40]. Hence:
P (E c ) = P (X < 0) + P (X > 40) (1056)

145
146CHAPTER 25. LECTURES 9 AND 10: JOINT DISTRIBUTIONS AND CONDITIONAL EXPECTATION

Since E and E c are complementary events, we have P (E) = 1 − P (E c ). Observe that the
complement event can be rewritten in terms of deviation from the mean µX = 20:

E c = {|X − 20| ≥ 20} (1057)

This reformulation allows us to apply Chebyshev’s inequality directly with k = 20:


20 20 1
P (|X − 20| ≥ 20) ≤ 2
= = (1058)
(20) 400 20
Therefore:
1 19
1 − P (E) ≤ =⇒ P (E) ≥ (1059)
20 20
Thus, even without knowing the exact distribution of X, we can conclude that:

19
P (0 ≤ X ≤ 40) ≥ (1060)
20
This bound illustrates the power of Chebyshev’s inequality when only limited information
about a random variable is available.

25.2 Jointly Discrete Random Variables


We now introduce the notion of jointly distributed random variables and define expectations
involving more than one random variable.
Let (X, Y ) be jointly discrete random variables with joint probability mass function:

pX,Y (x, y) = P (X = x, Y = y) (1061)

The joint PMF completely characterizes the probabilistic behavior of the pair (X, Y ). For any
real-valued function g(x, y), the expected value of g(X, Y ) is defined as:
XX
E[g(X, Y )] = g(x, y) pX,Y (x, y) (1062)
x y

Often, we are interested in the distribution of a single random variable obtained from the pair.
This leads to the concept of marginal distributions. The marginal PMF of X is obtained by
summing the joint PMF over all possible values of Y :
X
pX (x) = pX,Y (x, y) (1063)
y

Using the marginal PMF, expectations involving only X can be computed as:
X
E[g(X)] = g(x) pX (x) (1064)
x

25.3 Jointly Continuous Random Variables


We now consider the continuous analogue of jointly distributed random variables. If (X, Y )
are jointly continuous random variables with joint probability density function fX,Y (x, y),
then expectations are computed using double integrals:
Z ∞Z ∞
E[g(X, Y )] = g(x, y) fX,Y (x, y) dx dy (1065)
−∞ −∞

This definition extends naturally from the discrete case by replacing summations with
integrals.
25.4. LINEARITY OF EXPECTATION 147

25.4 Linearity of Expectation


One of the most fundamental and useful properties of expectation is its linearity. This
property holds regardless of whether the random variables involved are independent.

25.4.1 Discrete Case Proof


Suppose X, Y are discrete random variables with joint PMF pX,Y (x, y) = P (X = x, Y = y).
We now show that the expectation of a sum is equal to the sum of expectations.
XX
E[X + Y ] = (x + y) pX,Y (x, y)
x y
XX XX
= x pX,Y (x, y) + y pX,Y (x, y)
x y x y
X X X X
= x pX,Y (x, y) + y pX,Y (x, y) (1066)
x y y x

Recognizing the marginal probability mass functions:


X X
E[X + Y ] = x pX (x) + y pY (y)
x y

= E(X) + E(Y ) (1067)

This result does not require any independence assumption.

25.4.2 Continuous Case Proof


Suppose X, Y are continuous random variables with joint PDF fX,Y (x, y). Proceeding
analogously using integrals:
Z ∞Z ∞
E[X + Y ] = (x + y)fX,Y (x, y) dx dy
Z−∞
∞ Z −∞
∞ Z ∞Z ∞
= xfX,Y (x, y) dx dy + yfX,Y (x, y) dx dy
−∞ −∞ −∞ −∞
Z ∞ Z ∞  Z ∞ Z ∞ 
= x fX,Y (x, y) dy dx + y fX,Y (x, y) dx dy (1068)
−∞ −∞ −∞ −∞

Using marginal densities:


Z ∞ Z ∞
E[X + Y ] = xfX (x)dx + yfY (y)dy
−∞ −∞
= E(X) + E(Y ) (1069)

25.5 Covariance and Independence Example


This example illustrates the important distinction between independence and
uncorrelatedness.
Let X be defined as: 
1
 w.p. 13
X= 0 w.p. 13 (1070)

−1 w.p. 13

148CHAPTER 25. LECTURES 9 AND 10: JOINT DISTRIBUTIONS AND CONDITIONAL EXPECTATION

Define another random variable Y as:


(
1 if X = 0
Y = (1071)
0 otherwise
Note that when X = ±1, the value of Y is necessarily zero. We begin by computing the
expectation of X:      
1 1 1
E(X) = 1 +0 + (−1) =0 (1072)
3 3 3
We now check whether E(XY ) = E(X)E(Y ). Recall the following facts:
ˆ If X and Y are independent, then E(XY ) = E(X)E(Y ).
ˆ If E(XY ) = E(X)E(Y ), X and Y need not be independent (they are merely
uncorrelated, Cov(X, Y ) = 0).
Explicitly calculating XY :
(
1
0·1=0 w.p. 3 (when X = 0)
XY = 2
(1073)
(±1) · 0 = 0 w.p. 3 (when X = ±1)
Thus, XY = 0 with probability 1, and hence E(XY ) = 0. Since E(X)E(Y ) = 0 · E(Y ) = 0,
we have shown that Cov(X, Y ) = 0.
However, independence requires that P (X = x, Y = y) = P (X = x)P (Y = y) for all x, y.
Consider the event X = 1, Y = 1:
P (X = 1, Y = 1) = 0 (1074)
While the product of their marginals is:
1 1 1
P (X = 1)P (Y = 1) = × = (1075)
3 3 9
Since 0 ̸= 19 , we conclude that X and Y are dependent. This example proves that zero
covariance does not imply independence.

25.6 Variance of Linear Combinations


The variance of a linear combination of two random variables is given by:
Var(aX + bY ) = E[(aX + bY )2 ] − [E(aX + bY )]2
= E[a2 X 2 + b2 Y 2 + 2abXY ] − [aE(X) + bE(Y )]2
= a2 E(X 2 ) + b2 E(Y 2 ) + 2abE(XY ) − a2 (E(X))2 + b2 (E(Y ))2 + 2abE(X)E(Y )
 

= a2 {E(X 2 ) − [E(X)]2 } + b2 {E(Y 2 ) − [E(Y )]2 } + 2ab{E(XY ) − E(X)E(Y )}


(1076)
Hence:
Var(aX + bY ) = a2 Var(X) + b2 Var(Y ) + 2ab Cov(X, Y ) (1077)
If X and Y are independent, then Cov(X, Y ) = 0, and hence:
Var(X + Y ) = Var(X) + Var(Y ) (1078)
Var(X − Y ) = Var(X) + Var(Y ) (1079)
Finally, for sums of random variables:
 
Xn n
X n X
X n
Cov  Xi , Yj  = Cov(Xi , Yj ) (1080)
i=1 j=1 i=1 j=1
25.7. CONDITIONAL EXPECTATION 149

25.7 Conditional Expectation


Conditional expectation plays a central role in probability theory, particularly when problems
can be simplified by conditioning on an auxiliary random variable. Let X, Y be random
variables.
For discrete random variables, the conditional expectation of X given Y = y is defined as:
X
E[X | Y = y] = x P (X = x | Y = y) (1081)
x

For continuous random variables, the summation is replaced by an integral:


Z ∞
E[X | Y = y] = xfX|Y (x | y) dx (1082)
−∞

Define the function Z = g(y) = E(X | Y = y), which is a deterministic function of y.


Replacing y by the random variable Y , we obtain Z = g(Y ) = E(X | Y ), which is itself a
random variable.

25.8 Law of Total Expectation


Theorem 25.8.1 (Law of Total Expectation).

E[E(X | Y )] = E(X) (1083)

25.8.1 Proof (Discrete Case)


We begin by taking the expectation of the conditional expectation random variable
Z = E(X | Y ):
X
E{E(X | Y )} = E{Z} = g(y) P (Y = y) (1084)
y

Substituting the definition of g(y):


X
E{E(X | Y )} = E(X | Y = y) P (Y = y)
y
" #
X X
= x P (X = x | Y = y) P (Y = y)
y x
XX
= x P (X = x | Y = y)P (Y = y) (1085)
y x

Using the definition of joint probability P (X = x, Y = y) = P (X = x | Y = y)P (Y = y):


XX
E{E(X | Y )} = x P (X = x, Y = y)
y x
X X
= x P (X = x, Y = y)
x y
X
= x P (X = x) = E(X) (1086)
x

Hence proved.
150CHAPTER 25. LECTURES 9 AND 10: JOINT DISTRIBUTIONS AND CONDITIONAL EXPECTATION

25.8.2 Example 1: Tabular Data


The following example illustrates the law of total expectation using a simple tabular data set.
y Weight (X) Probability Y
1
140 5 1
1
y = 1 (Boys) 150 5 1
1
160 5 1
1
100 5 0
y = 0 (Girls) 1
120 5 0
Computing the expectation directly:
         
1 1 1 1 1
E(X) = 140 + 150 + 160 + 100 + 120 = 134 (1087)
5 5 5 5 5
Alternatively, we may compute E(X) by conditioning on gender Y :
3 2
P (Y = 1) = , P (Y = 0) = (1088)
5 5
E(X | Y = 0) = 110, E(X | Y = 1) = 150 (1089)
Applying the law of total expectation:
E(X) = E(X | Y = 0)P (Y = 0) + E(X | Y = 1)P (Y = 1)
   
2 3
= 110 + 150 = 44 + 90 = 134 (1090)
5 5

25.9 Example: Geometric Distribution Expectation


Let W ∼ Geom(p), where W denotes the number of trials required to observe the first success.
Our objective is to compute E(W ) using conditioning.
Define the auxiliary random variable Y :
(
1, if success occurs on the first trial,
Y = (1091)
0, if failure occurs on the first trial.
Thus, P (Y = 1) = p and P (Y = 0) = 1 − p = q. Using the law of total expectation:
E(W ) = E(W | Y = 0)P (Y = 0) + E(W | Y = 1)P (Y = 1) (1092)
If success occurs on the first trial, then exactly 1 trial was needed:
E(W | Y = 1) = 1 (1093)
If failure occurs on the first trial, one trial has elapsed, and the memoryless process restarts:
E(W | Y = 0) = 1 + E(W ) (1094)
Substituting these into the equation:
E(W ) = q[1 + E(W )] + p(1)
E(W ) = q + qE(W ) + p
E(W ) = 1 + qE(W ) (since p + q = 1)
E(W )(1 − q) = 1
E(W )p = 1 (1095)
Therefore:
1
E(W ) = (1096)
p
25.10. EXAMPLE: ELEVATOR STOPS 151

25.10 Example: Elevator Stops


Consider a building with a ground floor and N upper floors. Let Y be the number of people
entering the building through the ground floor, assuming Y ∼ Pois(λ). Let X be the number
of stops made by the elevator. We wish to compute E(X).
Using the law of total expectation:

X
E(X) = E(X | Y = n)P (Y = n) (1097)
n=0
−λ n
Where the PMF of Y is P (Y = n) = e n!λ .
Define indicator random variables for each floor i = 1, 2, . . . , N :
(
1, if the elevator stops at floor i,
Xi = (1098)
0, otherwise.
PN
The total number of stops is X = i=1 Xi . By linearity of expectation:
N
X
E(X | Y = n) = E(Xi | Y = n) (1099)
i=1

Since all floors are statistically identical, the expectation is the same for any floor:

E(X | Y = n) = N · E(X1 | Y = n) = N · P (X1 = 1 | Y = n) (1100)

Using the complement, P (X1 = 1 | Y = n) = 1 − P (X1 = 0 | Y = n).


The probability that the elevator does not stop at the first floor given n people entered means
all n people independently chose one of the remaining N − 1 floors:
N −1 n
 
P (X1 = 0 | Y = n) = (1101)
N
Substituting this back into the total expectation equation:

N − 1 n e−λ λn
X    
E(X) = N 1−
N n!
n=0
"∞ ∞ 
#
X e−λ λn X N − 1 n e−λ λn

=N − (1102)
n! N n!
n=0 n=0

The first sum is the total probability of a Poisson distribution, which equals 1. For the second
sum, we factor out e−λ and recognize the Taylor series for an exponential function:
∞  1 n

X λ(1 − )
E(X) = N − N e−λ N
n!
n=0
1
−λ
= N − Ne · eλ(1− N )
= N − N e−λ+λ−λ/N
= N − N e−λ/N (1103)

Therefore, the expected number of elevator stops is:


 
E(X) = N 1 − e−λ/N (1104)
152CHAPTER 25. LECTURES 9 AND 10: JOINT DISTRIBUTIONS AND CONDITIONAL EXPECTATION
Chapter 26

Lecture 11: Transformations of


Random Variables

26.1 Fundamental Theorem of Calculus (FTC)


It states that if f (x) is a continuous function on a closed interval [a, b], the FTC provides the
relationship between the Cumulative Distribution Function (CDF) and the Probability
Density Function (PDF) of the function.

1. The CDF is the integral of the PDF:


Z x
F (x) = f (t) dt (1105)
a

This calculates the total probability of a random variable X being less than or equal to
x by adding up all probability densities:
Z x
FX (x) = P (X ≤ x) = f (t) dt (1106)
a

2. The PDF is the derivative of the CDF:

fX (x) = F ′ (x) (1107)

26.1.1 Example 1: Square Root of an Exponential Variable


Let X be√a random variable of exponential distribution: X ∼ exp(λ), (λ > 0). Find the
PDF of X.

CDF of X:

F√X (x) = P ( X ≤ x) (1108)
= P (X ≤ x2 ) (1109)
Z x2
2
= FX (x ) = fX (t) dt (1110)
−∞


PDF of X: By FTC, we can find the PDF of X by differentiating the CDF:

d
f√X (x) = FX (x2 ) = fX (x2 ) · (2x) (1111)
dx

153
154 CHAPTER 26. LECTURE 11: TRANSFORMATIONS OF RANDOM VARIABLES

For an exponential random variable:


(
λe−λx x≥0
fX (x) = (1112)
0 otherwise

Substituting into the chain rule result:


2
fX (x2 )(2x) = 2xλe−λx , x>0 (1113)
√ 2
The PDF of X is: f√X (x) = 2λxe−λx .

26.1.2 Example 2: Transformation of a Uniform Variable


Let X be a continuous random variable with uniform distribution between closed interval
[0, 1]. Find the PDF of 2X.
X ∼ Uniform[0, 1] (1114)

CDF of 2X:

F2X (x) = P (2X ≤ x) = P (X ≤ x/2) (1115)


Z x/2
FX (x/2) = P (X ≤ x/2) = fX (t) dt (1116)
0

PDF of 2X:
d d
f2X (x) = F2X (x) = FX (x/2) (1117)
dx dx
Using the chain rule:
d 1
f2X (x) = fX (x/2) · (x/2) = fX (x/2) (1118)
dx 2
Since fX (t) = 1 for 0 < t < 1, then f2X (x) = 1/2 for 0 < x < 2.

The PDF of 2X = 1/2. (1119)

26.1.3 Example 3: Standardizing a Normal Distribution


Let X be a continuous random variable with normal distribution (µ, σ 2 ).

X ∼ N (µ, σ 2 ) (1120)
X−µ
Find the PDF of Z = σ , where µ is expectation, σ 2 is variance, and σ is standard deviation.

CDF of Z:

FZ (z) = P (Z ≤ z) (1121)
 
X −µ
=P ≤z (1122)
σ
= P [X ≤ µ + σz] (1123)
Z µ+σz
= fX (t) dt (1124)
−∞

Where fX (t) is the PDF of X ∼ N (µ, σ 2 ):


Z µ+σz
1 1 t−µ 2
FZ (z) = √ e− 2 ( σ ) dt (1125)
−∞ σ 2π
26.1. FUNDAMENTAL THEOREM OF CALCULUS (FTC) 155

PDF of Z: Let g(z) = µ + σz. By FTC and the chain rule:

d
fZ (z) = FZ (z) = fX (g(z)) · g ′ (z) (1126)
dz
d
g ′ (z) = (µ + σz) = σ (1127)
dz
Substituting g(z) into fX :
h i2
1 −1
(µ+σz)−µ
fX (g(z)) = √ e 2 σ
(1128)
σ 2π
1 1 2
= √ e− 2 z (1129)
σ 2π
Then, fZ (z) = σ · fX (g(z)):
1 1 2 1 1 2
fZ (z) = σ · √ e− 2 z = √ e− 2 z (1130)
σ 2π 2π
The result is the PDF of the standard normal distribution:

Z ∼ N (0, 1) (1131)
156 CHAPTER 26. LECTURE 11: TRANSFORMATIONS OF RANDOM VARIABLES
Chapter 27

Lecture 12: Jointly Distributed


Random Variables

27.1 1D vs 2D Probability Representation


In 1D, for a function y = f (x), the probability is represented as the area under the curve:
Z b
P [a ≤ X ≤ b] = f (x) dx (1132)
a

y
y = f (x) curve
Z t=b
= f (t)dt
a
(x, y)
y area under curve x=b
Z x=b
= f (x)dx
a

x
a bx
interval on x-axis

In 3D space, where z = f (x, y), we look at a region in the xy-plane. The input is 2D, and the
output exists in 3D space (x, y, z).

f (x, y)

S : z = f (x, y)

c
d y
a
x R
b
(x, y)

157
158 CHAPTER 27. LECTURE 12: JOINTLY DISTRIBUTED RANDOM VARIABLES

27.2 Jointly Continuous Random Variables


Definition 27.2.1. (X, Y ) is a jointly continuous RV if there exists a non-negative
function f (x, y) such that:
ZZ
P [(x, y) ∈ B] = f (x, y) dx dy (1133)
B

for some set B ⊆ R2 in the (x, y) plane.

27.2.1 Properties of Joint PDF


The function f (x, y) is a valid Probability Density Function (PDF) of (X, Y ) if it satisfies:

1. f (x, y) ≥ 0 ∀(x, y) ∈ R2
R∞ R∞
2. −∞ −∞ f (x, y) dx dy = 1

27.3 Joint CDF and its Relationship to PDF


The Joint Cumulative Distribution Function (CDF) is defined as:
Z b Z a
FX,Y (a, b) = P [X ≤ a, Y ≤ b] = f (u, v) du dv (1134)
−∞ −∞

To find the PDF from the CDF, we use partial differentiation:

∂2F
 
∂ ∂F
f (x, y) = = (1135)
∂x∂y ∂x ∂y

Y
c d
a
re
ct
an
g le

b
B = {a ≤ x < b, c ≤ y ≤ d}
X

When (u, v) is contained in B, along the x-axis, [u, v] ⊆ [a, b]:


Z x=v
P [X ∈ (u, v)] = fX (x) dx ←− marginal PDF of X (1136)
x=u
Z y=d
= f(x,y) (x, y) dy ←− (Keep x fixed as inner integral) (1137)
y=c
27.4. EVALUATING JOINT PDFS AND PROBABILITIES 159

Fact: The PDF of a continuous RV is unique. If X is a continuous RV, and it has a PDF
fX (x), this means:
Z
P [X ∈ A] = fX (x) dx for almost all sets A ⊆ R (1138)
A

For example:
Z b
P [X ∈ (a, b)] = fX (x) dx (1139)
a
(Note: If X is a discrete RV, then the PMF of X is unique.)

27.4 Evaluating Joint PDFs and Probabilities


27.4.1 Checking if a PDF is Valid
Suppose the joint PDF of (X, Y ) is defined as:
6 h 2 xy i
f (x, y) = x + for 0 < x < 1, 0<y<2 (1140)
7 2
To verify this is a valid PDF, we must check two conditions: 1. f (x, y) ≥ 0 within the bounds.
2. The double integral over the bounds equals 1.
Condition 1 (Non-negativity): For 0 < x < 1 and 0 < y < 2:
6 2 6
f (x, y) = x + xy ≥ 0 (1141)
7
|{z} |14{z }
≥0 ≥0

Thus, the first condition is satisfied.


Condition 2 (Total Volume equals 1): We evaluate the double integral:
Z y=2 Z x=1
f (x, y) dx dy (1142)
y=0 x=0

First, evaluate the inner integral with respect to x:


Z x=1 h 1
6 x3 x2 y

6 2 xy i
x + dx = + (1143)
x=0 7 2 7 3 4 0
 
6 1 y
= + (1144)
7 3 4
 
6 4 + 3y
= (1145)
7 12
4 + 3y y
= = (Simplified) (1146)
14 2
Next, evaluate the outer integral with respect to y:
Z y=2
1 2
Z
y
dy = y dy (1147)
y=0 2 2 0
 2
1 y2
= (1148)
2 2 0
 
1 4
= −0 =1 (1149)
2 2
Since the total volume is 1, the joint PDF is valid.
160 CHAPTER 27. LECTURE 12: JOINTLY DISTRIBUTED RANDOM VARIABLES

27.4.2 Finding Regional Probabilities


Find Probability P [X > Y ]
Let Exy be the event where X > Y . First, we find the region in the XY plane:
Exy = [X > Y ]. This is the area below the line x = y, bounded by 1 > x > y.
To find the probability (volume of the region), we set up the double integral:
Z x=1 Z y=x
P (X > Y ) = f (x, y) dy dx (1150)
x=0 y=0

Evaluate the inner integral with respect to y:


Z y=x  x
xy 2

6 2 xy  6 2
x + dy = x y+ (1151)
y=0 7 2 7 4 0
3 6x3 6x3
 
6 3 x
= x + = + (1152)
7 4 7 28

Evaluate the outer integral with respect to x:


Z x=1  3
6x3 6 1 x3
 Z  
6x 3
+ dx = x + dx (1153)
x=0 7 28 7 0 4
1
6 x4 x4

= + (1154)
7 4 16 0
 
6 1 1 6 6
= + = + (1155)
7 4 16 28 112
3 3 12 + 3 15
= + = = (1156)
14 56 56 56

27.4.3 Finding Conditional Probabilities


Find Probability P [Y > 1/2 | X < 1/2]
By definition of conditional probability:

P (Y > 1/2 ∩ X < 1/2)


P (A | B) = (1157)
P (X < 1/2)

1. Numerator: Z 1/2 Z 2
6  2 xy  69
x + dy dx = (1158)
0 1/2 7 2 448

2. Denominator: Z 1/2 Z 2
6  2 xy  5
x + dy dx = (1159)
0 0 7 2 28

Final Result:
69/448 69
= (1160)
5/28 80
Chapter 28

Lecture 13: Conditional Probability


(Continuous Case)

28.1 Joint PDF Analysis


Consider the following joint PDF:
6  2 xy 
fX,Y (x, y) = x + , 0 < x < 1, 0 < y < 2 (1161)
7 2
Outside the given region, the PDF is zero. The support of this PDF is a rectangle in the
xy-plane.

1.5
y

1 Support Region

0.5

0
0 0.2 0.4 0.6 0.8 1 1.2
x

28.2 Calculating Conditional Probabilities


1 1

We wish to find the probability P Y > 2 |X< 2 . Using the standard conditional
probability formula:
P (E ∩ F )
P (E|F ) = (1162)
P (F )
Where E = {Y > 12 } and F = {X < 12 }, we have:
P Y > 12 , X < 12
  
1 1
P Y > |X< = (1163)
2 2 P (X < 21 )

28.2.1 Graphical Representation of Events


The event X < 12 represents a vertical strip on the left side of the support, while Y > 21
represents a horizontal strip at the top. The intersection is the region where both conditions
are met.

161
162CHAPTER 28. LECTURE 13: CONDITIONAL PROBABILITY (CONTINUOUS CASE)

Intersection Region: Y > 1/2 | X < 1/2

1.5

y 1

0.5

0
0 0.2 0.4 0.6 0.8 1 1.2
x

28.3 Marginal Distributions and Expectation


28.3.1 Marginal PDF of X
To find fX (x), we integrate the joint PDF over all possible values of Y :
Z 2
fX (x) = fX,Y (x, y) dy (1164)
0
6 2  2 xy 
Z
= x + dy (1165)
7 0 2
6
= (2x2 + x) (1166)
7

2
fX (x)

0
0 0.2 0.4 0.6 0.8 1
x

28.3.2 Expectation of X
The expected value is computed by integrating the product of x and its marginal density:
Z 1
E(X) = xfX (x) dx (1167)
0
6 1
Z
= x(2x2 + x) dx (1168)
7 0
 
6 1 1 5
= + = (1169)
7 2 3 7
Chapter 29

Lecture 14: Independence and Sums


of Variables

29.1 Independence of Random Variables


Two random variables are independent if their joint distribution is the product of their
marginal distributions.
Continuous Case:
FX,Y (a, b) = FX (a)FY (b) (1170)

fX,Y (x, y) = fX (x)fY (y) (1171)

Discrete Case:
PX,Y (a, b) = PX (a)PY (b) (1172)

29.2 Example: Sum of Independent Poisson RVs


Let X ∼ Poisson(λ) and Y ∼ Poisson(β) be independent random variables. Define
Z = X + Y . We wish to find the PMF of Z.

29.2.1 Derivation via Convolution (PMF Method)


Using the Law of Total Probability and independence:

P (Z = r) = P (X + Y = r) (1173)
Xr
= P (X + Y = r | Y = y)P (Y = y) (1174)
y=0
r
X
= P (X = r − y) P (Y = y) (1175)
y=0

Substituting the Poisson PMFs:


r  r−y
 y

−λ λ −β β
X
P (Z = r) = e e (1176)
(r − y)! y!
y=0
r
X λr−y β y
= e−(λ+β) (1177)
(r − y)!y!
y=0

163
164 CHAPTER 29. LECTURE 14: INDEPENDENCE AND SUMS OF VARIABLES

Multiplying and dividing by r!:


r
−(λ+β) 1 X r!
P (Z = r) = e λr−y β y (1178)
r! (r − y)!y!
y=0

Recognizing the binomial coefficient and the binomial theorem:

(λ + β)r
P (Z = r) = e−(λ+β) (1179)
r!

Thus, Z ∼ Poisson(λ + β) .

29.2.2 Derivation via Moment Generating Functions (MGF)


t −1)
The MGF of a Poisson random variable is MX (t) = eλ(e . For independent variables, the
MGF of a sum is the product of individual MGFs:

MZ (t) = MX (t)MY (t) (1180)


λ(et −1) β(et −1)
=e e (1181)
(λ+β)(et −1)
=e (1182)

This result confirms that Z ∼ Poisson(λ + β).

29.3 Sum of Two Uniform Random Variables


Let X ∼ Uniform(0, 1) and Y ∼ Uniform(0, 1) be independent. Define A = X + Y .

29.3.1 Problem Setup


The joint PDF is fX,Y (x, y) = 1 inside the unit square. The CDF is FA (a) = P (X + Y ≤ a),
which corresponds to the area of the square below the line x + y = a.

29.3.2 Case 1: a < 0


Since X, Y ≥ 0, the sum cannot be negative.

FA (a) = 0 (1183)

29.3.3 Case 2: 0 ≤ a ≤ 1
The region is a triangle with vertices (0, 0), (a, 0), (0, a).
a Z a−x a
a2
Z Z
FA (a) = dy dx = (a − x) dx = (1184)
0 0 0 2

29.3.4 Case 3: 1 < a ≤ 2


It is easier to compute the complement area (the small triangle in the upper-right corner) and
subtract it from 1. The legs of the excluded triangle have length (2 − a).

(2 − a)2
FA (a) = 1 − (1185)
2
29.3. SUM OF TWO UNIFORM RANDOM VARIABLES 165

29.3.5 Final CDF and PDF


The complete CDF of A is:


0, a < 0,
2

a ,


0 ≤ a ≤ 1,

FA (a) = 2 (1186)
(2 − a)2
1 − , 1 < a ≤ 2,





 2
1, a > 2.

Differentiating the CDF piecewise gives the triangular PDF:



a,
 0 ≤ a ≤ 1,
fA (a) = 2 − a, 1 < a ≤ 2, (1187)

0, otherwise.

Triangular Distribution PDF


1.2

0.8
fA (a)

0.6

0.4

0.2

0
0 0.5 1 1.5 2
a
166 CHAPTER 29. LECTURE 14: INDEPENDENCE AND SUMS OF VARIABLES
Chapter 30

Lectures 15 and 16: Maximum


Likelihood Estimation

30.1 Introduction to Estimation


Estimation is a fundamental problem in statistics and signal processing. The goal is to infer
unknown parameters of a probability distribution based on observed data.

30.1.1 Estimation as an Inverse Problem


Let X be a random variable with probability density function (PDF)

fX (x; θ),

where θ is an unknown parameter. The forward process generates samples

X1 , X2 , . . . , XN ∼ fX (x; θ).

The estimation problem seeks to recover θ from the observed samples.

30.2 Parameters of Distributions


30.2.1 Bernoulli Distribution
If Xn ∼ Bernoulli(θ), the probability mass function (PMF) is

pXn (xn ; θ) = θxn (1 − θ)1−xn , xn ∈ {0, 1}.

30.2.2 Gaussian Distribution


If Xn ∼ N (µ, σ 2 ), then

(xn − µ)2
 
2 1
fXn (xn ; µ, σ ) = √ exp − .
2πσ 2 2σ 2

The parameter vector may be θ = (µ, σ) or simply θ = µ if σ is known.

167
168 CHAPTER 30. LECTURES 15 AND 16: MAXIMUM LIKELIHOOD ESTIMATION

30.3 Likelihood Function


30.3.1 Definition
Let X = [X1 , . . . , XN ]T be i.i.d. random variables with realizations x = [x1 , . . . , xN ]T . The
likelihood function is defined as
L(θ | x) = fX (x; θ).
For i.i.d. samples,
N
Y
L(θ | x) = fXn (xn ; θ).
n=1

30.3.2 Log-Likelihood
The log-likelihood simplifies optimization:
N
X
log L(θ | x) = log fXn (xn ; θ).
n=1

30.4 Examples of Log-Likelihood


30.4.1 Gaussian Case
For Xn ∼ N (µ, σ 2 ),
N
N 1 X
log L(µ, σ 2 | x) = − log(2πσ 2 ) − 2 (xn − µ)2 .
2 2σ
n=1

30.4.2 Bernoulli Case


For Bernoulli observations,
N
X
log L(θ | x) = xn log θ + (1 − xn ) log(1 − θ).
n=1
PN
Let S = n=1 xn , then
log L(θ | S) = S log θ + (N − S) log(1 − θ).

30.5 Maximum Likelihood Estimation


30.5.1 Definition
The maximum likelihood estimate (MLE) is

θ̂ML = arg max L(θ | x).


θ

30.5.2 MLE for Bernoulli Distribution


Differentiating the log-likelihood and setting it to zero yields
N
1 X
θ̂ML = xn .
N
n=1

Thus, the MLE is the sample mean.


30.6. APPLICATION: SOCIAL NETWORK ESTIMATION 169

30.5.3 MLE for Gaussian Mean


If σ 2 is known,
N
1 X
µ̂ML = xn .
N
n=1

As N increases, the estimate becomes more accurate due to the law of large numbers.

30.6 Application: Social Network Estimation


Consider an Erdős–Rényi graph where each edge is present independently with probability p.
Each adjacency matrix entry satisfies

Xij ∼ Bernoulli(p).

The log-likelihood is

N X
X N
log L(p | X) = [xij log p + (1 − xij ) log(1 − p)] .
i=1 j=1

P
Let S = i,j xij . The MLE is
S
p̂ML = .
N2

30.7 Application: Single-Photon Imaging


Each pixel follows a Poisson model:

Xn ∼ Poisson(λ).

A one-bit sensor records (


1, Xn ≥ 1
Yn =
0, Xn = 0

Then
P (Yn = 1) = 1 − e−λ , P (Yn = 0) = e−λ .

30.7.1 Log-Likelihood
N h
X i
log L(λ | y) = yn log(1 − e−λ ) − λ(1 − yn ) .
n=1
P
Let S = yn . The MLE is
 
S
λ̂ML = − log 1 − .
N

30.8 Estimation as an Inverse Problem


Estimation is the process of recovering unknown parameters of a probability distribution from
observed data.
170 CHAPTER 30. LECTURES 15 AND 16: MAXIMUM LIKELIHOOD ESTIMATION

Conceptual Diagram

θ fX (x; θ) X1 , . . . , XN

Estimator

Forward problem: θ → Xn
Inverse problem (estimation): Xn → θ

30.9 Likelihood and Log-Likelihood


Given i.i.d. samples X1 , . . . , XN with PDF f (x; θ), the likelihood is

N
Y
L(θ | x) = f (xn ; θ).
n=1

Taking logarithms,
N
X
log L(θ | x) = log f (xn ; θ).
n=1

Log-likelihood simplifies differentiation and numerical stability.

30.10 Sample Sum 1: Bernoulli MLE (Fully Worked)


Let Xn ∼ Bernoulli(θ).

p(xn ; θ) = θxn (1 − θ)1−xn

Step 1: Likelihood
N
Y
L(θ) = θxn (1 − θ)1−xn
n=1

Step 2: Log-Likelihood
N
X
log L(θ) = xn log θ + (1 − xn ) log(1 − θ)
n=1

Define
N
X
S= xn
n=1

Then
log L(θ) = S log θ + (N − S) log(1 − θ)
30.11. BERNOULLI LOG-LIKELIHOOD SHAPE (TIKZ PLOT) 171

Step 3: Optimization
d S N −S
log L(θ) = −
dθ θ 1−θ
Set derivative to zero:
S
θ̂ML =
N

θ̂ML = sample mean

30.11 Bernoulli Log-Likelihood Shape (TikZ Plot)

−40

−60
log L(θ)

−80

−100

−120
0 0.2 0.4 0.6 0.8 1
θ

Example shown for N = 50, S = 25. Peak occurs at θ = 0.5.

30.12 Sample Sum 2: Gaussian Mean Estimation


Assume
Xn ∼ N (µ, σ 2 ), σ 2 known

Log-Likelihood
N
1 X
log L(µ) = − 2 (xn − µ)2 + C

n=1

Derivative
N
d 1 X
log L(µ) = 2 (xn − µ)
dµ σ
n=1

Set to zero:

N N
X 1 X
(xn − µ) = 0 ⇒ µ̂ML = xn
N
n=1 n=1

µ̂ML = sample average


172 CHAPTER 30. LECTURES 15 AND 16: MAXIMUM LIKELIHOOD ESTIMATION

30.13 Likelihood Matching Intuition (Diagram)


PDF

Candidate PDF
Better Match

MLE selects the parameter whose PDF best aligns with the observed data.

30.14 Conclusion
Maximum Likelihood Estimation provides a principled and general framework for parameter
estimation. Many classical estimators such as the sample mean naturally arise as MLEs under
common distributional assumptions.
Chapter 31

Lectures 17 and 18: Limit Theorems

31.1 Introduction and Motivation


Probability theory is fundamentally concerned with understanding the behavior of random
phenomena. Individual observations are uncertain, unpredictable, and subject to randomness.
However, when many such observations are combined, remarkable regularities emerge.
Two of the most important results that explain this phenomenon are:
ˆ The Law of Large Numbers (LLN)
ˆ The Central Limit Theorem (CLT)
The LLN explains stability of averages, while the CLT explains the shape of fluctuations
around those averages. These results justify why averages are meaningful in statistics,
economics, finance, physics, and social sciences.

31.2 Probability Setup and Assumptions


Let
X1 , X2 , . . . , Xn (1188)
be a sequence of independent and identically distributed (i.i.d.) random variables
defined on a common probability space (Ω, F, P ).
We assume:
E(Xi ) = µ and Var(Xi ) = σ 2 < ∞ (1189)

Why These Assumptions Matter


ˆ Independence ensures that observations do not influence each other. Without
independence, averaging may not reduce variability.
ˆ Identical distribution ensures homogeneity: each observation comes from the same
population.
ˆ Finite mean ensures that the population average is well-defined.
ˆ Finite variance ensures that fluctuations are not too extreme.
31.3 Definition of the Sample Mean
The sample mean is defined as:
n
1X
Mn = Xi (1190)
n
i=1

173
174 CHAPTER 31. LECTURES 17 AND 18: LIMIT THEOREMS

Conceptual Meaning
ˆ M aggregates information from n random observations.
n

ˆ It is itself a random variable, not a constant.


ˆ Different samples of size n produce different values of M . n

ˆ As n increases, M uses more information from the population.


n

The central question of this chapter is: What happens to Mn as the sample size n becomes
large?

31.4 Expectation of the Sample Mean


We compute the expectation of Mn step by step:
n
!
1X
E(Mn ) = E Xi (1191)
n
i=1

Using linearity of expectation:


n
1X
E(Mn ) = E(Xi ) (1192)
n
i=1

Since E(Xi ) = µ for all i:


1
E(Mn ) = (nµ) = µ (1193)
n

Interpretation
ˆ The expected value of the sample mean equals the population mean.
ˆ The sample mean is an unbiased estimator.
ˆ On average, sampling neither inflates nor deflates the true mean.
31.5 Variance of the Sample Mean
Variance measures the spread or uncertainty of a random variable.
n
!
1X
Var(Mn ) = Var Xi (1194)
n
i=1

Using properties of variance and independence:


n
1 X
Var(Mn ) = 2 Var(Xi ) (1195)
n
i=1

Substituting Var(Xi ) = σ 2 :
1 2 σ2
Var(Mn ) = (nσ ) = (1196)
n2 n
31.6. WEAK LAW OF LARGE NUMBERS (WLLN) 175

Critical Insight
Var(Mn ) −−−→ 0 (1197)
n→∞

ˆ Averaging reduces randomness.


ˆ Larger samples lead to more precise estimates.
ˆ This shrinking variance drives the Law of Large Numbers.
31.6 Weak Law of Large Numbers (WLLN)
Formal Statement
P
Mn −
→µ as n → ∞ (1198)

Definition of Convergence in Probability


∀ϵ > 0, P (|Mn − µ| > ϵ) → 0 (1199)

Interpretation
ˆ Deviations from µ larger than ϵ become increasingly rare.
ˆ Convergence is probabilistic, not deterministic.
ˆ There is no claim that M (ω) converges for every ω.
n

31.7 Proof of WLLN Using Chebyshev’s Inequality


Chebyshev’s inequality states:

Var(Y )
P (|Y − E(Y )| ≥ ϵ) ≤ (1200)
ϵ2
Apply it to Y = Mn :
Var(Mn )
P (|Mn − µ| ≥ ϵ) ≤ (1201)
ϵ2
Substitute Var(Mn ) = σ 2 /n:
σ2
P (|Mn − µ| ≥ ϵ) ≤ (1202)
nϵ2
Taking limits:
lim P (|Mn − µ| ≥ ϵ) = 0 (1203)
n→∞

Thus:
P
Mn −
→µ (1204)

31.8 Event-Based Interpretation of WLLN


Define the events En = {|Mn − µ| ≤ ϵ}. Then:

P (En ) → 1 and P (Enc ) → 0 (1205)


176 CHAPTER 31. LECTURES 17 AND 18: LIMIT THEOREMS

Meaning
ˆ The probability mass concentrates around µ.
ˆ Large deviations vanish asymptotically.
31.9 Strong Law of Large Numbers (SLLN)
a.s.
Mn −−→ µ (1206)

Interpretation
ˆ Convergence occurs for almost every outcome.
ˆ Only a probability-zero set fails to converge.
ˆ SLLN is stronger than WLLN: SLLN ⇒ WLLN.
31.10 Bernoulli Example
Let Xi be a Bernoulli trials where success is 1 and failure is 0:
(
1 with probability p
Xi = (1207)
0 with probability 1 − p

Then:
E(Xi ) = p, Var(Xi ) = p(1 − p) (1208)
Sample mean:
n
1X
Mn = Xi (1209)
n
i=1

Interpretation
ˆ M equals the proportion of successes.
n

ˆ Randomness diminishes as sample size increases.


31.11 WLLN via Moment Generating Functions
Assume the MGF exists near zero:

MX (t) = E(etX ) (1210)

Using Taylor expansion:


MX (t) = 1 + µt + o(t) (1211)
MGF of Mn :  n

t
MMn (t) = MX → eµt (1212)
n
Thus, we confirm:
P
Mn −
→µ (1213)
31.12. CENTRAL LIMIT THEOREM (CLT) 177

31.12 Central Limit Theorem (CLT)


Define the standardized sample mean:

n(Mn − µ)
Zn = (1214)
σ
Then:
d
Zn −
→ N (0, 1) (1215)

Meaning
ˆ LLN explains convergence of averages.
ˆ CLT explains distribution of fluctuations.
ˆ Normality arises universally, regardless of the original distribution.
31.13 Final Summary
ˆ Averaging stabilizes randomness.
ˆ Variance shrinks at rate 1/n.
ˆ WLLN ensures probabilistic convergence.
ˆ SLLN ensures pathwise convergence.
ˆ CLT explains Gaussian fluctuations.
178 CHAPTER 31. LECTURES 17 AND 18: LIMIT THEOREMS
Chapter 32

Lectures 19 and 20: Sampling


Distributions (χ2, t, and F )

32.1 Sampling Setup


32.1.1 Population and Samples
Let the population be:
X ∼ N (µ, σ 2 ). (1216)

Take a random sample of size n:

X1 , X2 , . . . , Xn are i.i.d. r.v.s with Xi ∼ N (µ, σ 2 ). (1217)

32.1.2 After collecting a sample


After observing a sample, we get the realized values:

x1 , x2 , . . . , xn . (1218)

32.1.3 Sample Mean


n n
1X 1X
X= Xi . x= xi . (1219)
n n
i=1 i=1

32.1.4 Sample Variance


n n
1 X 1 X
S2 = (Xi − X)2 . s2 = (xi − x)2 . (1220)
n−1 n−1
i=1 i=1

32.1.5 Useful identity (DOF explanation)


n
X
(Xi − X) = 0 (1221)
i=1

which implies only n − 1 deviations are free (hence n − 1 degrees of freedom).

179
180 CHAPTER 32. LECTURES 19 AND 20: SAMPLING DISTRIBUTIONS (χ2 , t, AND F )

32.2 Distribution from Normal


32.2.1 Standard Normal
Define:
X −µ
Z= ∼ N (0, 1). (1222)
σ

32.2.2 Sampling distribution of X


σ2
 
X ∼ N µ, . (1223)
n
Hence
X −µ
√ ∼ N (0, 1). (1224)
σ/ n

32.3 Chi-square Distribution


32.3.1 Definition
Let
iid
Z1 , Z2 , . . . , Zn ∼ N (0, 1). (1225)
Define
n
X
X= Zi2 . (1226)
i=1

Then
X ∼ χ2n . (1227)

32.3.2 Chi-square from sample variance


Define
(n − 1)S 2
V = . (1228)
σ2
Then:
V ∼ χ2(n−1) . (1229)

32.3.3 Right tail critical value


Let χ2α,ν denote the upper-α critical value, i.e.

P χ2ν > χ2α,ν = α.



(1230)

32.3.4 Sketch (right tail)


pdf

area = α
x
χ2α,ν
32.4. STUDENT t DISTRIBUTION 181

32.3.5 Additive property


If
X1 ∼ χ2n1 , X2 ∼ χ2n2 , X1 , X2 independent, (1231)
then:
X1 + X2 ∼ χ2n1 +n2 . (1232)

32.3.6 Expansion idea


Let
X1 = Z12 + Z22 + · · · + Zn21 , X2 = Zn21 +1 + · · · + Zn21 +n2 . (1233)
Then
X1 + X2 = Z12 + · · · + Zn21 +n2 ∼ χ2n1 +n2 . (1234)

32.3.7 Example: distance in 3D


Let
D2 = x21 + x22 + x23 , (1235)
where xi is error in i-th coordinate. If we standardize
xi
Zi = ⇒ xi = 2Zi , (1236)
2
then
D2 = 4Z12 + 4Z22 + 4Z32 = 4(Z12 + Z22 + Z32 ) = 4χ23 . (1237)
Hence  
2 9
P(D > 9) = P(4χ23 2
> 9) = P χ3 > . (1238)
4

32.4 Student t Distribution


32.4.1 Definition
Let
Z ∼ N (0, 1), X ∼ χ2n (1239)
be independent. Then:
Z
Tn = p ∼ tn . (1240)
X/n

32.4.2 From sample mean when σ unknown


Since
X −µ (n − 1)S 2
√ ∼ N (0, 1), ∼ χ2n−1 , (1241)
σ/ n σ2
independent, we get:
X −µ
Tn−1 = √ ∼ tn−1 . (1242)
S/ n

32.4.3 Note: t density vs Normal


The t density has thicker tails than the normal density, indicating greater variability.
182 CHAPTER 32. LECTURES 19 AND 20: SAMPLING DISTRIBUTIONS (χ2 , t, AND F )

32.4.4 Symmetry property


t distribution is symmetric about 0:
d
Tn = −Tn . (1243)

32.4.5 Critical values and symmetry relation


If P(Tn > tα,n ) = α, then:

α = P(Tn > tα,n ) = P(−Tn < −tα,n ) = P(Tn < −tα,n ). (1244)

Hence:
P(Tn > −tα,n ) = 1 − α. (1245)
So we identify the relationship:
−tα,n = t1−α,n . (1246)

32.4.6 Sketch
[Image comparing a standard normal distribution curve with a Student’s t-distribution curve
to show the heavier tails]

pdf

α α
t
−tα,n 0 tα,n

32.4.7 As n → ∞
d
Tn −
→ Z ∼ N (0, 1). (1247)
Also,
χ2n Z 2 + · · · + Zn2
= 1 −−−→ 1. (1248)
n n n→∞

32.5 F Distribution
32.5.1 Definition
Let χ2n and χ2m be independent chi-square r.v.s with degrees of freedom n and m respectively.
Define
χ2 /n
Fn,m = 2n . (1249)
χm /m
Then:
Fn,m ∼ F (n, m). (1250)

32.5.2 F critical value (upper tail)


Define fα;n,m as:
P(Fn,m > fα;n,m ) = α. (1251)
32.6. TWO-POPULATION VARIANCE RATIO 183

32.5.3 Key reciprocal property derivation


Start with α = P(Fn,m > fα;n,m ). That is:
 2 
χn /n
α=P > fα;n,m . (1252)
χ2m /m
Invert inside probability:
χ2m /m
 
1
α=P 2
< . (1253)
χn /n fα;n,m
Hence
χ2m /m
 
1
1−α=P 2
≥ . (1254)
χn /n fα;n,m
χ2m /m
But χ2n /n
= Fm,n . Therefore:
 
1
P Fm,n ≥ = 1 − α. (1255)
fα;n,m

This implies the identity:


1
f1−α;m,n = . (1256)
fα;n,m

32.5.4 Sketch (right tail)


pdf

area = α
f
fα;n,m

32.6 Two-Population Variance Ratio


32.6.1 Setup
Consider two independent normal populations:

X ∼ N (µ1 , σ12 ), Y ∼ N (µ2 , σ22 ). (1257)

Sample sizes: n1 from X and n2 from Y . Let sample variances be S12 and S22 .

32.6.2 Chi-square forms


(n1 − 1)S12 (n2 − 1)S22
∼ χ2(n1 −1) , ∼ χ2(n2 −1) . (1258)
σ12 σ22

32.6.3 F ratio
Then:
S12 /σ12
∼ F (n1 − 1, n2 − 1). (1259)
S22 /σ22
184 CHAPTER 32. LECTURES 19 AND 20: SAMPLING DISTRIBUTIONS (χ2 , t, AND F )
Part III

Introduction to Machine Learning

185
Chapter 33

Lectures 1 and 2: Fundamentals of


Machine Learning and Hypothesis
Testing

Lecture 1

Fundamentals of Machine Learning


1. What is Machine Learning?
Machine Learning (ML) is a branch of data science in which algorithms learn patterns from
data and improve their performance through experience rather than explicit programming.
The primary goal of machine learning is to use historical data to make predictions or decisions
about future or unseen data. This makes ML particularly useful in complex environments
where fixed rules are difficult to define.
Common applications include prediction, classification, recommendation systems, and
automation.

2. Mathematical Setup
Machine learning aims to learn an unknown functional relationship between inputs and
outputs:

Y = f (X)
where
X = (x1 , x2 , . . . , xn ) ∈ Rn
represents the input feature vector, and

Y ∈R or Y ∈ {0, 1}

is the output variable.


The training dataset is given by:
D = {(x(i) , y (i) )}m
i=1

The objective is to estimate a function fˆ(·) that performs well not only on training data but
also on unseen data. This property is known as generalization.

187
188CHAPTER 33. LECTURES 1 AND 2: FUNDAMENTALS OF MACHINE LEARNING AND HYPOTHES

3. Loss Function and Learning Objective


Learning is performed by minimizing a loss function that measures prediction error:

L(y, ŷ)
Examples of loss functions:

ˆ Mean Squared Error (Regression)


ˆ Cross-Entropy Loss (Classification)
The learning problem can be written as the following optimization problem:

min E[L(Y, f (X))]


f

This expectation is taken over the underlying data distribution.

4. Types of Machine Learning


(a) Supervised Learning
In supervised learning, output labels are known.

Y = f (X)
Tasks:

1. Regression: Output is continuous, e.g., price, income.

2. Classification: Output is discrete.

Y ∈ {0, 1}, Y ∼ Bernoulli(p)

(b) Unsupervised Learning


Only input data is observed:
{x(i) }m
i=1

Objectives:

ˆ Pattern discovery
ˆ Dimensionality reduction
ˆ Feature extraction
Techniques:

1. Clustering

2. Dimensionality Reduction:

(x1 , . . . , xn ) → (z1 , . . . , zk ), k≪n

(c) Reinforcement Learning


An agent interacts with an environment and learns an optimal policy using rewards and
penalties. This approach is widely used in robotics, finance, and game playing.
189

5. Classification as a Geometric Problem


Binary encoding of classes:
Read = (1, 0), Write = (0, 1)

Inner product:
⟨(1, 0), (0, 1)⟩ = 0

Since the vectors are orthogonal, the classes are linearly separable. Classification corresponds
to finding a decision boundary in feature space.

Lecture 2

Statistical Hypothesis Testing


1. Statistical Hypothesis
A statistical hypothesis is a statement about population parameters such as mean, variance,
or proportion.

2. Null and Alternative Hypotheses


Let X be a random variable with mean µ.

H0 : µ = µ0

H1 : µ ̸= µ0

Types of alternatives:

ˆ Two-tailed: µ ̸= µ 0

ˆ Right-tailed: µ > µ 0

ˆ Left-tailed: µ < µ 0

3. Hypothesis Testing Steps


1. State H0 and H1

2. Choose significance level α

3. Compute test statistic

4. Determine rejection region or p-value

5. Make a decision
190CHAPTER 33. LECTURES 1 AND 2: FUNDAMENTALS OF MACHINE LEARNING AND HYPOTHES

Hypothesis Testing Graph


Critical Value t
H0 : µ = µ0
H1 : µ = µ1

Probability Density

β α

µ0 t µ1
x

Explanation:

ˆ Blue curve represents the null hypothesis distribution


ˆ Red dashed curve represents the alternative hypothesis distribution
ˆ Vertical line t is the critical value
ˆ Blue shaded area is the Type-I error (α)
ˆ Red shaded area is the Type-II error (β)
4. Errors in Hypothesis Testing
Reality / Decision Reject H0 Accept H0
H0 True Type-I Error (α) Correct
H0 False Correct (Power) Type-II Error (β)

α = P (Reject H0 | H0 true)
β = P (Accept H0 | H0 false)

5. p-Value
The p-value is the probability of observing a test statistic as extreme as the sample value
assuming H0 is true.
Decision rule:
Reject H0 if p-value < α

Central Limit Theorem (CLT)


Let
X1 , X2 , . . . , Xn ∼ i.i.d.(µ, σ 2 )
Sample mean:
n
1X
X̄n = Xi
n
i=1
191

Then,
X̄n − µ d
√ − → N (0, 1)
σ/ n
This holds regardless of the original distribution for large n.

Known vs Unknown Variance


X̄ − µ0
Z= √ ∼ N (0, 1)
σ/ n

X̄ − µ0
T = √ ∼ tn−1
s/ n

Hypothesis Testing Graph

H0
H1
Density

Threshold

µ0 t µ1
x

Explanation:

ˆ Left curve represents the null hypothesis


ˆ Right curve represents the alternative hypothesis
ˆ Overlapping regions correspond to Type-I and Type-II errors
Link Between Machine Learning and Hypothesis Testing
ˆ Binary classifier ↔ Hypothesis test
ˆ Decision boundary ↔ Critical value
ˆ False positive ↔ Type-I error
ˆ False negative ↔ Type-II error
192CHAPTER 33. LECTURES 1 AND 2: FUNDAMENTALS OF MACHINE LEARNING AND HYPOTHES
Chapter 34

Lectures 3 and 4: Hypothesis


Testing and Statistical Decisions

34.1 Introduction
Statistical inference provides a principled framework for reasoning under uncertainty. In many
real-world problems, we observe data generated by random mechanisms and seek to make
decisions about unknown population parameters. Hypothesis testing is one of the central
tools in this framework, allowing us to formally assess whether observed data are consistent
with a proposed model.
Special attention is given to the interpretation of Type I error, Type II error, and power
of a test, highlighting their roles as long-run frequency properties rather than statements of
absolute certainty. Graphical illustrations are used throughout to reinforce geometric intuition
behind rejection regions and probability mass under competing distributions.

34.2 Model Formulation


We observe a single random variable X with known variance.
Under the null hypothesis:
H0 : µ = 8 (1260)
X ∼ N (8, 16) (1261)
Thus,
µ0 = 8, σ 2 = 16, σ=4 (1262)
Under the alternative hypothesis:
H1 : µ = 25 (1263)
X ∼ N (25, 16) (1264)

34.2.1 Objective
Given one observation x, decide whether to reject H0 . Since µ1 = 25 > 8, large values of X
favor H1 . This is a right-tailed test.

34.3 Construction of Critical Region


We define the rejection region:
R = {x : x > C} (1265)

193
194CHAPTER 34. LECTURES 3 AND 4: HYPOTHESIS TESTING AND STATISTICAL DECISIONS

Choose C such that Type I error equals α:

P (X > C | H0 ) = α (1266)

34.4 Derivation of Critical Value


Under H0 :
X ∼ N (8, 16) (1267)
Standardize:
X −8
Z= ∼ N (0, 1) (1268)
4
Then:  
C −8
P (X > C) = P Z > (1269)
4
We require:  
C −8
P Z> =α (1270)
4
Thus:  
C −8
P Z≤ =1−α (1271)
4
Let
zα = Φ−1 (1 − α) (1272)
Then:
C −8
= zα
4
C = 8 + 4zα (1273)

34.5 Critical Values for Common α


34.5.1 α = 0.01
z0.01 = 2.33 (1274)
C = 8 + 4(2.33) = 17.32 (1275)

34.5.2 α = 0.05
z0.05 = 1.645 (1276)
C = 8 + 4(1.645) = 14.58 (1277)

34.5.3 α = 0.10
z0.10 = 1.282 (1278)
C = 8 + 4(1.282) = 13.13 (1279)

34.5.4 Observation
C0.01 > C0.05 > C0.10 (1280)
As α increases, rejection becomes easier.
34.6. DIAGRAM: 5% REJECTION REGION 195

34.6 Diagram: 5% Rejection Region


·10−2

Density 8

0
−5 0 5 10 15 20 25 30
x

Area to the right of C = 14.58 equals α = 0.05.

34.7 Example: Observed Value x = 15


Compare with critical values:

15 < 17.32 ⇒ Do not reject at 1% (1281)


15 > 14.58 ⇒ Reject at 5% (1282)
15 > 13.13 ⇒ Reject at 10% (1283)

Decision depends on chosen α.

34.8 p-value Derivation


Definition:
p = P (X ≥ x | H0 ) (1284)

Standardize:
15 − 8
Z= = 1.75 (1285)
4

p = P (Z > 1.75) (1286)

From tables:
Φ(1.75) = 0.9599 (1287)

p = 1 − 0.9599 = 0.0401 (1288)

34.8.1 Interpretation
Reject H0 whenever:
α > 0.0401 (1289)
196CHAPTER 34. LECTURES 3 AND 4: HYPOTHESIS TESTING AND STATISTICAL DECISIONS

34.9 Diagram: p-value


·10−2

Density 8

0
−5 0 5 10 15 20 25 30
x

Red region = probability of observing value at least as extreme as 15.

34.10 Power of the Test


Power:
Power = P (Reject H0 | H1 true) (1290)
Using α = 0.05, C = 14.58.
Under H1 :
X ∼ N (25, 16) (1291)

14.58 − 25
Z= = −2.605 (1292)
4

Power = P (Z > −2.605)


= 1 − Φ(−2.605) (1293)

Φ(−2.605) = 0.0046 (1294)

Power = 0.9954 (1295)

34.11 Diagram: Power of the Test


·10−2

8
Density

0
−5 0 5 10 15 20 25 30 35 40
x
34.12. CONCEPTUAL SUMMARY 197

Area right of C under green curve = Power.

34.12 Conceptual Summary


ˆ Type I Error = Reject true H 0

ˆ Type II Error = Fail to reject false H 0

ˆ Power = 1 − β
ˆ Smaller α  harder to reject
ˆ Greater separation of means  higher power
Statistical inference controls error probabilistically — it never yields absolute certainty.

34.13 Statistical Decision Framework


Hypothesis testing is a formal decision procedure for choosing between two competing
statistical models based on observed data.
We consider two simple hypotheses:
H0 : µ = 8 (1296)
H1 : µ = 25 (1297)
The data consist of a single observation X where

X ∼ N (µ, 16), σ 2 = 16, σ=4 (1298)

Thus under each hypothesis:

H0 : X ∼ N (8, 16) (1299)


H1 : X ∼ N (25, 16) (1300)

This is a simple vs simple hypothesis testing problem.

34.14 Decision Rule and Rejection Region


A test is defined through a rejection region R.
We choose a rule of the form:
R = {x : x > C} (1301)
Why right-tailed? Because under H1 , the mean is larger. Large values of X are more
consistent with H1 than with H0 .

34.15 Type I and Type II Errors


There are two types of errors:

34.15.1 Type I Error


Rejecting H0 when H0 is true.
α = P (X ∈ R | H0 ) (1302)
198CHAPTER 34. LECTURES 3 AND 4: HYPOTHESIS TESTING AND STATISTICAL DECISIONS

34.15.2 Type II Error


Failing to reject H0 when H1 is true.

β = P (X ∈
/ R | H1 ) (1303)

34.15.3 Power
Power = 1 − β (1304)
Power measures the probability of correctly rejecting H0 when H1 is true.

34.16 Simulation Interpretation of Type I Error


Suppose H0 is true and we repeatedly sample from

X ∼ N (8, 16) (1305)

If we fix
α = 0.01 (1306)
then by definition,
α = P (Reject H0 | H0 true) (1307)
Interpretation: If we repeat the experiment 1000 times while H0 is true,

Expected number of false rejections = 0.01 × 1000 = 10 (1308)

Thus:

ˆ About 990 times we accept H 0

ˆ About 10 times we incorrectly reject H 0

Type I error is therefore a long-run frequency property.

34.17 Type II Error Derivation


Now assume the alternative is true:
H1 : µ = 25 (1309)
Recall the rejection rule: Reject H0 if X > C. Thus,

β = P (Accept H0 | H1 true) (1310)

Since we accept H0 when X ≤ C,

β = P (X ≤ C | µ = 25) (1311)

Under H1 :
X ∼ N (25, 16) (1312)
Standardize:
X − 25
Z= (1313)
4
Therefore,  
C − 25
β=P Z≤ (1314)
4
34.18. GRAPHICAL ILLUSTRATION OF β AND POWER 199

34.17.1 Type II Error when α = 0.01


From earlier,
C = 17.32 (1315)

Thus,
C − 25 17.32 − 25
= = −1.92 (1316)
4 4
Hence,
β = P (Z ≤ −1.92) (1317)

From standard normal tables:


β = 0.0274 (1318)

34.17.2 Power of the Test Calculation


Power is defined as
Power = 1 − β (1319)

Thus,
Power = 1 − 0.0274 = 0.9726 (1320)

Power ≈ 97% (1321)

Interpretation: If the true mean is 25, the test correctly rejects H0 approximately 97% of the
time.

34.18 Graphical Illustration of β and Power


·10−2

8
Density

0
−5 0 5 10 15 20 25 30 35 40
x

Interpretation:

ˆ Area to right of C under H curve = α 0

ˆ Area to left of C under H curve = β


1

ˆ Area to right of C under H curve = Power


1
200CHAPTER 34. LECTURES 3 AND 4: HYPOTHESIS TESTING AND STATISTICAL DECISIONS

34.19 One-Sided Test


Current alternative:
H1 : µ = 25 > 8 (1322)
Thus rejection region is only on the right side:

R = {X > C} (1323)

This is called a right-tailed test.


Decision rule:

X > C ⇒ Reject H0
X ≤ C ⇒ Accept H0 (1324)

34.20 Two-Sided Test


If instead we specify:
H1 : µ ̸= 8 (1325)
then deviations in both directions matter.
We split significance level:
α α
α= + (1326)
2 2
Rejection region becomes:
X < C1 or X > C2 (1327)
Where

C1 = 8 − 4zα/2 (1328)
C2 = 8 + 4zα/2 (1329)

Thus we reject if
|Z| > zα/2 (1330)

34.21 Comparison: One-Sided vs Two-Sided


ˆ One-sided test places all α in one tail.
ˆ Two-sided test splits α equally.
ˆ One-sided test has higher power in a specified direction.
ˆ Two-sided test protects against deviations in both directions.
34.22 Summary of the Complete Testing Procedure
1. Specify H0 and H1

2. Fix significance level α

3. Derive critical value C

4. Compute test statistic

5. Compare with rejection region


34.23. TWO-SIDED HYPOTHESIS TEST FOR POPULATION MEAN 201

6. Optionally compute p-value


7. Compute β and power if required

Fundamental Insight: Hypothesis testing is not about proving a hypothesis true. It is


about controlling error probabilities under uncertainty.

34.23 Two-Sided Hypothesis Test for Population Mean


34.23.1 Problem Setup
We want to test whether the population mean is equal to a specified value.
H0 : µ = 8 (1331)
Ha : µ ̸= 8 (1332)
This is a two-sided (two-tailed) test because deviations on both sides of the hypothesized
mean are considered evidence against H0 .

34.23.2 Population Assumption


Let the population distribution be:
X ∼ N (µ, σ 2 ) (1333)
We take a random sample of size n:
X1 , X2 , . . . , Xn ∼ N (µ, σ 2 ) (1334)
The sample mean is:
n
1X
X̄ = Xi (1335)
n
i=1
Since each Xi is normally distributed, the sampling distribution of X̄ is:
σ2
 
X̄ ∼ N µ, (1336)
n

34.23.3 Test Statistic


If σ is known, the appropriate test statistic is:
X̄ − µ0
Z= √ (1337)
σ/ n
Under H0 , this follows:
Z ∼ N (0, 1) (1338)

34.23.4 Significance Level and Critical Region


Let the level of significance be α. For a two-sided test:
α = P (Reject H0 | H0 true) (1339)
This becomes:
α = P (Z < −zα/2 ) + P (Z > zα/2 ) (1340)
Decision Rule: Reject H0 if
|Z| > zα/2 (1341)
Otherwise, do not reject H0 .
202CHAPTER 34. LECTURES 3 AND 4: HYPOTHESIS TESTING AND STATISTICAL DECISIONS

34.23.5 Graphical Representation

Accept Region

Reject H0 Reject H0
−zα/2 0 zα/2

The shaded regions represent the rejection regions, each having probability α/2.

34.23.6 Theoretical Interpretation


A statistical hypothesis is a statement about the parameters of a population probability
distribution.
The objective of hypothesis testing is:

To develop a procedure for determining whether the observed sample data are
consistent with the given null hypothesis.

If the observed sample mean x̄ lies far from µ0 (measured in standard error units), then the
probability of such an observation under H0 is small. If this probability is less than α, we
reject H0 .

34.23.7 One Realization of the Random Variable


We observe one realization x of the random variable X. Using this observation, we compute
the test statistic:
x̄ − 8
z= √ (1342)
σ/ n
If:
|z| > zα/2 (1343)
then we conclude that the observed sample is not consistent with H0 . Otherwise, we do not
reject H0 .

34.23.8 Summary
ˆ Null Hypothesis: µ = 8
ˆ Alternative Hypothesis: µ ̸= 8
ˆ Distribution: X ∼ N (µ, σ )
2

ˆ Sampling Distribution: X̄ ∼ N (µ, σ /n)


2

ˆ Test Statistic: Z =X̄−µ


√0
σ/ n

ˆ Reject H if |Z| > z


0 α/2

This completes the theoretical and graphical explanation of the two-sided test for population
mean.
34.24. A NOTE ON HYPOTHESIS ”ACCEPTANCE” 203

34.24 A Note on Hypothesis ”Acceptance”


When accepting a given Hypothesis, we are not actually claiming it is true, but rather we are
saying that the resulting data appears to be consistent with the Hypothesis.

34.24.1 Significance Levels


Consider a population with distribution Fθ , where the parameter θ (mean) is unknown.
X ∼ Fθ =⇒ E(X) = θ, Var(X) = 1 (1344)

34.24.2 Null Hypothesis (H0 ) about θ


ˆ a) Simple Hypothesis:
H0 : θ = 1 (1345)
If (a) is true, the population distribution is completely specified:
X ∼ N (1, 1) (1346)

ˆ b) Composite Hypothesis:
H0 : θ ≤ 1 (1347)
If (b) is true, it does not specify a single population distribution because θ can take any
value ≤ 1.

34.25 Decision Rules and Rejection Regions


To test a specific H0 , we observe a population sample (x1 , x2 , . . . , xn ) of size n. These values
are realizations of the random variables X1 , X2 , . . . , Xn .

34.25.1 Rejection Region (Critical Region)


We decide whether or not to accept H0 based on where the sample vector lies in
n-dimensional space.
ˆ Reject H if (x , x , . . . , x ) ∈ C (Critical Region).
0 1 2 n

ˆ Accept H if (x , x , . . . , x ) ∈/ C.
0 1 2 n

34.25.2 Example: Testing H0 : θ = 1


For a population X ∼ N (θ, 1), the rejection region C is defined as:
 
1.96
C = (x1 , x2 , . . . , xn ) : |x̄ − 1| > √ (1348)
n
Where x̄ is the sample average:
n
1X
X̄ = Xi (1349)
n
i=1
Since Xi ∼ N (µ, σ 2 ), the distribution of the mean is:
σ 2 Under H0
   
1
X̄ ∼ N µ, −−−−−−→ X̄ ∼ N 1, (1350)
n n
204CHAPTER 34. LECTURES 3 AND 4: HYPOTHESIS TESTING AND STATISTICAL DECISIONS

34.25.3 Visualizing the Acceptance Region


1.96
The condition |x̄ − 1| > √
n
splits into two critical tails:

1.96
1. x̄ > 1 + √
n
(Right tail)
1.96
2. x̄ < 1 − √
n
(Left tail)

f (x̄)

Acceptance (95%)

CL (2.5%) CR (2.5%) x̄
1 − 1.96
√ 1 1+ 1.96

n n

Conclusion: H0 is rejected if the resulting observed data are very unlikely (i.e., fall in the
shaded tails) when H0 is true.
Chapter 35

Lectures 5 and 6: Hypothesis


Testing for Population Mean

35.1 Introduction
In statistical inference we frequently wish to determine whether observed data are consistent
with a specified value of an unknown population parameter. This decision problem is
formalized through statistical hypothesis testing.
The goal is not to prove a hypothesis true, but rather to determine whether the data provide
sufficient evidence to reject it.

35.2 Statistical Hypotheses


Definition 35.2.1 (Null and Alternative Hypotheses). Let θ be a parameter of a population
distribution.
The null hypothesis is a statement of the form

H0 : θ ∈ W

where W is a specified set of parameter values.


The alternative hypothesis describes values not contained in W .
A decision is based on a statistic computed from the sample:

T = T (X1 , . . . , Xn ).

If the observed value of T is unlikely under H0 , we reject H0 .

35.3 Errors in Hypothesis Testing


Definition 35.3.1 (Type I Error). Rejecting H0 when it is true.

α = P (Reject H0 | H0 True)

The constant α is called the significance level.


Definition 35.3.2 (Type II Error). Failing to reject H0 when it is false.

β = P (Accept H0 | H0 False)

Remark 35.3.1. The function β(θ) describes the probability of accepting H0 for each possible
true parameter value.

205
206CHAPTER 35. LECTURES 5 AND 6: HYPOTHESIS TESTING FOR POPULATION MEAN

35.4 Testing the Mean of a Normal Population


Suppose
X1 , . . . , Xn ∼ N (µ, σ 2 ), σ 2 known.

We test
H0 : µ = µ0 , HA : µ ̸= µ0 .

The sample mean is


n
1X
X̄ = Xi .
n
i=1

Theorem 35.4.1 (Distribution of the Sample Mean). If Xi ∼ N (µ, σ 2 ) then

σ2
 
X̄ ∼ N µ, .
n

35.5 Construction of the Test


Under H0 ,
σ2
 
X̄ ∼ N µ0 , .
n
Define the standardized statistic
X̄ − µ0
Z= √ .
σ/ n

Theorem 35.5.1 (Standard Normal Test Statistic). If H0 is true then

Z ∼ N (0, 1).

35.6 Rejection Region


We reject H0 when the observed value of Z is too large in magnitude.

Theorem 35.6.1 (Two-Sided Z Test). For significance level α, reject H0 if

|Z| > zα/2 .

Interpretation
The value zα/2 satisfies
α
P (Z > zα/2 ) = .
2
Thus the probability of rejecting a true null hypothesis equals α.

Rejection Region Diagram

Reject H0 Reject H0
z
35.7. ACCEPTANCE REGION 207

35.7 Acceptance Region


Values of Z within
−zα/2 ≤ Z ≤ zα/2
are considered consistent with H0 .

1−α

35.8 P-Value
Definition 35.8.1 (P-value). The p-value is the probability, under H0 , of observing a test
statistic at least as extreme as the observed value.

Decision rule:
Reject H0 if p-value < α.

35.9 Examples
Example 35.9.1. Test
H0 : µ = 168, HA : µ ̸= 168
with
n = 36, X̄ = 169.5, σ = 3.9.

Z = 2.307 > 1.96.

Conclusion: Reject H0 .

Example 35.9.2.
n = 5, X̄ = 8.5, σ = 2, µ0 = 8.

Z = 0.559.

Conclusion: Do not reject H0 .

35.10 Type II Error and OC Curve


When µ ̸= µ0 ,
σ2
 
X̄ ∼ N µ, .
n
Type II error probability:
 
X̄ − µ0
β(µ) = Pµ −zα/2 < √ < zα/2 .
σ/ n

Definition 35.10.1 (Operating Characteristic Curve). The function β(µ) is called the OC
curve.
208CHAPTER 35. LECTURES 5 AND 6: HYPOTHESIS TESTING FOR POPULATION MEAN

β(µ)

35.11 Statistical Hypothesis


A statistical hypothesis is a statement about an unknown parameter of a population
distribution.
In estimation problems we approximate the value of an unknown parameter. In hypothesis
testing we evaluate whether a proposed value of the parameter is plausible.

Definition 35.11.1 (Null Hypothesis). The null hypothesis H0 is a statement representing no


effect or no difference.

Definition 35.11.2 (Alternative Hypothesis). The alternative hypothesis Ha contradicts the


null hypothesis and usually contains an inequality.

Example:
H0 : θ = 75, Ha : θ ̸= 75.

35.12 Testing Equality of Two Population Means


Let
µ1 = mean of population 1, µ2 = mean of population 2.
Common hypotheses:

H0 : µ1 = µ2 , Ha : µ1 > µ2 , Ha : µ1 < µ2 , Ha : µ1 ̸= µ2 .

The alternative hypothesis always introduces a directional or non-equality claim.

35.13 Test Statistic


Definition 35.13.1. A test statistic is a numerical quantity computed from sample data used
to decide whether to reject H0 .

Statistical testing is a decision rule based on the observed value of the test statistic.

35.14 Acceptance and Rejection Regions


The range of possible values of the test statistic is divided into two parts:

ˆ Acceptance region  accept H 0

ˆ Rejection region  reject H0


35.15. TYPES OF ERRORS 209

35.15 Types of Errors


Definition 35.15.1 (Type I Error). Rejecting H0 when it is true.

α = P (Reject H0 |H0 true)

Definition 35.15.2 (Type II Error). Accepting H0 when it is false.

β = P (Accept H0 |H0 false)

35.16 Decision Table


Decision H0 True H0 False
Accept H0 Correct Type II error (β)
Reject H0 Type I error (α) Correct

35.17 Power of a Test


Definition 35.17.1. The power of a test is the probability of rejecting a false null hypothesis.

Power = 1 − β.

Goal of testing:
Minimize α and β.

In practice, α is fixed first, then power is maximized.

35.18 Likelihood Ratio Test


Let L(θ) be the likelihood function.

Definition 35.18.1 (Likelihood Ratio Statistic).

L(θ0 )
λ=
L(θ̂)

where θ̂ is the maximum likelihood estimator.

Reject H0 when λ is sufficiently small.

Remark 35.18.1. Among all tests with level α, the likelihood ratio test often has maximum
power.

35.19 Large Sample Z Test


For large samples,
Z ∼ N (0, 1).

Decision rules depend on the alternative hypothesis.


210CHAPTER 35. LECTURES 5 AND 6: HYPOTHESIS TESTING FOR POPULATION MEAN

35.20 Right-Tailed Test


H0 : θ = θ0 , Ha : θ > θ0 .
Reject H0 if
Z > zα .

35.21 Left-Tailed Test


H0 : θ = θ0 , Ha : θ < θ0 .
Reject H0 if
Z < −zα .

35.22 Two-Tailed Test


H0 : θ = θ0 , Ha : θ ̸= θ0 .
Reject H0 if
|Z| > zα/2 .

α/2 α/2

35.23 Introduction to Hypothesis Testing


A Statistical Hypothesis is a statement about an unknown parameter (θ) of a population
distribution. Unlike estimation, where we find an approximate value, hypothesis testing
determines if a specific statement about that parameter is True (T) or False (F).

35.23.1 Null vs. Alternative Hypotheses


ˆ Null Hypothesis (H ): A statement of ”no difference.” It is the status quo. Example:
0
H0 : θ = 75 or H0 : µ1 = µ2 .

ˆ Alternative Hypothesis (H or H1 ): A statement that contradicts H0 . It usually


a
contains an inequality. Example: Ha : θ > 75 (One-tailed) or Ha : µ1 ̸= µ2 (Two-tailed).

35.24 The Decision Framework


To test a hypothesis, we use a Test Statistic computed from sample data. The possible
values of this statistic are divided into two regions:

1. Acceptance Region (AR): If the test statistic falls here, we accept H0 .

2. Rejection Region (RR): If the test statistic falls here, we reject H0 and accept Ha .
35.25. LIKELIHOOD RATIO TEST (LRT) 211

35.24.1 Errors in Decision Making


Because decisions are based on samples, two types of errors can occur:

Decision H0 is True (Reality) H0 is False (Reality)


Accept H0 Correct Decision Type II Error (β)
Reject H0 Type I Error (α) Correct Decision

ˆ Type I Error (α): Rejecting H 0 when it is actually true. α is called the Significance
Level.

ˆ Type II Error (β): Accepting H when it is actually false.


0

ˆ Power of the Test (1 − β): The probability of correctly rejecting a false H . Goal:
0
Maximize (1 − β) for a fixed α.

Note: Type I error is considered more serious. We generally fix α to a very low
value (e.g., 0.05) and then try to minimize β.

35.25 Likelihood Ratio Test (LRT)


Developed by R.A. Fisher, the LRT is used to find a powerful test statistic (λ):

L(θ0 )
λ=
L(θ̂)

Where L(θ0 ) is the likelihood under H0 , and θ̂ is the Maximum Likelihood Estimate
(MLE).
212CHAPTER 35. LECTURES 5 AND 6: HYPOTHESIS TESTING FOR POPULATION MEAN
Chapter 36

Lectures 7 and 8: Likelihood Ratio


Test and Large Sample Z Test

36.1 Simple and Composite Alternative Hypotheses


We begin by distinguishing between simple and composite alternatives.

ˆH a : θ = θa
This is called a simple alternative hypothesis since the parameter θ has only one
specified value θa .
ˆH a : θ > θa
This is called a composite alternative hypothesis since many possible values of θ are
included.

Example

θ = 75 is a simple alternative hypothesis (1)


θ > 75 is a composite alternative hypothesis (2)

Hypothesis Structure

H0 : µ = 75 (simple) (3)
Ha : µ > 75 (composite) (4)
Remark: This is a case of simple vs composite testing.
Note: Mostly H0 is simple, but Ha can be either simple or composite.

36.2 Likelihood Ratio Test Statistic


We now define the likelihood ratio test statistic.

L(θ0 )
λ= (1351)
L(θ̂)
where:
ˆ L(θ ) is the likelihood under the null hypothesis H
0 0 : θ = θ0
ˆ L(θ̂) is the likelihood evaluated at the MLE of θ
213
214CHAPTER 36. LECTURES 7 AND 8: LIKELIHOOD RATIO TEST AND LARGE SAMPLE Z TEST

Interpretation
ˆ If λ is small, then L(θ ) is small relative to the maximum likelihood.
0

ˆ Hence, the observed data does not support H . 0

Parameter under Hypothesis

H0 : θ = θ0 (6)
θ̂ : MLE of θ (7)

Thus:

ˆ θ represents the value of the parameter under H


0 0

ˆ θ̂ is the estimate obtained from the data


36.3 Level of Significance
Let α denote the level of significance (size of the test).

α = P (λ ≤ λα ) (1352)
where λα is the critical value.
Thus, α represents the probability of rejecting H0 when it is true.

Decision Rule
λ ≤ λα ⇒ Reject H0 (1353)

Equivalently:

ˆ If the likelihood ratio is sufficiently small, reject H 0

ˆ Otherwise, do not reject H 0

Final Interpretation:

ˆ α is the level of significance (or size of the test)


ˆ It is fixed in advance before performing the test
36.4 Further Remarks on Likelihood Ratio Test
α = level of significance of test (1354)
or equivalently, the size of the test is fixed.
If
λ ≤ λα (1355)
then the sample does not support H0 : θ = θ0 .
Hence, we reject H0 and accept Ha : θ ̸= θ0 .
36.5. EXAMPLE SETUP (I.I.D SAMPLE) 215

Likelihood Ratio Form


Likelihood under Null L(θ0 )
λ= = (1356)
Likelihood at MLE L(θ̂)

If the ratio is small, then L(θ0 ) is small.

Ratio small ⇒ L(θ0 ) is small (1357)

Probability Interpretation
 
α = P λ ≤ λα (1358)

= P [Reject H0 ] = P [Accept Ha ] (1359)

36.5 Example Setup (i.i.d Sample)


Let
X = (X1 , X2 , . . . , Xn ) (1360)

be a sample of size n.

Observed sample:
x = (x1 , x2 , . . . , xn ) (1361)

where X1 , X2 , . . . , Xn are i.i.d.

Population Assumption

Population ∼ N (µ, 1) (1362)

i.e.,
X ∼ N (µ, 1) (1363)

Thus, variance is known:


σ2 = 1 (given) (1364)

36.6 Likelihood Ratio in this Case


L(θ0 )
λ= (1365)
L(θM LE )

or equivalently,
L(θ0 )
λ= (1366)
L(θ̂)

This completes the setup before deriving the likelihood explicitly.


216CHAPTER 36. LECTURES 7 AND 8: LIKELIHOOD RATIO TEST AND LARGE SAMPLE Z TEST

36.7 Likelihood Function and MLE (Detailed)


We consider the test:

H0 : µ = 0 (1)
Ha : µ > 0 (2)

L(θ0 )
λ= (3)
L(θ̂M LE )

Likelihood under H0
L(θ0 ) = L(µ = 0) (4)

General Likelihood
L(µ, σ 2 ) = f (x1 , x2 , . . . , xn | µ, σ 2 ) (5)

Here (x1 , x2 , . . . , xn ) are fixed data (not random variables).


n
Y
= f (xi | µ, σ 2 ) (6)
i=1

n 
xi −µ
2
Y 1 −1
= √ e 2 σ
(7)
i=1
σ 2π

Putting σ 2 = 1
n
Y 1 1 2
= √ e− 2 (xi −µ) (8)
i=1

n
Y 1 1 2
L(µ) = √ e− 2 (xi −µ) (9)
i=1

 n
n Y
1 1 2
= √ e− 2 (xi −µ) (10)
2π i=1
 n
1 1 Pn 2
= √ e− 2 i=1 (xi −µ) (11)

Expanded Form
 n  n
1 − 21 [(x1 −µ)2 +(x2 −µ)2 +···+(xn −µ)2 ] 1 1 P
(xi )2 ]
L(µ) = √ e = √ e− 2 [ (12)
2π 2π

Under H0 (Numerator)
 n
1 1 2 2 2
L(µ = 0) = √ e− 2 [x1 +x2 +···+xn ] (13)

36.7. LIKELIHOOD FUNCTION AND MLE (DETAILED) 217

Likelihood at MLE (Denominator)

L(θ̂M LE ) = L(µ̂M LE ) = L(x̄) (14)

 n
1 1 Pn 2
= √ e− 2 i=1 (xi −x̄) (15)

MLE of µ

µ̂M LE = x̄ (16)

Rewriting Likelihood
 n
1 1 Pn 2
L(µ) = √ e− 2 i=1 (xi −µ) (17)

Log Likelihood

ℓ(µ) = log L(µ) (18)

 n 
1 − 12
P
(xi −µ)2
= log √ e (19)

  n
1 1X
= n log √ − (xi − µ)2 (20)
2π 2
i=1

Derivative (Step-by-Step)

dℓ(µ) 1 X
=− ·2 (xi − µ)(−1) (21)
dµ 2

X
= (xi − µ) (22)

X X
= xi − µ (23)

X
= xi − nµ (24)

=0 (25)

µ̂M LE = x̄ (26)
218CHAPTER 36. LECTURES 7 AND 8: LIKELIHOOD RATIO TEST AND LARGE SAMPLE Z TEST

36.8 Derivation of Likelihood Ratio (Final Form)


Derivative Continuation
n
dℓ(µ) X
= (xi − µ) (27)

i=1

n
X n
X
= xi − µ (28)
i=1 i=1

n
X
= xi − nµ (29)
i=1

=0 (30)

n
X
xi = nµ (31)
i=1

n
1X
µ̂M LE = xi = x̄ (32)
n
i=1

Second Derivative
d2 ℓ(µ)
= −n < 0 (33)
dµ2

Hence, maximum is attained.

Let θ ≡ µ.

Likelihood Ratio
L(θ0 ) L(µ = 0)
λ= = (34)
L(θ̂M LE ) L(x̄)

Substitute Expressions
 n
1 1 Pn
x2i
L(0) = √ e− 2 i=1 (35)

 n
1 1 Pn 2
L(x̄) = √ e− 2 i=1 (xi −x̄) (36)

Form Ratio
1
x2i
P
e− 2
λ= 1 (37)
(xi −x̄)2
P
e− 2
36.9. LIKELIHOOD RATIO AND REJECTION REGION (DETAILED) 219

Expand Denominator Term


X X
(xi − x̄)2 = x2i + x̄2 − 2xi x̄

(38)

X X X
= x2i + x̄2 − 2x̄ xi (39)

X X
= x2i + nx̄2 − 2x̄ xi (40)

X
xi = nx̄ (41)

X
= x2i + nx̄2 − 2nx̄2 (42)

X
= x2i − nx̄2 (43)

Substitute Back
1
x2i
P
e− 2
λ= 1 (44)
x2i −nx̄2 )
P
e− 2 (
1 1
x2i x2i −nx̄2 )
P P
= e− 2 · e2( (45)

1 2
= e− 2 nx̄ (46)

Final Result
n 2
λ = e− 2 x̄ (47)

36.9 Likelihood Ratio and Rejection Region (Detailed)


From earlier:
n 2
λ = e− 2 x̄ (48)

Decision Rule
If λ ≤ λα reject H0 with level of significance α (49)

Substitution
n 2
e− 2 x̄ ≤ λα (50)

Taking Log
n
− x̄2 ≤ ln(λα ) (51)
2

2
x̄2 ≥ − ln(λα ) (52)
n
220CHAPTER 36. LECTURES 7 AND 8: LIKELIHOOD RATIO TEST AND LARGE SAMPLE Z TEST

Define Critical Quantity


Let
2
χ2α = − ln(λα ) (53)
n

x̄2 ≥ χ2α (54)

|x̄| ≥ χα (55)

If λ ≤ λα reject H0 (56)

2
x̄2 ≥ − ln(λα ) (57)
n
Let
2
χ2α = − ln(λα ) (58)
n

x̄2 ≥ χ2α (59)


Solve for x̄.

36.10 Recall
x2 ≥ y 2 (60)
Solve for x:

x2 − y 2 ≥ 0 (61)

(x − y)(x + y) ≥ 0 (62)

Case (1)
When both (x − y) ≥ 0 and (x + y) ≥ 0:

x≥y (63)
x ≥ −y (64)

Case (2)
When both (x − y) ≤ 0 and (x + y) ≤ 0:

x≤y (65)
x ≤ −y (66)
ˆ If x ≥ 0 and y ≥ 0, then x ≥ y
ˆ If x ≤ 0 and y ≤ 0, then −x ≥ −y
ˆ If x ≥ 0 and y ≤ 0, then x ≥ −y
ˆ If x ≤ 0 and y ≥ 0, then −x ≥ y
36.11. FINAL RESULT 221

Conclusion
|x| ≥ y (67)

36.11 Final Result


|x̄| ≥ χα (68)

Reject H0 if |x̄| ≥ χα (69)

(i) If x ≥ 0 and y ≥ 0, then


x≥y (70)

(ii) If x ≤ 0 and y ≤ 0, then


−x ≥ −y (71)

(iii) If x ≥ 0 and y ≤ 0, then


x ≥ −y (72)

Likelihood Ratio Form


L(µ = 0)
λ= <1 (73)
L(x̄)
since denominator is MLE.

Result from Earlier


x̄2 ≥ χ2α (74)

|x̄| ≥ |χα | (75)

Distribution of Sample Mean


σ2
 
X̄ ∼ N µ, (76)
n
Given σ 2 = 1,
 
1
X̄ ∼ N µ, (77)
n

Critical Quantity
 
2
χ2α = − ln(λα ) ≥ 0 (78)
n
Hence,

χ2α ≥ 0 (79)
Thus both χα and −χα satisfy the condition.
222CHAPTER 36. LECTURES 7 AND 8: LIKELIHOOD RATIO TEST AND LARGE SAMPLE Z TEST

Final Inequality
|x̄| ≥ χα or |x̄| ≥ −χα (80)

Sign-Based Breakdown
ˆ If x̄ > 0, then
x̄ ≥ χα (81)
ˆ If x̄ < 0, then
−x̄ ≥ χα ⇒ x̄ ≤ −χα (82)

Equivalent Form
x̄ ≥ χα or x̄ ≤ −χα (83)

Two Tail Test

H0 : µ = µ0 (48)
Ha : µ ̸= µ0 (49)

Area (1 − α)
Rejection Region Rejection Region
α/2 α/2

−zα/2 zα/2

Find from Normal Table.

α α
Size of Test = + =α (50)
2 2
Reject H0 with level α.
Acceptance Ha : µ ̸= 0

Left Tail Test

H0 : µ = µ0 (51)
Ha : µ < µ0 (52)

Left Tail Test


area = α

−zα

Rejection region: Z < −zα



36.11. FINAL RESULT 223

Right Tail Test

H0 : µ = µ0 (53)
Ha : µ > µ0 (54)

Right Tail Test


area = α

Rejection: Z > zα (55)

Two Tail Test

H0 : µ = µ0 (56)
Ha : µ ̸= µ0 (57)

Area (1 − α)
Rejection Region Rejection Region
α/2 α/2

−zα/2 zα/2

Find from Normal Table.

α α
Size of Test = + =α (58)
2 2
given.
Reject H0 with level α
Acceptance Ha : µ ̸= 0

Large Sample

n > 30 (Z - statistics) (59)

X̄ − µ
Z= √ ∼ N (0, 1) (60)
σ/ n

224CHAPTER 36. LECTURES 7 AND 8: LIKELIHOOD RATIO TEST AND LARGE SAMPLE Z TEST

Small Sample

n < 30 (61)

X̄ − µ0
t-statistic: T = √ (62)
S/ n

X̄ − µ
T = √ (63)
S/ n

under H0 : µ = µ0

Decision Rules
Right Tail

P (T > tα ) = α (fixed) (64)

got from t-table


Left Tail

P (T < −tα ) = α (65)

Two Tail

P (|T | > tα/2 ) = α (66)

(T > tα/2 ) or (T < −tα/2 ) (67)

area = α/2 area = α/2

Example

H0 : µ = 1 (68)
Ha : µ > 1 (69)

dof = (n − 1) = 19 (70)
36.11. FINAL RESULT 225

X̄ = 1.7, s = 1.8, n = 20 (71)

α = 5% (72)

Test Statistic:

X̄ − µ0
T = √ (73)
s/ n
under H0

1.7 − 1
t= √ = 1.739 (74)
1.8/ 20

t Tables

dof α = 0.05
19 t0.05,19 = 1.729

Reject H0

α = 0.05

1.729

tcal > t0.05,19 (75)

1.739 > 1.729 (76)

Reject H0 (77)

Large Sample

n > 30 (78)

Comparing µ1 and µ2 from Pop 1 and Pop 2


Test if Population means are same or different.

226CHAPTER 36. LECTURES 7 AND 8: LIKELIHOOD RATIO TEST AND LARGE SAMPLE Z TEST

Two tailed Test

H0 : (µ1 − µ2 ) = D0 (specified value of difference) (79)


Ha : (µ1 − µ2 ) ̸= D0 (80)

σ12
 
X̄1 ∼ N µ1 , (81)
n1

σ22
 
X̄2 ∼ N µ2 , (82)
n2

 2
σ2
 
σ1
[X̄1 − X̄2 ] ∼ N (µ1 − µ2 ), + 2 (83)
n1 n2

Var(X + Y ) = Var(X) + Var(Y ) (84)


Var(X − Y ) = Var(X) + Var(Y ) (85)

[X̄1 − X̄2 ] − (µ1 − µ2 )


q 2 ∼ N (0, 1) (86)
σ1 σ22
n1 + n2

[X̄1 − X̄2 ] − D0
Z= q 2 (87)
S1 S22
n1 + n2


Decision Rule

Reject if z > zα/2 or z < −zα/2 (88)

−zα/2 zα/2
36.11. FINAL RESULT 227

Left Tail Test

H0 : (µ1 − µ2 ) = D0 (89)
Ha : (µ1 − µ2 ) < D0 (90)

Reject if Z < −zα (91)

−zα


Note: For small sample n < 30, use t-test instead of Z-statistic.

Pop 1 → n1 (92)

Pop 2 → n2 (93)

dof = (n1 − 1) + (n2 − 1) = n1 + n2 − 2 (94)

Difference between Two population proportions

Large sample: n > 30 (95)

H0 : p = p0 (96)
Ha : p ̸= p0 (97)

p̂ − p0
Z= q under H0 (98)
pq
n

p̂ is estimation of p from data.

q0 = 1 − p 0 (99)


228CHAPTER 36. LECTURES 7 AND 8: LIKELIHOOD RATIO TEST AND LARGE SAMPLE Z TEST

X ∼ Bin(n, p), p = P (H) (100)

n
σ2
 
1X
X̄ ∼ N µ, , X̄ = Xi (101)
n n
i=1

Xi ∼ Bern(p) (102)

X ∼ Bin(n, p) (103)

(
1 with probability p = P (H)
Xi = (104)
0 with probability q = 1 − p = P (T )

n
X
X= Xi = (X1 + X2 + · · · + Xn ) (105)
i=1

n
X 1X
p̂ = = Xi (106)
n n
i=1

n
1X 1
E(p̂) = E(Xi ) = (np) = p (107)
n n
i=1

n
1 X 1 pq
Var(p̂) = 2 Var(Xi ) = 2 (npq) = (108)
n n n
i=1

Standardization using CLT:

p̂n − p n→∞
q −−−→ Z ∼ N (0, 1) (109)
pq CLT
n


36.11. FINAL RESULT 229

For Two Populations (2-sided test)

D0 = 0 (110)

H0 : (p1 − p2 ) = D0 (111)
Ha : (p1 − p2 ) ̸= D0 (112)

Test statistic:

(p̂1 − p̂2 ) − D0
Z= q (113)
p̂1 q̂1 p̂2 q̂2
n1 + n2

(Here q̂1 = 1 − p̂1 , q̂2 = 1 − p̂2 )


230CHAPTER 36. LECTURES 7 AND 8: LIKELIHOOD RATIO TEST AND LARGE SAMPLE Z TEST
Chapter 37

Lectures 9 and 10: Two-Population


Inference and Error Probabilities

Lecture 9: Inference for Two Population Proportions


Date: February 24, 2026 Topic: Large Sample Case (Z-Test)

1. Hypothesis Setup (Two-Sided Test)

This test is used to determine whether the difference between two population proportions (p1
and p2 ) is significantly different from a hypothesized value D0 .

H0 : p1 − p2 = D0
Ha : p1 − p2 ̸= D0

Note: In most practical cases, D0 = 0 (testing for no difference).

2. Population and Sample Information

Description Population 1 Population 2


Population Proportion p1 p2
Sample Size n1 n2
Number of Successes X1 X2
Sample Proportion p̂1 = X 1
n1 p̂2 = X 2
n2

Total Sample Size:


n = n1 + n2

3. Test Statistic (Z-Test)

Since this is a large sample case, the sampling distribution of (p̂1 − p̂2 ) is approximately
normal.

Case A: Testing for Equality (D0 = 0)


Use pooled proportion:
X1 + X2
p̄ = , q̄ = 1 − p̄
n1 + n2

231
232CHAPTER 37. LECTURES 9 AND 10: TWO-POPULATION INFERENCE AND ERROR PROBABILIT

(p̂1 − p̂2 )
Z=r  
p̄q̄ n11 + n12

Case B: Testing for Specific Difference (D0 ̸= 0)


Use unpooled standard error:
(p̂1 − p̂2 ) − D0
Z=q
p̂1 (1−p̂1 ) p̂2 (1−p̂2 )
n1 + n2

4. Decision Rule

For significance level α:

ˆ Reject H if |Z| > z


0 α/2

ˆ OR reject H if p-value < α


0

5. Key Assumptions (Large Sample Theory)

ˆ Success-Failure Condition:
n1 p̂1 ≥ 10, n1 (1 − p̂1 ) ≥ 10

n2 p̂2 ≥ 10, n2 (1 − p̂2 ) ≥ 10

ˆ Independence: Samples must be independent


ˆ Randomization: Data must be collected randomly
General Test Statistic for Two Proportions

When testing for a hypothesized difference D0 (where D0 ̸= 0), we use individual sample
proportions (unpooled case):

(p̂1 − p̂2 ) − D0
Z= r
p̂1 q̂1 p̂2 q̂2
+
n1 n2

Variable Definitions:

x1 , x2 : Number of successes in Sample 1 and Sample 2


n1 , n2 : Sample sizes
x1 x2
p̂1 = , p̂2 =
n1 n2
q̂1 = 1 − p̂1 , q̂2 = 1 − p̂2

Special Case: Testing for Equality (D0 = 0)


233

When H0 : p1 = p2 , we use the pooled proportion:


x1 + x2
p̂ = , q̂ = 1 − p̂
n1 + n2

Interpretation: Total successes divided by total observations.

Final Z-Statistic (Pooled Case):

p̂1 − p̂2
Z=s  
1 1
p̂q̂ +
n1 n2

Decision Criteria (Two-Sided Test)

ˆ Reject H if |Z| > z


0 α/2

ˆ Fail to reject H if |Z| ≤ z


0 α/2

Large Sample Assumptions

ˆ Success-Failure Condition:
n1 p̂1 ≥ 10, n1 q̂1 ≥ 10

n2 p̂2 ≥ 10, n2 q̂2 ≥ 10

ˆ Independence: Two samples must be independent


ˆ Random Sampling: Data must be collected randomly
234CHAPTER 37. LECTURES 9 AND 10: TWO-POPULATION INFERENCE AND ERROR PROBABILIT

Test Statistic under H0

(n − 1)s2
χ2 = ∼ χ2(n−1) d.o.f = (n − 1)
σ02

Note: The Chi-Square distribution is asymmetric (right-skewed) and defined only for
non-negative values.

One-Sided Tests

Area=α
Reject H0

χ2α
Right-Tailed Test

P (χ2 > χ2α ) = α

Reject H0 if χ2 > χ2α

Interpretation: Sample variance is significantly


greater.

Area=α

χ21−α χ2(1−α)
From Table
Left-Tailed Test

P (χ2 < χ21−α ) = α

Reject H0 if χ2 < χ21−α

Interpretation: Sample variance is significantly


smaller.

Two-Sided Test

Reject H0 if χ2 < χ21−α/2 or χ2 > χ2α/2


235

Ho Accept

Acceptance
Region
(1 − α)
Reject Area Reject H0
H0
χ21− α χ2α
2 2

Key Takeaways

ˆ Z and t distributions are symmetric


ˆ χ distribution is asymmetric (right-skewed)
2

ˆ Degrees of freedom = (n − 1)
ˆ Critical values must be looked up separately for each tail
236CHAPTER 37. LECTURES 9 AND 10: TWO-POPULATION INFERENCE AND ERROR PROBABILIT

Test for Ratio of Two Population Variances


 2
σ1
σ22

Step 1: Hypotheses

Two-Tailed Right-Tailed Left-Tailed

σ12 σ12 σ12


H0 : =1 H0 : =1 H0 : =1
σ22 σ22 σ22
σ12 σ12 σ12
Ha : ̸= 1 Ha : >1 Ha : <1
σ22 σ22 σ22

Step 2: Assumptions

ˆ Samples are independent


ˆ Both populations are normally distributed
Step 3: Test Statistic

S12
F =
S22

F ∼ F(n1 −1, n2 −1)

Degrees of Freedom:
df1 = n1 − 1 (numerator), df2 = n2 − 1 (denominator)

Important Convention:
Larger Variance
F = ≥1
Smaller Variance

Step 4: Decision Rules

Right-Tailed Test:
Reject H0 if F > Fα, (n1 −1, n2 −1)

Left-Tailed Test:
Reject H0 if F < F1−α, (n1 −1, n2 −1)

Two-Tailed Test:
Reject H0 if F < F1−α/2, (n1 −1, n2 −1) or F > Fα/2, (n1 −1, n2 −1)

Step 5: Key Properties


237

ˆ F ≥ 0 (always positive)
ˆ Right-skewed distribution
ˆ Depends on two degrees of freedom
ˆ Not symmetric (unlike Z and t)
Step 6: Mathematical Property

Reciprocal Property:
1
∼ F(n2 −1, n1 −1)
F

Step 7: Summary Table

Parameter Large Sample (n ≥ 30) Small Sample (n < 30)


Mean (µ) Z-test t-test
Variance (σ 2 ) Approx Z χ2 -test
Variance Ratio — F -test

Key Insight

The F -test is based on the Likelihood Ratio Test (LRT) and is the most powerful test
for comparing variances of normal populations.
238CHAPTER 37. LECTURES 9 AND 10: TWO-POPULATION INFERENCE AND ERROR PROBABILIT

Lecture–10: The p-value and Error Probabilities

Date: 26/02/2026

1. Understanding the p-value

Definition:
The p-value is the observed level of significance (α̂) computed from sample data.

p-value = P (Test statistic is at least as extreme as observed | H0 is true)

Interpretation:

ˆ Measures the evidence against H 0

ˆ Smaller p-value ⇒ stronger evidence against H 0

ˆ Represents the degree of disagreement between data and H 0

Decision Rule using p-value

(
p-value < α ⇒ Reject H0
p-value ≥ α ⇒ Fail to Reject H0

2. Type I and Type II Errors

α = P (Type I Error) = P (Reject H0 | H0 True)

β = P (Type II Error) = P (Fail to Reject H0 | Ha True)

Error Type Meaning Probability


Type I Error Reject H0 when true α
Type II Error Accept H0 when false β

Power of Test:
Power = 1 − β

3. Example: Finding α and β

Given:

ˆ X ∼ Exp(θ)
ˆ PDF: f (x) = θe −θx , x≥0

ˆ Hypotheses:
H0 : θ = 2, Ha : θ = 1
239

ˆ Critical Region: x ≥ 1
Step 1: Calculate α

α = P (X ≥ 1 | θ = 2)
Z ∞
= 2e−2x dx
1
∞
= −e−2x 1 = 0 − (−e−2 ) = e−2


∴ α = e−2 ≈ 0.1353

Step 2: Calculate β

β = P (X < 1 | θ = 1)

Z 1
= e−x dx
0
1
= −e−x 0 = (−e−1 ) − (−1) = 1 − e−1


∴ β = 1 − e−1 ≈ 0.6321

4. Key Takeaways

ˆ p-value is observed significance level


ˆ α is pre-fixed risk level
ˆ β measures missed detection
ˆ 1 − β is power of test
ˆ Decision: Compare p-value with α
Example: Calculation of α and β (Exponential Case)

Given:

ˆ X ∼ Exp(θ)
ˆ PDF:
f (x) = θe−θx , x≥0

ˆ Hypotheses:
H0 : θ = 2, Ha : θ = 1

ˆ Critical Region:
C = {x ≥ 1}
240CHAPTER 37. LECTURES 9 AND 10: TWO-POPULATION INFERENCE AND ERROR PROBABILIT

Step 1: Type I Error (α)

α = P (Reject H0 | H0 true)

= P (X ≥ 1 | θ = 2)

Z ∞
= 2e−2x dx
1

∞
= −e−2x 1 = 0 − (−e−2 ) = e−2


∴ α = e−2 ≈ 0.1353

Step 2: Type II Error (β)

β = P (Accept H0 | Ha true)

= P (X < 1 | θ = 1)

Z 1
= e−x dx
0

1
= −e−x 0 = (−e−1 ) − (−1) = 1 − e−1


∴ β = 1 − e−1 ≈ 0.6321

Step 3: Power of Test

Power = 1 − β = 1 − (1 − e−1 ) = e−1

∴ Power = e−1 ≈ 0.3679

Final Answer

α = e−2 , β = 1 − e−1 , Power = e−1


241

t-Test for Mean (Normal Population)


Test Statistic
If H0 is true, then:
X̄ − µ0
T = √ ∼ tn−1
S/ n
where
n
2 1 X
S = (Xi − X̄)2
n−1
i=1

(S 2 is an estimator of σ2)

Rejection Rule
Reject H0 if:
|T | > tα/2, n−1

Observed Test Statistic



n(X̄ − µ0 )
Tc =
S

p-value
p-value = P (|Tn−1 | > Tc )

Note
ˆ Normal distribution: N (0, 1)
ˆ t-distribution is symmetric
242CHAPTER 37. LECTURES 9 AND 10: TWO-POPULATION INFERENCE AND ERROR PROBABILIT

Power of Test
Power = (1 − β)
Example:
e−1 1
1− =
e e

Testing Hypothesis for Mean


H0 : µ = µ0
Ha : µ ̸= µ0

Case 1: σ Known (Small Sample)


Construct a two-sided test:

X̄ − µ0
Z= √ ∼ N (0, 1)
σ/ n
Reject H0 if:
|Z| > zα/2

Case 2: σ Unknown
Use t-test:
X̄ − µ0
T = √ ∼ tn−1
S/ n
Reject H0 if:
|T | > tα/2, n−1

Decision Rule (Two-sided Test)


ˆ Reject H in both tails of size α/2
0

ˆ Accept H in the middle region


0
Chapter 38

Lectures 11 and 12: Linear Algebra


Foundation - Subspaces and
Projections

38.1 Motivating Example: Why Do Some Systems Have


Infinite Solutions?
Consider the system:
x+y =1 2x + 2y = 2

This is the same equation scaled by 2, so the system has infinitely many solutions.

38.1.1 Matrix Form


Writing the system as A⃗x = ⃗b:
    
1 1 x 1
A⃗x = =
2 2 y 2

We can write A⃗x as a linear combination of columns:


     
1 1 1
x +y =
2 2 2

Notice both columns of A are identical, so:


   
1 1
(x + y) =
2 2

This gives x + y = 1, which has infinitely many solutions.

Key Insight: The infinite solutions arise from the fact that there are infinitely many ways
to express x + y = 1 in R.
To see this, look at the system:

x+y =1 and x − y = 0

The line x + y = 0 (the null space of A) induces infinite solutions in x + y = 1 by shifting.

243
244CHAPTER 38. LECTURES 11 AND 12: LINEAR ALGEBRA FOUNDATION - SUBSPACES AND PROJ

38.1.2 Geometric Interpretation


Given a basis {⃗c1 , ⃗c2 } for R2 , any vector ⃗b ∈ R2 can be expressed as:

x⃗c1 + y⃗c2 = ⃗b
   
1 1
For example, with ⃗c1 = and ⃗c2 = :
−1 −1
     
1 1 b
x +y = 1
−1 −1 b2

When ⃗b = 21 ⃗c1 + 21 ⃗c2 , we get:


 
1  1 b1
⃗c1 ⃗c2 =
2 2 b2

38.2 The Null Space


Definition (Null Space): The null space captures the linear dependencies (rela-
tionships) between the columns of a matrix A.
n o
N (A) = ⃗x ∈ Rn A⃗x = ⃗0m

For an m × n matrix A with columns ⃗c1 , ⃗c2 , . . . , ⃗cn and ⃗x ∈ Rn :

A⃗x = x1⃗c1 + x2⃗c2 + x3⃗c3 + x4⃗c4 + · · · = ⃗0m

A nonzero solution ⃗x means the columns are linearly dependent.

Connection to Infinite Solutions: The infinite solutions of x + y = 0 induce infinite


solutions for x + y = 1.
Precisely: if ⃗xh ∈ N (A) (a homogeneous solution) and ⃗xp is any particular solution to
A⃗x = ⃗b, then every solution has the form:

⃗x = ⃗xp + ⃗xh

38.3 The Four Fundamental Subspaces


For an m × n matrix A of rank r:

n o
C(A) = ⃗b ∈ Rm ⃗b = A⃗x (Column Space, m × 1)

R(A) = C(AT ) = ⃗z ∈ Rn ⃗z = ykT A



(Row Space, n × 1)
n o
N (A) = ⃗x ∈ Rn A⃗x = ⃗0m (Null Space, n × 1)
n o
N (AT ) = ⃗y ∈ Rm ⃗y T A = ⃗0n (Left Null Space, m × 1)

Intuitive descriptions:

ˆ C(A): Collection of all vectors that are linear combinations of the n columns of A
38.4. ORTHOGONALITY OF THE FUNDAMENTAL SUBSPACES 245

ˆ R(A): Collection of all vectors that are linear combinations of the n rows of A
ˆ N (A): Collection of all vectors that give linear dependencies among the n
columns of A

ˆ N (A T ):
Collection of all vectors that give linear dependencies among the n
columns of AT (i.e., rows of A)

38.3.1 Worked Example


   
1 1 T 1 2
Let A = , so A = , with:
2 2 1 2

⃗xT1 = (1, 1), ⃗xT2 = (2, 2) = 2⃗xT1

Since ⃗x2 = 2⃗x1 , we have 2⃗x1 − ⃗x2 = ⃗0, confirming a linear dependency.

C(A) = {(1, 1), (1, −1)}


T
C(A ) = R(A) = {(1, 1), (1, −1)}
N (A) = {⃗0} ⇒ unique solution
N (AT ) = {⃗0} ⇒ solution always exists
 
1
Inconsistency check: For A⃗x = :
4
   
 1 1
4 =x +y
2 2
 
1
Since ⃗b = ∈
/ C(A), this leads to an inconsistency: 1 × 2 ̸= 4.
4

38.3.2 Rank
rank(A) = r implies there are r independent rows and r independent columns.

Dimension summary:
Subspace Dimension Lives in
C(A) r Rm
R(A) = C(AT ) r Rn
N (A) n−r Rn
N (AT ) m−r Rm

38.4 Orthogonality of the Fundamental Subspaces


The diagram of the four subspaces shows:
A
Rn −
→ Rm

with N (A) ⊂ Rn (dimension n − r) and C(A) ⊂ Rm (dimension r).

38.4.1 Key Orthogonality Result


246CHAPTER 38. LECTURES 11 AND 12: LINEAR ALGEBRA FOUNDATION - SUBSPACES AND PROJ

If ⃗a ∈ R(A), then ⃗aT ⃗b = ⃗0m , i.e., R(A) ⊥ N (A).


Similarly, C(A) ⊥ N (AT ).

Proof sketch: Let ⃗y ∈ N (A) (i.e., A⃗y = ⃗0), and let ⃗xTi be the rows of A. Then:

y1 · ⃗xT1 ⃗b = 0
y2 · ⃗xT ⃗b = 0
2
..
.
yk · ⃗xTk ⃗b = 0

Summing:
y1 ⃗xT1 + y2 ⃗xT2 + y3 ⃗xT3 + · · · ⃗b = 0 ⃗aT ⃗b = 0

=⇒

38.5 Existence and Uniqueness of Solutions to A⃗x = ⃗b


Summary of Cases (for m × n matrix A of rank r):
Case (i): m = r, n = r (Square, Invertible)
N (AT ) = {⃗0m } and N (A) = {⃗0n } ⇒ Unique solution.
Case (ii): n > r, m = r (More unknowns than equations)
N (A) ̸= {⃗0} ⇒ ∞ solutions.
N (AT ) = {⃗0m } ⇒ Solution always exists.
Case (iii): n = r, m > r (More equations than unknowns)
N (AT ) = ∞ nonzero vectors ⇒ Solution may or may not exist.

a) If ⃗b ∈ C(A): solution exists and is unique.


b) If ⃗b ∈
/ C(A): no solution.

Case (iv): n > r, m > r


N (A) ⇒ ∞ solutions.
N (AT ) ⇒ may/may not have solution:

a) If ⃗b ∈ C(A): ∞ solutions.
b) If ⃗b ∈
/ C(A): no solution.

When do we use Regression? Cases 3b and 4b — when ⃗b ∈


/ C(A) and an exact solution
does not exist!

38.6 Projections
38.6.1 2D Setup
Goal: Project ⃗b onto ⃗a (i.e., find p⃗ ∈ span{⃗a} closest to ⃗b).
From the diagram: ⃗b = ⃗e + p⃗, so p⃗ = ⃗b − ⃗e.
Orthogonality condition: ⃗a ⊥ ⃗e, which means ⃗aT ⃗e = 0:

⃗aT ⃗b
⃗aT (⃗b − x⃗a) = 0 =⇒ x= (a scalar)
⃗aT ⃗a
38.7. PROPERTIES OF THE PROJECTION MATRIX 247

The projection:
" # 
⃗aT ⃗b ⃗a⃗aT ⃗

p⃗ = ⃗a · x = ⃗a T = T b
⃗a ⃗a ⃗a ⃗a

Projection Matrix (onto a vector ⃗a):

⃗a⃗aT
P =
⃗aT ⃗a

so that p⃗ = P⃗b.
Note the difference:

ˆ ⃗a · ⃗a is a scalar (1 × n times n × 1)
T

ˆ ⃗a · ⃗a is a matrix (n × 1 times 1 × n)
T

38.6.2 Numerical Example


 
2
Let ⃗a = :
3
 
2
⃗aT ⃗a = [2, 3] = 4 + 9 = 13 = ∥⃗a∥2
3
   
T 2 4 6
⃗a⃗a = [2, 3] =
3 6 9

38.7 Properties of the Projection Matrix

⃗a⃗aT
For P = :
⃗aT ⃗a
1. P is always a square matrix.

2. det(P ) = 0 and rank(P ) = 1.

3. P T = P (Symmetric).

4. C(P ) is a line through ⃗a.

5. P 2 = P (Idempotent / Projection Property).

Proof of (5):
Let p⃗ = P⃗b. Then:
P p⃗ = P (P⃗b) = P 2⃗b

Since projecting again does nothing (you’re already on the line):

P p⃗ = p⃗ = P⃗b =⇒ P 2⃗b = P⃗b =⇒ P2 = P

Long proof:
⃗a⃗aT ⃗a⃗aT ⃗a(⃗aT ⃗a)⃗aT ⃗a⃗aT
P2 = · = = =P ✓
⃗aT ⃗a ⃗aT ⃗a (⃗aT ⃗a)2 ⃗aT ⃗a
248CHAPTER 38. LECTURES 11 AND 12: LINEAR ALGEBRA FOUNDATION - SUBSPACES AND PROJ

38.7.1 Eigenvalues of P
From P ⃗x = λ⃗x, the two eigenvectors correspond to:

1. Perpendicular direction to ⃗a:


If ⃗b ⊥ ⃗a, then P⃗b = ⃗0 = 0 · ⃗b, so λ = 0.

2. Parallel direction to ⃗a:


If ⃗b ∥ ⃗a, then P⃗b = ⃗b, so λ = 1.

38.8 Projection onto a Column Space


38.8.1 Setup
Problem: We want to project ⃗b onto C(A), i.e., find p⃗ ∈ C(A) closest to ⃗b.
We use regression when A⃗x = ⃗b cannot be solved because ⃗b ∈/ C(A):

ˆ LHS: (A⃗x) ∈ C(A)


ˆ RHS: ⃗b ∈/ C(A) (due to errors/noise)
38.8.2 Derivation
Let p⃗ = A⃗xp be the projection. The error is ⃗e = ⃗b − p⃗, and ⃗e must belong to N (AT )
(perpendicular to C(A)):

AT ⃗e = ⃗0n =⇒ AT (⃗b − p⃗) = ⃗0n =⇒ AT (⃗b − A⃗xp ) = ⃗0n

Expanding:
AT ⃗b − AT A⃗xp = ⃗0 =⇒ AT A⃗xp = AT ⃗b

Normal Equations:
⃗xp = (AT A)−1 AT ⃗b

The projection onto C(A):

p⃗ = A⃗xp = A(AT A)−1 AT ⃗b = P⃗b

Projection Matrix onto C(A):

P = A(AT A)−1 AT

This has the same form as the 1D case but generalised to matrices.

38.8.3 Special Case: A is Square and Invertible


If A is n × n and invertible:

P = A(AT A)−1 AT = A · A−1 (AT )−1 · AT = I

So the projection matrix becomes the identity matrix — every vector is already in
C(A) = Rn .
38.9. LINEAR REGRESSION (FINALLY!) 249

38.9 Linear Regression (Finally!)


38.9.1 Setup
Fit a straight line b = c + dt through n data points (ti , bi ).
Example: 3 points (1, 1), (2, 2), (3, 3) — but the example in notes uses (1, 1), (2, 2), (3, 2).

38.9.2 Matrix Formulation


The model b = c + dt gives the overdetermined system A⃗x = ⃗b:
   
1 t1 b1
1 t2     b2 
c
A =  . .  , ⃗x = , ⃗b =  . 
   
.
. . . d  .. 
1 tn bn

Since we have more equations than unknowns (m > n) and ⃗b ∈


/ C(A) in general, we cannot
solve exactly. Instead, solve the normal equations:

AT A⃗xp = AT ⃗b
 

This gives the least squares solution ⃗xp = ˆ minimising ∥⃗b − A⃗x∥2 .
d

Summary: Regression is the projection of ⃗b onto the column space of A. The best-fit line
is b̂ = ĉ + dˆt, and the residual ⃗e = ⃗b − p⃗ is orthogonal to every column of A.

38.10 Quick Reference: Independence

Two vectors ⃗a and ⃗b are linearly independent if and only if:

x1⃗a + x2⃗b = ⃗0 ⇐⇒ x1 = x2 = 0

i.e., the only solution to x1⃗a + x2⃗b = ⃗0 is the trivial one.


250CHAPTER 38. LECTURES 11 AND 12: LINEAR ALGEBRA FOUNDATION - SUBSPACES AND PROJ
Chapter 39

Lectures 13 and 14: Projection and


Regression

24/03/26
Project : Project for human body (you)

ˆ 1 project in machine learning


ˆ 1 project in stochastic process
{where you can apply in machine learning and stochastic process} (1367)

→ Projection
2D
Given ⃗a and ⃗b

P⃗ =? (1368)

⃗b

⃗e = ⃗b − P⃗

⃗a
P⃗ ∈ C(A)

O
clp

⃗b ∈
/ C(A) (1369)

P⃗ ∈ C(A) (1370)

⃗a ∈ C(A) (1371)

251
252 CHAPTER 39. LECTURES 13 AND 14: PROJECTION AND REGRESSION

Here, column space of A is line along ⃗a.

x can be anything (1372)

P⃗ = x⃗a (1373)

⃗a ⊥ ⃗e (1374)

⃗aT ⃗e = 0 (1375)

You should think of ⃗a as a column space of a matrix.

Ax = ⃗b (1376)

(m × n)(n × 1) = (m × 1) (1377)

C(A) = collection of ∞ vectors got by taking linear combinations of n columns of A (1378)

= {⃗b | ⃗b = Ax} (1379)

R(A) = collection of ∞ vectors got by taking linear combinations of m rows of A (1380)

N (A) = collection of ∞ vectors that gives linear dependencies among n columns of A (1381)

N (AT ) = collection of ∞ vectors that gives linear dependencies among m rows of A (1382)

⃗aT ⃗e = 0 (1383)

⃗aT (⃗b − P⃗ ) = 0 (1384)

⃗aT (⃗b − x⃗a) = 0 (1385)

⃗aT ⃗b
x= (1386)
⃗aT ⃗a
" #

a T⃗
b
P⃗ = x⃗a = ⃗a T (1387)
⃗a ⃗a
 T
⃗a⃗a
P = T ⃗b
⃗ (1388)
⃗a ⃗a
253

⃗a⃗aT
 
P = T projection matrix (1389)
⃗a ⃗a

Take an example
 
2
⃗a = (1390)
3
 
 2
⃗aT ⃗a = 2 3 = 22 + 32 = 4 + 9 = 13 = |⃗a|2

(1391)
3
   
T 2   4 6
⃗a⃗a = 2 3 = (1392)
3 6 9

(2 × 1)(1 × 2) = 2 × 2 (1393)

→ Properties of P matrix

1. Square matrix

2. det(P ) = 0 ⇒ Rank(P ) = 1 (because 1 dependent column or 1 independent row)

3. P T = P i.e symmetric

4. C(P ) is line through ⃗a

C(P ) = {⃗b | ⃗b = P ⃗x} (1394)

5. Projection property: P 2 = P

⃗a⃗aT
 
1 4 6
P = T = (1395)
⃗a ⃗a 13 6 9
T
⃗a⃗aT (⃗a⃗aT )T

T
P = = (1396)
⃗aT ⃗a (⃗aT ⃗a)T

⃗a⃗aT
= =P (1397)
⃗aT ⃗a

P⃗ = P⃗b (1398)

P⃗b = P⃗ (1399)

P⃗b = P (P⃗b) = P P⃗b = P 2⃗b (1400)

P⃗ = P 2⃗b (1401)

⇒ P = P2 (1402)

OR
254 CHAPTER 39. LECTURES 13 AND 14: PROJECTION AND REGRESSION

⃗a⃗aT
P = (1403)
⃗aT ⃗a

⃗a⃗aT ⃗a⃗aT
  
2
P = PP = (1404)
⃗aT ⃗a ⃗aT ⃗a

⃗a(⃗aT ⃗a)⃗aT
= (1405)
(⃗aT ⃗a)(⃗aT ⃗a)
(scalar)

⃗a⃗aT
= =P (1406)
⃗aT ⃗a
→ What is the eigen values and eigen vectors of P ?

P ⃗x = λ⃗x (1407)
The direction does not change, is the beauty of eigen vector.
eg:
Case I
y

⃗b

x
⃗a

P⃗b = ⃗0 (1408)

= 0⃗b (1409)

λ=0 (1410)

⃗b ⊥ ⃗a (1411)
eigen vector / value for λ = 0
Case II

⃗b = P⃗b (1412)
255

P⃗b = 1⃗b (1413)

λ=1 (1414)

⃗b is along ⃗a (1415)

⃗a
⃗b

→ A⃗x = ⃗b

⃗b ∈
/ C(A) (1416)

(m × n)(n × 1) = (m × 1) (1417)
Project ⃗b onto C(A)
This is your model & data

A⃗x ∈ C(A) (1418)

⃗b ∈
/ C(A) (1419)
Data
You want to solve for A⃗x = ⃗b
You cannot solve for ⃗x
That’s why you have to go for regression

A⃗xp = p⃗ ∈ C(A) (1420)


solve ⃗xp

A = [⃗a1 ⃗a2 · · · ⃗an ] (1421)

⃗b ∈
/ C(A)

⃗e

C(A)
p⃗ ∈ C(A)

⃗e ⊥ C(A) ⇒ ⃗e ∈ N (AT ) (1422)

⃗eT p⃗ = 0 (1423)
⃗e is the error vector
256 CHAPTER 39. LECTURES 13 AND 14: PROJECTION AND REGRESSION

What is p⃗ now?
Given A, ⃗b

p⃗ = A⃗xp (1424)
⃗xp you have to calculate.
If ⃗e ∈ N (AT ), then what is the property of vector.

AT ⃗e = ⃗0n (1425)

(m × n)T (m × 1) = (n × 1) (1426)

AT (⃗b − p⃗) = ⃗0n (1427)

⇒ AT (⃗b − A⃗xp ) = ⃗0n (1428)

⇒ AT A⃗xp = AT ⃗b (1429)

⃗xp = (AT A)−1 AT ⃗b (1430)

So, now p⃗ is

p⃗ = A⃗xp = A(AT A)−1 AT ⃗b (1431)

P → projection matrix (1432)

p⃗ = P⃗b (1433)

Compare

P = ⃗a(⃗aT ⃗a)−1⃗aT (1434)

P = A(AT A)−1 AT (1435)

P = A(AT A)−1 AT (1436)

A (AT A)−1 |{z}


|{z} AT = |{z}
P (1437)
| {z }
m×n n×n n×m m×m

→ If A is (m × n) and invertible? (A−1 exist)


What is Projection Matrix P ?
Given that C(A) is Rn

C(A) = Rn (1438)
If take the vector which is not in column-space, it project that vector on column space.

P =I Identity matrix (1439)


257

All properties hold

P = PT, P2 = P (1440)

→ Example
3 points (1, 1), (2, 2), (3, 2)
(b, t) : Blood pressure

(3, 3)

(2, 2)
(3, 2)

(1, 1)

t (time (no. of hours you study))

b=t (1441)

b = C + Dt (1442)

Suppose the data was this instead


instead of (3, 2) it was (3, 3)

ˆ You can fit a straight line


ˆ Is not regression
ˆ You can simply solve the equation
1 = C + D(1) (1443)
2 = C + D(2) (1444)
3 = C + D(3) (1445)

C = 0, D=1 ⇒ b=t (1446)

When the points are not in straight line


we do regression

b ̸= t (1447)
inconsistent data

A⃗x = ⃗b (1448)

⃗b ∈
/ C(A) (1449)
258 CHAPTER 39. LECTURES 13 AND 14: PROJECTION AND REGRESSION

   
1 1   1
c ⃗b = 2
A = 1 2 , ⃗x = , (1450)
d
1 3 2

P = A(AT A)−1 AT (1451)

 
5 2 −1
1
P = 2 2 2 (3 × 3) (1452)
6
−1 2 5

A⃗xp = p⃗ (1453)

⃗xp =? (1454)

AT ⃗e = ⃗0 (1455)

AT (⃗b − p⃗) = AT (⃗b − A⃗xp ) = ⃗0 (1456)

AT ⃗b = AT A⃗xp (1457)

⃗xp = (AT A)−1 AT ⃗b (1458)

  2
c
⃗xp = = 31 (1459)
d 2

 
7
1
p⃗ = P⃗b = 10 (3 × 1) (1460)
6
13
 
−1
1
⃗e = ⃗b − p⃗ =  2  (1461)
6
−1

⃗eT p⃗ = ⃗e · p⃗ = 0 ⇒ ⃗e ⊥ p⃗ (1462)
Now solve

A⃗xp = p⃗ (1463)

   
1 1 7/6
1 2 ⃗xp = 10/6 (3 × 2)(2 × 1) = (3 × 1) (1464)
1 3 13/6
You should do by Row reduction method.
 
1 1 7/6
 1 2 10/6  (1465)
1 3 13/6
259

R3 → R3 − R1 (1466)

 
1 1 7/6
 1 2 10/6  (1467)
0 2 6/6

R2 → R2 − R1 (1468)

 
1 1 7/6
 0 1 3/6  (1469)
0 2 6/6

R3 → R3 − 2R2 (1470)

 
1 1 7/6
 0 1 3/6  (1471)
0 0 0

x + y = 7/6 (1472)

y = 3/6 = 1/2 (1473)

7 1
x= − (1474)
6 2

4 2
x= = (1475)
6 3
Normal eqn

AT A⃗xp = AT ⃗b (1476)

A⃗xp = p⃗ (1477)

A⃗xE = ⃗b × (1478)

(AT A)⃗xp = AT ⃗b (1479)

(2 × 3)(3 × 2)(2 × 1) = (2 × 1) (1480)

 
2 1
b = c + Dt = + t (1481)
3 2
260 CHAPTER 39. LECTURES 13 AND 14: PROJECTION AND REGRESSION

2.16
1.66 b3
p3
1.16 pb22
0.66 b1
p1

1 2 3

t bL
0 0.66 = 23
7
1 6 = 1.16 (1482)
5
2 3 = 1.66
13
3 6 = 2.16
 
−1
1 
⃗e = 2 = −1e1 + 2e2 − 1e3 (1483)
6
−1

⃗e = ⃗b − p⃗ (1484)

   
7/6 1.16
p⃗ = 10/6 = 1.67 (1485)
13/6 2.16
Chapter 40

Lectures 15 and 16: Fisher


Information, MLE, and Regression

40.1 Score Function and Fisher Information


40.1.1 Setup
Let X1 , X2 , . . . , Xn be i.i.d. random variables with common pdf f (x; θ), where θ ∈ Θ is an
unknown parameter.

Definition 1.1 — Log-Likelihood and Score

The log-likelihood for a single observation is:


n
X
ℓ(θ) = log f (xi ; θ)
i=1

The score function for a single observation is the derivative of the log-likelihood with
respect to θ:
∂ log f (x1 ; θ)
W =
∂θ

40.1.2 Properties of the Score Function


Theorem 40.1.1 (Mean of the Score). Under regularity conditions:
 
∂ log f (x; θ)
E[W ] = E =0 (1486)
∂θ
R∞
Proof. Since f (x; θ) is a valid pdf, −∞ f (x; θ) dx = 1. Differentiating both sides with respect
to θ: Z ∞ Z ∞
∂ ∂f (x; θ)
f (x; θ) dx = 0 =⇒ dx = 0.
∂θ −∞ −∞ ∂θ
Rewriting:
Z ∞  
1 ∂f (x; θ) ∂ log f
f (x; θ) dx = 0 =⇒ E = 0.
−∞ f (x; θ) ∂θ ∂θ
| {z }
= ∂ log f /∂θ

261
262CHAPTER 40. LECTURES 15 AND 16: FISHER INFORMATION, MLE, AND REGRESSION

Definition 1.2 — Fisher Information


The Fisher information in a single observation is:
2
I(θ) = Var(W ) = E W 2 − E[W ] = E W 2
    
(1487)

since E[W ] = 0. Equivalently:


" #
∂ log f (x; θ) 2
 
∂ log f (x; θ)
I(θ) = Var =E (1488)
∂θ ∂θ

40.1.3 Alternative Form of Fisher Information


Theorem 40.1.2 (Second-Derivative Form). Under regularity conditions:
 2 
∂ log f (x; θ)  2
I(θ) = −E = E W (1489)
∂θ2

Proof. Differentiating the identity E[∂ log f /∂θ] = 0 again:


   2   2 
∂ ∂ log f ∂ log f  2 ∂ log f
E = 0 =⇒ E + E W = 0 =⇒ I(θ) = −E .
∂θ ∂θ ∂θ2 ∂θ2

40.1.4 Fisher Information for a Sample of Size n


For n i.i.d. observations, define:
n
X ∂ log f (xi ; θ)
Z=
∂θ
i=1

Since the xi are i.i.d.:


n  
X ∂ log f (xi ; θ)
Var(Z) = Var = n I(θ) (1490)
∂θ
i=1

The total Fisher information in a sample of size n is n I(θ).

40.2 Cramér–Rao Lower Bound (CRLB)


40.2.1 Statement of the CRLB
Theorem 2.1 — Cramér–Rao Lower Bound
Let X1 , . . . , Xn be i.i.d. with pdf f (x; θ). Let Y = u(X1 , . . . , Xn ) be any unbiased
estimator of a function k(θ), with E[Y ] = k(θ). Then:

[k ′ (θ)]2
Var(Y ) ≥ (1491)
n I(θ)

where k ′ (θ) = dk/dθ and I(θ) is the Fisher information per observation.
40.3. UNBIASED ESTIMATORS AND EFFICIENCY 263

40.2.2 Proof of the CRLB


Proof. Let Z = ni=1 ∂ log f (xi ; θ)/∂θ be the total score (a random variable with E[Z] = 0).
P

Step 1. Compute k ′ (θ) = dE[Y ] /dθ:


Z Z n
∂ Y
k ′ (θ) = ··· u(x) f (xi ; θ) dx1 · · · dxn = E[Y · Z] = Cov(Y, Z)
∂θ
i=1

(using E[Z] = 0 so Cov(Y, Z) = E[Y Z] − E[Y ] E[Z] = E[Y Z]).


Step 2. By the Cauchy–Schwarz inequality:

[Cov(Y, Z)]2 ≤ Var(Y ) · Var(Z) =⇒ [k ′ (θ)]2 ≤ Var(Y ) · n I(θ).

Step 3. Rearranging gives (1491).

40.2.3 Correlation Coefficient Form


Define ρ = corr(Y, Z). Then:

Cov(Y, Z) k ′ (θ)
ρ= p =p (1492)
Var(Y ) · Var(Z) Var(Y ) · n I(θ)

Since ρ2 ≤ 1:
[k ′ (θ)]2 [k ′ (θ)]2
≤ 1 =⇒ Var(Y ) ≥ .
Var(Y ) · n I(θ) n I(θ)

40.2.4 Special Case: Estimating θ Itself


When k(θ) = θ so k ′ (θ) = 1, the CRLB becomes:
  1
Var θ̂ ≥ (1493)
n I(θ)

40.3 Unbiased Estimators and Efficiency


40.3.1 Unbiased Estimator of µ for Normal Distribution
iid
Let Xi ∼ N (µ, σ 2 ). Define the sample mean:

n
1X
θ̂ = X̄ = Xi (1494)
n
i=1

Proposition 40.3.1. X̄ is an unbiased estimator of µ:


h i  
E θ̂ = E X̄ = µ (1495)

Proposition 40.3.2 (Variance of the Sample Mean).

 nσ 2 σ2
Var X̄ = 2 = (1496)
n n
264CHAPTER 40. LECTURES 15 AND 16: FISHER INFORMATION, MLE, AND REGRESSION

40.3.2 CRLB for the Normal Mean


For N (µ, σ 2 ), the Fisher information is:
Step 1. Log-likelihood of a single observation:

1 (x − µ)2
ℓ(µ) = − log(2πσ 2 ) −
2 2σ 2

Step 2. First derivative (score):


x−µ
ℓ′ (µ) =
σ2
Step 3. Second derivative:
1
ℓ′′ (µ) = −
σ2
Step 4. Fisher information:
1
I(µ) = E −ℓ′′ (µ) = 2
 
σ
Step 5. CRLB:
1 σ2
CRLB = =
n I(µ) n

40.3.3 Efficiency of the Sample Mean

Key Result — Efficiency

Since Var X̄ = σ 2 /n = CRLB:




CRLB σ 2 /n
Efficiency of X̄ = = 2 =1 (1497)
Var X̄ σ /n

X̄ is the Best Unbiased Estimator (BLUE) for µ, since it achieves the CRLB with
efficiency = 1.

40.4 Maximum Likelihood Estimation (MLE)


40.4.1 Definition
Definition 4.1 — Maximum Likelihood Estimator

Given observations x1 , . . . , xn , the MLE θ̂MLE maximises the log-likelihood:


n
X
ℓ(θ) = log f (xi ; θ) (1498)
i=1

The MLE satisfies the score equations:


n
∂ℓ(θ) X ∂ log f (xi ; θ)
= =0 (1499)
∂θ θ=θ̂MLE ∂θ θ=θ̂MLE
i=1
40.5. PROBABILISTIC INTERPRETATION OF LINEAR REGRESSION 265

40.4.2 Asymptotic Property of MLE


Theorem 40.4.1 (Asymptotic Distribution of MLE). Under regularity conditions, as n → ∞:
 
d 1
θ̂MLE −→ N θ,
n I(θ)

The MLE asymptotically achieves the CRLB — it is asymptotically efficient.

40.5 Probabilistic Interpretation of Linear Regression


40.5.1 Setup
(j)
The training set consists of m examples: {(⃗x(j) , y (j) )}m x(j) ∈ Rn+1 (with x0 = 1)
j=1 , where ⃗
and y (j) ∈ R.
The parameter vector is:  
θ0
 θ1 
θ⃗ =  .  ∈ Rn+1
 
 .. 
θn
The hypothesis is the linear model:

hθ⃗ (⃗x) = θ⃗⊤ ⃗x = θ0 x0 + θ1 x1 + · · · + θn xn (1500)

40.5.2 Error Model


Assume the j-th training output satisfies:
iid
y (j) = θ⃗⊤ ⃗x(j) + ε(j) , ε(j) ∼ N (0, σ 2 ) (1501)

This is the homoskedasticity assumption: constant error variance.

40.5.3 Distribution of the Model


Given the error model:  
(j) (j) ⃗ ⃗⊤ (j) 2
y | ⃗x ; θ ∼ N θ ⃗x , σ (1502)

with:
h i
E y (j) | ⃗x(j) ; θ⃗ = θ⃗⊤ ⃗x(j) (1503)
 
Var y (j) | ⃗x(j) ; θ⃗ = σ 2 (1504)

40.5.4 PDF of the Model


The conditional pdf is:
1 (j) 2 2
fε (ε(j) ) = √ e−(ε ) /(2σ ) (1505)
σ 2π
Since ε(j) = y (j) − θ⃗⊤ ⃗x(j) :

⃗ = √1 e−[y(j) −θ⃗⊤ ⃗x(j) ]2 /(2σ2 )


f (y (j) | ⃗x(j) ; θ) (1506)
σ 2π
266CHAPTER 40. LECTURES 15 AND 16: FISHER INFORMATION, MLE, AND REGRESSION

40.5.5 MLE for Linear Regression


Assuming the m training examples are independent, the likelihood is:
m  m
Y 1 Pm ⃗⊤ ⃗
(j) −θ x(j) ]2 /(2σ 2 )
⃗ =
L(θ) f (y (j)
| ⃗x (j) ⃗ =
; θ) √ e− j=1 [y (1507)
j=1
σ 2π

The log-likelihood is:


m
⃗ ⃗ 1 X (j) ⃗⊤ (j) 2
ℓ(θ) = log L(θ) = log β − 2 y − θ ⃗x (1508)

j=1

where β collects the constant prefactor.

Key Result — MLE equals Least Squares

⃗ ⇐⇒ θ⃗ˆMLE = arg min J(θ)


max ℓ(θ) ⃗ (1509)
θ⃗ θ⃗

where the cost function is:


m
⃗ = 1 X 2
J(θ) hθ⃗ (⃗x(j) ) − y (j) (1510)
2
j=1

Maximising the log-likelihood is equivalent to minimising the sum of squared


errors (SSE).

40.6 Least Squares and the Normal Equation


40.6.1 The Cost Function
The i-th error (residual) for training example i:

e(i) = hθ⃗ (⃗x(i) ) − y (i)

The cost function over all m training examples:


m m
⃗ = J(θ0 , θ1 , . . . , θn ) = 1
X 2 1 X ⊤ (j) 2
J(θ) hθ⃗ (⃗x(j) ) − y (j) = θ⃗ ⃗x − y (j) (1511)
2 2
j=1 j=1

40.6.2 Gradient of the Cost


The gradient vector is:  
∂J/∂θ0
∂J(θ)⃗  ∂J/∂θ1 
=  .  ∈ R(n+1)×1 (1512)
 
∂ θ⃗  . .
∂J/∂θn
Each partial derivative:

⃗ m m
∂J(θ) X  (j) X  (j)
= θ⃗⊤ ⃗x(j) − y (j) xi = hθ⃗ (⃗x(j) ) − y (j) xi (1513)
∂θi
j=1 j=1
40.7. BATCH GRADIENT DESCENT FOR LINEAR REGRESSION 267

40.6.3 Normal Equation (Closed-Form Solution)


Setting the gradient to zero gives the Normal Equation:

A⊤ A ⃗xp = A⊤⃗b (1514)

where A is the design matrix, ⃗b is the output vector, and ⃗xp is the least-squares solution.

40.6.4 Numerical Example


Example 40.6.1 (5-mark question from notes). Given data points (1, 1), (2, 2), (3, 3) and
model b = 23 + 21 t:
Predicted values:
2
p1 = 3 + 12 (1) = 7
6 ≈ 1.167
2
p2 = 3 + 12 (2) = 5
3 ≈ 1.667
2
p3 = 3 + 12 (3) = 13
6 ≈ 2.167

SSE to minimise:

SSE[C, D] = (C + D − 1)2 + (C + 2D − 2)2 + (C + 3D − 2)2

First-order conditions:
∂ SSE
=0: 2(C + D − 1) · 1 + 2(C + 2D − 2) · 1 + 2(C + 3D − 2) · 1 = 0
∂C
which gives the system:
    
3 6 C 5
= ⇐⇒ A⊤ A ⃗xp = A⊤⃗b
6 14 D 11

40.7 Batch Gradient Descent for Linear Regression


40.7.1 Gradient Descent Update Rule
⃗ update each parameter simultaneously:
To minimise J(θ),

(k+1) (k) ∂J(θ⃗(k) )


θi ←− θi −α (1515)
∂θi
where α > 0 is the learning rate.

40.7.2 Expanded Batch Gradient Descent


Substituting (1513):
m
(k+1) (k)  (j)
X
hθ⃗ (⃗x(j) ) − y (j) xi

θi ←− θi −α (1516)
j=1

Stopping criterion: Repeat until


∂J ⃗
= 0n+1 (gradient converges to zero)
∂ θ⃗

40.7.3 Stochastic Gradient Descent (SGD)


Instead of summing over all m examples, update point by point:
268CHAPTER 40. LECTURES 15 AND 16: FISHER INFORMATION, MLE, AND REGRESSION

Algorithm — Stochastic Gradient Descent

Repeat {
for j = 1 to m:
for all i:
(k+1) (k)  (j)
←− θi − α hθ⃗ (⃗x(j) ) − y (j) xi

θi
}

SGD is faster and often better than batch GD because it updates parameters after each
training example rather than after a full pass over the data.

40.8 Logistic Regression — Binary Classification


40.8.1 Problem Setup
Setup — Binary Classification

ˆ Output: Y ∈ {0, 1} (discrete, two classes)


ˆ Since y ∈ [0, 1], the hypothesis must satisfy 0 ≤ h (⃗x) ≤ 1 θ⃗

ˆ Feature vector: ⃗x = (x , x , . . . , x ) with x = 1


0 1 n

0

ˆ Parameter vector: θ⃗ = (θ , θ , . . . , θ )
0 1 n

40.8.2 Sigmoid (Logistic) Function


Choose the hypothesis:
1
hθ⃗ (⃗x) = g(θ⃗⊤ ⃗x) = (1517)
1 + e−θ⃗⊤ ⃗x
1
where z = θ⃗⊤ ⃗x and g(z) = is the sigmoid (logistic) function.
1 + e−z
Key properties:
1 1
g(0) = 0
= = 0.5 (1518)
1+e 2
g(z) → 1 as z → +∞ (1519)
g(z) → 0 as z → −∞ (1520)
e−z 1 e−z
g ′ (z) = = · = g(z) (1 − g(z)) (1521)
(1 + e−z )2 1 + e−z 1 + e−z

1 1
g(z) = 1+e−z
g(z)

0.5

0
−6 −4 −2 0 2 4 6
z = θ⃗⊤ ⃗x
40.8. LOGISTIC REGRESSION — BINARY CLASSIFICATION 269

40.8.3 Probabilistic Interpretation


Assume:

⃗ = h⃗ (⃗x)
P [Y = 1 | ⃗x; θ] (1522)
θ
⃗ = 1 − h⃗ (⃗x)
P [Y = 0 | ⃗x; θ] (1523)
θ

So:
Y | ⃗x; θ⃗ ∼ Bernoulli hθ⃗ (⃗x)
  

40.8.4 Likelihood Function for Logistic Regression


Assume the training data {(⃗x(j) , y (j) )}m
j=1 are independent.
The likelihood function is:
m m
Y  Y y(j)  1−y(j)
⃗ = P Y (j) = y (j) | ⃗x(j) ; θ⃗ = hθ⃗ (⃗x(j) ) 1 − hθ⃗ (⃗x(j) )
 
L(θ) (1524)
j=1 j=1

The log-likelihood:
m n
X o
⃗ = log L(θ)
⃗ = y (j) log hθ⃗ (⃗x(j) ) + (1 − y (j) ) log 1 − hθ⃗ (⃗x(j) )
  
ℓ(θ) (1525)
j=1

40.8.5 Gradient of the Log-Likelihood


Theorem 8.1 — Gradient of Log-Likelihood

⃗ m
∂ℓ(θ) X  (j)
= y (j) − hθ⃗ (⃗x(j) ) xi (1526)
∂θi
j=1

Proof. For a single observation, let h = hθ⃗ (⃗x) = g(z), z = θ⃗⊤ ⃗x.
Step 1. Differentiate ℓ:
" #
∂ℓ X y (j) ∂h 1 − y (j) ∂h
= ·
(j) ) ∂θ
− ·
(j) ) ∂θ
∂θi
j
h ⃗
θ
(⃗
x i 1 − hθ⃗ (⃗
x i

(j)
Step 2. Use ∂g/∂θi = g(z)(1 − g(z)) xi :

∂h (j)
= hθ⃗ (⃗x(j) ) [1 − hθ⃗ (⃗x(j) )] xi
∂θi
Step 3. Substituting and simplifying:
" #
∂ℓ X y (j) (j) 1 − y (j) (j)
= · h(1 − h) xi − · h(1 − h) xi
∂θi h 1−h
j
X  (j)
= y (j) (1 − h) − (1 − y (j) )h xi
j
X  (j)
= y (j) − hθ⃗ (⃗x(j) ) xi
j
270CHAPTER 40. LECTURES 15 AND 16: FISHER INFORMATION, MLE, AND REGRESSION

40.8.6 Gradient Ascent for Logistic Regression


⃗ (gradient ascent):
To maximise ℓ(θ)

(k+1) (k) ∂ℓ(θ⃗(k) )


θi ←− θi +α (1527)
∂θi

Substituting (1526):
m
(k+1) (k)  (j)
X  (j)
θi ←− θi +α y − hθ⃗ (⃗x(j) ) xi (1528)
j=1

Comparison: Linear vs Logistic Update Rule

Linear Regression (minimise J):


(k+1) (k)  (j)
X
θi ← θi −α hθ⃗ (⃗x(j) ) − y (j) xi
j

Logistic Regression (maximise ℓ):


(k+1) (k)  (j)
X
θi ← θi +α y (j) − hθ⃗ (⃗x(j) ) xi
j

The update rules have the same form! The only differences are:

ˆ Sign: −α (gradient descent) vs +α (gradient ascent)


ˆ h (⃗x): linear function vs sigmoid of linear function
θ⃗

40.8.7 Stochastic and Batch Gradient Descent for Logistic Regression

Stochastic Gradient Descent (update point by point):

(k+1) (k)  (j)


− α hθ⃗ (⃗x(j) ) − y (j) xi

θi ←− θi (1529)

1
where hθ⃗ (⃗x(j) ) = (sigmoid)
e−θ⃗⊤ ⃗x
(j)
1+

Batch Gradient Descent (sum over all m examples):

m
(k+1) (k)  (j)
X
hθ⃗ (⃗x(j) ) − y (j) xi

θi ←− θi −α (1530)
j=1

For linear regression (same form but hθ⃗ (⃗x(j) ) = θ⃗⊤ ⃗x(j) ):

m
(k+1) (k)  (j)
X
θ⃗⊤ ⃗x(j) − y (j) xi

θi ←− θi −α (1531)
j=1
40.9. SUMMARY OF KEY FORMULAE 271

40.8.8 Stochastic Gradient Descent Algorithm

Algorithm — SGD for Logistic Regression (faster & better)

Repeat {
for j = 1 to m {
(k+1) (k)  (j)
− α hθ⃗ (⃗x(j) ) − y (j) xi

for all i: θi ←− θi
}
}

40.9 Summary of Key Formulae

Master Formula Sheet


Score function: W = ∂ log f (x; θ)/∂θ, E[W ] = 0
Fisher information: I(θ) = E W 2 = −E ∂ 2 log f /∂θ2
   

  [k ′ (θ)]2
CRLB: Var θ̂ ≥
n I(θ)
CRLB
Efficiency:   ≤1
eff(θ̂) =
Var θ̂
P
MLE score equation: i ∂ log f (xi ; θ)/∂θ = 0
m
⃗ = 1 X ⃗⊤ (j)
Linear regression cost: J(θ) [θ ⃗x − y (j) ]2
2
j=1

(j) ) (j)
− y (j) ]xi
P
Gradient: ∂J/∂θi = j [hθ⃗ (⃗
x
1
Sigmoid: g(z) = , g ′ (z) = g(z)[1 − g(z)]
1 + e−z
Logistic log-likelihood: ℓ = j {y (j) log h + (1 − y (j) ) log(1 − h)}
P

(j)
Logistic gradient: ∂ℓ/∂θi = j [y (j) − hθ⃗ (⃗x(j) )]xi
P

(j)
Batch GD (linear): θi ← θi − α j [h(⃗x(j) ) − y (j) ]xi
P

(j)
SGD (logistic): θi ← θi − α[hθ⃗ (⃗x(j) ) − y (j) ]xi
Normal equation: A⊤ A ⃗xp = A⊤⃗b
272CHAPTER 40. LECTURES 15 AND 16: FISHER INFORMATION, MLE, AND REGRESSION

Notation Reference

Symbol Meaning

θ, θ⃗ Scalar / vector parameter


θ̂MLE Maximum likelihood estimator
I(θ) Fisher information per observation
CRLB Cramér–Rao lower bound for variance
W Score function ∂ log f /∂θ

ℓ(θ) Log-likelihood

J(θ) Cost function (SSE) for regression
hθ⃗ (⃗x) Hypothesis function
g(z) Sigmoid (logistic) function
α Learning rate
m Number of training examples (j index)
n Number of features (i index for θi )
⃗x(j) Feature vector of j-th training example
y (j) Label of j-th training example
ε(j) Error/noise term for j-th example
Chapter 41

Lectures 17 and 18: Markov Chains


and Naive Bayes Classification

41.1 Introduction to Stochastic Processes


Definition 1.1: Stochastic Process
Let Xn be a random variable at discrete time n, where n = 0, 1, 2, . . . . The notation Xn = i
means the process is in state i at time n. The general conditional distribution defining the
process is P (Xn | Xn−1 , Xn−2 , . . . , X1 ), meaning the current state generally depends on all
past states.
Definition 1.2: Markov Chain (Markov Property)
A stochastic process {Xn } is called a Markov Chain if it satisfies the Markov Property, which
states that the future depends only on the present, NOT on the past.

P (Xn+1 = j | Xn = i, Xn−1 , . . . , X0 ) = P (Xn+1 = j | Xn = i) (1532)

41.2 Transition Matrices


Definition 1.3: Transition Probability Matrix
The one-step transition matrix P has entries pij = P (Xn+1 = j | Xn = i). Because the process
must transition to some state, each row sums to 1:
X
pij = 1 for all i (1533)
j

Example: A Two-State Chain


For a chain with states {0, 1}:  
0.5 0.5
P = (1534)
0.5 0.5

41.2.1 Weather Prediction Model


Suppose it rains today and will rain tomorrow with probability α. If it does not rain today, it
will rain tomorrow with probability β. Let Xn denote the weather on the n-th day:
ˆ 0: It rains on day n
ˆ 1: It does not rain on day n
Transition Probabilities:
ˆp 00 = P (Xn+1 = 0 | Xn = 0) = α

273
274CHAPTER 41. LECTURES 17 AND 18: MARKOV CHAINS AND NAIVE BAYES CLASSIFICATION

ˆp 01 = P (Xn+1 = 1 | Xn = 0) = 1 − α
ˆp 10 = P (Xn+1 = 0 | Xn = 1) = β
ˆp 11 = P (Xn+1 = 1 | Xn = 1) = 1 − β
Transition Matrix:  
α 1−α
P = (1535)
β 1−β

41.2.2 Non-Markovian Counter-example


Consider a process where Xn captures rain on both the (n − 1)-th and n-th day
simultaneously. The states are:
ˆ 0: Rained on both (n − 1)-th and n-th day
ˆ 1: Rained on n-th day but not on (n − 1)-th day
ˆ 2: Rained on (n − 1)-th day but not on n-th day
ˆ 3: Did not rain on either day
The 4 × 4 transition matrix recorded for this expanded state space:
 
0.7 0 0.3 0
0.5 0 0.5 0 
P = 0
 (1536)
0.4 0 0.6
0 0.2 0 0.8

41.2.3 Two-Step Transition Matrix


Question: Given that it rained on Monday and Tuesday, what is the probability it will rain
on Thursday?
Since Monday is day n, Tuesday is day n + 1, and Thursday is day n + 3, we need the two-step
transition matrix P (2) = P 2 . Starting from state 0 (rained on both consecutive days), the
two-step matrix is:  
0.49 0.12 0.21 0.18
0.35 0.20 0.15 0.30
P (2) = P 2 = 
0.20 0.12 0.20 0.48
 (1537)
0.10 0.16 0.10 0.64
Answer: Given it rained on Monday and Tuesday, we start in state 0. The probability of
(2)
raining on Thursday (two steps later, state 0) is P00 = 0.49. We can verify this via matrix
multiplication (row 0 of P 2 ):
X
(P 2 )00 = P0k Pk0 = (0.7 × 0.7) + (0 × 0.5) + (0.3 × 0) + (0 × 0) = 0.49 (1538)
k

41.3 Gambler’s Ruin Problem


41.3.1 Problem Statement
A gambler plays a sequence of games. Let Xn be the gambler’s wealth after the n-th game.
He quits playing when either:
ˆ He is broke (wealth = 0)
ˆ He attains a fortune of $N (wealth = N )
The state space is {0, 1, 2, . . . , N }.
41.3. GAMBLER’S RUIN PROBLEM 275

41.3.2 Transition Probabilities


At each game (for interior states i = 1, 2, . . . , N − 1):

ˆp i,i+1 = p (win: wealth increases by 1)

ˆp i,i−1 = 1 − p (lose: wealth decreases by 1)

States 0 and N are absorbing states:

ˆp 00 =1

ˆp NN =1

Transition Matrix:  
1 0 0 ··· 0 0
1 − p 0 p ··· 0 0
 
P = 0
 1−p 0 ··· 0 0  (1539)
 .. .. .. . . .. .. 
 . . . . . .
0 0 0 ··· 0 1
276CHAPTER 41. LECTURES 17 AND 18: MARKOV CHAINS AND NAIVE BAYES CLASSIFICATION
Chapter 42

Naive Bayes Classifier

42.1 Introduction and Setup


The Naive Bayes classifier is a generative probabilistic model used for classification. We focus
on email spam detection as the running example.

42.1.1 Notation
ˆ Y ∈ {0, 1}: email label (0 = good/ham, 1 = spam).
ˆ A fixed dictionary (vocabulary) of n words: V = {w , w , . . . , w }.
1 2 n

ˆ X⃗ = (X , X , . . . , X ) : a random vector representing email j.


(j)
1
(j)
2
(j) ⊤
n

ˆ Each feature: X = 1 if word i is present in email j, and 0 if absent.


(j)
i

Example: Vocabulary V = {cat, dog, sat, ran}. Email1 = ”dog ran”:

⃗x(1) = (0, 1, 0, 1)⊤ (1540)

42.2 The Classification Goal


⃗ = ⃗x). Given the word-vector of an
We want to compute the posterior probability P (Y = y | X
email, what is the probability it belongs to class y?

42.2.1 Bayes’ Theorem for Classification


⃗ = ⃗x | Y = y)P (Y = y)
P (X
⃗ = ⃗x) =
P (Y = y | X (1541)
P (X⃗ = ⃗x)

ˆ P (Y = y | X⃗ = ⃗x) is the Posterior


ˆ P (X⃗ = ⃗x | Y = y) is the Likelihood
ˆ P (Y = y) is the Prior
ˆ P (X⃗ = ⃗x) is the Normalization (evidence)
The denominator is the same for all classes and acts merely as normalization.

277
278 CHAPTER 42. NAIVE BAYES CLASSIFIER

42.3 The Generative Model


Definition 4.1: Generative Process
The Naive Bayes model assumes:
1. The class label Y is drawn from a prior: P (Y = y).
2. Given Y = y, the feature vector X⃗ is generated independently (The Naive
Assumption):
n
Y
⃗ = ⃗x | Y = y) =
P (X P (Xi = xi | Y = y) (1542)
i=1

42.3.1 ⃗ and Tractability


Distribution of X
Since each Xi ∈ {0, 1}, the vector ⃗x takes values in {0, 1}n . There are 2n possible email
vectors. If we modeled the full joint distribution, the total number of parameters required
would be:
2(2n − 1) + 1 (1543)
For a standard vocabulary of n = 40, 000 words, the number of parameters is ≈ 240000 , which
is completely intractable.
By applying the conditional independence assumption of Naive Bayes, the number of
parameters reduces drastically to:
2n + 1 (1544)
For n = 40, 000, this requires only 80, 001 parameters, making the model highly efficient and
tractable.

42.4 Worked Example: Spam Classifier


42.4.1 Dataset Setup
Vocabulary: V = {money, free, hello} (n = 3). There are 23 = 8 possible email vectors.

Email money free hello Label


e0 0 0 0
e1 1 0 0
e2 0 1 0
e3 1 1 0
e4 0 0 1
e5 0 1 1
e6 1 1 1
e7 1 1 1

42.4.2 Training Data


The training set contains 5 spam emails:

42.4.3 Class-Conditional Likelihoods


Based on the training data, the likelihoods are computed:
For a new email ⃗x∗ , the classifier assigns the label that maximizes the posterior:
⃗ = ⃗x∗ ) = arg max P (X
ŷ = arg max P (Y = y | X ⃗ = ⃗x∗ | Y = y)P (Y = y) (1545)
y∈{0,1} y
42.5. QUICK REFERENCE NOTATION TABLE 279

Email money free hello Label


g1 0 0 1 spam
g2 0 0 1 spam
g3 0 0 0 spam
g4 0 1 1 spam
g5 0 0 1 spam

Email vector ei P (ei | good) P (ei | spam)


e0 1/5 0
e1 0 1/5
e2 3/5 0
e3 0 2/5
e4 1/5 0
e5 0 0

42.5 Quick Reference Notation Table

Symbol Meaning
Xn Random variable (state) at time n
pij Transition probability from state i to state j
P Transition probability matrix (rows sum to 1)
P (k) = P k k-step transition matrix
N Target fortune in Gambler’s Ruin
p Win probability per game in Gambler’s Ruin
Y Email class label (0 = ham, 1 = spam)
⃗ (j)
X Feature vector (word indicator) for email j
(j)
Xi 1 if word wi present in email j; 0 otherwise
n Vocabulary size (number of words)
V Vocabulary set {w1 , . . . , wn }
P (Y = y) Prior probability of class y
P (X⃗ = ⃗x | Y = y) Likelihood of email given class y
280 CHAPTER 42. NAIVE BAYES CLASSIFIER
Chapter 43

Lectures 19 and 20: Asymptotic


Dynamics of Markov Chains

43.1 Introduction to Finite State Markov Chains


A finite state Markov chain is a stochastic process that transitions among a finite collection of
states, where each transition depends exclusively on the current state and not on the history
of states previously visited—a property known as the Markov property.
One of the most important, and often initially surprising, structural features of such chains is
that not all states need be of the same character. In particular, not every state in a finite state
Markov chain is necessarily transient; the state space may house a rich taxonomy of state
types, each with its own long-run behavior. Understanding this taxonomy is essential to
characterizing the asymptotic dynamics of the chain, and it begins with a careful study of
accessibility, communication, and recurrence.

43.2 Accessibility and Communication of States


Before classifying individual states, it is useful to establish when one state can be “reached”
from another. A state j is said to be accessible from state i if there exists some n > 0 for
which:
(n)
Pij > 0 (1546)
(n)
where Pij denotes the n-step transition probability of moving from state i to state j in
exactly n steps. In other words, starting from i, the process has a positive probability of
visiting j after a finite number of transitions.
Building on this notion of accessibility, two states i and j are said to communicate—written
i ↔ j—if each is accessible from the other:
i→j and j → i (1547)
Communication is an equivalence relation, and it partitions the state space into disjoint
communicating classes. This partitioning has far-reaching consequences: states within the
same communicating class share the same recurrence or transience properties, and the
long-run behavior of the chain can often be decomposed along these class boundaries.

43.3 Recurrent and Transient States


With a notion of accessibility in hand, it becomes meaningful to ask: once a chain leaves a
given state, will it necessarily return? This question leads to the fundamental dichotomy
between recurrent and transient states.

281
282CHAPTER 43. LECTURES 19 AND 20: ASYMPTOTIC DYNAMICS OF MARKOV CHAINS

43.3.1 Recurrent States


A state i is called recurrent if, starting from i, the process returns to i with probability one.
Formally, letting fi denote this return probability, a state is recurrent if and only if:

fi = 1 (1548)

An equivalent characterization, and one that connects the recurrence condition to the spectral
properties of the transition matrix, is given by the divergence of the series of n-step return
probabilities:

(n)
X
Pii = ∞ (1549)
n=1

The equivalence of these two conditions is a classical result in the theory of Markov chains.
P (n)
The divergence of n Pii signifies not merely that the process returns to state i, but that it
returns infinitely often with probability one—a much stronger statement about the chain’s
long-run behavior.

43.3.2 Transient States


In contrast, a state i is transient if there is a positive probability of never returning, that is:

fi < 1 (1550)

and equivalently:

(n)
X
Pii <∞ (1551)
n=1

The convergence of this sum reflects the fact that the process can only visit state i a finite
expected number of times. As time progresses, the probability of finding the chain in a
transient state diminishes to zero: the chain eventually “escapes” such states permanently.
This behavior stands in stark contrast to recurrent states, to which the chain perpetually
returns.

43.4 A Finer Classification: Positive and Null Recurrence


Recurrence alone does not tell the full story. While all recurrent states are guaranteed a
return, the expected time to that return may be finite or infinite. This distinction gives rise to
a further classification.

43.4.1 Positive Recurrent States


A state i is positive recurrent if the expected return time E[Ti ]—the expected number of
steps required to first revisit state i—is finite:

E[Ti ] < ∞ (1552)

Positive recurrence is the “well-behaved” form of recurrence. States with this property not
only guarantee a return, but do so on a timescale that is, on average, bounded. In finite state
Markov chains, every recurrent state is in fact positive recurrent, making this distinction most
relevant in chains with infinite state spaces.
43.5. PERIODICITY AND APERIODICITY 283

43.4.2 Null Recurrent States


A state is null recurrent if it is recurrent yet the expected return time is infinite:

E[Ti ] = ∞ (1553)

Null recurrent states represent an edge case in which the chain is guaranteed to return, but
the average waiting time grows without bound. Such states arise naturally in random walks
on infinite lattices, but they do not occur in finite state Markov chains.

43.5 Periodicity and Aperiodicity


Having distinguished states by their recurrence properties, we now turn to another
fundamental characteristic: the period of a state, which captures the temporal regularity with
which returns can occur.
(n)
A state i is said to have period d ∈ N if Pii = 0 whenever n is not divisible by d, and d is
(n)
the greatest such integer—that is, d = gcd{n ≥ 1 : Pii > 0}. Intuitively, if state i has period
d > 1, then the chain can only return to i at times that are multiples of d, imposing a cyclic
structure on its dynamics. This periodic behavior can complicate the analysis of long-run
convergence, since the chain’s distribution at time n depends critically on n mod d.

43.5.1 Aperiodic States


A state is said to be aperiodic if its period is d = 1. For an aperiodic state, there is no such
cyclic constraint: returns may occur at any sufficiently large time, and the chain does not
oscillate in a structured pattern. Aperiodicity is a prerequisite for the existence of a
well-defined limiting distribution.

43.6 Ergodic States and Ergodic Markov Chains


The concepts of positive recurrence and aperiodicity now combine to define the most
important class of states in the theory.
A state is said to be ergodic if it is simultaneously positive recurrent and aperiodic. An
ergodic state is, in a precise sense, the “best-behaved” type: it is guaranteed to be visited
infinitely often, the average return time is finite, and there are no cyclical regularities that
would prevent convergence of the transition probabilities.
This terminology naturally extends to Markov chains: a Markov chain is called ergodic if all
of its states are ergodic—that is, if the chain is irreducible (all states communicate), positive
recurrent, and aperiodic. Ergodicity is precisely the condition that guarantees the existence
and uniqueness of a stationary distribution and the convergence of transition probabilities to
that distribution from any initial state.

43.7 Irreducibility
A Markov chain is said to be irreducible if all of its states belong to a single communicating
class:
All states belong to a single communicating class. (1554)
Equivalently, every state is accessible from every other state. Irreducibility is a global
structural property: it prevents the chain from becoming “trapped” in a subset of the state
space. In an irreducible chain, properties such as recurrence, transience, and periodicity are
shared by all states, because they are class properties.
284CHAPTER 43. LECTURES 19 AND 20: ASYMPTOTIC DYNAMICS OF MARKOV CHAINS

To illustrate, consider a Markov chain with state space {0, 1, 2, 3} and transition matrix:

0 12 1

0 2
1 0 0 0
P =
0
 (1555)
1 0 0
0 1 0 0

One may verify that all four states communicate with one another—each is accessible from
every other via some finite sequence of transitions—and hence all states {0, 1, 2, 3} belong to a
single communicating class. The chain is therefore irreducible.

43.8 Limiting Probabilities and the Limiting Distribution


Having established the classification of states, we analyze the long-run behavior of the
(n)
transition probabilities. Define the n-step transition probability Pij as the probability of
being in state j after n steps, given that the process started in state i. A central question in
Markov chain theory is whether the limit:

(n)
lim Pij (1556)
n→∞

exists, and if so, whether it depends on the initial state i.


For an irreducible and ergodic Markov chain, a fundamental theorem guarantees that this
limit exists, is strictly positive, and is independent of the initial state i:

(n)
lim Pij = πj (1557)
n→∞

The independence from the initial state i is the hallmark of ergodic behavior. No matter
where the chain begins, after a long time it “forgets” its origins and settles into a distribution
determined solely by the structure of the transition matrix. The collection {πj } defines the
limiting distribution of the chain.

43.9 Stationary Distribution


Closely related to the limiting distribution is the concept of the stationary distribution. A
probability distribution π = (π0 , π1 , π2 , . . .) is said to be stationary for the Markov chain if it
satisfies the balance equations:
X
πj = πi Pij (1558)
i

together with the normalization condition:


X
πi = 1, πi ≥ 0 for all i (1559)
i

The balance equations can be written compactly in matrix form as π = πP , where π is treated
as a row vector. The interpretation is immediate: if the distribution of the chain at time n is
π, then its distribution at time n + 1 is also π. A chain started in its stationary distribution
remains there forever.
For an ergodic chain, the stationary distribution is unique, strictly positive at every state, and
coincides with the limiting distribution.
43.10. MEAN RECURRENCE TIME 285

43.10 Mean Recurrence Time


The mean recurrence time for state j, denoted mjj , is defined as the expected number of
steps required to return to state j, given that the chain starts in state j. This quantity is
precisely the expected first-passage time back to the originating state. A remarkable result
connects mjj directly to the stationary distribution:
1
mjj = (1560)
πj
This relationship reveals that the proportion of time the chain spends in state j in the long
run—captured by πj —is exactly the reciprocal of the average time between successive visits to
that state.

43.11 A Worked Example: Three-State Ergodic Chain


To ground the foregoing theory in a concrete computation, consider the 3 × 3 transition matrix:
 
0.2 0.5 0.3
P = 0.4 0.1 0.5 (1561)
0.3 0.6 0.1
One may verify that every pair of states communicates (irreducibility), each state has period 1
(aperiodicity), and all states are positive recurrent. The chain is therefore ergodic.
To find the stationary distribution, we solve the balance equations π = πP :
π0 = 0.2π0 + 0.4π1 + 0.3π2 (1562)
π1 = 0.5π0 + 0.1π1 + 0.6π2 (1563)
π2 = 0.3π0 + 0.5π1 + 0.1π2 (1564)
subject to the normalization constraint π0 + π1 + π2 = 1. Solving this system yields:
π0 = 0.307, π1 = 0.379, π2 = 0.313 (1565)
These values tell us that in the long run, the chain occupies state 1 most
frequently—approximately 37.9% of the time. Correspondingly, the mean recurrence times are
m00 ≈ 3.26, m11 ≈ 2.64, and m22 ≈ 3.19 steps.

43.12 The Two-State Markov Chain


As a canonical and analytically tractable special case, consider a two-state chain with state
space {0, 1} and transition matrix:
 
α 1−α
P = (1566)
β 1−β
where 0 ≤ α, β ≤ 1. The balance equations π = πP yield:
π0 = απ0 + βπ1 , π1 = (1 − α)π0 + (1 − β)π1 (1567)
with π0 + π1 = 1. Solving this system gives the closed-form stationary distribution:
β 1−α
π0 = , π1 = (1568)
1−α+β 1−α+β
This result is elegantly interpretable: π0 is governed by the rate of flow into state 0
(parameterized by β), relative to the total rate of transition between states (1 − α + β). The
mean recurrence times follow immediately: m00 = (1 − α + β)/β and
m11 = (1 − α + β)/(1 − α).
286CHAPTER 43. LECTURES 19 AND 20: ASYMPTOTIC DYNAMICS OF MARKOV CHAINS

43.13 Conclusion
The theory of Markov chains presented here traces a coherent path from the structural notion
of accessibility and communication, through the classification of states, to the dynamical
condition of ergodicity that underpins long-run convergence. The stationary distribution,
characterized by the balance equations π = πP , serves as the centerpiece of this theory: it is
simultaneously the unique invariant measure of the chain, the long-run frequency distribution
of state visits, and the reciprocal of the mean recurrence times.
43.13. CONCLUSION 287

partStochasatic Processes
288CHAPTER 43. LECTURES 19 AND 20: ASYMPTOTIC DYNAMICS OF MARKOV CHAINS
Chapter 44

Lectures 1 and 2: Introduction to


Stochastic Processes

44.1 Lecture 1: Foundations of Stochastic Processes


44.1.1 Definition
A Stochastic Process (SP) is an infinite collection of Random Variables:

X0 , X1 , X2 , . . . , Xn−1 , Xn , Xn+1 , . . . , X∞ (1569)

where:

ˆ X : Initial state at time t = 0.


0

ˆ X : Random Variable representing the process at time n.


n

44.1.2 Random Variable (RV)


An RV is a mapping from the sample space to the set of real numbers:

X:Ω→R (1570)
generates
sample space −−−−−→ Real number (1571)

”An RV is a bridge between numbers and things which are not numbers
(outcomes).”

44.1.3 Randomness
The randomness of the process at any time n comes from the outcome ω:

Xn (ω) where ω ∈ Ω (1572)

44.1.4 Example: Toss Coin Twice


Sample space Ω2 consists of 4 possible outcomes:

Ω = {H1 H2 , H1 T2 , T1 H2 , T1 T2 } (1573)
| {z } | {z } | {z } | {z }
ω1 ω2 ω3 ω4

Where P = P (H) and (1 − P ) = P (T ).

289
290CHAPTER 44. LECTURES 1 AND 2: INTRODUCTION TO STOCHASTIC PROCESSES

X CTCS Stochastic process

ω2 → Realization ω2

ω1

sample paths of CTCS SPX


ω3

t
0

44.1.5 Defining Different RVs on the Same Sample Space


Consider two different random variables, X (# of Heads) and Y (# of Tails):

ˆ X(ω ) = 2,
1 Y (ω1 ) = 0

ˆ X(ω ) = 1,
2 Y (ω2 ) = 1

ˆ X(ω ) = 1,
3 Y (ω3 ) = 1

ˆ X(ω ) = 0,
4 Y (ω4 ) = 2

Note: X(ω) ̸= Y (ω) for all ω ∈ Ω2 , even though they share the same Ω.

44.1.6 Classification of Stochastic Processes


Stochastic processes are classified by their Time Index (T ) and State Space (S):

States \Time Discrete Time Continuous Time


Discrete State DTDS: Discrete Time Dis- CTDS: Continuous Time Discrete
crete State (e.g., Random State (e.g., Poisson Process)
Walk)
Continuous State DTCS: Discrete Time Con- CTCS: Continuous Time Continu-
tinuous State ous State (e.g., Brownian Motion)

ˆ Discrete Time: n ∈ {0, 1, 2, . . . }


ˆ Continuous Time: t ∈ [0, ∞)
ˆ State Space (S): The set of all possible values (outcomes) of the process.
44.1.7 Sample Paths and Trajectories
A Realization or Sample Path is the sequence of values the process takes for a specific ω.

ˆ In Continuous Time, this is a curve X(t).


ˆ In Discrete Time, this is a sequence of points (n, X ). n
44.1. LECTURE 1: FOUNDATIONS OF STOCHASTIC PROCESSES 291

Discrete Time Discrete State (DTDS)

Let Sn (ω) be the price of a stock at time n ∈ N and outcome ω ∈ Ω:

ˆ t = 0, S 0 = 12 (Initial Info)

ˆ t = 1, S 1 ∈ {6, 24} (Partial Info)

ˆ t = 2, S 2 ∈ {3, 12, 48} (Full Info)

States
of SP (y)
48 ⋆ (Path or Trajectory 1) (ω1 )

24 ⋆

12⋆

6
3 (ω4 ) ⇒ Trajectory
0 1 2 Time (x)

Continuous Time Discrete State (CTDS)

Often modeled as a Continuous Time Markov Process (CTMP). A classic example is a


Counting Process (Discrete state, continuous time).
Discrete state of SP

0
9 AM 9:23 9:33 9:35 Continuous Time
one Realization of SP.

44.1.8 Filtrations and Paths

Let In represent the information available up to time n.

ˆ I : Known initial state.


0

ˆ I [X = x , X = x , X = x , X
3 0 0 1 1 2 2 3 = x3 ] represents the observed partial path of the
stochastic process.
292CHAPTER 44. LECTURES 1 AND 2: INTRODUCTION TO STOCHASTIC PROCESSES

ω1

ω2

ω3

ω4
Fixed X0 ω5

0 1 2 3 4 5
x1 x2 x3
observe

44.1.9 The Power Set and Realizations


The set of all possible events is the power set 2Ω2 = P(Ω2 ), containing 24 = 16 elements. For a
value x in the state space Sx :

ExX = [X = x] = {ω ∈ Ω2 | X(ω) = x} (1574)

Example:
E1 = [X = 1] = {ω2 } ∪ {ω3 } = {ω2 , ω3 } (1575)

44.2 Lecture 2: Probability and State Spaces


44.2.1 Stochastic Process Formal Definition
A Stochastic Process is defined as a mapping from the product of the time index set and the
sample space to the state space:
X : N × Ω → SY (1576)

where SY is the State Space of the process. The value of the process at time n is:

Yn = Yn (ω) ≡ Yn (n, ω) (1577)

ˆ If ω is fixed: Y (n, ω) represents a Time Series.


ˆ If n is fixed: Y (ω) is a simple Random Variable.
n

44.2.2 Probability Axioms


A probability function P must satisfy:

1. P : F → [0, 1]

2. P (Ω) = 1
P
3. For mutually exclusive events E1 , E2 , . . . , P (∪Ei ) = P (Ei ).
44.2. LECTURE 2: PROBABILITY AND STATE SPACES 293

44.2.3 Example 1: Discrete Sample Space


1
Let Ω = {1, 2, 3, . . . , n, . . . } with probability P (n) = 2n . Find the probability that the
outcome is even:
P (even) = P (2) + P (4) + P (6) + . . . (1578)
1 1 1
= 2 + 4 + 6 + ... (1579)
2 2 2
a
Using the geometric series sum formula S = 1−r :
1 1 1/4 1
a= , r= =⇒ P (even) = = ≈ 0.3333 (1580)
4 4 1 − 1/4 3

44.2.4 Example 2: Continuous Random Numbers


Pick two random numbers (x, y) uniformly in [0, 1] × [0, 1].
Find P [(x, y) = (0.5, 0.3)] = 0 (1581)
Since (x, y) is a continuous RV, the probability of any single discrete point value in a
continuous plane is exactly zero.
Find: P (x + y ≤ 1/2) for x, y ∈ [0, 1]. The valid area is a triangle with base = 1/2 and
height = 1/2.   
1 1 1 1 1
Area of ∆ = × base × height = = (1582)
2 2 2 2 8
y
1
Area of △ = 8
1

1
2

x
0 1 1
2

44.2.5 Conditional Probability


Given a sample space Ω with events A and B:
P (B|B) = 1 (because Ω is reduced to B) (1583)
P (A ∩ B)
P (A|B) = (1584)
P (B)

Radar Application Example


Let A = Plane is present and B = Blip in Radar. Given P (A) = 0.05, we want to find P (A|B)
(probability the plane is actually there given a blip).

A B


294CHAPTER 44. LECTURES 1 AND 2: INTRODUCTION TO STOCHASTIC PROCESSES

Using Bayes’ Theorem:


P (B|A)P (A)
P (A|B) = (1585)
P (B)
Given P (B|A) = 0.99 (True Positive), P (A) = 0.05, and assuming P (B) = 0.95:

P (A ∩ B) = 0.99 × 0.05 = 0.0495 (1586)


0.0495
P (A|B) = ≈ 0.052 (1587)
0.95
Chapter 45

Lectures 3 and 4: Bernoulli


Processes and Arrival Times

45.1 Introduction to the Bernoulli Process


A Bernoulli process is defined by a sequence of independent and identically distributed
(i.i.d.) trials. In this specific application, we are modeling the arrival of customers (or jobs)
within a discrete time framework.

45.2 System Modeling and Discretization


To analyze the arrivals, we define a specific ”window of time” and divide it into n smaller,
equal-sized units called time slots.

45.2.1 Parameters

ˆ n: The total number of i.i.d. time slots within the window.


ˆ p: The probability of a ”success” (an arrival) in any single given time slot, represented
as p = P [customer arrives in a given slot].

45.2.2 Assumptions

The model assumes that in each individual slot, either 1 or 0 customers arrive. This binary
outcome is analogous to tossing a coin for each slot, where an arrival is a ”head” and no
arrival is a ”tail.”

45.3 The Random Variable and Distribution


We define the random variable S as the total number of customers that arrived in the window
of n slots. Because the arrivals are independent and the probability p is constant for each slot,
the variable S follows a Binomial Distribution:

S ∼ Bin(n, p)

295
296CHAPTER 45. LECTURES 3 AND 4: BERNOULLI PROCESSES AND ARRIVAL TIMES

45.4 Mathematical Formulations


45.4.1 Probability Mass Function (PMF)
The probability that exactly k customers have arrived in the window is calculated using the
Binomial PMF. This formula accounts for the probability of k arrivals and n − k non-arrivals,
multiplied by the number of possible ways these arrivals can be ordered.
 
n k
P [S = k] = p (1 − p)(n−k) (1588)
k
Note: Equation 1588 is valid for k = 0, 1, . . . , n.

45.4.2 Statistical Moments


To understand the average behavior and the reliability of the system, we calculate the
Expected Value (Mean) and the Variance.

Expected Value

The expected value represents the average number of customers we expect to see in a window
of size n.
E[S] = np (1589)

Variance

The variance measures the spread or uncertainty regarding the number of arrivals around the
mean.
Var[S] = np(1 − p) (1590)

45.5 Summary of Results


The transition from a Bernoulli process to a Binomial distribution allows us to predict system
load over a fixed period. By using Equation 1588, we can determine the likelihood of specific
traffic volumes, while Equations 1589 and 1590 provide the fundamental metrics for system
performance and capacity planning.

45.6 Waiting Time for a Job


In a Bernoulli process, a common question is: For a given number of jobs, how much time did
it take for the jobs to arrive? This shifts our focus from the number of successes in a fixed
window to the time (or number of trials) required to achieve a success.

45.6.1 Defining the Random Variable T1


Let T1 be the random variable representing the number of trials until the 1st success (arrival
of a job) is observed.

ˆ p = P (Job arrival/Success)
ˆ 1 − p = P (No job/Failure)
45.7. THE GEOMETRIC DISTRIBUTION 297

45.7 The Geometric Distribution


If we observe the 1st success at trial t, it implies that the previous t − 1 trials were all failures.
Due to the independent and identically distributed (i.i.d.) nature of the trials, the probability
of this sequence is the product of their individual probabilities.

45.7.1 Probability Mass Function (PMF)


The probability that the first success occurs exactly on trial t is given by:

P [T1 = t] = (1 − p)t−1 p for t = 1, 2, 3, . . . (1591)

Because T1 follows this pattern, we say:

T1 ∼ geom(p)

45.7.2 Visualizing the Process


If T1 = 5, the sequence of events is:

(1 − p), (1 − p), (1 − p), (1 − p), p


| {z } |{z}
4 failures 1 success

45.8 Statistical Properties


The moments of the geometric distribution provide the average waiting time and the variance
(uncertainty) of that wait.

1
E[T1 ] = (1592)
p

1−p
Var(T1 ) = (1593)
p2
Memoryless Property: A critical consequence of the independence of trials is the
memoryless property, meaning the probability of a success in the next trial does not depend
on how many failures have already occurred.

45.9 Example: Lottery Ticket String of Losses


We consider a scenario of buying a lottery ticket every day and analyze the length of the first
string of losing days.

45.9.1 Scenario Analysis


Suppose we look at a ”string of losing days” L between two successes (wins).

ˆ If L is the number of failures between wins, then L can be 0, 1, 2, . . .


ˆ However, for T (the trial of the first success), the values must start at t = 1.
1
298CHAPTER 45. LECTURES 3 AND 4: BERNOULLI PROCESSES AND ARRIVAL TIMES

45.9.2 The Geometric Constraint


If we assume (L + 1) ∼ geom(p), then:

(L + 1) ∈ {1, 2, 3, . . . } =⇒ L ∈ {0, 1, 2, . . . }

In your analysis, you noted a distinction regarding the independence of these strings.
Specifically, if the starting point of the string is not independent of the last success, the
standard geometric assumptions may not apply directly to L + 1 in certain contexts.

45.10 Introduction to Total Waiting Time


Previously, we analyzed T1 , the time until the first success. We now extend this to Wk , which
represents the total waiting time for k successes (or k heads) in a sequence of i.i.d.
Bernoulli trials.

45.11 The Concept of the Losing Streak


Before deriving Wk , we consider the ”losing streak” L. In your notes, you highlight that a
string of failures occurs between successes.

ˆ L represents the length of the losing streak.


ˆ If L = 4, there are 4 consecutive zeros (0, 0, 0, 0) followed by a success (1).
ˆ Note on Independence: The ”last success” and the ”first failure” of a new streak are
independent trials in a Bernoulli process, but the streak itself is defined by the boundary
of these successes.

45.12 Derivation of the Negative Binomial Distribution


Suppose we want to find the probability that the k th arrival (success) occurs exactly at time t.

45.12.1 Logical Breakdown


For the k th success to happen exactly at time t, two conditions must be met:

1. Exactly (k − 1) successes must have occurred in the previous (t − 1) trials. This part
follows a Binomial Distribution: Bin(t − 1, p).

2. The trial at time t must be a success (probability p).

45.12.2 Probability Mass Function (PMF)


By multiplying the probability of (k − 1) successes in (t − 1) trials by the probability of a
success on the tth trial, we get the PMF for Wk :
 
t − 1 k−1
P [Wk = t] = p (1 − p)(t−1)−(k−1) · p (1594)
k−1
Simplifying the exponents (t − 1 − k + 1 = t − k) and combining the p terms:
 
t−1 k
P [Wk = t] = p (1 − p)t−k (1595)
k−1
Where:
45.13. WAITING TIME AS A SUM OF GEOMETRIC VARIABLES 299

ˆ t = k, k + 1, k + 2, . . . (You cannot have k successes in fewer than k trials).


ˆ p is the probability of a head (H); q = 1 − p is the probability of a tail (T ).
45.13 Waiting Time as a Sum of Geometric Variables
An intuitive way to view Wk is as the sum of independent waiting times for each individual
success. If T1 is the time for the first arrival, T2 is the time elapsed between the first and
second arrival, and so on:

Wk = T1 + T2 + · · · + Tk (1596)

Since each Ti ∼ geom(p), the total waiting time Wk is simply the sum of k i.i.d. Geometric
random variables.

ˆW 1 = T1 ∼ geom(p) when k = 1.

45.14 Total Waiting Time for k Arrivals


Having previously defined Ti as the inter-arrival time (the time between the (i − 1)-th and i-th
arrival), we now consider the total time required to observe k successes.

45.14.1 The Random Variable Wk


Let Wk be the total waiting time for k arrivals, where k is a fixed constant:

k
X
Wk = Ti (1597)
i=1

Since each Ti ∼ geom(p), Wk represents the sum of k i.i.d. geometric random variables.

45.14.2 Expectation and Variance


Using the properties of linearity of expectation and the variance of independent variables:

ˆ Expected Value:
k k
" #
X X k
E[Wk ] = E Ti = E[Ti ] = (1598)
p
i=1 i=1

ˆ Variance: " k # k
X X k(1 − p)
Var[Wk ] = Var Ti = Var(Ti ) = (1599)
p2
i=1 i=1

45.15 Splitting a Bernoulli Process


Splitting occurs when an original arrival process is divided into two or more sub-processes
based on a secondary probability.
300CHAPTER 45. LECTURES 3 AND 4: BERNOULLI PROCESSES AND ARRIVAL TIMES

45.15.1 The Mechanism


An original arrival process with probability p is split using a ”decision coin” with probability q.
ˆ Process A (Success): Occurs if there is an arrival (p) AND the split decision is
”High” (q).
ˆ Process B (Idle/Success B): Occurs if there is an arrival (p) AND the split decision
is ”Low” (1 − q).

45.15.2 Visual Representation of Splitting


The following diagram illustrates how the original arrival process (Server A) is partitioned
into two distinct streams (Server R and Server B) through independent i.i.d. coin tosses.

1 1

Server R (pq) 0 0 0 0 0 0

q 1 1 1 1

Original Process (p) 0 0 0 0


1 1
1−q

Server B (p(1 − q)) 0 0 0 0 0 0

Time

45.15.3 Sub-Process Probabilities


Each sub-process remains a Bernoulli process with modified probabilities:
ˆX R: Success with probability P (XR = 1) = pq.
ˆX B: Success with probability P (XB = 1) = p(1 − q).
The splitting occurs by tossing two independent coins, where each coin is i.i.d.

45.16 Merging of Bernoulli Processes


Merging occurs when we combine two separate, independent Bernoulli streams into a single
aggregate process. This is often used to model total arrivals from different sources (e.g., Men
and Women, or two different job types).

45.17 Tree Diagram Approach: Conditional Probabilities


The first part of the notes uses a decision tree to model admission and specialization.
ˆ Let p be the probability of Admission (Coin A).
ˆ Let q be the probability of selecting Finance given Admission (Coin I).
The joint probabilities for the branches are calculated as follows:
1. Admission and Finance: p · q
2. Admission and Business Analytics (BA): p · (1 − q)
45.18. MERGING INDEPENDENT STREAMS 301

45.18 Merging Independent Streams


When we merge two independent processes—Process M (Men) and Process W (Women)—we
are interested in the probability that at least one event occurs in a given time slot.

45.18.1 Individual Process Definitions


ˆ Process M: X M = 1 with probability p; 0 with probability 1 − p.

ˆ Process W: X W = 1 with probability q; 0 with probability 1 − q.

45.18.2 The Merged Probability (pM W )


The merged process XM W results in a ’1’ if there is an arrival from either Men, Women, or
both. Mathematically, the probability of a success in the merged process is defined as:

pM W = P [XM W = 1] = pq + p(1 − q) + (1 − p)q (1600)

Explanation of terms in Equation 1600:

ˆ pq: Both Men and Women arrive.


ˆ p(1 − q): Men arrive, but Women do not.
ˆ (1 − p)q: Women arrive, but Men do not.
45.18.3 Simplified Merged Equation
By simplifying the expression in Equation 1600, we arrive at the standard formula for the
union of two independent events:
pM W = p + q − pq (1601)
Equation 1601 represents the probability of the merged Bernoulli process.

45.19 Visual Interpretation


In the provided timeline diagrams, each vertical pulse represents a success (1) in a time slot.

ˆ Timeline 1 (Men): Shows arrivals with frequency p.


ˆ Timeline 2 (Women): Shows arrivals with frequency q.
ˆ Timeline 3 (Merged): Shows a pulse if either Timeline 1 or Timeline 2 has a pulse.
45.20 Conclusion
The merging of two independent Bernoulli processes with probabilities p and q results in a
new Bernoulli process with a success probability of p + q − pq. This demonstrates that the
merged process remains a Bernoulli process, provided the original streams are independent.

45.21 Alternative Derivation of Merged Probability


In the previous section, we summed the individual success cases. Here, we use the
Complement Rule to simplify the derivation of the merged probability pM W .
302CHAPTER 45. LECTURES 3 AND 4: BERNOULLI PROCESSES AND ARRIVAL TIMES

45.21.1 Algebraic Proof


The probability of at least one arrival in the merged process is the complement of the
probability that neither process has an arrival.

pM W = P [arrival in merged process]


= 1 − P [no arrival in merged process]
= 1 − [(1 − p)(1 − q)] (No men AND No women)
= 1 − [1 − q − p + pq]
= 1 − 1 + q + p − pq

pM W = p + q − pq (1602)
Equation 1602 confirms the previous result using a more efficient probabilistic approach.

45.22 Continuous Time: The Exponential Distribution


When we move from discrete trials to continuous time, the waiting time T for an event to
occur is modeled by the Exponential Distribution.

45.22.1 Probability Density Function (PDF)


For a random variable T ∼ Exp(λ) where λ > 0, the PDF is defined as:
(
λe−λt if t ≥ 0
fT (t) = (1603)
0 if t < 0

Key Concepts for Equation 1603:

ˆ λ (Rate): Represents the average number of events per unit time. It has units of . 1
Time

ˆ f (t): This is the density, not a probability. The probability is found by integrating this
T
function over an interval.

45.22.2 Cumulative Distribution Function (CDF)


The CDF represents the probability that the waiting time T is less than or equal to a specific
value t.
FT (t) = P [T ≤ t] = (1 − e−λt ) (1604)

45.23 Summary of Waiting Times


While the Geometric distribution models the number of discrete trials until a success, the
Exponential distribution models the actual duration of time until an event occurs. In
continuous time, T represents the ”waiting time for an event to occur,” assuming the events
occur continuously and independently at a constant average rate λ.
Chapter 46

Lectures 5 and 6: Geometric and


Exponential Distributions

46.1 Lecture 5: Sums of Geometric Random Variables


Let T1 , T2 , . . . be independent and identically distributed (i.i.d.) random variables following a
Geometric distribution:
Ti ∼ Geom(p) i.i.d. (1605)
We define the sum Wk as the time until the k-th success:
k
X
Wk = Ti (1606)
i=1

46.1.1 Mean and Variance


Using the linearity of expectation and the properties of independent variables:

ˆ Expectation: " k
# k k
X X X 1 k
E[Wk ] = E Ti = E[Ti ] = = (1607)
p p
i=1 i=1 i=1

ˆ Variance: " k # k
X X k(1 − p)
Var(Wk ) = Var Ti = Var(Ti ) = (1608)
p2
i=1 i=1

kq
Note: If we denote q = 1 − p, the variance can be written concisely as p2
.

46.2 Bernoulli Processes: Splitting and Merging


46.2.1 1. Splitting a Bernoulli Process
Consider an original arrival process A following a Bernoulli distribution with parameter p.
This process is split into two streams, R and B, based on a second independent Bernoulli trial
(like a coin flip) with parameter q.

ˆ Stream R (Server R): X R = 1 with probability pq.

ˆ Stream B (Server B): X B = 1 with probability p(1 − q).

303
304CHAPTER 46. LECTURES 5 AND 6: GEOMETRIC AND EXPONENTIAL DISTRIBUTIONS

Server R: pq
split prob q

Bern(p) Split

split prob 1 − q
Server B: p(1 − q)

46.2.2 2. Merging Bernoulli Processes


When two independent Bernoulli processes, Xm ∼ Bern(p) and Xw ∼ Bern(q), are merged, the
new process Xmw represents an arrival if either or both original processes have an arrival in a
given time slot.
The probability of an arrival in the merged process (Pmw ) is:

Pmw = P (Arrival in merged process)


= 1 − P (No arrival in either process)
= 1 − [(1 − p)(1 − q)]
= 1 − [1 − q − p + pq]
= p + q − pq (1609)

Bern(p)

Bern(q)

Merged: p + q − pq

46.3 The Exponential Distribution


For a continuous random variable T representing waiting time, we say T ∼ Exp(λ) where
λ > 0.

46.3.1 1. Probability Density Function (PDF)


(
λe−λt t≥0
fT (t) = (1610)
0 t<0
1
The parameter λ is the rate and has units of time .

46.3.2 2. Cumulative Distribution Function (CDF)


The CDF represents the probability that the event occurs by time t:

FT (t) = P (T ≤ t) = 1 − e−λt (1611)

Conversely, the probability that the event takes longer than t is:

P (T > t) = 1 − FT (t) = e−λt (1612)


46.4. LECTURE 6: FURTHER PROPERTIES OF THE EXPONENTIAL DISTRIBUTION305

46.3.3 3. Expectation (Mean)


The expected waiting time is the reciprocal of the rate:
1 1
E[T ] = =⇒ λ = (1613)
λ E[T ]

Intuition: λ and E[T ] are inversely related.

ˆ If λ is large, E[T ] is small (high frequency of events, short waits).


ˆ If λ is small, E[T ] is large (low frequency of events, long waits).
fT (t)

f (t) = λe−λt

46.4 Lecture 6: Further Properties of the Exponential


Distribution
For a random variable T ∼ Exp(λ), the PDF is f (t) = λe−λt for t ≥ 0.

46.4.1 Moments and Variance


The second moment is calculated via integration by parts:
Z ∞
2
E[T 2 ] = t2 λe−λt dt = 2 (1614)
0 λ

The variance is derived as:


 2
2 1 1
Var(T ) = E[T 2 ] − (E[T ])2 = − = 2 (1615)
λ2 λ λ

The Standard Deviation is exactly equal to the mean:


p 1
SD(T ) = Var(T ) = = E[T ] (1616)
λ

46.4.2 The Memoryless Property


The defining characteristic of the Exponential distribution is that the remaining time until an
event occurs does not depend on how much time has already passed:

P (T > t + ∆ | T > t) = P (T > ∆) (1617)


306CHAPTER 46. LECTURES 5 AND 6: GEOMETRIC AND EXPONENTIAL DISTRIBUTIONS

46.5 Competing Risks: Time to First Event


Assume two independent processes with rates λ1 and λ2 :

T1 ∼ Exp(λ1 ), T2 ∼ Exp(λ2 ) (1618)

46.5.1 (i) Distribution of the Minimum: T = min(T1 , T2 )


We find the survival function first:

P (T > t) = P (min(T1 , T2 ) > t) = P (T1 > t and T2 > t) (1619)

By independence:

P (T > t) = P (T1 > t)P (T2 > t) = e−λ1 t · e−λ2 t = e−(λ1 +λ2 )t (1620)

Conclusion: The minimum of two independent exponentials is also exponential.


T ∼ Exp(λ1 + λ2 ). The rate of the minimum is simply the sum of the individual rates.

46.5.2 (ii) Distribution of the Maximum: T = max(T1 , T2 )


The CDF of the maximum is given by:

FT (t) = P (max(T1 , T2 ) ≤ t) = P (T1 ≤ t)P (T2 ≤ t) = (1 − e−λ1 t )(1 − e−λ2 t ) (1621)

To find the PDF fT (t), we differentiate with respect to t:

fT (t) = λ1 e−λ1 t (1 − e−λ2 t ) + (1 − e−λ1 t )λ2 e−λ2 t (1622)

46.5.3 (iii) Probability of T1 occurring before T2


We calculate P (T1 < T2 ) by conditioning on the specific time T2 occurs:
Z ∞ Z ∞
P (T1 < T2 ) = P (T1 < t)fT2 (t)dt = (1 − e−λ1 t )λ2 e−λ2 t dt (1623)
0 0

Solving this integral yields a beautifully simple result:


λ1
P (T1 < T2 ) = (1624)
λ1 + λ2

T2

T2 = T1

T1 < T2

T2 < T1

T1

Figure 1: Region of integration for P (T1 < T2 )


46.6. SMALL INTERVAL APPROXIMATIONS 307

46.6 Small Interval Approximations


When λ is small, or the time interval h is very small (h → 0), we can approximate the
probability of an event occurring:

P (T < h) = 1 − e−λh (1625)


2
Using the Taylor expansion ex ≈ 1 + x + x2 :

λ2 h2
 
P (T < h) ≈ 1 − 1 − λh + = λh + o(h) (1626)
2

Where o(h) represents terms that go to zero faster than h. For very small h, the probability of
an event occurring scales linearly with h.

46.6.1 Probability of Two Events in a Small Interval h


For independent events T1 and T2 :

P (T1 ≤ h, T2 ≤ h) = P (T1 ≤ h)P (T2 ≤ h) ≈ (λ1 h)(λ2 h) = λ1 λ2 h2 (1627)

Since h2 is vastly smaller than h as h → 0, we say this joint probability is o(h). This means
the probability of two independent exponential events occurring in the same infinitesimally
small window is effectively zero.

Counting Process
Stochastic Process N (t) ≥ 0

(event count at time t)

N (t) is a counting process that counts number of events that have occurred by time t.

N (0) = 0; N (t) = 0, 1, 2, . . . (1628)

N (t) is non-decreasing; for s < t =⇒ N (s) ≤ N (t) (1629)

N (t) − N (s) = number of events in the interval (s, t) (1630)

Illustration:

N (s) = 3
continuous time
e1 e2 e3 s N (t) = 5

s = 7.6 t = 9.9
excluding s

N (t) = N (9.9) = 5 (1631)

N (t) − N (s) = N (9.9) − N (7.6) (1632)


308CHAPTER 46. LECTURES 5 AND 6: GEOMETRIC AND EXPONENTIAL DISTRIBUTIONS

=5−3 (1633)

=2 (1634)

Example 1:

N (t) = # of WhatsApp messages typed since the beginning of class

Example 2:
N (t) = # of economic crisis since 1947

Sample path or one realisation of counting process:

count N (t)

jump

jump

jump

t continuous time
t1 t2 t3

Random time where an event occurs and the process jumps.

N (t) is right continuous but not left continuous (1635)

lim N (t) = 2 = N (t2 ) (1636)


t→t+
2

lim N (t) = 1 ̸= N (t2 ) (1637)


t→t−
2

A particular path or realisation corresponds to ω ∈ Ω:

N (t) = N (t, ω) (1638)

N (ω) = Nt (1639)
46.6. SMALL INTERVAL APPROXIMATIONS 309

Independent Increments
Number of events in disjoint time intervals are independent.
Consider times s1 < t1 < s2 < t2 · · · are time intervals,

[s1 , t1 ] and [s2 , t2 ] are disjoint time intervals

[s1 , t1 ] [s2 , t2 ]

continuous time t
s1 t1 s2 t2
N (t1 ) − N (s1 ) N (t2 ) − N (s2 )

N (s2 ) − N (t1 )

Then the number of events in the 3 disjoint time intervals


[N (t1 ) − N (s1 )], [N (s2 ) − N (t1 )] and [N (t2 ) − N (s2 )]
are independent random variables.
Consider [s1 , t1 ] and [s2 , t2 ] as independent intervals:
 
P [N (t1 ) − N (s1 )] = k, [N (t2 ) − N (s2 )] = j
(1640)
= P {N (t1 ) − N (s1 ) = k} · P {N (t2 ) − N (s2 ) = j}

Note: Independent increments does not mean that N (t) is independent of N (s), since
N (t) ≥ N (s).

N (s) N (t)

s t

Stationary Increments
The probability distribution of the number of events depends only on the length of the time
interval and not on the location of that interval.
Consider time intervals [0, t] and [s, s + t]:

[N (t) − N (0)] [N (s + t) − N (s)]

N (t) N (s) N (s + t)

0 t s s+t

events occur in this time interval

P {N (t) = k} = P {N (s + t) − N (s) = k}, ∀k ∈ {0, 1, 2, 3, . . .} (1641)

P {N (t) − N (0) = k} (1642)


310CHAPTER 46. LECTURES 5 AND 6: GEOMETRIC AND EXPONENTIAL DISTRIBUTIONS

Bernoulli Process
Discrete time is divided into time slots. A Bernoulli trial (coin toss) occurs in each time slot.

X1 X2 X3 ··· Xi
Discrete time
Time slot Time slot Time slot Time slot
n=1 n=2 n=3 ··· n=i

Xi ∼ i.i.d (1643)

(
1 w.p. p
Xi = (1644)
0 w.p. 1 − p

Xi ∼ Bern(p) (1645)

The continuous time version of the Bernoulli process is the Poisson process.

n→∞
Discrete continuous time

Poisson Process
Continuous version of the Bernoulli Process.
Single probability experiment with ∞ stages.
Bernoulli Process: Discrete time in time slots. Number of arrivals of events in n time slots or
trials. Distribution (pmf) is Binomial.

A1 A2 A3 · · · Ak

0 T1 T2 T3 ···
Tk

(I) (II) (III)

T1 T2 T3 T4 T5 T6 T7 T8 T9 T10 T11 T12 T13 T14 T15

T1 = 4 T2 = 5 T3 = 6
Time of 1st arrival, inter-arrival time between 1st & 2nd arrival
46.6. SMALL INTERVAL APPROXIMATIONS 311

Ti ∼ Geom(p) i.i.d

Bernoulli Process: B1 , B2 , . . .
(
1 w.p. p
Bi = and Bi is i.i.d
0 w.p. (1 − p)

B1 B2 B3 ··· Bt−1 Bt Bt+1 Bt+2 Bt+3 · · ·

T
start observing BP from time T

Case 1: T depends only on past B1 , B2 , . . . , Bt−1 , Bt :

Xi ∼ Bern(p)

Case 2: T depends on future:


Xi ̸∼ Bern(p)

Events can occur anywhere in continuous time. The Poisson process is the continuous time
version of the Bernoulli process.

continuous time
0
τ τ

Intervals of same length τ have the same probability behaviour.


P (k, τ ) = probability of k arrivals in an interval of fixed duration τ :
X
P (k, τ ) = P (0, τ ) + P (1, τ ) + · · · = 1 (1646)
k

Properties of P (k, τ ):

1. Normalisation: X
P (k, τ ) = 1 for fixed τ (1647)
k

2. Time Homogeneity: P (k, τ ) depends only on the length τ of the time interval and not
on its location. (Analogous to the same success probability p across all time slots in the
Bernoulli process.)

3. Independent Increments: Numbers of arrivals in disjoint time intervals are


independent random variables. (Analogous to independence of disjoint time slots in the
Bernoulli process.)
312CHAPTER 46. LECTURES 5 AND 6: GEOMETRIC AND EXPONENTIAL DISTRIBUTIONS

Small interval probabilities. For a very small time interval δ, with λ denoting the rate
(intensity) of the arrival process:

1 − λδ,
 if k = 0 (zero arrivals)
P (k, δ) = λδ, if k = 1 (one arrival) (1648)

0, if k > 1 (two or more arrivals)

P (1, δ)
lim =λ (intensity of process) (1649)
δ→0 δ
 
E N (0, δ) = 1 · (λδ) + 0 · (1 − λδ) + 0 = λδ (1650)


1 w.p. λδ

N (0, δ) = 0 w.p. (1 − λδ) (1651)

2, 3, 4, . . . w.p. 0

E[# arrivals in [0, δ]]


=λ (1652)
δ

Observation: An interval of large length τ can be subdivided into n small intervals of length
δ, where nδ = τ .

δ δ δ ··· δ δ

0 1st 2nd 3rd (n − 1)th nth

τ τ
nδ = τ, n= , δ= (1653)
δ n

p = λδ = P (1, δ), δ→0 ⇒ n→∞ (1654)

Derivation of the Poisson PMF. In the Bernoulli process with n time slots each of success
probability p = λτ /n:
   k 
λτ n−k

n λτ
P [k arrivals in n time slots] = 1− (1655)
k n n
As n → ∞ (with np = λτ held fixed), expanding:
k 
λτ (n−k)
 
n! λτ
1−
k!(n − k)! n n

n − k + 1 (λτ )k λτ n λτ −k
     
n n−1
= · ··· · · 1− · 1−
n n n k! n n
The bracketed term → 1 and (1 − λτ /n)−k → 1 as n → ∞. For the exponential term, let
x = limn→∞ (1 − λτ /n)n :
46.6. SMALL INTERVAL APPROXIMATIONS 313

ln 1 − λτ
  
λτ n L’H −λτ n→∞
ln x = lim n ln 1 − = lim 1 = λτ
−−−→ −λτ
n→∞ n n→∞
n 1− n
 n
λτ
⇒ 1− −→ e−λτ
n
Therefore:

(λτ )k e−λτ
P (k, τ ) = , k = 0, 1, 2, . . . (1656)
k!

Nτ ∼ Pois(λτ ), Nt ∼ Pois(λt)

E(Nt ) = λt = Var(Nt )

Comparison: Binomial vs. Poisson

X ∼ Bin(n, p) Nτ ∼ Pois(λτ )
E(X) = np → λτ E(Nτ ) = λτ
Var(X) = np(1 − p) → λτ Var(Nτ ) = λτ

Example. You receive emails according to a Poisson process at rate λ = 5 msgs/hr. You
check your email every 30 minutes (τ = 12 hr), so λτ = 2.5.

(λτ )0 e−λτ
P [no new msgs] = P (0, τ ) = = e−2.5 ≈ 0.08
0!

E[# emails in 30 min] = λτ = 2.5 msgs

(λτ )1 e−λτ
P [exactly 1 new email in 30 min] = P (1, 12 ) = = 2.5 e−2.5 ≈ 0.205
1!
314CHAPTER 46. LECTURES 5 AND 6: GEOMETRIC AND EXPONENTIAL DISTRIBUTIONS
Chapter 47

Lecture 9: Poisson Distribution


Derivation

47.1 Poisson Distribution


The Poisson Distribution is a discrete probability distribution that models the number of
events occurring in a fixed interval of time or space, given that these events occur:

ˆ Independently
ˆ At a constant average rate
It is often called the “law of rare events”.

47.1.1 Probability Mass Function (PMF)


The probability of observing exactly x events is given by:

e−λ λx
P (X = x) = , x = 0, 1, 2, . . .
x!
where:

ˆ λ = average number of events (mean rate)


ˆ e = Euler’s constant (≈ 2.718)
47.1.2 Key Properties
ˆ Mean: E(X) = λ
ˆ Variance: Var(X) = λ
ˆ Standard Deviation: √λ
47.1.3 Assumptions of Poisson Distribution
1. Events occur independently

2. The average rate (λ) is constant

3. Two events cannot occur at the exact same instant

4. The probability of more than one event in a very small interval is negligible

315
316 CHAPTER 47. LECTURE 9: POISSON DISTRIBUTION DERIVATION

47.1.4 Relation with Binomial Distribution


Poisson distribution can be derived as a limiting case of the Binomial distribution when:

ˆ n→∞
ˆ p→0
ˆ np = λ (finite)
47.1.5 Cumulative Probability
x
X e−λ λk
P (X ≤ x) =
k!
k=0

47.1.6 Applications
ˆ Number of calls received at a call center per minute
ˆ Number of defects in a manufactured product
ˆ Number of accidents at a traffic signal
ˆ Number of customers arriving at a store
Let,
p = λδ = P (1 arrival in δ)
As δ → 0, n → ∞ such that nδ = τ (fixed).

Binomial Setup:
   k 
λτ n−k

n λτ
P (k arrivals in n time slots) = 1−
k n n

Expand the Binomial Coefficient:


 
n n!
=
k k!(n − k)!
n!
= n(n − 1)(n − 2) · · · (n − k + 1)
(n − k)!
So,  
n n(n − 1)(n − 2) · · · (n − k + 1)
=
k k!

Substitute into Probability:


k 
λτ n−k
 
n(n − 1) · · · (n − k + 1) λτ
P = 1−
k! n n

Rearranging:

λτ n
   
1 n(n − 1) · · · (n − k + 1) k
P = (λτ ) 1 −
k! nk n
47.1. POISSON DISTRIBUTION 317

As n → ∞:

n(n − 1) · · · (n − k + 1)
→1
nk
and

λτ n
 
1− → e−λτ
n

Final Result (Poisson Distribution):

e−λτ (λτ )k
P (k) = , k = 0, 1, 2, . . .
k!
Let: k = number of arrivals in a time interval of length τ

Limit Derivation:
 
λτ
lim n ln 1 −
n→∞ n
Using log approximation:

ln(1 − x) ≈ −x as x → 0
   
λτ λτ
⇒ lim n ln 1 − = lim n −
n→∞ n n→∞ n

= −λτ
  
λτ
⇒ exp lim n ln 1 − = e−λτ
n→∞ n

Poisson Distribution Result:


Let Nτ be the number of arrivals in time interval τ .

Nτ ∼ Pois(λτ )
Probability Mass Function (PMF):

e−λτ (λτ )k
P (Nτ = k) = , k = 0, 1, 2, . . .
k!

Poisson Process for Time t:

Nt ∼ Pois(λt)
Mean and Variance:

E(Nt ) = λt

Var(Nt ) = λt

Connection with Binomial Distribution:


318 CHAPTER 47. LECTURE 9: POISSON DISTRIBUTION DERIVATION

X ∼ Bin(n, p)
For large n and small p:

np = λτ

⇒ Bin(n, p) → Pois(λτ )

Final Insight:
The Poisson distribution arises as a limiting case of the Binomial distribution when:

ˆ n→∞
ˆ p→0
ˆ np = constant
Poisson Process Applications
Events such as email arrivals, calls, and accidents can be modeled using a Poisson process.

Variance Relationship:

Var(X) = np(1 − p)
As n → ∞ and p → 0:

np = λτ

⇒ Var(Nτ ) = λτ

Example: Email Arrivals


Given: Emails arrive according to a Poisson process with rate:

λ = 5 messages per hour

Time interval:
1
t = 30 minutes = hour
2

Step 1: Compute Mean Number of Arrivals

1
λt = 5 × = 2.5
2

E(number of emails) = 2.5

Step 2: Probability of No New Messages

e−λt (λt)0
P (X = 0) =
0!

= e−2.5
47.1. POISSON DISTRIBUTION 319

≈ 0.082

Step 3: Probability of Exactly One New Message

e−λt (λt)1
P (X = 1) =
1!

= (2.5)e−2.5

= 2.5 × 0.082

≈ 0.205

Final Results:

P (no new messages) ≈ 0.082

P (exactly one message) ≈ 0.205

Interpretation: There is a low probability (8.2%) of receiving no emails in 30 minutes, and


about a 20.5% chance of receiving exactly one email. Concept: Continuous-Time Arrival
Process
Events E1 , E2 , . . . , Ek occur over continuous time.

0 −→ Y1 −→ Y2 −→ · · · −→ Yk

Yk = time of the k th arrival

Interarrival Times:

T1 , T2 , . . . , Tk

Yk = T1 + T2 + · · · + Tk
Assumption:
Ti ∼ Exp(λ), Ti are i.i.d.

Number of Arrivals in Time Interval:

Nt = number of arrivals in (0, t)

Nt ∼ Pois(λt)

PDF of Yk :

fYk (t) = PDF of time of k th arrival


320 CHAPTER 47. LECTURE 9: POISSON DISTRIBUTION DERIVATION

Using definition of PDF:

fYk (t) δ ≈ P (t ≤ Yk ≤ t + δ)

This represents the probability that the k th arrival occurs in a small interval (t, t + δ).

Interpretation:
For Yk to lie in (t, t + δ):

ˆ Exactly (k − 1) arrivals occur before time t


ˆ One arrival occurs in (t, t + δ)
Continuous Random Variable Reminder:
Z b
P (a ≤ X ≤ b) = fX (x) dx
a

For a small interval:

P (a ≤ X ≤ a + ∆) ≈ fX (a) ∆

Graphical Representation:

fX (x)

x
ab

Key Insight:

k
X
Yk = Ti
i=1

Ti ∼ Exponential(λ)

⇒ Yk ∼ Gamma(k, λ)

Final Summary:

ˆ N ∼ Pois(λt)
t

ˆ Interarrival times T ∼ Exp(λ)


i

ˆ Arrival time Y ∼ Gamma(k, λ)


k
47.1. POISSON DISTRIBUTION 321

Poisson Process (Continuation)


Continuous Time Representation:
Events E1 , E2 , . . . , Ek occur along a continuous time axis.

0 −→ Y1 −→ Y2 −→ · · · −→ Yk

Yk = time of the k th arrival

Interarrival Times:

T1 , T2 , . . . , Tk

Y1 = T1 , Y2 = T1 + T2 , ..., Yk = T1 + T2 + · · · + Tk

k
X
Yk = Ti
i=1

Distribution of Interarrival Times:

Ti ∼ Exp(λ)

Ti are independent and identically distributed (i.i.d.)

Number of Arrivals in an Interval:

Nt = number of arrivals in (0, t)

Nt ∼ Pois(λt)

PDF of Yk :

fYk (t) =?

Using definition of PDF:

fYk (t) δ ≈ P (t ≤ Yk ≤ t + δ)

Interpretation:
For Yk to lie in (t, t + δ):

ˆ Exactly (k − 1) arrivals occur before time t


ˆ The k arrival occurs in (t, t + δ)
th
322 CHAPTER 47. LECTURE 9: POISSON DISTRIBUTION DERIVATION

Continuous Random Variable Reminder:


Z b
P (a ≤ X ≤ b) = fX (x) dx
a
For a very small interval ∆:

P (a ≤ X ≤ a + ∆) ≈ fX (a) ∆

Key Result:
k
X
Yk = Ti
i=1

Ti ∼ Exp(λ)

⇒ Yk ∼ Gamma(k, λ)

Summary:
ˆ N ∼ Pois(λt)
t

ˆ Interarrival times T ∼ Exp(λ)


i

ˆ Arrival time of k event: Y ∼ Gamma(k, λ)


th
k

Waiting Time Distribution (Erlang Distribution)


Let Yk be the waiting time until the k th arrival.

Step 1: Probability Expression

fYk (t) δ ≈ P (t ≤ Yk < t + δ)

= P (k − 1 arrivals in (0, t)) × P (1 arrival in (t, t + δ))

= P (N (t) = k − 1) · (λδ)

Using Poisson Distribution:

e−λt (λt)k−1
P (N (t) = k − 1) =
(k − 1)!

Substitute:

e−λt (λt)k−1
fYk (t) δ = λδ ·
(k − 1)!

Divide by δ:

λk tk−1 e−λt
fYk (t) = , t≥0
(k − 1)!
47.1. POISSON DISTRIBUTION 323

Final Result (Erlang Distribution):

Yk ∼ Erlang(λ, k)

λk tk−1 e−λt
fYk (t) = , k = 1, 2, 3, . . .
(k − 1)!

Special Case:
For k = 1:

fY1 (t) = λe−λt

Y1 ∼ Exponential(λ)

Lecture 10: Bernoulli vs Poisson Process


Basic Idea
ˆ Divide time into small intervals of length δ , δ , . . .
1 2

ˆ Arrival events occur in continuous time


ˆ For small interval δ, probability of one arrival:
p = λδ

ˆ As n → ∞, total time t = nδ
Nt ∼ Poisson(λt)

Comparison Table

Characteristic Poisson Process Bernoulli Process


(1) Time of Arrival Continuous Time Discrete Time
(2) Arrival Rate λ (constant) p (per trial)
(3) Number of Arrivals Poisson Distribution Binomial Distribution
(4) Interarrival Time Exponential (λ) Geometric (p)
(5) Time to k th Arrival Erlang (λ, k) Pascal (k, p)

Key Relations
ˆ N ∼ Poisson(λt)
t

ˆ Interarrival time:
Y1 ∼ Exponential(λ)

ˆ Waiting time for k th arrival:


Yk ∼ Erlang(λ, k)
324 CHAPTER 47. LECTURE 9: POISSON DISTRIBUTION DERIVATION

Example: Poisson Process


Let Nt ∼ Poisson(λt).

Given: λ = 1
Consider the time interval split:
[0, 5] = [0, 2] ∪ [2, 5]

Number of arrivals:
N[0,2] ∼ Poisson(2)

N[2,5] ∼ Poisson(3)

PMF of Poisson Distribution:

e−λt (λt)k
P (Nt = k) =
k!

So,
e−2 2k
P (N[0,2] = k) =
k!

e−3 3k
P (N[2,5] = k) =
k!

Mean number of arrivals:


E[Nt ] = λt

E[N[0,2] ] = 2, E[N[2,5] ] = 3

Total number of arrivals:


N[0,5] = N[0,2] + N[2,5]

N[0,5] ∼ Poisson(5)

Independence Property:
For disjoint intervals,
N[0,2] and N[2,5] are independent

⇒ P (N[0,2] ∩ N[2,5] ) = P (N[0,2] ) · P (N[2,5] )


47.1. POISSON DISTRIBUTION 325

Sum of Independent Poisson Random Variables


Let,
P1 ∼ Poisson(λ1 t), P2 ∼ Poisson(λ2 t)

Define:
P = P1 + P2

Goal: Find distribution of P

Using Convolution (Idea):


k
X
P (P = k) = P (P1 = i) P (P2 = k − i)
i=0

k
X e−λ1 t (λ1 t)i e−λ2 t (λ2 t)k−i
= ·
i! (k − i)!
i=0

k
X (λ1 t)i (λ2 t)k−i
= e−(λ1 +λ2 )t
i!(k − i)!
i=0

e−(λ1 +λ2 )t (λ1 + λ2 )k tk


=
k!

Using MGF:
MGF of Poisson(λt):
MP (t) = exp λt(es − 1)
 

For independent variables:


MP (s) = MP1 (s) · MP2 (s)

= exp[λ1 t(es − 1)] · exp[λ2 t(es − 1)]

= exp[(λ1 + λ2 )t(es − 1)]

Conclusion:

P1 + P2 ∼ Poisson((λ1 + λ2 )t)

General Result:
If
Pi ∼ Poisson(λi t), i = 1, 2, . . . , n
then
n n
! !
X X
Pi ∼ Poisson λi t
i=1 i=1

t −1) t −1) t −1)


Mp (t) = eλ1 (e · eλ2 (e = e[λ(e ]
326 CHAPTER 47. LECTURE 9: POISSON DISTRIBUTION DERIVATION

t
P ∼ Pois(λ) (Σ) λ1 + λ2 e[(λ1 +λ2 )(e −1)]

y
x+y =1

Through Convolution Method


∼→ indep
x1 , x 2 X = (x1 + x2 )
P [X = k] = P [(x1 + x2 ) = k]
indep RVs
Z = (X + Y )

Law of Total Probability


?
X
P [z = k] = P [(x + y) = k] = P [z = k, Y = y]
y
X
= P [z = k | Y = y] P (y = y)
y
X
⇒ P [(x + y) = k | Y = y] P (y = y) ,→ remove this as X & Y
y
X
P (z = k) ⇒ P [x = (k − y)] P (Y = y) are indep
y
RVs.
η
X
pz (k) = px (k − y) py (y)
y

⇒ pz = px ∗ py ↗ convolution

Convolution

Binomial Theorem

e−λ1 λx1 e−λ2 λy2


p1 (x) = p2 (y) =
x! y!

X λ(k−y) e−λ1 λy2 e−λ2


1
p2 (k) = ·
y
(k − y)! y!
!
(k−y)
X λ1 λy2
= e−(λ1 +λ2 )
y
(k − y)! y!

X λ(k−y) λy
= 1 2
· e−(λ1 +λ2 )
y
(k − y)! y!

X ∼ Pois(λ1 ), Y ∼ Pois(λ2 )
47.1. POISSON DISTRIBUTION 327

X k! λ(k−y) λy
(λ1 + λ2 )k = 1 2

y
(k − y)! y!

k ⩾ y, 0⩽k

(k − y)! y! λ1 + λ2 → p

λ2 ⩾ q

X n! p(n−y) q y
(p + q)n =
y
y! (n − y)!

By Binomial Theorem

X k! λ(k−y) λy 1 (λ1 + λ2 )k
1 2
· −→
y
(k − y)! y! k! k!
328 CHAPTER 47. LECTURE 9: POISSON DISTRIBUTION DERIVATION
Chapter 48

Lecture 11 and 12

Discrete Time Markov chain


Time is Discrete ⇒ Time Index n = 0, 1, 2, . . . , n
State Space also discrete S = set no. of states
The State space is a time dependent random state
⇒ State of System at time n

Xn ∈ S = {0, 1, 2, . . . , i, i + 1, . . . }

*can be countably finite / infinite


The states of system are:
X0 , X1 , X2 , . . . , Xn , . . .

an SP Xn is a DTMC if
 

⃗  Markov
P Xn+1 = j | Xn = i, X = ⃗x  −−−−→ P [Xn+1 = j | Xn = i] = pn,ij ∀i, j ∈ S

| {z } | {z } | n−1 {z n−1}
Future Present Past History

given the present, the future is conditionally independent of the past.

Typically in a DTMC time doesn’t exist outside any of these time intervals, (it is not
measured).

ˆ The state remains at x 0 till time 1, then it jumps to x1 , and so on.

⊛ Also, How you get to a certain state Xn = i, does not matter. Only the present state
determines the outcome of the future state.

BUT HOLD ON, THIS LOOKS LIKE GEOMETRIC MEMORYLESS PROP.

Where does the memory less property occur?


It occurs in the time period where there is a state change between various states.
The path we take to j doesn’t matter.

329
330 CHAPTER 48. LECTURE 11 AND 12

States
j

n
n n+1

→ If pn,ij is dependent of n → Time Inhomogeneous DTMC


→ If pn,ij is independent of n → Time Homogeneous DTMC

Suppose we are in 2 states: {H, A}


The states can be like:

0.2

0.8 H A 0.7

0.3

One step Transition prob.


 
0.8 0.2 Row sum = 1
P (1) = 
0.3 0.7 Row sum = 1

Each row is a different distribution, and the matrix captures that.


This is a model which can be said to be affected by today only.

What about a model affected by Today and Yesterday?

Model: Tomorrow Affected by Today and Yesterday.

Yesterday Today Tomorrow


Xn−1 −−−−−−→ Xn −−−−→ Xn+1 −−−−−−→

 
HH 0.6 0.4
 
AH  0.4 0.6 
 
 
 
HA 0.7 0.3 

 
AA 0.95 0.05

State Diagram:
331

0.85
HA AA 0.15
0.1

0.9 0.6

0.4
AH HH 0.25
0.75

Transition Matrix P :  
0.25 0.75 0 0 HH
 
 0 0 0.6 0.4  AH
 
P =



 0.1 0.9 0 0  HA
 
0 0 0.85 0.15 AA

*Columns correspond to HH, AH, HA, AA respectively from left to right.

Matrix Properties:

ˆ If row sum = 1 ⇒ Singly Stochastic


ˆ If row sum = 1 and col sum = 1 ⇒ Doubly Stochastic

The Drunkard’s Walk


A special Type of Random walk where a drunkard:

ˆ moves right with prob. P (H) = p


ˆ moves left with prob. P (T ) = q
0 is his house.
One qn: Can the drunkard ever come back home (make it back to zero)?

−2 −1 0 1 2

S = {−∞, . . . , −2, −1, 0, 1, 2, . . . }

Diagram
332 CHAPTER 48. LECTURE 11 AND 12

p p p p p
... i−1 i i+1 i+2 ...
q q q q q

S = {0, ±1, ±2, ±3, . . . , ±∞} |S| = ∞

If p < 1/2
If p > 1/2 If p = 1/2, moves about zero.

0 n n
0
0 n

60-70% prediction Rate?


Account for Climate
Condition, Pitch
Condition, etc

Random Walk with boundaries (Gambling)


In a 1D random walk, the stop conditions are either i = 0 or i = J.

1 1
p p p p p p p
0 1 2 ... i i+1 ... J −1 J
q = (1 − p)q = (1 − p) q = (1 − p)
q = (1 − p) q = (1 − p)

Gambler Gambler
Gets Wins
Ruined

States = (J + 1)

0 and J are called absorbing states as they enter these states and stay there forever.

P00 = 1

 Xn = State of gamblers at time n
Pi,i+1 = p
∀i ∈ {1, 2, 3, . . . , J − 1} Xn ∈ S
Pi,i−1 = q 

 n = 0, 1, 2, . . .

PJJ = 1
333

Gambler’s Ruin
→ You place Rs b bets

ˆ w.p p = P (H), you gain Rs b


ˆ w.p q = P (T ), you lose Rs b

Gambler’s Ruin (Continued)


ˆ You’re a gambler who places Rs. b bets
ˆ w.p p = P (H), you gain Rs. b
ˆ w.p q = P (T ), you lose Rs. b
ˆ The Initial wealth you start with is Rs. ib
q P (lose)
Bias factor α = p = P (win)

ˆ If α > 1, odds against us ⇒ q > p


ˆ If α < 1, odds favour us ⇒ q < p
(1) You keep gambling until:

a) You go broke (enter state 0)

b) Reach wealth of Rs. BJ

The question we wanna ask is:


Prob of Event Ei = Success Event = Gambler wins
Ei =
Gambler reaching Rs. BJ before going broke, provided he starts with initial wealth of Rs. ib

P (Ei ) =? ⇒ P (Xn = J | X0 = i) =?

W.L.O.G, we can assume that b = Rs. 1

Probability Foundations:
According to Law of Total Prob:

P (A) = P (A ∩ B) + P (A ∩ B c )

Multiplication Rule & Conditional Probability:

P (B) = P (B | A) × P (A)

P (A ∩ B | F ) = P (A | B, F )P (B | F )
334 CHAPTER 48. LECTURE 11 AND 12

Margin Notes on State Transitions:

P (Xn = J | X0 = i)
 
P (Xn = J ∩ X0 = i) = P Xn = J ∩ (X1 = 1 ∪ X1 = 2 · · · ∪ Xn = i)

Probability Rules applied:

P (A ∩ B | F ) = P (A | B, F )P (B | F )
P (B | F ) = P (B, A | F ) + P (B, Ac | F )

Applying This to Ei , using Law of Total Probability (LTP):

P (Ei ) = P [Xn = J, X1 = i + 1 | X0 = i]
+ P [Xn = J, X1 = i − 1 | X0 = i]

Expanding the first term:

P [Xn = J, X1 = i + 1 | X0 = i] = P [Xn = J | X1 = i + 1, X0 = i] × P [X1 = i + 1 | X0 = i]


| {z } | {z } | {z }
Future Present Past

*due to markov, it changes into P [Xn = J | X1 = i + 1]

Substituting back:

⇒ P (Ei ) = P [Xn = J | X1 = i + 1] × P [X1 = i + 1 | X0 = i]


| {z }
=p

+ P [Xn = J | X1 = i − 1] × P [X1 = i − 1 | X0 = i]
| {z }
=q

⇒ P (Ei ) = P [Xn = J | X1 = i + 1] ×p + P [Xn = J | X1 = i − 1] ×q


| {z } | {z }
δi+1 δi−1
due to memoryless property

δi = pδi+1 + qδi−1 for i = 1, 2, . . . , J − 1

α = q/p (+ doodle: ”DON’T KILL ME PLS”)

Boundary conditions to: δi = pδi+1 + qδi−1

i = 0 ⇒ δ0 = P [Xn = J | X0 = 0] = 0
i = J ⇒ δJ = P [Xn = J | X0 = J] = 1
335

Derivation:
Since p + q = 1, we can rewrite δi as δi (p + q):

⇒ δi (p + q) = pδi+1 + qδi−1
δi p + δi q = pδi+1 + qδi−1
p(δi+1 − δi ) = q(δi − δi−1 )
q
⇒ (δi+1 − δi ) = (δi − δi−1 )
p
⇒ (δi+1 − δi ) = α(δi − δi−1 )

Subbing in diff values of i:

i = 1 ⇒ (δ2 − δ1 ) = α(δ1 − δ0 )
i = 2 ⇒ (δ3 − δ2 ) = α(δ2 − δ1 ) ⇒ α2 (δ1 − δ0 )

i = 3 ⇒ (δ4 − δ3 ) = α(δ3 − δ2 ) ⇒ α3 (δ1 − δ0 )


..
. (add all of them)
i = J − 1 ⇒ (δJ − δJ−1 ) = α(δJ−1 − δJ−2 ) ⇒ αJ−1 (δ1 − δ0 )

ADDING ALL OF THEM:

⇒ α + α2 + α3 + · · · + αJ−1 (δ1 − δ0 )
 

For upto i − 1:
⇒ α + α2 + α3 + α4 + · · · + αi−1 (δ1 − δ0 )
 

Continuing from upto i − 1:

⇒ δi − δ1 = α + α2 + α3 + · · · + αi−1 (δ1 − δ0 )
 

Since δ0 = 0, we can rewrite this by moving δ1 to the RHS (effectively adding 1 to the
bracketed series):

δi = 1 + α + α2 + · · · + αi−1 (δ1 − 0)
 

δi = 1 + α + α2 + · · · + αi−1 δ1
 

Using the geometric series sum formula:

1 − αi
1 + α + α2 + · · · + αi−1 =
1−α

1 − αi
 
⇒ δi = δ1
1−α
336 CHAPTER 48. LECTURE 11 AND 12

Applying the boundary condition δJ = 1:

1 − αJ
   
1−α
δJ = 1 = δ1 ⇒ δ1 =
1−α 1 − αJ

Substitute δ1 back into the δi equation:

1 − αi
   
1−α
δi = × for α = q/p ̸= 1
1−α 1 − αJ

1−αi
δi = 1−αJ

When α = 1:
This implies q = p = 1/2 (a fair game with equal odds).

δi = 1 + 1 + 12 + · · · + 1i−1 δ1
 

δi = iδ1

F
Applying boundary condition δJ = 1:

δJ = 1 ⇒ Jδ1 = 1 ⇒ δ1 = 1/J
δi = iδ1 ⇒ δi = i(1/J)

δi = i/J when α = 1 ⇒ q = p = 1/2

HW:
Perform Gambler’s Ruin without
formula in Python for:
α = 1, α < 1, α > 1

Favour Gambler
When α < 1 ⇒ q < p ⇒ P (lose) < P (win)

1 − αi J→∞⇒αJ →0
δi = −−−−−−−−−→ (1 − αi )
1 − αJ Since α<1

ˆ As i → ∞ (initial wealth i is large): α → 0


i

ˆ Result → 1
337

ˆ Healthy Gambler wins w.p 1

When α > 1 ⇒ q > p ⇒ P (lose) > P (win)

αi − 1 J→∞ αi J→∞
δi = −−−−→ −−−→ 0
αJ − 1 αJ ≫1 αJ

ˆ Result → 0
ˆ Regular Gamblers lose w.p 1 (win w.p 0)

Assume Time Homogeneous DTMC on S


n-step transition prob:
(n)
Pij = P [Xn = j | X0 = i] ⇒ P (n)

To Xn+1
 
P11 P12 . . . P1m
 
P (n) = P21 P22 . . . P2m 
 
From Xn  ..

.. .. .. 

 . . . . 
 
Pi1 Pi2 . . . Pim
338 CHAPTER 48. LECTURE 11 AND 12
Chapter 49

Lecture 13 and 14: Poisson Process


Properties

49.1 Introduction
Stochastic processes serve as the mathematical foundation for modeling systems that evolve
over time under uncertainty, a concept central to modern financial econometrics and risk
management. This section explores two fundamental pillars of probability theory: the Poisson
Process and Discrete-Time Markov Chains (DTMC).The Poisson Process provides a bridge
between discrete event counts and continuous time. By assuming that inter-arrival times
follow an Exponential Distribution, we leverage the memoryless property, ensuring that the
probability of a future event is independent of the time elapsed since the last occurrence. As
we aggregate these waiting times, we transition from the simple Exponential PDF to the
Erlang (or Gamma) Distribution, which describes the time required for a specific number of
arrivals (r) to occur. This relationship is rigorously proven through the link between the
Poisson Cumulative Distribution Function and the survival probability of arrival
[Link] from counting processes to state-based evolution, we introduce the Markov
Property. Here, the ”memoryless” concept is applied to sequences of random variables where
the future state depends solely on the present, rendered independent of the historical path.
Whether analyzing arrival rates in a queue or the shifting states of a financial market, these
models provide the analytical rigor necessary to quantify randomness in complex,
time-dependent systems.

Relation between Exponential RV and Poisson Process (Stochastic Process)


Poisson Arrival Process
Example: Phone calls coming in.
Sequence of arrival events (x) that occurs at random (continuous) time.

w1 = waiting time of event 1 w3

t is continuous time
0 1 2 3
w2

Sequence of waiting time RVs w1 , w2 , w3 , . . .


Assume: The waiting times are independent identically distributed (i.i.d.).

339
340 CHAPTER 49. LECTURE 13 AND 14: POISSON PROCESS PROPERTIES

wi ∼ Exp(λ) i.i.d. i = 1, 2, 3, . . .
Memoryless Property: The process looks the same no matter when we start waiting.

Poisson Process: Fixed Time Interval and Properties


Fix time interval I = (t1 , t2 ).

t1 t2
N (I) = 2
w1

× × × × t
0 1 2 3
(t2 − t1 ) =I length(I)

N (I) = No. of events in the observation window I = (t1 , t2 )


↑ Discrete Counting RV (N ≥ 0)

Mean of no. of arrivals = λ · length(I) = λ(t2 − t1 )


↑ rate of arrival events.

If λ is high and mean is high ⇒ we see lot of events in I.



The inter-arrival times wi are small (since we are squeezing events).
1 1
E[wi ] = =⇒ λ =
λ E[wi ]
no. of events
λ units of unit time .

Conceptual Explanation: The Poisson Process


1. From Inter-arrival Times to Counting
The Poisson Process bridges the gap between continuous time and discrete events.
ˆ Inter-arrival Times (w ): These represent the duration between consecutive events.
i
We assume these are Exponentially distributed and independent.

ˆ Counting Variable (N (I)): Instead of looking at when things happen, we look at how
many things happen in a fixed window I. This is a discrete random variable.

2. The Role of the Arrival Rate (λ)


The parameter λ is the link between the two perspectives:
1 1
λ= =⇒ Rate = (1657)
E[wi ] Mean Waiting Time
If the rate λ is high, the events are ”squeezed” together, making the expected waiting time
E[wi ] very small.
49.1. INTRODUCTION 341

3. The Independent Increments Property


A stochastic process has independent increments if the number of events occurring in
disjoint (non-overlapping) time intervals are independent random variables.
The no. of arrival events in disjoint time intervals are independent.

e−λ(t2 −t1 ) · [λ(t2 − t1 )]k


P [N (t1 , t2 ) = k ] = , k = 0, 1, 2, . . .
| {z } k!
N (I)

This means that knowing how many events occurred in the interval (t1 , t2 ) provides zero
information about how many will occur in the future interval (t3 , t4 ).

Poisson Process: Waiting Times and Erlang Distribution


(λt3 )0 e−λt3
P [no arrival event in interval (0, t3 )] = 0! = e−λt3

w1
Nt3 = 0

× continuous time
t3

[Nt3 = 0] ≡ [w1 > t3 ] (1658)

P [w1 > t3 ] = e−λt3 =⇒ w1 ∼ Exp(λ)

Time until the r-th event


Let Tr be the waiting time until the r-th event:

Tr = (w1 + w2 + · · · + wr ) (1659)

We know that wi ∼ Exp(λ) i.i.d. for i = 1, 2, . . . , r. The PDF of Tr is given by the Erlang or
Gamma Distribution:

λr tr−1 e−λt
fTr (t) = (1660)
(r − 1)!

Survival Probability
The probability that the r-th event occurs after time t (the Survival Probability) is equivalent
to saying there have been fewer than r events by time t:

r−1
X
P [Tr > t] = P [Nt ≤ (r − 1)] = P [Nt = k] (1661)
k=0
342 CHAPTER 49. LECTURE 13 AND 14: POISSON PROCESS PROPERTIES

Rigorous Derivation: Relationship between Tr and Nt


1. Logical Case Analysis
To determine the probability that the rth arrival occurs after time t (Tr > t), we analyze the
state of the process at time t.

Case I: The (r − 1)th event occurs after t.


If even the (r − 1)th person hasn’t arrived by time t, then the rth person certainly
hasn’t arrived.
In this scenario: Nt < (r − 1)

Case II: The (r − 1)th event occurs before t, but the rth hasn’t.
If the (r − 1)th person arrived already, but we are still waiting for the rth arrival, then
at time t, exactly (r − 1) people have arrived.

In this scenario: Nt = (r − 1)

Conclusion: The event [Tr > t] is logically equivalent to the union of these cases:

[Tr > t] ⇐⇒ [Nt ≤ r − 1]

(r − 1)th rth
× × Time
0 t Tr

2. Probability Summation
Since Nt is a Poisson Random Variable with mean λt, the probability of the count being less
than or equal to r − 1 is the sum of the individual probabilities for k = 0, 1, . . . , r − 1:

r−1 −λt
X e (λt)k
P (Tr > t) = P (Nt ≤ r − 1) = (1662)
k!
k=0

3. Linking to CDF and PDF


The Cumulative Distribution Function (CDF) represents the probability that the event
happens before or at time t.

1. Survival Function: P (Tr > t) = 1 − FTr (t).


Pr−1 e−λt (λt)k
2. CDF Identity: From the summation above, FTr (t) = 1 − k=0 k! .

3. PDF Derivation: The Probability Density Function fTr (t) is the derivative of the
CDF:
r−1 −λt
" #
d d X e (λt)k
fTr (t) = FTr (t) = 1− (1663)
dt dt k!
k=0

This derivation proves that the sum of r i.i.d. exponential variables (the time Tr ) follows an
Erlang distribution, which can be expressed via the Poisson CDF.
49.1. INTRODUCTION 343

Derivation of the PDF fTr (t)


Starting from the CDF identity:
r−1 −λt
X e (λt)k
FTr (t) = 1 − P (Tr > t) = 1 −
k!
k=0
The PDF is the derivative of the CDF with respect to t:
r−1 −λt k k
" #
d X e λ t
fTr (t) = 1−
dt k!
k=0
Applying the derivative and pulling out constants:
r−1 k
X λ d k −λt
fTr (t) = − (t e )
k! dt
k=0

Using the Product Rule d


dt [u · v] = u′ v + uv ′ :
r−1 k h
X λ i
=− ktk−1 e−λt − λtk e−λt
k!
k=0
Distributing the negative sign and rearranging:
r−1 k h
X λ i
fTr (t) = λtk e−λt − ktk−1 e−λt
k!
k=0
Expanding the summation:
r−1 k+1 k −λt r−1 k k−1 −λt
X λ t e X λ kt e
= −
k! k!
k=0 k=0
Factoring out λ and simplifying k/k! = 1/(k − 1)!:
" r−1 r−1
#
X (λt)k e−λt X (λt)k−1 e−λt
=λ −
k! (k − 1)!
k=0 k=1
Note: In the second sum, the k = 0 term was 0, so the index shifts to k = 1.

Final Derivation: The Telescoping Sum Proof


Continuing from the derivative of the Poisson CDF, we simplify the expression using a
substitution method to visualize the cancellation of terms.

1. Substitution and Index Shifting


We define the Poisson probability term as zk :
(λt)k e−λt
zk = (1664)
k!
Substituting this into our differentiated expression, and shifting the index of the second
summation by setting p = k − 1:
 
Xr−1 r−2
X
fTr (t) = λ  zk − zp  (1665)
k=0 p=0

Explanation: By shifting the second index from k to p, we align the terms so they can be
compared directly.
344 CHAPTER 49. LECTURE 13 AND 14: POISSON PROCESS PROPERTIES

2. The Telescoping Cancellation


Expanding both summations shows that they are identical except for the final term of the first
sum:
fTr (t) = λ [(z0 + z1 + z2 + · · · + zr−2 + zr−1 ) − (z0 + z1 + · · · + zr−2 )] (1666)

Mathematical Logic

Every term from z0 to zr−2 in the first bracket is subtracted by the corresponding term
in the second bracket. This is known as a telescoping sum. Only the term for k = r − 1
survives.

3. Final Result: Erlang PDF


After cancellation, only zr−1 remains:

fTr (t) = λzr−1 (1667)


(λt)r−1 e−λt
=λ (1668)
(r − 1)!
λ · λr−1 tr−1 e−λt
= (1669)
(r − 1)!
λ t e−λt
r r−1
fTr (t) = (1670)
(r − 1)!

Summary: This proves that the time until the rth arrival in a Poisson Process with rate λ
follows the Erlang Distribution with parameters (r, λ).

Alternative Derivation: P [Tr ∈ (t, t + h]]


We are looking for the probability that the rth arrival occurs within a specific time window
(t, t + h]. This can be expressed by analyzing the count of events in two disjoint time intervals:
(0, t] and (t, t + h].

1. Logical Event Decomposition


For the rth event to fall between t and t + h, we must have seen some number of events k by
time t, and then exactly (r − k) events in the subsequent interval h.

ˆ Case 1: N t = 0 and N (t, t + h) = r. (Zero arrivals before t, all r arrivals happen in h).

ˆ Case 2: N t = 1 and N (t, t + h) = r − 1. (1 arrival before t, the remaining r − 1 happen


in h).

ˆ ...
ˆ Case r: N t = r − 1 and N (t, t + h) = 1. (The r − 1 event occurred before t, and exactly
the rth event occurs in h).

Nt = k N (t, t + h) = r − k
×
0 t th
r eventt + h
t h
49.1. INTRODUCTION 345

2. Summation of Disjoint Events


The probability of the union of these disjoint events is the sum of their individual probabilities:
r−1
X
P [Tr ∈ (t, t + h]] = P [Nt = k, N (t, t + h) = (r − k)] (1671)
k=0

3. Applying Independent Increments


Because the intervals (0, t] and (t, t + h] are disjoint, the number of arrivals in each are
independent:
r−1
X
P [Tr ∈ (t, t + h]] = P [Nt = k] · P [N (t, t + h) = (r − k)] (1672)
k=0

Final Alternative Derivation: The Small Interval Limit


Consider the case where the interval h → 0 (infinitesimally small). In this limit, the
probability of having more than one event in the interval (t, t + h] becomes zero.

1. Simplifying the Event


As h → 0, the only significant case for the rth event to fall in (t, t + h] is when exactly (r − 1)
events have occurred by time t, and exactly 1 event occurs in the small window h.

Nt = r − 1 N (t, t + h) = 1

1st 2nd (r − 1)th rth event


× × ... × ×
0 t t+h

2. Probability Calculation
Using the Independent Increments property:

P [Tr ∈ (t, t + h]] ≈ P [Nt = (r − 1)] · P [N (t, t + h) = 1] (1673)

Substituting the Poisson probabilities:

ˆ P [N = (r − 1)] =
t
e−λt (λt)r−1
(r−1)!

ˆ P [N (t, t + h) = 1] ≈ λh (The probability of 1 event in a very small interval)


3. Transition to PDF
The probability that the rth event occurs in dt = h is given by:
Z t+h
fTr (u) du ≈ fTr (t) · h (1674)
t

Equating both expressions:


e−λt (λt)r−1
fTr (t) · h = · λh (1675)
(r − 1)!
Dividing both sides by h:
λ · λr−1 tr−1 e−λt λr tr−1 e−λt
fTr (t) = = (1676)
(r − 1)! (r − 1)!
346 CHAPTER 49. LECTURE 13 AND 14: POISSON PROCESS PROPERTIES

Gamma Function and Erlang Relationship


The Gamma Function is defined as:
Z ∞
Γ(r) = tr−1 e−t dt
0

If r is an integer, then Γ(r) = (r − 1)!. This allows us to write the Gamma PDF (which is the
Erlang PDF for integer r) as:

λr tr−1 e−λt λr tr−1 e−λt


fΓ (t) = =
Γ(r) (r − 1)!

Discrete Time Markov Chains [DTMC] — 27/03/26


In a DTMC, time and state space are both discrete.

ˆ Time Index: n = 0, 1, 2, . . .
ˆ State Space (S): A discrete set of states S = {0, 1, 2, . . . , i, i + 1, . . . }.
ˆ Random State (X ): The state of the system at time n.
n

The Markov Property


A stochastic process {Xn } is a DTMC if the future is conditionally independent of the past,
given the present:

P [Xn+1 = j | Xn = i, Xn−1 = xn−1 , . . . , X0 = x0 ] = P [Xn+1 = j | Xn = i] = pij (1677)


Note: If this probability pij does not depend on the time index n, the DTMC is called Time
Homogeneous.
Chapter 50

Discrete Time Markov Chains


(DTMC)

Discrete Time Markov Chains (DTMC)

Time is discrete ⇒ time index n = 0, 1, 2, . . .


State space discrete ⇒ S = set of states is discrete.
Time dependent random state:

Xn ∈ S = {0, 1, 2, . . . , i, i + 1, . . . }

Xn = state of system at time n.


S can be finite or countable infinite.
Xn is a stochastic process.

X0 , X1 , X2 , . . . , Xn , . . .

DTMC Definition:
Xn is a DTMC if:

P (Xn+1 = j | Xn = i, Xn−1 , . . . , X0 ) = P (Xn+1 = j | Xn = i) = pij

Interpretation:

ˆ Future → X n+1

ˆ Present → X n

ˆ Past history → X n−1 , . . . , X0

Future is conditionally independent of past given present.


If pij does not depend on n ⇒ Time homogeneous DTMC.

347
348 CHAPTER 50. DISCRETE TIME MARKOV CHAINS (DTMC)

DTMC Sample Path Diagram


Xn states

xn+1 = j

xn = i

x2

Past history x1

x0
present future
Discrete Time
0 1 2 3 5 6 7
Transition is happening here

State Transition Illustration

The placement depends on


what you did at n
not before

n n+1

Example: Two State Model


Suppose system has two states:
S = {H, A}, |S| = 2
H = Happy, A = Angry
Transition diagram:
0.2 0.8
H −−→ A, H −−→ H
0.3 0.7
A −−→ H, A −−→ A
Transition matrix:  
0.8 0.2
P = 
0.3 0.7
349

Row sum = 1 ⇒ stochastic matrix.

P (Xn+1 = H | Xn = H) = 0.8
Tomorrow depends only on today.
Important Note:
ˆ Rows = present state
ˆ Columns = next state
ˆ Two different distributions possible
Higher Order Model Idea
Suppose tomorrow depends on today and yesterday.
Then not a Markov chain.
We expand state:

S = {HH, AH, HA, AA}, |S| = 4


Transition matrix:
 
0.25 0.75 0 0
 
 0 0 0.6 0.4 
 
P =
 

 0.1 0.9 0 0 
 
0 0 0.85 0.15
Consistency condition must hold.

Random Walk (Drunkard’s Walk)


Step right with probability p
Step left with probability q = 1 − p
State space:
S = {−∞, . . . , −1, 0, +1, +2, . . . , +∞}
Transitions:
i → i + 1 with p
i → i − 1 with q
Finite version:
S = {0, ±1, ±2, . . . }
Interpretation:
ˆ X = position at time n
n

ˆ Sample paths are realizations


Drift cases:
ˆ p< 1
2: downward drift
ˆ p> 1
2: upward drift
ˆ p= 1
2: zero drift (symmetric walk)
350 CHAPTER 50. DISCRETE TIME MARKOV CHAINS (DTMC)

Recurrence vs Transience
Infinite states.

ˆ Transient: drift away


ˆ Recurrent: return repeatedly
In 1D and 2D: return possible
Higher dimensions: return probability ≈ 0

Random Walk with Boundaries (Gambling)


Stops at:
i=0 or i = J

State space:
S = {0, 1, 2, . . . , J}, |S| = J + 1

Xn = wealth of gambler at time n

p00 = 1, pJJ = 1

pi,i+1 = p, pi,i−1 = q

pij = 0 otherwise

Game Description
Bet Rs b.
Win ⇒ gain Rs b
Lose ⇒ lose Rs b
Initial wealth = Rs ib
Bias factor:
q P (lose)
α= =
p P (win)

ˆ α > 1: against gambler


ˆ α < 1: favour gambler
ˆ α = 1: fair game
Game stops when:

ˆ reach 0 (ruin)
ˆ reach J (target wealth)
351

Probability of Winning
Let:
Ei = {reach J before 0}

si = P (Ei ) = P (Xn = J | X0 = i)
Using Law of Total Probability:

si = P (Ei |X1 = i + 1)P (X1 = i + 1|X0 = i) + P (Ei |X1 = i − 1)P (X1 = i − 1|X0 = i)
Using Markov property:

si = psi+1 + qsi−1
352 CHAPTER 50. DISCRETE TIME MARKOV CHAINS (DTMC)

Recurrence Relation
si = psi+1 + qsi−1
Let:
q
α=
p

si+1 − si = α(si − si−1 )


Iterating:

(s2 − s1 ) = α(s1 − s0 )

(s3 − s2 ) = α2 (s1 − s0 )

(si − si−1 ) = αi−1 (s1 − s0 )

Summation
si − s0 = (1 + α + α2 + · · · + αi−1 )(s1 − s0 )
Geometric series:

1 − αi
si = (s1 − s0 )
1−α
Boundary conditions:

s0 = 0, sJ = 1

1 − αJ
1= (s1 )
1−α
1−α
s1 =
1 − αJ
Final:

1 − αi
si =
1 − αJ

Special Cases
Case 1: α = 1

si = is1

1
1 = Js1 ⇒ s1 =
J

i
si =
J
Case 2: α < 1
50.1. INTRODUCTION TO DTMC 353

1 − αi
si =
1 − αJ
As J → ∞:

si → 1 − α i
If i → ∞, si → 1
Case 3: α > 1

αi − 1
si =
αJ − 1
As J → ∞:

si → 0

Winning probability → 0

Discrete Time Markov Chain (DTMC)


Lecture 17

sasi Gokul.P
April 2026

50.1 Introduction to DTMC


A Discrete Time Markov Chain (DTMC) is a mathematical framework used to model systems
that evolve over time in a probabilistic manner.
The evolution happens in discrete steps, meaning time moves as:
n = 0, 1, 2, 3, . . .
At each time step, the system occupies one of the possible states.
Core Idea: The most important assumption of a Markov chain is that the future behavior of
the system depends only on its present state and not on how it arrived there.
Markov Property:
P (Xn+1 = j | Xn = i, Xn−1 , . . . , X0 ) = P (Xn+1 = j | Xn = i)
Deep Explanation:
ˆ X = i means the system is currently in state i.
n

ˆ X = j means the system moves to state j in the next step.


n+1

ˆ The left side considers the entire past history.


ˆ The right side ignores the past and depends only on the present.
This property is what makes Markov chains simple and powerful.
Real-life intuition:
ˆ Weather tomorrow depends only on today’s weather.
ˆ Stock price movement depends mainly on current value.
354 CHAPTER 50. DISCRETE TIME MARKOV CHAINS (DTMC)

50.2 State Space


The set of all possible states is called the state space:

S = {1, 2, 3, . . . , n}
Detailed Explanation:

ˆ Each number represents a unique state.


ˆ The system must always be in one of these states.
ˆ No state outside this set is possible.
Example:

ˆ 1 = Sunny
ˆ 2 = Rainy
ˆ 3 = Cloudy
At any time n, the system takes one value from S.

50.3 Transition Probability


Pij = P (Xn+1 = j | Xn = i)
Detailed Explanation:

ˆ This represents the probability of moving from state i to state j in one step.
ˆ For a fixed current state i, there are multiple possible next states.
ˆ Each possible movement has an associated probability.
Important Understanding:

ˆ If P ij is high transition is likely.


ˆ If P ij is low  transition is rare.

50.4 Transition Probability Matrix


 
P11 P12 ··· P1n
 
 P21 P22 · · · P2n 
 
P =
 .. .. .. .. 

 . . . . 
 
Pn1 Pn2 · · · Pnn
Deep Explanation:

ˆ Row i  current state


ˆ Column j  next state
ˆ Entry (i, j)  probability of transition from i to j
50.5. INITIAL DISTRIBUTION 355

Fundamental Property:
n
X
Pij = 1
j=1

Why?

ˆ From state i, the system must go somewhere.


ˆ Total probability of all possibilities must be 1.
50.5 Initial Distribution
 
P (X0 = 1)
 
 P (X0 = 2) 
 
d(0) =

..



 . 

P (X0 = n)
Deep Explanation:

ˆ This describes where the system starts.


ˆ Each entry is the probability of starting in a state.
ˆ All probabilities sum to 1.
Cases:

ˆ Certain start  one value = 1


ˆ Uncertain start  probabilities spread
50.6 State Distribution at Time k
 
P (Xk = 1)
 
 P (Xk = 2) 
 
d(k) =

..



 . 

P (Xk = m)
Deep Explanation:

ˆ This vector shows the probabilities after k steps.


ˆ It evolves over time due to transitions.
50.7 Two-Step Transition Probability
(2)
Pij = P (Xn+2 = j | Xn = i)

Step 1: Think Physically


To go from time n to n + 2, the system must pass through time n + 1.
356 CHAPTER 50. DISCRETE TIME MARKOV CHAINS (DTMC)

Step 2: Consider All Paths

 k) × P (k  j)
m
(2)
X
Pij = P (i
k=1

Step 3: Formal Expansion


m
(2)
X
Pij = Pik Pkj
k=1

Deep Meaning:

ˆ Try all intermediate states


ˆ Multiply probabilities along path
ˆ Add all possibilities
Matrix Form
P (2) = P 2
Insight: Matrix multiplication naturally performs this calculation.

50.8 n-Step Transition Probability


P (k) = P k
Explanation:

ˆ Apply transition repeatedly


ˆ Each multiplication = one step
50.9 State Evolution Formula
d(k) = d(0) P k
Deep Explanation:

ˆ Start with initial probabilities


ˆ Apply transitions repeatedly
ˆ Observe how distribution changes
50.10 Chapman-Kolmogorov Equation
P (k+ℓ) = P (k) P (ℓ)
Deep Explanation:

ˆ Large transitions can be broken into smaller steps


ˆ Useful for simplifying complex problems
50.11. CLASSIFICATION OF STATES 357

50.11 Classification of States


Recurrent
Returns with probability 1.

Transient
May not return.

Positive Recurrent
Returns quickly (finite time).

Null Recurrent
Returns slowly (infinite expected time).

50.12 Final Intuition


Markov Chains help us understand:

ˆ How systems evolve over time


ˆ Long-term behavior
ˆ Stability and steady states
358 CHAPTER 50. DISCRETE TIME MARKOV CHAINS (DTMC)
Chapter 51

Lecture 18: Markov Chains in


matrix form

Problem
Given Markov Chain:

1 0.3

0.2
R3 T1 0.2 R1
0.2

0.6 0.7 0.7

T2 0.2 R2

0.4

Transient state

P (Xn = R3 | X0 = R3 ) = 1

fi = Prob. that starting at i, DTMC will ever return to state i


Given: {Xn }
 
fi = P X0 = i ∪ · · · ∪ Xn = i ∪ Xn+1 = i | X0 = i

 
= P ∪n≥1 {Xn = i} | X0 = i
If i is recurrent ⇒ fi = 1
If i is transient fi < 1 ⇒ (1 − fi ) > 0 probability of escaping
⇒ I want show R1 is recurrent ⇒ fR1 = 1
First you have to define an event:
We are starting at X0 = R1

n=0: P (X1 = R1 | X0 = R1 ) = 0.3

359
360 CHAPTER 51. LECTURE 18: MARKOV CHAINS IN MATRIX FORM

n=1: P (X2 = R1 , X1 ̸= R1 | X0 = R1 ) = (0.7)(0.6) + (0.7)(0.4)(0.6)

n=2: P (X3 = R1 , X2 ̸= R1 , X1 ̸= R1 | X0 = R1 ) = (0.7)(0.4)(0.6)

n=n: P (Xn+1 = R1 , Xn ̸= R1 , . . . , X1 ̸= R1 | X0 = R1 ) = (0.7)(0.4)n−1 (0.6)


Now moving to calculate fi

fi = P (return)


X
fi = P (Xn+1 = R1 , Xn ̸= R1 , . . . , X2 ̸= R1 , X1 ̸= R1 | X0 = R1 )
n=0

X
= 0.3 + (0.7)(0.4)n−1 (0.6)
n=1

X
= 0.3 + (0.7)(0.6) (0.4)n−1
n=1

(infinite geometric series)


 
1
= 0.3 + (0.7)(0.6)
1 − 0.4
 
1
= 0.3 + (0.7)(0.6)
0.6

= 0.3 + 0.7

= 1.00

⇒ fR1 = 1
⇒ Now what about fT1 ?

n=0: P (X1 = T1 | X0 = T1 ) = 0

n=1: P (X2 = T1 , X1 ̸= T1 | X0 = T1 ) = (0.6)(0.6)


X
= (0.6)(0.6)
n=1

= 0.36

fT1 = (0.6)(0.6) = 0.36

⇒ Prob. of escaping from T1 = 1 − 0.36 = 0.64 > 0


361

fT2 = (0.6)(0.6) = 0.36


Escape from T1 :

(T1 → R1 ) ∪ (T1 → R3 ) ∪ (T1 → T2 → R2 ) ∪ (T1 → T2 → R3 )

P (escaping from T1 ) = 0.2 + 0.2 + (0.6)(0.2) + (0.6)(0.2)

= 0.4 + 0.12 + 0.12

= 0.64

⇒ Transient states

IFriendhsip btw states


i and j are 2 states. i, j ∈ S that are friends with each other if they communicate with each
other in both ways.

i↔j
Two way communication

i→j and j → i
One way communication
(n)
(i → j) : pij > 0 for some n

(m)
(j → i) : pji > 0 for some m
Two way communication is an equivalence relation on a set of states S.
It has following properties:
(i) Reflexive:
(0)
i ↔ i since pii = 1
(ii) Symmetric:
If i ↔ j then j ↔ i
(n) (m)
i ↔ j ⇒ pij > 0 and pji > 0

(m) (n)
⇒ pji > 0 and pij > 0
(iii) Transitive: for i, j, k ∈ S
If i ↔ j and j ↔ k then i ↔ k

i → j, j→k

(n) (m)
⇒ pij > 0, pjk > 0
Take t = n + m
362 CHAPTER 51. LECTURE 18: MARKOV CHAINS IN MATRIX FORM

(n+m) (n) (m)


X
pik = pij pjk
j∈S

(n+m)
⇒ pik >0
Similarly,

k → j, j→i

(m) (n)
⇒ pkj > 0, pji > 0
Take t = n + m
(n+m) (m) (n)
X
pki = pkj pji
j∈S

(n+m)
⇒ pki >0

∴ i↔k
As we did before

S = {T1 , T2 , R1 , R2 , R3 }

{T1 , T2 }, {R3 }, {R1 , R2 }


There are 3 equivalence classes on S.
Calculate the no. of visits in state i:
Number of visits to state i given that X0 = i

X
Ni = I(Xn = i | X0 = i)
n=1

= I(X1 = i | X0 = i) + I(X2 = i | X0 = i) + · · ·
( (n)
1 if Xn = i w.p. pii
I(Xn = i | X0 = i) = (n)
0 if Xn ̸= i w.p. 1 − pii

What is the probability of revisiting state i exactly n times?

P (Ni = n) = fin (1 − fi )
(Geometric)

fi < 1 ⇒ i is transient

Geometric distribution:

P (Y = k) = (1 − p)k−1 p

1
E(Y ) =
p
363

Ni ∼ geom(1 − fi )

1
E[Ni ] =
1 − fi

If i is recurrent ⇒ fi = 1

1 1
E[Ni ] = = =∞
1−1 0

Ni ∼ geom(1 − fi )

1
E[Ni ] = , fi < 1
1 − fi
Now from this definition,

X
Ni = I(Xn = i | X0 = i)
n=1

" #
X
E[Ni ] = E I(Xn = i | X0 = i)
n=1

X
= E[I(Xn = i | X0 = i)]
n=1

X
= P (Xn = i | X0 = i)
n=1

(n)
X
= pii
n=1

1 X (n)
⇒ E[Ni ] = = pii
1 − fi
n=1

If i is transient:

1
E[Ni ] = <∞
1 − fi

(n)
X
⇒ pii < ∞
n=1

If i is recurrent:

(n)
X
E[Ni ] = pii = ∞
n=1

Theorem
If state i is recurrent and i ↔ j then state j is recurrent.
Proof:
364 CHAPTER 51. LECTURE 18: MARKOV CHAINS IN MATRIX FORM

i ↔ j ⇒ i → j and j → i

(n) (m)
pij > 0 and pji > 0 for some n, m
Then for any n > 0 we have
To show j ∈ S is recurrent:

(t)
X
pjj = ∞
t=1

l steps
n steps m steps
j i j

Choose t ∈ (l + n + m)

(l+n+m) (n) (l) (m)
X
pjj ≥ pji pii pij
n=1

>0 >0 >0


" #
(n) (n) (m)
X
= pji pii pij
n=1

= ∞ since i is recurrent

⇒∞

That means j is recurrent


sectionStochastic Processes — Fundamentals

51.0.1 Introduction and Notation


Definition 1.1 — Stochastic Process
Let Xn be a random variable at discrete time n, where

n = 0, 1, 2, . . .

The notation Xn = i means the process is in state i at time n.


The general conditional distribution

P (Xn | Xn−1 , Xn−2 , . . . , X1 )

defines the full stochastic process: the current state depends on all past states.
365

Definition 1.2 — Markov Chain (Markov Property)

A stochastic process {Xn } is called a Markov Chain if it satisfies the Markov Property:

P (Xn+1 | Xn ) (future depends only on present, NOT on past)

That is,
P (Xn+1 = j | Xn = i, Xn−1 , . . . , X0 ) = P (Xn+1 = j | Xn = i).

51.0.2 Transition Matrix


Definition 1.3 — Transition Probability Matrix

The one-step transition matrix P has entries

pij = P (Xn+1 = j | Xn = i),


X
where each row sums to 1: pij = 1 for all i.
j

Example from notes: A two-state chain with states {0, 1}:

future 0 future 1
 
present 0 0.5 0.5
P =
present 1 0.5 0.5

Both rows sum to 0.5 + 0.5 = 1.

Problem Statement
Suppose it rains today and will rain tomorrow with probability α. If it does not rain
today, it will rain tomorrow with probability β.
Let Xn denote whether it rains on the n-th day:
(
0 rains on day n
Xn =
1 does not rain on day n

51.0.3 Transition Probabilities

P (Xn+1 = 0 | Xn = 0) = α (rain → rain)


P (Xn+1 = 1 | Xn = 0) = 1 − α (rain → no rain)
P (Xn+1 = 0 | Xn = 1) = β (no rain → rain)
P (Xn+1 = 1 | Xn = 1) = 1 − β (no rain → no rain)

These are also written in shorthand as:

p00 = α, p01 = 1 − α, p10 = β, p11 = 1 − β.


366 CHAPTER 51. LECTURE 18: MARKOV CHAINS IN MATRIX FORM

51.0.4 Transition Matrix


0 1
 
0 α 1−α
P =
1 β 1−β
1−α

State 0 State 1
α 1−β
Rain No Rain

51.0.5 Non-Markovian Counter-example from Notes


The notes also record a non-Markovian process where Xn captures rain on both the
(n − 1)-th and n-th day simultaneously. The states are:

State Meaning

0 Rained on both (n − 1)-th and n-th day


1 Rained on n-th day but not on (n − 1)-th day
2 Rained on (n − 1)-th day but not on n-th day
3 Did not rain on either day

The 4 × 4 transition matrix recorded in the notes:

0 1 2 3
 
0 0.7 0 0.3 0
1 0.5 0 0.5 0 
P =  
2  0 0.4 0 0.6 
3 0 0.2 0 0.8

51.0.6 Two-Step Transition Matrix


Question: Given that it rained on Monday and Tuesday, what is the probability it will rain
on Thursday?
Since Monday = day n, Tuesday = day n + 1, Thursday = day n + 3, we need the two-step
transition matrix P (2) = P 2 .
Starting from state 0 (rained on both consecutive days), the two-step matrix:

0 1 2 3
 
0 0.43 0.12 0.21 0.18
1  0.35 0.20 0.15 0.30 
P (2) = P2 =  
2  0.20 0.12 0.20 0.48 
3 0.10 0.16 0.10 0.64

Answer: Given it rained on Monday and Tuesday, we start in state 0. The probability
of raining on Thursday (two steps later, state 0) is:
(2)
P00 = 0.43
51.1. NUMERICAL 2 — GAMBLER’S RUIN PROBLEM 367

Verification of matrix multiplication (row 0 of P 2 ):


X
(P 2 )0j = P0k Pkj
k

For (P 2 )00 :

= 0.7 × 0.7 + 0 × 0.5 + 0.3 × 0 + 0 × 0 = 0.49 + 0 + 0 + 0 = 0.49

51.1 Numerical 2 — Gambler’s Ruin Problem


51.1.1 Problem Statement
Gambler’s Ruin
A gambler plays a sequence of games. Let Xn = gambler’s wealth after the n-th game.
He quits playing when either:

ˆ He is broke (wealth = 0), or


ˆ He attains a fortune of $N (wealth = N ).
States: {0, 1, 2, . . . , N }.

51.1.2 Transition Probabilities


At each game (for interior states i = 1, 2, . . . , N − 1):

pi,i+1 = p (win: wealth increases by 1) (1678)


pi,i−1 = 1 − p (lose: wealth decreases by 1) (1679)

States 0 and N are absorbing states:

p00 = 1, pN N = 1.

51.1.3 Transition Matrix


0 1 2 ··· N −1 N
 
0 1 0 0 ··· 0 0
1 1 − p
 0 p ··· 0 0 

2  0 1 − p 0 p ··· 0 
P = 3  0
 
 0 1−p 0 p ··· 

..  .. .. 
.  . . 
N 0 0 0 ··· 0 1
notation: pi,i+1 = p and pi,i−1 = 1 − p for i = 1, 2, . . . , N − 1.

51.2 Naive Bayes Classifier


51.2.1 Introduction
The Naive Bayes classifier is a generative probabilistic model used for classification. The
notes focus on email spam detection as the running example.

51.2.2 Setup and Notation


368 CHAPTER 51. LECTURE 18: MARKOV CHAINS IN MATRIX FORM

Setup

ˆ Y ∈ {0, 1}: email label — 0 = good (ham), 1 = spam.


ˆ A fixed dictionary (vocabulary) of n words: {w , w , . . . , w }.
1 2 n

ˆ X⃗ = (X , X , . . . , X ) : a random vector representing email j.


(j)
1
(j)
2
(j) ⊤
n

ˆ Each feature:
(
(j) 1 if word i (wi ) is present in email j
Xi =
0 if word i (wi ) is absent from email j

Example. Vocabulary: V = {cat, dog, sat, ran}.


Email1 = “dog ran”:
⃗x(1) = (0, 1, 0, 1)⊤

(cat absent, dog present, sat absent, ran present.)


Email2 (from notes):
⃗x(2) = (0, 1, 0, 1)⊤

51.2.3 The Classification Goal


We want to compute the posterior probability:

h i
⃗ (j) = ⃗x(j) = ?
P Y = y (j) X

i.e., given the word-vector of email j, what is the probability it belongs to class y?

51.2.4 Bayes’ Theorem (Derivation)


By Bayes’ theorem:

Theorem 4.1 — Bayes’ Theorem for Classification


h i
h i P X ⃗ = ⃗x Y = y P [Y = y]
P Y =y X ⃗ = ⃗x = h i
P X⃗ = ⃗x

⃗ = ⃗x]
P [Y = y | X Posterior
⃗ = ⃗x | Y = y]
P [X Likelihood
P [Y = y] Prior
⃗ = ⃗x]
P [X Normalisation (evidence)

⃗ = ⃗x] is the same for all classes and acts as normalisation.


The denominator P [X
51.2. NAIVE BAYES CLASSIFIER 369

51.2.5 Generative Model

Definition 4.1 — Generative Process


The Naive Bayes model assumes:

1. The class label Y is drawn from a prior: P [Y = y].


⃗ is generated independently:
2. Given Y = y, the feature vector X
n
Y
⃗ = ⃗x | Y = y] =
P [X P [Xi = xi | Y = y] (Naive/independence assumption)
i=1

Diagrammatically (from notes):

generates
Y −−−−−→ X1 , X2 , . . . , Xn

51.2.6 ⃗
Distribution of X

Since each Xi ∈ {0, 1}, the vector ⃗x takes values in {0, 1}n . There are 2n possible email
vectors.

⃗x follows a Multinomial distribution over 2n outcomes.


Let P2k denote the probability of the last outcome. Then:

P2n = 1 − (P1 + P2 + P3 + · · · + P2n −1 )

Total number of parameters (for both classes y ∈ {0, 1}):

2 (2n − 1) + 1

For n = 40,000 words: number of parameters ≈ 240000 — completely intractable.

Why Naive Bayes reduces this. By the conditional independence assumption, the
number of parameters reduces to:

2n
|{z} + |{z}
1 = 2n + 1
class-conditional prior

For n = 40,000: only 80,001 parameters — tractable.

51.2.7 Worked Example — Spam Classifier

Dataset Setup

Vocabulary: V = {money, free, hello} (n = 3). So 23 = 8 possible emails.


370 CHAPTER 51. LECTURE 18: MARKOV CHAINS IN MATRIX FORM

Email money free hello Label

e0 0 0 0 —
e1 1 0 0 —
e2 0 1 0 —
e3 1 1 0 —
e4 0 0 1 —
e5 0 1 1 —
e6 1 1 1 —
e7 1 1 1 —

Training Data
From the notes, the training set contains 5 spam emails:

Email money free hello Label

g1 0 0 1 spam
g2 0 0 1 spam
g3 0 0 0 spam
g4 0 1 1 spam
g5 0 0 1 spam

And the good (ham) emails have the class-conditional probabilities as listed in Section 4.7.

Class-Conditional Likelihoods from Notes

Email vector ei P (ei | good) P (ei | spam)

e0 1/5 0
e1 0 1/5
e2 3/5 0
e3 0 2/5
e4 1/5 0
e5 0 0

Applying Bayes’ Theorem


For a new email ⃗x∗ , the classifier assigns:
⃗ = ⃗x∗ ]
ŷ = arg max P [Y = y | X
y∈{0,1}

Expanding:
⃗ = ⃗x∗ | Y = y] P [Y = y]
ŷ = arg max P [X
y
51.3. QUICK REFERENCE — NOTATION TABLE 371

⃗ = ⃗x∗ ] is the same for both classes and can be dropped).


(the denominator P [X

51.2.8 Parameter Counting — Tractability Analysis

Model Parameters n = 40,000

Full joint 2(2n − 1) + 1 ≈ 240000 (intractable)


Naive Bayes 2n + 1 80,001 (tractable)

The ratio of parameters is:

2(2n − 1) + 1 2n n→∞
≈ −−−→ ∞
2n + 1 n

51.3 Quick Reference — Notation Table

Symbol Meaning

Xn Random variable (state) at time n


pij Transition probability from state i to state j in one step
P Transition probability matrix (rows sum to 1)
P (k) = P k k-step transition matrix
α Probability of rain given it rained today
β Probability of rain given it did not rain today
N Target fortune in Gambler’s Ruin
p Win probability per game in Gambler’s Ruin
Y Email class label (0=ham, 1=spam)
⃗ (j)
X Feature vector (word indicator) for email j
(j)
Xi 1 if word wi present in email j; 0 otherwise
n Vocabulary size (number of words)
V Vocabulary set {w1 , . . . , wn }
P [Y = y] Prior probability of class y
⃗ = ⃗x | Y = y]
P [X Likelihood of email ⃗x given class y
⃗ = ⃗x]
P [Y = y | X Posterior probability (output of classifier)

You might also like