0% found this document useful (0 votes)
13 views361 pages

TIA Exam P Probability Course Handouts

The document provides lesson handouts for the TIA Exam P online course, authored by David Revelle, PhD. It covers various topics in discrete probability, including fundamentals, conditional probability, discrete moments, combinatorics, key distributions, and deductibles and limits. Each section is detailed with subtopics and page numbers for easy navigation.

Uploaded by

vicvo123456789
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
13 views361 pages

TIA Exam P Probability Course Handouts

The document provides lesson handouts for the TIA Exam P online course, authored by David Revelle, PhD. It covers various topics in discrete probability, including fundamentals, conditional probability, discrete moments, combinatorics, key distributions, and deductibles and limits. Each section is detailed with subtopics and page numbers for easy navigation.

Uploaded by

vicvo123456789
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

TIA Exam P Online Course

Lesson Handouts

David Revelle, PhD

© 2025 The Infinite Actuary, LLC


Exam P Lesson Handouts

A. Discrete Probability 3
A.1 Fundamentals of Probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
A.1.1 Fundamentals of Probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
A.1.2 Complements . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
A.1.3 Venn Diagrams . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
A.1.4 De Morgan’s Laws . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17
A.1.5 Inclusion-Exclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22
A.2 Conditional Probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
A.2.1 Conditional Probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
A.2.2 Independence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29
A.2.3 Sequences of Events . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32
A.2.4 Bayes’ Theorem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38
A.3 Discrete Moments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44
A.3.1 Mode . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44
A.3.2 Median . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 48
A.3.3 Moments: Expected Value/Mean . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 52
A.3.4 Tools for Finding Means . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 55
A.3.5 Means: Survival Function Approach . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 58
A.3.6 Variance: Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 62
A.3.7 Variance: Tools . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 65
A.3.8 Discrete Uniform Random Variables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69
A.4 Combinatorics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 73
A.4.1 Permutations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 73
A.4.2 Combinations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 78
A.4.3 The Binomial Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83
A.4.4 Multinomial Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 88
A.4.5 Hypergeometric Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 91
A.5 Key Discrete Distributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 96
A.5.0a Geometric Series . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 96
A.5.0b Taylor Series for exp(x) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 100
A.5.1 The Geometric Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 102
A.5.2 Memoryless Property . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 106
A.5.3 Negative Binomial Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 110
A.5.4 The Poisson Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 114
A.5.5 Sums of Independent Poissons . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 119
A.6 Deductibles and Limits . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 122
A.6.1 Deductibles . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 122
A.6.2 Policy Limits . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 126
A.6.3 Calculator Approach to Deductible . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 130
A.7 Discrete Review . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 133
A.7.1 Discrete Review . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 133

B. Continuous Probability 139

2
A.1.1 Fundamentals of Probability Exam P Handouts – Page 3

B.0 1-d Calculus Review 139


B.0.1 Differentiation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 139
B.0.2 Integration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 160
B.0.3 Integration by Parts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 181
B.1 Densities and CDFs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 202
B.1.1 Continuous Distributions: Overview . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 202
B.1.2 Densities and CDFs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 207
B.1.3 Mixed Distributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 213
B.2 Continuous Moments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 220
B.2.1 Moments of Continuous Distributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 220
B.2.2 Moments of Mixed Distributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 225
B.2.3 The Survival Function Approach . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 230
B.3 Key Continuous Distributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 237
B.3.1 Continuous Uniform - Basics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 237
B.3.2 Continuous Uniform - Exam Concepts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 243
B.3.3 Exponential Random Variables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 249
B.3.4 Gamma, Exponential, and Poisson . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 256
B.3.5 Beta and Pareto Distributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 264
B.4 Normal Approximations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 270
B.4.1 Normal Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 270
B.4.2 Interpolation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 277
B.4.3 The Central Limit Theorem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 283
B.4.4 Continuity Correction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 289
B.4.5 Lognormal Distributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 294
B.5 Continuous Review . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 299
B.5.1 Deductibles: Calculus Approach . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 299
B.5.2 Deductibles: Cases Approach . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 304
B.5.3 Review of Other Continuous Ideas . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 307

C. Multi-Variate Probability 310


C.1 Joint Distributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 310
C.1.1 Joint Distributions and CDFs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 310
C.1.2 Marginal and Conditional Distributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 315
C.2 Joint Moments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 322
C.2.1 Joint Moments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 322
C.2.2 Covariances and Correlations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 327
C.2.3 Conditional Moments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 335
C.3 Order Statistics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 342
C.3.1 Order Statistics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 342
C.3.2 General Order Stats . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 349
C.4 Multivariate Review . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 356
C.4.1 Multivariate Review . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 356
A.1.1 Fundamentals of Probability Exam P Handouts – Page 4

Fundamentals of Probability 1

Basic Principles
Unions and Intersections
Exercise

Basic Principles of Probability 2

1. 0 ≤ P[A] ≤ 1
2. P[S] = 1 P[∅] = 0
3. If A1 ∩ A2 = ∅, then P[A1 ∪ A2 ] = P[A1 ] + P[A2 ]
Example 1
Roll a fair six sided die
S = “Sample Space” = all possible outcomes = {1, 2, 3, 4, 5, 6}
4
A = {3, 4, 5, 6} P[A] =
6
A1 = {3, 4, 5} A2 = {6} A1 ∪ A2 = A A1 ∩ A2 = ∅
3 1
P[A1 ] = P[A2 ] = P[A1 ] + P[A2 ] = P[A]
6 6

4
A.1.1 Fundamentals of Probability Exam P Handouts – Page 5

Basic Principles of Probability 3

1. 0 ≤ P[A] ≤ 1
2. P[S] = 1 P[∅] = 0
3. If A1 ∩ A2 = ∅, then P[A1 ∪ A2 ] = P[A1 ] + P[A2 ]
Example 2
Pick a point “at random” from S, where total area of S = 1
P[A] = area of A A1 ∩ A2 = ∅, A1 ∪ A2 = A
Area of A = area of A1 + area of A2
P[A] = P[A1 ] + P[A2 ]

A
S

A1 A2

Unions and Intersections 4

1. 0 ≤ P[A] ≤ 1
2. P[S] = 1 P[∅] = 0
3. P[A ∪ B] = P[A] + P[B] − P[A ∩ B]
Example 1
Roll a fair six sided die
S = “Sample Space” = all possible outcomes = {1, 2, 3, 4, 5, 6}
3
A = “roll an odd number” = {1, 3, 5} P[A] =
6
3
B = “roll a 3 or less” = {1, 2, 3} P[B] =
6
A ∪ B = {1, 2, 3, 5} A ∩ B = {1, 3}
4 3 3 2
P[A ∪ B] = = + − = P[A] + P[B] − P[A ∩ B]
6 6 6 6

5
A.1.1 Fundamentals of Probability Exam P Handouts – Page 6

Unions and Intersections 5

1. 0 ≤ P[A] ≤ 1
2. P[S] = 1 P[∅] = 0
3. P[A ∪ B] = P[A] + P[B] − P[A ∩ B]
Example 2
Pick a point “at random” from S, where total area of S = 1
P[A] = area of A.
Area of A ∪ B = area of A + area of B - area of A ∩ B
P[A ∪ B] = P[A] + P[B] − P[A ∩ B]

A B
S

Exercise 6

The probability that a visit to a primary care physician’s (PCP) office results in either
lab work or referral to a specialist is 60%. Of those coming to a PCP’s office, 35% are
referred to specialists and 50% require lab work. Determine the probability that a visit
to a PCP’s office results in both lab work and referral to a specialist.

6
A.1.1 Fundamentals of Probability Exam P Handouts – Page 7

Exercise 6

The probability that a visit to a primary care physician’s (PCP) office results in either
lab work or referral to a specialist is 60%. Of those coming to a PCP’s office, 35% are
referred to specialists and 50% require lab work. Determine the probability that a visit
to a PCP’s office results in both lab work and referral to a specialist.

Let L denote the event that the trip requires lab work, and S the event that it results
in a referral to a specialist.
We are given:
P[L ∪ S] = 0.60 P[S] = 0.35 P[L] = 0.50
and we want P[L ∩ S]
P[L ∪ S] = P[S] + P[L] − P[L ∩ S]
0.60 = 0.35 + 0.50 − P[L ∩ S]
P[L ∩ S] = 0.25

7
A.1.2 Complements Exam P Handouts – Page 8

Complements 1

Recap
Complements
Exercise

Recap 2

Previously, saw that

P[A ∪ B] = P[A] + P[B] − P[A ∩ B]

Also had special case: If A ∩ B = ∅

P[A ∪ B] = P[A] + P[B]

Want now to look at another important special case.

8
A.1.2 Complements Exam P Handouts – Page 9

Complements 3

S
A
A0

Ac = A0 = everything in S but not in A.


A ∩ A0 = ∅ A ∪ A0 = S

E.g., if we roll a fair six sided die and A = roll a 3 or higher,


then S = {1, 2, 3, 4, 5, 6}, A = {3, 4, 5, 6} and A0 = {1, 2}

Complements 4

S
A B

A0

A ∩ A0 = ∅ A ∪ A0 = S
P[A] + P[A0 ] = P[A ∪ A0 ] = P[S] = 1
P[A0 ] = 1 − P[A]

Also, (A0 )0 = A so P[A] = 1 − P[A0 ]


B = (A ∩ B) ∪ (A0 ∩ B) and (A ∩ B) ∩ (A0 ∩ B) = ∅
so P[B] = P[A ∩ B] + P[A0 ∩ B]

9
A.1.2 Complements Exam P Handouts – Page 10

Exercise 5

In a group of patients with shoulder injuries, 20% visit a physical therapist but not a
chiropractor, and 75% visit at least one of these. The probability that a patient visits a
chiropractor exceeds by 0.1 the probability of visiting a physical therapist. Find the
probability that a randomly chosen member of this group visits a physical therapist.

Exercise 5

In a group of patients with shoulder injuries, 20% visit a physical therapist but not a
chiropractor, and 75% visit at least one of these. The probability that a patient visits a
chiropractor exceeds by 0.1 the probability of visiting a physical therapist. Find the
probability that a randomly chosen member of this group visits a physical therapist.

Let T = visit a physical therapist, and C = visit a chiropractor.

P[T ∩ C 0 ] = 0.20 P[T ∪ C ] = 0.75


P[C ] = 0.1 + P[T ]
P[T ∪ C ] = P[C ] + P[T ] − P[T ∩ C ]
0.75 = 0.1 + P[T ] + P[T ] − (P[T ] − P[T ∩ C 0 ])
0.75 = 0.1 + P[T ] + P[T ∩ C 0 ]
P[T ] = 0.75 − 0.1 − 0.2
P[T ] = 0.45

10
A.1.3 Venn Diagrams Exam P Handouts – Page 11

Venn Diagrams 1

Venn Diagrams
Examples
Mutually Exclusive Events
Exercises

Venn Diagrams 2

Venn diagrams are a visual way to represent unions and intersections.

They allow us to easily deal with more complicated cases than what we have seen so
far.

11
A.1.3 Venn Diagrams Exam P Handouts – Page 12

Example 1 3

A survey of a group’s viewing habits found the following:


1. 31% watched gymnastics, 25% watched baseball, and 21% watched soccer
2. 11% watched gymnastics and baseball, 8% watched baseball
and soccer, and 9% watched gymnastics and soccer
3. 5% watched all three sports.
Find the percentage of the group that watched none of the three sports.
G .11 − .05 B
.16 .11
= 0.06

.04 .05 .03

.21 − .12 = .09


S

Example 1 4

G .11 − .05 B
.16 .11
= 0.06

.04 .05 .03

.21 − .12 = .09


S

We want the probability that someone watches none of these sports, so we want

1 − 0.16 − 0.06 − 0.11 − 0.04 − 0.05 − 0.03 − 0.09 = 0.46

12
A.1.3 Venn Diagrams Exam P Handouts – Page 13

Example 2 5

In a company’s health care plan, employees may choose exactly two of the coverages
A, B, and C, or they may choose none of them. The proportions of the employees that
choose A, B, and C are 1/4, 1/3, and 5/12, respectively. Find the probability that a
randomly chosen employee will choose none of A, B, or C.

A B
0 x 0

1 1
4 −x 0 3 −x

0
C

SOA #15; S.01.31 6

A B
0 x 0

1 1
4 −x 0 3 −x

0
C

  
5 1 1 7 1
P[C ] = = + −x −x = − 2x, x=
12 4 3 12 12
   
0 1 1
P[(A ∪ B ∪ C ) ] = 1 − x − −x − −x
4 3
1 3−1 4−1 1
=1− − − =
12 12 12 2

13
A.1.3 Venn Diagrams Exam P Handouts – Page 14

Example 2 7

Alternatively, P[A] + P[B] + P[C ] counts each person who buys coverage twice since
each person buys exactly 2 coverages. P[A ∪ B ∪ C ] counts each person who buys
coverage once, so

P[A] + P[B] + P[C ] = 2 · P[A ∪ B ∪ C ]


1 1 5
+ + = 2 · P[A ∪ B ∪ C ]
4 3 12
1
= P[A ∪ B ∪ C ]
2
1 1
P[(A ∪ B ∪ C )0 ] = 1 − =
2 2

Mutually Exclusive Events 8

A and B are said to be mutually exclusive (or mutually disjoint) events if


A ∩ B = AB = ∅. Note that this means that P[A ∩ B] = P[AB] = 0.

From the point of view of a Venn diagram, we can draw them as non-overlapping sets.

A B

If we have 4 or more events, typically some of them will be mutually exclusive to


simplify any Venn diagrams we need to draw.

14
A.1.3 Venn Diagrams Exam P Handouts – Page 15

Exercise 9

In a survey of gaming habits:


1. 30% played TLOU, 25% played Celeste, 10% played FFVII and 29% played BG3.
2. No one played both TLOU and FFVII, and no one played both Celeste and BG3.
3. 13% played only Celeste, 26% played only BG3, but no one played only FFVII.
Find the percentage of the group that played none of the four.

Exercise 9

In a survey of gaming habits:


1. 30% played TLOU, 25% played Celeste, 10% played FFVII and 29% played BG3.
2. No one played both TLOU and FFVII, and no one played both Celeste and BG3.
3. 13% played only Celeste, 26% played only BG3, but no one played only FFVII.
Find the percentage of the group that played none of the four.
T

0.03 − x 0.25 0.02 + x

B 0.26 0.13 C

x 0 0.1 − x

15
A.1.3 Venn Diagrams Exam P Handouts – Page 16

Exercise (cont) 10

0.03 − x 0.25 0.02 + x

B 0.26 0.13 C

x 0 0.1 − x

F
We want the probability that someone plays none of these 4, which is

1 − (0.03 − x + 0.25 + 0.02 + x) − 0.26 − 0.13 − (x + 0.1 − x)


= 1 − 0.30 − 0.26 − 0.13 − 0.1
= 0.21

16
A.1.4 De Morgan’s Laws Exam P Handouts – Page 17

De Morgan’s Laws 1

Key Words
De Morgan’s Laws
Exercise

Key Words 2

For us, A or B means union (A ∪ B), i.e., A ∪ B means A or B or both.


In words, or is the inclusive or.

A and B means intersection (A ∩ B, also denoted as AB)


A or B but not both means we want to start with A ∪ B but exclude the intersection.

A B

Algebraically, that means we want (A ∩ B 0 ) ∪ (A0 ∩ B)


What about neither A nor B?

17
A.1.4 De Morgan’s Laws Exam P Handouts – Page 18

De Morgan’s Laws: 2 Sets 3

A∪B A0 ∩ B 0 = (A ∪ B)0

A B

A B

The left hand diagram is A or B, while the right hand diagram is neither A nor B.

De Morgan’s Laws: 3 Sets 4

A∪B ∪C A0 ∩ B 0 ∩ C 0 = (A ∪ B ∪ C )0

A B
A B

C C

18
A.1.4 De Morgan’s Laws Exam P Handouts – Page 19

De Morgan’s Laws 5

General version of De Morgan’s Laws:

" k
#0 " k
#
[ \
Ai = A0i
i=1 i=1
" k
#0 " k
#
\ [
Ai = A0i
i=1 i=1

If you bring the complement inside the brackets, unions and intersections get flipped.

Example 6

An auto insurance company has 10,000 policyholders. Each policyholder is classified as


1. young or old; and
2. male or female;
2,000 policyholders are neither young nor male, and 6,000 are either young or female.
How many policyholders are young?

Let Y denote young policyholders, and M male policyholders.


Neither young nor male is (Y ∪ M)0 = Y 0 ∩ M 0 .
We are given P[(Y ∪ M)0 ] = P[Y 0 ∩ M 0 ] = 0.2 and P[Y ∪ M 0 ] = 0.6.

We want P[Y ].

19
A.1.4 De Morgan’s Laws Exam P Handouts – Page 20

Example 7

P[Y 0 ∩ M 0 ] = 0.2
P[Y ∪ M 0 ] = 0.6 Y M
P[(Y ∪ M 0 )0 ] = 1 − 0.6 0.4
0
P[Y ∩ M] = 0.4
P[Y 0 ] = P[Y 0 ∩ M 0 ] + P[Y 0 ∩ M] 0.2
P[Y 0 ] = 0.2 + 0.4 = 0.6
P[Y ] = 1 − 0.6 = 0.4

Exercise 8

You are given:


• P[A0 ∪ B 0 ∪ C ] = 0.7
• P[A ∩ B ∩ C ] = 0.1
Find P[A ∩ B]

20
A.1.4 De Morgan’s Laws Exam P Handouts – Page 21

Exercise 8

You are given:


• P[A0 ∪ B 0 ∪ C ] = 0.7
• P[A ∩ B ∩ C ] = 0.1
Find P[A ∩ B]

P[A ∩ B ∩ C ] = 0.1
0
A ∩ B ∩ C 0 = A0 ∪ B 0 ∪ C
P[A ∩ B ∩ C 0 ] = 1 − 0.7 = 0.3
P[A ∩ B] = P[A ∩ B ∩ C ] + P[A ∩ B ∩ C 0 ]
= 0.1 + 0.3
= 0.4

21
A.1.5 Inclusion-Exclusion Exam P Handouts – Page 22

Inclusion-Exclusion 1

Review of Key Formulas


Inclusion-Exclusion
Exercise

Formulas from before 2

S = sample space, P[S] = 1, P[∅] = 0

P[A ∪ B] = P[A] + P[B] − P[A ∩ B]

Special case: if A1 ∩ A2 = ∅ then P[A1 ∩ A2 ] = 0


so P[A1 ∪ A2 ] = P[A1 ] + P[A2 ] − 0 = P[A1 ] + P[A2 ]

Generalization: if Ai ∩ Aj = ∅ for all i 6= j then


" k
# k
[ X
P Ai = P[Ai ]
i=1 i=1

Complements: A ∩ A0 = ∅, A ∪ A0 = S
1 = P[S] = P[A] + P[A0 ]
P[A] = 1 − P[A0 ] P[A0 ] = 1 − P[A]

22
A.1.5 Inclusion-Exclusion Exam P Handouts – Page 23

Inclusion-Exclusion 3

P[A ∪ B] = P[A] + P[B] − P[AB]


P[A ∪ B ∪ C ] = P[A] + P[B] + P[C ]
− P[AB] − P[AC ] − P[BC ]
+ P[ABC ]

A B

Inclusion-Exclusion 4

P[A ∪ B∪C ∪ D] = P[A] + P[B] + P[C ] + P[D]


− P[AB] − P[AC ] − P[AD] − P[BC ] − P[BD] − P[CD]
+ P[ABC ] + P[ABD] + P[ACD] + P[BCD]
− P[ABCD]

23
A.1.5 Inclusion-Exclusion Exam P Handouts – Page 24

Exercise 5

A survey of a group’s viewing habits found the following:


1. 31% watched gymnastics, 25% watched baseball, and 21% watched soccer
2. 11% watched gymnastics and baseball, 8% watched baseball and soccer, and 9%
watched gymnastics and soccer
3. 5% watched all three sports.
Find the percentage of the group that watched none of the three sports.

Exercise 5

A survey of a group’s viewing habits found the following:


1. 31% watched gymnastics, 25% watched baseball, and 21% watched soccer
2. 11% watched gymnastics and baseball, 8% watched baseball and soccer, and 9%
watched gymnastics and soccer
3. 5% watched all three sports.
Find the percentage of the group that watched none of the three sports.

Again, let G , B and S denote those who watch gymnastics, baseball, and soccer
respectively. Then our inclusion-exclusion formula gives us
P[G ∪ B ∪ S] = P[G ] + P[B] + P[S] − P[GB] − P[BS] − P[GS] + P[GBS]
= 0.31 + 0.25 + 0.21 − 0.11 − 0.08 − 0.09 + 0.05
= 0.54
P[(G ∪ B ∪ S)0 ] = 1 − 0.54 = 0.46

24
A.2.1 Conditional Probability Exam P Handouts – Page 25

Conditional Probability 1

Die Rolling Example


Examples
Definitions
Key Words
Exercise

Die Rolling Example 2

Suppose you roll a fair 6-sided die. Given that the result is odd, what is the probability
that it is 3 or less?
# of ways to roll 3 or less and odd
P[roll 3 or less | odd] =
# ways to roll an odd number
#{1, 3} 2
= =
#{1, 3, 5} 3

# of ways to roll 3 or less and odd/6


P[roll 3 or less | odd] =
# ways to roll an odd number/6
P[roll 3 or less and odd]
=
P[odd]
2/6 2
= =
3/6 3

25
A.2.1 Conditional Probability Exam P Handouts – Page 26

Example 1 3

In a group of 635 men who died in 1999, 160 of the men died from causes related to
heart disease. Moreover, 275 of the 635 men had at least one parent who suffered from
heart disease, and, of these 275 men, 95 died from causes related to heart disease.

Find the probability that a man randomly selected from this group died of causes not
related to heart disease and that neither of his parents suffered from heart disease.

At least 1 parent
95 180 275
with heart disease
Neither parent
65 295 360
with heart disease
HD No HD
160 475
Answer: 295/635

Example 2 4

In a group of 635 men who died in 1999, 160 of the men died from causes related to
heart disease. Moreover, 275 of the 635 men had at least one parent who suffered from
heart disease, and, of these 275 men, 95 died from causes related to heart disease.

Find the probability that a man randomly selected from this group died of causes not
related to heart disease, given that neither of his parents suffered from heart disease.

At least 1 parent
95 180 275
with heart disease
Neither parent
65 295 360
with heart disease
HD No HD
160 475
Answer: 295/360

26
A.2.1 Conditional Probability Exam P Handouts – Page 27

Definitions 5

Definition (Conditional Probability)


The conditional probability of A given B is

P[A ∩ B] P[AB]
P[A | B] = =
P[B] P[B]

Note that rearranging terms gives

P[B] · P[A | B] = P[AB] = P[A] · P[B | A]

Key words 6

“Given that a person is in college, there is an 80% chance that they use Facebook”

means P[Facebook | college] = 0.80.

But they might not always use the word “given.”

“80% of college students use Facebook”

also means P[Facebook | college] = 0.80.

as does “College students have an 80% chance of using Facebook, while non-college
students . . . ”

Language that restricts possible outcomes to one group/case means condition on that
case.

27
A.2.1 Conditional Probability Exam P Handouts – Page 28

Exercise 7

The blood pressure (high, low, or normal) and heartbeats (regular or irregular) of a random
sample of patients are measured. Of the patients,
1. 36% have high blood pressure and 16% have low blood pressure.
2. 21% have an irregular heartbeat.
3. Of those with an irregular heartbeat, one-third have high blood pressure.
4. Of those with normal blood pressure, one-eighth have an irregular heartbeat.
What portion have a regular heartbeat and low blood pressure?

Exercise 7

The blood pressure (high, low, or normal) and heartbeats (regular or irregular) of a random
sample of patients are measured. Of the patients,
1. 36% have high blood pressure and 16% have low blood pressure.
2. 21% have an irregular heartbeat.
3. Of those with an irregular heartbeat, one-third have high blood pressure.
(1/3) · 0.21 = 0.07
4. Of those with normal blood pressure, one-eighth have an irregular heartbeat.
(1/8) · 0.48 = 0.06
What portion have a regular heartbeat and low blood pressure?
0.08 0.42 0.29 0.79 regular

0.08 0.06 0.07 0.21 irregular


0.16 0.48 0.36
low normal high

28
A.2.2 Independence Exam P Handouts – Page 29

Independence 1

Independence
Examples
Exercise

Independence 2

Previously saw

P[A ∩ B] P[AB]
P[A | B] = =
P[B] P[B]
P[AB] = P[B] · P[A | B] = P[A] · P[B | A]

Definition (Independence)
A and B are independent if P[AB] = P[A] · P[B].

Intuitively, independence means P[A] = P[A | B] and P[B] = P[B | A] so knowing if A


or B occurred gives no information on whether or not the other event occurred.

29
A.2.2 Independence Exam P Handouts – Page 30

Examples 3

1. Suppose A and B are independent events with P[A] = 0.6 and P[AB] = 0.3. Find
P[B] and P[A | B].
P[AB] = P[A] · P[B] by independence
0.3 = 0.6 · P[B]
0.3
P[B] = = 0.5
0.6
P[A | B] = P[A] = 0.6 by independence
2. A and B are events such that P[A] = 0.4, P[B] = 0.1 and P[AB] = 0.05. Are they
independent? What is P[B | A]?
P[A] · P[B] = 0.4 · 0.1 = 0.04 6= 0.05 = P[AB]
so not independent
P[AB] 0.05
P[B | A] = = = 0.125
P[A] 0.4

Exercise 4

If P[A] = 0.2 and P[B] = 0.3, find P[A ∪ B] if:


a) A and B are independent
b) A and B are mutually exclusive

30
A.2.2 Independence Exam P Handouts – Page 31

Exercise 4

If P[A] = 0.2 and P[B] = 0.3, find P[A ∪ B] if:


a) A and B are independent
b) A and B are mutually exclusive

a) P[A ∪ B] = P[A] + P[B] − P[A ∩ B]


= 0.2 + 0.3 − 0.2 · 0.3
= 0.44
b) P[A ∪ B] = P[A] + P[B] − P[A ∩ B]
= 0.2 + 0.3 − 0
= 0.50

31
A.2.3 Sequences of Events Exam P Handouts – Page 32

Sequences of Events 1

Probability of a Flush
Urn Problem
Exercises

Key Idea 2

The definition of conditional probability was that

P[AB]
P[B | A] =
P[A]

We can use that to find P[AB]. Clearing the denominator gives

P[A] · P[B | A] = P[AB]

This equation can be thought of as a sequence of events: first we need A to occur, and
then second we need B to also occur, taking into account the fact that A occurred.

32
A.2.3 Sequences of Events Exam P Handouts – Page 33

Probability of a flush 3

Find the probability of having a flush after being dealt five cards from a standard deck?
(Flush: at least 5 cards of one same suit. Standard deck: 4 suits, each with 13 cards)
We want: P[all 5 cards have the same suit].
Method 1:

P[all 5 cards have the same suit] = P[all 5 are spades]


+ P[all 5 are hearts] + P[all 5 are diamonds]
+ P[all 5 are clubs] = 4 · P[all 5 are spades]
13 12 11 10 9
=4· · · · · = 0.198%
52 51 50 49 48
52 12 11 10 9
Method 2: · · · · = 0.198%
52 51 50 49 48

Probability of a flush: Variation 1 4

If exactly three of the first 5 cards dealt are spades, what is the probability of being
dealt a flush in the first 7 cards?

P[next 2 cards are spades]


= P[6th card is a spade] · P[7th is a spade | 6th is a spade]
13 − 3 13 − 4
= ·
52 − 5 52 − 6
10 9
= ·
47 46
= 0.0416
= 4.16%

33
A.2.3 Sequences of Events Exam P Handouts – Page 34

Probability of a flush: Variation 2 5

If exactly four of the first 5 cards dealt are spades, what is the probability of being
dealt a flush in the first 7 cards?

We want P[at least one of next two is a spade]

Method 1:

P[at least one of next two is a spade] = 1 − P[neither is a spade]


38 37
=1− ·
47 46
= 1 − 0.65
= 0.35

Probability of a flush: Variation 2 Tree Diagram 6

If exactly four of the first 5 cards dealt are spades, what is the probability of being
dealt a flush in the first 7 cards?

Method 2: Draw a tree diagram


9/47 spade

start 9/46 spade

non-
38/47
spade
non-
37/46
spade
9 38 9
So P[flush] = + · = 0.35
47 47 46

34
A.2.3 Sequences of Events Exam P Handouts – Page 35

Urn Problem 7

An urn contains 10 balls: 4 red and 6 blue. A second urn contains 16 red balls and an
unknown number of blue balls. A single ball is drawn from each urn. The probability
that both balls are different colors is 0.528.

Calculate the number of blue balls in the second urn.

b/(16+b)
blue
6/10 blue
16/(16+b) red
b/(16+b)
blue
4/10 red
16/(16+b) red

Urn Problem 8

b/(16+b)
blue
6/10 blue
16/(16+b) red
b/(16+b)
blue
4/10 red
16/(16+b) red

6 16 4 b
0.528 = · + ·
10 16 + b 10 16 + b
96 + 4b
=
160 + 10b
b= 9

35
A.2.3 Sequences of Events Exam P Handouts – Page 36

Exercise 1 9

A family has two children, and they are not twins. Given that at least one of the
children is a boy, what is the probability that both children are boys?

Exercise 1 9

A family has two children, and they are not twins. Given that at least one of the
children is a boy, what is the probability that both children are boys?
1/2
boy
1/2 boy
girl
1/2
1/2
boy
1/2 girl
girl
1/2

1/4
P[2 boys | at least 1 boy] = = 1/3
3/4

36
A.2.3 Sequences of Events Exam P Handouts – Page 37

Exercise 2 10

A family has two children, and they are not twins. Given that the oldest child is a boy,
what is the probability that both children are boys?

Exercise 2 10

A family has two children, and they are not twins. Given that the oldest child is a boy,
what is the probability that both children are boys?
1/2
boy
1/2 boy
girl
1/2
1/2
boy
1/2 girl
girl
1/2

1/4
P[2 boys | oldest is a boy] = = 1/2
2/4

37
A.2.4 Bayes’ Theorem Exam P Handouts – Page 38

Bayes’ Theorem 1

Example 1
Statement of Bayes’ Theorem
Example 2
Exercises

Bayes’ Theorem 2

Often want P[A | B] but are given P[B | A]

Can use

P[AB] P[A] · P[B | A]


P[A | B] = =
P[B] P[B]

Typically then find P[B] by breaking into cases.

38
A.2.4 Bayes’ Theorem Exam P Handouts – Page 39

Example 1 3

An auto insurance company insures drivers of all ages. An actuary compiled the
following statistics on the company’s insured drivers:

Age of Probability Portion of Company’s


Driver of Accident Insured Drivers
16-20 0.06 0.08
21-30 0.03 0.15
31-65 0.02 0.49
66-99 0.04 0.28

A randomly selected driver that the company ensures has an accident. Calculate the
probability that the driver was 31-65.

We are given P[Accident | Age], are reversing condition to P[Age | Accident]

Example 1 (cont.) 4

Age of Probability Portion of Company’s


Driver of Accident Insured Drivers
16-20 0.06 0.08
21-30 0.03 0.15
31-65 0.02 0.49
66-99 0.04 0.28
P[age 31-65, accident]
P[age 31-65 | accident] =
P[accident]
P[age 31-65] · P[accident | age 31-65] (0.49)(0.02)
= =
P[accident] P[accident]
X
P[accident] = P[accident, age group]
age groups

= (.08)(.06) + (.15)(.03) + (.49)(.02) + (.28)(.04) = 0.0303


0.49 · 0.02
P[age 31-65 | accident] = = 0.3234
0.0303

39
A.2.4 Bayes’ Theorem Exam P Handouts – Page 40

The Law of Total Probability 5

Theorem (Law of Total Probability)


If A1 , A2 . . . , Ak are disjoint and P[A1 ] + P[A2 ] + · · · + P[Ak ] = 1 then
• P[B] = P[BA1 ] + P[BA2 ] + · · · + P[BAk ]
• P[B] = P[A1 ] · P[B | A1 ] + · · · + P[Ak ] · P[B | Ak ]

The sets A1 , . . . , Ak are called a partition of the sample space. We will often refer to
them as a list of all possible cases.

In the previous example, the age groups were the Ai , and B was the event of an
accident.

Bayes’ Theorem 6

Theorem (Bayes’ Theorem)


Suppose A1 , . . . , Ak are a partition of the sample space. Then

P[A1 B]
P[A1 | B] =
P[B]
P[A1 ] · P[B | A1 ]
=
P
k
P[BAi ]
i=1
P[A1 ] · P[B | A1 ]
P[A1 | B] =
P
k
P[Ai ] · P[B | Ai ]
i=1

X
The final denominator is P[case] · P[B | case]
cases

40
A.2.4 Bayes’ Theorem Exam P Handouts – Page 41

Example 2 7

Life insurance policy holders are categorized as standard, preferred, and ultra-preferred.
Of a company’s policyholders, 50% are standard, 40% are preferred, and 10% are
ultra-preferred. The probability of dying in the next year is 0.010 for each standard
policyholder, 0.005 for preferred policyholders, and 0.001 for ultra-preferred.

A policyholder dies in the next year. What is the probability that the deceased
policyholder was standard?

Let S denote someone who is standard.

P[S and died]


P[S | died] =
P[died]
0.50 · 0.010
=
0.50 · 0.010 + 0.40 · 0.005 + 0.10 · 0.001
0.005
= = 70.4%
0.0071

Exercise 1 8

Taxicabs in Crobuzon are all either green or blue. On Tuesday, a taxicab got into an
accident. A witness to the accident thought that the cab involved was blue, and
further tests showed that the witness has an 80% chance of correctly identifying the
color of a taxicab, independently of its color.

If 100% of the taxicabs on the streets on Tuesday were green, what was the probability
that the taxicab involved in the accident was blue?

41
A.2.4 Bayes’ Theorem Exam P Handouts – Page 42

Exercise 1 8

Taxicabs in Crobuzon are all either green or blue. On Tuesday, a taxicab got into an
accident. A witness to the accident thought that the cab involved was blue, and
further tests showed that the witness has an 80% chance of correctly identifying the
color of a taxicab, independently of its color.

If 100% of the taxicabs on the streets on Tuesday were green, what was the probability
that the taxicab involved in the accident was blue?

None of the taxicabs were blue that night. The witness is wrong and the cab was not
blue. So the answer is 0% (and in particular, not 80%)!

Exercise 2 9

Taxicabs in Crobuzon are all either green or blue. On Tuesday, a taxicab got into an
accident. A witness to the accident thought that the cab involved was blue, and
further tests showed that the witness has an 80% chance of correctly identifying the
color of a taxicab, independently of its color.

If 85% of the taxicabs on the streets on Tuesday were green, what was the probability
that the taxicab involved in the accident was blue?

42
A.2.4 Bayes’ Theorem Exam P Handouts – Page 43

Exercise 2 9

Taxicabs in Crobuzon are all either green or blue. On Tuesday, a taxicab got into an
accident. A witness to the accident thought that the cab involved was blue, and
further tests showed that the witness has an 80% chance of correctly identifying the
color of a taxicab, independently of its color.

If 85% of the taxicabs on the streets on Tuesday were green, what was the probability
that the taxicab involved in the accident was blue?

P[Blue cab and witness said blue]


P[Blue cab | witness said blue] =
P[witness said blue]
0.15 · 0.80
=
0.15 · 0.80 + 0.85 · 0.20
0.12
= = 41%
0.29

43
A.3.1 Mode Exam P Handouts – Page 44

Mode 1

Overview
Random Variables
Mode
Example
Exercise

Overview 2

3 ways of describing “typical” value of a random variable:

• Mode = most common value


• Median = “middle” value
• Mean = average value

Median and mode are relatively easy to find, usually appear as part of a longer
problem. Mean can be either a step or an entire problem.

Will focus on mode in this lesson.

44
A.3.1 Mode Exam P Handouts – Page 45

Random Variables 3

X is a random variable if it is a number whose value depends on chance. Formally,

X : S → R where S = sample space.

Typically we will use capital letters for random variables and lower case letters for
possible (non-random!) values.

X is a discrete random variable if we can list all the possible values.


X
1= P[X = x] for discrete variables.
x

Mode 4

For a discrete random variable X , y is the mode of X if P[X = y ] ≥ P[X = x] for all x.

I.e., the mode y is the input that maximizes P[X = y ]

It is possible that the max is not unique so X can have multiple modes.

45
A.3.1 Mode Exam P Handouts – Page 46

Mode Example 5

Suppose I roll an otherwise fair 7 sided die whose faces are 1, 1, 1, 2, 4, 4, and 6. Find
the mode.

Let X be the result of the roll


3
P[X = 1] =
7
1
P[X = 2] =
7
2
P[X = 4] =
7
1
P[X = 6] =
7
When y = 1, P[X = y ] reaches its max of 3/7, so 1 is the mode.

Exercise 6

Find the mode of a Poisson random variable with mean 2.8, meaning that
2.8n −2.8
P[N = n] = ·e for n = 0, 1, 2, . . .
n!
where n! = n(n − 1)(n − 2) . . . (2)(1)

46
A.3.1 Mode Exam P Handouts – Page 47

Exercise 6

Find the mode of a Poisson random variable with mean 2.8, meaning that
2.8n −2.8
P[N = n] = ·e for n = 0, 1, 2, . . .
n!
where n! = n(n − 1)(n − 2) . . . (2)(1)

Using the TI-30XS MultiView Data function:


n P[N = n]
0 0.0608
1 0.1703
2 0.2384
3 0.2225
4 0.1557
5 0.0872
6 0.0407
So the mode is 2.

47
A.3.2 Median Exam P Handouts – Page 48

Medians 1

Definitions
Example
Percentiles
Exercise

Median 2

Definition (Cumulative Distribution Function)


F (x) = P[X ≤ x] is the cumulative distribution function of X .

For example, F (2) = P[X ≤ 2]

Definition (Median)
Exam definition: The median of X is the smallest m such that
P[X ≤ m] = F (m) ≥ 1/2.

Remark: there are more complicated definitions because the median is not uniquely
defined for some random variables. Such cases will not appear on the exam, this
simplified definition is equivalent when the median is uniquely defined.

48
A.3.2 Median Exam P Handouts – Page 49

Median Example 3

Suppose I roll an otherwise fair 7 sided die whose faces are 1, 1, 1, 2, 4, 4, and 6. Find
the median.

Let X be the result of the roll


3 3
P[X = 1] = P[X ≤ 1] =
7 7
1 4
P[X = 2] = P[X ≤ 2] =
7 7
2 6
P[X = 4] = P[X ≤ 4] =
7 7
1 7
P[X = 6] = P[X ≤ 6] =
7 7
When y = 2, P[X ≤ y ] first reaches / exceeds 1/2, so 2 is the median.

Percentiles 4

Definition (Percentiles)
The 100% · p th percentile πp is the smallest possible x such that P[X ≤ x] ≥ p.
Example: If X is our die rolling example from before:

x 0 1 2 4 6
F (x) 0 3/7 4/7 6/7 1

5th percentile = 1 since P[N ≤ 1] ≥ 0.05 but P[N ≤ 0] < 0.05


10th percentile = 1 since P[N ≤ 1] ≥ 0.10 but P[N ≤ 0] < 0.10
Median = 50th percentile = 2
90th percentile = 6 since P[N ≤ 4] < 0.9 but P[N ≤ 6] ≥ 0.9

49
A.3.2 Median Exam P Handouts – Page 50

More Complex Example 5

Roll a fair 6-sided die. What is the median outcome?


n P[N = n] F (n) = P[N ≤ n]
1 1/6 1/6
2 1/6 2/6
3 1/6 3/6=1/2
4 1/6 4/6
5 1/6 5/6
6 1/6 1
By our definition, 3 is the first time F (n) ≥ 1/2 so the median is 3.
But F (x) = 1/2 for any x such that 3 ≤ x < 4 (e.g., F (3.5) = 1/2.)
Some texts say that any x with F (x) = 1/2 is a median, giving infinitely many
medians (i.e., anything in the interval 3 ≤ x < 4).
This situation will not occur on the exam.

Exercise 6

Suppose that P[N = n] = n/15 for n = 1, 2, 3, 4 or 5. Find the median of N.

50
A.3.2 Median Exam P Handouts – Page 51

Exercise 6

Suppose that P[N = n] = n/15 for n = 1, 2, 3, 4 or 5. Find the median of N.

Evaluating the CDF:

P[N ≤ 1] = P[N = 1] = 1/15 < 1/2


P[N = 2] = 2/15
P[N ≤ 2] = 1/15 + 2/15 = 3/15 < 1/2
P[N = 3] = 3/15
P[N ≤ 3] = 3/15 + 3/15 = 6/15 < 1/2
P[N = 4] = 4/15
P[N ≤ 4] = 6/15 + 4/15 = 10/15 ≥ 1/2

So the median is 4.

51
A.3.3 Moments: Expected Value/Mean Exam P Handouts – Page 52

Moments: Expected Value / Mean 1

Example
Defintions
Example
Exercise

Example 2

In a group of 10 people, I owe $4 to two of them, $2 to one of them, $1 to two of


them, and nothing to the others. On average, how much do I owe to these 10 people?

Method 0: The average debt is

Total Amount Owed 4·2+2·1+1·2+0·5 12


= = = 1.20
Number of People 10 10
Method 1: Pick one of the 10 people at random. Let X be the amount owed to that
person.
X 2 1 2 5
E[X ] = x · P[X = x] = 4 · +2· +1· +0·
x
10 10 10 10
= 0.8 + 0.2 + 0.2 + 0 = 1.2

52
A.3.3 Moments: Expected Value/Mean Exam P Handouts – Page 53

Definitions 3

Definition (Expected Value)


If X is a discrete random variable then
X
E[X ] = x · P[X = x]
x

Definition (Generalizations)

  X 2
E X2 = x · P[X = x]
x
X
E[g (X )] = g (x) · P[X = x]
x

Example 4

An insurance policy pays 100 per day for up to 3 days of hospitalization and 50 per day
of hospitalization thereafter. Find the expected payment for hospitalization if the
number of days of hospitalization, X , is a discrete random variable with

 6 − k for k = 1, 2, 3, 4, 5
P(X = k) = 15
0 otherwise
Let g (k) = payment for k days in the hospital.
X
E[g (X )] = g (k) · P[X = k]

k 1 2 3 4 5
g (k) 100 200 300 350 400
P[X = k] 5/15 4/15 3/15 2/15 1/15
5 4 3 2 1
E[g (X )] = 100 · + 200 · + 300 · + 350 · + 400 · = 220
15 15 15 15 15

53
A.3.3 Moments: Expected Value/Mean Exam P Handouts – Page 54

Exercise 5

Suppose that P[N = n] = n/15 for n = 1, 2, 3, 4 or 5. Find E[N].

Exercise 5

Suppose that P[N = n] = n/15 for n = 1, 2, 3, 4 or 5. Find E[N].


X
E[N] = n · P[N = n]
n
1 2 3 4 5
=1· +2· +3· +4· +5·
15 15 15 15 15
55
=
15
11
= = 3.67
3
Recall that the median was 4. The mode is 5. The median and mode must be possible
values of the variable, the expected value doesn’t have to be.

54
A.3.4 Tools for Finding Means Exam P Handouts – Page 55

Tools for Finding Means 1

Overview
Law of Total Expectation
Linear Combinations
Exercise

Overview 2

Previously had
X
E[X ] = x · P[X = x]
x
X
E[g (X )] = g (x) · P[X = x]
x

Sometimes can relate to other means to find what we want.

Will look at 2 methods that do so.

55
A.3.4 Tools for Finding Means Exam P Handouts – Page 56

Law of Total Expectation 3

Definition (Law of Total Expectation)


n
[
If A1 , . . . , An are a partition of the sample space, that is, S = Ai and Ai ∩ Aj = ∅
i=1
for i 6= j, then
n
X
E[X ] = E[X | Ai ] · P[Ai ]
i=1

This often happens with the Ai being high / medium / low risk groups

The usual definition is the situation with Ai = {X = i}

Linear Combinations 4

Note that
X
E[aX + b] = (ax + b) · P[X = x]
x
X X
= ax · P[X = x] + b · P[X = x]
x x
!
X X
=a· x · P[X = x] +b· P[X = x]
x x
E[aX + b] = a · E[X ] + b

That is, means of linear functions distribute as you would expect

In general, E[aX + bY ] = a E[X ] + b E[Y ] for any two random variables X and Y .

56
A.3.4 Tools for Finding Means Exam P Handouts – Page 57

Exercise 5

40% of an insurer’s claims are from high risk individuals, with an average claim amount
of 500, and 60% are from low risk individuals with an average claim amount of 100.
Due to administrative costs, a claim of X costs the insurer 1.05X + 10. What is the
expected cost to the insurer of a randomly selected claim?

Exercise 5

40% of an insurer’s claims are from high risk individuals, with an average claim amount
of 500, and 60% are from low risk individuals with an average claim amount of 100.
Due to administrative costs, a claim of X costs the insurer 1.05X + 10. What is the
expected cost to the insurer of a randomly selected claim?

E[Cost] = E[1.05X + 10] = 1.05 E[X ] + 10


= 1.05 (0.4 · 500 + 0.6 · 100) + 10
= 1.05 · 260 + 10
= 283

57
A.3.5 Means: Survival Function Approach Exam P Handouts – Page 58

Means: Survival Function Approach 1

Initial Example Revisited


Survival Function Method
Exercises

Initial Example Revisited 2

In a group of 10 people, I owe $4 to two of them, $2 to one of them, $1 to two of


them, and nothing to the others. On average, how much do I owe to these 10 people?

Saw how to use the definition previously.

Method 2: I need to repay the debt.


• Step 1: Pay $1 to everyone who is owed money.
• Step 2: Pay $1 more to everyone who is still owed money (i.e., people who were
initially owed $2 or $4). etc.
Total payment: $5 in step 1, $3 in step 2, $2 in step 3 and $2 in step 4. After step 4,
everyone has been paid back.

Total payment = # owed at least $1 + # owed at least $2 + . . . = 5+3+2+2 =12

Average per person = (total payment) / (# people) = 12/10 = 1.20 as before.

58
A.3.5 Means: Survival Function Approach Exam P Handouts – Page 59

Survival function method 3

If N ≥ 0 and N is an integer valued variable,

E[N] = P[N > 0] + P[N > 1] + P[N > 2] + . . .


X∞
E[N] = P[N > n]
n=0
or letting k = n + 1 so P[N > n] = P[N ≥ n + 1] = P[N ≥ k]
X∞
E[N] = P[N ≥ k]
k=1

The name is because P[N > n] = 1 − F (n) is called the survival function.
The discrete version doesn’t extend easily to finding E[g (X )] (but the continuous
version of the survival function method will).

Exercise 1 4

Following a certain type of surgery, patients are hospitalized for N days, with
5−k
P[N ≥ k] = for k = 0, 1, 2, 3, 4 or 5. Find E[N] using the survival method.
5

59
A.3.5 Means: Survival Function Approach Exam P Handouts – Page 60

Exercise 1 4

Following a certain type of surgery, patients are hospitalized for N days, with
5−k
P[N ≥ k] = for k = 0, 1, 2, 3, 4 or 5. Find E[N] using the survival method.
5

X
E[N] = P[N > n]
n=0
= P[N > 0] + P[N > 1] + P[N > 2] + P[N > 3] + P[N > 4]
= P[N ≥ 1] + · · · + P[N ≥ 5]
4 3 2 1
= + + + +0
5 5 5 5
= 2

Exercise 2 5

Following a certain type of surgery, patients are hospitalized for N days, with
5−k
P[N ≥ k] = for k = 0, 1, 2, 3, 4 or 5. Find E[N] using the definition.
5

60
A.3.5 Means: Survival Function Approach Exam P Handouts – Page 61

Exercise 2 5

Following a certain type of surgery, patients are hospitalized for N days, with
5−k
P[N ≥ k] = for k = 0, 1, 2, 3, 4 or 5. Find E[N] using the definition.
5
P[N = k] = P[N ≥ k] − P[N ≥ k + 1]
5−k 5 − (k + 1)
= −
5 5
1
= for k = 0, 1, 2, 3 or 4
5
1 1 1 1 1
E[N] = 0 · + 1 · + 2 · + 3 · + 4 ·
5 5 5 5 5
10
=
5
= 2

61
A.3.6 Variance: Definition Exam P Handouts – Page 62

Variance 1

Basic Definition
Exercise

Example 2

Suppose that X , Y and Z are three random variables such that

P[X = 2] = 1
1 1 1
P[Y = 1] = P[Y = 2] = P[Y = 3] =
3 3 3
1 1
P[Z = 1] = P[Z = 3] =
2 2
Then E[X ] = E[Y ] = E[Z ] = 2. But intuitively, Y is more likely to differ from the
mean than X , and Z is even more likely to do so.
The variance of a variable is a way to quantify how much it differs from its mean.
Definition (Variance)
Var[X ] = E[(X − µX )2 ] where µX = E[X ]

62
A.3.6 Variance: Definition Exam P Handouts – Page 63

Example 3

We want to find the variance of X , Y , and Z , where

P[X = 2] = 1
1 1 1
P[Y = 1] = P[Y = 2] = P[Y = 3] =
3 3 3
1 1
P[Z = 1] = P[Z = 3] =
2 2

Var[X ] = E[(X − 2)2 ] = 1 · (2 − 2)2 = 0


1 1 1 2
Var[Y ] = · (1 − 2)2 + · (2 − 2)2 + · (3 − 2)2 =
3 3 3 3
1 1
Var[Z ] = · (1 − 2)2 + (3 − 2)2 = 1
2 2

Terminology 4

h i
E X k is the kth moment of X . Sometimes it is called the kth raw moment of X .

µ = E[X ] = mean = average = first moment of X .


 
E X 2 is the second (raw) moment of X .
 
Var[X ] = E (X − µ)2 = σ 2 = 2nd central moment of X .
 
E (X − µ)k = kth central moment.
 
E (X − a)k = kth moment about a.

63
A.3.6 Variance: Definition Exam P Handouts – Page 64

Exercise 5

An insurance policy pays 100 per day for up to 3 days of hospitalization and 50 per day
for each day of hospitalization thereafter. The number of days of hospitalization, X , is
a discrete random variable with probability function
6−k
P(X = k) = for k = 1, 2, 3, 4, 5 and 0 otherwise.
15
The mean payment amount is 220. Find the variance of a payment for hospitalization.

Exercise 5

An insurance policy pays 100 per day for up to 3 days of hospitalization and 50 per day
for each day of hospitalization thereafter. The number of days of hospitalization, X , is
a discrete random variable with probability function
6−k
P(X = k) = for k = 1, 2, 3, 4, 5 and 0 otherwise.
15
The mean payment amount is 220. Find the variance of a payment for hospitalization.

Let Y denote the payment amount.


k 1 2 3 4 5
Y 100 200 300 350 400
P[X = k] 5/15 4/15 3/15 2/15 1/15
h i 5(−120)2 4(−20)2 3 · 802 2 · 1302 1802
2
E (Y − 220) = + + + +
15 15 15 15 15
= 10,600

64
A.3.7 Variance: Tools Exam P Handouts – Page 65

Variance 1

Main Method for Finding Variance


Example Revisited
Linear Transformations
Standard Deviation
Exercise

A faster way (usually) to find the variance: 2

Let µ = E[X ]. Then


 
Var[X ] = E (X − µ)2
 
= E X 2 − 2µX + µ2
 
= E X 2 − 2µ · E[X ] + µ2
 
= E X 2 − 2µ · µ + µ2
 
= E X 2 − µ2
  2
Var[X ] = E X 2 − E[X ]
2
E[X 2 ] = Var[X ] + E[X ]

Remark: from the definition, Var[X ] ≥ 0, with Var[X ] = 0 if and only if P[X = µ] = 1,
i.e, X is constant.  
This implies that E X 2 ≥ (E[X ])2 .

65
A.3.7 Variance: Tools Exam P Handouts – Page 66

Example Revisited 3

An insurance policy pays 100 per day for up to 3 days of hospitalization and 50 per day
thereafter. The number of days of hospitalization, X , has probability function
6−k
P(X = k) = for k = 1, 2, 3, 4, 5 and 0 otherwise.
15
Find the variance of a payment Y for hospitalization.
k 1 2 3 4 5
Y 100 200 300 350 400
P[X = k] 5/15 4/15 3/15 2/15 1/15

5 4 3 2 1
E[Y ] = 100 · + 200 · + 300 · + 350 · + 400 · = 220
15 15 15 15 15
5 4 3 2 1
E[Y 2 ] = 1002 · + 2002 · + 3002 · + 3502 · + 4002 · = 59,000
15 15 15 15 15
Var[Y ] = 59,000 − 2202 = 10,600

Multiplying by a Constant 4

Suppose we know Var[X ]. What is Var(3X )? What is Var(c X )?


 
Var[3X ] = E (3X )2 − (E[3X ])2
 
= E 9X 2 − (3E[X ])2
 
= 9 E(X 2 ) − (E[X ])2
= 9 · Var[X ]
Var[cX ] = c 2 Var[X ]

66
A.3.7 Variance: Tools Exam P Handouts – Page 67

Linear Transformations 5

Suppose we know Var[X ]. What is Var[X + b]? What is Var[aX + b]?


 
Var[X + b] = E (X + b − E[X + b])2
 
= E (X + b − E[X ] − b)2
 
= E (X − E[X ])2
= Var[X ]

I.e., shifting by a constant doesn’t change distances from the mean so variance is same.

Var[aX + b] = Var[aX ]
= a2 Var[X ]

Standard Deviation and Coefficient of Variation 6

The standard deviation SD[X ] satisfies:


p
SD[X ] = σX = Var[X ]
Var[cX ] = c 2 Var[X ]
SD[cX ] = |c| · SD[X ]
Exam problems may ask about the Coefficient of Variation CV[X ]
σ SD[X ]
CV[X ] = =
µ E[X ]
SD[cX ]
CV[cX ] =
E[cX ]
c · SD[X ]
= if c > 0
c · E[X ]
CV[cX ] = CV[X ]
Note: If E[X ] is held constant and σ increases, then CV[X ] also increases.

67
A.3.7 Variance: Tools Exam P Handouts – Page 68

Exercise 7

A random variable X satisfies E[X ] = 5 and SD[X ] = 3. Find


a) E[X 2 ]
b) Var[2X + 6]
c) CV[2X + 6]

Exercise 7

A random variable X satisfies E[X ] = 5 and SD[X ] = 3. Find


a) E[X 2 ]
b) Var[2X + 6]
c) CV[2X + 6]

E[X 2 ] = Var[X ] + (E[X ])2 = 32 + 52 = 34


Var[2X + 6] = Var[2X ] = 22 Var[X ] = 4 · 9 = 36
E[2X + 6] = 2E[X ] + 6 = 2 · 5 + 6 = 16
SD[2X + 6]
CV[2X + 6] =
E[2X + 6]

36 3
= =
16 8

68
A.3.8 Discrete Uniform Random Variables Exam P Handouts – Page 69

Discrete uniform random variables 1

Standard Uniform
Generalizations
Exercise

Standard Uniform: Expected Value and Variance 2

Definition (Discrete Uniform)


1
X is a (discrete) uniform on {1, 2, . . . , n − 1, n} if P[X = i] = for those n choices.
n
n+1
Note that the average of 1 and n is , as is the average of 2 and n − 1, and 3 and
2
n − 2 and . . . , so
n+1
E[X ] =
2

n2 − 1
We will see later that Var[X ] =
12

69
A.3.8 Discrete Uniform Random Variables Exam P Handouts – Page 70

Example 3

Suppose X is uniform on {3, 4, 5, 6, 7, 8}. What is E[X ]?

As before, we find E[X ] by pairing extremes (3 and 8, 4 and 7, ...). Each pair has an
3+8 11 11
average value of = so E[X ] =
2 2 2
Alternatively, X is not a standard uniform because it starts at 3, not 1. But X − 2 is a
standard uniform on {1, 2, . . . , 6}

1+6 7
E[X − 2] = =
2 2
E[X ] = E[X −2] + 2
7 11
= +2=
2 2

Example 4

Suppose X is uniform on {3, 4, 5, 6, 7, 8}. What is Var[X ]?

As before, X − 2 is a standard uniform on {1, 2, . . . , 6}

n2 − 1
Var[X − 2] =
12
36 − 1
=
12
35
=
12
Var[X ] = Var[X − 2]
35
=
12

70
A.3.8 Discrete Uniform Random Variables Exam P Handouts – Page 71

General case 5

Suppose that X is uniform on {a, a + 1, . . . , b}.


a+b
Then E[X ] = which is the average of the first and last possible values.
2
Var[X ] = Var[X − (a − 1)] since shifting by constants doesn’t change the variance.

But X − (a − 1) is uniform on {1, 2, . . . , b − (a − 1)} so

[b − (a − 1)]2 − 1
Var[X ] =
12
(# of values in range)2 − 1
=
12

Exercise 6

The number of losses N is uniformly distributed on {5, 6, . . . , 20}. Each loss results in
a payment of 100. Find the mean and standard deviation of the payment amount.

71
A.3.8 Discrete Uniform Random Variables Exam P Handouts – Page 72

Exercise 6

The number of losses N is uniformly distributed on {5, 6, . . . , 20}. Each loss results in
a payment of 100. Find the mean and standard deviation of the payment amount.

Let X denote the payment amount. X = 100N.

E[X ] = 100 · E[N]


5 + 20
= 100 · = 1,250
2
Var[X ] = 1002 · Var[N]
2 (20 − 4)2 − 1
= 100 · = 212,500
r 12
255
SD[X ] = 100 ·
12
= 461

72
A.4.1 Permutations Exam P Handouts – Page 73

Permutations 1

Complete Lists
Partial Lists
Exercises

Complete Lists 2

How many possible rankings are there of a group of 3 people named A, B and C ?

A B C

AB AC BA BC CA CB

ABC ACB BAC BCA CAB CBA

There are 3 choices of who comes first, 2 people left who can be second, and then only
1 choice remaining for third. The total number of possibilities is 3 · 2 · 1

73
A.4.1 Permutations Exam P Handouts – Page 74

Complete Lists 3

How many possible rankings are there of a group of 12 people?

12 11 10 . . . 3 2 1
There are 12 choices of who can be first.
There are 11 people left who can be second.
There are 10 people left who can be third.
.. .. .. ..
. . . .
There are 3 people left who can be tenth.
There are 2 people who can be eleventh, and that leaves just one who can be last.

Complete Lists 4

Key idea: at each step, multiply the number of choices for the new step with the
choices so far. That gives an answer of

12 · 11 · 10 · · · 2 · 1
= 12!
n! = n · (n − 1) · (n − 2) · · · 2 · 1
1! = 1
0! = 1

Note: The MultiView’s “prb” menu has an n! function


Each order is called a permutation, n! is the number of permutations of n objects.

74
A.4.1 Permutations Exam P Handouts – Page 75

Partial Lists 5

A contest with 12 people gives out 3 distinct prizes. How many ways are there to give
out these prizes?

12 11 10

So there are 12 · 11 · 10 ways to do this.

This is an example of a partial permutation, or a k-permutation.

There is a formula for the number of k-permutations of n objects, but it is better to


think through them from scratch and multiply the number of choices at each step.

Exam questions will have complications so that formula doesn’t apply but ideas do.

Exercise 1 6

How many 3 digit numbers are there with all even digits?

75
A.4.1 Permutations Exam P Handouts – Page 76

Exercise 1 6

How many 3 digit numbers are there with all even digits?

4 · 5 · 5 = 100

The first digit can’t be 0 because then our number would be at most 2 digits. So there
are 4 choices of first digit (namely 2, 4, 6 or 8)

The second and third digits can be 0, so there are 5 choices for each of them (namely
0, 2, 4, 6 or 8).

Terminology: because I am allowed to repeat digits, I am choosing digits with


replacement.

Exercise 2 7

How many 3 digit numbers are there with all even digits and no repeated digits?

76
A.4.1 Permutations Exam P Handouts – Page 77

Exercise 2 7

How many 3 digit numbers are there with all even digits and no repeated digits?

4 · 4 · 3 = 48

The first digit can’t be 0 because then our number would be at most 2 digits. There
are 4 choices of first digit (namely 2, 4, 6 or 8)

The second digit can be 0, which should give 5 choices (0, 2, 4, 6 or 8). But we
cannot repeat the first digit, leaving only 5 − 1 = 4 choices.

The third digit can be 0, giving 5 potential choices (0, 2, 4, 6 or 8). But we cannot
repeat either of the first 2 digits, leaving only 5 − 2 = 3 choices.

Terminology: because I am not allowed to repeat digits, I am choosing digits without


replacement.

77
A.4.2 Combinations Exam P Handouts – Page 78

Combinations 1

Example
Combinations
Partitions
Exercises

Partial Lists 2

A contest with 12 people gives out 3 distinct prizes. Previously counted the ways to
give out these prizes.

12 11 10

Giving 12 · 11 · 10 ways to do this.

How many ways are there to give out the prizes if all 3 are the same?

The key difference in the second version is whether the top 3 in order are A, B, C or B,
A, C, or C, B, A, or ... we give out the same prizes and there is no difference.

The answer will be less because each rearrangement of the top 3 gives the same prizes.

78
A.4.2 Combinations Exam P Handouts – Page 79

Partial Lists 3

A contest with 12 people gives out 3 prizes. How many ways are there to give out the
prizes if all 3 are the same?

In writing 12 · 11 · 10, each top 3 was counted 3 · 2 · 1 times (since that is how many
orders there are of 3 things), so our answer is
12 · 11 · 10
3·2·1
This comes up often enough to deserve notation.
 
12
Answer is
3
Can be found on calculator by 12 nCr 3 under ‘prb’ menu

Key Ideas 4

n! = n · (n − 1) · · · 3 · 2 · 1
= # ways to arrange n items in a list
To choose r items from a group of n, if all items we choose are equal then the number
of ways to do so is:
n · (n − 1) · (n − 2) · · · · (n − r + 1) n!
=
r · (r − 1) · (r − 2) · · · · (2) · (1) (n − r )!r !

 
n! n
= is called n choose r .
(n − r )!r !   r 
n n
Notes: = = nCr on calculators
    r n − r
n n n!
= = = 1 since 0! = 1 as there is only 1 way to choose the entire set.
0 n 0! · n!

79
A.4.2 Combinations Exam P Handouts – Page 80

Partitions 5

18 people are to be divided into 3 groups, one with 8 people, one with 6, and one with
4. How many such divisions are possible?

Standard method: Initially assume that we have a complete rank, and then divide by
the amount of “overcounting”

So first we rank all 18 people, and let group A be the top 8, group B the next 6, and
group C the bottom 4.
A B C

8 slots 6 slots 4 slots


There are 18! ways to rank everyone, but within each group people are equal, so we
18!
over counted by 8! · 6! · 4! Our final answer is thus
8! · 6! · 4!

Partitions 6

18 people are to be divided into 3 groups, one with 8 people, one with 6, and one with
4. How many such divisions are possible?

Second approach:

Step 1: Pick the 


group
 with 8 people.
18 18!
We can do so in = ways.
8 8! · 10!
Step 2: From the remaining 18 − 8 = 10 people, pick the group with 6 people. This
also determines
 the group with 4 people.
10 10!
There are = ways to do so.
6 6! · 4!    
18 10 18! 10! 18!
That gives a final answer of · = · =
8 6 8! · 10! 6! · 4! 8! · 6! · 4!

80
A.4.2 Combinations Exam P Handouts – Page 81

Exercise 1 7

4 distinct numbers are picked from the integers {1, 2, . . . , 30}. How many ways are
there to draw them such that all 4 are divisible by 3?

Exercise 1 7

4 distinct numbers are picked from the integers {1, 2, . . . , 30}. How many ways are
there to draw them such that all 4 are divisible by 3?

As 30 is the 10th multiple of 3, 10 of the possible numbers are divisible by 3.

We want to choose 4 of those 10, which can be done in


 
10 10 · 9 · 8 · 7
= = 210
4 4!

different ways.

81
A.4.2 Combinations Exam P Handouts – Page 82

Exercise 2 8

4 distinct numbers are picked from the integers {1, 2, . . . , 30}. How many ways are
there to draw them such that 3 are divisible by 5 and the other is divisible by 7?

Exercise 2 8

4 distinct numbers are picked from the integers {1, 2, . . . , 30}. How many ways are
there to draw them such that 3 are divisible by 5 and the other is divisible by 7?

6 numbers in the list are multiples of 5, while 4 are multiples of 7.

5 · 7 = 35 > 30 so none are multiples of both.


 
6 6·5·4
There are = = 20 ways to choose the 3 that are divisible by 5.
3 3!  
4
For each pick of those 3, there are = 4 ways to choose the one divisible by 7.
1
That gives a total of 20 · 4 = 80 ways to do both.

82
A.4.3 The Binomial Distribution Exam P Handouts – Page 83

The Binomial Distribution 1

Example
Bernoulli random variables
Binomial distributions
Exercises

Example 2

Avery is practicing free throws. If they make each shot with probability 0.7 and each
shot is independent, what is the probability that they make the next 4 shots and then
miss the 2 after that? What is the probability that they make exactly 4 of the next 6
shots?

4 makes then 2 misses is (P[Make])4 · (P[Miss])2 = 0.74 · 0.32


 
6
To make exactly 4 of 6, there are ways to choose which 4 shots are successful.
4  
4 2 6
Each order has probability 0.7 · 0.3 for a final answer of · 0.74 · 0.32
4

83
A.4.3 The Binomial Distribution Exam P Handouts – Page 84

Bernoulli random variables 3

A Bernoulli(p) random variable, aka a Bernoulli 0-1 random variable, is a variable that
can only be 0 or 1. Sometimes the case X = 1 is called a success, and X = 0 a failure.

P[X = 1] = p
P[X = 0] = 1 − p = q
E[X ] = 1 · p + 0 · (1 − p) = p
 2
E X = 12 · p + 02 · (1 − p) = p
 
Var[X ] = E X 2 − (E[X ])2
Var[X ] = p − p 2 = p(1 − p)
Var[X ] = pq

Bernoulli variance trick 4

If X is a random variable that can only take on two values a and b, with
P[X = b] = p P[X = a] = 1 − p = q
then the mean and variance are
E[X ] = a · q + b · p = a + p · (b − a)
 2
E X = a2 · q + b 2 · p
 
Var[X ] = E X 2 − (E[X ])2
Var[X ] = (b − a)2 · p · q
Or: X = (b − a)Y + a Y ∼ Bernoulli(p)
E[X ] = (b − a)E[Y ] + a = p(b − a) + a
Var[X ] = (b − a)2 Var[Y ]
Var[X ] = (b − a)2 · p · q
Final result arguably worth memorizing especially on later exams

84
A.4.3 The Binomial Distribution Exam P Handouts – Page 85

Binomial distributions 5

X is a binomial (n, p) random variable if X is the number of successes in n


independent trials, each of which is a success with the same probability p.

In the free throw example, n = 6, p = 0.7 and X was the number of made free throws.
 
n k
P[X = k] = p (1 − p)n−k
k

The key pieces that you need are


• A fixed number of trials
• Trials are independent
• Success probability is the same in all trials

Binomial Distributions 6

To find the mean and variance, note that


n
X
X = Xi , Xi ∼ independent, Bernoulli(p)
i=1
Xn n
X
E[X ] = E [Xi ] = p = np
i=1 i=1
Xn
Var[X ] = Var [Xi ] by independence
i=1
Xn
Var[X ] = p(1 − p)
i=1
Var[X ] = np(1 − p) = npq

85
A.4.3 The Binomial Distribution Exam P Handouts – Page 86

Exercise 1 7

A commuter airline sells 32 tickets for a flight on a plane that has 30 seats. The
probability that any particular passenger will not show up for a flight is 0.10,
independent of other passengers. Find the probability that more passengers show up
for the flight than there are seats available.

Exercise 1 7

A commuter airline sells 32 tickets for a flight on a plane that has 30 seats. The
probability that any particular passenger will not show up for a flight is 0.10,
independent of other passengers. Find the probability that more passengers show up
for the flight than there are seats available.

Let N = # of passengers that show up. N ∼ Binomial(n = 32, p = 0.9).


 
32
P[N = 32] = · (0.9)32 = 0.932
32
 
32
P[N = 31] = · (0.9)31 (0.1)1 = 32(0.9)31 (0.1)1
31
P[N > 30] = P[N = 32] + P[N = 31]
= 0.932 + 32(0.9)31 (0.1)1 = 0.156

86
A.4.3 The Binomial Distribution Exam P Handouts – Page 87

Exercise 2 8

An airline sells 32 tickets for a flight. The probability that any particular passenger will
not show up for a flight is 0.10, independent of other passengers. What are the mean
and variance of the number of passengers who show up?

Exercise 2 8

An airline sells 32 tickets for a flight. The probability that any particular passenger will
not show up for a flight is 0.10, independent of other passengers. What are the mean
and variance of the number of passengers who show up?

The number of passengers N who arrive is binomial with n = 32 and p = 0.9, so

E[N] = 32 · 0.9 = 28.8


Var[N] = 32 · 0.9 · 0.1 = 2.88
SD[N] = 1.69

87
A.4.4 Multinomial Distribution Exam P Handouts – Page 88

Multinomial Distribution 1

Example
Multinomial Distribution
Exercise

Example 2

Accidents are categorized into three groups: minor, moderate, and severe. These occur
with probabilities 0.5 for minor, 0.4 for moderate, and 0.1 for severe.
Two accidents occur independently in one month. Find the probability that neither
accident is severe and at most one is moderate.

This is not binomial because each trial / accident has more than 2 possible outcomes.
But many of the ideas remain the same.

We want P[1 minor, 1 moderate] + P[2 minor]


   
2 2
= (0.5)(0.4) + (0.5)2 = 2(0.4)(0.5) + (0.5)2 = 0.65
1 2

Choose which accident is minor

88
A.4.4 Multinomial Distribution Exam P Handouts – Page 89

Multinomial Distribution: 3 Outcomes 3

That example had 3 possible outcomes instead of 2.


That made it an example of the multinomial distribution.

Suppose there are n trials with P


3 possible outcomes. Let the probabilities of those
outcomes be p1 , p2 and p3 (so pi = 1) and let Xi be the number of trials that have
outcome i. Then

   
n n − k1 n − k1 − k2 k1 k2 k3
P[X1 = k1 , X2 = k2 , X3 = k3 ] = p1 p2 p3
k1 k2 k3
 
n! (n − k1 )! k3 k1 k2 k3
= · p p p
k1 !(n − k1 )! k2 !(n − k1 − k2 )! k3 1 2 3
n!
= p1k1 p2k2 p3k3
k1 !k2 !k3 !

Multinomial Distribution: General Case 4

Suppose that there are n independent trials, each with the same r possible outcomes.
Let p1 , p2 , . . . , pr be the probabilities of the outcomes, and Xi the number of trials
resulting in the i-th outcome. Then
n!
P[X1 = k1 , X2 = k2 , . . . , Xr = kr ] = p1k1 p2k2 . . . prkr
k1 !k2 ! . . . kr !
As with binomial, need:
• A fixed number of trials
• Different trials are independent
• All trials have same distribution
Difference is now can have > 2 possibilities.

89
A.4.4 Multinomial Distribution Exam P Handouts – Page 90

Exercise 5

Accidents are categorized as minor, moderate, or severe. The probability that a given
accident is minor is 0.5, that it is moderate is 0.4, and that it is severe is 0.1.

Four accidents occur independently in one month. Find the probability that there is at
least one accident of each type.

Exercise 5

Accidents are categorized as minor, moderate, or severe. The probability that a given
accident is minor is 0.5, that it is moderate is 0.4, and that it is severe is 0.1.

Four accidents occur independently in one month. Find the probability that there is at
least one accident of each type.

There are 3 cases:


4!
P[2 minor, 1 moderate, 1 severe] = (0.5)2 (0.4)(0.1)
2!1!1!
4!
P[1 minor, 2 moderate, 1 severe] = (0.5)(0.4)2 (0.1)
1!2!1!
4!
P[1 minor, 1 moderate, 2 severe] = (0.5)(0.4)(0.1)2
1!1!2!
P[Total] = 0.12 + 0.096 + 0.024
= 0.24

90
A.4.5 Hypergeometric Distribution Exam P Handouts – Page 91

Hypergeometric Distribution 1

Hypergeometric Distribution
Binomial vs Hypergeometric
Examples
Exercises

Non-Independent Draws 2

A crate of 10 electrical components has 4 defective components. If 3 components are


randomly selected, find the probability that at most one of them is defective.

Two cases: none of the 3 are defective, or 1 of the three is defective. Those are
mutually disjoint, so we can sum their probabilities, giving

P[none are defective] + P[1 is defective]


    
6 6 4
3 2 1
= +  
10 10
3 3
20 15 · 4 2
= + =
120 120 3

91
A.4.5 Hypergeometric Distribution Exam P Handouts – Page 92

Hypergeometric Distribution 3

Hypergeometric: have G good pieces out of N total. Choose n, and want the
probability of choosing g good items.
  
G N −G
Number of ways to choose exactly g good items g n−g
=  
Number of ways to choose n total items N
n

This is not a binomial distribution: knowing whether or not the first item is good
gives information about whether or not the second one will be good.

If n > g , then it isn’t possible to have n successes, while that would be possible with a
binomial.

Binomial vs Hypergeometric 4

Key words to tell hypergeometric vs binomial: Is sample with or without replacement?

A binomial needs each trial to be independent. This occurs when sampling with
replacement.

A hypergeometric distribution comes up when we are sampling without replacement


from a finite population.

92
A.4.5 Hypergeometric Distribution Exam P Handouts – Page 93

Examples 5

When packing for a trip, I draw 6 socks without replacement from a drawer that
contains 16 black socks and 4 white socks.

What is the probability that I will draw 6 white socks? What is the probability that I
will draw 4 black socks and 2 white socks?

This is hypergeometric since the different draws are not independent.


Let W denote the number of white socks that are drawn.

P[W = 6] = 0 since there are only 4 white socks in the drawer.


   
4 16
·
2 4
P[W = 2] =   = 0.282
20
6

Binomial variation 6

Suppose that I draw 6 socks with replacement from a drawer that contains 16 black
socks and 4 white socks.

What is the probability that I will draw 6 white socks? What is the probability that I
will draw 4 black socks and 2 white socks?

This is binomial distribution since the different draws are independent.


Let W denote the number of white socks that are drawn.

P[W = 6] = 0.26
 
6
P[W = 2] = (0.2)2 (0.8)4
2
= 0.246

93
A.4.5 Hypergeometric Distribution Exam P Handouts – Page 94

Exercise 1 7

I draw 6 socks without replacement from a drawer that contains 10 black socks, 6
brown socks, and 4 white socks. Find the probability that I will draw 2 socks of each
color.

Exercise 1 7

I draw 6 socks without replacement from a drawer that contains 10 black socks, 6
brown socks, and 4 white socks. Find the probability that I will draw 2 socks of each
color.

Draws are not independent b/c sampling without replacement.


Let Bl, Br and W equal # of black, brown, and white socks drawn.
   
10 6 4
2 2 2
P[Bl = Br = W = 2] =   = 0.104
20
6

94
A.4.5 Hypergeometric Distribution Exam P Handouts – Page 95

Exercise 2 8

I randomly select 6 socks from a drawer. Each sock has a 50% chance of being black,
a 30% chance of being brown, and a 20% chance of being white, independently of the
other socks. Find the probability that I will draw 2 socks of each color.

Exercise 2 8

I randomly select 6 socks from a drawer. Each sock has a 50% chance of being black,
a 30% chance of being brown, and a 20% chance of being white, independently of the
other socks. Find the probability that I will draw 2 socks of each color.

This is multinomial distribution because the draws are independent.

Let Bl, Br and W equal # of black, brown, and white socks drawn. Then
  
6 4
P[Bl = Br = W = 2] = (0.5)2 (0.3)2 (0.2)2
2 2
6!
= (0.5)2 (0.3)2 (0.2)2
2!2!2!
= 0.081

95
A.5.0a Geometric Series Exam P Handouts – Page 96

Geometric Series 1

Overview
Geometric Series
Generalizations

Overview 2

Many discrete distributions can equal any non-negative integer.

Deriving the mean / variance of these requires using infinite series.

Exam questions will not require derivations, just the results.

The material in this lesson is optional, will only be used for theory.

96
A.5.0a Geometric Series Exam P Handouts – Page 97

Geometric Series 3

Suppose that |r | < 1. Then



X
S= ar n
n=0
S = a + ar + ar 2 + . . .
& & &
Sr = ar + ar 2 + ar 3 + . . .
S − Sr = a
a
S=
1−r

Geometric Series 4

In words, the sum of geometric series is


first term
S=
1 − ratio

X ar 10
ar n =
1−r
n=10

For a partial sum (the final result holds for any r ),


m
X ∞
X ∞
X
n n
ar = ar − ar n
n=0 n=0 n=m+1
a ar m+1
= −
1−r 1−r
first term − first missing term
=
1−r

97
A.5.0a Geometric Series Exam P Handouts – Page 98

Example 5

17
X e n+2
Suppose we want to find 5 · 3n −n .
2 3
n=3
n = 3 is the first term, and n = 18 is the first missing term.
The ratio is e/(23 3−1 ) = 3e/8 so we get
17
X e n+2 first term − first missing term
5· =
23n 3−n 1 − ratio
n=3

5e 3+2 5e 18+2
9 3−3
− 3·18 −18
= 2 2 3
3e
1−
8

n · ar n 6

What if instead we wanted



X
S= n · ar n , |r | < 1
n=0
S = 0 · a + 1 · ar + 2 · ar 2 + . . .
& & &
Sr = 0 · ar + 1 · ar 2 + 2 · ar 3 + . . .
S − Sr = 0 · a + 1 · ar + 1 · ar 2 + . . .
ar
S(1 − r ) =
1−r
ar
S=
(1 − r )2

Will use this to find mean of a geometric random variable

98
A.5.0a Geometric Series Exam P Handouts – Page 99

n2 · ar n 7

Harder still:

X
S= n2 · ar n , |r | < 1
n=0
S = 0 · a + 12 · ar + 22 · ar 2 + . . .
2

& & &


Sr = 02 · ar + 12 · ar 2 + 22 · ar 3 + . . .
S − Sr = 1 · ar + 3 · ar 2 + 5 · ar 3 + . . .

X X∞ X∞
n n
S(1 − r ) = (2n − 1)ar = 2 nar − ar n
n=1 n=1 n=1
ar ar
S(1 − r ) = 2 · −
(1 − r )2 1 − r
ar (1 + r )
S=
(1 − r )3

Calculus Approach 8

Alternatively,

X a
ar n =
1−r
n=0

X ∞
X
d d n d a
ar n = ar =
dr dr dr 1 − r
n=0 n=0

X a
nar n−1 =
(1 − r )2
n=0
X∞
ar
nar n =
(1 − r )2
n=0

X ar (1 + r )
Taking a second derivative leads to n2 ar n =
(1 − r )3

99
A.5.0b Taylor Series for exp(x) Exam P Handouts – Page 100

Taylor Series for e x 1

The Taylor Series for e x


Examples

The Taylor Series for e x 2

When we study the Poisson distribution, we will need facts about e x .

x x2 x3 xn
e =1+x + + + ··· + + ···
 2 3! n!
d x2 x3
Note : 1+x + + + ···
dx 2 3!
!
x x2
= 0+1+2· +3· + ···
2! 3!
x2
= 1+x + + · · · = ex
2!
u 2 u3
u
e =1+u+ + + ···
2 3!
−x x2 x3
e =1−x + − + ···
2 3!

100
A.5.0b Taylor Series for exp(x) Exam P Handouts – Page 101

Examples 3


X ∞
X
5n e tn −5 −5 (5 e t )n
·e =e
n! n!
n=0 n=0
X∞
−5 un
=e u = 5e t
n!
n=0
−5 t)
=e · e = e −5 · e (5 e
u

t −5
= e 5e

Missing Terms 4


X 2n
Suppose we want to find
n!
n=2

X 2n
We know = e2
n!
n=0
Our sum is missing the n = 0 and n = 1 terms, so

X ∞
X  
2n 2n 20 21
= − +
n! n! 0! 1!
n=2 n=0
 
2 1 2
=e − +
1 1
= e2 − 3

101
A.5.1 The Geometric Distribution Exam P Handouts – Page 102

The Geometric Distribution 1

Rolling a die until the first six


Geometric random variables
Exercise

Rolling a die until the first six 2

Suppose I roll a die until I get a 6. Let N be the total number of rolls. What is the
distribution of N?
1
P[N > 0] = 1 P[N = 1] =
6 

5 5 1
P[N > 1] = P[N = 2] = ·
6 6 6
 2  2
5 5 1
P[N > 2] = P[N = 3] = ·
6 6 6
.. ..
. .
 k  n−1
5 5 1
P[N > k] = P[N = n] = ·
6 6 6

102
A.5.1 The Geometric Distribution Exam P Handouts – Page 103

Expected Value 3


X ∞
X
E[N] = n · P[N = n] = P[N > k]
n=1 k=0
X∞  n−1 ∞  k
X
1 5 5
= n· · =
6 6 6
n=1 k=0

The survival approach is nicer here because it gives us a geometric series.


1
E[N] = =6
5
1−
6

Variance 4


X
 2

E N = n2 · P[N = n]
n=1
X∞  n−1 X∞  n
5 2 1 2 1 6 5
= n · · = n · · ·
6 6 6 5 6
n=0 n=0
 
1 6 5 11
· · ·
ar (1 + r ) 6 5 6 6
= =  
(1 − r )3 5 3
1−
6
1 5 11
· ·
= 5 6 6 = 66
1
63
Var[N] = 66 − 62 = 30

103
A.5.1 The Geometric Distribution Exam P Handouts – Page 104

Geometric starting at 1 5

X is a geometric random variable on {1, 2, . . .} with parameter p if X is the number of


trials up to, and including, the first success.

P[X = n] = p · (1 − p)n−1
1
E[X ] = , memorize or use survival method
p
  2 1
E X2 = 2 − (don’t memorize)
p p
   2
2 1 1 1 1
Var[X ] = − − = −
p2 p p p2 p
1−p
Var[X ] = (memorize)
p2

Geometrics Starting at 0 6

Y is a geometric random variable on 0, 1, 2, . . . if Y counts the number of failures


before the first success.
P[Y = n] = p · (1 − p)n
Y = X − 1 since the # of trials until the first success, including the success, is always
exactly one plus the number of failures. That means that
1
E[Y ] = E[X ] − 1 = −1
p
1−p
E[Y ] =
p
Var[Y ] = Var[X ]
1−p
Var[Y ] = = E[Y ] · E[X ]
p2

104
A.5.1 The Geometric Distribution Exam P Handouts – Page 105

Exercise 7

Let N be the number of visits (possibly 0) that a randomly chosen insured makes to the
doctor in a year. If N has a geometric distribution with mean 3, what is the probability
that a randomly chosen insured makes at least 2 visits to the doctor in a year?

Exercise 7

Let N be the number of visits (possibly 0) that a randomly chosen insured makes to the
doctor in a year. If N has a geometric distribution with mean 3, what is the probability
that a randomly chosen insured makes at least 2 visits to the doctor in a year?

We are told that N can be 0, so N is a geometric on {0, 1, 2, . . . }


1−p
E[N] =
p
1−p 1
3= , p=
p 4
P[N ≥ 2] = 1 − P[N = 0] − P[N = 1]
= 1 − p − (1 − p) · p
= 1 − 2p + p 2
9
= (1 − p)2 =
16

105
A.5.2 Memoryless Property Exam P Handouts – Page 106

Memoryless Property 1

Intuitive Example
Algebra Approach
Memoryless Property
Variance
Exercise

“Nines” 2

In each round of the dice game “Nines” I roll two fair six-sided dice. The game ends if
either a 7 or 9 is rolled, and continues to the next round on any other outcome.

If I play a game of Nines, what is the expected number of rounds that I will play?

6 4 5
The game ends on a given round with probability + = .
36 36 18

The length is therefore a geometric random variable (starting at 1) with p = 5/18.

Expected game length is 1/p = 18/5 rounds.

106
A.5.2 Memoryless Property Exam P Handouts – Page 107

Watching “Nines” 3

If I watch someone play Nines, find the expected number of rounds that I will watch.

There is the same as the previous example, so the answer is still 18/5.

Now suppose I start watching a game after the 3rd round. How many more rounds will
I watch?

Intuitively, this is no different than before. Each round I watch will end the game with
probability 5/18, the number of rounds I watch is a geometric starting at 1, and the
answer is still 18/5.

Algebra Approach 4

Let N = game length.

Start watching after 3 rounds, so watch for N − 3 rounds.

The game lasts for more than 3 rounds, so we are given N > 3.

P[N = k + 3, N > 3]
P[N − 3 = k | N > 3] =
P[N > 3]
(1 − p)k+3−1 · p
=
(1 − p)3
= (1 − p)k−1 · p
= P[N = k]

so (N − 3 | N > 3) and (the original) N have the same distribution.


In particular, E[N − 3 | N > 3] = E[N] = 18/5.

107
A.5.2 Memoryless Property Exam P Handouts – Page 108

Memoryless Property 5

If N is a geometric with parameter p, then (N − k | N > k) is a geometric starting at


1 with the same p.

• (N − k | N > k) is at least 1, hence starting at 1


• Holds whether N starts at 0 or 1
• Geometrics are only discrete distribution with this property

Variance 6

A game of Nines lasts for at least 4 rounds. What are the mean and variance of the
length of the game?

We are given that N ≥ 4, so N > 3.

E[N | N > 3] = E[N−3+3 | N > 3]


= E[N−3 | N > 3]+3
18
= E[N]+3 = +3 = 6.6
5
Var[N | N > 3] = Var[N−3+3 | N > 3]
= Var[N−3 | N > 3]
1−p
= Var[N] =
p2
= 9.36

108
A.5.2 Memoryless Property Exam P Handouts – Page 109

Exercise 7

Suppose X satisfies P[X = k] = 0.2(0.8)k for k = 0, 1, 2, . . . . Find E[X | X > 6] and


Var[X | X > 6].

Exercise 7

Suppose X satisfies P[X = k] = 0.2(0.8)k for k = 0, 1, 2, . . . . Find E[X | X > 6] and


Var[X | X > 6].

X can be 0, P[X = k] decays geometrically, so X is a geometric starting at 0.


P[X = 0] = p = 0.2(0.8)0 , so p = 0.2
E[X | X > 6] = E[X − 6 + 6 | X > 6]
= E[X − 6 | X > 6] + 6
= E[Geo starting at 1, p = 0.2] + 6
1
= + 6 = 11
0.2
Var[X | X > 6] = Var[X − 6 | X > 6]
= Var[Geo starting at 1, p = 0.2]
1 − 0.2
= = 20
0.22

109
A.5.3 Negative Binomial Distribution Exam P Handouts – Page 110

Negative Binomial Distribution 1

Rolling a die until the third six


Negative binomial
Example
Exercise

Rolling a die until the third six 2

Roll a die until the third time that a 6 is rolled. Let N denote the number of non-sixes
(failures) that we roll. What is the distribution of N?

If N = n, then roll n + 3 was a 6.

The first (n + 3) − 1 rolls had 3 − 1 = 2 sixes. Since there were exactly 3 sixes in the
first n + 3 rolls, there were n non-sixes.
   2  n
(n + 3) − 1 1 5 1
P[N = n] =
3−1 6 6 6
   3  n
n + (3 − 1) 1 5
=
3−1 6 6
   3  n
n + (3 − 1) 1 5
=
n 6 6

110
A.5.3 Negative Binomial Distribution Exam P Handouts – Page 111

Rolling a die until the third six 3

What are the mean and variance of N?

We can use a trick to find them.

Let N1 = # failures before first 6


N2 = # number of failures before 2nd 6
N = N3 = # number of failures before 3rd 6
N = (N1 − 0) + (N2 − N1 ) + (N3 − N2 )

N1 − 0 is the number of failures before the first six.


N2 − N1 is the number of failures after the first six, but before the second six.
N3 − N2 is the number of failures after the second six, but before the third six.

Rolling a die until the third six comes up 4

Key point: # of failures between sixes is a geometric that starts at 0.


N is the sum of 3 independent geometrics on {0, 1, 2, . . . } and

E[N] = E[N1 ] + E[N2 − N1 ] + E[N3 − N2 ]


E[N] = 3E[N1 ] = 3 · (6 − 1) = 15
  
1
Var[N] = 3 Var of a geometric p =
6
1−p 1
=3· 2
, p=
p 6
5

Var[N] =   6 = 90
1 2
6

111
A.5.3 Negative Binomial Distribution Exam P Handouts – Page 112

Negative binomial 5

N is a negative binomial random variable with parameters r and p if it is the sum of r


independent geometric random variables starting at 0. It is the number of failures
before the r -th success.
 
n + (r − 1)
P[N = n] = · p r · (1 − p)n
r −1
 
n + (r − 1)
= · p r · (1 − p)n
n
1−p
E[N] = r ·
p
1−p r (1 − p)
Var[N] = r · =
p2 p2

Example 6

An insurance policy covers accidents at a manufacturing plant. The probability that


one or more accidents will occur during any given month is 3/5. The number of
accidents that occur in any given month is independent of the number of accidents
that occur in all other months.

Find the probability that June will be the fourth month in 2025 in which at least one
accident occurs.

Having an accident = “success,” p = 3/5.


Want r = 4th success in 6th try, n = 6 − 4 = 2 “failures”.
   4  2
2 + (4 − 1) 3 2
· ·
4−1 5 5

112
A.5.3 Negative Binomial Distribution Exam P Handouts – Page 113

Exercise 7

Let N be the sum of r independent geometrics on {0, 1, 2, . . . }. Suppose that


E[N] = 12 and Var[N] = 60. Find the probability that N is no more than 2.

Exercise 7

Let N be the sum of r independent geometrics on {0, 1, 2, . . . }. Suppose that


E[N] = 12 and Var[N] = 60. Find the probability that N is no more than 2.
r (1 − p)
E[N] = = 12
p
r (1 − p) r (1 − p) 1
Var[N] = = ·
p2 p p
1 1
60 = 12 · ⇒ p= , r =3
p 5
P[N ≤ 2] = P[N = 0] + P[N = 1] + P[N = 2]
     
0+3−1 3 1+2 3 2+2
= p + (1 − p)p + (1 − p)2 p 3
3−1 2 2
   
3 4
= p3 + (1 − p)p 3 + (1 − p)2 p 3
2 2
= 0.05792

113
A.5.4 The Poisson Distribution Exam P Handouts – Page 114

The Poisson distribution 1

The Poisson distribution


Mean and variance
Key Points
Exercises

The Poisson distribution 2

X is a Poisson(λ) random variable if

−λ λn
P[X = n] = e · , n = 0, 1, 2, . . .
n!

To remember this, recall that the Taylor series for e λ is

λ2 λn
eλ = 1 + λ + + ··· + + ···
2! n!
The e −λ term is the constant needed to make the probabilities sum to 1.

Poisson variables arise in nature by counting the number of occurrences of unusual


events if the number of occurrences in disjoint time intervals are independent.

114
A.5.4 The Poisson Distribution Exam P Handouts – Page 115

Poisson Means 3


X ∞
X
−λ λn
E[N] = n · P[N = n] = n·e ·
n!
n=0 n=1

X λn−1
=λ e −λ ·
(n − 1)!
n=1
X∞
−λ λm
=λ e · by letting m = n − 1
m!
m=0

X
= λ · 1 since the sum is P[N = m] = 1
m=0

Point: The mean of a Poisson(λ) is λ.

Variance 4

∞ ∞
 2 X −λ λ
n X λ · λn−1
E N = 2
n ·e · = n · e −λ ·
n! (n − 1)!
n=0 n=1

X λm
=λ (m + 1) · e −λ · by letting m = n − 1
m!
m=0

! ∞
!
X e −λ · λm X e −λ · λm
=λ m· +λ
m! m!
m=0 m=0

! ∞
!
X X
=λ m · P[N = m] +λ P[N = m]
m=0 m=0
= λ · λ + λ · 1 = λ2 + λ

Var[N] = λ2 + λ − λ2 = λ = E[N]

115
A.5.4 The Poisson Distribution Exam P Handouts – Page 116

Key Points 5

The derivations are harder than what you will be doing on the exam.

What you need to know so far:

• The probability mass function

−λ λk
P[N = k] = e ·
k!

• E[N] = Var[N] = λ

Exercise 1 6

Policyholders are three times as likely to file two claims as to file four claims.
If the number of claims filed has a Poisson distribution, find the variance of the
number of claims filed.

116
A.5.4 The Poisson Distribution Exam P Handouts – Page 117

Exercise 1 6

Policyholders are three times as likely to file two claims as to file four claims.
If the number of claims filed has a Poisson distribution, find the variance of the
number of claims filed.

Let N be the number of claims. Var[N] = λ, so we need to find λ.

P[N = 2] = 3 · P[N = 4]
−λ λ2 −λ λ
4
e · = 3e ·
2 4!
4! λ4
= 2
2·3 λ
4 = λ2
λ=2 or − 2 but λ > 0
so λ = Var[N] = 2

Exercise 2 7

The number of annual losses has a Poisson distribution with second moment equal to
12. Find the probability that the number of annual losses is at least 2.

117
A.5.4 The Poisson Distribution Exam P Handouts – Page 118

Exercise 2 7

The number of annual losses has a Poisson distribution with second moment equal to
12. Find the probability that the number of annual losses is at least 2.

Let N be the number of annual losses.

E[N 2 ] = Var[N] + (E[N])2


12 = λ + λ2
0 = λ2 + λ − 12
0 = (λ − 3)(λ + 4)
λ=3 because λ > 0
P[N ≥ 2] = 1 − P[N = 0] − P[N = 1]
= 1 − e −3 − 3e −3
= 0.80

118
A.5.5 Sums of Independent Poissons Exam P Handouts – Page 119

Sums of Independent Poissons 1

Example
Sums of Poisson
Exercise

Example 2

If X is Poisson with mean 1.7 and Y is an independent Poisson with mean 1.3, find:
a) E[X + Y ]
b) Var[X + Y ]
c) P[X + Y = 2]

E[X + Y ] = E[X ] + E[Y ] = 1.7 + 1.3 = 3


Var[X + Y ] = Var[X ] + Var[Y ] by independence
= 1.7 + 1.3 = 3
P[X + Y = 2] = P[X = 0, Y = 2] + P[X = 1, Y = 1] + P[X = 2, Y = 0]
1.32 −1.3 1.72 −1.7 −1.3
= e −1.7 · e + 1.7e −1.7 · 1.3e −1.3 + e ·e
2 2
32
= 4.5e −3 = e −3
2
Note: answers are consistent with X + Y ∼ Poisson(3)

119
A.5.5 Sums of Independent Poissons Exam P Handouts – Page 120

Sums of Poisson 3

If N ∼ Poisson (λ), M ∼ Poisson (µ) and they are independent, find P[N + M = n].
n
X
P[N + M = n] = P[N = k] · P[M = n − k]
k=0
Xn
λk −µ µn−k
= e −λ · ·e ·
k! (n − k)!
k=0
n
!
X λk · µn−k 1
= e −(λ+µ) · n! ·
k!(n − k)! n!
k=0
n
X  
e −(λ+µ) n
= λk · µn−k ·
n! k
k=0
(λ + µ)n
= e −(λ+µ) ·
n!
= P[Poisson(λ + µ) = n]

Key Point 4

If N1 , . . . , Nk are independent Poissons, then their sum is Poisson.


P P
E[ Ni ] = E[Ni ].

Revisiting our first example: If X is Poisson with mean 1.7 and Y is an independent
Poisson with mean 1.3, find P[X + Y = 2]

X + Y ∼ Poisson(λ = 1.3 + 1.7 = 3)


32 −3
P[X + Y = 2] = e
2

120
A.5.5 Sums of Independent Poissons Exam P Handouts – Page 121

Exercise 5

The number of accidents per day at a busy intersection has a Poisson distribution with
mean 0.5 during a workday and 0.3 during a weekend day. If the number of accidents
on different days is independent, what is the probability that there will be exactly three
accidents at the intersection during a week?

Exercise 5

The number of accidents per day at a busy intersection has a Poisson distribution with
mean 0.5 during a workday and 0.3 during a weekend day. If the number of accidents
on different days is independent, what is the probability that there will be exactly three
accidents at the intersection during a week?

The sum of independent Poissons is Poisson, so the number of accidents per week is a
Poisson with mean

λ = 5 · 0.5 + 2 · 0.3 = 3.1


λ3
P[N = 3] = e −λ ·
3!
3.13
= e −3.1 ·
6
= 0.224

121
A.6.1 Deductibles Exam P Handouts – Page 122

Deductibles 1

Deductibles
Finding Expected Payments
Exercises

Deductibles 2

Suppose that X represents the amount of a loss.

If there is a deductible of(d, then the resulting payment is


0 X ≤d
Payment = (X − d)+ =
X −d X >d
The uncovered cost to the insured is (
X X ≤d
Uncovered Cost = min{X , d} = X ∧ d =
d X >d
In words, note that

Total Loss = Insurance Payment + Uncovered Cost to Insured


X = (X − d)+ + (X ∧ d)

122
A.6.1 Deductibles Exam P Handouts – Page 123

Example 3

Suppose that loss amounts are uniform on {1, 2, 3, 4, 5} and that there is a deductible
of 2. What is the expected payment? What is the probability that the uncovered loss
will be 2?

x P[X = x] Payment Uncovered Loss


1 1/5 0 1
2 1/5 0 2
3 1/5 1 2
4 1/5 2 2
5 1/5 3 2

1 1 1 1 1
E[Payment] = ·0+ ·0+ ·1+ ·2+ ·3
5 5 5 5 5
= 1.2
4
P[Uncovered Loss = 2] =
5

Finding the Expected Payment 4

There are often fewer possible values for the uncovered loss than for the payment,
which means that it is often easier to find E[X ∧ d] than E[(X − d)+ ]. We can take
advantage of this as follows:

X = (X − d)+ + (X ∧ d)
E[X ] = E[(X − d)+ ] + E[X ∧ d]
E[(X − d)+ ] = E[X ] − E[X ∧ d]
Warning: This only works for first moments:
X 2 6= (X − d)2+ + (X ∧ d)2
E[X 2 ]6= E[(X − d)2+ ] + E[(X ∧ d)2 ]
Exam questions about second moments are very rare.

123
A.6.1 Deductibles Exam P Handouts – Page 124

Exercise 1 5

A farm is insured against tornado damage. During tornado season, each week has
either 0 or 1 tornados, with a probability of 0.3 of having a tornado. The policy pays
$100 per tornado, with an annual deductible of $50.

Tornado season is 8 weeks long and the number of tornados in different weeks are
independent. Find the expected annual insurance payment.

Exercise 1 5

A farm is insured against tornado damage. During tornado season, each week has
either 0 or 1 tornados, with a probability of 0.3 of having a tornado. The policy pays
$100 per tornado, with an annual deductible of $50.

Tornado season is 8 weeks long and the number of tornados in different weeks are
independent. Find the expected annual insurance payment.

Let N be the number of storms, and X = 100N the total loss. The uncovered loss is
either 0 (if there are no tornados) or 50 (if there is at least 1 tornado). So
E[X ∧ 50] = 0 · P[N = 0] + 50 · P[N ≥ 1]

= 0 + 50 · 1 − 0.78
= 47.12
E[X ] = 100 E[N] = 100 · 8 · 0.3 = 240
E[Payment] = 240 − 47.12 = 192.88

124
A.6.1 Deductibles Exam P Handouts – Page 125

Exercise 2 6

The number of annual losses N is a geometric on {0, 1, 2, . . . } with mean 2. Losses are
insured for $100 each, with an annual deductible of $150. Find the expected annual
payment.

Exercise 2 6

The number of annual losses N is a geometric on {0, 1, 2, . . . } with mean 2. Losses are
insured for $100 each, with an annual deductible of $150. Find the expected annual
payment.
1−p
E[N] = 2 =
p
p = 1/3
E[Uncovered Loss] = 0 · P[N = 0] + 100 · P[N = 1] + 150 · P[N ≥ 2]
= 0 · p + 100 · (1 − p)p + 150 · (1 − p)2
= 88.9
E[Payment] = E[Total Loss] − E[Uncovered Loss]
= 2 · 100 − 88.9
= 111.1

125
A.6.2 Policy Limits Exam P Handouts – Page 126

Policy Limits 1

Policy Limits
Exercises

Policy Limits 2

Another way for the payment to be less than the total loss is to have a policy limit.

Let X be the loss amount, and u the policy limit. With no deductible,
(
X X ≤u
Payment =
u u<X

In this case, Payment = min{X , u} = X ∧ u.

126
A.6.2 Policy Limits Exam P Handouts – Page 127

Limits and Deductibles 3

With both a deductible of d and a limit of u, then there are different types of limits.

Exam questions will be explicit about how the limit works.

Examples in the study note all have a payment limit, meaning that u is the maximum
payment allowed. In that case:


0 X ≤d
Payment = X − d d < X ≤ d + u


u d +u <X
The expected payment is also called the net premium or the benefit premium.

Exercise 1 4

The number of annual losses is Poisson with mean 2.4. Each loss results in 50 in
damages. Total annual claims are insured with a payment limit of 75. Find the
expected annual payment.

127
A.6.2 Policy Limits Exam P Handouts – Page 128

Exercise 1 4

The number of annual losses is Poisson with mean 2.4. Each loss results in 50 in
damages. Total annual claims are insured with a payment limit of 75. Find the
expected annual payment.

Let N be the number of losses. The payment is 0 when N = 0, 50 when N = 1, and


75 when N ≥ 2, so

E[Payment] = 0 · P[N = 0] + 50 · P[N = 1] + 75 · P[N ≥ 2]



= 0 + 50 · 2.4e −2.4 + 75 1 − e −2.4 − 2.4e −2.4
= 62.75

Exercise 2 5

Loss amounts X have a binomial distribution with n = 5 and p = 0.4. If there is a


deductible of 1 and a payment limit of 3, find the expected payment for a randomly
selected loss.

128
A.6.2 Policy Limits Exam P Handouts – Page 129

Exercise 2 5

Loss amounts X have a binomial distribution with n = 5 and p = 0.4. If there is a


deductible of 1 and a payment limit of 3, find the expected payment for a randomly
selected loss.
x 0 1 2 3 4 5
P[X = x] 0.65 5 · 0.4 · 0.64 10 · 0.42 0.63 10 · 0.43 0.62 5 · 0.44 0.6 0.45
Payment 0 0 1 2 3 3
Uncovered Loss 0 1 1 1 1 2

E[Payment] = 0.3456 · 1 + 0.2304 · 2 + (0.0768 + 0.0102) · 3


= 1.06752
E[Uncovered Loss] = 1 · (1 − 0.65 − 0.45 ) + 2 · 0.45 = 0.93248
E[Payment] = E[X ] − E[Uncovered Loss]
= 5 · 0.4 − 0.93248 = 1.06752

129
A.6.3 Calculator Approach to Deductible Exam P Handouts – Page 130

Calculator Approach to Deductibles 1

Examples Revisited

Example 2

Suppose that loss amounts are uniform on {1, 2, 3, 4, 5} and that there is a deductible
of 2. What is the expected payment?

Let X be a random loss amount.


When X exceeds the deductible, the payment is X − 2.
When X is below the deductible, the payment is 0.

Put x in L1 , P[X = x] in L2 and payment in L3 . For the payment, we will start with a
formula (L3 = L1 − 2) and then edit to take into account the deductible.

E[Payment] = 1.2

130
A.6.3 Calculator Approach to Deductible Exam P Handouts – Page 131

Example 3

A farm is insured against damage from tornados. Each week during tornado season has
either 0 or 1 tornados, with a probability of 0.3 of having a tornado. The insurance
policy pays $100 per tornado, with an annual deductible of $50.
If tornado season is 8 weeks long and the number of tornados in different weeks are
independent, what is the expected annual insurance payment?

Let N be the number of storms. N ∼ Binomial(8, 0.3)

Put N in L1 , probabilities in L2 .

The payment (L3 ) is 100N − 50 when that is positive, and is 0 otherwise.

E[Payment] = 193

Example 4

The number of annual losses N is geometric on {0, 1, 2, . . . } with mean 2. Each loss is
insured for $100, with annual deductible of $150. Find the expected annual payment.

Can’t enter all possible values, but E[N] is only 2. Turns out going up to 10 isn’t
enough, 15 gets pretty close.
1−p
E[N] = 2 =
p
1
2p = 1 − p ⇒p=
3
 k
1 2
P[N = k] = p(1 − p)k = ·
3 3
Payment is 100N − 150 when positive, 0 otherwise.

Again, k in L1, P[N = k] in L2, and payment in L3. E[Payment] = 111.1


Only get 99 if stopping at k = 10. Need n to be closer to 1.

131
A.6.3 Calculator Approach to Deductible Exam P Handouts – Page 132

Example 5

Loss amounts X have a binomial distribution with n = 5 and p = 0.4. Suppose that
there is a deductible of 1 and a payment limit of 3. What is the resulting expected
payment?

When X is above the deductible and the limit hasn’t yet been reached, the payment is
X − 1. Below the deductible the payment is 0 and above the limit it is 3.

This time we will edit L3 for both the deductible and limit.

E[Payment] = 1.07

132
A.7.1 Discrete Review Exam P Handouts – Page 133

Discrete Review 1

Basic Formulas
Moments
Combinatorics
Key Distributions

General Probability Rules 2

• 0 ≤ P[A] ≤ 1
• P[State Space] = 1, P[∅] = 0
• P[A ∪ B] = P[A] + P[B] − P[A ∩ B]
If A and B are mutually exclusive, then P[A ∩ B] = 0 and P[A ∪ B] = P[A] + P[B]
• A ∪ A0 = State Space, A ∩ A0 = ∅
• P[Ac ] = P[A0 ] = 1 − P[A], P[A] = 1 − P[A0 ]
P[A ∩ B]
• P[A | B] =
P[B]
• P[A ∩ B] = P[AB] = P[A] · P[B | A]
• A and B are independent if and only if P[AB] = P[A] · P[B].
In that case, P[A | B] = P[A]

133
A.7.1 Discrete Review Exam P Handouts – Page 134

Examples 3

If A and B are independent with P[A ∪ B] = 0.58 and P[A] = 0.3, find P[A0 ∪ B 0 ].

P[A ∪ B] = P[A] + P[B] − P[A ∩ B]


P[A ∪ B] = P[A] + P[B] − P[A] · P[B]
0.58 = 0.3 + P[B] − 0.3P[B]
P[B] = 0.4
P[A0 ∪ B 0 ] = P[A0 ] + P[B 0 ] − P[A0 ∩ B 0 ]
= 0.7 + 0.6 − 0.7 · 0.6
= 0.88
or: A0 ∪ B 0 = (A ∩ B)0
P[A0 ∪ B 0 ] = 1 − P[AB]
= 1 − 0.3 · 0.4
= 0.88

Examples 4

If P[A | B] = 3P[B | A], P[A ∪ B] = 7P[AB] and P[B] = 0.1, then what is P[AB]?

P[A | B] = 3P[B | A]
P[AB] P[AB]
=3·
P[B] P[A]
P[A] = 3 · P[B] = 0.3
P[A ∪ B] = P[A] + P[B] − P[AB]
7P[AB] = 0.3 + 0.1 − P[AB]
0.4
P[AB] = = 0.05
8

134
A.7.1 Discrete Review Exam P Handouts – Page 135

Moments, etc. 5

Let X be a random variable.


• y is the mode of X if P[X = y ] ≥ P[X = x] for all x
• m is the median of X if m is the smallest value such that
P[X ≤ m] = F (m) ≥ 1/2
P
• E[X ] = x · P[X = x]
x
P
• E[X 2 ] =x 2 · P[X = x]
x
P
• E[g (X )] = g (x) · P[X = x]
x
 2  
• Var[X ] = E X − (E[X ])2 = E (X − µ)2
 
• E X 2 = Var[X ] + (E[X ])2

Example 6

Suppose that P[X = x] = cx for x = 1, 2, 3, 4 or 5. Find the median and mode of X .

x P[X = x] P[X ≤ x]
1 c 1/15
2 2c 3/15
3 3c 6/15
4 4c 10/15
5 5c 15/15
X 1
1= P[X = x] = c + 2c + 3c + 4c + 5c, c=
x
15

The mode is 5 since that maximizes P[X = x] and the median is 4 since that is the
first time F (x) = P[X ≤ x] exceeds 1/2.

135
A.7.1 Discrete Review Exam P Handouts – Page 136

Example 7

P[X = x] = x/15 for x = 1, 2, 3, 4 or 5. Find the variance of X .

1 2 5
E[X ] = 1 · +2· + ··· + 5 ·
15 15 15
55 11
E[X ] = =
15 3
 2 1 2 5
E X = 12 · + 22 · + · · · + 52 ·
15 15 15
 2  225
E X = = 15
15
 2
11 14
Var[X ] = 15 − =
3 9

Breaking Down Into Cases 8

Let A1 , A2 , . . . , An be a partition of the sample space, i.e., Ai ∩ Aj = ∅ for i 6= j and


S
Ai = S.
i
Then
P
• P[B] = P[Ai ] · P[B | Ai ]
i
P
• E[X ] = P[Ai ] · E[X | Ai ]
i
P[Ai B] P[Ai ] · P[B | Ai ]
• P[Ai | B] = = P
P[B] P[Aj ] · P[B | Aj ]
j

136
A.7.1 Discrete Review Exam P Handouts – Page 137

Combinatorics 9

n! = n(n − 1)(n − 2) . . . 2 · 1
= # of ways to order n objects
1! = 1
0! = 1
 
n n!
=
k (n − k)!k!
 
n
=
n−k
= # of ways to choose k objects from a set of n

Example 10

I randomly select two socks from a drawer that has 4 black socks and 6 brown socks,
and put them into a bag that initially had 3 black and 3 brown socks. I then randomly
select a sock from the bag. Find the probability that both socks from the drawer were
brown given that the sock taken from the bag was brown.
Let X be the number of brown socks taken from the drawer and Y the number of
brown socks drawn from the bag.
P[X = 2, Y = 1]
P[X = 2 | Y = 1] =
P[Y = 1]
6

2 5
10
 ·
2
8
= 6 
6 4
 4

2 5 1 1 4 2 3
10
 · + 10
 · + 10
 ·
2
8 2
8 2
8
15 · 5 25
= =
15 · 5 + 24 · 4 + 6 · 3 63

137
A.7.1 Discrete Review Exam P Handouts – Page 138

Key Distributions 11

Type of Variable Key Properties


All possibilities
Uniform
are equally likely
Bernoulli 0 or 1
Sum of Bernoullis
Binomial Number of successes
in n independent trials
Geometric Number of failures
on {0, 1, 2, . . . } before first success
Geometric Number of trials
on {1, 2, 3, . . . } including first success
Number of failures
Negative Binomial before r -th success
Sum of geometrics
Sum of independent Poissons
Poisson
is Poisson

Key Distributions 12

Variable P[X = x] E[X ] Var[X ]


Uniform 1/n n+1 n2 − 1
on {1, 2, . . . , n} 1≤x ≤n 2 12
p, x = 1
Bernoulli(p) p p(1 − p)
 1− p, x = 0
n x
Binomial(n, p) p (1 − p)n−x np np(1 − p)
x
Geometric(p) 1 1−p 1−p
p(1 − p)x −1=
on {0, 1, 2, . . . } p p p2
Geometric(p) 1 1−p
p(1 − p)x−1
on {1, 2, 3, . . . }   p p2
Negative x + (r − 1) r (1 − p) r (1 − p)
· p r · (1 − p)x
Binomial r −1 p p2
−λ λx
Poisson e · λ λ
x!

138
B.0.1 Differentiation Exam P Handouts – Page 139

B.0 One Dimensional Calculus - Outline


B.0.1 One-Dimensional Derivatives
Definition of Derivative
Basic Formulas
Chain Rule
Product Rule
Absolute Values
Sines and Cosines
Further Examples

B.0.2 1-Dimensional integrals

B.0.3 Integration By Parts

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 1 / 42

Definition of Derivative (Background Only)


The average rate of change
from x to x + h is
total change
=
length of interval
f (x + h) − f (x)
f (x + h) =
h

f (x)
The derivative is the
instantaneous rate of change
x x+h
f (x + h) − f (x)
= lim
h→0 h
f (x + dx) − f (x) df
= =
dx dx

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 2 / 42


139
B.0.1 Differentiation Exam P Handouts – Page 140

Example
The definition is typically cumbersome to use.
For example,

d 2 (x + h)2 − x2
x = lim
dx h→0 h 
x2 + 2xh + h2 − x2
= lim
h→0 h
2xh + h2
= lim
h→0 h
= lim (2x + h)
h→0
= 2x

Instead of always doing this, in practice people use a smaller


number of key formulas.

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 3 / 42

Basic Formulas

d
f (x) = f 0 (x)
dx
d
a = 0 for any constant a
dx
d n
x = nxn−1
dx
d 1
[ln(x)] =
dx x
d x
e = ex
dx
d
[f (x) + g(x)] = f 0 (x) + g 0 (x)
dx
d d
[c f (x)] = c f (x) = cf 0 (x)
dx dx

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 4 / 42


140
B.0.1 Differentiation Exam P Handouts – Page 141

Basic Formulas: Examples

d 3
x = 3x2
dx
d 2
3y = 3 · 2y 1 = 6y
dy
d 
5et + 3t4 = 5et + 3 · 4t3
dt
d 2 d −3
= 2s
ds s3 ds
= −6s−4
−6
= 4
s

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 5 / 42

Chain Rule

Theorem (Chain Rule)


If f and u are differentiable functions,
d du
[f (u)] = f 0 (u) ·
dx dx

Examples

d n du
u = nun−1
dx dx
d u du
e = eu ·
dx dx

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 6 / 42


141
B.0.1 Differentiation Exam P Handouts – Page 142

Chain Rule: Examples


d  
Suppose we want to find exp −2t + t2 . Let u = −2t + t2 .
dt

d   d
exp −2t + t2 = exp [u]
dt dt
du
= exp [u] ·
 dx 
= exp −2t + t2 · (−2 + 2t)
d 5 4
2x2 + 5x + 3 = 5 2x2 + 5x + 3 ·(4x + 5)
dx
d     
exp 5et − 5 + 3t = exp 5et − 5 + 3t · 5et + 3
dt

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 7 / 42

Chain Rule: Examples


d x d x
Suppose we want to find 2 . We know how to find e ,
dx dx
so let’s rewrite 2x in terms of ex .
 x
ln(2) x ln(2)
2=e ,2 = e = ex ln(2)
d x d x ln(2)
2 = e = ln(2) · ex ln(2) = (ln 2) · 2x
dx dx
More generally, for any a,
d x
a =(ln a) · ax
dx
d x3 −3x d u
2 = 2 u = x3 − 3x
dx dx
du
= 2u · ln(2) ·
 dx  
3
= 2 x −3x · ln(2) · 3x2 − 3

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 8 / 42


142
B.0.1 Differentiation Exam P Handouts – Page 143

Product Rule
What about the derivatives of products or quotients of two
functions?
d  dv du
u·v =u· + ·v
dx dx dx
tiny (uv)0 = u · v 0 + u0 · v
dv u · dv
Quotients can be done by
rewriting them as products
v uv du · v
 
d 1 d 1 1
u· =u· + u0 ·
dx v dx v v
u du −u v 0 u0
= +
v2 v
−uv + u0 v
0
=
v2

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 9 / 42

Product Rule Examples

d  2 3x 
x e = x2 · 3e3x + 2x · e3x
dx
d e−x d  −x −3 
= e ·x
dx x3 dx
−3 1
= e−x · 4 + (−e−x ) · 3
x x
d h 3 i 3 d 4x
x2 + 3x + 5 · e4x = x2 + 3x + 5 · e
dx dx
d h 2 3 i 4x
+ x + 3x + 5 ·e
dx
3
= x2 + 3x + 5 · 4e4x
+ 3 · (x2 + 3x + 5)2 · (2x + 3) · e4x

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 10 / 42


143
B.0.1 Differentiation Exam P Handouts – Page 144

Absolute Values

y
|x| |x| = x if x ≥ 0
|x| = −x if x < 0
(
d 1 x>0
|x| =
x dx −1 x<0

d|x|
Note that is
dx
undefined if x = 0.

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 11 / 42

Sines and Cosines

cos(x) d
cos(x) = − sin(x)
dx
x

sin(x)

x d
sin(x) = cos(x)
dx

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 12 / 42


144
B.0.1 Differentiation Exam P Handouts – Page 145

Further Examples

d 2x + 5 (−1)(2x − 3) 2
= (2x + 5) +
dx x2 − 3x + 4 (x2 − 3x + 4)2 x2 − 3x + 4

d 2 2 2
(x + 2)ex −5x = (x + 2)(2x − 5)ex −5x + 1 · ex −5x
dx

d 2 −3x2 2 2
x e = x2 (−6x)e−3x + 2x · e−3x
dx

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 13 / 42

Further Examples

d d
sin |x + 2| = sin(x + 2) if x + 2 > 0
dx dx
= cos(x + 2) x > −2
d d
sin |x + 2| = sin(−x − 2) if x + 2 < 0
dx dx
= − cos(−x − 2) x < −2

Key point: We get two cases based on whether or not what is


inside the absolute value is positive. If x + 2 > 0 then what is
inside the absolute value is positive, so |x + 2| = x + 2, while if
x + 2 < 0 then what is inside the absolute value is negative so
|x + 2| = −x − 2.

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 14 / 42


145
B.0.1 Differentiation Exam P Handouts – Page 146

B.0 One Dimensional Calculus - Outline


B.0.1 One-Dimensional Derivatives

B.0.2 1-Dimensional integrals


What is an Integral?
The Fundamental Theorem of Calculus
Common Formulas
Substitution
Other Formulas

B.0.3 Integration By Parts

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 15 / 42

Definition of an Integral

f (x)

(x, f (x))

x
a x x + dx b

Zb
f (x) dx = area under curve
a
In some sense, f (x) dx is the area of an infinitely thin rectangle
and the integral says that the area under the curve is the sum
of the areas of infinitely many of these thin rectangles.
B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 16 / 42
146
B.0.1 Differentiation Exam P Handouts – Page 147

Geometric Examples
Often we can use geometry to find the integral/area under the
curve.

f (x) f (x)

x x
a a

Z a Z a
a2
2dx = 2a xdx =
0 0 2

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 17 / 42

Geometric Examples

In that example,
1. The integral of a constant was a linear function.
2. The integral of a line was a quadratic function.
So in these two examples, when we integrated the power of a
polynomial increased by 1.

When we differentiate,
1. The derivative of a linear function is a constant.
2. The derivative of a quadratic function is linear.
More generally, when we differentiate the power of a polynomial
decreases by 1. That is the opposite of when we integrate.

Hmmm.......isn’t that an interesting coincidence?

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 18 / 42


147
B.0.1 Differentiation Exam P Handouts – Page 148

The Fundamental Theorem of Calculus


f (t)

t
a x x+h

R
x+h Rx
Zx f (t) dt − f (t) dt
d d a a
F (x) = f (t) dt = lim
dx dx h→0 h
a
f (x) · h
= lim = f (x)
h→0 h

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 19 / 42

The Fundamental Theorem of Calculus

Theorem (Fundamental Theorem of Calculus)

Zx
d
f (t) dx = f (x)
dx
a

Generalization: If v and u are functions,


Zv Zv
d dv d dv du
f (t) dt = f (v) f (t) dt = f (v) − f (u)
dx dx dx dx dx
a u

In words, the Fundamental Theorem of Calculus says that


derivatives and integrals are inverse operations.

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 20 / 42


148
B.0.1 Differentiation Exam P Handouts – Page 149

Evaluating Integrals
The Fundamental Theorem of Calculus says that derivatives
and integrals are inverse operations. To find the integral of
f (x), we need to find a function whose derivative is f (x).

Examples
Z 5 5
1 52 02 25
x dx = x2 = − =
0 2 0 2 2 2
Zb b
n 1
x dx = · xn+1
n+1 a
a
bn+1 an+1
= −
n+1 n+1

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 21 / 42

Common Formulas

Z
a dx = ax + C
Z
x2
x dx = +C
2
Z
n xn+1
x dx = + C for n 6= −1
n+1
Z
1
ebx dx = ebx + C
b
Z Z
ax dx = ex ln a dx
1 x
= a +C
ln a

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 22 / 42


149
B.0.1 Differentiation Exam P Handouts – Page 150

Examples

Z 5 5
4 x5
3x dx = 3 ·
−2 5 −2
55 (−2)5
=3· −3·
Z 5 5
∞ ∞
3 3 1
dx = ·
2 x4 x3 −3 2
−1 −1
= 3− 3
∞ 2
1 1
=0+ =
8 8

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 23 / 42

Substitution

When we differentiate, we often have nested functions and need


to use the chain rule. For example,
d x2 2
e = ex · 2x
dx
d 2
(x + 3)5 = 5(x2 + 3)4 · 2x
dx
In both those examples, the 2x factor comes from the chain
rule. Often when we are doing integration, we will have a term
that we need to somehow recognize as a chain rule factor. If we
can do that, we can do a substitution to do the chain rule
“backwards.”

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 24 / 42


150
B.0.1 Differentiation Exam P Handouts – Page 151

Substitution

Suppose we want to integrate 2x · e(x ) . Let u = x2 . Then


2

du
= 2x so du = 2x dx and we get
dx

Z
x=b Z 2
u=b

2x e(x ) dx =
2
eu du
x=a u=a2
u=b2 x=b
= e(x )
u 2
=e
u=a2 x=a
b2 a2
=e −e

Note the limits! Either we convert back to x at the end, or we


change the limits to be in terms of u.

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 25 / 42

Substitution Examples

Z ∞
2
xe−2x dx u = 2x2 du = 4x dx
2
x = 2 u = 2 · 22 = 8
x = ∞ u = 2 · ∞2 = ∞
Z∞
du
= e−u ·
4
8

1 −1 −8
= · (−1) · e−u =0− e
4 8 4
1
= e−8
4

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 26 / 42


151
B.0.1 Differentiation Exam P Handouts – Page 152

Other Formulas

Z
dx
= ln x + C
x
Z Z
cf (x) dx = c f (x) dx
Z Z Z
 
f (x) + g(x) dx = f (x) dx + g(x) dx
Z
d
cos x dx = sin x + C because sin x = cos x
dx
Z
d
sin x dx = − cos x + C because cos x = − sin x
dx

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 27 / 42

Examples

u=x+5 x = 2, u = 2 + 5 = 7
du = dx x = 5, u = 5 + 5 = 10

Z 5 Z 10
3x 3(u − 5)
dx = du
2 (x + 5)2 7 u2
Z 10
3 15
= − du
7 u u2
 
15 10
= 3 ln u +
u 7
   
15 15
= 3 ln 10 + − 3 ln 7 +
10 7

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 28 / 42


152
B.0.1 Differentiation Exam P Handouts – Page 153

Examples

Z π π
(1 + cos t)dt = t + sin t
0 0

= (π + 0) − (0 + 0) = π
Z5 Z0 Z5
|x| dx = −x dx + x dx
−2 −2 0
0 5
−x2 x2 (−2)2 25 29
= + = + =
2 −2 2 0 2 2 2
Zx3
d 3 −5
e5t−5 dt = e5x · 3x2 − e5(−2x)−5 (−2)
dx
−2x

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 29 / 42

B.0 One Dimensional Calculus - Outline


B.0.1 One-Dimensional Derivatives

B.0.2 1-Dimensional integrals

B.0.3 Integration By Parts


Integration By Parts
Tabular integration
The Gamma trick

B.0 One Dimensional Calculus B.0.3 Integration By Parts 30 / 42


153
B.0.1 Differentiation Exam P Handouts – Page 154

Integration By Parts

Integration by parts is doing the product rule backwards.


d dv du
uv = u · + ·v
Z dx Z dx dx
Z
d (uv) = uv = u dv + v du
Z Z
u · dv = uv − v du

B.0 One Dimensional Calculus B.0.3 Integration By Parts 31 / 42

Integration By Parts

Example
Z Z
x ex dx = u dv

u=x dv = ex dx
du = dx v = ex
so
Z Z
x ex dx = uv − vdu
Z
x
= xe − ex dx

= xex − ex + C

B.0 One Dimensional Calculus B.0.3 Integration By Parts 32 / 42


154
B.0.1 Differentiation Exam P Handouts – Page 155

Logarithms
You can use integration by parts to handle functions whose
derivatives are easier to find than their integrals.

Z Z
ln x dx = u dv

u = ln x dv = dx
dx
du = v=x
x
so
Z Z
ln x dx = uv − vdu
Z
dx
= (ln x)x − x
x
= x ln x − x + C

B.0 One Dimensional Calculus B.0.3 Integration By Parts 33 / 42

Logarithms

Z Z
x ln x dx = u dv

u = ln x dv = x dx
dx x2
du = v=
x 2
so
Z Z
xln x dx = uv − vdu
Z 2
x2 x
= ln x − dx
2 2x
x2 x2
= ln x − +C
2 4

B.0 One Dimensional Calculus B.0.3 Integration By Parts 34 / 42


155
B.0.1 Differentiation Exam P Handouts – Page 156

Which Part to Use

How do we choose u and dv? We need dv to be something that


is easy to integrate and we need u to be something that is easy
to differentiate.
Ideally we also want u to become simpler when you differentiate.

Easy to Easy to
Differentiate Integrate
Logs Polynomials Exponentials

B.0 One Dimensional Calculus B.0.3 Integration By Parts 35 / 42

Integration by parts

Iterated Parts
Z
x2 e2x dx Let u = x2 dv = e2x dx
1
du = 2x dx v = e2x
Z 2
1 2x
2 1 2x
=x · e − e 2x dx
2 2
Z
1 2 2x
= x e − x e2x dx
2
And to find this, we have to repeat integration by parts.

B.0 One Dimensional Calculus B.0.3 Integration By Parts 36 / 42


156
B.0.1 Differentiation Exam P Handouts – Page 157

Tabular integration
Tabular integration is a way to organize our work when doing
repeated integration by parts. To integrate x2 e2x ,

Derivative column Integral column


x2 + e2x
1 2x
2x e
− 2
1 2x
2 e
+ 4
1 2x
0 − e
8

Z
1 1 1
x2 e2x dx = x2 · e2x − 2x · · e2x + 2 · e2x − 0
2 4 8

B.0 One Dimensional Calculus B.0.3 Integration By Parts 37 / 42

The Gamma trick


If we have a definite integral from 0 to infinity, we often can
skip using integration by parts.
If b > 0 and a is an integer, then
Z∞
a!
xa · e−bx dx =
ba+1
0

Example
Z ∞
2! 1
x2 e−2x dx = =
0 22+1 4

−b= −2 b=2
a= 2

B.0 One Dimensional Calculus B.0.3 Integration By Parts 38 / 42


157
B.0.1 Differentiation Exam P Handouts – Page 158

The Gamma trick Z ∞


An example both ways: 4x2 e−x/3 dx
0
Derivative column Integral column
4x2 + e−x/3

8x − − 3 e−x/3

8 + 9 e−x/3

0 − 27 e−x/3
Z
So 4x2 e−x/3 dx is

(4x2 )(−3 e−x/3 ) − (8x)(9 e−x/3 ) + 8 · (−27 e−x/3 )

= (−12x2 − 72x − 8 · 27) e−x/3

B.0 One Dimensional Calculus B.0.3 Integration By Parts 39 / 42

The Gamma trick


Z ∞ ∞
4x2 e−x/3 = (−12x2 − 72x − 8 · 27)e−x/3
0 0
=8 · 27

or we can let b = 1/3 and a = 2 in our formula to get


Z ∞
a!
xa e−bx = a+1
b
Z 0∞
2!
x2 e−x/3 = = 2 · 27
0 (1/3)2+1
Z ∞
2!
4x2 e−x/3 =4 · 2+1 = 8 · 27
0 1
3

B.0 One Dimensional Calculus B.0.3 Integration By Parts 40 / 42


158
B.0.1 Differentiation Exam P Handouts – Page 159

Other Definite Integrals


Z ∞
4x2 e−x/3 We want u = 0 when x = 1
1
u = x − 1, x=u+1
Z∞
= 4(u + 1)2 e(−u−1)/3 du
0
Z∞

= 4 e−1/3 u2 + 2u + 1 e−u/3 du
0

Z∞
a!
and now we plug into xa e−bx =
ba+1
0
" #
2! 1 1
= 4 e−1/3  +2·  +
1 3 1 2 1
3 3 3

B.0 One Dimensional Calculus B.0.3 Integration By Parts 41 / 42

The Gamma trick


Idea of Proof:

Let u = xa and dv = e−bx dx


Z ∞
∞ Z ∞  
a −bx −1 −bx a−1 −1
x e dx = x · a
e − ax e−bx dx
0 b 0 b
0
Z
a ∞ a−1 −bx
=0−0+ x e dx
b 0
a (a − 1)!
= · a−1+1
b b
a!
= a+1
b
It also is related to E [X a ] when X is an exponential random
variable as well as the density of a Gamma random variable.

B.0 One Dimensional Calculus B.0.3 Integration By Parts 42 / 42


159
B.0.2 Integration Exam P Handouts – Page 160

B.0 One Dimensional Calculus - Outline


B.0.1 One-Dimensional Derivatives
Definition of Derivative
Basic Formulas
Chain Rule
Product Rule
Absolute Values
Sines and Cosines
Further Examples

B.0.2 1-Dimensional integrals

B.0.3 Integration By Parts

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 1 / 42

Definition of Derivative (Background Only)


The average rate of change
from x to x + h is
total change
=
length of interval
f (x + h) − f (x)
f (x + h) =
h

f (x)
The derivative is the
instantaneous rate of change
x x+h
f (x + h) − f (x)
= lim
h→0 h
f (x + dx) − f (x) df
= =
dx dx

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 2 / 42


160
B.0.2 Integration Exam P Handouts – Page 161

Example
The definition is typically cumbersome to use.
For example,

d 2 (x + h)2 − x2
x = lim
dx h→0 h 
x2 + 2xh + h2 − x2
= lim
h→0 h
2xh + h2
= lim
h→0 h
= lim (2x + h)
h→0
= 2x

Instead of always doing this, in practice people use a smaller


number of key formulas.

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 3 / 42

Basic Formulas

d
f (x) = f 0 (x)
dx
d
a = 0 for any constant a
dx
d n
x = nxn−1
dx
d 1
[ln(x)] =
dx x
d x
e = ex
dx
d
[f (x) + g(x)] = f 0 (x) + g 0 (x)
dx
d d
[c f (x)] = c f (x) = cf 0 (x)
dx dx

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 4 / 42


161
B.0.2 Integration Exam P Handouts – Page 162

Basic Formulas: Examples

d 3
x = 3x2
dx
d 2
3y = 3 · 2y 1 = 6y
dy
d 
5et + 3t4 = 5et + 3 · 4t3
dt
d 2 d −3
= 2s
ds s3 ds
= −6s−4
−6
= 4
s

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 5 / 42

Chain Rule

Theorem (Chain Rule)


If f and u are differentiable functions,
d du
[f (u)] = f 0 (u) ·
dx dx

Examples

d n du
u = nun−1
dx dx
d u du
e = eu ·
dx dx

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 6 / 42


162
B.0.2 Integration Exam P Handouts – Page 163

Chain Rule: Examples


d  
Suppose we want to find exp −2t + t2 . Let u = −2t + t2 .
dt

d   d
exp −2t + t2 = exp [u]
dt dt
du
= exp [u] ·
 dx 
= exp −2t + t2 · (−2 + 2t)
d 5 4
2x2 + 5x + 3 = 5 2x2 + 5x + 3 ·(4x + 5)
dx
d     
exp 5et − 5 + 3t = exp 5et − 5 + 3t · 5et + 3
dt

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 7 / 42

Chain Rule: Examples


d x d x
Suppose we want to find 2 . We know how to find e ,
dx dx
so let’s rewrite 2x in terms of ex .
 x
ln(2) x ln(2)
2=e ,2 = e = ex ln(2)
d x d x ln(2)
2 = e = ln(2) · ex ln(2) = (ln 2) · 2x
dx dx
More generally, for any a,
d x
a =(ln a) · ax
dx
d x3 −3x d u
2 = 2 u = x3 − 3x
dx dx
du
= 2u · ln(2) ·
 dx  
3
= 2 x −3x · ln(2) · 3x2 − 3

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 8 / 42


163
B.0.2 Integration Exam P Handouts – Page 164

Product Rule
What about the derivatives of products or quotients of two
functions?
d  dv du
u·v =u· + ·v
dx dx dx
tiny (uv)0 = u · v 0 + u0 · v
dv u · dv
Quotients can be done by
rewriting them as products
v uv du · v
 
d 1 d 1 1
u· =u· + u0 ·
dx v dx v v
u du −u v 0 u0
= +
v2 v
−uv + u0 v
0
=
v2

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 9 / 42

Product Rule Examples

d  2 3x 
x e = x2 · 3e3x + 2x · e3x
dx
d e−x d  −x −3 
= e ·x
dx x3 dx
−3 1
= e−x · 4 + (−e−x ) · 3
x x
d h 3 i 3 d 4x
x2 + 3x + 5 · e4x = x2 + 3x + 5 · e
dx dx
d h 2 3 i 4x
+ x + 3x + 5 ·e
dx
3
= x2 + 3x + 5 · 4e4x
+ 3 · (x2 + 3x + 5)2 · (2x + 3) · e4x

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 10 / 42


164
B.0.2 Integration Exam P Handouts – Page 165

Absolute Values

y
|x| |x| = x if x ≥ 0
|x| = −x if x < 0
(
d 1 x>0
|x| =
x dx −1 x<0

d|x|
Note that is
dx
undefined if x = 0.

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 11 / 42

Sines and Cosines

cos(x) d
cos(x) = − sin(x)
dx
x

sin(x)

x d
sin(x) = cos(x)
dx

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 12 / 42


165
B.0.2 Integration Exam P Handouts – Page 166

Further Examples

d 2x + 5 (−1)(2x − 3) 2
= (2x + 5) +
dx x2 − 3x + 4 (x2 − 3x + 4)2 x2 − 3x + 4

d 2 2 2
(x + 2)ex −5x = (x + 2)(2x − 5)ex −5x + 1 · ex −5x
dx

d 2 −3x2 2 2
x e = x2 (−6x)e−3x + 2x · e−3x
dx

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 13 / 42

Further Examples

d d
sin |x + 2| = sin(x + 2) if x + 2 > 0
dx dx
= cos(x + 2) x > −2
d d
sin |x + 2| = sin(−x − 2) if x + 2 < 0
dx dx
= − cos(−x − 2) x < −2

Key point: We get two cases based on whether or not what is


inside the absolute value is positive. If x + 2 > 0 then what is
inside the absolute value is positive, so |x + 2| = x + 2, while if
x + 2 < 0 then what is inside the absolute value is negative so
|x + 2| = −x − 2.

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 14 / 42


166
B.0.2 Integration Exam P Handouts – Page 167

B.0 One Dimensional Calculus - Outline


B.0.1 One-Dimensional Derivatives

B.0.2 1-Dimensional integrals


What is an Integral?
The Fundamental Theorem of Calculus
Common Formulas
Substitution
Other Formulas

B.0.3 Integration By Parts

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 15 / 42

Definition of an Integral

f (x)

(x, f (x))

x
a x x + dx b

Zb
f (x) dx = area under curve
a
In some sense, f (x) dx is the area of an infinitely thin rectangle
and the integral says that the area under the curve is the sum
of the areas of infinitely many of these thin rectangles.
B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 16 / 42
167
B.0.2 Integration Exam P Handouts – Page 168

Geometric Examples
Often we can use geometry to find the integral/area under the
curve.

f (x) f (x)

x x
a a

Z a Z a
a2
2dx = 2a xdx =
0 0 2

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 17 / 42

Geometric Examples

In that example,
1. The integral of a constant was a linear function.
2. The integral of a line was a quadratic function.
So in these two examples, when we integrated the power of a
polynomial increased by 1.

When we differentiate,
1. The derivative of a linear function is a constant.
2. The derivative of a quadratic function is linear.
More generally, when we differentiate the power of a polynomial
decreases by 1. That is the opposite of when we integrate.

Hmmm.......isn’t that an interesting coincidence?

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 18 / 42


168
B.0.2 Integration Exam P Handouts – Page 169

The Fundamental Theorem of Calculus


f (t)

t
a x x+h

R
x+h Rx
Zx f (t) dt − f (t) dt
d d a a
F (x) = f (t) dt = lim
dx dx h→0 h
a
f (x) · h
= lim = f (x)
h→0 h

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 19 / 42

The Fundamental Theorem of Calculus

Theorem (Fundamental Theorem of Calculus)

Zx
d
f (t) dx = f (x)
dx
a

Generalization: If v and u are functions,


Zv Zv
d dv d dv du
f (t) dt = f (v) f (t) dt = f (v) − f (u)
dx dx dx dx dx
a u

In words, the Fundamental Theorem of Calculus says that


derivatives and integrals are inverse operations.

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 20 / 42


169
B.0.2 Integration Exam P Handouts – Page 170

Evaluating Integrals
The Fundamental Theorem of Calculus says that derivatives
and integrals are inverse operations. To find the integral of
f (x), we need to find a function whose derivative is f (x).

Examples
Z 5 5
1 52 02 25
x dx = x2 = − =
0 2 0 2 2 2
Zb b
n 1
x dx = · xn+1
n+1 a
a
bn+1 an+1
= −
n+1 n+1

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 21 / 42

Common Formulas

Z
a dx = ax + C
Z
x2
x dx = +C
2
Z
n xn+1
x dx = + C for n 6= −1
n+1
Z
1
ebx dx = ebx + C
b
Z Z
ax dx = ex ln a dx
1 x
= a +C
ln a

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 22 / 42


170
B.0.2 Integration Exam P Handouts – Page 171

Examples

Z 5 5
4 x5
3x dx = 3 ·
−2 5 −2
55 (−2)5
=3· −3·
Z 5 5
∞ ∞
3 3 1
dx = ·
2 x4 x3 −3 2
−1 −1
= 3− 3
∞ 2
1 1
=0+ =
8 8

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 23 / 42

Substitution

When we differentiate, we often have nested functions and need


to use the chain rule. For example,
d x2 2
e = ex · 2x
dx
d 2
(x + 3)5 = 5(x2 + 3)4 · 2x
dx
In both those examples, the 2x factor comes from the chain
rule. Often when we are doing integration, we will have a term
that we need to somehow recognize as a chain rule factor. If we
can do that, we can do a substitution to do the chain rule
“backwards.”

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 24 / 42


171
B.0.2 Integration Exam P Handouts – Page 172

Substitution

Suppose we want to integrate 2x · e(x ) . Let u = x2 . Then


2

du
= 2x so du = 2x dx and we get
dx

Z
x=b Z 2
u=b

2x e(x ) dx =
2
eu du
x=a u=a2
u=b2 x=b
= e(x )
u 2
=e
u=a2 x=a
b2 a2
=e −e

Note the limits! Either we convert back to x at the end, or we


change the limits to be in terms of u.

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 25 / 42

Substitution Examples

Z ∞
2
xe−2x dx u = 2x2 du = 4x dx
2
x = 2 u = 2 · 22 = 8
x = ∞ u = 2 · ∞2 = ∞
Z∞
du
= e−u ·
4
8

1 −1 −8
= · (−1) · e−u =0− e
4 8 4
1
= e−8
4

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 26 / 42


172
B.0.2 Integration Exam P Handouts – Page 173

Other Formulas

Z
dx
= ln x + C
x
Z Z
cf (x) dx = c f (x) dx
Z Z Z
 
f (x) + g(x) dx = f (x) dx + g(x) dx
Z
d
cos x dx = sin x + C because sin x = cos x
dx
Z
d
sin x dx = − cos x + C because cos x = − sin x
dx

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 27 / 42

Examples

u=x+5 x = 2, u = 2 + 5 = 7
du = dx x = 5, u = 5 + 5 = 10

Z 5 Z 10
3x 3(u − 5)
dx = du
2 (x + 5)2 7 u2
Z 10
3 15
= − du
7 u u2
 
15 10
= 3 ln u +
u 7
   
15 15
= 3 ln 10 + − 3 ln 7 +
10 7

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 28 / 42


173
B.0.2 Integration Exam P Handouts – Page 174

Examples

Z π π
(1 + cos t)dt = t + sin t
0 0

= (π + 0) − (0 + 0) = π
Z5 Z0 Z5
|x| dx = −x dx + x dx
−2 −2 0
0 5
−x2 x2 (−2)2 25 29
= + = + =
2 −2 2 0 2 2 2
Zx3
d 3 −5
e5t−5 dt = e5x · 3x2 − e5(−2x)−5 (−2)
dx
−2x

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 29 / 42

B.0 One Dimensional Calculus - Outline


B.0.1 One-Dimensional Derivatives

B.0.2 1-Dimensional integrals

B.0.3 Integration By Parts


Integration By Parts
Tabular integration
The Gamma trick

B.0 One Dimensional Calculus B.0.3 Integration By Parts 30 / 42


174
B.0.2 Integration Exam P Handouts – Page 175

Integration By Parts

Integration by parts is doing the product rule backwards.


d dv du
uv = u · + ·v
Z dx Z dx dx
Z
d (uv) = uv = u dv + v du
Z Z
u · dv = uv − v du

B.0 One Dimensional Calculus B.0.3 Integration By Parts 31 / 42

Integration By Parts

Example
Z Z
x ex dx = u dv

u=x dv = ex dx
du = dx v = ex
so
Z Z
x ex dx = uv − vdu
Z
x
= xe − ex dx

= xex − ex + C

B.0 One Dimensional Calculus B.0.3 Integration By Parts 32 / 42


175
B.0.2 Integration Exam P Handouts – Page 176

Logarithms
You can use integration by parts to handle functions whose
derivatives are easier to find than their integrals.

Z Z
ln x dx = u dv

u = ln x dv = dx
dx
du = v=x
x
so
Z Z
ln x dx = uv − vdu
Z
dx
= (ln x)x − x
x
= x ln x − x + C

B.0 One Dimensional Calculus B.0.3 Integration By Parts 33 / 42

Logarithms

Z Z
x ln x dx = u dv

u = ln x dv = x dx
dx x2
du = v=
x 2
so
Z Z
xln x dx = uv − vdu
Z 2
x2 x
= ln x − dx
2 2x
x2 x2
= ln x − +C
2 4

B.0 One Dimensional Calculus B.0.3 Integration By Parts 34 / 42


176
B.0.2 Integration Exam P Handouts – Page 177

Which Part to Use

How do we choose u and dv? We need dv to be something that


is easy to integrate and we need u to be something that is easy
to differentiate.
Ideally we also want u to become simpler when you differentiate.

Easy to Easy to
Differentiate Integrate
Logs Polynomials Exponentials

B.0 One Dimensional Calculus B.0.3 Integration By Parts 35 / 42

Integration by parts

Iterated Parts
Z
x2 e2x dx Let u = x2 dv = e2x dx
1
du = 2x dx v = e2x
Z 2
1 2x
2 1 2x
=x · e − e 2x dx
2 2
Z
1 2 2x
= x e − x e2x dx
2
And to find this, we have to repeat integration by parts.

B.0 One Dimensional Calculus B.0.3 Integration By Parts 36 / 42


177
B.0.2 Integration Exam P Handouts – Page 178

Tabular integration
Tabular integration is a way to organize our work when doing
repeated integration by parts. To integrate x2 e2x ,

Derivative column Integral column


x2 + e2x
1 2x
2x e
− 2
1 2x
2 e
+ 4
1 2x
0 − e
8

Z
1 1 1
x2 e2x dx = x2 · e2x − 2x · · e2x + 2 · e2x − 0
2 4 8

B.0 One Dimensional Calculus B.0.3 Integration By Parts 37 / 42

The Gamma trick


If we have a definite integral from 0 to infinity, we often can
skip using integration by parts.
If b > 0 and a is an integer, then
Z∞
a!
xa · e−bx dx =
ba+1
0

Example
Z ∞
2! 1
x2 e−2x dx = =
0 22+1 4

−b= −2 b=2
a= 2

B.0 One Dimensional Calculus B.0.3 Integration By Parts 38 / 42


178
B.0.2 Integration Exam P Handouts – Page 179

The Gamma trick Z ∞


An example both ways: 4x2 e−x/3 dx
0
Derivative column Integral column
4x2 + e−x/3

8x − − 3 e−x/3

8 + 9 e−x/3

0 − 27 e−x/3
Z
So 4x2 e−x/3 dx is

(4x2 )(−3 e−x/3 ) − (8x)(9 e−x/3 ) + 8 · (−27 e−x/3 )

= (−12x2 − 72x − 8 · 27) e−x/3

B.0 One Dimensional Calculus B.0.3 Integration By Parts 39 / 42

The Gamma trick


Z ∞ ∞
4x2 e−x/3 = (−12x2 − 72x − 8 · 27)e−x/3
0 0
=8 · 27

or we can let b = 1/3 and a = 2 in our formula to get


Z ∞
a!
xa e−bx = a+1
b
Z 0∞
2!
x2 e−x/3 = = 2 · 27
0 (1/3)2+1
Z ∞
2!
4x2 e−x/3 =4 · 2+1 = 8 · 27
0 1
3

B.0 One Dimensional Calculus B.0.3 Integration By Parts 40 / 42


179
B.0.2 Integration Exam P Handouts – Page 180

Other Definite Integrals


Z ∞
4x2 e−x/3 We want u = 0 when x = 1
1
u = x − 1, x=u+1
Z∞
= 4(u + 1)2 e(−u−1)/3 du
0
Z∞

= 4 e−1/3 u2 + 2u + 1 e−u/3 du
0

Z∞
a!
and now we plug into xa e−bx =
ba+1
0
" #
2! 1 1
= 4 e−1/3  +2·  +
1 3 1 2 1
3 3 3

B.0 One Dimensional Calculus B.0.3 Integration By Parts 41 / 42

The Gamma trick


Idea of Proof:

Let u = xa and dv = e−bx dx


Z ∞
∞ Z ∞  
a −bx −1 −bx a−1 −1
x e dx = x · a
e − ax e−bx dx
0 b 0 b
0
Z
a ∞ a−1 −bx
=0−0+ x e dx
b 0
a (a − 1)!
= · a−1+1
b b
a!
= a+1
b
It also is related to E [X a ] when X is an exponential random
variable as well as the density of a Gamma random variable.

B.0 One Dimensional Calculus B.0.3 Integration By Parts 42 / 42


180
B.0.3 Integration by Parts Exam P Handouts – Page 181

B.0 One Dimensional Calculus - Outline


B.0.1 One-Dimensional Derivatives
Definition of Derivative
Basic Formulas
Chain Rule
Product Rule
Absolute Values
Sines and Cosines
Further Examples

B.0.2 1-Dimensional integrals

B.0.3 Integration By Parts

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 1 / 42

Definition of Derivative (Background Only)


The average rate of change
from x to x + h is
total change
=
length of interval
f (x + h) − f (x)
f (x + h) =
h

f (x)
The derivative is the
instantaneous rate of change
x x+h
f (x + h) − f (x)
= lim
h→0 h
f (x + dx) − f (x) df
= =
dx dx

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 2 / 42


181
B.0.3 Integration by Parts Exam P Handouts – Page 182

Example
The definition is typically cumbersome to use.
For example,

d 2 (x + h)2 − x2
x = lim
dx h→0 h 
x2 + 2xh + h2 − x2
= lim
h→0 h
2xh + h2
= lim
h→0 h
= lim (2x + h)
h→0
= 2x

Instead of always doing this, in practice people use a smaller


number of key formulas.

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 3 / 42

Basic Formulas

d
f (x) = f 0 (x)
dx
d
a = 0 for any constant a
dx
d n
x = nxn−1
dx
d 1
[ln(x)] =
dx x
d x
e = ex
dx
d
[f (x) + g(x)] = f 0 (x) + g 0 (x)
dx
d d
[c f (x)] = c f (x) = cf 0 (x)
dx dx

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 4 / 42


182
B.0.3 Integration by Parts Exam P Handouts – Page 183

Basic Formulas: Examples

d 3
x = 3x2
dx
d 2
3y = 3 · 2y 1 = 6y
dy
d 
5et + 3t4 = 5et + 3 · 4t3
dt
d 2 d −3
= 2s
ds s3 ds
= −6s−4
−6
= 4
s

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 5 / 42

Chain Rule

Theorem (Chain Rule)


If f and u are differentiable functions,
d du
[f (u)] = f 0 (u) ·
dx dx

Examples

d n du
u = nun−1
dx dx
d u du
e = eu ·
dx dx

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 6 / 42


183
B.0.3 Integration by Parts Exam P Handouts – Page 184

Chain Rule: Examples


d  
Suppose we want to find exp −2t + t2 . Let u = −2t + t2 .
dt

d   d
exp −2t + t2 = exp [u]
dt dt
du
= exp [u] ·
 dx 
= exp −2t + t2 · (−2 + 2t)
d 5 4
2x2 + 5x + 3 = 5 2x2 + 5x + 3 ·(4x + 5)
dx
d     
exp 5et − 5 + 3t = exp 5et − 5 + 3t · 5et + 3
dt

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 7 / 42

Chain Rule: Examples


d x d x
Suppose we want to find 2 . We know how to find e ,
dx dx
so let’s rewrite 2x in terms of ex .
 x
ln(2) x ln(2)
2=e ,2 = e = ex ln(2)
d x d x ln(2)
2 = e = ln(2) · ex ln(2) = (ln 2) · 2x
dx dx
More generally, for any a,
d x
a =(ln a) · ax
dx
d x3 −3x d u
2 = 2 u = x3 − 3x
dx dx
du
= 2u · ln(2) ·
 dx  
3
= 2 x −3x · ln(2) · 3x2 − 3

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 8 / 42


184
B.0.3 Integration by Parts Exam P Handouts – Page 185

Product Rule
What about the derivatives of products or quotients of two
functions?
d  dv du
u·v =u· + ·v
dx dx dx
tiny (uv)0 = u · v 0 + u0 · v
dv u · dv
Quotients can be done by
rewriting them as products
v uv du · v
 
d 1 d 1 1
u· =u· + u0 ·
dx v dx v v
u du −u v 0 u0
= +
v2 v
−uv + u0 v
0
=
v2

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 9 / 42

Product Rule Examples

d  2 3x 
x e = x2 · 3e3x + 2x · e3x
dx
d e−x d  −x −3 
= e ·x
dx x3 dx
−3 1
= e−x · 4 + (−e−x ) · 3
x x
d h 3 i 3 d 4x
x2 + 3x + 5 · e4x = x2 + 3x + 5 · e
dx dx
d h 2 3 i 4x
+ x + 3x + 5 ·e
dx
3
= x2 + 3x + 5 · 4e4x
+ 3 · (x2 + 3x + 5)2 · (2x + 3) · e4x

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 10 / 42


185
B.0.3 Integration by Parts Exam P Handouts – Page 186

Absolute Values

y
|x| |x| = x if x ≥ 0
|x| = −x if x < 0
(
d 1 x>0
|x| =
x dx −1 x<0

d|x|
Note that is
dx
undefined if x = 0.

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 11 / 42

Sines and Cosines

cos(x) d
cos(x) = − sin(x)
dx
x

sin(x)

x d
sin(x) = cos(x)
dx

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 12 / 42


186
B.0.3 Integration by Parts Exam P Handouts – Page 187

Further Examples

d 2x + 5 (−1)(2x − 3) 2
= (2x + 5) +
dx x2 − 3x + 4 (x2 − 3x + 4)2 x2 − 3x + 4

d 2 2 2
(x + 2)ex −5x = (x + 2)(2x − 5)ex −5x + 1 · ex −5x
dx

d 2 −3x2 2 2
x e = x2 (−6x)e−3x + 2x · e−3x
dx

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 13 / 42

Further Examples

d d
sin |x + 2| = sin(x + 2) if x + 2 > 0
dx dx
= cos(x + 2) x > −2
d d
sin |x + 2| = sin(−x − 2) if x + 2 < 0
dx dx
= − cos(−x − 2) x < −2

Key point: We get two cases based on whether or not what is


inside the absolute value is positive. If x + 2 > 0 then what is
inside the absolute value is positive, so |x + 2| = x + 2, while if
x + 2 < 0 then what is inside the absolute value is negative so
|x + 2| = −x − 2.

B.0 One Dimensional Calculus B.0.1 One-Dimensional Derivatives 14 / 42


187
B.0.3 Integration by Parts Exam P Handouts – Page 188

B.0 One Dimensional Calculus - Outline


B.0.1 One-Dimensional Derivatives

B.0.2 1-Dimensional integrals


What is an Integral?
The Fundamental Theorem of Calculus
Common Formulas
Substitution
Other Formulas

B.0.3 Integration By Parts

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 15 / 42

Definition of an Integral

f (x)

(x, f (x))

x
a x x + dx b

Zb
f (x) dx = area under curve
a
In some sense, f (x) dx is the area of an infinitely thin rectangle
and the integral says that the area under the curve is the sum
of the areas of infinitely many of these thin rectangles.
B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 16 / 42
188
B.0.3 Integration by Parts Exam P Handouts – Page 189

Geometric Examples
Often we can use geometry to find the integral/area under the
curve.

f (x) f (x)

x x
a a

Z a Z a
a2
2dx = 2a xdx =
0 0 2

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 17 / 42

Geometric Examples

In that example,
1. The integral of a constant was a linear function.
2. The integral of a line was a quadratic function.
So in these two examples, when we integrated the power of a
polynomial increased by 1.

When we differentiate,
1. The derivative of a linear function is a constant.
2. The derivative of a quadratic function is linear.
More generally, when we differentiate the power of a polynomial
decreases by 1. That is the opposite of when we integrate.

Hmmm.......isn’t that an interesting coincidence?

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 18 / 42


189
B.0.3 Integration by Parts Exam P Handouts – Page 190

The Fundamental Theorem of Calculus


f (t)

t
a x x+h

R
x+h Rx
Zx f (t) dt − f (t) dt
d d a a
F (x) = f (t) dt = lim
dx dx h→0 h
a
f (x) · h
= lim = f (x)
h→0 h

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 19 / 42

The Fundamental Theorem of Calculus

Theorem (Fundamental Theorem of Calculus)

Zx
d
f (t) dx = f (x)
dx
a

Generalization: If v and u are functions,


Zv Zv
d dv d dv du
f (t) dt = f (v) f (t) dt = f (v) − f (u)
dx dx dx dx dx
a u

In words, the Fundamental Theorem of Calculus says that


derivatives and integrals are inverse operations.

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 20 / 42


190
B.0.3 Integration by Parts Exam P Handouts – Page 191

Evaluating Integrals
The Fundamental Theorem of Calculus says that derivatives
and integrals are inverse operations. To find the integral of
f (x), we need to find a function whose derivative is f (x).

Examples
Z 5 5
1 52 02 25
x dx = x2 = − =
0 2 0 2 2 2
Zb b
n 1
x dx = · xn+1
n+1 a
a
bn+1 an+1
= −
n+1 n+1

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 21 / 42

Common Formulas

Z
a dx = ax + C
Z
x2
x dx = +C
2
Z
n xn+1
x dx = + C for n 6= −1
n+1
Z
1
ebx dx = ebx + C
b
Z Z
ax dx = ex ln a dx
1 x
= a +C
ln a

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 22 / 42


191
B.0.3 Integration by Parts Exam P Handouts – Page 192

Examples

Z 5 5
4 x5
3x dx = 3 ·
−2 5 −2
55 (−2)5
=3· −3·
Z 5 5
∞ ∞
3 3 1
dx = ·
2 x4 x3 −3 2
−1 −1
= 3− 3
∞ 2
1 1
=0+ =
8 8

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 23 / 42

Substitution

When we differentiate, we often have nested functions and need


to use the chain rule. For example,
d x2 2
e = ex · 2x
dx
d 2
(x + 3)5 = 5(x2 + 3)4 · 2x
dx
In both those examples, the 2x factor comes from the chain
rule. Often when we are doing integration, we will have a term
that we need to somehow recognize as a chain rule factor. If we
can do that, we can do a substitution to do the chain rule
“backwards.”

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 24 / 42


192
B.0.3 Integration by Parts Exam P Handouts – Page 193

Substitution

Suppose we want to integrate 2x · e(x ) . Let u = x2 . Then


2

du
= 2x so du = 2x dx and we get
dx

Z
x=b Z 2
u=b

2x e(x ) dx =
2
eu du
x=a u=a2
u=b2 x=b
= e(x )
u 2
=e
u=a2 x=a
b2 a2
=e −e

Note the limits! Either we convert back to x at the end, or we


change the limits to be in terms of u.

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 25 / 42

Substitution Examples

Z ∞
2
xe−2x dx u = 2x2 du = 4x dx
2
x = 2 u = 2 · 22 = 8
x = ∞ u = 2 · ∞2 = ∞
Z∞
du
= e−u ·
4
8

1 −1 −8
= · (−1) · e−u =0− e
4 8 4
1
= e−8
4

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 26 / 42


193
B.0.3 Integration by Parts Exam P Handouts – Page 194

Other Formulas

Z
dx
= ln x + C
x
Z Z
cf (x) dx = c f (x) dx
Z Z Z
 
f (x) + g(x) dx = f (x) dx + g(x) dx
Z
d
cos x dx = sin x + C because sin x = cos x
dx
Z
d
sin x dx = − cos x + C because cos x = − sin x
dx

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 27 / 42

Examples

u=x+5 x = 2, u = 2 + 5 = 7
du = dx x = 5, u = 5 + 5 = 10

Z 5 Z 10
3x 3(u − 5)
dx = du
2 (x + 5)2 7 u2
Z 10
3 15
= − du
7 u u2
 
15 10
= 3 ln u +
u 7
   
15 15
= 3 ln 10 + − 3 ln 7 +
10 7

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 28 / 42


194
B.0.3 Integration by Parts Exam P Handouts – Page 195

Examples

Z π π
(1 + cos t)dt = t + sin t
0 0

= (π + 0) − (0 + 0) = π
Z5 Z0 Z5
|x| dx = −x dx + x dx
−2 −2 0
0 5
−x2 x2 (−2)2 25 29
= + = + =
2 −2 2 0 2 2 2
Zx3
d 3 −5
e5t−5 dt = e5x · 3x2 − e5(−2x)−5 (−2)
dx
−2x

B.0 One Dimensional Calculus B.0.2 1-Dimensional integrals 29 / 42

B.0 One Dimensional Calculus - Outline


B.0.1 One-Dimensional Derivatives

B.0.2 1-Dimensional integrals

B.0.3 Integration By Parts


Integration By Parts
Tabular integration
The Gamma trick

B.0 One Dimensional Calculus B.0.3 Integration By Parts 30 / 42


195
B.0.3 Integration by Parts Exam P Handouts – Page 196

Integration By Parts

Integration by parts is doing the product rule backwards.


d dv du
uv = u · + ·v
Z dx Z dx dx
Z
d (uv) = uv = u dv + v du
Z Z
u · dv = uv − v du

B.0 One Dimensional Calculus B.0.3 Integration By Parts 31 / 42

Integration By Parts

Example
Z Z
x ex dx = u dv

u=x dv = ex dx
du = dx v = ex
so
Z Z
x ex dx = uv − vdu
Z
x
= xe − ex dx

= xex − ex + C

B.0 One Dimensional Calculus B.0.3 Integration By Parts 32 / 42


196
B.0.3 Integration by Parts Exam P Handouts – Page 197

Logarithms
You can use integration by parts to handle functions whose
derivatives are easier to find than their integrals.

Z Z
ln x dx = u dv

u = ln x dv = dx
dx
du = v=x
x
so
Z Z
ln x dx = uv − vdu
Z
dx
= (ln x)x − x
x
= x ln x − x + C

B.0 One Dimensional Calculus B.0.3 Integration By Parts 33 / 42

Logarithms

Z Z
x ln x dx = u dv

u = ln x dv = x dx
dx x2
du = v=
x 2
so
Z Z
xln x dx = uv − vdu
Z 2
x2 x
= ln x − dx
2 2x
x2 x2
= ln x − +C
2 4

B.0 One Dimensional Calculus B.0.3 Integration By Parts 34 / 42


197
B.0.3 Integration by Parts Exam P Handouts – Page 198

Which Part to Use

How do we choose u and dv? We need dv to be something that


is easy to integrate and we need u to be something that is easy
to differentiate.
Ideally we also want u to become simpler when you differentiate.

Easy to Easy to
Differentiate Integrate
Logs Polynomials Exponentials

B.0 One Dimensional Calculus B.0.3 Integration By Parts 35 / 42

Integration by parts

Iterated Parts
Z
x2 e2x dx Let u = x2 dv = e2x dx
1
du = 2x dx v = e2x
Z 2
1 2x
2 1 2x
=x · e − e 2x dx
2 2
Z
1 2 2x
= x e − x e2x dx
2
And to find this, we have to repeat integration by parts.

B.0 One Dimensional Calculus B.0.3 Integration By Parts 36 / 42


198
B.0.3 Integration by Parts Exam P Handouts – Page 199

Tabular integration
Tabular integration is a way to organize our work when doing
repeated integration by parts. To integrate x2 e2x ,

Derivative column Integral column


x2 + e2x
1 2x
2x e
− 2
1 2x
2 e
+ 4
1 2x
0 − e
8

Z
1 1 1
x2 e2x dx = x2 · e2x − 2x · · e2x + 2 · e2x − 0
2 4 8

B.0 One Dimensional Calculus B.0.3 Integration By Parts 37 / 42

The Gamma trick


If we have a definite integral from 0 to infinity, we often can
skip using integration by parts.
If b > 0 and a is an integer, then
Z∞
a!
xa · e−bx dx =
ba+1
0

Example
Z ∞
2! 1
x2 e−2x dx = =
0 22+1 4

−b= −2 b=2
a= 2

B.0 One Dimensional Calculus B.0.3 Integration By Parts 38 / 42


199
B.0.3 Integration by Parts Exam P Handouts – Page 200

The Gamma trick Z ∞


An example both ways: 4x2 e−x/3 dx
0
Derivative column Integral column
4x2 + e−x/3

8x − − 3 e−x/3

8 + 9 e−x/3

0 − 27 e−x/3
Z
So 4x2 e−x/3 dx is

(4x2 )(−3 e−x/3 ) − (8x)(9 e−x/3 ) + 8 · (−27 e−x/3 )

= (−12x2 − 72x − 8 · 27) e−x/3

B.0 One Dimensional Calculus B.0.3 Integration By Parts 39 / 42

The Gamma trick


Z ∞ ∞
4x2 e−x/3 = (−12x2 − 72x − 8 · 27)e−x/3
0 0
=8 · 27

or we can let b = 1/3 and a = 2 in our formula to get


Z ∞
a!
xa e−bx = a+1
b
Z 0∞
2!
x2 e−x/3 = = 2 · 27
0 (1/3)2+1
Z ∞
2!
4x2 e−x/3 =4 · 2+1 = 8 · 27
0 1
3

B.0 One Dimensional Calculus B.0.3 Integration By Parts 40 / 42


200
B.0.3 Integration by Parts Exam P Handouts – Page 201

Other Definite Integrals


Z ∞
4x2 e−x/3 We want u = 0 when x = 1
1
u = x − 1, x=u+1
Z∞
= 4(u + 1)2 e(−u−1)/3 du
0
Z∞

= 4 e−1/3 u2 + 2u + 1 e−u/3 du
0

Z∞
a!
and now we plug into xa e−bx =
ba+1
0
" #
2! 1 1
= 4 e−1/3  +2·  +
1 3 1 2 1
3 3 3

B.0 One Dimensional Calculus B.0.3 Integration By Parts 41 / 42

The Gamma trick


Idea of Proof:

Let u = xa and dv = e−bx dx


Z ∞
∞ Z ∞  
a −bx −1 −bx a−1 −1
x e dx = x · a
e − ax e−bx dx
0 b 0 b
0
Z
a ∞ a−1 −bx
=0−0+ x e dx
b 0
a (a − 1)!
= · a−1+1
b b
a!
= a+1
b
It also is related to E [X a ] when X is an exponential random
variable as well as the density of a Gamma random variable.

B.0 One Dimensional Calculus B.0.3 Integration By Parts 42 / 42


201
B.1.1 Continuous Distributions: Overview Exam P Handouts – Page 202

Continuous Distributions: Overview 1

Discrete vs Continuous
Mixed Distributions
Exercises

Comparing Uniforms 2

Let N be an integer uniformly chosen from {1, 2, . . . , 100} and let X be a real number
uniformly chosen from (0, 100). Then

1
P[N = n] = n = 1, 2, 3, . . . , 100
100
n
P[N ≤ n] = n = 1, 2, 3, . . . , 100
100

P[X = x] = 0 for all x


x
P[X ≤ x] = 0 ≤ x ≤ 100
100
Point: The cumulative distribution function F (x) = P[X ≤ x] still makes sense for
continuous distributions, and will still be useful.

P[X = x] = 0 for all x for a purely continuous function, and so we’ll need a slightly
different idea.

202
B.1.1 Continuous Distributions: Overview Exam P Handouts – Page 203

Discrete vs Continuous 3

For discrete random variables, we often summed expressions that involved P[X = x],
such as X
E[X ] = x · P[X = x]

In the continuous case, the sums will become integrals and the “density” of X ,
denoted f (x), will replace P[X = x] in most formulas. For example,
Z
E[X ] = x · f (x)dx

We will define f (x) in the next lesson.

Mixed Distributions 4

Not all distributions are purely discrete or purely continuous.

A mixed distribution is some of each.

On the exam, mixed distributions often come from adding deductibles and limits to
continuous loss amounts.

203
B.1.1 Continuous Distributions: Overview Exam P Handouts – Page 204

Mixed Example 5

Losses X are uniformly distributed on (0, 100). Let Y be the payment amount after a
deductible of 30 is applied to the loss.

The deductible of 30 means that Y = 0 if the loss X is less than 30, and Y = X − 30
if the loss exceeds 30. That means that
30
P[Y = 0] = P[X ≤ 30] =
100
P[Y = y ] = 0 for y > 0
y + 30
P[Y ≤ y ] = P[X ≤ y + 30] = for 0 < y < 70
100
so Y has a discrete piece (at 0), a continuous piece (from 0 to 70), and the cdf makes
sense everywhere so we can still use it to study Y .

Exercise 1 6

If N is uniform on {1, 2, 3, 4, 5}, and X is uniform on (0, 5), find P[N ≤ 2.3] and
P[X ≤ 2.3].

204
B.1.1 Continuous Distributions: Overview Exam P Handouts – Page 205

Exercise 1 6

If N is uniform on {1, 2, 3, 4, 5}, and X is uniform on (0, 5), find P[N ≤ 2.3] and
P[X ≤ 2.3].

P[N ≤ 2.3] = P[N = 1 or 2]


2
=
5
2.3
P[X ≤ 2.3] =
5

Exercise 2 7

Loss amounts are uniform on the interval (0, 6), and insured with a deductible of 1.6.
Find the probabilities that:
a) the payment for a randomly chosen loss is 0
b) the payment for a randomly chosen loss is less than 2.

205
B.1.1 Continuous Distributions: Overview Exam P Handouts – Page 206

Exercise 2 7

Loss amounts are uniform on the interval (0, 6), and insured with a deductible of 1.6.
Find the probabilities that:
a) the payment for a randomly chosen loss is 0
b) the payment for a randomly chosen loss is less than 2.

Let X denote the amount of a randomly chosen loss, and Y the corresponding
payment.
1.6
P[Y = 0] = P[X ≤ 1.6] =
6.0
P[Y ≤ 2] = P[X ≤ 2 + 1.6]
3.6
= P[X ≤ 3.6] =
6.0

206
B.1.2 Densities and CDFs Exam P Handouts – Page 207

Densities and CDFs 1

Definitions of CDFs and Densities


Percentiles
Exercises

Definitions: CDF 2

The cumulative distribution function (CDF) of X is given by

F (x) = FX (x) = P[X ≤ x]

This applies to all random variables, whether they have discrete, continuous, or mixed
distributions.

207
B.1.2 Densities and CDFs Exam P Handouts – Page 208

Definitions: Density 3

If X is purely continuous, then the density f (x) is given by

d
f (x) = F (x)
dx
Z x
The cdf is then F (x) = f (y )dy
−∞
X
For discrete X , F (x) = P[X ≤ x] = P[X = y ]
y ≤x

In most formulas, f (y ) dy will take the place of P[X = y ].

In some sense, f (y ) dy “ = ”P[y < X ≤ y + dy ]

Properties of the CDF 4

The cumulative distribution function satisfies


1. 0 ≤ F (x) ≤ 1
2. If x ≤ y then F (x) ≤ F (y )
3. lim F (x) = 1
x→∞
4. lim F (x) = 0
x→−∞

Since F (x) = P[X ≤ x], P[X > x] = 1 − F (x)

1 − F (x) is called the survival function

208
B.1.2 Densities and CDFs Exam P Handouts – Page 209

Properties of Densities 5

For continuous X , the density fX (x) satisfies:


1. f (x) ≥ 0
2. There need not be an upper bound for f (x)
Z ∞
3. f (x)dx = 1
−∞
Z b
4. f (x)dx = P[a < X ≤ b] = F (b) − F (a)
a

Example 6

Continuous Uniform
Suppose that X is uniform on (0, 0.1). What are F (x) and f (x)?



0x
 x <0
F (x) = = 10x 0 ≤ x ≤ 0.1

 0.1
1 0.1 < x


0 x <0
f (x) = 10 0 < x < 0.1


0 0.1 < x

209
B.1.2 Densities and CDFs Exam P Handouts – Page 210

Percentiles and Medians 7

x is a k-th percentile of X if F (x) = k%.

The median is the 50-th percentile, so F (x) = 0.5 at the median.

On the exam, continuous distributions will usually only have one point x such that
F (x) = k%. In that case the k-th percentile is uniquely defined and we say that x is
the k-th percentile.

Exercise 1 8

The CDF of X satisfies:




0 x <1
F (x) = (x − 1) − 14 (x − 1)2 1≤x <3


1 3≤x

Find P[X ≤ 2], P[1.5 < X ≤ 2] and f (x).

210
B.1.2 Densities and CDFs Exam P Handouts – Page 211

Exercise 1 8

The CDF of X satisfies:




0 x <1
F (x) = (x − 1) − 14 (x − 1)2 1≤x <3


1 3≤x
Find P[X ≤ 2], P[1.5 < X ≤ 2] and f (x).
1 1 3
P[X ≤ 2] = F (2) = (2 − 1) − (2 − 1)2 = 1 − =
4 4 4
P[1.5 < X ≤ 2] = F (2) − F (1.5)
3
= − [(1.5 − 1) − 0.25(0.5)2 ] = 0.3125
4


0 x <1
f (x) = F 0 (x) = 1 − 12 (x − 1) 1 < x < 3


0 3<x

Exercise 2 9

A continuous random variable Y has density f (y ) = 2/y 3 for 1 < y < ∞ and f (y ) = 0
otherwise. Find a formula for the CDF F (y ), and find P[Y ≤ 4 | Y > 2].

211
B.1.2 Densities and CDFs Exam P Handouts – Page 212

Exercise 2 9

A continuous random variable Y has density f (y ) = 2/y 3 for 1 < y < ∞ and f (y ) = 0
otherwise. Find a formula for the CDF F (y ), and find P[Y ≤ 4 | Y > 2].

Key point: The CDF is the definite integral, not any arbitrary anti-derivative. f (y ) = 0
for y < 1, so we can start our integral at 1. More precisely, for y > 1,
Z y Z 1 Z y
2
F (y ) = f (t) dt = 0 dt + 3
dt
−∞ −∞ 1 t
Z y
2 −1 y −1 −1 1
= 3
dt = = − = 1 −
1 t t2 1 y2 12 y2

P[2 < Y ≤ 4] F (4) − F (2)


P[Y ≤ 4 | Y > 2] = =
P[Y > 2] 1 − F (2)
(1 − (1/4) ) − (1 − (1/2)2 )
2 (1/22 ) − (1/42 ) 3
= 2
= 2
=
1 − (1 − (1/2) ) 1/2 4

212
B.1.3 Mixed Distributions Exam P Handouts – Page 213

Mixed distributions 1

Deductibles and/or Benefit Limits


A Possibility of No Loss
When F (x) isn’t continuous
Exercises

Deductibles and Benefit Limits 2

Suppose an insurance policy has a deductible of d and a payment limit of u. A


customer / insured has a loss of L.
• Insured is responsible for the first d of loss.
• Insurance company / insurer pays for portion of loss that exceeds d up to a total
payment u
• Insured is responsible for the rest.

If L < d Insurance payment = 0


If d ≤ L ≤ d + u Insurance payment = L − d
If d + u < L Insurance payment = u

Remarks:
• Payment always refers to payment made by insurer
• The premium is paid by insured to insurer
• Exam P study note only covers payment limits, other types are on later exams

213
B.1.3 Mixed Distributions Exam P Handouts – Page 214

Deductible example 3

Suppose that loss amounts X have density f (x) = 0.02x, 0 < x < 10. If there is a
deductible of 2 and a maximum payment of 6, then what is the probability of a
payment of 5 or less? What is the probability of a payment of 6?

P[Payment ≤ 5] = P[X ≤ 5 + 2]
Z 7
= 0.02x dx
0
= 0.01 · 72 = 0.49

P[Payment = 6] = P[X ≥ 6 + 2] = P[X ≥ 8]


Z 10
= 0.02x dx
8

= 0.01 · 102 − 82 = 0.36

A Possibility of No Loss 4

Losses, if they occur, are uniformly distributed on the interval (100, 500)
If there is a 60% probability of no loss and a 40% probability of exactly one loss, what
is the cdf of the total loss amount?

Let L = loss. P[L = 0] = 0.6

P[L ≤ 33] = P[L = 0] = 0.6. In fact, P[L ≤ `] = 0.6 for all ` < 100


 0 `<0


0.6 0 ≤ ` < 100
FL (`) =

 ?? 100 ≤ ` < 500


1 ` ≥ 500

214
B.1.3 Mixed Distributions Exam P Handouts – Page 215

A Possibility of No Loss 5

If 100 ≤ ` < 500,

P[L ≤ `] = P[no loss] + P[L ≤ `, have a loss]


= P[no loss] + P[loss] · P[L ≤ ` | have a loss]
FL (`) = 0.6 + 0.4 · P[L ≤ ` | have a loss]
` − 100
= 0.6 + 0.4 ·
500 − 100
` − 100
= 0.6 + 0.4 ·
400

0 100 500

A Possibility of No Loss 6

So we have found that




0 `<0



0.6 0 ≤ ` < 100
FL (`) = ` − 100

0.6 + 0.4 · 100 ≤ ` < 500

 400

1 500 ≤ `

F (l)
1

0.6
P[L = 0] = 0.6
`
100 500

215
B.1.3 Mixed Distributions Exam P Handouts – Page 216

When F (x) isn’t continuous 7

Z ∞
If X is purely continuous, then f (x) dx = 1, and F (x) is continuous.
−∞

If F (x) is defined piecewise, it may not be continuous and we may have a mixed
distribution.

Possible jumps are at the beginning and end of cases.

When F (x) isn’t continuous 8

Suppose X has CDF



0 x <0 F (x)



 1 1

 0≤x <
3 2 1
F (x) =

 x2 + 1 1

 ≤x <1 5/8

 2 2

 Jump = P[X = 1/2]
1 1≤x 1/3
1 5 Jump = P[X = 0]
F (0) = F (1/2) = x
3 8
1/2 1

There is no jump at 1 even though it is the endpoint of a case because formula for
F (x) leading up to 1 matches F (1).

216
B.1.3 Mixed Distributions Exam P Handouts – Page 217

When F (x) isn’t continuous 9

Jumps are points with non-zero probability:


F (x) 1
P[X = 0] = − 0
  3
1 1 5 1
P X = = −
2 8 3
5/8
Jump = P[X = 1/2]
1/3 for (1/2) < x < 1 (continuous piece):
Jump = P[X = 0] d
f (x) = F (x)
x dx
1/2 1 d x2 + 1
=
dx 2
=x

Exercise 1 10

An insurance policy pays for a random loss X subject to a deductible of d. The loss
amount is a continuous random variable with density function
(
2x for 0 < x < 1
f (x) =
0 otherwise

For a random loss X , the probability that the insurance payment is less then 0.3 is
equal to 0.49. Find d.

217
B.1.3 Mixed Distributions Exam P Handouts – Page 218

Exercise 1 10

An insurance policy pays for a random loss X subject to a deductible of d. The loss
amount is a continuous random variable with density function
(
2x for 0 < x < 1
f (x) =
0 otherwise

For a random loss X , the probability that the insurance payment is less then 0.3 is
equal to 0.49. Find d.

P[Payment ≤ 0.3] = P[Loss ≤ 0.3 + d]


Z 0.3+d Z 0.3+d
0.49 = f (x) dx = 2x dx
0 0
0.3+d
0.49 = x 2 = (0.3 + d)2
0
d = 0.7 − 0.3 = 0.4

Exercise 2 11

A random variable X has CDF




0 x <1

(x − 1)2
F (x) = 1≤x <3

 5

1 3≤x

Find P[X = 1], P[X = 3] and f (x) for 1 < x < 3.

218
B.1.3 Mixed Distributions Exam P Handouts – Page 219

Exercise 2 11

A random variable X has CDF




0 x <1

(x − 1)2
F (x) = 1≤x <3

 5

1 3≤x

Find P[X = 1], P[X = 3] and f (x) for 1 < x < 3.

(1 − 1)2
P[X = 1] = F (1) − lim F (x) = −0=0
x↑1 5
(3 − 1)2 1
P[X = 3] = F (3) − lim F (x) = 1 − =
x↑3 5 5
d d (x − 1)2 2(x − 1)
f (x) = F (x) = = for 1 < x < 3
dx dx 5 5

219
B.2.1 Moments of Continuous Distributions Exam P Handouts – Page 220

Moments of Continuous Distributions 1

Definition
Continuous Example
Exercises

Definition of Moments 2

In the purely discrete case, we had


X
E[X ] = x · P[X = x]
x

In the purely continuous case, this becomes


Z
E[X ] = x · f (x)dx
x

220
B.2.1 Moments of Continuous Distributions Exam P Handouts – Page 221

Definition of Moments 3

For other moments, the discrete formula


X
E[g (X )] = g (x) · P[X = x]
x

turns into
Z
E[g (X )] = g (x) · f (x)dx
x

In particular,

Var[X ] = E[X 2 ] − (E[X ])2


Z Z 
2
= x · f (x)dx − x · f (x)dx 2

Example 4

A random variable X has density 3x 2 for 0 < x < 1. Find its mean and variance.

Z 1
E[X ] = x · f (x) dx
0
Z 1 Z 1
2
= x · 3x dx = 3x 3 dx
0 0
1
3 4
= x
4 0
3
=
4

221
B.2.1 Moments of Continuous Distributions Exam P Handouts – Page 222

Example: Variance 5

For the variance,


Z 1
 2

E X = x 2 f (x) dx
0
Z 1 Z 1
2 2
= x · 3x dx = 3x 4 dx
0 0
3 1 3
= x5 =
5 0 5
 
Var[X ] = E X 2 − (E[X ])2
 2
3 3
= −
5 4
3
=
80

Exercise 1 6

If Y has density f (y ) = 1 − 0.5y for 0 < y < 2, and f (y ) = 0 otherwise, find E[Y ].

222
B.2.1 Moments of Continuous Distributions Exam P Handouts – Page 223

Exercise 1 6

If Y has density f (y ) = 1 − 0.5y for 0 < y < 2, and f (y ) = 0 otherwise, find E[Y ].

Z 2
E[Y ] = y f (y ) dy
0
Z 2
= y (1 − 0.5y ) dy
0
Z 2
= y − 0.5y 2 dy
0
2
y2
0.5y 3
= −
2 3 0
4 2
=2− =
3 3

Exercise 2 7

If Y has density f (y ) = 1 − 0.5y for 0 < y < 2, and f (y ) = 0 otherwise, find Var[Y ].

223
B.2.1 Moments of Continuous Distributions Exam P Handouts – Page 224

Exercise 2 7

If Y has density f (y ) = 1 − 0.5y for 0 < y < 2, and f (y ) = 0 otherwise, find Var[Y ].
2
E[Y ] = from before
3
Z 2 Z 2
2 2
E[Y ] = y f (y ) dy = y 2 (1 − 0.5y ) dy
0 0
Z 2
= y 2 − 0.5y 3 dy
0
2
y3
y4
= −
3 8 0
8 16 2
= − =
3 8 3
 2
2 2 2 2 2
Var[Y ] = E[Y ] − (E[Y ]) = − =
3 3 9

224
B.2.2 Moments of Mixed Distributions Exam P Handouts – Page 225

Moments of Mixed Distributions 1

Mixed Distributions Idea


Example
Exercises

Mixed distributions 2

X
For discrete variables, E[X ] = x · P[X = x]
x Z
For purely continuous variables, E[X ] = x · f (x) dx

For mixed distributions:


1. Use the discrete formula on discrete piece
2. Use the continuous formula on the continuous piece
3. Then sum the two parts

225
B.2.2 Moments of Mixed Distributions Exam P Handouts – Page 226

Example 3

Find the mean and variance of X if




0 for x < 1
 2
x − 2x + 2
F (x) = for 1 ≤ x < 2

 2

1 for x ≥ 2

F (x)

1 P[X = 1] = F (1) − lim F (x)


x↑1
1 1
= −0
2 1 2
2 1
x =
2
1 2

Example: Mean 4

For 1 < x < 2,


d d x 2 − 2x + 2
f (x) = F (x) = =x −1
dx dx 2
Z 2
E[X ] = 1 · P[X = 1] + x · (x − 1)dx
1
Z 2
1
=1· + (x 2 − x) dx
2 1
 3 2
1 x x2
= + −
2 3 2 1
   
1 8 4 1 1
= + − − −
2 3 2 3 2
4
=
3

226
B.2.2 Moments of Mixed Distributions Exam P Handouts – Page 227

Example: Variance 5

For the variance:


Z 2
2
 2
E X = 1 · P[X = 1] + x 2 · (x − 1) dx
1
Z 2
1
= 12 · + (x 3 − x 2 ) dx
2 1
 4 2
1 x x3
= + −
2 4 3 1
   
1 16 8 1 1
= + − − −
2 4 3 4 3
23
=
12
 2
23 4 5
Var(X ) = − =
12 3 36

Exercise 1 6

Suppose that X is a mixed random variable such that P[X = 3] = 0.5 and X has
density f (x) = x for 0 < x < 1, and 0 otherwise. Find E[X ].

227
B.2.2 Moments of Mixed Distributions Exam P Handouts – Page 228

Exercise 1 6

Suppose that X is a mixed random variable such that P[X = 3] = 0.5 and X has
density f (x) = x for 0 < x < 1, and 0 otherwise. Find E[X ].

Z 1
E[X ] = 3 · P[X = 3] + x f (x) dx
0
Z 1
= 3 · 0.5 + x · x dx
0
1
3 x3
= +
2 3 0
3 1
= +
2 3
11
=
6

Exercise 2 7

Suppose that X is a mixed random variable such that P[X = 3] = 0.5 and X has
density f (x) = x for 0 < x < 1, and 0 otherwise. Find E[X 2 ].

228
B.2.2 Moments of Mixed Distributions Exam P Handouts – Page 229

Exercise 2 7

Suppose that X is a mixed random variable such that P[X = 3] = 0.5 and X has
density f (x) = x for 0 < x < 1, and 0 otherwise. Find E[X 2 ].

Z 1
2 2
E[X ] = 3 · P[X = 3] + x 2 · f (x) dx
0
Z 1
9
= + x 3 dx
2 0
1
9 x4
= +
2 4 0
9 1
= +
2 4
19
=
4

229
B.2.3 The Survival Function Approach Exam P Handouts – Page 230

The Survival Function Approach 1

Using the survival function


Generalizations
Pros and Cons
Example Revisited
Other Examples
Exercises

Using the survival function 2

Suppose that P[X ≥ 0] = 1 and X is continuous. Then


Z ∞
E(X ) = x · f (x) dx
0

Using integration by parts,

u=x dv = f (x) dx
du = dx v = F (x)−1
Z ∞ Z ∞

E(X ) = x · f (x) dx = uv + (−v )du
0 0 0
at 0, uv = 0 · [F (0) − 1] = 0
at ∞, v = F (∞) − 1 = 1 − 1 = 0, uv = 0
Z ∞ Z ∞
E(X ) = 0 + [1 − F (x)] dx = P[X > x]dx
0 0

230
B.2.3 The Survival Function Approach Exam P Handouts – Page 231

Generalizations: Discrete Distributions 3

So for continuous, non-negative X ,


Z ∞
E[X ] = P[X > x] dx
0

This actually holds for all non-negative random variables, including discrete and mixed
distributions. For discrete distributions,
Z n+1
P[X > x]dx = P[X > n]
n
Z ∞ ∞
X
P[X > x]dx = P[X > n]
0 n=0

See A.3.2 for an intuitive interpretation of this.

Generalizations: E[g (x)] 4

For any non-negative variable (continuous, mixed, or discrete), if g (0) = 0 then


Z ∞
E[g (X )] = g 0 (x) P[X > x]dx
0

For example, if g (x) = x 2 then g 0 (x) = 2x and we get


Z ∞
 2
E X = 2x · P[X > x]dx
0
Z ∞
 3
E X = 3x 2 · P[X > x]dx
0

Unfortunately this is rarely useful, main application is with mixed distributions.

231
B.2.3 The Survival Function Approach Exam P Handouts – Page 232

Pros and Cons of Survival Method 5

Advantages of survival method:


• Often saves some steps, especially if F (x) is given but f (x) is not.
• Often faster for mixed distributions.
• Often gives nicer integrals (e.g., if f (x) = e −x , integrating x · f (x)dx requires
integration by parts, but the survival method does not.)

Disadvantages:
• Because the integral starts at 0, it sometimes will be messier.
• If f (x) is directly given, finding P[X > x] can require an extra step.

Example Revisited 6

Find the mean and variance of X if




0 for x < 1
 2
x − 2x + 2
F (x) = for 1 ≤ x < 2

 2

1 for x ≥ 2
Z ∞ Z ∞
E(X ) = P[X > x] dx = 1 − F (x) dx
0 0
Z 1 Z 2 Z ∞
2x − x 2
= (1 − 0) dx + dx + (1 − 1) dx
0 1 2 2
 2 2
x x3
=1+ − +0
2 6 1
   
4 8 1 1 4
=1+ − − − +0=
2 6 2 6 3

232
B.2.3 The Survival Function Approach Exam P Handouts – Page 233

Example: Variance 7

Z ∞ Z ∞
2
 d 2
E X = (x ) · P[X > x]dx = 2x · [1 − F (x)]dx
0 dx 0
Z 1 Z 2 Z ∞
2x − x 2
= 2x · 1 dx + 2x · dx + 0 dx
0 1 2 2
Z 1 Z 2
= 2x dx + (2x 2 − x 3 ) dx
0 1
1 3
 4
2
2x x
= x2 + −
0 3 4 1
   
2 · 8 16 2 1 23
= (1 − 0) + − − − =
3 4 3 4 12
 2
23 4 5
Var(X ) = − =
12 3 36

Example 8

3 · 1003
Suppose X has density f (x) = for 0 < x < ∞ and 0 otherwise. Find E[X ].
(x + 100)4
We have 3 possible approaches:
Z∞
x · 3 · 1003
1. Find dx using u-substitution with u = x + 100
(x + 100)4
0
Z∞
x · 3 · 1003
2. Find dx using integration by parts
(x + 100)4
0
3. Use the survival method
Remark: The last two approaches are mathematically identical. The survival method is
just automating the integration by parts.

233
B.2.3 The Survival Function Approach Exam P Handouts – Page 234

Survival Approach 9

3 · 1003
f (x) =
(x + 100)4
Z ∞
3 · 1003
P[X > x] = dt
x (t + 100)4
−1003 ∞ 1003
= =
(t + 100)3 x (x + 100)3
Z ∞
1003
E[X ] = dx
0 (x + 100)3

−1 1003
=
2 (x + 100)2 0
1003
= = 50
2 · 1002

Exercise 1 10

Use the survival approach to find E[X ] if the cdf of X is



 1003
1− 3 x > 100
F (x) = x

0 x ≤ 100

234
B.2.3 The Survival Function Approach Exam P Handouts – Page 235

Exercise 1 10

Use the survival approach to find E[X ] if the cdf of X is



 1003
1− 3 x > 100
F (x) = x

0 x ≤ 100

Z ∞ Z ∞
E[X ] = P[X > x] dx = [1 − F (x)] dx
0 0
Z 100 Z ∞  
1003
= (1 − 0)dx + 1− 1− 3 dx
0 100 x
Z ∞
1003
= 100 + 3
dx
100 x
−1 1003 ∞
= 100 + = 150
2 x 2 100

Exercise 2 11

Use the density to find E[X ] if the cdf of X is



 1003
1− 3 x > 100
F (x) = x

0 x ≤ 100

235
B.2.3 The Survival Function Approach Exam P Handouts – Page 236

Exercise 2 11

Use the density to find E[X ] if the cdf of X is



 1003
1− 3 x > 100
F (x) = x

0 x ≤ 100

3 · 1003
f (x) = , x > 100
x4
Z ∞ Z ∞
3 · 1003 3 · 1003
E[X ] = x· dx = dx
100 x4 100 x3
−1 3 · 1003 ∞
= · = 150
2 x2 100

236
B.3.1 Continuous Uniform - Basics Exam P Handouts – Page 237

Continuous Uniforms: Basics 1

Definitions
CDF and Density
Mean and Variance
Exercises

Continuous Uniforms 2

Pick a point X uniformly between 0 and 10

0 10
A

length of A
P[X ∈ A] =
total length
for 0 < x < 10 can find the CDF and thus density by
x
P[X ≤ x] = F (x) =
10
d 1 1
f (x) = F (x) = =
dx 10 total length

237
B.3.1 Continuous Uniform - Basics Exam P Handouts – Page 238

CDF and Density 3

More generally, if X is uniform on S,

length (or area) of A


P[X ∈ A] =
length (or area) of S
1
density =
length of S

If X is uniform on (a, b)

X
a b
1 x −a
f (x) = F (x) =
b−a b−a

Mean and variance of Uniform(0, 1) 4

If X ∼ uniform (0, 1)

1
f (x) = =1
1−0
Z 1 1
1 1
E[X ] = x · 1 dx = x 2 =
0 2 0 2
Z 1 1
 2 x3 1
E X = x 2 · 1 dx = =
0 3 0 3
2 2
Var[X ] = E[X ] − (E[X ])
1 1 1
= − =
3 4 12

238
B.3.1 Continuous Uniform - Basics Exam P Handouts – Page 239

General Mean 5

To go from a Uniform(0, 1) to a general uniform,

If X ∼ Uniform(a, b)
X − a ∼ Uniform(0, b − a)
X −a
∼ Uniform(0, 1)
 b − a

X −a 1
E =
b−a 2
1 1
(E[X ] − a) =
b−a 2
b−a b+a
E[X ] = +a=
2 2
i.e., mean = average of endpoints.

General variance 6

For the variance:


 
X −a 1
Var = Var[Uniform(0, 1)] =
b−a 12
 
X −a 1
Var = · Var[X − a]
b−a (b − a)2
1 1
= · Var[X ]
12 (b − a)2
(b − a)2
= Var[X ]
12
length2
= Var[X ]
12

239
B.3.1 Continuous Uniform - Basics Exam P Handouts – Page 240

Discrete vs Uniform 7

So for a continuous uniform on (a, b),

a+b
E[X ] = = Average of endpoints
2
(b − a)2 (Length of Interval)2
Var[X ] = =
12 12
while for a discrete uniform on a, a + 1, . . . , b
a+b
E[X ] = = Average of endpoints
2
(Number of Possible Values)2 − 1
Var[X ] =
12

Exercise 1 8

If N is uniform on {7, 8, 9, 10, 11, 12, 13}, find the mean and variance of N.

240
B.3.1 Continuous Uniform - Basics Exam P Handouts – Page 241

Exercise 1 8

If N is uniform on {7, 8, 9, 10, 11, 12, 13}, find the mean and variance of N.

The key point is that the set of possible values is explicitly listed out as finite, so we
have a discrete uniform.
7 + 13
E[N] = = 10
2
(# possible values)2 − 1
Var[N] =
12
2
7 −1
= =4
12

Exercise 2 9

If X is uniform on (7, 13), find the mean and variance of X .

241
B.3.1 Continuous Uniform - Basics Exam P Handouts – Page 242

Exercise 2 9

If X is uniform on (7, 13), find the mean and variance of X .

Now we are uniform on a continuous interval, so have a continuous uniform.


7 + 13
E[X ] = = 10
2
(13 − 7)2
Var[X ] = =3
12

242
B.3.2 Continuous Uniform - Exam Concepts Exam P Handouts – Page 243

Continuous Uniforms: Exam Style Questions 1

Mean and Variance Review


Moments of Mixtures
Examples
Exercises

Mean and Variance Review 2

Recall that for a continuous uniform on (a, b),

a+b
E[X ] = = Average of endpoints
2
(b − a)2 (Length of Interval)2
Var[X ] = =
12 12
while for a discrete uniform on a, a + 1, . . . , b
a+b
E[X ] = = Average of endpoints
2
(Number of Possible Values)2 − 1
Var[X ] =
12
But exam questions won’t just involve a simple uniform distribution and instead will
modify it somehow.

243
B.3.2 Continuous Uniform - Exam Concepts Exam P Handouts – Page 244

Moments with Mixtures 3

Key idea: Raw moments (eg. mean, 2nd moment) can be broken up into pieces.

Danger: Variance cannot be broken up into cases without an extra correction term.
X
E[X ] = x · P[X = x]
x

and from the law of total probability: if A1 , A2 . . . is a list of all possible cases
X
E[X ] = E[X | X ∈ Ai ] · P[X ∈ Ai ]
all Ai
X
2
E[X ] = E[X 2 | X ∈ Ai ] · P[X ∈ Ai ]
all Ai
X
E[g (X )] = E[g (X ) | X ∈ Ai ] · P[X ∈ Ai ]
all Ai

We will use this with cases based on deductibles.

Example 1 4

Losses X have a uniform distribution on [0, 100].

Losses are insured with a deductible. At what level must a deductible be set in order
for the expected payment to be 40% of what it would be with no deductible?

X ∼ U(0, 100)
d = deductible
Y = Payment after deductible
(
0 if X ≤ d
Y =
X − d if X > d
0 + 100
E[X ] = = 50
2
and we need d such that E[Y ] = 0.4 · 50 = 20

244
B.3.2 Continuous Uniform - Exam Concepts Exam P Handouts – Page 245

Example 1: Calculus Approach 5

The slower approach is to use calculus and set up the relevant integral for E[Y ]:

P[Y ≤ y ] = P[X ≤ y + d]
y +d
=
100
1
fY (y ) = for y > 0
100
Z
100−d
1
E[Y ] = 0 · P[Y = 0] + y· dy
100
0
100−d
y2 1 (100 − d)2
= · =
2 100 0 200

Example 1: Mixture Approach 6

A faster way to find E[Y ] is:

E[Y ] = E[Y | X ≤ d] · P[X ≤ d] + E[Y | X > d] · P[X > d]


if X ≤ d, Y = 0
if X > d, Y ∼ Uniform(d − d = 0, 100 − d)
100 − d
E[Y ] = 0 · P[X ≤ d] + · P[X > d]
2
100 − d 100 − d (100 − d)2
20 = · =
2 100 200
2
20 · 200 = (100 − d)

20 · 200 = 100 − d
p
d = 100 − 4,000 = 36.75

245
B.3.2 Continuous Uniform - Exam Concepts Exam P Handouts – Page 246

Example 2 7

A homeowner insures their home against storm damage with an insurance policy with a
deductible of 50 florins. In the event of storm damage, repair costs are modeled by a
uniform random variable on the interval (0, 300).

Find the standard deviation of the insurance payment in the event that the home
receives storm damage.

X = loss, Y = payment
if X ≤ 50, Y = 0
if X > 50, Y ∼ U(50 − 50, 300 − 50) = U(0, 250)

Example 2 Solution 8

0 + 250
E[Y ] = 0 · P[X ≤ 50] + · P[X > 50]
2
300 − 50 5
= 0 + 125 · = 125 · = 104.17
 2  2 300 6  
E Y = E Y | X ≤ 50 · P[X ≤ 50] + E Y 2 | X > 50 · P[X > 50]
5
= 0 + · E[U 2 , U ∼ Uniform(0, 250)]
6
for any random variable U, E[U 2 ] = Var[U] + (E[U])2
"  2 #
5 250 2 250
= · + = 17,361
6 12 2
Var[Y ] = 17,361 − 104.172 = 6,510
SD[Y ] = 80.7

246
B.3.2 Continuous Uniform - Exam Concepts Exam P Handouts – Page 247

Exercise 1 9

Loss amounts are uniform on (0, 20), and insured with a deductible of 3 and a payment
limit of 12. Find the expected payment amount on a randomly selected loss.

Exercise 1 9

Loss amounts are uniform on (0, 20), and insured with a deductible of 3 and a payment
limit of 12. Find the expected payment amount on a randomly selected loss.

Let X denote the loss and Y the payment.




0 X ≤3
Y = (X − 3) 3 < X ≤ 15


12 X > 3 + 12 = 15
E[Y ] = 0 · P[X < 3] + E[U(3 − 3, 15 − 3)] · P[3 < X ≤ 15] + 12 · P[X > 15]
0 + 12 15 − 3 20 − 15
=0+ · + 12 ·
2 20 20
= 0 + 6 · 0.6 + 12 · 0.25 = 6.6

247
B.3.2 Continuous Uniform - Exam Concepts Exam P Handouts – Page 248

Exercise 2 10

Loss amounts are uniform on (0, 20), and insured with a deductible of 3 and a payment
limit of 12. Find the variance of the payment for a randomly selected loss.

Exercise 2 10

Loss amounts are uniform on (0, 20), and insured with a deductible of 3 and a payment
limit of 12. Find the variance of the payment for a randomly selected loss.

Again let X denote the loss and Y the payment. We already know E[Y ] = 6.6, so
finding the 2nd moment will give us the variance.
E[Y 2 | 3 < X ≤ 15] = E[U 2 , U ∼ Unif.(0, 12)] = Var[U] + (E[U])2
 
(12 − 0)2 0 + 12 2
= + = 12 + 36 = 48
12 2
E[Y 2 ] = E[Y 2 | X ≤ 3] · P[X ≤ 3] + E[Y 2 | 3 < X ≤ 15] · P[3 < X ≤ 15]
+ E[Y 2 | 15 < X ] · P[15 < X ]
3 15 − 3 20 − 15
=0· + 48 · + 122 ·
20 20 20
= 0 + 48 · 0.6 + 144 · 0.25 = 64.8
Var[Y ] = E[Y 2 ] − (E[Y ])2 = 64.8 − 6.62 = 21.24

248
B.3.3 Exponential Random Variables Exam P Handouts – Page 249

Exponential random variables 1

Density and CDF


Mean and variance
Memoryless property
Example
Exercises

Density and CDF of exponentials 2

X is an exponential random variable with mean θ if

FX (x) = 1 − e −x/θ
1 − F (x) = e −x/θ

Sometimes will be called a rate λ = 1/θ exponential instead.


Exam questions may also give density rather than name.

FX (x) = 1 − e −x/θ = 1 − e −λ x
d 1
f (x) = F (x) = e −x/θ = λ e −λ x for x > 0
dx θ
Exponentials often are used to model waiting times (e.g., time between hits of a web
page, time between rain drops, etc.)

249
B.3.3 Exponential Random Variables Exam P Handouts – Page 250

Mean of an exponential 3

1 −x/θ
f (x) = e F (x) = 1 − e −x/θ for 0 < x < ∞
θ
P[X > x] = 1 − F (x) = e −x/θ = survival function
Z∞ Z∞
E[X ] = x · f (x) dx = P[X > x] dx
0 0
Z∞
E[X ] = e −x/θ dx
0

= −θ e −x/θ
0
−∞ −
= −θ e − θ e −0

Variance 4


Variance = E X 2 − [E(X )]2
Z∞ Z∞
1
E(X 2 ) = x 2 · f (x)dx = x 2 e −x/θ dx
θ
0 0

Using Tabular Integration


1 −x/θ
x2 + e
θ
2x − − e −x/θ
2 + θ e −x/θ
0 − θ2 e −x/θ
↑ ↑
differentiate left integrate right

250
B.3.3 Exponential Random Variables Exam P Handouts – Page 251

Variance 5

Plugging in:
      ∞
E(X 2 ) = x 2 − e −x/θ − 2x θe −x/θ + 2 − θ2 e −x/θ
0

= −0 − 0 − 0 − − 0 − 0 − 2θ2 e −0
= 2θ2

Var(X ) = 2 θ2 − θ2
= θ2

You won’t need to rederive this. Just remember Var(X ) = [E(X )]2 for exponentials

Variance of Geometrics 6

If X is exponential, then Var[X ] = (E[X ])2 .

We can use this to remember the variance formula for geometrics. Geometrics are the
discrete analog to exponentials, and

1 (1 − p)
If N ∼ Geo{1, 2, 3, . . . } then E[N] = , Var[N] =
p p2
1−p (1 − p)
If N ∼ Geo{0, 1, 2, . . . } then E[N] = , Var[N] =
p p2
For a geometric, instead of Variance = Mean2 we have
Var[Geo.] = E[Geo. starting at 0] · E[Geo. starting at 1]

251
B.3.3 Exponential Random Variables Exam P Handouts – Page 252

Memoryless property 7

Suppose that X is an exponential random variable with mean θ

P[X > x] = e −x/θ

What is P[X > x + a | X > a]?

P[X > x + a, X > a]


P[X > x + a | X > a] =
P[X > a]
e −(x+a)/θ
=
e −a/θ
= e (−x−a)/θ e a/θ = e −x/θ

i.e., P[X > x + a | X > a] = P[X > x]

Applications of the Memoryless Property 8

Since P[X > x + a | X > a] = P[X > x], we have P[X − a > x | X > a] = P[X > x].

In other words, given that X > a, X − a has the same distribution as the original
variable X . For example, if the time between buses is exponential with mean 15
minutes, the amount of time I need to wait (X − a) is an exponential with mean 15
minutes no matter how long it has been (a minutes) since the last bus.

For us, key application is payment amounts X − d with a deductible d conditioned on


a payment being made (i.e., given X > d) have same distribution as losses X .

E[X − a | X > a] = E[X ] = θ


E[X | X > a] = E[X − a | X > a] + a = θ + a
Var[X | X > a] = Var[X − a | X > a] = Var[X ] = θ2

252
B.3.3 Exponential Random Variables Exam P Handouts – Page 253

Example 9

Loss amounts are exponential with rate 0.02. If losses are insured with a deductible of
10, find the probability of a loss exceeding 40 given that a positive payment is made.

Let X denote the loss amount. X is exponential with mean θ = 1/λ = 1/0.02 = 50.

P[X > 40 | X > 10] = P[X − 10 > 30 | X > 10]


= P[X > 30] by memoryless property
= e −30/θ = e −3/5

Note: Exponentials are the only continuous distribution that follow the memoryless
property.

Exercise 1 10

Losses are exponential with mean 50, and are insured with a deductible of 10. Find the
median loss amount given that a positive payment is made.

253
B.3.3 Exponential Random Variables Exam P Handouts – Page 254

Exercise 1 10

Losses are exponential with mean 50, and are insured with a deductible of 10. Find the
median loss amount given that a positive payment is made.

We want P[X ≤ x | X > 10] = 0.5. We know (X − 10 | X > 10) ∼ Exponential(50)

P[X ≤ x | X > 10] = 1 − P[X > x | X > 10]


0.5 = 1 − P[X − 10 > x − 10 | X > 10]
1 − 0.5 = P[X > x − 10] = e −(x−10)/50
−(x − 10)
ln(0.5) =
50
x = −50 ln(0.5) + 10
= 44.66

Exercise 2 11

Losses have density f (x) = 0.1e −0.1x for x > 0, and 0 otherwise. If losses are insured
with a deductible of 3, find the expected payment for a randomly selected loss.

254
B.3.3 Exponential Random Variables Exam P Handouts – Page 255

Exercise 2 11

Losses have density f (x) = 0.1e −0.1x for x > 0, and 0 otherwise. If losses are insured
with a deductible of 3, find the expected payment for a randomly selected loss.

Let X denote our loss, and Y the payment. X is exponential with mean 1/0.1 = 10,
and

E[Y ] = E[Y | X ≤ 3] · P[X ≤ 3] + E[Y | X > 3] · P[X > 3]


 
−3/10
=0· 1−e + E[X − 3 | X > 3] · e −3/10
= 0 + 10 · e −3/10
= 7.41

255
B.3.4 Gamma, Exponential, and Poisson Exam P Handouts – Page 256

Gamma, Exponential & Poisson 1

Density and Link with Exponential


The Gamma CDF
Mean and Variance
Examples
Exercises
General α (Optional)
The Gamma CDF – Poisson Connection (Optional)

The Gamma Density 2

Suppose that X ∼ exp(θ). Then

1 −x/θ
f (x) = ·e for x > 0
θ
Z∞
the 1/θ in front is so f (x) dx = 1
0

To generalize this, we can add a factor of x α−1 in front of the exponential, so

f (x) = c x α−1 e −x/θ

where c is chosen to make this a density. The resulting distribution is called a


Gamma(α, θ) distribution.

When α is an integer, it turns out that a Gamma is the sum of α iid exponentials.

256
B.3.4 Gamma, Exponential, and Poisson Exam P Handouts – Page 257

The Gamma Density: Integer α 3

Key points to remember: α = 1 is an exponential. Lead term has factor of x α−1 .

Unlikely to need α = 3. Shouldn’t need α > 3. For x > 0,


1
α=1: f (x) = e −x/θ
θ
x −x/θ
α=2: f (x) = e
θ2

1 x 2 −x/θ
α=3: f (x) = · e
2 θ3
General form: for x > 0, α an integer:

1 x α−1 −x/θ
f (x) = · e
(α − 1)! θα

The Gamma CDF 4

If α = 1 then X is exponential, so CDF is F (x) = 1 − e −x/θ .

For other integer α, the CDF is still nice.

Again, unlikely to need α > 2 and would be shocking to need α > 3.


α = 1 : F (x) = 1 − e −x/θ
x
α = 2 : F (x) = 1 − e −x/θ − e −x/θ
θ
−x/θ x −x/θ  x 2 1 −x/θ
α = 3 : F (x) = 1 − e − e − · e
θ θ 2
α−1
X
P[X < x] = P[X ≤ x] = F (x) = 1 − P[Poisson(x/θ) = i]
i=0
α−1
X
P[X ≥ x] = P[X > x] = P[Poisson(x/θ) = i]
i=0

257
B.3.4 Gamma, Exponential, and Poisson Exam P Handouts – Page 258

Mean and Variance 5

Let X ∼ exp (θ) and Y ∼ Gamma (α, θ)

If α is an integer, then Y is sum of α iid exp(θ) variables.

E[Y ] = α · E[X ] = αθ
Var[Y ] = α · Var[X ] = αθ2

If α is an integer, these formulas are basic properties of sums.

If α is not an integer, it turns out that they still hold.

Example 1 6

If X is Gamma distributed with mean 10 and variance 50, find P[X > 10].

E[X ] = αθ = 10
Var[X ] = αθ2 = 50
Var [X ]
=θ=5
E[X ]
10
α= =2
5
10 −10/θ
P[X > 10] = e −10/θ + e
θ
= e −2 + 2e −2
= 0.406

258
B.3.4 Gamma, Exponential, and Poisson Exam P Handouts – Page 259

Example 2 7

A company has two electric generators. The time until failure for each generator
follows an exponential distribution with mean 10. The company will begin using the
second generator immediately after the first one fails.

What is the probability that the total time that the generators produce electricity is
less than 30 hours?

X1 ∼ exp(10), X2 ∼ exp(10)
Y = X1 + X2 ∼ Gamma(α = 2, θ = 10)
Want: P[Y ≤ 30]
   
30
= P Poisson ≥2
10
= 1 − e −3 − 3e −3
= 0.80

Exercise 1 8

A random variable X has density f (x) = 4xe −2x for 0 < x < ∞, and 0 otherwise. Find
the variance of X .

259
B.3.4 Gamma, Exponential, and Poisson Exam P Handouts – Page 260

Exercise 1 8

A random variable X has density f (x) = 4xe −2x for 0 < x < ∞, and 0 otherwise. Find
the variance of X .

The density is of the form (constant)·(variable to a power)·exp(- c· variable).

That makes it a Gamma. We get α from the power, θ from what is in the exponential.

Looking at the power: x = x 1 = x α−1


1=α−1
α=2
Looking inside exponent: − 2x = −x/θ
θ = 1/2
 2
1 1
Var[X ] = αθ2 = 2 · =
2 2

Exercise 2 9

An insured has 3 losses. If loss amounts are independent and exponentially distributed
with mean 5, find the probability that the sum of the 3 losses is no more than 11.2.

260
B.3.4 Gamma, Exponential, and Poisson Exam P Handouts – Page 261

Exercise 2 9

An insured has 3 losses. If loss amounts are independent and exponentially distributed
with mean 5, find the probability that the sum of the 3 losses is no more than 11.2.

The sum of independent exponentials is a Gamma.


α = 3 = # of variables in sum.
θ = 5 = mean of each exponential.

Total is Gamma(α = 3, θ = 5).


 2
−11.2/5 11.2 −11.2/5 1 11.2
F (11.2) = 1 − e − e − e −11.2/5 = 0.388
5 2 5

General α (Optional) 10

X is a gamma (α, θ) random variable if for x > 0 the density is

x α−1 −x/θ
1
e · for α an integer
θα
(α − 1)!
x α−1
1
· α e −x/θ for general α
Γ(α)
θ
Z ∞
where Γ(α) is the number such that f (x) dx = 1
0

Facts about Γ(α):


1. Γ(α) = (α − 1) · Γ(α − 1) if α − 1 > 0
2. Γ(α) = (α − 1)! for α a positive integer
√ 1√
3. Γ(1/2) = π, Γ(3/2) = (3/2 − 1) · Γ(1/2) = π, etc.
2

261
B.3.4 Gamma, Exponential, and Poisson Exam P Handouts – Page 262

Exponential Waiting Times 11

How else can we visualize the Poisson term in the CDF?

The following is not on the exam syllabus, but may be helpful for remembering the
equations.

X ∼ exp(θ) often used to model time between visits to a webpage

Exponential Waiting Times 12

Suppose N = # of visits by time t.


X1 = time of first visit
X1 + X2 = time of 2nd visit
X1 + · · · + Xn = time of nth visit
θ = average time between visits
t
Intuitively, E(N) =
θ
t 
In fact, N ∼ Poisson
θ
X1 X2 X3

0 X1 X1 + X2 X1 + X2 + X3 t

262
B.3.4 Gamma, Exponential, and Poisson Exam P Handouts – Page 263

First 2 Waiting Times 13


t  t
Claim: N ∼ Poisson , that is N has mean λ =
θ θ

P[N = 0] = P[No hits before time t]


= P[first hit after time t]
= P[X > t] = e −t/θ = e −λ

P[N = 1] = P[Exactly 1 visit by time t]


Integrating over possible cases x for when that visit occurred:
Zt Zt
1 −x/θ (−t+x)/θ
P[N = 1] = f (x) · P[X2 > t − x] dx = e e dx
θ
0 0
Zt
1 −t/θ t
= e dx = e −t/θ = λ e −λ
θ θ
0

General Result 14

So the claim holds for P[N = 0] and P[N = 1].

To complete the proof, you can proceed by induction. We have seen the base case
and integration by parts goes from n to n + 1.

Neither mathematical induction nor Poisson processes are on syllabus, so we will omit
the details.

263
B.3.5 Beta and Pareto Distributions Exam P Handouts – Page 264

Beta and Pareto Distributions 1

Syllabus Coverage
Beta
Pareto
Single Parameter Pareto
Exercises

Syllabus Coverage 2

Beta and Pareto distributions are explicitly listed on syllabus.

But recommended readings aren’t consistent on notation for Beta or even definition of
Pareto.

Historically, both have been tested purely by having density given to avoid ambiguities.
Presumably this will continue in future.

Point: All problems can be done from first principles, material in this lesson is optional

264
B.3.5 Beta and Pareto Distributions Exam P Handouts – Page 265

Beta 3

X is Beta(a, b) if f (x) = cx a−1 (1 − x)b−1 for 0 < x < 1, and 0 otherwise.

Don’t need to memorize c b/c density must be given.


(a + b − 1)!
Fact: c =
(a − 1)!(b − 1)!

Key idea: Most likely b = 2 or 3, can expand (1 − x)b−1 in doing integrals.

Moments (Optional):
a
E[X ] =
a+b
a(a + 1)
E[X 2 ] =
(a + b)(a + b + 1)
Easy to do from scratch, tested rarely, not worth memorizing anything.

Beta Example 4

Find E[X ] given f (x) = 6x(1 − x) for 0 < x < 1 and 0 otherwise.
Z 1
E[X ] = x · f (x) dx
0
Z 1
= x · 6x(1 − x) dx
0
Z 1
= 6x 2 − 6x 3 dx
0
6 6
= −
3 4
6 1
= =
12 2
Or a = b = 2 and E[X ] = a/(a + b) = 2/(2 + 2) = 1/2

265
B.3.5 Beta and Pareto Distributions Exam P Handouts – Page 266

Pareto Distributions 5

Suppose X
R∞≥ 0 but unbounded.
We need 0 f (x) dx = 1, which requires lim f (x) = 0.
x→∞

An exponential distribution does this with f (x) = λe −λx


But x −p also goes to 0. Can it be basis for a density?

Problem: x −p is horrible at 0.

2 solutions to avoid division by 0 at 0.


• Have f (x) ∝ (x + θ)−p for x > 0, and 0 otherwise
• Have f (x) ∝ x −p but require x > θ > 0
Some texts call first distribution Pareto, second single parameter Pareto. Others call
second distribution Pareto, don’t consider first. We will use Pareto / Single parameter
Pareto names.

Pareto Distribution 6

For α > 0, θ > 0, X is Pareto(α, θ) if f (x) = 0 for x < 0 and for x > 0,
αθα
f (x) =
(x + θ)α+1
Don’t need to memorize! Density or CDF must be given.
Have seen examples in previous lessons.

If α > 1 then
Z ∞ Z ∞
α
1 − F (x) = f (t) dt = θα ·α+1
dt
x x (t + θ)

α −1 θα
=θ · =
(t + θ)α x (x + θ)α
Z ∞ Z ∞
θα
E[X ] = [1 − F (x)] dx = dx
0 0 (x + θ)α

−θα θ
= =
(α − 1)(x + θ)α−1 0 α−1

266
B.3.5 Beta and Pareto Distributions Exam P Handouts – Page 267

Single Parameter Pareto 7

A single parameter Pareto avoids division by 0 by starting at θ > 0.

X is a single parameter Pareto(α, θ) if for x > θ,


αθα
f (x) = α+1
x
and f (x) = 0 for x < θ.

The simpler denominator makes calculations easier. E.g.,


Z ∞
αθα
E[X ] = x · α+1 dx
x
Zθ ∞ α
αθ
= dx
θ xα
−α θα ∞ αθ
= · α−1 =
α−1 x θ α−1

Exercise 1 8

If f (x) = 12x 2 (1 − x) for 0 < x < 1 and 0 otherwise, find Var[X ]

267
B.3.5 Beta and Pareto Distributions Exam P Handouts – Page 268

Exercise 1 8

If f (x) = 12x 2 (1 − x) for 0 < x < 1 and 0 otherwise, find Var[X ]


Z 1 Z 1
2
E[X ] = x · 12x (1 − x) dx = 12x 3 − 12x 4 dx
0 0
12 12 12 3
= − = =
4 5 20 5
Z 1 Z 1
2 2 2
E[X ] = x · 12x (1 − x) dx = 12x 4 − 12x 5 dx
0 0
12 12 12 2
= − = =
5 6 30 5
2 2
Var[X ] = E[X ] − (E[X ])
 2
2 3 1
= − =
5 5 25

Exercise 1: Alternate Solution 9

Alternatively, f (x) = 12x 2 (1 − x) is a Beta with x a−1 = x 2 so a − 1 = 2 and a = 3.


(1 − x)b−1 = (1 − x)1 so b = 2.
a
E[X ] =
a+b
3 3
= =
3+2 5
a(a + 1)
E[X 2 ] =
(a + b)(a + b + 1)
3·4 12 2
= = =
5·6 30 5
 2
2 3
Var[X ] = −
5 5
2 9 1
= − =
5 25 25

268
B.3.5 Beta and Pareto Distributions Exam P Handouts – Page 269

Exercise 2 10

2 · 1002
Y has density f (y ) = for 0 < y < ∞, and f (y ) = 0 otherwise. Find the
(y + 100)3
75th percentile of Y .

Exercise 2 10

2 · 1002
Y has density f (y ) = for 0 < y < ∞, and f (y ) = 0 otherwise. Find the
(y + 100)3
75th percentile of Y .

Let t be the 75th percentile. So F (t) = 0.75 and 1 − F (t) = 1 − 0.75 = 0.25.
Z ∞
1 − F (t) = f (y ) dy
t
Z ∞
2 · 1002
0.25 = 3
dy
t (y + 100)

−1002 1002
0.25 = =
(y + 100)2 t (t + 100)2
100
0.5 =
t + 100
0.5t + 50 = 100
t = 100

269
B.4.1 Normal Distribution Exam P Handouts – Page 270

Normal Distribution 1

Density (not tested)


The cdf (very important) Using tables
Linear Interpolation
SOA # 80
Sums of independent normals
SOA # 18

The normal density 2

f (x)

1 2
For a standard normal, f (x) = √ e −x /2

Z∞ makes std. deviation is 1
makes f (x) dx = 1
−∞

270
B.4.1 Normal Distribution Exam P Handouts – Page 271

Other Normals 3

A standard normal Z has mean 0 and variance 1.


What if we want Y to be normal with mean µ and variance σ 2 ?
We can construct such a Y by rescaling Z :

Z ∼ N(0, 1)
Y =σ·Z +µ
E[Y ] = σ · E[Z ] + µ = µ
Var[Y ] = Var[σZ ] = σ 2 · Var[Z ] = σ 2

It turns out that Y is still normal, so

Y = σZ + µ ∼ Normal(µ, σ 2 )

Other Normals 4

What is the density of Y ∼ N(µ, σ 2 )? (Optional)


1 2
fZ (z) = √ · e −z /2

Y = σ Z + µ y = g (z) = σ z + µ
y −µ
z = g −1 (y ) =
σ
If Y = g (Z )
 d −1
fY (y )= fZ g −1 (y ) g (y )
dy
" (   )#
1 1 y −µ 2 1
fY (y ) = √ exp − ·
2π 2 σ σ
1 2 /(2σ 2 )
=√ e −(y −µ)
2πσ 2

271
B.4.1 Normal Distribution Exam P Handouts – Page 272

The Normal cdf 5

f (z)

z
Suppose that Z is a standard normal, so

Z ∼ Normal(0, 1)
Φ(z) = P[Z ≤ z]

denotes the CDF of Z (so it is the shaded area). Fact: No elementary formula for
Φ(z) exists, so we will look it up on tables.

The Normal cdf 6

f (z)

area = area =
P[Z ≤ −z] P[Z > z]

z
−z z
Tables for Φ(z) only include z ≥ 0. To find the cdf for negative values, we compare
Φ(−z) with Φ(z).

P[Z > z] = P[Z ≤ −z]


1 − Φ(z) = Φ(−z)

272
B.4.1 Normal Distribution Exam P Handouts – Page 273

NORMAL DISTRIBUTION TABLE 7

Entries represent the area under the standardized normal distribution from −∞ to z,
i.e., Φ(z) = Pr(Z ≤ z) is the cdf. The value of z to the first decimal is given in the
left column. The second decimal place is given in the top row.

z 0.00 0.01 0.02 0.03 0.04 ···


0.0 0.5000 0.5040 0.5080 0.5120 0.5160
0.1 0.5398 0.5438 0.5478 0.5517 0.5557
0.2 0.5793 0.5832 0.5871 0.5910 0.5948
0.3 0.6179 0.6217 0.6255 0.6293 0.6331
.. ..
. .

Φ(0.12)= 0.5478
Φ(−0.33)= 1 − Φ(0.33)
= 1 − 0.6293

Probabilities Example 8

X is normal with mean µ = 2 and variance σ 2 = 9.


Find P[X > 3.86], P[X > 1.49] and the 95th percentile of X .

Idea: X = σZ + µ for a standard normal Z , so Z = (X − µ)/σ.


Will convert to Z and use tables.
 
X −µ 3.86 − 2
P[X > 3.86] = P > = P[Z > 0.62]
σ 3
= 1 − Φ(0.62)
= 1 − 0.7324 = 0.2676
 
X −µ 1.49 − 2
P[X > 1.49] = P > = P[Z > −0.17]
σ 3
= 1 − Φ(−0.17) = Φ(0.17)
= 0.5675

273
B.4.1 Normal Distribution Exam P Handouts – Page 274

Percentile Example 9

X is normal with mean µ = 2 and variance σ 2 = 9. Find the 95th percentile of X .

First, let’s find z, the 95th percentile of Z , then convert to t, the 95th percentile of X .

0.95 = P[Z ≤ z] = Φ(z)


z = Φ−1 (0.95)
z = 1.6449
X = σZ + µ
t = σz + µ = 3 · 1.6449 + 2
t = 6.935

Note the list of select inverse values at bottom of table.

Sums of Normals 10

If X and Y are independent normals, then it turns out that X + Y is also normal.

E(X + Y ) = E(X ) + E(Y )


Var(X + Y ) = Var(X ) + Var(Y )
X + Y ∼ N(µX + µY , σX2 + σY2 )

If we want averages,
 
X +Y E(X ) + E(Y ) 1 
∼ Normal , Var(X ) + Var(Y )
2 2 4

More generally, for any c, we have



c X ∼ Normal c E(X ), c 2 Var(X )

274
B.4.1 Normal Distribution Exam P Handouts – Page 275

Exercise 1 11

X is normal with mean −2.47 and variance 1.69. Find P[X > 0]

Exercise 1 11

X is normal with mean −2.47 and variance 1.69. Find P[X > 0]
 
X −µ 0 − (−2.47)
P[X > 0] = P > √
σ 1.69
= P[Z > 1.9]
= 1 − Φ(1.9)
= 1 − 0.9713
= 0.0287

To avoid sign errors, can do reasonability check. We want the probability that X is
much bigger than its mean, so should have a small value as the answer.

275
B.4.1 Normal Distribution Exam P Handouts – Page 276

Exercise 2 12

If X and Y are independent normal random variables with E[X ] = 0.6, E[Y ] = 3.2,
Var[X ] = 1.08 and Var[Y ] = 1.17, find the 75th percentile of the average of X and Y .

Exercise 2 12

If X and Y are independent normal random variables with E[X ] = 0.6, E[Y ] = 3.2,
Var[X ] = 1.08 and Var[Y ] = 1.17, find the 75th percentile of the average of X and Y .

Let W = (X + Y )/2 and t the 75th percentile of W . Also let Z be a standard


normal, and z the 75th percentile of Z .
E[X ] + E[Y ] 0.6 + 3.2
E[W ] = = = 1.9
 2  2
X +Y 1
Var[W ] = Var = Var[X + Y ]
2 4
1 2.25
Var[W ] = (1.08 + 1.17) =
4
p 4
SD[W ] = 2.25/4 = 0.75
z = 0.6745
t = SD[W ] · z + E[W ] = 0.75 · 0.6745 + 1.9 = 2.4

276
B.4.2 Interpolation Exam P Handouts – Page 277

Interpolation 1

Basic Idea
Linear Interpolation
Inverse Values
Exercises

Example 2

Let X ∼ N(µ = 4, σ 2 = 2.2). Find P[X > 3.3] to the nearest thousandth.
 
X −µ 3.3 − 4
P[X > 3.3] = P > √
σ 2.2
 
−0.7
=1−P Z ≤ √
2.2
= 1 − Φ(−0.472) = Φ(0.472)

In our tables, Φ(0.47) = 0.6808 and Φ(0.48) = 0.6844.

What is Φ(0.472)?

To nearest hundredth, it is 0.68.

To nearest thousandth, it is at least 0.681, at most 0.684, but can’t immediately tell.

277
B.4.2 Interpolation Exam P Handouts – Page 278

Exam Rules 3

Question wants Φ(0.472) to the nearest thousandth.

Know 0.6808 = Φ(0.47) < Φ(0.472) < Φ(0.48) = 0.6844

True value (via computer, not tables) Φ(0.472) = 0.6815 ≈ 0.682

Exam rule: Out of values that are possible based on tables, here 0.681, 0.682, 0.683
and 0.684, the only one included as an answer choice will be true value of 0.682.

Exam strategy options:


1. Accept that you may not match exam choices on normal distribution questions
2. Guesstimate. 0.472 is closer to 0.47 than 0.48. Φ(0.472) will be in the middle,
but closer to Φ(0.47), so probably 0.682
3. Use linear interpolation to get better approximation
You can use any of these approaches.

Linear Interpolation Example 4

From tables, Φ(0.47) = 0.6808, Φ(0.48) = 0.6844. Want Φ(0.472).

Φ(x)
0.6844

Φ(t)
0.6808
0.47 t 0.48 x
Comparing slopes:

Φ(t) − 0.6808 0.6844 − 0.6808



t − 0.47 0.48 − 0.47
0.472 − 0.47
Φ(0.472) ≈ · (0.6844 − 0.6808) + 0.6808
0.48 − 0.47
Φ(0.472) ≈ 0.2(0.6844 − 0.6808) + 0.6808 = 0.6815

278
B.4.2 Interpolation Exam P Handouts – Page 279

Linear Interpolation: Formulas 5

From tables, have Φ(a) and Φ(a + 0.01). Want Φ(a + t).

Φ(x)

Φ(a + 0.01)

Φ(a + t)
Φ(a)
x
a a+t a + 0.01

Φ(a + t) − Φ(a) Φ(a + 0.01) − Φ(a)


=
(t + a) − a 0.01
t
Φ(a + t) = Φ(a) + [Φ(a + 0.01) − Φ(a)]
0.01

Percentile Example 6

Still have X ∼ N(µ = 4, σ 2 = 2.2). Find the 13th percentile of X .


For a standard normal, the 13th percentile z is:
z = Φ−1 (0.13) = −Φ−1 (1 − 0.13) = −Φ−1 (0.87).
From tables, Φ(1.12) = 0.8686, Φ(1.13) = 0.8708, so −1.13 < z < −1.12.
Let t denote the 13th percentile of X

X = σZ + µ = 2.2Z + 4

t = 2.2z + 4
√ √
− 1.13 2.2 + 4 < t < −1.12 2.2 + 4
2.324 < t < 2.339

Φ(1.13) = 0.8708 is closer to 0.87 than Φ(1.12) = 0.8686, so exact value is closer to
2.324, which corresponds to z = 1.13, than 2.339. Exact value is around 2.33.

279
B.4.2 Interpolation Exam P Handouts – Page 280

Inverse Values: Linear Interpolation 7


We want t = 4 − 2.2 · Φ−1 (0.87)

y − 0.8686 0.8708 − 0.8686



Φ(x) Φ−1 (y ) − 1.12 1.13 − 1.12
0.8708 0.87 − 0.8686 0.8708 − 0.8686

Φ−1 (0.87) − 1.12 1.13 − 1.12
y
0.87 − 0.8686 Φ−1 (0.87) − 1.12
0.8686 =
0.8708 − 0.8686 1.13 − 1.12
1.12 1.13 14 −1
Φ (0.87) − 1.12
=
x 22 0.01
Φ−1 (y ) 14
Φ−1 (0.87) = 1.12 + · 0.01 = 1.1264
√ 22
t = 4 − 2.2 · 1.1264 = 2.329

Exercise 1 8

If X is normal with mean 0.6 and variance 1.3, find P[X ≤ 0.8]

280
B.4.2 Interpolation Exam P Handouts – Page 281

Exercise 1 8

If X is normal with mean 0.6 and variance 1.3, find P[X ≤ 0.8]
 
X −µ 0.8 − 0.6
P[X ≤ 0.8] = P ≤ √
σ 1.3
= P[Z ≤ 0.1754]
P[Z ≤ 0.17] ≈ 0.5675
P[Z ≤ 0.18] ≈ 0.5714
P[Z ≤ 0.1754] ≈ 0.5675 + 0.54(0.5714 − 0.5675)
= 0.5696

Exercise 2 9

Find the 60th percentile of X if E[X ] = 81.2 and Var[X ] = 33.8.

281
B.4.2 Interpolation Exam P Handouts – Page 282

Exercise 2 9

Find the 60th percentile of X if E[X ] = 81.2 and Var[X ] = 33.8.

Let z denote the 60th percentile of a standard normal, and t the 60th percentile of X .

z = Φ−1 (0.6)
Φ−1 (0.5987) = 0.25
Φ−1 (0.6026) = 0.26
0.6 − 0.5987
Φ−1 (0.6) = 0.25 + (0.26 − 0.25)
0.6026 − 0.5987
13
= 0.25 + · 0.01 = 0.25333
39
x = σz + µ

= 33.8 · z + 81.2 = 82.67

282
B.4.3 The Central Limit Theorem Exam P Handouts – Page 283

The Central Limit Theorem 1

Basic CLT
Normal Approximation
Example
Averages
Exercises

CLT: Central Limit Theorem 2

Key idea: Sums of a large number of random variables are often approximately normal
Theorem (CLT: Exam version)
If X1 , . . . , Xn are iid random variables, then

(X1 + · · · Xn ) − n E[X1 ]
p ⇒ N(0, 1)
n Var[X1 ]

Alternatively, let Sn = X1 + · · · Xn
E[Sn ] = nE[X1 ] = nµ
p √
SD[Sn ] = nVar[X1 ] = σ n
Sn − n µ
√ ⇒ N(0, 1)

283
B.4.3 The Central Limit Theorem Exam P Handouts – Page 284

Main Application: Normal Approximations 3

This lets us approximate the distribution of sums of random variables.

Let Sn = X1 + · · · + Xn . Finding the exact distribution of Sn is hard, but


 
Sn − n µ x −nµ
P[Sn ≤ x] = P √ ≤ √
σ n σ n
↑ ↑
Std normal “z-value”
 
x −nµ
≈Φ √
σ n

and we can find that on a normal table

Example 4

An insurance company pays claims on 625 losses. Losses are independent and
exponentially distributed with mean 3. Find the approximate probability that the total
payment is between 1800 and 2010.
Sn = X1 + · · · + X625 Xi ∼ Exponential(3)
E[Xi ] = 3 Var[Xi ] = 32 = 9
E[Sn ] = 625 · 3 = 1,875
Var[Sn ] = 625 · 9 = 5,625
SD[Sn ] = 75
 
1800 − 1875 Sn − E[Sn ] 2010 − 1875
P[1800 < Sn < 2010] = P < <
75 SD[Sn ] 75
= P[−1 < Z < 1.8]
= Φ(1.8) − Φ(−1) = Φ(1.8) − [1 − Φ(1)]
= 0.9641 − [1 − 0.8413] = 0.8054

284
B.4.3 The Central Limit Theorem Exam P Handouts – Page 285

Sums versus Products 5

Suppose X1 , X2 , . . . , X100 are iid random variables with P[Xi = 1] = P[Xi = −1] = 1/2.
Think of Xi as the amount that we win in bet number i.

100 X1 is the payoff if we bet $100 on the first bet.


It is either −100 or 100.
100
X 100
X
Xi is the net payoff after all 100 bets. We will win some and lose some, so Xi
i=1 i=1
will probably be close to 0. In particular, it is closer to 0 than 100X1 and therefore has
smaller variance.

Var(100 X1 ) = 1002 Var(X1 ) = 1002


P
Var( Xi ) = 100 Var(X ) = 100

Averages 6

Suppose that X1 , . . . , Xn are iid random variables, and let X be their average. What is
the distribution of X ?

S = X1 + X2 + · · · + Xn
S 1
X = = (X1 + X2 + · · · + Xn )
n n
1
E[X ] = (E(X1 ) + E(X2 ) + · · · + E(Xn ))
n
1
= · n · E[X ]
n
= E[X ]

285
B.4.3 The Central Limit Theorem Exam P Handouts – Page 286

Averages 7

 
1
Var[X ] = Var (X1 + · · · + Xn )
n
1
= 2 (Var(X1 ) + · · · + Var(Xn ))
n
1
= 2 · nVarX
n
Var(X )
=
n
Since S is roughly normal, so is S/n, so we have that
 
Var(X )
X ≈ N E(X ),
n

This comes up often enough that it is worth knowing.

Exercise 1 8

Losses have mean 4 and standard deviation 3. If losses are independent, use a normal
approximation to estimate the probability that the sum of 30 losses is at least 100.

286
B.4.3 The Central Limit Theorem Exam P Handouts – Page 287

Exercise 1 8

Losses have mean 4 and standard deviation 3. If losses are independent, use a normal
approximation to estimate the probability that the sum of 30 losses is at least 100.

Let Xi denote a loss amount, S the sum, and Z a standard normal.


E[S] = 30 · E[X ] = 30 · 4 = 120
Var[S] = 30 · Var[X ] = 30 · 32 = 270
√ √
SD[S] = SD[X ] 30 = 270 = 16.43
 
S − E[S] 100 − 120
P[S ≥ 100] = P ≥
SD[S] 16.43
≈ P[Z > −1.217] = 1 − Φ(−1.217) = Φ(1.217)
≈ Φ(1.21) + 0.7[Φ(1.22) − Φ(1.21)]
= 0.8869 + 0.7[0.8888 − 0.8869] = 0.8882
Could have said 0.89 without interpolation.

Exercise 2 9

Losses are independent, each with density 0.2e −0.2x for x > 0 and 0 otherwise. Losses
are insured with a deductible of 5. The first 60 randomly selected positive payments are
averaged. Using a normal approximation, estimate the 88th percentile of the average.

287
B.4.3 The Central Limit Theorem Exam P Handouts – Page 288

Exercise 2 9

Losses are independent, each with density 0.2e −0.2x for x > 0 and 0 otherwise. Losses
are insured with a deductible of 5. The first 60 randomly selected positive payments are
averaged. Using a normal approximation, estimate the 88th percentile of the average.

X is exponential with mean 1/0.2 = 5. By the memoryless property, the payments Yi


have the same distribution as the losses, so are also exponential with mean 5.
E[Y ] = E[Y ] = 5
Var[Y ] = Var[Y ]/60 = 52 /60 = 0.4167

SD[Y ] = 0.4167 = 0.6455
Φ(1.17) = 0.8790
Φ(1.18) = 0.8810
10
Φ−1 (0.88) = 1.17 + · 0.01 = 1.175
20
88th %-ile of Y = 0.6455 · 1.175 + 5 = 5.76

288
B.4.4 Continuity Correction Exam P Handouts – Page 289

Continuity Correction 1

Example
Formulas
Example
Exercises

Example: Binomial Version 2

Suppose X is binomial with n = 25 and p = 0.2. Find:


a) E[X ] and Var[X ]
b) Give an expression for the probability that X is at least 8
c) Given an expression for the probability that X is more than 8.

E[X ] = np = 25 · 0.2 = 5
Var[X ] = np(1 − p) = 25 · 0.2 · 0.8 = 4
X25 X25  
25
P[X ≥ 8] = P[X = k] = 0.2k (1 − 0.2)25−k
k
k=8 k=8
X25 X25  
25
P[X > 8] = P[X = k] = 0.2k (1 − 0.2)25−k
k
k=9 k=9

By computer, P[X ≥ 8] = 0.109 and P[X > 8] = 0.047

289
B.4.4 Continuity Correction Exam P Handouts – Page 290

Example: Normal Version 3

Suppose W is normal with mean 5 and variance 4. Find:


a) The probability that W is at least 8
b) The probability that W is more than 8

a) is asking for P[W ≥ 8] while b) wants P[W > 8].

But normals are continuous, so P[W = 8] = 0 and P[W ≥ 8] = P[W > 8].
 
W − E[W ] 8−5
P[W > 8] = P > √
SD[W ] 4
= 1 − Φ(1.5)
= 1 − 0.9332
= 0.0668

Example: Finding Approximation 4

In those examples, X and W had the same mean and variance, but
P[X ≥ 8] = 0.1091 and P[X > 8] = 0.0468
P[W ≥ 8] = 0.0668 = P[W > 8] = 0.0668.

Shouldn’t they be closer if we are using normals as approximations?

The problem is X must be an integer, while W does not.


W can be 7.7, 8.233, 8.6479, etc.

A better approximation to P[X ≥ 8] is P[W rounds to something ≥ 8].


I.e., P[X ≥ 8] ≈ P[W > 7.5]. P[X > 8] = P[W > 8.5], etc.

P[W > 7.5] = 1 − Φ[(7.5 − 5)/2] = 0.1056


P[W > 8.5] = 1 − Φ[(8.5 − 5)/2] = 0.0401

Still not perfect (b/c n is small) but much better.

290
B.4.4 Continuity Correction Exam P Handouts – Page 291

Continuity Correction 5

A continuity correction is used with normal approximations of integer valued variables.

It corrects for the normal being continuous while the variable we care about is not.

Used when approximating a sum Sn of discrete variables with a normal W .

P[Sn ∈ A] ≈ P[W rounds to something ∈ A].

1. P[Sn < x] ≈ P[W < x − 0.5]


2. P[Sn ≤ x] ≈ P[W < x + 0.5]
3. P[Sn ≥ x] ≈ P[W > x − 0.5]
4. P[Sn > x] ≈ P[W > x + 0.5]
Try to think these through rather than memorize.

Example 6

A roulette player betting on black has an 18/38 probability of winning on each spin of
the wheel. Using a normal approximation with a continuity correction, find the
approximate probability of betting on black and winning in at least 45 out of 100 spins.
 
18
Sn = # wins ∼ Binomial 100,
38
18
E[Sn ] = 100 · ≈ 47.368
r 38
18 20
SD[Sn ] = 100 · · ≈ 4.993
38 38
P[Sn ≥ 45] ≈ P[W > 44.5]
 
W − E[W ] 44.5 − 47.368
=P >
SD[W ] 4.993
= P[Z > −0.5745] = 1 − Φ(−0.5745)
= 1 − (1 − Φ(0.5745)) = Φ(0.5745) ≈ 0.72

291
B.4.4 Continuity Correction Exam P Handouts – Page 292

Exercise 1 7

Using a normal approximation with a continuity correction, estimate P[X ≤ 4] if X is


Poisson with mean 2.3.

Exercise 1 7

Using a normal approximation with a continuity correction, estimate P[X ≤ 4] if X is


Poisson with mean 2.3.

Let W be normal with the same mean and variance as X .

E[X ] = 2.3
p √
SD[X ] = Var[X ] = 2.3
P[X ≤ 4] ≈ P[W < 4.5]
 
4.5 − 2.3
=Φ √
2.3
= Φ(1.45)
= 0.9265

292
B.4.4 Continuity Correction Exam P Handouts – Page 293

Exercise 2 8

Using a normal approximation with a continuity correction, estimate P[X = 4] if X is


Poisson with mean 2.3. How does that compare to the true value?

Exercise 2 8

Using a normal approximation with a continuity correction, estimate P[X = 4] if X is


Poisson with mean 2.3. How does that compare to the true value?

Let W be normal with the same mean and variance as X .



E[X ] = 2.3 SD[X ] = 2.3
P[X = 4] ≈ P[3.5 < W < 4.5]
= P[W ≤ 4.5] − P[W ≤ 3.5]
   
4.5 − 2.3 3.5 − 2.3
=Φ √ −Φ √
2.3 2.3
= Φ(1.45) − Φ(0.79)
= 0.9265 − 0.7852 = 0.1413
−2.3 2.34
Actual P[X = 4] = e · = 0.1169
4!

293
B.4.5 Lognormal Distributions Exam P Handouts – Page 294

Lognormal Distributions 1

Example of lognormal
Definition of lognormal
Moments
Mean and Variance of lognormal
Exercises

Example of lognormal 2

Losses in a year have a lognormal distribution, Y = e X , where X is a normal random


variable with mean 3 and variance 0.5. What is the probability that losses exceed 80?

P[Y > 80] = P[e X > 80]


= P[X > ln 80] = P[X > 4.38]
 
X − E(X ) 4.38 − 3
=P > √
SD(X ) 0.5
↑ ↑
Std normal z-value
 
4.38 − 3
=1−Φ √
.5
= 1 − Φ(1.95)
= 1 − 0.97 = 0.03

294
B.4.5 Lognormal Distributions Exam P Handouts – Page 295

Lognormal Defintion 3

Definition (lognormal)
Y is a lognormal if Y = e X , X is normal
In words, the log of a lognormal distribution is normal.

Notes: if X is normal, then the range of X is −∞ to ∞


So e X is well-defined, but ln(X ) is not.

Lognormal Moments 4

Suppose that X is a Normal(µ, σ 2 ) random variable and Y = e X is the corresponding


lognormal. Then
2 )/2
E[Y ] = E[e X ] = e µ+(σ
E[Y ]6= e µ
h i
2 X 2
E[Y ] = E (e )
= E[e 2X ]
2
= e 2µ+2σ
More generally (optional formula, may help with memorizing)
2 /2
E[e tX ] = e tµ+(tσ)

Key point: For finding moments, we use these formulas about the lognormal. For
finding probabilities, we take logs and work with the underlying normal.

295
B.4.5 Lognormal Distributions Exam P Handouts – Page 296

Mean and Variance of Lognormal 5

Losses Y have a lognormal distribution with mean 50 and variance 400. What is the
probability that losses exceed 75?

Y = eX
E(Y ) = E(e X )
2
50 = e µ+σ /2
  
2 2 2X
E Y = 50 + 400 = E e
2
= e 2µ+2σ
σ2
ln(50) = µ +
 2
ln 50 + 400 = 2µ + 2σ 2
2

σ 2 = 0.148 µ = 3.84

Remark: To solve, double ln[E (Y )] equation and subtract from 2nd moment equation.

Mean and Variance of Lognormal 6

Losses Y have a lognormal distribution with mean 50 and variance 400. What is the
probability that losses exceed 75?

σ 2 = 0.148 µ = 3.84
P[Y > 75] = P[X > ln(75)]
 
X − 3.84 ln(75) − 3.84
=P >
0.1481/2 0.1481/2
= 1 − Φ(1.245)
≈ 1 − [0.8925 + 0.5(0.8944 − 0.8925)] ≈ 0.11

296
B.4.5 Lognormal Distributions Exam P Handouts – Page 297

Exercise 1 7

Suppose that X is normal with P[X > 5] = 0.5 and P[X > 8] = 0.05. Find E[e 2X ].

Exercise 1 7

Suppose that X is normal with P[X > 5] = 0.5 and P[X > 8] = 0.05. Find E[e 2X ].

µ=5
8−5
= 1.645
σ
σ = 1.8237
e X ∼ LN(µ = 5, σ 2 = 3.3259)
E[e 2X ] = E[(e X )2 ]
2
= e 2µ+2σ
= e 16.6518
= 17,052,000

297
B.4.5 Lognormal Distributions Exam P Handouts – Page 298

Exercise 2 8

X is lognormal with mean 10 and variance 200. Find P[X > 10].

Exercise 2 8

X is lognormal with mean 10 and variance 200. Find P[X > 10].
2 /2
E[X ] = 10 = e µ+σ
2
E[X 2 ] = 102 + 200 = e 2µ+2σ
2µ + σ 2 = 2 ln(10)
2µ + 2σ 2 = ln(300)
σ 2 = ln(300) − 2 ln(10) = 1.099
µ = 1.753
P[X > 10] = P[ln(X ) > ln(10)]
 
ln(X ) − µ ln(10) − 1.753
=P > √
σ 1.099
= 1 − Φ(0.52) = Φ(−0.52) = 0.3015
Warning: Lognormals are very sensitive to rounding. Carry as many digits as possible.

298
B.5.1 Deductibles: Calculus Approach Exam P Handouts – Page 299

Deductibles: Calculus Approach 1

Percentiles
Expected Values
Payment Limits

Losses Only 2

Loss amounts X have density f (x) = 0.02x for 0 < x < 10 and 0 otherwise. Find the
10th and 90th percentiles of X .

Z x
P[X ≤ x] = 0.02t dt
0
= 0.01x 2

Let s and t denote the 10th and 90th percentiles of X . Then

P[X ≤ s] = 0.10
0.01s 2 = 0.10
s = 3.16
0.01t 2 = 0.90
t = 9.49

299
B.5.1 Deductibles: Calculus Approach Exam P Handouts – Page 300

Deductible Example 3

Loss amounts have density f (x) = 0.02x for 0 < x < 10 and 0 otherwise. Losses are
insured subject to a deductible of 4. Let Y denote the payment amount corresponding
to a randomly selected loss. Find the 90th percentile of Y .

We want to find y such that P[Y ≤ y ] = 0.90.


Z y +4
P[Y ≤ y ] = P[X ≤ y + 4] = 0.02x dx
0
= 0.01(y + 4)2
0.90 = 0.01(y + 4)2
y = 5.49

Note that this equals the 90th percentile of X minus the deductible of 4.

Deductible Example 4

Loss amounts have density f (x) = 0.02x for 0 < x < 10 and 0 otherwise. Losses are
insured subject to a deductible of 4. Let Y denote the payment amount corresponding
to a randomly selected loss. Find the 10th percentile of Y .

Trying the approach from last time:

P[Y ≤ y ] = P[X ≤ y + 4] = 0.01(y + 4)2


0.10 = 0.01(y + 4)2
y = −0.83 which isn’t possible

This happened because P[Y = 0] = P[X ≤ 4] = 0.16 > 0.10, so the 16th percentile,
and every smaller percentile, of Y is 0.

The 10th percentile of X was 3.16. A loss of 3.16 results in a payment of 0, which is
the 10th percentile of Y . The p-th percentile of Y is the payment corresponding to
the p-th percentile of X .

300
B.5.1 Deductibles: Calculus Approach Exam P Handouts – Page 301

Conditional Payment Example 5

Loss amounts have density f (x) = 0.02x for 0 < x < 10 and 0 otherwise. Losses are
insured subject to a deductible of 4. Find the 90th percentile of a loss that exceeds the
deductible.

We are conditioning on X > 4, so


P[4 < X ≤ t]
0.9 = P[X ≤ t | X > 4] =
P[X > 4]
0.01 · t 2 − 0.01 · 42
=
1 − 0.01 · 42
t = 9.57
Or, using the survival function
P[X > t]
1 − 0.9 = P[X > t | X > 4] =
P[X > 4]
1 − 0.01 · t 2
0.1 = , t = 9.57
1 − 0.01 · 42

Deductible: Expected Payment 6

Loss amounts have density f (x) = 0.02x for 0 < x < 10 and 0 otherwise. Losses are
insured subject to a deductible of 4.

What is the expected payment amount?


Let Y denote the payment. So Y = X − 4 for X > 4, and Y = 0 for X ≤ 4.
Z 4 Z 10
E[Y ] = 0 · fX (x) dx + (x − 4) · fX (x) dx
0 4
Z 10
= 0 · P[X ≤ 4] + (x − 4)(0.02x) dx
4
Z 10
=0+ 0.02x 2 − 0.08x dx
4
0.02 3 10
= x − 0.04x 2 = 2.88
3 4

301
B.5.1 Deductibles: Calculus Approach Exam P Handouts – Page 302

Adding a Payment Limit 7

Loss amounts have density f (x) = 0.02x for 0 < x < 10, and f (x) = 0 otherwise.
Losses are insured subject to a deductible of 2 and a maximum payment limit of 5.
Find the cdf of the payment amount.

Let Y denote the payment. Then Y = 0 if X ≤ 2, so P[Y = 0] = P[X ≤ 2].


P[Y ≤ y ] = P[X ≤ y + 2] for 0 < y < 5.
P[Y ≤ 5] = 1 since 5 is the maximum possible payment.

 F (y )

0 y <0

 1
0.04 y =0 Jump = P[Y = 5] = 0.51
FY (y ) = 0.49


0.01(y + 2)2 0<y <5

1 0.04
y >5 Jump = P[Y = 0] = 0.04
y
5

Expected Payment with Limit: Densities 8

Loss amounts have density f (x) = 0.02x for 0 < x < 10, and f (x) = 0 otherwise.
Losses are insured subject to a deductible of 2 and a maximum payment limit of 5.
Find the expected payment amount.

Y has a mixed distribution: P[Y = 0] = 0.04 and P[Y = 5] = P[X > 7] = 0.51.
R5
E[Y ] = 0 · P[Y = 0] + 5 · P[Y = 5] + 0 y · fY (y )dy
So one approach would be to find fY (y ).

But sticking with X :


Z ∞
E[Y ] = payment · fX (x) dx
0
Z 2 Z 7 Z 10
= 0 · 0.02x dx + (x − 2) · 0.02x dx + 5 · 0.02x dx
0 2 7
0.02 3  
=0+ 7 − 23 − 0.02 72 − 22 + 5 · 0.51 = 3.8833
3

302
B.5.1 Deductibles: Calculus Approach Exam P Handouts – Page 303

Expected Payment with Limit: Survival Function 9

Loss amounts have density f (x) = 0.02x for 0 < x < 10, and f (x) = 0 otherwise.
Losses are insured subject to a deductible of 2 and a maximum payment limit of 5.
Find the expected payment amount.

Using the survival method, P[Y > y ] = P[X > y + 2] for 0 < y < 5. P[Y > 5] = 0.
Z ∞ Z 5 Z ∞
P[Y > y ] dy = P[X > y + 2] dy + 0 dy
0 0 5
Z 7
= P[X > x] dx
2
Z 7
= (1 − 0.01x 2 ) dx
2
0.01 · 73 0.01 · 23
=5− +
3 3
= 3.8833

303
B.5.2 Deductibles: Cases Approach Exam P Handouts – Page 304

Deductibles: Cases Approach 1

Overview
Uniform Examples
Exponential Examples

Splitting Into Cases 2

Let X be a loss. With a deductible of d, there are two cases for the payment Y :
(
0 X ≤d
Y =
X −d X >d
E[Y ] = E[Y | X ≤ d] · P[X ≤ d] + E[Y | X > d] · P[X > d]
= 0 · P[X ≤ d] + E[X − d | X > d] · P[X > d]

Sometimes E[X − d | X > d] can be found without calculus.


a−d
If X ∼ U(0, a), then (X − d | X > d) ∼ U(0, a − d) and E[X − d | X > d] = ;
2
if X ∼ exp(θ) then (X − d | X > d) ∼ exp(θ) and E[X − d | X > d] = θ.

304
B.5.2 Deductibles: Cases Approach Exam P Handouts – Page 305

Uniform Example 3

An industrial worker’s loss due to injury is uniformly distributed on (0, 100). An


insurance company will reimburse the worker for 60% of the amount of the loss that
exceeds a deductible of 20. What is the expected reimbursement for a randomly
selected loss?

There are two cases: either the loss exceeds 20, or it doesn’t. Let X denote the loss,
and Y the reimbursement.
E[Y ] = E[Y | X > 20] · P[X > 20] + E[Y | X ≤ 20] · P[X ≤ 20]
8 2
= E[0.6 · U(0, 80)] · +0·
10 10
= 0.6 · 40 · 0.8
= 19.2
Z 20 Z 100
1 1
or: E[Y ] = 0· dx + 0.6 · (x − 20) · dx
0 100 20 100

Adding Limits 4

Losses X are uniform on (0, 100). An insurance company will pay a reimbursement of
60% of the amount of the loss that exceeds a deductible of 20, up to a maximum
payment of 30. Find the expected reimbursement for a randomly selected loss.

Again let Y = Payment. We now have 3 cases:


no payment (X < 20), a payment of 0.6(X − 20), or a max payment of 30.
The max kicks in when 0.6(X − 20) = 30 or X = 70.
E[Y ] = 0 · P[X ≤ 20] + E[Y | 20 < X ≤ 70] · P[20 < X ≤ 70]
+ 30 · P[X > 70]
2 5 3
=0· + E[0.6 · U(0, 50)] · + 30 ·
10 10 10
= 0 + 0.6 · 25 · 0.5 + 9 = 16.5
Z 20 Z 70 Z 100
0 0.6(x − 20) 30
Or: E[Y ] = dx + dx + dx
0 100 20 100 70 100

305
B.5.2 Deductibles: Cases Approach Exam P Handouts – Page 306

Exponential Example 5

Losses X are exponentially distributed with mean 50. An insurance company will pay a
reimbursement of 60% of the amount of the loss that exceeds a deductible of 20. Find
the mean and variance of a randomly selected reimbursement Y .

E[Y ] = 0 · P[X ≤ 20] + E[Y | X > 20] · P[X > 20]


= 0 + 0.6 · E[X − 20 | X > 20] · P[X > 20]
= 0.6 · 50 · e −20/50 = 20.1
E[Y 2 ] = E[02 ] · P[X ≤ 20] + E[(0.6(X − 20))2 | X > 20] · P[X > 20]
= 0 + 0.62 E[(X − 20)2 | X > 20] · P[X > 20]
= 0.36 · (502 + 502 ) · e −20/50 = 1,206.6
Var[Y ] = 1,206.6 − 20.12 = 802.2

306
B.5.3 Review of Other Continuous Ideas Exam P Handouts – Page 307

Review of Other Continuous Ideas 1

Key Formulas
Distribution Review

Densities and CDFs 2

F (x) = P[X ≤ x]
F (∞) = 1 F (−∞) = 0

If X is continuous, then F (x) is continuous and


d
f (x) =F (x)
dx
Z b
P[a < X ≤ b] = f (x) dx = F (b) − F (a)
a

If X has a mixed distribution, then F (x) has some jumps. Jump sizes determine
probabilities, i.e.,

P[X = a] = F (a) − lim F (x)


x↑a

307
B.5.3 Review of Other Continuous Ideas Exam P Handouts – Page 308

Moments 3

If X is continuous, then
Z
E[X ] = x · f (x)dx
Z
E[g (X )] = g (x) · f (x)dx

For mixed distributions, sum the discrete and continuous pieces

If X is non-negative, and g (0) = 0, then


Z ∞
E[X ] = P[X > x]dx
Z0 ∞
E[g (X )] = g 0 (x) · P[X > x]dx
0

These last two equations hold for mixed distributions as well.

Normal Approximation 4

Suppose X1 , X2 , . . . , Xn are i.i.d. (independent, identically distrubted)

Let µ = E[Xi ] and σ 2 = Var[Xi ]

Let S = X1 + X2 + . . . Xn

The Central Limit Theorem says S is approximately normal.

S − E[S] S − n E[X ] S − nµ
= p = √ ≈ N(0, 1)
SD[S] n Var[X ] σ n

308
B.5.3 Review of Other Continuous Ideas Exam P Handouts – Page 309

Key Distributions 5

Uniform(a, b) Exponential Gamma,α = 2 Gamma, integer α


1 1 −x/θ x −x/θ x α−1
f (x) e e e −x/θ
b−a θ θ2 α
θ (α − 1)!
x −a x P
α−1 (x/θ)i −x/θ
F (x) 1− e −x/θ 1− e −x/θ − e −x/θ 1− e
b−a θ i=0 i!
a+b
E[X ] θ αθ αθ
2
(b − a)2
Var[X ] θ2 αθ2 αθ2
12

309
C.1.1 Joint Distributions and CDFs Exam P Handouts – Page 310

Joint Distributions and CDFs 1

Joint Distributions
Joint CDFs
Exercises

Joint Distribution 2

P
In 1-dimension, for a random variable X we had Pr[X = x] = 1
x

In 2 dimensions, suppose we have 2 random variables X and Y .


The joint probability mass function is Pr[X = x, Y = y ].
PP
Total probability is still 1 ⇒ Pr[X = x, Y = y ] = 1
x y

Note: Current syllabus only includes discrete joint distributions.


Continuous joint distributions no longer tested due to calculus requirements.

310
C.1.1 Joint Distributions and CDFs Exam P Handouts – Page 311

Example 3

The joint distribution of X and Y is given by

X
0 1 2
1 0.1 0.2 0.3
Y
2 0.1 0.1 0.2

P[X = 0, Y = 1] = 0.1
P[X = 2, Y = 2] = 0.2

Total probability needs to sum to 1:


X
P[X = x, Y = y ] = 0.1 + 0.2 + 0.3 + 0.1 + 0.1 + 0.2
x,y

=1

Uniform Example 4

Suppose X and Y are integer valued random variables with 1 ≤ X ≤ 4 and


1 ≤ Y ≤ X . If every possible outcome is equally likely, find P[X = Y = 3].

Possible outcomes are:


{X = 1, Y = 1}
{X = 2, Y = 1}, {X = 2, Y = 2}
{X = 3, Y = 1}, {X = 3, Y = 2}, {X = 3, Y = 3}
{X = 4, Y = 1}, {X = 4, Y = 2}, {X = 4, Y = 3}, {X = 4, Y = 4}

That is 10 possible outcomes, all with the same probability.

Total probability = 1 = 10 · P[X = Y = 3]


so P[X = Y = 3] = 0.1

311
C.1.1 Joint Distributions and CDFs Exam P Handouts – Page 312

Joint CDFs 5

In 1-d, the cdf of X is FX (x) = P[X ≤ x]

In 2-d, the joint cdf of X and Y is


FX ,Y (x, y ) = P[X ≤ x, Y ≤ y ]

Properties:
1. 0 ≤ F (x, y ) ≤ 1
2. F (x, ∞) = P[X ≤ x, Y ≤ ∞] = P[X ≤ x] = FX (x)
3. F (∞, y ) = P[X ≤ ∞, Y ≤ y ] = P[Y ≤ y ] = FY (y )
4. F (∞, ∞) = 1
5. F (−∞, y ) = 0 = F (x, −∞)

CDF Example 6

The joint distribution of X and Y is given by

X
0 1 2
1 0.1 0.2 0.3
Y
2 0.1 0.1 0.2

Some values of the CDF:

FX ,Y (0, 1) = P[X ≤ 0, Y ≤ 1] = 0.1


FX ,Y (1, 1) = P[X ≤ 1, Y ≤ 1]
= 0.1 + 0.2 = 0.3
FX ,Y (1, 2) = P[X ≤ 1, Y ≤ 2]
= 0.1 + 0.2 + 0.1 + 0.1 = 0.5

312
C.1.1 Joint Distributions and CDFs Exam P Handouts – Page 313

Exercise 1 7

P[X = x, Y = y ] = c(x + y ) for x and y integers such that 1 ≤ x ≤ 3 and 1 ≤ y ≤ x.


Find c.

Exercise 1 7

P[X = x, Y = y ] = c(x + y ) for x and y integers such that 1 ≤ x ≤ 3 and 1 ≤ y ≤ x.


Find c.

The point is that the total probability must equal 1.

If x = 1 then y = 1.
If x = 2 then y can be 1 or 2
If x = 3, then y can be 1, 2, or 3.
That gives 6 possible cases to sum, and

1 = c(1 + 1) + c(2 + 1) + c(2 + 2) + c(3 + 1) + c(3 + 2) + c(3 + 3)


 
1 = c 2 + 3 + 4 + 4 + 5 + 6 = 24c
1
c=
24

313
C.1.1 Joint Distributions and CDFs Exam P Handouts – Page 314

Exercise 2 8

X and Y are integer valued random variables. If you know that FX ,Y (3, 3) = 1,
FX ,Y (3, 2) = 0.7, FX ,Y (2, 3) = 0.6, FX ,Y (2, 2) = 0.4, FX ,Y (1, 3) = 0.3, and
FX ,Y (1, 2) = 0.2, find P[X = 2].

Exercise 2 8

X and Y are integer valued random variables. If you know that FX ,Y (3, 3) = 1,
FX ,Y (3, 2) = 0.7, FX ,Y (2, 3) = 0.6, FX ,Y (2, 2) = 0.4, FX ,Y (1, 3) = 0.3, and
FX ,Y (1, 2) = 0.2, find P[X = 2].

Since X is integer valued, P[X = 2] = P[X ≤ 2] − P[X ≤ 1] = FX (2) − FX (1).

FX (x) = FX ,Y (x, ∞) so we wish we could plug in y = ∞. But FX ,Y (3, 3) = 1, so


X ≤ 3 and Y ≤ 3. In particular, 3 is the largest Y can be, so FX (x) = FX ,Y (x, 3), and

P[X = 2] = FX (2) − FX (1)


= FX ,Y (2, 3) − FX ,Y (1, 3)
= 0.6 − 0.3 = 0.3

314
C.1.2 Marginal and Conditional Distributions Exam P Handouts – Page 315

Marginal and Conditional Distributions 1

Marginal Distributions
Discrete Example
Conditional Distributions
Key Points
Exercises

Marginal distributions 2

For discrete variables,


X
P [X = x] = P [X = x, Y = y ]
y

is marginal distribution of X .

In words, it’s the distribution of X without knowing Y .

315
C.1.2 Marginal and Conditional Distributions Exam P Handouts – Page 316

SOA/CAS 1999 Practice Exam # 13 3

Let X and Y be discrete random variables with joint probability function


2x + y
p(x, y ) =
12
for (x, y ) = (0, 1), (0, 2), (1, 2), (1, 3), and 0 otherwise. Determine the marginal
probability function of X .

X has two possible values, 0 or 1. Finding the marginal probability function means
finding P[X = 0] and P[X = 1].

SOA/CAS 1999 Practice Exam # 13 4

2x + y
p(x, y ) =
12
for (x, y ) = (0, 1), (0, 2), (1, 2), (1, 3), and 0 otherwise.

P[X = 0] = P[(X , Y ) = (0, 1)] + P[(X , Y ) = (0, 2)]


2·0+1 2·0+2 3 1
= + = =
12 12 12 4
P[X = 1] = P[(X , Y ) = (1, 2)] + P[(X , Y ) = (1, 3)]
2·1+2 2·1+3 9 3
= + = =
12 12 12 4
P[X = x] = 0 otherwise

To check our work, note that P[X = 0] + P[X = 1] = 1

316
C.1.2 Marginal and Conditional Distributions Exam P Handouts – Page 317

Variation on Discrete Example 5

Let X and Y be discrete random variables with joint probability function


2x + y
p(x, y ) =
12
for (x, y ) = (0, 1), (0, 2), (1, 2), (1, 3), and 0 otherwise. Find P[X = 0 | Y = 2]

P[X = 0, Y = 2]
P[X = 0 | Y = 2] =
P[Y = 2]
2·0+2 2
P[X = 0, Y = 2] = =
12 12
P[Y = 2] = P[X = 0, Y = 2] + P[X = 1, Y = 2]
2·0+2 2·1+2 6
= + =
12 12 12
2/12 1
P[X = 0 | Y = 2] = =
6/12 3

Conditional distributions 6

For discrete variables,


P[X = x, Y = y ]
P[X = x | Y = y ] =
P [Y = y ]
P[X = x, Y = y ]
=P
P[X = x, Y = y ]
x

Note that the role of the denominator


P is so that the conditional distribution is itself a
probability distribution, i.e., P[X = x | Y = y ] = 1.
x

317
C.1.2 Marginal and Conditional Distributions Exam P Handouts – Page 318

Independence 7

X and Y are independent if P[X = x | Y = y ] = P[X = x]


For discrete variables,
P[X = x, Y = y ]
P[X = x | Y = y ] =
P [Y = y ]
P[X = x, Y = y ] = P[Y = y ] · P[X = x | Y = y ]
= P[Y = y ] · P[X = x] if X and Y are independent

In particular, X and Y are independent if and only if the joint probability mass
function factors as a function of x times a function of y and the range of {X , Y } is a
discrete rectangle.

Example 8

The joint distribution of X and Y is:


P[Y = 1] = 0.1 + 0.2 + 0.3 = 0.6
X P[Y = 2] = 0.1 + 0.1 + 0.2 = 0.4
0 1 2
1 0.1 0.2 0.3
Y 0.1 1
2 0.1 0.1 0.2
P[X = 0 | Y = 1] = =
0.6 6
P[X = 0] = 0.1 + 0.1 = 0.2 0.2 2
P[X = 1 | Y = 1] = =
0.6 6
P[X = 1] = 0.2 + 0.1 = 0.3 0.3 3
P[X = 2 | Y = 1] = =
P[X = 2] = 0.3 + 0.2 = 0.5 0.6 6
P[Y = 1 | X = 2] = 0.6
P[Y = 2 | X = 2] = 0.4

318
C.1.2 Marginal and Conditional Distributions Exam P Handouts – Page 319

Key Points 9

The marginal distribution P[X = x] of X can only involve x. It cannot involve any
other variable.

The conditional distribution P[X = x | Y = y ] can involve x and y .

If X and Y are independent, P[X = x | Y = y ] = P[X = x]. In particular, the


conditional distribution of X will only involve x. Even the range cannot depend on y .

Exercise 1 10

Let X and Y be discrete random variables with joint probability function

x 2y
p(x, y ) =
23
for (x, y ) = (1, 1), (1, 2), (2, 2), (2, 3), and 0 otherwise. Determine the marginal
probability function of X .

319
C.1.2 Marginal and Conditional Distributions Exam P Handouts – Page 320

Exercise 1 10

Let X and Y be discrete random variables with joint probability function

x 2y
p(x, y ) =
23
for (x, y ) = (1, 1), (1, 2), (2, 2), (2, 3), and 0 otherwise. Determine the marginal
probability function of X .

P[X = 1] = P[X = 1, Y = 1] + P[X = 1, Y = 2]


1 3
= (12 · 1 + 12 · 2) =
23 23
P[X = 2] = P[X = 2, Y = 2] + P[X = 2, Y = 3]
1 20
= (22 · 2 + 22 · 3) =
23 23
To check, note that P[X = 1] + P[X = 2] = (3/23) + (20/23) = 1

Exercise 2 11

Let X and Y be discrete random variables with joint probability function

x 2y
p(x, y ) =
23
for (x, y ) = (1, 1), (1, 2), (2, 2), (2, 3), and 0 otherwise. Find P[X = 1 | Y = 2]

320
C.1.2 Marginal and Conditional Distributions Exam P Handouts – Page 321

Exercise 2 11

Let X and Y be discrete random variables with joint probability function


x 2y
p(x, y ) =
23
for (x, y ) = (1, 1), (1, 2), (2, 2), (2, 3), and 0 otherwise. Find P[X = 1 | Y = 2]

P[X = 1, Y = 2]
P[X = 1 | Y = 2] =
P[Y = 2]
P[X = 1, Y = 2]
=
P[X = 1, Y = 2] + P[X = 2, Y = 2]
1 2
23 (1 · 2)
= 1 2 2
23 (1 · 2 + 2 · 2)
2 1
= =
2+8 5

321
C.2.1 Joint Moments Exam P Handouts – Page 322

Joint Moments 1

Examples
Definitions
Exercises

Discrete Example 2

Let X denote that number of years I will use my current computer, and Y the number
of years that I will use my tablet. The joint distribution of X and Y is
X
1 2 3
1 0.05 0.16 0.19
Y 2 0.10 0.13 0.12
3 0.07 0.10 0.08
What is the average number of years that I will use the tablet?
One approach is to first find the marginal distribution of Y .

P[Y = 1] = 0.05 + 0.16 + 0.19 = 0.40


P[Y = 2] = 0.10 + 0.13 + 0.12 = 0.35
P[Y = 3] = 0.07 + 0.10 + 0.08 = 0.25
E[Y ] = 1 · 0.40 + 2 · 0.35 + 3 · 0.25 = 1.85

322
C.2.1 Joint Moments Exam P Handouts – Page 323

Discrete Example 3

The joint distribution of X and Y is


X
1 2 3
1 0.05 0.16 0.19
Y 2 0.10 0.13 0.12
3 0.07 0.10 0.08
What is E(Y )?

A second approach is to sum over all 9 cases.

E[Y ] = 1 · 0.05 + 1 · 0.16 + 1 · 0.19


+ 2 · 0.10 + 2 · 0.13 + 2 · 0.12
+ 3 · 0.07 + 3 · 0.10 + 3 · 0.08
= 1.85

Discrete Example 4

The joint distribution of X and Y is


X
1 2 3
1 0.05 0.16 0.19
Y 2 0.10 0.13 0.12
3 0.07 0.10 0.08
What is E(X + Y )?

When we are dealing with X + Y , the first approach no longer works (easily), but we
can still use the second.

E[X + Y ] = (1 + 1) · 0.05 + (2 + 1) · 0.16 + (3 + 1) · 0.19


+ (1 + 2) · 0.10 + (2 + 2) · 0.13 + (3 + 2) · 0.12
+ (1 + 3) · 0.07 + (2 + 3) · 0.10 + (3 + 3) · 0.08
= 4.02

323
C.2.1 Joint Moments Exam P Handouts – Page 324

Definitions 5

For discrete random variables X and Y ,


XX
E[g (X , Y )] = g (x, y ) · P[X = x, Y = y ]
x y

Some special cases


XX
E[X ] = x · P[X = x, Y = y ]
x y
XX
E[Y 2 ] = y 2 · P[X = x, Y = y ]
x y

Exercise 1 6

Let X and Y be discrete random variables with joint probability function


2x + y
p(x, y ) =
12
for (x, y ) = (0, 1), (0, 2), (1, 2), (1, 3), and 0 otherwise. Find E[X ].

324
C.2.1 Joint Moments Exam P Handouts – Page 325

Exercise 1 6

Let X and Y be discrete random variables with joint probability function


2x + y
p(x, y ) =
12
for (x, y ) = (0, 1), (0, 2), (1, 2), (1, 3), and 0 otherwise. Find E[X ].
E[X ] = 0 · P[X = 0, Y = 1] + 0 · P[X = 0, Y = 2]
+ 1 · P[X = 1, Y = 2] + 1 · P[X = 1, Y = 3]
2·1+2 2·1+3
=0+0+ +
12 12
9 3
= =
12 4
Or we could say:
1 3
E[X ] = 0 · P[X = 0] + 1 · P[X = 1] = 0 ·
+1·
4 4
with P[X = 0] and P[X = 1] coming from an earlier lesson.

Exercise 2 7

Let X and Y be discrete random variables with joint probability function


2x + y
p(x, y ) =
12
for (x, y ) = (0, 1), (0, 2), (1, 2), (1, 3), and 0 otherwise. Find E[(X + 1)Y ].

325
C.2.1 Joint Moments Exam P Handouts – Page 326

Exercise 2 7

Let X and Y be discrete random variables with joint probability function


2x + y
p(x, y ) =
12
for (x, y ) = (0, 1), (0, 2), (1, 2), (1, 3), and 0 otherwise. Find E[(X + 1)Y ].

Key difference: now we want the mean of something that involves both X and Y . We
must use the joint probability mass.

E[(X + 1)Y ] = (0 + 1) · 1 · P[X = 0, Y = 1] + (0 + 1) · 2 · P[X = 0, Y = 2]


+ 2 · 2 · P[X = 1, Y = 2] + 2 · 3 · P[X = 1, Y = 3]
2·0+1 2·0+2 2·1+2 2·1+3
=1· +2· +4· +6·
12 12 12 12
1 + 4 + 16 + 30 51 17
= = =
12 12 4

326
C.2.2 Covariances and Correlations Exam P Handouts – Page 327

Covariances and Correlations 1

Definitions
Example
Using independence
SOA # 75; S.03.15
Exercises

Covariance and correlation 2

h i  
Var[X ] = E (X − µX ) = E X 2 − (E[X ])2
2

Cov(X , Y ) = E [(X − µX ) (Y − µY )]
= E(XY ) − E(X ) · E(Y )
Note: Cov(X , b) = E[bX ] − bE[X ] = 0

Covariance is “bilinear,” i.e.,

Cov (aX + bY , cZ ) = Cov (aX , cZ ) + Cov (bY , cZ )


= ac Cov (X , Z ) + bc Cov (Y , Z )

327
C.2.2 Covariances and Correlations Exam P Handouts – Page 328

Variance of Sums 3

Cov(X , X ) = E(X · X ) − E(X ) · E(X ) = Var(X )


Var(X + Y ) = Cov (X + Y , X + Y )
= Cov (X , X ) + 2Cov (X , Y ) + Cov (Y , Y )
= Var(X ) + 2Cov (X , Y ) + Var(Y )
Likewise Var(aX + bY ) = a2 Var(X ) + 2ab Cov(X , Y ) + b 2 Var(Y )
Key Example Var(X − Y ) = Var(X + (−1)Y )
= Var(X ) − 2Cov(X , Y ) + Var(Y )
If X and Y are independent then E(XY ) = E(X )E(Y ) so Cov(X , Y ) = 0. That means
that if X and Y are independent,
Var(X + Y ) = Var(X ) + Var(Y )
Var(X − Y ) = Var(X ) + Var(Y )

Correlation 4

Definition (Correlation)
The correlation of X and Y is given by

Cov (X , Y )
Corr (X , Y ) =
SD(X ) · SD(Y )

Note that if X and Y are independent then Corr(X , Y ) = 0 since Cov(X , Y ) = 0.

328
C.2.2 Covariances and Correlations Exam P Handouts – Page 329

Properties of Correlation 5

Suppose that Y = a X + b. Then

Cov(X , Y ) = Cov(X , a X + b)
= a · Cov(X , X ) + Cov(X , b)
= a · Var[X ]

So if a > 0, then Cov(X , Y ) > 0, and if a < 0, Cov(X , Y ) < 0

SD(Y ) = |a| SD(X )


(
a · VarX 1 a>0
so Corr(X , Y ) = 2
=
[SD(X )] |a| −1 a<0

Fact : −1 ≤ Corr (X , Y ) ≤ 1

Example 6

X and Y are integer valued random variables with 0 < X ≤ 3 and 1 ≤ Y ≤ 2.


P[X = x, Y = y ] = kxy when positive, for some constant k. Find Cov[X , Y ].

1 = k[1 · 1 + 1 · 2 + 2 · 1 + 2 · 2 + 3 · 1 + 3 · 2] = 18k
k = 1/18
1 · (1 · 1 + 1 · 2) + 2 · (2 · 1 + 2 · 2) + 3 · (3 · 1 + 3 · 2) 42 7
E[X ] = = =
18 18 3
1 · (1 · 1 + 2 · 1 + 3 · 1) + 2 · (1 · 2 + 2 · 2 + 3 · 2) 30 5
E[Y ] = = =
18 18 3
(1 · 1) · (1 · 1) + (1 · 2)2 + (2 · 1)2 + (2 · 2)2 + (3 · 1)2 + (3 · 2)2 70 35
E[XY ] = = =
18 18 9
35 7 5
Cov[X , Y ] = E[XY ] − E[X ]E[Y ] = − · =0
9 3 3

329
C.2.2 Covariances and Correlations Exam P Handouts – Page 330

When are random variables independent? 7

Theorem
X and Y are independent if both
1. The support of (X , Y ), i.e., the points such that P[X = x, Y = y ] > 0, is a
rectangular lattice
2. P[X = x, Y = y ] = P[X = x] · P[Y = y ]

Remark: The Soviet notation for independence was X ⊥ Y , partly because of this
requirement that the support be a rectangle.

Applications of Independence: 8

If X and Y are independent, then E[XY ] = (E[X ]) · (E[Y ]).

More generally,

E[g (X ) h(Y )] = E[g (X )] · E[(h(Y )]


   
so E X 2 Y 3 = E[X 2 ] E[Y 3 ] , etc.

Cov(X , Y ) = E[XY ] − (E[X ]) · (E[Y ])


= (E[X ]) · (E[Y ]) − (E[X ]) · (E[Y ]) = 0

Note: Independence implies that Cov(X , Y ) = 0, but the converse is false.

330
C.2.2 Covariances and Correlations Exam P Handouts – Page 331

Example of Independence 9

X and Y are integer valued random variables with 0 < X ≤ 3 and 1 ≤ Y ≤ 2.


P[X = x, Y = y ] = kxy when positive, for some constant k. Find Cov[X , Y ].

We could have simply observed that the support is a rectangle, and f (x, y ) factors as a
function of x times a function of y , so X and Y are independent and therefore
Cov(X , Y ) = 0.

SOA # 75; S.03.15 10

An insurance policy pays a total medical benefit consisting of two parts for each claim.
Let X represent the part of the benefit that is paid to the surgeon, and let Y represent
the part that is paid to the hospital. The variance of X is 5,000, the variance of Y is
10,000, and the variance of the total benefit, X + Y , is 17,000.

Due to increasing medical costs, the company that issues the policy decides to increase
X by a flat amount of 100 per claim and to increase Y by 10% per claim. Calculate
the variance of the total benefit after these revisions have been made.

Total benefit = X + 100 + 1.1Y .


Var(X + 100 + 1.1 Y ) = Var(X + 1.1 Y )
= Cov(X + 1.1 Y , X + 1.1 Y )
= Var(X ) + 2 · 1.1Cov(X , Y ) + 1.12 Var(Y )
= 5,000 + 2.2 Cov(X , Y ) + 1.21 · 10,000

331
C.2.2 Covariances and Correlations Exam P Handouts – Page 332

SOA # 75; S.03.15 11

So we just need to find Cov(X , Y ) and plug it in.

Var(X + Y ) = Var(X ) + 2 Cov(X , Y ) + Var(Y )


17,000 = 5,000 + 2 Cov (X , Y ) + 10,000
1,000 = Cov(X , Y )

Var(X + 1.1Y ) = 5,000 + 2.2 Cov(X , Y ) + (1.21) · 10,000


= 5,000 + 2,200 + 12,100
= 19,300

Exercise 1 12

Suppose P[X = Y = 0] = 1/2 and P[X = 1, Y = 1] = P[X = −1, Y = 1] = 1/4.


Find Cov[X , Y ]. Are X and Y independent?

332
C.2.2 Covariances and Correlations Exam P Handouts – Page 333

Exercise 1 12

Suppose P[X = Y = 0] = 1/2 and P[X = 1, Y = 1] = P[X = −1, Y = 1] = 1/4.


Find Cov[X , Y ]. Are X and Y independent?

Finding covariance first,

E[X ] = 0 · 1/2 + 1 · 1/4 + (−1) · 1/4 = 0


E[Y ] = 0 · 1/2 + 1 · (1/4 + 1/4) = 1/2
E[XY ] = 0 · 0 · 1/2 + 1 · 1 · 1/4 + (−1) · 1 · 1/4 = 0
Cov[X , Y ] = E[XY ] − E[X ]E[Y ]
= 0 − 0 · (1/2) = 0
Not independent because support isn’t rectangular. Or
P[X = 0] = 1/2 = P[Y = 0]
P[X = Y = 0] = 1/2 6= P[X = 0] · P[Y = 0]

Exercise 2 13

X and Y are random variables with Corr(X , Y ) = 0.6, Var[X ] = 64 and Var[Y ] = 100.
Find Var[2X − 3Y ].

333
C.2.2 Covariances and Correlations Exam P Handouts – Page 334

Exercise 2 13

X and Y are random variables with Corr(X , Y ) = 0.6, Var[X ] = 64 and Var[Y ] = 100.
Find Var[2X − 3Y ].

Cov[X , Y ]
Corr(X , Y ) =
SD[X ]SD[Y ]
Cov[X , Y ]
0.6 = √ √
64 100
48 = Cov[X , Y ]
Var[2X − 3Y ] = 22 Var[X ] + 2(2)(−3)Cov[X , Y ] + (−3)2 Var[Y ]
= 4 · 64 − 12 · 48 + 9 · 100
= 580

334
C.2.3 Conditional Moments Exam P Handouts – Page 335

Conditional Moments 1

Example
Double Expectation Theorem
Examples
Law of Total Variation
Exercises

Example 2

X
0 1 2
1 0.1 0.2 0.3
Y
2 0.1 0.1 0.2
Previously we saw
1 2 3
P[X = 0 | Y = 1] = P[X = 1 | Y = 1] = P[X = 2 | Y = 1] =
6 6 6

Conditional Moments:
X
E[X | Y = y ] = x · P[X = x | Y = y ]
all x
1 2 3 4
E[X | Y = 1] = 0 · +1· +2· =
6 6 6 3
0.1 0.1 0.2 5
E[X | Y = 2] = 0 · +1· +2· =
0.4 0.4 0.4 4

335
C.2.3 Conditional Moments Exam P Handouts – Page 336

Example Continued 3

4 5
E[X | Y = 1] = E[X | Y = 2] =
3 4
Key point: E[X | Y ] is a function of Y . As a result, it is also a random variable.

P[E[X | Y ] = 4/3] = P[Y = 1] = 0.6


P[E[X | Y ] = 5/4] = P[Y = 2] = 0.4
4 5
E[E[X | Y ]] = 0.6 · + 0.4 ·
3 4
= 1.3
= E[X ]

That is not a coincidence.


Remember: E[X | Y ] is random. E[X ] is a (non-random) number.

Double Expectation Theorem 4

Note that the conditional expectation of X is a function of y :


X
E[X | Y = y ] = x · P[X = x | Y = y ]
all x

When E[X | Y = y ] is a nice function of y , the unconditional expectations are also


closely related:
Theorem (Double Expectation)

E[X ] = E[E[X | Y = y ]]
X
E[X ] = E[X | Y = y ] · P[Y = y ]
all y

This is useful when the distribution of X is complicated, but the conditional


distribution (X | Y ) and distribution of Y are both nice.

336
C.2.3 Conditional Moments Exam P Handouts – Page 337

Example 5

Let N be the value rolled by a fair six-sided die. Suppose that I then flip N
independent fair coins. What is the expected number of heads? What is the variance
in the number of heads?

The key to this example is that if we know N then it is easy to find the first and
second moment
N
E[Heads | N] =
2
So by double expectation,

E[Heads] = E[E[Heads | N]]


 
N 7/2 7
=E = =
2 2 4

Example (Cont) 6

For the second moment, if we know N then the number of heads (H) is binomial with
N trials and p = 1/2.
That means that
N N
E[H | N] = , Var[H | N] =
2 4
E[H | N] = Var[H | N] + (E[H | N])2
2

N N2
= +
4 4
2 2
E[H ] = E[E[H | N]]
 
N N2
=E +
4 4

337
C.2.3 Conditional Moments Exam P Handouts – Page 338

Example (cont) 7

Since N is uniform on {1, 2, . . . , 6},

1+6
E[N] =
2
2
6 −1 35
Var[N] = =
12 12
 2
35 7 91
E[N 2 ] = + =
12 2 6
 
N N2
E[H 2 ] = E +
4 4
7 91 14
= + =
2·4 6·4 3
 2
14 7 77
Var[H] = − =
3 4 48

Law of Total Variation 8

E[X ] = E[E[X | Y ]]
E[X 2 ] = E[E[X 2 | Y ]]

but the analogous statement is not true for variances which is why we used
Var[H] = E[H 2 ] − (E[H])2 .
For the variance you need an additional term:
Definition (Law of Total Variation)

Var(X ) = E[Var(X | Y )] + Var[E(X | Y )]

This is heavily tested on later exams, but we only care about one specific application.

338
C.2.3 Conditional Moments Exam P Handouts – Page 339

Variances of Random Sums 9

Suppose S = X1 + · · · + XN , where X1 , X2 , . . . are i.i.d. and N is an independent


integer valued random variable. What is Var[S]?
Because we would understand it if N was fixed, we will use total variation.

E[S | N] = NE[X ]
Var[S | N] = N · Var[X ]
Var[S] = E[Var[S | N]] + Var[E[S | N]]
= E[NVar[X ]] + Var[NE[X ]]
= E[N]Var[X ] + Var[N](E[X ])2

The final result is perhaps important enough to be worth memorizing.

Exercise 1 10

The number of losses N is Poisson with mean 3. Loss amounts are mutually
independent, and also independent of the number of losses, and have mean 5 and
variance 20. What is the expected value of the sum of all the losses?

339
C.2.3 Conditional Moments Exam P Handouts – Page 340

Exercise 1 10

The number of losses N is Poisson with mean 3. Loss amounts are mutually
independent, and also independent of the number of losses, and have mean 5 and
variance 20. What is the expected value of the sum of all the losses?

Let X denote a loss amount.

E[S] = E[E[S | N]]


= E[NE[X ]]
= E[N]E[X ]
= 3 · 5 = 15

Exercise 2 11

The number of losses N is Poisson with mean 3. Loss amounts are mutually
independent, and also independent of the number of losses, and have mean 5 and
variance 20. What is the variance of the sum of all the losses?

340
C.2.3 Conditional Moments Exam P Handouts – Page 341

Exercise 2 11

The number of losses N is Poisson with mean 3. Loss amounts are mutually
independent, and also independent of the number of losses, and have mean 5 and
variance 20. What is the variance of the sum of all the losses?

E[N] = 3
Var[N] = E[N] = 3
E[X ] = 5
Var[X ] = 20
Var[S] = E[N]Var[X ] + Var[N](E[X ])2
= 3 · 20 + 3 · 52
= 135

341
C.3.1 Order Statistics Exam P Handouts – Page 342

Order Statistics 1

Maximums
Minimums
Discrete Example
General Formulas
Exercises

Maximum Example 2

Claim amounts for flood damage are independent random variables with common
density function 
 4 for x > 1
f (x) = x 5
0 otherwise
where x is the amount of a claim in thousands.
Suppose 3 such claims X1 , X2 , X3 will be made. Find the CDF and density of the
largest of the 3 claims.

Let M denote the maximum loss amount. The key idea is that M ≤ x if and only if
each of the 3 claims are at most x.

342
C.3.1 Order Statistics Exam P Handouts – Page 343

Maximum Example (cont.) 3

4
For one claim, f (x) = for x > 1
x5
P[M ≤ x] = P[all 3 losses ≤ x]
= (P[X1 ≤ x])3
 x 3
Z
4 
= dt
t5
1
 3
1 x
= − 4
t 1
 
1 3
= 1− 4 x >1
x

Max Example (Density) 4

What about the density? For x > 1,

FM (x) = P[all 3 losses ≤ x]


 
1 3
= 1− 4 , x >1
x
 
1 2 4
fM (x) = 3 1 − 4
x x5

and the density is 0 for x < 1 as the loss amounts must all be at least 1 so their
maximum must also be at least 1. (And note FM (1) = 0)

343
C.3.1 Order Statistics Exam P Handouts – Page 344

Minimum Example 5

Claim amounts are independent random variables with common density function

 4 for x > 1
f (x) = x 5
0 otherwise
where x is the amount of a claim in thousands.
Suppose 3 such claims will be made. What is the expected value of the smallest of the
three claims?

Let Y denote the minimum loss amount. For minimums, the key idea is Y > x only if
all the individual losses are > x. Note inequalities are reversed from maxes.
We can find E[Y ] in one of two ways:
Z∞ Z∞
E[Y ] = P[Y > y ] dy = y · fY (y ) dy
0 1

Minimum Example (Cont). 6

4
For one claim, f (x) = for x > 1
x5
P[Y > y ] = P[all 3 losses > y ]
∞ 3
Z
4 
= dt
t5
y
!3
1 ∞
= − 4
t y
 3
1
= y >1
y4
1
= 12 y >1
y

344
C.3.1 Order Statistics Exam P Handouts – Page 345

Minimum Example: Expected Value Computation 7

1
P[Y > y ] = y >1
y 12
Z 1 Z ∞
1
E[Y ] = 1 dx + dy
0 1 y 12
1 12
=1+ =
11 11
1
FY (y ) = 1 − 12 y > 1
y
12
fY (y ) = 13 y > 1
y
Z ∞ ∞
12 −12 12
E[Y ] = y · 13 dy = =
1 y 11y 11 1 11

Discrete Example 8

If I roll a fair die 5 times, what is the probability that the maximum roll is 4?
This is harder because we could have exactly one 4 and four smaller rolls or two 4s and
three smaller, or ...
But dealing with inequalities is easier.

P[Max roll ≤ 4] = P[All rolls are ≤ 4]


 5
4
=
6
P[Max roll = 4] = P[Max roll ≤ 4] − P[Max roll ≤ 3]
 5  5
4 3
= − = 0.100
6 6

345
C.3.1 Order Statistics Exam P Handouts – Page 346

General Max/Min Formulas 9

Suppose X1 , . . . , Xn are iid random variables. Let Y1 = min{X1 , . . . , Xn } and let


Yn = max{X1 , . . . , Xn }. Then

P[max{X1 , . . . , Xn } ≤ x] = P[Yn ≤ x] = [FX (x)]n


P[min{X1 , . . . , Xn } > x] = P[Y1 > x] = (P[X > x])n

If X1 , . . . , Xn are iid discrete random variables, then

P[Yn = x] = P[Yn ≤ x] − P[Yn ≤ x − 1]


= [FX (x)]n − [FX (x − 1)]n
P[Y1 = x] = P[Y1 ≥ x] − P[Y1 ≥ x + 1]
= (P[X ≥ x])n − (P[X ≥ x + 1])n

Exercise 1 10

Suppose that X1 , X2 , X3 , X4 are iid exponential random variables, each with mean 3.
Find the probability that at least one of them exceeds 5.

346
C.3.1 Order Statistics Exam P Handouts – Page 347

Exercise 1 10

Suppose that X1 , X2 , X3 , X4 are iid exponential random variables, each with mean 3.
Find the probability that at least one of them exceeds 5.

Since we are talking about maxes, it is easier to work with the CDF. Let M denote the
maximum of our Xi . Then

P[M > 5] = 1 − P[M ≤ 5]


= 1 − P[X1 ≤ 5, X2 ≤ 5, . . . ]
= 1 − [FX (5)]4
= 1 − (1 − e −5/3 )4
= 0.567

Exercise 2 11

Let N1 , N2 , . . . , N5 be 5 iid Poisson random variables with mean 1.2. Find the
probability that the maximum of these 5 variables is 2.

347
C.3.1 Order Statistics Exam P Handouts – Page 348

Exercise 2 11

Let N1 , N2 , . . . , N5 be 5 iid Poisson random variables with mean 1.2. Find the
probability that the maximum of these 5 variables is 2.

P[max{N1 , . . . N5 } = 2] = P[max{N1 , . . . N5 } ≤ 2] − P[max{N1 , . . . N5 } ≤ 1]


= (P[N1 ≤ 2])5 − (P[N1 ≤ 1])5
 5  5
= e −1.2 (1 + 1.2 + 1.22 /2) − e −1.2 (1 + 1.2)
= 0.87955 − 0.66265
= 0.3985

348
C.3.2 General Order Stats Exam P Handouts – Page 349

General Order Stats 1

Medians
Distribution and density of order stats
Exercises

Finding the Median 2

Suppose we have some data, and want to know whether or not the median m of the
sample is ≤ 4.

Sample: 1.4, 5.3, 3.8 : m = 3.8 ≤ 4


Sample: 3.7, 2.3, 1.8 : m = 2.3 ≤ 4
Sample: 4.4, 5.3, 3.8 : m = 4.4 > 4
Sample: 4.1, 6.2, 5.5 : m = 5.5 > 4
Sample: 2.6, 1.5, 3.2 : m = 2.6 ≤ 4

Note that m ≤ 4 if 2 or 3 of the data points are ≤ 4, but m > 4 if 0 or 1 data point
are ≤ 4.

349
C.3.2 General Order Stats Exam P Handouts – Page 350

Example 3

Claim amounts for wind damage to insured homes are independent random variables
with common density function

 3 for x > 1
f (x) = x 4
0 otherwise

where x is the amount of a claim in thousands.


Suppose 3 such claims will be made. Find the density and CDF of the median of the 3
claims.

Let Y denote the median. Y ≤ y if either 2 or 3 claims are ≤ y .

Median Example: CDF 4

3
f (x) = , x > 1, Y = Median
x4
Z y
3 1
P[X ≤ y ] = 4
dx = 1 −
1 x y3
P[Y ≤ y ] = P[at least 2 claims ≤ y ]
= P[exactly 2 claims ≤ y ] + P[all 3 claims ≤ y ]
    
3 1 2 1 1 3
= 1− 3 + 1− 3
2 y y3 y
 2  
1 3 1
= 1− 3 3
+1− 3
y y y
 2  
1 2
= 1− 3 1+ 3
y y

350
C.3.2 General Order Stats Exam P Handouts – Page 351

Median Example: Density 5

   
1 2 2
FY (y ) = 1 − 3 1+ 3
y y
       
1 3 2 1 2 2 · (−3)
fY (y ) = 2 1 − 3 1+ 3 + 1− 3
y y4 y y y4
  
6 1 2 −1
= 4 1− 3 1+ 3 −1− 3
y y y y
 
18 1
= 7 1− 3
y y

Distributions of order stats 6

Suppose that we have n data points, denoted as X1 , X2 , . . . , Xn , where the Xi are iid
random variables. We can sort the n data points from smallest to largest:

X(1) < X(2) < X(3) < · · · < X(n)


Y1 < Y2 < Y3 < . . . < Yn

The i-th smallest data point is called the i-th order statistic, and is denoted either as
X(i) or Yi .

351
C.3.2 General Order Stats Exam P Handouts – Page 352

Distributions of order stats 7

So Y1 is the smallest data value, aka the minimum.

P[Y1 ≤ y ] = P[at least one Xi is ≤ y ]


 
= 1 − P all Xi > y
n
= 1 − 1 − P(X1 ≤ y )
n
= 1 − 1 − FX (y )

Yn is the largest data value (aka the max), and

P[Yn ≤ y ] = P[all n points are ≤ y ]


 n
= FX (y )

Distributions of Order Stats 8

The other order statistics are messier (and are tested less):

   n
P Yn ≤ y = FX (y )
   
P Yn−1 ≤ y = P at least n − 1 are ≤ y
   
= P exactly n − 1 are ≤ y + P all n are ≤ y
 
n  n−1    n
= FX (y ) 1 − FX (y ) + FX (y )
n−1
   
P Yn−2 ≤ y = P at least n − 2 are ≤ y
   
= P exactly n − 2 ≤ y + P at least n − 1 are ≤ y
 
n  n−2 2  
= FX (y ) 1 − FX (y ) + P Yn−1 ≤ y
n−2
..
.

352
C.3.2 General Order Stats Exam P Handouts – Page 353

Densities of Order Stats 9

The density of a general order stat has a nicer formula:

fi (y ) = density of Yi

fi (y )dy = P y ≤ Yi ≤ y + dy ]

For that to happen, we need i − 1 data values to be less than y , one that is between y
and y + dy , and n − i that are greater than y
 
n n−i i−1
fi (y )dy = i · 1 − FX (y ) · FX (y ) · P[y ≤ X ≤ y + dy ]
n−i
n!  n−i
fi (y ) = i · 1 − FX (y ) FX (y )i−1 · fX (y )
i!(n − i)!
n!  n−i
= 1 − FX (y ) FX (y )i−1 · fX (y )
(i − 1)!(n − i)!

Exercise 1 10

Suppose that X1 , X2 , X3 , X4 , X5 are iid exponential random variables, each with mean
4. Find the density of the median of those variables.

353
C.3.2 General Order Stats Exam P Handouts – Page 354

Exercise 1 10

Suppose that X1 , X2 , X3 , X4 , X5 are iid exponential random variables, each with mean
4. Find the density of the median of those variables.

Let Y3 denote the median order statistic. From our formula,


 
5
f3 (y ) = 3 · [1 − FX (y )]5−3 FX (y )3−1 fX (y )
5−3
1
= 3 · 10[e −y /4 ]2 (1 − e −y /4 )2 e −y /4
4
15
= e −3y /4 (1 − e −y /4 )2
2

Exercise 2 11

Let W1 , W2 , . . . , W4 be 4 iid Poisson random variables with mean 3.2. Find the
probability that the minimum of these 4 variables is 2.

354
C.3.2 General Order Stats Exam P Handouts – Page 355

Exercise 2 11

Let W1 , W2 , . . . , W4 be 4 iid Poisson random variables with mean 3.2. Find the
probability that the minimum of these 4 variables is 2.

P[min{W1 , . . . , W4 } = 2] = P[min{W1 , . . . , W4 } ≥ 2] − P[min{W1 , . . . , W4 } ≥ 3]


= (P[W1 ≥ 2])4 − (P[W1 ≥ 3])4
= (1 − P[W1 = 0] − P[W1 = 1])4 − (P[W1 ≥ 3])4
= (1 − exp(−3.2)(1 + 3.2))4
− (1 − exp(−3.2)(1 + 3.2 + 3.22 /2))4
= 0.82884 − 0.62014
= 0.324

355
C.4.1 Multivariate Review Exam P Handouts – Page 356

Multivariate Review 1

Joint Distributions
Marginal Distributions
Conditional Distributions
Uniforms
Moments
Order Stats

Joint Distributions 2

X
P[(X , Y ) ∈ A] = P[(X , Y ) = (x, y )]
(x,y )∈A
X
P[(X , Y ) = (x, y )] = 1
x,y

F (x, y ) = FX ,Y (x, y ) = P[X ≤ x, Y ≤ y ]


FX ,Y (x, ∞) = FX (x)
FX ,Y (∞, y ) = FY (y )

356
C.4.1 Multivariate Review Exam P Handouts – Page 357

Marginal Distributions 3

In the discrete case


X
P[X = x] = P[(X , Y ) = (x, y )]
y
X
P[Y = y ] = P[(X , Y ) = (x, y )]
x

Essentially we are breaking the probability up into all the possible cases, with the
second variable counting the cases.

The marginal of X can only depend on x, not y . Likewise, the marginal of Y can only
depend on y , not x.

Conditional Distributions 4

In the discrete case


P[X = x, Y = y ]
P[Y = y | X = x] =
P[X = x]
P[X = x, Y = y ]
= P
P[X = x, Y = y ]
y

P[X = x, Y = y ] = P[X = x] · P[Y = y | X = x]

X and Y are independent if 1) the joint distribution factors and 2) the support is a
rectangle.
The conditional distribution of Y given X = x can depend on both x and y .

357
C.4.1 Multivariate Review Exam P Handouts – Page 358

Example 5

(x 2 + 1)(y + 1)
The joint distribution of X and Y is given by P[X = x, Y = y ] = for
73
x and y integers such that 0 ≤ y ≤ |x| ≤ 2, and is 0 otherwise. Find the marginal
distribution of Y .
P[Y = 2] = P[Y = 2, X = 2] + P[Y = 2, X = −2]
1 2 
= (2 + 1)(2 + 1) + ((−2)2 + 1)(2 + 1)
73
15 + 15 30
= =
73 73 
P[Y = 1] = 2 P[Y = 1, X = 2] + P[Y = 1, X = 1]
2 2 
= (2 + 1)(1 + 1) + (12 + 1)(1 + 1)
73
2 28
= (10 + 4) =
73 73
73 − 28 − 30 15
P[Y = 0] = 1 − P[Y = 1] − P[Y = 2] = =
73 73

Example 6

(x 2 + 1)(y + 1)
The joint distribution of X and Y is given by P[X = x, Y = y ] = for
73
x and y integers such that 0 ≤ y ≤ |x| ≤ 2, and is 0 otherwise. Find the conditional
probability that X = 1 given Y = 1.

P[X = 1, Y = 1]
P[X = 1 | Y = 1] =
P[Y = 1]
1 2
73 [(1 + 1)(1 + 1)]
= 28
73
4
=
28
1
=
7

358
C.4.1 Multivariate Review Exam P Handouts – Page 359

Uniforms 7

(X , Y ) are jointly uniform if all possible (x, y ) are equally likely, i.e., the joint
probability function is constant.

If (X , Y ) are jointly uniform, then the conditional distributions are uniform. The
marginals need not be uniform.

If P[(X , Y ) = (x, y )], when positive, only depends on x, then the conditional
distribution of Y is constant and hence uniform. If P[(X , Y ) = (x, y )] only depends on
y , then the conditional distribution of X is uniform.

Moments 8

X
E[g (X , Y )] = g (x, y ) · P[(X , Y ) = (x, y )]
x,y
X
E[X ] = x · P[(X , Y ) = (x, y )]
x,y
  X
2
E Y = y 2 · P[(X , Y ) = (x, y )]
x,y
X
E[Y | X = x] = E[Y | X ] = y · P[Y = y | X = x]
h i
E Y k = E[E[Y k | X ]]
Var[Y ] = E[Var[Y | X ]] + Var[E[Y | X ]]
PN
If S = i=1 Xi , where Xi are iid and N is independent,
Var[S] = E[N] · Var[X ] + Var[N] · (E[X ])2

359
C.4.1 Multivariate Review Exam P Handouts – Page 360

Example 9

The joint distribution of X and Y is P[X = x, Y = y ] = x/14 for x and y integers


with 1 ≤ y ≤ x ≤ 3. Find E[X ] and E[Y ].
3 X
X x
x 1 2 2 3
E[X ] = x· =1· +2· +2· +3· ·3
14 14 14 14 14
x=1 y =1
1 + 8 + 27 18
= =
14 7

P[X = x, Y = y ] doesn’t involve y , so (Y | X = x) is uniform on {1, . . . , x}


1+X
E[Y | X ] =
2  
1+X
E[Y ] = E[E[Y | X ]] = E
2
1 1 18 25
= + · =
2 2 7 14

Covariance 10

   
Var[X ] = σ 2 = E X 2 − (E[X ])2 = E (X − µX )2
Cov(X , Y ) = E[XY ] − E[X ] · E[Y ] = E[(X − µX )(Y − µY )]
Var[X + Y ] = Var[X ] + 2Cov(X , Y ) + Var[Y ]
Var[aX + bY ] = a2 Var[X ] + 2ab Cov(X , Y ) + b 2 Var[Y ]
Var[X − Y ] = Var[X ] − 2Cov(X , Y ) + Var[Y ]
Cov(X , Y )
Corr(X , Y ) = ρ =
SD(X ) · SD(Y )

If X and Y are independent, then Cov(X , Y ) = 0.

360
C.4.1 Multivariate Review Exam P Handouts – Page 361

Example 11

Suppose that X and Y are random variables with Corr(X , Y ) = 0.3, E[X ] = E[Y ] = 1
and SD[X ] = SD[Y ] = 2. What are E[2X − 3Y ] and Var[2X − 3Y ]?

E[2X − 3Y ] = 2 · E[X ] − 3 · E[Y ]


= −1
Var[2X − 3Y ] = 4Var[X ] − 12Cov[X , Y ] + 9Var[Y ]
= 4 · 4 − 12 · 0.3 · 2 · 2 + 9 · 4
= 37.6

Order Stats 12

Suppose X1 , . . . , Xn are iid and integer valued.


Then for integer k,

P[max{X1 , . . . , Xn } = k] = P[max{X1 , . . . , Xn } ≤ k] − P[max{X1 , . . . , Xn } ≤ k − 1]


n n
= P[X ≤ k] − P[X ≤ k − 1]
P[min{X1 , . . . , Xn } = k] = P[min{X1 , . . . , Xn } ≥ k] − P[min{X1 , . . . , Xn } ≥ k + 1]
n n
= P[X ≥ k] − P[X ≥ k + 1]

361

You might also like