Approximation of Random Variables
[Link]@[Link]
Mathematics Department
INSAT - Tunis - Tunisia
January 23, 2023
[Link]@[Link] 1
Table of contents
1 Introduction
2 Convergence of random variables
Convergence almost surely
Convergence in probability
Convergence in mean and in mean square
Convergence in distribution
3 Limit Theorems I
Confidence Interval for µ, I
4 Limit Theorems II
Confidence Interval for µ, II
[Link]@[Link] 2
Introduction
Plan
1 Introduction
2 Convergence of random variables
Convergence almost surely
Convergence in probability
Convergence in mean and in mean square
Convergence in distribution
3 Limit Theorems I
Confidence Interval for µ, I
4 Limit Theorems II
Confidence Interval for µ, II
[Link]@[Link] 3
Introduction
What should you know after this lecture?
1 Different modes of convergence of a sequence of random variables.
2 The Strong Law of Large Number (SLLN).
3 The Central Limit Theorem (CLT).
4 Some results about approximation of random variables.
[Link]@[Link] 4
Introduction
Let {Xn }n∈N be a sequence of random variables, and X another given
random variable, all defined on the same probability space (Ω, F, P).
When we let n → +∞, a question arises, how would we treat the
parameter ω ∈ Ω ?
???
Xn (ω) −−−−→ X (ω)
n→+∞
.
[Link]@[Link] 5
Convergence of random variables
Plan
1 Introduction
2 Convergence of random variables
Convergence almost surely
Convergence in probability
Convergence in mean and in mean square
Convergence in distribution
3 Limit Theorems I
Confidence Interval for µ, I
4 Limit Theorems II
Confidence Interval for µ, II
[Link]@[Link] 6
Convergence of random variables
Convergence almost surely
Definition
A given sequence of random variable {Xn }n∈N converges almost surely to
another given random variable X , if
n o
P ω ∈ Ω : Xn (ω) −−−−→ X (ω) = 1,
n→+∞
we denote
a.s.
Xn −−−−→ X .
n→+∞
[Link]@[Link] 7
Convergence of random variables
Convergence in Probability
Definition
Let {Xn }n∈N be a sequence of random variables. It is said to converge in
probability to a given random variable X , if for all ϵ > 0
n o
P | Xn (ω) − X (ω) |> ϵ −−−−→ 0,
n→+∞
we denote
P
Xn −−−−→ X .
n→+∞
[Link]@[Link] 8
Convergence of random variables
Example
Let {X1 , X2 , · · · } be a sequence of independent random variables such that
1
Xn ∼ B( α ), where α > 0.
n
Study the convergence of {Xn }n∈N to 0 almost surely and in
probability.
[Link]@[Link] 9
Convergence of random variables
Solution#1
\
P{Xn = 0 : n ≥ 1} = P {Xn = 0}
n≥1
Y
= P{Xn = 0}
n≥1
Y 1
= (1 − )
nα
n≥1
Y 1 X 1
(1 − ) < ∞ ⇔ < ∞.
nα nα
n≥1 n∈N
So we have convergence ”a.s.” for only α > 1.
[Link]@[Link] 10
Convergence of random variables
Solution#2
Let 0 < ϵ < 1
n o n o
P |Xn | > ϵ = P Xn = 1
1
= −−−−→ 0
nα n→+∞
And so we have convergence in probability for all α > 0.
[Link]@[Link] 11
Convergence of random variables
Convergence in mean and in mean square
Definition
Let {Xn }n∈N be a sequence of random variables and X another given
random variable.
i. {Xn }n∈N converges in mean to X , if
E {| Xn − X |} −−−−→ 0,
n→+∞
and we denote
m.
Xn −−−−→ X .
n→+∞
ii. {Xn }n∈N converges in mean square to X , if
E | Xn − X |2 −−−−→ 0,
n→+∞
and we denote
m.s.
Xn −−−−→ X .
[Link]@[Link] n→+∞ 12
Convergence of random variables
Prposition
If a given sequence of random variables converges in mean square then it
converges in mean.
Proof
q
E|X | ≤ E [|X |2 ].
[Link]@[Link] 13
Convergence of random variables
Example
We re-consider the example above and we study the convergence in mean
and in mean square.
1 ∀α>0
E|Xn | = E |Xn |2 = α −−−−→ 0.
n n→+∞
[Link]@[Link] 14
Convergence of random variables
Convergence in Distribution
Definition
Let {Xn }n∈N be a sequence of random variables and X another given
random variable. The sequence {Xn }n∈N is said to converge in distribution
to X , if for any continuous and bounded function ψ on R
E ψ(Xn ) −−−−→E ψ(X ) ,
n→+∞
we denote
D
Xn −−−−→ X .
n→+∞
[Link]@[Link] 15
Convergence of random variables
Convergence in distribution: Characterization
Theorem
Let {Xn }n∈N be a sequence of random variables and X another given
random variable, denote by FXn and FX their respective cumulative
functions. Then
D
Xn −−−−→ X
n→+∞
if and only if
FXn (x) −−−−→ FX (x)
n→+∞
for all x ∈ R at which FX is continuous.
[Link]@[Link] 16
Convergence of random variables
Convergence in distribution: Discrete case
Theorem
Let {Xn }n∈N be a sequence of discrete random variables and X another
given discrete random variable, suppose moreover that they havce all the
same range {xk }. Then
D
Xn −−−−→ X
n→+∞
if and only if
∀k
P{Xn = xk } −−−−→ P{X = xk }.
n→+∞
[Link]@[Link] 17
Convergence of random variables
Example
Let λ > 0 and {X1 , X2 , · · · } a sequence of discrete random variables such
that
λ
Xn ∼ B n; , for n > λ.
n
Show that {Xn }n∈N converges in distribution to the Poisson distribution
P(λ).
[Link]@[Link] 18
Convergence of random variables
Solution
We just need to show that
∀k λk
P{Xn = k} −−−−→ e −λ
n→+∞ k!
k
λ n−k
λ
lim P{Xn = k} = lim C k 1−
n→∞ n→∞ n n n
λ n−k
k n! 1
= λ lim 1−
n→∞ k!(n − k)! nk n
k
λ n(n − 1)(n − 2)...(n − k + 1)
= lim × ···
k! n→∞ nk
" #!
λ n λ −k
· · · × lim 1− 1−
n→∞ n n
λk −λ
= e .
k!
[Link]@[Link] 19
Convergence of random variables
Convergence diagram
a.s. Convergence
⇒
⇒ Convergence in P ⇒ Converg. in D
m.s. Convergence
⇒
m. Convergence
[Link]@[Link] 20
Limit Theorems I
Plan
1 Introduction
2 Convergence of random variables
Convergence almost surely
Convergence in probability
Convergence in mean and in mean square
Convergence in distribution
3 Limit Theorems I
Confidence Interval for µ, I
4 Limit Theorems II
Confidence Interval for µ, II
[Link]@[Link] 21
Limit Theorems I
Random Sample
Let X be a given random variable. A random vector (X1 , · · · , Xn ) is a
random sample of X if all Xi ’s are mutually independent and identically
distributed following all the same distribution as X .
1 A given statistical series {x1 · · · , xn } can be seen as a realization of
the random sample of X , (X1 , · · · , Xn ), this means that:
▶ we have realized each Xi apart and we have obtained xi , and this for all
i = 1, · · · , n.
2 Agiven function of (X1 , · · · , Xn )
T = ϕ(X1 , · · · , Xn )
is called a statistic of X .
[Link]@[Link] 22
Limit Theorems I
Examples of statistics of X
1 The sample mean
n
1X
X = Xi ,
n
i=1
2 The sample variance
n
1 X
S2 = (Xi − X )2 ,
n−1
i=1
3 The minimum and the maximum
Xmin = min Xi Xmax = max Xi .
i i
[Link]@[Link] 23
Limit Theorems I
SLLN: Strong Law of Large Numbers
Theorem
Suppose X a given random variable such that E(|X |) < ∞. If {Xn }n∈N is
a sequence of independent and identically distribution ”iid” random
variable that follow all the same distribution as X , then
n
1X a.e
X = Xi −−−−→ E(X ).
n n→+∞
i=1
Notes
1 The SLLN justify the well known and intuitive point estimate for the
expectation of X , E(X ) = µ by the sample mean x of any n
independent realizations of X .
2 Unfortunately this estimate is not always enough since it depends too
much on sample’s variation ”the n independent realizations of X ”
[Link]@[Link] 24
Limit Theorems I
The Central Limit Theorem (CLT)
Theorem
Suppose X a given random variable with a finite variance
V(X ) = σ 2 < ∞. If {Xn }n∈N is a sequence of iid random variables that
follow all the same distribution as X , then
X − E(X ) X −µ L
q = σ −−−−→ N(0, 1). (1)
√ n→+∞
V(X ) n
Remarks
1 The CLT ensure that asymptotically speaking the statistic X is
normally distributed σ
X ∼ N µ, √ (2)
n
and this, whatever is the initial distribution’s nature of X (incredible
but true!)
[Link]@[Link] 25
Limit Theorems I
Remarks (continued)
2. If X ∼ N(µ, σ) we have equality in both equations (3) and (2) and
this is without letting n goes to infinity.
3. The CLT gives an idea about how the sample mean statistic X
approach the expectation µ, actually on can say that the order of this
1
approximation is ”most likely” of order i.e
2
( )
σ X −µ
P |X − µ| ≤ √ = P −2.57 ≤ σ ≤ 2.57 = 99%. (3)
n √
n
[Link]@[Link] 26
Limit Theorems I
What to do with this? Is it so important?
Suppose σ known and we have a number of realizations of X , for
instance, {x1 , x2 , · · · , x100 } and so a realization of X 100 = x 100
Equation (3) can be read: in 99% of cases one can assume that
σ σ σ σ
µ ∈ [x 100 − √ , x 100 + √ ] = [x 100 − , x 100 + ] (4)
100 100 10 10
Yes indeed, this is actually very important!
[Link]@[Link] 27
Limit Theorems I
When σ is known
We use Theorem 2, let α ∈ [0, 1], a confidence interval of level 1 − α (or
of risk α) of the population mean µ is
σ σ
x̄ − zα/2 √ ≤ µ ≤ x̄ + zα/2 √
n n
or equivalently
σ σ
µ ∈ [x̄ − zα/2 √ , x̄ + zα/2 √ ]
n n
where zα/2 satisfies
α
ϕ(zα/2 ) = 1 −
2
and ϕ the normal CDF.
[Link]@[Link] 28
Limit Theorems I
[Link]@[Link] 29
Limit Theorems I
Example
The number of defects in a sample of electric bulbs produced in a given
factory is x = 1 |0 |1 |3 |2 |0 |1 |2 |0.
If we suppose that the population variance is given by σ = 0.1. An
estimation of the defect rate of this factory with a risk level of 5% is given
by
σ σ 0.1 0.1
µ ∈ [x̄ − z1− α2 √ , x̄ + z1− α2 √ ] = [x̄ − z2.5% √ , x̄ + z2.5% √ ]
n n 9 9
where the sample mean is computed x = 1.11 and z2.5% = 1.96, and so we
get
µ ∈ [1.04, 1.17].
[Link]@[Link] 30
Limit Theorems II
Plan
1 Introduction
2 Convergence of random variables
Convergence almost surely
Convergence in probability
Convergence in mean and in mean square
Convergence in distribution
3 Limit Theorems I
Confidence Interval for µ, I
4 Limit Theorems II
Confidence Interval for µ, II
[Link]@[Link] 31
Limit Theorems II
The sample variance
Let (X1 , · · · , Xn ) be a random sample of a given X , we define the
following statistics
n
1 X
S ∗2 = (Xi − X )2 (5)
n−1
i=1
Proposition
1 E(S ∗2 ) = σ 2
the statistic S ∗2 is unbiased for σ 2 ,
µ4 n−3 2
2 V(S ∗2 ) = − µ where µk sets for the k th moment of X
n n(n − 1) 2
i.e µk = E(X k ).
[Link]@[Link] 32
Limit Theorems II
We apply the SLLN and the CLT to obtain
Theorem
a.e
1 S ∗2 −−−−→ σ 2
n→+∞
(n − 1)S ∗2 D
2
2
−−−−→ χ2n−1
σ n→+∞
X −µ D
3 √ −−−−→ Tn−1
∗
s / n − 1 n→+∞
1
Theorem 3 justify the in formula (5) and especially in implemented
n−1
formula of the sample variance in all statistical software
n
∗2 1 X
s = (xi − x)2
n−1
i=1
[Link]@[Link] 33
Limit Theorems II
When σ is unknown
We use 3. of Theorem 3, let α ∈ [0, 1], a confidence interval of level
1 − α (or of risk α) of the population mean µ is
s∗ s∗
x̄ − tα/2 √ ≤ µ ≤ x̄ + tα/2 √
n−1 n−1
or equivalently
s∗ s∗
µ ∈ [x̄ − tα/2 √ , x̄ + tα/2 √ ]
n−1 n−1
where tα/2 satisfies
α
ϕt (tα/2 ) = 1 −
2
and ϕt the CDF of t-student distribution.
[Link]@[Link] 34
Limit Theorems II
Example
We reconsider the same example as for the case when σ was known. We
need to compute the sample variance that is sx = 1.05 . We get the
following confidence interval
sx sx
µ ∈ [x̄ − tα/2 √ , x̄ + tα/2 √ ] =
n−1 n−1
1.05 1.05
[1.11 − 2.26 √ , 1.11 + 2.26 √ ] = [0.27, 1.95].
8 8
Note that when we don’t know σ we loose accuracy and this is expected
since we do approximate σ by s ∗ .
[Link]@[Link] 35