Note
Note
Obisesan
DISTRIBUTION THEORY II
1 Basic Identities
n
X n(n + 1)
r =
r=1
2
n
X n(n + 1)(2n + 1)
r2 =
r=1
6
n 2
X n(n + 1)
r3 =
r=1
2
n
X
4 n(n + 1)(2n + 1)(3n2 + 3n + 1)
r =
r=1
30
n
X n X n
(1 + t)n = tk
i.e =1
k k
k=0
n
X n
ak bn−k = (a + b)n
k
k=0
X n
n
(1 − t) = (−1)k tk
k
n
n
X n
2 =
k
k=0
n
X a b a+b
=
k n−k n
k=0
1
Distribution Theory II: STA 412 Dr. [Link]
Solution:
−3
2(1 + x)
x>0
Hence; f (x) =
0 x≤0
DOC: CDF This is denoted by F(x) and it is the probability that X takes on values less
than or equal to x i.e. F(x)=P(X ≤ x)
Discrete: In case of discrete random variable, it is clear that
X
F (c) = f (x)
x≤C
2
Distribution Theory II: STA 412 Dr. [Link]
P
Where F (c) = f (x) means sum the values of f (x) for all values of x less than or equal to C.
x≤C
Relationship Between pdf and cdf i.e. f(x) and F(x)
(i) If X is a discrete random variable of interest and the possible values of X are in increasing
order, x1 < x2 < x3 < . . .. Then f (x) < F (x).
Continuous
Z ∞
F (x) = P (x ≤ x) = f (t)dt
−∞
F 0 (x) = f (x)
3
Distribution Theory II: STA 412 Dr. [Link]
Solution:
(i)
f (x) = 4xc
Z 1
f (x)dx = 1
0
Z 1
4xc = 1
Z0 1
4 xc = 1
0
c+1 1
x
4 = 1
c+1 0
1 0
4 − = 1
c+1 c+1
4
= 1
c+1
c+1 = 4
c = 4−1
c = 3
(ii)
Z x
F (x) = P (X ≤ x) = f (t)dt
−∞
Z x Z x
c
4t dt = 4 tc dt c=s
0 0
Z x 4 x
3 t x
→ 4 t dt = 4 = t4 0 = x4
0 4 0
F (x) = x4
0; x≤0
T hus F (x) = x4 ; 0≤x≤1
1;4
1≥x
Distribution Theory II: STA 412 Dr. [Link]
1 3 x
QUESTION: Given f (x) = 4 4
, x = 0, 1, 2, · · · , x > 0 find the distribution
function F (x) of x Solution:
F (x) = P (X ≤ x)
Z x Z x t
1 3
= f (t)dt = dt
−∞ 0 4 4
Z t
1 x 3
= dt
4 0 4
" t+1 #x
3
1 4
=
4 t+1
" t 0#x
3 3
1 4 4
=
4 t+1
0
" 3 t #x
3 1 4
=
4 4 t+1
0
" 3 x #
3 1 4 1
= −
4 4 x+1 0+1
" x #
3
3 4
= −1
16 x + 1
0 x<0
3 x
3 (4)
F (x) = 16 x+1
−1 0≤x<∞
∞≤x
1
x
QUESTION: Given f(x) = 15
, x = 1, 2, 3, 4, 5, Find F(x) of x
5
Distribution Theory II: STA 412 Dr. [Link]
SOLUTION:
x
f (x) =
15
P (x ≤ 0) = f (0) = 0
1
P (x ≤ 1) = f (1) =
15
1 2 3
P (x ≤ 2) = f (1) + f (2) = + =
15 15 15
1 2 3 6
P (x ≤ 3) = f (1) + f (2) + f (3) = + + =
15 15 15 15
1 2 3 4 10
P (x ≤ 4) = f (1) + f (2) + f (3) + f (4) = + + + =
15 15 15 15 15
1 2 3 4 5 15
P (x ≤ 4) = f (1) + f (2) + f (3) + f (4) + f (5) = + + + + = =1
15 15 15 15 15 15
Hence;
0; −∞ < x < 1
1
15
1≤x<2
3
2≤x<3
15
F (x) =
36
3≤x<4
15
10
15
4≤x<5
1 5≤x<∞
x
1 3
f (x) = , x = 0, 1, 2, · · ·
4 4
The sequence of f(x) implies
2 3
1 1 3 1 3 1 3
, , ,···
4 4 4 4 4 4 4
6
Distribution Theory II: STA 412 Dr. [Link]
2 MATHEMATICAL EXPECTATION
QUESTION: Show that E(C) =C for C is a constant
SOLUTION:
Assuming a discrete approach
X
E(C) = Cf (x)
R
X
= C f (x)
R
= Cx1
= C
7
Distribution Theory II: STA 412 Dr. [Link]
= C
X Z ∞
f or f (x) = 1 and f (x)dx = 1
R −∞
P
Question: Show that E[C1 U1 (x) + C2 U2 (x)] = R [C1 U1 (x) + C2 U2 (x)] f (x)
SOLUTION:
8
Distribution Theory II: STA 412 Dr. [Link]
9
Distribution Theory II: STA 412 Dr. [Link]
f (x) = px q 1−x , x = 0, 1
The Bernoulli distribution follows a random process or stochastic process for the outcome
of any specific trial is determined by chance
µ=p
σ 2 = pq
√
Standard deviation σ = pq
i. The experiment contains n repeated trials ( n is defined before the experiment begins).
ii. The result of every can only be classified into one of two mutually exclusive categories.
iii. The probability of success does not change from trial to trial.
10
Distribution Theory II: STA 412 Dr. [Link]
iv. In all cases, as n gets larger the distribution gets closer to being a symmetric, bell-
shaped distribution.
µ = np
σ 2 = npq
√
σ = npq
i. The time interval can be divided into many sub-intervals so small that the probability
of the event occuring in any one sub-interval is almost zero.
ii. The probability of more than one occurence in any sub-interval is negligible.
iii. The occurences of the event are independent i.e. The occurrence of an event in any
interval space of time has no effect on the probability of a second occurrence of the
event in the same or any other interval.
iv. The probability of the single occurrence of thw event in a given interval is proportional
to the lenght of the interval.
v. The probability of an occurence in any of the sub -interval (or the mean rate of
occurence) remains constant throughtout the entire time consideration.
λx e−λ
f (x) = x = 0, 1, 2, · · · , ∞, λ > 0
x!
where λ is the average number of occurrences of the random event in the interval. The poisson
distribution can be used when n is very large and p is very small.
11
Distribution Theory II: STA 412 Dr. [Link]
f (x) = px q 1−x , x = 0, 1
X1
E(x) = xf (x)
x=0
= 0p0 q 1−0 + 1p1 q 1−1
= 0(q) + p
= p=µ
1
X
2
E(x − µ) = (x − µ)f (x)
x=0
= (0 − p)2 p0 q 1−0 + (1 − p)2 p1 q 1−1
= p2 q + q 2 p
= pq[p + q]
= pq(1) = pq
= σ2
12
Distribution Theory II: STA 412 Dr. [Link]
= np
T o obtain σ 2
σ 2 = E(x2 ) − (E(x))2
13
Distribution Theory II: STA 412 Dr. [Link]
n
X n!
= x(x − 1) px q n−x
x=0
(n − x)!x!
n
X x(x − 1)n!
= px q n−x
x=0
x(x − 1)(n − x)!x(x − 1)(x − 2)!
n
X n(n − 1)(n − 2)! x n−x
= p q
x=0
(n − x)!(x − 2)!
n−2
X (n − 2)!
= n(n − 1) px q n−x
x=2
(x − 2)!(n − x)!
Let x − 2 = y x = y + 2
n−2
X (n − 2)!
= n(n − 1) py+2 q n−(y+2)
y=0
y![n − (y + 2)]!
n−2
X (n − 2)!
= n(n − 1) p2 py q n−(y+2)
y=0
y![n − (y + 2)]!
n−2
2
X n−2
= n(n − 1)p py q n−2−y
y
y=2
= n(n − 1)p2
σ 2 = E(x2 ) − [E(x)]2
= n2 p2 − np2 + np − n2 p2
= np − np2
= np[1 − p]
= npq
14
Distribution Theory II: STA 412 Dr. [Link]
λx e−λ
f (x) = x = 0, 1, 2, . . . , ∞ λ > 0
x!
∞
X
E(x) = xf (x)
x=0
∞
X λx e−λ
= x
x=0
x!
∞
−λ
X λx
= e x
x=0
x!
∞
X λx
= e−λ x
x=0
x(x − 1)!
∞
X λx
= e−λ
x=0
(x − 1)!
T hen Let x − 1 = y
∞ ∞
X λy+1 X λy λ
E(x) = e−λ = e−λ
x=0
y! x=0
y!
∞
X λy
= e−λ λ
x=0
y!
0
e = λ
15
Distribution Theory II: STA 412 Dr. [Link]
To obtain variance σ 2
σ 2 = E(x2 ) − [E(x)]2
∴ σ2
= E(x(x − 1) + E(x) − [E(x)]2
∞
X
E(x(x − 1) = x(x − 1)f (x)
x=0
∞
X λx e−λ
= x(x − 1)
x=0
x!
∞
X λx e−λ
= x(x − 1)
x=0
x(x − 1)(x − 2)!
∞
X λx−2 λ2 e−λ
=
x=2
(x − 2)!
∞
2
X λt e−λ
= λ
x=2
t!
= λ2 x1 = λ2
∞
X λt e−λ
f or =1
t=0
t!
∴ σ 2 = E(x(x − 1) + E(x) − [E(x)]2
= λ2 + λ − λ2
σ = λ
Given λ = np
n
show that as n increases p(x) = px q n−x , x = 0, 1, 2, · · · to poisson prob-
x
ability
16
Distribution Theory II: STA 412 Dr. [Link]
λ
p =
n
λ
p(x, n, p) = p x, n,
n
n
= px q n−x
x
x n−x
n λ λ
= 1− , x = 0, 1, . . . , n
x n n
n −x
n(n − 1)(n − 2) · · · (n − x + 1) x λ λ
= λ 1− 1−
x!nx n n
x
n −x
n(n − 1)(n − 2) · · · (n − x + 1) λ λ λ
= 1− 1−
nx x! n n
n(n − 1)(n − 2) · · · (n − x + 1)
But ≡1
nx
−x
λ
lim 1 − ≡1
n→∞ n
n
λ
lim 1 − = e−λ
n→∞ n
If there exist a lot of m + n items of which m are defective and the remaining n of them are not
defective. A sample of r items is drawn randomly without replacement.
Let X denote the number of defective items that is observed in the sample. The random
variable x is the hypergeometric random variable with parameter
m + n and m. Then number of
m
ways of selecting x defective items from m defective is . The number of ways of sellecting
x
n
r −x non-defective items from non-defective items is . The total no of way of selecting
r − x
m n
x defective and r − x non-defective item is .
x r−x
Therefore the probability of observing x defective items in a sample of r items (Probability
17
Distribution Theory II: STA 412 Dr. [Link]
density function) is
m n
x r−x x = 0, 1, 2, · · ·
f or x≤m
m+n r−x≤n
r
PROOF:
m n
x r−x
f (x) = x = 0, 1, 2, · · · , r
m+n
r
m n
r
X x r−x
E(x) = x·
x=0
m+n
r
r m! n
X x!(m−x)!
Cr−x
= x m+n C
x=0 r
r xm(m−1)!) n
X x(x−1)!(m−x)!
Cr−x
= m+n C
x=0 r
m(m−1)!) n
r
X (x−1)!(m−x)!
Cr−x
= m+n C
x=0 r
= Let x − 1 = y; x = y + 1
m−x = m−y−1
= (m − 1) − y
r (m−1)!) n
X y!(m−1−y)!
Cr−x
= m m+n C
x=0 r
m m−1
= m+n C
Cy+r−1−y
r
m
= m+n C
·m+n−1 Cr−1
r
18
Distribution Theory II: STA 412 Dr. [Link]
m (m + n − 1)!
= (m+n)!
X
(r − 1)!(m + n − 1 − r + 1)!
r!(m+n−r)!
mr!(m + n − r)! (m + n − 1)1)
= x
(m + n)! (r − 1)!(m + n − r)!
mr(r − 1)! (m + n − 1)!
= X
(m + n)(m + n − 1)! (r − 1)!
mr
=
m+n
19
Distribution Theory II: STA 412 Dr. [Link]
m
Let = pm,n −→ p 0<p<1
m+n
givenm, n −→ ∞
m n
r−x
x r
then −→ px q r−x
m+n x
r
m n
m!
x r−x x!(m−x)! n!
= (m+n)! X
m+n (r − x)!(n − r + x)!
r!(m+n−r)!
r
m! r!(m + n − r)! n!
= X X
x!(m − x)! (m + n)! (r − x)!(n − r + x)!
r! m!n!(m + n − r)!
= ·
x!(m − x)! (m + n)!(m − x)!(n − r + x)!
r m(m − 1)(m − 2) · · · (m − x+!)(m − x)!n!(m + n − r)!
=
x (m + n)!(m − x)!(n − r + x)!
r m(m − 1) · · · (m − x + 1)n!(m + n − r)!
=
x (m + n)!(n − r + x)!
r m(m − 1) · · · (m − x + 1)n!(m + n − r)!
= !
x (m + n)(m + n − 1) · · · (m + n − r + 1)(m + n − r)(n − r + x)!
r m(m − 1) · · · (m − x + 1)n!
=
x (m + n)(m + n − 1) · · · (m + n − r + 1)(n − r + x1)!
r m(m − 1) · · · (m − x + 1)n(n − 1)! · · · (n − r + x + 1)(n − r + x)!
=
x (m + n)(m + n − 1) · · · (m + n − r + 1)(n − r + x1)!
r m(m − 1) · · · (m − x + 1)n(n − 1)! · · · (n − r + x + 1)
=
x (m + n)(m + n − 1) · · · (m + n − (r − 1))
Divide through by m + n
" m m 1
m x−1
n
n 1
n r−x−1
#
r m+n m+n
− m+n
· · · m+n
− m+n
X m+n m+n
− m+n
· · · m+n
− m+n
= m+n
m+n r−1
x m+n
· · · m+n
− m+n
" m m 1
m x−1
n
n 1
n r−x−1
#
r m+n m+n
− m+n
· · · m+n
− m+n
X m+n
− · · · −
= r−1
m+n m+n m+n m+n
x 1 · · · 1 − m+n
21
Distribution Theory II: STA 412 Dr. [Link]
m
Since p =
m+n
m m+n−m n
1−p = 1− = =
m+n m+n m+n
" 1
x−1
1
r−x−1
#
r p p − m+n · · · p − m+n X q q − m+n ··· q − m+n
T aking the limm+n−→∞ r−1
x 1 · · · 1 − m+n
r
= px q r−x
x
This is the distribution where process deal with an experiment of Bernoulli trials sequence
deserved utill the first success occurs
Note: r − 1 denote the total number of trials before the first success.
Thus: r is the total number of trial at exactly when first success occurs.
POC: r is not predetermined for BINOMIAL exberiment.
r−1
p(1 − p)
r = 1, 2, · · · ∞
f (x) =
0 elsewhere
∞
X
PROOF:E(r) = rp(1 − p)r−1
r=1
∞
X
= rpq r−1
r=1
X∞
= p rq r−1
r=1
∞
∂
X
q + q2 + q3 + · · ·
= p
r=1
∂q
∂
1 + q + q2 + · · ·
= pq
∂q
But 1 + q + q 2 + · · · is a geometric series
a 1
S∞ + = f orr = q
1−r 1−q
22
Distribution Theory II: STA 412 Dr. [Link]
∂ 1
= pq
∂q 1 − q
∂ q
= p
∂q 1 − q
(1 − q)(1) − q(−1)
= p
(1 − q)2
1−q+q
= p
(1 − q)2
1
= p 2
p
1
=
p
T o obtain σ 2
σ 2 = E(r2 ) − [E(r)]2
∞
X ∞
X
E(r2 ) = r2 pq r−1 = p r2 qr − 1
r=1 r=1
∞
X ∂
= p (rq r )
r=1
∂q
∞
X ∂ r
= p rq
r=1
∂q
∞
∂ X r
= p rq
∂q r=1
∂
q + 2q 2 + 3q 3 + · · ·
= p
∂q
∂
= p q 1 + 2q + 3q 2 · · ·
∂q
1
But 1 + 2q + 3q 2 · · · =
(1 − q)2
∂ q
= p
∂q (1 − q)2
(1 − q)2 (1) − q(−2(1 − q))
= p
(1 − q)4
(1 − q)[1 − q + 2q]
= p
(1 − q)4
(1 − q)(1 + q)
= p
(1 − q)(1 + q)3
23
Distribution Theory II: STA 412 Dr. [Link]
1+q
= p
p3
1+q
=
p2
1+q 1
∴ σ2 = 2
− 2
p p
1+q−1
=
p2
q
=
p2
Consider a succession of Bernoulli trials, let p(r) denotes the probability that exactly r+k(k > 0)
trials are needed to produce k successes. This will so happen when the last trial i.e(r + k)th
trials is a success with probability p and the previous trial (r + k − 1) must (k − 1) successes
r+k−1
with probability ck−1 pk−1 q r where q = 1 − p p( r) = p(k − 1) successes in (x+k-1) trials X
prob (r + k)th trial
r+k−1
= ck−1 pk−1 q r · p
r+k−1
= ck−1 pk−1+1 q r
r+k−1
= ck−1 pk q r
Show that
r+k−1 k r k r −k
Ck−1 p q if r = 0, 1, 2, · · · = p (−1) qr
r
SOLUTION:
r+k−1
Ck−1 pk q r
(r + k − 1)!
= pk q r
(k − 1)!(r + k − 1 − k + 1)!
24
Distribution Theory II: STA 412 Dr. [Link]
(r + k − 1)! k r
= p q
(k − 1)!r!
(k + r − 1)(k + r − 2) · · · (k + r + 1 − (r + 1))(k − 1)! k r
= p q
(k − 1)!r!
(k + r − 1)(k + r − 2) · · · k k r
= p q
r!
k(k + 1) · · · (k + r − 2)(k + r − 1)pk q r
=
r!
(−k)(−k − 1) · · · (−k − r + 2)(−k − r + 1) k r
= (−1)r p q
r!
−k
= pk (−1)r qr
r
∞ ∞
X
k
X
r −k
POC: p(r) = p (−1) qr
r
r≥0 r=0
= p (1 − q)−k
k
= pk p−k
= pk−k
= p0 = 1
25
Distribution Theory II: STA 412 Dr. [Link]
5 UNIFORM DISTRIBUTION
Let X ∼ U (a, b)
Then
1
b−a
a<x<b
f (x) =
0, otherwise
Z b
1
E(x) = x· dx
a b−a
Z b
1
= xdx
b−a a
2 b
1 x 1 2
b − a2
= =
b − a 2 a 2(b − a)
(b − 1)(b + a)
=
2(b − a)
b+a
=
2
26
Distribution Theory II: STA 412 Dr. [Link]
To obtain σ 2
σ 2 = E(x2 ) − [E(x)]2
Z b
2 1
E(x ) = x2 dx
a b−a
Z b 3 b
1 2 1 x
= x dx =
b−a a b−a 3 a
1 (b − a)(b2 + ab + a2 )
= (b3 − a3 ) =
3(b − a) 3(b − a)
2 2
b + ab + a
=
2
2
b + ab + a2
2
2 b−a
σ = −
3 2
4b + 4ab + 4a − 3b2 − 6ab − a2
2 2
=
12
2 2
b − 2ab + a
=
12
2
(b − a)
σ2 =
12
6 EXPONENTIAL DISTRIBUTION
Let X ∼ Exponential(λ)
27
Distribution Theory II: STA 412 Dr. [Link]
λ ∞ −y
Z
= ye dy
λ2 0
λ
= Γ2
λ2
1
= (2 − 1)!
λ
1
∴ E(x) =
λ
To obtain σ 2
σ 2 = E(x2 ) − [E(x)]2
Z ∞ Z ∞
−λx
2
E(x ) = 2
x λe =λ x2 e−λx
0 0
y dy dy
Let y = λx, x = = λ dx =
λ dx λ
By Substitution
Z ∞ 2
y 1
E(x2 ) = λ e−y · dy
0 λ λ
Z ∞
λ
= 3
y 2 e−y dy
λ 0
λ
= Γ3
λ3
1
= (3 − 1)!
λ2
2
E(x2 ) =
λ2
2 1 2−1
∴ σ2 = 2
− 2 =
λ λ λ2
1
σ2 =
λ2
28
Distribution Theory II: STA 412 Dr. [Link]
POC:
In a general context
Z ∞
r
E(x ) = xr λe−λx dx
0
y dy dy
Let y = λx, x= = λ dx =
λ dx λ
By Substitution
Z ∞ r
y 1
r
E(x ) = λ e−y · dy
0 λ λ
Z ∞
λ
= r
y r e−y dy
λλ 0
1
= Γr + 1
λr
7 GAMMA DISTRIBUTION:
This is generally known as a distribution frequently used in waiting time modeling
It has pdf as x
xα−1 e− β
f (x) = x > 0, α > 0, β > 0
Γαβ α
Proof: Mean and Variance
In a general case since the mean and variance of any distribution can be obtain by observing
0 0 0
the first and second moment about the origin i.e. µ1 , µ2 , Hence it is adviceable to establish µt
for concerning and to avoid error as the basis for estimation.
29
Distribution Theory II: STA 412 Dr. [Link]
Thus:
x
∞
xr xα−1 e− β
Z
r 0
E(x ) = µr = dx
Γαβ α
∞ Z0
1 α−1 r − βx
= x x e
Γαβ α 0
Z ∞
1 α+r−1 − βx
= x e dx
Γαβ α 0
x dy 1
Let y = ; = ; dx = βdy and x = yβ
β dx β
By substitution
Z ∞
1
= α
(yβ)α+r−1 e−y · βdy
Γαβ 0
Z ∞
1
= β α+r−1
·β y α+r−1 e−y dy
Γαβ α 0
1
= β α+r Γ(α + r)
Γαβ α
β α+r Γ(α + r)
=
Γαβ α
β r Γ(α + r)
=
Γα
β
Hence : E(x) = = αβ
Γα
β2 β2
E(x2 ) = Γ(α + 2) = · (α + 1)Γ(α + 1)
· α
= β 2 (α + 1)α
Thus:
σ2 = E(x2 ) − [E(x)]2
= β 2 (α + 1)α − α2 β 2
= (αβ 2 + β 2 )α − α2 β 2
= αβ 2
8 ERLANG DISTRIBUTION
λr xr−1 e−λx
f (x) = 0 < x, r = 1, 2, · · ·
(r − 1)!
30
Distribution Theory II: STA 412 Dr. [Link]
To obtain the mean and variance using the general approach of moment about the origin:
Z ∞
k λr xr−1 e−λx
E(x ) = xk dx
0 (r − 1)!
Z ∞
λr
= xk xr−1 e−λx dx
(r − 1)! 0
Z ∞
λr
= xk+r−1 e−λx dx
(r − 1)! 0
y dy dy
Let y = λx, x = , = λ dx =
λ dx λ
By Substitution
Z ∞ k+r−1
λr y 1
= e−y · dy
(r − 1)! 0 λ λ
r Z ∞
λ 1 1
= · k+r−1 · y k+r−1 e−y dy
(r − 1)! λ λ 0
λr 1
= · k−1 r Γ(k + r)
(r − 1)! λ λ · λ
1
= k
Γ(k + r)
λ Γr
Hence W hen k = 1 to get
Γ(r + 1)
E(x) =
λΓr
rΓr
=
λΓr
r
=
λ
W hen k = 2
Γ(r + 2)
E(x2 ) =
λ2 Γr
(r + 1)rΓr
=
λ2 Γr
r(1 + r)
=
λ2
∴ σ = E(x2 ) − [E(x)]2
2
2 r(1 + r) h r i2
σ = −
λ2 λ
r + r2 − r2
λ2
r
σ2 =
λ2
31
Distribution Theory II: STA 412 Dr. [Link]
9 WEIBULL DISTRIBUTION
β x β−1 −( σx )β
f (x) = e 0 < x, 0 < β, 0 < σ
σ σ
Z ∞
rβ
x β−1 x β
r
E(x ) = x e−( σ ) dx
0 σ σ
∞
xr xβ−1 −( σx )β
Z
β
= e dx
σ 0 σ β−1
Z ∞
β x β
r+β−1 −( σ ) dx
= x e
σ · σ β · σ −1 0
Z ∞
β r+β−1 −( σx β
) dx
= x e
σβ 0
x β 1 x 1 dy x β−1 1 β x β−1
Lety = ∴ y β = ; x = σy β ; =β · =
σ σ dx σ σ σ σ
32
Distribution Theory II: STA 412 Dr. [Link]
σ
dx = dy
x β−1
β σ
By Substitution
Z ∞
β 1
r+β−1 σ
r
E(x ) = σy β e−y dy
σβ x β−1
0 β σ
Z ∞
β r β 1 σ
= σ r+β−1 y β β β e−y 1 β−1 dy
σβ 0 β
β σyσ
r β 1
β · σ r+β−1 σ ∞ y β + β − β −y
Z
= 1 e dy
σβ · β 0 y · y− β
Z ∞ βr ββ −1
r y y y β −y
= σ 1 e dy
0 y 1− β
Z ∞
r
+β − β1 1
= σ r
y β β − 1− e−y dy
β
Z0 ∞
(r+β−1−1+1)
= σr y β e−y dy
Z0 ∞
r
= σr y β e−y dy
0
r r
= σ Γ 1+
β
W hen r = 1
1
⇒ E(x) = σΓ 1 +
β
W hen r = 2
2 2 2
⇒ E(x ) = σ Γ 1 +
β
2
2 2 1
V ar(x) = σ Γ 1 + − σΓ 1 +
β β
2 !
2 1
V ar(x) = σ 2 Γ 1 + − Γ 1+
β β
10 MAXWELL DISTRIBUTION
x2
2 x2 e− 2a2
r
Given f (x) = x > 0; a > 0
π a3
33
Distribution Theory II: STA 412 Dr. [Link]
34
Distribution Theory II: STA 412 Dr. [Link]
3
22 a
E(x) = √
π
2
21+ 2 a2
2 2+3
E(x ) = √ Γ
π 2
2 2
2a 5
= √ Γ
π 2
2
4a 5
= √ Γ
π 2
" 3 #2
4a2 5 22 a
∴ V ar(x) = √ Γ − √
π 2 π
4a2 Γ 25
8a2
= √ −
π π
" #
5
Γ 2
= 4a2 √ 2 −
π π
5 3 3 3 1 1
But Γ = Γ = · Γ
2 2 2 2 2 2
" #
3 1 1
· Γ 2
V ar(x) = 4a2 2 2 1 2 −
Γ 2 π
√
1
N ote : that π = Γ
2
2 3 2
T hus V (x) = 4a −
4 π
2 8
V ar(x) = a 3−
π
36
Distribution Theory II: STA 412 Dr. [Link]
11 PARETO DISTRIBUTION
kβ k
The probability density function is given by f (x) = xk+1
,k > β, x > 0
By general approach moment
∞
kβ k
Z
r r
E(x ) = x . k+1 dx
β x
Z ∞
= kβ k xr .x−k−1 dx
β
Z ∞
k
= kβ xr−k−1 dx
β
r−k
x
= kβ k [ ]∞
r−k β
kβ k
= [0 − β r−k ]
r−k
kβ k 1
= ( k−r )
k−r β
kβ k 1
= ( k −r )
k−r β β
k
= (β −r )−1
k−r
kβ r
=
k−r
kβ
E(x) = since r = 1
k−1
37
Distribution Theory II: STA 412 Dr. [Link]
kβ 2
Also the E(x2 ) = k−2
12 NORMAL DISTRIBUTION
The probability density function for the normal distribution is given as
1 x−µ 2
f (x) = √ e(− σ ) , −∞ < x < ∞, µ0 , σ 2 > 0
2πσ 2
38
Distribution Theory II: STA 412 Dr. [Link]
To obtain the mean and variance using the general moment approach about the origin.
Z ∞
r 1 x−µ 2
E(x ) = xr √ e(− σ ) dx
2πσ 2
−∞
Z ∞
1 x−µ 2
= √ xr e(− σ ) dx
σ 2π −∞
x−µ 2
Let y 2 = ( )
σ
x−µ x µ
y = = − ;
σ σ σ
x = µ + σy
dy 1
= ; dx = σdy
dx σ
By substitution, we have
Z ∞
1 1
r
E(x ) = √ (µ + σy)r e− 2 y 2 σdy
σ 2π −∞
Z ∞
σ 1
= √ (µ + σy)r e− 2 y 2 dy
σ 2π −∞
Z ∞
1 1
= √ (µ + σy)r e− 2 y 2 dy
2π −∞
Z ∞
1 1 2
E(x) = √ (µ + σy)e− 2 y dy i.e when r = 1
2π −∞
Z ∞ Z ∞
1 − y2 y
= √ [µ e dy + σ e− 2 dy]
2π
Z ∞ −∞ Z −∞
∞
1 −y 1 y
= µ √ e 2 dy + σ √ e− 2 dy
−∞ 2π −∞ 2π
= µ(1) + σ(0) = µ
39
Distribution Theory II: STA 412 Dr. [Link]
NOTE: The mean of a standardized normal distribution is zero and the variance is one. This
implies that given
1 y
f (x) = √ e− 2
Z 2π
∞
1 y
E(y) = y √ e− 2 dy = 0
−∞ 2π
T hus, var(y) = E(y ) − (E(y))2 ⇒ 1 = E(y 2 ) − 02
2
E(y 2 ) = 1
Z ∞
1 1
2
E(x ) = √ (µ + σy)2 e− 2 y 2 dy
2π −∞
Z ∞
1 y2
= √ (µ2 + 2µσy + σ 2 y 2 )e − dy
2π −∞ 2
Z ∞ Z ∞ Z ∞
1 −y 1 −y 1 y
= µ 2
√ e 2 dy + 2µσ y √ e 2 dy + σ 2
y 2 √ e− 2 dy
−∞ 2π −∞ 2π −∞ 2π
2 2
= µ (1) + 2µσ(0) + σ (1)
E(x2 ) = µ2 + σ 2
13 BETA DISTRIBUTION
The beta probability density function is given as
a−1
x
f (x) = (1 − x)b−1 , a > 0, b > 0, 0 < x < 1
β(a, b)
where Z 1
Γ(a)(b) ΓaΓb
β(a, b) = = = xa−1 (1 − x)b−1 dx
Γ(a + b) Γa + b 0
40
Distribution Theory II: STA 412 Dr. [Link]
To obtain the mean and the variance of the beta distribution, the general moment is given as
Z 1
r xa−1
E(x ) = xr (1 − x)b−1 dx
0 β(a, b)
Z 1
1
= xr xa−1 (1 − x)b−1 dx
β(a, b) 0
Z 1
1
= xr+a−1 (1 − x)b−1 dx
β(a, b) 0
1
= β(a + r, b)
β(a, b)
1 Γ(a + r)Γb
= ΓaΓb ∗
Γ(a+b)
Γ(a + r + b)
Γ(a + b) Γ(a + r + b)
= ∗
ΓaΓb Γ(a + r)Γb
Γ(a + b)Γ(a + r)
=
ΓaΓ(a + r + b)
T hen; E(x) ⇒ r=1
Γ(a + b)Γ(a + 1) Γ(a + b).aΓa
E(x) = =
ΓaΓ(a + 1 + b) Γa((a + b))Γ(a + b)
a
=
a+b
Γ(a + b)Γ(a + 2)
E(x2 ) =
ΓaΓ(a + 2 + b)
Γ(a + b)(a + 1)aΓa
=
Γa(a + b + 1)(a + b)Γ(a + b)
a2 + a
=
a2 + ab + 2ab + a + b
a2 + a a2
var(x) = −
a2 + ab + 2ab + a + b a2 + 2ab + b2
41
Distribution Theory II: STA 412 Dr. [Link]
14 RAYLEIGH DISTRIBUTION
−αx2
2αxe
f or x > 0
f (x) =
0 elsewhere
Z ∞
2
E(x) = x2αxe−αx dx
0
Z ∞
2
= 2α x2 e−αx dx
0
p
Let y = αx2 ; x = y/α = (y/α)1/2 ;
dy/dx = 2αx;
1 1 1
dx = dy = dy = dy
2αx (y/α)1/2 2α1/2 y 1/2
Z ∞
1
⇒ 2α ((y/α)1/2 )2 e−y 1/2 1/2 dy
2α y
Z0 ∞
y 1
= 2α . 1/2 1/2 e−y dy
0 α 2α y
Z ∞
2
= y 1/2 e−y dy
2α1/2 0
2
= Γ(3/2)
2α1/2
2
= 1/2
2α Γ(1/2)
√
π 1p
= 1/2
= (π/α)
Z2α ∞
2
2
E(x2 ) = 2αx3 e−αx dx
0
Let y = αx2 ; dy = 2ax dx
1
dx = dy
2ax
and x = (y/α)1/2
42
Distribution Theory II: STA 412 Dr. [Link]
By substitution, we have
Z ∞
1
2
E(x ) = 2α ((y/α)1/2 )3 e−y 1/2 1/2 dy
0 2α y
Z ∞
2
= y 3/2−1/2 e−y dy
2α α3/2 0
1/2
α ∞ −y
Z
1 1
= 2
ye dy = Γ2 =
α 0 α α
2
1 1p
var(x) = − (π/α)
α 2
1 π 1h πi
= − = 1−
α 4α α 4
1 MX (0) = 1
2. MX (t) ≤ 1
NOTE:
dn MX (t)
n
|t=0 = E(X n ); n = 1, 2, 3..., if E|X n | < ∞
dt
43
Distribution Theory II: STA 412 Dr. [Link]
n
f (x) = Cx px q n−x ; x = 0, 1, 2, ..., n
MX (t) = E(etx )
Xn n
X
= etx nCx px q n−x = nCx (pet )x q n−x
i=0 i=0
t n
= (pe + q)
0
MX (t) = n(pet + q)n−1 .pet
0
MX (0) = n(p + q)n−1 .pe( 0) = np...mean
00
MX (t) = n(n − 1)(pet + q)n−2 (pet )2 + n(pet + q)n−1 .pet
00
MX (0) = n(n − 1)(pe0 + q)n−2 (pe0 )2 + n(pe0 + q)n−1 .pe0
= n(n − 1)p2 + np
= np(1 − p)
= npq.
44
Distribution Theory II: STA 412 Dr. [Link]
MX00 (0) = λ + λ2
V ar(x) = λ + λ2 − λ2
=λ
1
f (x) = ;a < x < b
b−a
0 elsewhere
MX (t) = E(etx )
Z b
1
= dx
a b−a
Z b
1 1 etx b
= etx dx = [ ]
b−a a b−a t a
1 etb − eta
= [etb − eta ] =
t(b − a) t(b − a)
2 3
(tb) (tb)
But etb = 1 + b + + + ...
2! 3!
(ta)2 (ta)3
eta = 1 + a + + + ...
2! 3!
T hus
2 2 3 3 4 4 t2 a2 t3 a3 t4 b4
[1 + b + t 2b + t 6b + t 24a + ...] − [1 + a + 2
+ 6
+ 24
+ ...]
Mx (t) =
t(b − a)
t2 (b2 −t2 ) 3 3 3 4 4 4
t(b − a) + 2
+ t (b 6−t ) + t (b24−t ) + ...
=
t(b − a)
t(b − a) t (b − a)(b + a) t3 (b − a)(b2 + ab + a2 ) t4 (b − a)(b + a)(b2 + a2 )
2
= + + +
t(b − a) 2t(b − a) 6t(b − a) 24t(b − a)
2 2 2 3 2 2
t(b + a) t (b + ab + a ) t ((b + a)(b + a )
= 1+ + +
2 6 24
(b + a) t b + ab + a ) t ((b + a)(b + a2 )
( 2 2 2 2
MX0 (t) = + +
2 3 8
b+a
MX0 (0) =
2
45
Distribution Theory II: STA 412 Dr. [Link]
Z ∞
MX (t) = E(e ) = tx
λe−λx etx dx
0
Z ∞
= λ e−λx+tx dx
Z0 ∞
= λ e−(λ+t)x dx
0 −(λ+t)x ∞
e
= λ
t−λ 0
λ −(λ+t)∞
− e−(λ+t)0
= [e
t−λ
λ 1
= =λ
λ−t λ−t
−1
= λ(λ − t)
λ
MX0 (t) = λ(t − λ)−2 =
(λ − t)2
λ 1
MX0 (0) = =
λ2 λ
46
Distribution Theory II: STA 412 Dr. [Link]
2λ
MX00 (t) = 2λ(λ − t)−3 =
(λ − t)3
2λ 2
MX00 (0) = =
λ3 λ2
2 1 2 2−1
V ar(x) = − ( ) =
λ2 λ λ
1
= .
λ2
λr xr−1 e−λx
f (x) = ; x > 0; r = 1, 2...
Γr
MX (t) = E(etx )
Z ∞ r r−1 −λx
λx e
= dx
0 Γr
λr ∞ r−1 −λx+tx
Z
= x e dx
Γr 0
λr ∞ r−1 −(λ+t)x
Z
= x e dx
Γr 0
dy 1
let y = (λ − t)x ; = λ − t; dx = dy;
dx λ−t
y
x =
λ−t
by substitution, we have
λr ∞ y r−1 −y 1
Z
MX (t) = ( ) e dy
Γr 0 λ − t λ−t
Z ∞
λr
= r −1
y r−1 e−y dy
Γr(λ − t)(λ − t) (λ − t) 0
λr
= Γr
Γr(λ − t)r
λr
= = λr (λ − t)r
(λ − t)r
MX0 (t) = rλr (λ − t)−r−1
47
Distribution Theory II: STA 412 Dr. [Link]
λr λ−r r2 + r
= r(r + 1) =
λ2 λ2
2 2
r +r r
V ar = 2
− 2
λ λ
2 2
r + −r
=
λ2
r
= .
λ2
MX0 (0) = rλ
V ar(x) = r2 λ2 + rλ2 − r2 λ2
= rλ2 .
48
Distribution Theory II: STA 412 Dr. [Link]
µ2 µ2 (2µσ 2 t) (σ 2 t)2
= e− σ2 + σ2 + 2σ 2
+
2σ 2
σ 2 t2
= eµt+ 2
σ 2 t2
MX0 (t) = µ + σ 2 t eµt+ 2
kβ k
f (x) = , k > β, x > 0
xk+1
49
Distribution Theory II: STA 412 Dr. [Link]
∞
xr−k
k
= kβ
r−k β
k
kβ
= [0 − β r−k ]
r−k
kβ k
1
=
k − r β k−r
kβ k
1
=
k − r β k β −r
k
= (β −r )−1
k−r
kβ r
E(xr ) =
k−r
kβ
E(x) = since r = 1
k−1
Also the
E(x2 ) = kβ 2 k − 2
50
Distribution Theory II: STA 412 Dr. [Link]
Then var(x)implies
kβ 2 kβ
var(x) = −( )
k−2 k−1
kβ 2 (k − 1)2 − k 2 β 2 (k − 2)
=
(k − 1)2 (k − 2)
kβ (k − 2k + 1) − k 3 β 2 − 2k 2 β 2
2 2
=
(k − 1)2 (k − 2)
k 3 β 2 − 2k 2 β 2 + kβ 2 − k 3 β 2 − 2k 2 β 2
=
(k − 1)2 (k − 2)
kβ 2
V (x) =
(k − 1)2 (k − 2)
EXAMPLE 1
Given
1
M (t) = (1 + et )2 ,
4
find the CGF.
solution:
51
Distribution Theory II: STA 412 Dr. [Link]
1
M (t) = (1 + et )2
4
C(t) = lnM (t)
1 1 1
= ln (1 + et )2 = ln + ln(1 + et )2 = ln + 2ln(1 + et )
4 4 4
t
d 1 2e
= C(t) = .2et =
dt 1 + et (1 + et )
d
= C(t)|t=0 = 1 = µ
dt
d2 (1 + et )2et − 2et et
= C(t) =
dt (1 + et )2
2et
σ2 =
(1 + et )2
d2
⇒ C(t)|t=0
dt
2
=
4
1
σ2 =
2
EXAMPLE 2
t
Find the CGF if MGF = e2(e −1)
solution:
t
MX (t) = e2(e −1)
t
C(t) = lne2(e −1) = 2et − 2
d
= C(t) = 2et
dt
d
= C(t)|t=0 = 2 = µ
dt
d2
= C(t : |t=0 = 2et |t=0 = 2 = σ 2
dt
52
Distribution Theory II: STA 412 Dr. [Link]
EXAMPLE 3
σ 2 t2
Given MX (t) = eµt+ 2 Find the CGF.
solution:
since
σ 2 t2
MX (t) = eµt+ 2
d
C(t)|t=0 = µ + σ 2 t|t=0 = µ
dt
d2
= C(t)|t=0
dt
= σ 2 |t=0
= σ2
EXAMPLE 4
Solution :
= µ
2
d (pet )npet − npet .pet
C(t)|t=0 = |t=0
dt (pet + q)2
np − np2
= = np(1 − p)
1
= npq
= σ2
53
Distribution Theory II: STA 412 Dr. [Link]
EXAMPLE 5
Given
t
MX (t) = eλ(e −1)
t
C(t) = lneλ(e −1)
= λet − λ
d
C(t) = λet
dt
d
C(t)|t=0 = λ = µ
dt
d2
C(t)|t=0 = λet |t=0
dt
= λ
= σ2
17 MARKOV INEQUALITY
If X is a random variable and U (X) is a non- negative real- value function, then for any positive
constant
E(U (X))
C > 0, P (U (X) ≥ C) ≤
C
PROOF:
If A = {x|u(x) ≥ c} then for a continuous random variable
Z ∞
E(u(x)) = u(x)f (x)dx
Z−∞ Z
= u(x)f (x) + u(x)f (x)dx
ZA Ac
≥ u(x)f (x)
ZA
≥ f (x)dx
A
= cp(xA)
= cp(u(x) ≥ c)
54
Distribution Theory II: STA 412 Dr. [Link]
18 CHEBYSHEV INEQUALITY
If the random variable X has a mean µ and variance σ 2 then for every
1
k ≥ 1, P (|x − µ| ≥ kσ) ≤
k2
PROOF:
Let f (x) denote the pdf of X , then
σ 2 = E(x − µ)2
X
= (x − µ)2 f (x)
xR
X X
= (x − µ)2 f (x) + (x − µ)2 f (x)
xA xAc
A = {X : |x − µ| ≥ kσ}
X
⇒ σ2 ≥ (x − µ)2 f (x)
xR
since in A, |x − µ| ≥ kσ
By substitution,
X
⇒ σ2 ≥ (kσ)2 ≥ kσ
xA
X
= k2σ2 f (x)
xA
X
But f (x) = P (xA)
xA
σ 2 ≥ k 2 σ 2 P (xA)
σ 2 ≥ k 2 σ 2 P (|x − µ| ≥ kσ)
1 ≥ k 2 P (|x − µ| ≥ kσ)
1
≥ P (|x − µ| ≥ kσ)
k2
1
P (|x − µ| ≥ kσ) ≤
k2
55
Distribution Theory II: STA 412 Dr. [Link]
and
ε = kσ,
then
ε
k =
σ
σ2
P (|x − µ| ≥ kσ) ≤ 2 .
ε
PROOF:
Ms (t) = (M (t))n
where s = Σxi
MW (t) = E(etw )
t Σxi − nµ
= E e √
σ n
t s − nµ
= E e √
σ n
t s nµ
= E e √ − √
σ n σ n
√ h st i
− µtσ n √
= e E eσ n
√ h tx in
µt n √
= e− σ E e σ n
√
n
− µtσ n t
= e M √
σ n
56
Distribution Theory II: STA 412 Dr. [Link]
√ "
2
#
µt n µt (µ2 + σ 2 )t2 1 µt (µ2 + σ 2 )t2
⇒ lnMW (t) = − +n √ + − [ √ + + ···
σ σ n σ2n 2 σ n ]
√ √
nµt nµt (µ2 + σ 2 )t2 µ2 t2 µt3 (µ2 + σ 2 )
= − + + − − √ + ···
σ nσ σ2n 2σ 2 2 nσ 3
t2 µt3
= − 2
2 (µ + σ 2 )
t2
= since ∼ N (0, 1) and n → ∞
2
20 F-DISTRIBUTION
Consider two independent chi-square random variables U and V having r1 and r2 degrees of
U/r1
freedom respectively; Then F = V /r2
has an F- distribution with r1 and r2 degrees of freedom.
Its pdf is defined as folow
h(u,v)=
r1 r2 u+v
1 r1
−1 −u 1 r2
−1 − v2 u 2 −1 v 2
−1
e− 2 u/r1
r1 u 2 e2 . r2 v 2 e = r1 +r2 f=
Γ( r21 )2 2 Γ( r22 )2 2 Γ( r21 )Γ( r22 )2 2 v/r2
57
Distribution Theory II: STA 412 Dr. [Link]
zr1 f r1
zr1
|J| = r2 r2 =
0 1 r2
T hen,
r21 −1 h
f zr1
i
− 12
r2
h i
f zr1 −1 +z zr1
r2
z 2 e r2
r2
g(f, z) = r1 +r2
Γ( r21 )Γ( r22 )2 2
r21 r1 r1 +r2 z
h
f r1
i
r1 −1 −1 +1
r2
f 1 z 2 e 2 r2
= r1 +r2
Γ( r21 )Γ( r22 )2 2
Z ∞
g(f ) = g(f, z)dz
−∞
r21 h
f r1
i
r1 r1
−1
r1 +r2
−1 − z2 +1
Z ∞
r2
f 1 z 2 e r2
= r1 +r2 dz
Γ( r21 )Γ( r22 )2 2
0
z f r1 dy 1 f r1 2y 2dy
Let y = +1 ; = + 1 ; z = f r1 ; dz = f r1
2 r2 dx 2 r2 r
+ 1 r2
+1
2
By substitution we have
rr1 r1 +r
2
2 −1
r2
r1 −1 2y
Z ∞ r2
2
f 2
f r1 e−y . f r12
2
+1 2
+1
g(f ) = dy r1 +r2
0 Γ( r21 )Γ( r22 )2 2
rr1 r
r1 2 2 −1 r1 +r2 −1
r2
f 2 2 2 2 Z ∞
r1 +r2
−1 −y
g(f ) = r1 +r2
r +r
1 2 2 f r1 −1 f r1 y 2 e dy
r1 r2 f r
Γ( 2 )Γ( 2 )2 2 1
+ 1 + 1 + 1 0
2 2 2
rr1 r r1 +r2
2
r1
r2
2
f 2 −1 2 2 2−1 .2 r1 + r2
= r1 +r2 .Γ( )
r1 r2 r1 +r2
f r1
2 2
Γ( 2 )Γ( 2 )2 2
2
+1
rr1 r
2
r1
r2
2
f 2 −1 r1 + r2
= r1 +r2 .Γ( ); 0 < f < ∞
r1 r2 f r1
2 2
Γ( 2 )Γ( 2 ) 2 + 1
58
Distribution Theory II: STA 412 Dr. [Link]
2 3 1 3−x
3!
x!(3−x)!
3 3
, x = 0, 1, 2, 3, . . .
f (x) =
0 elsewhere.
y = x2 = u(x)
x = w(y)
√
= y
3! 2 1 √
g(y) = √ √ ( )3 ( )3− y , y = 0, 1, 4, 9, . . .
y!(3 − y)! 3 3
0 elsewhere.
Let X be a random variable of the continuous type having pdf f(x). Let the one-dimensional
space where f (x) > 0. Consider the random variable y = u(x), where y = u(x) defines a one-
to-one transformation that maps the set onto the set B.
Let the inverse ofy = u(x) be denoted by x = w(y) and let the derivative dx
dy
= w0 (y) be contin-
uous and equal to zero for all point y in B.
Then the pdf of the random variable Y = U (X)is given by g(y)f (w(y))|w0 (y)|, yB
0 elsewhwere.
Where dx
dy
= w0 (y) is the jacobian (denoted by J) of the transformation.
2x
0<x<1
Let f (x)
0, elsewhere.
Define the random variable of Y = 8X 3
59
Distribution Theory II: STA 412 Dr. [Link]
SOLUTION:
y = 8x3
y
x3 =
8
y 1
x = ( )3
8
1
y3
= 1
83
1 1
= y3
2
⇒ w(y) T hen
1 1
f (w(y)) = f ( y 3 )
2
1 1
= 2 y3
2
1
= y3
To obtain |J|
dx
|J| =
dy
1 −2
= y 3
6
g(y) = (f (w(y)))|J|
1 1 2
= y 3 . y− 3
6
1 −1
= y 3
6
T he Support is as f ollow.
1 1
0 < y3 < 1
2
1
0 < y3 < 2
0<y<8
0, elsewhere
60
Distribution Theory II: STA 412 Dr. [Link]
SOLUTION:
y = u(x) = −2lnx
y
− = lnx
2
y
e− 2 = x
y
x = w(y) = e− 2
dx 1 y 1 y
= − e− 2 = e− 2 = |J|
dy 2 2
f (w(y)) = 1
1 y
g(y) = f (w(y))|J| = e− 2
2
y
0 < e− 2 < 1
0 < y < ∞.
x3
f (x) = , 0 < x < 2.
4
61
Distribution Theory II: STA 412 Dr. [Link]
SOLUTION:
y = x2
= u(x)
√
x = y
= w(y)
√ 3
y
f (w(y)) =
4
3
y 2
=
4
dx 1 −1
= y 2
dy 2
g(y) = f (w(y))|J|
3
y 2 1 −1 1
= . y 2 = y;
4 2 8
√
0< y<2
0<y<4
22 JOINT TRANSFORMATION
This method of finding the pdf of a function of one random variable of the continuous type will
be extended to functions of two random variables of this type.
Let y1 − u1 (x1 , x2 ) and y2 − u2 (x1 , x2 ) define a one-to-one transformation that maps a (two-
dimensional) set A in the x1 , x2 , plane unto a two-dimensional set B in the y1 , y2 - plane . If
we express x1 , x2 in terms of y1 , y2 we can make x1 = w(y1 , y2 ). The determinant is of the order 2
dx1 dx1
dy1 dy2
dx2 dx2
dy1 dy2
62
Distribution Theory II: STA 412 Dr. [Link]
1 r x
f (x) = x 2 −1 e− 2 , 0 ≤ x < ∞,
Γ(r/2)
then X has a chi-square distribution with r degrees of freedom abbreviated by saying X is χ2(r) .
Find (i) the joint pdf of Y1 = X1 and Y2 = X2 + X1 for 0 < y1 < y2 < ∞
(ii) the marginal distribution of each y1 and y2
(iii) Are Y1 and Y2 independent?
SOLUTION:
1 − x1 1 x2
(1)f (x1 ) = e 2 , f (x2 ) = e− 2
2 2
1 − x1 +x2
f (x1 )f (x2 ) = e 2
4
y1 = u(x1 , x2 ) = x1 then x1 = w(y1 , y2 ) = y1
1 0
= =1
−1 1
1 − y2 1 y2
g(y1 , y2 ) = e 2 .1 = e− 2
4 4
(ii)
Z
g1 (y1 ) = g(y1 , y2 )dy2 , 0 < y1 < y2 < ∞
Z ∞
1 y2 2 y2 1 − y2
= e− 2 dy2 = − e− 2 |∞ y1 = e
2 ,0 < y < y
1 2
4 y1 4 2
Z
g2 (y2 ) = g(y1 , y2 )dy1 , 0 < y1 < y2 < ∞
Z y2
1 y2
= e− 2 dy1
4 0
63
Distribution Theory II: STA 412 Dr. [Link]
1 − y2 y2
= y1 e 2 |0
4
1 − y2
= y2 e 2
4
23 t-Distribution
W ∼ N (1, 0) i.e a standard normal distribution with mean 0 and variance 1
V ∼ χ2(r) is a chi-square distribution with d.f r
i.e.
1 −w2
W =√ e 2 −∞<w <∞
2π
r −v
V 2 −1 e 2
V = r 0<v<∞
Γ 2r 2 2
Let W and V be Independent; Then
W
T = pv
r
−w2 r −1 −v
√1 e 2 V e 2
2
r −∞<w <∞
2π Γ( r2 )2 2
0<v<∞
0 otherwise
To Proof
w
t = pv & u=v
r
√
t u
w = √
r
√
dw u
|J| = = √
dt v
64
Distribution Theory II: STA 412 Dr. [Link]
1 2z −z 2
1+ tr
g(t) = √ r 2 e dz
2πrΓ 2r 2 2 1 + tr
0
h i
(r+1)
Γ 2
= √ r+1 ∞<t<∞
2
πr Γ 2r 1 + tr 2
65
Distribution Theory II: STA 412 Dr. [Link]
Let A denote an n x n real symmetric matrix which is positive definite. Let µ denote the
n x 1 matrix such that µ0 , the transpose of µ, is µ0 =[µ1 , µ2 , . . . , µn ],where each µi is a real
[Link], let x denote the n x 1 matrix such that X 0 = [x1 , x2 , . . . , xn ].We shall show
that if C is an approximately chosen positive constant, the non-negative function
(X − µ)0 A(X − µ)
f (x1 , x2 , . . . , xn ) = C exp − ;
2
−∞ < x1 < ∞, i = 1, 2, . . . , n,
is a joint p.d.f of n random variable X1 , X2 , . . . , Xn that are of the continuous type show that
Z ∞ Z ∞
(1) ... f (x1 , x2 , . . . , xn )dx1 dx2 . . . dxn = 1
−∞ ∞
let t denote the n x 1 matrix such that t0 = [t1 , t2 , . . . , tn ], wheret1 , t2 , . . . , tn are arbitrary real
numbers. We shall evaluate the integral
Z ∞ ∞
(X − µ)0 A(X − µ)
Z
0
(2) C ... exp t x − dx1 , . . . dxn ,
−∞ ∞ 2
Because the real symmetric matrix A is positive definite, the n characteristics numbers (proper
values, latent roots, or eigenvalues)a1 , a2 , . . . , an of A are positive. There exists an approximately
chosen n x n real orthogonal matrix L(L0 = L−1 , where L−1 is the inverse of L) such that
a1 0 . . . 0
0 a2 . . . 0
L0 AL = .. ..
..
. . .
0 0 ... an
66
Distribution Theory II: STA 412 Dr. [Link]
Moreover, n
ai zi2
P
z 0 (L0 AL)z
i
exp − −
= exp
2 2
. Then the integral(4) may be written as the product of n integrals in the following manner:
n Z ∞
ai zi2
Y
0 0
5) C exp (W L µ) exp wi zi − dzi
i=1 −∞ 2
r Z ∞ exp w z − ai zi2
n
2π
Y i i 2
= C exp (W 0 L0 µ) p dzi .
i=1
a i −∞ 2π/a i
The integral that involves zi can be treated as the moment-generating function, with the more
familiar symbols t rel replaced by wi , of a distribution which is n(0,1/a).Thus the right-hand
member of equation(5) is equal to
n r
wi2
0 0
Y 2π
6) C exp (W L µ) exp
i=1
ai 2ai
67
Distribution Theory II: STA 412 Dr. [Link]
s !
n
(2π)n X wi2
= C exp (W 0 L0 µ) exp .
a1 a2 . . . an 1
2a
Now, because L−1 = L0 , we have
0 0 −1 1 1 1
(L AL) = L A L = diag , ...., .
a1 a2 an
n
wi2
= W 0 (L0 A−1 L)W = (LW ) A−1 (LW ) = t0 A−1 t. Moreover, the determinant
P
Thus, ai
1
A−1 of A−1 is
1
A−1 = L0 A−1 L = .
a1 a2 . . . an
Accordingly, the right-hand member of equation (6), which is equal to integral (2) may be written
as
−1
!
t0A t
q
t0
7) Ce µ (2π)n A−1 exp
2
If, in this function, we set t1 = t2 = · · · = tn = 0, we have the value of the left-hand members of
Equation (1). Thus, we
q
C (2π)n A−1 = 1
h 0
i
Accordingly, the function f (x1 , x2 , . . . , xn ) = 1
r exp − (X−µ) 2A(X−µ) ,
A−1
(2π)n/2
t0 A−1 t
0
M (t1 , t2 , . . . , tn ) = exp t µ + .
2
σi it2i
M (0, . . . , 0, ti , 0, . . . , 0) = exp ti µi +
2
68
Distribution Theory II: STA 412 Dr. [Link]
But this is the moment-generating function of bivariate normal distribution, so that aσij is the
covariance of the random variables X − i and Xj . Thus the matrix µ, where µ0 = [µi , µ2 . . . , µn ]
is the matrix of the means of the random variables X1 , . . . , Xn . Moreover, the elements on
the principal diagonal of A−1 are, respectively, the variances σii = σi2 , i = 1, 2, . . . , n, and the
elements not on the principal diagonal of A−1 are, respectively, the covariances σij = ρij σi , σj , i 6=
j, of the random variables X1 , X2 , . . . , Xn . We call the matrix A−1 ,which is given by
σ11 , σ12 , . . . , σ1n
σ12 , σ22 , . . . , σ2n
.. ,
.. ..
. . .
σ1n , σ2n . . . , σnn
the covariance matrix of the Multivariate normal distribution and henceforth we shall denote
this matrix by symbol V, In terms of the positive definite covariance matrix V, the multivariate
normal p.d.f. is written
(X − µ)0 V −1 (X − µ)
1
q exp − , −∞ < x1 < ∞,
(2π)n/2 V 2
t0 V t
0
exp t µ +
2
69
Distribution Theory II: STA 412 Dr. [Link]
Now the expectation (8) exists for all real values of [Link] we can replace t0 in expectation(8)
by tc0 and obtain
c0 V Ct2
0
M (t) = exp tc µ +
2
Exercies
1. Let X1 , X2 , . . . , Xn have a multivariate normal distribution with positive definite covariance
matrix V. prove that these random variable are mutually stochastically independent if and
only if V is a diagonal matrix.
70
Distribution Theory II: STA 412 Dr. [Link]
from in the Xi − µi and Q is seen to be, apart from the coefficient − 21 ,the random variable
which is defined by the exponent on the number e in the joint p.d.f. of X1 , X2 , . . . , Xn .We shall
now show that this result can be generalized.
(X − µ)0 V −1 (X − µ)
1
q exp − ,
(2π)n/2 V 2
Where, as usual, the covariance matrix V is positive definite. We shall show that the random
variable Q (a quadratic form in the Xi − µi ), which is defined by (X − µ)0 V −1 (X − µ), is χ2 (n).
We have for the moment-generating function M(t) of Q the integral
∞ ∞
(X − µ)0 V −1 (X − µ)
Z Z
1 0 −1
... q X exp t(X − µ) V (X − µ) − dx1 . . . dxn
−∞ −∞ (2π)n/2 V 2
∞ ∞
(X − µ)0 V −1 (X − µ)(1 − 2t)
Z Z
1
= ... q X exp − dx1 . . . dxn
−∞ −∞ (2π)n/2 V 2
With V −1 positive definite, the integral is seen to exist for all real values of t < 12 . Moreover,
(1 − 2t)n | V −1 |, it follows that
(X − µ)0 V −1 (X − µ)(1 − 2t)
1
p exp −
(2π)n/2 | V | /(1 − 2t)n 2
The remarkable fact that the random variable which is defined by (X − µ)0 (X − µ)0 is (χ2 n)
stimulates a number of questions about quadratic forms in normally distributed variables. We
71
Distribution Theory II: STA 412 Dr. [Link]
should like to treat this problem in complete generally, but limitations of space forbid this, and
we find it necessary to restrict ourselves to some special cases.
Let X1 , X2 , . . . , Xn denote a random sample of size n from a distribution which is n(0, σ 2 )
,σ 2 > 0 . Let X 0 = [X1 , X2 , . . . , Xn ] and let A denote an arbitrary n X n real symmetric matrix.
We shall investigate the distribution of the quadratic from X 0 AX. For instance, we know that
n
X 0 In X/σ 2 = X 0 X/σ 2 = Xi2 /σ 2 is χ2 (n) .First we shall find the moment-generating function
P
1
of X 0 AX/σ 2 , Then we shall investigate the conditions which must be imposed upon the real
symmetric matrix A if X 0 AX/σ 2 is to have a chi-square [Link] moment-generating
function is given by
∞ ∞n 0
X 0X
Z Z
1 tX AX
M (t) = ... √ exp − dx1 , . . . , dxn
−∞ −∞ σ 2π σ2 2σ 2
Z ∞ Z ∞ n
X 0 (I − 2tA)X
1
= ... √ exp − dx1 , . . . , dxn ,
−∞ −∞ σ 2π 2σ 2
Where I = In The matrix I-2tA is positive definite if we take | t | sufficiently small say
| t |< h, h > 0. Moreover, we can treat
X 0 (I − 2tA)X
1
p exp −
(2π)n/2 | (I − 2tA)−1 σ 2 | 2σ 2
72
Distribution Theory II: STA 412 Dr. [Link]
(1 − 2ta1 ) 0 ... 0
0 (1 − 2ta2 ) . . . 0
L0 (I − 2tA)L =
.. .. ..
. . .
0 0 ... (1 − 2tan )
Then
n
Y
(1 − 2tai ) =| L0 (I − 2tA)L | = | I − 2tA |
i=1
let r, 0 < r ≤ n, denote the rank of the real symmetric matrix [Link] exactly r of the
real numbers a1 , a2 . . . , an , say a1 , a2 . . . , ar are not zero and exactly n-r of these numbers, say
ar+1 , . . . , an , are zero. Thus we can write thr moment-generating function of X 0 AX/σ 2 as
Now that we have found, in suitable form, the moment-generating function variable, let us
turn to the equation of the conditions that must be imposed if X 0 AX/σ 2 is to have a chi-square
distribution. Assume that X 0 AX/σ 2 is χ2 (k). Then
or, equivalently .
Because the positive integers r and h are the degree of these polynomials, and because these
polynomials are equal for infinitely many values of t, we have K = r,the rank of [Link],
the uniqueness of the factorization of a polynomial implies that a1 = a2 = · · · = ar = 1. If each
non-zero characteristics numbers of a real symmetric matrix is one, the matrix is idempotent,
73
Distribution Theory II: STA 412 Dr. [Link]
that is, A2 = A, and conversely. Accordingly if X 0 AX/σ 2 has a chi-square distribution , then
A2 = A and the random variable is χ2 (r) , where r is the rank of A. Conversely, if A has exactly
r characteristics numbers are equal to one, and the remaining n - r characteristics numbers that
are equal to zero. Thus the moment-generating function of X 0 AX/σ 2 is given by (1 − 2t)−r/2 ,
t < 12 , and X 0 AX/σ 2 is χ2 (r).this established the following theorem.
Theorem 1
. Let Q denote a random variable which is a quadratic form in the items of a random sample
of size n from a distribution which is n(0, σ 2 ). Let A denote the symmetric matrix of Q and let
r, 0 < r ≤ n, denote the rank of A. Then Q/σ 2 is χ2 (r) if and only ifA2 = A.
, the condition A2 = A and aij = 0 are necessary and sufficient condition that Q/σ 2 be
P
i.j
central χ2 (r). Moreover,the theorem may be extended to a quadratic from in random variables
which have a multivariate normal distribution with positive definite covariance matrix V;here
the necessary and sufficient condition that Q have chi-square distribution is AVA = A.
Exercises
1. Let Q = X1 X2 − X3 X4 , where X1 X2 , X3 X4 , is a random sample of size 4 from a distri-
bution which is n(0,σ 2 ).Show that Q/σ 2 does not have a chi-square distribution. Find the
moment-generating function of Q/σ 2 .
2. Let A be a symmetric matrix. prove that each of the non-zero characteristic number of A
is equal to one if and only if A2 = A. Hint. LetL be an orthogonal matrix such that L0 AL
74
Distribution Theory II: STA 412 Dr. [Link]
is idempotent.
n Z ∞ Z ∞
t1 X 0 AX t1 X 0 BX X 0X
1
M (t1 , t2 ) = √ ... exp + − dx1 . . . dxn
σ 2π −∞ −∞ σ2 σ2 2σ 2
n Z ∞ Z ∞
X 0 (I − 2t1 A − 2t2 B)X
1
= √ ... exp − dx1 . . . dxn
σ 2π −∞ −∞ 2σ 2
The matrix I-2t1 A − 2t2 B is positive definite if be take | t1 | and | t2 | sufficiently small, say
| t1 | < h1 | t2 | < h2 where h1 , h2 > 0. Then, we have
75
Distribution Theory II: STA 412 Dr. [Link]
Let us assume that X 0 AX/σ 2 and X 0 BX/σ 2 are stochastically independent (so that likewise
areX 0 AX and X 0 BX) and prove that AB=0. Thus we assume that
Let r> 0 denote the rank of A and let a1 , a2 . . . , ar denote the r nonzero characteristic
number of A. There exists an orthogonal matrix L such that
..
a1 0 . . . 0 ..
.
0 a2 . . . 0
. 0
..
. .. .. .. C 11 . 0
..
0 . . .
L AL = = . . . . . . . . . = C
0 .
. ..
0 . . . a r .
0 .
. . . . . . . . . . . . . . . . . .
0 0
76
Distribution Theory II: STA 412 Dr. [Link]
of the minor of order r in the upper left-hand corner, namely , | I − 2t1 C1 1 − 2t2 D11 |, and
the minor of order n-r in the lower right-hand corner, namely, | In−r − 2t2 D22 |.Moreover,
this product is the only term in the expansion of the determinant that involves (−2t1 )r .
Thus the coefficient (−2t1 )r in the left-hand members of Equation(3) is a a1 a2 . . . ar |
I − 2t1 C1 1 − 2t2 D11 |. If we equate these coefficients of (−2t1 )r , we have, for all t2 , | t2 |
< h2
Since AB = 0. Thus ,
Since the moment-generating function of the joint distribution X 0 AX/σ 2 and X 0 BX/σ 2 is
given by:
M (t1 , t2 ) =| I − 2t1 A − 2t2 B |−1/2 , | t1 |< hi i = 1, 2,
We have :
M (t1 , t2 ) = M (t1 , 0)M (0, t2 )
77
Distribution Theory II: STA 412 Dr. [Link]
Theorem 2.
Let Q1 and Q1 denote random variables which are quadratic forms in the items of a random
sample of size n from a distribution which is n(0,σ 2 ).Let A and B denote respectively the
real symmetric matrices of Q1 and Q2 . The random variables Q1 and Q2 are stochastically
independent if and only if AB = 0.
Remark
Theorem 2 remain valid if the random samples is from a distribution which is n(µ, σ 2 ),
whatever be the real value of µ. Moreover, Theorem 2 may be extended to quadratic forms in
a random variables that have a joint multivariate normal distribution with a positive definite
covariance matrixV. The necessary and sufficient condition for the stochastic independence
of two such quadratic forms with symmetric matrices A and B then becomes AVB = 0.
In our Theorem 2, we have V = σ 2 I, so that AVB = Aσ 2 IB = σ 2 AB = 0
Theorem 3
proof : Take first the case of k=2 and let the real symmetric matrices of Q, Q1 , and Q2 be
denoted, respectively, by A, A1 , A2 . We are given that Q = Q1 + Q2 or, equivalently, that
A = A1 + A2 . We are also given that Q/σ 2 is χ2 and that Q1 /σ 2 is χ2 (ri ). In accordance
with Theorem 1, we have A2 = A and A2i . since Q2 ≥ 0,each of the matrices A, Ai andA2
78
Distribution Theory II: STA 412 Dr. [Link]
(5)
Ir 0 Gr 0 Hr 0
= +
0 0 0 0 0 0
79
Distribution Theory II: STA 412 Dr. [Link]
Remark
In our statement of Theorem 3 we took X1 , X2 . . . , Xn to be items of a random sample from
distribution which is n(0, σ 2 ).We did this because our proof of Theorem 2 was restricted to that
case. In fact, if Q0 , Q01 . . . Q0k are quadratic forms in any normal variables (including multivariate
normal variables), if Q0 = Q01 + · · · + Q0k if Q0 , Q01 . . . Q0k−1 are central or non central chi-square,
and if Q0K is nonnegative, then Q01 . . . Q0k are mutually stochastically independent and Q0k is
either central or non central chi-square.
Theorem 4
Let X1 , X2 . . . , Xn denote a random sample from a distribution which is n(0, σ 2 ).Let the sum of
the squares of these items be written in the form
n
X
Xi2 = Q1 + Q2 + · · · + Qk ,
i
k k k
Xi2 =
P P P
proof: First assume the two conditions rj = n and Qj to be satisfied. The
1 1 1
latter equation implies that I = A1 + A2 + · · · + Ak . Let Bi = I − Ai . That is, Bi is the sum
of the sum of the matrices A1 , . . . , Ak exclusive of Ai . Let Ri denote the rank of Bi . Since
the rank of the sum several matrices is less than or equal to the sum of the ranks, we have
Pk
Ri ≤ rj − ri = n − ri . However, I = Ai + Bi , so that n ≤ ri + Ri and n − ri ≤ Ri . Hence
1
Ri = n − ri . The characteristics numbers of Bi are the roots of the equation | Bi − λI | = 0.
Since Bi = I - Ai , this equation can be written as | Bi − λI | = 0. Thus, we have | Ai − (1 − λ)I |
= 0. But each root of the last equation is one minus a characteristic number of Ai . since Bi
has exactly n − Ri = ri characteristic numbers that are zero, then Ai has exactly r − i charac-
teristics numbers that are equal to one. However, ri is the rank of Ai . Thus, each of ri nonzero
80
Distribution Theory II: STA 412 Dr. [Link]
n
Xi2 = Qi + Q2 + . . . , +Qk let Qi , Q2 . . . , Qk be
P
To complete the proof of Theorem 4, take
1
k
mutually stochastically independent, and let Qj /σ 2 be χ2 (ri ), j = 1, 2, . . . , k. Then Qj /σ 2 is
P
n 1
k n n
2
Qj /σ 2 = Xi2 /σ 2 is χ2 (n). Thus, rj = n and the proof is complete.
P P P P
χ rj . But
1 1 1 1
Exercise
(1) Let X1 , X2 , . . . , Xn denote a random sample of size n from a distribution which is n(0, σ 2 ).
n
prove that Xi2 and every quadratic form, which is non identically zero in X1 , X2 , . . . , Xn
P
1
are stochastically dependent.
(2) X1 , X2 , X3 , X4 denote a random sample of size 4 from a distribution which is n(0, σ 2 ). Let
4
ai Xi , where a1 , a2 , a3 and a4 are real constants. If Y 2 and Q = X1 X2 − X3 X4 are
P
Y=
1
stochastically independent, determine a1 , a2 , a3 and a4
Limit Theorem
27 Introduction
This chapter is principally concerned with the limiting behavior of the sum of independent
random variable as the number of summands becomes large. The results presented here are
both intrinsically interested and useful in statistics, since many commonly computed statistical
quantities, such as average, can be represented as sums.
81
Distribution Theory II: STA 412 Dr. [Link]
n
1X
X̄n = Xi
2 i=1
1
The law of large numbers sates that X̄n approaches 2
in a sense that is specified by the following
theorem.
82
Distribution Theory II: STA 412 Dr. [Link]
n
1X
E(X̄n ) = E(Xi ) = µ
n i=1
Since the Xi are independent ,
n
¯ 1 X σ2
V ar(Xn ) = 2 V ar(Xi ) =
n i=1 n
The desired result now follows immediately from Chebyshev’s inequality, which sates that
V ar(X¯n ) σ2
P ( X̄n − µ > ) ≤ 2
= n2
−→ 0, as n −→ ∞
In the case of a fair coin toss, the Xi are Bernoulli random variables with p=1/2, E(Xi ) = 1/2
and Var(Xi )=1/4. if tossed 10,000 times
¯ ) = 2.5 × 10−5
V ar(X10,000
and the standard deviation of the average is the square root of the variance, 0.005. The
proportion observed by Kerrich, 0.5067, is thus a little more than one standard deviation away
from its expected value of 0.5, consistent with Chebyshev’s inequality. Chebyshev’s inequality
can be written in the formP ( X̄n − µ > kσ) ≤ 1/k 2
if a sequence of random variable P (|Zn − α|) > ) approaches zero as n approaches infinity,
for any > 0 and where α is some scale, then Zn is said to converge in probability to α.
There is another mode of convergence , called strong convergence or almost sure convergence,
83
Distribution Theory II: STA 412 Dr. [Link]
which asserts more Zn is said to converge almost surely to α if for every > 0,,|Zn − α| >
only a finite number of times with probability 1; that is, beyond some point is random. The
version of the law ot large numbers stated and proved earlier asserts that X̄n converges to µ
in [Link] version is usually called the weak law of large numbers. Under the same
assumptions, a strong law of large numbers, which assert that X̄n converges almost surely to µ,
can also be proved, bt we not do so.
We now consider some example that illustrate the utility of the law of large numbers.
EXAMPLE I
Monte Carlo Integration Suppose that we wish to calculate
Z 1
I(f ) = f (x)dx
0
where the integration cannot be done by elementary means or evaluated using tables of integrals.
The most common approach is to use a numerical method in which the integral is approximated
by a sum; Various schemes and computer packages exist for doing this. Another method, called
the Monte Carlo method, works in the following way. Generate independent uniform random
variables on [0,1]- that is, X1 , X2 , . . . , Xn - and compute
n
ˆ )= 1
X
I(f f (Xi )
n i=1
By the law of large numbers,this should be close to E[f, (Xi )], which is simply
Z 1
E[f (X)] = f (x)dx = I(f )
0
This simple scheme can be easily modified in order to change the range of integration and in
other ways. Compare to the standard numerical methods, it is not especially efficient in one
dimension, but becomes increasingly efficient as the dimensionality of the integral grows.
As a concrete example, let us consider the evaluation of
Z 1
a x2
I(f ) = √ e− 2 dx
2π 0
The integration is that of the standard normal density, which cannot be evaluation in closed form.
From the table of the normal distribution ,an accurate numerical approximation is I(f)=.3413.
84
Distribution Theory II: STA 412 Dr. [Link]
If 1000 points X1 , . . . , X1 000, uniformly distributed over the interval 0 ≤ x ≤ 1, are generated
using a pseudorandom number generator, the integral is then integrated by
1000
X
ˆ )= 1 1 2
I(f √ e−Xi /2
1000 2π i=1
85
Distribution Theory II: STA 412 Dr. [Link]
σ2
V arX̄ =
n
n
can be estimated from the data assess the precision of X̄, First, note that n−1 Xi2 converges
P
i=1
to E(X 2 ), from the law of large number. Second, it can be shown that if Zn converges to α in
probability and g is a continuos function, then
g(Zn ) −→ g(x)
X̄ 2 −→ [E(X)]2
n
Finally, since n−1 Xi2 converges to E(X 2 ) and X̄ 2 converges to [E(X)]2 , with a little additional
P
i=1
argument it can be shown that
n
1X 2
Xi − X̄ 2 −→ E(X 2 ) − [E(X)]2 = V ar(X)
n i=1
n
More generally, it follows from the law of large numbers that the sample moments n−1 Xir ,
P
i=1
converge in probability to the moment of X,E(X r )
EXAMPLE III
A muscle or nerve cell membrane contains a very large number of channels;when open, these
channels allow ions to pass through. Individual channels seem to open and close randomly, and
it is often assumed that in an equilibrium situations the channels open and close independently
of each other and that only a very small fraction are open at any one time. Suppose then that
the probability of a channel is open is p, a very small number, that there are m channels in
86
Distribution Theory II: STA 412 Dr. [Link]
all, and that the amountof current flowing through and individual channel is c. The number of
channel open at a particular time is N, a binomial random variable with m trials and probability
p of success on each trial. The total amount of current is S = cN and can be measured. We
then have
V ar(S) = c2 mp(1 − p)
and
V ar(S)
= c(1 − p) ≈ c
E(S)
since p is small, Thus, through independent measurements,S1 , . . . , Sn . we can estimate E(S) and
Var(S) and therefore c, the amount of current flowing through a single channel, without knowing
how many channels are there.
DEFINITION
87
Distribution Theory II: STA 412 Dr. [Link]
EXAMPLE 1
We will show that the poisson distribution can be approximated by the normal distribution
for large values of λ. Which shows that as λ increases, the probability mass function of the
poisson distribution becomes more symmetric and more shaped
Let λ1 , λ2 , . . . be an increasing sequence with λn −→ ∞, and let {Xn } be a sequence of poisson
random variables with the corresponding parameters. We know that E(Xn ) = V ar(Xn ) = λn .
if we wish to approximate the poisson distribution function by a normal distribution function ,
the normal must have the same mean and variance as the poisson does. In addition, if we wish
to approve a limiting result, we run into the difficulty that the mean and variance are tending to
infinity. This difficulty is dealt with by Standardizing the random variable-that is by letting
Xn − E(Xn )
Zn = p
V ar(Xn )
Xn λn
= √
λn
We then have E(Zn ) = 0 and V ar(Zn ) = 1, and we will show that the mgf of Zn converges to
the mgf of the standard normal distribution.
88
Distribution Theory II: STA 412 Dr. [Link]
The mgf of Xn is
et −1
MXn (t) = eλn
the mgf of Zn is
√
−t λn t
MZn (t) = e MXn √
λn
√ √
λn λn(et / λn−1)
= e−t e
t2
lim logMZn (t) =
n−→∞ 2
or
2 /2
lim MZn (t) = et
n−→∞
89
Distribution Theory II: STA 412 Dr. [Link]
Let X be a poisson random variable with mean 900. We find P(X¿950) by standardizing:
X − 900 950 − 900
P (X > 950) = P √ > √
900 900
5
≈ 1−φ
3
= .04779
Where φ is the standard normal cdf. For comparison, the exact probability is .04712
We now turn to the central limit theorem, which is concern with a limiting property of sums
of random variables. If X1 , X2 , . . . is a sequence of independent random variable with mean µ
and variance σ 2 , and if
n
X
Sn = Xi
i=1
We know from the law of large numbers that Sn /n converges to µ in probability. This followed
from the fact that
σ2
Sn 1
V ar = 2 V ar(Sn ) = −→ 0
n n n
The central limit theorem is concerned not with the fact that the ratio Sn /n converges yo µ
but with how it fluctuates around µ. To analyze these fluctuations, we standardize:
Sn − nµ
Zn = √
σ n
You should verify that Zn has mean 0 and variance 1. The central limit theorem states that
the distribution of Zn converges to the standard normal distribution
THEOREM II Central Limit Theorem
Let X1 X2 , . . . be a sequence of independent random variable having mean 0 and variance
σ 2 and the common distribution function F and moment generating function M define in a
neighborhood of zero. let
n
X
Sn = Xi
i=1
Then
Sn
lim P √ ≤x = φ(x). −∞<x<∞
n−→∞ σ n
90
Distribution Theory II: STA 412 Dr. [Link]
Proof
√
Let Zn = Sn /(σ n) we will show that the mgf of Zn tend to the mgf of standard normal
distribution. Since Sn is a sum of independent random variables.
and
n
t
MZn (t) = M √
σ n
M(s)has a Taylor series expansion about zero :
1
M (s) = M (O) + sM 0 (0) + s2 M 00 (0) + s
2
0 √
Where s /s2 −→ [Link] E(X)=0, M (O) = 0, and M 00 (0) = σ 2 . As n −→ ∞, t/(σ n) −→ o,
and 2
t 1 t
M √ = 1 + σ2 √ + n
σ n 2 σ n
wheren /(t2 /(nσ 2 )) −→ 0 as n −→ ∞. we thus have
n
t2
MZn (t) = 1 + + n
2n
2 /2
MZn (t) −→ et as n −→ ∞
91
Distribution Theory II: STA 412 Dr. [Link]
By using characteristic functions instead, we could modify the proof so that it would only be
necessary that the first and second moment exist. Further generalization weaken the assumption
that the Xi have the same distribution and apply to linear combinations of independent random
variables. Central limit theorem have also been proved that weaken the independent assump-
tion and allow the Xi to be dependent but not to dependent. For practical purpose,especially
in statistics, the limiting result in it self is not of primary interest. Statisticians are more inter-
ested in its uses and approximation with finite value of [Link] is impossible to give a concise and
definitive statement of how good the approximation is, but some general guideline are available,
and examining special case can give in sight. How far the approximation becomes good depends
on the distribution of the summands, the Xi .If the distribution is fairly symmetric and has tails
that die off rapidly, the approximation becomes good and relatively small values of n. If the
distribution is very skewed or if the tail die down very slowly, a larger value of n is needed for a
good approximation. The following example deals with two special cases.
EXAMPLE
Suppose that X1 , . . . , Xn are repeated, independent measurement of quantity, µ, and that
E(Xi ) = µ and V ar(Xi ) = σ 2 . The law of large numbers tells us that X̄ converge to µ in
probability, so we can hope X̄ is close to µ if n is large. Chebyshev’s inequality allows us to
bound the probability of an error of a given size, but the central limit theorem gives a more
sharper approximation to the actual error. Suppose that we wish to find P (|X ¯− µ| < c) for
some constant c. To use the central limit theorem to approximate this probability, we first
standardize, using E(X) = µ and V arX̄ = σ 2 n
For example, suppose that 16 measurements are taken with σ = 1. The probability that the
92
Distribution Theory II: STA 412 Dr. [Link]
This can be turn around. That is, given c and γ, n can be found such that
P (|X̄ − µ| < c) ≥ γ
= 0.0228
The probability is rather small, so the fairness of the coin is called into question.
93
Distribution Theory II: STA 412 Dr. [Link]
The distribution of size of grains of a particulate matter is found to be quite skewed, with
a slowly decreasing right tail. A distribution called the lognormal is sometimes fit to such
distribution, and X is said to follow a lognormal distribution if log X as a normal distribution.
The central limit theorem gives a theoretical rationale for the use of lognormal distribution in
some situations.
Suppose that a particle of initial y0 is subjected to repeated impact, that on each impact a
proportion, Xi , of the particle remains, and that the Xi are modeled as independent random
variable having the same distribution. After the first impact, the size of the particle is Y1 =
X1 y − 0; after the second impact, the size is Y2 = X2 X1 y0 ; and after the nth impact, the size is
Yn = Xn Xn−1 . . . X2 X1 y0
T hen
n
X
LogYn = Logy0 + logXi
i=1
A similar construction is relevant to the theory of finance. Suppose that an initial investment
of value v0 is made and that returns occur in discrete time, for example, daily. If the return on
the first day is R1 , then the value becomes V1 = R1 v0 . After day two the value is V2 = R2 R1 v0 ,
and after day n the value is
Vn = Rn Rn−1 . . . R1 v0
T he log value is
n
X
LogVn = logv0 + logRi
i=1
If the return are independent random variable with the same distribution, then the distribution
of log Vn is approximately normally distribution.
94
Distribution Theory II: STA 412 Dr. [Link]
DWFINITION
If ∪1 , ∪2 , . . . , ∪n are independent chi-square variables with 1 degree of freedom, the distribution
of V = ∪1 + ∪2 + · · · + ∪n is called the chi-square distribution with n degree of freedom and is
denoted by χ2n
We know that the sum of independent gamma random variables that have the same value of λ
follows a gamma distribution, and therefore the chi-square distribution with n degrees of freedom
1
is a gamma distribution with α = n and λ = 2
Its density is
1
f (v) = v (n/2)−1 e−v/2 , v ≤ 0
2n/2 Γ(n/2)
Its moment generating function is
M (t) = (1 − 2t)−n/2
Also, E(V ) = n and V ar(V ) = 2n . To indicate that V follows a chi-square distribution with
n degrees of freedom , we write V ∼ χ2n . A notable consequence of the definition of chi-square
95
Distribution Theory II: STA 412 Dr. [Link]
distribution is that if U and V are independent and U ∼ χ2n and V ∼ χ2m then U + V ∼ χ2m+n
we now turn to the t distribution.
DEFINITION
p
If Z ∼ N (0, 1) and ∪ ∼ χ2n and Z and ∪ are dependent, then the distribution of Z/ U/n is
called the t distribution with n degrees of freedom.
PROPOSITION I
The density function of a t distribution with n degrees of freedom is
−(n+1)/2
t2
Γ[(n + 1)/2]
f (t) = √ 1+
nπΓ(n/2) n
Proof
p
This is proved by a standard method. The density function of U/n is straight forward to ob-
tain, and the density function of the quotient of two independent random variable was derived
From the density of the proposition f (t) = f (−t), so the t distribution is symmetric about zero.
As the number of degrees of freedom approaches infinity , the t distribution tends to the stan-
dard normal distribution; in fact for more than 20 and 30 degrees of freedom, the distribution
are very close
DEFINITION
Let U and V be independent chi-square random variable with m and n degrees of freedom,
respectively. The distribution of
U/m
W =
V /n
PROPOSITION II
The density function of W is given by
Γ[(m + n)2] m m/2 m/2−1 m −(m+n)/2
f (w) = w 1+ w , w≤0
Γ(m/2)Γ(n, 2) n n
96
Distribution Theory II: STA 412 Dr. [Link]
Proof
W is the ratio of two independent variables, and its density
It can be shown that, for n¿2, E(W) exist and equal n/(n − 2) from the definition of the t and
F distributions, it follows that the square of a tn random variable follows an F1.n distribution .
These are called sample mean and sample variance,respectively. First note that because X̄ is a
linear combination of independent normal random variables, it is normally distributed with
E(X̄) = µ
σ2
V ar(X̄) =
n
THEOREM
The random variable X̄ and the vector of random variables
(X1 − X̄, X2 − X̄, . . . , Xn − X̄) are independent.
Proof
At the level of this course, it is difficult to give a proof that provide sufficient insight into why
this result is true;a rigorous proof essentially depends on geometric properties of the multivariate
97
Distribution Theory II: STA 412 Dr. [Link]
Factors into the product of two moment generating functions-one of the X̄ and the other of
(X1 − X̄), . . . (Xn − X̄). the factoring implies that the random variables are independent of each
other and is accomplish through some algebraic trickery. First we observe that since
n
X n
X
ti (Xi − X̄) = ti Xi − nX̄ t̄
i=1 i=1
then
n n h
X X s i
sX̄ + ti (Xi − X̄) = + (ti − t̄)
i=1 i=1
n
Xn
= ai Xi where
i=1
s
a1 = + (ti − t̄)
n
f urther we observe that
X n
ai = s
i=1
n n
X s2 X
a2i = + (ti − t̄)2
i=1
n i=1
now we have
98
Distribution Theory II: STA 412 Dr. [Link]
The first factor is the mgf of X̄. Since the mgf of the vector
(X1 − X̄, . . . , Xn − X̄) can be obtain by setting s=0 in M, the second factor is this mgf.
COROLLARY
X̄ and S 2 are independently distributed
Proof
This follow immediately since S 2 is a function of vector
(X − 1 − X̄, . . . Xn − X̄), which is dependent of X̄
The next theorem gives the marginal function of S 2
THEOREM
The distribution of(n − 1)S 2 /σ 2 is the chi-square distribution with n-1 degrees of freedom.
Proof
We first note that
99
Distribution Theory II: STA 412 Dr. [Link]
n 2
1 X 2 Xi − µ
(Xi − µ) = ∼ χ2n
σ 2 i=1 σ2
Also
n n
1 X
2 1 X
(Xi − µ) = [(Xi − X̄) + (X̄ − µ)]2
σ 2 i=1 σ 2 i=1
n
P
Expanding the square and using the fact that (Xi − X̄) = 0, we obtain
i=1
n n 2
1 X 2 1 X 2 X̄ − µ
(Xi − µ) = 2 (Xi − X̄) + √
σ 2 i=1 σ i=1 σ/ n
This is a relation of the form W=U+V. Since U and V are independent by corollary,Mw (t) =
Mu (t)Mv (t). W and V both follow chi-square distributions,so
Mw (t)
Mu (t) =
Mv (t)
(1 − 2t)−n/2
=
1 − 2t)−1/2
= 1 − 2t)−(n−1)/2
The last expression is the mgf of a random variable with a χ2n−1 distribution.
COROLLARY
Let X̄ and S 2 be given as the beginning of the first corollary. Then
X̄ − µ
√ ∼ tn − 1
S/ n
Proof
We simply express the given ratio in a different form:
X̄−µ
√
X̄ − µ σ/ n
√ =p
S/ n S 2 /σ 2
The later is the ratio of an N(0,1)random variable to the square root of an independent random
variable with a χ2n−1 distribution divided by its degrees of [Link] from the definition,the
ratio follows a t distribution with n-1 degrees of freedom.
100