0% found this document useful (0 votes)
3 views100 pages

Note

The document covers various concepts in distribution theory, including probability density functions (pdf), cumulative distribution functions (cdf), and mathematical expectations. It provides examples and solutions for finding constants in pdfs, calculating cdfs for discrete and continuous random variables, and demonstrates properties of expected values. The document is structured as a lecture by Dr. K.O. Obisesan for a course titled Distribution Theory II.

Uploaded by

faruqtaiwo80
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views100 pages

Note

The document covers various concepts in distribution theory, including probability density functions (pdf), cumulative distribution functions (cdf), and mathematical expectations. It provides examples and solutions for finding constants in pdfs, calculating cdfs for discrete and continuous random variables, and demonstrates properties of expected values. The document is structured as a lecture by Dr. K.O. Obisesan for a course titled Distribution Theory II.

Uploaded by

faruqtaiwo80
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Distribution Theory II: STA 412 Dr. K.O.

Obisesan

DISTRIBUTION THEORY II
1 Basic Identities
n
X n(n + 1)
r =
r=1
2
n
X n(n + 1)(2n + 1)
r2 =
r=1
6
n  2
X n(n + 1)
r3 =
r=1
2
n
X
4 n(n + 1)(2n + 1)(3n2 + 3n + 1)
r =
r=1
30
n
X n  X n 
(1 + t)n = tk
i.e =1
k k
k=0
n  
X n
ak bn−k = (a + b)n
k
k=0
X n 
n
(1 − t) = (−1)k tk
k
n  
n
X n
2 =
k
k=0
n      
X a b a+b
=
k n−k n
k=0

Question: Given a pdf as below, find the value of C



−3
C(1 + x)
 x>0
f (x) =

0 x≤0

1
Distribution Theory II: STA 412 Dr. [Link]

Solution:

f (x) = C [1 + x]−3 x>0


Z ∞
But f (x)dx = 1
0
Z ∞
⇒ C[1 + x]−3 dx = 1
0
Z ∞
C [1 + x]−3 dx = 1
0
dy
Let y = 1 + x : = 1 : dy = dx
dx
T hus y > 1 + 0 ; y > 1(Range of y)
Z ∞
⇒C y −3 dy = 1
1
 −2 ∞
y
C = 1
−2 1
−C −2 ∞

y 1 = 1
2
−C  2
∞ − 1−2

= 1
2
−C
[0 − 1)] = 1
2
C
= 1
2
C = 2


−3
2(1 + x)
 x>0
Hence; f (x) =

0 x≤0

DOC: CDF This is denoted by F(x) and it is the probability that X takes on values less
than or equal to x i.e. F(x)=P(X ≤ x)
Discrete: In case of discrete random variable, it is clear that
X
F (c) = f (x)
x≤C

2
Distribution Theory II: STA 412 Dr. [Link]

P
Where F (c) = f (x) means sum the values of f (x) for all values of x less than or equal to C.
x≤C
Relationship Between pdf and cdf i.e. f(x) and F(x)

(i) If X is a discrete random variable of interest and the possible values of X are in increasing
order, x1 < x2 < x3 < . . .. Then f (x) < F (x).

(ii) For any i > 1 F (xi ) > F (xi−1 )

(iii) If x < x1 then F (x) = 0


P
(iv) F (x) = x≤x (f xi )

Continuous

Z ∞
F (x) = P (x ≤ x) = f (t)dt
−∞
F 0 (x) = f (x)

Question: Given a pdf f (x) = 4xc 0≤x<1

(i) Obtain the valu of c and

(ii) Find the distribution function F (x) = P (x ≤ x)

3
Distribution Theory II: STA 412 Dr. [Link]

Solution:

(i)

f (x) = 4xc
Z 1
f (x)dx = 1
0
Z 1
4xc = 1
Z0 1
4 xc = 1
0
 c+1 1
x
4 = 1
c+1 0
 
1 0
4 − = 1
c+1 c+1
4
= 1
c+1
c+1 = 4

c = 4−1

c = 3

(ii)
Z x
F (x) = P (X ≤ x) = f (t)dt
−∞

Z x Z x
c
4t dt = 4 tc dt c=s
0 0

Z x  4 x
3 t  x
→ 4 t dt = 4 = t4 0 = x4
0 4 0

F (x) = x4



 0; x≤0





T hus F (x) = x4 ; 0≤x≤1



1;4



1≥x
Distribution Theory II: STA 412 Dr. [Link]

1 3 x
 
QUESTION: Given f (x) = 4 4
, x = 0, 1, 2, · · · , x > 0 find the distribution
function F (x) of x Solution:

F (x) = P (X ≤ x)
Z x Z x    t
1 3
= f (t)dt = dt
−∞ 0 4 4
Z  t
1 x 3
= dt
4 0 4
" t+1 #x
3
1 4
=
4 t+1
" t 0#x
3 3
1 4 4
=
4 t+1
0
  " 3 t #x
3 1 4
=
4 4 t+1
0
  " 3 x #
3 1 4 1
= −
4 4 x+1 0+1
" x #
3
3 4
= −1
16 x + 1


 0 x<0





  3 x 
3 (4)
F (x) = 16 x+1
−1 0≤x<∞







∞≤x

1
x
QUESTION: Given f(x) = 15
, x = 1, 2, 3, 4, 5, Find F(x) of x

5
Distribution Theory II: STA 412 Dr. [Link]

SOLUTION:

x
f (x) =
15
P (x ≤ 0) = f (0) = 0
1
P (x ≤ 1) = f (1) =
15
1 2 3
P (x ≤ 2) = f (1) + f (2) = + =
15 15 15
1 2 3 6
P (x ≤ 3) = f (1) + f (2) + f (3) = + + =
15 15 15 15
1 2 3 4 10
P (x ≤ 4) = f (1) + f (2) + f (3) + f (4) = + + + =
15 15 15 15 15
1 2 3 4 5 15
P (x ≤ 4) = f (1) + f (2) + f (3) + f (4) + f (5) = + + + + = =1
15 15 15 15 15 15
Hence;


 0; −∞ < x < 1





1




 15
1≤x<2





3
2≤x<3


 15

F (x) =
 36
3≤x<4




 15




10




 15
4≤x<5





1 5≤x<∞

QUESTION: Alternative solution to question 2

   x
1 3
f (x) = , x = 0, 1, 2, · · ·
4 4
The sequence of f(x) implies
         2    3
1 1 3 1 3 1 3
, , ,···
4 4 4 4 4 4 4

6
Distribution Theory II: STA 412 Dr. [Link]

Since the sequence is geometric in nature// Thus:


 n
1 − rn
  
r −1
Sn = a or a
r−1 1−r
3
But r = <1
4
1
a =
4"
n #
1 1 − 34
Sn =
4 1 − 43
" n #
1 1 − 34
= 1
4 4
  n 
1 4 3
= x 1−
4 1 4
 n
3
Sn = 1 −
4
 x
3
F (x) = Sx = 1 −
4


 0 x<0




 x
F (x) = 1 − 34 0≤x<∞






1 ∞≥x

2 MATHEMATICAL EXPECTATION
QUESTION: Show that E(C) =C for C is a constant
SOLUTION:
Assuming a discrete approach
X
E(C) = Cf (x)
R
X
= C f (x)
R
= Cx1

= C

7
Distribution Theory II: STA 412 Dr. [Link]

Assuming a continous approach


Z ∞
E(C) = Cf (x)dx
−∞
= Cx1

= C
X Z ∞
f or f (x) = 1 and f (x)dx = 1
R −∞

P
Question: Show that E[C1 U1 (x) + C2 U2 (x)] = R [C1 U1 (x) + C2 U2 (x)] f (x)
SOLUTION:

8
Distribution Theory II: STA 412 Dr. [Link]

Assuming a discrete random variable


X
E [C1 U1 (x) + C2 U2 (x)] = [C1 U1 (x) + C2 U2 (x)] f (x)
R
X
= [C1 U1 (x)f (x) + C2 U2 (x)f (x)]
R
Since C1 and C2 are constant
X X
= C1 U1 (x)f (x) + C2 U2 (x)f (x)
R R
X X
= C1 U1 (x)f (x) + C2 U2 (x)f (x)
R R
= C1 E(U1 (x)) + C2 E(u2 (x))
X
f or E(U1 (x)) = U1 (x)f (x) and
R
X
E(U2 (x)) = U2 (x)f (x)
R

If Assumed Continous random variable

E [C1 U1 (x) + C2 U2 (x)]


Z ∞
= [C1 U1 (x) + C2 U2 (x)) f (x)dx
−∞
Z ∞ Z ∞
= C1 U1 (x)f (x)dx + C2 U2 (x)f (x)dx
−∞ −∞
Z ∞ Z ∞
= C1 U1 (x)f (x)dx + C2 U2 (x)f (x)dx
−∞ −∞
= C1 E(U1 (x)) + C2 E(U2 (x))
Z ∞
f or E(U1 (x)) = U1 (x)f (x)dx and
−∞
Z ∞
E(U2 (x)) = U2 (x)f (x)dx
−∞

9
Distribution Theory II: STA 412 Dr. [Link]

3 Discrete Random Variable


1. BERNOLLI DISTRIBUTION: This is a process or distribution which when single
trial of experiment results in only one of two mutually exclusive outcomes. e.g. dead or
alive, yes or no. It is also called Bernoulli trial.
It has a probability distribution function as

f (x) = px q 1−x , x = 0, 1

The Bernoulli distribution follows a random process or stochastic process for the outcome
of any specific trial is determined by chance
µ=p
σ 2 = pq

Standard deviation σ = pq

2. BINOMIAL DISTRIBUTION: This is a distribution that define the process of subse-


quent bernolli trial andthe outcome
 of any trial is independent of the previous trial.
n
It is defined as f (x) = px q n−x , x = 0, 1, 2, · · · , n
x

> ASSUMPTIONS OF BINOMIAL DISTRIBUTION

i. The experiment contains n repeated trials ( n is defined before the experiment begins).

ii. The result of every can only be classified into one of two mutually exclusive categories.

iii. The probability of success does not change from trial to trial.

iv. The result of any trial is independent of any other.

> SHAPE OF BINOMIAL DISTRIBUTION


The shape of the binomial distribution depends on the two parameters n and p

i. If p < 0.5 and n is small → Rightly skewed

ii. If p > 0.5 and n is small → Leftly skewed

iii. If p = 0.5 → Symmetric

10
Distribution Theory II: STA 412 Dr. [Link]

iv. In all cases, as n gets larger the distribution gets closer to being a symmetric, bell-
shaped distribution.
µ = np
σ 2 = npq

σ = npq

3. POISSON DISTRIBUTION: This has to do with occurrence that can be described by


a discrete random variable. The Poisson distribution can be used to find the probability
that a certain number of event will occur in a given period if time provided that the
following criteria are satisfied;

i. The time interval can be divided into many sub-intervals so small that the probability
of the event occuring in any one sub-interval is almost zero.

ii. The probability of more than one occurence in any sub-interval is negligible.

iii. The occurences of the event are independent i.e. The occurrence of an event in any
interval space of time has no effect on the probability of a second occurrence of the
event in the same or any other interval.

iv. The probability of the single occurrence of thw event in a given interval is proportional
to the lenght of the interval.

v. The probability of an occurence in any of the sub -interval (or the mean rate of
occurence) remains constant throughtout the entire time consideration.

λx e−λ
f (x) = x = 0, 1, 2, · · · , ∞, λ > 0
x!
where λ is the average number of occurrences of the random event in the interval. The poisson
distribution can be used when n is very large and p is very small.

11
Distribution Theory II: STA 412 Dr. [Link]

4 DERIVATION OF THE MEAN AND VARIANCE OF


THE ABOVE DISCRETE DISTRIBUTION AND SOME
OTHER ONES
4.1 BERNOULLI DISTRIBUTION

f (x) = px q 1−x , x = 0, 1
X1
E(x) = xf (x)
x=0
= 0p0 q 1−0 + 1p1 q 1−1

= 0(q) + p

= p=µ

1
X
2
E(x − µ) = (x − µ)f (x)
x=0
= (0 − p)2 p0 q 1−0 + (1 − p)2 p1 q 1−1

= p2 q + q 2 p

= pq[p + q]

= pq(1) = pq

= σ2

12
Distribution Theory II: STA 412 Dr. [Link]

4.2 BINOMIAL DISTRIBUTION



n
f (x) = px q n−x , x = 0, 1, 2, · · · , n
x
n
X
E(x) = xf (x)
x=0
n
X n!
E(x) = x px q n−x
x=0
(n − x)!x!
n
X n!
= px q n−x
x=0
(n − x)!(x − 1)!
n
X n(n − 1)!
= px q n−x
x=0
(n − x)!(x − 1)!
n−1
X (n − 1)!
= n p1 px−1 q n−x
x=1
(n − x)!(x − 1)!
n−1
X (n − 1)!
= np px−1 q n−x
x=1
(n − x)!(x − 1)!
let y = x − 1 and x = y + 1
n−1
X (n − 1)!
= np py q n−(y+1)
x=1
(n − (y + 1))!(y + 1 − 1)!
n−1
X (n − 1)!
= np py q (n−1)−y
y=1
(n − 1) − y)!(y!
n−1
X n − 1 
= np py q n−1−y
y
y=1
= np X 1

= np

T o obtain σ 2

σ 2 = E(x2 ) − (E(x))2

But E(x2 ) = E(x(x − 1)) + E(x)


n
X
E(x(x − 1)) = x(x − 1)f (x)
x=0

13
Distribution Theory II: STA 412 Dr. [Link]

n
X n!
= x(x − 1) px q n−x
x=0
(n − x)!x!
n
X x(x − 1)n!
= px q n−x
x=0
x(x − 1)(n − x)!x(x − 1)(x − 2)!
n
X n(n − 1)(n − 2)! x n−x
= p q
x=0
(n − x)!(x − 2)!
n−2
X (n − 2)!
= n(n − 1) px q n−x
x=2
(x − 2)!(n − x)!
Let x − 2 = y x = y + 2
n−2
X (n − 2)!
= n(n − 1) py+2 q n−(y+2)
y=0
y![n − (y + 2)]!
n−2
X (n − 2)!
= n(n − 1) p2 py q n−(y+2)
y=0
y![n − (y + 2)]!
n−2  
2
X n−2
= n(n − 1)p py q n−2−y
y
y=2

= n(n − 1)p2

σ 2 = E(x2 ) − [E(x)]2

= E(x(x − 1)) + E(x) + [E(x)]2

= n(n − 1)p2 + np − (np)2

= n2 p2 − np2 + np − n2 p2

= np − np2

= np[1 − p]

= npq

14
Distribution Theory II: STA 412 Dr. [Link]

4.3 POISSON DISTRIBUTION:

λx e−λ
f (x) = x = 0, 1, 2, . . . , ∞ λ > 0
x!

X
E(x) = xf (x)
x=0

X λx e−λ
= x
x=0
x!

−λ
X λx
= e x
x=0
x!

X λx
= e−λ x
x=0
x(x − 1)!

X λx
= e−λ
x=0
(x − 1)!
T hen Let x − 1 = y
∞ ∞
X λy+1 X λy λ
E(x) = e−λ = e−λ
x=0
y! x=0
y!

X λy
= e−λ λ
x=0
y!
0
e = λ

15
Distribution Theory II: STA 412 Dr. [Link]

To obtain variance σ 2

σ 2 = E(x2 ) − [E(x)]2

But E(x)2 = E(x(x − 1) + x)

∴ σ2
= E(x(x − 1) + E(x) − [E(x)]2

X
E(x(x − 1) = x(x − 1)f (x)
x=0

X λx e−λ
= x(x − 1)
x=0
x!

X λx e−λ
= x(x − 1)
x=0
x(x − 1)(x − 2)!

X λx−2 λ2 e−λ
=
x=2
(x − 2)!

2
X λt e−λ
= λ
x=2
t!
= λ2 x1 = λ2

X λt e−λ
f or =1
t=0
t!
∴ σ 2 = E(x(x − 1) + E(x) − [E(x)]2

= λ2 + λ − λ2

σ = λ

POISSON DISTRIBUTION AS AN APPROXIMATELY TO THE BINOMIAL


DISTRIBUTION

Given λ = np  
n
show that as n increases p(x) = px q n−x , x = 0, 1, 2, · · · to poisson prob-
x
ability

16
Distribution Theory II: STA 412 Dr. [Link]

λ
p =
n 
λ
p(x, n, p) = p x, n,
n
 
n
= px q n−x
x
   x  n−x
n λ λ
= 1− , x = 0, 1, . . . , n
x n n
 n  −x
n(n − 1)(n − 2) · · · (n − x + 1) x λ λ
= λ 1− 1−
x!nx n n
x
 n  −x
n(n − 1)(n − 2) · · · (n − x + 1) λ λ λ
= 1− 1−
nx x! n n
n(n − 1)(n − 2) · · · (n − x + 1)
But ≡1
nx
 −x
λ
lim 1 − ≡1
n→∞ n
 n
λ
lim 1 − = e−λ
n→∞ n

4.4 Hypergeometric Distribution:

If there exist a lot of m + n items of which m are defective and the remaining n of them are not
defective. A sample of r items is drawn randomly without replacement.
Let X denote the number of defective items that is observed in the sample. The random
variable x is the hypergeometric random variable with  parameter
 m + n and m. Then number of
m
ways of selecting x defective items from m defective is . The number of ways of sellecting
 x 
n
r −x non-defective items from non-defective items is . The total no of way of selecting
  r − x
m n
x defective and r − x non-defective item is .
x r−x
Therefore the probability of observing x defective items in a sample of r items (Probability

17
Distribution Theory II: STA 412 Dr. [Link]

density function) is
  
m n
x r−x x = 0, 1, 2, · · ·
  f or x≤m
m+n r−x≤n
r

PROOF:
  
m n
x r−x
f (x) =   x = 0, 1, 2, · · · , r
m+n
r
  
m n
r
X x r−x
E(x) = x·  
x=0
m+n
r
r m! n
X x!(m−x)!
Cr−x
= x m+n C
x=0 r

r xm(m−1)!) n
X x(x−1)!(m−x)!
Cr−x
= m+n C
x=0 r
m(m−1)!) n
r
X (x−1)!(m−x)!
Cr−x
= m+n C
x=0 r

= Let x − 1 = y; x = y + 1

m−x = m−y−1

= (m − 1) − y
r (m−1)!) n
X y!(m−1−y)!
Cr−x
= m m+n C
x=0 r

m m−1
= m+n C
Cy+r−1−y
r
m
= m+n C
·m+n−1 Cr−1
r

18
Distribution Theory II: STA 412 Dr. [Link]

m (m + n − 1)!
= (m+n)!
X
(r − 1)!(m + n − 1 − r + 1)!
r!(m+n−r)!
mr!(m + n − r)! (m + n − 1)1)
= x
(m + n)! (r − 1)!(m + n − r)!
mr(r − 1)! (m + n − 1)!
= X
(m + n)(m + n − 1)! (r − 1)!
mr
=
m+n

To obtain σ 2 = E(X 2 ) − [E(X)]2

19
Distribution Theory II: STA 412 Dr. [Link]

But E(X 2 ) = E(x(x − 1)) + E(x)


  
m n
r
X x r−x
E(x(x − 1)) = x(x − 1)  
x=0
m+n
r
m(m−1)(m−2)! C
r x(x − 1)
X x(x−1)(x−2)!(m−x)! r−x
= m+n C
x=0 r
(m−2)! n
r
X (x−2)!(m−x)!
Cr−x
= m(m − 1) m+n C
x=2 r
r n
m(m − 1) X (m − 2)!
= m+n C
= Cr−x
r x=2
(x − 2)!(m − x)!
Let y = x − 2; x = y + 2; m − x = m − 2 − y
r
m(m − 1) X (m − 2)! n
= m+n C
Cr−x
r y=2
y!(m − 2 − y)!
m(m − 1)r(r − 1)(r − 2)! (m + n − 2)!
= X
(m + n)(m + n − 1)(m + n − 2)! (r − 2)!
m(m − 1)r(r − 1)
=
(m + n(m + n − 1)
Hence

E(X 2 ) = E(x(x − 1)) + E(x)


m(m − 1)r(r − 1)(r − 2)! (m + n − 2)! mr
= X +
(m + n)(m + n − 1)(m + n − 2)! (r − 2)! m+n
m(m − 1)r(r − 1)(r − 2)! (m + n − 2)! mr m2 r2
σ2 = X + −
(m + n)(m + n − 1)(m + n − 2)! (r − 2)! m + n (m + n)2
 
mr (m − 1)(r − 1) mr
= +1−
m+n m+n−1 m+n
 
mr (m − 1)(r − 1)(m + n) + (m + n)(m + n − 1) − mr(m + n − 1)
=
m+n (m + n − 1)(m + n)
 2
m r − m − rn + m + mn + n2 − m2 r
2 2

mr
=
m+n (m + n)(m + n − 1)
mr
=
m+n 
mr n(n + m − r)
=
m + n (m + n)(m + n − 1)
  
mr n n+m−r
=
m+n m+n m+n−1
  
mr m m+n−r
= 1− 20
m+n m+n m+n−1
m m+n−m n
N ote : 1 − = =
m+n m+n m+n
Distribution Theory II: STA 412 Dr. [Link]

4.4.1 Binomial Distribution as an Approximately to the Hypergeometric Distri-


bution

m
Let = pm,n −→ p 0<p<1
m+n
givenm, n −→ ∞
  
m n
r−x
 
x r
then   −→ px q r−x
m+n x
r
  
m n
m!
x r−x x!(m−x)! n!
  = (m+n)! X
m+n (r − x)!(n − r + x)!
r!(m+n−r)!
r
m! r!(m + n − r)! n!
= X X
x!(m − x)! (m + n)! (r − x)!(n − r + x)!
r! m!n!(m + n − r)!
= ·
x!(m − x)! (m + n)!(m − x)!(n − r + x)!
 
r m(m − 1)(m − 2) · · · (m − x+!)(m − x)!n!(m + n − r)!
=
x (m + n)!(m − x)!(n − r + x)!
 
r m(m − 1) · · · (m − x + 1)n!(m + n − r)!
=
x (m + n)!(n − r + x)!
 
r m(m − 1) · · · (m − x + 1)n!(m + n − r)!
= !
x (m + n)(m + n − 1) · · · (m + n − r + 1)(m + n − r)(n − r + x)!
 
r m(m − 1) · · · (m − x + 1)n!
=
x (m + n)(m + n − 1) · · · (m + n − r + 1)(n − r + x1)!
 
r m(m − 1) · · · (m − x + 1)n(n − 1)! · · · (n − r + x + 1)(n − r + x)!
=
x (m + n)(m + n − 1) · · · (m + n − r + 1)(n − r + x1)!
 
r m(m − 1) · · · (m − x + 1)n(n − 1)! · · · (n − r + x + 1)
=
x (m + n)(m + n − 1) · · · (m + n − (r − 1))
Divide through by m + n
 " m  m 1
 m x−1
 n
 n 1
 n r−x−1
#
r m+n m+n
− m+n
· · · m+n
− m+n
X m+n m+n
− m+n
· · · m+n
− m+n
= m+n
 m+n r−1

x m+n
· · · m+n
− m+n
 " m  m 1
 m x−1
 n
 n 1
 n r−x−1
#
r m+n m+n
− m+n
· · · m+n
− m+n
X m+n
− · · · −
= r−1
 m+n m+n m+n m+n
x 1 · · · 1 − m+n

21
Distribution Theory II: STA 412 Dr. [Link]

m
Since p =
m+n
m m+n−m n
1−p = 1− = =
m+n m+n m+n
 " 1
 x−1
 1
 r−x−1
#
r p p − m+n · · · p − m+n X q q − m+n ··· q − m+n
T aking the limm+n−→∞ r−1

x 1 · · · 1 − m+n
 
r
= px q r−x
x

4.5 Geometric Distribution:

This is the distribution where process deal with an experiment of Bernoulli trials sequence
deserved utill the first success occurs
Note: r − 1 denote the total number of trials before the first success.
Thus: r is the total number of trial at exactly when first success occurs.
POC: r is not predetermined for BINOMIAL exberiment.


r−1
p(1 − p)
 r = 1, 2, · · · ∞
f (x) =

0 elsewhere


X
PROOF:E(r) = rp(1 − p)r−1
r=1

X
= rpq r−1
r=1
X∞
= p rq r−1
r=1

∂ 
X
q + q2 + q3 + · · ·

= p
r=1
∂q
∂ 
1 + q + q2 + · · ·

= pq
∂q
But 1 + q + q 2 + · · · is a geometric series
a 1
S∞ + = f orr = q
1−r 1−q

22
Distribution Theory II: STA 412 Dr. [Link]

 
∂ 1
= pq
∂q 1 − q
 
∂ q
= p
∂q 1 − q
 
(1 − q)(1) − q(−1)
= p
(1 − q)2
 
1−q+q
= p
(1 − q)2
 
1
= p 2
p
1
=
p

T o obtain σ 2

σ 2 = E(r2 ) − [E(r)]2

X ∞
X
E(r2 ) = r2 pq r−1 = p r2 qr − 1
r=1 r=1

X ∂
= p (rq r )
r=1
∂q

X ∂ r
= p rq
r=1
∂q

∂ X r
= p rq
∂q r=1
∂ 
q + 2q 2 + 3q 3 + · · ·

= p
∂q
∂ 
= p q 1 + 2q + 3q 2 · · ·

∂q
1
But 1 + 2q + 3q 2 · · · =
(1 − q)2
 
∂ q
= p
∂q (1 − q)2
(1 − q)2 (1) − q(−2(1 − q))
 
= p
(1 − q)4
 
(1 − q)[1 − q + 2q]
= p
(1 − q)4
 
(1 − q)(1 + q)
= p
(1 − q)(1 + q)3
23
Distribution Theory II: STA 412 Dr. [Link]

 
1+q
= p
p3
1+q
=
p2
1+q 1
∴ σ2 = 2
− 2
p p
1+q−1
=
p2
q
=
p2

4.6 Negative Binomial Distribution:

Consider a succession of Bernoulli trials, let p(r) denotes the probability that exactly r+k(k > 0)
trials are needed to produce k successes. This will so happen when the last trial i.e(r + k)th
trials is a success with probability p and the previous trial (r + k − 1) must (k − 1) successes
r+k−1
with probability ck−1 pk−1 q r where q = 1 − p p( r) = p(k − 1) successes in (x+k-1) trials X
prob (r + k)th trial

r+k−1
= ck−1 pk−1 q r · p
r+k−1
= ck−1 pk−1+1 q r
r+k−1
= ck−1 pk q r

Show that

 
r+k−1 k r k r −k
Ck−1 p q if r = 0, 1, 2, · · · = p (−1) qr
r
SOLUTION:

r+k−1
Ck−1 pk q r
(r + k − 1)!
= pk q r
(k − 1)!(r + k − 1 − k + 1)!

24
Distribution Theory II: STA 412 Dr. [Link]

(r + k − 1)! k r
= p q
(k − 1)!r!
(k + r − 1)(k + r − 2) · · · (k + r + 1 − (r + 1))(k − 1)! k r
= p q
(k − 1)!r!
(k + r − 1)(k + r − 2) · · · k k r
= p q
r!
k(k + 1) · · · (k + r − 2)(k + r − 1)pk q r
=
r!
(−k)(−k − 1) · · · (−k − r + 2)(−k − r + 1) k r
= (−1)r p q
  r!
−k
= pk (−1)r qr
r

∞ ∞  
X
k
X
r −k
POC: p(r) = p (−1) qr
r
r≥0 r=0

= p (1 − q)−k
k

= pk p−k

= pk−k

= p0 = 1

4.7 Multinomial Distribution:

This is the probability distribution of outcomes from a multinomial experiment.


Suppose a multinomial experiment consists of n trials and each trial can result in any of k possible
outcomes E1 , E2 , · · · Ek . Suppose that each of outcomes occur with probabilities p1 , p2 , · · · pk .
Yherefore, the probability that E1 occurs n1 times, E2 occurs n2 times and Ek occurs nk times
is  
n!
p= (pn1 1 , pn2 2 · · · pnk k ) where n = n1 + n2 + · · · + nk
n1 !, n2 !, · · · , nk !

25
Distribution Theory II: STA 412 Dr. [Link]

4.7.1 Multivariate Distribution (Hypergeometric Distribution WOR)


    
M1 M2 Mk
······
x1 x2 xk
f (x1 , x2 , · · · ; n, M1 , M2 , · · · ) =  
N
n
k
X k
X
N= Mi n= xi
i=1 i=1
Example:
P
Let x1 = 3, x2 = 2, x3 = 2, x4 = 1x5 = 1 xi = n
n!
f (x) = (p1 )x1 (p2 )x2 (p3 )x3 (p4 )x4 (p5 )x5
x1 !x2 !x3 !x4 !x5 !
 2  3  2  1  1
9! 1 1 1 1 1
=
2!3!2! 6 6 6 6 6

CONTIUOUS PROBABILITY DISTRIBUTIONS

5 UNIFORM DISTRIBUTION
Let X ∼ U (a, b)
Then

1
 b−a
 a<x<b
f (x) =

0, otherwise

Z b
1
E(x) = x· dx
a b−a
Z b
1
= xdx
b−a a
 2 b
1 x 1  2
b − a2

= =
b − a 2 a 2(b − a)
(b − 1)(b + a)
=
2(b − a)
b+a
=
2

26
Distribution Theory II: STA 412 Dr. [Link]

To obtain σ 2

σ 2 = E(x2 ) − [E(x)]2
Z b
2 1
E(x ) = x2 dx
a b−a
Z b  3 b
1 2 1 x
= x dx =
b−a a b−a 3 a
1 (b − a)(b2 + ab + a2 )
= (b3 − a3 ) =
3(b − a) 3(b − a)
2 2
b + ab + a
=
2
2
b + ab + a2
2

2 b−a
σ = −
3 2
4b + 4ab + 4a − 3b2 − 6ab − a2
2 2
=
12
2 2
b − 2ab + a
=
12
2
(b − a)
σ2 =
12

6 EXPONENTIAL DISTRIBUTION
Let X ∼ Exponential(λ)

f (x) = λe−λx λ > 0, x > 0


Z ∞ Z ∞
−λx
E(x) = xλe =λ xe−λx
0 0
y dy dy
Let y = λx, x = = λ dx =
λ dx λ
By Substitution
Z ∞ 
y −y 1
E(x) = λ e · dy
0 λ λ

27
Distribution Theory II: STA 412 Dr. [Link]

λ ∞ −y
Z
= ye dy
λ2 0
λ
= Γ2
λ2
1
= (2 − 1)!
λ

1
∴ E(x) =
λ

To obtain σ 2

σ 2 = E(x2 ) − [E(x)]2
Z ∞ Z ∞
−λx
2
E(x ) = 2
x λe =λ x2 e−λx
0 0
y dy dy
Let y = λx, x = = λ dx =
λ dx λ
By Substitution
Z ∞  2
y 1
E(x2 ) = λ e−y · dy
0 λ λ
Z ∞
λ
= 3
y 2 e−y dy
λ 0
λ
= Γ3
λ3
1
= (3 − 1)!
λ2

2
E(x2 ) =
λ2
2 1 2−1
∴ σ2 = 2
− 2 =
λ λ λ2
1
σ2 =
λ2

28
Distribution Theory II: STA 412 Dr. [Link]

POC:
In a general context

Z ∞
r
E(x ) = xr λe−λx dx
0
y dy dy
Let y = λx, x= = λ dx =
λ dx λ
By Substitution
Z ∞  r
y 1
r
E(x ) = λ e−y · dy
0 λ λ
Z ∞
λ
= r
y r e−y dy
λλ 0
1
= Γr + 1
λr

7 GAMMA DISTRIBUTION:
This is generally known as a distribution frequently used in waiting time modeling
It has pdf as x
xα−1 e− β
f (x) = x > 0, α > 0, β > 0
Γαβ α
Proof: Mean and Variance
In a general case since the mean and variance of any distribution can be obtain by observing
0 0 0
the first and second moment about the origin i.e. µ1 , µ2 , Hence it is adviceable to establish µt
for concerning and to avoid error as the basis for estimation.

29
Distribution Theory II: STA 412 Dr. [Link]

Thus:
x

xr xα−1 e− β
Z
r 0
E(x ) = µr = dx
Γαβ α
∞ Z0
1 α−1 r − βx
= x x e
Γαβ α 0
Z ∞
1 α+r−1 − βx
= x e dx
Γαβ α 0
x dy 1
Let y = ; = ; dx = βdy and x = yβ
β dx β
By substitution
Z ∞
1
= α
(yβ)α+r−1 e−y · βdy
Γαβ 0
Z ∞
1
= β α+r−1
·β y α+r−1 e−y dy
Γαβ α 0
1
= β α+r Γ(α + r)
Γαβ α
β α+r Γ(α + r)
=
Γαβ α
β r Γ(α + r)
=
Γα
β
Hence : E(x) = = αβ
Γα
β2 β2
E(x2 ) = Γ(α + 2) = · (α + 1)Γ(α + 1)
· α
= β 2 (α + 1)α

Thus:

σ2 = E(x2 ) − [E(x)]2

= β 2 (α + 1)α − α2 β 2

= (αβ 2 + β 2 )α − α2 β 2

= αβ 2

8 ERLANG DISTRIBUTION
λr xr−1 e−λx
f (x) = 0 < x, r = 1, 2, · · ·
(r − 1)!

30
Distribution Theory II: STA 412 Dr. [Link]

To obtain the mean and variance using the general approach of moment about the origin:
Z ∞
k λr xr−1 e−λx
E(x ) = xk dx
0 (r − 1)!
Z ∞
λr
= xk xr−1 e−λx dx
(r − 1)! 0
Z ∞
λr
= xk+r−1 e−λx dx
(r − 1)! 0
y dy dy
Let y = λx, x = , = λ dx =
λ dx λ
By Substitution
Z ∞  k+r−1
λr y 1
= e−y · dy
(r − 1)! 0 λ λ
r Z ∞
λ 1 1
= · k+r−1 · y k+r−1 e−y dy
(r − 1)! λ λ 0
λr 1
= · k−1 r Γ(k + r)
(r − 1)! λ λ · λ
1
= k
Γ(k + r)
λ Γr
Hence W hen k = 1 to get
Γ(r + 1)
E(x) =
λΓr
rΓr
=
λΓr
r
=
λ

W hen k = 2
Γ(r + 2)
E(x2 ) =
λ2 Γr
(r + 1)rΓr
=
λ2 Γr
r(1 + r)
=
λ2
∴ σ = E(x2 ) − [E(x)]2
2

2 r(1 + r) h r i2
σ = −
λ2 λ
r + r2 − r2
λ2
r
σ2 =
λ2
31
Distribution Theory II: STA 412 Dr. [Link]

9 WEIBULL DISTRIBUTION
β  x β−1 −( σx )β
f (x) = e 0 < x, 0 < β, 0 < σ
σ σ
Z ∞

 x β−1 x β
r
E(x ) = x e−( σ ) dx
0 σ σ

xr xβ−1 −( σx )β
Z
β
= e dx
σ 0 σ β−1
Z ∞
β x β
r+β−1 −( σ ) dx
= x e
σ · σ β · σ −1 0
Z ∞
β r+β−1 −( σx β
) dx
= x e
σβ 0
 x β 1 x 1 dy  x β−1 1 β  x β−1
Lety = ∴ y β = ; x = σy β ; =β · =
σ σ dx σ σ σ σ

32
Distribution Theory II: STA 412 Dr. [Link]

σ
dx = dy
x β−1

β σ
By Substitution
Z ∞
β  1
r+β−1 σ
r
E(x ) = σy β e−y dy
σβ x β−1

0 β σ
Z ∞
β r β 1 σ
= σ r+β−1 y β β β e−y  1 β−1 dy
σβ 0 β
β σyσ
r β 1
β · σ r+β−1 σ ∞ y β + β − β −y
Z
= 1 e dy
σβ · β 0 y · y− β
Z ∞ βr ββ −1
r y y y β −y
= σ 1 e dy
0 y 1− β
Z ∞  
r
+β − β1 1
= σ r
y β β − 1− e−y dy
β
Z0 ∞
(r+β−1−1+1)
= σr y β e−y dy
Z0 ∞
r
= σr y β e−y dy
0
 
r r
= σ Γ 1+
β

W hen r = 1
 
1
⇒ E(x) = σΓ 1 +
β
W hen r = 2
 
2 2 2
⇒ E(x ) = σ Γ 1 +
β
    2
2 2 1
V ar(x) = σ Γ 1 + − σΓ 1 +
β β
    2 !
2 1
V ar(x) = σ 2 Γ 1 + − Γ 1+
β β

10 MAXWELL DISTRIBUTION
x2
2 x2 e− 2a2
r
Given f (x) = x > 0; a > 0
π a3

33
Distribution Theory II: STA 412 Dr. [Link]

Show that f (x) is a probability density function


Solution:
Z ∞
f (x)dx = 1
0
x2
2 x2 e− 2a2
Z ∞
r
= dx
0 π a3
q
2 Z ∞
π x2
= x2 e− 2a2 dx
a3 0
x2 2 2
p dy x a2
Let = y ; 2a y = x x = a 2y ; = ; dx = dy
2a2 dx a2 x
By substitution
q
2 Z ∞
π p 2 −y a2
= a 2y e · dy
a3 0 x
q
2 Z ∞
2
π 2 −y a
= a 2ye √ dy
a3 0 a 2y
q
2 Z ∞
π 1
= 3
·a 3
(2y)1− 2 e−y dy
a
√ Z ∞ 0
2 √ 1
= √ 2 · y 2 e−y dy
π 0
Z ∞
2 1
= √ y 2 e−y dy
π 0
 
2 3
= √ ·Γ
π 2
 
2 1 1
= √ · ·Γ
π 2 2

 
1
N ote : π = Γ
2

Z ∞
f (x)dx = 1
0

34
Distribution Theory II: STA 412 Dr. [Link]

To obtain the mean and variance of MAXWELL DISTRIBUTION


By general Approach
x2
2 x2 e− 2a2
Z ∞
r
r r
E(x ) = x dx
0 π a3
q
2 Z ∞
π x2
= xr · x2 e− 2a2 dx
a3 0
q
2 Z ∞
π x2
= xr+2 e− 2a2 dx
a3 0
F rom previous solution , By substitution, we have :
q
2 Z ∞
r π p r+2 −y a2
E(x ) = a 2y e dy
a3 0 x
q
2 Z ∞
π r+2
√ r+2 −y a2
= a ( 2) e √ dy
a3 0 a 2y
q
2 Z ∞
π r+3 r+2
− 12 −y
= a (2y) 2 e dy
a3
q 0
2 Z ∞
π r+3 r+1
= 3
a (2y) 2 e−y dy
a 0
q
2 Z ∞
π r r+1 r+1
= 3
3
·a ·a 2 2 y 2 e−y dy
a 0
r  
2 r r+1 r+3
= ·a ·2 Γ 2
π 2
1  
22 r r 1 r+3
= √ · a · 22 · 22 Γ
π 2
1+ r2 r
 
2 a r+3
E(xr ) = √ Γ
π 2
1
21+ 2 a
 
1+3
Hence; E(x) = √ Γ i.e when r = 1
π 2
3  
22 a 4
= √ Γ
π 2
3
22 a
= √ Γ(2)
π
3
22 a
= √ ·1
π
35
Distribution Theory II: STA 412 Dr. [Link]

3
22 a
E(x) = √
π
2
21+ 2 a2
 
2 2+3
E(x ) = √ Γ
π 2
2 2
 
2a 5
= √ Γ
π 2
2
 
4a 5
= √ Γ
π 2

  " 3 #2
4a2 5 22 a
∴ V ar(x) = √ Γ − √
π 2 π
4a2 Γ 25

8a2
= √ −
π π
" #
5

Γ 2
= 4a2 √ 2 −
π π
     
5 3 3 3 1 1
But Γ = Γ = · Γ
2 2 2 2 2 2
" #
3 1 1

· Γ 2
V ar(x) = 4a2 2 2 1 2 −
Γ 2 π

 
1
N ote : that π = Γ
2
 
2 3 2
T hus V (x) = 4a −
4 π
 
2 8
V ar(x) = a 3−
π

36
Distribution Theory II: STA 412 Dr. [Link]

11 PARETO DISTRIBUTION
kβ k
The probability density function is given by f (x) = xk+1
,k > β, x > 0
By general approach moment

kβ k
Z
r r
E(x ) = x . k+1 dx
β x
Z ∞
= kβ k xr .x−k−1 dx
β
Z ∞
k
= kβ xr−k−1 dx
β
r−k
x
= kβ k [ ]∞
r−k β
kβ k
= [0 − β r−k ]
r−k
kβ k 1
= ( k−r )
k−r β
kβ k 1
= ( k −r )
k−r β β
k
= (β −r )−1
k−r
kβ r
=
k−r

Thus, the mean of the pareto distribution is obtained as :


E(x) = since r = 1
k−1

37
Distribution Theory II: STA 412 Dr. [Link]

kβ 2
Also the E(x2 ) = k−2

Then var(x) implies


2
kβ 2


var(x) = −
k−2 k−1
kβ (k − 1)2 − k 2 β 2 (k − 2)
2
=
(k − 1)2 (k − 2)
kβ 2 (k 2 − 2k + 1) − k 3 β 2 − 2k 2 β 2
=
(k − 1)2 (k − 2)
k 3 β 2 − 2k 2 β 2 + kβ 2 − k 3 β 2 − 2k 2 β 2
=
(k − 1)2 (k − 2)
kβ 2
=
(k − 1)2 (k − 2)

12 NORMAL DISTRIBUTION
The probability density function for the normal distribution is given as

1 x−µ 2
f (x) = √ e(− σ ) , −∞ < x < ∞, µ0 , σ 2 > 0
2πσ 2

38
Distribution Theory II: STA 412 Dr. [Link]

To obtain the mean and variance using the general moment approach about the origin.
Z ∞
r 1 x−µ 2
E(x ) = xr √ e(− σ ) dx
2πσ 2
−∞
Z ∞
1 x−µ 2
= √ xr e(− σ ) dx
σ 2π −∞
x−µ 2
Let y 2 = ( )
σ
x−µ x µ
y = = − ;
σ σ σ
x = µ + σy
dy 1
= ; dx = σdy
dx σ
By substitution, we have
Z ∞
1 1
r
E(x ) = √ (µ + σy)r e− 2 y 2 σdy
σ 2π −∞
Z ∞
σ 1
= √ (µ + σy)r e− 2 y 2 dy
σ 2π −∞
Z ∞
1 1
= √ (µ + σy)r e− 2 y 2 dy
2π −∞
Z ∞
1 1 2
E(x) = √ (µ + σy)e− 2 y dy i.e when r = 1
2π −∞
Z ∞ Z ∞
1 − y2 y
= √ [µ e dy + σ e− 2 dy]

Z ∞ −∞ Z −∞

1 −y 1 y
= µ √ e 2 dy + σ √ e− 2 dy
−∞ 2π −∞ 2π
= µ(1) + σ(0) = µ

39
Distribution Theory II: STA 412 Dr. [Link]

NOTE: The mean of a standardized normal distribution is zero and the variance is one. This
implies that given

1 y
f (x) = √ e− 2
Z 2π

1 y
E(y) = y √ e− 2 dy = 0
−∞ 2π
T hus, var(y) = E(y ) − (E(y))2 ⇒ 1 = E(y 2 ) − 02
2

E(y 2 ) = 1
Z ∞
1 1
2
E(x ) = √ (µ + σy)2 e− 2 y 2 dy
2π −∞
Z ∞
1 y2
= √ (µ2 + 2µσy + σ 2 y 2 )e − dy
2π −∞ 2
Z ∞ Z ∞ Z ∞
1 −y 1 −y 1 y
= µ 2
√ e 2 dy + 2µσ y √ e 2 dy + σ 2
y 2 √ e− 2 dy
−∞ 2π −∞ 2π −∞ 2π
2 2
= µ (1) + 2µσ(0) + σ (1)

E(x2 ) = µ2 + σ 2

Thus, the variance of the x i.e

var(x) = E(x2 ) − (E(x))2 = µ2 + σ 2 − µ2 = σ 2 .

13 BETA DISTRIBUTION
The beta probability density function is given as
 a−1 
x
f (x) = (1 − x)b−1 , a > 0, b > 0, 0 < x < 1
β(a, b)

where Z 1
Γ(a)(b) ΓaΓb
β(a, b) = = = xa−1 (1 − x)b−1 dx
Γ(a + b) Γa + b 0

40
Distribution Theory II: STA 412 Dr. [Link]

To obtain the mean and the variance of the beta distribution, the general moment is given as
Z 1
r xa−1
E(x ) = xr (1 − x)b−1 dx
0 β(a, b)
Z 1
1
= xr xa−1 (1 − x)b−1 dx
β(a, b) 0
Z 1
1
= xr+a−1 (1 − x)b−1 dx
β(a, b) 0
1
= β(a + r, b)
β(a, b)
1 Γ(a + r)Γb
= ΓaΓb ∗
Γ(a+b)
Γ(a + r + b)
Γ(a + b) Γ(a + r + b)
= ∗
ΓaΓb Γ(a + r)Γb
Γ(a + b)Γ(a + r)
=
ΓaΓ(a + r + b)
T hen; E(x) ⇒ r=1
Γ(a + b)Γ(a + 1) Γ(a + b).aΓa
E(x) = =
ΓaΓ(a + 1 + b) Γa((a + b))Γ(a + b)
a
=
a+b
Γ(a + b)Γ(a + 2)
E(x2 ) =
ΓaΓ(a + 2 + b)
Γ(a + b)(a + 1)aΓa
=
Γa(a + b + 1)(a + b)Γ(a + b)
a2 + a
=
a2 + ab + 2ab + a + b
a2 + a a2
var(x) = −
a2 + ab + 2ab + a + b a2 + 2ab + b2

a4 + 2a3 b + a2 b2 + a3 + 2a2 b + ab2 − a4 − a2 b2 − 2a3 b − a3 − a2 b


=
(a + b + 1)(a + b)(a + b)2
a2 b + ab2
=
(a + b + 1)(a + b)(a + b)2
ab(a + b)
=
(a + b + 1)(a + b)(a + b)2
ab
= .
(a + b + 1)(a + b)2

41
Distribution Theory II: STA 412 Dr. [Link]

14 RAYLEIGH DISTRIBUTION

−αx2
2αxe
 f or x > 0
f (x) =

0 elsewhere

Z ∞
2
E(x) = x2αxe−αx dx
0
Z ∞
2
= 2α x2 e−αx dx
0
p
Let y = αx2 ; x = y/α = (y/α)1/2 ;

dy/dx = 2αx;
1 1 1
dx = dy = dy = dy
2αx (y/α)1/2 2α1/2 y 1/2
Z ∞
1
⇒ 2α ((y/α)1/2 )2 e−y 1/2 1/2 dy
2α y
Z0 ∞
y 1
= 2α . 1/2 1/2 e−y dy
0 α 2α y
Z ∞
2
= y 1/2 e−y dy
2α1/2 0
2
= Γ(3/2)
2α1/2
2
= 1/2
2α Γ(1/2)

π 1p
= 1/2
= (π/α)
Z2α ∞
2
2
E(x2 ) = 2αx3 e−αx dx
0
Let y = αx2 ; dy = 2ax dx
1
dx = dy
2ax
and x = (y/α)1/2

42
Distribution Theory II: STA 412 Dr. [Link]

By substitution, we have

Z ∞
1
2
E(x ) = 2α ((y/α)1/2 )3 e−y 1/2 1/2 dy
0 2α y
Z ∞
2
= y 3/2−1/2 e−y dy
2α α3/2 0
1/2

α ∞ −y
Z
1 1
= 2
ye dy = Γ2 =
α 0 α α
 2
1 1p
var(x) = − (π/α)
α 2
1 π 1h πi
= − = 1−
α 4α α 4

15 MOMENT GENERATING FUNCTION


The folowing are some characteristics of the moment generating function

1 MX (0) = 1

2. MX (t) ≤ 1

3. MX is unif ormly continuous

4. MX+d (t) = E(et(X+d) ) = E(etX+td ) = etd E(etX ) = etd MX (t)

5. McX (t) = E(ecXt ) = MX (ct) where c is a constant

6. McX+d (t) = E(et(cX+d) ) = E(etcX+td ) = etd E(ectX ) = etd MX (ct)

NOTE:
dn MX (t)
n
|t=0 = E(X n ); n = 1, 2, 3..., if E|X n | < ∞
dt

43
Distribution Theory II: STA 412 Dr. [Link]

15.1 BINOMIAL DISTRIBUTION

n
f (x) = Cx px q n−x ; x = 0, 1, 2, ..., n

MX (t) = E(etx )
Xn n
X
= etx nCx px q n−x = nCx (pet )x q n−x
i=0 i=0
t n
= (pe + q)
0
MX (t) = n(pet + q)n−1 .pet
0
MX (0) = n(p + q)n−1 .pe( 0) = np...mean
00
MX (t) = n(n − 1)(pet + q)n−2 (pet )2 + n(pet + q)n−1 .pet
00
MX (0) = n(n − 1)(pe0 + q)n−2 (pe0 )2 + n(pe0 + q)n−1 .pe0

= n(n − 1)p2 + np

V ar(x) = n(n − 1)p2 + np − (np)2 = n2 p2 − np2 − n2 p2 + np = np − np2

= np(1 − p)

= npq.

15.2 POISON DISTRIBUTION


λx e−λ
f (x) = x!
;x = 0, 1, 2, ...
0elsewhere.

λx e−λ
MX (t) = E(etx ) =
P
x!
x=0
∞ ∞
λx
e−λ etx = (λet )x
P P
= x!
x=0 x=0

= e−λ (λet )x
P
x=0
t 0 λet )1 λet )2
= e−λ [ λe0!) + 1!
+ 2!
+ ...]
t t t
= e−λ eλe = e−(λ−λe ) = e−λ[1−e ]
t
MX0 (t) = λet [e−λ (1 − et )] = λet eλe −λ
MX0 (0) = λ
t t
MX00 (t) = λet eλe −λ + (λet )2 eλe −λ

44
Distribution Theory II: STA 412 Dr. [Link]

MX00 (0) = λ + λ2
V ar(x) = λ + λ2 − λ2

15.3 UNIFORM DISTRIBUTION

1
f (x) = ;a < x < b
b−a
0 elsewhere

MX (t) = E(etx )
Z b
1
= dx
a b−a
Z b
1 1 etx b
= etx dx = [ ]
b−a a b−a t a
1 etb − eta
= [etb − eta ] =
t(b − a) t(b − a)
2 3
(tb) (tb)
But etb = 1 + b + + + ...
2! 3!
(ta)2 (ta)3
eta = 1 + a + + + ...
2! 3!
T hus
2 2 3 3 4 4 t2 a2 t3 a3 t4 b4
[1 + b + t 2b + t 6b + t 24a + ...] − [1 + a + 2
+ 6
+ 24
+ ...]
Mx (t) =
t(b − a)
t2 (b2 −t2 ) 3 3 3 4 4 4
t(b − a) + 2
+ t (b 6−t ) + t (b24−t ) + ...
=
t(b − a)
t(b − a) t (b − a)(b + a) t3 (b − a)(b2 + ab + a2 ) t4 (b − a)(b + a)(b2 + a2 )
2
= + + +
t(b − a) 2t(b − a) 6t(b − a) 24t(b − a)
2 2 2 3 2 2
t(b + a) t (b + ab + a ) t ((b + a)(b + a )
= 1+ + +
2 6 24
(b + a) t b + ab + a ) t ((b + a)(b + a2 )
( 2 2 2 2
MX0 (t) = + +
2 3 8
b+a
MX0 (0) =
2

45
Distribution Theory II: STA 412 Dr. [Link]

(b2 + ab + a2 ) t((b + a)(b2 + a2 )


MX00 (t) = +
3 4
(b2 + ab + a2 )
MX00 (0) =
3
(b + ab + a2 ) (b + a)2
2
V ar(x) = −
3 4
(b2 + ab + a2 ) (b2 + ab + a2 )
= −
3 4
4b2 + 4ab + 4a2 − 3b2 − 6ab − 3a2
=
12
b2 − 2ab + a2
=
12
(b − a)2
=
12

15.4 EXPONENTIAL DISTRIBUTION



−λx
λe ;
 x>0 t<λ
f (x) =

0 elsewhere0

Z ∞
MX (t) = E(e ) = tx
λe−λx etx dx
0
Z ∞
= λ e−λx+tx dx
Z0 ∞
= λ e−(λ+t)x dx
0 −(λ+t)x ∞
e
= λ
t−λ 0
λ  −(λ+t)∞
− e−(λ+t)0

= [e
t−λ  
λ 1
= =λ
λ−t λ−t
−1
= λ(λ − t)
λ
MX0 (t) = λ(t − λ)−2 =
(λ − t)2
λ 1
MX0 (0) = =
λ2 λ

46
Distribution Theory II: STA 412 Dr. [Link]


MX00 (t) = 2λ(λ − t)−3 =
(λ − t)3
2λ 2
MX00 (0) = =
λ3 λ2
2 1 2 2−1
V ar(x) = − ( ) =
λ2 λ λ
1
= .
λ2

15.5 ERLANG DISTRIBUTION

λr xr−1 e−λx
f (x) = ; x > 0; r = 1, 2...
Γr
MX (t) = E(etx )
Z ∞ r r−1 −λx
λx e
= dx
0 Γr
λr ∞ r−1 −λx+tx
Z
= x e dx
Γr 0
λr ∞ r−1 −(λ+t)x
Z
= x e dx
Γr 0
dy 1
let y = (λ − t)x ; = λ − t; dx = dy;
dx λ−t
y
x =
λ−t
by substitution, we have
λr ∞ y r−1 −y 1
Z
MX (t) = ( ) e dy
Γr 0 λ − t λ−t
Z ∞
λr
= r −1
y r−1 e−y dy
Γr(λ − t)(λ − t) (λ − t) 0
λr
= Γr
Γr(λ − t)r
λr
= = λr (λ − t)r
(λ − t)r
MX0 (t) = rλr (λ − t)−r−1

MX0 (0) = rλr λ−r−1


rλr λ−r r
= =
λ λ
MX (t) = r(r + 1)λr (λ − t)−r−2
00

47
Distribution Theory II: STA 412 Dr. [Link]

λr λ−r r2 + r
= r(r + 1) =
λ2 λ2
2 2
r +r r
V ar = 2
− 2
λ λ
2 2
r + −r
=
λ2
r
= .
λ2

15.6 GAMMA DISTRIBUTION


x
xr−1 e− λ
f (x) =
λr Γr
MX (t) = E(etx )
Z ∞
1 x
= r
xr−1 e− λ etx dx
λ Γr 0
Z ∞
1 1
= r
xr−1 e−[ λ −t]x dx
λ Γr 0
1 1 − λt λ yλ
let y = ( − t)x; ; dx = dy; x =
λ λ 1 − λt 1 − λt
by substitution, we have
Z ∞
1 yλ r−1 −y λ
MX (t) = r
( ) e dy
λ Γr 0 1 − λt 1 − λt
Z ∞
λλr−1
= y r−1 e−y dy
λr Γr(1 − λt)(1 − λt)r−1 0
λλr λ−1
=
λr Γr(1 − λt)r (1 − λt)−1
1
= = (1 − λ)−r
(1 − λt)r
MX0 (t) = rλ(1 − λt)−1−r

MX0 (0) = rλ

MX00 (t) = r(r + 1)λ2 (1 − λt)−r−2

MX0 (0) = r(r + 1)λ2 = r2 λ2 + rλ2

V ar(x) = r2 λ2 + rλ2 − r2 λ2

= rλ2 .

48
Distribution Theory II: STA 412 Dr. [Link]

15.7 MOMENT GENERATING FUNCTION OF A NORMAL DIS-


TRIBUTION
1 1 2
f (x) = √ e− 2σ2 (x−µ) ; −∞ < x < ∞ 0 elsewhere
2πσ 2
Z ∞
1 1 2
tx
MX (t) = E(e ) = √ e− 2σ2 (x−µ) etx dx
2πσ 2
−∞
Z ∞
1 1 2
= √ e− 2σ2 (x−µ) +tx dx
2πσ 2 Z−∞

1 1 2 2
= √ e− 2σ2 [(x−µ) −2σ tx] dx
2πσ 2 Z−∞

1 1 2 2 2
= √ e− 2σ2 [x −2xµ+µ −2σ tx] dx
2πσ 2 Z−∞

1 1 2 2 2
= √ e− 2σ2 [x −2(µ+2σ t)x+µ ] dx
2πσ 2 −∞ Z

1 µ2 1 2 2
= √ e− σ 2 e− 2σ2 [x −2(µ+2σ t)x] dx
2πσ 2 Z−∞

1 µ2 1 2 2 2 2 2 2
= √ e− σ 2 e− 2σ2 [x −2(µ+2σ t)x+(µ+σ t) −(µ+σ t) ] dx
2πσ 2 −∞
Z ∞
1 − 2µ2 (µ+σ 2 t)2 1 2 2
= √ e σ e 2σ 2
e− 2σ2 [x−(µ+σ t)] dx
2πσ 2 Z ∞ −∞
µ2 (µ+σ 2 t)2 1 1 2 2
= e− σ2 + 2σ2 √ e− 2σ2 [x−(µ+σ t)] dx
−∞ 2πσ 2
µ2 (µ+σ 2 t)2
= e− σ2 + 2σ 2

µ2 µ2 (2µσ 2 t) (σ 2 t)2
= e− σ2 + σ2 + 2σ 2
+
2σ 2

σ 2 t2
= eµt+ 2

σ 2 t2
MX0 (t) = µ + σ 2 t eµt+ 2


15.8 PARETO DISTRIBUTION

The probability density function is given by

kβ k
f (x) = , k > β, x > 0
xk+1

49
Distribution Theory II: STA 412 Dr. [Link]

By general approach moment



kβ k
Z
r
E(x ) = xr . k+1 dx
β x
Z ∞
= kβ k xr .x−k−1 dx
Zβ ∞
= kβ k xr−k−1 dx
β

∞
xr−k

k
= kβ
r−k β
k

= [0 − β r−k ]
r−k
kβ k
 
1
=
k − r β k−r
kβ k
 
1
=
k − r β k β −r
k
= (β −r )−1
k−r
kβ r
E(xr ) =
k−r

Thus, the mean of the pareto distribution is obtained as


E(x) = since r = 1
k−1

Also the
E(x2 ) = kβ 2 k − 2

50
Distribution Theory II: STA 412 Dr. [Link]

Then var(x)implies

kβ 2 kβ
var(x) = −( )
k−2 k−1
kβ 2 (k − 1)2 − k 2 β 2 (k − 2)
=
(k − 1)2 (k − 2)
kβ (k − 2k + 1) − k 3 β 2 − 2k 2 β 2
2 2
=
(k − 1)2 (k − 2)
k 3 β 2 − 2k 2 β 2 + kβ 2 − k 3 β 2 − 2k 2 β 2
=
(k − 1)2 (k − 2)
kβ 2
V (x) =
(k − 1)2 (k − 2)

16 CUMMULANT GENERATING FUNCTION


The cummulant generating function is denoted by

CX (t) = lnMX (t) ⇒ MX (t) = eCX (t)

If CX (t) were known , then


d M 0 (t)
C(t) =
dt M (t)

d2 M 00 (t)M (t) − [M 0 (t)]2


C(t) =
dt [M (t)]2
0
d M (0)
⇒ C(t)|t=0 = =µ
dt M (0)
d2 M 00 (t)M (0) − [M 0 (0)]2
C(t)|t=0 = = M2 − M12 = σ 2 since M (0) = 1
dt [M (0)]2

EXAMPLE 1
Given
1
M (t) = (1 + et )2 ,
4
find the CGF.
solution:

51
Distribution Theory II: STA 412 Dr. [Link]

1
M (t) = (1 + et )2
4
C(t) = lnM (t)
1 1 1
= ln (1 + et )2 = ln + ln(1 + et )2 = ln + 2ln(1 + et )
4 4 4
t
d 1 2e
= C(t) = .2et =
dt 1 + et (1 + et )
d
= C(t)|t=0 = 1 = µ
dt
d2 (1 + et )2et − 2et et
= C(t) =
dt (1 + et )2
2et
σ2 =
(1 + et )2
d2
⇒ C(t)|t=0
dt
2
=
4
1
σ2 =
2

EXAMPLE 2
t
Find the CGF if MGF = e2(e −1)
solution:

t
MX (t) = e2(e −1)
t
C(t) = lne2(e −1) = 2et − 2
d
= C(t) = 2et
dt
d
= C(t)|t=0 = 2 = µ
dt
d2
= C(t : |t=0 = 2et |t=0 = 2 = σ 2
dt

52
Distribution Theory II: STA 412 Dr. [Link]

EXAMPLE 3

σ 2 t2
Given MX (t) = eµt+ 2 Find the CGF.
solution:
since
σ 2 t2
MX (t) = eµt+ 2

C(t) = lnMX (t)


σ 2 t2
= lneµt+ 2

d
C(t)|t=0 = µ + σ 2 t|t=0 = µ
dt
d2
= C(t)|t=0
dt
= σ 2 |t=0

= σ2

EXAMPLE 4

Given MX (t) = (P et + q)n , f ind the CGF

Solution :

C(t) = lnMX (t) = ln(P et + q)n = nln(pet + q)


d n
C(t)|t=0 = t
.pet |t=0
dt pe + q
np
=
p+q
= np

= µ
2
d (pet )npet − npet .pet
C(t)|t=0 = |t=0
dt (pet + q)2
np − np2
= = np(1 − p)
1
= npq

= σ2

53
Distribution Theory II: STA 412 Dr. [Link]

EXAMPLE 5
Given

t
MX (t) = eλ(e −1)
t
C(t) = lneλ(e −1)

= λet − λ
d
C(t) = λet
dt
d
C(t)|t=0 = λ = µ
dt
d2
C(t)|t=0 = λet |t=0
dt
= λ

= σ2

17 MARKOV INEQUALITY
If X is a random variable and U (X) is a non- negative real- value function, then for any positive
constant
E(U (X))
C > 0, P (U (X) ≥ C) ≤
C
PROOF:
If A = {x|u(x) ≥ c} then for a continuous random variable
Z ∞
E(u(x)) = u(x)f (x)dx
Z−∞ Z
= u(x)f (x) + u(x)f (x)dx
ZA Ac

≥ u(x)f (x)
ZA

≥ f (x)dx
A
= cp(xA)

= cp(u(x) ≥ c)

54
Distribution Theory II: STA 412 Dr. [Link]

18 CHEBYSHEV INEQUALITY
If the random variable X has a mean µ and variance σ 2 then for every

1
k ≥ 1, P (|x − µ| ≥ kσ) ≤
k2

PROOF:
Let f (x) denote the pdf of X , then

σ 2 = E(x − µ)2
X
= (x − µ)2 f (x)
xR
X X
= (x − µ)2 f (x) + (x − µ)2 f (x)
xA xAc
A = {X : |x − µ| ≥ kσ}
X
⇒ σ2 ≥ (x − µ)2 f (x)
xR
since in A, |x − µ| ≥ kσ

By substitution,
X
⇒ σ2 ≥ (kσ)2 ≥ kσ
xA
X
= k2σ2 f (x)
xA
X
But f (x) = P (xA)
xA
σ 2 ≥ k 2 σ 2 P (xA)

σ 2 ≥ k 2 σ 2 P (|x − µ| ≥ kσ)

1 ≥ k 2 P (|x − µ| ≥ kσ)
1
≥ P (|x − µ| ≥ kσ)
k2
1
P (|x − µ| ≥ kσ) ≤
k2

55
Distribution Theory II: STA 412 Dr. [Link]

and

ε = kσ,

then
ε
k =
σ
σ2
P (|x − µ| ≥ kσ) ≤ 2 .
ε

19 CENTRAL LIMIT THEOREM


The central limit theorem states thet if X̄ is the mean of random samples, X1 , X2 , ..., Xn of size
n from a distribution with a finite mean µ and finite positive varianceσ 2 then the distribution
x̄−µ
of W = √σ
is N(0,1)in the limit as n → ∞
n

PROOF:

Let Mx (t) = M (t)

Ms (t) = (M (t))n

where s = Σxi

MW (t) = E(etw )
  
t Σxi − nµ
= E e √
σ n
  
t s − nµ
= E e √
σ n
 
t s nµ
= E e √ − √
σ n σ n
√ h st i
− µtσ n √
= e E eσ n
√ h tx in
µt n √
= e− σ E e σ n

  n
− µtσ n t
= e M √
σ n

56
Distribution Theory II: STA 412 Dr. [Link]

T aking the log of both sides,


√  
µt n t
lnMW (t) = − + n ln M √
σ σ n
2 3
x x
Since, ln x = 1+x+ + + ···
2! 3!
M 00 (0)t2
MW (t) = 1 + M 0 (0)t + + ...
2!
Ex2 t2
= 1+µ+ + ...
2
(σ 2 + µ2 )t2
= 1+µ+ + ···
√ 2
(µ2 + σ 2 )t2

µt n µt
ln MW (t) = − + n ln 1 + √ + + ....
σ σ n σ2n
x 2 x3
But ln (1 + x) = x− + − ···
2 3

√ "
2
#
µt n µt (µ2 + σ 2 )t2 1 µt (µ2 + σ 2 )t2
⇒ lnMW (t) = − +n √ + − [ √ + + ···
σ σ n σ2n 2 σ n ]
√ √
nµt nµt (µ2 + σ 2 )t2 µ2 t2 µt3 (µ2 + σ 2 )
= − + + − − √ + ···
σ nσ σ2n 2σ 2 2 nσ 3
t2 µt3
= − 2
2 (µ + σ 2 )
t2
= since ∼ N (0, 1) and n → ∞
2

20 F-DISTRIBUTION
Consider two independent chi-square random variables U and V having r1 and r2 degrees of
U/r1
freedom respectively; Then F = V /r2
has an F- distribution with r1 and r2 degrees of freedom.
Its pdf is defined as folow
h(u,v)=
r1 r2 u+v
1 r1
−1 −u 1 r2
−1 − v2 u 2 −1 v 2
−1
e− 2 u/r1
r1 u 2 e2 . r2 v 2 e = r1 +r2 f=
Γ( r21 )2 2 Γ( r22 )2 2 Γ( r21 )Γ( r22 )2 2 v/r2

57
Distribution Theory II: STA 412 Dr. [Link]

u/r1 f zr1 du zr1 du f r1


Let v = z; f = z/r2
;u = r2
; df = ;
r2 dz
= r2

zr1 f r1
zr1
|J| = r2 r2 =
0 1 r2
T hen,
 r21 −1 h
f zr1
i
− 12
 r2
h i
f zr1 −1 +z zr1
r2
z 2 e r2
r2
g(f, z) = r1 +r2
Γ( r21 )Γ( r22 )2 2

  r21 r1 r1 +r2 z
h
f r1
i
r1 −1 −1 +1
r2
f 1 z 2 e 2 r2

= r1 +r2
Γ( r21 )Γ( r22 )2 2
Z ∞
g(f ) = g(f, z)dz
−∞
  r21 h
f r1
i
r1 r1
−1
r1 +r2
−1 − z2 +1
Z ∞
r2
f 1 z 2 e r2

= r1 +r2 dz
Γ( r21 )Γ( r22 )2 2
0
   
z f r1 dy 1 f r1 2y 2dy
Let y = +1 ; = + 1 ; z = f r1 ; dz = f r1
2 r2 dx 2 r2 r
+ 1 r2
+1
2

By substitution we have

  rr1   r1 +r
2
2 −1
r2
r1 −1 2y
Z ∞ r2
2
f 2
f r1 e−y . f r12
2
+1 2
+1
g(f ) = dy r1 +r2
0 Γ( r21 )Γ( r22 )2 2

  rr1 r
r1 2 2 −1 r1 +r2 −1
r2
f 2 2 2 2 Z ∞
r1 +r2
−1 −y
g(f ) = r1 +r2
r +r
 1 2 2 f r1 −1 f r1 y 2 e dy
r1 r2 f r

Γ( 2 )Γ( 2 )2 2 1
+ 1 + 1 + 1 0
2 2 2
  rr1 r r1 +r2
2
r1
r2
2
f 2 −1 2 2 2−1 .2 r1 + r2
= r1 +r2 .Γ( )
r1 r2 r1 +r2
f r1
 2 2
Γ( 2 )Γ( 2 )2 2
2
+1
  rr1 r
2
r1
r2
2
f 2 −1 r1 + r2
= r1 +r2 .Γ( ); 0 < f < ∞
r1 r2 f r1
 2 2
Γ( 2 )Γ( 2 ) 2 + 1

58
Distribution Theory II: STA 412 Dr. [Link]

21 TRANSFORMATION OF RANDOM VARIABLES


Let u(x) be a real value function of a random variable x. If the equation y = u(x) can be
uniquely solved , say x = w(y), then we have a one-to-0ne transformation.
Let X has the binomial pdf

2 3 1 3−x

3!
 
 x!(3−x)!
 3 3
, x = 0, 1, 2, 3, . . .
f (x) =

0 elsewhere.

Find the pdf of y = x2


Solution:

y = x2 = u(x)

x = w(y)

= y
3! 2 1 √
g(y) = √ √ ( )3 ( )3− y , y = 0, 1, 4, 9, . . .
y!(3 − y)! 3 3
0 elsewhere.

Let X be a random variable of the continuous type having pdf f(x). Let the one-dimensional
space where f (x) > 0. Consider the random variable y = u(x), where y = u(x) defines a one-
to-one transformation that maps the set onto the set B.
Let the inverse ofy = u(x) be denoted by x = w(y) and let the derivative dx
dy
= w0 (y) be contin-
uous and equal to zero for all point y in B.
Then the pdf of the random variable Y = U (X)is given by g(y)f (w(y))|w0 (y)|, yB
0 elsewhwere.
Where dx
dy
= w0 (y) is the jacobian (denoted by J) of the transformation.

2x
 0<x<1
Let f (x)

0, elsewhere.

Define the random variable of Y = 8X 3

59
Distribution Theory II: STA 412 Dr. [Link]

SOLUTION:

y = 8x3
y
x3 =
8
y 1
x = ( )3
8
1
y3
= 1
83
1 1
= y3
2
⇒ w(y) T hen
1 1
f (w(y)) = f ( y 3 )
2
1 1
= 2 y3
2
1
= y3

To obtain |J|
dx
|J| =
dy
1 −2
= y 3
6
g(y) = (f (w(y)))|J|
1 1 2
= y 3 . y− 3
6

1 −1
= y 3
6
T he Support is as f ollow.
1 1
0 < y3 < 1
2
1
0 < y3 < 2

0<y<8

0, elsewhere

60
Distribution Theory II: STA 412 Dr. [Link]

Let the pdf


f (x) = 1, 0 < x < 1 0 elsewhere.

Find the pdf of


Y = −2lnX

SOLUTION:

y = u(x) = −2lnx
y
− = lnx
2
y
e− 2 = x
y
x = w(y) = e− 2
dx 1 y 1 y
= − e− 2 = e− 2 = |J|
dy 2 2
f (w(y)) = 1
1 y
g(y) = f (w(y))|J| = e− 2
2
y
0 < e− 2 < 1

0 < y < ∞.

Let the pdf of X be defined by

x3
f (x) = , 0 < x < 2.
4

Find the pdf of


Y = X2

61
Distribution Theory II: STA 412 Dr. [Link]

SOLUTION:

y = x2

= u(x)

x = y

= w(y)
√ 3
y
f (w(y)) =
4
3
y 2
=
4
dx 1 −1
= y 2
dy 2
g(y) = f (w(y))|J|
3
y 2 1 −1 1
= . y 2 = y;
4 2 8

0< y<2

0<y<4

22 JOINT TRANSFORMATION
This method of finding the pdf of a function of one random variable of the continuous type will
be extended to functions of two random variables of this type.
Let y1 − u1 (x1 , x2 ) and y2 − u2 (x1 , x2 ) define a one-to-one transformation that maps a (two-
dimensional) set A in the x1 , x2 , plane unto a two-dimensional set B in the y1 , y2 - plane . If
we express x1 , x2 in terms of y1 , y2 we can make x1 = w(y1 , y2 ). The determinant is of the order 2

dx1 dx1
dy1 dy2
dx2 dx2
dy1 dy2

is called the jacobian of the transformation denoted by J .

Let X1 and X − 2 denote a random sample from a distribution χ2(2) .


Note that the pdf of X is

62
Distribution Theory II: STA 412 Dr. [Link]

1 r x
f (x) = x 2 −1 e− 2 , 0 ≤ x < ∞,
Γ(r/2)
then X has a chi-square distribution with r degrees of freedom abbreviated by saying X is χ2(r) .
Find (i) the joint pdf of Y1 = X1 and Y2 = X2 + X1 for 0 < y1 < y2 < ∞
(ii) the marginal distribution of each y1 and y2
(iii) Are Y1 and Y2 independent?
SOLUTION:

1 − x1 1 x2
(1)f (x1 ) = e 2 , f (x2 ) = e− 2
2 2
1 − x1 +x2
f (x1 )f (x2 ) = e 2
4
y1 = u(x1 , x2 ) = x1 then x1 = w(y1 , y2 ) = y1

y2 = u2 (x1 , x2 ) = x2 + x1 then x2 = w(y1 , y2 ) = y2 − y1


1 − y1 +y2 −y1 1 y2
f (w1 (y1 , y2 ))f (w2 (y1 , y2 )) = e 2 = e− 2
4 4
dx1 dx1
dy1 dy2
|J| = dx2 dx2
dy1 dy2

1 0
= =1
−1 1
1 − y2 1 y2
g(y1 , y2 ) = e 2 .1 = e− 2
4 4
(ii)
Z
g1 (y1 ) = g(y1 , y2 )dy2 , 0 < y1 < y2 < ∞
Z ∞
1 y2 2 y2 1 − y2
= e− 2 dy2 = − e− 2 |∞ y1 = e
2 ,0 < y < y
1 2
4 y1 4 2
Z
g2 (y2 ) = g(y1 , y2 )dy1 , 0 < y1 < y2 < ∞
Z y2
1 y2
= e− 2 dy1
4 0

63
Distribution Theory II: STA 412 Dr. [Link]

1 − y2 y2
= y1 e 2 |0
4
1 − y2
= y2 e 2
4

23 t-Distribution
W ∼ N (1, 0) i.e a standard normal distribution with mean 0 and variance 1
V ∼ χ2(r) is a chi-square distribution with d.f r
i.e.
1 −w2
W =√ e 2 −∞<w <∞

r −v
V 2 −1 e 2
V =  r 0<v<∞
Γ 2r 2 2
Let W and V be Independent; Then
W
T = pv
r
 −w2 r −1 −v
√1 e 2 V e 2
2
 r −∞<w <∞
 2π Γ( r2 )2 2

 0<v<∞

0 otherwise

To Proof

w
t = pv & u=v
r

t u
w = √
r

dw u
|J| = = √
dt v

64
Distribution Theory II: STA 412 Dr. [Link]

The joint p.d.f of T & U is given by


 √ 
u
g(t, u) = h t √ |J|
u r
r  √
U 2 −1 t2
 
−u u
= √ r exp 1 − √
2πΓ 2r 2 2

2 r r
The marginal p.d.f of T is
Z ∞
g(t) = g(t, u)du
−∞
r+1

U ( 2 −1) t2
  
−u
Z
= √  r exp 1+ du
0 2πrΓ 2r 2 2 2 r
h  2 i
U 1 + tr
Let Z = then
2
Z ∞ " #( r+1
2 )
−1 2
!

1 2z −z 2
1+ tr
g(t) = √ r 2 e dz
2πrΓ 2r 2 2 1 + tr

0
h i
(r+1)
Γ 2
= √ r+1 ∞<t<∞
2
πr Γ 2r 1 + tr 2


If w is N(0,1), If V is χ2(r) and W & V are Independent, Therefore


W
T =q
V
r

Normal Distribution Theory


24 The Multivariate Normal Distribution
we shall investigate a joint distribution of n random variables that will called multivariate nor-
mal distribution. This investigation assumes that the student is familiar with elementary matrix
algebra, with real symmetric quadratic forms, and with orthogonal [Link]
the expression quadratic form means a quadratic form in prescribed number of variables whose
matrix is real and symmetric. All symbols which represent matrices will be set in boldface type.

65
Distribution Theory II: STA 412 Dr. [Link]

Let A denote an n x n real symmetric matrix which is positive definite. Let µ denote the
n x 1 matrix such that µ0 , the transpose of µ, is µ0 =[µ1 , µ2 , . . . , µn ],where each µi is a real
[Link], let x denote the n x 1 matrix such that X 0 = [x1 , x2 , . . . , xn ].We shall show
that if C is an approximately chosen positive constant, the non-negative function

(X − µ)0 A(X − µ)
 
f (x1 , x2 , . . . , xn ) = C exp − ;
2
−∞ < x1 < ∞, i = 1, 2, . . . , n,

is a joint p.d.f of n random variable X1 , X2 , . . . , Xn that are of the continuous type show that
Z ∞ Z ∞
(1) ... f (x1 , x2 , . . . , xn )dx1 dx2 . . . dxn = 1
−∞ ∞

let t denote the n x 1 matrix such that t0 = [t1 , t2 , . . . , tn ], wheret1 , t2 , . . . , tn are arbitrary real
numbers. We shall evaluate the integral
Z ∞ ∞
(X − µ)0 A(X − µ)
Z  
0
(2) C ... exp t x − dx1 , . . . dxn ,
−∞ ∞ 2

and then we shall subsequently set t1 = t2 = · · · = tn = 0, and thus establish Equation(1).


First we change the variables of integration in integral (2) from x1 , x2 , . . . , xn to y1 , y2 , . . . , yn by
writing X − µ = y, where y 0 = [y1 , y2 , . . . , yn ]. The Jacobian of the transformation is one and
the n-dimensional x-space is mapped onto an n-dimensional y-space, so that integral (2) may be
written as
∞ ∞
y 0 Ay
Z Z  
0 0
(3) C exp (t µ) ... exp t y − dy1 , . . . , dyn
−∞ ∞ 2

Because the real symmetric matrix A is positive definite, the n characteristics numbers (proper
values, latent roots, or eigenvalues)a1 , a2 , . . . , an of A are positive. There exists an approximately
chosen n x n real orthogonal matrix L(L0 = L−1 , where L−1 is the inverse of L) such that
 
a1 0 . . . 0
 0 a2 . . . 0
L0 AL =  .. ..
 
.. 
. . .
0 0 ... an

66
Distribution Theory II: STA 412 Dr. [Link]

For a suitable ordering of a1 , a2 , . . . , an . We shall sometimes write L0 AL = diag[a1 , a2 , . . . , an ].


In integral(3), we shall change the variables of integration from y1 , y2 , . . . , yn to z1 , z2 , . . . , zn by
writing y = Lz where z 0 =[ z1 , z2 , . . . , zn ]. The Jacobian of the transformation is the determinant
of the orthogonal matrix L, Since L0 L = In , where In is the matrix of order n, we have the
determinant | L0 L |= 1 and | L |2 = [Link] the absolute value of the Jacobian is [Link]
the n-dimensional y-space is mapped onto an n-dimensional z-space. The integral (3) becomes
Z ∞ Z ∞
z 0 (L0 AL)z
 
0 0
4) C exp (t µ) ... exp t Lz − dz1 , . . . , dzn .
−∞ ∞ 2
It is computationally convenient to write, momentarily, t0 L = W 0 where W 0 = [w1 , w2 , . . . , wn ].
Then !
n
X
exp [t0 Lz] = exp [W 0 z] = exp wi zi .
i

Moreover,  n 
ai zi2 
P
z 0 (L0 AL)z
 
i

exp − −
= exp  
2 2 

. Then the integral(4) may be written as the product of n integrals in the following manner:

n Z ∞
ai zi2
Y   
0 0
5) C exp (W L µ) exp wi zi − dzi
i=1 −∞ 2
  
r Z ∞ exp w z − ai zi2

n
 2π
Y i i 2
= C exp (W 0 L0 µ) p dzi  .
i=1
a i −∞ 2π/a i

The integral that involves zi can be treated as the moment-generating function, with the more
familiar symbols t rel replaced by wi , of a distribution which is n(0,1/a).Thus the right-hand
member of equation(5) is equal to

n r
wi2
 
0 0
Y 2π
6) C exp (W L µ) exp
i=1
ai 2ai

67
Distribution Theory II: STA 412 Dr. [Link]

s !
n
(2π)n X wi2
= C exp (W 0 L0 µ) exp .
a1 a2 . . . an 1
2a
Now, because L−1 = L0 , we have


0 0 −1 1 1 1
(L AL) = L A L = diag , ...., .
a1 a2 an
n
wi2
= W 0 (L0 A−1 L)W = (LW ) A−1 (LW ) = t0 A−1 t. Moreover, the determinant
P
Thus, ai
1
A−1 of A−1 is
1
A−1 = L0 A−1 L = .
a1 a2 . . . an
Accordingly, the right-hand member of equation (6), which is equal to integral (2) may be written
as
−1
!
t0A t
q
t0
7) Ce µ (2π)n A−1 exp
2

If, in this function, we set t1 = t2 = · · · = tn = 0, we have the value of the left-hand members of
Equation (1). Thus, we
q
C (2π)n A−1 = 1
h 0
i
Accordingly, the function f (x1 , x2 , . . . , xn ) = 1
r exp − (X−µ) 2A(X−µ) ,
A−1
(2π)n/2

−∞ < xi < ∞, i = 1, 2, . . . , n is a joint p.d.f. of n random variables X1 , X2 , . . . , Xn that are of


the continuous type. such a p.d.f. is called a nonsingular multivariate normal p.d.f.
We have now proved f (x1 , x2 , . . . , xn ) is a p.d.f. However, we have proved more that. Because
f (x1 , x2 , . . . , xn ) is a p.d.f., integral(2) is the moment-generating function of the multivariate
normal distribution is given by

t0 A−1 t
 
0
M (t1 , t2 , . . . , tn ) = exp t µ + .
2

is the moment-generating function of Xi , i = 1, 2, . . . , n. Then

σi it2i
 
M (0, . . . , 0, ti , 0, . . . , 0) = exp ti µi +
2

68
Distribution Theory II: STA 412 Dr. [Link]

is the moment-generating function of Xi , i = 1, 2, . . . , n. Thus, Xi is n(µi , σii ),i = 1, 2, . . . , n.


Moreover, with i 6= j, we see that M (0, . . . , 0, ti , 0, . . . , 0, tj , 0, . . . , 0), the moment-generating
function of Xi and Xj , is equal to

σii t2 i + 2σij ti tj + σjj t2j


 
exp ti µi + tj µj +
2

But this is the moment-generating function of bivariate normal distribution, so that aσij is the
covariance of the random variables X − i and Xj . Thus the matrix µ, where µ0 = [µi , µ2 . . . , µn ]
is the matrix of the means of the random variables X1 , . . . , Xn . Moreover, the elements on
the principal diagonal of A−1 are, respectively, the variances σii = σi2 , i = 1, 2, . . . , n, and the
elements not on the principal diagonal of A−1 are, respectively, the covariances σij = ρij σi , σj , i 6=
j, of the random variables X1 , X2 , . . . , Xn . We call the matrix A−1 ,which is given by
 
σ11 , σ12 , . . . , σ1n
 σ12 , σ22 , . . . , σ2n 
..  ,
 
 .. ..
 . . . 
σ1n , σ2n . . . , σnn

the covariance matrix of the Multivariate normal distribution and henceforth we shall denote
this matrix by symbol V, In terms of the positive definite covariance matrix V, the multivariate
normal p.d.f. is written

(X − µ)0 V −1 (X − µ)
 
1
q exp − , −∞ < x1 < ∞,
(2π)n/2 V 2

Where i = 1, 2 . . . , n, and the moment- generating function of this distribution is given by

t0 V t
 
0
exp t µ +
2

for all real values of t.

Example [Link] X1 , X2 , . . . , Xn have a multivariate normal distribution with matrix µ of


means and positive definite covariance matrix V. if we let X 0 = [X1 , X2 , . . . , Xn ], then the

69
Distribution Theory II: STA 412 Dr. [Link]

moment-generating function M(t1 , t2 , . . . , tn ) of this joint distribution of probability is


t0 V t
 
t0X 0
8) E(e ) = exp t µ + .
2
n
consider a linear function Y of X1 , X2 , . . . , Xn which is defined by Y = c0 X =
P
ci Xi , where
1
C 0 c = [c1 , c2 , . . . , xn ] and the several ci are real and not zero. We wish to find the p.d.f of Y.
The moment-generating function M(t) of the distribution of Y is given by
0
M (t) = E(etY ) = E(ete X ).

Now the expectation (8) exists for all real values of [Link] we can replace t0 in expectation(8)
by tc0 and obtain
c0 V Ct2
 
0
M (t) = exp tc µ +
2

Thus the random variable Y is n (c0 µ, c0 V C).

Exercies
1. Let X1 , X2 , . . . , Xn have a multivariate normal distribution with positive definite covariance
matrix V. prove that these random variable are mutually stochastically independent if and
only if V is a diagonal matrix.

2. Let n=2 and take  


σ12 ρσ1 σ2
V =
ρσ1 σ2 σ22
Determine | V |, V −1 , and (X − µ). compare the bivariate normal p.d.f with the multi-
variate normal p.d.f. when n=2

25 The Distribution of certain Quadratic Forms


Let Xi , i = 1, 2, 3 . . . , n ,denote mutually stochastically independence random variable with are
n
n(µi , σi2 ) , i = 1, 2, . . . , n., respectively. Then Q= (Xi − µ)2 /σi2 is χ2 (n) .Now Q is a quadratic
P
1

70
Distribution Theory II: STA 412 Dr. [Link]

from in the Xi − µi and Q is seen to be, apart from the coefficient − 21 ,the random variable
which is defined by the exponent on the number e in the joint p.d.f. of X1 , X2 , . . . , Xn .We shall
now show that this result can be generalized.

Let X1 , X2 . . . , Xn have a multivariate normal distribution with p.d.f.

(X − µ)0 V −1 (X − µ)
 
1
q exp − ,
(2π)n/2 V 2

Where, as usual, the covariance matrix V is positive definite. We shall show that the random
variable Q (a quadratic form in the Xi − µi ), which is defined by (X − µ)0 V −1 (X − µ), is χ2 (n).
We have for the moment-generating function M(t) of Q the integral

∞ ∞
(X − µ)0 V −1 (X − µ)
Z Z  
1 0 −1
... q X exp t(X − µ) V (X − µ) − dx1 . . . dxn
−∞ −∞ (2π)n/2 V 2

∞ ∞
(X − µ)0 V −1 (X − µ)(1 − 2t)
Z Z  
1
= ... q X exp − dx1 . . . dxn
−∞ −∞ (2π)n/2 V 2

With V −1 positive definite, the integral is seen to exist for all real values of t < 12 . Moreover,
(1 − 2t)n | V −1 |, it follows that
(X − µ)0 V −1 (X − µ)(1 − 2t)
 
1
p exp −
(2π)n/2 | V | /(1 − 2t)n 2

can be treated as a multivariate normal p.d.f. If we multiply our integrand by (1 − 2t)n/2 , we


have this multivariate p.d.f. Thus the moment-generating function of Q is given by
1 1
M (t) = , t< ,
(1 − 2t)n/2 2
and Q is (χ2 n), as we wished to show.

The remarkable fact that the random variable which is defined by (X − µ)0 (X − µ)0 is (χ2 n)
stimulates a number of questions about quadratic forms in normally distributed variables. We

71
Distribution Theory II: STA 412 Dr. [Link]

should like to treat this problem in complete generally, but limitations of space forbid this, and
we find it necessary to restrict ourselves to some special cases.
Let X1 , X2 , . . . , Xn denote a random sample of size n from a distribution which is n(0, σ 2 )
,σ 2 > 0 . Let X 0 = [X1 , X2 , . . . , Xn ] and let A denote an arbitrary n X n real symmetric matrix.
We shall investigate the distribution of the quadratic from X 0 AX. For instance, we know that
n
X 0 In X/σ 2 = X 0 X/σ 2 = Xi2 /σ 2 is χ2 (n) .First we shall find the moment-generating function
P
1
of X 0 AX/σ 2 , Then we shall investigate the conditions which must be imposed upon the real
symmetric matrix A if X 0 AX/σ 2 is to have a chi-square [Link] moment-generating
function is given by

∞ ∞n  0
X 0X
Z Z  
1 tX AX
M (t) = ... √ exp − dx1 , . . . , dxn
−∞ −∞ σ 2π σ2 2σ 2
Z ∞ Z ∞ n
X 0 (I − 2tA)X
 
1
= ... √ exp − dx1 , . . . , dxn ,
−∞ −∞ σ 2π 2σ 2

Where I = In The matrix I-2tA is positive definite if we take | t | sufficiently small say
| t |< h, h > 0. Moreover, we can treat

X 0 (I − 2tA)X
 
1
p exp −
(2π)n/2 | (I − 2tA)−1 σ 2 | 2σ 2

as a multivariate normal p.d.f. Now | (I − 2tA)−1 σ 2 |1/2 = σ n / | (I − 2tA)−1 |1/2 . If we multiply


our integrand by | (I −2tA)−1 |1/2 , we have this multivariate p.d.f. Hence the moment-generating
function of X 0 AX/σ 2 is given by

(1) M (t) =| (I − 2tA)−1 |1/2 , | t |< h

It proves useful to express this moment-generating function in a different form. To do this,


let a1 , a2 . . . , an denote the characteristics numbers of A and L denote an n X n orthogonal
matrix such that L0 AL = diag[a1 , a2 . . . , an ].Thus,

72
Distribution Theory II: STA 412 Dr. [Link]

 
(1 − 2ta1 ) 0 ... 0
 0 (1 − 2ta2 ) . . . 0 
L0 (I − 2tA)L = 
 
.. .. .. 
 . . . 
0 0 ... (1 − 2tan )
Then
n
Y
(1 − 2tai ) =| L0 (I − 2tA)L | = | I − 2tA |
i=1

Accordingly we can write M(t), as given in Equation(1), in the form


" n
#−1/2
Y
(2) M (t) = (1 − 2tai ) , | t |< h
i=1

let r, 0 < r ≤ n, denote the rank of the real symmetric matrix [Link] exactly r of the
real numbers a1 , a2 . . . , an , say a1 , a2 . . . , ar are not zero and exactly n-r of these numbers, say
ar+1 , . . . , an , are zero. Thus we can write thr moment-generating function of X 0 AX/σ 2 as

M (t) = [(1 − 2ta1 )(1 − 2ta2 ) . . . (1 − 2tar )]−1/2

Now that we have found, in suitable form, the moment-generating function variable, let us
turn to the equation of the conditions that must be imposed if X 0 AX/σ 2 is to have a chi-square
distribution. Assume that X 0 AX/σ 2 is χ2 (k). Then

M (t) = [(1 − 2ta1 )(1 − 2ta2 ) . . . (1 − 2tar )]−1/2 = (1 − 2t)−k/2 ,

or, equivalently .

(1 − 2ta1 )(1 − 2ta2 ) . . . (1 − 2tar ) = (1 − 2t)−k/2 | t |< h

Because the positive integers r and h are the degree of these polynomials, and because these
polynomials are equal for infinitely many values of t, we have K = r,the rank of [Link],
the uniqueness of the factorization of a polynomial implies that a1 = a2 = · · · = ar = 1. If each
non-zero characteristics numbers of a real symmetric matrix is one, the matrix is idempotent,

73
Distribution Theory II: STA 412 Dr. [Link]

that is, A2 = A, and conversely. Accordingly if X 0 AX/σ 2 has a chi-square distribution , then
A2 = A and the random variable is χ2 (r) , where r is the rank of A. Conversely, if A has exactly
r characteristics numbers are equal to one, and the remaining n - r characteristics numbers that
are equal to zero. Thus the moment-generating function of X 0 AX/σ 2 is given by (1 − 2t)−r/2 ,
t < 12 , and X 0 AX/σ 2 is χ2 (r).this established the following theorem.

Theorem 1
. Let Q denote a random variable which is a quadratic form in the items of a random sample
of size n from a distribution which is n(0, σ 2 ). Let A denote the symmetric matrix of Q and let
r, 0 < r ≤ n, denote the rank of A. Then Q/σ 2 is χ2 (r) if and only ifA2 = A.

Remarks. If the normal distribution in Theorem 1 is n(µ, σ 2 ), the condition A2 = A


remains a necessary and sufficient condition that Q/σ 2 have a chi-square distribution, In general,
however,Q/σ 2 is not χ2 (r) but, instead Q/σ 2 has a non central chi-square distribution if A2 = A.
The number of degrees of freedom is r, the rank of A and the non centrality parameter is
µ0 Aµ/σ 2 , where µ0 = [µ, µ, . . . , µ] . Since µ0 Aµ = µ2
P
aij , Where a A = [ai j], then if µ 6= 0
i.j

, the condition A2 = A and aij = 0 are necessary and sufficient condition that Q/σ 2 be
P
i.j
central χ2 (r). Moreover,the theorem may be extended to a quadratic from in random variables
which have a multivariate normal distribution with positive definite covariance matrix V;here
the necessary and sufficient condition that Q have chi-square distribution is AVA = A.

Exercises
1. Let Q = X1 X2 − X3 X4 , where X1 X2 , X3 X4 , is a random sample of size 4 from a distri-
bution which is n(0,σ 2 ).Show that Q/σ 2 does not have a chi-square distribution. Find the
moment-generating function of Q/σ 2 .

2. Let A be a symmetric matrix. prove that each of the non-zero characteristic number of A
is equal to one if and only if A2 = A. Hint. LetL be an orthogonal matrix such that L0 AL

74
Distribution Theory II: STA 412 Dr. [Link]

is idempotent.

a2ij is equal to the sum of the


PP
3. Let A =[ai j] be a real symmetric matrix. prove that
j i
squares of the characteristic numbers of A,Hint. if L is an orthogonal matrix show that
aij = tr (A2 ) = tr (L0 A2 L) = tr [(L0 A2 L)(L0 A2 L)],
PP 2
j i

26 The independence of Certain Quadratic Forms


We have previously investigated the stochastic independence of linear functions of normally
distributed variables. In this section we shall prove some theorem about stochastic indepen-
dence of quadratic [Link] shall confine our attention to normally distributed variables that
constitute a random sample of size n from a distribution that is n(0, σ 2 ).

Let X1 , X2 , . . . , Xn denote a random variable sample of size n from a distribution which


is n(0, σ 2 ). Let A and B denote two real symmetric matrices, each of order n. Let X 0 =
[X1 , X2 , . . . , Xn ] and consider the two quadratic forms X 0 AX and X 0 BX. We wish to show
that these quadratic forms are stochastically independent if and only if AB = 0, the zero ma-
trix. We shall first compute the moment-generating function M(t1 , t2 ) of the joint distribution
of X 0 AX/σ 2 and X 0 BX/σ 2

n Z ∞ Z ∞
t1 X 0 AX t1 X 0 BX X 0X
  
1
M (t1 , t2 ) = √ ... exp + − dx1 . . . dxn
σ 2π −∞ −∞ σ2 σ2 2σ 2
n Z ∞ Z ∞
X 0 (I − 2t1 A − 2t2 B)X
  
1
= √ ... exp − dx1 . . . dxn
σ 2π −∞ −∞ 2σ 2

The matrix I-2t1 A − 2t2 B is positive definite if be take | t1 | and | t2 | sufficiently small, say
| t1 | < h1 | t2 | < h2 where h1 , h2 > 0. Then, we have

M (t1 , t2 ) | I − 2t1 A2t2 B |−1/2 ,| t1 | < h1 | t2 |

75
Distribution Theory II: STA 412 Dr. [Link]

Let us assume that X 0 AX/σ 2 and X 0 BX/σ 2 are stochastically independent (so that likewise
areX 0 AX and X 0 BX) and prove that AB=0. Thus we assume that

(1) M (t1 , t2 ) = M (t1 , 0)M (0, t2 )


for all t1 andt2 for which | t1 | < hi , i = 1, 2. Identity (1) is equivalent to the identity

(2) | I − 2t1 A − 2t2 B | = | I − 2t1 A | | I − 2t2 B |, | t1 | < hi i=1,2.

Let r> 0 denote the rank of A and let a1 , a2 . . . , ar denote the r nonzero characteristic
number of A. There exists an orthogonal matrix L such that
..
 
 a1 0 . . . 0 ..
. 
 0 a2 . . . 0

. 0 
 .. 
 . .. .. .. C 11 . 0
 ..

0 . . .
L AL =   = . . . . . . . . . = C
  

 0 .
.  ..
 0 . . . a r . 
 0 .
. . . . . . . . . . . . . . . . . .
0 0

for a suitable ordering of a1 , a2 . . . , ai Then L0 BL may be written in the identically parti-


tioned form  . 
D11 .. D12
L0 BL =  . . . . . . ... = D
 
.
D21 .. D22
The identity (2) may be written as

(20 ) | L0 | | I − 2t1 A 2t2 B | | L | = | L0 | | I − 2t1 A | | L | | L0 | I − 2t2 B | | L |,


or as

(3) | I − 2t1 C − 2t2 D | = | I − 2t1 C | | I − 2t1 D |.

The coefficient of (−2t1 )r in the right member of Equation(3) is seen by inspection to be


a1 a2 . . . ar | I −2t1 D |.It is not so easy to conceive of expanding this determinant in terms of
minors of order r formed from the first r columns. One term in this expansion is the product

76
Distribution Theory II: STA 412 Dr. [Link]

of the minor of order r in the upper left-hand corner, namely , | I − 2t1 C1 1 − 2t2 D11 |, and
the minor of order n-r in the lower right-hand corner, namely, | In−r − 2t2 D22 |.Moreover,
this product is the only term in the expansion of the determinant that involves (−2t1 )r .
Thus the coefficient (−2t1 )r in the left-hand members of Equation(3) is a a1 a2 . . . ar |
I − 2t1 C1 1 − 2t2 D11 |. If we equate these coefficients of (−2t1 )r , we have, for all t2 , | t2 |
< h2

(4) | I − 2t2 D | =| I − 2t1 C1 1 − 2t2 D22 |


Equation(4) implies that the nonzero characteristics numbers of the matrices D and D22
are the same. Recall that the sum of the squares of the characteristics numbers of a
symmetric matrix is equal to the sum of the squares of the elements of matrix. Thus the
sum of the elements of D22 .Since the elements of the matrix D are real,it follows that each
of the element of D22 .Since the elements of D11 , D12 , and D21 is zero. Accordingly, we can
write D in the form
..  
. 0 0
0
D = L BL = . . .
... ...
 
..
0 . D22
0 0 0
Thus CD = L ALL BL = 0 and L ABL = 0 and AB = 0, as we assume that AB = 0.
We are to show that X 0 AX/σ 2 X 0 BX/σ 2 are stochastically independent. We have, for all
real values of t1 and t2 ,

(I − 2t1 A)(I − 2t2 B) = I − 2t1 A − 2t2 B

Since AB = 0. Thus ,

| I − 2t1 A − 2t2 B |=| I − 2t1 A || I − 2t1 B |

Since the moment-generating function of the joint distribution X 0 AX/σ 2 and X 0 BX/σ 2 is
given by:
M (t1 , t2 ) =| I − 2t1 A − 2t2 B |−1/2 , | t1 |< hi i = 1, 2,

We have :
M (t1 , t2 ) = M (t1 , 0)M (0, t2 )

and the proof of the following theorem is complete.

77
Distribution Theory II: STA 412 Dr. [Link]

Theorem 2.

Let Q1 and Q1 denote random variables which are quadratic forms in the items of a random
sample of size n from a distribution which is n(0,σ 2 ).Let A and B denote respectively the
real symmetric matrices of Q1 and Q2 . The random variables Q1 and Q2 are stochastically
independent if and only if AB = 0.

Remark

Theorem 2 remain valid if the random samples is from a distribution which is n(µ, σ 2 ),
whatever be the real value of µ. Moreover, Theorem 2 may be extended to quadratic forms in
a random variables that have a joint multivariate normal distribution with a positive definite
covariance matrixV. The necessary and sufficient condition for the stochastic independence
of two such quadratic forms with symmetric matrices A and B then becomes AVB = 0.
In our Theorem 2, we have V = σ 2 I, so that AVB = Aσ 2 IB = σ 2 AB = 0

Theorem 3

Let Q = Q1 + · · · + Qk−1 + Qk , where Q1 , . . . , Qk−1 , Qk are K+1 random variables that


are quadratic forms in the items of a random sample of size n from a distribution which
is n(0,σ 2 ). Let Q/σ 2 be χ2 (r), let Q1 /σ 2 be χ2 (ri ), i = 1, 2, . . . , k − 1, and let Qk be non-
[Link] the random variables Q1 , . . . , Qk−1 , Qk are mutually stochastically indepen-
dent and hence, Qk /σ 2 is χ2 (rk = r − r1 − . . . − rk−1 )

proof : Take first the case of k=2 and let the real symmetric matrices of Q, Q1 , and Q2 be
denoted, respectively, by A, A1 , A2 . We are given that Q = Q1 + Q2 or, equivalently, that
A = A1 + A2 . We are also given that Q/σ 2 is χ2 and that Q1 /σ 2 is χ2 (ri ). In accordance
with Theorem 1, we have A2 = A and A2i . since Q2 ≥ 0,each of the matrices A, Ai andA2

78
Distribution Theory II: STA 412 Dr. [Link]

is positive [Link] A2 = A, we can find an orthogonal matrix L such that


 
0 I 0
L AL =
0 0

If then we multiply both members of A = A1 + A2 on the left by L0 and on the right by L,


we have  
I 0
= L0 A1 L + L0 A2 L
0 0
Now each of A1 and A1 and hence each of L0 A1 L and L0 A2 L is positive semidefinite. Recall
that, if a real symmetric matrix id positive semidefinite, each element on the principal
diagonal is postive or zero. Moreover, if an element on the principal diagonal is zero, then
all elements in the row and all elements in the column are zero. Thus L0 A1 L + L0 A2 L can
be written as

(5)      
Ir 0 Gr 0 Hr 0
= +
0 0 0 0 0 0

Since A2i = Ai , we have  


0 2 0 Gr 0
(L A1 L) = L A1 L =
0 0
If we multiply both members of equation (5) on the left by the matrix L0 A1 L,we have
     
Gr 0 Gr 0 Gr Hr 0
= +
0 0 0 0 0 0

or equivalently,L0 A1 L = L0 A1 L + (L0 A1 L) (L0 A2 L). Thus (L0 A1 L) X (L0 A2 L) = 0 and A1 A2


= 0. in accordance with theorem 2, Q1 and Q2 are stochastically independent. This stochastic
independence immediately implies that Q2 /σ 2 is χ2 (r2 = r − r1 ). This completes the proof when
k =2. For K ≥ 2, the proof may be made by induction. We shall merely indicate how this can
be done by using k =3. Take A = A1 + A2 + A3 where A2 = A, A21 = A1 ,A22 = A2 and A3 is
positive semidefinite. Write A = A1 + (A2 + A3 ) = A1 + B1 , say. Now A2 = A, A21 = A1 and B1
is positive semidefinite. In accordance with the case of k=2, we have A1 B1 = 0, so that B12 =
B1 . With B1 = A2 + A3 , where B12 = B1 , A22 = A2 , it follows from the case of k=2 that A2 A3
= 0 and A23 = A3 . If we regroup by writing A = A2 (A1 + A3 ), we obtain A1 A3 =0, and so on.

79
Distribution Theory II: STA 412 Dr. [Link]

Remark
In our statement of Theorem 3 we took X1 , X2 . . . , Xn to be items of a random sample from
distribution which is n(0, σ 2 ).We did this because our proof of Theorem 2 was restricted to that
case. In fact, if Q0 , Q01 . . . Q0k are quadratic forms in any normal variables (including multivariate
normal variables), if Q0 = Q01 + · · · + Q0k if Q0 , Q01 . . . Q0k−1 are central or non central chi-square,
and if Q0K is nonnegative, then Q01 . . . Q0k are mutually stochastically independent and Q0k is
either central or non central chi-square.

Theorem 4
Let X1 , X2 . . . , Xn denote a random sample from a distribution which is n(0, σ 2 ).Let the sum of
the squares of these items be written in the form
n
X
Xi2 = Q1 + Q2 + · · · + Qk ,
i

Q1 is a quadratic form in X1 , X2 . . . , Xn , with matrix Aj which has rank rj , j = 1, 2, . . . , k. The


random variables Q1 , Q, 2 . . . , Qk are mutually stochastically independent and Qj /σ 2 is χ2 (rj ),
k
P
j = 1, 2, . . . , k if and only if rj = n.
1

k k k
Xi2 =
P P P
proof: First assume the two conditions rj = n and Qj to be satisfied. The
1 1 1
latter equation implies that I = A1 + A2 + · · · + Ak . Let Bi = I − Ai . That is, Bi is the sum
of the sum of the matrices A1 , . . . , Ak exclusive of Ai . Let Ri denote the rank of Bi . Since
the rank of the sum several matrices is less than or equal to the sum of the ranks, we have
Pk
Ri ≤ rj − ri = n − ri . However, I = Ai + Bi , so that n ≤ ri + Ri and n − ri ≤ Ri . Hence
1
Ri = n − ri . The characteristics numbers of Bi are the roots of the equation | Bi − λI | = 0.
Since Bi = I - Ai , this equation can be written as | Bi − λI | = 0. Thus, we have | Ai − (1 − λ)I |
= 0. But each root of the last equation is one minus a characteristic number of Ai . since Bi
has exactly n − Ri = ri characteristic numbers that are zero, then Ai has exactly r − i charac-
teristics numbers that are equal to one. However, ri is the rank of Ai . Thus, each of ri nonzero

80
Distribution Theory II: STA 412 Dr. [Link]

characteristic numbers of Ai is one. That is, A2i = Ai and thus Qi /σ 2 is χ2 (ri ) , i = 1, 2, . . . , K .


In accordance with Theorem 3, the random variables Qi , Q2 . . . , Qk are mutually stochastically
independent.

n
Xi2 = Qi + Q2 + . . . , +Qk let Qi , Q2 . . . , Qk be
P
To complete the proof of Theorem 4, take
1
k
mutually stochastically independent, and let Qj /σ 2 be χ2 (ri ), j = 1, 2, . . . , k. Then Qj /σ 2 is
P
n  1
k n n
2
Qj /σ 2 = Xi2 /σ 2 is χ2 (n). Thus, rj = n and the proof is complete.
P P P P
χ rj . But
1 1 1 1

Exercise
(1) Let X1 , X2 , . . . , Xn denote a random sample of size n from a distribution which is n(0, σ 2 ).
n
prove that Xi2 and every quadratic form, which is non identically zero in X1 , X2 , . . . , Xn
P
1
are stochastically dependent.

(2) X1 , X2 , X3 , X4 denote a random sample of size 4 from a distribution which is n(0, σ 2 ). Let
4
ai Xi , where a1 , a2 , a3 and a4 are real constants. If Y 2 and Q = X1 X2 − X3 X4 are
P
Y=
1
stochastically independent, determine a1 , a2 , a3 and a4

Limit Theorem
27 Introduction
This chapter is principally concerned with the limiting behavior of the sum of independent
random variable as the number of summands becomes large. The results presented here are
both intrinsically interested and useful in statistics, since many commonly computed statistical
quantities, such as average, can be represented as sums.

81
Distribution Theory II: STA 412 Dr. [Link]

28 The Law Of Large Numbers


It is commonly believed that if a fair coin is tossed many times and the proportion of heads is
calculated, that proportion will be close to 12 . John Kerrich, a South African mathematician,
tested this belief empirically while detained as a prisoner during the world war II. He tossed a
coin 10,000 times and observed 5067 heads. The law of large numbers is a mathematical formu-
lation of this belief. The successive tosses of the coin are modeled as independent random trials.
The random variable Xi takes on the value 0 0r 1 according to whether the ith trial results ina
tail or head, and the proportion of the heads in n trials is

n
1X
X̄n = Xi
2 i=1
1
The law of large numbers sates that X̄n approaches 2
in a sense that is specified by the following
theorem.

82
Distribution Theory II: STA 412 Dr. [Link]

THEOREM A Law of large numbers

Let X1 , X2 , . . . , Xi . . . be a sequence of independent random variables with E(Xi )µ and


n
V ar(Xi ) = σ 2 .LetX̄n = n−1
P
Xi . Then, for any  > 0,
i=1
P ( X̄n − µ > ) −→ 0 as n −→ ∞
proof
We first find E(X̄n ) and V ar(X̄n ):

n
1X
E(X̄n ) = E(Xi ) = µ
n i=1
Since the Xi are independent ,

n
¯ 1 X σ2
V ar(Xn ) = 2 V ar(Xi ) =
n i=1 n
The desired result now follows immediately from Chebyshev’s inequality, which sates that
V ar(X¯n ) σ2
P ( X̄n − µ > ) ≤ 2
= n2
−→ 0, as n −→ ∞
In the case of a fair coin toss, the Xi are Bernoulli random variables with p=1/2, E(Xi ) = 1/2
and Var(Xi )=1/4. if tossed 10,000 times

¯ ) = 2.5 × 10−5
V ar(X10,000

and the standard deviation of the average is the square root of the variance, 0.005. The
proportion observed by Kerrich, 0.5067, is thus a little more than one standard deviation away
from its expected value of 0.5, consistent with Chebyshev’s inequality. Chebyshev’s inequality
can be written in the formP ( X̄n − µ > kσ) ≤ 1/k 2

if a sequence of random variable P (|Zn − α|) > ) approaches zero as n approaches infinity,
for any  > 0 and where α is some scale, then Zn is said to converge in probability to α.
There is another mode of convergence , called strong convergence or almost sure convergence,

83
Distribution Theory II: STA 412 Dr. [Link]

which asserts more Zn is said to converge almost surely to α if for every  > 0,,|Zn − α| > 
only a finite number of times with probability 1; that is, beyond some point is random. The
version of the law ot large numbers stated and proved earlier asserts that X̄n converges to µ
in [Link] version is usually called the weak law of large numbers. Under the same
assumptions, a strong law of large numbers, which assert that X̄n converges almost surely to µ,
can also be proved, bt we not do so.
We now consider some example that illustrate the utility of the law of large numbers.
EXAMPLE I
Monte Carlo Integration Suppose that we wish to calculate

Z 1
I(f ) = f (x)dx
0
where the integration cannot be done by elementary means or evaluated using tables of integrals.
The most common approach is to use a numerical method in which the integral is approximated
by a sum; Various schemes and computer packages exist for doing this. Another method, called
the Monte Carlo method, works in the following way. Generate independent uniform random
variables on [0,1]- that is, X1 , X2 , . . . , Xn - and compute
n
ˆ )= 1
X
I(f f (Xi )
n i=1
By the law of large numbers,this should be close to E[f, (Xi )], which is simply
Z 1
E[f (X)] = f (x)dx = I(f )
0

This simple scheme can be easily modified in order to change the range of integration and in
other ways. Compare to the standard numerical methods, it is not especially efficient in one
dimension, but becomes increasingly efficient as the dimensionality of the integral grows.
As a concrete example, let us consider the evaluation of

Z 1
a x2
I(f ) = √ e− 2 dx
2π 0
The integration is that of the standard normal density, which cannot be evaluation in closed form.
From the table of the normal distribution ,an accurate numerical approximation is I(f)=.3413.

84
Distribution Theory II: STA 412 Dr. [Link]

If 1000 points X1 , . . . , X1 000, uniformly distributed over the interval 0 ≤ x ≤ 1, are generated
using a pseudorandom number generator, the integral is then integrated by
 1000
X
ˆ )= 1 1 2
I(f √ e−Xi /2
1000 2π i=1

Which produce for one realization of the Xi the value.3417

85
Distribution Theory II: STA 412 Dr. [Link]

EXAMPLE II Repeated Measurements


Suppose that repeated independent unbiased measurements, X1 , . . . , Xn , of a quantity are made.
If n is large, the law of large number says that X̄ will be close to the true value, µ, of the quantity,
but how close X̄ is depends not only on n but on the variance of the measurement error, σ 2 , as
can be seen in the proof of Theorem A.
Fortunately, σ 2 can be estimated and therefore

σ2
V arX̄ =
n
n
can be estimated from the data assess the precision of X̄, First, note that n−1 Xi2 converges
P
i=1
to E(X 2 ), from the law of large number. Second, it can be shown that if Zn converges to α in
probability and g is a continuos function, then

g(Zn ) −→ g(x)

Which implies that

X̄ 2 −→ [E(X)]2
n
Finally, since n−1 Xi2 converges to E(X 2 ) and X̄ 2 converges to [E(X)]2 , with a little additional
P
i=1
argument it can be shown that
n
1X 2
Xi − X̄ 2 −→ E(X 2 ) − [E(X)]2 = V ar(X)
n i=1
n
More generally, it follows from the law of large numbers that the sample moments n−1 Xir ,
P
i=1
converge in probability to the moment of X,E(X r )
EXAMPLE III
A muscle or nerve cell membrane contains a very large number of channels;when open, these
channels allow ions to pass through. Individual channels seem to open and close randomly, and
it is often assumed that in an equilibrium situations the channels open and close independently
of each other and that only a very small fraction are open at any one time. Suppose then that
the probability of a channel is open is p, a very small number, that there are m channels in

86
Distribution Theory II: STA 412 Dr. [Link]

all, and that the amountof current flowing through and individual channel is c. The number of
channel open at a particular time is N, a binomial random variable with m trials and probability
p of success on each trial. The total amount of current is S = cN and can be measured. We
then have

E(S) = cE(N ) = cmp

V ar(S) = c2 mp(1 − p)

and
V ar(S)
= c(1 − p) ≈ c
E(S)

since p is small, Thus, through independent measurements,S1 , . . . , Sn . we can estimate E(S) and
Var(S) and therefore c, the amount of current flowing through a single channel, without knowing
how many channels are there.

29 Convergence in Distribution and the


Central Limit Theorem
In application, we often want to find P(a¡X¡b) when we do not know the cdf of X precisely; it
is sometimes possible to do this by approximating FX . The approximation is often arrived at
by some sort of limiting argument. The most famous limit theorem in probability theory is the
central limit theorem, which is the main topic of this section. Before discussing the central limit
theorem, we develop some introductory terminology, theory, and examples.

DEFINITION

Let X1 , X2 , . . . be a sequence of random variables with cumulative distribution functions


F1 , F2 , . . . and let x be a random variable with distribution function F. We say that Xn converges
in distribution to X if
lim Fn (x) = F (x)
n−→∞

87
Distribution Theory II: STA 412 Dr. [Link]

at every point at which F is continuous.


Moment generating function are often useful for establishing the convergence of distribution
functions. We know that a distribution function Fn is uniquely determine by its mgf, Mn . The
following theorem, which will give without proof, states that this unique determination holds
for limits as well.

THEOREM 1 Continuity Theorem

Let Fn be a sequence of cumulative distribution function with the corresponding moment


generating function Mn . Let F be a cumulative distribution function M. If Mn (t) −→ M (t) for
all t in an open interval containing zero, then Fn (x) −→ F (x) at all continuity points of F.

EXAMPLE 1

We will show that the poisson distribution can be approximated by the normal distribution
for large values of λ. Which shows that as λ increases, the probability mass function of the
poisson distribution becomes more symmetric and more shaped
Let λ1 , λ2 , . . . be an increasing sequence with λn −→ ∞, and let {Xn } be a sequence of poisson
random variables with the corresponding parameters. We know that E(Xn ) = V ar(Xn ) = λn .
if we wish to approximate the poisson distribution function by a normal distribution function ,
the normal must have the same mean and variance as the poisson does. In addition, if we wish
to approve a limiting result, we run into the difficulty that the mean and variance are tending to
infinity. This difficulty is dealt with by Standardizing the random variable-that is by letting

Xn − E(Xn )
Zn = p
V ar(Xn )
Xn λn
= √
λn

We then have E(Zn ) = 0 and V ar(Zn ) = 1, and we will show that the mgf of Zn converges to
the mgf of the standard normal distribution.

88
Distribution Theory II: STA 412 Dr. [Link]

The mgf of Xn is
et −1
MXn (t) = eλn

the mgf of Zn is

 
−t λn t
MZn (t) = e MXn √
λn
√ √
λn λn(et / λn−1)
= e−t e

It will be easier to work with the log of this expression.


√ √
logMZn (t) = −t λn + λn(et/ λn − 1)

xk
Using the power series expansion ex =
P
k!
, we see that
k=0

t2
lim logMZn (t) =
n−→∞ 2
or
2 /2
lim MZn (t) = et
n−→∞

The last expression is the mgf of the standard normal distribution.


We have that the standardized Poisson random variable converges in distribution to a standard
normal variable as λ approaches infinity. Practically, we wish to use this limiting result as a
basis for an approximation for large but finite values of λn . How adequate the approximation is
for λ = 100, say , is a matter of theoretical and/or empirical investigation. It turns out that the
approximation is increasingly good for large values of λ and that λ does not have to be all that
large.
EXAMPLE 2
A certain type of particle is emitted at a rate of 900 per hour. What is the probability that
more than 950 particle will be emitted in a given hour if the counts form a poisson process?

89
Distribution Theory II: STA 412 Dr. [Link]

Let X be a poisson random variable with mean 900. We find P(X¿950) by standardizing:
 
X − 900 950 − 900
P (X > 950) = P √ > √
900 900
 
5
≈ 1−φ
3
= .04779

Where φ is the standard normal cdf. For comparison, the exact probability is .04712

We now turn to the central limit theorem, which is concern with a limiting property of sums
of random variables. If X1 , X2 , . . . is a sequence of independent random variable with mean µ
and variance σ 2 , and if
n
X
Sn = Xi
i=1
We know from the law of large numbers that Sn /n converges to µ in probability. This followed
from the fact that
σ2
 
Sn 1
V ar = 2 V ar(Sn ) = −→ 0
n n n
The central limit theorem is concerned not with the fact that the ratio Sn /n converges yo µ
but with how it fluctuates around µ. To analyze these fluctuations, we standardize:
Sn − nµ
Zn = √
σ n
You should verify that Zn has mean 0 and variance 1. The central limit theorem states that
the distribution of Zn converges to the standard normal distribution
THEOREM II Central Limit Theorem
Let X1 X2 , . . . be a sequence of independent random variable having mean 0 and variance
σ 2 and the common distribution function F and moment generating function M define in a
neighborhood of zero. let
n
X
Sn = Xi
i=1
Then

 
Sn
lim P √ ≤x = φ(x). −∞<x<∞
n−→∞ σ n

90
Distribution Theory II: STA 412 Dr. [Link]

Proof


Let Zn = Sn /(σ n) we will show that the mgf of Zn tend to the mgf of standard normal
distribution. Since Sn is a sum of independent random variables.

Msn (t) = [M (t)]n

and

  n
t
MZn (t) = M √
σ n
M(s)has a Taylor series expansion about zero :

1
M (s) = M (O) + sM 0 (0) + s2 M 00 (0) + s
2
0 √
Where s /s2 −→ [Link] E(X)=0, M (O) = 0, and M 00 (0) = σ 2 . As n −→ ∞, t/(σ n) −→ o,
and    2
t 1 t
M √ = 1 + σ2 √ + n
σ n 2 σ n
wheren /(t2 /(nσ 2 )) −→ 0 as n −→ ∞. we thus have
n
t2

MZn (t) = 1 + + n
2n

It can be shown that if an −→ a, then


 an n a
lim 1+ =
n−→∞ n

From this result it follows that

2 /2
MZn (t) −→ et as n −→ ∞

where exp(t2 /2) is the mgf of the standard normal distribution.


Theorem II is one of the simplest version of the central limit theorem;there are many central
limit theorems of various degree of abstraction and generality. We have proved theorem II under
the assumption that the moment-generating function exist, Which is a rather strong assumption.

91
Distribution Theory II: STA 412 Dr. [Link]

By using characteristic functions instead, we could modify the proof so that it would only be
necessary that the first and second moment exist. Further generalization weaken the assumption
that the Xi have the same distribution and apply to linear combinations of independent random
variables. Central limit theorem have also been proved that weaken the independent assump-
tion and allow the Xi to be dependent but not to dependent. For practical purpose,especially
in statistics, the limiting result in it self is not of primary interest. Statisticians are more inter-
ested in its uses and approximation with finite value of [Link] is impossible to give a concise and
definitive statement of how good the approximation is, but some general guideline are available,
and examining special case can give in sight. How far the approximation becomes good depends
on the distribution of the summands, the Xi .If the distribution is fairly symmetric and has tails
that die off rapidly, the approximation becomes good and relatively small values of n. If the
distribution is very skewed or if the tail die down very slowly, a larger value of n is needed for a
good approximation. The following example deals with two special cases.

EXAMPLE
Suppose that X1 , . . . , Xn are repeated, independent measurement of quantity, µ, and that
E(Xi ) = µ and V ar(Xi ) = σ 2 . The law of large numbers tells us that X̄ converge to µ in
probability, so we can hope X̄ is close to µ if n is large. Chebyshev’s inequality allows us to
bound the probability of an error of a given size, but the central limit theorem gives a more
sharper approximation to the actual error. Suppose that we wish to find P (|X ¯− µ| < c) for
some constant c. To use the central limit theorem to approximate this probability, we first
standardize, using E(X) = µ and V arX̄ = σ 2 n

P (|X̄ − µ| < c) = P (−c < X − ¯µ < c)


 
−c X̄ − µ c
= P √ < √ < √
σ/ n σ/ n σ/ n
 √   √ 
c c c n
≈ φ −φ −
σ σ

For example, suppose that 16 measurements are taken with σ = 1. The probability that the

92
Distribution Theory II: STA 412 Dr. [Link]

average deviates from µ by less than 0.5 is approximately

P (|X̄ − µ| < 0.5) = φ(0.5 × 4) − φ(0.5 × 4) = 0.954

This can be turn around. That is, given c and γ, n can be found such that

P (|X̄ − µ| < c) ≥ γ

EXAMPLE: Normal approximation to Binomial Distribution


Since a binomial random variable is the sum of Bernoulli random variables, its distribution
can be approximated by a normal distribution. The approximation is best when the binomial
diatribution is symmetric-that is, when P = 12 . A frequently used rule of thumb is that the
approximation is reasonable when np¿5 and n(1-p)¿5. The approximation is especially useful
for large values of n, for which tables are not readily available. Suppose that a coin is toss 100
times and lands head up 60 times. Should will be surprised and doubt that the coin is fair ?
To answer this question, we note that if the coin is fair. the number of heads, X, is a binomial
1
random with n=100 trials and probability of success P = 2
, so that E(X) = np = 50 and
Var(X)=np(1-p)=25. We could calculate P(X=60), which will be a small number . But because
there are so many outcomes, P(X=50) is also a small number, so this calculation will not really
answer the question . Instead we calculate the probability of a deviation as extreme as or
more extreme than 60 if the coin is fair; that is we calculate P (X ≥ 60). To approximate this
probability from the normal distribution, we standardize:
 
X − 50 60 − 50
P (X ≥ 60) = P ≥
5 5
≈ 1 − φ(2)

= 0.0228

The probability is rather small, so the fairness of the coin is called into question.

EXAMPLE: Particle Size Distribution

93
Distribution Theory II: STA 412 Dr. [Link]

The distribution of size of grains of a particulate matter is found to be quite skewed, with
a slowly decreasing right tail. A distribution called the lognormal is sometimes fit to such
distribution, and X is said to follow a lognormal distribution if log X as a normal distribution.
The central limit theorem gives a theoretical rationale for the use of lognormal distribution in
some situations.
Suppose that a particle of initial y0 is subjected to repeated impact, that on each impact a
proportion, Xi , of the particle remains, and that the Xi are modeled as independent random
variable having the same distribution. After the first impact, the size of the particle is Y1 =
X1 y − 0; after the second impact, the size is Y2 = X2 X1 y0 ; and after the nth impact, the size is

Yn = Xn Xn−1 . . . X2 X1 y0

T hen
n
X
LogYn = Logy0 + logXi
i=1

and the central limit theorem apply to log Yn

A similar construction is relevant to the theory of finance. Suppose that an initial investment
of value v0 is made and that returns occur in discrete time, for example, daily. If the return on
the first day is R1 , then the value becomes V1 = R1 v0 . After day two the value is V2 = R2 R1 v0 ,
and after day n the value is

Vn = Rn Rn−1 . . . R1 v0

T he log value is
n
X
LogVn = logv0 + logRi
i=1

If the return are independent random variable with the same distribution, then the distribution
of log Vn is approximately normally distribution.

94
Distribution Theory II: STA 412 Dr. [Link]

Distribution Derived From the Normal Distribution


30 Introduction
These chapter assembles some results concerning three probability distribution derived fro the
normal distribution -the χ2 ,t, and F distribution. These distribution occur in many statistical
problem .

31 χ2, t, and F Distribution


DEFINITION
If Z is a standard normal random variable, the distribution of U = Z 2 is called the chi-square
distribution with 1 degree of freedom.
The chi-square distribution with 1 degree of freedom is denoted χ2 . It is useful to note that if
X ∼ N (X − µ)/σ ∼ N (0, 1) and therefore [(X − µ)/σ]2 ∼ χ21

DWFINITION
If ∪1 , ∪2 , . . . , ∪n are independent chi-square variables with 1 degree of freedom, the distribution
of V = ∪1 + ∪2 + · · · + ∪n is called the chi-square distribution with n degree of freedom and is
denoted by χ2n
We know that the sum of independent gamma random variables that have the same value of λ
follows a gamma distribution, and therefore the chi-square distribution with n degrees of freedom
1
is a gamma distribution with α = n and λ = 2
Its density is

1
f (v) = v (n/2)−1 e−v/2 , v ≤ 0
2n/2 Γ(n/2)
Its moment generating function is

M (t) = (1 − 2t)−n/2

Also, E(V ) = n and V ar(V ) = 2n . To indicate that V follows a chi-square distribution with
n degrees of freedom , we write V ∼ χ2n . A notable consequence of the definition of chi-square

95
Distribution Theory II: STA 412 Dr. [Link]

distribution is that if U and V are independent and U ∼ χ2n and V ∼ χ2m then U + V ∼ χ2m+n
we now turn to the t distribution.

DEFINITION
p
If Z ∼ N (0, 1) and ∪ ∼ χ2n and Z and ∪ are dependent, then the distribution of Z/ U/n is
called the t distribution with n degrees of freedom.

PROPOSITION I
The density function of a t distribution with n degrees of freedom is
−(n+1)/2
t2

Γ[(n + 1)/2]
f (t) = √ 1+
nπΓ(n/2) n
Proof
p
This is proved by a standard method. The density function of U/n is straight forward to ob-
tain, and the density function of the quotient of two independent random variable was derived
From the density of the proposition f (t) = f (−t), so the t distribution is symmetric about zero.
As the number of degrees of freedom approaches infinity , the t distribution tends to the stan-
dard normal distribution; in fact for more than 20 and 30 degrees of freedom, the distribution
are very close

DEFINITION
Let U and V be independent chi-square random variable with m and n degrees of freedom,
respectively. The distribution of
U/m
W =
V /n

PROPOSITION II
The density function of W is given by
Γ[(m + n)2]  m m/2 m/2−1  m −(m+n)/2
f (w) = w 1+ w , w≤0
Γ(m/2)Γ(n, 2) n n

96
Distribution Theory II: STA 412 Dr. [Link]

Proof
W is the ratio of two independent variables, and its density
It can be shown that, for n¿2, E(W) exist and equal n/(n − 2) from the definition of the t and
F distributions, it follows that the square of a tn random variable follows an F1.n distribution .

The Sample Mean and Sample Variance


Let X1 . . . , Xn be independent N (µ, σ 2 ) random variables. We sometime refer to them as a
sample from a normal [Link] this case, we will find the joint and marginal distribution
of
n
1X
X̄ = Xi
2 i=1
n
1 X
S2 = (Xi − X̄)2
n − 1 i=1

These are called sample mean and sample variance,respectively. First note that because X̄ is a
linear combination of independent normal random variables, it is normally distributed with

E(X̄) = µ
σ2
V ar(X̄) =
n

As a preliminary to showing that X̄ and S 2 are independently distributed. we establish the


following theorem

THEOREM
The random variable X̄ and the vector of random variables
(X1 − X̄, X2 − X̄, . . . , Xn − X̄) are independent.
Proof
At the level of this course, it is difficult to give a proof that provide sufficient insight into why
this result is true;a rigorous proof essentially depends on geometric properties of the multivariate

97
Distribution Theory II: STA 412 Dr. [Link]

normal distribution. We present a proof based on moment generating functions; in particular,we


will show that the joint moment-generating function

M (s, t1 , . . . , tn ) = Eexp[sX̄ + t1 (X1 barX) + · · · + tn (Xn − X̄)]

Factors into the product of two moment generating functions-one of the X̄ and the other of
(X1 − X̄), . . . (Xn − X̄). the factoring implies that the random variables are independent of each
other and is accomplish through some algebraic trickery. First we observe that since
n
X n
X
ti (Xi − X̄) = ti Xi − nX̄ t̄
i=1 i=1

then

n n h
X X s i
sX̄ + ti (Xi − X̄) = + (ti − t̄)
i=1 i=1
n
Xn
= ai Xi where
i=1
s
a1 = + (ti − t̄)
n
f urther we observe that
X n
ai = s
i=1
n n
X s2 X
a2i = + (ti − t̄)2
i=1
n i=1
now we have

M (s, t1 , . . . , tn ) = Mxi...Xn (a1 , . . . , an )

98
Distribution Theory II: STA 412 Dr. [Link]

and since the Xi are independent normal variables, we have


n
Y
M (s, t1 , . . . , tn ) = Mxi (ai )
i=1
n
σ2
Y  
= exp µai + a2i
i=1
2
nn
!
σ2 X 2
X
= exp µ ai a
i=1
2 i=1 i
" n
#
σ 2 s2 σ2 X
 
2
= exp µs + + (ti − t̄)
2 n 2 i=1
" n
#
σ2 2 σ2 X
 
2
= exp µs + s exp (ti − t̄)
2n 2 i=1

The first factor is the mgf of X̄. Since the mgf of the vector
(X1 − X̄, . . . , Xn − X̄) can be obtain by setting s=0 in M, the second factor is this mgf.

COROLLARY
X̄ and S 2 are independently distributed

Proof
This follow immediately since S 2 is a function of vector
(X − 1 − X̄, . . . Xn − X̄), which is dependent of X̄
The next theorem gives the marginal function of S 2

THEOREM
The distribution of(n − 1)S 2 /σ 2 is the chi-square distribution with n-1 degrees of freedom.
Proof
We first note that

99
Distribution Theory II: STA 412 Dr. [Link]

n  2
1 X 2 Xi − µ
(Xi − µ) = ∼ χ2n
σ 2 i=1 σ2
Also
n n
1 X
2 1 X
(Xi − µ) = [(Xi − X̄) + (X̄ − µ)]2
σ 2 i=1 σ 2 i=1
n
P
Expanding the square and using the fact that (Xi − X̄) = 0, we obtain
i=1

n n  2
1 X 2 1 X 2 X̄ − µ
(Xi − µ) = 2 (Xi − X̄) + √
σ 2 i=1 σ i=1 σ/ n
This is a relation of the form W=U+V. Since U and V are independent by corollary,Mw (t) =
Mu (t)Mv (t). W and V both follow chi-square distributions,so

Mw (t)
Mu (t) =
Mv (t)
(1 − 2t)−n/2
=
1 − 2t)−1/2
= 1 − 2t)−(n−1)/2

The last expression is the mgf of a random variable with a χ2n−1 distribution.
COROLLARY
Let X̄ and S 2 be given as the beginning of the first corollary. Then

X̄ − µ
√ ∼ tn − 1
S/ n
Proof
We simply express the given ratio in a different form:
 
X̄−µ

X̄ − µ σ/ n
√ =p
S/ n S 2 /σ 2
The later is the ratio of an N(0,1)random variable to the square root of an independent random
variable with a χ2n−1 distribution divided by its degrees of [Link] from the definition,the
ratio follows a t distribution with n-1 degrees of freedom.

100

You might also like