0% found this document useful (0 votes)
4 views28 pages

Multivariate Normal Distributions

Distributions Notes
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views28 pages

Multivariate Normal Distributions

Distributions Notes
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Multivariate Normal Distributions

Variance – Covariance Matrix —


The 𝑝 × 𝑝 symmetric matrix, Σ = (𝜎𝑖𝑗 )
𝑝×𝑝
𝜎11 𝜎12 …. 𝜎1𝑝
𝜎21 𝜎22 …. 𝜎2𝑝
= (…. …. …. ….)
𝜎𝑝1 𝜎𝑝2 …. 𝜎𝑝𝑝
𝑉 (𝑋1 ) 𝐶𝑜𝑣(𝑋1 𝑋2 ) …. 𝐶𝑜𝑣(𝑋1 𝑋𝑝 )
= 𝐶𝑜𝑣(𝑋2 𝑋1 ) 𝑉 (𝑋2 ) …. 𝐶𝑜𝑣(𝑋2 𝑋𝑝 )
…. …. …. ….
𝐶𝑜𝑣(𝑋𝑝 𝑋1 ) 𝐶𝑜𝑣(𝑋𝑝 𝑋2 ) …. 𝑉(𝑋𝑝 )
( ⸝
)
= 𝐸(𝑋̰ − 𝐸 (𝑋̰ ))(𝑋̰ − 𝐸 (𝑋̰ ))

= 𝐸(𝑋̰ − 𝜇̰ )(𝑋̰ − 𝜇̰ )
= 𝑉 (𝑋̰ ) is known as the variance – covariance matrix of the
vector X̰ or dispersion matrix of X̰ or simply the covariance matrix of X̰. It is denoted by
𝑉 (𝑋̰ ) or 𝐷 (𝑋̰ )

We have, 𝑉 (𝑋̰ ) = 𝐸(𝑋̰ − 𝜇̰ )(𝑋̰ − 𝜇̰ )
𝑋1 − 𝜇1
𝑋 −𝜇
= 𝐸 ( 2… . 2) (𝑋1 − 𝜇1 𝑋2 − 𝜇2 … . 𝑋𝑝 − 𝜇𝑝 )1×𝑝
𝑋𝑝 − 𝜇𝑝 𝑝×1
𝐸 (𝑋1 − 𝜇1)2 𝐸 (𝑋1 − 𝜇1)(𝑋2 − 𝜇2) …. 𝐸 (𝑋1 − 𝜇1)(𝑋𝑝 − 𝜇𝑝 )
𝐸 (𝑋2 − 𝜇2 )(𝑋1 − 𝜇1) 𝐸 (𝑋2 − 𝜇2)2 …. 𝐸 (𝑋2 − 𝜇2)(𝑋𝑝 − 𝜇𝑝 )
=
…. …. …. ….
2
(𝐸(𝑋𝑝 − 𝜇𝑝 )(𝑋1 − 𝜇1) 𝐸(𝑋𝑝 − 𝜇𝑝 )(𝑋2 − 𝜇2) …. 𝐸(𝑋𝑝 − 𝜇𝑝 ) )
Result :—
𝑉 (𝐴. 𝑋̰ ) = 𝐴. 𝑉 (𝑋̰ )𝐴⸝
𝑋1
𝑋
where 𝑋̰ 𝑝×1 = (…2.); 𝐴 = (𝑎𝑖𝑗 )𝑛×𝑝 → constant matrix
𝑋𝑝
Let 𝑌̰ = 𝐴𝑚×𝑝 𝑋̰ 𝑝×1
By definition, 𝑉 (𝑌̰ ) = 𝑉 (𝐴𝑋̰ )
= 𝐸 [𝐴𝑋̰ − 𝐸 (𝐴𝑋̰ )][𝐴𝑋̰ − 𝐸 (𝐴𝑋̰ )]⸝
= 𝐴𝐸 [𝑋̰ − 𝐸 (𝑋̰ )][𝑋̰ ⸝ 𝐴⸝ − 𝐸 (𝑋̰ )⸝ 𝐴⸝]
= 𝐴𝐸 [𝑋̰ − 𝐸 (𝑋̰ )][𝑋̰ ⸝ − {𝐸 (𝑋̰ )}⸝ ]𝐴⸝
= 𝐴𝐸 [𝑋̰ − 𝐸 (𝑋̰ )][𝑋̰ − 𝐸 (𝑋̰ )]⸝ 𝐴⸝
= 𝐴𝑉 (𝑋̰ )𝐴⸝
𝑋1 𝑌1
𝑋 𝑌
Similarly, if 𝑋̰ 𝑝×1 = ( 2 ); 𝑌̰𝑞×1 = (…2.)
….
𝑝 𝑌𝑞
1|D R C
𝐴𝑛×𝑝 = 𝑎𝑖𝑗 and 𝐵𝑛×𝑞 = 𝑏𝑖𝑗
Cov(𝐴𝑋̰ , 𝐵𝑌̰ ) = 𝐴𝐶𝑜𝑣 (𝑋̰ , 𝑌̰ )𝐵 ⸝
By definition, Cov(𝐴𝑋̰ , 𝐵𝑌̰ ) = 𝐸 [𝐴𝑋̰ − 𝐸 (𝐴𝑋̰ )][𝐵𝑌̰ − 𝐸 (𝐵𝑌̰ )]⸝
= 𝐸 [𝐴{𝑋̰ − 𝐸 (𝑋̰ )}][𝐵 {𝑌̰ − 𝐸 (𝑌̰ )}]⸝
= 𝐸 [𝐴{𝑋̰ − 𝐸 (𝑋̰ )}][{𝑌̰ − 𝐸 (𝑌̰ )}⸝𝐵 ⸝ ]
= 𝐴. 𝐸 [𝑋̰ − 𝐸 (𝑋̰ )][𝑌̰ − 𝐸 (𝑌̰ )]⸝ 𝐵 ⸝
= 𝐴𝐶𝑜𝑣(𝑋̰ , 𝑌̰ )𝐵 ⸝
Multivariate Normal Distribution —
The 𝑝 – component vector 𝑋̰ 𝑝×1 is said to follow multivariate (or more precisely 𝑝 variate)
normal distribution with (parameter) mean vector 𝜇̰𝑝×1 = 𝐸 (𝑋̰ ) and variance – covariance
matrix Σ = 𝑉 (𝑋̰ ) (assumed as positive definite) then its pdf is given by
1 ⸝
1 − (𝑥̰ −𝜇̰) Σ−1 (𝑥̰ −𝜇̰)
𝑓 (𝑥̰ ) = 𝑓(𝑥1 , 𝑥2 , … . , 𝑥𝑝 ) = 𝑝 1𝑒 2
(2𝜋)2 |Σ|2
and is written as 𝑥̰ ~ 𝑁𝑝 (𝜇̰ , Σ) or 𝑁(𝜇̰ , Σ)
Derivation of the 𝒑 – variate Normal Distribution —
Let X ~ N(μ, 𝜎 2 )
The density function of the univariate normal distribution is given by
1 𝑥−𝜇 2 1
1 − ( ) − 𝛼(𝑥−𝛽)2
𝑓(𝑥 ) = 𝑒 2 𝜎 = 𝑘. 𝑒 2
√2𝜋𝜎
1 1
writing 𝑘 = ,α= and μ = β
𝜎√2𝜋 𝜎2
1
− (𝑥−𝛽)𝛼(𝑥−𝛽)
We get 𝑓 (𝑥 ) = 𝑘𝑒 2 ……(1)
Now, let us suppose that we have 𝑝 random variables 𝑋1 , 𝑋2 , …., 𝑋𝑝 which are jointly
normally distributed. Then the density function of the multivariate normal distribution of
𝑋1 , 𝑋2 , …., 𝑋𝑝 (i.e., the joint density of 𝑋1 , 𝑋2 , …., 𝑋𝑝 has an analogous form.
𝑋1
𝑋2
The scalar variable X is replaced by a vector 𝑋̰ 𝑝×1 = (… .) of 𝑝 component.
𝑋𝑝
𝑏1
𝑏
The scalar constant β is replaced by a vector 𝑏̰ 𝑝×1 = (…2.) of constant element.
𝑏𝑝
The positive constant α is replaced by a positive definite symmetric matrix 𝐴𝑝×𝑝 = (𝑎𝑖𝑗 )𝑝×𝑝
𝑎11 𝑎12 … . 𝑎1𝑝
𝑎21 𝑎22 … . 𝑎2𝑝
= (…. …. …. ….)
𝑎𝑝1 𝑎𝑝2 … . 𝑎𝑝𝑝
The square 𝛼 (𝑥 − 𝛽 )2 is replaced by the quadratic form (𝑥̰ − 𝑏̰ )⸝ 𝐴(𝑥̰ − 𝑏̰ ) =
∑𝑝𝑖=1 ∑𝑝𝑗=1 𝑎𝑖𝑗 (𝑥𝑖 − 𝑏𝑖 )(𝑥𝑗 − 𝑏𝑗 )
Thus, the multivariate normal density function of 𝑋1 , 𝑋2 , …., 𝑋𝑝 is given by
1
− (𝑥̰ −𝑏̰ )⸝ 𝐴(𝑥̰ −𝑏̰ )
𝑓(𝑥1 , 𝑥2 , … . , 𝑥𝑝 ) = 𝑘𝑒 2 ……(2)
2|D R C
where 𝑘 (> 0) is chosen so that the integral of (2) over the entire 𝑝 dimensional space of 𝑋1 ,
1
∞ ∞ −2(𝑥̰ −𝑏̰ )⸝ 𝐴(𝑥̰ −𝑏̰ )
𝑋2 , …., 𝑋𝑝 is unity i.e., ∫−∞ … . ∫−∞ 𝑘𝑒 𝑑𝑥1 𝑑𝑥2
… . 𝑑𝑥𝑝 = 1 ……(3)
Since A is a positive definite matrix, therefore there exist a non-singular matrix 𝐶𝑝×𝑝 = (𝑐𝑖𝑗 )
such that 𝐶 ⸝𝐴𝐶 = I ……(4)
Let 𝑋̰ − 𝑏̰ = 𝐶𝑝×𝑝 𝑌𝑝×1 ……(5)
𝑦1
𝑦2
where 𝑌̰𝑝×1 = (… .)
𝑦𝑝
−1 (
or 𝑦 = 𝑐 𝑥̰ − 𝑏̰ )
or 𝑥1 − 𝑏1 = 𝑐11 𝑦1 + 𝑐12𝑦2 +. … + 𝑐1𝑝 𝑦𝑝
𝑥2 − 𝑏2 = 𝑐21 𝑦1 + 𝑐22 𝑦2 +. … + 𝑐2𝑝 𝑦𝑝
---------------------------------------------
𝑥𝑝 − 𝑏𝑝 = 𝑐𝑝1𝑦1 + 𝑐𝑝2 𝑦2 +. … + 𝑐𝑝𝑝 𝑦𝑝
𝛿𝑥1 𝛿𝑥1 𝛿𝑥1
….
𝛿𝑦1 𝛿𝑦2 𝛿𝑦𝑝
𝛿(𝑥1 ,𝑥2 ,….,𝑥𝑝 )
|𝛿𝑥2 𝛿𝑥2 𝛿𝑥2 |
Now, J = J(𝑋̰ → 𝑌̰ ) = = 𝑚𝑜𝑑 ….
𝛿𝑦1 𝛿𝑦2 𝛿𝑦𝑝
𝛿(𝑦1 ,𝑦2 ,….,𝑦𝑝 )
|… . …. …. … .|
𝛿𝑥𝑝 𝛿𝑥𝑝 𝛿𝑥𝑝
….
𝛿𝑦1 𝛿𝑦2 𝛿𝑦𝑝
𝑐11 𝑐12 …. 𝑐1𝑝
𝑐21 𝑐22 …. 𝑐2𝑝
= 𝑚𝑜𝑑 | … . …. …. ….|
𝑐𝑝1 𝑐𝑝2 …. 𝑐𝑝𝑝
= 𝑚𝑜𝑑|𝐶|
Thus, from (3) we have
1
∞ ∞ (𝐶𝑌̰ )⸝ 𝐴(𝐶𝑌̰ )
𝑘 ∫−∞ … . ∫−∞ 𝑒 −2 . 𝐽 𝑑𝑦1 𝑑𝑦2 … . 𝑑𝑦𝑝 = 1
1
∞ ∞ −2𝑌̰ ⸝ 𝑌̰
or 𝑘. 𝑚𝑜𝑑|𝐶| ∫−∞ … . ∫−∞ 𝑒 𝑑𝑦1 𝑑𝑦2 … . 𝑑𝑦𝑝 = 1 {from (4), 𝐶 ⸝𝐴𝐶 = I} ……(6)
From (4), we have 𝐶 ⸝𝐴𝐶 = I
⇒ |𝐶 ⸝𝐴𝐶| = 1
⇒ |𝐶 ⸝||𝐴||𝐶| = 1
⇒ |𝐶|2 |𝐴| = 1
1
⇒ |𝐶|2 = | |
𝐴
1
⇒ 𝑚𝑜𝑑|𝐶| = 1
|𝐴|2
From (6), we have
1 𝑝
𝑘 ∞ ∞ −2 ∑𝑖=1 𝑦𝑖 2
1 ∫−∞ … . ∫−∞ 𝑒 𝑑𝑦1 𝑑𝑦2 … . 𝑑𝑦𝑝 = 1
|𝐴|2
1 1 1
𝑘 ∞ ∞ −2𝑦1 2 −2𝑦2 2 2
⇒ 1 ∫−∞ … . ∫−∞ 𝑒 𝑒 … . 𝑒 −2𝑦𝑝 𝑑𝑦1 𝑑𝑦2 … . 𝑑𝑦𝑝 = 1
|𝐴|2

3|D R C
1 2
𝑘 𝑝 ∞
− 𝑦𝑖
⇒ 1 ∏𝑖=1 [∫−∞ 𝑒 2 𝑑𝑦𝑖 ] = 1
|𝐴|2
𝑘
⇒ 1 ∏𝑝𝑖=1[√2𝜋] = 1
|𝐴|2
𝑝
𝑘
⇒ 1 (2𝜋) 2 = 1
|𝐴|2
1
|𝐴|2
⇒𝑘= 𝑝
(2𝜋)2
Thus, the multivariate normal density function is given by from (2)
1
1
|𝐴|2 − (𝑥̰ −𝑏̰ )⸝ 𝐴(𝑥̰ −𝑏̰ )
𝑓(𝑥1 , 𝑥2 , … . , 𝑥𝑝 ) = 𝑝𝑒 2 ……(7)
(2𝜋)2
We shall now show the significance of 𝑏̰ and A by finding the first and second moments of
𝑋1 , 𝑋2 , …., 𝑋𝑝 .
From (6) we note that the joint density function of 𝑦1 , 𝑦2 , …., 𝑦𝑝 is given by
1 ⸝
𝑔(𝑦1 , 𝑦2 , … . , 𝑦𝑝 ) = 𝑘𝑚𝑜𝑑|𝐶|𝑒 −2𝑌̰ 𝑌̰
1
1 ∑𝑝 2
= 𝑝 . 𝑒 −2 𝑖=1 𝑦𝑖
(2𝜋) 2
1 2
1
= 𝑝 ∏𝑝𝑖=1 𝑒 −2𝑦𝑖
(2𝜋) 2
1 2
𝑝 1
= ∏𝑖=1 [ 𝑒 −2𝑦𝑖 ]
√2𝜋
𝑝
= ∏𝑖=1 𝑔𝑖 (𝑦𝑖 )
where 𝑦𝑖 ~ N(0, 1)
Now, 𝑦𝑖 (𝑖 = 1, 2, …., 𝑝) are independently distributed each according to N(0, 1) i.e., they
are independent standard normal variate.
𝐸 (𝑦1 ) 0
𝐸 (𝑦2 ) 0
Thus, 𝐸(𝑦̰) = ( ) = ( ) = 0̰
…. ….
𝐸(𝑦𝑝 ) 0
1 0 …. 0
0 1 …. 0
and 𝑉(𝑦̰) = ( ) = 𝐼𝑝
…. …. …. ….
0 0 …. 1
From (5), we have 𝑥̰ − 𝑏̰ = 𝑐𝑦̰
Taking expectation on both side
𝐸 (𝑥̰ − 𝑏̰ ) = 𝐸(𝑐𝑦̰)
⇒ 𝐸 (𝑥̰ ) − 𝑏̰ = 𝑐𝐸(𝑦̰) = 0
⇒ 𝑏̰ = 𝐸 (𝑥̰ ) = 𝜇̰ (say)
Taking variance – covariance matrix on both sides
𝑉 (𝑥̰ − 𝑏̰ ) = 𝑉(𝑐𝑦̰)
⇒ 𝑉 (𝑥̰ ) = 𝐶𝑉(𝑦̰)𝐶 ⸝ = 𝐶𝐼𝑝 𝐶 ⸝ = 𝐶𝐶 ⸝
Since from (4) 𝐶 ⸝𝐴𝐶 = I

4|D R C
⇒ 𝐴 = (𝐶 ⸝)−1𝐶 −1
⇒ 𝐴 = (𝐶𝐶 ⸝)−1
⇒ 𝐶𝐶 ⸝ = 𝐴−1
or 𝐴−1 = 𝑉 (𝑥̰ ) = Σ (say)
Therefore, 𝐴 = Σ −1
Since A is positive definite matrix, therefore 𝐴−1 = Σ is also a positive definite matrix.
Thus, the multivariate or more precisely 𝑝 – variate normal density function is given by
{from (7)} –
1 ⸝
1 − (𝑥̰ −𝜇̰) Σ−1 (𝑥̰ −𝜇̰)
𝑓(𝑥1 , 𝑥2 , … . , 𝑥𝑝 ) = 𝑝 1 𝑒 2
(2𝜋)2 |Σ|2
Characteristic Function —
⸝ 1 ⸝
If 𝑋̰ ~ 𝑁𝑝 (𝜇̰ , Σ) then its characteristic function is given by 𝜙𝑋̰ (𝑡̰ ) = 𝑒 𝑖𝑡̰ 𝜇̰−2𝑡̰ Σ𝑡̰ where 𝑡̰ is a
𝑝 component vector or dummy vector.
Proof – In a univariate case if X ~ N(μ, 𝜎 2 ) then the characteristic function of X is define
as 𝜙𝑋 (𝑡) = 𝐸[𝑒 𝑖𝑡𝑋 ]
Similarly, for the multivariate case if 𝑋̰ ~ 𝑁𝑝 (𝜇̰ , Σ) then its characteristic function is define
as 𝜙𝑋̰ (𝑡̰ ) = 𝐸[𝑒 𝑖(𝑡1 𝑋1 +𝑡2 𝑋2 +.…+𝑡𝑝𝑋𝑝) ]

= 𝐸[𝑒 𝑖𝑡̰ 𝑋̰ ] where 𝑡 ⸝ = 𝑡1, 𝑡2 , …., 𝑡𝑝
Since 𝑋̰ ~ 𝑁𝑝 (𝜇̰ , Σ) therefore we have the pdf of 𝑋̰
1 ⸝ −1
1 − (𝑥̰ −𝜇̰) Σ (𝑥̰ −𝜇) ̰
𝑓(𝑥̰ ) = 𝑝 1𝑒 2
(2𝜋)2 |Σ|2
Since Σ −1 is a positive definite matrix therefore there exist a non-singular matrix 𝐶 =
(𝑐𝑖𝑗 )𝑝×𝑝 such that 𝐶 ⸝Σ −1𝐶 = 𝐼𝑝 ……(1)
⇒ (𝐶 ⸝)−1𝐶 ⸝Σ −1𝐶𝐶 −1 = (𝐶 ⸝ )−1𝐼𝑝 𝐶 −1
⇒ Σ −1 = (𝐶𝐶 ⸝)−1
⇒ Σ = 𝐶𝐶 ⸝ ……(2)
Again, from (1) taking determinant on both sides we get
|𝐶 ⸝Σ −1𝐶| = |𝐼𝑝 |
⇒ |𝐶 ⸝||Σ −1||𝐶| = 1
1
⇒ |𝐶|2 = | −1 |
Σ
1
1
Therefore, 𝑚𝑜𝑑|𝐶| = 1 = |Σ|2
|Σ−1 |2
𝑦1
𝑦2
Let 𝑋̰ − 𝜇̰ = 𝐶𝑌̰ where 𝑌̰ = (… .)
𝑦𝑝
1
Therefore, the Jacobian of transformation is |J| = 𝑚𝑜𝑑|𝐶| = |Σ|2
Therefore, the pdf of vector 𝑌̰ is
1
1 − (𝐶𝑌̰ )⸝ Σ−1 (𝐶𝑌̰ )
𝑔(𝑌̰ ) = 𝑝 1𝑒 2 |𝐽|
(2𝜋)2 |Σ|2

5|D R C
1 1
1 −2𝑌̰ ⸝ 𝐶 ⸝ Σ−1 𝐶𝑌̰
= 𝑝 1 𝑒 |Σ|2
(2𝜋)2 |Σ|2
1
1 − 𝑌̰ ⸝ 𝐼𝑝 𝑌̰
= 𝑝 𝑒 2
(2𝜋)2
1
1 ∑𝑝 2
= 𝑝 𝑒 −2 𝑗=1 𝑦𝑗
(2𝜋)2
1 2
𝑝 1
= ∏𝑗=1 [ 𝑒 −2𝑦𝑗 ]
√2𝜋
𝑝
= ∏𝑗=1 𝑔𝑖 (𝑦𝑗 )
where 𝑦𝑗 ~ N(0, 1)
Now, the characteristic function of 𝑌̰ with dummy vector 𝑢̰ is given by

𝜙𝑌̰ (𝑢̰ ) = 𝐸[𝑒 𝑖𝑢̰ 𝑌̰ ]
∑𝑝
= 𝐸 [𝑒 𝑖 𝑗=1 𝑢𝑗 𝑦𝑗 ]
𝑝
= 𝐸[∏𝑗=1 𝑒 𝑖𝑢𝑗 𝑦𝑗 ]
𝑝
= ∏𝑗=1 𝐸[𝑒 𝑖𝑢𝑗 𝑦𝑗 ]
1 2 2
𝑝
= ∏𝑗=1 𝜙𝑦𝑗 (𝑢𝑗 ) {Since 𝜙𝑋 (𝑡) = 𝑒 𝑖𝜇𝑡−2𝑡 𝜎
}
1
𝑝 −2𝑢𝑗 2
= ∏𝑗=1 𝑒
1
∑𝑝 2
= 𝑒 −2 𝑗=1 𝑢𝑗
1
−2𝑢̰ ⸝ 𝑢̰
=𝑒 ……(3)
Thus, the characteristic function of 𝑋̰ is given by

𝜙𝑋̰ (𝑡̰ ) = 𝐸[𝑒 𝑖𝑡̰ 𝑋̰ ]
⸝ (𝜇+𝐶𝑌̰ )
= 𝐸 [𝑒 𝑖𝑡̰ ̰ ]
𝑖𝑡̰ ⸝ 𝜇 𝑖𝑡̰ ⸝ 𝐶𝑌̰
= 𝐸[𝑒 ̰ . 𝑒 ]
⸝ ⸝
= 𝑒 𝑖𝑡̰ 𝜇̰ . 𝐸[𝑒 𝑖𝑡̰ 𝐶𝑌̰ ]
⸝ ⸝
= 𝑒 𝑖𝑡̰ 𝜇̰ 𝐸[𝑒 𝑖𝑢̰ 𝑌̰ ] where 𝑢̰ ⸝ = 𝑡̰ ⸝𝐶

= 𝑒 𝑖𝑡̰ 𝜇̰ 𝜙𝑌̰ (𝑢̰ )
⸝ 1 ⸝
= 𝑒 𝑖𝑡̰ 𝜇̰ . 𝑒 −2𝑢̰ 𝑢̰
{using (3)}
1 ⸝ ⸝
𝑖𝑡̰ ⸝ 𝜇 ̰ −2𝑡̰ 𝐶𝐶 𝑡̰
=𝑒 𝑒
1
𝑖𝑡̰ ⸝ 𝜇̰ −2𝑡̰ ⸝ Σ𝑡̰
=𝑒 𝑒
1
𝑖𝑡̰ ⸝ 𝜇̰−2𝑡̰ ⸝ Σ𝑡̰
=𝑒
which is the required characteristic function of 𝑋̰ .
Marginal Distribution —
Let 𝑋̰ 𝑝×1 = 𝑁𝑝 (𝜇̰ , Σ) and let 𝑋̰ , 𝜇̰ and Σ be partitioned as

6|D R C
𝑥1
𝑥2
…. 𝑋̰ 𝑘×1 (1)
𝑋̰ 𝑝×1 = 𝑥 =( )
𝑘+1 𝑋̰ (𝑝−𝑘)×1 (2)
….
( 𝑥𝑝 )
𝜇1
𝜇2
….
𝜇̰𝑘×1 (1)
𝜇̰𝑝×1 = 𝜇𝑘 =( )
𝜇𝑘+1 𝜇̰(𝑝−𝑘)×1(2)
….
( 𝜇𝑝 )
𝜎11 𝜎12 …. 𝜎1𝑘 𝜎1(𝑘+1) …. 𝜎1𝑝
𝜎21 𝜎22 …. 𝜎2𝑘 𝜎2(𝑘+1) …. 𝜎2𝑝
…. …. …. …. …. …. ….
Σ= 𝜎𝑘1 𝜎𝑘2 …. 𝜎𝑘𝑘 𝜎𝑘(𝑘+1) …. 𝜎𝑘𝑝
𝜎(𝑘+1)1 𝜎(𝑘+1)2 … . 𝜎(𝑘+1)𝑘 𝜎(𝑘+1)(𝑘+1) …. 𝜎(𝑘+1)𝑝
…. …. …. …. …. …. ….
( 𝜎𝑝1 𝜎𝑝2 …. 𝜎𝑝𝑘 𝜎𝑝(𝑘+1) …. 𝜎𝑝𝑝 )
Σ11𝑘×𝑘 Σ12 𝑘×(𝑝−𝑘)
Σ𝑝×𝑝 = ( )
Σ21 (𝑝−𝑘)×𝑘 Σ22 (𝑝−𝑘)×(𝑝−𝑘)
Such that Σ11𝑘×𝑘 = 𝑉(𝑋̰ (1) )
Σ22 (𝑝−𝑘)×(𝑝−𝑘) = 𝑉(𝑋̰ (2) )
Σ12𝑘×(𝑝−𝑘) = 𝐶𝑜𝑣(𝑋̰ (1) , 𝑋̰ (2) )
Σ21 (𝑝−𝑘)×𝑘 = 𝐶𝑜𝑣(𝑋̰ (2) , 𝑋̰ (1) )
Σ12 = Σ21
then the marginal distribution of sub vector 𝑋̰ (1) is
𝑋̰ 𝑘×1(1) ~ 𝑁𝑘 (𝜇̰ (1) , Σ11 )
and marginal distribution of sub vector 𝑋̰ (2) is
𝑋̰ (𝑝−𝑘)×1(2) ~ 𝑁𝑝−𝑘 (𝜇̰ (2) , Σ22)
Proof – The characteristic function of 𝑋̰ is defined as

𝜙𝑋̰ (𝑡̰ ) = 𝐸[𝑒 𝑖𝑡̰ 𝑋̰ ] ……(1)
𝑡1
𝑡2
where 𝑡̰ = (… .) is a dummy vector
𝑡𝑝

7|D R C
𝑡1
𝑡2
….
𝑡̰ 𝑘×1 (1)
Now, let us partition 𝑡̰ as 𝑡̰ 𝑝×1 = 𝑡𝑘 =( (2) )
𝑡𝑘+1 𝑡̰ (𝑝−𝑘)×1
….
( 𝑡𝑝 )
(1) ⸝ (1)
Now, the characteristic function of the vector 𝑋̰ (1) is 𝜙𝑋̰ (1) (𝑡̰ (1) ) = 𝐸 [𝑒 𝑖𝑡̰ 𝑋̰
] ……(2)

𝑡̰ (1) 𝑋̰ (1)
𝑖( (2) ) ( (2) )
Further (1) can be rewritten as 𝜙𝑋̰ (𝑡̰ ) = 𝐸 [𝑒 𝑡̰ 𝑋̰ ]

(1)
⸝ ⸝ 𝑋̰
𝑖(𝑡̰ (1) 𝑡̰ (2) )( (2) )
= 𝐸 [𝑒 𝑋̰ ]
⸝ ⸝
𝑖{𝑡̰ (1) 𝑋̰ (1) +𝑡̰ (2) 𝑋̰ (2) }
= 𝐸 [𝑒 ] ……(3)
On comparing (2) and (3) it is noted that the characteristic function of 𝑋̰ (1) can be obtain
from the characteristic function 𝑋̰ by putting 𝑡̰ (2) = 0̰
Thus, 𝜙𝑋̰ (1) (𝑡̰ (1) ) = [𝜙𝑋̰ (𝑡̰ )] (2) ……(4)
𝑡̰ =0̰
Now, since 𝑋̰ ~ 𝑁𝑝 (𝜇̰ , Σ)
Therefore, the characteristic function of 𝑋̰ is given by
⸝ 𝜇−1𝑡̰ ⸝ Σ𝑡̰
𝜙𝑋̰ (𝑡̰ ) = 𝑒 𝑖𝑡̰ ̰ 2
⸝ ⸝
𝑡̰ (1) 𝜇̰(1) 1 𝑡̰ (1) Σ Σ12 𝑡̰ (1)
𝑖( (2) ) ( (2) )−2( (2) ) ( 11 )( )
Σ21 Σ22 𝑡̰ (2)
= 𝑒 𝑡̰ 𝜇̰ 𝑡̰
(1)
⸝ ⸝ 𝜇̰ 1 ⸝ ⸝ Σ11 Σ12 𝑡̰ (1)
𝑖(𝑡̰ (1) 𝑡̰ (2) )( (2) )−2(𝑡̰ (1) 𝑡̰ (2) )(Σ )( )
𝜇̰ 21 Σ22 𝑡̰ (2)
=𝑒 ……(5)
(2)
Setting 𝑡̰ = 0̰ we get from (5)
(1)
⸝ 𝜇̰ 1 ⸝ Σ11 Σ12 𝑡̰ (1)
𝑖(𝑡̰ (1) 0̰ )( (2) )−2(𝑡̰ (1) 0̰ )(Σ )( )
𝜇̰ 21 Σ22 0̰
[𝜙𝑋̰ (𝑡̰ )] =𝑒
𝑡̰ (2) =0̰
⸝ 1
𝑖𝑡̰ (1) 𝜇̰(1) − (𝑡̰ (1) Σ11
⸝ ⸝ 𝑡̰ (1) )
(1) 2 𝑡̰ (1) Σ12 )(
⇒ 𝜙𝑋̰ (1) (𝑡̰ )=𝑒 0̰
(1) (1) (1)⸝ (1) 1 ⸝
= 𝑒 𝑖𝑡̰ 𝜇̰ −2𝑡̰ Σ11𝑡̰
which is the characteristic function of a 𝑘 – variate normal distribution with parameter 𝜇̰ (1) ,
Σ11. Hence, by uniqueness theorem of characteristic function we get 𝑋̰ 𝑘×1 (1) ~
𝑁𝑘 (𝜇̰ (1) , Σ11)
Similarly, it can be shown that 𝑋̰ (𝑝−𝑘)×1 (2) ~ 𝑁𝑝−𝑘 (𝜇̰ (2) , Σ22 )

8|D R C
Thus, if 𝑋̰ 𝑝×1 has a multivariate normal distribution then any sub vector of 𝑋̰ is also a
multivariate normal with parameters – means, variances and covariances obtained by taking
proper component of 𝜇̰ and Σ respectively.
Theorem 1 :— If 𝑋̰ ~ 𝑁𝑝 (𝜇̰ , Σ) then 𝑍̰ = 𝐷𝑋̰ ~ 𝑁𝑝 (𝐷𝜇̰ , 𝐷Σ𝐷 ⸝ ) where 𝐷 is a 𝑝 × 𝑝 non-
singular matrix of constant element.
Proof – Since 𝑋̰ ~ 𝑁𝑝 (𝜇̰ , Σ) therefore the pdf of 𝑋̰ is given by
1 ⸝
1 (𝑥̰ −𝜇̰) Σ−1 (𝑥̰ −𝜇̰)
𝑓(𝑥̰ ) = 𝑝 1 𝑒 −2
(2𝜋)2 |Σ|2
Given that 𝑍̰ = 𝐷𝑋̰ ⇒ 𝑋̰ = 𝐷 −1𝑍̰
Then the Jacobian of transformation is |J| = |𝐷|−1
1
=| |
𝐷
1
= √|
𝐷|2
1
= √|
𝐷||𝐷⸝ |
|Σ|
= √|
𝐷||Σ||𝐷⸝ |
1
|Σ|2
= 1
|𝐷||Σ||𝐷⸝ |2
Now, the pdf of 𝑍̰ is
1 ⸝
1 − (𝐷−1 𝑍̰ −𝜇̰) Σ−1 (𝐷−1 𝑍̰ −𝜇̰)
𝑔(𝑍̰ ) = 𝑝 1𝑒 2 |𝐽|
(2𝜋)2 |Σ|2
1
1 ⸝
1 −2[𝐷−1 (𝑍̰ −𝐷𝜇̰)] Σ−1 [𝐷−1 (𝑍̰ −𝐷𝜇̰)] |Σ|2
= 𝑝 1 𝑒 . 1
(2𝜋) 2 |Σ|2 |𝐷||Σ||𝐷⸝ |2
1 ⸝
1 − (𝑍̰ −𝐷𝜇̰) (𝐷⸝ )−1 Σ−1 𝐷−1 (𝑍̰ −𝐷𝜇̰) 1
= 𝑝 𝑒 2 1
(2𝜋) 2 |𝐷||Σ||𝐷⸝ |2
1 ⸝
1 − (𝑍̰ −𝐷𝜇̰) (𝐷Σ𝐷⸝ )−1 (𝑍̰ −𝐷𝜇̰)
= 𝑝 1 𝑒 2
(2𝜋) 2 |𝐷||Σ||𝐷⸝ |2
which is the pdf of 𝑝 – variate normal distribution with mean vector 𝐷𝜇̰ and variance –
covariance matrix 𝐷Σ𝐷 ⸝ .
Therefore, 𝑍̰ = 𝐷𝑋̰ ~ 𝑁𝑝 (𝐷𝜇̰ , 𝐷Σ𝐷 ⸝)
Theorem 2 :— Let 𝑋̰ ~ 𝑁𝑝 (𝜇̰ , Σ) then show that for any 𝑟 × 𝑝 matrix 𝐶, 𝐶𝑋̰ ~ 𝑁𝑟 (𝐶𝜇̰ , 𝐶Σ𝐶 ⸝)
Proof – Since 𝑋̰ ~ 𝑁𝑝 (𝜇̰ , Σ) therefore the characteristic function of 𝑋̰ with dummy vector 𝑡̰
⸝ 1 ⸝
is given by 𝜙𝑋̰ (𝑡̰ ) = 𝑒 𝑖𝑡̰ 𝜇̰−2𝑡̰ Σ𝑡̰ ……(1)
The characteristic function of 𝑍̰ 𝑟×1 = 𝐶𝑟×𝑝 𝑋̰ 𝑝×1 is given by
𝑡1
⸝ 𝑡
𝜙𝑍̰ (𝑡̰ ) = 𝐸[𝑒 𝑖𝑡̰ 𝑍̰ ] where 𝑡̰ 𝑟×1 = (…2.)
𝑡𝑟
⸝ 𝐶𝑋̰
= 𝐸[𝑒 𝑖𝑡̰ ]
9|D R C

= 𝐸[𝑒 𝑖𝑢̰ 𝑋̰ ] where 𝑢̰ ⸝ = 𝑡̰ ⸝𝐶 and 𝑢̰ = 𝐶 ⸝𝑡̰
= 𝜙𝑋̰ (𝑢̰ )
⸝ 𝜇−1𝑢̰ ⸝ Σ𝑢̰
= 𝑒 𝑖𝑢̰ ̰ 2
1
𝑖𝑡̰ ⸝ 𝐶𝜇̰−2𝑡̰ ⸝ 𝐶 Σ𝐶 ⸝ 𝑡̰
=𝑒
1
𝑖𝑡̰ ⸝ (𝐶𝜇̰)−2𝑡̰ ⸝ (𝐶 Σ𝐶 ⸝ )𝑡̰
=𝑒
which is the characteristic function of 𝑟 – variate normal distribution with mean 𝐶𝜇̰ and
variance – covariance matrix 𝐶Σ𝐶 ⸝. Hence, by uniqueness theorem of characteristic function
we have 𝑍̰ 𝑟×1 = 𝐶𝑟×𝑝 𝑋̰ 𝑝×1 ~ 𝑁𝑟 (𝐶𝜇̰ , 𝐶Σ𝐶 ⸝)
Example 1 – Prove that in a multivariate normal distribution any subset of the variable is
distributed independently of the complementary subset if and only if every variable in the
former set is uncorrelated with every variable in the other.
Solution – Let 𝑋̰ 𝑝×1 ~ 𝑁𝑝 (𝜇̰ , Σ) and 𝑋̰ , 𝜇̰ and Σ are partitioned as
𝑋̰ 𝑘×1(1) 𝜇̰𝑘×1(1)
𝑋̰ 𝑝×1 = ( ) and 𝜇̰𝑝×1 = ( )
𝑋̰ (𝑝−𝑘)×1(2) 𝜇̰(𝑝−𝑘)×1(2)
Σ11 𝑘×𝑘 Σ12𝑘×(𝑝−𝑘)
Σ𝑝×𝑝 = ( )
Σ21 (𝑝−𝑘)×𝑘 Σ22(𝑝−𝑘)×(𝑝−𝑘)
Sufficient Part :—
Let Σ12 = 0 = Σ21
We are required to prove that 𝑋̰ (1) and 𝑋̰ (2) are independent.
Σ 0
Since Σ = ( 11 )
0 Σ22
Therefore, |Σ| = Σ11Σ22
Therefore, 𝑚𝑜𝑑|Σ| = |Σ11||Σ22 |
−1 Σ11 −1 0
Also, Σ = ( )
0 Σ22−1
1 ⸝ −1
1 − (𝑥̰ −𝜇̰) Σ (𝑥̰ −𝜇) ̰
Now, the pdf of 𝑋̰ is 𝑓 (𝑥̰ ) = 𝑝 1𝑒 2
(2𝜋)2 |Σ|2
1 ⸝
1 − (𝑥̰ −𝜇̰) Σ−1 (𝑥̰ −𝜇̰)
= 𝑘 1 𝑝−𝑘 1 𝑒 2 ……(1)
(2𝜋)2 |Σ11 |2 (2𝜋) 2 |Σ22 |2

Now, (𝑥̰ − 𝜇̰ ) Σ −1(𝑥̰ − 𝜇̰ )
(1) (1) ⸝ (2) (2) ⸝ Σ11 −1 (𝑥̰ (1) − 𝜇̰ (1) )
0
= [(𝑥̰ − 𝜇̰ ) (𝑥̰ − 𝜇̰ ) ][ ][ ]
Σ22−1 (𝑥̰ (2) − 𝜇̰ (2) )
0
(1) (1) ⸝ −1 (2) (2) ⸝ −1
(𝑥̰ (1) − 𝜇̰ (1) )
= [(𝑥̰ − 𝜇̰ ) Σ11 (𝑥̰ − 𝜇̰ ) Σ22 ] [ (2) ]
(𝑥̰ − 𝜇̰ (2) )
⸝ ⸝
= [(𝑥̰ (1) − 𝜇̰ (1) ) Σ11−1(𝑥̰ (1) − 𝜇̰ (1) ) + (𝑥̰ (2) − 𝜇̰ (2) ) Σ22 −1(𝑥̰ (2) − 𝜇̰ (2) )]
From (1) and (2) we have
1 ⸝ ⸝
1 −2[(𝑥̰ (1)
−𝜇̰(1) ) Σ11 −1 (𝑥̰ (1) −𝜇̰(1) )+(𝑥̰ (2) −𝜇̰(2) ) Σ22 −1 (𝑥̰ (2) −𝜇̰(2) )]
𝑓 (𝑥̰ ) = 𝑘 1 𝑝−𝑘 1𝑒
(2𝜋)2 |Σ11 |2 (2𝜋) 2 |Σ22 |2

10 | D R C
1 ⸝ 1 ⸝
1 −2[(𝑥̰ (1) −𝜇̰(1) ) Σ11 −1 (𝑥̰ (1) −𝜇̰(1) )] 1 −2[(𝑥̰ (2) −𝜇̰(2) ) Σ22 −1 (𝑥̰ (2) −𝜇̰(2) )]
=[ 𝑘 1 𝑒 ][ 𝑝−𝑘 1 𝑒 ]
(2𝜋)2 |Σ11 |2 (2𝜋) 2 |Σ22 |2

= 𝑓1 (𝑥̰ (1) ). 𝑓2 (𝑥̰ (2) )


where 𝑋̰ 𝑘×1 (1) ~ 𝑁𝑘 (𝜇̰ (1) , Σ11 )
and 𝑋̰ (𝑝−𝑘)×1(2) ~ 𝑁𝑝−𝑘 (𝜇̰ (2) , Σ22)
(1) (2)
Therefore, 𝑋̰ and 𝑋̰ are independent.
Necessary Part :—
Let 𝑋̰ (1) and 𝑋̰ (2) are independent. We are required to show that correlation between 𝑋̰ (1)
and 𝑋̰ (2) is 0.
Let 𝑋𝑖 be any variate from 𝑋̰ (1) and 𝑋𝑗 be any variate from 𝑋̰ (2)
Therefore, 𝜎𝑖𝑗 = 𝐸 (𝑋𝑖 − 𝜇𝑖 )(𝑋𝑗 − 𝜇𝑗 )
∞ ∞
= ∫−∞ … . ∫−∞(𝑥𝑖 − 𝜇𝑖 )(𝑥𝑗 − 𝜇𝑗 )𝑓(𝑥̰ ) 𝑑𝑥̰
{Since 𝑥̰ (1) and 𝑥̰ (2) are independent}
∞ ∞ ∞ ∞
= ∫−∞ … . ∫−∞(𝑥𝑖 − 𝜇𝑖 )𝑓(𝑥̰ (1) ) 𝑑𝑥̰ (1) . ∫−∞ … . ∫−∞(𝑥𝑗 − 𝜇𝑗 )𝑓(𝑥̰ (2) ) 𝑑𝑥̰ (2)
=0
Now, 𝜎𝑖𝑗 = 𝜎𝑖 𝜎𝑗 𝜌𝑖𝑗 where 𝜌𝑖𝑗 is the correlation between 𝑋𝑖 and 𝑋𝑗
and 𝜎𝑖 ≠ 𝜎𝑗 ≠ 0
Therefore, 𝜎𝑖𝑗 = 0 ⇒ 𝜌𝑖𝑗 = 0
i.e., one set of variables is uncorrelated with the variables of the other set.
Therefore, if one set of variables is uncorrelated with the variables of the other sets are
independent.
5 2 3
Example 2 – Let 𝑋̰ = (𝑋1 𝑋2 𝑋3 )⸝ and Σ = 𝑉 (𝑋̰ ) = (2 3 0). Find the variance of
3 0 2
𝑋1 − 2𝑋2 − 𝑋3 .
Solution – We are given
𝜎12 = 5, 𝜎2 2 = 3 and 𝜎32 = 2
𝜎12 = 𝜎21 = 2, 𝜎13 = 𝜎31 = 3 and 𝜎23 = 𝜎32 = 0
Now, 𝑉 (𝑋1 − 2𝑋2 − 𝑋3 )
= 𝑉 (𝑋1 ) + 4𝑉 (𝑋2 ) + 𝑉 (𝑋3 ) − 4𝐶𝑜𝑣 (𝑋1 𝑋2 ) + 2𝐶𝑜𝑣(𝑋1 𝑋3 ) − 4𝐶𝑜𝑣 (𝑋2 𝑋3 )
=5+4×3+2−4×2+2×3−4×0
= 5 + 12 + 2 − 8 + 6
= 17
1 −1 0
Example 3 – Let 𝑋̰ ~ 𝑁3 (𝜇̰ , Σ) find the distribution of 𝐴𝑋̰ where 𝐴2×3 = ( )
−1 1 −1
Solution – We know that if 𝑋̰ ~ 𝑁3 (𝜇̰ , Σ) then
𝐴2×3 𝑋̰ 3×1 ~ 𝑁2 (𝐴𝜇̰ , 𝐴Σ𝐴⸝)
𝜇1
1 −1 0 𝜇1 − 𝜇2
Let 𝐴𝜇̰ = ( ) (𝜇2 ) = (−𝜇 + 𝜇 − 𝜇 )
−1 1 −1 𝜇 1 2 3
3

11 | D R C
𝜎12 𝜎12 𝜎13 1 −1
⸝ 1 −1 0
then 𝐴Σ𝐴 = ( ) (𝜎21 𝜎2 2 𝜎23 ) (−1 1)
−1 1 −1
𝜎31 𝜎32 𝜎32 0 −1
𝜎12 − 𝜎21 𝜎12 − 𝜎2 2 𝜎13 − 𝜎23 1 −1
=( ) (−1 1 )
−𝜎1 2 + 𝜎21 − 𝜎31 −𝜎12 + 𝜎2 2 − 𝜎32 −𝜎13 + 𝜎23 − 𝜎32
0 −1
2 2 2 2
𝜎1 − 𝜎21 − 𝜎12 + 𝜎2 −𝜎1 + 𝜎21 + 𝜎12 − 𝜎2 − 𝜎13 + 𝜎23
=( )
−𝜎1 + 2𝜎12 − 𝜎2 + 𝜎32 − 𝜎31 𝜎1 2 + 𝜎22 + 𝜎3 2 − 2𝜎12 − 2𝜎23 + 2𝜎13
2 2

Example 4 – Let 𝑋̰ = (𝑋1 , 𝑋2 , 𝑋3 , 𝑋4 , 𝑋5 )⸝ ~ 𝑁5 (𝜇̰ , Σ). Find the distribution of (𝑋2 𝑋4 )⸝


𝑋1
𝑋2
𝑋2 0 1 0 0 0
Solution – Let 𝑌̰2×1 = ( ) =( ) 𝑋3
𝑋4 2×1 0 0 0 1 0 2×5 𝑋
4
(𝑋5 )5×1

i.e., 𝑌̰ = 𝐶𝑋̰ ~ 𝑁2 (𝐶𝜇̰ , 𝐶Σ𝐶 )
0 1 0 0 0
where 𝐶2×5 = ( )
0 0 0 1 0 2×5
𝜇1
𝜇2
Let 𝜇̰ = 𝜇3
𝜇4
(𝜇 5 )
𝜇1
𝜇2
0 1 0 0 0 𝜇 𝜇2
Therefore, 𝐶𝜇̰ = ( ) 3 = (𝜇 )
0 0 0 1 0 𝜇 4
4
(𝜇 5 )
2
𝜎1 𝜎12 𝜎13 𝜎14 𝜎15
𝜎21 𝜎2 2 𝜎23 𝜎24 𝜎25
Let Σ = 𝜎31 𝜎32 𝜎3 2 𝜎34 𝜎35
𝜎41 𝜎42 𝜎43 𝜎42 𝜎45
(𝜎51 𝜎52 𝜎53 𝜎54 𝜎5 2)
𝜎12 𝜎12 𝜎13 𝜎14 𝜎15 0 0
2
𝜎21 𝜎2 𝜎23 𝜎24 𝜎25 1 0
⸝ 0 1 0 0 0 2
Therefore, 𝐶Σ𝐶 = ( ) 𝜎31 𝜎32 𝜎3 𝜎34 𝜎35 0 0
0 0 0 1 0
𝜎41 𝜎42 𝜎43 𝜎4 2 𝜎45 0 1
2 (0 0 )
(𝜎51 𝜎52 𝜎53 𝜎54 𝜎5 )
0 0
𝜎21 𝜎2 2 𝜎23 𝜎24 𝜎25 1 0
=( ) 0 0
𝜎41 𝜎42 𝜎43 𝜎42 𝜎45
0 1
(0 0 )
12 | D R C
𝜎2 2 𝜎24
=( )
𝜎42 𝜎4 2 2×2

Example 5 – Let 𝑋̰ = (𝑋1 , 𝑋2 , … . , 𝑋𝑝 ) be a vector of random variable. Define 𝑋1 = 𝑌1 , 𝑌𝑖 =
𝑋𝑖 − 𝑋𝑖−1 (𝑖 = 1, 2, …., 𝑝). If 𝑌𝑖 ’s are mutually independent each with unit variance then
find the dispersion matrix of 𝑋̰ .
Solution – Given 𝑌1 = 𝑋1
𝑌𝑖 = 𝑋𝑖 − 𝑋𝑖−1 (𝑖 = 1, 2, …., 𝑝)
and 𝑌𝑖 ’s are mutually independent.
Now, 𝑋1 = 𝑌1
⇒ 𝑉 (𝑋1 ) = 𝑉 (𝑌1) = 1 ……(1)
2
Let 𝑉 (𝑋𝑖 ) = 𝜎𝑖 ; 𝐸 (𝑋𝑖 ) = 𝜇𝑖 ; 𝐶𝑜𝑣(𝑋𝑖 𝑋𝑗 ) = 𝜎𝑖𝑗
Also, 𝐶𝑜𝑣 (𝑌1 𝑌2 ) = 0
⇒ 𝐸 [𝑋1 (𝑋2 − 𝑋1 )] − 𝐸 (𝑋1 )𝐸 (𝑋2 − 𝑋1 ) = 0
⇒ 𝐸 (𝑋1 𝑋2 ) − 𝐸(𝑋1 2) − 𝐸 (𝑋1 )𝐸 (𝑋2 ) + {𝐸 (𝑋1 )}2 = 0
⇒ 𝐸 (𝑋1 𝑋2 ) − 𝐸 (𝑋1 )𝐸 (𝑋2 ) = 𝐸(𝑋1 2) − {𝐸 (𝑋1 )}2
⇒ 𝐶𝑜𝑣(𝑋1 𝑋2 ) = 𝑉 (𝑋1 ) = 1
i.e., 𝜎12 = 𝜎21 = 1 ……(2)
𝐶𝑜𝑣(𝑌1 𝑌3 ) = 0
⇒ 𝐸 (𝑌1 𝑌3 ) − 𝐸 (𝑌1 )𝐸 (𝑌3 ) = 0
⇒ 𝐸 [𝑋1 (𝑋3 − 𝑋2 )] − 𝐸 (𝑋1 )𝐸 (𝑋3 − 𝑋2 ) = 0
⇒ 𝐸 (𝑋1 𝑋3 ) − 𝐸 (𝑋1 𝑋2 ) − 𝐸 (𝑋1 )𝐸 (𝑋3 ) + 𝐸 (𝑋1 )𝐸 (𝑋2 ) = 0
⇒ 𝐸 (𝑋1 𝑋3 ) − 𝐸 (𝑋1 )𝐸 (𝑋3 ) = 𝐸 (𝑋1 𝑋2 ) − 𝐸 (𝑋1 )𝐸 (𝑋2 )
⇒ 𝐶𝑜𝑣 (𝑋1 𝑋3 ) = 𝐶𝑜𝑣(𝑋1 𝑋2 ) = 1
i.e., 𝜎13 = 𝜎31 = 1 ……(3)
Similarly, it can be shown that 𝜎14 = 𝜎15 = …. = 𝜎1𝑝 = 1 ……(4)
Now, 𝑌2 = 𝑋2 − 𝑋1
𝑉 (𝑌2 ) = 𝑉 (𝑋2 − 𝑋1 )
⇒ 1 = 𝑉 (𝑋2 ) + 𝑉 (𝑋1 ) − 2𝐶𝑜𝑣 (𝑋1 𝑋2 )
⇒ 𝑉 (𝑋2 ) = 1 − 1 + 2 = 2 ……(5)
𝐶𝑜𝑣(𝑌2 𝑌3 ) = 0
⇒ 𝐸 (𝑌2 𝑌3 ) − 𝐸 (𝑌2 )𝐸 (𝑌3 ) = 0
⇒ 𝐸 [(𝑋2 − 𝑋1 )(𝑋3 − 𝑋2 )] − 𝐸 (𝑋2 − 𝑋1 )𝐸 (𝑋3 − 𝑋2 ) = 0
⇒ 𝐸 (𝑋2 𝑋3 ) − 𝐸(𝑋2 2) − 𝐸 (𝑋1 𝑋3 ) + 𝐸 (𝑋1 𝑋2 ) − {𝐸 (𝑋2 ) − 𝐸 (𝑋1 )}{𝐸 (𝑋3 ) − 𝐸 (𝑋2 )} = 0
⇒ {𝐸 (𝑋2 𝑋3 ) − 𝐸 (𝑋2 )𝐸 (𝑋3 )} − {𝐸 (𝑋1 𝑋3 ) − 𝐸 (𝑋1 )𝐸 (𝑋3 )} + {𝐸 (𝑋1 𝑋2 ) −
2
𝐸 (𝑋1 )𝐸 (𝑋2 )} − {𝐸(𝑋2 2 ) − (𝐸 (𝑋2 )) } = 0
⇒ 𝐶𝑜𝑣 (𝑋2 𝑋3 ) − 𝐶𝑜𝑣(𝑋1 𝑋3 ) + 𝐶𝑜𝑣(𝑋1 𝑋2 ) − 𝑉 (𝑋2 ) = 0
⇒ 𝐶𝑜𝑣 (𝑋2 𝑋3 ) = 1 − 1 + 2
⇒ 𝐶𝑜𝑣 (𝑋2 𝑋3 ) = 2
i.e., 𝜎23 = 𝜎32 = 2 ……(6)
Similarly, it can be shown that 𝜎24 = 𝜎25 = 𝜎26 = …. = 𝜎2𝑝 = 2 ……(7)
Again, 𝑌3 = 𝑋3 − 𝑋2
𝑉 (𝑌3 ) = 𝑉 (𝑋3 − 𝑋2 )
13 | D R C
⇒ 1 = 𝑉 (𝑋3 ) + 2 − 2 × 2
⇒ 𝑉 (𝑋3 ) = 3
Similarly, it can be shown that 𝜎34 = 𝜎35 = …. = 𝜎3𝑝 = 3 ……(8)
Therefore, the variance – covariance matrix is given by
𝜎12 𝜎12 𝜎13 … . 𝜎1𝑝 1 1 …. …. 1
2
1 2 …. …. 2
Σ= 𝜎21 𝜎2 𝜎 23 … . 𝜎2𝑝 = 1 2 3 …. 3
…. …. …. …. …. …. …. …. …. ….
𝜎𝑝1 𝜎𝑝2 𝜎𝑝3 … . 𝜎𝑝 2
( ) (1 2 3 …. 𝑝 )
𝑝(𝑝+1)
Therefore, trace of Σ = 1 + 2 + 3+. … + 𝑝 =
2
Bivariate Normal Distribution as a Particular Case of Multivariate Normal —
The 𝑝 – variate normal density is given by
1 ⸝
1 −2(𝑥̰ −𝜇̰) Σ−1 (𝑥̰ −𝜇̰)
𝑓 (𝑥̰ ) = 𝑝 1 𝑒 ……(1)
(2𝜋)2 |Σ|2
Let us take 𝑝 = 2 and consider the parameters
𝜇1 = 𝐸 (𝑋1 ), 𝜇2 = 𝐸 (𝑋2 ), 𝜎12 = 𝑉 (𝑋1 ), 𝜎22 = 𝑉 (𝑋2 )
𝐶𝑜𝑣(𝑋1 𝑋2 ) 𝜎
and 𝜌12 = Correlation (𝑋1 𝑋2 ) = = 12 ……(2)
𝜎1 𝜎2 𝜎1 𝜎2
𝜎1 2 𝜎12
If 𝑝 = 2 then the variance – covariance matrix Σ becomes Σ = ( )
𝜎21 𝜎2 2
2 2
Also, |Σ| = 𝜎1 𝜎2 − 𝜎12𝜎21
= 𝜎12 𝜎2 2 − 𝜎122
𝜎12 2
= 𝜎12 𝜎2 2 (1 − )
𝜎1 2 𝜎2 2
= 𝜎1 𝜎2 1 − 𝜌12 2 ) {using (2)}
2 2(
……(3)
−1 1 𝜎22 −𝜎12
Therefore, Σ = | | ( )
Σ −𝜎
21 𝜎1 2
1 𝜎2 2 −𝜎12
= 2 2 (1−𝜌 2) ( ) {Since 𝜎12 = 𝜎21 } ……(4)
𝜎1 𝜎2 12 −𝜎12 𝜎1 2
⸝ −1
Again, (𝑥̰ − 𝜇̰ ) Σ (𝑥̰ − 𝜇̰ )
𝑥1 − 𝜇1 ⸝ 1 𝜎2 2 −𝜎12 𝑥1 − 𝜇1
= (𝑥 − 𝜇 ) 2 2 (1−𝜌 2) ( ) (𝑥 − 𝜇 )
2 2 𝜎1 𝜎2 12 −𝜎12 𝜎12 2 2
2
1 𝜎 −𝜎12 𝑥1 − 𝜇1
= (𝑥1 − 𝜇1 𝑥2 − 𝜇2 ) ( 2 ) (𝑥 − 𝜇 )
𝜎1 2 𝜎2 2 (1−𝜌12 2 ) −𝜎12 𝜎1 2 2 2
=
1 𝑥 −𝜇
(𝜎22 (𝑥1 − 𝜇1) − 𝜎12 (𝑥2 − 𝜇2 ) −𝜎12(𝑥1 − 𝜇1) + 𝜎12 (𝑥2 − 𝜇2 )) (𝑥1 − 𝜇1 )
𝜎1 2 𝜎2 2 (1−𝜌12 2 ) 2 2
1 2(
= )2
(𝜎 𝑥1 − 𝜇1 − 𝜎12 (𝑥1 − 𝜇1)(𝑥2 − 𝜇2) − 𝜎12(𝑥1 − 𝜇1)(𝑥2 −
𝜎1 2 𝜎2 2 (1−𝜌12 2 ) 2
𝜇2) + 𝜎12 (𝑥2 − 𝜇2)2 )
1
= 2 2 (1−𝜌 2 ) (𝜎22 (𝑥1 − 𝜇1 )2 − 2𝜎12 (𝑥1 − 𝜇1)(𝑥2 − 𝜇2 ) + 𝜎1 2 (𝑥2 − 𝜇2)2)
𝜎1 𝜎2 12
1 (𝑥1 −𝜇1 )2 𝜎12 (𝑥2 −𝜇2 )2
= (1−𝜌 2) { − 2 (𝑥1 − 𝜇1 )( 𝑥2 − 𝜇 2 ) + } ……(5)
12 𝜎1 2 𝜎1 2 𝜎2 2 𝜎2 2

14 | D R C
Thus, from (1) using (3) and (5) we get
1 (𝑥 −𝜇 )2 𝜎 (𝑥 −𝜇 )2
1 − 2 { 1 21 −2 212 2 (𝑥1 −𝜇1 )(𝑥2 −𝜇2 )+ 2 22 }
𝑓(𝑥1𝑥2 ) = 𝑒 2(1−𝜌12 ) 𝜎1 𝜎1 𝜎2 𝜎2
2𝜋𝜎1 𝜎2 √1−𝜌12 2
which is the pdf of the bivariate normal distribution with parameter 𝜇1 , 𝜇2 , 𝜎12 , 𝜎2 2, 𝜌12
Theorem 3 :— If 𝑋̰ ~ 𝑁𝑝 (𝜇̰ , Σ) then the quadratic form in the exponent of multivariate

normal density function i.e., 𝑄 = (𝑥̰ − 𝜇̰ ) Σ −1(𝑥̰ − 𝜇̰ ) follows central 𝜒 2 distribution with
𝑝 df.
Proof – Since 𝑋̰ ~ 𝑁𝑝 (𝜇̰ , Σ)
1 ⸝ −1
1
Therefore, 𝑓 (𝑥̰ ) = 𝑝 1 𝑒 −2(𝑥̰ −𝜇̰) Σ (𝑥̰ −𝜇) ̰
(2𝜋)2 |Σ|2
Since Σ −1 is a positive definite matrix, therefore there exist a non-singular matrix 𝐶 such
that 𝐶 ⸝Σ −1 𝐶 = 𝐼𝑝
𝑦1
𝑦2
Now, let 𝑥̰ − 𝜇̰ = 𝐶𝑌̰ where 𝑌̰ = (… .)
𝑦𝑝
1
then the Jacobian of transformation is |J| = 𝑚𝑜𝑑|𝐶| = |Σ| 2

Then the pdf of 𝑌̰ becomes


1
1 − (𝐶𝑌̰ )⸝ Σ−1 (𝐶𝑌̰)
𝑔(𝑦̰) = 𝑝 1𝑒 2 . |𝐽|
(2𝜋)2 |Σ|2
1 1
1 −2𝑌̰ ⸝ 𝐶 ⸝ Σ−1 𝐶𝑌̰
= 𝑝 1 𝑒 . |Σ| 2
(2𝜋)2 |Σ|2
1
1 − 𝑌̰ ⸝ 𝑌̰
= 𝑝 𝑒 2
(2𝜋)2
1
1 ∑𝑝 2
= 𝑝 𝑒 −2 𝑖=1 𝑦𝑖
(2𝜋)2
1 2
𝑝 1
= ∏𝑖=1 [ 𝑒 − 2𝑦𝑖 ]
√2𝜋
𝑝
∏𝑖=1 𝑔𝑖 (𝑦𝑖 )
= where 𝑦𝑖 ~ N(0, 1)
⸝ −1
Now, 𝑄 = (𝑥̰ − 𝜇̰ ) Σ (𝑥̰ − 𝜇̰ )
= (𝐶𝑌̰ )⸝Σ −1(𝐶𝑌̰ )
= 𝑌̰ ⸝ 𝐶 ⸝Σ −1𝐶𝑌̰
= 𝑌̰ ⸝ 𝑌̰
𝑝
= ∑𝑖=1 𝑦𝑖 2 ~ 𝜒𝑝 2

Thus, 𝑄 = (𝑥̰ − 𝜇̰ ) Σ −1(𝑥̰ − 𝜇̰ ) follows 𝜒 2 distribution with 𝑝 df.
𝑋 − 𝑋2
Example 1 – Let 𝑋̰ ~ 𝑁𝑝 (𝜇̰ , Σ). Find the distribution of ( 1 )
𝑋2 − 𝑋3
𝑋 − 𝑋2 1 −1 0 𝑋 − 𝑋2
Solution – Let 𝑌̰2×1 = ( 1 )=( ) ( 1 ) = 𝐶𝑋̰
𝑋2 − 𝑋3 0 1 −1 2×3 𝑋2 − 𝑋3 3×1
1 −1 0
where 𝐶 = ( )
0 1 −1
Therefore, 𝑌̰ = 𝐶𝑋̰ ~ 𝑁3 (𝐶𝜇̰ , 𝐶Σ𝐶 ⸝)
15 | D R C
𝜇1 𝜎12
𝜎12 𝜎13
Let 𝜇̰ = (𝜇2 ) and Σ = (𝜎21
𝜎22 𝜎23 )
𝜇3 𝜎32 𝜎3 2
𝜎31
𝜇1
1 −1 0 𝜇1 − 𝜇2
Now, 𝐶𝜇̰ = ( ) (𝜇2) = (𝜇 − 𝜇 )
0 1 −1 𝜇 2 3
3
𝜎12 𝜎12 𝜎13 1 0
⸝ 1 −1 0 2
and 𝐶Σ𝐶 = ( ) (𝜎21 𝜎2 𝜎23 ) (−1 1)
0 1 −1
𝜎31 𝜎32 𝜎3 2 0 −1
𝜎12 − 𝜎21 𝜎12 − 𝜎22 𝜎13 − 𝜎23 1 0
=( ) (−1 1 )
𝜎21 − 𝜎31 𝜎2 2 − 𝜎32 𝜎23 − 𝜎3 2
0 −1
𝜎12 − 𝜎12 − 𝜎21 + 𝜎2 2 𝜎12 − 𝜎2 2 − 𝜎13 + 𝜎23
=( )
𝜎21 − 𝜎31 − 𝜎2 2 + 𝜎32 𝜎2 2 − 𝜎32 − 𝜎23 + 𝜎3 2
Some Important Results —
Result 1 :—
a) If 𝑋1 (𝐾1 𝑋1 ) and 𝑋2 (𝐾2 𝑋1 ) are independent then 𝐶𝑜𝑣 (𝑋1 𝑋2 ) = 0 a 𝐾1 × 𝐾2 matrix of
zeroes.
𝑋 𝜇1 Σ Σ12
b) If [ 1 ] is 𝑁𝐾1 +𝐾2 ([𝜇 ] , [ 11 ]) then 𝑋1 and 𝑋2 are independent if and only if Σ12
𝑋2 2 Σ21 Σ22
=0
Result 2 :—
𝑋 𝜇1 Σ Σ12
Let [ 1 ] be distributed as 𝑁𝑝 (𝜇̰ , Σ) with 𝜇̰ = [𝜇 ] and Σ = [ 11 ] and |Σ22 | > 0, then
𝑋2 2 Σ21 Σ22
the conditional distribution of 𝑋̰ 1 given 𝑋̰ 2 = 𝑥2 is normal and has mean = 𝜇1 +
Σ12Σ22 −1(𝑥2 − 𝜇2 ) and covariance = Σ11 − Σ12Σ22 −1Σ21
Result 3 :—

Let 𝑋̰ be distributed as 𝑁𝑝 (𝜇̰ , Σ) with |Σ| > 0 then (𝑥̰ − 𝜇̰ ) Σ −1(𝑥̰ − 𝜇̰ ) is distributed as 𝜒𝑝 2
i.e., chi-square distribution with 𝑝 df.
In particular, if 𝑋̰ ~ 𝑁𝑝 (0, Σ) then 𝑋̰ ⸝ Σ −1𝑋̰ ~ 𝜒𝑝 2
4 0 −1

Example 2 – If 𝑋̰ ~ 𝑁3 (𝜇̰ , Σ) with 𝜇̰ = 1 −1 2 and Σ = ( 0 6 0 ). Find
( )
−1 0 3
𝑋2
a) the marginal distribution of ( )
𝑋3
b) the conditional distribution of 𝑋1 given that 𝑋3 = 𝑥3
c) the conditional distribution of 𝑋1 given that 𝑋2 = 𝑥2 and 𝑋3 = 𝑥3
𝑋̰ 𝜇̰1 Σ Σ12
Solution – Let 𝑋̰ = ( 1 ) ~ 𝑁𝑝 (𝜇̰ , Σ) with 𝜇̰ = (𝜇 ), Σ = ( 11 ) and Σ22 > 0
𝑋̰ 2 ̰2 Σ21 Σ22
𝑋 𝜇2 𝜎22 𝜎23 −1 6 0
a) The marginal distribution of 𝑋̰ 2 = ( 2 ) is 𝑁2 ([𝜇 ] , [𝜎 𝜎 ]) = 𝑁2 [( ) , ( )]
𝑋3 3 32 33 2 0 3
b) Now, 𝑋̰ 1⁄𝑋̰ 2 = 𝑥2 ~ 𝑁[𝜇1 + Σ12Σ22 −1(𝑥2 − 𝜇2 ), Σ11 − Σ12Σ22 −1Σ21 ]
Here, 𝑋̰ 1 given 𝑋̰ 3 = 𝑥3 ~ 𝑁[1 + (−1) × 3−1 (𝑥2 − 2), 4 − (−1) × 3−1(−1)]

16 | D R C
1 1
~ 𝑁 [1 − (𝑥2 − 2), 4 − ]
3 3
5 𝑥2 11
~𝑁[ − , ]
3 3 3
c) Again, 𝑋̰ 1⁄𝑋̰ 2, 𝑋̰ 3 = 𝑥3
−1 𝑥 + 1 −1
6 0 2 6 0 0
~ 𝑁 [1 + (0 −1) ( ) ( ) , 4 − (0 −1) ( ) ( )]
0 3 𝑥3 − 2 0 3 −1
−1 𝑥 + 1
0.17 0 0.17 0 −1 0
~ 𝑁 [1 + (0 −1) ( ) ( 2 ) , 4 − (0 −1) ( ) ( )]
0 0.33 𝑥3 − 2 0 0.33 −1
𝑥 +1 0
~ 𝑁 [1 + (0 −0.33) ( 2 ) , 4 − (0 −0.33) ( )]
𝑥3 − 2 −1
~ 𝑁[1 + (−0.33)(𝑥3 − 2), 4 − 0.33]
~ 𝑁[1 − 0.33𝑥3 , 3.67]
Sampling Distribution for Mean Vector and Variance – Covariance Matrix —
𝑋1
𝑋2
Let 𝑋̰ 𝑝×1 = (… .) ~ 𝑁𝑝 (𝜇̰ , Σ) and let 𝑋̰ 𝛼 (𝛼 = 1, 2, …., 𝑁) be a random sample of size 𝑁
𝑋𝑝
from 𝑁𝑝 (𝜇̰ , Σ). It is written as 𝑋̰ 𝛼 (𝛼 = 1, 2, …., 𝑁) ~ 𝑁𝑝 (𝜇̰ , Σ)
𝑋11 𝑋12 𝑋1𝛼 𝑋1𝑝
𝑋21 𝑋22 𝑋 𝑋
Thus, 𝑋̰ 1 = ( … . ), 𝑋̰ 2 = ( … . ), …., 𝑋̰ 𝛼 = ( 2𝛼 ), …., 𝑋̰ 𝑝 = ( 2𝑝 )
…. ….
𝑋𝑝1 𝑋𝑝2 𝑋𝛼 𝑋𝑝𝑝
Sample Mean Vector :—
1 𝑁
𝑋1𝛼 ∑ 𝑋 𝑋̅1
𝑁 𝛼=1 1𝛼
1 𝑁
1
It is define as 𝑋̰̅ = ∑𝑁 𝑋̰ =
1 𝑁
∑ (
𝑋2𝛼
) = ∑𝛼=1 𝑋2𝛼 = 𝑋̅2
𝑁
𝑁 𝛼=1 𝛼 𝑁 𝛼=1 …. …. ….
𝑋𝑝𝛼 1 𝑁 (𝑋̅𝑝 )
(𝑁 ∑𝛼=1 𝑋𝑝𝛼 )
Matrix of Sum of Squares and Cross Products :—
The 𝑝 × 𝑝 symmetric matrix 𝐴𝑝×𝑝 = ∑𝑁 ̅ ̅ ⸝
𝛼=1(𝑋̰ 𝛼 − 𝑋̰ )𝑝×1 (𝑋̰ 𝛼 − 𝑋̰ ) 1×𝑝 is known as the matrix
of sum of squares and cross products of deviation about the sample mean.
𝑋1𝛼 − 𝑋̅1
𝑋2𝛼 − 𝑋̅2
Therefore, 𝐴𝑝×𝑝 = ∑𝑁 𝛼=1 (𝑋1𝛼 − 𝑋̅1 𝑋2𝛼 − 𝑋̅2 … . 𝑋𝑝𝛼 − 𝑋̅𝑝 )1×𝑝
….
(𝑋𝑝𝛼 − 𝑋̅𝑝 )𝑝×1
∑𝑁 ̅ 2
𝛼=1(𝑋1𝛼 − 𝑋1 ) ∑𝑁 ̅ ̅
𝛼=1(𝑋1𝛼 − 𝑋1 )(𝑋2𝛼 − 𝑋2 ) … . ∑𝑁 ̅ ̅
𝛼=1(𝑋1𝛼 − 𝑋1 ) (𝑋𝑝𝛼 − 𝑋𝑝 )
𝑁 ̅ ̅ ∑𝛼=1(𝑋2𝛼 − 𝑋̅2 )
𝑁 2
… . ∑𝛼=1(𝑋2𝛼 − 𝑋̅2 )(𝑋𝑝𝛼 − 𝑋̅𝑝 )
𝑁
= ∑𝛼=1(𝑋1𝛼 − 𝑋1 )(𝑋2𝛼 − 𝑋2 )
…. …. …. ….
∑ 𝑁 ( ̅ )( ̅ ∑ 𝑁 ( ̅ ̅
𝛼=1 𝑋2𝛼 − 𝑋2 )(𝑋𝑝𝛼 − 𝑋𝑝 ) ∑ 𝑁 ̅ 2
( 𝛼=1 𝑋1𝛼 − 𝑋1 𝑋𝑝𝛼 − 𝑋𝑝 ) …. 𝛼=1(𝑋𝑝𝛼 − 𝑋𝑝 ) )
𝑎11 𝑎12 …. 𝑎1𝑝
𝑎21 𝑎22 …. 𝑎2𝑝
= (…. …. …. ….)
𝑎𝑝1 𝑎𝑝2 …. 𝑎𝑝𝑝
17 | D R C
where 𝑎𝑖𝑗 = ∑𝑁 ̅ ̅ 𝑁 ̅ 2
𝛼=1(𝑋𝑖𝛼 − 𝑋𝑖 )(𝑋𝑗𝛼 − 𝑋𝑗 ) and 𝑎𝑖𝑖 = ∑𝛼=1(𝑋𝑖𝛼 − 𝑋𝑖 )
𝑝(𝑝+1)
The matrix A is also known as WISHART matrix. It is a symmetric matrix, and it has
2
distinct element.
Theorem 1 :— Let 𝑋̰ 𝛼 (𝛼 = 1, 2, …., 𝑁) be 𝑁 independent observation from 𝑁𝑝 (𝜇̰ , Σ). Then
𝐴
the maximum likelihood estimation for 𝜇̰ and Σ are given by 𝜇̰̂ = 𝑋̰̅ and Σ̂ = where 𝐴 =
𝑁
∑𝑁 ( ̅ )( ̅ ) ⸝
𝛼=1 𝛼𝑋̰ − 𝑋̰ 𝑋̰ 𝛼 − 𝑋̰
Proof – Since 𝑋̰ 𝛼 (𝛼 = 1, 2, …., 𝑁) are 𝑁 independent observation from 𝑁𝑝 (𝜇̰ , Σ) therefore
the density function of the 𝛼 𝑡ℎ observation 𝑋̰ 𝛼 is given by
1 ⸝
1 − (𝑋̰𝛼 −𝜇̰) Σ−1 (𝑋̰𝛼 −𝜇̰)
𝑓 (𝑋̰ 𝛼 ) = 𝑝 1𝑒 2
(2𝜋)2 |Σ|2
The likelihood function of the sample is
𝐿(𝜇̰ , Σ) = ∏𝑁
𝛼=1 𝑓(𝑋̰ 𝛼 )
1 ⸝
1 − (𝑋̰𝛼 −𝜇̰) Σ−1 (𝑋̰𝛼 −𝜇̰)
= ∏𝑁
𝛼=1 𝑝 1𝑒 2
(2𝜋)2 |Σ|2
1 ⸝ −1
1 − 2 ∑𝑁
𝛼=1(𝑋̰𝛼 −𝜇̰) Σ (𝑋̰𝛼 −𝜇̰)
= 𝑁𝑝 𝑁 𝑒 ……(1)
(2𝜋) 2 |Σ| 2
−1
Let a function Ψ = Σ then
𝑁
1 ⸝
Ψ2 − 2 ∑𝑁
𝛼=1(𝑋̰𝛼 −𝜇̰) Ψ(𝑋̰𝛼 −𝜇̰)
𝐿(𝜇̰ , Ψ) = 𝑁𝑝 𝑒 ……(2)
(2𝜋) 2

where ∑𝑁 𝛼=1(𝑋̰ 𝛼 − 𝜇̰ ) Ψ(𝑋̰ 𝛼 − 𝜇̰ )

= ∑𝑁 ̅ ̅ ̅ ̅
𝛼=1(𝑋̰ 𝛼 − 𝑋̰ + 𝑋̰ − 𝜇̰ ) Ψ(𝑋̰ 𝛼 − 𝑋̰ + 𝑋̰ − 𝜇̰ )

= ∑𝑁 ̅ ⸝ ̅ 𝑁 ̅ ⸝ 𝑁 ̅
𝛼=1(𝑋̰ 𝛼 − 𝑋̰ ) Ψ(𝑋̰ 𝛼 − 𝑋̰ ) + ∑𝛼=1(𝑋̰ 𝛼 − 𝑋̰ ) Ψ(𝑋̰ 𝛼 − 𝜇̰ ) + ∑𝛼=1(𝑋̰ − 𝜇̰ ) Ψ(𝑋̰ 𝛼 −

𝑋̰̅ ) + ∑𝑁 ̅ ̅
𝛼=1(𝑋̰ − 𝜇̰ ) Ψ(𝑋̰ − 𝜇̰ )

= ∑𝑁 ̅ ⸝ ̅ ̅ ̅
𝛼=1 𝑡𝑟 (𝑋̰ 𝛼 − 𝑋̰ ) Ψ(𝑋̰ 𝛼 − 𝑋̰ ) + 𝑁(𝑋̰ − 𝜇̰ ) Ψ(𝑋̰ − 𝜇̰ )

= ∑𝑁 ̅ ̅ ⸝ ̅ ̅
𝛼=1 𝑡𝑟 (𝑋̰ 𝛼 − 𝑋̰ )(𝑋̰ 𝛼 − 𝑋̰ ) Ψ + 𝑁(𝑋̰ − 𝜇̰ ) Ψ(𝑋̰ − 𝜇̰ )

= 𝑡𝑟 ∑𝑁 ̅ ̅ ⸝ ̅ ̅
𝛼=1(𝑋̰ 𝛼 − 𝑋̰ )(𝑋̰ 𝛼 − 𝑋̰ ) Ψ + 𝑁(𝑋̰ − 𝜇̰ ) Ψ(𝑋̰ − 𝜇̰ )

= 𝑡𝑟. 𝐴Ψ + 𝑁(𝑋̰̅ − 𝜇̰ ) Ψ(𝑋̰̅ − 𝜇̰ ) ……(3)
Putting (3) in (2) we get
𝑁𝑝 𝑁 1 𝑁 ⸝
log 𝐿(𝜇̰ , Ψ) = − log(2𝜋) + log|Ψ| − 𝑡𝑟. 𝐴Ψ − (𝑋̰̅ − 𝜇̰ ) Ψ(𝑋̰̅ − 𝜇̰ ) ……(4)
2 2 2 2

We see that log 𝐿(𝜇̰ , Ψ) will be maximum with respect to 𝜇 if (𝑋̰̅ − 𝜇̰ ) Ψ(𝑋̰̅ − 𝜇̰ ) is

minimum. Since (𝑋̰̅ − 𝜇̰ ) Ψ(𝑋̰̅ − 𝜇̰ ) is a positive definite quadratic form, its minimum
value is zero and it will be only when 𝑋̰̅ − 𝜇̰ = 0 or when 𝜇̰̂ = 𝑋̰̅ .
Hence, the MLE for 𝜇̰ is given by 𝜇̰̂ = 𝑋̰̅ = the sample mean vector.
𝑁𝑝 𝑁 1
Replacing 𝜇̰ by 𝜇̰̂ in (4) we get log 𝐿(𝜇̰ , Ψ) = − log(2𝜋) + log|Ψ| − 𝑡𝑟. 𝐴Ψ
2 2 2
𝛿 𝑁 𝛿 1 𝛿
then the MLE for Ψ is given by 𝐿(𝜇̰̂ , Ψ) = 0 ⇒ log|Ψ| − 𝑡𝑟. 𝐴Ψ = 0
𝛿Ψ 2 𝛿Ψ 2 𝛿Ψ

18 | D R C
𝑎11 𝑎12 …. 𝑎1𝑝 Ψ11 Ψ12 …. Ψ1𝑝
𝑎21 𝑎22 …. 𝑎2𝑝 Ψ21 Ψ22 …. Ψ2𝑝
Now, 𝐴Ψ = ( … . …. …. …. ) ( )
…. …. …. ….
𝑎𝑝1 𝑎𝑝2 …. 𝑎𝑝𝑝 Ψ𝑝1 Ψ𝑝2 …. Ψ𝑝𝑝
𝑎11 Ψ11 + 𝑎12 Ψ12+. … + 𝑎1𝑝 Ψ1𝑝 …. …. ….
…. 𝑎21 Ψ12 + 𝑎22 Ψ22 +. … + 𝑎2𝑝 Ψ𝑝2 …. ….
= (
…. …. …. ….
)
…. …. … . 𝑎𝑝1 Ψ1𝑝 + 𝑎𝑝2 Ψ2𝑝 +. … + 𝑎𝑝𝑝 Ψ𝑝𝑝
Therefore, 𝑡𝑟. 𝐴Ψ = (𝑎11Ψ11 + 𝑎12 Ψ12 +. … + 𝑎1𝑝 Ψ1𝑝 ) + (𝑎21 Ψ12 + 𝑎22 Ψ22 +. … +
𝑎2𝑝 Ψ𝑝2 )+. … + (𝑎𝑝1 Ψ1𝑝 + 𝑎𝑝2 Ψ2𝑝 +. … + 𝑎𝑝𝑝 Ψ𝑝𝑝 )
𝑝 𝑝
= ∑𝑖=1 ∑𝑗=1 𝑎𝑖𝑗 Ψ𝑗𝑖
𝛿
Again, we know that log|𝑋| = (𝑋 −1)⸝
𝛿𝑥
𝑁 1 𝛿
Therefore, from (5) we get (Ψ −1)⸝ − ∑𝑝𝑖=1 ∑𝑝𝑗=1 𝑎𝑖𝑗 Ψ𝑗𝑖 = 0
2 2 𝛿Ψ
𝑁 −1 1 ⸝
⇒ Ψ − 𝐴 =0
2 2
𝑁 𝐴
⇒ Ψ −1
= {Since 𝐴⸝ = 𝐴}
2 2
−1 𝐴
⇒Ψ =
𝑁
𝐴
Therefore, Σ̂ = where 𝐴 = 𝑁
− 𝑋̰̅ )(𝑋̰ 𝛼 − 𝑋̰̅ )⸝
∑𝛼=1(𝑋̰ 𝛼
𝑁
Distribution of the Sample Mean Vector :—
Theorem 2 :— Let 𝑋̰ 𝛼 (𝛼 = 1, 2, …., 𝑁) be 𝑁 independent observation from 𝑁𝑝 (𝜇̰ , Σ) and
1
let 𝑋̰̅ = ∑𝑁 ̂ 𝐴 where 𝐴 = ∑𝑁
𝛼=1 𝑋̰ 𝛼 and Σ =
̅ ̅ ⸝ ̅
𝛼=1(𝑋̰ 𝛼 − 𝑋̰ )(𝑋̰ 𝛼 − 𝑋̰ ) then 𝑋̰ is distributed as
𝑁 𝑁
Σ
𝑁𝑝 (𝜇̰ , ) and is independent of the distribution of Σ̂, the MLE of Σ.
𝑁

Further, 𝐴 = 𝑁Σ̂ is distributed as ∑𝑁−1
𝛼=1 𝑍̰ 𝛼 𝑍̰ 𝛼 where 𝑍̰ 𝛼 (𝛼 = 1, 2, …., 𝑁) are independently
distributed each according to 𝑁𝑝 (0, Σ).
Proof – We are given that 𝑋̰ 𝛼 (𝛼 = 1, 2, …., 𝑁) ~ 𝑁𝑝 (𝜇̰ , Σ)
1 1 1
Let 𝐵 = (𝑏𝛼𝛽 ) be an orthogonal matrix with the last row given by ( , ,…., )
𝑁×𝑁 √𝑁 √𝑁 √𝑁
Let us define 𝑍𝛼 = ∑𝑁 𝛽=1 𝑏𝛼𝛽 𝑋̰ 𝛽 where (𝛼 = 1, 2, …., 𝑁)
Then in particular for 𝛼 = 𝑁 we have 𝑍̰ 𝑁 = ∑𝑁
𝛽=1 𝑏𝑁𝛽 𝑋̰ 𝛽 ……(1)
1
= ∑𝑁
𝛽=1 𝑋̰ 𝛽
√𝑁
1
= . 𝑁𝑋̰̅
√𝑁
= √𝑁𝑋̰̅
𝑍̰
⇒ 𝑋̰̅ = 𝑁 ……(2)
√𝑁
Now we will use the following 2 lemma’s which states that –
Lemma 1 :— 𝑋̰ 𝛼 (𝛼 = 1, 2, …., 𝑁) are independently distributed with 𝑁𝑝 (𝜇̰ , Σ). Let 𝐶 =
(𝑐𝛼𝛽 ) be an orthogonal matrix.
𝑁×𝑁
Let 𝑌̰𝛼 = ∑𝑁
𝛽=1 𝑐𝛼𝛽 𝑋̰ 𝛽 (𝛼 = 1, 2, …., 𝑁)
then 𝑌̰𝛼 ~ 𝑁𝑝 (𝜈̰ 𝛼 , Σ) where ∑𝑁
𝛽=1 𝑐𝛼𝛽 𝜇𝛽

19 | D R C
and 𝑌̰𝛼 (𝛼 = 1, 2, …., 𝑁) are independently distributed.
⸝ ⸝
Lemma 2 :— Under the same assumption ∑𝑁 𝑁
𝛼=1 𝑌̰𝛼 𝑌̰𝛼 = ∑𝛼=1 𝑋̰ 𝛼 𝑋̰ 𝛼
Now, by lemma 1, 𝑍̰ 𝛼 (𝛼 = 1, 2, …., 𝑁) are independently distributed each according to
multivariate normal with same dispersion matrix Σ but different mean vector.
Now, from (1) we have
𝐸 (𝑍̰ 𝑁 ) = 𝐸[∑𝑁𝛽=1 𝑏𝑁𝛽 𝑋̰ 𝛽 ]
𝑁
= ∑𝛽=1 𝑏𝑁𝛽 𝜇̰
= (𝑏𝑁1 + 𝑏𝑁2+. … + 𝑏𝑁𝑁 )𝜇̰
1 1 1
=( + +. … + ) 𝜇̰
√𝑁 √𝑁 √𝑁
𝑁
= 𝜇̰
√𝑁
= √𝑁𝜇̰
Therefore, 𝐸 [𝑋̰̅ ] = 𝐸[𝑍̰ 𝑁 ⁄√𝑁 ]
𝐸[𝑍̰ 𝑁 ]
=
√𝑁
√𝑁𝜇̰
=
√𝑁
= 𝜇̰
𝑍̰ 𝑁
𝑉 (𝑋̰̅ ) = 𝑉 ( )
√𝑁
1
= 𝑉 (𝑍̰ 𝑁 )
𝑁
1
= Σ
𝑁
Σ
=
𝑁
Σ
Therefore, 𝑋̰̅ ~ 𝑁𝑝 (𝜇̰ , )
𝑁
Now, 𝐴 = ∑𝛼=1(𝑋̰ 𝛼 − 𝑋̰̅ )(𝑋̰ 𝛼 − 𝑋̰̅ )⸝
𝑁
⸝ ̅ ̅⸝
= ∑𝑁 𝛼=1 𝑋̰ 𝛼 𝑋̰ 𝛼 − 𝑁𝑋̰ 𝑋̰
⸝ ̅ ̅⸝
= ∑𝑁 𝛼=1 𝑍̰ 𝛼 𝑍̰ 𝛼 − 𝑁𝑋̰ 𝑋̰
⸝ ⸝
= ∑𝑁 ̅
𝛼=1 𝑍̰ 𝛼 𝑍̰ 𝛼 − (√𝑁𝑋̰ )(√𝑁𝑋̰ )
̅
⸝ ⸝
= ∑𝑁 𝛼=1 𝑍̰ 𝛼 𝑍̰ 𝛼 − 𝑍̰ 𝛼 𝑍̰ 𝛼

⇒ 𝐴 = ∑𝑁−1 𝛼=1 𝑍̰ 𝛼 𝑍̰ 𝛼
Since, 𝑍̰ 𝑁 is independent of 𝑍̰ 1, 𝑍̰ 2 , …., 𝑍̰ 𝛼 therefore 𝑋̰̅ and 𝐴 are independently distributed.
Now, for 𝛼 = 1, 2, …., 𝑁 − 1 (i.e., 𝛼 = 𝑁)
𝐸 [𝑍̰ 𝛼 ] = 𝐸[∑𝑁 𝛽=1 𝑏𝛼𝛽 𝑋̰ 𝛽 ]
= ∑𝑁
𝛽=1 𝑏𝛼𝛽 𝜇̰
= (∑𝑁 𝛽=1 𝑏𝛼𝛽 𝑏𝑁𝛽 )√𝑁𝜇̰
=0
i.e., 𝐴 = ∑𝛼=1 𝑍̰ 𝛼 𝑍̰ 𝛼 ⸝ = ∑𝑛𝛼=1 𝑍̰ 𝛼 𝑍̰ 𝛼 ⸝ where 𝑍̰ 𝛼 (𝛼 = 1, 2, …., 𝑁) follows independent
𝑁−1

𝑁𝑝 (0, Σ)
NOTE : The quantity 𝑛 = 𝑁 − 1 is known as the df of the WISHART matrix.
Sample Covariance Matrix :—
20 | D R C
If 𝑋̰ 𝛼 (𝛼 = 1, 2, …., 𝑁) be a random vector from 𝑁𝑝 (𝜇̰ , Σ) then we know that the MLE for
𝐴 ⸝
Σ is given by Σ̂ = where 𝐴 = ∑𝑁 ̅ ̅ ⸝ 𝑁−1
𝛼=1(𝑋̰ 𝛼 − 𝑋̰ )(𝑋̰ 𝛼 − 𝑋̰ ) = ∑𝛼=1 𝑍̰ 𝛼 𝑍̰ 𝛼 where 𝑍̰ 𝛼 (𝛼 = 1, 2,
𝑁
…., 𝑁) ~ 𝑁𝑝 (0, Σ) then we have
𝐴
𝐸(Σ̂) = 𝐸 [ ]
𝑁
1
= 𝐸 [𝐴]
𝑁
1 ⸝
= 𝐸 [∑𝑁−1
𝛼=1 𝑍̰ 𝛼 𝑍̰ 𝛼 ]
𝑁
1 ⸝
= ∑𝑁−1
𝛼=1 𝐸 [𝑍̰ 𝛼 𝑍̰ 𝛼 ]
𝑁
1
= ∑𝑁−1
𝛼=1 𝑉 (𝑍̰ 𝛼 ) { 𝑉 (𝑍̰ 𝛼 ) = 𝐸 [𝑍̰ 𝛼 − 𝐸 (𝑍̰ 𝛼 )][𝑍̰ 𝛼 − 𝐸 (𝑍̰ 𝛼 )]⸝}
𝑁
1
= ∑𝑁−1
𝛼=1 Σ
𝑁
𝑁−1
= Σ≠Σ
𝑁
𝐴
Therefore, Σ = is not an unbiased estimator of Σ̂.
𝑁
Now, from (1) we have
𝑁−1
𝐸( Σ) = Σ
𝑁
𝐴 𝐴
⇒𝐸( )=Σ {𝑆𝑖𝑛𝑐𝑒 Σ̂ = }
𝑁−1 𝑁
(
⇒𝐸 𝑆 =Σ )
where we define the matrix S, which is given by
𝐴 1
S= = ∑𝑁 (𝑋̰ − 𝑋̰̅ )(𝑋̰ 𝛼 − 𝑋̰̅ )⸝
𝑁−1 𝑁−1 𝛼=1 𝛼
The matrix S is an unbiased estimator of Σ and is known as the sample covariance matrix.
Hotelling 𝑻𝟐 —
In univariate case, if 𝑋 ~ 𝑁 (𝜇, 𝜎 2 ), 𝜎 2 being unknown and if we desire to test
𝐻0 : 𝜇 = 𝜇0 against
𝐻1 : 𝜇 = 𝜇1
𝑥̅ −𝜇
then the test statistic is 𝑡 = 𝑠
⁄ 𝑛

which follows 𝑡 distribution with 𝑛 − 1 df.
where 𝑠 2 is the sampling variance
𝑥̅ is the sample mean and
𝑛 is the number of observation in the sample.
𝑛(𝑥̅ −𝜇)2
Squaring the test statistic 𝑡 2 = = 𝑛(𝑥̅ − 𝜇 )(𝑠 2 )−1(𝑥̅ − 𝜇 )
𝑠2
In multivariate case, if 𝑋̰ 𝑝×1 ~ 𝑁𝑝 (𝜇̰ , Σ), Σ being unknown then the multivariate analogous
of square of ‘𝑡’ is given by 𝑇 2 = 𝑁(𝑋̰̅ − 𝜇 )𝑆 −1(𝑋̰̅ − 𝜇 ) where 𝑥̰ ̅ is the sample mean vector
1
and S is the sample covariance matrix given by S = ∑𝑁 ̅ ̅ ⸝
𝛼=1(𝑋̰ 𝛼 − 𝑋̰ )(𝑋̰ 𝛼 − 𝑋̰ ) and 𝐸 (𝑆 )
𝑁−1
= Σ.
The statistic 𝑇 2 is called Hotelling 𝑇 2 statistic based on (𝑁 − 1) df in the honour of Harold
Hotelling, a pioneer in multivariate analysis who first introduce its sampling variance.

21 | D R C
NOTE : Hotelling 𝑇 2 statistic is the multivariate analogous of the univariate 𝑡 – test which
is used to test the hypothesis concerning mean when variance is unknown. In other words,
we use Hotelling 𝑇 2 statistic to test the null hypothesis about mean vector of a multivariate
normal distribution whose variance – covariance matrix is unknown.
Application of 𝑻𝟐 —
i) One Sample Case :—
𝑇 2 can be used in one sample problem to test the hypothesis that 𝐻0 : 𝜇̰ = 𝜇̰0 when the
variance – covariance matrix Σ is unknown.
Let 𝑋̰ 𝛼 (𝛼 = 1, 2, …., 𝑁) be a random vector from 𝑁𝑝 (𝜇̰ , Σ), Σ being unknown and we
are to test the hypothesis 𝐻0 : 𝜇̰ = 𝜇̰0
𝑇2 𝑁−𝑝
For this we can use the test statistic . ~ F with 𝑝 and (𝑁 − 𝑝) df.
𝑁−1 𝑝

where 𝑇 2 = 𝑁(𝑋̰̅ − 𝜇̰0) 𝑆 −1(𝑋̰̅ − 𝜇̰0) ……(1)
𝜇10
𝜇20
where 𝜇̰0 = ( … . )
𝜇𝑝0 𝑝×1
1
𝑋1𝛼 ∑𝑁
𝛼=1 𝑋1𝛼 𝑋̅1
𝑁
1
1 1 𝑋 ∑𝑁
𝛼=1 𝑋2𝛼
𝑋̅2
𝑋̰̅𝑝×1 = ∑𝑁
𝛼=1 𝑋̰ 𝛼 = ∑𝑁
𝛼=1 ( …2𝛼. ) = 𝑁 =
𝑁 𝑁 …. ….
𝑋𝑝𝛼 1 𝑁 (𝑋̅𝑝 )
( ∑𝛼=1 𝑋𝑝𝛼 )
𝑁
1
and S = ∑𝑁 (𝑋̰ − 𝑋̰̅ )(𝑋̰ 𝛼 − 𝑋̰ ̅ )⸝
𝑁−1 𝛼=1 𝛼
The calculated value of 𝑇 2 from (1) is compared against the table value of 𝑇 2 given by
𝑝(𝑁−1)
𝑇2 = . 𝐹𝛼 (𝑝, 𝑁 − 𝑝) ……(2)
𝑁−𝑝
at 𝛼% level of significance.
If calculated 𝑇 2 is less than table value of 𝑇 2 obtained by using (2), we may accept the
null hypothesis otherwise we reject it.
ii) Two Sample Case :—
𝑇 2 can also be used to test the null hypothesis about the equality of two population mean
vector where the variance – covariance matrices are assumed to be equal but unknown.
Let 𝑋̰ 𝛼 (1) (𝛼 = 1, 2, …., 𝑁) be a random vector from 𝑁𝑝 (𝜇̰ (1) , Σ) and 𝑋̰ 𝛼 (2) (𝛼 = 1, 2,
…., 𝑁) be a random vector from 𝑁𝑝 (𝜇̰ (2) , Σ) where the common variance – covariance
matrix is unknown.
Suppose we want to test the null hypothesis 𝐻0 : 𝜇̰ (1) = 𝜇̰ (2)
1 𝑁1 1
Now, 𝑋̰̅ (1) = ∑𝛼=1 𝑋̰ 𝛼 (1) is distributed as 𝑁𝑝 (𝜇̰ (1) , Σ)
𝑁1 𝑁1
1 𝑁2 (2) 1
̅ (2)
and 𝑋̰ = ∑𝛼=1 𝑋̰ 𝛼 (2) is distributed as 𝑁𝑝 (𝜇̰ , Σ)
𝑁2 𝑁2
𝑁1 𝑁2
Consequently, √ (𝑋̰̅ (1) − 𝑋̰̅ (2) ) is distributed as 𝑁𝑝 (0̰ , Σ) under 𝐻0 : 𝜇̰ (1) = 𝜇̰ (2)
𝑁1 +𝑁2

22 | D R C
1 𝑁 ⸝ 𝑁 ⸝
Let 𝑆 = 𝑁 1
[∑𝛼=1(𝑋̰𝛼 (1) − 𝑋̰̅ (1) )(𝑋̰𝛼 (1) − 𝑋̰̅ (1) ) + ∑𝛼=1
2
(𝑋̰𝛼 (2) − 𝑋̰̅ (2) )(𝑋̰𝛼 (2) − 𝑋̰̅ (2) ) ]
1 +𝑁2 −2
𝑁 +𝑁 −2
then (𝑁1 + 𝑁2 − 2)𝑆 is distributed as ∑𝛼=1 1 2
𝑍̰ 𝛼 𝑍̰ 𝛼 ⸝ where 𝑍̰ 𝛼 ~ 𝑁𝑝 (0̰ , Σ)
𝑁 𝑁 ⸝
Then 𝑇 2 = 1 2 (𝑋̰̅ (1) − 𝑋̰̅ (2) ) 𝑆 −1 (𝑋̰̅ (1) − 𝑋̰̅ (2) ) ……(1)
𝑁1 +𝑁2
is distributed as 𝑇 2 with (𝑁1 + 𝑁2 − 2) df. The required test statistic for testing the null
𝑇2 𝑁1 +𝑁2 −𝑝−1
hypothesis is . ~ F distribution with 𝑝 and 𝑁1 + 𝑁2 − 𝑝 − 1 df.
𝑁1 +𝑁2 −2 𝑝
We compare the calculated value of 𝑇 2 from (1) against the table value of 𝑇 2 given by
𝑝(𝑁1 +𝑁2 −2)
𝑇2 = . 𝐹𝛼 (𝑝, 𝑁1 + 𝑁2 − 𝑝 − 1)
(𝑁1 +𝑁2 −𝑝−1)
conclusion are drawn in usual manner.
Theorem 1 :— Hotelling 𝑇 2 is invariant under change in the unit of measurement.
Proof – Let 𝑋̰ 𝑝×1 ~ 𝑁𝑝 (𝜇̰ , Σ)
We want to test the null hypothesis 𝐻0 : 𝜇̰ = 𝜇̰0 where Σ is unknown.
For this purpose, we use the Hotelling 𝑇 2 statistic given by

𝑇 2 = 𝑁(𝑋̰̅ − 𝜇̰0 ) 𝑆 −1 (𝑋̰̅ − 𝜇̰0) ……(1)
1 𝑁
∑ 𝑋 𝑋̅1 𝜇10
𝑁 𝛼=1 1𝛼
1 𝑁
𝑋̅2 𝜇20
where 𝑋̰̅ = 𝑁 ∑𝛼=1 𝑋2𝛼 = and 𝜇̰0 = ( … . )
…. ….
1 𝑁 (𝑋̅𝑝 ) 𝜇𝑝0
(𝑁 ∑𝛼=1 𝑋𝑝𝛼 )
1
and 𝑆𝑝×𝑝 = ∑𝑁𝛼=1 (𝑋̰ 𝛼 − 𝑋̰̅ )(𝑋̰ 𝛼 − 𝑋̰̅ )⸝
𝑁−1
Now, let us consider the transformation 𝑌̰𝑝×1 = 𝐶𝑝×𝑝 𝑋̰ 𝑝×1 + 𝐷̰𝑝×1 where 𝐶𝑝×𝑝 is a non-
singular matrix.
The Hotelling 𝑇 2 statistic corresponding to random vector will be given by
𝑇𝑌̰ 2 = 𝑁(𝑌̰̅ − 𝜈̰ 0)⸝ 𝑆𝑌̰ −1(𝑌̰̅ − 𝜈̰ 0) ……(2)
where 𝜈̰ 0 = 𝐸 [𝑌̰ ] = 𝐸 [𝐶𝑋̰ + 𝐷̰ ] = 𝐶𝐸 [𝑋̰ ] + 𝐸 [𝐷̰ ] = 𝐶𝜇̰0 + 𝐷̰ ……(3)
Again, 𝑌̰𝛼 = 𝐶𝑋̰ 𝛼 + 𝐷̰
1 1 1
⇒ ∑𝑁 𝛼=1 𝑌̰𝛼 = 𝐶 ∑𝑁 𝛼=1 𝑋̰ 𝛼 + ∑𝑁 𝐷̰
𝑁 𝑁 𝑁 𝛼=1
⇒ 𝑌̰̅ = 𝐶𝑋̰̅ + 𝐷̰ ……(4)
Therefore, (𝑌̰̅ − 𝜈̰ 0 ) = 𝐶𝑋̰̅ + 𝐷̰ − 𝐶𝜇̰0 − 𝐷̰ = 𝐶(𝑋̰̅ − 𝜇̰0) ……(5)
1
Now, 𝑆𝑌̰ = ∑𝑁 ̅
𝛼=1(𝑌̰𝛼 − 𝑌̰ )(𝑌̰𝛼 − 𝑌̰ )
̅ ⸝
𝑁−1
1 ⸝
= ∑𝑁 ̅ ̅
𝛼=1(𝐶𝑋̰ 𝛼 + 𝐷̰ − 𝐶𝑋̰ − 𝐷̰ )(𝐶𝑋̰ 𝛼 + 𝐷̰ − 𝐶𝑋̰ − 𝐷̰ )
𝑁−1
1
= ∑𝑁 ̅ ̅ ⸝
𝛼=1{𝐶 (𝑋̰ 𝛼 − 𝑋̰ )}{𝐶 (𝑋̰ 𝛼 − 𝑋̰ )}
𝑁−1
1
=𝐶{ ∑𝑁 ̅ ̅ ⸝ ⸝
𝛼=1(𝑋̰ 𝛼 − 𝑋̰ )(𝑋̰ 𝛼 − 𝑋̰ ) } 𝐶
𝑁−1

= 𝐶𝑆𝐶 ……(6)
Putting (5) and (6) in (2) we get
𝑇𝑌̰ 2 = 𝑁(𝑌̰̅ − 𝜈̰ 0)⸝ 𝑆𝑌̰ −1(𝑌̰̅ − 𝜈̰ 0)

= 𝑁{𝐶(𝑋̰̅ − 𝜇̰0)} (𝐶𝑆𝐶 ⸝)−1{𝐶(𝑋̰̅ − 𝜇̰0 )}

23 | D R C

= 𝑁(𝑋̰̅ − 𝜇̰0) 𝐶 ⸝(𝐶 ⸝)−1 𝑆 −1𝐶 −1𝐶(𝑋̰̅ − 𝜇̰0 )

= 𝑁(𝑋̰̅ − 𝜇̰0) 𝑆 −1(𝑋̰̅ − 𝜇̰0)
= 𝑇 2 {based on 𝑋̰ }
Therefore, 𝑇 2 is invariant under change in the unit of measurement.
Principal Component Analysis —
The principal component analysis is used to construct new variables and replace the original
variables with as few new variables as possible. In other words, we want to reduce the
number of variables. In doing this naturally some information will have to be sacrificed.
However, one should keep in mind that this lost information is kept to a minimum.
If one wants to replace the entire set of 𝑝 variables by one new variable, one has to take the
first principal component, if one can afford to consider more than one new variable, one has
to take the desired number of principal component. Naturally the greater the number of
principal component chosen, less will be the information lost and greater will be the
performance of the new variables in explaining the internal relationship of the original
variables.
In most of the social sciences an investigator usually collects observation on a large number
of variables, as he does not know initially which variable are more important and useful for
his investigation. His next problem is to reduce this data, and here the principal component
analysis comes into picture. He can usually condense the whole information into a
manageable number of new variables and consider this variables in detail.
The principal components are linear combination of random variable which have special
properties in terms of variances. For example, the first principal component is the normalize
linear combination with maximum variances. It can be shown that the principal component
turn out to be the characteristic vectors of the variance – covariance matrix Σ.
Mathematically speaking, if 𝑉 (𝑋̰ ) = Σ, then the various principal components are given by

𝑢1 = 𝛽̰ (1) 𝑋̰

𝑢2 = 𝛽̰ (2) 𝑋̰
--------------

𝑢𝑝 = 𝛽̰ (𝑝) 𝑋̰

where 𝛽̰ (𝑖) (𝑖 = 1, 2, …., 𝑝) are the normalized characteristic vectors corresponding to the
𝑢1
𝑢2
largest characteristic root 𝜆(𝑖) of the characteristic equation |Σ − λI| = 0 and 𝑢̰ = (… .) is
𝑢𝑝
known as the vector of the principal component.
Multiple Correlation and Regression —
When the values of one variable are associated with or influenced by another variable, Karl
Pearson’s coefficient of correlation can be use as a measure of linear relationship between
them but sometimes there is inter-relation between many variables and the value of one
variable maybe influenced by many others. For example, the production of particular crop
(𝑋1 ) depends upon quality of seeds (𝑋2 ), fertility of soil (𝑋3 ), fertilizers used (𝑋4 ) and so
on. Whenever we are interested on studying the joint effect of a group of variable upon a

24 | D R C
variable not included in the group, our study is that of multiple correlation and multiple
regression.
Properties of Residuals —
1. ∑ 𝑋2 𝑋1.23 = 0
∑ 𝑋3 𝑋1.23 = 0
2. ∑ 𝑋1.2𝑋1.23 = ∑ 𝑋1 𝑋1.23
∑ 𝑋1.232 = ∑ 𝑋1 𝑋1.23
3. ∑ 𝑋1.2𝑋3.12 = 0
𝜔
4. Variance of residuals 𝜎1.232 = 𝑉 (𝑋1.23) = 𝜎1 2
𝜔11
1 𝑟12 𝑟13
𝜔 = |𝑟21 1 𝑟23|
𝑟31 𝑟32 1
and 𝜔11 is the cofactor of the element in the 1st row and 1st column of 𝜔.
Multiple Correlation Coefficient of 𝑿𝟏 on 𝑿𝟐 and 𝑿𝟑 —
Suppose that in trivariate distribution each of the variable 𝑋1 , 𝑋2 , 𝑋3 has N observations.
The multiple correlation coefficient of 𝑋1 on 𝑋2 and 𝑋3 denoted by 𝑅1.23 is the simple
correlation coefficient between 𝑋1 and the joint effect of 𝑋2 and 𝑋3 on 𝑋1 . After joint effect
of 𝑋2 and 𝑋3 on 𝑋1 is measured by the estimated value of regression plane of 𝑋1 on 𝑋2 and
𝑋3 i.e., 𝑒1.23 = 𝑏12.3𝑋2 + 𝑏13.2𝑋3 ……(1)
The residual is given by 𝑋1.23 = 𝑋1 − 𝑏12.3𝑋2 − 𝑏13.2𝑋3 = 𝑋1 − 𝑒1.23 {using (1)}
⇒ 𝑒1.23 = 𝑋1 − 𝑋1.23 ……(2)
where 𝑏12.3 and 𝑏13.2 are known as partial regression coefficient 𝑋1 on 𝑋2 and 𝑋2 on 𝑋3
respectively without lose of generality, we assume that variable 𝑋1 , 𝑋2 , 𝑋3 have been
measured from respective means so that 𝐸 (𝑋1 ) = 𝐸 (𝑋2 ) = 𝐸 (𝑋3 ) = 0
Then 𝐸 (𝑋1.23) = 0 and 𝐸 (𝑒1.23) = 0
By the definition of multiple correlation we have
𝐶𝑜𝑣(𝑋1 ,𝑒1.23 )
𝑅1.23 = 𝐶𝑜𝑟(𝑋1 , 𝑒1.23) = ……(3)
√𝑉(𝑋1 )√𝑉(𝑒1.23 )
Now, 𝐶𝑜𝑣 (𝑋1 , 𝑒1.23 ) = 𝐸 [𝑋1 − 𝐸 (𝑋1 )][𝑒1.23 − 𝐸 (𝑒1.23)]
= 𝐸 [𝑋1 𝑒1.23]
1
= ∑ 𝑋1 𝑒1.23
𝑁
1
= ∑ 𝑋1 (𝑋1 − 𝑋1.23) {using (2)}
𝑁
1 1
= ∑ 𝑋1 2 − ∑ 𝑋1 𝑋1.23
𝑁 𝑁
1 1
= ∑ 𝑋1 − ∑ 𝑋1.232
2
{using property number (2) of residual}
𝑁 𝑁
= 𝑉 (𝑋1 ) − 𝑉 (𝑋1.23 )
= 𝜎1 2 − 𝜎1.232 ……(4)
Again, 𝑉 (𝑒1.23) = 𝐸 [𝑒1.23 − 𝐸 (𝑒1.23 )]2
= 𝐸 (𝑒1.23 2)
= 𝐸 [(𝑋1 − 𝑋1.23 )2]
1
= ∑(𝑋1 − 𝑋1.23)2
𝑁
1
= ∑(𝑋1 2 + 𝑋1.23 2 − 2𝑋1 𝑋1.23)
𝑁
25 | D R C
1 1 1
= ∑ 𝑋1 2 + ∑ 𝑋1.232 − 2 ∑ 𝑋1 𝑋1.23
𝑁 𝑁 𝑁
1 1 1
= ∑ 𝑋1 + ∑ 𝑋1.23 − 2 ∑ 𝑋1.23 2
2 2
{using the property of residual}
𝑁 𝑁 𝑁
1 1
= ∑ 𝑋1 2 − ∑ 𝑋1.232
𝑁 𝑁
= 𝜎1 − 𝜎1.23 2
2
……(5)
2
Also 𝑉 (𝑋1 ) = 𝜎1 ……(6)
Thus, from (3) using (4), (5) and (6) we get
𝐶𝑜𝑣(𝑋1 ,𝑒1.23 ) 𝜎1 2 −𝜎1.23 2
𝑅1.23 = =
√𝑉(𝑋1 )√𝑉(𝑒1.23 ) √𝜎1 2 √𝜎1 2 −𝜎1.23 2
(𝜎 2 −𝜎 2 )2 𝜎 2 −𝜎 2 𝜎 2
𝑅1.23 2 = 21(𝜎 2 1.23 2) = 1 21.23 = 1 − 1.232 ……(7)
𝜎1 1 −𝜎1.23 𝜎1 𝜎1
But we know that from the property of residual the variance of 𝑋1.23 is given by 𝑉 (𝑋1.23) =
𝜔
𝜎1.232 = 𝜎12 ……(8)
𝜔11
1 𝑟12
𝑟13
where 𝜔 = |𝑟21 𝑟23| = 1(1 − 𝑟23 2 ) − 𝑟12(𝑟12 − 𝑟23𝑟31 ) + 𝑟13(𝑟32𝑟12 − 𝑟31 )
1
𝑟31 𝑟32
1
= 1 − 𝑟232 − 𝑟122 + 𝑟12𝑟23𝑟31 + 𝑟12𝑟23𝑟31 − 𝑟132
= 1 − 𝑟122 − 𝑟232 − 𝑟132 + 2𝑟12 𝑟23𝑟13
𝜔11 = cofactor of element in the 1st row and 1st column of 𝜔 = 1 − 𝑟232
From (7) and (8) we get
𝜔
𝜎1 2 𝜔
2 11
𝑅1.23 = 1 − 2
𝜎1
𝜔
=1−
𝜔11
1−𝑟12 2 −𝑟23 2 −𝑟13 2 +2𝑟12 𝑟23 𝑟13
=1−
1−𝑟23 2
1−𝑟23 −1+𝑟12 +𝑟23 2 +𝑟13 2 −2𝑟12 𝑟23 𝑟13
2 2
=
1−𝑟23 2
𝑟12 2 +𝑟13 2 −2𝑟12 𝑟23 𝑟13
𝑅1.23 = √
1−𝑟23 2
𝑟21 2 +𝑟23 2 −2𝑟21 𝑟23 𝑟13
Similarly, 𝑅2.13 = √
1−𝑟13 2
𝑟31 2 +𝑟32 2 −2𝑟12 𝑟32 𝑟31
and 𝑅3.12 = √
1−𝑟12 2
Remark: Since 𝑅1.23 is the simple correlation between 𝑋1 and 𝑒1.23 it must lie between −1
and +1. But since 𝑅1.23 is a non-negative quantity we conclude that 𝑅1.23 lies between 0
and 1.
Partial Correlation Coefficient —
Sometimes the correlation between 2 variables 𝑋1 and 𝑋2 maybe partly due to the correlation
of a 3rd variable 𝑋3 with both 𝑋1 and 𝑋2 in such a situation if we want to find the correlation
between 𝑋1 and 𝑋2 , it would be if the effect of 𝑋3 on each of 𝑋1 and 𝑋2 where eliminated.
The correlation between 𝑋1 and 𝑋2 after eliminating the linear effect of 𝑋3 on each of them

26 | D R C
is called partial correlation. Thus, the partial correlation coefficient between 𝑋1 and 𝑋2 after
𝐶𝑜𝑣(𝑋1.3 ,𝑋2.3 )
eliminating the linear effect of 𝑋3 is given by 𝑟12.3 = ……(1)
√𝑉(𝑋1.23 )√𝑉(𝑋2.3 )
where 𝑋1.3 = 𝑋1 − 𝑏13 𝑋3 and 𝑋2.3 = 𝑋2 − 𝑏23 𝑋3
Now, 𝐶𝑜𝑣 (𝑋1.3, 𝑋2.3) = 𝐸 [𝑋1.3 − 𝐸 (𝑋1.3 )][𝑋2.3 − 𝐸 (𝑋2.3)]
= 𝐸 [𝑋1.3𝑋2.3]
1
= ∑ 𝑋1.3𝑋2.3
𝑁
1
= ∑{(𝑋1 − 𝑏13𝑋3 )(𝑋2 − 𝑏23 𝑋3 )}
𝑁
1
= ∑{𝑋1 𝑋2 − 𝑏23 𝑋1 𝑋3 − 𝑏13 𝑋2 𝑋3 + 𝑏13 𝑏23 𝑋3 2 }
𝑁
1 1 1 1
= ∑ 𝑋1 𝑋2 − 𝑏23 ∑ 𝑋1 𝑋3 − 𝑏13 ∑ 𝑋2 𝑋3 + 𝑏13 𝑏23 ∑ 𝑋3 2
𝑁 𝑁 𝑁 𝑁
= 𝐶𝑜𝑣(𝑋1 , 𝑋2 ) − 𝑏23 𝐶𝑜𝑣(𝑋1 , 𝑋3 ) − 𝑏13 𝐶𝑜𝑣 (𝑋2 , 𝑋3 ) + 𝑏13 𝑏23 𝑉 (𝑋3 )
𝜎 𝜎 𝜎 𝜎
= 𝑟12𝜎1𝜎2 − 𝑟23 2 𝑟13 𝜎1𝜎3 − 𝑟13 1 𝑟23𝜎2 𝜎3 + 𝑟13 1 𝑟23 2 𝜎3 2
𝜎3 𝜎3 𝜎3 𝜎3
𝜎2 𝜎1
[∵ 𝑏23 = 𝑟23 𝑎𝑛𝑑 𝑏13 = 𝑟13 ]
𝜎3 𝜎3
= 𝑟12𝜎1𝜎2 − 𝑟23𝑟13𝜎1𝜎2 − 𝑟13𝑟23𝜎1𝜎2 + 𝑟13𝑟23𝜎1𝜎2
= 𝜎1𝜎2 (𝑟12 − 𝑟13𝑟23 ) ……(2)
Again, 𝑉 (𝑋1.3) = 𝐸 {𝑋1.3 − 𝐸 (𝑋1.3)}2
= 𝐸(𝑋1.32) {Since, 𝐸 (𝑋1.3) = 0}
1
= ∑ 𝑋1.32
𝑁
1
= ∑ 𝑋1 𝑋1.3 {using the properties of residuals}
𝑁
1
= ∑ 𝑋1 (𝑋1 − 𝑏13 𝑋3 )
𝑁
1
= ∑ 𝑋1 2 − 𝑏13 ∑ 𝑋1 𝑋3
𝑁
1
= ∑ 𝑋1 2 − 𝑏13 𝐶𝑜𝑣 (𝑋1 , 𝑋3 )
𝑁
𝜎1
= 𝜎12 − 𝑟13 𝑟 𝜎𝜎
𝜎3 13 1 3
= 𝜎12 − 𝑟132 𝜎1 2
= 𝜎12 (1 − 𝑟13 2) ……(3)
Similarly, we obtain 𝑉 (𝑋2.3) = 𝜎2 2(1 − 𝑟232 ) ……(4)
Thus, from (1) using (2), (3) and (4) we get
𝜎1 𝜎2 (𝑟12 −𝑟13 𝑟23 ) 𝑟12 −𝑟13 𝑟23
𝑟12.3 = =
√𝜎1 (1−𝑟13 2 )√𝜎2 2 (1−𝑟23 2 )
2 √(1−𝑟13 2 )(1−𝑟23 2 )
𝑟23 −𝑟21 𝑟31
Similarly, 𝑟23.1 =
√(1−𝑟21 2 )(1−𝑟31 2 )
𝑟13 −𝑟12 𝑟32
and 𝑟13.2 =
√(1−𝑟12 2 )(1−𝑟32 2 )
Multiple Correlation Coefficient in Terms of Total and Partial Correlation Coefficient

1 − 𝑅1.232 = (1 − 𝑟122 )(1 − 𝑟13.22)
𝑟12 2 +𝑟13 2 −2𝑟12 𝑟23 𝑟13
Proof – We know that 𝑅1.232 = 2
1−𝑟23
𝑟12 2 +𝑟13 2 −2𝑟12 𝑟23 𝑟13
Now, 1 − 𝑅1.232 = 1 −
1−𝑟23 2
27 | D R C
1−𝑟23 2 −𝑟12 2 −𝑟13 2 +2𝑟12 𝑟23 𝑟13
= ……(1)
1−𝑟23 2
𝑟13 −𝑟12 𝑟23
Again, we have 𝑟13.2 =
√(1−𝑟12 2 )(1−𝑟23 2 )
(𝑟 −𝑟12 𝑟23 )2
𝑟13.22 = (1−𝑟13 2)(1−𝑟 2
12 23 )
2 2 2
𝑟 −𝑟12 𝑟23 −2𝑟13 𝑟12 𝑟23
Now, 1 − 𝑟13.22 = 1 − 13 (1−𝑟 2 2
12 )(1−𝑟23 )
1−𝑟23 −𝑟12 −𝑟12 𝑟23 −𝑟13 2 +𝑟12 2 𝑟23 2 +2𝑟13 𝑟12 𝑟23
2 2 2 2
= (1−𝑟12 2 )(1−𝑟23 2 )
1−𝑟23 −𝑟12 −𝑟13 2 +2𝑟13 𝑟12 𝑟23
2 2
= (1−𝑟12 2 )(1−𝑟23 2 )
1−𝑟23 2 −𝑟12 2 −𝑟13 2 +2𝑟13 𝑟12 𝑟23
Again, (1 − 𝑟122)(1 − 𝑟13.22 ) = (1 − 𝑟122) × (1−𝑟12 2 )(1−𝑟23 2 )
2 2 2
1−𝑟23 −𝑟12 −𝑟13 +2𝑟13 𝑟12 𝑟23
= ……(2)
1−𝑟23 2
Thus, from (1) and (2) we get 1 − 𝑅1.23 2 = (1 − 𝑟122)(1 − 𝑟13.22)

28 | D R C

You might also like