Multivariate Normal Distributions
Multivariate Normal Distributions
3|D R C
1 2
𝑘 𝑝 ∞
− 𝑦𝑖
⇒ 1 ∏𝑖=1 [∫−∞ 𝑒 2 𝑑𝑦𝑖 ] = 1
|𝐴|2
𝑘
⇒ 1 ∏𝑝𝑖=1[√2𝜋] = 1
|𝐴|2
𝑝
𝑘
⇒ 1 (2𝜋) 2 = 1
|𝐴|2
1
|𝐴|2
⇒𝑘= 𝑝
(2𝜋)2
Thus, the multivariate normal density function is given by from (2)
1
1
|𝐴|2 − (𝑥̰ −𝑏̰ )⸝ 𝐴(𝑥̰ −𝑏̰ )
𝑓(𝑥1 , 𝑥2 , … . , 𝑥𝑝 ) = 𝑝𝑒 2 ……(7)
(2𝜋)2
We shall now show the significance of 𝑏̰ and A by finding the first and second moments of
𝑋1 , 𝑋2 , …., 𝑋𝑝 .
From (6) we note that the joint density function of 𝑦1 , 𝑦2 , …., 𝑦𝑝 is given by
1 ⸝
𝑔(𝑦1 , 𝑦2 , … . , 𝑦𝑝 ) = 𝑘𝑚𝑜𝑑|𝐶|𝑒 −2𝑌̰ 𝑌̰
1
1 ∑𝑝 2
= 𝑝 . 𝑒 −2 𝑖=1 𝑦𝑖
(2𝜋) 2
1 2
1
= 𝑝 ∏𝑝𝑖=1 𝑒 −2𝑦𝑖
(2𝜋) 2
1 2
𝑝 1
= ∏𝑖=1 [ 𝑒 −2𝑦𝑖 ]
√2𝜋
𝑝
= ∏𝑖=1 𝑔𝑖 (𝑦𝑖 )
where 𝑦𝑖 ~ N(0, 1)
Now, 𝑦𝑖 (𝑖 = 1, 2, …., 𝑝) are independently distributed each according to N(0, 1) i.e., they
are independent standard normal variate.
𝐸 (𝑦1 ) 0
𝐸 (𝑦2 ) 0
Thus, 𝐸(𝑦̰) = ( ) = ( ) = 0̰
…. ….
𝐸(𝑦𝑝 ) 0
1 0 …. 0
0 1 …. 0
and 𝑉(𝑦̰) = ( ) = 𝐼𝑝
…. …. …. ….
0 0 …. 1
From (5), we have 𝑥̰ − 𝑏̰ = 𝑐𝑦̰
Taking expectation on both side
𝐸 (𝑥̰ − 𝑏̰ ) = 𝐸(𝑐𝑦̰)
⇒ 𝐸 (𝑥̰ ) − 𝑏̰ = 𝑐𝐸(𝑦̰) = 0
⇒ 𝑏̰ = 𝐸 (𝑥̰ ) = 𝜇̰ (say)
Taking variance – covariance matrix on both sides
𝑉 (𝑥̰ − 𝑏̰ ) = 𝑉(𝑐𝑦̰)
⇒ 𝑉 (𝑥̰ ) = 𝐶𝑉(𝑦̰)𝐶 ⸝ = 𝐶𝐼𝑝 𝐶 ⸝ = 𝐶𝐶 ⸝
Since from (4) 𝐶 ⸝𝐴𝐶 = I
4|D R C
⇒ 𝐴 = (𝐶 ⸝)−1𝐶 −1
⇒ 𝐴 = (𝐶𝐶 ⸝)−1
⇒ 𝐶𝐶 ⸝ = 𝐴−1
or 𝐴−1 = 𝑉 (𝑥̰ ) = Σ (say)
Therefore, 𝐴 = Σ −1
Since A is positive definite matrix, therefore 𝐴−1 = Σ is also a positive definite matrix.
Thus, the multivariate or more precisely 𝑝 – variate normal density function is given by
{from (7)} –
1 ⸝
1 − (𝑥̰ −𝜇̰) Σ−1 (𝑥̰ −𝜇̰)
𝑓(𝑥1 , 𝑥2 , … . , 𝑥𝑝 ) = 𝑝 1 𝑒 2
(2𝜋)2 |Σ|2
Characteristic Function —
⸝ 1 ⸝
If 𝑋̰ ~ 𝑁𝑝 (𝜇̰ , Σ) then its characteristic function is given by 𝜙𝑋̰ (𝑡̰ ) = 𝑒 𝑖𝑡̰ 𝜇̰−2𝑡̰ Σ𝑡̰ where 𝑡̰ is a
𝑝 component vector or dummy vector.
Proof – In a univariate case if X ~ N(μ, 𝜎 2 ) then the characteristic function of X is define
as 𝜙𝑋 (𝑡) = 𝐸[𝑒 𝑖𝑡𝑋 ]
Similarly, for the multivariate case if 𝑋̰ ~ 𝑁𝑝 (𝜇̰ , Σ) then its characteristic function is define
as 𝜙𝑋̰ (𝑡̰ ) = 𝐸[𝑒 𝑖(𝑡1 𝑋1 +𝑡2 𝑋2 +.…+𝑡𝑝𝑋𝑝) ]
⸝
= 𝐸[𝑒 𝑖𝑡̰ 𝑋̰ ] where 𝑡 ⸝ = 𝑡1, 𝑡2 , …., 𝑡𝑝
Since 𝑋̰ ~ 𝑁𝑝 (𝜇̰ , Σ) therefore we have the pdf of 𝑋̰
1 ⸝ −1
1 − (𝑥̰ −𝜇̰) Σ (𝑥̰ −𝜇) ̰
𝑓(𝑥̰ ) = 𝑝 1𝑒 2
(2𝜋)2 |Σ|2
Since Σ −1 is a positive definite matrix therefore there exist a non-singular matrix 𝐶 =
(𝑐𝑖𝑗 )𝑝×𝑝 such that 𝐶 ⸝Σ −1𝐶 = 𝐼𝑝 ……(1)
⇒ (𝐶 ⸝)−1𝐶 ⸝Σ −1𝐶𝐶 −1 = (𝐶 ⸝ )−1𝐼𝑝 𝐶 −1
⇒ Σ −1 = (𝐶𝐶 ⸝)−1
⇒ Σ = 𝐶𝐶 ⸝ ……(2)
Again, from (1) taking determinant on both sides we get
|𝐶 ⸝Σ −1𝐶| = |𝐼𝑝 |
⇒ |𝐶 ⸝||Σ −1||𝐶| = 1
1
⇒ |𝐶|2 = | −1 |
Σ
1
1
Therefore, 𝑚𝑜𝑑|𝐶| = 1 = |Σ|2
|Σ−1 |2
𝑦1
𝑦2
Let 𝑋̰ − 𝜇̰ = 𝐶𝑌̰ where 𝑌̰ = (… .)
𝑦𝑝
1
Therefore, the Jacobian of transformation is |J| = 𝑚𝑜𝑑|𝐶| = |Σ|2
Therefore, the pdf of vector 𝑌̰ is
1
1 − (𝐶𝑌̰ )⸝ Σ−1 (𝐶𝑌̰ )
𝑔(𝑌̰ ) = 𝑝 1𝑒 2 |𝐽|
(2𝜋)2 |Σ|2
5|D R C
1 1
1 −2𝑌̰ ⸝ 𝐶 ⸝ Σ−1 𝐶𝑌̰
= 𝑝 1 𝑒 |Σ|2
(2𝜋)2 |Σ|2
1
1 − 𝑌̰ ⸝ 𝐼𝑝 𝑌̰
= 𝑝 𝑒 2
(2𝜋)2
1
1 ∑𝑝 2
= 𝑝 𝑒 −2 𝑗=1 𝑦𝑗
(2𝜋)2
1 2
𝑝 1
= ∏𝑗=1 [ 𝑒 −2𝑦𝑗 ]
√2𝜋
𝑝
= ∏𝑗=1 𝑔𝑖 (𝑦𝑗 )
where 𝑦𝑗 ~ N(0, 1)
Now, the characteristic function of 𝑌̰ with dummy vector 𝑢̰ is given by
⸝
𝜙𝑌̰ (𝑢̰ ) = 𝐸[𝑒 𝑖𝑢̰ 𝑌̰ ]
∑𝑝
= 𝐸 [𝑒 𝑖 𝑗=1 𝑢𝑗 𝑦𝑗 ]
𝑝
= 𝐸[∏𝑗=1 𝑒 𝑖𝑢𝑗 𝑦𝑗 ]
𝑝
= ∏𝑗=1 𝐸[𝑒 𝑖𝑢𝑗 𝑦𝑗 ]
1 2 2
𝑝
= ∏𝑗=1 𝜙𝑦𝑗 (𝑢𝑗 ) {Since 𝜙𝑋 (𝑡) = 𝑒 𝑖𝜇𝑡−2𝑡 𝜎
}
1
𝑝 −2𝑢𝑗 2
= ∏𝑗=1 𝑒
1
∑𝑝 2
= 𝑒 −2 𝑗=1 𝑢𝑗
1
−2𝑢̰ ⸝ 𝑢̰
=𝑒 ……(3)
Thus, the characteristic function of 𝑋̰ is given by
⸝
𝜙𝑋̰ (𝑡̰ ) = 𝐸[𝑒 𝑖𝑡̰ 𝑋̰ ]
⸝ (𝜇+𝐶𝑌̰ )
= 𝐸 [𝑒 𝑖𝑡̰ ̰ ]
𝑖𝑡̰ ⸝ 𝜇 𝑖𝑡̰ ⸝ 𝐶𝑌̰
= 𝐸[𝑒 ̰ . 𝑒 ]
⸝ ⸝
= 𝑒 𝑖𝑡̰ 𝜇̰ . 𝐸[𝑒 𝑖𝑡̰ 𝐶𝑌̰ ]
⸝ ⸝
= 𝑒 𝑖𝑡̰ 𝜇̰ 𝐸[𝑒 𝑖𝑢̰ 𝑌̰ ] where 𝑢̰ ⸝ = 𝑡̰ ⸝𝐶
⸝
= 𝑒 𝑖𝑡̰ 𝜇̰ 𝜙𝑌̰ (𝑢̰ )
⸝ 1 ⸝
= 𝑒 𝑖𝑡̰ 𝜇̰ . 𝑒 −2𝑢̰ 𝑢̰
{using (3)}
1 ⸝ ⸝
𝑖𝑡̰ ⸝ 𝜇 ̰ −2𝑡̰ 𝐶𝐶 𝑡̰
=𝑒 𝑒
1
𝑖𝑡̰ ⸝ 𝜇̰ −2𝑡̰ ⸝ Σ𝑡̰
=𝑒 𝑒
1
𝑖𝑡̰ ⸝ 𝜇̰−2𝑡̰ ⸝ Σ𝑡̰
=𝑒
which is the required characteristic function of 𝑋̰ .
Marginal Distribution —
Let 𝑋̰ 𝑝×1 = 𝑁𝑝 (𝜇̰ , Σ) and let 𝑋̰ , 𝜇̰ and Σ be partitioned as
6|D R C
𝑥1
𝑥2
…. 𝑋̰ 𝑘×1 (1)
𝑋̰ 𝑝×1 = 𝑥 =( )
𝑘+1 𝑋̰ (𝑝−𝑘)×1 (2)
….
( 𝑥𝑝 )
𝜇1
𝜇2
….
𝜇̰𝑘×1 (1)
𝜇̰𝑝×1 = 𝜇𝑘 =( )
𝜇𝑘+1 𝜇̰(𝑝−𝑘)×1(2)
….
( 𝜇𝑝 )
𝜎11 𝜎12 …. 𝜎1𝑘 𝜎1(𝑘+1) …. 𝜎1𝑝
𝜎21 𝜎22 …. 𝜎2𝑘 𝜎2(𝑘+1) …. 𝜎2𝑝
…. …. …. …. …. …. ….
Σ= 𝜎𝑘1 𝜎𝑘2 …. 𝜎𝑘𝑘 𝜎𝑘(𝑘+1) …. 𝜎𝑘𝑝
𝜎(𝑘+1)1 𝜎(𝑘+1)2 … . 𝜎(𝑘+1)𝑘 𝜎(𝑘+1)(𝑘+1) …. 𝜎(𝑘+1)𝑝
…. …. …. …. …. …. ….
( 𝜎𝑝1 𝜎𝑝2 …. 𝜎𝑝𝑘 𝜎𝑝(𝑘+1) …. 𝜎𝑝𝑝 )
Σ11𝑘×𝑘 Σ12 𝑘×(𝑝−𝑘)
Σ𝑝×𝑝 = ( )
Σ21 (𝑝−𝑘)×𝑘 Σ22 (𝑝−𝑘)×(𝑝−𝑘)
Such that Σ11𝑘×𝑘 = 𝑉(𝑋̰ (1) )
Σ22 (𝑝−𝑘)×(𝑝−𝑘) = 𝑉(𝑋̰ (2) )
Σ12𝑘×(𝑝−𝑘) = 𝐶𝑜𝑣(𝑋̰ (1) , 𝑋̰ (2) )
Σ21 (𝑝−𝑘)×𝑘 = 𝐶𝑜𝑣(𝑋̰ (2) , 𝑋̰ (1) )
Σ12 = Σ21
then the marginal distribution of sub vector 𝑋̰ (1) is
𝑋̰ 𝑘×1(1) ~ 𝑁𝑘 (𝜇̰ (1) , Σ11 )
and marginal distribution of sub vector 𝑋̰ (2) is
𝑋̰ (𝑝−𝑘)×1(2) ~ 𝑁𝑝−𝑘 (𝜇̰ (2) , Σ22)
Proof – The characteristic function of 𝑋̰ is defined as
⸝
𝜙𝑋̰ (𝑡̰ ) = 𝐸[𝑒 𝑖𝑡̰ 𝑋̰ ] ……(1)
𝑡1
𝑡2
where 𝑡̰ = (… .) is a dummy vector
𝑡𝑝
7|D R C
𝑡1
𝑡2
….
𝑡̰ 𝑘×1 (1)
Now, let us partition 𝑡̰ as 𝑡̰ 𝑝×1 = 𝑡𝑘 =( (2) )
𝑡𝑘+1 𝑡̰ (𝑝−𝑘)×1
….
( 𝑡𝑝 )
(1) ⸝ (1)
Now, the characteristic function of the vector 𝑋̰ (1) is 𝜙𝑋̰ (1) (𝑡̰ (1) ) = 𝐸 [𝑒 𝑖𝑡̰ 𝑋̰
] ……(2)
⸝
𝑡̰ (1) 𝑋̰ (1)
𝑖( (2) ) ( (2) )
Further (1) can be rewritten as 𝜙𝑋̰ (𝑡̰ ) = 𝐸 [𝑒 𝑡̰ 𝑋̰ ]
(1)
⸝ ⸝ 𝑋̰
𝑖(𝑡̰ (1) 𝑡̰ (2) )( (2) )
= 𝐸 [𝑒 𝑋̰ ]
⸝ ⸝
𝑖{𝑡̰ (1) 𝑋̰ (1) +𝑡̰ (2) 𝑋̰ (2) }
= 𝐸 [𝑒 ] ……(3)
On comparing (2) and (3) it is noted that the characteristic function of 𝑋̰ (1) can be obtain
from the characteristic function 𝑋̰ by putting 𝑡̰ (2) = 0̰
Thus, 𝜙𝑋̰ (1) (𝑡̰ (1) ) = [𝜙𝑋̰ (𝑡̰ )] (2) ……(4)
𝑡̰ =0̰
Now, since 𝑋̰ ~ 𝑁𝑝 (𝜇̰ , Σ)
Therefore, the characteristic function of 𝑋̰ is given by
⸝ 𝜇−1𝑡̰ ⸝ Σ𝑡̰
𝜙𝑋̰ (𝑡̰ ) = 𝑒 𝑖𝑡̰ ̰ 2
⸝ ⸝
𝑡̰ (1) 𝜇̰(1) 1 𝑡̰ (1) Σ Σ12 𝑡̰ (1)
𝑖( (2) ) ( (2) )−2( (2) ) ( 11 )( )
Σ21 Σ22 𝑡̰ (2)
= 𝑒 𝑡̰ 𝜇̰ 𝑡̰
(1)
⸝ ⸝ 𝜇̰ 1 ⸝ ⸝ Σ11 Σ12 𝑡̰ (1)
𝑖(𝑡̰ (1) 𝑡̰ (2) )( (2) )−2(𝑡̰ (1) 𝑡̰ (2) )(Σ )( )
𝜇̰ 21 Σ22 𝑡̰ (2)
=𝑒 ……(5)
(2)
Setting 𝑡̰ = 0̰ we get from (5)
(1)
⸝ 𝜇̰ 1 ⸝ Σ11 Σ12 𝑡̰ (1)
𝑖(𝑡̰ (1) 0̰ )( (2) )−2(𝑡̰ (1) 0̰ )(Σ )( )
𝜇̰ 21 Σ22 0̰
[𝜙𝑋̰ (𝑡̰ )] =𝑒
𝑡̰ (2) =0̰
⸝ 1
𝑖𝑡̰ (1) 𝜇̰(1) − (𝑡̰ (1) Σ11
⸝ ⸝ 𝑡̰ (1) )
(1) 2 𝑡̰ (1) Σ12 )(
⇒ 𝜙𝑋̰ (1) (𝑡̰ )=𝑒 0̰
(1) (1) (1)⸝ (1) 1 ⸝
= 𝑒 𝑖𝑡̰ 𝜇̰ −2𝑡̰ Σ11𝑡̰
which is the characteristic function of a 𝑘 – variate normal distribution with parameter 𝜇̰ (1) ,
Σ11. Hence, by uniqueness theorem of characteristic function we get 𝑋̰ 𝑘×1 (1) ~
𝑁𝑘 (𝜇̰ (1) , Σ11)
Similarly, it can be shown that 𝑋̰ (𝑝−𝑘)×1 (2) ~ 𝑁𝑝−𝑘 (𝜇̰ (2) , Σ22 )
8|D R C
Thus, if 𝑋̰ 𝑝×1 has a multivariate normal distribution then any sub vector of 𝑋̰ is also a
multivariate normal with parameters – means, variances and covariances obtained by taking
proper component of 𝜇̰ and Σ respectively.
Theorem 1 :— If 𝑋̰ ~ 𝑁𝑝 (𝜇̰ , Σ) then 𝑍̰ = 𝐷𝑋̰ ~ 𝑁𝑝 (𝐷𝜇̰ , 𝐷Σ𝐷 ⸝ ) where 𝐷 is a 𝑝 × 𝑝 non-
singular matrix of constant element.
Proof – Since 𝑋̰ ~ 𝑁𝑝 (𝜇̰ , Σ) therefore the pdf of 𝑋̰ is given by
1 ⸝
1 (𝑥̰ −𝜇̰) Σ−1 (𝑥̰ −𝜇̰)
𝑓(𝑥̰ ) = 𝑝 1 𝑒 −2
(2𝜋)2 |Σ|2
Given that 𝑍̰ = 𝐷𝑋̰ ⇒ 𝑋̰ = 𝐷 −1𝑍̰
Then the Jacobian of transformation is |J| = |𝐷|−1
1
=| |
𝐷
1
= √|
𝐷|2
1
= √|
𝐷||𝐷⸝ |
|Σ|
= √|
𝐷||Σ||𝐷⸝ |
1
|Σ|2
= 1
|𝐷||Σ||𝐷⸝ |2
Now, the pdf of 𝑍̰ is
1 ⸝
1 − (𝐷−1 𝑍̰ −𝜇̰) Σ−1 (𝐷−1 𝑍̰ −𝜇̰)
𝑔(𝑍̰ ) = 𝑝 1𝑒 2 |𝐽|
(2𝜋)2 |Σ|2
1
1 ⸝
1 −2[𝐷−1 (𝑍̰ −𝐷𝜇̰)] Σ−1 [𝐷−1 (𝑍̰ −𝐷𝜇̰)] |Σ|2
= 𝑝 1 𝑒 . 1
(2𝜋) 2 |Σ|2 |𝐷||Σ||𝐷⸝ |2
1 ⸝
1 − (𝑍̰ −𝐷𝜇̰) (𝐷⸝ )−1 Σ−1 𝐷−1 (𝑍̰ −𝐷𝜇̰) 1
= 𝑝 𝑒 2 1
(2𝜋) 2 |𝐷||Σ||𝐷⸝ |2
1 ⸝
1 − (𝑍̰ −𝐷𝜇̰) (𝐷Σ𝐷⸝ )−1 (𝑍̰ −𝐷𝜇̰)
= 𝑝 1 𝑒 2
(2𝜋) 2 |𝐷||Σ||𝐷⸝ |2
which is the pdf of 𝑝 – variate normal distribution with mean vector 𝐷𝜇̰ and variance –
covariance matrix 𝐷Σ𝐷 ⸝ .
Therefore, 𝑍̰ = 𝐷𝑋̰ ~ 𝑁𝑝 (𝐷𝜇̰ , 𝐷Σ𝐷 ⸝)
Theorem 2 :— Let 𝑋̰ ~ 𝑁𝑝 (𝜇̰ , Σ) then show that for any 𝑟 × 𝑝 matrix 𝐶, 𝐶𝑋̰ ~ 𝑁𝑟 (𝐶𝜇̰ , 𝐶Σ𝐶 ⸝)
Proof – Since 𝑋̰ ~ 𝑁𝑝 (𝜇̰ , Σ) therefore the characteristic function of 𝑋̰ with dummy vector 𝑡̰
⸝ 1 ⸝
is given by 𝜙𝑋̰ (𝑡̰ ) = 𝑒 𝑖𝑡̰ 𝜇̰−2𝑡̰ Σ𝑡̰ ……(1)
The characteristic function of 𝑍̰ 𝑟×1 = 𝐶𝑟×𝑝 𝑋̰ 𝑝×1 is given by
𝑡1
⸝ 𝑡
𝜙𝑍̰ (𝑡̰ ) = 𝐸[𝑒 𝑖𝑡̰ 𝑍̰ ] where 𝑡̰ 𝑟×1 = (…2.)
𝑡𝑟
⸝ 𝐶𝑋̰
= 𝐸[𝑒 𝑖𝑡̰ ]
9|D R C
⸝
= 𝐸[𝑒 𝑖𝑢̰ 𝑋̰ ] where 𝑢̰ ⸝ = 𝑡̰ ⸝𝐶 and 𝑢̰ = 𝐶 ⸝𝑡̰
= 𝜙𝑋̰ (𝑢̰ )
⸝ 𝜇−1𝑢̰ ⸝ Σ𝑢̰
= 𝑒 𝑖𝑢̰ ̰ 2
1
𝑖𝑡̰ ⸝ 𝐶𝜇̰−2𝑡̰ ⸝ 𝐶 Σ𝐶 ⸝ 𝑡̰
=𝑒
1
𝑖𝑡̰ ⸝ (𝐶𝜇̰)−2𝑡̰ ⸝ (𝐶 Σ𝐶 ⸝ )𝑡̰
=𝑒
which is the characteristic function of 𝑟 – variate normal distribution with mean 𝐶𝜇̰ and
variance – covariance matrix 𝐶Σ𝐶 ⸝. Hence, by uniqueness theorem of characteristic function
we have 𝑍̰ 𝑟×1 = 𝐶𝑟×𝑝 𝑋̰ 𝑝×1 ~ 𝑁𝑟 (𝐶𝜇̰ , 𝐶Σ𝐶 ⸝)
Example 1 – Prove that in a multivariate normal distribution any subset of the variable is
distributed independently of the complementary subset if and only if every variable in the
former set is uncorrelated with every variable in the other.
Solution – Let 𝑋̰ 𝑝×1 ~ 𝑁𝑝 (𝜇̰ , Σ) and 𝑋̰ , 𝜇̰ and Σ are partitioned as
𝑋̰ 𝑘×1(1) 𝜇̰𝑘×1(1)
𝑋̰ 𝑝×1 = ( ) and 𝜇̰𝑝×1 = ( )
𝑋̰ (𝑝−𝑘)×1(2) 𝜇̰(𝑝−𝑘)×1(2)
Σ11 𝑘×𝑘 Σ12𝑘×(𝑝−𝑘)
Σ𝑝×𝑝 = ( )
Σ21 (𝑝−𝑘)×𝑘 Σ22(𝑝−𝑘)×(𝑝−𝑘)
Sufficient Part :—
Let Σ12 = 0 = Σ21
We are required to prove that 𝑋̰ (1) and 𝑋̰ (2) are independent.
Σ 0
Since Σ = ( 11 )
0 Σ22
Therefore, |Σ| = Σ11Σ22
Therefore, 𝑚𝑜𝑑|Σ| = |Σ11||Σ22 |
−1 Σ11 −1 0
Also, Σ = ( )
0 Σ22−1
1 ⸝ −1
1 − (𝑥̰ −𝜇̰) Σ (𝑥̰ −𝜇) ̰
Now, the pdf of 𝑋̰ is 𝑓 (𝑥̰ ) = 𝑝 1𝑒 2
(2𝜋)2 |Σ|2
1 ⸝
1 − (𝑥̰ −𝜇̰) Σ−1 (𝑥̰ −𝜇̰)
= 𝑘 1 𝑝−𝑘 1 𝑒 2 ……(1)
(2𝜋)2 |Σ11 |2 (2𝜋) 2 |Σ22 |2
⸝
Now, (𝑥̰ − 𝜇̰ ) Σ −1(𝑥̰ − 𝜇̰ )
(1) (1) ⸝ (2) (2) ⸝ Σ11 −1 (𝑥̰ (1) − 𝜇̰ (1) )
0
= [(𝑥̰ − 𝜇̰ ) (𝑥̰ − 𝜇̰ ) ][ ][ ]
Σ22−1 (𝑥̰ (2) − 𝜇̰ (2) )
0
(1) (1) ⸝ −1 (2) (2) ⸝ −1
(𝑥̰ (1) − 𝜇̰ (1) )
= [(𝑥̰ − 𝜇̰ ) Σ11 (𝑥̰ − 𝜇̰ ) Σ22 ] [ (2) ]
(𝑥̰ − 𝜇̰ (2) )
⸝ ⸝
= [(𝑥̰ (1) − 𝜇̰ (1) ) Σ11−1(𝑥̰ (1) − 𝜇̰ (1) ) + (𝑥̰ (2) − 𝜇̰ (2) ) Σ22 −1(𝑥̰ (2) − 𝜇̰ (2) )]
From (1) and (2) we have
1 ⸝ ⸝
1 −2[(𝑥̰ (1)
−𝜇̰(1) ) Σ11 −1 (𝑥̰ (1) −𝜇̰(1) )+(𝑥̰ (2) −𝜇̰(2) ) Σ22 −1 (𝑥̰ (2) −𝜇̰(2) )]
𝑓 (𝑥̰ ) = 𝑘 1 𝑝−𝑘 1𝑒
(2𝜋)2 |Σ11 |2 (2𝜋) 2 |Σ22 |2
10 | D R C
1 ⸝ 1 ⸝
1 −2[(𝑥̰ (1) −𝜇̰(1) ) Σ11 −1 (𝑥̰ (1) −𝜇̰(1) )] 1 −2[(𝑥̰ (2) −𝜇̰(2) ) Σ22 −1 (𝑥̰ (2) −𝜇̰(2) )]
=[ 𝑘 1 𝑒 ][ 𝑝−𝑘 1 𝑒 ]
(2𝜋)2 |Σ11 |2 (2𝜋) 2 |Σ22 |2
11 | D R C
𝜎12 𝜎12 𝜎13 1 −1
⸝ 1 −1 0
then 𝐴Σ𝐴 = ( ) (𝜎21 𝜎2 2 𝜎23 ) (−1 1)
−1 1 −1
𝜎31 𝜎32 𝜎32 0 −1
𝜎12 − 𝜎21 𝜎12 − 𝜎2 2 𝜎13 − 𝜎23 1 −1
=( ) (−1 1 )
−𝜎1 2 + 𝜎21 − 𝜎31 −𝜎12 + 𝜎2 2 − 𝜎32 −𝜎13 + 𝜎23 − 𝜎32
0 −1
2 2 2 2
𝜎1 − 𝜎21 − 𝜎12 + 𝜎2 −𝜎1 + 𝜎21 + 𝜎12 − 𝜎2 − 𝜎13 + 𝜎23
=( )
−𝜎1 + 2𝜎12 − 𝜎2 + 𝜎32 − 𝜎31 𝜎1 2 + 𝜎22 + 𝜎3 2 − 2𝜎12 − 2𝜎23 + 2𝜎13
2 2
14 | D R C
Thus, from (1) using (3) and (5) we get
1 (𝑥 −𝜇 )2 𝜎 (𝑥 −𝜇 )2
1 − 2 { 1 21 −2 212 2 (𝑥1 −𝜇1 )(𝑥2 −𝜇2 )+ 2 22 }
𝑓(𝑥1𝑥2 ) = 𝑒 2(1−𝜌12 ) 𝜎1 𝜎1 𝜎2 𝜎2
2𝜋𝜎1 𝜎2 √1−𝜌12 2
which is the pdf of the bivariate normal distribution with parameter 𝜇1 , 𝜇2 , 𝜎12 , 𝜎2 2, 𝜌12
Theorem 3 :— If 𝑋̰ ~ 𝑁𝑝 (𝜇̰ , Σ) then the quadratic form in the exponent of multivariate
⸝
normal density function i.e., 𝑄 = (𝑥̰ − 𝜇̰ ) Σ −1(𝑥̰ − 𝜇̰ ) follows central 𝜒 2 distribution with
𝑝 df.
Proof – Since 𝑋̰ ~ 𝑁𝑝 (𝜇̰ , Σ)
1 ⸝ −1
1
Therefore, 𝑓 (𝑥̰ ) = 𝑝 1 𝑒 −2(𝑥̰ −𝜇̰) Σ (𝑥̰ −𝜇) ̰
(2𝜋)2 |Σ|2
Since Σ −1 is a positive definite matrix, therefore there exist a non-singular matrix 𝐶 such
that 𝐶 ⸝Σ −1 𝐶 = 𝐼𝑝
𝑦1
𝑦2
Now, let 𝑥̰ − 𝜇̰ = 𝐶𝑌̰ where 𝑌̰ = (… .)
𝑦𝑝
1
then the Jacobian of transformation is |J| = 𝑚𝑜𝑑|𝐶| = |Σ| 2
16 | D R C
1 1
~ 𝑁 [1 − (𝑥2 − 2), 4 − ]
3 3
5 𝑥2 11
~𝑁[ − , ]
3 3 3
c) Again, 𝑋̰ 1⁄𝑋̰ 2, 𝑋̰ 3 = 𝑥3
−1 𝑥 + 1 −1
6 0 2 6 0 0
~ 𝑁 [1 + (0 −1) ( ) ( ) , 4 − (0 −1) ( ) ( )]
0 3 𝑥3 − 2 0 3 −1
−1 𝑥 + 1
0.17 0 0.17 0 −1 0
~ 𝑁 [1 + (0 −1) ( ) ( 2 ) , 4 − (0 −1) ( ) ( )]
0 0.33 𝑥3 − 2 0 0.33 −1
𝑥 +1 0
~ 𝑁 [1 + (0 −0.33) ( 2 ) , 4 − (0 −0.33) ( )]
𝑥3 − 2 −1
~ 𝑁[1 + (−0.33)(𝑥3 − 2), 4 − 0.33]
~ 𝑁[1 − 0.33𝑥3 , 3.67]
Sampling Distribution for Mean Vector and Variance – Covariance Matrix —
𝑋1
𝑋2
Let 𝑋̰ 𝑝×1 = (… .) ~ 𝑁𝑝 (𝜇̰ , Σ) and let 𝑋̰ 𝛼 (𝛼 = 1, 2, …., 𝑁) be a random sample of size 𝑁
𝑋𝑝
from 𝑁𝑝 (𝜇̰ , Σ). It is written as 𝑋̰ 𝛼 (𝛼 = 1, 2, …., 𝑁) ~ 𝑁𝑝 (𝜇̰ , Σ)
𝑋11 𝑋12 𝑋1𝛼 𝑋1𝑝
𝑋21 𝑋22 𝑋 𝑋
Thus, 𝑋̰ 1 = ( … . ), 𝑋̰ 2 = ( … . ), …., 𝑋̰ 𝛼 = ( 2𝛼 ), …., 𝑋̰ 𝑝 = ( 2𝑝 )
…. ….
𝑋𝑝1 𝑋𝑝2 𝑋𝛼 𝑋𝑝𝑝
Sample Mean Vector :—
1 𝑁
𝑋1𝛼 ∑ 𝑋 𝑋̅1
𝑁 𝛼=1 1𝛼
1 𝑁
1
It is define as 𝑋̰̅ = ∑𝑁 𝑋̰ =
1 𝑁
∑ (
𝑋2𝛼
) = ∑𝛼=1 𝑋2𝛼 = 𝑋̅2
𝑁
𝑁 𝛼=1 𝛼 𝑁 𝛼=1 …. …. ….
𝑋𝑝𝛼 1 𝑁 (𝑋̅𝑝 )
(𝑁 ∑𝛼=1 𝑋𝑝𝛼 )
Matrix of Sum of Squares and Cross Products :—
The 𝑝 × 𝑝 symmetric matrix 𝐴𝑝×𝑝 = ∑𝑁 ̅ ̅ ⸝
𝛼=1(𝑋̰ 𝛼 − 𝑋̰ )𝑝×1 (𝑋̰ 𝛼 − 𝑋̰ ) 1×𝑝 is known as the matrix
of sum of squares and cross products of deviation about the sample mean.
𝑋1𝛼 − 𝑋̅1
𝑋2𝛼 − 𝑋̅2
Therefore, 𝐴𝑝×𝑝 = ∑𝑁 𝛼=1 (𝑋1𝛼 − 𝑋̅1 𝑋2𝛼 − 𝑋̅2 … . 𝑋𝑝𝛼 − 𝑋̅𝑝 )1×𝑝
….
(𝑋𝑝𝛼 − 𝑋̅𝑝 )𝑝×1
∑𝑁 ̅ 2
𝛼=1(𝑋1𝛼 − 𝑋1 ) ∑𝑁 ̅ ̅
𝛼=1(𝑋1𝛼 − 𝑋1 )(𝑋2𝛼 − 𝑋2 ) … . ∑𝑁 ̅ ̅
𝛼=1(𝑋1𝛼 − 𝑋1 ) (𝑋𝑝𝛼 − 𝑋𝑝 )
𝑁 ̅ ̅ ∑𝛼=1(𝑋2𝛼 − 𝑋̅2 )
𝑁 2
… . ∑𝛼=1(𝑋2𝛼 − 𝑋̅2 )(𝑋𝑝𝛼 − 𝑋̅𝑝 )
𝑁
= ∑𝛼=1(𝑋1𝛼 − 𝑋1 )(𝑋2𝛼 − 𝑋2 )
…. …. …. ….
∑ 𝑁 ( ̅ )( ̅ ∑ 𝑁 ( ̅ ̅
𝛼=1 𝑋2𝛼 − 𝑋2 )(𝑋𝑝𝛼 − 𝑋𝑝 ) ∑ 𝑁 ̅ 2
( 𝛼=1 𝑋1𝛼 − 𝑋1 𝑋𝑝𝛼 − 𝑋𝑝 ) …. 𝛼=1(𝑋𝑝𝛼 − 𝑋𝑝 ) )
𝑎11 𝑎12 …. 𝑎1𝑝
𝑎21 𝑎22 …. 𝑎2𝑝
= (…. …. …. ….)
𝑎𝑝1 𝑎𝑝2 …. 𝑎𝑝𝑝
17 | D R C
where 𝑎𝑖𝑗 = ∑𝑁 ̅ ̅ 𝑁 ̅ 2
𝛼=1(𝑋𝑖𝛼 − 𝑋𝑖 )(𝑋𝑗𝛼 − 𝑋𝑗 ) and 𝑎𝑖𝑖 = ∑𝛼=1(𝑋𝑖𝛼 − 𝑋𝑖 )
𝑝(𝑝+1)
The matrix A is also known as WISHART matrix. It is a symmetric matrix, and it has
2
distinct element.
Theorem 1 :— Let 𝑋̰ 𝛼 (𝛼 = 1, 2, …., 𝑁) be 𝑁 independent observation from 𝑁𝑝 (𝜇̰ , Σ). Then
𝐴
the maximum likelihood estimation for 𝜇̰ and Σ are given by 𝜇̰̂ = 𝑋̰̅ and Σ̂ = where 𝐴 =
𝑁
∑𝑁 ( ̅ )( ̅ ) ⸝
𝛼=1 𝛼𝑋̰ − 𝑋̰ 𝑋̰ 𝛼 − 𝑋̰
Proof – Since 𝑋̰ 𝛼 (𝛼 = 1, 2, …., 𝑁) are 𝑁 independent observation from 𝑁𝑝 (𝜇̰ , Σ) therefore
the density function of the 𝛼 𝑡ℎ observation 𝑋̰ 𝛼 is given by
1 ⸝
1 − (𝑋̰𝛼 −𝜇̰) Σ−1 (𝑋̰𝛼 −𝜇̰)
𝑓 (𝑋̰ 𝛼 ) = 𝑝 1𝑒 2
(2𝜋)2 |Σ|2
The likelihood function of the sample is
𝐿(𝜇̰ , Σ) = ∏𝑁
𝛼=1 𝑓(𝑋̰ 𝛼 )
1 ⸝
1 − (𝑋̰𝛼 −𝜇̰) Σ−1 (𝑋̰𝛼 −𝜇̰)
= ∏𝑁
𝛼=1 𝑝 1𝑒 2
(2𝜋)2 |Σ|2
1 ⸝ −1
1 − 2 ∑𝑁
𝛼=1(𝑋̰𝛼 −𝜇̰) Σ (𝑋̰𝛼 −𝜇̰)
= 𝑁𝑝 𝑁 𝑒 ……(1)
(2𝜋) 2 |Σ| 2
−1
Let a function Ψ = Σ then
𝑁
1 ⸝
Ψ2 − 2 ∑𝑁
𝛼=1(𝑋̰𝛼 −𝜇̰) Ψ(𝑋̰𝛼 −𝜇̰)
𝐿(𝜇̰ , Ψ) = 𝑁𝑝 𝑒 ……(2)
(2𝜋) 2
⸝
where ∑𝑁 𝛼=1(𝑋̰ 𝛼 − 𝜇̰ ) Ψ(𝑋̰ 𝛼 − 𝜇̰ )
⸝
= ∑𝑁 ̅ ̅ ̅ ̅
𝛼=1(𝑋̰ 𝛼 − 𝑋̰ + 𝑋̰ − 𝜇̰ ) Ψ(𝑋̰ 𝛼 − 𝑋̰ + 𝑋̰ − 𝜇̰ )
⸝
= ∑𝑁 ̅ ⸝ ̅ 𝑁 ̅ ⸝ 𝑁 ̅
𝛼=1(𝑋̰ 𝛼 − 𝑋̰ ) Ψ(𝑋̰ 𝛼 − 𝑋̰ ) + ∑𝛼=1(𝑋̰ 𝛼 − 𝑋̰ ) Ψ(𝑋̰ 𝛼 − 𝜇̰ ) + ∑𝛼=1(𝑋̰ − 𝜇̰ ) Ψ(𝑋̰ 𝛼 −
⸝
𝑋̰̅ ) + ∑𝑁 ̅ ̅
𝛼=1(𝑋̰ − 𝜇̰ ) Ψ(𝑋̰ − 𝜇̰ )
⸝
= ∑𝑁 ̅ ⸝ ̅ ̅ ̅
𝛼=1 𝑡𝑟 (𝑋̰ 𝛼 − 𝑋̰ ) Ψ(𝑋̰ 𝛼 − 𝑋̰ ) + 𝑁(𝑋̰ − 𝜇̰ ) Ψ(𝑋̰ − 𝜇̰ )
⸝
= ∑𝑁 ̅ ̅ ⸝ ̅ ̅
𝛼=1 𝑡𝑟 (𝑋̰ 𝛼 − 𝑋̰ )(𝑋̰ 𝛼 − 𝑋̰ ) Ψ + 𝑁(𝑋̰ − 𝜇̰ ) Ψ(𝑋̰ − 𝜇̰ )
⸝
= 𝑡𝑟 ∑𝑁 ̅ ̅ ⸝ ̅ ̅
𝛼=1(𝑋̰ 𝛼 − 𝑋̰ )(𝑋̰ 𝛼 − 𝑋̰ ) Ψ + 𝑁(𝑋̰ − 𝜇̰ ) Ψ(𝑋̰ − 𝜇̰ )
⸝
= 𝑡𝑟. 𝐴Ψ + 𝑁(𝑋̰̅ − 𝜇̰ ) Ψ(𝑋̰̅ − 𝜇̰ ) ……(3)
Putting (3) in (2) we get
𝑁𝑝 𝑁 1 𝑁 ⸝
log 𝐿(𝜇̰ , Ψ) = − log(2𝜋) + log|Ψ| − 𝑡𝑟. 𝐴Ψ − (𝑋̰̅ − 𝜇̰ ) Ψ(𝑋̰̅ − 𝜇̰ ) ……(4)
2 2 2 2
⸝
We see that log 𝐿(𝜇̰ , Ψ) will be maximum with respect to 𝜇 if (𝑋̰̅ − 𝜇̰ ) Ψ(𝑋̰̅ − 𝜇̰ ) is
⸝
minimum. Since (𝑋̰̅ − 𝜇̰ ) Ψ(𝑋̰̅ − 𝜇̰ ) is a positive definite quadratic form, its minimum
value is zero and it will be only when 𝑋̰̅ − 𝜇̰ = 0 or when 𝜇̰̂ = 𝑋̰̅ .
Hence, the MLE for 𝜇̰ is given by 𝜇̰̂ = 𝑋̰̅ = the sample mean vector.
𝑁𝑝 𝑁 1
Replacing 𝜇̰ by 𝜇̰̂ in (4) we get log 𝐿(𝜇̰ , Ψ) = − log(2𝜋) + log|Ψ| − 𝑡𝑟. 𝐴Ψ
2 2 2
𝛿 𝑁 𝛿 1 𝛿
then the MLE for Ψ is given by 𝐿(𝜇̰̂ , Ψ) = 0 ⇒ log|Ψ| − 𝑡𝑟. 𝐴Ψ = 0
𝛿Ψ 2 𝛿Ψ 2 𝛿Ψ
18 | D R C
𝑎11 𝑎12 …. 𝑎1𝑝 Ψ11 Ψ12 …. Ψ1𝑝
𝑎21 𝑎22 …. 𝑎2𝑝 Ψ21 Ψ22 …. Ψ2𝑝
Now, 𝐴Ψ = ( … . …. …. …. ) ( )
…. …. …. ….
𝑎𝑝1 𝑎𝑝2 …. 𝑎𝑝𝑝 Ψ𝑝1 Ψ𝑝2 …. Ψ𝑝𝑝
𝑎11 Ψ11 + 𝑎12 Ψ12+. … + 𝑎1𝑝 Ψ1𝑝 …. …. ….
…. 𝑎21 Ψ12 + 𝑎22 Ψ22 +. … + 𝑎2𝑝 Ψ𝑝2 …. ….
= (
…. …. …. ….
)
…. …. … . 𝑎𝑝1 Ψ1𝑝 + 𝑎𝑝2 Ψ2𝑝 +. … + 𝑎𝑝𝑝 Ψ𝑝𝑝
Therefore, 𝑡𝑟. 𝐴Ψ = (𝑎11Ψ11 + 𝑎12 Ψ12 +. … + 𝑎1𝑝 Ψ1𝑝 ) + (𝑎21 Ψ12 + 𝑎22 Ψ22 +. … +
𝑎2𝑝 Ψ𝑝2 )+. … + (𝑎𝑝1 Ψ1𝑝 + 𝑎𝑝2 Ψ2𝑝 +. … + 𝑎𝑝𝑝 Ψ𝑝𝑝 )
𝑝 𝑝
= ∑𝑖=1 ∑𝑗=1 𝑎𝑖𝑗 Ψ𝑗𝑖
𝛿
Again, we know that log|𝑋| = (𝑋 −1)⸝
𝛿𝑥
𝑁 1 𝛿
Therefore, from (5) we get (Ψ −1)⸝ − ∑𝑝𝑖=1 ∑𝑝𝑗=1 𝑎𝑖𝑗 Ψ𝑗𝑖 = 0
2 2 𝛿Ψ
𝑁 −1 1 ⸝
⇒ Ψ − 𝐴 =0
2 2
𝑁 𝐴
⇒ Ψ −1
= {Since 𝐴⸝ = 𝐴}
2 2
−1 𝐴
⇒Ψ =
𝑁
𝐴
Therefore, Σ̂ = where 𝐴 = 𝑁
− 𝑋̰̅ )(𝑋̰ 𝛼 − 𝑋̰̅ )⸝
∑𝛼=1(𝑋̰ 𝛼
𝑁
Distribution of the Sample Mean Vector :—
Theorem 2 :— Let 𝑋̰ 𝛼 (𝛼 = 1, 2, …., 𝑁) be 𝑁 independent observation from 𝑁𝑝 (𝜇̰ , Σ) and
1
let 𝑋̰̅ = ∑𝑁 ̂ 𝐴 where 𝐴 = ∑𝑁
𝛼=1 𝑋̰ 𝛼 and Σ =
̅ ̅ ⸝ ̅
𝛼=1(𝑋̰ 𝛼 − 𝑋̰ )(𝑋̰ 𝛼 − 𝑋̰ ) then 𝑋̰ is distributed as
𝑁 𝑁
Σ
𝑁𝑝 (𝜇̰ , ) and is independent of the distribution of Σ̂, the MLE of Σ.
𝑁
⸝
Further, 𝐴 = 𝑁Σ̂ is distributed as ∑𝑁−1
𝛼=1 𝑍̰ 𝛼 𝑍̰ 𝛼 where 𝑍̰ 𝛼 (𝛼 = 1, 2, …., 𝑁) are independently
distributed each according to 𝑁𝑝 (0, Σ).
Proof – We are given that 𝑋̰ 𝛼 (𝛼 = 1, 2, …., 𝑁) ~ 𝑁𝑝 (𝜇̰ , Σ)
1 1 1
Let 𝐵 = (𝑏𝛼𝛽 ) be an orthogonal matrix with the last row given by ( , ,…., )
𝑁×𝑁 √𝑁 √𝑁 √𝑁
Let us define 𝑍𝛼 = ∑𝑁 𝛽=1 𝑏𝛼𝛽 𝑋̰ 𝛽 where (𝛼 = 1, 2, …., 𝑁)
Then in particular for 𝛼 = 𝑁 we have 𝑍̰ 𝑁 = ∑𝑁
𝛽=1 𝑏𝑁𝛽 𝑋̰ 𝛽 ……(1)
1
= ∑𝑁
𝛽=1 𝑋̰ 𝛽
√𝑁
1
= . 𝑁𝑋̰̅
√𝑁
= √𝑁𝑋̰̅
𝑍̰
⇒ 𝑋̰̅ = 𝑁 ……(2)
√𝑁
Now we will use the following 2 lemma’s which states that –
Lemma 1 :— 𝑋̰ 𝛼 (𝛼 = 1, 2, …., 𝑁) are independently distributed with 𝑁𝑝 (𝜇̰ , Σ). Let 𝐶 =
(𝑐𝛼𝛽 ) be an orthogonal matrix.
𝑁×𝑁
Let 𝑌̰𝛼 = ∑𝑁
𝛽=1 𝑐𝛼𝛽 𝑋̰ 𝛽 (𝛼 = 1, 2, …., 𝑁)
then 𝑌̰𝛼 ~ 𝑁𝑝 (𝜈̰ 𝛼 , Σ) where ∑𝑁
𝛽=1 𝑐𝛼𝛽 𝜇𝛽
19 | D R C
and 𝑌̰𝛼 (𝛼 = 1, 2, …., 𝑁) are independently distributed.
⸝ ⸝
Lemma 2 :— Under the same assumption ∑𝑁 𝑁
𝛼=1 𝑌̰𝛼 𝑌̰𝛼 = ∑𝛼=1 𝑋̰ 𝛼 𝑋̰ 𝛼
Now, by lemma 1, 𝑍̰ 𝛼 (𝛼 = 1, 2, …., 𝑁) are independently distributed each according to
multivariate normal with same dispersion matrix Σ but different mean vector.
Now, from (1) we have
𝐸 (𝑍̰ 𝑁 ) = 𝐸[∑𝑁𝛽=1 𝑏𝑁𝛽 𝑋̰ 𝛽 ]
𝑁
= ∑𝛽=1 𝑏𝑁𝛽 𝜇̰
= (𝑏𝑁1 + 𝑏𝑁2+. … + 𝑏𝑁𝑁 )𝜇̰
1 1 1
=( + +. … + ) 𝜇̰
√𝑁 √𝑁 √𝑁
𝑁
= 𝜇̰
√𝑁
= √𝑁𝜇̰
Therefore, 𝐸 [𝑋̰̅ ] = 𝐸[𝑍̰ 𝑁 ⁄√𝑁 ]
𝐸[𝑍̰ 𝑁 ]
=
√𝑁
√𝑁𝜇̰
=
√𝑁
= 𝜇̰
𝑍̰ 𝑁
𝑉 (𝑋̰̅ ) = 𝑉 ( )
√𝑁
1
= 𝑉 (𝑍̰ 𝑁 )
𝑁
1
= Σ
𝑁
Σ
=
𝑁
Σ
Therefore, 𝑋̰̅ ~ 𝑁𝑝 (𝜇̰ , )
𝑁
Now, 𝐴 = ∑𝛼=1(𝑋̰ 𝛼 − 𝑋̰̅ )(𝑋̰ 𝛼 − 𝑋̰̅ )⸝
𝑁
⸝ ̅ ̅⸝
= ∑𝑁 𝛼=1 𝑋̰ 𝛼 𝑋̰ 𝛼 − 𝑁𝑋̰ 𝑋̰
⸝ ̅ ̅⸝
= ∑𝑁 𝛼=1 𝑍̰ 𝛼 𝑍̰ 𝛼 − 𝑁𝑋̰ 𝑋̰
⸝ ⸝
= ∑𝑁 ̅
𝛼=1 𝑍̰ 𝛼 𝑍̰ 𝛼 − (√𝑁𝑋̰ )(√𝑁𝑋̰ )
̅
⸝ ⸝
= ∑𝑁 𝛼=1 𝑍̰ 𝛼 𝑍̰ 𝛼 − 𝑍̰ 𝛼 𝑍̰ 𝛼
⸝
⇒ 𝐴 = ∑𝑁−1 𝛼=1 𝑍̰ 𝛼 𝑍̰ 𝛼
Since, 𝑍̰ 𝑁 is independent of 𝑍̰ 1, 𝑍̰ 2 , …., 𝑍̰ 𝛼 therefore 𝑋̰̅ and 𝐴 are independently distributed.
Now, for 𝛼 = 1, 2, …., 𝑁 − 1 (i.e., 𝛼 = 𝑁)
𝐸 [𝑍̰ 𝛼 ] = 𝐸[∑𝑁 𝛽=1 𝑏𝛼𝛽 𝑋̰ 𝛽 ]
= ∑𝑁
𝛽=1 𝑏𝛼𝛽 𝜇̰
= (∑𝑁 𝛽=1 𝑏𝛼𝛽 𝑏𝑁𝛽 )√𝑁𝜇̰
=0
i.e., 𝐴 = ∑𝛼=1 𝑍̰ 𝛼 𝑍̰ 𝛼 ⸝ = ∑𝑛𝛼=1 𝑍̰ 𝛼 𝑍̰ 𝛼 ⸝ where 𝑍̰ 𝛼 (𝛼 = 1, 2, …., 𝑁) follows independent
𝑁−1
𝑁𝑝 (0, Σ)
NOTE : The quantity 𝑛 = 𝑁 − 1 is known as the df of the WISHART matrix.
Sample Covariance Matrix :—
20 | D R C
If 𝑋̰ 𝛼 (𝛼 = 1, 2, …., 𝑁) be a random vector from 𝑁𝑝 (𝜇̰ , Σ) then we know that the MLE for
𝐴 ⸝
Σ is given by Σ̂ = where 𝐴 = ∑𝑁 ̅ ̅ ⸝ 𝑁−1
𝛼=1(𝑋̰ 𝛼 − 𝑋̰ )(𝑋̰ 𝛼 − 𝑋̰ ) = ∑𝛼=1 𝑍̰ 𝛼 𝑍̰ 𝛼 where 𝑍̰ 𝛼 (𝛼 = 1, 2,
𝑁
…., 𝑁) ~ 𝑁𝑝 (0, Σ) then we have
𝐴
𝐸(Σ̂) = 𝐸 [ ]
𝑁
1
= 𝐸 [𝐴]
𝑁
1 ⸝
= 𝐸 [∑𝑁−1
𝛼=1 𝑍̰ 𝛼 𝑍̰ 𝛼 ]
𝑁
1 ⸝
= ∑𝑁−1
𝛼=1 𝐸 [𝑍̰ 𝛼 𝑍̰ 𝛼 ]
𝑁
1
= ∑𝑁−1
𝛼=1 𝑉 (𝑍̰ 𝛼 ) { 𝑉 (𝑍̰ 𝛼 ) = 𝐸 [𝑍̰ 𝛼 − 𝐸 (𝑍̰ 𝛼 )][𝑍̰ 𝛼 − 𝐸 (𝑍̰ 𝛼 )]⸝}
𝑁
1
= ∑𝑁−1
𝛼=1 Σ
𝑁
𝑁−1
= Σ≠Σ
𝑁
𝐴
Therefore, Σ = is not an unbiased estimator of Σ̂.
𝑁
Now, from (1) we have
𝑁−1
𝐸( Σ) = Σ
𝑁
𝐴 𝐴
⇒𝐸( )=Σ {𝑆𝑖𝑛𝑐𝑒 Σ̂ = }
𝑁−1 𝑁
(
⇒𝐸 𝑆 =Σ )
where we define the matrix S, which is given by
𝐴 1
S= = ∑𝑁 (𝑋̰ − 𝑋̰̅ )(𝑋̰ 𝛼 − 𝑋̰̅ )⸝
𝑁−1 𝑁−1 𝛼=1 𝛼
The matrix S is an unbiased estimator of Σ and is known as the sample covariance matrix.
Hotelling 𝑻𝟐 —
In univariate case, if 𝑋 ~ 𝑁 (𝜇, 𝜎 2 ), 𝜎 2 being unknown and if we desire to test
𝐻0 : 𝜇 = 𝜇0 against
𝐻1 : 𝜇 = 𝜇1
𝑥̅ −𝜇
then the test statistic is 𝑡 = 𝑠
⁄ 𝑛
√
which follows 𝑡 distribution with 𝑛 − 1 df.
where 𝑠 2 is the sampling variance
𝑥̅ is the sample mean and
𝑛 is the number of observation in the sample.
𝑛(𝑥̅ −𝜇)2
Squaring the test statistic 𝑡 2 = = 𝑛(𝑥̅ − 𝜇 )(𝑠 2 )−1(𝑥̅ − 𝜇 )
𝑠2
In multivariate case, if 𝑋̰ 𝑝×1 ~ 𝑁𝑝 (𝜇̰ , Σ), Σ being unknown then the multivariate analogous
of square of ‘𝑡’ is given by 𝑇 2 = 𝑁(𝑋̰̅ − 𝜇 )𝑆 −1(𝑋̰̅ − 𝜇 ) where 𝑥̰ ̅ is the sample mean vector
1
and S is the sample covariance matrix given by S = ∑𝑁 ̅ ̅ ⸝
𝛼=1(𝑋̰ 𝛼 − 𝑋̰ )(𝑋̰ 𝛼 − 𝑋̰ ) and 𝐸 (𝑆 )
𝑁−1
= Σ.
The statistic 𝑇 2 is called Hotelling 𝑇 2 statistic based on (𝑁 − 1) df in the honour of Harold
Hotelling, a pioneer in multivariate analysis who first introduce its sampling variance.
21 | D R C
NOTE : Hotelling 𝑇 2 statistic is the multivariate analogous of the univariate 𝑡 – test which
is used to test the hypothesis concerning mean when variance is unknown. In other words,
we use Hotelling 𝑇 2 statistic to test the null hypothesis about mean vector of a multivariate
normal distribution whose variance – covariance matrix is unknown.
Application of 𝑻𝟐 —
i) One Sample Case :—
𝑇 2 can be used in one sample problem to test the hypothesis that 𝐻0 : 𝜇̰ = 𝜇̰0 when the
variance – covariance matrix Σ is unknown.
Let 𝑋̰ 𝛼 (𝛼 = 1, 2, …., 𝑁) be a random vector from 𝑁𝑝 (𝜇̰ , Σ), Σ being unknown and we
are to test the hypothesis 𝐻0 : 𝜇̰ = 𝜇̰0
𝑇2 𝑁−𝑝
For this we can use the test statistic . ~ F with 𝑝 and (𝑁 − 𝑝) df.
𝑁−1 𝑝
⸝
where 𝑇 2 = 𝑁(𝑋̰̅ − 𝜇̰0) 𝑆 −1(𝑋̰̅ − 𝜇̰0) ……(1)
𝜇10
𝜇20
where 𝜇̰0 = ( … . )
𝜇𝑝0 𝑝×1
1
𝑋1𝛼 ∑𝑁
𝛼=1 𝑋1𝛼 𝑋̅1
𝑁
1
1 1 𝑋 ∑𝑁
𝛼=1 𝑋2𝛼
𝑋̅2
𝑋̰̅𝑝×1 = ∑𝑁
𝛼=1 𝑋̰ 𝛼 = ∑𝑁
𝛼=1 ( …2𝛼. ) = 𝑁 =
𝑁 𝑁 …. ….
𝑋𝑝𝛼 1 𝑁 (𝑋̅𝑝 )
( ∑𝛼=1 𝑋𝑝𝛼 )
𝑁
1
and S = ∑𝑁 (𝑋̰ − 𝑋̰̅ )(𝑋̰ 𝛼 − 𝑋̰ ̅ )⸝
𝑁−1 𝛼=1 𝛼
The calculated value of 𝑇 2 from (1) is compared against the table value of 𝑇 2 given by
𝑝(𝑁−1)
𝑇2 = . 𝐹𝛼 (𝑝, 𝑁 − 𝑝) ……(2)
𝑁−𝑝
at 𝛼% level of significance.
If calculated 𝑇 2 is less than table value of 𝑇 2 obtained by using (2), we may accept the
null hypothesis otherwise we reject it.
ii) Two Sample Case :—
𝑇 2 can also be used to test the null hypothesis about the equality of two population mean
vector where the variance – covariance matrices are assumed to be equal but unknown.
Let 𝑋̰ 𝛼 (1) (𝛼 = 1, 2, …., 𝑁) be a random vector from 𝑁𝑝 (𝜇̰ (1) , Σ) and 𝑋̰ 𝛼 (2) (𝛼 = 1, 2,
…., 𝑁) be a random vector from 𝑁𝑝 (𝜇̰ (2) , Σ) where the common variance – covariance
matrix is unknown.
Suppose we want to test the null hypothesis 𝐻0 : 𝜇̰ (1) = 𝜇̰ (2)
1 𝑁1 1
Now, 𝑋̰̅ (1) = ∑𝛼=1 𝑋̰ 𝛼 (1) is distributed as 𝑁𝑝 (𝜇̰ (1) , Σ)
𝑁1 𝑁1
1 𝑁2 (2) 1
̅ (2)
and 𝑋̰ = ∑𝛼=1 𝑋̰ 𝛼 (2) is distributed as 𝑁𝑝 (𝜇̰ , Σ)
𝑁2 𝑁2
𝑁1 𝑁2
Consequently, √ (𝑋̰̅ (1) − 𝑋̰̅ (2) ) is distributed as 𝑁𝑝 (0̰ , Σ) under 𝐻0 : 𝜇̰ (1) = 𝜇̰ (2)
𝑁1 +𝑁2
22 | D R C
1 𝑁 ⸝ 𝑁 ⸝
Let 𝑆 = 𝑁 1
[∑𝛼=1(𝑋̰𝛼 (1) − 𝑋̰̅ (1) )(𝑋̰𝛼 (1) − 𝑋̰̅ (1) ) + ∑𝛼=1
2
(𝑋̰𝛼 (2) − 𝑋̰̅ (2) )(𝑋̰𝛼 (2) − 𝑋̰̅ (2) ) ]
1 +𝑁2 −2
𝑁 +𝑁 −2
then (𝑁1 + 𝑁2 − 2)𝑆 is distributed as ∑𝛼=1 1 2
𝑍̰ 𝛼 𝑍̰ 𝛼 ⸝ where 𝑍̰ 𝛼 ~ 𝑁𝑝 (0̰ , Σ)
𝑁 𝑁 ⸝
Then 𝑇 2 = 1 2 (𝑋̰̅ (1) − 𝑋̰̅ (2) ) 𝑆 −1 (𝑋̰̅ (1) − 𝑋̰̅ (2) ) ……(1)
𝑁1 +𝑁2
is distributed as 𝑇 2 with (𝑁1 + 𝑁2 − 2) df. The required test statistic for testing the null
𝑇2 𝑁1 +𝑁2 −𝑝−1
hypothesis is . ~ F distribution with 𝑝 and 𝑁1 + 𝑁2 − 𝑝 − 1 df.
𝑁1 +𝑁2 −2 𝑝
We compare the calculated value of 𝑇 2 from (1) against the table value of 𝑇 2 given by
𝑝(𝑁1 +𝑁2 −2)
𝑇2 = . 𝐹𝛼 (𝑝, 𝑁1 + 𝑁2 − 𝑝 − 1)
(𝑁1 +𝑁2 −𝑝−1)
conclusion are drawn in usual manner.
Theorem 1 :— Hotelling 𝑇 2 is invariant under change in the unit of measurement.
Proof – Let 𝑋̰ 𝑝×1 ~ 𝑁𝑝 (𝜇̰ , Σ)
We want to test the null hypothesis 𝐻0 : 𝜇̰ = 𝜇̰0 where Σ is unknown.
For this purpose, we use the Hotelling 𝑇 2 statistic given by
⸝
𝑇 2 = 𝑁(𝑋̰̅ − 𝜇̰0 ) 𝑆 −1 (𝑋̰̅ − 𝜇̰0) ……(1)
1 𝑁
∑ 𝑋 𝑋̅1 𝜇10
𝑁 𝛼=1 1𝛼
1 𝑁
𝑋̅2 𝜇20
where 𝑋̰̅ = 𝑁 ∑𝛼=1 𝑋2𝛼 = and 𝜇̰0 = ( … . )
…. ….
1 𝑁 (𝑋̅𝑝 ) 𝜇𝑝0
(𝑁 ∑𝛼=1 𝑋𝑝𝛼 )
1
and 𝑆𝑝×𝑝 = ∑𝑁𝛼=1 (𝑋̰ 𝛼 − 𝑋̰̅ )(𝑋̰ 𝛼 − 𝑋̰̅ )⸝
𝑁−1
Now, let us consider the transformation 𝑌̰𝑝×1 = 𝐶𝑝×𝑝 𝑋̰ 𝑝×1 + 𝐷̰𝑝×1 where 𝐶𝑝×𝑝 is a non-
singular matrix.
The Hotelling 𝑇 2 statistic corresponding to random vector will be given by
𝑇𝑌̰ 2 = 𝑁(𝑌̰̅ − 𝜈̰ 0)⸝ 𝑆𝑌̰ −1(𝑌̰̅ − 𝜈̰ 0) ……(2)
where 𝜈̰ 0 = 𝐸 [𝑌̰ ] = 𝐸 [𝐶𝑋̰ + 𝐷̰ ] = 𝐶𝐸 [𝑋̰ ] + 𝐸 [𝐷̰ ] = 𝐶𝜇̰0 + 𝐷̰ ……(3)
Again, 𝑌̰𝛼 = 𝐶𝑋̰ 𝛼 + 𝐷̰
1 1 1
⇒ ∑𝑁 𝛼=1 𝑌̰𝛼 = 𝐶 ∑𝑁 𝛼=1 𝑋̰ 𝛼 + ∑𝑁 𝐷̰
𝑁 𝑁 𝑁 𝛼=1
⇒ 𝑌̰̅ = 𝐶𝑋̰̅ + 𝐷̰ ……(4)
Therefore, (𝑌̰̅ − 𝜈̰ 0 ) = 𝐶𝑋̰̅ + 𝐷̰ − 𝐶𝜇̰0 − 𝐷̰ = 𝐶(𝑋̰̅ − 𝜇̰0) ……(5)
1
Now, 𝑆𝑌̰ = ∑𝑁 ̅
𝛼=1(𝑌̰𝛼 − 𝑌̰ )(𝑌̰𝛼 − 𝑌̰ )
̅ ⸝
𝑁−1
1 ⸝
= ∑𝑁 ̅ ̅
𝛼=1(𝐶𝑋̰ 𝛼 + 𝐷̰ − 𝐶𝑋̰ − 𝐷̰ )(𝐶𝑋̰ 𝛼 + 𝐷̰ − 𝐶𝑋̰ − 𝐷̰ )
𝑁−1
1
= ∑𝑁 ̅ ̅ ⸝
𝛼=1{𝐶 (𝑋̰ 𝛼 − 𝑋̰ )}{𝐶 (𝑋̰ 𝛼 − 𝑋̰ )}
𝑁−1
1
=𝐶{ ∑𝑁 ̅ ̅ ⸝ ⸝
𝛼=1(𝑋̰ 𝛼 − 𝑋̰ )(𝑋̰ 𝛼 − 𝑋̰ ) } 𝐶
𝑁−1
⸝
= 𝐶𝑆𝐶 ……(6)
Putting (5) and (6) in (2) we get
𝑇𝑌̰ 2 = 𝑁(𝑌̰̅ − 𝜈̰ 0)⸝ 𝑆𝑌̰ −1(𝑌̰̅ − 𝜈̰ 0)
⸝
= 𝑁{𝐶(𝑋̰̅ − 𝜇̰0)} (𝐶𝑆𝐶 ⸝)−1{𝐶(𝑋̰̅ − 𝜇̰0 )}
23 | D R C
⸝
= 𝑁(𝑋̰̅ − 𝜇̰0) 𝐶 ⸝(𝐶 ⸝)−1 𝑆 −1𝐶 −1𝐶(𝑋̰̅ − 𝜇̰0 )
⸝
= 𝑁(𝑋̰̅ − 𝜇̰0) 𝑆 −1(𝑋̰̅ − 𝜇̰0)
= 𝑇 2 {based on 𝑋̰ }
Therefore, 𝑇 2 is invariant under change in the unit of measurement.
Principal Component Analysis —
The principal component analysis is used to construct new variables and replace the original
variables with as few new variables as possible. In other words, we want to reduce the
number of variables. In doing this naturally some information will have to be sacrificed.
However, one should keep in mind that this lost information is kept to a minimum.
If one wants to replace the entire set of 𝑝 variables by one new variable, one has to take the
first principal component, if one can afford to consider more than one new variable, one has
to take the desired number of principal component. Naturally the greater the number of
principal component chosen, less will be the information lost and greater will be the
performance of the new variables in explaining the internal relationship of the original
variables.
In most of the social sciences an investigator usually collects observation on a large number
of variables, as he does not know initially which variable are more important and useful for
his investigation. His next problem is to reduce this data, and here the principal component
analysis comes into picture. He can usually condense the whole information into a
manageable number of new variables and consider this variables in detail.
The principal components are linear combination of random variable which have special
properties in terms of variances. For example, the first principal component is the normalize
linear combination with maximum variances. It can be shown that the principal component
turn out to be the characteristic vectors of the variance – covariance matrix Σ.
Mathematically speaking, if 𝑉 (𝑋̰ ) = Σ, then the various principal components are given by
⸝
𝑢1 = 𝛽̰ (1) 𝑋̰
⸝
𝑢2 = 𝛽̰ (2) 𝑋̰
--------------
⸝
𝑢𝑝 = 𝛽̰ (𝑝) 𝑋̰
⸝
where 𝛽̰ (𝑖) (𝑖 = 1, 2, …., 𝑝) are the normalized characteristic vectors corresponding to the
𝑢1
𝑢2
largest characteristic root 𝜆(𝑖) of the characteristic equation |Σ − λI| = 0 and 𝑢̰ = (… .) is
𝑢𝑝
known as the vector of the principal component.
Multiple Correlation and Regression —
When the values of one variable are associated with or influenced by another variable, Karl
Pearson’s coefficient of correlation can be use as a measure of linear relationship between
them but sometimes there is inter-relation between many variables and the value of one
variable maybe influenced by many others. For example, the production of particular crop
(𝑋1 ) depends upon quality of seeds (𝑋2 ), fertility of soil (𝑋3 ), fertilizers used (𝑋4 ) and so
on. Whenever we are interested on studying the joint effect of a group of variable upon a
24 | D R C
variable not included in the group, our study is that of multiple correlation and multiple
regression.
Properties of Residuals —
1. ∑ 𝑋2 𝑋1.23 = 0
∑ 𝑋3 𝑋1.23 = 0
2. ∑ 𝑋1.2𝑋1.23 = ∑ 𝑋1 𝑋1.23
∑ 𝑋1.232 = ∑ 𝑋1 𝑋1.23
3. ∑ 𝑋1.2𝑋3.12 = 0
𝜔
4. Variance of residuals 𝜎1.232 = 𝑉 (𝑋1.23) = 𝜎1 2
𝜔11
1 𝑟12 𝑟13
𝜔 = |𝑟21 1 𝑟23|
𝑟31 𝑟32 1
and 𝜔11 is the cofactor of the element in the 1st row and 1st column of 𝜔.
Multiple Correlation Coefficient of 𝑿𝟏 on 𝑿𝟐 and 𝑿𝟑 —
Suppose that in trivariate distribution each of the variable 𝑋1 , 𝑋2 , 𝑋3 has N observations.
The multiple correlation coefficient of 𝑋1 on 𝑋2 and 𝑋3 denoted by 𝑅1.23 is the simple
correlation coefficient between 𝑋1 and the joint effect of 𝑋2 and 𝑋3 on 𝑋1 . After joint effect
of 𝑋2 and 𝑋3 on 𝑋1 is measured by the estimated value of regression plane of 𝑋1 on 𝑋2 and
𝑋3 i.e., 𝑒1.23 = 𝑏12.3𝑋2 + 𝑏13.2𝑋3 ……(1)
The residual is given by 𝑋1.23 = 𝑋1 − 𝑏12.3𝑋2 − 𝑏13.2𝑋3 = 𝑋1 − 𝑒1.23 {using (1)}
⇒ 𝑒1.23 = 𝑋1 − 𝑋1.23 ……(2)
where 𝑏12.3 and 𝑏13.2 are known as partial regression coefficient 𝑋1 on 𝑋2 and 𝑋2 on 𝑋3
respectively without lose of generality, we assume that variable 𝑋1 , 𝑋2 , 𝑋3 have been
measured from respective means so that 𝐸 (𝑋1 ) = 𝐸 (𝑋2 ) = 𝐸 (𝑋3 ) = 0
Then 𝐸 (𝑋1.23) = 0 and 𝐸 (𝑒1.23) = 0
By the definition of multiple correlation we have
𝐶𝑜𝑣(𝑋1 ,𝑒1.23 )
𝑅1.23 = 𝐶𝑜𝑟(𝑋1 , 𝑒1.23) = ……(3)
√𝑉(𝑋1 )√𝑉(𝑒1.23 )
Now, 𝐶𝑜𝑣 (𝑋1 , 𝑒1.23 ) = 𝐸 [𝑋1 − 𝐸 (𝑋1 )][𝑒1.23 − 𝐸 (𝑒1.23)]
= 𝐸 [𝑋1 𝑒1.23]
1
= ∑ 𝑋1 𝑒1.23
𝑁
1
= ∑ 𝑋1 (𝑋1 − 𝑋1.23) {using (2)}
𝑁
1 1
= ∑ 𝑋1 2 − ∑ 𝑋1 𝑋1.23
𝑁 𝑁
1 1
= ∑ 𝑋1 − ∑ 𝑋1.232
2
{using property number (2) of residual}
𝑁 𝑁
= 𝑉 (𝑋1 ) − 𝑉 (𝑋1.23 )
= 𝜎1 2 − 𝜎1.232 ……(4)
Again, 𝑉 (𝑒1.23) = 𝐸 [𝑒1.23 − 𝐸 (𝑒1.23 )]2
= 𝐸 (𝑒1.23 2)
= 𝐸 [(𝑋1 − 𝑋1.23 )2]
1
= ∑(𝑋1 − 𝑋1.23)2
𝑁
1
= ∑(𝑋1 2 + 𝑋1.23 2 − 2𝑋1 𝑋1.23)
𝑁
25 | D R C
1 1 1
= ∑ 𝑋1 2 + ∑ 𝑋1.232 − 2 ∑ 𝑋1 𝑋1.23
𝑁 𝑁 𝑁
1 1 1
= ∑ 𝑋1 + ∑ 𝑋1.23 − 2 ∑ 𝑋1.23 2
2 2
{using the property of residual}
𝑁 𝑁 𝑁
1 1
= ∑ 𝑋1 2 − ∑ 𝑋1.232
𝑁 𝑁
= 𝜎1 − 𝜎1.23 2
2
……(5)
2
Also 𝑉 (𝑋1 ) = 𝜎1 ……(6)
Thus, from (3) using (4), (5) and (6) we get
𝐶𝑜𝑣(𝑋1 ,𝑒1.23 ) 𝜎1 2 −𝜎1.23 2
𝑅1.23 = =
√𝑉(𝑋1 )√𝑉(𝑒1.23 ) √𝜎1 2 √𝜎1 2 −𝜎1.23 2
(𝜎 2 −𝜎 2 )2 𝜎 2 −𝜎 2 𝜎 2
𝑅1.23 2 = 21(𝜎 2 1.23 2) = 1 21.23 = 1 − 1.232 ……(7)
𝜎1 1 −𝜎1.23 𝜎1 𝜎1
But we know that from the property of residual the variance of 𝑋1.23 is given by 𝑉 (𝑋1.23) =
𝜔
𝜎1.232 = 𝜎12 ……(8)
𝜔11
1 𝑟12
𝑟13
where 𝜔 = |𝑟21 𝑟23| = 1(1 − 𝑟23 2 ) − 𝑟12(𝑟12 − 𝑟23𝑟31 ) + 𝑟13(𝑟32𝑟12 − 𝑟31 )
1
𝑟31 𝑟32
1
= 1 − 𝑟232 − 𝑟122 + 𝑟12𝑟23𝑟31 + 𝑟12𝑟23𝑟31 − 𝑟132
= 1 − 𝑟122 − 𝑟232 − 𝑟132 + 2𝑟12 𝑟23𝑟13
𝜔11 = cofactor of element in the 1st row and 1st column of 𝜔 = 1 − 𝑟232
From (7) and (8) we get
𝜔
𝜎1 2 𝜔
2 11
𝑅1.23 = 1 − 2
𝜎1
𝜔
=1−
𝜔11
1−𝑟12 2 −𝑟23 2 −𝑟13 2 +2𝑟12 𝑟23 𝑟13
=1−
1−𝑟23 2
1−𝑟23 −1+𝑟12 +𝑟23 2 +𝑟13 2 −2𝑟12 𝑟23 𝑟13
2 2
=
1−𝑟23 2
𝑟12 2 +𝑟13 2 −2𝑟12 𝑟23 𝑟13
𝑅1.23 = √
1−𝑟23 2
𝑟21 2 +𝑟23 2 −2𝑟21 𝑟23 𝑟13
Similarly, 𝑅2.13 = √
1−𝑟13 2
𝑟31 2 +𝑟32 2 −2𝑟12 𝑟32 𝑟31
and 𝑅3.12 = √
1−𝑟12 2
Remark: Since 𝑅1.23 is the simple correlation between 𝑋1 and 𝑒1.23 it must lie between −1
and +1. But since 𝑅1.23 is a non-negative quantity we conclude that 𝑅1.23 lies between 0
and 1.
Partial Correlation Coefficient —
Sometimes the correlation between 2 variables 𝑋1 and 𝑋2 maybe partly due to the correlation
of a 3rd variable 𝑋3 with both 𝑋1 and 𝑋2 in such a situation if we want to find the correlation
between 𝑋1 and 𝑋2 , it would be if the effect of 𝑋3 on each of 𝑋1 and 𝑋2 where eliminated.
The correlation between 𝑋1 and 𝑋2 after eliminating the linear effect of 𝑋3 on each of them
26 | D R C
is called partial correlation. Thus, the partial correlation coefficient between 𝑋1 and 𝑋2 after
𝐶𝑜𝑣(𝑋1.3 ,𝑋2.3 )
eliminating the linear effect of 𝑋3 is given by 𝑟12.3 = ……(1)
√𝑉(𝑋1.23 )√𝑉(𝑋2.3 )
where 𝑋1.3 = 𝑋1 − 𝑏13 𝑋3 and 𝑋2.3 = 𝑋2 − 𝑏23 𝑋3
Now, 𝐶𝑜𝑣 (𝑋1.3, 𝑋2.3) = 𝐸 [𝑋1.3 − 𝐸 (𝑋1.3 )][𝑋2.3 − 𝐸 (𝑋2.3)]
= 𝐸 [𝑋1.3𝑋2.3]
1
= ∑ 𝑋1.3𝑋2.3
𝑁
1
= ∑{(𝑋1 − 𝑏13𝑋3 )(𝑋2 − 𝑏23 𝑋3 )}
𝑁
1
= ∑{𝑋1 𝑋2 − 𝑏23 𝑋1 𝑋3 − 𝑏13 𝑋2 𝑋3 + 𝑏13 𝑏23 𝑋3 2 }
𝑁
1 1 1 1
= ∑ 𝑋1 𝑋2 − 𝑏23 ∑ 𝑋1 𝑋3 − 𝑏13 ∑ 𝑋2 𝑋3 + 𝑏13 𝑏23 ∑ 𝑋3 2
𝑁 𝑁 𝑁 𝑁
= 𝐶𝑜𝑣(𝑋1 , 𝑋2 ) − 𝑏23 𝐶𝑜𝑣(𝑋1 , 𝑋3 ) − 𝑏13 𝐶𝑜𝑣 (𝑋2 , 𝑋3 ) + 𝑏13 𝑏23 𝑉 (𝑋3 )
𝜎 𝜎 𝜎 𝜎
= 𝑟12𝜎1𝜎2 − 𝑟23 2 𝑟13 𝜎1𝜎3 − 𝑟13 1 𝑟23𝜎2 𝜎3 + 𝑟13 1 𝑟23 2 𝜎3 2
𝜎3 𝜎3 𝜎3 𝜎3
𝜎2 𝜎1
[∵ 𝑏23 = 𝑟23 𝑎𝑛𝑑 𝑏13 = 𝑟13 ]
𝜎3 𝜎3
= 𝑟12𝜎1𝜎2 − 𝑟23𝑟13𝜎1𝜎2 − 𝑟13𝑟23𝜎1𝜎2 + 𝑟13𝑟23𝜎1𝜎2
= 𝜎1𝜎2 (𝑟12 − 𝑟13𝑟23 ) ……(2)
Again, 𝑉 (𝑋1.3) = 𝐸 {𝑋1.3 − 𝐸 (𝑋1.3)}2
= 𝐸(𝑋1.32) {Since, 𝐸 (𝑋1.3) = 0}
1
= ∑ 𝑋1.32
𝑁
1
= ∑ 𝑋1 𝑋1.3 {using the properties of residuals}
𝑁
1
= ∑ 𝑋1 (𝑋1 − 𝑏13 𝑋3 )
𝑁
1
= ∑ 𝑋1 2 − 𝑏13 ∑ 𝑋1 𝑋3
𝑁
1
= ∑ 𝑋1 2 − 𝑏13 𝐶𝑜𝑣 (𝑋1 , 𝑋3 )
𝑁
𝜎1
= 𝜎12 − 𝑟13 𝑟 𝜎𝜎
𝜎3 13 1 3
= 𝜎12 − 𝑟132 𝜎1 2
= 𝜎12 (1 − 𝑟13 2) ……(3)
Similarly, we obtain 𝑉 (𝑋2.3) = 𝜎2 2(1 − 𝑟232 ) ……(4)
Thus, from (1) using (2), (3) and (4) we get
𝜎1 𝜎2 (𝑟12 −𝑟13 𝑟23 ) 𝑟12 −𝑟13 𝑟23
𝑟12.3 = =
√𝜎1 (1−𝑟13 2 )√𝜎2 2 (1−𝑟23 2 )
2 √(1−𝑟13 2 )(1−𝑟23 2 )
𝑟23 −𝑟21 𝑟31
Similarly, 𝑟23.1 =
√(1−𝑟21 2 )(1−𝑟31 2 )
𝑟13 −𝑟12 𝑟32
and 𝑟13.2 =
√(1−𝑟12 2 )(1−𝑟32 2 )
Multiple Correlation Coefficient in Terms of Total and Partial Correlation Coefficient
—
1 − 𝑅1.232 = (1 − 𝑟122 )(1 − 𝑟13.22)
𝑟12 2 +𝑟13 2 −2𝑟12 𝑟23 𝑟13
Proof – We know that 𝑅1.232 = 2
1−𝑟23
𝑟12 2 +𝑟13 2 −2𝑟12 𝑟23 𝑟13
Now, 1 − 𝑅1.232 = 1 −
1−𝑟23 2
27 | D R C
1−𝑟23 2 −𝑟12 2 −𝑟13 2 +2𝑟12 𝑟23 𝑟13
= ……(1)
1−𝑟23 2
𝑟13 −𝑟12 𝑟23
Again, we have 𝑟13.2 =
√(1−𝑟12 2 )(1−𝑟23 2 )
(𝑟 −𝑟12 𝑟23 )2
𝑟13.22 = (1−𝑟13 2)(1−𝑟 2
12 23 )
2 2 2
𝑟 −𝑟12 𝑟23 −2𝑟13 𝑟12 𝑟23
Now, 1 − 𝑟13.22 = 1 − 13 (1−𝑟 2 2
12 )(1−𝑟23 )
1−𝑟23 −𝑟12 −𝑟12 𝑟23 −𝑟13 2 +𝑟12 2 𝑟23 2 +2𝑟13 𝑟12 𝑟23
2 2 2 2
= (1−𝑟12 2 )(1−𝑟23 2 )
1−𝑟23 −𝑟12 −𝑟13 2 +2𝑟13 𝑟12 𝑟23
2 2
= (1−𝑟12 2 )(1−𝑟23 2 )
1−𝑟23 2 −𝑟12 2 −𝑟13 2 +2𝑟13 𝑟12 𝑟23
Again, (1 − 𝑟122)(1 − 𝑟13.22 ) = (1 − 𝑟122) × (1−𝑟12 2 )(1−𝑟23 2 )
2 2 2
1−𝑟23 −𝑟12 −𝑟13 +2𝑟13 𝑟12 𝑟23
= ……(2)
1−𝑟23 2
Thus, from (1) and (2) we get 1 − 𝑅1.23 2 = (1 − 𝑟122)(1 − 𝑟13.22)
28 | D R C