Unit 2: Multivariate Statistical Methods
Wishart Distribution —
Let X be a random matrix of order 𝑛 × 𝑝 and let 𝑋 ′ 𝐻𝑋 where 𝐻 = [𝐼 − 𝑛−1 1𝐼] be a
quadratic form, the matrix 𝑋 ′ 𝐻𝑋 is known as Wishart matrix, the elements of which are sum
of squares and sum of products. Since, X is a matrix of order 𝑛 × 𝑝, 𝑋 ′ 𝐻𝑋 is a symmetric
𝑝(𝑝+1)
matrix of order 𝑝 × 𝑝 which contains distinct values of sum of squares and sum of
2
𝑝(𝑝+1)
products. The distribution of the Wishart matrix or the joint distribution of the distinct
2
′
elements of 𝑋 𝐻𝑋 is known as Wishart distribution and it’s density function denoted by
𝑛−𝑝−1 1
𝑊𝑝 (𝑛, 𝐼) is given by 𝑊𝑝 (𝑛, 𝐼) = 𝑘(𝑛, 𝑝)|𝐴| 2 𝑒 −2𝑡𝑟𝐴 ; |𝐴| > 0 where 𝑘(𝑛, 𝑝) =
1
𝑛𝑝 𝑝(𝑝−1) and 𝐴 = 𝑋 ′ 𝐻𝑋
𝑝 𝑛−𝑝+1
2 ⁄2 ∏𝑖=1 Γ( 2 )𝜋 4
The Wishart distribution is a generalisation of 𝜒 2 distribution. If 𝑋̰ ~ 𝑁𝑝 (𝜇̰, Σ) and 𝐴 =
∑𝑁 ̅ ̅ ′
𝛼=1(𝑋̰ 𝛼 − 𝑋̰ )(𝑋̰ 𝛼 − 𝑋̰ ) , then Wishart distribution is given by
1 −1 𝑋
𝑛−𝑝−1 − 𝑡𝑟Σ
𝑒 2
𝑊𝑝 = 𝑘(𝑛, 𝑝)|𝐴| 2 𝑛 ; |𝐴| > 0
|Σ| ⁄2
1
where 𝑘(𝑛, 𝑝) = 𝑛𝑝 𝑝(𝑝−1)
𝑝 𝑛−𝑗+1
2 2 𝜋 4 ∏𝑗=1 Γ( 2 )
NOTE : Let 𝑋̰ ~ 𝑁𝑝 (0̰, 𝐼) where 𝑋̰ is a 𝑛 × 𝑝 random matrix and let 𝐴 = 𝑋̰ ′ 𝑋̰ , then 𝐴 is
known as standard Wishart matrix and the distribution of A is known as Standard Wishart
distribution.
Derivation of Wishart Distribution:—
Let 𝑋1 , 𝑋2 , …., 𝑋𝑝 be 𝑝 vectors each of 𝑛 observations (𝑛 > 𝑝). Let us perform the Gram
Schmidt Orthogonalization on these vectors such that
𝑌𝑖 = 𝑏 𝑖1 𝑋1 + 𝑏 𝑖2 𝑋2 +. … + 𝑏 𝑖𝑖 𝑋𝑖 ; 𝑖 = 1, 2, …., 𝑝 ……(1)
𝑖1 𝑖2 𝑖𝑖
The coefficients 𝑏 , 𝑏 , …., 𝑏 are so chosen such that 𝑌𝑖 is orthogonal 𝑌1 , 𝑌2 , …., 𝑌𝑖−1
and is of unit length. The orthogonal transformation can be written as 𝑌 = 𝐵−1 𝑋 ⇒ 𝑋 = 𝐵𝑌
𝑏11 0 …. 0 𝑏11 0 …. 0
where 𝐵−1 = (𝑏
21
𝑏 22 … . 0 ) and 𝐵 = (𝑏21 𝑏22 … . 0 )
…. …. …. …. …. …. …. ….
𝑏 𝑝1 𝑏 𝑝2 … . 𝑏 𝑝𝑝 𝑏𝑝1 𝑏𝑝2 … . 𝑏𝑝𝑝
We know that if 𝑋̰ ~ 𝑁𝑝 (0̰, 𝐼), then 𝐴 = 𝑋̰ 𝑋̰ ′ is the Wishart matrix.
Therefore, 𝐴 = 𝑋̰ 𝑋̰ ′
= (𝐵𝑌)(𝐵𝑌)′
= 𝐵𝑌𝑌 ′ 𝐵′
= 𝐵𝐵′ {Since, 𝑌 is orthogonal}
where 𝐵 is called the Bartlet decomposition of 𝐴.
Now, 𝑋𝑖 = 𝑏𝑖1 𝑌1 + 𝑏𝑖2 𝑌2 +. … + 𝑏𝑖𝑖 𝑌𝑖 ……(2)
′
Pre multiplying both sides of (2) by 𝑌𝑗 , we get
𝑌𝑗 ′ 𝑋𝑖 = 𝑏𝑖1 𝑌𝑗 ′ 𝑌1 + 𝑏𝑖2 𝑌𝑗 ′ 𝑌2 +. … + 𝑏𝑖𝑖 𝑌𝑗 ′ 𝑌𝑖
0, 𝑖𝑓 𝑖 ≠ 𝑗
⇒ 𝑌𝑗 ′ 𝑋𝑖 = 𝑏𝑖𝑗 {Since, 𝑌 is orthogonal and 𝑌𝑗 ′ 𝑌𝑖 = { }
1, 𝑖𝑓 𝑖 = 𝑗
Again, pre multiplying both sides of (2) by 𝑋𝑖 ′ , we get
1|DRC
𝑋𝑖 ′ 𝑋𝑖 = 𝑋𝑖 ′ (𝑏𝑖1 𝑌1 + 𝑏𝑖2 𝑌2 +. … + 𝑏𝑖𝑖 𝑌𝑖 )
⇒ 𝑋𝑖 ′ 𝑋𝑖 = (𝑏𝑖1 𝑌1 + 𝑏𝑖2 𝑌2 +. … + 𝑏𝑖𝑖 𝑌𝑖 )′ (𝑏𝑖1 𝑌1 + 𝑏𝑖2 𝑌2 +. … + 𝑏𝑖𝑖 𝑌𝑖 )
⇒ 𝑋𝑖 ′ 𝑋𝑖 = 𝑏𝑖1 2 + 𝑏𝑖2 2 +. … + 𝑏𝑖𝑖 2 {Since, 𝑌 is orthogonal}
′
⇒ 𝑋𝑖 𝑋𝑖 = 𝑎𝑖𝑖 {Since, 𝐴 = 𝑋𝑋 ′ }
Therefore, 𝑏𝑖𝑖 2 = 𝑋𝑖 ′ 𝑋𝑖 − ∑𝑖−1 𝑗=1 𝑏𝑖𝑗
2
𝑖 = 1, 2, … . , 𝑝
If 𝑋̰ ~ 𝑁𝑝 (0̰, 𝐼), then each of 𝑋𝑖𝑙 ( ) is distributed according to 𝑁1 (0, 1). By
𝑙 = 1(1)𝑛
transformation of 𝑋𝑖 to 𝑏𝑖1 , 𝑏𝑖2 , …., 𝑏𝑖,𝑖−1 we can write
𝑏𝑖1 𝑌1 ′
𝑏𝑖2 ′
( )=( 𝑌2 ) 𝑋𝑖
…. ….
𝑏𝑖,𝑖−1 𝑌𝑖−1 ′
This transformation is incomplete since only (𝑖 − 1) new variables are obtained from this
transformation.
But 𝑏𝑖1 , 𝑏𝑖2 , …., 𝑏𝑖,𝑖−1 are distributed as 𝑁(0,1) and hence 𝑏𝑖𝑖 2 = 𝑋𝑖 ′ 𝑋𝑖 − ∑𝑖−1 2
𝑗=1 𝑏𝑖𝑗 ~ 𝜒
2
with 𝑛 − (𝑖 − 1) df. Thus, the distribution of 𝐵 is a mixture of normal variates and 𝜒 2
variates and is given by
1 2
1
∏𝑝𝑖=1 ∏𝑖−1
𝑗=1 (
𝑝
𝑒 −2𝑏𝑖𝑗 𝑑𝑏𝑖𝑗 ) ∏𝑘=1(𝜒𝑛−𝑘+1 2 𝑏𝑘𝑘 2 𝑑𝑏𝑘𝑘 2 ); −∞ < 𝑏𝑖𝑗 < ∞ and 0 < 𝑏𝑘𝑘 2 < ∞
√2𝜋
……(3)
′
Since, 𝐴 = 𝐵𝐵
𝑝
⇒ |𝐴| = |𝐵𝐵′ | = ∏𝑖=1 𝑏𝑖𝑖 2
𝑝 2
and 𝑡𝑟(𝐴) = 𝑡𝑟(𝐵𝐵′ ) = ∑𝑖=1 ∑𝑖−1𝑗=1 𝑏𝑖𝑗
1 𝑝
The Jacobian of transformation is 𝐽(𝐵 → 𝐴) = = 2𝑝 ∏𝑖=1 𝑏𝑖𝑖 𝑖−𝑝−1
𝐽(𝐴→𝐵)
The second part of (3) can be written as
1 𝑛−𝑖−1
1 2
∏𝑝𝑖=1 { 𝑛−𝑖+1 𝑒 −2𝑏𝑖𝑖 (𝑏𝑖𝑖 2 ) 2
𝑑𝑏𝑖𝑖 2 }
𝑛−𝑖+1
2 2 Γ( )
2
𝑛−𝑝−1 1
The pdf of 𝐴 becomes 𝑊(𝑛, 𝑝) = 𝑘(𝑛, 𝑝)|𝐴| 2 𝑒 −2𝑡𝑟𝐴 𝑑𝐴; |𝐴| > 0 where 𝐴 = 𝑋𝑋 ′ and
1
𝑘(𝑛, 𝑝) = 𝑛𝑝 𝑝(𝑝−1)
𝑝 𝑛−𝑖+1
22𝜋 4 ∏𝑖=1 Γ( )
2
Here, 𝑛 and 𝑝 are the parameters of the distribution and 𝑛 is the df of Wishart distribution.
Distribution of Wishart Matrix for 𝑿 ~ 𝑵𝒑 (𝝁̰, 𝚺) —
Let us draw a random sample of size 𝑛 from 𝑁𝑝 (𝜇̰, Σ). The Wishart matrix is defined by
A = (𝑋̰ − 𝜇𝐼𝑛 )(𝑋̰ − 𝜇𝐼𝑛 )′ ……(1)
which is a matrix of the sum of the squares and sum of products measured from the
population mean.
Let us consider the lower triangular transformation 𝑌 = 𝐶 −1 (𝑋̰ − 𝜇𝐼𝑛 ) ……(2)
′
such that 𝐶𝐶 = Σ where 𝐶 is a lower triangular matrix. Then by this transformation, we get
𝑌 ~ 𝑁𝑝 (0,1)
𝑛−𝑝−1 1
Hence, 𝐷 = 𝑌𝑌 ′ ~ 𝑊𝑝 (𝑛, 𝐼) and 𝑓(𝐷) = 𝑘(𝑛, 𝑝)|𝐷| 2 𝑒 −2𝑡𝑟(𝐷) ……(3)
2|DRC
Here, 𝐷 = 𝑌𝑌 ′
⇒ 𝐷 = {𝐶 −1 (𝑋̰ − 𝜇𝐼𝑛 )}{𝐶 −1 (𝑋̰ − 𝜇𝐼𝑛 )}′ {using (2)}
⇒ 𝐷 = 𝐶 −1 (𝑋̰ − 𝜇𝐼𝑛 )(𝑋̰ − 𝜇𝐼𝑛 )′ (𝐶 −1 )′
⇒ 𝐷 = 𝐶 −1 𝐴(𝐶 −1 )′ {using (1)}
Also, |𝐷| = |𝐶 −1 𝐴(𝐶 −1 )′ |
= |𝐴||(𝐶𝐶 ′ )−1 |
= |𝐴||𝐶𝐶 ′ |−1
= |𝐴||Σ|−1 {Since, 𝐶𝐶 ′ = Σ}
|𝐴|
= |Σ| ……(4)
and 𝑡𝑟(𝐷) = 𝑡𝑟(𝐶 −1 𝐴(𝐶 −1 )′ )
= 𝑡𝑟{(𝐶𝐶 ′ )−1 𝐴}
= 𝑡𝑟Σ −1 𝐴 ……(5)
The Jacobian of transformation is 𝐽(𝐷 → 𝐴) = |𝐶 −1 |𝑝+1
= |𝐶|−(𝑝+1)
Now, 𝐶𝐶 ′ = Σ
⇒ |𝐶𝐶 ′ | = |Σ|
⇒ |𝐶|2 = |Σ|
1
⇒ |𝐶| = |Σ|2
1
(𝑝+1)
Therefore, 𝐽(𝐷 → 𝐴) = |𝐶|−2 ……(6)
Using (4), (5) and (6) in equation (3), we get the pdf of A as
𝑛−𝑝−1
|𝐴| 1 −1 𝐴
𝑒 −2𝑡𝑟Σ
2
𝑓(𝐴) = 𝑘(𝑛, 𝑝) (|Σ| )
1
where 𝑘(𝑛, 𝑝) = 𝑛𝑝 𝑝(𝑝−1)
𝑝 𝑛−𝑖+1
2 2 𝜋 4 ∏𝑖=1 Γ( 2 )
Properties of Wishart Distribution —
1. If A ~ 𝑊𝑝 (𝐴|𝑛, Σ) and B is 𝑞 × 𝑝 matrix, then 𝐵𝐴𝐵′ ~ 𝑊𝑝 (𝐵𝐴𝐵′ |𝑛, 𝐵Σ𝐵′ )
Proof: Let 𝑋̰ ~ 𝑁𝑝 (0̰, Σ)
Therefore, A = 𝑋𝑋 ′ ~ 𝑊𝑝 (𝐴|𝑛, Σ)
Now, 𝐵𝐴𝐵′ = 𝐵𝑋𝑋 ′ 𝐵′
= (𝐵𝑋)(𝐵𝑋)′
= 𝑌𝑌 ′ (say), where 𝑌 = 𝐵𝑋
Now, 𝐸(𝑌) = 𝐸(𝐵𝑋)
= 𝐵𝐸(𝑋)
=0 {Since, 𝑋 ~ 𝑁𝑝 (0, Σ). Thus, 𝐸(𝑋) = 0}
′]
Again, 𝑉𝑎𝑟(𝑌) = 𝐸[{𝑌 − 𝐸(𝑌)}{𝑌 − 𝐸(𝑌)}
= 𝐸[𝑌𝑌 ′ ] {Since, 𝐸(𝑌) = 0}
′]
= 𝐸[(𝐵𝑋)(𝐵𝑋)
= 𝐸[𝐵𝑋𝑋 ′ 𝐵′ ]
= 𝐵𝐸[𝑋𝑋 ′ ]𝐵′
= 𝐵Σ𝐵′
Therefore, 𝑌 ~ 𝑁𝑝 (0, 𝐵Σ𝐵′ )
3|DRC
Hence, 𝑌𝑌 ′ ~ 𝑊𝑝 (𝑌𝑌 ′ |𝑛, 𝐵Σ𝐵′ )
⇒ 𝐵𝐴𝐵′ ~ 𝑊𝑝 (𝐵𝐴𝐵′ |𝑛, 𝐵Σ𝐵′ )
Hence, proved.
𝑎̰ ′ 𝐴𝑎̰
2. If A ~ 𝑊𝑝 (𝑛, Σ) and 𝑎̰ is a 𝑝 × 1 vector then 𝑈 = is distributed as 𝜒 2 variate with
𝑎̰ ′ Σ𝑎̰
𝑛 df and is independent of A.
|𝐴|
3. If A ~ 𝑊𝑝 (𝑛, Σ), then |Σ| is distributed as the product of 𝑝 independent 𝜒 2 variates with
𝑛, 𝑛 − 1, …., 𝑛 − 𝑝 + 1 df respectively.
𝜎𝑝𝑝
4. If A ~ 𝑊𝑝 (𝑛, Σ), then is distributed as a 𝜒 2 variate with 𝑛 − 𝑝 + 1 df where 𝜎𝑝𝑝 and
𝑎𝑝𝑝
𝑎𝑝𝑝 are respectively the last terms of Σ −1 and 𝐴−1
Σ−1
5. If A ~ 𝑊𝑝 (𝑛, Σ), then 𝐸(𝐴) = 𝑛Σ and 𝐸(𝐴−1 ) =
𝑛−𝑝−1
𝑎̰ ′ Σ−1 𝑎̰ 2
6. If A ~ 𝑊𝑝 (𝑛, Σ) and 𝑎̰ is a 𝑝 × 1 vector then 𝑈 ′ = ′ −1 ~ 𝜒(𝑛−𝑝+1)
𝑎̰ 𝐴 𝑎̰
7. Characteristic function of Wishart distribution
Proof: Let 𝑋̰ be a data matrix form 𝑁𝑝 (0, Σ) where the size of the sample is 𝑛 = 𝑁 − 1
Let us define Wishart matrix A = 𝑋𝑋 ′ (say)
𝑝(𝑝+1)
We are to find the characteristic function of A is of elements viz., 𝑎11 , 𝑎12 , ….,
2
𝑎1𝑃 , 𝑎22 , 𝑎23 , …., 𝑎2𝑃 , …., 𝑎𝑃𝑃
Therefore, 𝜙 = 𝜙𝑋̰ (𝑡11 , 𝑡12 , …., 𝑡1𝑃 , 𝑡22 , …., 𝑡2𝑃 , …., 𝑡𝑃𝑃 )
= 𝐸[𝑒 𝑖𝑡11𝑎11+𝑖𝑡12𝑎12+.…+𝑖𝑡1𝑃𝑎1𝑃+𝑖𝑡22𝑎22+.…+𝑖𝑡2𝑃𝑎2𝑃 +.…+𝑖𝑡𝑃𝑃𝑎𝑃𝑃 ]
1
= 𝐸 [exp { ∑𝑃𝑗=1 ∑𝑅𝑘=1 𝑖(1 + 𝛿𝑗𝑘 )𝑡𝑗𝑘 𝑎𝑗𝑘 }]
2
1, 𝑖𝑓 𝑖 = 𝑗
where 𝛿𝑗𝑘 is Kronecker’s delta defined by 𝛿𝑗𝑘 = {
0, 𝑖𝑓 𝑖 ≠ 𝑗
1
Also, 𝜙 = 𝐸 [exp { (𝑡𝑟 𝑖𝐴𝑀)}] where 𝐴 = (𝑎𝑗𝑘 ) and 𝑀 = (1 + 𝛿𝑗𝑘 )𝑡𝑗𝑘 = 𝑀𝑗𝑘
2
∞ ∞ 1 1 1
⇒𝜙= ∫−∞ … . ∫−∞ 𝑛𝑝 𝑛 exp {− 𝑡𝑟(Σ −1 𝑋𝑋 ′ ) + 𝑡𝑟(𝑖𝐴𝑀)} 𝑑𝑥̰
(2𝜋) 2 |Σ| 2 2 2
1 ∞ ∞ 1
= 𝑛𝑝 𝑛∫
−∞
… . ∫−∞
exp {− 𝑡𝑟𝐴(Σ −1 − 𝑖𝑀)} 𝑑𝑥̰
(2𝜋) 2 |Σ| 2 2
1 ∞ ∞ 1
= 𝑛𝑝 𝑛 ∫−∞ … . ∫−∞ exp {− 2 𝑡𝑟𝐴((Σ−1 − 𝑖𝑀)−1 )−1 } 𝑑𝑥̰
(2𝜋) 2 |Σ| 2
𝑛𝑝 𝑛
1
= 𝑛𝑝 𝑛 (2𝜋) 2 |Σ
−1
− 𝑖𝑀|−2
(2𝜋) 2 |Σ| 2
𝑛
−
|Σ−1 −𝑖𝑀| 2
= 𝑛
|Σ| 2
𝑛
= |𝐼 − 𝑖Σ𝑀|−2
which is the characteristic function of Wishart distribution.
8. The sum of independent Wishart variate is again a Wishart variate.
If 𝐴𝑖 ~ 𝑊𝑝 (𝑛𝑖 , Σ); 𝑖 = 1, 2, …., 𝑘 (which are independent) then the sum A = ∑𝑘𝑖=1 𝐴𝑖 ~
𝑊𝑝 (𝑛 = ∑𝑘𝑖=1 𝑛𝑖 , Σ)
4|DRC
Proof: Since, 𝐴𝑖 ~ 𝑊𝑝 (𝑛𝑖 , Σ) therefore, the characteristic function is
𝑛𝑖
𝜙𝐴 (𝑡) = |𝐼 − 𝑖Σ𝑀|− 2 ; 𝑖 = 1, 2, …., 𝑘
Now, the characteristic function of A = ∑𝑘𝑖=1 𝐴𝑖
Therefore, 𝜙𝐴 (𝑡) = 𝜙∑𝑘 𝐴𝑖 (𝑡) = 𝜙𝐴1 +𝐴2 +.…+𝐴𝑘 (𝑡)
𝑖=1
= ∏𝑘𝑖=1 𝜙𝐴𝑖 (𝑡) {Since, 𝐴𝑖 ’s are independent}
𝑛
− 𝑖
= ∏𝑘𝑖=1|𝐼 − 𝑖Σ𝑀| 2
𝑛1 𝑛2 𝑛𝑘
= |𝐼 − 𝑖Σ𝑀|− 2 |𝐼 − 𝑖Σ𝑀|− 2 … . |𝐼 − 𝑖Σ𝑀|− 2
𝑘 𝑛𝑖
∑
= |𝐼 − 𝑖Σ𝑀|− 𝑖=1 2
𝑛
= |𝐼 − 𝑖Σ𝑀|−2 where 𝑛 = ∑𝑘𝑖=1 𝑛𝑖
which is the characteristic function of Wishart distribution with 𝑛 = ∑𝑘𝑖=1 𝑛𝑖 df
Therefore, A ~ 𝑊𝑝 (𝑛, Σ)
9. If A ~ 𝑊𝑝 (𝑛, 𝐼) and B is a 𝑝 × 𝑞 matrix such that 𝐵′ 𝐵 = 𝐼𝑞 then B ~ 𝑊𝑝 (𝑛, 𝐼)
1 1
10. If A ~ 𝑊𝑝 (𝑛, Σ) and B = Σ −2 𝐴Σ 2 then B ~ 𝑊𝑝 (𝑛, 𝐼)
11. If A is a Wishart matrix, then the principal diagonal sub matrices of A are also Wishart.
12. If A ~ 𝑊𝑝 (𝑛, Σ), 𝜙 = 𝐶 −1 Σ(𝐶 −1 )′ and B = 𝐶𝐴𝐶 ′ then B ~ 𝑊𝑝 (𝑛, 𝜙)
Hotelling 𝑻𝟐 —
In the univariate case, if X ~ 𝑁(𝜇, 𝜎 2 ) where 𝜎 2 is unknown and if we are to test the
𝑥̅ −𝜇
hypothesis 𝐻0 : 𝜇 = 𝜇0 against 𝐻1 : 𝜇 ≠ 𝜇0 then our test statistic will be 𝑡 = 𝑠 which
⁄ 𝑛
√
follows student’s 𝑡 distribution with 𝑛 − 1 df.
Here, 𝑥̅ is the sample mean 𝑠 2 is the sampling variance and 𝑛 is the number of observation
in the sample.
2
2 𝑥̅ −𝜇 (𝑥̅ −𝜇)2
Now, 𝑡 = (𝑠 ) = = 𝑛(𝑥̅ − 𝜇)2 𝑠 −2 = 𝑛(𝑥̅ − 𝜇)(𝑠 2 )−1 (𝑥̅ − 𝜇)
⁄ 𝑛 𝑠 2 ⁄𝑛
√
In multivariate case, if 𝑋̰ 𝑝×1 ~ 𝑁𝑝 (𝜇̰, Σ), Σ being unknown then the multivariate analogous
′
of 𝑡 2 is given by 𝑇 2 = 𝑁(𝑋̰̅ − 𝜇̰) 𝑆 −1 (𝑋̰̅ − 𝜇̰)
1 1
where 𝑋̰̅ = ∑𝑁 𝛼=1 𝑋̰ 𝛼 and 𝑆 = (𝑋̰ 𝛼 − 𝑋̰̅ )(𝑋̰ 𝛼 − 𝑋̰̅ )′
𝑁 𝑁−1
are the sample mean and variance covariance matrix respectively.
Also, 𝐸(𝑋̰̅ ) = 𝜇 and 𝐸(𝑆) = Σ
This is generalised 𝑇 2 statistic.
The distribution of this statistic was derived by Harold Hotelling and therefore, it is also
known as Hotelling 𝑇 2 statistics.
Derivation of distribution of Hotelling 𝑻𝟐 :—
Let 𝑆 ~ 𝑁𝑝 (𝑘, Σ) and 𝑑̰ ~ 𝑁𝑝 (𝛿, 𝐶 −1 Σ)
Suppose 𝑆 and 𝑑 are independent. The Hotelling 𝑇 2 statistic is given by
𝑇 2 = 𝐶𝑘𝑑̰ ′ 𝑆 −1 𝑑̰
𝑑̰ ′ 𝑆 −1 𝑑̰
=𝑘 𝐶𝑑̰ ′ Σ−1 𝑑̰ ……(1)
𝑑̰ ′ Σ−1 𝑑̰
5|DRC
𝑑̰ ′ Σ−1 𝑑̰ 2
For any given 𝑑̰ , we have ~ 𝜒(𝑘−𝑝+1) {Wishart property (6)} ……(2)
𝑑̰ ′ 𝑆 −1 𝑑̰
𝑑̰ ′ Σ−1 𝑑̰
The distribution (2) does not involve 𝑑̰ . Thus, the ratio and 𝑑̰ ′ Σ −1 𝑑̰ are distributed
𝑑̰ ′ 𝑆 −1 𝑑̰
independently of 𝑑̰ .
Since, 𝑑̰ ~ 𝑁𝑝 (𝛿, 𝐶 −1 Σ)
2
Therefore, 𝐶𝑑̰ ′ Σ−1 𝑑̰ ~ 𝜒(𝑝,𝐶𝜏 2) ……(3)
where 𝜏 2 = 𝛿̰ ′ Σ −1 𝛿̰
𝑇2
The distribution (2) and (3) are independently distributed. Thus, is the ratio of a non-
𝑘
central 𝜒 variate with 𝑝 df and non-centrality parameter 𝐶𝜏 to an independent central 𝜒 2
2 2
variate with 𝑘 − 𝑝 + 1 df.
𝑘−𝑝+1 𝑇2
Therefore, × = 𝐹 ′ (𝑝, 𝑘 − 𝑝 + 1, 𝐶𝜏 2 ) which follows singly non-central F
𝑝 𝑘
distribution with (𝑝, 𝑘 − 𝑝 + 1) df and non-central parameter 𝐶𝜏 2
If 𝛿̰ = 0̰ , then both the numerator and denominator in equation (1) are central 𝜒 2 variate and
𝑘−𝑝+1 𝑇 2
hence . ~ 𝐹(𝑝, 𝑘 − 𝑝 + 1)
𝑝 𝑘
Application of Hotelling 𝑻𝟐 —
1. One Sample Case:—
The Hotelling 𝑇 2 statistic can be used in one sample problem to test the hypothesis that
the mean vector 𝜇̰ = 𝜇̰0 when the dispersion matrix Σ is unknown. Let 𝑋̰ is distributed
according to 𝑁𝑝 (𝜇̰, Σ). For testing 𝐻0 : 𝜇̰ = 𝜇̰0 when Σ is unknown, the test statistic is
𝑇 2 𝑁−𝑝 ′
given by . ~ 𝐹(𝑝,𝑁−𝑝) where 𝑇 2 = 𝑁(𝑋̰̅ − 𝜇̰0 ) 𝑆 −1 (𝑋̰̅ − 𝜇̰0 )
𝑁−1 𝑝
𝜇10 𝑋̅1
𝜇20 1 𝑋̅ 1
Here, 𝜇̰0 = [ … . ], 𝑋̰̅ = ∑𝑁
𝛼=1 𝑋̰ 𝛼 = 2 and 𝑆 = ∑𝑁𝛼=1(𝑋̰ 𝛼 − 𝑋̰̅ )(𝑋̰ 𝛼 − 𝑋̰̅ )′
𝑁 …. 𝑁−1
𝜇𝑝0 ̅
[𝑋𝑝 ]
The calculated value of 𝑇 is compared with the tabulated value of 𝑇 2 =
2
(𝑁−1)𝑝
𝐹𝛼(𝑝,𝑁−𝑝) at 𝛼 level of significance for (𝑝, 𝑁 − 𝑝) df and conclusion is drawn in
𝑁−𝑝
the usual manner.
2. Two Sample Case:—
The Hotelling 𝑇 2 statistic can also be used for testing the null hypothesis about the
equality of two population means where the covariance matrices are assumed to be equal
but unknown.
Let 𝑋̰ 𝛼 (1) (𝛼 = 1, 2, …., 𝑁1 ) be a random sample from 𝑁𝑝 (𝜇̰1 , Σ) and 𝑋̰ 𝛼 (2) (𝛼 = 1, 2,
1 𝑁1
…., 𝑁2 ) be a random sample from 𝑁𝑝 (𝜇̰2 , Σ) then 𝑋̰̅ (1) = ∑𝛼=1 𝑋̰ 𝛼 (1) is distributed as
𝑁1
Σ 1 Σ
𝑁𝑝 (𝜇̰1 , ) and 𝑋̰ ̅ (2) = ∑𝑁2 𝑋̰ (2) is distributed as 𝑁𝑝 (𝜇̰2 , )
𝑁1 𝑁2 𝛼=1 𝛼 𝑁2
𝑁1 𝑁2
Thus, √
𝑁1 +𝑁2
(𝑋̰̅ (1) − 𝑋̰̅ (2) ) ~ 𝑁𝑝 (0, Σ) under the null hypothesis.
6|DRC
1 𝑁1 (1) ′ 𝑁2
Let 𝑆 = [∑𝛼=1(𝑋̰ 𝛼 − 𝑋̰̅ (1) )(𝑋̰ 𝛼 (1) − 𝑋̰̅ (1) ) + ∑𝛼=1(𝑋̰ 𝛼
(2)
− 𝑋̰̅ (2) )(𝑋̰ 𝛼 (2) −
𝑁1 +𝑁2 −2
̅ (2) ′
𝑋̰ )]
𝑁1 +𝑁2 −2
Thus, (𝑁1 + 𝑁2 − 2)𝑆 is distributed as ∑𝛼=1 𝑍𝛼 𝑍𝛼 ′ where 𝑍𝛼 ~ 𝑁(0, Σ)
𝑁 𝑁 ′
Therefore, 1 2 (𝑋̰̅ (1) − 𝑋̰̅ (2) ) 𝑆 −1 (𝑋̰̅ (1) − 𝑋̰̅ (2) ) = 𝑇 2 is distributed as Hotelling 𝑇 2
𝑁1 +𝑁2
with 𝑁1 + 𝑁2 − 2 df and the required test statistic for testing 𝐻0 : 𝜇̰ (1) = 𝜇̰ (2) is
𝑇2 𝑁1 +𝑁2 −𝑝−1
. ~ 𝐹(𝑝,𝑁1 +𝑁2−𝑝−1)
𝑁1 +𝑁2 −2 𝑝
The calculated value of 𝑇 2 is compared with the tabulated value of 𝑇 2 =
𝑝(𝑁1 +𝑁2 −2)
𝐹𝛼(𝑝,𝑁1 +𝑁2−𝑝−1) and the conclusion is drawn in the usual manner.
𝑁1 +𝑁2 −𝑝−1
Mahalonobis’ 𝑫𝟐 —
The Mahalonobis distance between 2 mean vectors 𝜇̰ (1) and 𝜇̰ (2) is defined as
′
𝐷2 = (𝜇̰ (1) − 𝜇̰ (2) ) Σ −1 (𝜇̰ (1) − 𝜇̰ (2) ) ……(1)
Let us suppose that there is only 1 mean vector of 𝑝 elements then the Mahalonobis distance
of 𝜇̰ from the origin 0̰ is given by
𝐷2 = 𝜇̰′ Σ −1 𝜇̰ ……(2)
𝜇̰ (1)
Let us partition the mean vector 𝜇̰ as 𝜇̰ = [ (2) ] where 𝜇̰ (1) has 𝑘 elements and 𝜇̰ (2) has 𝑝 −
𝜇̰
𝑘 elements.
Σ Σ12
Similarly, let us partition Σ as Σ = [ 11 ] where Σ11 is of order 𝑘 × 𝑘, Σ12 is of order
Σ21 Σ22
𝑘 × (𝑝 − 𝑘), Σ21 is of order (𝑝 − 𝑘) × 𝑘 and Σ22 is of order (𝑝 − 𝑘) × (𝑝 − 𝑘)
Then from the partition of 𝜇̰ and Σ, we have
′
𝐷𝑝 2 = 𝜇̰ (1) Σ11 −1 𝜇̰ (1) + 𝜇̰2.1 ′ Σ22.1 −1 𝜇̰2.1
= 𝐷𝑘 2 + 𝜇̰2.1 ′ Σ22.1 −1 𝜇̰2.1 ……(3)
where 𝜇̰2.1 = 𝜇̰ (2) − Σ12 Σ11 −1 𝜇̰ (1)
Σ22.1 = Σ22 − Σ21 Σ11 −1 Σ12
𝐷𝑘 2 is the Mahalonobis distance of 𝜇̰ (1) from the origin 0̰ i.e., the Mahalonobis
distance for the first 𝑘 elements/ variables
If 𝜇̰2.1 = 0 then 𝐷2.1 2 = 𝜇̰2.1 ′ Σ22.1 −1 𝜇̰2.1 = 0 then 𝐷𝑝 2 = 𝐷𝑘 2
The decomposition of Mahalonobis distance given in equation (3) is called the population
mean vector 𝜇̰. For obtaining the sample Mahalonobis distance. Let us assume that
𝑑̰ ~ 𝑁𝑝 (𝜇̰, Σ) and A ~ 𝑊𝑝 (𝑛, Σ)
then the sample Mahalonobis distance is given by 𝐷𝑝 2 = 𝑛𝑑̰ ′ 𝐴−1 𝑑̰
𝑑̰ (1)
Now, let us partition 𝑑̰ = [ (2) ] where 𝑑̰ (1) contains 𝑘 elements 𝑑̰ (2) contains 𝑝 − 𝑘
𝑑̰
elements.
7|DRC
𝐴 𝐴12
Similarly, let us partition A = [ 11 ] where 𝐴11 is of order 𝑘 × 𝑘, 𝐴12 is of order
𝐴21 𝐴22
𝑘 × (𝑝 − 𝑘), 𝐴21 is of order (𝑝 − 𝑘) × 𝑘, and 𝐴22 is of order (𝑝 − 𝑘) × (𝑝 − 𝑘)
Then from the partition of 𝑑̰ and A, we have
𝐷𝑝 2 = 𝐷𝑘 2 + 𝑛𝑑̰ 21 ′ 𝐴22.1 −1 𝑑̰ 2.1 ……(4)
′
where 𝐷𝑘 2 = 𝑛𝑑̰ (1) 𝐴11 −1 𝑑̰ (1)
𝑑̰ 2.1 = 𝑑̰ (2) − 𝐴12 𝐴11 −1 𝑑̰ (1)
𝐴22.1 −1 = 𝐴22 − 𝐴12 𝐴11 −1 𝐴21
From the sample Mahalonobis distance given by (4), the Hotelling 𝑇 2 can be written as
𝐷𝑝 2 −𝐷𝑘 2 1
= 𝑇2 ; if 𝑑̰ 2.1 = 0
𝑛+𝐷𝑘 2
𝑛−𝑘 (𝑝−𝑘,𝑛−𝑘)
𝐷𝑝 2 −𝐷𝑘 2 𝑝−𝑘
Thus, ~ 𝐹
𝑛+𝐷𝑘 2 𝑛−𝑝+1 (𝑝−𝑘,𝑛−𝑝+1)
𝑑̰ 2.1 = 𝑑̰ (2) − 𝐴12 𝐴11 −1 𝑑̰ (1)
𝐴22.1 −1 = 𝐴22 − 𝐴12 𝐴11 −1 𝐴21
Example 1 – If 𝑋1 and 𝑋2 are independent data matrices of order 𝑛1 × 𝑝 and 𝑛2 × 𝑝
respectively and if the 𝑛𝑖 rows of 𝑋𝑖 ; 𝑖 = 1, 2 are i.i.d. as 𝑁𝑝 (𝜇̰𝑖 , Σ𝑖 ) then for 𝜇̰1 = 𝜇̰2 and Σ1
𝑛 𝑛 2
= Σ2 . Show that 1 2 𝐷2 ~ 𝑇(𝑝,𝑛 1 +𝑛2 −2)
𝑛1 +𝑛2
Solution – Since, 𝑋̰ 𝑖 ~ 𝑁𝑝 (𝜇̰𝑖 , Σ𝑖 )
1
Therefore, 𝑋̰̅𝑖 ~ 𝑁𝑝 (𝜇̰𝑖 , Σ𝑖 )
𝑛
Let us consider the variable 𝑑 = 𝑋̰̅1 − 𝑋̰̅2
Then 𝐸(𝑑) = 𝜇̰1 − 𝜇̰2
1 1
and 𝑉(𝑑) = Σ1 + Σ2
𝑛1 𝑛2
Therefore, 𝑑 is multivariate normal.
𝑛 +𝑛
If 𝜇̰1 = 𝜇̰2 and Σ1 = Σ2 = Σ (say) then 𝑑 ~ 𝑁 (0, 1 2 Σ)
𝑛 1 𝑛2
𝑛1 +𝑛2
Let = 𝐶 and 𝑀𝑖 = 𝑁𝑖 𝑆𝑖 (𝑖 = 1, 2) where 𝑆𝑖 is the MLE of Σ𝑖
𝑛1 𝑛2
Then 𝑀𝑖 ~ 𝑊𝑝 (𝑛𝑖 − 1, Σ𝑖 )
If Σ1 = Σ2 = Σ, then the pooled estimate of Σ is
𝑛 𝑆 +𝑛 𝑆 𝑀 +𝑀 𝑀
𝑆𝑝 = 1 1 2 2 = 1 2 = {Since, 𝑀 = 𝑀1 + 𝑀2 }
𝑛1 +𝑛2 −2 𝑛1 +𝑛2 −2 𝑛1 +𝑛2 −2
Therefore, 𝑀 = (𝑛1 + 𝑛2 − 2)𝑆𝑝
Thus, 𝐶𝑀 ~ 𝑊𝑝 (𝑛1 + 𝑛2 − 2, 𝐶𝑀)
Moreover, 𝑀 is independent of 𝑑.
Since, 𝑋̅𝑖 is independent of 𝑆𝑖 . Also, the 2 samples are independent. Thus, from the
definition of Hotelling 𝑇 2 statistic, we can write
2
(𝑛1 + 𝑛2 − 2)𝑑′ (𝐶𝑀)−1 𝑑 ~ 𝑇(𝑝,𝑛 1 +𝑛2 −2)
𝑛1 𝑛2
⇒ (𝑛1 + 𝑛2 − 2)(𝑋̰̅1 − 𝑋̰̅2 ) ′ (𝑛1 + 𝑛2 − 2)−1 𝑆𝑝 −1 (𝑋̰̅1 − 𝑋̰̅2 ) ~ 𝑇(𝑝,𝑛
2
1 +𝑛2 −2)
𝑛1 +𝑛2
𝑛1 𝑛 2 −1 2
⇒ (𝑋̰̅1 − 𝑋̰̅2 𝑆𝑝)′ (𝑋̰̅1 − 𝑋̰̅2 ) ~ 𝑇(𝑝,𝑛 1 +𝑛2 −2)
𝑛1 +𝑛2
8|DRC
𝑛1 𝑛 2 2
⇒ 𝐷2 ~ 𝑇(𝑝,𝑛 1 +𝑛2 −2)
𝑛1 +𝑛2
Example 2 – Prove that Hotelling’s 𝑇 2 is invariant under change in the unit of measurements
of 𝑋’s
Solution – Let 𝑋̰ ~ 𝑁𝑝 (𝜇̰, Σ) where Σ is unknown. For testing the null hypothesis 𝜇̰ = 𝜇̰0 the
′
test statistic is 𝑇 2 = 𝑁(𝑋̰̅ − 𝜇̰0 ) 𝑆 −1 (𝑋̰̅ − 𝜇̰0 )
𝑋̅1 𝜇10
1 𝑋̅ 𝜇20 1
where 𝑋̰̅ = ∑𝑁 𝛼=1 𝑋̰ 𝛼 = 2 , 𝜇̰0 = [ … . ] and 𝑆𝑝×𝑝 = ∑𝑁 ̅ ̅ ′
𝛼=1(𝑋̰ 𝛼 − 𝑋̰ )(𝑋̰ 𝛼 − 𝑋̰ )
𝑁 …. 𝑁−1
[𝑋̅𝑝 ] 𝜇𝑝0
Let us consider the transformation 𝑌̰𝑝×1 = 𝐶𝑝×𝑝 𝑋̰ 𝑝×1 + 𝐷𝑝×1 where C is non-singular. Then
the Hotelling 𝑇 2 statistic for 𝑌̰ is of the form
𝑇𝑌 2 = 𝑁(𝑌̰̅ − 𝜈̰ 0 )′ 𝑆𝑌 −1 (𝑌̰̅ − 𝜈̰ 0 ) where 𝜈̰ 0 = 𝐶𝜇̰0 + 𝐷
Let 𝑋̰ 𝛼 (𝛼 = 1, 2, …., 𝑁) be a random sample from 𝑁(𝜇̰, Σ)
Then 𝑌̰ = 𝐶𝑋̰ + 𝐷
𝑌1𝛼 𝑋1𝛼
𝑌 𝑋
⇒ [ …2𝛼. ] = 𝐶 [ …2𝛼. ] + 𝐷
𝑌𝑝𝛼 𝑋𝑝𝛼
1 𝑁 1 𝑁
∑𝛼=1 𝑌1𝛼 ∑𝛼=1 𝑋1𝛼
𝑁 𝑁
1 1
⇒ ∑𝑁 𝑁
𝛼=1 𝑌2𝛼 = 𝐶 𝑁 ∑𝛼=1 𝑋2𝛼 + 𝐷
𝑁
…. ….
1 𝑁 1 𝑁
[𝑁 ∑𝛼=1 𝑌𝑝𝛼 ] [𝑁 ∑𝛼=1 𝑋𝑝𝛼 ]
⇒ 𝑌̰̅ = 𝐶𝑋̰̅ + 𝐷
Therefore, 𝑌̰̅ − 𝜈̰ 0 = 𝐶𝑋̰̅ + 𝐷 − (𝐶𝜇̰0 + 𝐷)
= 𝐶𝑋̰̅ − 𝐶𝜇̰0
= 𝐶(𝑋̰̅ − 𝜇̰0 )
1
and 𝑆𝑌 = ∑𝑁 ̅ ̅ ′
𝛼=1(𝑌̰𝛼 − 𝑌̰ )(𝑌̰𝛼 − 𝑌̰ )
𝑁−1
1
= ∑𝑁 ̅ ̅
𝛼=1(𝐶𝑋̰ 𝛼 + 𝐷 − 𝐶𝑋̰ − 𝐷)(𝐶𝑋̰ 𝛼 + 𝐷 − 𝐶𝑋̰ − 𝐷)
′
𝑁−1
1
= ∑𝑁 ̅ ̅ ′ ′
𝛼=1 𝐶(𝑋̰ 𝛼 − 𝑋̰ )(𝑋̰ 𝛼 − 𝑋̰ ) 𝐶
𝑁−1
1
=𝐶[ ∑𝑁 (𝑋̰ − 𝑋̰̅ )(𝑋̰ 𝛼 − 𝑋̰̅ )′ ] 𝐶 ′
𝑁−1 𝛼=1 𝛼
= 𝐶𝑆𝐶 ′
Therefore, 𝑇𝑌 2 = 𝑁(𝑌̰̅ − 𝜈̰ 0 )′ 𝑆𝑌 −1 (𝑌̰̅ − 𝜈̰ 0 )
′
= 𝑁{𝐶(𝑋̰̅ − 𝜇̰0 )} (𝐶𝑆𝐶 ′ )−1 {𝐶(𝑋̰̅ − 𝜇̰0 )}
′
= 𝑁(𝑋̰̅ − 𝜇̰0 ) 𝐶 ′ (𝐶 ′ )−1 𝑆 −1 𝐶 −1 𝐶(𝑋̰̅ − 𝜇̰0 )
′
= 𝑁(𝑋̰̅ − 𝜇̰0 ) 𝑆 −1 (𝑋̰̅ − 𝜇̰0 )
= 𝑇2
Thus, Hotelling 𝑇 2 is invariant under change in the unit of measurement.
9|DRC