0% found this document useful (0 votes)
16 views73 pages

Multivariate Probability Distribution

Uploaded by

nanditag822
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
16 views73 pages

Multivariate Probability Distribution

Uploaded by

nanditag822
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

MULTIVARIATE PROBABILITY DISTRIBUTION

Multivariate Analysis || Lecture Notes || Anup Kumar Giri


MULTIVARIATE PROBABILITY DISTRIBUTION

MULTIVARIATE PROBABILITY DISTRIBUTION

RANDOM VECTOR

Let us consider (Ω, ℱ, 𝑃) be a probability space of a random experiment.


𝑇
A collection or vector 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) defined on (Ω, ℱ, 𝑃) and maps Ω into ℛ𝑝 i.e.,
~

𝑇
𝑋(𝜔) = (𝑋1 (𝜔), 𝑋2 (𝜔), ⋯ , 𝑋𝑝 (𝜔)) ; 𝜔 ∈ Ω,
~

is called a p-dimensional random vector if the inverse image of each p-dimensional interval
𝑇
𝐼(𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) = {𝑥 = (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) : 𝑥 ∈ ℛ𝑝 or 𝑥𝑖 ∈ ℛ, 𝑖 = 1(1)𝑝}
~ ~

is also in ℱ, that is, if

𝑋 −1 (𝐼) = {𝜔 ∶ 𝑋1 (𝜔) ≤ 𝑥1 , 𝑋2 (𝜔) ≤ 𝑥2 , ⋯ , 𝑋𝑝 (𝜔) ≤ 𝑥𝑝 } ∈ ℱ ; for all 𝑥𝑖 ∈ ℛ, 𝑖 = 1(1)𝑝.


~

𝑇
Suppose 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 are 𝑝 one dimensional random variable. The vector 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) is a p-
~

dimensional random vector.


𝑇 𝑇
Let the vector 𝑥 = (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ∈ ℛ𝑝 be realisation of 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) .
~ ~

𝑇
The random vector 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) may be discrete, continuous or mixed.
~

DISTRIBUTION FUNCTION:
𝑇
The cumulative distribution function (c.d.f.) or d.f. of a random vector 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) is defined as
~

𝐹𝑋 (𝑥 ) = 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) = 𝑃(𝑋1 ≤ 𝑥1 , 𝑋2 ≤ 𝑥2 , ⋯ , 𝑋𝑝 ≤ 𝑥𝑝 ).
~ ~ ~

PROPERTIES OF DISTRIBUTION FUNCTION:

A function 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) is said to be the joint distribution function of a p-dimensional random


~

vector if and only if it satisfies the following conditions:

(𝑖) monotonically non-decreasing with respect to all the realisations or arguments 𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ,

(𝑖𝑖) continuous from the right with respect to all the realisations or arguments 𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ,

(𝑖𝑖𝑖) 0 ≤ 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ≤ 1,
~

(𝑖𝑣) lim 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) = 𝐹𝑋 (−∞, 𝑥2 , ⋯ , 𝑥𝑝 ) = ⋯ = 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , −∞) = 0, ∀ 𝑖 = 1(1)𝑝,


𝑥𝑖 →−∞ ~ ~ ~

or, lim 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) = 0, ∀ 1 ≤ 𝑖1 < 𝑖2 < ⋯ < 𝑖𝑘 ≤ 𝑝, 𝑘 = 1(1)𝑝,


𝑥𝑖1 ,𝑥𝑖2 ,⋯,𝑥𝑖𝑘 → −∞ ~

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 1


MULTIVARIATE PROBABILITY DISTRIBUTION

(𝑣) lim 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) = 𝐹𝑋 (∞, ∞, ⋯ , ∞) = 1,


𝑥1 ,𝑥2 ,⋯,𝑥𝑝 → ∞ ~ ~

𝑇
(𝑣𝑖) For all 𝑥 = (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ∈ ℛ𝑝 and for all ℎ𝑖 > 0 ( 𝑖 = 1(1)𝑝), the following inequality holds.
~

𝐹𝑋 (𝑥1 + ℎ1 , 𝑥2 + ℎ2 , ⋯ , 𝑥𝑝 + ℎ𝑝 ) − ∑ 𝐹𝑋 (𝑥1 + ℎ1 , ⋯ , 𝑥𝑖−1 + ℎ𝑖−1 , 𝑥𝑖 , 𝑥𝑖+1 + ℎ𝑖+1 ⋯ , 𝑥𝑝 + ℎ𝑝 )


~ ~
𝑖=1

+ ∑ 𝐹𝑋 (𝑥1 + ℎ1 , ⋯ , 𝑥𝑖−1 + ℎ𝑖−1 , 𝑥𝑖 , 𝑥𝑖+1 + ℎ𝑖+1 , ⋯ , 𝑥𝑗−1 + ℎ𝑗−1 , 𝑥𝑗 , 𝑥𝑗+1 + ℎ𝑗+1 ⋯ , 𝑥𝑝 + ℎ𝑝 )


~
1≤𝑖<𝑗≤𝑝

+ ⋯ + (−1)𝑝 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ≥ 0
~

A. DISCRETE RANDOM VECTOR:


𝑇
If 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) is discrete, the joint probability mass function (p.m.f.) of 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 is given by
~

𝑝𝑋 (𝑥 ) = 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) = 𝑃(𝑋1 = 𝑥1 , 𝑋2 = 𝑥2 , ⋯ , 𝑋𝑝 = 𝑥𝑝 )
~ ~ ~

which satisfy the conditions

(𝑖) 𝑝𝑋 (𝑥 ) ≥ 0 ∀ 𝑥 ∈ ℛ𝑝 𝑖. 𝑒., 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ≥ 0 ∀ 𝑥𝑖 ∈ ℛ, 𝑖 = 1(1)𝑝


~ ~ ~ ~

(𝑖𝑖) ∑ 𝑝𝑋 (𝑥 ) = 1 𝑖. 𝑒., ∑ ∑ ⋯ ∑ 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) = 1.


~ ~ ~
𝑥 ∈ℛ𝑝 𝑥1 ∈ℛ 𝑥2 ∈ℛ 𝑥𝑝 ∈ℛ
~

B. CONTINUOUS RANDOM VECTOR:


𝑇
A p-dimensional random vector 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) is said to be continuous random vector if there exists a
~

non-negative integrable function 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) such that


~

𝑥𝑝 𝑥𝑝−1 𝑥1

𝐹𝑋 (𝑥 ) = 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) = ∫ ∫ ⋯ ∫ 𝑓𝑋 (𝑦1 , 𝑦2 , ⋯ , 𝑦𝑝 ) 𝑑𝑦1 𝑑𝑦2 ⋯ 𝑑𝑦𝑝 ∀ 𝑥𝑖 ∈ ℛ, 𝑖 = 1(1)𝑝


~ ~ ~ ~
−∞ −∞ −∞

𝑇
The function 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) is called the probability density function p.d.f. of 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) , which
~ ~

satisfies

(𝑖) 𝑓𝑋 (𝑥 ) ≥ 0 ∀ 𝑥 ∈ ℛ𝑝 𝑖. 𝑒., 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ≥ 0 ∀ 𝑥𝑖 ∈ ℛ, 𝑖 = 1(1)𝑝


~ ~ ~ ~

∞ ∞ ∞

(𝑖𝑖) ∫ 𝑓𝑋 (𝑥 ) 𝑑𝑥 = 1 𝑖. 𝑒., ∫ ∫ ⋯ ∫ 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) 𝑑𝑥1 𝑑𝑥2 ⋯ 𝑑𝑥𝑝 = 1.


~ ~ ~ ~
𝑥 ∈ℛ𝑝 −∞ −∞ −∞
~

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 2


MULTIVARIATE PROBABILITY DISTRIBUTION

If the distribution function of a p-dimensional random vector 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) is absolutely continuous and
~

the probability density function (p.d.f.) 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) is continuous, we have


~

𝜕 𝑝 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 )
~
= 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) (almost everywhere), and
𝜕𝑥1 𝜕𝑥2 ⋯ 𝜕𝑥𝑝 ~

𝑥𝑝 𝑥𝑝−1 𝑥1

𝐹𝑋 (𝑥 ) = 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) = ∫ ∫ ⋯ ∫ 𝑓𝑋 (𝑦1 , 𝑦2 , ⋯ , 𝑦𝑝 ) 𝑑𝑦1 𝑑𝑦2 ⋯ 𝑑𝑦𝑝 ∀ 𝑥𝑖 ∈ ℛ, 𝑖 = 1(1)𝑝


~ ~ ~ ~
−∞ −∞ −∞

𝑇
The cumulative distribution function (c.d.f.) or d.f. of a random vector 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) is given by
~

𝐹𝑋 (𝑥 ) = 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) = 𝑃(𝑋1 ≤ 𝑥1 , 𝑋2 ≤ 𝑥2 , ⋯ , 𝑋𝑝 ≤ 𝑥𝑝 )
~ ~ ~

∑ ∑ ⋯ ∑ 𝑝𝑋 (𝑦1 , 𝑦2 , ⋯ , 𝑦𝑝 ) ; if 𝑋 is discrete
~ ~
𝑦1 ≤𝑥1 𝑦2 ≤𝑥2 𝑦𝑝 ≤𝑥𝑝
= 𝑥𝑝 𝑥𝑝−1 𝑥1

∫ ∫ ⋯ ∫ 𝑓𝑋 (𝑦1 , 𝑦2 , ⋯ , 𝑦𝑝 ) 𝑑𝑦1 𝑑𝑦2 ⋯ 𝑑𝑦𝑝 ; if 𝑋 is continuous.


~ ~
{−∞ −∞ −∞

where 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) and 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) are the joint probability mass function (p.m.f.) and probability
~ ~

density function of 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 , respectively.

MARGINAL DISTRIBUTION:
𝑇
A. The marginal cumulative distribution function (c.d.f.) of a random vector 𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑞 ) , (𝑞 <
~

𝑝) is given by

𝐹𝑋1 ,𝑋2 ,⋯,𝑋𝑞 (𝑥 (1) ) = 𝐹𝑋 (1) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑞 ) = 𝑃(𝑋1 ≤ 𝑥1 , 𝑋2 ≤ 𝑥2 , ⋯ , 𝑋𝑞 ≤ 𝑥𝑞 )


~ ~

= 𝑃(𝑋1 ≤ 𝑥1 , 𝑋2 ≤ 𝑥2 , ⋯ , 𝑋𝑞 ≤ 𝑥𝑞 , 𝑋𝑞+1 < ∞, ⋯ , 𝑋𝑝 < ∞)


= lim 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 )
𝑥𝑞+1 ,𝑥𝑞+2 ,⋯,𝑥𝑝 → ∞ ~

= 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑞 , ∞, ⋯ , ∞)
~

∑ ∑ ⋯ ∑ ∑ ⋯ ∑ 𝑝𝑋 (𝑦1 , 𝑦2 , ⋯ , 𝑦𝑞 , 𝑦𝑞+1 , ⋯ , 𝑦𝑝 ) ; if 𝑋 is discrete


~ ~
𝑦1 ≤𝑥1 𝑦2 ≤𝑥2 𝑦𝑞 ≤𝑥𝑞 𝑦𝑞+1 ∈ℛ 𝑦𝑝 ∈ℛ
= ∞ ∞ 𝑥𝑞 𝑥1

∫ ⋯ ∫ ∫ ⋯ ∫ 𝑓𝑋 (𝑦1 , 𝑦2 , ⋯ , 𝑦𝑞 , 𝑦𝑞+1 , ⋯ , 𝑦𝑝 ) 𝑑𝑦1 ⋯ 𝑑𝑦𝑞 𝑑𝑦𝑞+1 ⋯ 𝑑𝑦𝑝 ; if 𝑋 is continuous.


~ ~
{−∞ −∞ −∞ −∞

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 3


MULTIVARIATE PROBABILITY DISTRIBUTION

where 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) and 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) are the joint probability mass function (p.m.f.) and probability
~ ~

density function of 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 , respectively.

𝑇
In general, the marginal c.d.f. of random vector 𝑋 (𝑘) = (𝑋𝑖1 , 𝑋𝑖2 , ⋯ , 𝑋𝑖𝑘 ) , (1 ≤ 𝑖1 < 𝑖2 < ⋯ < 𝑖𝑘 ≤ 𝑝, 𝑘 =
~

1(1)𝑝) is given by

𝐹𝑋𝑖1 ,𝑋𝑖2 ,⋯,𝑋𝑖 (𝑥𝑖1 , 𝑥𝑖2 , ⋯ , 𝑥𝑖𝑘 ) = 𝐹𝑋 (𝑘) (𝑥𝑖1 , 𝑥𝑖2 , ⋯ , 𝑥𝑖𝑘 )
𝑘 ~

= 𝑃(𝑋𝑖1 ≤ 𝑥𝑖1 , 𝑋𝑖2 ≤ 𝑥𝑖2 , ⋯ , 𝑋𝑖𝑘 ≤ 𝑥𝑖𝑘 )


= 𝑃(𝑋1 < ∞, ⋯ , 𝑋𝑖1 ≤ 𝑥𝑖1 , ⋯ , 𝑋𝑖2 ≤ 𝑥𝑖2 , ⋯ , 𝑋𝑖𝑘 ≤ 𝑥𝑖𝑘 , ⋯ , 𝑋𝑝 < ∞)
= lim 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 )
𝑥𝑗 → +∞ ~
𝑗≠𝑖1 ,𝑖2 ,⋯,𝑖𝑘

= 𝐹𝑋 ( ∞, ⋯ , ∞, 𝑥𝑖1 , ∞, ⋯ , ∞, 𝑥𝑖2 , ∞, ⋯ , ∞, 𝑥𝑖𝑘 , ∞, ⋯ , ∞)


~

B. The marginal probability mass function (p.m.f.) 𝑝𝑋1 ,𝑋2 ,⋯,𝑋𝑞 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑞 ) of 𝑋 (1) =
~
𝑇
(𝑋1 , 𝑋2 , ⋯ , 𝑋𝑞 ) , (𝑞 < 𝑝) is given by

𝑝𝑋1 ,𝑋2 ,⋯,𝑋𝑞 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑞 ) = ∑ ∑ ⋯ ∑ 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑞 , 𝑦𝑞+1 , ⋯ , 𝑦𝑝 ) ; if 𝑋 is discrete, and


~ ~
𝑦𝑞+1 ∈ ℛ 𝑦𝑞+2 ∈ ℛ 𝑦𝑝 ∈ ℛ

𝑇
the marginal probability density function (p.d.f.) 𝑓𝑋1 ,𝑋2 ,⋯,𝑋𝑞 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑞 ) of 𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑞 ) , (𝑞 <
~

𝑝) is given by
∞ ∞

𝑓𝑋1 ,𝑋2 ,⋯,𝑋𝑞 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑞 ) = ∫ ⋯ ∫ 𝑓𝑋 (𝑥1 , ⋯ , 𝑥𝑞 , 𝑦𝑞+1 , ⋯ , 𝑦𝑝 ) 𝑑𝑦𝑞+1 ⋯ 𝑑𝑦𝑝 ; if 𝑋 is continuous,


~ ~
−∞ −∞

where 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) and 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) are the joint probability mass function (p.m.f.) and probability
~ ~

density function of 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 , respectively.

CONDITIONAL DISTRIBUTION:
𝑇
A. The conditional cumulative distribution function (c.d.f.) of a random vector 𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑞 ) , (𝑞 < 𝑝)
~

given that 𝑋 (2) = 𝑥 (2) or 𝑋𝑞+1 = 𝑥𝑞+1 , 𝑋𝑞+2 = 𝑥𝑞+2 , ⋯ , 𝑋𝑝 = 𝑥𝑝 is given by


~ ~

𝐹𝑋 (1)|𝑋 (2) (𝑥 (1) |𝑥 (2) ) = 𝐹(𝑋 , ⋯ , 𝑋 |𝑋 , ⋯ , 𝑋 ) (𝑥1 , ⋯ , 𝑥𝑞 |𝑥𝑞+1 , ⋯ , 𝑥𝑝 )


~ ~ ~ ~ 1 𝑞 𝑞+1 𝑝

𝑃(𝑋1 ≤ 𝑥1 , ⋯ , 𝑋𝑞 ≤ 𝑥𝑞 , 𝑋𝑞+1 ≤ 𝑥𝑞+1 , ⋯ , 𝑋𝑝 ≤ 𝑥𝑝 )


=
𝑃(𝑋𝑞+1 ≤ 𝑥𝑞+1 , 𝑋𝑞+2 ≤ 𝑥𝑞+2 , ⋯ , 𝑋𝑝 ≤ 𝑥𝑝 )
𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 )
~ ~
= =
lim 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) 𝐹𝑋 (∞, ⋯ , ∞, 𝑥𝑞+1 , 𝑥𝑞+2 , ⋯ , 𝑥𝑝 )
𝑥1 → ∞,⋯,𝑥𝑞 → ∞ ~ ~

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 4


MULTIVARIATE PROBABILITY DISTRIBUTION

∑𝑦1≤𝑥1 ⋯ ∑𝑦𝑞 ≤𝑥𝑞 ∑𝑦𝑞+1≤𝑥𝑞+1 ⋯ ∑𝑦𝑝 ≤𝑥𝑝 𝑝𝑋 (𝑦1 , 𝑦2 , ⋯ , 𝑦𝑞 , 𝑦𝑞+1 , ⋯ , 𝑦𝑝 )


~
; if 𝑋 is discrete
∑𝑦1∈ ℛ ⋯ ∑𝑦𝑞∈ ℛ ∑𝑦𝑞+1≤𝑥𝑞+1 ⋯ ∑𝑦𝑝≤𝑥𝑝 𝑝𝑋 (𝑦1 , 𝑦2 , ⋯ , 𝑦𝑞 , 𝑦𝑞+1 , ⋯ , 𝑦𝑝 ) ~
~
= 𝑥𝑝 𝑥𝑞+1 𝑥𝑞 𝑥1
∫−∞ ⋯ ∫−∞ ∫−∞ ⋯ ∫−∞ 𝑓𝑋 (𝑦1 , 𝑦2 , ⋯ , 𝑦𝑝 ) 𝑑𝑦1 ⋯ 𝑑𝑦𝑞 𝑑𝑦𝑞+1 ⋯ 𝑑𝑦𝑝
~
𝑥𝑝 𝑥𝑞+1 ∞ ∞ ; if 𝑋 is continuous.
{∫−∞ ⋯ ∫−∞ ∫−∞ ⋯ ∫−∞ 𝑓𝑋 (𝑦1 , 𝑦2 , ⋯ , 𝑦𝑝 ) 𝑑𝑦1 ⋯ 𝑑𝑦𝑞 𝑑𝑦𝑞+1 ⋯ 𝑑𝑦𝑝
~
~

where 𝑝𝑿 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) and 𝑓𝑿 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) are the joint probability mass function (p.m.f.) and probability density
function of 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 , respectively.

𝑇 𝑇
In general, the conditional c.d.f. of random vector 𝑋 (𝑘) = (𝑋𝑖1 , 𝑋𝑖2 , ⋯ , 𝑋𝑖𝑘 ) given that 𝑋 (𝑙) = (𝑋𝑗1 , 𝑋𝑗2 , ⋯ , 𝑋𝑗𝑙 ) = 𝑥 (𝑙)
~ ~ ~

or (𝑋𝑗1 = 𝑥𝑗1 , 𝑋𝑗2 = 𝑥𝑗2 , ⋯ , 𝑋𝑗𝑙 = 𝑥𝑗𝑙 ); (1 ≤ 𝑖1 < ⋯ < 𝑖𝑘 ≤ 𝑝, 1 ≤ 𝑗1 < ⋯ < 𝑗𝑙 ≤ 𝑝; 𝑖𝑡 ≠ 𝑗𝑠 ) is given by

𝐹𝑋 (𝑘)|𝑋 (𝑙) (𝑥 (𝑘) |𝑥 (𝑙) ) = 𝐹𝑋 (𝑥𝑖1 , 𝑥𝑖2 , ⋯ , 𝑥𝑖𝑘 |𝑥𝑗1 , 𝑥𝑗2 , ⋯ , 𝑥𝑗𝑙 )
~ ~ ~ ~ 𝑖1 ,⋯,𝑋𝑖𝑘 |𝑋𝑗1 ,⋯,𝑋𝑗𝑙

lim
𝑥𝑡 → +∞
𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 )
~
𝑃(𝑋𝑖1 ≤ 𝑥𝑖1 , ⋯ , 𝑋𝑖𝑘 ≤ 𝑥𝑖𝑘 , 𝑋𝑗1 ≤ 𝑥𝑗1 , ⋯ , 𝑋𝑗𝑙 ≤ 𝑥𝑗𝑙 ) 𝑡≠𝑖1 ,𝑖2 ,⋯,𝑖𝑘 ,𝑗1 ,𝑗2 ,⋯,𝑗𝑙
= =
𝑃(𝑋𝑗1 ≤ 𝑥𝑗1 , ⋯ , 𝑋𝑗𝑙 ≤ 𝑥𝑗𝑙 ) 𝑥𝑡 → +∞
lim 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 )
~
𝑡≠𝑗1 ,𝑗2 ,⋯,𝑗𝑙

B. The conditional probability mass function (p.m.f.) 𝑝(𝑋 , ⋯ , 𝑋 |𝑋 , ⋯ , 𝑋 ) (𝑥1 , ⋯ , 𝑥𝑞 |𝑥𝑞+1 , ⋯ , 𝑥𝑝 ) of a random
1 𝑞 𝑞+1 𝑝
𝑇
vector 𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑞 ) , (𝑞 < 𝑝) given that 𝑋 (2) = 𝑥 (2) or 𝑋𝑞+1 = 𝑥𝑞+1 , 𝑋𝑞+2 = 𝑥𝑞+2 , ⋯ , 𝑋𝑝 = 𝑥𝑝 is given by
~ ~ ~

𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑞 , 𝑥𝑞+1 , ⋯ , 𝑥𝑝 )
𝑝 (1) (2) (𝑥 (1) |𝑥 (2) ) = 𝑝(𝑋 , ⋯ , 𝑋 |𝑋 , ⋯ , 𝑋 ) (𝑥1 , ⋯ , 𝑥𝑞 |𝑥𝑞+1 , ⋯ , 𝑥𝑝 ) =
(𝑋 |𝑋 ) ~ ~ 1 𝑞 𝑞+1 𝑝 𝑝𝑋𝑞+1,𝑋𝑞+2,⋯,𝑋𝑝 (𝑥𝑞+1 , ⋯ , 𝑥𝑝 )
~ ~

𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑞 , 𝑥𝑞+1 , ⋯ , 𝑥𝑝 )
~
= ; if 𝑿 is discrete, and
∑𝑦1∈ ℛ ∑𝑦2 ∈ ℛ ⋯ ∑𝑦𝑞 ∈ ℛ 𝑝𝑋 (𝑦1 , 𝑦2 , ⋯ , 𝑦𝑞 , 𝑥𝑞+1 , ⋯ , 𝑥𝑝 )
~

the marginal probability density function (p.d.f.) 𝑓(𝑋 , ⋯ , 𝑋 |𝑋 , ⋯ , 𝑋 ) (𝑥1 , ⋯ , 𝑥𝑞 |𝑥𝑞+1 , ⋯ , 𝑥𝑝 ) of a random vector
1 𝑞 𝑞+1 𝑝
𝑇
𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑞 ) , (𝑞 < 𝑝) given that 𝑋 (2) = 𝑥 (2) or 𝑋𝑞+1 = 𝑥𝑞+1 , 𝑋𝑞+2 = 𝑥𝑞+2 , ⋯ , 𝑋𝑝 = 𝑥𝑝 is given by
~ ~ ~

𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑞 , 𝑥𝑞+1 , ⋯ , 𝑥𝑝 )
𝑓 (1) (2) (𝑥 (1) |𝑥 (2) ) = 𝑓(𝑋 , ⋯ , 𝑋 |𝑋 , ⋯ , 𝑋 ) (𝑥1 , ⋯ , 𝑥𝑞 |𝑥𝑞+1 , ⋯ , 𝑥𝑝 ) = ~
(𝑋 |𝑋 ) ~ ~ 1 𝑞 𝑞+1 𝑝 𝑓𝑋𝑞+1,𝑋𝑞+2,⋯,𝑋𝑝 (𝑥𝑞+1 , ⋯ , 𝑥𝑝 )
~ ~

𝑓𝑋 (𝑥1 , ⋯ , 𝑥𝑞 , 𝑥𝑞+1 , ⋯ , 𝑥𝑝 )
~
= ∞ ∞ ; if 𝑋 is continuous,
∫−∞ ⋯ ∫−∞ 𝑓𝑋 (𝑦1 , ⋯ , 𝑦𝑞 , 𝑥𝑞+1 , ⋯ , 𝑥𝑝 )𝑑𝑦1 ⋯ 𝑑𝑦𝑞 ~
~

where 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) and 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) are the joint probability mass function (p.m.f.) and probability density
~ ~

function (p.d.f.) of 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 , respectively.

STATISTICAL INDEPENDENCE

𝑇 𝑇
The random vector 𝑋 (𝑘) = (𝑋𝑖1 , 𝑋𝑖2 , ⋯ , 𝑋𝑖𝑘 ) is independent of 𝑋 (𝑙) = (𝑋𝑗1 , 𝑋𝑗2 , ⋯ , 𝑋𝑗𝑙 ) ; (1 ≤ 𝑖1 < ⋯ < 𝑖𝑘 ≤ 𝑝, 1 ≤
~ ~
𝑇
𝑗1 < ⋯ < 𝑗𝑙 ≤ 𝑝; 𝑖𝑡 ≠ 𝑗𝑠 ) iff the conditional d.f. of random vector 𝑋 (𝑘) = (𝑋𝑖1 , 𝑋𝑖2 , ⋯ , 𝑋𝑖𝑘 ) given that 𝑋 (𝑙) =
~ ~

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 5


MULTIVARIATE PROBABILITY DISTRIBUTION

𝑇
(𝑋𝑗1 , 𝑋𝑗2 , ⋯ , 𝑋𝑗𝑙 ) = 𝑥 (𝑙) or (𝑋𝑗1 = 𝑥𝑗1 , 𝑋𝑗2 = 𝑥𝑗2 , ⋯ , 𝑋𝑗𝑙 = 𝑥𝑗𝑙 ) is the marginal d.f. of random vector 𝑋 (𝑘) =
~ ~
𝑇
(𝑋𝑖1 , 𝑋𝑖2 , ⋯ , 𝑋𝑖𝑘 ) i.e.,

𝐹𝑋 (𝑘)|𝑋 (𝑙) (𝑥 (𝑘) |𝑥 (𝑙) ) = 𝐹𝑋 (𝑘) (𝑥𝑖1 , 𝑥𝑖2 , ⋯ , 𝑥𝑖𝑘 ) = 𝐹𝑋 (𝑘) (𝑥 (𝑘) )
~ ~ ~ ~ ~ ~ ~

𝑇 𝑇 𝑇
Any subset 𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑞 ) ; (𝑞 < 𝑝) of 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) is independent of 𝑋 (2) = (𝑋𝑞+1 , 𝑋𝑞+2 , ⋯ , 𝑋𝑝 ) iff
~ ~ ~

𝐹 (1) (2) (𝑥 (1) |𝑥 (2) ) = 𝐹𝑋 (1) (𝑥 (1) )


(𝑋 |𝑋 ) ~ ~ ~ ~
~ ~

𝑜𝑟, 𝐹𝑋 (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) = 𝐹𝑋 (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑞 , ∞, ∞, ⋯ , ∞) ∙ 𝐹𝑋 (∞, ∞, ⋯ , ∞, 𝑋𝑞+1 , 𝑋𝑞+2 , ⋯ , 𝑋𝑝 )


~ ~ ~

= 𝐹𝑋 (1) (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑞 ) ∙ 𝐹𝑋 (2) (𝑋𝑞+1 , 𝑋𝑞+2 , ⋯ , 𝑋𝑝 )


~ ~

MUTUALLY INDEPENDENT:

A. A collection of jointly distributed random variables (RVs) 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 is said to be mutually or completely or

stochastically independent if and only if the joint cumulative distribution function (c.d.f.) 𝐹𝑋 (𝑥 ) of the random vector
~ ~
𝑇
𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) can be expressed as
~

𝑝
𝑇
𝐹𝑋 (𝑥 ) = 𝐹(𝑋1,𝑋2,⋯,𝑋𝑝) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) = ∏ 𝐹𝑋𝑖 (𝑥𝑖 ) ; ∀ 𝑥 = (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ∈ ℛ𝑝 ,
~ ~ ~
𝑖=1

where 𝐹𝑋𝑖 (𝑥𝑖 ) be the marginal cumulative distribution function (c.d.f.) of 𝑋𝑖 ; 𝑖 = 1(1)𝑝.

B. A collection of jointly distributed discrete random variables (RVs) 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 is said to be mutually or

completely or stochastically independent if and only if the joint probability mass function (p.m.f.) 𝑝𝑋 (𝑥 ) of the random
~ ~
𝑇
vector 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) can be expressed as
~

𝑝
𝑇
𝑝𝑋 (𝑥 ) = 𝑝(𝑋1,𝑋2,⋯,𝑋𝑝) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) = ∏ 𝑝𝑋𝑖 (𝑥𝑖 ) ; ∀ 𝑥 = (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ∈ ℛ𝑝 ,
~ ~ ~
𝑖=1

where 𝑝𝑋𝑖 (𝑥𝑖 ) be the marginal probability mass function (p.m.f.) of 𝑋𝑖 ; 𝑖 = 1(1)𝑝.

A collection of jointly distributed continuous random variables (RVs) 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 is said to be mutually or completely

or stochastically independent if and only if the joint probability density function (p.d.f.) 𝑓𝑋 (𝑥 ) of the random vector 𝑋 =
~ ~ ~
𝑇
(𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) can be expressed as

𝑝
𝑇
𝑓𝑋 (𝑥 ) = 𝑓(𝑋1,𝑋2,⋯,𝑋𝑝) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) = ∏ 𝑓𝑋𝑖 (𝑥𝑖 ) ; ∀ 𝑥 = (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ∈ ℛ𝑝 ,
~ ~ ~
𝑖=1

where 𝑓𝑋𝑖 (𝑥𝑖 ) be the marginal probability density function (p.d.f.) of 𝑋𝑖 ; 𝑖 = 1(1)𝑝.

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 6


MULTIVARIATE PROBABILITY DISTRIBUTION

PAIRWISE INDEPENDENT:

A. A collection of jointly distributed random variables (RVs) 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 is said to be pairwise independent if and
only if every pair of them are independent i.e.,

𝐹(𝑋𝑖,𝑋𝑗 ) (𝑥𝑖 , 𝑥𝑗 ) = 𝐹𝑋𝑖 (𝑥𝑖 ) × 𝐹𝑋𝑗 (𝑥𝑗 ) ; ∀ 𝑖 ≠ 𝑗, 𝑥𝑖 , 𝑥𝑗 ∈ ℛ, 𝑖, 𝑗 = 1(1)𝑝

where 𝐹(𝑋𝑖,𝑋𝑗 ) (𝑥𝑖 , 𝑥𝑗 ) be the joint cumulative distribution function (c.d.f.) of the random vector (𝑋𝑖 , 𝑋𝑗 ) and 𝐹𝑋𝑖 (𝑥𝑖 ) be

the marginal cumulative distribution function (c.d.f.) of 𝑋𝑖 ; (𝑖 ≠ 𝑗), 𝑖, 𝑗 = 1(1)𝑝.

B. A collection of jointly distributed discrete random variables (RVs) 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 is said to be pairwise independent
if and only if every pair of them are independent i.e.,

𝑝(𝑋𝑖,𝑋𝑗 ) (𝑥𝑖 , 𝑥𝑗 ) = 𝑝𝑋𝑖 (𝑥𝑖 ) × 𝑝𝑋𝑗 (𝑥𝑗 ) ; ∀ 𝑖 ≠ 𝑗, 𝑥𝑖 , 𝑥𝑗 ∈ ℛ, 𝑖, 𝑗 = 1(1)𝑝

where 𝑝(𝑋𝑖,𝑋𝑗 ) (𝑥𝑖 , 𝑥𝑗 ) be the joint probability mass function (p.m.f.) of the random vector (𝑋𝑖 , 𝑋𝑗 ) and 𝑝𝑋𝑖 (𝑥𝑖 ) be the

marginal probability mass function (p.m.f.) of 𝑋𝑖 ; (𝑖 ≠ 𝑗), 𝑖, 𝑗 = 1(1)𝑝.

A collection of jointly distributed continuous random variables (RVs) 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 is said to be pairwise independent if
and only if every pair of them are independent i.e.,

𝑓(𝑋𝑖,𝑋𝑗 ) (𝑥𝑖 , 𝑥𝑗 ) = 𝑓𝑋𝑖 (𝑥𝑖 ) × 𝑓𝑋𝑗 (𝑥𝑗 ) ; ∀ 𝑖 ≠ 𝑗, 𝑥𝑖 , 𝑥𝑗 ∈ ℛ, 𝑖, 𝑗 = 1(1)𝑝

where 𝑓(𝑋𝑖 ,𝑋𝑗 ) (𝑥𝑖 , 𝑥𝑗 ) be the joint probability density function (p.d.f.) of the random vector (𝑋𝑖 , 𝑋𝑗 ) and 𝑓𝑋𝑖 (𝑥𝑖 ) be the

marginal probability density function (p.d.f.) of 𝑋𝑖 ; (𝑖 ≠ 𝑗), 𝑖, 𝑗 = 1(1)𝑝.

NOTE:

1. A collection of jointly distributed discrete random variables (RVs) 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 is said to be pairwise independent
if and only if every pair of them are independent i.e., 𝐹(𝑋 |𝑋 ) (𝑥𝑖 |𝑥𝑗 ) = 𝐹𝑋𝑖 (𝑥𝑖 ) ; ∀ 𝑖 ≠ 𝑗, 𝑥𝑖 , 𝑥𝑗 ∈ ℛ, 𝑖, 𝑗 = 1(1)𝑝,
𝑖 𝑗

where 𝐹(𝑋 |𝑋 ) (𝑥𝑖 |𝑥𝑗 ) be the conditional cumulative distribution function (c.d.f.) of (𝑋𝑖 |𝑋𝑗 ) and 𝐹𝑋𝑖 (𝑥𝑖 ) be the marginal
𝑖 𝑗

cumulative distribution function (c.d.f.) of 𝑋𝑖 ; (𝑖 ≠ 𝑗), 𝑖, 𝑗 = 1(1)𝑝.

2. If 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 are (complete or mutually) independent, every subset {(𝑋𝑖1 , 𝑋𝑖2 , ⋯ , 𝑋𝑖𝑘 ); 1 ≤ 𝑖1 < ⋯ < 𝑖𝑘 ≤ 𝑝}
of (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) is also independent.

EXAMPLE: Pairwise independent does not imply mutually independent

Let us consider the joint p.m.f. of (𝑋1 , 𝑋2 , 𝑋3 ) be

3
; 𝑖𝑓 (𝑥1 , 𝑥2 , 𝑥3 ) ∈ {(0, 0, 0), (0, 1, 1), (1, 0, 1), (1, 1, 0)}
𝑃(𝑋1 = 𝑥1 , 𝑋2 = 𝑥2 , 𝑋3 = 𝑥3 ) = { 16
1
; 𝑖𝑓 (𝑥1 , 𝑥2 , 𝑥3 ) ∈ {(0, 0, 1), (0, 1, 0), (1, 0, 0), (1, 1, 1)}
16

Now, we have

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 7


MULTIVARIATE PROBABILITY DISTRIBUTION

1
𝑃(𝑋𝑖 = 𝑥𝑖 , 𝑋𝑗 = 𝑥𝑗 ) = ; 𝑖𝑓 (𝑥𝑖 , 𝑥𝑗 ) ∈ {(0, 0), (0, 1), (1, 0), (1, 1)}; 𝑖 ≠ 𝑗, 𝑖, 𝑗 = 1, 2, 3
4

1
𝑃(𝑋𝑖 = 𝑥𝑖 ) = ; 𝑥𝑖 = 0, 1; 𝑖 = 1, 2, 3
2

∴ 𝑃(𝑋𝑖 = 𝑥𝑖 , 𝑋𝑗 = 𝑥𝑗 ) = 𝑃(𝑋𝑖 = 𝑥𝑖 )𝑃(𝑋𝑗 = 𝑥𝑗 ) ; ∀ 𝑖 ≠ 𝑗, 𝑖, 𝑗 = 1, 2, 3

The random vectors (𝑋1 , 𝑋2 ), (𝑋1 , 𝑋3 ), and (𝑋2 , 𝑋3 ) are pairwise independent.

Since 𝑃(𝑋1 = 𝑥1 , 𝑋2 = 𝑥2 , 𝑋3 = 𝑥3 ) ≠ 𝑃(𝑋1 = 𝑥1 )𝑃(𝑋2 = 𝑥2 )𝑃(𝑋3 = 𝑥3 ), the random variables 𝑋1 , 𝑋2 , 𝑋3 are not
mutually independent.

THEOREM: Let 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 be independent random variables and 𝑔𝑖 (𝑋𝑖 ) be function of 𝑋𝑖 such that 𝑔𝑖 (𝑋𝑖 ) be a

random variable, 𝑖 = 1(1)𝑝. Then 𝑔1 (𝑋1 ), 𝑔2 (𝑋2 ), ⋯ , 𝑔𝑝 (𝑋𝑝 ) are independent random variables.

MOMENTS:

𝑇 𝑇
Let us define, Mean vector of 𝑋 = 𝐸 (𝑋) = (𝐸(𝑋1 ), 𝐸(𝑋2 ), ⋯ , 𝐸(𝑋𝑝 )) = (𝜇1 , 𝜇2 , ⋯ , 𝜇𝑝 ) = 𝜇 (say)
~ ~ ~

∑ 𝑥 ∙ 𝑝𝑋 (𝑥 ) = ∑ ∑ ⋯ ∑ 𝑥 ∙ 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ; if 𝑋 is discrete
~ ~ ~ ~ ~ ~
𝑥∈ℛ𝑝 𝑥1 ∈ℛ 𝑥2 ∈ℛ 𝑥𝑝 ∈ℛ
~
= ∞ ∞ ∞

∫ 𝑥 ∙ 𝑓𝑋 (𝑥 ) 𝑑𝑥 = ∫ ∫ ⋯ ∫ 𝑥 ∙ 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) 𝑑𝑥1 𝑑𝑥2 ⋯ 𝑑𝑥𝑝 ; if 𝑋 is continuous.


~ ~ ~ ~ ~ ~ ~
{𝑥∈ℛ𝑝
~
−∞ −∞ −∞

provided 𝐸|𝑋𝑖 | < ∞, ∀ 𝑖 = 1(1)𝑝.

𝑇
Variance-Covariance matrix of 𝑋 = Dispersion matrix of 𝑋 = 𝐷 (𝑋) = 𝐸 {(𝑋 − 𝜇 ) (𝑋 − 𝜇 ) } = Σ (say)
~ ~ ~ ~ ~ ~ ~

(𝑋1 − 𝜇1 ) 𝜎11 𝜎12 ⋯ 𝜎1𝑝


(𝑋2 − 𝜇2 ) 𝜎21 𝜎22 ⋯ 𝜎2𝑝
=𝐸 ( ) ((𝑋1 − 𝜇1 ) (𝑋2 − 𝜇2 ) ⋯ (𝑋𝑝 − 𝜇𝑝 )) = ( ⋮ ⋮ ⋱ ⋮ ) = ((𝜎𝑖𝑗 ))𝑝×𝑝

{ (𝑋𝑝 − 𝜇𝑝 ) } 𝜎𝑝1 𝜎𝑝2 ⋯ 𝜎𝑝𝑝

𝑇
∑ {(𝑥 − 𝜇 ) (𝑥 − 𝜇 ) } ∙ 𝑝𝑋 (𝑥 ) ; if 𝑋 is discrete
~ ~ ~ ~ ~ ~ ~
𝑥∈ℛ𝑝
~
=
𝑇
∫ {(𝑥 − 𝜇 ) (𝑥 − 𝜇 ) } ∙ 𝑓𝑋 (𝑥 ) 𝑑𝑥 ; if 𝑋 is continuous.
~ ~ ~ ~ ~ ~ ~ ~
{𝑥∈ℛ𝑝
~

provided 𝐸 |(𝑋𝑖 − 𝐸(𝑋𝑖 )) (𝑋𝑗 − 𝐸(𝑋𝑗 ))| < ∞, ∀ 𝑖, 𝑗 = 1(1)𝑝.

𝜌11 𝜌12 ⋯ 𝜌1𝑝


𝜌21 𝜌22 ⋯ 𝜌2𝑝
Correlation matrix of 𝑋 = 𝐶𝑜𝑟𝑟 (𝑋) = ((𝜌𝑖𝑗 ))
~ ~ 𝑝×𝑝
=( ⋮ ⋮ ⋱ ⋮ ) = 𝑅 (say),
𝜌𝑝1 𝜌𝑝2 ⋯ 𝜌𝑝𝑝

𝐶𝑜𝑣(𝑋𝑖 , 𝑋𝑗 ) = 𝐸{(𝑋𝑖 − 𝜇𝑖 )(𝑋𝑗 − 𝜇𝑗 )} if 𝑖 ≠ 𝑗


where 𝜇𝑖 = 𝐸(𝑋𝑖 ); 𝑖 = 1(1)𝑝, 𝜎𝑖𝑗 = { 𝑖, 𝑗 = 1(1)𝑝,
𝑉𝑎𝑟(𝑋𝑖 ) = 𝐸(𝑋𝑖 − 𝜇𝑖 )2 if 𝑖 = 𝑗,

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 8


MULTIVARIATE PROBABILITY DISTRIBUTION

𝐶𝑜𝑣(𝑋𝑖 , 𝑋𝑗 )
𝐶𝑜𝑟𝑟(𝑋𝑖 , 𝑋𝑗 ) = if 𝑖 ≠ 𝑗
𝜌𝑖𝑗 = √𝑉𝑎𝑟(𝑋𝑖 ) ∙ √𝑉𝑎𝑟(𝑋𝑗 ) and 𝜎𝑖𝑗 = 𝜌𝑖𝑗 𝜎𝑖 𝜎𝑗 ; 𝑖, 𝑗 = 1(1)𝑝.

{𝐶𝑜𝑟𝑟(𝑋𝑖 , 𝑋𝑗 ) = 1 if 𝑖 = 𝑗,

The variance-covariance matrix Σ and the correlation matrix 𝑅 of 𝑋 are always symmetric.
~

Let us define 𝑅𝑖𝑗 = cofactor of 𝜌𝑖𝑗 in 𝑅 ; 𝑖, 𝑗 = 1(1)𝑝.

We know that minor of 𝜌𝑖𝑗 in 𝑅 = (−1)𝑖+𝑗 × cofactor of 𝜌𝑖𝑗 in 𝑅 = (−1)𝑖+𝑗 𝑅𝑖𝑗 .

𝑇
Let 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) be random vector defined on a probability space (Ω, ℱ, 𝑃), and let 𝑔(∙) be a Borel-measurable
~

function. Then the expectation of 𝑔 (𝑋) = 𝑔(𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) is defined as


~

𝐸 {𝑔 (𝑋)} = 𝐸{𝑔(𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 )}
~

∑ 𝑔 (𝑥 ) ∙ 𝑝𝑋 (𝑥 ) = ∑ ∑ ⋯ ∑ 𝑔 (𝑥 ) ∙ 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ; if 𝑋 is discrete
~ ~ ~ ~ ~ ~
𝑥∈ℛ𝑝 𝑥1 ∈ℛ 𝑥2 ∈ℛ 𝑥𝑝 ∈ℛ
~
= ∞ ∞ ∞

∫ 𝑔 (𝑥 ) ∙ 𝑓𝑋 (𝑥 ) 𝑑𝑥 = ∫ ∫ ⋯ ∫ 𝑔 (𝑥 ) ∙ 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) 𝑑𝑥1 𝑑𝑥2 ⋯ 𝑑𝑥𝑝 ; if 𝑋 is continuous.


~ ~ ~ ~ ~ ~ ~
{~𝑥∈ℛ𝑝 −∞ −∞ −∞

provided 𝐸|𝑔(𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 )| < ∞.

The 𝑘-th moment of 𝑔 (𝑋) is defined as


~

∑ 𝑔𝑘 (𝑥 ) ∙ 𝑝𝑋 (𝑥 ) ; if 𝑋 is discrete
~ ~ ~ ~
𝑥∈ℛ𝑝
~
𝐸 {𝑔𝑘 (𝑋)} =
~
∫ 𝑔𝑘 (𝑥 ) ∙ 𝑓𝑋 (𝑥 ) 𝑑𝑥 ; if 𝑋 is continuous, provided 𝐸 |𝑔𝑘 (𝑋 )| < ∞. .
~ ~ ~ ~ ~ ~
{~𝑥∈ℛ𝑝

A. The joint moments about origin of order (ℎ1 + ℎ2 + ⋯ + ℎ𝑝 ) for non-negative integers (ℎ1 , ℎ2 , ⋯ , ℎ𝑝 ) is defined
as

′ ℎ ℎ ℎ
𝜇(ℎ1 ,ℎ2 ,⋯,ℎ𝑝 )
= 𝐸 (𝑋1 1 𝑋2 2 ⋯ 𝑋𝑝 𝑝 )

ℎ ℎ ℎ
∑ ∑ ⋯ ∑ 𝑥1 1 𝑥2 2 ⋯ 𝑥𝑝 𝑝 ∙ 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ; if 𝑋 is discrete
~ ~
𝑥1 ∈ℛ 𝑥2 ∈ℛ 𝑥𝑝 ∈ℛ
= ∞ ∞ ∞
ℎ ℎ ℎ
∫ ∫ ⋯ ∫ 𝑥1 1 𝑥2 2 ⋯ 𝑥𝑝 𝑝 ∙ 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) 𝑑𝑥1 𝑑𝑥2 ⋯ 𝑑𝑥𝑝 ; if 𝑋 is continuous.
~ ~
{ −∞ −∞ −∞

ℎ ℎ ℎ
provided 𝐸 |𝑋1 1 𝑋2 2 ⋯ 𝑋𝑝 𝑝 | < ∞.

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 9


MULTIVARIATE PROBABILITY DISTRIBUTION

MOMENTS ABOUT ORIGIN:


′ ′
𝜇(1,0,⋯,0) = 𝐸(𝑋1 ) = 𝜇1 𝑜𝑟, 𝜇(0,0,⋯,1,⋯,0) = 𝐸(𝑋𝑖 ) = 𝜇𝑖

′ ′
𝜇(1,1,0,⋯,0) = 𝐸(𝑋1 𝑋2 ) = 𝐶𝑜𝑣(𝑋1 , 𝑋2 ) + 𝐸(𝑋1 )𝐸(𝑋2 ) = 𝜎12 + 𝜇1 𝜇2 ; 𝜇(0,⋯,1,⋯,1,⋯,0) = 𝐸(𝑋𝑖 𝑋𝑗 ) = 𝜎𝑖𝑗 + 𝜇𝑖 𝜇𝑗 ; 𝑖 ≠ 𝑗

′ ′
𝜇(2,0,⋯,0) = 𝐸(𝑋12 ) = 𝑉𝑎𝑟(𝑋1 ) + 𝐸 2 (𝑋1 ) = 𝜎11 + 𝜇12 ; 𝜇(0,0,⋯,2,⋯,0) = 𝐸(𝑋𝑖2 ) = 𝑉𝑎𝑟(𝑋𝑖 ) + 𝐸 2 (𝑋𝑖 ) = 𝜎𝑖𝑖 + 𝜇𝑖2

CENTRAL MOMENTS:

𝜇(1,0,⋯,0) = 𝐸(𝑋1 − 𝜇1 ) = 0 𝑜𝑟, 𝜇(0,0,⋯,1,⋯,0) = 𝐸(𝑋𝑖 − 𝜇𝑖 ) = 0

𝜇(1,1,0,⋯,0) = 𝐸{(𝑋1 − 𝜇1 )(𝑋2 − 𝜇2 )} = 𝐶𝑜𝑣(𝑋1 , 𝑋2 ) = 𝜎12 ; 𝜇(0,⋯,1,⋯,1,⋯,0) = 𝐸{(𝑋𝑖 − 𝜇𝑖 )(𝑋𝑗 − 𝜇𝑗 )} = 𝜎𝑖𝑗 ; 𝑖 ≠ 𝑗

𝜇(2,0,⋯,0) = 𝐸(𝑋1 − 𝜇1 )2 = 𝑉𝑎𝑟(𝑋1 ) = 𝜎11 ; 𝜇(0,0,⋯,2,⋯,0) = 𝐸(𝑋𝑖 − 𝜇𝑖 )2 = 𝑉𝑎𝑟(𝑋𝑖 ) = 𝜎𝑖𝑖

B. MARGINAL MOMENTS:

𝑇 𝑇
The joint moments of 𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑞 ) ; (𝑞 < 𝑝), a subset of random vector 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) , of order
~ ~

(ℎ1 + ℎ2 + ⋯ + ℎ𝑞 ) for non-negative integers (ℎ1 , ℎ2 , ⋯ , ℎ𝑞 ) is defined as (computed from the marginal distribution)

′ ℎ ℎ ℎ ℎ ℎ ℎ ℎ ℎ ℎ
𝜇(ℎ1 ,ℎ2 ,⋯,ℎ𝑞 )
= 𝐸 (𝑋1 1 𝑋2 2 ⋯ 𝑋𝑞 𝑞 ) = 𝐸 (𝑋1 1 𝑋2 2 ⋯ 𝑋𝑞 𝑞 𝑋𝑞+1
0 0
𝑋𝑞+2 ⋯ 𝑋𝑝0 ) ; provided 𝐸 |𝑋1 1 𝑋2 2 ⋯ 𝑋𝑞 𝑞 | < ∞.

ℎ ℎ ℎ
∑ ∑ ⋯ ∑ 𝑥1 1 𝑥2 2 ⋯ 𝑥𝑞 𝑞 ∙ 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ; if 𝑋 is discrete
~ ~
𝑥1 ∈ℛ 𝑥2 ∈ℛ 𝑥𝑝 ∈ℛ
= ∞ ∞ ∞
ℎ ℎ ℎ
∫ ∫ ⋯ ∫ 𝑥1 1 𝑥2 2 ⋯ 𝑥𝑞 𝑞 ∙ 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) 𝑑𝑥1 𝑑𝑥2 ⋯ 𝑑𝑥𝑝 ; if 𝑋 is continuous.
~ ~
{ −∞ −∞ −∞

ℎ ℎ ℎ
∑ ∑ ⋯ ∑ 𝑥1 1 𝑥2 2 ⋯ 𝑥𝑞 𝑞 ∙ 𝑝𝑋 (1) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑞 ) ; if 𝑋 (1) is discrete
~ ~
𝑥1 ∈ℛ 𝑥2 ∈ℛ 𝑥𝑞 ∈ℛ
= ∞ ∞ ∞
ℎ ℎ ℎ
∫ ∫ ⋯ ∫ 𝑥1 1 𝑥2 2 ⋯ 𝑥𝑞 𝑞 ∙ 𝑓𝑋 (1) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑞 ) 𝑑𝑥1 𝑑𝑥2 ⋯ 𝑑𝑥𝑞 ; if 𝑋 (1) is continuous.
~ ~
{−∞ −∞ −∞

C. CONDITIONAL MOMENTS:

𝑇
Let 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) be random vector defined on a probability space (Ω, ℱ, 𝑃), and let 𝑔(∙) be a Borel-measurable
~

function. Then the conditional expectation of 𝑔 (𝑋 (1) ) = 𝑔(𝑋1 , 𝑋2 , ⋯ , 𝑋𝑞 ) given 𝑋 (2) = 𝑥 (2) ((𝑋𝑞+1 , 𝑋𝑞+2 , ⋯ , 𝑋𝑝 ) =
~ ~ ~

(𝑥𝑞+1 , 𝑥𝑞+2 , ⋯ , 𝑥𝑝 )) is defined as

𝐸 {𝑔 (𝑋 (1) ) |𝑥 (2) } = 𝐸{𝑔(𝑋1 , 𝑋2 , ⋯ , 𝑋𝑞 )|(𝑥𝑞+1 , 𝑥𝑞+2 , ⋯ , 𝑥𝑝 )} =


~ ~

∑ ∑ ⋯ ∑ 𝑔(𝑥1 , 𝑥2 , ⋯ , 𝑥𝑞 ) ∙ 𝑝 (1) (2) (𝑥 (1) |𝑥 (2) ) ; if 𝑋 is discrete and 𝑝𝑋 (2) (𝑥 (2) ) > 0
(𝑋 |𝑋 ) ~ ~ ~ ~ ~
𝑥1 ∈ℛ 𝑥2 ∈ℛ 𝑥𝑞 ∈ℛ ~ ~
∞ ∞ ∞

∫ ∫ ⋯ ∫ 𝑔(𝑥1 , 𝑥2 , ⋯ , 𝑥𝑞 ) ∙ 𝑓 (1) (2) (𝑥 (1) |𝑥 (2) ) 𝑑𝑥1 𝑑𝑥2 ⋯ 𝑑𝑥𝑞 ; if 𝑋 is continuous and 𝑓𝑋 (2) (𝑥 (2) ) > 0.
(𝑋 |𝑋 ) ~ ~ ~ ~ ~
{−∞ −∞ −∞ ~ ~

D. MOMENTS UNDER INDEPENDENCE:

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 10


MULTIVARIATE PROBABILITY DISTRIBUTION

𝑇
If the random vector 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) are mutually independent, the joint moments about origin of order
~

(ℎ1 + ℎ2 + ⋯ + ℎ𝑝 ) for non-negative integers (ℎ1 , ℎ2 , ⋯ , ℎ𝑝 ) is defined as

′ ℎ ℎ ℎ
𝜇(ℎ1 ,ℎ2 ,⋯,ℎ𝑝 )
= 𝐸 (𝑋1 1 𝑋2 2 ⋯ 𝑋𝑝 𝑝 )

𝑝
ℎ ℎ ℎ ℎ
∑ ∑⋯ ∑ 𝑥1 1 𝑥2 2 ⋯ 𝑥𝑝 𝑝 ∙ 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) = ∏ { ∑ 𝑥𝑖 𝑖 ∙ 𝑝𝑋𝑖 (𝑥𝑖 )} ; if 𝑋 is discrete
~ ~
𝑥1 ∈ℛ 𝑥2 ∈ℛ 𝑥𝑝 ∈ℛ 𝑖=1 𝑥𝑖 ∈ℛ
= ∞ ∞ ∞ 𝑝 ∞
ℎ ℎ ℎ ℎ
∫ ∫⋯ ∫ 𝑥1 1 𝑥2 2 ⋯ 𝑥𝑝 𝑝 ∙ 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) 𝑑𝑥1 𝑑𝑥2 ⋯ 𝑑𝑥𝑝 = ∏ { ∫ 𝑥𝑖 𝑖 ∙ 𝑓𝑋𝑖 (𝑥𝑖 ) 𝑑𝑥𝑖 } ; if 𝑋 is continuous.
~ ~
{−∞ −∞ −∞ 𝑖=1 −∞

𝑝
ℎ ℎ
= ∏ 𝐸(𝑋𝑖 𝑖 ) ; provided 𝐸|𝑋𝑖 𝑖 | < ∞, ∀ 𝑖 = 1(1)𝑝.
𝑖=1

CHARACTERISTIC FUNCTION:

𝑇 𝑇
Let 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) be random vector defined on a probability space (Ω, ℱ, 𝑃) and 𝑡 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 ) ∈ ℛ𝑝 . Then
~ ~

characteristic function of 𝑋 or the joint c.f. of the random variables 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 is denoted by 𝜙𝑋 (𝑡 ) and defined as
~ ~ ~

𝑝
𝑖(𝑡 𝑇 𝑋) 𝑖𝑡1 𝑋1 +𝑖𝑡2 𝑋2 + ⋯+𝑖𝑡𝑝 𝑋𝑝
𝜙𝑋 (𝑡 ) = 𝜙𝑋 (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 ) = 𝐸 {𝑒 ~ ~ } = 𝐸{𝑒 } = 𝐸 {exp (𝑖 ∑ 𝑡𝑗 𝑋𝑗 )} ; 𝑖 = √−1
~ ~ ~
𝑗=1

𝑝 𝑝

= 𝐸 {cos (∑ 𝑡𝑗 𝑋𝑗 )} + 𝑖𝐸 {sin (∑ 𝑡𝑗 𝑋𝑗 )}
𝑗=1 𝑗=1

𝑝
𝑖(𝑡 𝑇 𝑋 ) 𝑖 ∑𝑗=1 𝑡𝑗𝑋𝑗
∑ 𝑒 ~ ~ ∙ 𝑝𝑋 (𝑥 ) = ∑ ∑ ⋯ ∑ 𝑒 ∙ 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ; if 𝑋 is discrete
~ ~ ~ ~
𝑥∈ℛ𝑝 𝑥1 ∈ℛ 𝑥2 ∈ℛ 𝑥𝑝 ∈ℛ
~
= ∞ ∞ ∞
𝑝
𝑖(𝑡 𝑇 𝑋 ) 𝑖 ∑𝑗=1 𝑡𝑗𝑋𝑗
∫ 𝑒 ~ ~ ∙ 𝑓𝑋 (𝑥 ) 𝑑𝑥 = ∫ ∫ ⋯ ∫ 𝑒 ∙ 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) 𝑑𝑥1 𝑑𝑥2 ⋯ 𝑑𝑥𝑝 ; if 𝑋 is continuous.
~ ~ ~ ~ ~
{~𝑥∈ℛ𝑝 −∞ −∞ −∞

𝑇
The characteristic function of 𝑋 i.e., 𝜙𝑋 (𝑡 ) always exists for all 𝑡 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 ) ∈ ℛ𝑝 .
~ ~ ~ ~

PROPERTIES OF CHARACTERISTIC FUNCTION:

Some properties of c.f.’s are

(𝑖) 𝜙𝑋 (0, 0, ⋯ ,0) = 𝜙𝑋 (0) = 1 (𝑖𝑖) 𝜙− 𝑋 (𝑡 ) = 𝜙𝑋 (𝑡 ) (𝑖𝑖𝑖) |𝜙𝑋 (𝑡 )| ≤ 1


~ ~ ~ ~ ~ ~ ~ ~ ~

(𝑖𝑣) 𝜙𝑋 (𝑡 ) is uniformly continuous in 𝑛-dimensional Euclidian space.


~ ~

(𝑣) If 𝜙𝑋 (𝑡 ) = 𝜙𝑌 (𝑡 ) for all 𝑡 ∈ ℛ𝑝 , then 𝑋 and 𝑌 have the same distribution.


~ ~ ~ ~ ~ ~ ~

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 11


MULTIVARIATE PROBABILITY DISTRIBUTION

𝑝 𝑝
𝑖 ∑𝑗=1 𝑡𝑗 𝑏𝑗 𝑖 ∑𝑗=1 𝑡𝑗𝑎𝑗 𝑋𝑗 𝑖(𝑏 𝑇 𝑡 )
(𝑣𝑖) 𝜙𝑎𝑇 𝑋+𝑏 (𝑡 ) = 𝜙𝑎1 𝑋1 +𝑏1,⋯,𝑎𝑝 𝑋𝑝+𝑏𝑝 (𝑡 ) = 𝑒 ∙ 𝐸 {𝑒 }=𝑒 ~ ~ ∙ 𝜙𝑋 (𝑎1 𝑡1 , 𝑎2 𝑡2 , ⋯ , 𝑎𝑝 𝑡𝑝 )
~ ~ ~ ~ ~ ~

(𝑣𝑖𝑖) The random variables 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 are (mutually) independent iff 𝜙𝑋 (𝑡 ) = 𝜙𝑋1 (𝑡1 ) ∙ 𝜙𝑋2 (𝑡2 ) ⋯ 𝜙𝑋𝑝 (𝑡𝑝 ).
~ ~


(𝑣𝑖𝑖𝑖) If 𝐸 |𝑋1ℎ1 𝑋2ℎ2 ⋯ 𝑋𝑝 𝑝 | < ∞, the joint moments about origin 𝜇(ℎ

of order (ℎ1 + ℎ2 + ⋯ + ℎ𝑝 ) for non-
1 ,ℎ2 ,⋯,ℎ𝑝 )

negative integers (ℎ1 , ℎ2 , ⋯ , ℎ𝑝 ) is given by

𝜕 (ℎ1 +ℎ2 +⋯+ℎ𝑝) 𝑝


∑𝑗=1 ℎ𝑗 ℎ ℎ ℎ ′
ℎ ℎ ℎ
𝜙𝑋 (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 )| =𝑖 𝐸 (𝑋1 1 𝑋2 2 ⋯ 𝑋𝑝 𝑝 ) = 𝑖 (ℎ1 +ℎ2 +⋯+ℎ𝑝 ) 𝜇(ℎ1 ,ℎ2 ,⋯,ℎ𝑝 )
; 𝑖 = √−1
𝜕𝑡1 1 𝜕𝑡2 2 ⋯ 𝜕𝑡𝑝 𝑝 ~
𝑡1 =𝑡2 =⋯=𝑡𝑝 =0

𝜕ℎ
In particular, 𝜙𝑋 (𝑡 )| = 𝑖 ℎ 𝐸(𝑋𝑗ℎ ); 𝑗 = 1, 2, ⋯ , 𝑝.
𝜕𝑡𝑗ℎ ~ ~
𝑡 =0
~ ~

Let 𝑔(∙) be a Borel-measurable function. Then characteristic function of 𝑔 (𝑋) = 𝑔(𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) is defined as
~

𝑖𝑡𝑔(𝑋)
𝜙𝑔(𝑋) (𝑡) = 𝐸 {𝑒 ~ } = 𝐸 {exp (𝑖𝑡𝑔 (𝑋))}.
~ ~

MOMENT GENERATING FUNCTION:

𝑇 𝑇
Let 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) be random vector defined on a probability space (Ω, ℱ, 𝑃) and 𝑡 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 ) ∈ ℛ𝑝 . Then
~ ~

the moment generating function (𝑚. 𝑔. 𝑓. ) of 𝑋 or the 𝑚. 𝑔. 𝑓. of the joint distribution of (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) is defined as
~

𝑝
𝑡 𝑇𝑋 𝑡1 𝑋1 +𝑡2 𝑋2 + ⋯+𝑡𝑝 𝑋𝑝 }
𝑀𝑋 (𝑡 ) = 𝑀𝑋 (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 ) = 𝐸 (𝑒 ~ ~ ) = 𝐸{𝑒 = 𝐸 {exp (∑ 𝑡𝑗 𝑋𝑗 )}
~ ~ ~
𝑗=1

𝑝
𝑡 𝑇𝑋 ∑𝑗=1 𝑡𝑗 𝑋𝑗
∑ 𝑒~ ~ ∙ 𝑝𝑋 (𝑥 ) = ∑ ∑ ⋯ ∑ 𝑒 ∙ 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ; if 𝑋 is discrete
~ ~ ~ ~
𝑥∈ℛ𝑝 𝑥1 ∈ℛ 𝑥2 ∈ℛ 𝑥𝑝 ∈ℛ
~
= ∞ ∞ ∞
𝑝
𝑡 𝑇𝑋 ∑𝑗=1 𝑡𝑗 𝑋𝑗
∫ 𝑒~ ~ ∙ 𝑓𝑋 (𝑥 ) 𝑑𝑥 = ∫ ∫ ⋯ ∫ 𝑒 ∙ 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) 𝑑𝑥1 𝑑𝑥2 ⋯ 𝑑𝑥𝑝 ; if 𝑋 is continuous.
~ ~ ~ ~ ~
{~𝑥∈ℛ𝑝 −∞ −∞ −∞

𝑝
𝑡 𝑇𝑋 ∑𝑗=1 𝑡𝑗𝑋𝑗
provided 𝐸 |𝑒 ~ ~ | = 𝐸 |𝑒 | < ∞, for |𝑡𝑗 | ≤ ℎ𝑗 ; 𝑗 = 1, 2, ⋯ , 𝑝, for some ℎ𝑗 > 0 ; 𝑗 = 1, 2, ⋯ , 𝑝.

PROPERTIES OF MOMENT GENERATING FUNCTION:

Some properties of m.g.f.’s are

(𝑖) 𝑀𝑋 (0, 0, ⋯ ,0) = 𝑀𝑋 (0) = 1


~ ~ ~

(𝑖𝑖) If 𝑀𝑋 (𝑡 ) = 𝑀𝑌 (𝑡 ) for all 𝑡 ∈ ℛ𝑝 , then 𝑋 and 𝑌 have the same distribution. The m.g.f. 𝑀𝑋 (𝑡 ) uniquely
~ ~ ~ ~ ~ ~ ~ ~ ~
𝑇
determines the joint distribution of 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) , and conversely, if the m.g.f. exists, it is unique.
~

𝑡 𝑇 (𝑎 𝑇 𝑋+𝑏 )
(𝑖𝑖𝑖) 𝑀𝑎𝑇𝑋+𝑏 (𝑡 ) = 𝑀𝑎1 𝑋1+𝑏1,⋯,𝑎𝑝𝑋𝑝 +𝑏𝑝 (𝑡 ) = 𝐸 {𝑒 ~ ~ ~ ~ } = 𝐸{𝑒 𝑡1(𝑎1𝑋1+𝑏1)+ ⋯+𝑡𝑝(𝑎𝑝𝑋𝑝 +𝑏𝑝) }
~ ~ ~ ~ ~

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 12


MULTIVARIATE PROBABILITY DISTRIBUTION

𝑝
∑𝑗=1 𝑡𝑗𝑏𝑗 𝑏𝑇 𝑡
=𝑒 ∙ 𝐸{𝑒 𝑡1𝑎1𝑋1+ ⋯+𝑡𝑝𝑎𝑝 𝑋𝑝 } = 𝑒 ~ ~ ∙ 𝑀𝑋 (𝑎1 𝑡1 , 𝑎2 𝑡2 , ⋯ , 𝑎𝑝 𝑡𝑝 )
~

(𝑖𝑣) The random variables 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 are (mutually) independent iff

𝑀𝑋 (𝑡 ) = 𝑀𝑋1 (𝑡1 ) ∙ 𝑀𝑋2 (𝑡2 ) ⋯ 𝑀𝑋𝑝 (𝑡𝑝 ).


~ ~

ℎ ℎ ℎ ′
(𝑣) If 𝐸 |𝑋1 1 𝑋2 2 ⋯ 𝑋𝑝 𝑝 | < ∞, the joint moments about origin 𝜇(ℎ1 ,ℎ2 ,⋯,ℎ𝑝 )
of order (ℎ1 + ℎ2 + ⋯ + ℎ𝑝 ) for non-

negative integers (ℎ1 , ℎ2 , ⋯ , ℎ𝑝 ) is given by

′ ℎ ℎ ℎ 𝜕 (ℎ1+ℎ2 +⋯+ℎ𝑝 )
𝜇(ℎ1 ,ℎ2 ,⋯,ℎ𝑝 )
= 𝐸 (𝑋1 1 𝑋2 2 ⋯ 𝑋𝑝 𝑝 ) = ℎ ℎ ℎ
𝑀𝑋 (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 )|
𝜕𝑡1 1 𝜕𝑡2 2 ⋯ 𝜕𝑡𝑝 𝑝 ~
𝑡1 =𝑡2 =⋯=𝑡𝑝 =0

𝜕ℎ
In particular, 𝑀𝑋 (𝑡 )| = 𝐸(𝑋𝑗ℎ ); 𝑗 = 1, 2, ⋯ , 𝑝.
𝜕𝑡𝑗ℎ ~ ~
𝑡 =0
~ ~

𝑇
(𝑣𝑖) The moment generating function of 𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑞 ) ; (𝑞 < 𝑝), a subset of random vector 𝑋 =
~ ~
𝑇 𝑇
(𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) is given by 𝑀𝑋 (1) (𝑡 (1) ) = 𝑀𝑋 (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑞 , 0, 0, ⋯ , 0 ) ; 𝑡 (1) = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑞 ) ∈ ℛ𝑞 .
~ ~ ~ ~

Let 𝑔(∙) be Borel-measurable functions. Then moment generating function of 𝑔 (𝑋) = 𝑔(𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) is defined as
~ ~ ~ ~

𝑡 𝑇 𝑔(𝑋 )
𝑀𝑔(𝑋) (𝑡 ) = 𝐸 {𝑒 ~ ~ ~ } = 𝐸 {exp (𝑡 𝑇 𝑔 (𝑋))}.
~ ~ ~ ~
~ ~

Using the uniqueness property of moment generating function we may find the distribution of 𝑌 = 𝑔(𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ). We
~ ~

compute the m.g.f. of 𝑌 = 𝑔(𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ). If this m.g.f. is one of the known kind, 𝑌 must have this kind of distribution.
~ ~ ~

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 13


MULTIVARIATE PROBABILITY DISTRIBUTION

MULTINOMIAL DISTRIBUTION

INTRODUCTION

Recall the Binomial Distribution. It models the number of "successes" in a fixed number 𝑛 of independent
Bernoulli trials, where each trial has only two possible outcomes (e.g., Success/Failure, Head/Tail, Yes/No)
with constant probabilities 𝑝 and (1 − 𝑝).

But what if each trial has more than two possible outcomes?
o Rolling a die: 6 outcomes {1, 2, 3, 4, 5, 6}
o Survey question with responses: {Strongly Agree, Agree, Neutral, Disagree, Strongly Disagree}
o Drawing balls from an urn with replacement: {Red, Green, Blue}
o Classifying a product from a production line: {Perfect, Minor Flaw, Major Flaw, Unusable}

The Multinomial Distribution is designed for precisely this scenario. It describes the probabilities of the counts
for each category across a fixed number of independent trials.

Consider an experiment with the following characteristics:


1. The experiment consists of 𝑛 fixed, independent trials.
2. Each trial results in exactly one of 𝑘 possible, mutually exclusive, and exhaustive outcomes:
𝐴1 , 𝐴2 , ⋯ , 𝐴𝑘 (say).
3. The probability of outcome 𝐴𝑖 occurring on any given trial is 𝑝𝑖 , where 0 < 𝑝𝑖 < 1 for all 𝑖 = 1(1)𝑘.
4. These probabilities (𝑝𝑖 ) are constant across all trials, and ∑𝑘𝑖=1 𝑝𝑖 = 1.

Let the random variables 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘 denote the number of times each outcome 𝐴1 , 𝐴2 , ⋯ , 𝐴𝑘 occurs in the
𝑛 trials.

Note that these counts must sum to the total number of trials: ∑𝑘𝑖=1 𝑋𝑖 = 𝑛.

The vector of counts 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘 )𝑇 follows a Multinomial Distribution with parameters 𝑛 and 𝑝 =
~ ~

)𝑇
(𝑝1 , 𝑝2 , ⋯ , 𝑝𝑘 . We denote this as: 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘 )𝑇 ~ 𝑀𝑢𝑙𝑡𝑖𝑛𝑜𝑚𝑖𝑎𝑙(𝑛, 𝑝1 , 𝑝2 , ⋯ , 𝑝𝑘 ).
~

The probability of observing exactly 𝑥1 outcomes of type 𝐴1 , 𝑥2 outcomes of type 𝐴2 , ⋯, and 𝑥𝑘 outcomes of
type 𝐴𝑘 , where ∑𝑘𝑖=1 𝑥𝑖 = 𝑛 and each 𝑥𝑖 ≥ 0 (= 0(1)𝑛) is an integer, is given by the 𝑝. 𝑚. 𝑓.:
𝑘
𝑛! 𝑥 𝑥 𝑥
∙ 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑘 𝑘 ; 𝑥𝑖 = 0(1)𝑛, ∑ 𝑥𝑖 = 𝑛,
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑘 !
𝑖=1
𝑝𝑋 (𝑥 ) = 𝑝(𝑋1 ,𝑋2 ,⋯,𝑋𝑘 ) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑘 ) = 𝑘
~ ~
0 < 𝑝𝑖 < 1, ∑ 𝑝𝑖 = 1, 𝑖 = 1(1)𝑘
𝑖=1
{ 0; otherwise

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 14


MULTIVARIATE PROBABILITY DISTRIBUTION

DERIVATION OF THE 𝒑. 𝒎. 𝒇.:

Component 1: Probability of a specific sequence.

Consider one specific sequence of 𝑛 trial outcomes that results in 𝑥1 𝐴1′ 𝑠, 𝑥2 𝐴′2 𝑠, ⋯ , 𝑥𝑘 𝐴𝑘 ′𝑠.

For example, if 𝑛 = 5, 𝑘 = 3, 𝑥1 = 2, 𝑥2 = 1, 𝑥3 = 2, one such sequence could be (𝐴₁, 𝐴₃, 𝐴₁, 𝐴₂, 𝐴₃). Since
the trials are independent, the probability of this specific sequence is the product of the individual trial
probabilities: 𝑝1 ∙ 𝑝3 ∙ 𝑝1 ∙ 𝑝2 ∙ 𝑝3 = 𝑝12 ∙ 𝑝21 ∙ 𝑝32 .
𝑥 𝑥 𝑥
In general, the probability of any one specific sequence with the desired counts 𝑥1 , 𝑥2 , ⋯ , 𝑥𝑘 is: 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑘 𝑘 .

Component 2: Number of possible sequences.

How many different sequences of trial outcomes yield exactly 𝑥1 𝐴1′ 𝑠, 𝑥2 𝐴′2 𝑠, ⋯ , 𝑥𝑘 𝐴𝑘 ′𝑠? This is a
combinatorial problem of arranging 𝑛 items where there are 𝑥1 identical items of type 1, 𝑥2 identical items of
type 2, ... , 𝑥𝑘 identical items of type 𝑘. The number of distinct permutations is given by the multinomial
𝑛!
coefficient:
𝑥1 !𝑥2 !⋯𝑥𝑘 !

Combining the components:

The total probability 𝑃(𝑋1 = 𝑥1 , 𝑋2 = 𝑥2 , ⋯ , 𝑋𝑘 = 𝑥𝑘 ) is the sum of the probabilities of all sequences that
𝑥 𝑥 𝑥
result in these counts. Since each such sequence has the same probability (𝑝1 1 𝑝2 2 ⋯ 𝑝𝑘 𝑘 ) and there are
𝑛!
𝑥1 !𝑥2 !⋯𝑥𝑘 !
such sequences, the total probability is:

𝑛! 𝑥 𝑥 𝑥
(Number of sequences) × (Probability of one sequence) = ∙ 𝑝 1 𝑝 2 ⋯ 𝑝𝑘 𝑘
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑘 ! 1 2

DEFINITION

Suppose a random experiment is repeated 𝑛 times. Each replication of the experiment terminates in one of the
𝑘 mutually exclusive and exhaustive events 𝐴1 , 𝐴2 , ⋯ , 𝐴𝑘 . Let 𝑝𝑖 be the probability that the experiment
terminates in 𝐴𝑖 ; 𝑖 = 1(1)𝑘 and 𝑝𝑖 ; 𝑖 = 1(1)𝑘 remains constant for 𝑛 repetitions. We are assume that the 𝑛
replications are independent. Let 𝑋𝑖 ; 𝑖 = 1(1)𝑘 be the number of times 𝐴𝑖 occurs in 𝑛 replications.

The joint p.m.f. of (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘 ) is given by

𝑘
𝑛! 𝑥 𝑥 𝑥
∙ 𝑝 1 𝑝 2 ⋯ 𝑝𝑘 𝑘 ; 𝑥𝑖 = 0(1)𝑛, ∑ 𝑥𝑖 = 𝑛,
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑘 ! 1 2
𝑖=1
𝑝𝑋 (𝑥 ) = 𝑝(𝑋1 ,𝑋2 ,⋯,𝑋𝑘 ) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑘 ) = 𝑘
~ ~
0 < 𝑝𝑖 < 1, ∑ 𝑝𝑖 = 1, 𝑖 = 1(1)𝑘
𝑖=1
{ 0; otherwise

Then the random vector 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘 )𝑇 is said to follow the multinomial distribution (singular) with
~

parameters (𝑛, 𝑝1 , 𝑝2 , ⋯ , 𝑝𝑘 ).

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 15


MULTIVARIATE PROBABILITY DISTRIBUTION

NOTATION:

𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘 )𝑇 ~ 𝑀𝑢𝑙𝑡𝑖𝑛𝑜𝑚𝑖𝑎𝑙(𝑛, 𝑝1 , 𝑝2 , ⋯ , 𝑝𝑘 ),
~

𝑘 𝑘

where 0 < 𝑝𝑖 < 1, 𝑖 = 1(1)𝑘, ∑ 𝑋𝑖 = 𝑛 and ∑ 𝑝𝑖 = 1.


𝑖=1 𝑖=1

EXAMPLE:

If a fair die is thrown n times and 𝑋𝑖 ; 𝑖 = 1(1)6 is the number of times the 𝑖 th face appears. Then
1 1 1
(𝑋1 , 𝑋2 , ⋯ , 𝑋6 )𝑇 ~ 𝑀𝑢𝑙𝑡𝑖𝑛𝑜𝑚𝑖𝑎𝑙 (𝑛, , , ⋯ , ).
6 6 6

VERIFICATION OF 𝒑. 𝒎. 𝒇.

Since 𝑛 > 0, 𝑥𝑖 ≥ 0, 0 < 𝑝𝑖 < 1; ∀ 𝑖 = 1(1)𝑘,

𝑝𝑋 (𝑥 ) ≥ 0; ∀ 𝑥 ∈ 𝑹𝑘 and
~ ~ ~

𝑛 𝑛 𝑛 𝑛 𝑛 𝑛
𝑛! 𝑥 𝑥 𝑥
∑ ∑ ⋯ ∑ 𝑝𝑋 (𝑥 ) = ∑ ∑ ⋯∑ ∙ 𝑝 1 𝑝 2 ⋯ 𝑝𝑘 𝑘
~ ~ 𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑘 ! 1 2
𝑥1 =0 𝑥2 =0 𝑥𝑘 =0 𝑥1 =0 𝑥2 =0 𝑥𝑘 =0
∑𝑘
𝑖=1 𝑥𝑖 =𝑛 ∑𝑘
𝑖=1 𝑥𝑖 =𝑛

𝑛
𝑘 𝑘

= (𝑝1 + 𝑝2 + ⋯ + 𝑝𝑘 )𝑛 = (∑ 𝑝𝑖 ) = 1 [∵ ∑ 𝑝𝑖 = 1]
𝑖=1 𝑖=1

∴ 𝑝𝑋 (𝑥 ) is the p.m.f. of 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘 )𝑇


~ ~ ~

MOMENT GENERATING FUNCTION:

The moment generating function (𝑚. 𝑔. 𝑓. ) of 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘 )𝑇 or the 𝑚. 𝑔. 𝑓. of the joint distribution of
~

(𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘 ) for all 𝑡 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑘 )𝑇 ∈ ℛ𝑘 is given by


~

𝑡 𝑇𝑋
𝑀𝑋 (𝑡 ) = 𝑀(𝑋1 ,𝑋2 ,⋯,𝑋𝑘) (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑘 ) = 𝐸 (𝑒 ~ ~ ) = 𝐸{𝑒 𝑡1 𝑋1 +𝑡2 𝑋2 +⋯+𝑡𝑘 𝑋𝑘 }
~ ~

𝑛 𝑛 𝑛
𝑛! 𝑥 𝑥 𝑥
= ∑ ∑ ⋯ ∑ 𝑒 𝑡1 𝑥1 +𝑡2 𝑥2 +⋯+𝑡𝑘 𝑥𝑘 ∙ 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑘 𝑘
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑘 !
𝑥1 =0 𝑥2 =0 𝑥𝑘 =0
∑𝑘
𝑖=1 𝑥𝑖 =𝑛

𝑛 𝑛 𝑛
𝑛!
= ∑ ∑ ⋯∑ ∙ (𝑝1 𝑒 𝑡1 )𝑥1 (𝑝2 𝑒 𝑡2 )𝑥2 ⋯ (𝑝𝑘 𝑒 𝑡𝑘 )𝑥𝑘
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑘 !
𝑥1 =0 𝑥2 =0 𝑥𝑘 =0
∑𝑘
𝑖=1 𝑥𝑖 =𝑛

𝑛
𝑘

= (𝑝1 𝑒 𝑡1 + 𝑝2 𝑒 𝑡2 + ⋯ + 𝑝𝑘 𝑒 𝑡𝑘 )𝑛 = (∑ 𝑝𝑖 𝑒 𝑡𝑖 )
𝑖=1

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 16


MULTIVARIATE PROBABILITY DISTRIBUTION

MOMENTS:

Let 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘 )𝑇 ~ 𝑀𝑢𝑙𝑡𝑖𝑛𝑜𝑚𝑖𝑎𝑙(𝑛, 𝑝1 , 𝑝2 , ⋯ , 𝑝𝑘 ).


~

DERIVATION USING INDICATOR VARIABLES:

Let 𝐼𝛼𝑖 be an indicator random variable for trial 𝛼; ( 𝛼 = 1(1)𝑛) and category 𝑖; (𝑖 = 1(1)𝑘), i.e.

1, if trial 𝛼 results in outcome 𝐴𝑖


𝐼𝛼𝑖 = { ; 𝑖 = 1(1)𝑘, 𝛼 = 1(1)𝑛
0, otherwise

By definition, 𝑃(𝐼𝛼𝑖 = 1) = 𝑝𝑖 ; 𝑖 = 1(1)𝑘.

We have 𝐼𝛼𝑖 ~ 𝐵𝑒𝑟𝑛𝑜𝑢𝑙𝑙𝑖(𝑝𝑖 ), 𝐸[𝐼𝛼𝑖 ] = 𝑝𝑖 and 𝑉𝑎𝑟[𝐼𝛼𝑖 ] = 𝑝𝑖 (1 − 𝑝𝑖 ) ; 𝑖 = 1(1)𝑘, 𝛼 = 1(1)𝑛

The expected value of an indicator variable is just the probability of it being 1:

𝐸[𝐼𝛼𝑖 ] = 1 × 𝑃(𝐼𝛼𝑖 = 1) + 0 × 𝑃(𝐼𝛼𝑖 = 0) = 𝑝𝑖 ; 𝑖 = 1(1)𝑘.

The total count for category 𝑖, 𝑋𝑖 , is the sum of the indicators for that category across all 𝑛 trials:

𝑋𝑖 = ∑ 𝐼𝛼𝑖 ; 𝑖 = 1(1)𝑘.
𝛼=1

𝑛 𝑛 𝑛

∴ 𝐸(𝑋𝑖 ) = 𝐸 {∑ 𝐼𝛼𝑖 } = ∑ 𝐸(𝐼𝛼𝑖 ) = ∑ 𝑝𝑖 = 𝑛𝑝𝑖 ; 𝑖 = 1(1)𝑘.


𝛼=1 𝛼=1 𝛼=1

Since 𝐼𝛼𝑖 is a Bernoulli random variable (taking values 1 or 0), its variance is given by 𝑝𝑖 (1 − 𝑝𝑖 ) where 𝑝𝑖 is
the probability of success (being 1). 𝑉𝑎𝑟[𝐼𝛼𝑖 ] = 𝑝𝑖 (1 − 𝑝𝑖 ). Alternatively,

2
𝑉𝑎𝑟(𝐼𝛼𝑖 ) = 𝐸(𝐼𝛼𝑖 ) − 𝐸 2 (𝐼𝛼𝑖 ) = {12 ∙ 𝑃(𝐼𝛼𝑖 = 1) + 02 ∙ 𝑃(𝐼𝛼𝑖 = 0)} − 𝑝𝑖2 = 𝑝𝑖 − 𝑝𝑖2 = 𝑝𝑖 (1 − 𝑝𝑖 ); 𝑖 = 1(1)𝑘

Since the trials are independent, the indicator variables 𝐼𝛼𝑖 (outcome 𝑖 on trial 𝛼) and 𝐼𝛼′ 𝑖 (outcome 𝑖 on trial
𝛼 ′ ) are independent for different trials (𝛼 ≠ 𝛼 ′ ).

𝑛 𝑛 𝑛

𝑉𝑎𝑟(𝑋𝑖 ) = 𝑉𝑎𝑟 {∑ 𝐼𝛼𝑖 } = ∑ 𝑉𝑎𝑟(𝐼𝛼𝑖 ) = ∑ 𝑝𝑖 (1 − 𝑝𝑖 ) = 𝑛𝑝𝑖 (1 − 𝑝𝑖 ) ; 𝑖 = 1(1)𝑘.


𝛼=1 𝛼=1 𝛼=1

𝑛 𝑛 𝑛 𝑛

𝐶𝑜𝑣(𝑋𝑖 , 𝑋𝑗 ) = 𝐶𝑜𝑣 {∑ 𝐼𝛼𝑖 , ∑ 𝐼𝛼′ 𝑗 } = ∑ ∑ 𝐶𝑜𝑣(𝐼𝛼𝑖 , 𝐼𝛼′ 𝑗 ) ; 𝑖≠𝑗


𝛼=1 𝛼 ′ =1 𝛼=1 𝛼 ′ =1

Now, consider the covariance term 𝐶𝑜𝑣(𝐼𝛼𝑖 , 𝐼𝛼′ 𝑗 ) under two cases based on the trial indices 𝛼 and 𝛼 ′ :

Case 1: 𝛼 ≠ 𝛼 ′ (Different trials)

Since trials are independent, the outcome of trial 𝛼 (𝐼𝛼𝑖 ) is independent of the outcome of trial 𝛼 ′ (𝐼𝛼′ 𝑗 ).

Independent random variables have zero covariance. Therefore, 𝐶𝑜𝑣(𝐼𝛼𝑖 , 𝐼𝛼′ 𝑗 ) = 0 when 𝛼 ≠ 𝛼 ′ .

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 17


MULTIVARIATE PROBABILITY DISTRIBUTION

Case 2: 𝛼 = 𝛼 ′ (Same trial)

We are considering the covariance between the indicator for category 𝑖 and the indicator for category 𝑗 within
the same trial 𝛼. Since 𝑖 ≠ 𝑗, these represent two different, mutually exclusive outcomes for the same trial, it's
impossible for a single trial to result in both outcomes simultaneously.

∴ 𝐸(𝐼𝛼𝑖 , 𝐼𝛼′𝑗 ) = 0; 𝛼 = 𝛼 ′ and 𝐶𝑜𝑣(𝐼𝛼𝑖 , 𝐼𝛼′ 𝑗 ) = 𝐸(𝐼𝛼𝑖 , 𝐼𝛼′ 𝑗 ) − 𝐸(𝐼𝛼𝑖 )𝐸(𝐼𝛼′ 𝑗 ) = 0 − 𝑝𝑖 𝑝𝑗 = −𝑝𝑖 𝑝𝑗 ; 𝑖 ≠ 𝑗

𝑛 𝑛 𝑛 𝑛

𝐶𝑜𝑣(𝑋𝑖 , 𝑋𝑗 ) = 𝐶𝑜𝑣 {∑ 𝐼𝛼𝑖 , ∑ 𝐼𝛼′ 𝑗 } = ∑ ∑ 𝐶𝑜𝑣(𝐼𝛼𝑖 , 𝐼𝛼′ 𝑗 )


𝛼=1 𝛼 ′ =1 𝛼=1 𝛼 ′ =1

𝑛 𝑛

= ∑ 𝐶𝑜𝑣(𝐼𝛼𝑖 , 𝐼𝛼𝑗 ) = ∑(−𝑝𝑖 𝑝𝑗 ) = −𝑛𝑝𝑖 𝑝𝑗 ; 𝑖 ≠ 𝑗


𝛼=1 𝛼=1

DERIVATION USING MGF:

𝑛−1
𝑘
𝜕 𝜕
𝑀𝑋 (𝑡 ) = 𝑀 (𝑡 , 𝑡 , ⋯ , 𝑡𝑘 ) = 𝑛𝑝𝑖 𝑒 𝑡𝑖 (∑ 𝑝𝑖 𝑒 𝑡𝑖 ) ; 𝑖 = 1(1)𝑘
𝜕𝑡𝑖 ~ ~ 𝜕𝑡𝑖 (𝑋1 ,𝑋2 ,⋯,𝑋𝑘 ) 1 2
𝑖=1

𝜕 𝜕
𝐸(𝑋𝑖 ) = 𝑀𝑋 (𝑡 )| = 𝑀 (𝑡 , 𝑡 , ⋯ , 𝑡𝑘 )| ; 𝑖 = 1(1)𝑘
𝜕𝑡𝑖 ~ ~ 𝑡 =0 𝜕𝑡𝑖 (𝑋1 ,𝑋2 ,⋯,𝑋𝑘 ) 1 2 𝑡 1 =𝑡2 =⋯=𝑡𝑘 =0
~ ~

𝑛−1 𝑛−1
𝑘 𝑘

= 𝑛𝑝𝑖 𝑒 𝑡𝑖 (∑ 𝑝𝑖 𝑒 𝑡𝑖 ) | = 𝑛𝑝𝑖 (∑ 𝑝𝑖 ) = 𝑛𝑝𝑖


𝑖=1 𝑖=1
𝑡1 =𝑡2=⋯=𝑡𝑘 =0

𝐸(𝑋𝑖 ) = 𝑛𝑝𝑖 ; 𝑖 = 1(1)𝑘

𝜕2 𝜕2
𝑀𝑋 (𝑡 ) = 𝑀(𝑋1 ,𝑋2 ,⋯,𝑋𝑘 ) (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑘 )
𝜕𝑡𝑖2 ~ ~ 𝜕𝑡𝑖2
𝑛−2 𝑛−1
𝑘 𝑘

= 𝑛(𝑛 − 1)𝑝𝑖2 𝑒 2𝑡𝑖 (∑ 𝑝𝑖 𝑒 𝑡𝑖 ) + 𝑛𝑝𝑖 𝑒 𝑡𝑖 (∑ 𝑝𝑖 𝑒 𝑡𝑖 ) ; 𝑖 = 1(1)𝑘


𝑖=1 𝑖=1

𝜕2 𝜕2
𝐸(𝑋𝑖2 ) = 𝑀𝑋 (𝑡 )| = 𝑀(𝑋1 ,𝑋2 ,⋯,𝑋𝑘 ) (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑘 )| ; 𝑖 = 1(1)𝑘
𝜕𝑡𝑖2 ~ ~ 𝑡 =0 𝜕𝑡𝑖2 𝑡 1 =𝑡2 =⋯=𝑡𝑘 =0
~ ~

𝑛−2 𝑛−1
𝑘 𝑘

= [𝑛(𝑛 − 1)𝑝𝑖2 𝑒 2𝑡𝑖 (∑ 𝑝𝑖 𝑒 𝑡𝑖 ) + 𝑛𝑝𝑖 𝑒 𝑡𝑖 (∑ 𝑝𝑖 𝑒 𝑡𝑖 ) ]


𝑖=1 𝑖=1
𝑡1 =𝑡2 =⋯=𝑡𝑘 =0

= 𝑛(𝑛 − 1)𝑝𝑖2 + 𝑛𝑝𝑖

𝑉𝑎𝑟(𝑋𝑖 ) = 𝐸(𝑋𝑖2 ) − 𝐸 2 (𝑋𝑖 ) = 𝑛(𝑛 − 1)𝑝𝑖2 + 𝑛𝑝𝑖 − 𝑛2 𝑝𝑖2 = 𝑛𝑝𝑖 (1 − 𝑝𝑖 ) ; 𝑖 = 1(1)𝑘

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 18


MULTIVARIATE PROBABILITY DISTRIBUTION

𝜕2 𝜕2
𝑀𝑋 (𝑡 ) = 𝑀 (𝑡 , 𝑡 , ⋯ , 𝑡𝑘 )
𝜕𝑡𝑖 𝜕𝑡𝑗 ~ ~ 𝜕𝑡𝑖 𝜕𝑡𝑗 (𝑋1 ,𝑋2 ,⋯,𝑋𝑘 ) 1 2

𝑛−2
𝑘

= 𝑛(𝑛 − 1)𝑝𝑖 𝑝𝑗 𝑒 𝑡𝑖 𝑒 𝑡𝑗 (∑ 𝑝𝑖 𝑒 𝑡𝑖 ) ; 𝑖, 𝑗 = 1(1)𝑘, 𝑖 ≠ 𝑗


𝑖=1

𝜕2 𝜕2
𝐸(𝑋𝑖 𝑋𝑗 ) = 𝑀 (𝑡 )| = 𝑀 (𝑡 , 𝑡 , ⋯ , 𝑡𝑘 )| ; 𝑖, 𝑗 = 1(1)𝑘, 𝑖 ≠ 𝑗
𝜕𝑡𝑖 𝜕𝑡𝑗 𝑋~ ~
𝑡 =0
𝜕𝑡𝑖 𝜕𝑡𝑗 (𝑋1 ,𝑋2 ,⋯,𝑋𝑘) 1 2 𝑡 1 =𝑡2 =⋯=𝑡𝑘 =0
~ ~

𝑛−2
𝑘

= [𝑛(𝑛 − 1)𝑝𝑖 𝑝𝑗 𝑒 𝑡𝑖 𝑒 𝑡𝑗 (∑ 𝑝𝑖 𝑒 𝑡𝑖 ) ] = 𝑛(𝑛 − 1)𝑝𝑖 𝑝𝑗


𝑖=1
𝑡1 =𝑡2 =⋯=𝑡𝑘 =0

𝐶𝑜𝑣(𝑋𝑖 , 𝑋𝑗 ) = 𝐸(𝑋𝑖 𝑋𝑗 ) − 𝐸(𝑋𝑖 )𝐸(𝑋𝑗 ) = 𝑛(𝑛 − 1)𝑝𝑖 𝑝𝑗 − 𝑛2 𝑝𝑖 𝑝𝑗 = −𝑛𝑝𝑖 𝑝𝑗 ; 𝑖, 𝑗 = 1(1)𝑘, 𝑖 ≠ 𝑗

𝐶𝑜𝑣(𝑋𝑖 , 𝑋𝑗 ) −𝑛𝑝𝑖 𝑝𝑗
𝜌(𝑋𝑖 , 𝑋𝑗 ) = 𝐶𝑜𝑟𝑟(𝑋𝑖 , 𝑋𝑗 ) = =
√𝑉𝑎𝑟(𝑋𝑖 )√𝑉𝑎𝑟( 𝑋𝑗 ) √𝑛𝑝𝑖 (1 − 𝑝𝑖 )√𝑛𝑝𝑗 (1 − 𝑝𝑗 )

𝑝𝑖 𝑝𝑗
= −√ ; 𝑖, 𝑗 = 1(1)𝑘, 𝑖 ≠ 𝑗
(1 − 𝑝𝑖 )(1 − 𝑝𝑗 )

The variance-covariance matrix or dispersion matrix of 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘 )𝑇 is given by


~

𝑛𝑝1 (1 − 𝑝1 ) −𝑛𝑝1 𝑝2 ⋯ −𝑛𝑝1 𝑝𝑘


−𝑛𝑝1 𝑝2 𝑛𝑝2 (1 − 𝑝2 ) ⋯ −𝑛𝑝2 𝑝𝑘
𝚺𝑘×𝑘 = 𝐷𝑖𝑠𝑝 (𝑋) = 𝐷(𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘 ) = ( )
~ ⋮ ⋮ ⋱ ⋮
−𝑛𝑝1 𝑝𝑘 −𝑛𝑝2 𝑝𝑘 ⋯ 𝑛𝑝𝑘 (1 − 𝑝𝑘 ) 𝑘×𝑘

The correlation matrix of 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘 )𝑇 is given by


~

𝑝1 𝑝2 𝑝1 𝑝𝑘
1 −√ ⋯ −√
(1 − 𝑝1 )(1 − 𝑝2 ) (1 − 𝑝1 )(1 − 𝑝𝑘 )

𝑝2 𝑝1 𝑝2 𝑝𝑘
𝑹𝑘×𝑘 = 𝐶𝑜𝑟𝑟 (𝑋) = −√ 1 ⋯ −√
~
(1 − 𝑝2 )(1 − 𝑝1 ) (1 − 𝑝2 )(1 − 𝑝𝑘 )
⋮ ⋮ ⋱ ⋮
𝑝𝑘−1 𝑝1 𝑝𝑘 𝑝2
−√ −√ ⋯ 1
(1 − 𝑝𝑘 )(1 − 𝑝1 ) (1 − 𝑝𝑘 )(1 − 𝑝2 )
( )𝑘×𝑘

INTUITION FOR NEGATIVE COVARIANCE:

The covariance is negative because the total number of trials 𝑛 is fixed (∑𝑘𝑖=1 𝑋𝑖 = 𝑛). If the count for one
category, 𝑋𝑖 , is higher than expected, the counts for the other categories (𝑋𝑗 ; for 𝑗 ≠ 𝑖) must tend to be lower
to compensate, and vice versa. This inverse relationship is captured by the negative covariance.

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 19


MULTIVARIATE PROBABILITY DISTRIBUTION

MARGINAL DISTRIBUTION

The marginal MGF of 𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑠 )𝑇 is given by


~

𝑛
𝑘

𝑀𝑋 (1) (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑠 ) = 𝑀(𝑋1 ,𝑋2 ,⋯,𝑋𝑘 ) (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑠 , 0, 0, ⋯ , 0) = [(∑ 𝑝𝑖 𝑒 𝑡𝑖 ) ]


~
𝑖=1
𝑡𝑗 =0; 𝑗=𝑠+1(1)𝑘

𝑛
𝑠 𝑘

= (∑ 𝑝𝑖 𝑒 𝑡𝑖 + ∑ 𝑝𝑖 )
𝑖=1 𝑖=𝑠+1

𝑠 𝑠 𝑛
𝑡𝑖
= {∑ 𝑝𝑖 𝑒 + (1 − ∑ 𝑝𝑖 )}
𝑖=1 𝑖=1

𝑠 𝑠

← MGF of 𝑀𝑢𝑙𝑡𝑖𝑛𝑜𝑚𝑖𝑎𝑙(𝑛, 𝑝1 , 𝑝2 , ⋯ , 𝑝𝑠 ), where ∑ 𝑥𝑖 ≤ 𝑛 , 0 < 𝑝𝑖 < 1 and ∑ 𝑝𝑖 < 1


𝑖=1 𝑖=1

By uniqueness property of MGF, we get


𝑠 𝑠
(1)
𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑠 )𝑇 ~ 𝑀𝑢𝑙𝑡𝑖𝑛𝑜𝑚𝑖𝑎𝑙(𝑛, 𝑝1 , 𝑝2 , ⋯ , 𝑝𝑠 ), where ∑ 𝑥𝑖 ≤ 𝑛 , 0 < 𝑝𝑖 < 1 and ∑ 𝑝𝑖 < 1.
~
𝑖=1 𝑖=1

The marginal MGF of 𝑋𝑖 ; 𝑖 = 1(1)𝑘 is given by


𝑛
𝑛
𝑘 𝑘
𝑡𝑖 𝑡𝑖
𝑀𝑋𝑖 (𝑡𝑖 ) = 𝑀(𝑋1 ,𝑋2 ,⋯,𝑋𝑘 ) (0, 0, ⋯ , 0, 𝑡𝑖 , 0, ⋯ , 0) = [(∑ 𝑝𝑖 𝑒 ) ] = 𝑝𝑖 𝑒 + ∑ 𝑝𝑗
𝑖=1 𝑗=1
𝑡𝑗 =0; 𝑗≠𝑖
𝑗≠𝑖
( )

= (𝑝1 + ⋯ + 𝑝𝑖−1 + 𝑝𝑖 𝑒 𝑡𝑖 + 𝑝𝑖+1 + ⋯ + 𝑝𝑘 )𝑛


𝑛
= (𝑝𝑖 𝑒 𝑡𝑖 + (1 − 𝑝𝑖 )) ; 𝑖 = 1(1)𝑘

← MGF of 𝐵𝑖𝑛(𝑛, 𝑝𝑖 )

By uniqueness property of MGF, we get

𝑋𝑖 ~ 𝐵𝑖𝑛(𝑛, 𝑝𝑖 ) ; 𝑖 = 1(1)𝑘.

ALTERNATIVE

The marginal p.m.f. of 𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑠 )𝑇 is given by


~

𝑝𝑋 (1) (𝑥 (1) ) = 𝑝(𝑋1 ,𝑋2 ,⋯,𝑋𝑠 ) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑠 ) = ∑ ∑ ⋯ ∑ 𝑝𝑋 (𝑥 )


~ ~ ~ ~
𝑥𝑗 =0(1)𝑛, 𝑗=𝑠+1(1)𝑘, ∑𝑘
𝑗=1 𝑥𝑗 =𝑛

𝑛! 𝑥 𝑥 𝑥
= ∑ ∑ ⋯ ∑ ∙ 𝑝 1 𝑝 2 ⋯ 𝑝𝑘 𝑘
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑘 ! 1 2
𝑥𝑗 =0(1)𝑛, 𝑗=𝑠+1(1)𝑘, ∑𝑘
𝑗=1 𝑥𝑗 =𝑛

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 20


MULTIVARIATE PROBABILITY DISTRIBUTION

𝑛! 𝑥 𝑥 𝑥 (𝑛 − ∑𝑠𝑖=1 𝑥𝑖 )! 𝑥 𝑥 𝑥
= ∙ 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑠 𝑠 ∙ ∑ ∑ ⋯ ∑ ∙ 𝑝 𝑠+1 𝑝 𝑠+2 ⋯ 𝑝𝑘 𝑘
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! (𝑛 − ∑𝑠𝑖=1 𝑥𝑖 )! 𝑥𝑠+1 ! 𝑥𝑠+2 ! ⋯ 𝑥𝑘 ! 𝑠+1 𝑠+2
𝑥𝑗 =0(1)𝑛, 𝑗=𝑠+1(1)𝑘, ∑𝑘
𝑗=1 𝑥𝑗 =𝑛

𝑛! 𝑥 𝑥 𝑥 𝑠
= ∙ 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑠 𝑠 ∙ (𝑝𝑠+1 + 𝑝𝑠+2 + ⋯ + 𝑝𝑘 )(𝑛−∑𝑖=1 𝑥𝑖)
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! (𝑛 − ∑𝑠𝑖=1 𝑥𝑖 )!

𝑠 (𝑛−∑𝑠𝑖=1 𝑥𝑖 )
𝑛! 𝑥 𝑥 𝑥
= ∙ 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑠 𝑠 ∙ (1 − ∑ 𝑝𝑖 )
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! (𝑛 − ∑𝑠𝑖=1 𝑥𝑖 )!
𝑖=1

𝑠 𝑠

← 𝑝. 𝑚. 𝑓. of 𝑀𝑢𝑙𝑡𝑖𝑛𝑜𝑚𝑖𝑎𝑙(𝑛, 𝑝1 , 𝑝2 , ⋯ , 𝑝𝑠 ), where ∑ 𝑥𝑖 ≤ 𝑛 , 0 < 𝑝𝑖 < 1 and ∑ 𝑝𝑖 < 1


𝑖=1 𝑖=1

Thus, 𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑠 )𝑇 ~ 𝑀𝑢𝑙𝑡𝑖𝑛𝑜𝑚𝑖𝑎𝑙(𝑛, 𝑝1 , 𝑝2 , ⋯ , 𝑝𝑠 );


~
𝑠 𝑠

where ∑ 𝑥𝑖 ≤ 𝑛 , 0 < 𝑝𝑖 < 1 and ∑ 𝑝𝑖 < 1.


𝑖=1 𝑖=1

The marginal p.m.f. of 𝑋𝑖 ; 𝑖 = 1(1)𝑘 is given by

𝑝𝑋𝑖 (𝑥𝑖 ) = ∑ ∑ ⋯ ∑ 𝑝𝑋 (𝑥 )
~ ~
𝑥𝑗 =0(1)𝑛, 𝑥𝑗 ≠𝑥𝑖 , ∑𝑘
𝑗=1 𝑥𝑗 =𝑛

𝑛! 𝑥 𝑥 𝑥
= ∑ ∑ ⋯ ∑ ∙ 𝑝 1 𝑝 2 ⋯ 𝑝𝑘 𝑘 ; 𝑖 = 1(1)𝑘
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑘 ! 1 2
𝑥𝑗 =0(1)𝑛, 𝑥𝑗 ≠𝑥𝑖 , ∑𝑘
𝑗=1 𝑥𝑗 =𝑛

𝑛! 𝑥 (𝑛 − 𝑥𝑖 )! 𝑥 𝑥𝑖−1 𝑥𝑖+1 𝑥
= ∙ 𝑝𝑖 𝑖 ∙ ∑ ∑ ⋯ ∑ ∙ 𝑝1 1 ⋯ 𝑝𝑖−1 𝑝𝑖+1 ⋯ 𝑝𝑘 𝑘
𝑥𝑖 ! (𝑛 − 𝑥𝑖 )! 𝑥1 ! ⋯ 𝑥𝑖−1 ! 𝑥𝑖+1 ! ⋯ 𝑥𝑘 !
𝑥𝑗 =0(1)(𝑛−𝑥𝑖 ), 𝑥𝑗 ≠𝑥𝑖 , ∑𝑘
𝑗=1,𝑗≠𝑖 𝑥𝑗 =𝑛−𝑥𝑖

𝑛! 𝑥
= ∙ 𝑝 𝑖 ∙ (𝑝1 + ⋯ + 𝑝𝑖−1 + 𝑝𝑖+1 + ⋯ + 𝑝𝑘 )(𝑛−𝑥𝑖)
𝑥𝑖 ! (𝑛 − 𝑥𝑖 )! 𝑖
𝑘
𝑛! 𝑥
= ∙ 𝑝 𝑖 ∙ (1 − 𝑝𝑖 )(𝑛−𝑥𝑖) [∵ ∑ 𝑝𝑖 = 1]
𝑥𝑖 ! (𝑛 − 𝑥𝑖 )! 𝑖
𝑖=1

← 𝑝. 𝑚. 𝑓. of 𝐵𝑖𝑛(𝑛, 𝑝𝑖 ) ; 𝑖 = 1(1)𝑘

Hence, 𝑋𝑖 ~ 𝐵𝑖𝑛(𝑛, 𝑝𝑖 ) ; 𝑖 = 1(1)𝑘.

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 21


MULTIVARIATE PROBABILITY DISTRIBUTION

CONDITIONAL DISTRIBUTION

The conditional distribution of 𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑠 )𝑇 given that 𝑋 (2) = 𝑥 (2) , where 𝑋 (2) =
~ ~ ~ ~

(𝑋𝑠+1 , 𝑋𝑠+2 , ⋯ , 𝑋𝑘 )𝑇 [ i. e. given that (𝑋𝑠+1 = 𝑥𝑠+1 , 𝑋𝑠+2 = 𝑥𝑠+2 , ⋯ , 𝑋𝑘 = 𝑥𝑘 )] is given by

𝑝 (1) (2) (𝑥 (1) ) = 𝑝((𝑋 , 𝑋 , ⋯ , 𝑋 )|(𝑥 , 𝑥 , ⋯ , 𝑥 )) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑠 )


(𝑋 |𝑋 = 𝑥 (2) ) ~ 1 2 𝑠 𝑠+1 𝑠+2 𝑘
~ ~ ~

𝑝𝑋 (𝑥 ) 𝑝(𝑋1 ,𝑋2 ,⋯,𝑋𝑘) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑘 )


~ ~
= =
𝑝𝑋 (2) (𝑥 (2) ) 𝑝(𝑋𝑠+1 ,𝑋𝑠+2 ,⋯,𝑋𝑘) (𝑥𝑠+1 , 𝑥𝑠+2 , ⋯ , 𝑥𝑘 )
~ ~

𝑛! 𝑥1 𝑥2 𝑥𝑘
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑘 ! ∙ 𝑝1 𝑝2 ⋯ 𝑝𝑘
=
𝑛! 𝑥𝑠+1 𝑥𝑠+2 𝑥 (𝑛−∑𝑘
𝑖=𝑠+1 𝑥𝑖 )
𝑘 ∙ 𝑝𝑠+1 𝑝𝑠+2 ⋯ 𝑝𝑘 𝑘 (1 − ∑𝑘𝑖=𝑠+1 𝑝𝑖 )
𝑥𝑠+1 ! 𝑥𝑠+2 ! ⋯ 𝑥𝑘 ! (𝑛 − ∑𝑖=𝑠+1 𝑥𝑖 )!

𝑥 𝑥 𝑥 𝑘 𝑠 𝑘
(𝑛 − ∑𝑘𝑖=𝑠+1 𝑥𝑖 )! 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑠 𝑠
= ∙ ∑𝑠 𝑥
[∵ ∑ 𝑥𝑖 = 𝑛, ∑ 𝑥𝑖 = 𝑛 − ∑ 𝑥𝑖 ]
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! (1 − ∑𝑘 𝑝 ) 𝑖=1 𝑖
𝑖=𝑠+1 𝑖 𝑖=1 𝑖=1 𝑖=𝑠+1

𝑥1 𝑥2 𝑥𝑠
(𝑛 − ∑𝑘𝑖=𝑠+1 𝑥𝑖 )! 𝑝1 𝑝2 𝑝𝑠
= ∙{ } { } ⋯{ }
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! (1 − ∑𝑘𝑖=𝑠+1 𝑝𝑖 ) (1 − ∑𝑘𝑖=𝑠+1 𝑝𝑖 ) (1 − ∑𝑘𝑖=𝑠+1 𝑝𝑖 )
𝑥1 𝑥2 𝑥𝑠
(𝑛 − ∑𝑘𝑖=𝑠+1 𝑥𝑖 )! 𝑝1 𝑝2 𝑝𝑠
= ∙{ 𝑠 } { } ⋯{ }
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! ∑𝑖=1 𝑝𝑖 ∑𝑠𝑖=1 𝑝𝑖 ∑𝑠𝑖=1 𝑝𝑖

(𝑛 − ∑𝑘𝑖=𝑠+1 𝑥𝑖 )! 𝑥1 𝑥2 𝑥
= ∙ 𝑄1 𝑄2 ⋯ 𝑄𝑠 𝑠
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 !
𝑘 𝑘 𝑠
𝑝𝑖
← 𝑝. 𝑚. 𝑓. of 𝑀𝑢𝑙𝑡𝑖𝑛𝑜𝑚𝑖𝑎𝑙 [{𝑛 − ∑ 𝑥𝑖 } , 𝑄1 , 𝑄2 , ⋯ , 𝑄𝑠 ] , where ∑ 𝑥𝑖 = 𝑛, 𝑄𝑖 = and ∑ 𝑄𝑖 = 1
∑𝑠𝑖=1 𝑝𝑖
𝑖=𝑠+1 𝑖=1 𝑖=1

The conditional distribution of 𝑋 (1) = (𝑋1 , 𝑋2 )𝑇 given that 𝑋3 = 𝑥3 , is given by


~

𝑛! 𝑥1 𝑥2 𝑥 3
𝑝(𝑋1 ,𝑋2 ,𝑋3 ) (𝑥1 , 𝑥2 , 𝑥3 ) 𝑥1 ! 𝑥2 ! 𝑥3 ! ∙ 𝑝1 𝑝2 𝑝3
𝑝(𝑋1 ,𝑋2 )| 𝑋3 =𝑥3 (𝑥1 , 𝑥2 ) = =
𝑝𝑋3 (𝑥3 ) 𝑛! 𝑥
∙ 𝑝 3 (1 − 𝑝3 )(𝑛−𝑥3 )
𝑥3 ! (𝑛 − 𝑥3 )! 3

𝑥 𝑥 3
(𝑛 − 𝑥3 )! 𝑝1 1 𝑝2 2
= ∙ [∵ ∑ 𝑥𝑖 = 𝑛, 𝑥1 + 𝑥2 = 𝑛 − 𝑥3 ]
𝑥1 ! 𝑥2 ! (1 − 𝑝3 )𝑥1 +𝑥2
𝑖=1

(𝑛 − 𝑥3 )! 𝑝1 𝑥1 𝑝2 𝑥2
= ∙( ) ( )
𝑥1 ! 𝑥2 ! 1 − 𝑝3 1 − 𝑝3

The conditional distribution of 𝑋1 given that 𝑋 (2) = 𝑥 (2) , where 𝑋 (2) = (𝑋2 , 𝑋3 , ⋯ , 𝑋𝑘 )𝑇
~ ~ ~

[ i. e. The conditional distribution of 𝑋1 given that (𝑋2 = 𝑥2 , 𝑋3 = 𝑥3 , ⋯ , 𝑋𝑘 = 𝑥𝑘 )] is given by

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 22


MULTIVARIATE PROBABILITY DISTRIBUTION

𝑝𝑋 (𝑥 ) 𝑝(𝑋1 ,𝑋2 ,⋯,𝑋𝑘 ) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑘 )


~ ~
𝑝 (2) (𝑥 ) = 𝑝(𝑋 |(𝑥 , 𝑥 , ⋯ , 𝑥 )) (𝑥1 ) = =
(𝑋1 |𝑋 = 𝑥 (2)) 1 1 2 3 𝑘 𝑝(𝑋 (2) ) (𝑥 (2) ) 𝑝(𝑋2 ,𝑋3 ,⋯,𝑋𝑘 ) (𝑥2 , 𝑥3 , ⋯ , 𝑥𝑘 )
~ ~
~ ~

𝑛! 𝑥1 𝑥2 𝑥𝑘
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑘 ! ∙ 𝑝1 𝑝2 ⋯ 𝑝𝑘
=
𝑛! 𝑥2 𝑥3 𝑥𝑘 𝑘 (𝑛−∑𝑘
𝑖=2 𝑥𝑖 )
𝑘 ∙ 𝑝2 𝑝3 ⋯ 𝑝𝑘 (1 − ∑ 𝑖=2 𝑝𝑖 )
𝑥2 ! 𝑥3 ! ⋯ 𝑥𝑘 ! (𝑛 − ∑𝑖=2 𝑥𝑖 )!

𝑥 𝑘 𝑘
(𝑛 − ∑𝑘𝑖=2 𝑥𝑖 )! 𝑝1 1
= ∙ 𝑥1 [∵ ∑ 𝑥𝑖 = 𝑛, 𝑥1 = 𝑛 − ∑ 𝑥𝑖 ]
𝑥1 ! (1 − ∑𝑘𝑖=2 𝑝𝑖 ) 𝑖=1 𝑖=2

𝑘
(𝑛 − ∑𝑘𝑖=2 𝑥𝑖 )!
= [∵ ∑ 𝑝𝑖 = 1]
𝑥1 !
𝑖=1

SINGULARITY OF VARIANCE-COVARIANCE MATRIX (DISPERSION MATRIX):

The variance-covariance matrix or dispersion matrix of 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘 )𝑇 is given by


~

𝑛𝑝1 (1 − 𝑝1 ) −𝑛𝑝1 𝑝2 ⋯ −𝑛𝑝1 𝑝𝑘


−𝑛𝑝1 𝑝2 𝑛𝑝2 (1 − 𝑝2 ) ⋯ −𝑛𝑝2 𝑝𝑘
𝚺𝑘×𝑘 = 𝐷𝑖𝑠𝑝 (𝑋) = 𝐷(𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘 ) = ( )
~ ⋮ ⋮ ⋱ ⋮
−𝑛𝑝1 𝑝𝑘 −𝑛𝑝2 𝑝𝑘 ⋯ 𝑛𝑝𝑘 (1 − 𝑝𝑘 ) 𝑘×𝑘

The determinant of variance-covariance matrix or dispersion matrix is given by

𝑛𝑝1 (1 − 𝑝1 ) −𝑛𝑝1 𝑝2 ⋯ −𝑛𝑝1 𝑝𝑘


−𝑛𝑝1 𝑝2 𝑛𝑝2 (1 − 𝑝2 ) ⋯ −𝑛𝑝2 𝑝𝑘
|𝚺𝑘×𝑘 | = |𝐷𝑖𝑠𝑝 (𝑋)| = | |
~ ⋮ ⋮ ⋱ ⋮
−𝑛𝑝1 𝑝𝑘 −𝑛𝑝2 𝑝𝑘 ⋯ 𝑛𝑝𝑘 (1 − 𝑝𝑘 )

| 𝑛𝑝1 (1 − ∑ 𝑝𝑖 ) −𝑛𝑝1 𝑝2 ⋯ −𝑛𝑝1 𝑝𝑘 |


𝑖=1
𝑘

= |𝑛𝑝2 (1 − ∑ 𝑝𝑖 ) 𝑛𝑝2 (1 − 𝑝2 ) ⋯ −𝑛𝑝2 𝑝𝑘 |


𝑖=1
⋮ ⋮ ⋱ ⋮
𝑘
|𝑛𝑝 (1 − ∑ 𝑝 ) −𝑛𝑝2 𝑝𝑘 ⋯ 𝑛𝑝𝑘 (1 − 𝑝𝑘 )|
𝑘 𝑖
𝑖=1

0 −𝑛𝑝1 𝑝2 ⋯ −𝑛𝑝1 𝑝𝑘
0 𝑛𝑝2 (1 − 𝑝2 ) ⋯ −𝑛𝑝2 𝑝𝑘
=| |=0
⋮ ⋮ ⋱ ⋮
0 −𝑛𝑝2 𝑝𝑘 ⋯ 𝑛𝑝𝑘 (1 − 𝑝𝑘 )

Since |𝚺𝑘×𝑘 | = 0, the probability distribution of (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘 ) is singular and 𝑅𝑎𝑛𝑘(𝚺𝑘×𝑘 ) = (𝑘 − 1).

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 23


MULTIVARIATE PROBABILITY DISTRIBUTION

MULTINOMIAL DISTRIBUTION OF (𝑿𝟏 , 𝑿𝟐 , ⋯ , 𝑿𝒌−𝟏 )𝑻 :

To avoid the difficulties associated with singularity, we consider the joint probability distribution of (𝑘 − 1)
random variables.
Let 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘−1 )𝑇 ~ 𝑀𝑢𝑙𝑡𝑖𝑛𝑜𝑚𝑖𝑎𝑙(𝑛, 𝑝1 , 𝑝2 , ⋯ , 𝑝𝑘−1 );
~
𝑘−1 𝑘−1

where 0 < 𝑝𝑖 < 1, ∑ 𝑝𝑖 < 1 and ∑ 𝑥𝑖 ≤ 𝑛, 𝑖 = 1(1)(𝑘 − 1).


𝑖=1 𝑖=1

The joint p.m.f. of (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘−1 ) is given by

𝑝𝑋 (𝑥 ) = 𝑝(𝑋1 ,𝑋2 ,⋯,𝑋𝑘−1 ) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑘−1 )


~ ~

(𝑛−∑𝑘−1
𝑖=1 𝑥𝑖 )
𝑘−1 𝑘−1
𝑛! 𝑥 𝑥 𝑥𝑘−1
∙ 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑘−1 (1 − ∑ 𝑝𝑖 ) ; 𝑥𝑖 = 0(1)𝑛, ∑ 𝑥𝑖 ≤ 𝑛,
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑘−1 ! (𝑛 − ∑𝑘−1
𝑖=1 𝑥𝑖 )! 𝑖=1 𝑖=1
= 𝑘−1

0 < 𝑝𝑖 < 1, ∑ 𝑝𝑖 < 1, 𝑖 = 1(1)𝑘 − 1


𝑖=1
{ 0; otherwise

(𝑛−∑𝑘−1
𝑖=1 𝑥𝑖 )
𝑘−1 𝑘−1 𝑘−1
𝑛! 𝑥
∙ ∏(𝑝𝑖 𝑖 ) (1 − ∑ 𝑝𝑖 ) ; 𝑥𝑖 = 0(1)𝑛, ∑ 𝑥𝑖 ≤ 𝑛,
∏𝑘−1 𝑘−1
𝑖=1 (𝑥𝑖 !) (𝑛 − ∑𝑖=1 𝑥𝑖 )! 𝑖=1 𝑖=1 𝑖=1
= 𝑘−1

0 < 𝑝𝑖 < 1, ∑ 𝑝𝑖 < 1, 𝑖 = 1(1)𝑘 − 1


𝑖=1
{ 0; otherwise
𝑛 𝑛 𝑛

∑ ∑ ⋯ ∑ 𝑝𝑋 (𝑥 )
~ ~
𝑥1 =0 𝑥2 =0 𝑥𝑘−1 =0
∑𝑘−1
𝑖=1 𝑥𝑖 ≤𝑛

(𝑛−∑𝑘−1
𝑖=1 𝑥𝑖 )
𝑛 𝑛 𝑛 𝑘−1
𝑛! 𝑥 𝑥 𝑥
= ∑ ∑ ⋯ ∑ ∙ 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑘−1
𝑘−1
(1 − ∑ 𝑝𝑖 )
𝑥 ! 𝑥 ! ⋯ 𝑥𝑘−1 ! (𝑛 − ∑𝑘−1
𝑖=1 𝑥𝑖 )!
𝑥1 =0 𝑥2 =0 𝑥𝑘−1 =0 1 2 𝑖=1
∑𝑘−1
𝑖=1 𝑥𝑖 ≤𝑛
𝑛
𝑘−1

= [(1 − ∑ 𝑝𝑖 ) + 𝑝1 + 𝑝2 + ⋯ + 𝑝𝑘−1 ] = (1)𝑛 = 1


𝑖=1

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 24


MULTIVARIATE PROBABILITY DISTRIBUTION

MOMENT GENERATING FUNCTION:

The moment generating function (𝑚. 𝑔. 𝑓. ) of 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘−1 )𝑇 or the 𝑚. 𝑔. 𝑓. of the joint distribution
~

of (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘−1 ) for all 𝑡 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑘−1 )𝑇 ∈ ℛ𝑘−1 is given by


~

𝑡 𝑇𝑋 𝑘−1
𝑀𝑋 (𝑡 ) = 𝑀(𝑋1 ,𝑋2 ,⋯,𝑋𝑘−1 ) (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑘−1 ) = 𝐸 (𝑒 ~ ~ ) = 𝐸(𝑒 𝑡1 𝑋1 +𝑡2 𝑋2 +⋯+𝑡𝑘−1 𝑋𝑘−1 ) = 𝐸 (𝑒 ∑𝑖=1 𝑡𝑖𝑥𝑖 )
~ ~

𝑛 𝑛 𝑛 𝑘−1 𝑠 (𝑛−∑𝑘−1
𝑖=1 𝑥𝑖 )
∑𝑘−1
𝑛! 𝑥
= ∑ ∑ ⋯ ∑ (𝑒 𝑖=1 𝑡𝑖 𝑥𝑖 ) ∙ ∏(𝑝𝑖 𝑖 ) (1 − ∑ 𝑝𝑖 )
𝑥2 =0
∏𝑘−1 𝑘−1
𝑖=1 (𝑥𝑖 !) (𝑛 − ∑𝑖=1 𝑥𝑖 )!
𝑥1 =0 𝑥𝑘−1 =0 𝑖=1 𝑖=1
∑𝑘−1
𝑖=1 𝑥𝑖 ≤𝑛

𝑛 𝑛 𝑛 𝑘−1 𝑠 (𝑛−∑𝑘−1
𝑖=1 𝑥𝑖 )
𝑛!
= ∑ ∑ ⋯ ∑ ∙ ∏(𝑝𝑖 𝑒 𝑡𝑖 )𝑥𝑖 (1 − ∑ 𝑝𝑖 )
𝑥2 =0
∏𝑘−1 𝑘−1
𝑖=1 (𝑥𝑖 !) (𝑛 − ∑𝑖=1 𝑥𝑖 )!
𝑥1 =0 𝑥𝑘−1 =0 𝑖=1 𝑖=1
∑𝑘−1
𝑖=1 𝑥𝑖 ≤𝑛

𝑛
𝑘−1

= [(1 − ∑ 𝑝𝑖 ) + 𝑝1 𝑒 𝑡1 + 𝑝2 𝑒 𝑡2 + ⋯ + 𝑝𝑘−1 𝑒 𝑡𝑘−1 ]


𝑖=1

𝑛
𝑘−1 𝑘−1

= {∑ 𝑝𝑖 𝑒 𝑡𝑖 + (1 − ∑ 𝑝𝑖 )}
𝑖=1 𝑖=1

ALTERNATIVE:

𝑀(𝑋1 ,𝑋2 ,⋯,𝑋𝑘−1 ) (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑘−1 ) = 𝑀(𝑋1 ,𝑋2 ,⋯,𝑋𝑘−1 ) (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑘−1 , 0)
𝑛 𝑛 𝑛
𝑘 𝑘−1 𝑘−1 𝑘−1 𝑘
𝑡𝑖 𝑡𝑖 𝑡𝑖
= [(∑ 𝑝𝑖 𝑒 ) ] = (∑ 𝑝𝑖 𝑒 + 𝑝𝑘 ) = {∑ 𝑝𝑖 𝑒 + (1 − ∑ 𝑝𝑖 )} [∵ ∑ 𝑝𝑖 = 1]
𝑖=1 𝑖=1 𝑖=1 𝑖=1 𝑖=1
𝑡𝑘−1 =0

MOMENTS:
𝑛−1
𝑘−1 𝑘−1
𝜕 𝜕
𝑀𝑋 (𝑡 ) = 𝑀 (𝑡 , 𝑡 , ⋯ , 𝑡𝑘−1 ) = 𝑛𝑝𝑖 𝑒 𝑡𝑖 {∑ 𝑝𝑖 𝑒 𝑡𝑖 + (1 − ∑ 𝑝𝑖 )} ; 𝑖 = 1(1)(𝑘 − 1)
𝜕𝑡𝑖 ~ ~ 𝜕𝑡𝑖 (𝑋1 ,𝑋2 ,⋯,𝑋𝑘−1 ) 1 2
𝑖=1 𝑖=1

𝜕 𝜕
𝐸(𝑋𝑖 ) = 𝑀𝑋 (𝑡 )| = 𝑀(𝑋1 ,𝑋2 ,⋯,𝑋𝑘−1 ) (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑘−1 )| ; 𝑖 = 1(1)(𝑘 − 1)
𝜕𝑡𝑖 ~ ~ 𝑡 =0 𝜕𝑡𝑖 𝑡
~ ~ 1 =𝑡2 =⋯=𝑡𝑘−1 =0

𝑛−1
𝑘−1 𝑘−1

= 𝑛𝑝𝑖 𝑒 𝑡𝑖 {∑ 𝑝𝑖 𝑒 𝑡𝑖 + (1 − ∑ 𝑝𝑖 )} | = 𝑛𝑝𝑖 (1)𝑛−1 = 𝑛𝑝𝑖


𝑖=1 𝑖=1
𝑡1 =𝑡2 =⋯=𝑡𝑘−1 =0

𝐸(𝑋𝑖 ) = 𝑛𝑝𝑖 ; 𝑖 = 1(1)(𝑘 − 1)

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 25


MULTIVARIATE PROBABILITY DISTRIBUTION

𝜕2 𝜕2
𝑀 (𝑡 ) = 𝑀(𝑋1 ,𝑋2 ,⋯,𝑋𝑘−1 ) (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑘−1 )
𝜕𝑡𝑖2 𝑋~ ~ 𝜕𝑡𝑖2

𝑛−2 𝑛−1
𝑘−1 𝑘−1 𝑘−1 𝑘−1

= 𝑛(𝑛 − 1)𝑝𝑖2 𝑒 2𝑡𝑖 {∑ 𝑝𝑖 𝑒 𝑡𝑖 + (1 − ∑ 𝑝𝑖 )} + 𝑛𝑝𝑖 𝑒 𝑡𝑖 {∑ 𝑝𝑖 𝑒 𝑡𝑖 + (1 − ∑ 𝑝𝑖 )} ; 𝑖 = 1(1)(𝑘 − 1)


𝑖=1 𝑖=1 𝑖=1 𝑖=1

𝜕2 𝜕2
𝐸(𝑋𝑖2 ) = 𝑀 (𝑡 )| = 𝑀(𝑋1 ,𝑋2 ,⋯,𝑋𝑘−1 ) (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑘−1 )|
𝜕𝑡𝑖2 𝑋~ ~ 𝑡 =0 𝜕𝑡𝑖2 𝑡 1 =𝑡2 =⋯=𝑡𝑘−1 =0
~ ~

𝑛−2 𝑛−1
𝑘−1 𝑘−1 𝑘−1 𝑘−1

= [𝑛(𝑛 − 1)𝑝𝑖2 𝑒 2𝑡𝑖 {∑ 𝑝𝑖 𝑒 𝑡𝑖 + (1 − ∑ 𝑝𝑖 )} + 𝑛𝑝𝑖 𝑒 𝑡𝑖 {∑ 𝑝𝑖 𝑒 𝑡𝑖 + (1 − ∑ 𝑝𝑖 )} ]


𝑖=1 𝑖=1 𝑖=1 𝑖=1
𝑡1 =𝑡2 =⋯=𝑡𝑘−1 =0

= 𝑛(𝑛 − 1)𝑝𝑖2 + 𝑛𝑝𝑖 ; 𝑖 = 1(1)(𝑘 − 1)

𝑉𝑎𝑟(𝑋𝑖 ) = 𝐸(𝑋𝑖2 ) − 𝐸 2 (𝑋𝑖 ) = 𝑛(𝑛 − 1)𝑝𝑖2 + 𝑛𝑝𝑖 − 𝑛2 𝑝𝑖2 = 𝑛𝑝𝑖 (1 − 𝑝𝑖 ) ; 𝑖 = 1(1)(𝑘 − 1)

𝜕2 𝜕2
𝑀𝑋 (𝑡 ) = 𝑀 (𝑡 , 𝑡 , ⋯ , 𝑡𝑘−1 )
𝜕𝑡𝑖 𝜕𝑡𝑗 ~ ~ 𝜕𝑡𝑖 𝜕𝑡𝑗 (𝑋1 ,𝑋2 ,⋯,𝑋𝑘−1 ) 1 2

𝑛−2
𝑘−1 𝑘−1

= 𝑛(𝑛 − 1)𝑝𝑖 𝑝𝑗 𝑒 𝑡𝑖 𝑒 𝑡𝑗 {∑ 𝑝𝑖 𝑒 𝑡𝑖 + (1 − ∑ 𝑝𝑖 )} ; 𝑖, 𝑗 = 1(1)(𝑘 − 1), 𝑖 ≠ 𝑗


𝑖=1 𝑖=1

𝜕2 𝜕2
𝐸(𝑋𝑖 𝑋𝑗 ) = 𝑀 (𝑡 )| = 𝑀(𝑋1 ,𝑋2 ,⋯,𝑋𝑘−1 ) (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑘−1 )|
𝜕𝑡𝑖 𝜕𝑡𝑗 𝑋~ ~ 𝜕𝑡𝑖 𝜕𝑡𝑗
𝑡 =0 𝑡 1 =𝑡2 =⋯=𝑡𝑘 =0
~ ~

𝑛−2
𝑘−1 𝑘−1

= [𝑛(𝑛 − 1)𝑝𝑖 𝑝𝑗 𝑒 𝑡𝑖 𝑒 𝑡𝑗 {∑ 𝑝𝑖 𝑒 𝑡𝑖 + (1 − ∑ 𝑝𝑖 )} ]
𝑖=1 𝑖=1
𝑡1 =𝑡2=⋯=𝑡𝑘−1 =0

= 𝑛(𝑛 − 1)𝑝𝑖 𝑝𝑗 ; 𝑖, 𝑗 = 1(1)(𝑘 − 1), 𝑖 ≠ 𝑗

𝐶𝑜𝑣(𝑋𝑖 , 𝑋𝑗 ) = 𝐸(𝑋𝑖 𝑋𝑗 ) − 𝐸(𝑋𝑖 )𝐸(𝑋𝑗 ) = 𝑛(𝑛 − 1)𝑝𝑖 𝑝𝑗 − 𝑛2 𝑝𝑖 𝑝𝑗 = −𝑛𝑝𝑖 𝑝𝑗 ; 𝑖, 𝑗 = 1(1)(𝑘 − 1), 𝑖 ≠ 𝑗

𝐶𝑜𝑣(𝑋𝑖 , 𝑋𝑗 ) −𝑛𝑝𝑖 𝑝𝑗
𝜌(𝑋𝑖 , 𝑋𝑗 ) = 𝐶𝑜𝑟𝑟(𝑋𝑖 , 𝑋𝑗 ) = =
√𝑉𝑎𝑟(𝑋𝑖 )√𝑉𝑎𝑟( 𝑋𝑗 ) √𝑛𝑝𝑖 (1 − 𝑝𝑖 )√𝑛𝑝𝑗 (1 − 𝑝𝑗 )

𝑝𝑖 𝑝𝑗
= −√ ; 𝑖, 𝑗 = 1(1)(𝑘 − 1), 𝑖 ≠ 𝑗
(1 − 𝑝𝑖 )(1 − 𝑝𝑗 )

The variance-covariance matrix or dispersion matrix of 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘−1 )𝑇 is given by


~

𝑛𝑝1 (1 − 𝑝1 ) −𝑛𝑝1 𝑝2 ⋯ −𝑛𝑝1 𝑝𝑘−1


−𝑛𝑝1 𝑝2 𝑛𝑝2 (1 − 𝑝2 ) ⋯ −𝑛𝑝2 𝑝𝑘−1
𝚺(𝑘−1)×(𝑘−1) = 𝐷𝑖𝑠𝑝 (𝑋) = ( )
~ ⋮ ⋮ ⋱ ⋮
−𝑛𝑝1 𝑝𝑘−1 −𝑛𝑝2 𝑝𝑘−1 ⋯ 𝑛𝑝𝑘−1 (1 − 𝑝𝑘−1 ) (𝑘−1)×(𝑘−1)

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 26


MULTIVARIATE PROBABILITY DISTRIBUTION

The correlation matrix of 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘−1 )𝑇 is given by


~

𝑹(𝑘−1)×(𝑘−1) = 𝐶𝑜𝑟𝑟 (𝑋)


~

𝑝1 𝑝2 𝑝1 𝑝𝑘−1
1 −√ ⋯ −√
(1 − 𝑝1 )(1 − 𝑝2 ) (1 − 𝑝1 )(1 − 𝑝𝑘−1 )

𝑝2 𝑝1 𝑝2 𝑝𝑘−1
= −√ 1 ⋯ −√
(1 − 𝑝2 )(1 − 𝑝1 ) (1 − 𝑝2 )(1 − 𝑝𝑘−1 )
⋮ ⋮ ⋱ ⋮
𝑝𝑘−1 𝑝1 𝑝𝑘−1 𝑝2
−√ −√ ⋯ 1
(1 − 𝑝𝑘−1 )(1 − 𝑝1 ) (1 − 𝑝𝑘−1 )(1 − 𝑝2 )
( )(𝑘−1)×(𝑘−1)

MARGINAL DISTRIBUTION (From (𝑿𝟏 , 𝑿𝟐 , ⋯ , 𝑿𝒌−𝟏 )𝑻)

The marginal MGF of 𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑠 )𝑇 is given by


~

𝑀𝑋 (1) (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑠 ) = 𝑀(𝑋1 ,𝑋2 ,⋯,𝑋𝑘−1 ) (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑠 , 0, 0, ⋯ , 0)


~

𝑛
𝑘−1 𝑘−1

= [{∑ 𝑝𝑖 𝑒 𝑡𝑖 + (1 − ∑ 𝑝𝑖 )} ]
𝑖=1 𝑖=1
𝑡𝑗 =0; 𝑗=(𝑠+1)(1)(𝑘−1)

𝑛 𝑛
𝑠 𝑘−1 𝑘−1 𝑠 𝑠
𝑡𝑖 𝑡𝑖
= {∑ 𝑝𝑖 𝑒 + ∑ 𝑝𝑖 + (1 − ∑ 𝑝𝑖 )} = {∑ 𝑝𝑖 𝑒 + (1 − ∑ 𝑝𝑖 )}
𝑖=1 𝑖=𝑠+1 𝑖=1 𝑖=1 𝑖=1

𝑠 𝑠

← MGF of 𝑀𝑢𝑙𝑡𝑖𝑛𝑜𝑚𝑖𝑎𝑙(𝑛, 𝑝1 , 𝑝2 , ⋯ , 𝑝𝑠 ), where ∑ 𝑥𝑖 ≤ 𝑛 , 0 < 𝑝𝑖 < 1 and ∑ 𝑝𝑖 < 1, 𝑖 = 1(1)𝑠


𝑖=1 𝑖=1

By uniqueness property of MGF, we get


𝑠 𝑠
(1)
𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑠 )𝑇 ~ 𝑀𝑢𝑙𝑡𝑖𝑛𝑜𝑚𝑖𝑎𝑙(𝑛, 𝑝1 , 𝑝2 , ⋯ , 𝑝𝑠 ), where ∑ 𝑥𝑖 ≤ 𝑛 , 0 < 𝑝𝑖 < 1 and ∑ 𝑝𝑖 < 1.
~
𝑖=1 𝑖=1

The marginal MGF of 𝑋𝑖 ; 𝑖 = 1(1)(𝑘 − 1) is given by


𝑛
𝑘−1 𝑘−1

𝑀𝑋𝑖 (𝑡𝑖 ) = 𝑀(𝑋1 ,𝑋2 ,⋯,𝑋𝑘−1 ) (0, 0, ⋯ , 0, 𝑡𝑖 , 0, ⋯ , 0) = [{∑ 𝑝𝑖 𝑒 𝑡𝑖 + (1 − ∑ 𝑝𝑖 )} ]


𝑖=1 𝑖=1
𝑡𝑗 =0; 𝑗≠𝑖

𝑛
𝑘−1 𝑘−1
𝑡𝑖
= 𝑝𝑖 𝑒 + ∑ 𝑝𝑗 + (1 − ∑ 𝑝𝑖 )
𝑗=1 𝑖=1
{ 𝑗≠𝑖 }

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 27


MULTIVARIATE PROBABILITY DISTRIBUTION

𝑛
𝑘−1

= {𝑝1 + ⋯ + 𝑝𝑖−1 + 𝑝𝑖 𝑒 𝑡𝑖 + 𝑝𝑖+1 + ⋯ + 𝑝𝑘−1 + (1 − ∑ 𝑝𝑖 )}


𝑖=1

𝑛
= (𝑝𝑖 𝑒 𝑡𝑖 + (1 − 𝑝𝑖 )) ; 𝑖 = 1(1)(𝑘 − 1) ← MGF of 𝐵𝑖𝑛(𝑛, 𝑝𝑖 )

By uniqueness property of MGF, we get

𝑋𝑖 ~ 𝐵𝑖𝑛(𝑛, 𝑝𝑖 ) ; 𝑖 = 1(1)(𝑘 − 1).

ALTERNATIVE

The marginal p.m.f. of 𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑠 )𝑇 is given by


~

𝑝𝑋 (1) (𝑥 (1) ) = 𝑝(𝑋1 ,𝑋2 ,⋯,𝑋𝑠 ) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑠 ) = ∑ ∑ ⋯ ∑ 𝑝𝑋 (𝑥 )


~ ~ ~ ~
𝑥𝑗 =0(1)𝑛, 𝑗=(𝑠+1)(1)(𝑘−1), ∑𝑘−1
𝑗=1 𝑥𝑗 ≤𝑛

(𝑛−∑𝑘−1
𝑖=1 𝑥𝑖 )
𝑘−1 𝑘−1
𝑛! 𝑥
= ∑ ∑ ⋯ ∑ ∙ ∏(𝑝𝑖 𝑖 ) (1 − ∑ 𝑝𝑖 )
∏𝑘−1
𝑖=1 (𝑥𝑖 !) (𝑛 − ∑𝑘−1
𝑖=1 𝑥𝑖 )!
𝑥𝑗 =0(1)𝑛, 𝑗=(𝑠+1)(1)(𝑘−1), ∑𝑘−1 𝑖=1 𝑖=1
𝑗=1 𝑥𝑗 ≤𝑛

𝑛! 𝑥 𝑥 𝑥
= ∙ 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑠 𝑠
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! (𝑛 − ∑𝑠𝑖=1 𝑥𝑖 )!
(𝑛−∑𝑘−1
𝑖=1 𝑥𝑖 )
𝑘−1 𝑘−1
(𝑛 − ∑𝑠𝑖=1 𝑥𝑖 )! 𝑥𝑖
∙ ∑ ∑ ⋯ ∑ ∙ ∏ (𝑝𝑖 ) (1 − ∑ 𝑝𝑖 )
∏𝑘−1
𝑖=𝑠+1(𝑥𝑖 !) (𝑛 − ∑𝑘−1
𝑖=1 𝑥𝑖 )!
𝑥𝑗 =0(1)𝑛, 𝑗=(𝑠+1)(1)(𝑘−1), ∑𝑘−1
𝑥𝑗 ≤𝑛 𝑖=𝑠+1 𝑖=1
𝑗=1

(𝑛−∑𝑠𝑖=1 𝑥𝑖 )
𝑘−1
𝑛! 𝑥 𝑥 𝑥
= ∙ 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑠 𝑠 ∙ {(1 − ∑ 𝑝𝑖 ) + 𝑝𝑠+1 + 𝑝𝑠+2 + ⋯ + 𝑝𝑘−1 }
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! (𝑛 − ∑𝑠𝑖=1 𝑥𝑖 )!
𝑖=1

𝑠 (𝑛−∑𝑠𝑖=1 𝑥𝑖 )
𝑛! 𝑥 𝑥 𝑥
= ∙ 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑠 𝑠 ∙ (1 − ∑ 𝑝𝑖 )
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! (𝑛 − ∑𝑠𝑖=1 𝑥𝑖 )!
𝑖=1

𝑠 𝑠

← 𝑝. 𝑚. 𝑓. of 𝑀𝑢𝑙𝑡𝑖𝑛𝑜𝑚𝑖𝑎𝑙(𝑛, 𝑝1 , 𝑝2 , ⋯ , 𝑝𝑠 ), where ∑ 𝑥𝑖 ≤ 𝑛 , 0 < 𝑝𝑖 < 1 and ∑ 𝑝𝑖 < 1, 𝑖 = 1(1)𝑠


𝑖=1 𝑖=1

Thus, 𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑠 )𝑇 ~ 𝑀𝑢𝑙𝑡𝑖𝑛𝑜𝑚𝑖𝑎𝑙(𝑛, 𝑝1 , 𝑝2 , ⋯ , 𝑝𝑠 );


~
𝑠 𝑠

where ∑ 𝑥𝑖 ≤ 𝑛 , 0 < 𝑝𝑖 < 1 and ∑ 𝑝𝑖 < 1.


𝑖=1 𝑖=1

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 28


MULTIVARIATE PROBABILITY DISTRIBUTION

The marginal p.m.f. of 𝑋𝑖 ; 𝑖 = 1(1)(𝑘 − 1) is given by

𝑝𝑋𝑖 (𝑥𝑖 ) = ∑ ∑ ⋯ ∑ 𝑝𝑋 (𝑥 ) ; 𝑖 = 1(1)(𝑘 − 1)


~ ~
𝑥𝑗 =0(1)𝑛, 𝑥𝑗 ≠𝑥𝑖 , ∑𝑘−1
𝑗=1 𝑥𝑗 ≤𝑛

(𝑛−∑𝑘−1
𝑖=1 𝑥𝑖 )
𝑘−1 𝑘−1
𝑛! 𝑥
= ∑ ∑ ⋯ ∑ ∙ ∏(𝑝𝑖 𝑖 ) (1 − ∑ 𝑝𝑖 )
∏𝑘−1
𝑖=1 (𝑥𝑖 !) (𝑛 − ∑𝑘−1
𝑖=1 𝑥𝑖 )!
𝑥𝑗 =0(1)𝑛, 𝑥𝑗 ≠𝑥𝑖 , ∑𝑘−1 𝑖=1 𝑖=1
𝑗=1 𝑥𝑗 ≤𝑛

𝑛! 𝑥 (𝑛 − 𝑥𝑖 )!
= ∙𝑝 𝑖 ∙ ∑ ∑ ⋯ ∑
𝑥𝑖 ! (𝑛 − 𝑥𝑖 )! 𝑖 𝑥1 ! ⋯ 𝑥𝑖−1 ! 𝑥𝑖+1 ! ⋯ 𝑥𝑘−1 ! (𝑛 − ∑𝑘−1
𝑖=1 𝑥𝑖 )!
𝑥𝑗 =0(1)(𝑛−𝑥𝑖 ), 𝑥𝑗 ≠𝑥𝑖 , ∑𝑘−1
𝑗=1,𝑗≠𝑖 𝑥𝑗 ≤𝑛−𝑥𝑖

(𝑛−∑𝑘−1
𝑖=1 𝑥𝑖 )
𝑘−1
𝑥 𝑥 𝑥 𝑥
∙ 𝑝1 1 ⋯ 𝑝𝑖−1
𝑖−1
𝑝𝑖+1
𝑖+1
⋯ 𝑝𝑘−1
𝑘−1
(1 − ∑ 𝑝𝑖 )
𝑖=1

(𝑛−𝑥𝑖 )
𝑘−1
𝑛! 𝑥
= ∙ 𝑝 𝑖 ∙ {(1 − ∑ 𝑝𝑖 ) 𝑝1 + ⋯ + 𝑝𝑖−1 + 𝑝𝑖+1 + ⋯ + 𝑝𝑘−1 }
𝑥𝑖 ! (𝑛 − 𝑥𝑖 )! 𝑖
𝑖=1

𝑘
𝑛! 𝑥
= ∙ 𝑝 𝑖 ∙ (1 − 𝑝𝑖 )(𝑛−𝑥𝑖) [∵ ∑ 𝑝𝑖 = 1]
𝑥𝑖 ! (𝑛 − 𝑥𝑖 )! 𝑖
𝑖=1

← 𝑝. 𝑚. 𝑓. of 𝐵𝑖𝑛(𝑛, 𝑝𝑖 ) ; 𝑖 = 1(1)(𝑘 − 1)

Hence, 𝑋𝑖 ~ 𝐵𝑖𝑛(𝑛, 𝑝𝑖 ) ; 𝑖 = 1(1)(𝑘 − 1).

CONDITIONAL DISTRIBUTION

The conditional distribution of 𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑠 )𝑇 given that 𝑋 (2) = 𝑥 (2) , where 𝑋 (2) =
~ ~ ~ ~

(𝑋𝑠+1 , 𝑋𝑠+2 , ⋯ , 𝑋𝑘−1 )𝑇 [ i. e. given that (𝑋𝑠+1 = 𝑥𝑠+1 , 𝑋𝑠+2 = 𝑥𝑠+2 , ⋯ , 𝑋𝑘−1 = 𝑥𝑘−1 )] is given by

𝑝 (1) (2) (𝑥 (1) ) = 𝑝((𝑋 , 𝑋 , ⋯ , 𝑋 )|(𝑥 , 𝑥 , ⋯ , 𝑥 )) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑠 )


(𝑋 |𝑋 = 𝑥 (2) ) ~ 1 2 𝑠 𝑠+1 𝑠+2 𝑘−1
~ ~ ~

𝑝𝑋 (𝑥 ) 𝑝(𝑋1 ,𝑋2 ,⋯,𝑋𝑘−1 ) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑘−1 )


~ ~
= =
𝑝(𝑋 (2) ) (𝑥 (2) ) = 𝑝(𝑋𝑠+1 ,𝑋𝑠+2 ,⋯,𝑋𝑘−1 ) (𝑥𝑠+1 , 𝑥𝑠+2 , ⋯ , 𝑥𝑘−1 )
~ ~

𝑛! 𝑘−1 𝑥𝑖 𝑘−1 (𝑛−∑𝑘−1


𝑖=1 𝑥𝑖 )
𝑘−1 𝑘−1 ∙ ∏ 𝑖=1 (𝑝𝑖 ) (1 − ∑𝑖=1 𝑖𝑝 )
∏𝑖=1 (𝑥𝑖 !) (𝑛 − ∑𝑖=1 𝑥𝑖 )!
=
𝑛! 𝑥𝑠+1 𝑥𝑠+2 𝑥𝑘−1 (𝑛−∑𝑘−1
𝑖=𝑠+1 𝑥𝑖 )
𝑘−1 ∙ 𝑝𝑠+1 𝑝𝑠+2 ⋯ 𝑝𝑘−1 (1 − ∑𝑘−1
𝑖=𝑠+1 𝑝𝑖 )

𝑥𝑠+1 ! 𝑥𝑠+2 ! ⋯ 𝑥𝑘−1 ! (𝑛 − 𝑖=𝑠+1 𝑥𝑖 )!

𝑥 𝑥 𝑥 (𝑛−∑𝑘−1 𝑥𝑖 )
(𝑛 − ∑𝑘−1𝑖=𝑠+1 𝑥𝑖 )! 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑠 𝑠 (1 − ∑𝑘−1
𝑖=1 𝑝𝑖 )
𝑖=1

= ∙
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! (𝑛 − ∑𝑘−1
𝑖=1 𝑥𝑖 )! (1 − ∑𝑘−1
{(𝑛−∑𝑘−1 𝑠
𝑖=1 𝑥𝑖 )+∑𝑖=1 𝑥𝑖 }
𝑖=𝑠+1 𝑝𝑖 )

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 29


MULTIVARIATE PROBABILITY DISTRIBUTION

𝑥 𝑥 𝑥 (𝑛−∑𝑘−1 𝑥𝑖 )
(𝑛 − ∑𝑘−1𝑖=𝑠+1 𝑥𝑖 )! 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑠 𝑠 {(1 − ∑𝑘−1 𝑠
𝑖=𝑠+1 𝑝𝑖 ) − ∑𝑖=1 𝑝𝑖 }
𝑖=1

= ∙ ∙
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! (𝑛 − ∑𝑘−1 (∑𝑠 𝑥𝑖 ) (𝑛−∑𝑘−1
𝑖=1 𝑥𝑖 )! (1 − ∑𝑘−1 𝑝𝑖 ) 𝑖=1 (1 − ∑𝑘−1 𝑖=1 𝑥𝑖 )
𝑖=𝑠+1 𝑖=𝑠+1 𝑝𝑖 )

(𝑛 − ∑𝑘−1
𝑖=𝑠+1 𝑥𝑖 )!
=
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! ((𝑛 − ∑𝑘−1 𝑠
𝑖=𝑠+1 𝑥𝑖 ) − ∑𝑖=1 𝑥𝑖 ) !
𝑥1 𝑥2 𝑥𝑠
𝑝1 𝑝2 𝑝𝑠
∙{ } { } ⋯{ }
(1 − ∑𝑘−1
𝑖=𝑠+1 𝑝𝑖 ) (1 − ∑𝑘−1
𝑖=𝑠+1 𝑝𝑖 ) (1 − ∑𝑘−1
𝑖=𝑠+1 𝑝𝑖 )

𝑠 ((𝑛−∑𝑘−1 𝑠
𝑖=𝑠+1 𝑥𝑖 )−∑𝑖=1 𝑥𝑖 )
𝑝𝑖
∙ {1 − ∑ ( )}
𝑖=1
(1 − ∑𝑘−1
𝑖=𝑠+1 𝑝𝑖 )

𝑠 ((𝑛−∑𝑘−1 𝑠
𝑖=𝑠+1 𝑥𝑖 )−∑𝑖=1 𝑥𝑖 )
(𝑛 − ∑𝑘−1
𝑖=𝑠+1 𝑥𝑖 )! 𝑥 𝑥 𝑥
= ∙ 𝑄1 1 𝑄2 2 ⋯ 𝑄𝑠 𝑠 ∙ {1 − ∑ 𝑄𝑖 }
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! ((𝑛 − ∑𝑘−1
𝑖=𝑠+1 𝑥𝑖 ) − ∑𝑠𝑖=1 𝑥𝑖 ) ! 𝑖=1

𝑘−1

← 𝑝. 𝑚. 𝑓. of 𝑀𝑢𝑙𝑡𝑖𝑛𝑜𝑚𝑖𝑎𝑙 [{𝑛 − ∑ 𝑥𝑖 } , 𝑄1 , 𝑄2 , ⋯ , 𝑄𝑠 ] ;
𝑖=𝑠+1
𝑘−1 𝑠
𝑝𝑖
where ∑ 𝑥𝑖 ≤ 𝑛, 𝑄𝑖 = ( ) and ∑ 𝑄𝑖 < 1; 𝑖 = 1(1)𝑠
(1 − ∑𝑘−1
𝑖=𝑠+1 𝑝𝑖 )
𝑖=1 𝑖=1

The conditional distribution of 𝑋 (1) = (𝑋1 , 𝑋2 )𝑇 given that 𝑋3 = 𝑥3 , is given by


~

𝑝(𝑋1 ,𝑋2 ,𝑋3 ) (𝑥1 , 𝑥2 , 𝑥3 )


𝑝(𝑋1 ,𝑋2 )| 𝑋3 =𝑥3 (𝑥1 , 𝑥2 ) =
𝑝𝑋3 (𝑥3 )

𝑛! 𝑥 𝑥 𝑥
∙ 𝑝 1 𝑝 2 𝑝 3 (1 − 𝑝1 − 𝑝2 − 𝑝3 )(𝑛−𝑥1 −𝑥2 −𝑥3 )
𝑥1 ! 𝑥2 ! 𝑥3 ! (𝑛 − 𝑥1 − 𝑥2 − 𝑥3 )! 1 2 3
=
𝑛! 𝑥
∙ 𝑝 3 (1 − 𝑝3 )(𝑛−𝑥3 )
𝑥3 ! (𝑛 − 𝑥3 )! 3

𝑥 𝑥 ((𝑛−𝑥3 )−𝑥1 −𝑥2 )


(𝑛 − 𝑥3 )! 𝑝1 1 𝑝2 2 ((1 − 𝑝3 ) − 𝑝1 − 𝑝2 )
= ∙
𝑥1 ! 𝑥2 ! ((𝑛 − 𝑥3 ) − 𝑥1 − 𝑥2 )! (1 − 𝑝3 )𝑥1 +𝑥2 (1 − 𝑝3 )((𝑛−𝑥3 )−𝑥1 −𝑥2 )

(𝑛 − 𝑥3 )! 𝑝1 𝑥1 𝑝2 𝑥2 𝑝1 𝑝2 ((𝑛−𝑥3 )−𝑥1 −𝑥2 )


= ∙( ) ( ) ∙ (1 − − )
𝑥1 ! 𝑥2 ! ((𝑛 − 𝑥3 ) − 𝑥1 − 𝑥2 )! 1 − 𝑝3 1 − 𝑝3 1 − 𝑝3 1 − 𝑝3

(𝑛 − 𝑥3 )! 𝑥 𝑥
= ∙ 𝑄1 1 𝑄2 2 ∙ (1 − 𝑄1 − 𝑄2 )((𝑛−𝑥3 )−𝑥1 −𝑥2 )
𝑥1 ! 𝑥2 ! ((𝑛 − 𝑥3 ) − 𝑥1 − 𝑥2 )!

3
𝑝𝑖
← 𝑝. 𝑚. 𝑓. of 𝑀𝑢𝑙𝑡𝑖𝑛𝑜𝑚𝑖𝑎𝑙((𝑛 − 𝑥3 ), 𝑄1 , 𝑄2 ), where ∑ 𝑥𝑖 ≤ 𝑛, 𝑄𝑖 = ( ) and 𝑄1 + 𝑄2 < 1; 𝑖 = 1, 2
1 − 𝑝3
𝑖=1

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 30


MULTIVARIATE PROBABILITY DISTRIBUTION

The conditional distribution of 𝑋1 given that 𝑋 (2) = 𝑥 (2) , where 𝑋 (2) = (𝑋2 , 𝑋3 , ⋯ , 𝑋𝑘−1 )𝑇
~ ~ ~

[ i. e. The conditional distribution of 𝑋1 given that (𝑋2 = 𝑥2 , 𝑋3 = 𝑥3 , ⋯ , 𝑋𝑘−1 = 𝑥𝑘−1 )] is given by

𝑝𝑋 (𝑥 ) 𝑝(𝑋1 ,𝑋2 ,⋯,𝑋𝑘−1 ) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑘−1 )


~ ~
𝑝 (2) (𝑥 ) = 𝑝(𝑋 |(𝑥 , 𝑥 , ⋯ , 𝑥 )) (𝑥1 ) = =
(𝑋1 |𝑋 = 𝑥 (2)) 1 1 2 3 𝑘−1 𝑝(𝑋 (2) ) (𝑥 (2) ) 𝑝(𝑋2 ,𝑋3 ,⋯,𝑋𝑘−1 ) (𝑥2 , 𝑥3 , ⋯ , 𝑥𝑘−1 )
~ ~
~ ~

𝑛! 𝑘−1 𝑥𝑖 𝑘−1 (𝑛−∑𝑘−1


𝑖=1 𝑥𝑖 )
𝑘−1 𝑘−1 ∙ ∏ 𝑖=1 (𝑝𝑖 ) (1 − ∑ 𝑖=1 𝑝 𝑖 )
∏𝑖=1 (𝑥𝑖 !) (𝑛 − ∑𝑖=1 𝑥𝑖 )!
=
𝑛! 𝑥2 𝑥3 𝑥𝑘−1 𝑘−1 (𝑛−∑𝑘−1
𝑖=2 𝑥𝑖 )
𝑘−1 ∙ 𝑝2 𝑝 3 ⋯ 𝑝𝑘−1 (1 − ∑𝑖=2 𝑖 𝑝 )
𝑥2 ! 𝑥3 ! ⋯ 𝑥𝑘−1 ! (𝑛 − ∑𝑖=2 𝑥𝑖 )!

𝑥 ((𝑛−∑𝑘−1
𝑖=2 𝑥𝑖 )−𝑥1 )
(𝑛 − ∑𝑘−1𝑖=2 𝑥𝑖 )!
𝑝1 1 ((1 − ∑𝑘−1
𝑖=2 𝑝𝑖 ) − 𝑝1 )
= ∙
𝑥1 ! (𝑛 − ∑𝑘−1
𝑖=1 𝑥𝑖 )!
𝑥1 ((𝑛−∑𝑘−1
𝑖=2 𝑥𝑖 )−𝑥1 )
(1 − ∑𝑘−1 𝑘−1
𝑖=2 𝑝𝑖 ) (1 − ∑𝑖=2 𝑝𝑖 )

𝑥1 ((𝑛−∑𝑘−1
𝑖=2 𝑥𝑖 )−𝑥1 )
(𝑛 − ∑𝑘−1
𝑖=2 𝑥𝑖 )! 𝑝1 𝑝1
= ∙{ } ∙ {1 − }
𝑥1 ! ((𝑛 − ∑𝑘−1
𝑖=2 𝑥𝑖 ) − 𝑥1 ) !
(1 − ∑𝑘−1
𝑖=2 𝑝𝑖 ) (1 − ∑𝑘−1
𝑖=2 𝑝𝑖 )

𝑘−1
𝑝1
← 𝑝. 𝑚. 𝑓. of 𝐵𝑖𝑛𝑜𝑚𝑖𝑎𝑙 ({𝑛 − ∑ 𝑥𝑖 } , { })
𝑖=2
(1 − ∑𝑘−1
𝑖=2 𝑝𝑖 )

𝑘−1
𝑝1
← 𝑝. 𝑚. 𝑓. of 𝐵𝑖𝑛𝑜𝑚𝑖𝑎𝑙 ({𝑛 − ∑ 𝑥𝑖 } , 𝑄1 ) , where 𝑄1 = { }
𝑖=2
(1 − ∑𝑘−1
𝑖=2 𝑝𝑖 )

𝑘−1
(2) (2) 𝑝1
Now, 𝐸 (𝑋1 |𝑋 =𝑥 ) = 𝐸(𝑋1 |(𝑥2 , 𝑥3 , ⋯ , 𝑥𝑘−1 )) = {𝑛 − ∑ 𝑥𝑖 } ∙ { }
~ ~
𝑖=2
(1 − ∑𝑘−1
𝑖=2 𝑝𝑖 )

𝑝1 𝑝1 𝑝1 𝑝1
= 𝑛{ }−{ } 𝑥2 − { } 𝑥3 − ⋯ − { } 𝑥𝑘−1
(1 − ∑𝑘−1
𝑖=2 𝑝𝑖 ) (1 − ∑𝑘−1
𝑖=2 𝑝𝑖 ) (1 − ∑𝑘−1
𝑖=2 𝑝𝑖 ) (1 − ∑𝑘−1
𝑖=2 𝑝𝑖 )

← 𝐴 𝑙𝑖𝑛𝑒𝑎𝑟 𝑓𝑢𝑛𝑐𝑡𝑖𝑜𝑛 𝑜𝑓 𝑥2 , 𝑥3 , ⋯ , 𝑥𝑘−1


𝑘−1
(2) (2) 𝑝1 𝑝1
𝑉𝑎𝑟 (𝑋1 |𝑋 =𝑥 ) = 𝑉𝑎𝑟(𝑋1 |(𝑥2 , 𝑥3 , ⋯ , 𝑥𝑘−1 )) = {𝑛 − ∑ 𝑥𝑖 } { } {1 − }
~ ~
𝑖=2
(1 − ∑𝑘−1
𝑖=2 𝑝𝑖 ) (1 − ∑𝑘−1
𝑖=2 𝑝𝑖 )

The regression coefficient of 𝑋1 on 𝑋2 , eliminating the effects of (𝑋3 , ⋯ , 𝑋𝑘−1 ), is given by

𝑝1
𝛽12.34⋯𝑘−1 = − { }
(1 − 𝑝2 − ∑𝑘−1
𝑖=3 𝑝𝑖 )

Similarly, the regression coefficient of 𝑋2 on 𝑋1 , eliminating the effects of (𝑋3 , ⋯ , 𝑋𝑘−1 ), is given by

𝑝2
𝛽21.34⋯𝑘−1 = − { }
(1 − 𝑝1 − ∑𝑘−1
𝑖=3 𝑝𝑖 )

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 31


MULTIVARIATE PROBABILITY DISTRIBUTION

The partial correlation coefficient of 𝑋1 and 𝑋2 , eliminating the effects of (𝑋3 , ⋯ , 𝑋𝑘−1 ), is given by

2 𝑝1 𝑝2
𝜌12.34⋯𝑘−1 =[ ]
(1 − 𝑝2 − ∑𝑖=3 𝑝𝑖 )(1 − 𝑝1 − ∑𝑘−1
𝑘−1
𝑖=3 𝑝𝑖 )

1
𝑝1 𝑝2 2
𝜌12.34⋯𝑘−1 = − [ ] ; [∵ 𝛽12.34⋯𝑘−1 , 𝛽21.34⋯𝑘−1 < 0]
(1 − 𝑝2 − ∑𝑘−1 𝑘−1
𝑖=3 𝑝𝑖 )(1 − 𝑝1 − ∑𝑖=3 𝑝𝑖 )

MULTIPLE CORRELATION COEFFICIENT:

The multiple correlation coefficient of 𝑋1 on 𝑋 (2) = (𝑋2 , 𝑋3 , ⋯ , 𝑋𝑘−1 )𝑇 is given by


~

1
|𝚺| 2
𝜌1.23⋯(𝑘−1) = 𝐶𝑜𝑟𝑟(𝑥1 , 𝑋1.23⋯(𝑘−1) ) = (1 − )
𝜎11 |𝚺2 |

The variance-covariance matrix or dispersion matrix of 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘−1 )𝑇 is given by


~

𝑛𝑝1 (1 − 𝑝1 ) −𝑛𝑝1 𝑝2 −𝑛𝑝1 𝑝3 ⋯ −𝑛𝑝1 𝑝𝑘−1


−𝑛𝑝2 𝑝1 𝑛𝑝2 (1 − 𝑝2 ) −𝑛𝑝2 𝑝3 ⋯ −𝑛𝑝2 𝑝𝑘−1
𝚺(𝑘−1)×(𝑘−1) = −𝑛𝑝3 𝑝1 −𝑛𝑝3 𝑝2 𝑛𝑝3 (1 − 𝑝3 ) ⋯ −𝑛𝑝3 𝑝𝑘−1
⋮ ⋮ ⋮ ⋱ ⋮
( −𝑛𝑝 𝑝
𝑘−1 1 −𝑛𝑝 𝑝
𝑘−1 2 −𝑛𝑝 𝑘−1 𝑝3 ⋯ 𝑛𝑝𝑘−1 (1 − 𝑝𝑘−1 ))(𝑘−1)×(𝑘−1)

The determinant of variance-covariance matrix or dispersion matrix is given by

𝑛𝑝1 (1 − 𝑝1 ) −𝑛𝑝1 𝑝2 −𝑛𝑝1 𝑝3 ⋯ −𝑛𝑝1 𝑝𝑘−1


−𝑛𝑝2 𝑝1 𝑛𝑝2 (1 − 𝑝2 ) −𝑛𝑝2 𝑝3 ⋯ −𝑛𝑝2 𝑝𝑘−1
| |
|𝚺(𝑘−1)×(𝑘−1) | = −𝑛𝑝3 𝑝1 −𝑛𝑝3 𝑝2 𝑛𝑝3 (1 − 𝑝3 ) ⋯ −𝑛𝑝3 𝑝𝑘−1
| |
⋮ ⋮ ⋮ ⋱ ⋮
−𝑛𝑝𝑘−1 𝑝1 −𝑛𝑝𝑘−1 𝑝2 −𝑛𝑝𝑘−1 𝑝3 ⋯ 𝑛𝑝𝑘−1 (1 − 𝑝𝑘−1 )

𝑘−1

𝑝1 {1 − ∑ 𝑝𝑖 } −𝑝1 𝑝2 −𝑝1 𝑝3 ⋯ −𝑝1 𝑝𝑘−1


| |
𝑖=1
𝑘−1

| 𝑝2 {1 − ∑ 𝑝𝑖 } 𝑝2 (1 − 𝑝2 ) −𝑝2 𝑝3 ⋯ −𝑝2 𝑝𝑘−1 |


𝑖=1
= 𝑛𝑘−1 𝑘−1

| 𝑝3 {1 − ∑ 𝑝𝑖 } −𝑝3 𝑝2 𝑝3 (1 − 𝑝3 ) ⋯ −𝑝3 𝑝𝑘−1 |


𝑖=1
⋮ ⋮ ⋮ ⋱ ⋮
𝑘−1
| |
𝑝𝑘−1 {1 − ∑ 𝑝𝑖 } −𝑝𝑘−1 𝑝2 −𝑝𝑘−1 𝑝3 ⋯ 𝑝𝑘−1 (1 − 𝑝𝑘−1 )
𝑖=1

1 −𝑝2 −𝑝3 ⋯ −𝑝𝑘−1


𝑘−1
1 (1 − 𝑝2 ) −𝑝3 ⋯ −𝑝𝑘−1
| |
= 𝑛𝑘−1 𝑝1 𝑝2 𝑝3 ⋯ 𝑝𝑘−1 {1 − ∑ 𝑝𝑖 } 1 −𝑝2 (1 − 𝑝3 ) ⋯ −𝑝𝑘−1
| |
𝑖=1 ⋮ ⋮ ⋮ ⋱ ⋮
1 −𝑝2 −𝑝3 ⋯ (1 − 𝑝𝑘−1 )

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 32


MULTIVARIATE PROBABILITY DISTRIBUTION

1 −𝑝2 −𝑝3 ⋯ −𝑝𝑘−1


𝑘−1
0 1 0 ⋯ 0 𝑅𝑖′ = 𝑅𝑖 − 𝑅1 ,
= 𝑛𝑘−1 𝑝1 𝑝2 𝑝3 ⋯ 𝑝𝑘−1 {1 − ∑ 𝑝𝑖 } ||0 0 1 ⋯ 0 || ;
𝑖 = 2(1)(𝑘 − 1)
𝑖=1 ⋮ ⋮ ⋮ ⋱ ⋮
0 0 0 ⋯ 1
𝑘−1
𝑘−1
=𝑛 𝑝1 𝑝2 𝑝3 ⋯ 𝑝𝑘−1 {1 − ∑ 𝑝𝑖 }
𝑖=1

𝑘−1
𝑘−1
|𝚺(𝑘−1)×(𝑘−1) | = 𝑛 𝑝1 𝑝2 𝑝3 ⋯ 𝑝𝑘−1 {1 − ∑ 𝑝𝑖 } > 0 ;
𝑖=1
𝑘−1

since 0 < 𝑝𝑖 < 1, 𝑖 = 1(1)(𝑘 − 1), ∑ 𝑝𝑖 < 1 and 𝑛 (> 0) is a positive integer.
𝑖=1

|𝚺2 (𝑘−2)×(𝑘−2) | = Determinant obtained by deleting the1st row and 1st column of 𝚺(𝑘−1)×(𝑘−1)

𝑛𝑝2 (1 − 𝑝2 ) −𝑛𝑝2 𝑝3 −𝑛𝑝2 𝑝4 ⋯ −𝑛𝑝2 𝑝𝑘−1


−𝑛𝑝3 𝑝2 𝑛𝑝3 (1 − 𝑝3 ) −𝑛𝑝3 𝑝4 ⋯ −𝑛𝑝3 𝑝𝑘−1
| |
= −𝑛𝑝 𝑝
4 2 −𝑛𝑝 𝑝
4 3 𝑛𝑝4 (1 − 𝑝4 ) ⋯ −𝑛𝑝4 𝑝𝑘−1 |
|
⋮ ⋮ ⋮ ⋱ ⋮
−𝑛𝑝𝑘−1 𝑝2 −𝑛𝑝𝑘−1 𝑝3 −𝑛𝑝𝑘−1 𝑝4 ⋯ 𝑛𝑝𝑘−1 (1 − 𝑝𝑘−1 )
𝑘−1 𝑘−1
𝑘−2 𝑘−1
=𝑛 𝑝2 𝑝3 𝑝4 ⋯ 𝑝𝑘−1 {1 − ∑ 𝑝𝑖 } [Similar as |𝚺| = 𝑛 𝑝1 𝑝2 ⋯ 𝑝𝑘−1 {1 − ∑ 𝑝𝑖 }]
𝑖=2 𝑖=1

1
|𝚺| 2
∴ 𝜌1.23⋯(𝑘−1) = (1 − )
𝜎11 |𝚺2 |

1
𝑛𝑘−1 𝑝1 𝑝2 𝑝3 ⋯ 𝑝𝑘−1 {1 − ∑𝑘−1
𝑖=1 𝑝𝑖 }
2
= [1 − ]
𝑛𝑝1 (1 − 𝑝1 ) ∙ 𝑛𝑘−2 𝑝2 𝑝3 𝑝4 ⋯ 𝑝𝑘−1 {1 − ∑𝑘−1
𝑖=2 𝑝𝑖 }

1
{1 − ∑𝑘−1
𝑖=1 𝑝𝑖 }
2
= [1 − ]
(1 − 𝑝1 ){1 − ∑𝑘−1
𝑖=2 𝑝𝑖 }

1
1 − ∑𝑘−1 𝑘−1 𝑘−1
𝑖=2 𝑝𝑖 − 𝑝1 + 𝑝1 ∑𝑖=2 𝑝𝑖 − 1 + ∑𝑖=1 𝑝𝑖
2
=[ ]
(1 − 𝑝1 ){1 − ∑𝑘−1
𝑖=2 𝑝𝑖 }

1
𝑝1 ∑𝑘−1
𝑖=2 𝑝𝑖
2
=[ ]
(1 − 𝑝1 ){1 − ∑𝑘−1
𝑖=2 𝑝𝑖 }

The multiple correlation coefficient of 𝑋1 on 𝑋 (2) = (𝑋2 , 𝑋3 , ⋯ , 𝑋𝑘−1 )𝑇 is given by


~

1
𝑝1 ∑𝑘−1
𝑖=2 𝑝𝑖
2
𝜌1.23⋯(𝑘−1) =[ ]
(1 − 𝑝1 ){1 − ∑𝑘−1
𝑖=2 𝑝𝑖 }

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 33


MULTIVARIATE PROBABILITY DISTRIBUTION

PARTIAL CORRELATION COEFFICIENT:

The partial correlation coefficient between 𝑋1 and 𝑋2 , eliminating the effect of 𝑋 (3) = (𝑋3 , 𝑋4 , ⋯ , 𝑋𝑘−1 )𝑇 is
~

given by
𝚺12
𝜌12.34⋯(𝑘−1) = − ,
√𝚺11 ∙ 𝚺22

where 𝚺𝑖𝑗 = cofactor of 𝜎𝑖𝑗 in 𝚺(𝑘−1)×(𝑘−1) ; 𝑖, 𝑗 = 1(1)(𝑘 − 1).

𝚺12 = cofactor of 𝜎12 in 𝚺(𝑘−1)×(𝑘−1)

−𝑛𝑝2 𝑝1 −𝑛𝑝2 𝑝3 −𝑛𝑝2 𝑝4 ⋯ −𝑛𝑝2 𝑝𝑘−1


−𝑛𝑝3 𝑝1 𝑛𝑝3 (1 − 𝑝3 ) −𝑛𝑝3 𝑝4 ⋯ −𝑛𝑝3 𝑝𝑘−1
1+2 | 𝑛𝑝4 (1 − 𝑝4 ) −𝑛𝑝4 𝑝𝑘−1 ||
= (−1) | −𝑛𝑝4 𝑝1 −𝑛𝑝4 𝑝3 ⋯
⋮ ⋮ ⋮ ⋱ ⋮
−𝑛𝑝𝑘−1 𝑝1 −𝑛𝑝𝑘−1 𝑝3 −𝑛𝑝𝑘−1 𝑝4 ⋯ 𝑛𝑝𝑘−1 (1 − 𝑝𝑘−1 )

1 −𝑝3 −𝑝4 ⋯ −𝑝𝑘−1


1 (1 − 𝑝3 ) −𝑝4 ⋯ −𝑝𝑘−1
| |
= (−1)3 (−1) ∙ 𝑛𝑘−2 𝑝1 𝑝2 𝑝3 𝑝4 ⋯ 𝑝𝑘−1 1 −𝑝3 (1 − 𝑝4 ) ⋯ −𝑝𝑘−1 |
|
⋮ ⋮ ⋮ ⋱ ⋮
1 −𝑝3 −𝑝4 ⋯ (1 − 𝑝𝑘−1 )

1 −𝑝3 −𝑝4 ⋯ −𝑝𝑘−1


0 1 0 ⋯ 0 𝑅𝑖′ = 𝑅𝑖 − 𝑅1 ,
= 𝑛𝑘−2 𝑝1 𝑝2 𝑝3 𝑝4 ⋯ 𝑝𝑘−1 ||0 0 1 ⋯ 0 || ;
𝑖 = 2(1)(𝑘 − 2)
⋮ ⋮ ⋮ ⋱ ⋮
0 0 0 ⋯ 1

= 𝑛𝑘−2 𝑝1 𝑝2 𝑝3 𝑝4 ⋯ 𝑝𝑘−1

𝚺11 = cofactor of 𝜎11 in 𝚺(𝑘−1)×(𝑘−1)

𝑛𝑝2 (1 − 𝑝2 ) −𝑛𝑝2 𝑝3 −𝑛𝑝2 𝑝4 ⋯ −𝑛𝑝2 𝑝𝑘−1


−𝑛𝑝3 𝑝2 𝑛𝑝3 (1 − 𝑝3 ) −𝑛𝑝3 𝑝4 ⋯ −𝑛𝑝3 𝑝𝑘−1
| |
= 𝑛𝑝4 (1 − 𝑝4 )
| −𝑛𝑝4 𝑝2 −𝑛𝑝4 𝑝3 ⋯ −𝑛𝑝4 𝑝𝑘−1 |
⋮ ⋮ ⋮ ⋱ ⋮
−𝑛𝑝𝑘−1 𝑝2 −𝑛𝑝𝑘−1 𝑝3 −𝑛𝑝𝑘−1 𝑝4 ⋯ 𝑛𝑝𝑘−1 (1 − 𝑝𝑘−1 )
𝑘−1

{1 − 𝑝2 − ∑ 𝑝𝑖 } −𝑝3 −𝑝4 ⋯ −𝑝𝑘−1


| |
𝑖=3
𝑘−1

− 𝑝2 − ∑ 𝑝𝑖 } (1 − 𝑝3 ) −𝑝4 ⋯ −𝑝𝑘−1
|{1 |
𝑖=3
𝑘−2
=𝑛 𝑝2 𝑝3 𝑝4 ⋯ 𝑝𝑘−1 𝑘−1

|{1 − 𝑝2 − ∑ 𝑝𝑖 } −𝑝3 (1 − 𝑝4 ) ⋯ −𝑝𝑘−1 |


𝑖=3
⋮ ⋮ ⋮ ⋱ ⋮
𝑘−1
| |
{1 − 𝑝2 − ∑ 𝑝𝑖 } −𝑝3 −𝑝4 ⋯ (1 − 𝑝𝑘−1 )
𝑖=3

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 34


MULTIVARIATE PROBABILITY DISTRIBUTION

1 −𝑝3 −𝑝4 ⋯ −𝑝𝑘−1


1 𝑘−1 (1 − 𝑝3 ) −𝑝4 ⋯ −𝑝𝑘−1
𝑘−2 | |
= 𝑛 𝑝2 𝑝3 𝑝4 ⋯ 𝑝𝑘−1 {1 − 𝑝2 − ∑ 𝑝𝑖 } 1 −𝑝3 (1 − 𝑝4 ) ⋯ −𝑝𝑘−1
| |
𝑖=3 ⋮ ⋮ ⋮ ⋱ ⋮
1 −𝑝3 −𝑝4 ⋯ (1 − 𝑝𝑘−1 )

1 −𝑝3 −𝑝4 ⋯ −𝑝𝑘−1


𝑘−1
0 1 0 ⋯ 0 𝑅𝑖′ = 𝑅𝑖 − 𝑅1 ,
= 𝑛𝑘−2 𝑝2 𝑝3 𝑝4 ⋯ 𝑝𝑘−1 {1 − 𝑝2 − ∑ 𝑝𝑖 } ||0 0 1 ⋯ 0 || ;
𝑖 = 2(1)(𝑘 − 2)
𝑖=3 ⋮ ⋮ ⋮ ⋱ ⋮
0 0 0 ⋯ 1
𝑘−1
𝑘−2
=𝑛 𝑝2 𝑝3 𝑝4 ⋯ 𝑝𝑘−1 {1 − 𝑝2 − ∑ 𝑝𝑖 } > 0 ;
𝑖=3
𝑘−1

since 0 < 𝑝𝑖 < 1, 𝑖 = 1(1)(𝑘 − 1), ∑ 𝑝𝑖 < 1 and 𝑛 (> 0) is a positive integer.
𝑖=1

Similarly, 𝚺22 = cofactor of 𝜎11 in 𝚺(𝑘−1)×(𝑘−1)

𝑛𝑝1 (1 − 𝑝1 ) −𝑛𝑝1 𝑝3 −𝑛𝑝1 𝑝4 ⋯ −𝑛𝑝1 𝑝𝑘−1


−𝑛𝑝3 𝑝1 (1
𝑛𝑝3 − 𝑝3 ) −𝑛𝑝3 𝑝4 ⋯ −𝑛𝑝3 𝑝𝑘−1
| |
= −𝑛𝑝 𝑝
4 1 −𝑛𝑝 𝑝
4 3 𝑛𝑝4 (1 − 𝑝4 ) ⋯ −𝑛𝑝4 𝑝𝑘−1 |
|
⋮ ⋮ ⋮ ⋱ ⋮
−𝑛𝑝𝑘−1 𝑝1 −𝑛𝑝𝑘−1 𝑝3 −𝑛𝑝𝑘−1 𝑝4 ⋯ 𝑛𝑝𝑘−1 (1 − 𝑝𝑘−1 )

𝑘−1
𝑘−2
=𝑛 𝑝1 𝑝3 𝑝4 ⋯ 𝑝𝑘−1 {1 − 𝑝1 − ∑ 𝑝𝑖 } > 0 ;
𝑖=3
𝑘−1

since 0 < 𝑝𝑖 < 1, 𝑖 = 1(1)(𝑘 − 1), ∑ 𝑝𝑖 < 1 and 𝑛 (> 0) is a positive integer.
𝑖=1

𝚺12
∴ 𝜌12.34⋯(𝑘−1) = −
√𝚺11 ∙ 𝚺22

𝑛𝑘−2 𝑝1 𝑝2 𝑝3 𝑝4 ⋯ 𝑝𝑘−1
=−
√𝑛𝑘−2 𝑝2 𝑝3 𝑝4 ⋯ 𝑝𝑘−1 {1 − 𝑝2 − ∑𝑘−1
𝑖=3 𝑝𝑖 } ∙ √𝑛
𝑘−2 𝑝 𝑝 𝑝 ⋯ 𝑝
1 3 4
𝑘−1
𝑘−1 {1 − 𝑝1 − ∑𝑖=3 𝑝𝑖 }

1
𝑝1 𝑝2 2
= −[ ]
(1 − 𝑝2 − ∑𝑘−1 𝑘−1
𝑖=3 𝑝𝑖 )(1 − 𝑝1 − ∑𝑖=3 𝑝𝑖 )

The partial correlation coefficient between 𝑋1 and 𝑋2 , eliminating the effect of 𝑋 (3) = (𝑋3 , 𝑋4 , ⋯ , 𝑋𝑘−1 )𝑇 is
~

given by
1
𝑝1 𝑝2 2
𝜌12.34⋯𝑘−1 = − [ ]
(1 − 𝑝2 − ∑𝑘−1 𝑘−1
𝑖=3 𝑝𝑖 )(1 − 𝑝1 − ∑𝑖=3 𝑝𝑖 )

2 𝑝1 𝑝2
𝜌12.34⋯𝑘−1 =[ ]
(1 − 𝑝2 − ∑𝑖=3 𝑝𝑖 )(1 − 𝑝1 − ∑𝑘−1
𝑘−1
𝑖=3 𝑝𝑖 )

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 35


MULTIVARIATE PROBABILITY DISTRIBUTION

ALTERNATIVE

The conditional distribution of 𝑋1 given that (𝑋2 = 𝑥2 , 𝑋3 = 𝑥3 , ⋯ , 𝑋𝑘−1 = 𝑥𝑘−1 ) is given by


𝑘−1
𝑝1
𝑋1 |(𝑥2 , 𝑥3 , ⋯ , 𝑥𝑘−1 ) ~ 𝐵𝑖𝑛𝑜𝑚𝑖𝑎𝑙 ({𝑛 − ∑ 𝑥𝑖 } , { })
(1 − ∑𝑘−1
𝑖=2 𝑝𝑖 )
𝑖=2

𝑘−1
𝑝1
∴ 𝐸(𝑋1 |(𝑥2 , 𝑥3 , ⋯ , 𝑥𝑘−1 )) = {𝑛 − ∑ 𝑥𝑖 } ∙ { }
𝑖=2
(1 − ∑𝑘−1
𝑖=2 𝑝𝑖 )

𝑝1 𝑝1 𝑝1 𝑝1
= 𝑛{ }−{ } 𝑥2 − { } 𝑥3 − ⋯ − { } 𝑥𝑘−1
(1 − ∑𝑘−1
𝑖=2 𝑝𝑖 ) (1 − ∑𝑘−1
𝑖=2 𝑝𝑖 ) (1 − ∑𝑘−1
𝑖=2 𝑝𝑖 ) (1 − ∑𝑘−1
𝑖=2 𝑝𝑖 )

← A linear function of 𝑥2 , 𝑥3 , ⋯ , 𝑥𝑘−1

The conditional distribution of 𝑋2 given that (𝑋1 = 𝑥1 , 𝑋3 = 𝑥3 , ⋯ , 𝑋𝑘−1 = 𝑥𝑘−1 ) is given by

𝑘−1
𝑝1
𝑋2 |(𝑥1 , 𝑥3 , ⋯ , 𝑥𝑘−1 ) ~ 𝐵𝑖𝑛𝑜𝑚𝑖𝑎𝑙 ({𝑛 − 𝑥1 − ∑ 𝑥𝑖 } , { })
𝑖=3
(1 − 𝑝1 − ∑𝑘−1
𝑖=3 𝑝𝑖 )

𝑘−1
𝑝1
∴ 𝐸(𝑋2 |(𝑥1 , 𝑥3 , ⋯ , 𝑥𝑘−1 )) = {𝑛 − 𝑥1 − ∑ 𝑥𝑖 } ∙ { }
(1 − 𝑝1 − ∑𝑘−1
𝑖=3 𝑝𝑖 )
𝑖=3

𝑝1 𝑝1 𝑝1
= 𝑛{ 𝑘−1
}−{ 𝑘−1
} 𝑥2 − { } 𝑥3 − ⋯
(1 − 𝑝1 − ∑𝑖=3 𝑝𝑖 ) (1 − 𝑝1 − ∑𝑖=3 𝑝𝑖 ) (1 − 𝑝1 − ∑𝑘−1
𝑖=3 𝑝𝑖 )

𝑝1
−{ } 𝑥𝑘−1
(1 − 𝑝1 − ∑𝑘−1
𝑖=3 𝑝𝑖 )

← A linear function of 𝑥1 , 𝑥3 , ⋯ , 𝑥𝑘−1

The regression coefficient of 𝑋1 on 𝑋2 , eliminating the effects of (𝑋3 , ⋯ , 𝑋𝑘−1 ), is given by

𝑝1
𝛽12.34⋯𝑘−1 = − { }
(1 − 𝑝2 − ∑𝑘−1
𝑖=3 𝑝𝑖 )

Similarly, the regression coefficient of 𝑋2 on 𝑋1 , eliminating the effects of (𝑋3 , ⋯ , 𝑋𝑘−1 ), is given by

𝑝2
𝛽21.34⋯𝑘−1 = − { }
(1 − 𝑝1 − ∑𝑘−1
𝑖=3 𝑝𝑖 )

The partial correlation coefficient between 𝑋1 and 𝑋2 , eliminating the effects of (𝑋3 , ⋯ , 𝑋𝑘−1 ), is given by

2 𝑝1 𝑝2
𝜌12.34⋯𝑘−1 =[ ]
(1 − 𝑝2 − ∑𝑘−1 𝑘−1
𝑖=3 𝑝𝑖 )(1 − 𝑝1 − ∑𝑖=3 𝑝𝑖 )

1
𝑝1 𝑝2 2
𝜌12.34⋯𝑘−1 = − [ ] ; [∵ 𝛽12.34⋯𝑘−1 , 𝛽21.34⋯𝑘−1 < 0]
(1 − 𝑝2 − ∑𝑘−1 𝑘−1
𝑖=3 𝑝𝑖 )(1 − 𝑝1 − ∑𝑖=3 𝑝𝑖 )

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 36


MULTIVARIATE NORMAL DISTRIBUTION

Multivariate Analysis || Lecture Notes || Anup Kumar Giri


MULTIVARIATE NORMAL DISTRIBUTION

MULTIVARIATE NORMAL DISTRIBUTION

𝑇
Suppose 𝑿 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) be a p-dimensional random vector of p-random variables.

𝑎11 𝑎12 ⋯ 𝑎1𝑝


𝑇 𝑎21 𝑎22 ⋯ 𝑎2𝑝
Let 𝒃 = (𝑏1 , 𝑏2 , ⋯ , 𝑏𝑝 ) be a real vector and 𝑨 = ((𝑎𝑖𝑗 )) = ( ⋮ ⋮ ⋱ ⋮ ) ; 𝑖, 𝑗 = 1(1)𝑝 be a 𝑝 × 𝑝 real
𝑎𝑝1 𝑎𝑝2 ⋯ 𝑎𝑝𝑝
symmetric and positive definite (p.d.) matrix.

𝑇
The p-dimensional random vector 𝑿 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) have multivariate normal distribution if their joint probability
density function can be written in the form

1
𝑓𝑿 (𝒙) = 𝐾 × exp {− (𝒙 − 𝒃)𝑇 𝑨(𝒙 − 𝒃)} ; 𝒙 ∈ ℝ𝑝 , 𝒃 ∈ ℝ𝑝 , 𝑨 is 𝑝. 𝑑.
2
𝑝 𝑝
1
= 𝐾 × exp [− ∑ ∑{𝑎𝑖𝑗 (𝑥𝑖 − 𝑏𝑖 )(𝑥𝑗 − 𝑏𝑗 )}]
2
𝑖=1 𝑗=1

𝑇
where 𝒙 = (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ∈ ℝ𝑝 and 𝐾(> 0) is a suitable constant.

The constant 𝐾(> 0) is such that

∞ ∞ ∞

∫ ∫ ⋯ ∫ 𝑓𝑿 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) 𝑑𝑥1 𝑑𝑥2 ⋯ 𝑑𝑥𝑝 = 1 (or, ∫ 𝑓𝑿 (𝒙)𝑑𝒙 = 1)


−∞ −∞ −∞ ℝ𝑝

∞ ∞ ∞
1
⟹ 𝐾 × ∫ ∫ ⋯ ∫ exp {− (𝒙 − 𝒃)𝑇 𝑨(𝒙 − 𝒃)} 𝑑𝑥1 𝑑𝑥2 ⋯ 𝑑𝑥𝑝 = 1
2
−∞ −∞ −∞

Since 𝑨 is a positive definite, there exist a non-singular matrix 𝑪 such that 𝑪𝑪𝑇 = 𝑨.

Let us consider the transformation 𝒀 = 𝑪𝑇 (𝑿 − 𝒃).

The jacobian of the transformation is

𝑿 1 1 −1 1 −1
|𝐽| = |𝐽 ( )| = 𝑇 = = (√|𝑪𝑪𝑇 |) = = (√|𝑨|) .
𝒀 |𝑪 | √|𝑪𝑪𝑇 | √|𝑨|

𝑇
Now, (𝑿 − 𝒃)𝑇 𝑨(𝑿 − 𝒃) = (𝑿 − 𝒃)𝑇 𝑪𝑪𝑇 (𝑿 − 𝒃) = (𝑪𝑇 (𝑿 − 𝒃)) (𝑪𝑇 (𝑿 − 𝒃)) = 𝒀𝑇 𝒀.

1
∴ 𝐾 × ∫ exp (− (𝒙 − 𝒃)𝑇 𝑨(𝒙 − 𝒃)) 𝑑𝒙 = 1
2
ℝ𝑝

−1 1 𝑇
⟹ 𝐾 ⋅ (√|𝑨|) × ∫ exp (− 𝒚 𝒚) 𝑑𝒚 = 1
2
ℝ𝑝

∞ ∞ ∞ 𝑝
−1 1
⟹ 𝐾 ⋅ (√|𝑨|) × ∫ ∫ ⋯ ∫ exp (− ∑ 𝑦𝑖2 ) 𝑑𝑦1 𝑑𝑦2 ⋯ 𝑑𝑦𝑝 = 1
2
−∞ −∞ −∞ 𝑖=1

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 1


MULTIVARIATE NORMAL DISTRIBUTION

∞ ∞ ∞ 𝑝
−1 1 2
⟹ 𝐾 ⋅ (√|𝑨|) × ∫ ∫ ⋯ ∫ {∏ exp (− 𝑦 )} 𝑑𝑦1 𝑑𝑦2 ⋯ 𝑑𝑦𝑝 = 1
2 𝑖
−∞ −∞ −∞ 𝑖=1

𝑝 ∞
−1 1 2
⟹ 𝐾 ⋅ (√|𝑨|) × ∏ { ∫ exp (− 𝑦 ) 𝑑𝑦𝑖 } = 1
2 𝑖
𝑖=1 −∞

𝑝
−1
⟹ 𝐾 ⋅ (√|𝑨|) × ∏(√2𝜋) = 1 [Since 𝑌𝑖 ~𝑖𝑖𝑑 𝑁1 (0, 1) ; 𝑖 = 1(1)𝑝]
𝑖=1

√|𝑨|
⟹ 𝐾= 𝑝 .
(2𝜋)2

The joint density function (p.d.f.) is given by

√|𝑨| 1
𝑓𝑿 (𝒙) = 𝑝 × exp {− (𝒙 − 𝒃)𝑇 𝑨(𝒙 − 𝒃)} ; 𝒙 ∈ ℝ𝑝 , 𝒃 ∈ ℝ𝑝 , 𝑨 is 𝑝. 𝑑.
(2𝜋)2 2

IN TERMS OF MOMENTS:

Let us define,

𝑇 𝑇
Mean vector of 𝑿 = 𝐸(𝑿) = (𝐸(𝑋1 ), 𝐸(𝑋2 ), ⋯ , 𝐸(𝑋𝑝 )) = (𝜇1 , 𝜇2 , ⋯ , 𝜇𝑝 ) = 𝝁 (𝑠𝑎𝑦) (= ∫ 𝒙 ∙ 𝑓𝑿 (𝒙)𝑑𝒙),
ℝ𝑝

Variance − Covariance matrix of 𝑿 = Dispersion matrix of 𝑿 = 𝑫(𝑿) = 𝐸{(𝑿 − 𝝁)(𝑿 − 𝝁)𝑇 }

(𝑋1 − 𝜇1 )
(𝑋2 − 𝜇2 )
=𝐸 ( ) ((𝑋1 − 𝜇1 ) (𝑋2 − 𝜇2 ) ⋯ (𝑋𝑝 − 𝜇𝑝 ))

{ (𝑋𝑝 − 𝜇𝑝 ) }

𝜎11 𝜎12 ⋯ 𝜎1𝑝


𝜎21 𝜎22 ⋯ 𝜎2𝑝
=( ⋮ ⋮ ⋱ ⋮ ) = ((𝜎𝑖𝑗 ))𝑝×𝑝 = 𝚺 (say) (= ∫{(𝒙 − 𝝁)(𝒙 − 𝝁)𝑇 } ∙ 𝑓𝑿 (𝒙)𝑑𝒙) ,
𝜎𝑝1 𝜎𝑝2 ⋯ 𝜎𝑝𝑝 ℝ𝑝

𝜌11 𝜌12 ⋯ 𝜌1𝑝


𝜌21 𝜌22 ⋯ 𝜌2𝑝
Correlation matrix of 𝑿 = 𝐶𝑜𝑟𝑟(𝑿) = ((𝜌𝑖𝑗 ))
𝑝×𝑝
=( ⋮ ⋮ ⋱ ⋮ ) = 𝑹 (say),
𝜌𝑝1 𝜌𝑝2 ⋯ 𝜌𝑝𝑝

where 𝜇𝑖 = 𝐸(𝑋𝑖 ); 𝑖 = 1(1)𝑝,

𝐶𝑜𝑣(𝑋𝑖 , 𝑋𝑗 ) = 𝐸{(𝑋𝑖 − 𝜇𝑖 )(𝑋𝑗 − 𝜇𝑗 )} if 𝑖 ≠ 𝑗


𝜎𝑖𝑗 = { 𝑖, 𝑗 = 1(1)𝑝,
𝑉𝑎𝑟(𝑋𝑖 ) = 𝐸(𝑋𝑖 − 𝜇𝑖 )2 if 𝑖 = 𝑗,

𝐶𝑜𝑣(𝑋𝑖 , 𝑋𝑗 )
𝐶𝑜𝑟𝑟(𝑋𝑖 , 𝑋𝑗 ) = if 𝑖 ≠ 𝑗
𝜌𝑖𝑗 = √𝑉𝑎𝑟(𝑋𝑖 ) ∙ √𝑉𝑎𝑟(𝑋𝑗 ) and 𝜎𝑖𝑗 = 𝜌𝑖𝑗 𝜎𝑖 𝜎𝑗 ; 𝑖, 𝑗 = 1(1)𝑝.

{𝐶𝑜𝑟𝑟(𝑋𝑖 , 𝑋𝑗 ) = 1 if 𝑖 = 𝑗,

Let us define 𝑅𝑖𝑗 = cofactor of 𝜌𝑖𝑗 in 𝑹 ; 𝑖, 𝑗 = 1(1)𝑝.

We know that minor of 𝜌𝑖𝑗 in 𝑹 = (−1)𝑖+𝑗 × cofactor of 𝜌𝑖𝑗 in 𝑹 = (−1)𝑖+𝑗 𝑅𝑖𝑗 .

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 2


MULTIVARIATE NORMAL DISTRIBUTION

We determine the value of 𝒃 and 𝑨 in terms of moments of 𝑿.

𝑇 1
𝝁 = 𝐸(𝑿) = (𝜇1 , 𝜇2 , ⋯ , 𝜇𝑝 ) = ∫ 𝒙 ∙ 𝑓𝑿 (𝒙)𝑑𝒙 = 𝐾 × ∫ 𝒙 ∙ exp (− (𝒙 − 𝒃)𝑇 𝑨(𝒙 − 𝒃)) 𝑑𝒙 , and
2
ℝ𝑝 ℝ𝑝

𝚺 = 𝑫(𝑿) = 𝐸{(𝑿 − 𝝁)(𝑿 − 𝝁)𝑇 } = ∫{(𝒙 − 𝝁)(𝒙 − 𝝁)𝑇 } ∙ 𝑓𝑿 (𝒙)𝑑𝒙


ℝ𝑝

Since 𝑨 is a positive definite, there exist a non-singular matrix 𝑪 such that 𝑪𝑪𝑇 = 𝑨.

Let us consider the transformation 𝒀 = 𝑪𝑇 (𝑿 − 𝒃).

The jacobian of the transformation is

𝑿 1 1 −1 1 −1
|𝐽| = |𝐽 ( )| = = = (√|𝑪𝑪𝑇 |) = = (√|𝑨|) .
𝒀 ||𝑪𝑇 || √|𝑪𝑪𝑇 | √|𝑨|

𝑇
Now, (𝑿 − 𝒃)𝑇 𝑨(𝑿 − 𝒃) = (𝑿 − 𝒃)𝑇 𝑪𝑪𝑇 (𝑿 − 𝒃) = (𝑪𝑇 (𝑿 − 𝒃)) (𝑪𝑇 (𝑿 − 𝒃)) = 𝒀𝑇 𝒀.

1
𝝁 = 𝐸(𝑿) = 𝐾 × ∫ 𝑥 ∙ exp (− (𝒙 − 𝒃)𝑇 𝑨(𝒙 − 𝒃)) 𝑑𝒙
2
ℝ𝑝

1 𝑇 −1
= 𝐾 × ∫ {(𝑪𝑇 )−1 𝒚 + 𝒃} ∙ exp (− 𝒚 𝒚) ∙ (√|𝑨|) 𝑑𝒚
2
ℝ𝑝

(𝑪𝑇 )−1 1 𝑇 𝒃 1 𝑇 √|𝑨|


= 𝑝 × ∫ 𝒚 ∙ exp (− 𝒚 𝒚) 𝑑𝒚 + 𝑝 × ∫ exp (− 𝒚 𝒚) 𝑑𝒚 [Since 𝐾 = 𝑝]
(2𝜋) 2 2 (2𝜋)2 2 (2𝜋)2
ℝ𝑝 ℝ𝑝

= (𝑪𝑇 )−1 ∙ 𝐸(𝒀) + 𝒃 = 𝒃 [Since 𝒀 ~ 𝑵𝒑 (𝟎, 𝑰𝒑 ) 𝑖. 𝑒. , 𝑌𝑖′ s are 𝑖𝑖𝑑 𝑁(0,1); 𝑖 = 1(1)𝑝 and 𝐸(𝒀) = 𝟎]

∴ 𝒀 = 𝑪𝑇 (𝑿 − 𝝁)

𝚺 = 𝑫(𝑿) = 𝐸{(𝑿 − 𝝁)(𝑿 − 𝝁)𝑇 }

= 𝐸{(𝑪𝑇 )−1 𝒀𝑇 𝒀𝑪−1 }

= (𝑪𝑇 )−1 𝐸(𝒀𝑇 𝒀)𝑪−1

= (𝑪𝑇 )−1 𝑰𝒑 𝑪−1 = (𝑪𝑪𝑇 )−1 = 𝑨−1 ; [Since 𝒀 ~ 𝑵𝒑 (𝟎, 𝑰𝒑 ); 𝐷(𝒀) = 𝑰𝒑 ]

∴ 𝚺 = 𝑨−1 𝑖. 𝑒. , 𝑨 = 𝚺 −1 .

1
∴ 𝐾= 𝑝 , 𝒃=𝝁 and 𝑨 = 𝚺 −1 .
(2𝜋) 2 √|𝚺|

We denote 𝑿 ~ 𝑵𝒑 (𝝁, 𝚺) and the density function (p.d.f.) of 𝑿 is given by

1 1
𝑓𝑿 (𝒙) = 𝑝 × exp {− (𝒙 − 𝝁)𝑇 𝚺 −1 (𝒙 − 𝝁)} ; 𝒙 ∈ ℝ𝑝 , 𝝁 ∈ ℝ𝑝 , 𝚺 is 𝑝. 𝑑.
(2𝜋)2 √|𝚺| 2

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 3


MULTIVARIATE NORMAL DISTRIBUTION

NOTE:

𝑥−𝜇 2
The term ( ) = (𝑥 − 𝜇)(𝜎 2 )−1 (𝑥 − 𝜇) in the exponent of the univariate normal density function measures the square
𝜎

of the distance from 𝑥 to 𝜇 in standard deviation units. This can be generalized for a 𝑝 × 1 vector 𝒙𝑝×1 of observations
on several variables as (𝒙 − 𝝁)𝑇 𝚺 −1 (𝒙 − 𝝁).

The 𝑝 × 1 vector 𝝁 represents the expected value of the random vector 𝑿, and the 𝑝 × 𝑝 matrix 𝚺 is the variance-
covariance matrix of 𝑿. We shall assume that the symmetric matrix 𝚺 is positive definite, so the expression
(𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) is the square of the generalized distance from 𝒙 to 𝝁, or the Mahalanobis distance.
𝑝
The constant (2𝜋) 2 √|𝚺| makes the volume under the surface of the multivariate density function unity for any 𝑝.

MODE:

The 𝑝-variate normal density has a maximum value when the squared distance (𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) = 0, i.e.,
when 𝑿 = 𝝁. Thus, 𝝁 is the point of maximum density, or mode, as well as the expected value or mean of 𝑿.

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 4


MULTIVARIATE NORMAL DISTRIBUTION

RESULT 1:

The random vector 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) iff 𝒀 = 𝑪−1 (𝑿 − 𝝁) ~ 𝑵𝒑 (𝟎, 𝑰𝒑 ) i. e. , 𝑌𝑖 ′s are iid 𝑁(0, 1), where 𝑪𝑝×𝑝 is
non-singular matrix such that 𝑪𝑪𝑇 = 𝚺.

PROOF:
𝑇
Suppose 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺).

The density function (p.d.f.) of 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) is given by

1 1
𝑓𝑿 (𝒙) = 𝑝 × exp {− (𝒙 − 𝝁)𝑇 𝚺 −1 (𝒙 − 𝝁)} ; 𝒙 ∈ ℝ𝑝 , 𝝁 ∈ ℝ𝑝 , 𝚺 is 𝑝. 𝑑.
(2𝜋)2 √|𝚺| 2

Since 𝚺 is a positive definite, there exist a non-singular matrix 𝑪 such that 𝑪𝑪𝑇 = 𝚺.

Let us consider the transformation 𝒀 = 𝑪−1 (𝑿 − 𝝁) or, 𝑿 = 𝑪𝒀 + 𝝁.

The jacobian of the transformation is

𝑿
|𝐽| = |𝐽 ( )| = ||𝑪|| = √|𝑪𝑪𝑇 | = √|𝚺|.
𝒀
𝑇
∴ (𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) = (𝑿 − 𝝁)𝑇 (𝑪𝑪𝑇 )−1 (𝑿 − 𝝁) = (𝑪−1 (𝑿 − 𝝁)) (𝑪−1 (𝑿 − 𝝁)) = 𝒀𝑇 𝒀 [∵ (𝑪−1 )𝑇 = (𝑪𝑇 )−1 ]

The density function (p.d.f.) of 𝒀𝑝×1 is given by

1 1
𝑔𝒀 (𝒚) = 𝑝 × exp (− 𝒚𝑇 𝒚) × √|𝚺| ; 𝒚 ∈ ℝ𝑝 , 𝚺 is 𝑝. 𝑑.
(2𝜋)2 √|𝚺| 2

𝑝
1 1 2
= 𝑝 ∙ exp (− ∑ 𝑦𝑖 )
(2𝜋)2 2
𝑖=1

𝑝
1 1
= ∏{ ∙ exp (− 𝑦𝑖2 )}
√2𝜋 2
𝑖=1

Hence, 𝒀 = 𝑪−1 (𝑿 − 𝝁) ~ 𝑵𝒑 (𝟎, 𝑰𝒑 ) i.e.,𝑌𝑖′ s are independent and identically distributed with 𝑌𝑖 ~ 𝑁1 (0, 1); 𝑖 = 1(1)𝑝.

Suppose 𝒀𝑝×1 ~ 𝑵𝒑 (𝟎, 𝑰𝒑 ) i. e. , 𝑌𝑖 ′s are iid 𝑁(0, 1), where 𝑪𝑝×𝑝 is non-singular matrix such that 𝑪𝑪𝑇 = 𝚺.

𝑇
The density function (p.d.f.) of 𝒀𝑝×1 = (𝑌1 , 𝑌2 , ⋯ , 𝑌𝑝 ) ~ 𝑵𝒑 (𝟎, 𝑰𝒑 ) is given by

𝑝 𝑝
1 1 𝑇 1 1 2
1 1
𝑓𝒀 (𝒚) = 𝑝 × exp {− 𝒚 𝑦} = 𝑝 ∙ exp (− ∑ 𝑦𝑖 ) = ∏ { ∙ exp (− 𝑦𝑖2 )} ; 𝒚 ∈ ℝ𝑝
(2𝜋) 2 2 (2𝜋) 2 2 √2𝜋 2
𝑖=1 𝑖=1

Since 𝚺 is a positive definite, there exist a non-singular matrix 𝑪 such that 𝑪𝑪𝑇 = 𝚺.

Let us consider the transformation 𝑿 = 𝑪𝒀 + 𝝁, where 𝑪 is a non-singular matrix such that 𝑪𝑪𝑇 = 𝚺.

The jacobian of the transformation 𝒀 = 𝑪−1 (𝑿 − 𝝁) is

𝒀 1 1 1
|𝐽| = |𝐽 ( )| = ||𝑪−1 || = = = .
𝑿 ||𝑪|| √|𝑪𝑪 | √|𝚺|
𝑇

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 5


MULTIVARIATE NORMAL DISTRIBUTION

𝑇
∴ 𝒀𝑇 𝒀 = (𝑪−1 (𝑿 − 𝝁)) (𝑪−1 (𝑿 − 𝝁)) = (𝑿 − 𝝁)𝑇 (𝑪𝑪𝑇 )−1 (𝑿 − 𝝁) = (𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) [∵ (𝑪−1 )𝑇 = (𝑪𝑇 )−1 ].

The density function (p.d.f.) of 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) is given by

1 1 1
𝑓𝑿 (𝒙) = 𝑝 × exp {− (𝒙 − 𝝁)𝑇 𝚺 −1 (𝒙 − 𝝁)} × ; 𝒙 ∈ ℝ𝑝 , 𝝁 ∈ ℝ𝑝 , 𝚺 is 𝑝. 𝑑.
(2𝜋)2 2 √|𝚺|

1 1
= 𝑝 × exp {− (𝒙 − 𝝁)𝑇 𝚺 −1 (𝒙 − 𝝁)}
(2𝜋)2 √|𝚺| 2

Hence, 𝑿𝑝×1 = 𝑪𝒀 + 𝝁 ~ 𝑵𝒑 (𝝁, 𝚺), where 𝒀 ~ 𝑵𝒑 (𝟎, 𝑰𝒑 ) i. e. , 𝑌𝑖 ′s are iid 𝑁(0, 1) and 𝑪𝑝×𝑝 is non-singular matrix such
that 𝑪𝑪𝑇 = 𝚺.

OBSERVATION 1:

Suppose 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺). If 𝚺 = 𝐃𝐢𝐚𝐠(𝜎𝑖𝑖 ; 𝑖 = 1(1)𝑝) = 𝐃𝐢𝐚𝐠(𝜎11 , 𝜎22 , ⋯ , 𝜎𝑝𝑝 ), then 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 are all
mutually independent and 𝑋𝑖 ~ 𝑁1 (𝜇𝑖 , 𝜎𝑖𝑖 ); 𝑖 = 1(1)𝑝.

PROOF:
𝑇
Suppose 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺). The density function (p.d.f.) of 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) is given by

1 1
𝑓𝑿 (𝒙) = 𝑝 × exp {− (𝒙 − 𝝁)𝑇 𝚺 −1 (𝒙 − 𝝁)} ; 𝒙 ∈ ℝ𝑝 , 𝝁 ∈ ℝ𝑝 , 𝚺 is 𝑝. 𝑑.
(2𝜋)2 √|𝚺| 2

𝜎11 0 ⋯ 0
0 𝜎22 ⋯ 0 0 if 𝑖 ≠ 𝑗
Since 𝚺 = 𝐃𝐢𝐚𝐠(𝜎11 , 𝜎22 , ⋯ , 𝜎𝑝𝑝 ) = ( ⋱ ⋮ ), where 𝜎𝑖𝑗 = 𝐶𝑜𝑣(𝑋𝑖 , 𝑋𝑗 ) = { 𝜎𝑖𝑖 if 𝑖 = 𝑗, 𝑖, 𝑗 = 1(1)𝑝.
⋮ ⋮
0 0 ⋯ 𝜎𝑝𝑝

𝑝
−1
1 1 1
∴ 𝚺 = 𝐃𝐢𝐚𝐠 ( , ,⋯, ) and |𝚺| = ∏ 𝜎𝑖𝑖 .
𝜎11 𝜎22 𝜎𝑝𝑝
𝑖=1

𝑝
𝑇 −1 (𝑿
1
Now, (𝑿 − 𝝁) 𝚺 − 𝝁) = ∑ (𝑋 − 𝜇𝑖 )2 .
𝜎𝑖𝑖 𝑖
𝑖=1

𝑇
The density function (p.d.f.) of 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺) is given by

𝑝
1 1 1
𝑓𝑿 (𝒙) = × exp {− ∑ (𝑋𝑖 − 𝜇𝑖 )2 }
𝑝 𝑝 2 𝜎𝑖𝑖
(√2𝜋) √∏𝑖=1 𝜎𝑖𝑖 𝑖=1

𝑝 𝑝
1 1
= ∏[ × exp {− (𝑋𝑖 − 𝜇𝑖 )2 }] = ∏{𝑓𝑋𝑖 (𝑥𝑖 )} ;
√2𝜋√𝜎𝑖𝑖 2 𝜎 𝑖𝑖
𝑖=1 𝑖=1

where 𝑓𝑋𝑖 (𝑥𝑖 ) be the 𝑝. 𝑑. 𝑓. of 𝑋𝑖 ~ 𝑁1 (𝜇𝑖 , 𝜎𝑖𝑖 ); 𝑖 = 1(1)𝑝.

Hence, 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 are all mutually independent and 𝑋𝑖 ~ 𝑁1 (𝜇𝑖 , 𝜎𝑖𝑖 ); 𝑖 = 1(1)𝑝.

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 6


MULTIVARIATE NORMAL DISTRIBUTION

OBSERVATION 2:

If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝜎 2 𝑰𝒑 ), then 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 are all mutually independent and 𝑋𝑖 ~ 𝑁1 (𝜇𝑖 , 𝜎 2 ); 𝑖 = 1(1)𝑝.

PROOF:

0 if 𝑖 ≠ 𝑗
Same as OBSERVATION 1, where 𝜎𝑖𝑗 = 𝐶𝑜𝑣(𝑋𝑖 , 𝑋𝑗 ) = { 𝑖, 𝑗 = 1(1)𝑝.
𝜎𝑖𝑖 = 𝜎 2 if 𝑖 = 𝑗,

MOMENT GENERATING FUNCTION (M.G.F.):

𝑇
The M.G.F. of 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺), for a set of arbitrary real vector 𝒕𝑝×1 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 ) is given by

𝑴𝑿 (𝒕) = 𝐸{exp(𝒕𝑇 𝑿)}

1 1
= 𝑝 × ∫ exp(𝒕𝑇 𝒙) ∙ exp {− (𝒙 − 𝝁)𝑇 𝚺 −1 (𝒙 − 𝝁)} 𝑑𝒙
(2𝜋)2 √|𝚺| 2
ℝ𝑝

1 1
= 𝑝 × ∫ exp {− (𝒙𝑇 𝚺 −1 𝒙 − 2𝒙𝑇 𝚺 −1 𝝁 + 𝝁𝑇 𝚺 −1 𝝁 − 2𝒕𝑇 𝒙)} 𝑑𝒙 [∵ 𝒙𝑇 𝚺 −1 𝝁 = 𝝁𝑇 𝚺 −1 𝒙]
(2𝜋)2 √|𝚺| 2
ℝ𝑝

1 1
= 𝑝 × ∫ exp [− {𝒙𝑇 𝚺 −1 𝒙 − 2𝒙𝑇 𝚺 −1 (𝝁 + 𝚺𝒕) + 𝝁𝑇 𝚺 −1 𝝁}] 𝑑𝒙 [∵ 𝒕𝑇 𝒙 = 𝒙𝑇 𝒕]
(2𝜋)2 √|𝚺| 2
ℝ𝑝

1 1
= 𝑝 × ∫ exp [− {𝒙𝑇 𝚺 −1 𝒙 − 2𝒙𝑇 𝚺 −1 (𝝁 + 𝚺𝒕) + (𝝁 + 𝚺𝒕)𝑇 𝚺 −1 (𝝁 + 𝚺𝒕)}]
(2𝜋)2 √|𝚺| 2
ℝ𝑝
1
× exp [− {𝝁𝑇 𝚺 −1 𝝁 − (𝝁 + 𝚺𝒕)𝑇 𝚺 −1 (𝝁 + 𝚺𝒕)}] 𝑑𝒙
2

1
= exp [− (𝝁𝑇 𝚺 −1 𝝁 − 𝝁𝑇 𝚺 −1 𝝁 − 𝝁𝑇 𝚺 −1 𝚺𝒕 − 𝒕𝑇 𝚺 𝑇 𝚺 −1 𝝁 − 𝒕𝑇 𝚺 𝑇 𝚺 −1 𝚺𝒕)]
2

1 1 𝑇
× 𝑝 ∫ exp [− (𝒙 − (𝝁 + 𝚺𝒕)) 𝚺 −1 (𝒙 − (𝝁 + 𝚺𝒕))] 𝑑𝒙
(2𝜋)2 √|𝚺| ℝ𝑝 2

1
= exp (𝒕𝑇 𝝁 + 𝒕𝑇 𝚺𝒕) × ∫ 𝑓𝑿 {𝒙|𝑿 ~ 𝑵𝒑 ((𝝁 + 𝚺𝒕), 𝚺)}𝑑𝒙
2
ℝ𝑝

1
= exp (𝒕𝑇 𝝁 + 𝒕𝑇 𝚺𝒕)
2
𝑇 𝝁+1𝒕𝑇 𝚺𝒕
= 𝑒𝒕 2 .

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 7


MULTIVARIATE NORMAL DISTRIBUTION

NOTE:
1 1 𝑇
1) The moment generating function for (𝑿 − 𝝁) is 𝑴𝑿−𝝁 (𝒕) = exp ( 𝒕𝑇 𝚺𝒕) = 𝑒 2𝒕 𝚺𝒕
.
2

2) Two important properties of moment generating functions.

(i) If two random vectors have the same moment generating function, they have the same density.
(ii) Random vectors are independent if and only if their joint moment generating function factors into the
𝑇
product of their separate moment generating functions; that is, if 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) and any real
𝑇
𝒕𝑝×1 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 ) , then 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 are independent if and only if

𝑴𝑿 (𝒕) = 𝑀𝑋1 (𝑡1 )𝑀𝑋2 (𝑡2 ) ⋯ 𝑀𝑋𝑝 (𝑡𝑝 ).

CHARACTERISTIC FUNCTION (C.F.):

𝑇
The C.F. of 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺), for a set of arbitrary real number 𝒕𝑝×1 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 ) is given by

1
𝝓𝑿 (𝒕) = 𝐸{exp(𝑖𝒕𝑇 𝑿)} = exp (𝑖𝒕𝑇 𝝁 − 𝒕𝑇 𝚺𝒕) .
2

RESULT 2:

If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) and 𝒀𝑝×1 = 𝑪𝑿 + 𝒃 where 𝑪𝑝×𝑝 is non-singular matrix and 𝒃 is 𝑝 × 1 vector, then
𝒀 ~ 𝑵𝑝 (𝑪𝝁 + 𝒃, 𝑪𝚺𝑪𝑇 ).

PROOF:
𝑇
Suppose 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺). The density function (p.d.f.) of 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) is given by

1 1
𝑓𝑿 (𝒙) = 𝑝 × exp {− (𝒙 − 𝝁)𝑇 𝚺 −1 (𝒙 − 𝝁)} ; 𝒙 ∈ ℝ𝑝 , 𝝁 ∈ ℝ𝑝 , 𝚺 is 𝑝. 𝑑.
(2𝜋)2 √|𝚺| 2

Let us consider the transformation 𝒀𝑝×1 = 𝑪𝑝×𝑝 𝑿𝑝×1 + 𝒃𝑝×1 i.e., 𝑿 = 𝑪−1 (𝒀 − 𝒃).

Since 𝚺 is a positive definite and 𝑪𝑝×𝑝 is non-singular matrix, 𝑪𝚺𝑪𝑇 is also positive definite.

The jacobian of the transformation is

𝑿 1 1 |𝚺| √|𝚺|
|𝐽| = |𝐽 ( )| = ||𝑪−1 || = = =√ = .
𝒀 ||𝑪|| √|𝑪||𝑪𝑇 | |𝑪| ∙ |𝚺| ∙ |𝑪𝑇 | √|𝑪𝚺𝑪𝑇 |

Now, (𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) = (𝑪−1 (𝒀 − 𝒃) − 𝝁)𝑇 𝚺 −1 (𝑪−1 (𝒀 − 𝒃) − 𝝁)

= (𝑪−1 𝒀 − 𝑪−1 𝒃 − 𝝁)𝑇 𝚺 −1 (𝑪−1 𝒀 − 𝑪−1 𝒃 − 𝝁)

𝑇
= (𝑪−1 (𝒀 − (𝑪𝝁 + 𝒃))) 𝚺 −1 (𝑪−1 (𝒀 − (𝑪𝝁 + 𝒃)))

𝑇
= (𝒀 − (𝑪𝝁 + 𝒃)) (𝑪𝚺𝑪𝑇 )−1 (𝒀 − (𝑪𝝁 + 𝒃)) [∵ (𝑪−1 )𝑇 = (𝑪𝑇 )−1 ]

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 8


MULTIVARIATE NORMAL DISTRIBUTION

The density function (p.d.f.) of 𝒀𝑝×1 is given by

1 1 𝑇 √|𝚺|
𝑔𝒀 (𝒚) = 𝑝 ∙ exp {− (𝒚 − (𝑪𝝁 + 𝒃)) (𝑪𝚺𝑪𝑇 )−1 (𝒚 − (𝑪𝝁 + 𝒃))} ∙ ; 𝒚 ∈ ℝ𝑝 , 𝑪𝚺𝑪𝑇 is 𝑝. 𝑑.
(2𝜋)2 √|𝚺| 2 √|𝑪𝚺𝑪𝑇 |

1 1 𝑇
= 𝑝 ∙ exp {− (𝒚 − (𝑪𝝁 + 𝒃)) (𝑪𝚺𝑪𝑇 )−1 (𝒚 − (𝑪𝝁 + 𝒃))}
(2𝜋)2 √|𝑪𝚺𝑪𝑇 | 2

⟶ Density function of 𝑵𝑝 (𝑪𝝁 + 𝒃, 𝑪𝚺𝑪𝑇 )

Hence, 𝒀𝑝×1 = 𝑪𝑿 + 𝒃 ~ 𝑵𝑝 (𝑪𝝁 + 𝒃, 𝑪𝚺𝑪𝑇 ) where 𝑪 is 𝑝 × 𝑝 non-singular matrix and 𝒃 is 𝑝 × 1 vector.

ALTERNATIVE

Let 𝒀𝑝×1 = 𝑪𝑿 + 𝒃.

𝑇
For any real 𝒕𝑝×1 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 ) , the moment generating function (M.G.F.) of 𝒀 is given by

𝑴𝒀 (𝒕) = 𝐸{exp(𝒕𝑇 𝒀)}

= 𝐸[exp{𝒕𝑇 (𝑪𝑿 + 𝒃)}]

= exp(𝒕𝑇 𝒃) ∙ 𝐸{exp(𝒕𝑇 𝑪𝑿)}

= exp(𝒕𝑇 𝒃) ∙ 𝐸{exp(𝑺𝑇 𝑿)} ; [where 𝑺𝑝×1 = 𝑪𝑇 𝒕 ]

1 1
= exp(𝒕𝑇 𝒃) ∙ exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺) ; [∵ If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺), 𝑴𝑿 (𝑺) = exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺)]
2 2

1
= exp (𝒕𝑇 (𝑪𝝁 + 𝒃) + 𝒕𝑇 (𝑪𝚺𝑪𝑇 )𝒕)
2

⟶ M.G.F. of 𝑵𝑝 (𝑪𝝁 + 𝒃, 𝑪𝚺𝑪𝑇 )

Since 𝚺 is a positive definite and 𝑪𝑝×𝑝 is non-singular matrix, 𝑪𝚺𝑪𝑇 is also positive definite.

By uniqueness property of moment generating function (M.G.F.), 𝑴𝒀 (𝒕) is the M.G.F. of 𝑵𝑝 (𝑪𝝁 + 𝒃, 𝑪𝚺𝑪𝑇 ) i.e., the
distribution of 𝒀𝑝×1 = 𝑪𝑿 + 𝒃 is 𝑵𝑝 (𝑪𝝁 + 𝒃, 𝑪𝚺𝑪𝑇 ).

Hence, 𝒀𝑝×1 = 𝑪𝑿 + 𝒃 ~ 𝑵𝑝 (𝑪𝝁 + 𝒃, 𝑪𝚺𝑪𝑇 ).

OBSERVATION 3:

If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) and 𝒀𝑝×1 = 𝑪𝑿 where 𝑪𝑝×𝑝 is non-singular matrix, then 𝒀 ~ 𝑵𝑝 (𝑪𝝁, 𝑪𝚺𝑪𝑇 ).

PROOF:

Same as RESULT 1, where 𝒃𝑝×1 = 𝟎𝑝×1 = (0, 0, ⋯ , 0)𝑇 .

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 9


MULTIVARIATE NORMAL DISTRIBUTION

OBSERVATION 4:

If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝜎 2 𝑰𝒑 ) and 𝒀𝑝×1 = 𝑳𝑿 where 𝑳𝑝×𝑝 is an orthogonal matrix, then 𝒀 ~ 𝑵𝑝 (𝑳𝝁, 𝜎 2 𝑰𝒑 ).

PROOF:

Let 𝒀𝑝×1 = 𝑳𝑿 and 𝚺 = 𝜎 2 𝑰𝒑 .

𝑇
For any real 𝒕𝑝×1 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 ) , the M.G.F. of 𝒀 is given by

𝑴𝒀 (𝒕) = 𝐸{exp(𝒕𝑇 𝒀)}

= 𝐸{exp(𝒕𝑇 𝑳𝑿)}

= 𝐸{exp(𝑺𝑇 𝑿)} ; [where 𝑺𝑝×1 = 𝑳𝑇 𝒕 ]

1 1
= exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺) ; [∵ If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺), 𝑴𝑿 (𝑺) = exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺)]
2 2

1
= exp (𝒕𝑇 𝑳𝝁 + 𝒕𝑇 (𝑳𝚺𝑳𝑇 )𝒕)
2

⟶ M.G.F. of 𝑵𝑝 (𝑳𝝁, 𝑳𝚺𝑳𝑇 )

Since 𝑳𝑝×𝑝 is an orthogonal matrix and 𝚺 = 𝜎 2 𝑰𝒑 , 𝑳𝚺𝑳𝑇 = 𝜎 2 (𝑳𝑰𝒑 𝑳𝑇 ) = 𝜎 2 (𝑳𝑳𝑇 ) = 𝜎 2 𝑰𝒑 .

By uniqueness property of moment generating function (M.G.F.), 𝑴𝒀 (𝒕) is the M.G.F. of 𝑵𝑝 (𝑳𝝁, 𝑳𝚺𝑳𝑇 ) i.e., the
distribution of 𝒀𝑝×1 = 𝑳𝑿 is 𝑵𝑝 (𝑳𝝁, 𝜎 2 𝑰𝒑 ).

Hence, 𝒀𝑝×1 = 𝑳𝑿 ~ 𝑵𝑝 (𝑳𝝁, 𝜎 2 𝑰𝒑 ).

NOTE:

1. From OBSERVATION 2, we know that, if 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝜎 2 𝑰𝒑 ), then 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 are all mutually independent and
𝑋𝑖 ~ 𝑁1 (𝜇𝑖 , 𝜎 2 ); 𝑖 = 1(1)𝑝. Since 𝒀𝑝×1 = 𝑳𝑿 ~ 𝑵𝑝 (𝑳𝝁, 𝜎 2 𝑰𝒑 ), 𝑌1 , 𝑌2 , ⋯ , 𝑌𝑝 are all mutually independent with
common variance 𝜎 2 .

The OBSERVATION 4 states that mutually independent normal variables with the same variance remain mutually
independent normal variables with the same variance under orthogonal transformation. This invariance property
is a characterisation of normal distribution.

2. If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝜎 2 𝑰𝒑 ) and 𝒀𝑝×1 = 𝑳𝑇 𝑿 where 𝑳𝑝×𝑝 is an orthogonal matrix, then 𝒀 ~ 𝑵𝑝 (𝑳𝑇 𝝁, 𝜎 2 𝑰𝒑 ).

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 10


MULTIVARIATE NORMAL DISTRIBUTION

MARGINAL AND CONDITIONAL DISTRIBUTION:

RESULT 3:

(1)
𝑇 𝑿𝑚×1
Suppose 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) = ( (2) ) ~ 𝑵𝒑 (𝝁, 𝚺),
𝑿(𝑝−𝑚)×1

(1) 𝑚×𝑚 𝑚×(𝑝−𝑚)


𝝁𝑚×1 𝚺11 𝚺12 𝑇
where 𝝁𝑝×1 = ( (2) ),𝚺 = ( (𝑝−𝑚)×𝑚 (𝑝−𝑚)×(𝑝−𝑚)
) and 𝚺12 = 𝚺21 .
𝝁(𝑝−𝑚)×1 𝚺21 𝚺22

(1)
The marginal distribution of any subset 𝑿𝑚×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑚 )𝑇 ; (𝑚 < 𝑝) of 𝑿 is 𝑚-dimensional multivariate
(2) 𝑇
normal and the marginal distribution of subset 𝑿(𝑝−𝑚)×1 = (𝑋𝑚+1 , 𝑋𝑚+2 , ⋯ , 𝑋𝑝 ) of 𝑿 is (𝑝 − 𝑚)-dimensional

multivariate normal. The marginal distributions of 𝑿(1) and 𝑿(2) , respectively, are given by

𝑿(1) ~ 𝑵𝒎 (𝝁(1) , 𝚺11 ) and 𝑿(2) ~ 𝑵(𝒑−𝒎) (𝝁(2) , 𝚺22 ).

The conditional distribution of 𝑿(1) given for fixed 𝑿(2) = 𝒙(2) is given by

(𝑿(1) |𝑿(2) = 𝒙(2) ) ~ 𝑵𝒎 ({𝝁(1) + 𝚺12 𝚺22


−1
(𝒙(2) − 𝝁(2) )}, 𝚺11.2 ), −1
where 𝚺11.2 = 𝚺11 − 𝚺12 𝚺22 𝚺21 .

The conditional distribution of 𝑿(2) given for fixed 𝑿(1) = 𝒙(1) is given by

(𝑿(2) |𝑿(1) = 𝒙(1) ) ~ 𝑵(𝒑−𝒎) ({𝝁(2) + 𝚺21 𝚺11


−1
(𝒙(1) − 𝝁(1) )}, 𝚺22.1 ), where 𝚺22.1 = 𝚺22 − 𝚺21 𝚺11
−1
𝚺12 .

PROOF:
𝑇
Suppose 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺). The density function (p.d.f.) of 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) is given by

1 1
𝑓𝑿 (𝒙) = 𝑝 × exp {− (𝒙 − 𝝁)𝑇 𝚺 −1 (𝒙 − 𝝁)} ; 𝒙 ∈ ℝ𝑝 , 𝝁 ∈ ℝ𝑝 , 𝚺 is 𝑝. 𝑑.
(2𝜋)2 √|𝚺| 2

Let us define

(1) (1) 𝑚×𝑚 𝑚×(𝑝−𝑚)


𝑿𝑚×1 𝝁𝑚×1 𝚺11 𝚺12 𝑇
𝑿𝑝×1 = ( (2) ), 𝝁𝑝×1 = ( (2) ), 𝚺=( (𝑝−𝑚)×𝑚 (𝑝−𝑚)×(𝑝−𝑚)
) and 𝚺12 = 𝚺21 ,
𝑿(𝑝−𝑚)×1 𝝁(𝑝−𝑚)×1 𝚺21 𝚺22

where 𝐸(𝑿(1) ) = 𝝁(1) , 𝐸(𝑿(2) ) = 𝝁(2) ,

𝑇
𝐸 {(𝑿(1) − 𝝁(1) )(𝑿(1) − 𝝁(1) ) } = 𝚺11 ,

𝑇
𝐸 {(𝑿 (1) − 𝝁(1) )(𝑿(2) − 𝝁(2) ) } = 𝚺12 , and

𝑇
𝐸 {(𝑿(2) − 𝝁(2) )(𝑿(2) − 𝝁(2) ) } = 𝚺22 .

Let us consider the transformation

(1) (1) (2)


𝒀𝑚×1 = 𝑿𝑚×1 + 𝑩𝑚×(𝑝−𝑚) 𝑿(𝑝−𝑚)×1 ,

(2) (2)
𝒀(𝑝−𝑚)×1 = 𝑿(𝑝−𝑚)×1 , where 𝑩𝑚×(𝑝−𝑚) is such that 𝐶𝑜𝑣(𝒀(1) , 𝒀(2) ) = 𝟎.

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 11


MULTIVARIATE NORMAL DISTRIBUTION

𝑇
Now, 𝐶𝑜𝑣(𝒀(1) , 𝒀(2) ) = 0 ⟹ 𝐸 [(𝒀(1) − 𝐸(𝒀(1) )) (𝒀(2) − 𝐸(𝒀(2) )) ] = 𝟎

𝑇
⟹ 𝐸 [{(𝑿(1) − 𝝁(1) ) + 𝑩(𝑿(2) − 𝝁(2) )}(𝑿(2) − 𝝁(2) ) ] = 𝟎

𝑇 𝑇
⟹ 𝐸 {(𝑿(1) − 𝝁(1) )(𝑿(2) − 𝝁(2) ) } + 𝑩 𝐸 {(𝑿(2) − 𝝁(2) )(𝑿(2) − 𝝁(2) ) } = 𝟎

−1
⟹ 𝚺12 + 𝑩 𝚺22 = 𝟎 ⟹ 𝑩 = −𝚺12 𝚺22 .

Hence, the transformation becomes

(1)
𝒀𝑚×1 𝑰 −1
−𝚺12 𝚺22
−1
𝒀𝑝×1 = 𝑷𝑝×𝑝 𝑿𝑝×1 i. e., 𝑿 = 𝑷 𝒀, where 𝒀𝑝×1 = ( (2) ) and 𝑷𝑝×𝑝 = ( 𝑚 ).
𝒀(𝑝−𝑚)×1 𝟎 (𝑝−𝑚)×𝑚 𝑰(𝑝−𝑚)

𝑰𝑚 −1
−𝚺12 𝚺22 𝝁(1)
Now, 𝑷𝝁 = ( ) ( (2) )
𝟎(𝑝−𝑚)×𝑚 𝑰(𝑝−𝑚) 𝝁

𝝁(1) − 𝚺12 𝚺22


−1 (2)
𝝁 𝜽(1)
=( ) = ( ) ; (say), where 𝜽(1) = 𝝁(1) − 𝚺12 𝚺22
−1 (2)
𝝁 ,
𝝁(2) 𝝁(2)

−1 −1 𝑇
𝑰 −𝚺12 𝚺22 𝚺 𝚺12 𝑰𝑚 −𝚺12 𝚺22
and 𝑷𝚺𝑷𝑇 = ( 𝑚 ) ( 11 )( )
𝟎(𝑝−𝑚)×𝑚 𝑰(𝑝−𝑚) 𝚺21 𝚺22 𝟎(𝑝−𝑚)×𝑚 𝑰(𝑝−𝑚)

−1
𝚺 − 𝚺12 𝚺22 𝚺21 𝟎𝑚×(𝑝−𝑚) 𝑰𝑚 𝟎𝑚×(𝑝−𝑚)
= ( 11 ) ( −1 )
𝚺21 𝚺22 −𝚺22 𝚺21 𝑰(𝑝−𝑚)

−1
𝚺11 − 𝚺12 𝚺22 𝚺21 𝟎𝑚×(𝑝−𝑚)
=( )
𝟎(𝑝−𝑚)×𝑚 𝚺22

𝚺11.2 𝟎𝑚×(𝑝−𝑚) −1
=( ) ; (say), where 𝚺11.2 = 𝚺11 − 𝚺12 𝚺22 𝚺21 .
𝟎(𝑝−𝑚)×𝑚 𝚺22

The jacobian of the transformation is

𝑿 1
|𝐽| = |𝐽 ( )| = ||𝑷−1 || = =1.
𝒀 ||𝑷||

𝚺 𝟎 −1 −1
𝚺11.2 𝟎
∴ (𝑷𝚺𝑷𝑇 )−1 = ( 11.2 ) ⟺ (𝑷𝑇 )−1 𝚺 −1 𝑷−1 = ( −1 ).
𝟎 𝚺22 𝟎 𝚺22

𝚺 𝟎
and |𝑷𝚺𝑷𝑇 | = | 11.2 | ⟹ |𝑷||𝚺||𝑷𝑇 | = |𝚺11.2 ||𝚺22 |
𝟎 𝚺22

⟹ |𝚺| = |𝚺11.2 ||𝚺22 |; [∵ |𝑷| = |𝑷𝑇 | = 1]

∴ (𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) = (𝑷−1 𝒀 − 𝝁)𝑇 𝚺 −1 (𝑷−1 𝒀 − 𝝁)

𝑇
= (𝑷−1 (𝒀 − 𝑷𝝁)) 𝚺 −1 (𝑷−1 (𝒀 − 𝑷𝝁))

= (𝒀 − 𝑷𝝁)𝑇 (𝑷𝚺𝑷𝑇 )−1 (𝒀 − 𝑷𝝁), [∵ (𝑷−1 )𝑇 = (𝑷𝑇 )−1 ]

−1𝑇
𝒀(1) − 𝜽(1) 𝚺11.2 𝟎 𝒀(1) − 𝜽(1)
=( ) ( −1 ) ( (2) )
𝒀(2) − 𝝁(2) 𝟎 𝚺22 𝒀 − 𝝁(2)

𝑇 𝑇
= (𝒀(1) − 𝜽(1) ) 𝚺11.2
−1
(𝒀(1) − 𝜽(1) ) + (𝒀(2) − 𝝁(2) ) 𝚺22
−1
(𝒀(2) − 𝝁(2) ).

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 12


MULTIVARIATE NORMAL DISTRIBUTION

𝑇
The density function (p.d.f.) of 𝒀𝑝×1 = (𝒀(1) , 𝒀(2) ) is given by

1 1 (1) 𝑇 𝑇
𝑔𝒀 (𝒚) = 𝑝 exp [− {(𝒚 − 𝜽(1) ) 𝚺11.2
−1
(𝒚(1) − 𝜽(1) ) + (𝒚(2) − 𝝁(2) ) 𝚺22
−1
(𝒚(2) − 𝝁(2) )}] ;
(2𝜋)2 √|𝚺11.2 ||𝚺22 | 2

[𝒚 ∈ ℝ𝑝 , 𝜽(1) ∈ ℝ𝑚 , 𝝁(2) ∈ ℝ(𝑝−𝑚) , 𝚺11.2 and 𝚺22 are 𝑝. 𝑑. ]

1 1 (1) 𝑇
=( 𝑚 exp {− (𝒚 − 𝜽(1) ) 𝚺11.2
−1
(𝒚(1) − 𝜽(1) )})
(2𝜋) 2 √|𝚺11.2 | 2

1 1 (2) 𝑇
×( 𝑝−𝑚 exp {− (𝒚 − 𝝁(2) ) 𝚺22
−1
(𝒚(2) − 𝝁(2) )})
(2𝜋) 2 √|𝚺22 | 2

Hence, 𝒀(1) and 𝒀(2) are independently distributed with

𝒀(1) ~ 𝑵𝒎 (𝜽(1) , 𝚺11.2 ) i. e., 𝒀(1) ~ 𝑵𝒎 ((𝝁(1) − 𝚺12 𝚺22


−1 (2) (𝚺 −1
𝝁 ), 11 − 𝚺12 𝚺22 𝚺21 )) , and

𝒀(2) ~ 𝑵(𝒑−𝒎) (𝝁(2) , 𝚺22 ).

ALTERNATIVE

𝑇
For any real 𝒕𝑝×1 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 ) , the M.G.F. of 𝒀 is given by

𝑴𝒀 (𝒕) = 𝐸{exp(𝒕𝑇 𝒀)} = 𝐸{exp(𝒕𝑇 𝑷𝑿)} = 𝐸{exp(𝑺𝑇 𝑿)} ; [where 𝑺𝑝×1 = 𝑷𝑇 𝒕 ]

1 1
= exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺) ; [∵ If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺), 𝑴𝑿 (𝑺) = exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺)]
2 2

1
= exp (𝒕𝑇 𝑷𝝁 + 𝒕𝑇 (𝑷𝚺𝑷𝑇 )𝒕) = 𝑀. 𝐺. 𝐹. of 𝑵𝑝 (𝑷𝝁, 𝑷𝚺𝑷𝑇 )
2

By uniqueness property of moment generating function (M.G.F.), 𝑴𝒀 (𝒕) is the M.G.F. of 𝑵𝑝 (𝑷𝝁, 𝑷𝚺𝑷𝑇 ) i.e., the
distribution of 𝒀𝑝×1 is 𝑵𝑝 (𝑷𝝁, 𝑷𝚺𝑷𝑇 ). Hence, 𝒀𝑝×1 ~ 𝑵𝑝 (𝑷𝝁, 𝑷𝚺𝑷𝑇 ).

∴ 𝐶𝑜𝑣(𝒀(1) , 𝒀(2) ) = 𝟎 𝑖. 𝑒. , 𝐶𝑜𝑣(𝒀(1) , 𝑿(2) ) = 𝟎 [∵ 𝒀(2) = 𝑿(2) ]

Hence, 𝒀(1) and 𝑿(2) are independently distributed with

𝒀(1) = 𝑿(1) − 𝚺12 𝚺22


−1 (2)
𝑿 ~ 𝑵𝒎 ((𝝁(1) − 𝚺12 𝚺22
−1 (2) (𝚺 −1
𝝁 ), 11 − 𝚺12 𝚺22 𝚺21 )) , and

𝒀(2) = 𝑿(2) ~ 𝑵(𝒑−𝒎) (𝝁(2) , 𝚺22 ).

For any real 𝒕𝑚×1 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑚 )𝑇 , the M.G.F. of 𝑿(1) is given by

𝑴𝑿(1) (𝒕) = 𝐸{exp(𝒕𝑇 𝑿(1) )} = 𝐸[exp{𝒕𝑇 (𝒀(1) + 𝚺12 𝚺22


−1 (2)
𝑿 )}]

= 𝐸[exp{𝒕𝑇 𝒀(1) }] × 𝐸[exp{𝒕𝑇 𝚺12 𝚺22


−1 (2)
𝑿 }]; [∵ 𝒀(1) and 𝑿(2) are independent]

1
= exp {𝒕𝑇 (𝝁(1) − 𝚺12 𝚺22
−1 (2)
𝝁 ) + 𝒕𝑇 (𝚺11 − 𝚺12 𝚺22
−1
𝚺21 )𝒕}
2
−1 (2)
1
× exp {𝒕𝑇 𝚺12 𝚺22 𝝁 + 𝒕𝑇 𝚺12 𝚺22
−1 −1
𝚺22 𝚺22 𝚺21 𝒕}
2

1 𝑇
= exp (𝒕𝑇 𝝁(1) + 𝒕 𝚺11 𝒕) = 𝑀. 𝐺. 𝐹. of 𝑵𝒎 (𝝁(1) , 𝚺11 )
2

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 13


MULTIVARIATE NORMAL DISTRIBUTION

By uniqueness property of moment generating function (M.G.F.), 𝑴𝑿(1) (𝒕) is the M.G.F. of 𝑵𝒎 (𝝁(1) , 𝚺11 ) i.e., the
distribution of 𝑿(1) is 𝑵𝒎 (𝝁(1) , 𝚺11 ). Hence, 𝑿(1) ~ 𝑵𝒎 (𝝁(1) , 𝚺11 ).

∴ 𝑿(1) ~ 𝑵𝒎 (𝝁(1) , 𝚺11 ) and 𝑿(2) ~ 𝑵(𝒑−𝒎) (𝝁(2) , 𝚺22 ).

The marginal density function (p.d.f.) of 𝑿(2) ~ 𝑵(𝒑−𝒎) (𝝁(2) , 𝚺22 ) is given by

1 1 𝑇 −1
𝑓𝑿(2) (𝒙(2) ) = 𝑝−𝑚 × exp {− (𝒙(2) − 𝝁(2) ) 𝚺22 (𝒙(2) − 𝝁(2) )}.
(2𝜋) 2 √|𝚺22 | 2

The joint density function (p.d.f.) of (𝑿(1) , 𝑿(2) ) can be written as

1 1 𝑇 −1
𝑓𝑿 (𝒙(1) , 𝒙(2) ) = 𝑚 𝑝−𝑚 × exp {− (𝒙(2) − 𝝁(2) ) 𝚺22 (𝒙(2) − 𝝁(2) )}
(2𝜋) 2 (2𝜋) 2 √|𝚺11.2 |√|𝚺22 | 2

1 𝑇 −1
× exp {− (𝒙(1) − {𝝁(1) + 𝚺12 𝚺22
−1
(𝒙(2) − 𝝁(2) )}) 𝚺11.2 (𝒙(1) − {𝝁(1) + 𝚺12 𝚺22
−1
(𝒙(2) − 𝝁(2) )})}.
2
−1
where 𝚺11.2 = 𝚺11 − 𝚺12 𝚺22 𝚺21 .

The conditional density function (p.d.f.) of 𝑿(1) given for fixed 𝑿(2) = 𝒙(2) is given by

(1)
1 1 𝑇 −1
𝑓 (1) (2) (𝒙 )= 𝑚 × exp {− (𝒙(1) − 𝝀(1) ) 𝚺11.2 (𝒙(1) − 𝝀(1) )},
( 𝑿 |𝒙 ) 2
(2𝜋) 2 √|𝚺11.2 |

where 𝝀(1) = 𝝁(1) + 𝚺12 𝚺22


−1
(𝒙(2) − 𝝁(2) ) and 𝚺11.2 = 𝚺11 − 𝚺12 𝚺22
−1
𝚺21 .

The conditional distribution of 𝑿(1) given for fixed 𝑿(2) = 𝒙(2) is given by

(𝑿(1) |𝑿(2) = 𝒙(2) ) ~ 𝑵𝒎 ({𝝁(1) + 𝚺12 𝚺22


−1
(𝒙(2) − 𝝁(2) )}, 𝚺11.2 ); where 𝚺11.2 = 𝚺11 − 𝚺12 𝚺22
−1
𝚺21 .

The marginal density function (p.d.f.) of 𝑿(1) ~ 𝑵𝒎 (𝝁(1) , 𝚺11 ) is given by

1 1 𝑇 −1
𝑓𝑿(1) (𝒙(1) ) = 𝑚 × exp {− (𝒙(1) − 𝝁(1) ) 𝚺11 (𝒙(1) − 𝝁(1) )}.
(2𝜋) 2 √|𝚺11 | 2

Similarly, the joint density function (p.d.f.) of (𝑿(1) , 𝑿(2) ) can be written as

1 1 𝑇 −1
𝑓𝑿 (𝒙(1) , 𝒙(2) ) = 𝑚 𝑝−𝑚 × exp {− (𝒙(1) − 𝝁(1) ) 𝚺11 (𝒙(1) − 𝝁(1) )}
(2𝜋) 2 (2𝜋) 2 √|𝚺11 |√|𝚺22.1 | 2

1 𝑇 −1
× exp {− (𝒙(1) − {𝝁(2) + 𝚺21 𝚺11
−1
(𝒙(1) − 𝝁(1) )}) 𝚺22.1 (𝒙(1) − {𝝁(2) + 𝚺21 𝚺11
−1
(𝒙(1) − 𝝁(1) )})}.
2
−1
where 𝚺22.1 = 𝚺22 − 𝚺21 𝚺11 𝚺12 .

The conditional density function (p.d.f.) of 𝑿(2) given for fixed 𝑿(1) = 𝒙(1) is given by

(2)
1 1 𝑇 −1
𝑓 (2) (1) (𝒙 )= 𝑝−𝑚 × exp {− (𝒙(2) − 𝝀(2) ) 𝚺22.1 (𝒙(2) − 𝝀(2) )},
( 𝑿 |𝒙 ) 2
(2𝜋) 2 √|𝚺22.1 |

where 𝝀(2) = 𝝁(2) + 𝚺21 𝚺11


−1
(𝒙(1) − 𝝁(1) ) and 𝚺22.1 = 𝚺22 − 𝚺21 𝚺11
−1
𝚺12 .

The conditional distribution of 𝑿(2) given for fixed 𝑿(1) = 𝒙(1) is given by

(𝑿(2) |𝑿(1) = 𝒙(1) ) ~ 𝑵(𝒑−𝒎) ({𝝁(2) + 𝚺21 𝚺11


−1
(𝒙(1) − 𝝁(1) )}, 𝚺22.1 ), where 𝚺22.1 = 𝚺22 − 𝚺21 𝚺11
−1
𝚺12 .

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 14


MULTIVARIATE NORMAL DISTRIBUTION

NOTE:

The marginal distribution of each component of 𝑿 is univariate normal. Converse is not true in general. The
converse is true if the components of 𝑿 are all independent and normal.

EXAMPLE:

Suppose the random variable 𝑋1 and 𝑋2 have joint distribution function

𝐹(𝑋1,𝑋2) (𝑥1 , 𝑥2 ) = Φ(𝑥1 )Φ(𝑥2 ){1 + 𝛼(1 − Φ(𝑥1 ))(1 − Φ(𝑥2 ))}

where |𝛼| < 1, Φ(𝑥) be the distribution function of standard normal distribution. It can be shown that the marginal
distribution of 𝑋1 and 𝑋2 are standard normal.

RESULT 4:

If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) and 𝑩𝑚×𝑝 is a 𝑚 × 𝑝 (𝑚 ≤ 𝑝) matrix with 𝑅𝑎𝑛𝑘(𝑩) = 𝑚 and 𝒅 is 𝑚 × 1 vector, then the 𝑚
linear combinations 𝑩𝑿 ~ 𝑵𝒎 (𝑩𝝁 + 𝒅, 𝑩𝚺𝑩𝑇 ).

PROOF:
(1)
𝒀𝑚×1 𝑩𝑚×𝑝 𝒅𝑚×1
Let us consider the transformation 𝒀𝑝×1 = ( ) =( )𝑿 + (𝟎 ) or, 𝒀 = 𝑷𝑿 + 𝒃,
(2)
𝒀(𝑝−𝑚)×1 𝑪(𝑝−𝑚)×𝑝 𝑝×1 (𝑝−𝑚)×1

𝑩𝑚×𝑝 (1) (2)


where 𝑷𝑝×𝑝 = ( ), 𝒀𝑚×1 = 𝑩𝑚×𝑝 𝑿𝑝×1 + 𝒅𝑚×1 , 𝒀(𝑝−𝑚)×1 = 𝑪(𝑝−𝑚)×𝑝 𝑿𝑝×1 and 𝑪 is such that 𝑷 is non-
𝑪(𝑝−𝑚)×𝑝
singular.

From RESULT 2, we know that, if 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) and 𝒀𝑝×1 = 𝑷𝑿 + 𝒃 where 𝑷𝑝×𝑝 is non-singular matrix and 𝒃 is 𝑝 ×
1 vector, then 𝒀 ~ 𝑵𝑝 (𝑷𝝁 + 𝒃, 𝑷𝚺𝑷𝑇 ).

(1)
𝒀𝑚×1 𝑩𝝁 + 𝒅 𝑇
𝑩𝚺𝑪𝑇 ))
Hence, 𝒀𝑝×1 = ( ) ~ 𝑵𝑝 (( ) , (𝑩𝚺𝑩𝑇
(2)
𝒀(𝑝−𝑚)×1 𝑪𝝁 𝑪𝚺𝑩 𝑪𝚺𝑪𝑇

(1)
From RESULT 3, we know that, the marginal distribution of any subset 𝒀𝑚×1 = (𝑌1 , 𝑌2 , ⋯ , 𝑌𝑚 )𝑇 ; (𝑚 < 𝑝) of 𝒀 is 𝑚-
(2) 𝑇
dimensional multivariate normal and the marginal distribution of subset 𝒀(𝑝−𝑚)×1 = (𝑌𝑚+1 , 𝑌𝑚+2 , ⋯ , 𝑌𝑝 ) of 𝒀 is
(𝑝 − 𝑚)-dimensional multivariate normal.

The marginal distributions of 𝒀(1) and 𝒀(2) , respectively, are given by

𝒀(1) = 𝑩𝑿 + 𝒅 ~ 𝑵𝒎 (𝑩𝝁 + 𝒅, 𝑩𝚺𝑩𝑇 ) and 𝒀(2) = 𝑪𝑿 ~ 𝑵(𝒑−𝒎) (𝑪𝝁, 𝑪𝚺𝑪𝑇 ).

ALTERNATIVE

Let 𝒀𝑚×1 = 𝑩𝑚×𝑝 𝑿𝑝×1 + 𝒅𝑚×1 . For any real 𝒕𝑚×1 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑚 )𝑇 , the moment generating function (M.G.F.) of 𝒀 is
given by

𝑴𝒀 (𝒕) = 𝐸{exp(𝒕𝑇 𝒀)} = 𝐸[exp{𝒕𝑇 (𝑩𝑿 + 𝒅)}]

= exp(𝒕𝑇 𝒅) ∙ 𝐸{exp(𝒕𝑇 𝑩𝑿)}

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 15


MULTIVARIATE NORMAL DISTRIBUTION

= exp(𝒕𝑇 𝒅) ∙ 𝐸{exp(𝑺𝑇 𝑿)} ; [where 𝑺𝑝×1 = 𝑩𝑇 𝒕 ]

1 1
= exp(𝒕𝑇 𝒅) ∙ exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺) ; [∵ If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺), 𝑴𝑿 (𝑺) = exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺)]
2 2

1
= exp (𝒕𝑇 (𝑩𝝁 + 𝒅) + 𝒕𝑇 (𝑩𝚺𝑩𝑇 )𝒕)
2

⟶ M.G.F. of 𝑵𝑚 (𝑩𝝁 + 𝒅, 𝑩𝚺𝑩𝑇 )

Since 𝚺 is a positive definite and 𝑩𝑚×𝑝 is a 𝑚 × 𝑝 (𝑚 ≤ 𝑝) matrix with 𝑎𝑛𝑘(𝑩) = 𝑚 , 𝑩𝚺𝑩𝑇 is also positive definite.

By uniqueness property of moment generating function (M.G.F.), 𝑴𝒀 (𝒕) is the M.G.F. of 𝑵𝑚 (𝑩𝝁 + 𝒅, 𝑩𝚺𝑩𝑇 ) i.e., the
distribution of 𝒀𝑚×1 = 𝑩𝑿 + 𝒅 is 𝑵𝑚 (𝑩𝝁 + 𝒅, 𝑩𝚺𝑩𝑇 ). Hence, 𝒀𝑚×1 = 𝑩𝑿 + 𝒅 ~ 𝑵𝑚 (𝑩𝝁 + 𝒅, 𝑩𝚺𝑩𝑇 ).

OBSERVATION 5:

If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) and 𝑩𝑚×𝑝 is a 𝑚 × 𝑝 (𝑚 ≤ 𝑝) matrix with 𝑅𝑎𝑛𝑘(𝑩) = 𝑚, then 𝑩𝑿 ~ 𝑵𝒎 (𝑩𝝁, 𝑩𝚺𝑩𝑇 ).

PROOF:

Same as RESULT 4, where 𝒅𝑚×1 = 𝟎𝑚×1 = (0, 0, ⋯ , 0)𝑇 .

ALTERNATIVE

Let 𝒀𝑚×1 = 𝑩𝑚×𝑝 𝑿𝑝×1 . For any real 𝒕𝑚×1 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑚 )𝑇 , the moment generating function (M.G.F.) of 𝒀 is given by

𝑴𝒀 (𝒕) = 𝐸{exp(𝒕𝑇 𝒀)} = 𝐸[exp{𝒕𝑇 (𝑩𝑿)}] = 𝐸{exp(𝒕𝑇 𝑩𝑿)}

= 𝐸{exp(𝑺𝑇 𝑿)} ; [where 𝑺𝑝×1 = 𝑩𝑇 𝒕 ]

1 1
= exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺) ; [∵ If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺), 𝑴𝑿 (𝑺) = exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺)]
2 2

1
= exp (𝒕𝑇 𝑩𝝁 + 𝒕𝑇 (𝑩𝚺𝑩𝑇 )𝒕)
2

⟶ M.G.F. of 𝑵𝑚 (𝑩𝝁, 𝑩𝚺𝑩𝑇 )

Since 𝚺 is a positive definite and 𝑩𝑚×𝑝 is a 𝑚 × 𝑝 (𝑚 ≤ 𝑝) matrix with 𝑎𝑛𝑘(𝑩) = 𝑚 , 𝑩𝚺𝑩𝑇 is also positive definite.

By uniqueness property of moment generating function (M.G.F.), 𝑴𝒀 (𝒕) is the M.G.F. of 𝑵𝑚 (𝑩𝝁, 𝑩𝚺𝑩𝑇 ) i.e., the
distribution of 𝒀𝑚×1 = 𝑩𝑿 is 𝑵𝑚 (𝑩𝝁, 𝑩𝚺𝑩𝑇 ). Hence, 𝒀𝑚×1 = 𝑩𝑿 ~ 𝑵𝑚 (𝑩𝝁, 𝑩𝚺𝑩𝑇 ).

OBSERVATION 6:

If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) and 𝑪𝑝×𝑚 is a 𝑚 × 𝑝 (𝑚 ≤ 𝑝) matrix with 𝑅𝑎𝑛𝑘(𝑩) = 𝑚, then 𝑪𝑇 𝑿 ~ 𝑵𝒎 (𝑪𝑇 𝝁, 𝑪𝑇 𝚺𝑪).

PROOF:

Same as RESULT 4, where 𝑩 = 𝑪𝑇 and 𝒅𝑚×1 = 𝟎𝑚×1 = (0, 0, ⋯ , 0)𝑇 .

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 16


MULTIVARIATE NORMAL DISTRIBUTION

RESULT 5:

𝑇
The random vector 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) follows 𝑝-dimensional multivariate normal distribution if and only
𝑝 𝑇
if any linear combination 𝒍𝑇 𝑿 = ∑𝑖=1 𝑙𝑖 𝑋𝑖 follows univariate normal distribution, where 𝒍𝑝×1 = (𝑙1 , 𝑙2 , ⋯ , 𝑙𝑝 ) , i.e., the
𝑝
random vector 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) if and only if 𝒍𝑇 𝑿 = ∑𝑖=1 𝑙𝑖 𝑋𝑖 ~ 𝑁1 (𝒍𝑇 𝝁, 𝒍𝑇 𝚺𝒍).

PROOF:
𝑝 𝑇
Suppose 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) and 𝑌 = 𝒍𝑇 𝑿 = ∑𝑖=1 𝑙𝑖 𝑋𝑖 , where 𝒍𝑝×1 = (𝑙1 , 𝑙2 , ⋯ , 𝑙𝑝 ) .

For any real scalar 𝑡 ∈ ℝ, the moment generating function (M.G.F.) of 𝑌 is given by

𝑀𝑌 (𝑡) = 𝐸{exp(𝑡 𝑌)} = 𝐸[exp{𝑡 𝒍𝑇 𝑿}] = 𝐸{exp(𝑺𝑇 𝑿)} ; [where 𝑺𝑇𝑝×1 = 𝑡 𝒍𝑇 ]

1 1
= exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺) ; [∵ If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺), 𝑴𝑿 (𝑺) = exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺)]
2 2

1
= exp (𝑡 𝒍𝑇 𝝁 + 𝑡 2 (𝒍𝑇 𝚺𝒍))
2

𝑝 𝑝 𝑝
1
= exp (𝑡 ∑ 𝑙𝑖 𝜇𝑖 + 𝑡 2 ∑ ∑ 𝑙𝑖 𝑙𝑗 𝜎𝑖𝑗 ) ⟶ 𝑀. 𝐺. 𝐹. of 𝑵1 (𝒍𝑇 𝝁, 𝒍𝑇 𝚺𝒍)
2
𝑖=1 𝑖=1 𝑗=1

Since 𝚺 is a positive definite and 𝒍𝑇 𝚺 𝒍 is positive.

By uniqueness property of moment generating function (M.G.F.), 𝑀𝑌 (𝑡) is the M.G.F. of 𝑁1 (𝒍𝑇 𝝁, 𝒍𝑇 𝚺𝒍) i.e., the
𝑝
distribution of 𝑌 = 𝒍𝑇 𝑿 = ∑𝑖=1 𝑙𝑖 𝑋𝑖 is 𝑵1 (𝒍𝑇 𝝁, 𝒍𝑇 𝚺𝒍).

𝑝 𝑝 𝑝 𝑝 𝑝
𝑇 (𝒍𝑇 𝑇 𝑇
Hence, 𝑌 = 𝒍 𝑿 = ∑ 𝑙𝑖 𝑋𝑖 ~ 𝑁1 𝝁, 𝒍 𝚺𝒍) i. e. , 𝑌 = 𝒍 𝑿 = ∑ 𝑙𝑖 𝑋𝑖 ~ 𝑁1 (∑ 𝑙𝑖 𝜇𝑖 , ∑ ∑ 𝑙𝑖 𝑙𝑗 𝜎𝑖𝑗 )
𝑖=1 𝑖=1 𝑖=1 𝑖=1 𝑗=1

𝑇
Suppose 𝑌 = 𝒍𝑇 𝑿 follows univariate normal distribution for any real 𝒍𝑝×1 = (𝑙1 , 𝑙2 , ⋯ , 𝑙𝑝 ) .

Let 𝐸(𝑿) = 𝝁 and dispersion matrix of 𝑿 be 𝐷𝑖𝑠𝑝(𝑿) = 𝚺.

For any real scalar 𝑡 ∈ ℝ, the moment generating function (M.G.F.) of 𝑌 is given by

1
𝑀𝑌 (𝑡) = 𝐸{exp(𝑡 𝑌)} = exp (𝑡𝑏 + 𝑡 2 𝜎 2 ),
2

where 𝑏 = 𝐸(𝑌) = 𝐸(𝒍𝑇 𝑿) = 𝒍𝑇 𝝁, (say) and 𝜎 2 = 𝑉𝑎𝑟(𝑌) = 𝑉𝑎𝑟(𝒍𝑇 𝑿) = 𝒍𝑇 𝚺𝒍 , (say).

1
∴ 𝑀𝑌 (𝑡) = 𝐸{exp(𝑡 𝑌)} = 𝐸{exp(𝑡 𝒍𝑇 𝑿)} = exp (𝑡 𝒍𝑇 𝝁 + 𝑡 2 (𝒍𝑇 𝚺𝒍)).
2

1 𝑇
For 𝑡 = 1, 𝑀𝑌 (1) = 𝐸{exp(𝑌)} = exp (𝑡 𝒍𝑇 𝝁 + 𝑡 2 (𝒍𝑇 𝚺𝒍)) = 𝐸{exp(𝒍𝑇 𝑿)} = 𝑴𝑿 (𝒍), for real 𝒍 = (𝑙1 , 𝑙2 , ⋯ , 𝑙𝑝 ) ,
2
𝑇
which is the M.G.F. of 𝑵𝒑 (𝝁, 𝚺) for real 𝒍𝑝×1 = (𝑙1 , 𝑙2 , ⋯ , 𝑙𝑝 ) .

By uniqueness property of moment generating function (M.G.F.), 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺).

Hence, 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺).

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 17


MULTIVARIATE NORMAL DISTRIBUTION

NOTE:

Suppose 𝑿 = (𝑋1 , 𝑋2 )𝑇 and the marginal distribution of 𝑋1 and 𝑋2 are univariate normal. The random vector 𝑿
follows a bivariate normal distribution iff 𝑙1 𝑋1 + 𝑙2 𝑋2 follows univariate normal for every real 𝑙1 and 𝑙2 .

OBSERVATION 7:

If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺), then 𝑋𝑖 ~ 𝑁1 (𝜇𝑖 , 𝜎𝑖𝑖 ); 𝑖 = 1(1)𝑝.

PROOF:
𝑇 1 if 𝑘 = 𝑖
Same as RESULT 5, where 𝒍𝑝×1 = (𝑙1 , 𝑙2 , ⋯ , 𝑙𝑝 ) = (0, 0, ⋯ 0, 1,0, ⋯ , 0)𝑇 i. e., 𝑙𝑘 = { ; 𝑘 = 1(1)𝑝.
0 if 𝑘 ≠ 𝑖

OBSERVATION 8:

𝑇 𝜇𝑖 𝜎𝑖𝑖 𝜎𝑖𝑗
If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺), then (𝑋𝑖 , 𝑋𝑗 ) ~ 𝑵𝟐 (𝜇 , (𝜎 𝜎𝑗𝑗 )) ; 𝑖, 𝑗 = 1(1)𝑝 (𝑖 ≠ 𝑗).
𝑗 𝑗𝑖

PROOF:

𝑇 1 if 𝑘 = 𝑖, 𝑗
Same as RESULT 5, where 𝒍𝑝×1 = (𝑙1 , 𝑙2 , ⋯ , 𝑙𝑝 ) = (0, ⋯ , 1, 0, ⋯ ,0, 1,0, ⋯ , 0)𝑇 i. e., 𝑙𝑘 = { ; 𝑘 = 1(1)𝑝.
0 if 𝑘 ≠ 𝑖, 𝑗

RESULT 6:

(1) (1) 𝑚×𝑚 𝑚×(𝑝−𝑚)


𝑿𝑚×1 𝝁𝑚×1 𝚺11 𝚺12
Suppose 𝑿𝑝×1 = ( (2) ) ~ 𝑵𝒑 (𝝁, 𝚺), where 𝝁𝑝×1 = ( (2) ) , 𝚺 = ( (𝑝−𝑚)×𝑚 (𝑝−𝑚)×(𝑝−𝑚)
)
𝑿(𝑝−𝑚)×1 𝝁(𝑝−𝑚)×1 𝚺21 𝚺22
𝑇
and 𝚺12 = 𝚺21 .

Then 𝑿(1) and 𝑿(2) are independently distributed iff 𝚺12 = 𝚺21
𝑇
= 𝟎𝑚×(𝑝−𝑚) [i. e., 𝐶𝑜𝑣(𝑿(1) , 𝑿(2) ) = 𝟎].

PROOF:
𝑇
Suppose 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺).

The density function (p.d.f.) of 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) is given by

1 1
𝑓𝑿 (𝒙) = 𝑝 × exp {− (𝒙 − 𝝁)𝑇 𝚺 −1 (𝒙 − 𝝁)} ; 𝒙 ∈ ℝ𝑝 , 𝝁 ∈ ℝ𝑝 , 𝚺 is 𝑝. 𝑑.
(2𝜋)2 √|𝚺| 2

Suppose 𝐶𝑜𝑣(𝑿(1) , 𝑿(2) ) = 𝚺12 = 𝚺21


𝑇
= 𝟎𝑚×(𝑝−𝑚) .

𝚺11 𝚺12 𝚺 𝟎
Now, 𝚺 = ( ) = ( 11 ).
𝚺21 𝚺22 𝟎 𝚺22
−1
𝚺11 𝟎
Hence, |𝚺| = |𝚺11 | ∙ |𝚺22 | and 𝚺 −1 = ( −1 ).
𝟎 𝚺22

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 18


MULTIVARIATE NORMAL DISTRIBUTION

𝑇
𝑿(1) − 𝝁(1) 𝚺 −1 𝟎 𝑿(1) − 𝝁(1)
∴ (𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) = ( (2) (2)
) ( 11 −1 ) ( (2) )
𝑿 −𝝁 𝟎 𝚺22 𝑿 − 𝝁(2)

𝑇 𝑇
= (𝑿(1) − 𝝁(1) ) 𝚺11
−1
(𝑿(1) − 𝝁(1) ) + (𝑿(2) − 𝝁(2) ) 𝚺22
−1
(𝑿(2) − 𝝁(2) ).

𝑇 𝑇
The density function (p.d.f.) of 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) = (𝑿(1) , 𝑿(2) ) becomes

1 1 𝑇 −1
𝑓𝑿 (𝒙) = [ 𝑚 × exp {− (𝑿(1) − 𝝁(1) ) 𝚺11 (𝑿(1) − 𝝁(1) )}]
(2𝜋) 2 √|𝚺11 | 2

1 1 𝑇 −1
×[ 𝑝−𝑚 × exp {− (𝒙(2) − 𝝁(2) ) 𝚺22 (𝒙(2) − 𝝁(2) )}] ;
(2𝜋) 2 √|𝚺22 | 2

[𝒙(1) , 𝝁(1) ∈ ℝ𝑚 ; 𝒙(2) , 𝝁(2) ∈ ℝ(𝑝−𝑚) ; 𝚺11 and 𝚺22 are 𝑝. 𝑑. ]

Hence, 𝑿(1) and 𝑿(2) are independently distributed with

𝑿(1) ~ 𝑵𝒎 (𝝁(1) , 𝚺11 ) and 𝑿(2) ~ 𝑵(𝒑−𝒎) (𝝁(2) , 𝚺22 ).

Suppose 𝑿(1) and 𝑿(2) are independently distributed.

(1) (2)
Since 𝑿(1) and 𝑿(2) are independently distributed, 𝑋𝑖 and 𝑋𝑗 are independent; 𝑖 = 1(1)𝑚, 𝑗 = 𝑚 + 1(1)𝑝.

∴ 𝑓𝑿 (𝒙) = 𝑓𝑿(1) (𝒙(1) ) ∙ 𝑓𝑿(2) (𝒙(2) ) = 𝑓𝑿(1) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑚 ) ∙ 𝑓𝑿(2) (𝑥𝑚+1 , 𝑥𝑚+2 , ⋯ , 𝑥𝑝 ).

(1) (2) (1) (1) (2) (2)


Now, 𝜎𝑖𝑗 = 𝐶𝑜𝑣(𝑋𝑖 , 𝑋𝑗 ) = 𝐸 [{𝑋𝑖 − 𝐸(𝑋𝑖 )}{𝑋𝑗 − 𝐸(𝑋𝑗 )}] ; for all 𝑖 = 1(1)𝑚, 𝑗 = 𝑚 + 1(1)𝑝.

(1) (1) (2) (2)


= 𝐸{(𝑋𝑖 − 𝜇𝑖 )(𝑋𝑗 − 𝜇𝑗 )}

∞ ∞ ∞
(1) (1) (2) (2)
= ∫ ∫ ⋯ ∫ (𝑥𝑖 − 𝜇𝑖 )(𝑥𝑗 − 𝜇𝑗 )𝑓𝑿 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) 𝑑𝑥1 𝑑𝑥2 ⋯ 𝑑𝑥𝑝
−∞ −∞ −∞

(1) (1) (2) (2)


= ∫ (𝑥𝑖 − 𝜇𝑖 )(𝑥𝑗 − 𝜇𝑗 )𝑓𝑿(1) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑚 ) ∙ 𝑓𝑿(2) (𝑥𝑚+1 , 𝑥𝑚+2 , ⋯ , 𝑥𝑝 ) 𝑑𝑥1 𝑑𝑥2 ⋯ 𝑑𝑥𝑝
ℝ𝑝

(1) (1)
= { ∫(𝑥𝑖 − 𝜇𝑖 ) 𝑓𝑿(1) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑚 ) 𝑑𝑥1 𝑑𝑥2 ⋯ 𝑑𝑥𝑚 }
ℝ𝑚

(2) (2)
×{ ∫ (𝑥𝑗 − 𝜇𝑗 ) 𝑓𝑿(2) (𝑥𝑚+1 , 𝑥𝑚+2 , ⋯ , 𝑥𝑝 ) 𝑑𝑥𝑚+1 𝑑𝑥𝑚+2 ⋯ 𝑑𝑥𝑝 }
ℝ(𝑝−𝑚)

(1) (1) (2) (2)


= 𝐸(𝑋𝑖 − 𝜇𝑖 ) × 𝐸(𝑋𝑗 − 𝜇𝑗 )

=0

∴ 𝜎𝑖𝑗 = 0; for all 𝑖 = 1(1)𝑚, 𝑗 = 𝑚 + 1(1)𝑝

𝑇
⟺ 𝚺12 = 𝚺21 = 𝟎𝑚×(𝑝−𝑚) .

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 19


MULTIVARIATE NORMAL DISTRIBUTION

OBSERVATION 9:

𝑇 𝑇
If 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) = (𝑿(1) , 𝑿(2) , ⋯ , 𝑿(𝑟) ) ~ 𝑵𝒑 (𝝁, 𝚺) and 𝐶𝑜𝑣(𝑿(𝑖) , 𝑿(𝑗) ) = 𝟎; 𝑖 ≠ 𝑗 (𝑖, 𝑗 = 1(1)𝑝),

then 𝑿(𝑖) s ( 𝑖 = 1(1)𝑝) are mutually independent and not just pairwise independent.

OBSERVATION 10:

𝑇
If 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝜎 2 𝑰𝒑 ), then 𝒂𝑇 𝑿 and 𝒃𝑇 𝑿 are independent iff 𝒂𝑇 𝒃 = 0.

PROOF:
𝑇 𝑇 𝑝
Let 𝒂𝑝×1 = (𝑎1 , 𝑎2 , ⋯ , 𝑎𝑝 ) and 𝒃𝑝×1 = (𝑏1 , 𝑏2 , ⋯ , 𝑏𝑝 ) be two 𝑝 × 1 vectors with 𝒂𝑇 𝒃 = ∑𝑖=1 𝑎𝑖 𝑏𝑖 = 0.

𝑌 𝑇 𝑇
Let us consider the transformation 𝑌1 = 𝒂𝑇 𝑿, 𝑌2 = 𝒃𝑇 𝑿 and 𝒀2×1 = ( 1 ) = (𝒂𝑇 𝑿) = (𝒂𝑇 ) 𝑿 = 𝑷2×𝑝 𝑿.
𝑌2 𝒃 𝑿 𝒃

For any real 𝒕2×1 = (𝑡1 , 𝑡2 )𝑇 , the moment generating function (M.G.F.) of 𝒀 is given by

𝑴𝒀 (𝒕) = 𝐸{exp(𝒕𝑇 𝒀)} = 𝐸{exp(𝒕𝑇 𝑷𝑿)}

= 𝐸{exp(𝑺𝑇 𝑿)} ; [where 𝑺𝑝×1 = 𝑷𝑇 𝒕 ]

1 1
= exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺) ; [∵ If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺), 𝑴𝑿 (𝑺) = exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺)]
2 2
1
= exp (𝒕𝑇 𝑷𝝁 + 𝒕𝑇 (𝑷𝚺𝑷𝑇 )𝒕) ⟶ M.G.F. of 𝑵2 (𝑷𝝁, 𝑷𝚺𝑷𝑇 )
2

Since 𝚺 is a positive definite and 𝑷2×𝑝 is a 2 × 𝑝 (2 ≤ 𝑝) matrix with 𝑎𝑛𝑘(𝑷) = 2 , 𝑷𝚺𝑷𝑇 is also positive definite.

By uniqueness property of moment generating function (M.G.F.), 𝑴𝒀 (𝒕) is the M.G.F. of 𝑵2 (𝑷𝝁, 𝑷𝚺𝑷𝑇 ).

𝒂𝑇 𝝁 𝑇
2 𝒂 𝒂 0 )).
Hence, 𝒀2×1 = 𝑷𝑿 ~ 𝑵2 (( 𝑇 ),𝜎 (
𝒃 𝝁 0 𝒃𝑇 𝒃

𝑇 𝒂𝑇 𝝁 ∑𝑝 𝑎𝑖 𝜇𝑖
Now, 𝑷𝝁 = (𝒂𝑇 ) 𝝁 = ( 𝑇 ) = ( 𝑖=1 )
𝒃 𝒃 𝝁 ∑𝑝𝑖=1 𝑏𝑖 𝜇𝑖

𝑝
𝑇
and 𝑷𝚺𝑷 = 𝜎 𝑷𝑷 = 𝜎 (𝒂𝑇 𝒂
𝑇 2 𝑇 2 𝒂 𝑇 𝒃 ) = (𝜎 2 𝒂 𝑇 𝒂 0 ) 𝑇
[∵ 𝒂 𝒃 = ∑ 𝑎𝑖 𝑏𝑖 = 0]
𝒃 𝒂 𝒃𝑇 𝒃 0 𝜎 2 𝒃𝑇 𝒃
𝑖=1

−1
2 𝑇 (𝜎 2 𝒂𝑇 𝒂)−1 0
∴ (𝑷𝚺𝑷𝑇 )−1 = (𝜎 𝒂 𝒂 0 ) ⟺ (𝑷𝑇 )−1 𝚺 −1 𝑷−1 = ( ).
0 2 𝑇
𝜎 𝒃 𝒃 0 (𝜎 2 𝒃𝑇 𝒃)−1
2 𝑇
and |𝑷𝚺𝑷𝑇 | = |𝜎 𝒂 𝒂 0 | ⟹ |𝑷||𝚺||𝑷𝑇 | = (𝜎 2 𝒂𝑇 𝒂) × (𝜎 2 𝒃𝑇 𝒃)
0 𝜎 2 𝒃𝑇 𝒃
𝑇
𝑌 − 𝒂𝑇 𝝁 (𝜎 2 𝒂𝑇 𝒂)−1 0 𝑌1 − 𝒂𝑇 𝝁
∴ (𝒀 − 𝑷𝝁)𝑇 (𝑷𝚺𝑷𝑇 )−1 (𝒀 − 𝑷𝝁) = ( 1 𝑇 ) ( −1 ) ( )
𝑌2 − 𝒃 𝝁 0 (𝜎 2 𝑇
𝒃 𝒃) 𝑌2 − 𝒃𝑇 𝝁

= (𝑌1 − 𝒂𝑇 𝝁)𝑇 (𝜎 2 𝒂𝑇 𝒂)−1 (𝑌1 − 𝒂𝑇 𝝁) + (𝑌2 − 𝒃𝑇 𝝁)𝑇 (𝜎 2 𝒃𝑇 𝒃)−1 (𝑌2 − 𝒃𝑇 𝝁).

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 20


MULTIVARIATE NORMAL DISTRIBUTION

The joint density function (p.d.f.) of 𝒀2×1 = (𝑌1 , 𝑌2 )𝑇 can be written as

1 1
𝑓𝒀 (𝑦1 , 𝑦2 ) = 2 × exp {− (𝒀 − 𝑷𝝁)𝑇 (𝑷𝚺𝑷𝑇 )−1 (𝒀 − 𝑷𝝁)}
(2𝜋)2 √|𝑷𝚺𝑷𝑇 | 2

1 1
=[ × exp {− (𝑌1 − 𝒂𝑇 𝝁)𝑇 (𝜎 2 𝒂𝑇 𝒂)−1 (𝑌1 − 𝒂𝑇 𝝁)}]
√2𝜋√𝜎 2 𝒂𝑇 𝒂 2
1 1
×[ × exp {− (𝑌2 − 𝒃𝑇 𝝁)𝑇 (𝜎 2 𝒃𝑇 𝒃)−1 (𝑌2 − 𝒃𝑇 𝝁)}]
√2𝜋√𝜎 2 𝒃𝑇 𝒃 2

Hence, 𝑌1 = 𝒂𝑇 𝑿 and 𝑌2 = 𝒃𝑇 𝑿 are independent with 𝑌1 = 𝒂𝑇 𝑿 ~ 𝑁1 (𝒂𝑇 𝝁, 𝜎 2 𝒂𝑇 𝒂) and 𝑌2 = 𝒃𝑇 𝑿 ~ 𝑁1 (𝒃𝑇 𝝁, 𝜎 2 𝒃𝑇 𝒃).

Suppose 𝑌1 = 𝒂𝑇 𝑿 and 𝑌2 = 𝒃𝑇 𝑿 are independent.

Now, 𝐶𝑜𝑣(𝑌1 , 𝑌2 ) = 0 ⟹ 𝐶𝑜𝑣(𝒂𝑇 𝑿, 𝒃𝑇 𝑿) = 𝒂𝑇 𝚺𝒃 = 𝜎 2 𝒂𝑇 𝑰𝒑 𝒃 = 𝜎 2 𝒂𝑇 𝒃 = 0.

∴ 𝒂𝑇 𝒃 = 0 [∵ 𝜎 2 > 0]

RESULT 7:

Suppose 𝑿𝑖 ~ 𝑵𝒑 (𝝁𝒊 , 𝚺𝐢 ); 𝑖 = 1(1)𝑘 be 𝑘 independent 𝑝-dimensional multivariate normal vectors.

Then, 𝒀 = ∑𝑘𝑖=1 𝑿𝑖 ~ 𝑵𝒑 (𝝁, 𝚺), where 𝝁 = ∑𝑘𝑖=1 𝝁𝑖 and 𝚺 = ∑𝑘𝑖=1 𝚺𝑖

PROOF:

Suppose 𝑿𝑖 ~ 𝑵𝒑 (𝝁𝒊 , 𝚺𝐢 ); 𝑖 = 1(1)𝑘 be 𝑘 independent 𝑝-dimensional multivariate normal vectors and 𝒀𝑝×1 = ∑𝑘𝑖=1 𝑿𝑖 .

𝑇
For any real 𝒕𝑝×1 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 ) , the moment generating function (M.G.F.) of 𝒀 is given by

𝑘 𝑘
𝑇 𝑇
𝑴𝒀 (𝒕) = 𝐸{exp(𝒕 𝒀)} = 𝐸 [exp {𝒕 (∑ 𝑿𝑖 )}] = 𝐸 {exp (∑ 𝒕𝑇 𝑿𝑖 )}
𝑖=1 𝑖=1

= ∏[𝐸{exp(𝒕𝑇 𝑿𝑖 )}] ; [∵ 𝑿′𝑖 s are independent]


𝑖=1

𝑘
1 1
= ∏ [exp (𝒕𝑇 𝝁𝒊 + 𝒕𝑇 𝚺𝒊 𝒕)] ; [∵ If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁𝑿 , 𝚺𝐗 ), 𝑴𝑿 (𝒕) = exp (𝒕𝑇 𝝁𝑿 + 𝒕𝑇 𝚺𝐗 𝒕)]
2 2
𝑖=1

𝑘 𝑘
1
𝑇
= exp (𝒕 ∑ 𝝁𝒊 + 𝒕𝑇 (∑ 𝚺𝒊 ) 𝒕)
2
𝑖=1 𝑖=1

1
= exp (𝒕𝑇 𝝁 + 𝒕𝑇 𝚺 𝒕) ⟶ M.G.F. of 𝑵𝑝 (𝝁, 𝚺)
2

Since 𝚺𝒊 ; 𝑖 = 1(1)𝑘 are positive definite matrix, 𝚺 = ∑𝑘𝑖=1 𝚺𝑖 is also positive definite.

By uniqueness property of moment generating function (M.G.F.), 𝑴𝒀 (𝒕) is the M.G.F. of 𝑵𝑝 (𝝁, 𝚺) i.e., the distribution
of 𝒀𝑝×1 = ∑𝑘𝑖=1 𝑿𝑖 is 𝑵𝒑 (𝝁, 𝚺).

Hence, 𝒀𝑝×1 = ∑𝑘𝑖=1 𝑿𝑖 ~ 𝑵𝒑 (𝝁, 𝚺).

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 21


MULTIVARIATE NORMAL DISTRIBUTION

RESULT 8:

𝑇
Suppose 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺). Then 𝑸 = (𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) ~ 𝝌𝟐𝒑 .

PROOF:
𝑇
Suppose 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺). The density function (p.d.f.) of 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) is given by

1 1
𝑓𝑿 (𝒙) = 𝑝 × exp {− (𝒙 − 𝝁)𝑇 𝚺 −1 (𝒙 − 𝝁)} ; 𝒙 ∈ ℝ𝑝 , 𝝁 ∈ ℝ𝑝 , 𝚺 is 𝑝. 𝑑.
(2𝜋)2 √|𝚺| 2

Since 𝚺 is a positive definite, there exist a non-singular matrix 𝑪 such that 𝑪𝑪𝑇 = 𝚺.

Let us consider the transformation 𝒀 = 𝑪−1 (𝑿 − 𝝁) or, 𝑿 = 𝑪𝒀 + 𝝁.

The jacobian of the transformation is

𝑿
|𝐽| = |𝐽 ( )| = ||𝑪|| = √|𝑪𝑪𝑇 | = √|𝚺|.
𝒀
𝑇
Now, (𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) = (𝑿 − 𝝁)𝑇 (𝑪𝑪𝑇 )−1 (𝑿 − 𝝁) = (𝑪−1 (𝑿 − 𝝁)) (𝑪−1 (𝑿 − 𝝁)) = 𝒀𝑇 𝒀.

The density function (p.d.f.) of 𝒀𝑝×1 is given by

1 1
𝑔𝒀 (𝒚) = 𝑝 × exp (− 𝒚𝑇 𝒚) × √|𝚺| ; 𝒚 ∈ ℝ𝑝 , 𝚺 is 𝑝. 𝑑.
(2𝜋)2 √|𝚺| 2

𝑝
1 1 2
= 𝑝 ∙ exp (− ∑ 𝑦𝑖 )
(2𝜋)2 2
𝑖=1

𝑝
1 1
= ∏{ ∙ exp (− 𝑦𝑖2 )}
√2𝜋 2
𝑖=1

Hence, 𝒀 = 𝑪−1 (𝑿 − 𝝁) ~ 𝑵𝒑 (𝟎, 𝑰𝒑 ) i.e.,𝑌𝑖′ s are independent and identically distributed with 𝑌𝑖 ~ 𝑁1 (0, 1); 𝑖 = 1(1)𝑝.

Now, 𝑸 = (𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) = 𝒀𝑇 𝒀 = ∑ 𝑌𝑖2 = sum of squares of 𝑝 independent standard normal variates


𝑖=1

∴ 𝑸 = (𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) = 𝒀𝑇 𝒀 ~ 𝝌𝟐𝒑 .

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 22


MULTIVARIATE NORMAL DISTRIBUTION

RESULT 9:

Suppose 𝚺 is positive definite, so that 𝚺 −1 exists and (𝜆, 𝒆) is an eigenvalue-eigenvector pair for 𝚺. Then, 𝚺𝒆 = 𝜆𝒆
1 1
implies 𝚺 −1 𝒆 = 𝒆, where ( , 𝒆) is the corresponding eigenvalue-eigenvector pair for 𝚺 −1 . Also, 𝚺 −1 is positive definite.
𝜆 𝜆

PROOF:

Suppose 𝚺 is positive definite, so that 𝚺 −1 exists and (𝜆, 𝒆) is an eigenvalue-eigenvector pair for 𝚺 corresponding to the
1
pair ( , 𝒆) for 𝚺 −1 .
𝜆

Since 𝚺 is positive definite and 𝒆 ≠ 𝟎 an eigenvector, we have

𝒆𝑇 𝚺𝒆 > 0 ⟹ 𝒆𝑇 𝚺𝒆 = 𝒆𝑇 (𝚺𝒆) = 𝒆𝑇 (𝜆𝒆) = 𝜆𝒆𝑇 𝒆 = 𝜆 > 0.

1
Now, 𝒆 = 𝚺 −1 (𝚺𝒆) = 𝚺 −1 (𝜆𝒆) = 𝜆𝚺 −1 𝒆 ⟹ 𝒆 = 𝚺 −1 𝒆.
𝜆
1
Hence, ( , 𝒆) is an eigenvalue-eigenvector pair for 𝚺 −1 .
𝜆

1
Suppose ( , 𝒆𝒊 ) ; 𝑖 = 1(1)𝑝 are eigenvalue-eigenvector pair for 𝚺 −1 .
𝜆𝑖

𝑝 𝑝
1 1 1
For any vector 𝒙𝑝×1 , 𝒙𝑇 𝚺 −1 𝒙 = 𝒙𝑇 (∑ 𝒆𝒊 𝒆𝑇𝒊 ) 𝒙 = ∑ ( ) (𝒙𝑇 𝒆𝒊 )𝟐 ≥ 0 ; [∵ ( ) (𝒙𝑇 𝒆𝒊 )𝟐 ≥ 0; 𝑖 = 1(1)𝑝]
𝜆𝑖 𝜆𝑖 𝜆𝑖
𝑖=1 𝑖=1

𝑝
𝑇 −1
1
∴ 𝒙≠𝟎 ⟹ 𝒙 𝚺 𝒙 = ∑ ( ) (𝒙𝑇 𝒆𝒊 )𝟐 > 0.
𝜆𝑖
𝑖=1

Now, 𝒙𝑇 𝒆𝒊 = 0 for all 𝑖 = 1(1)𝑝 only if 𝒙 = 𝟎.

Hence, 𝚺 −1 is positive definite.

DIAGONALIZATION:

𝑇
Suppose 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺) and 𝚺 is positive definite with eigen-vectors 𝒆𝟏 , 𝒆𝟐 , ⋯ , 𝒆𝒑
corresponding to the eigen-values 𝜆1 , 𝜆2 , ⋯ , 𝜆𝑝 ; (𝜆𝑖 > 0).

0 if 𝑖 ≠ 𝑗
Now, 𝚺𝒆𝒊 = 𝜆𝑖 𝒆𝒊 ; 𝑖 = 1(1)𝑝 and 𝒆𝑇𝒊 𝒆𝒋 = { .
1 if 𝑖 = 𝑗

𝑇
Let us define 𝑼𝑝×𝑝 = (𝒆𝟏 , 𝒆𝟐 , ⋯ , 𝒆𝒑 ) . Then, 𝑼𝑇 𝑿 ~ 𝑵𝒑 (𝑼𝑇 𝝁, 𝑼𝑇 𝚺𝑼).

𝜆 if 𝑖 = 𝑗
But, 𝒆𝑇𝒋 𝚺𝒆𝒊 = 𝜆𝑖 𝒆𝑇𝒋 𝒆𝒊 = { 𝑖
0 if 𝑖 ≠ 𝑗.

Hence, 𝑼𝑇 𝚺𝑼 = 𝑫𝒊𝒂𝒈(𝜆1 , 𝜆2 , ⋯ , 𝜆𝑝 ).

Thus, given 𝚺, we can always construct an orthogonal matrix 𝑼𝑝×𝑝 such that if 𝒀 = 𝑼𝑇 𝑿, then 𝑌1 , 𝑌2 , ⋯ , 𝑌𝑝 are
independent normal variables with variances 𝜆1 , 𝜆2 , ⋯ , 𝜆𝑝 respectively.

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 23


MULTIVARIATE NORMAL DISTRIBUTION

CONTOUR:

For the density of a 𝑝-dimensional normal variable, the paths of 𝒙 values yielding a constant height for the density
are ellipsoids. The multivariate normal density is constant on surfaces where the square of the distance
(𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) is constant. These paths are called contours:

Constant probability density contour = {𝑿: (𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) = 𝑐 2 } = surface of an ellipsoid centered at 𝝁.

The family of ellipsoids obtained by varying 𝑐 (𝑐 > 0) has the same centre 𝝁, their shapes and orientation are determined
by 𝚺 and their sizes for a given 𝚺 are determined by 𝑐.

The axes of each ellipsoid of constant density are in the direction of the eigenvectors of 𝚺 −1 and their lengths are
proportional to the reciprocals of the square roots of the eigenvalues of 𝚺 −1.

Contours of constant density for the 𝑝-dimensional normal distribution are ellipsoids defined by 𝒙 such that
(𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) = 𝑐 2 , where 𝑐 is a constant.

𝑝
These ellipsoids are centered at 𝝁 and have axes ±𝑐√𝜆𝑖 𝒆𝒊 , where ∑𝑖=1 𝒆𝒊 = 𝜆𝑖 𝒆𝒊 ; 𝑖 = 1(1)𝑝.

In particular, (𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) = 𝑝 + 2 is called the concentration ellipsoid of 𝑿.

FIGURE: The 50% and 90% contours for the bivariate normal distributions with 𝜎11 = 𝜎22 (Johnson, Wichern).

OBSERVATION 11:

The solid ellipsoid of 𝒙 values satisfying (𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) ≤ 𝝌𝟐𝒑 (𝛼), has probability (1 − 𝛼) i.e.,

𝑷{(𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) ≤ 𝝌𝟐𝒑 (𝛼)} = 1 − 𝛼.

OBSERVATION 12:

The probability density function defined by the uniform distribution

𝑝
𝚪 ( + 1)
[ 2 ]; if (𝒙 − 𝝁)𝑇 𝚺 −1 (𝒙 − 𝝁) ≤ 𝑝 + 2 , 𝒙 ∈ ℝ𝑝 , 𝝁 ∈ ℝ𝑝 , 𝚺 is 𝑝. 𝑑.
𝑝
𝑓𝑿 (𝒙) = {(𝑝 + 2)𝜋}2 √|𝚺|
{ 0; otherwise,

has the same mean 𝐸(𝑿) = 𝝁 and the same covariance matrix 𝐸{(𝑿 − 𝝁)(𝑿 − 𝝁)𝑇 } = 𝚺 as the p-variate distribution.
This will be called the ellipsoid of concentration for the given distribution.

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 24


MULTIVARIATE NORMAL DISTRIBUTION

REGRESSION AND CORRELATION:

CASE I:

Let us define

(1) (1) 𝑚×𝑚 𝑚×(𝑝−𝑚)


𝑿𝑚×1 𝝁𝑚×1 𝚺11 𝚺12 𝑇
𝑿𝑝×1 = ( (2) ), 𝝁𝑝×1 = ( (2) ), 𝚺=( (𝑝−𝑚)×𝑚 (𝑝−𝑚)×(𝑝−𝑚)
) and 𝚺12 = 𝚺21 ,
𝑿(𝑝−𝑚)×1 𝝁(𝑝−𝑚)×1 𝚺21 𝚺22

where 𝐸(𝑿(1) ) = 𝝁(1) , 𝐸(𝑿(2) ) = 𝝁(2) ,

𝑇
𝐸 {(𝑿(1) − 𝝁(1) )(𝑿(1) − 𝝁(1) ) } = 𝚺11 ,

𝑇
𝐸 {(𝑿 (1) − 𝝁(1) )(𝑿(2) − 𝝁(2) ) } = 𝚺12 , and

𝑇
𝐸 {(𝑿(2) − 𝝁(2) )(𝑿(2) − 𝝁(2) ) } = 𝚺22 .

We know that, the conditional distribution of 𝑿(1) given for fixed 𝑿(2) = 𝒙(2) is given by

(𝑿(1) |𝑿(2) = 𝒙(2) ) ~ 𝑵𝒎 ({𝝁(1) + 𝚺12 𝚺22


−1
(𝒙(2) − 𝝁(2) )}, 𝚺11.2 ), −1
where 𝚺11.2 = 𝚺11 − 𝚺12 𝚺22 𝚺21 ,

(1)
The conditional expectation and variance-covariance matrix of 𝑿𝑚×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑚 )𝑇 given the fixed values
(2) 𝑇
𝒙(𝑝−𝑚)×1 = (𝑥𝑚+1 , 𝑥𝑚+2 , ⋯ , 𝑥𝑝 ) of the predictor variables are, respectively,

𝐸(𝑿(1) |𝑿(2) = 𝒙(2) ) = 𝝁(1) + 𝚺12 𝚺22


−1
(𝒙(2) − 𝝁(2) ) and 𝑉𝑎𝑟(𝑿(1) |𝑿(2) = 𝒙(2) ) = 𝚺11.2 = 𝚺11 − 𝚺12 𝚺22
−1
𝚺21 .

This conditional expected value, considered as a function of 𝑥𝑚+1 , 𝑥𝑚+2 , ⋯ , 𝑥𝑝 is called the multivariate regression of
the vector 𝑿(1) on 𝑿(2) = 𝒙(2) . It is composed of 𝑚 univariate linear regressions.

The vector

(1)
𝑿1.(2) = 𝝁(1) + 𝚺12 𝚺22
−1
(𝒙(2) − 𝝁(2) )

is called the vector of linear regression functions of 𝑿(1) on 𝑿(2) = 𝒙(2) .

The 𝑖-th component of the conditional mean vector 𝐸(𝑋𝑖 |𝑿(2) = 𝒙(2) ) is the linear regression of 𝑋𝑖 on 𝑿(2) = 𝒙(2) i.e.,

𝐸(𝑋𝑖 |𝑿(2) = 𝒙(2) ) = 𝜇𝑖 + 𝚺𝑋 𝑿(2) 𝚺22


−1
(𝒙(2) − 𝝁(2) ) = 𝜇𝑖 + 𝛔𝑇(𝑖) 𝚺22
−1
(𝒙(2) − 𝝁(2) ) = 𝑋𝑖.(2) (say),
𝑖

which minimizes the mean square error for the prediction of 𝑋𝑖 , where 𝛔𝑇(𝑖) = 𝚺𝑋 𝑿(2) ; 𝑖 = 1(1)𝑚 is the 𝑖-th row of 𝚺12 .
𝑖

(𝑝−𝑚)×1 𝑇 𝑇
The vector 𝜷(𝑖) −1
= (𝚺𝑋 𝑿(2) 𝚺22 ) = (𝛔𝑇(𝑖) 𝚺22
−1
) is the vector of regression coefficients, where 𝛔𝑇(𝑖) ; 𝑖 = 1(1)𝑚 is
𝑖

the 𝑖-th row of 𝚺12 .

−1
The 𝑚 × (𝑝 − 𝑚) matrix 𝑩 = 𝚺12 𝚺22 = (𝜷𝑚×1 𝑚×1 𝑚×1
(1) , 𝜷(2) , ⋯ , 𝜷(𝑝−𝑚) ) is called the matrix of regression coefficients.

The element of 𝑖-th row and 𝑘-th column of 𝑩𝑚×(𝑝−𝑚) , denoted by,

𝛽𝑖𝑘.(𝑚+1)⋯(𝑘−1)(𝑘+1)⋯𝑝 ; 𝑖 = 1(1)𝑚, 𝑘 = (𝑚 + 1)(1)𝑝

𝑇
is the regression coefficient of 𝑋𝑖 on 𝑋𝑘 = 𝑥𝑘 , and it is also the 𝑘-th element of (𝛔𝑇(𝑖) 𝚺22
−1
) .

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 25


MULTIVARIATE NORMAL DISTRIBUTION

The correlation between 𝑋𝑖 and its best linear predictor 𝑋𝑖.(2) = 𝐸(𝑋𝑖 |𝑿(2) = 𝒙(2) ) ; 𝑖 = 1(1)𝑚 is called the population
multiple correlation coefficient.

The population multiple correlation coefficient 𝑅𝑖.(2) or 𝜌𝑖.(2) is given by

𝑅𝑖.(2) = 𝜌𝑖.(2) = 𝐶𝑜𝑟𝑟(𝑋𝑖 , 𝑋𝑖.(2) ) ; 𝑖 = 1(1)𝑚

= 𝐶𝑜𝑟𝑟[𝑋𝑖 , {𝜇𝑖 + 𝛔𝑇(𝑖) 𝚺22


−1
(𝑿(2) − 𝝁(2) )}]

𝑉𝑎𝑟(𝑋𝑖.(2) ) 𝑉𝑎𝑟(𝛔𝑇(𝑖) 𝚺22


−1 (2)
𝑿 )
= +√ = +√
𝑉𝑎𝑟(𝑋𝑖 ) 𝜎𝑖𝑖

𝛔𝑇(𝑖) 𝚺22
−1
𝝈(𝑖)
= +√ ; [∵ 𝑿(2) ~ 𝑵(𝒑−𝒎) (𝝁(2) , 𝚺22 )]
𝜎𝑖𝑖

The square of the population multiple correlation coefficient, 𝑅𝑖.(2) or 𝜌𝑖.(2) , is called the population coefficient of
determination. The multiple correlation coefficient is a positive square root, so 0 ≤ 𝜌𝑖.(2) ≤ 1.

(1) (1)
The error of prediction vector 𝜺1.(2) = 𝑿(1) − 𝝁(1) − 𝚺12 𝚺22
−1
(𝑿(2) − 𝝁(2) ) = (𝑿(1) − 𝑿1.(2) ) has the expected squares
and cross-products (error-variation) matrix
𝑇
𝚺11.2 = 𝐸[𝑿(1) − 𝝁(1) − 𝚺12 𝚺22
−1
(𝑿(2) − 𝝁(2) )][𝑿(1) − 𝝁(1) − 𝚺12 𝚺22
−1
(𝑿(2) − 𝝁(2) )]

−1 −1 −1 −1
= 𝚺11 − 𝚺12 𝚺22 𝚺21 − 𝚺12 𝚺22 𝚺21 + 𝚺12 𝚺22 𝚺22 𝚺22 𝚺21
−1
= 𝚺11 − 𝚺12 𝚺22 𝚺21 .

(1) (1)
It can be shown that 𝜺1.(2) = (𝑿(1) − 𝑿1.(2) ) ~ 𝑵𝒎 (𝟎, 𝚺11.2 ).

Consider the pair of errors

𝜀𝑖.(2) = 𝑋𝑖 − 𝜇𝑖 − 𝛔𝑇(𝑖) 𝚺22


−1
(𝑿(2) − 𝝁(2) )

𝜀𝑘.(2) = 𝑋𝑘 − 𝜇𝑘 − 𝛔𝑇(𝑘) 𝚺22


−1
(𝑿(2) − 𝝁(2) )

obtained from the best linear predictors to predict 𝑋𝑖 and 𝑋𝑘 on 𝑿(2) = 𝒙(2) ; 𝑖, 𝑘 = 1(1)𝑚.

−1
Let 𝜎𝑖𝑘.(2) = 𝜎𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 = 𝐶𝑜𝑣(𝜀𝑖.(2) , 𝜀𝑘.(2) ) be the (𝑖, 𝑘)-th element of 𝚺11.2 = 𝚺11 − 𝚺12 𝚺22 𝚺21 ; 𝑖, 𝑘 = 1(1)𝑚.

The correlation between 𝜀𝑖.(2) and 𝜀𝑘.(2) ; 𝑖, 𝑘 = 1(1)𝑚 is called the population partial correlation coefficient between
𝑋𝑖 and 𝑋𝑘 eliminating the effects of 𝑿(2) . The population partial correlation coefficient between 𝑋𝑖 and 𝑋𝑘 eliminating
the effects of 𝑿(2) , denoted by 𝜌𝑖𝑘.(2) or, 𝜌𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 , is given by

𝜌𝑖𝑘.(2) = 𝜌𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 = 𝐶𝑜𝑟𝑟(𝜀𝑖.(2) , 𝜀𝑘.(2) ) ; 𝑖, 𝑘 = 1(1)𝑚

= 𝐶𝑜𝑟𝑟[{𝑋𝑖 − 𝜇𝑖 − 𝛔𝑇(𝑖) 𝚺22


−1
(𝑿(2) − 𝝁(2) )}, {𝑋𝑘 − 𝜇𝑘 − 𝛔𝑇(𝑘) 𝚺22
−1
(𝑿(2) − 𝝁(2) )}]

𝜎𝑖𝑘.(2) 𝜎𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝
= =
√𝜎𝑖𝑖.(2) ∙ √𝜎𝑘𝑘.(2) √𝜎𝑖𝑖.(𝑚+1)(𝑚+2)⋯𝑝 ∙ √𝜎𝑘𝑘.(𝑚+1)(𝑚+2)⋯𝑝

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 26


MULTIVARIATE NORMAL DISTRIBUTION

CASE II:

Let us define

𝑋1 𝜇1 𝜎11 𝑇
𝛔(1)
𝑿𝑝×1 = ( (2) ), 𝝁𝑝×1 = (𝝁(2) ), 𝚺=( (𝑝−1)×(𝑝−1) ) ,
𝑿(𝑝−1)×1 (𝑝−1)×1 𝝈(1) 𝚺22

𝑇
where 𝐸(𝑋1 ) = 𝜇1 , 𝐸(𝑿(2) ) = 𝝁(2) = (𝜇2 , 𝜇3 , ⋯ , 𝜇𝑝 ) , 𝐸{(𝑋𝑖 − 𝜇𝑖 )(𝑋𝑗 − 𝜇𝑗 )} = 𝜎𝑖𝑗 ; 𝑖, 𝑗 = 1(1)𝑝

𝑇
𝐸 {(𝑋1 − 𝜇1 )(𝑿(2) − 𝝁(2) ) } = 𝛔𝑇(1) = (𝜎12 , 𝜎13 , ⋯ , 𝜎1𝑝 ) , and

𝑇
𝐸 {(𝑿(2) − 𝝁(2) )(𝑿(2) − 𝝁(2) ) } = 𝚺22 .

We know that, the conditional distribution of 𝑋1 given for fixed 𝑿(2) = 𝒙(2) is given by

(𝑋1 |𝑿(2) = 𝒙(2) ) ~ 𝑁1 ({𝜇1 + 𝛔𝑇(1) 𝚺22


−1
(𝒙(2) − 𝝁(2) )}, 𝜎11.(2) ), 𝑇
where 𝜎11.(2) = 𝜎11 − 𝛔(1) −1
𝚺22 𝝈(1) ,

(2) 𝑇
The conditional expectation and variance of 𝑋1 given the fixed values 𝒙(𝑝−1)×1 = (𝑥2 , 𝑥3 , ⋯ , 𝑥𝑝 ) of the predictor
variables are, respectively,

𝐸(𝑋1 |𝑿(2) = 𝒙(2) ) = 𝜇1 + 𝛔𝑇(1) 𝚺22


−1 𝑇
(𝒙(2) − 𝝁(2) ) and 𝑉𝑎𝑟(𝑋1 |𝑿(2) = 𝒙(2) ) = 𝜎11.(2) = 𝜎11 − 𝛔(1) −1
𝚺22 𝝈(1) .

This conditional expected value, considered as a function of 𝑥2 , 𝑥3 , ⋯ , 𝑥𝑝 is called the multivariate regression of 𝑋1 on
𝑿(2) = 𝒙(2) . Hence, the function of 𝑥2 , 𝑥3 , ⋯ , 𝑥𝑝 ,

𝑋1.23⋯𝑝 = 𝑋1.(2) = 𝐸(𝑋1 |𝑿(2) = 𝒙(2) ) = 𝜇1 + 𝛔𝑇(1) 𝚺22


−1
(𝒙(2) − 𝝁(2) )

is the linear regression function, which minimizes the mean square error for the prediction of 𝑋1 on 𝑿(2) = 𝒙(2) .

𝑇 −1 𝑇
−1
The vector 𝜷(𝑝−1)×1 = (𝛔(1) 𝚺22 ) = 𝚺22 𝝈(1) is the vector of regression coefficients.

The 𝑘-th element of 𝜷(𝑝−1)×1 , denoted by, 𝛽1𝑘.23⋯(𝑘−1)(𝑘+1)⋯𝑝 is the regression coefficient of 𝑋1 on 𝑋𝑘 = 𝑥𝑘 ; 𝑘 = 2(1)𝑝.

The correlation between 𝑋1 and its best linear predictor 𝑋1.23⋯𝑝 = 𝑋1.(2) = 𝐸(𝑋1 |𝑿(2) = 𝒙(2) ) is called the population
multiple correlation coefficient.

The population multiple correlation coefficient 𝑅 (𝑅1.(2) or 𝜌1.(2) or 𝑅1.23⋯𝑝 ) is given by

𝑅 = 𝑅1.23⋯𝑝 = 𝑅1.(2) = 𝜌1.(2) = 𝐶𝑜𝑟𝑟(𝑋1 , 𝑋1.23⋯𝑝 ) = 𝐶𝑜𝑟𝑟(𝑋1 , 𝑋1.(2) )

= 𝐶𝑜𝑟𝑟[𝑋1 , {𝜇1 + 𝛔𝑇(1) 𝚺22


−1
(𝒙(2) − 𝝁(2) )}]

𝑉𝑎𝑟(𝑋1.(2) ) 𝑉𝑎𝑟(𝛔𝑇(1) 𝚺22


−1 (2)
𝑿 )
= +√ = +√
𝑉𝑎𝑟(𝑋1 ) 𝜎11

𝛔𝑇(1) 𝚺22
−1
𝝈(1)
= +√ ; [∵ 𝑿(2) ~ 𝑵(𝒑−𝟏) (𝝁(2) , 𝚺22 )]
𝜎11

The square of the population multiple correlation coefficient, 𝑅 (𝑅1.(2) or 𝜌1.(2) or 𝑅1.23⋯𝑝 ), is called the population
coefficient of determination. The multiple correlation coefficient is a positive square root, so 0 ≤ 𝑅 ≤ 1.

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 27


MULTIVARIATE NORMAL DISTRIBUTION

𝑇
The error of prediction 𝜀1.23⋯𝑝 = 𝜀1.(2) = (𝑋1 − 𝑋1.(2) ) = 𝑋1 − 𝜇1 − 𝛔(1) −1
𝚺22 (𝒙(2) − 𝝁(2) ) has the expected squares and
cross-products (error-variation) matrix

𝜎11.(2) = 𝜎11.23⋯𝑝 = 𝑉𝑎𝑟(𝜀1.(2) ) = 𝑉𝑎𝑟(𝜀1.23⋯𝑝 )

𝑇 2
= 𝐸{𝑋1 − 𝜇1 − 𝛔(1) −1
𝚺22 (𝒙(2) − 𝝁(2) )}

= 𝜎11 − 𝛔𝑇(1) 𝚺22


−1
𝝈(1) − 𝛔𝑇(1) 𝚺22
−1
𝝈(1) + 𝛔𝑇(1) 𝚺22
−1 −1
𝚺22 𝚺22 𝝈(1)

= 𝜎11 − 𝛔𝑇(1) 𝚺22


−1
𝝈(1) .

It can be shown that 𝜀1.(2) = (𝑋1 − 𝑋1.(2) ) ~ 𝑁1 (0, 𝜎11.(2) ).

Suppose

𝑋1 𝜇1 𝜎11 𝜎12 𝛔𝑇(1)


𝑋 𝜇
𝑿𝑝×1 =( 2 ), 𝝁𝑝×1 =( 2
(3)
), 𝚺 = ( 𝜎21 𝜎22 𝛔𝑇(2) ) ,
(3)
𝑿(𝑝−2)×1 𝝁(𝑝−2)×1 𝝈(1) 𝝈(2) 𝚺33

𝑇
where 𝐸(𝑋𝑖 ) = 𝜇𝑖 ; 𝑖 = 1, 2, 𝐸(𝑿(3) ) = 𝝁(3) = (𝜇3 , 𝜇4 , ⋯ , 𝜇𝑝 ) , 𝐸{(𝑋𝑖 − 𝜇𝑖 )(𝑋𝑗 − 𝜇𝑗 )} = 𝜎𝑖𝑗 ; 𝑖, 𝑗 = 1(1)𝑝,

𝑇
𝐸 {(𝑋1 − 𝜇1 )(𝑿(3) − 𝝁(3) ) } = 𝛔𝑇(1) = (𝜎13 , 𝜎14 , ⋯ , 𝜎1𝑝 ),

𝑇
𝐸 {(𝑋2 − 𝜇2 )(𝑿(3) − 𝝁(3) ) } = 𝛔𝑇(2) = (𝜎23 , 𝜎24 , ⋯ , 𝜎2𝑝 ), and

𝑇 (𝑝−2)×(𝑝−2)
𝐸 {(𝑿(3) − 𝝁(3) )(𝑿(3) − 𝝁(3) ) } = 𝚺33 .

Consider the pair of errors

𝜀1.(3) = 𝜀1.34⋯𝑝 = 𝑋1 − 𝜇1 − 𝛔𝑇(1) 𝚺33


−1
(𝑿(3) − 𝝁(3) )

𝜀2.(3) = 𝜀2.34⋯𝑝 = 𝑋2 − 𝜇2 − 𝛔𝑇(2) 𝚺33


−1
(𝑿(3) − 𝝁(3) )

obtained from the best linear predictors to predict 𝑋1 and 𝑋2 on 𝑿(3) = 𝒙(3) .

The variances of the errors are given by

𝑇 −1
𝜎11.(3) = 𝜎11.34⋯𝑝 = 𝑉𝑎𝑟(𝜀1.(3) ) = 𝑉𝑎𝑟(𝜀1.34⋯𝑝 ) = 𝜎11 − 𝛔(1) 𝚺33 𝝈(1) ,

𝑇 −1
𝜎22.(3) = 𝜎22.34⋯𝑝 = 𝑉𝑎𝑟(𝜀2.(3) ) = 𝑉𝑎𝑟(𝜀2.34⋯𝑝 ) = 𝜎22 − 𝛔(2) 𝚺33 𝝈(2) , and

𝜎12.(3) = 𝜎12.34⋯𝑝 = 𝐶𝑜𝑣(𝜀1.(3) , 𝜀2.(3) ) = 𝐶𝑜𝑣(𝜀1.34⋯𝑝 , 𝜀2.34⋯𝑝 )

The correlation between 𝜀1.(3) and 𝜀2.(3) is called the population partial correlation coefficient between 𝑋1 and 𝑋2
eliminating the effects of 𝑿(3) . The population partial correlation coefficient between 𝑋1 and 𝑋2 eliminating the effects
𝑇
of 𝑿(3) = (𝑋3 , 𝑋4 , ⋯ , 𝑋𝑝 ) , denoted by 𝜌12.(3) or, 𝜌12.34⋯𝑝 , is given by

𝜌12.(3) = 𝜌12.34⋯𝑝 = 𝐶𝑜𝑟𝑟(𝜀1.(3) , 𝜀2.(3) ) = 𝐶𝑜𝑟𝑟(𝜀1.34⋯𝑝 , 𝜀2.34⋯𝑝 )

= 𝐶𝑜𝑟𝑟[{𝑋1 − 𝜇1 − 𝛔𝑇(1) 𝚺33


−1
(𝑿(3) − 𝝁(3) )}, {𝑋2 − 𝜇2 − 𝛔𝑇(2) 𝚺33
−1
(𝑿(3) − 𝝁(3) )}]

𝜎12.(3) 𝜎12.34⋯𝑝
= =
√𝜎11.(3) ∙ √𝜎22.(3) √𝜎11.34⋯𝑝 ∙ √𝜎22.34⋯𝑝

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 28


MULTIVARIATE NORMAL DISTRIBUTION

TEST OF HYPOTHESES FOR MULTIPLE CORRELATION COEFFICIENT:

CASE I:
2
Suppose we want to test the null hypothesis 𝐻0 ∶ 𝜌𝑖.(2) = 0, where 𝜌𝑖.(2) is the population multiple correlation coefficient
𝑇
between 𝑋𝑖 ; 𝑖 = 1(1)𝑚 and 𝑿(2) = (𝑋𝑚+1 , 𝑋𝑚+2 , ⋯ , 𝑋𝑝 ) .

𝑇
Let 𝑅𝑖.(2) ; 𝑖 = 1(1)𝑚 be the sample multiple correlation coefficient between 𝑋𝑖 and 𝑿(2) = (𝑋𝑚+1 , 𝑋𝑚+2 , ⋯ , 𝑋𝑝 )
𝑇
based on a sample [(𝑋1𝑘 , 𝑋2𝑘 , ⋯ , 𝑋𝑝𝑘 ) ; 𝑘 = 1(1)𝑁] of size 𝑁 from 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺).

𝑇
The sample multiple correlation coefficient 𝑅𝑖.(2) between 𝑋𝑖 and 𝑿(2) = (𝑋𝑚+1 , 𝑋𝑚+2 , ⋯ , 𝑋𝑝 ) is given by

𝑇 −1
𝑇 −1 (2) 𝐬(𝑖) 𝐒22 𝒔(𝑖)
𝑅𝑖.(2) = 𝐶𝑜𝑟𝑟(𝑋𝑖 , 𝑋𝑖.(2) ) = 𝐶𝑜𝑟𝑟 [𝑋𝑖 , {𝑋𝑖 + 𝐬(𝑖) 𝐒22 (𝑿(2) − 𝑿 )}] = +√ ; 𝑖 = 1(1)𝑚
𝑠𝑖𝑖

2 −1
The null hypothesis 𝐻0 ∶ 𝜌𝑖.(2) = 0 is equivalent to 𝐻0 ∶ 𝜷𝑖.(2) = 𝟎, since 𝜷𝑖.(2) = 𝚺22 𝝈(𝑖) and 𝝈(𝑖) = 𝟎.

If the population multiple correlation coefficient 𝜌𝑖.(2) = 0, the statistic

2
𝑅𝑖.(2) {𝑁 − (𝑝 − 𝑚) − 1}
𝐹= 2 × ~ 𝐹(𝑝−𝑚),{𝑁−(𝑝−𝑚)−1}
1− 𝑅𝑖.(2) (𝑝 − 𝑚)

is distributed as central 𝐹-distribution with (𝑝 − 𝑚) and {𝑁 − (𝑝 − 𝑚) − 1} degrees of freedom.

2
To test 𝐻0 ∶ 𝜌𝑖.(2) = 0 or equivalently 𝐻0 ∶ 𝜷𝑖.(2) = 𝟎, we use the likelihood ratio test based on 𝐹 and we reject 𝐻0 if

2
𝑅𝑖.(2) {𝑁 − (𝑝 − 𝑚) − 1}
(Observed ) 𝐹 = 2 × ≥ 𝐹(𝑝−𝑚),{𝑁−(𝑝−𝑚)−1} (𝛼)
1− 𝑅𝑖.(2) (𝑝 − 𝑚)

where 𝐹(𝑝−𝑚),{𝑁−(𝑝−𝑚)−1} (𝛼) is the upper significance point of 𝐹(𝑝−𝑚),{𝑁−(𝑝−𝑚)−1} distribution corresponding to the level

of significance 𝛼 i.e., 𝑃𝐻0 (𝐹 ≥ 𝐹(𝑝−𝑚),{𝑁−(𝑝−𝑚)−1} (𝛼)) = 𝛼.

CASE II:
2
2
Suppose we want to test the null hypothesis 𝐻0 ∶ 𝜌1.23⋯𝑝 = 0 (or 𝑅1.23⋯𝑝 = 0), where 𝜌1.23⋯𝑝 (or 𝑅1.23⋯𝑝 or 𝑅) is the
𝑇
population multiple correlation coefficient between 𝑋1 and 𝑿(2) = (𝑋2 , 𝑋3 , ⋯ , 𝑋𝑝 ) .

𝑇
Let 𝑅1.23⋯𝑝 or 𝑅 be the sample multiple correlation coefficient between 𝑋1 and 𝑿(2) = (𝑋2 , 𝑋3 , ⋯ , 𝑋𝑝 ) based on a
𝑇
sample [(𝑋1𝑘 , 𝑋2𝑘 , ⋯ , 𝑋𝑝𝑘 ) ; 𝑘 = 1(1)𝑁] of size 𝑁 from 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺).

𝑇
The sample multiple correlation coefficient 𝑅1.23⋯𝑝 or 𝑅 between 𝑋1 and 𝑿(2) = (𝑋2 , 𝑋3 , ⋯ , 𝑋𝑝 ) is given by

𝑇 −1
𝑇
(2) 𝐬(1) 𝐒22 𝒔(1)
𝑅 = 𝑅1.23⋯𝑝 = 𝐶𝑜𝑟𝑟(𝑋1 , 𝑋1.23⋯𝑝 ) = 𝐶𝑜𝑟𝑟 [𝑋1 , {𝑋1 + 𝐬(1) −1
𝐒22 (𝑿(2) − 𝑿 )}] = +√ .
𝑠11

2 −1
The null hypothesis 𝐻0 ∶ 𝜌1.23⋯𝑝 = 0 is equivalent to 𝐻0 ∶ 𝜷1.23⋯𝑝 = 𝟎, since 𝜷 = 𝜷1.23⋯𝑝 = 𝚺22 𝝈(1) and 𝝈(1) = 𝟎.

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 29


MULTIVARIATE NORMAL DISTRIBUTION

If the population multiple correlation coefficient 𝜌1.23⋯𝑝 = 0, the statistic

𝑅2 (𝑁 − 𝑝)
𝐹= × ~ 𝐹(𝑝−1),(𝑁−𝑝)
1−𝑅 2 (𝑝 − 1)

is distributed as central 𝐹-distribution with (𝑝 − 1) and (𝑁 − 𝑝) degrees of freedom.

2
To test 𝐻0 ∶ 𝜌1.23⋯𝑝 = 0 or equivalently 𝐻0 ∶ 𝜷 = 𝟎, we use the likelihood ratio test based on 𝐹 and we reject 𝐻0 if

𝑅2 (𝑁 − 𝑝)
(Observed ) 𝐹 = × ≥ 𝐹(𝑝−1),(𝑁−𝑝) (𝛼)
1 − 𝑅2 (𝑝 − 1)

where 𝐹(𝑝−1),(𝑁−𝑝) (𝛼) is the upper significance point of 𝐹(𝑝−1),(𝑁−𝑝) distribution corresponding to the level of

significance 𝛼 i.e., 𝑃𝐻0 (𝐹 ≥ 𝐹(𝑝−1),(𝑁−𝑝) (𝛼)) = 𝛼.

NOTE:

We have
𝑇 −1 𝑇 −1 𝑇 −1
2
𝐬(1) 𝐒22 𝒔(1) 2
𝐬(1) 𝐒22 𝒔(1) 𝑠11 − 𝐬(1) 𝐒22 𝒔(1) 𝑠11.(2)
𝑅 = or, 1−𝑅 =1− = =
𝑠11 𝑠11 𝑠11 𝑠11
𝑇 −1
𝑅2 𝐬(1) 𝐒22 𝒔(1)
∴ 2
=
1−𝑅 𝑠11.(2)

When 𝜌1.23⋯𝑝 = 0 or, 𝑅1.23⋯𝑝 = 0, that is, when 𝜷1.23⋯𝑝 = 𝟎 or when 𝛽1𝑘.23⋯(𝑘−1)(𝑘+1)⋯𝑝 = 0 for each 𝑘 = 2(1)𝑝,
𝑇 𝑁−𝑝 𝑇
𝑠11.(2) = 𝑠11 − 𝐬(1) −1
𝐒22 𝒔(1) is distributed as ∑𝑘=1 𝑍𝑘2 and 𝐬(1) −1
𝐒22 𝒔(1) is distributed as ∑𝑁−1 2
𝑘=𝑁−𝑝+1 𝑍𝑘 , where

𝑍1 , 𝑍2 , ⋯ , 𝑍𝑁−1 are independent, each with distribution 𝑁1 (0, 𝜎11.(2) ).

𝑇 −1 𝑇 −1
𝑠11.(2) 𝑠11..23⋯𝑝 𝐬(1) 𝐒22 𝒔(1) 𝐬(1) 𝐒22 𝒔(1)
Then = and = are distributed independently as 𝜒 2 -variables with (𝑁 − 𝑝) and (𝑝 − 1)
𝜎11.(2) 𝜎11..23⋯𝑝 𝜎11.(2) 𝜎11..23⋯𝑝

degrees of freedom, respectively. Thus


𝑇 −1
𝑅2 (𝑁 − 𝑝) 𝐬(1) 𝐒22 𝒔(1) ⁄𝜎11.(2) (𝑁 − 𝑝)
× = × ~ 𝐹(𝑝−1),(𝑁−𝑝)
1−𝑅 2 (𝑝 − 1) 𝑠11.(2) ⁄𝜎11.(2) (𝑝 − 1)

has the central 𝐹-distribution with (𝑝 − 1) and (𝑁 − 𝑝) degrees of freedom.

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 30


MULTIVARIATE NORMAL DISTRIBUTION

TEST OF HYPOTHESES FOR PARTIAL CORRELATION COEFFICIENT:

CASE I:

Suppose we want to test the null hypothesis 𝐻0 ∶ 𝜌𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 = 𝜌0 (Given) against two-sided alternative hypothesis
𝐻1 ∶ 𝜌𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 ≠ 𝜌0 , where 𝜌𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 ; 𝑖, 𝑘 = 1(1)𝑚 is the population partial correlation coefficient
𝑇
between 𝑋𝑖 and 𝑋𝑘 ; 𝑖, 𝑘 = 1(1)𝑚 eliminating the effects of 𝑿(2) = (𝑋𝑚+1 , 𝑋𝑚+2 , ⋯ , 𝑋𝑝 ) .

Suppose 𝑟𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 (or 𝑟𝑖𝑘.(2) ) is the sample partial correlation coefficient based on a sample
𝑇
[(𝑋1𝑘 , 𝑋2𝑘 , ⋯ , 𝑋𝑝𝑘 ) ; 𝑘 = 1(1)𝑁] of size 𝑁 from 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺) with population partial
correlation coefficient 𝜌𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 (or 𝜌𝑖𝑘.(2) ); 𝑖, 𝑘 = 1(1)𝑚.

−1
Let 𝑠𝑖𝑘.(2) = 𝑠𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 = 𝐶𝑜𝑣(𝑒𝑖.(2) , 𝑒𝑘.(2) ) be the (𝑖, 𝑘)-th element of 𝐒11.2 = 𝐒11 − 𝐒12 𝐒22 𝐒21 ; 𝑖, 𝑘 = 1(1)𝑚.
𝑇
The sample partial correlation coefficient between 𝑋𝑖 and 𝑋𝑘 eliminating the effects of 𝑿(2) = (𝑋𝑚+1 , 𝑋𝑚+2 , ⋯ , 𝑋𝑝 ) ,
denoted by 𝑟𝑖𝑘.(2) or, 𝑟𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 , is given by

𝑠𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝
𝑟𝑖𝑘.(2) = 𝑟𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 = 𝐶𝑜𝑟𝑟(𝑒𝑖.(2) , 𝑒𝑘.(2) ) = ; 𝑖, 𝑘 = 1(1)𝑚.
√𝑠𝑖𝑖.(𝑚+1)(𝑚+2)⋯𝑝 ∙ √𝑠𝑘𝑘.(𝑚+1)(𝑚+2)⋯𝑝

We use Fisher’s 𝑧 statistic for an approximate significance test of 𝐻0 ∶ 𝜌𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 = 𝜌0 against two-sided
alternatives based on a sample [(𝑋1𝑘 , 𝑋2𝑘 , ⋯ , 𝑋𝑝𝑘 ) ; 𝑘 = 1(1)𝑁] of size 𝑁.

Let

1 1 + 𝑟𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 1 1 + 𝜌0
𝑧= 𝑙𝑛 and 𝜉0 = 𝑙𝑛 .
2 1 − 𝑟𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 2 1 − 𝜌0

Then 𝐻0 is rejected if

|√𝑁 − (𝑝 − 𝑚) − 3 × (𝑧 − 𝜉0 )| > 𝜏𝛼
2

where 𝜏𝛼 is the upper significance point of standard normal distribution corresponding to the level of significance 𝛼 i.e.,
2

𝛼
𝑃𝐻0 (𝑍 ≥ 𝜏𝛼 |𝑍 ~ 𝑁(0, 1)) = .
2 2

NOTE:

Since the distribution of a sample partial correlation 𝑟𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 based on a sample of 𝑁 from a distribution
with population correlation 𝜌𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 equal to a certain value, 𝜌, say, is the same as the distribution of a
simple correlation 𝑟 based on a sample of size 𝑁 − (𝑝 − 𝑚) from a distribution with the corresponding population
correlation of 𝜌, all statistical inference procedures for the simple population correlation can be used for the partial
correlation. The procedure for the partial correlation is exactly the same except that 𝑁 is replaced by 𝑁 − (𝑝 − 𝑚).

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 31


MULTIVARIATE NORMAL DISTRIBUTION

CASE II:

Suppose we want to test 𝐻0 ∶ 𝜌12.34⋯𝑝 = 𝜌0 (Given) against two-sided alternative 𝐻1 ∶ 𝜌12.34⋯𝑝 ≠ 𝜌0 , where 𝜌12.34⋯𝑝 is
𝑇
the population partial correlation coefficient between 𝑋1 and 𝑋2 eliminating the effects of 𝑿(3) = (𝑋3 , 𝑋4 , ⋯ , 𝑋𝑝 ) .

Suppose 𝑟12.34⋯𝑝 is the sample partial correlation coefficient based on a sample [(𝑋1𝑘 , 𝑋2𝑘 , 𝑋3𝑘 , ⋯ , 𝑋𝑝𝑘 ); 𝑘 = 1(1)𝑁] of
𝑇
size 𝑁 from 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺) with population partial correlation coefficient 𝜌12.34⋯𝑝 .

Let 𝑠𝑖𝑘.(3) = 𝑠𝑖𝑘.34⋯𝑝 = 𝐶𝑜𝑣(𝑒𝑖.(3) , 𝑒𝑘.(3) ) = 𝐶𝑜𝑣(𝑒𝑖.34⋯𝑝 , 𝑒𝑘.34⋯𝑝 ) ; 𝑖, 𝑘 = 1, 2.

𝑇
The sample partial correlation coefficient between 𝑋1 and 𝑋2 eliminating the effects of 𝑿(3) = (𝑋3 , 𝑋4 , ⋯ , 𝑋𝑝 ) , denoted
by 𝑟12.(3) or, 𝑟12.34⋯𝑝 , is given by

𝑠12.(3) 𝑠12.34⋯𝑝
𝑟12.(3) = 𝑟12.34⋯𝑝 = 𝐶𝑜𝑟𝑟(𝑒1.34⋯𝑝 , 𝑒2.34⋯𝑝 ) = = .
√𝑠11.(3) ∙ √𝑠22.(3) √𝑠11.34⋯𝑝 ∙ √𝑠22.34⋯𝑝

We use Fisher’s 𝑧 statistic for an approximate significance test of 𝐻0 ∶ 𝜌12.34⋯𝑝 = 𝜌0 against two-sided alternatives based
on a sample [(𝑋1𝑘 , 𝑋2𝑘 , ⋯ , 𝑋𝑝𝑘 ) ; 𝑘 = 1(1)𝑁] of size 𝑁.

Let

1 1 + 𝑟12.34⋯𝑝 1 1 + 𝜌0
𝑧= 𝑙𝑛 and 𝜉0 = 𝑙𝑛 .
2 1 − 𝑟12.34⋯𝑝 2 1 − 𝜌0

Then 𝐻0 is rejected if

|√𝑁 − (𝑝 − 2) − 3 × (𝑧 − 𝜉0 )| > 𝜏𝛼
2

where 𝜏𝛼 is the upper significance point of standard normal distribution corresponding to the level of significance 𝛼 i.e.,
2

𝛼
𝑃𝐻0 (𝑍 ≥ 𝜏𝛼 |𝑍 ~ 𝑁(0, 1)) = .
2 2

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 32


MULTIVARIATE NORMAL DISTRIBUTION

RESULT 10:

Let us define

𝑌 𝜇𝑌 𝜎𝑌𝑌 𝛔𝑇(𝑌)
𝒁(𝑝+1)×1 = (𝑿 ) , 𝝁(𝑝+1)×1 = (𝝁𝑝×1 ) , 𝚺 (𝑝+1)×(𝑝+1) = ( 𝑝×𝑝 ) ,
𝑝×1 (𝑋) 𝝈(𝑌) 𝚺𝑋𝑋

𝑇 𝑇
𝑝×1
where 𝐸(𝑌) = 𝜇𝑌 , 𝐸(𝑿𝑝×1 ) = 𝝁(𝑋) = (𝜇𝑋1 , 𝜇𝑋2 , ⋯ , 𝜇𝑋𝑝 ) = (𝜇1 , 𝜇2 , ⋯ , 𝜇𝑝 ) , 𝑉𝑎𝑟(𝑌) = 𝜎𝑌𝑌 ,

𝑇
𝛔𝑇(𝑌) = (𝜎𝑌𝑋1 , 𝜎𝑌𝑋2 , ⋯ , 𝜎𝑌𝑋𝑝 ) , 𝜎𝑌𝑋𝑖 = 𝐶𝑜𝑣(𝑌, 𝑋𝑖 ); 𝑖 = 1(1)𝑝, and 𝐸 {(𝑿 − 𝝁(𝑿) )(𝑿 − 𝝁(𝑋) ) } = 𝚺𝑋𝑋 .

𝑌
Suppose 𝒁(𝑝+1)×1 = (𝑿 ) ~ 𝑵(𝒑+𝟏) (𝝁, 𝚺). Then, the conditional distribution of 𝑌 given for fixed 𝑿 = 𝒙 is given by
𝑝×1

𝑇
(𝑌|𝑿 = 𝒙) ~ 𝑵𝟏 ({𝜇𝑌 + 𝛔(𝑌) −1
𝚺𝑋𝑋 (𝒙 − 𝝁(𝑋) )}, 𝜎𝑌𝑌.(𝑋) ), where 𝜎𝑌𝑌.(𝑋) = 𝜎𝑌𝑌 − 𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌) .

The linear predictor (𝛽0 + 𝜷𝑇 𝑿) with coefficients 𝛽0 = 𝜇𝑌 − 𝛔𝑇(𝑌) 𝚺𝑋𝑋


−1
𝝁(𝑋) and 𝜷𝑇 = 𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
has minimum mean square
error among all linear predictors of the response 𝑌 with mean square error

2
𝐸(𝑌 − 𝛽0 − 𝜷𝑇 𝑿)2 = 𝐸 (𝑌 − 𝜇𝑌 − 𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
(𝑿 − 𝝁(𝑋) )) = 𝜎𝑌𝑌 − 𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌) = 𝜎𝑌𝑌.(𝑋) ,

that is, the linear predictor 𝑌𝑌.1 2 ⋯𝑝 = 𝐸(𝑌|𝑥1 , 𝑥, ⋯ , 𝑥𝑝 ) = 𝜇𝑌 + 𝛔𝑇(𝑌) 𝚺𝑋𝑋


−1
(𝑿 − 𝝁(𝑋) ) has minimum mean square error.

Also, 𝑌𝑌.1 2 ⋯𝑝 = 𝐸(𝑌|𝑥1 , 𝑥, ⋯ , 𝑥𝑝 ) = 𝛽0 + 𝜷𝑇 𝑿 = 𝜇𝑌 + 𝛔𝑇(𝑌) 𝚺𝑋𝑋


−1
(𝒙 − 𝝁(𝑋) ) is the linear predictor having maximum
correlation with 𝑌; i.e.,

𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌)
𝜌𝑌.(𝑋) = 𝐶𝑜𝑟𝑟(𝑌, (𝛽0 + 𝜷𝑇 𝑿)) = max{𝐶𝑜𝑟𝑟(𝑌, (𝑏0 + 𝒃𝑇 𝑿))} = +√ .
𝑏0 ,𝒃 𝜎𝑌𝑌

PROOF:
𝑌
Suppose 𝒁(𝑝+1)×1 = (𝑿 ) ~ 𝑵(𝒑+𝟏) (𝝁, 𝚺).
𝑝×1

From RESULT 3, the conditional distribution of 𝑌 given for fixed 𝑿 = 𝒙 is given by

(𝑌|𝑿 = 𝒙) ~ 𝑵𝟏 ({𝜇𝑌 + 𝛔𝑇(𝑌) 𝚺𝑋𝑋


−1
(𝒙 − 𝝁(𝑋) )}, 𝜎𝑌𝑌.(𝑋) ), where 𝜎𝑌𝑌.(𝑋) = 𝜎𝑌𝑌 − 𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌) .

The mean square error of 𝑌 for linear predictor 𝑏0 + 𝒃𝑇 𝑿 is given by

2
𝐸(𝑌 − 𝑏0 − 𝒃𝑇 𝑿)2 = 𝐸{(𝑌 − 𝜇𝑌 ) − 𝒃𝑇 (𝑿 − 𝝁(𝑋) ) + (𝜇𝑌 − 𝑏0 − 𝒃𝑇 𝝁(𝑋) )}

2 2
= 𝐸(𝑌 − 𝜇𝑌 )2 + 𝐸{𝒃𝑇 (𝑿 − 𝝁(𝑋) )} + 𝐸(𝜇𝑌 − 𝑏0 − 𝒃𝑇 𝝁(𝑋) ) − 2𝐸{𝒃𝑇 (𝑿 − 𝝁(𝑋) )(𝑌 − 𝜇𝑌 )}

2
= 𝜎𝑌𝑌 + 𝒃𝑇 𝚺𝑋𝑋 𝒃 + (𝜇𝑌 − 𝑏0 − 𝒃𝑇 𝝁(𝑋) ) − 2𝒃𝑇 𝝈(𝑌)

2
= 𝜎𝑌𝑌 − 𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌) + (𝜇𝑌 − 𝑏0 − 𝒃𝑇 𝝁(𝑋) ) + 𝒃𝑇 𝚺𝑋𝑋 𝒃 − 2𝒃𝑇 𝝈(𝑌) + 𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌)

2 𝑇
= 𝜎𝑌𝑌 − 𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌) + (𝜇𝑌 − 𝑏0 − 𝒃𝑇 𝝁(𝑋) ) + (𝒃 − 𝚺𝑋𝑋
−1 −1
𝝈(𝑌) ) 𝚺𝑋𝑋 (𝒃 − 𝚺𝑋𝑋 𝝈(𝑌) )

−1
The mean square error is minimized by taking 𝒃 = 𝚺𝑋𝑋 𝝈(𝑌) = 𝜷 and then choosing 𝑏0 = 𝜇𝑌 + 𝒃𝑇 𝝁(𝑋) = 𝛽0 to make the
last and third term zero.

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 33


MULTIVARIATE NORMAL DISTRIBUTION

Thus, the minimum mean square error is the variance of the conditional distribution of 𝑌 given for fixed 𝑿 = 𝒙,

2
𝐸(𝑌 − 𝛽0 − 𝜷𝑇 𝑿)2 = 𝐸 (𝑌 − 𝜇𝑌 − 𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
(𝒙 − 𝝁(𝑋) )) = 𝜎𝑌𝑌 − 𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌) = 𝜎𝑌𝑌.(𝑋) .

The mean of the conditional distribution

𝑌𝑌.1 2 ⋯𝑝 = 𝐸(𝑌|𝑿 = 𝒙) = 𝐸(𝑌|𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) = 𝛽0 + 𝜷𝑇 𝑿 = 𝜇𝑌 + 𝛔𝑇(𝑌) 𝚺𝑋𝑋


−1
(𝒙 − 𝝁(𝑋) )

is the best linear predictor of 𝑌 when the population 𝑵(𝒑+𝟏) (𝝁, 𝚺).

Now, 𝑉𝑎𝑟(𝑏0 + 𝒃𝑇 𝑿) = 𝑉𝑎𝑟(𝒃𝑇 𝑿) = 𝒃𝑇 𝚺𝑋𝑋 𝒃 and 𝐶𝑜𝑣(𝑌, (𝑏0 + 𝒃𝑇 𝑿)) = 𝐶𝑜𝑣(𝑌, 𝒃𝑇 𝑿) = 𝒃𝑇 𝝈(𝑌)

2
2 {𝒃𝑇 𝝈(𝑌) }
∴ {𝐶𝑜𝑟𝑟(𝑌, (𝑏0 + 𝒃𝑇 𝑿))} = ; for all 𝑏0 and 𝒃
𝜎𝑌𝑌 (𝒃𝑇 𝚺𝑋𝑋 𝒃)

Since 𝚺𝑋𝑋 is a positive definite, there exist a non-singular matrix 𝑪 such that 𝑪𝑪𝑇 = 𝚺𝑋𝑋 .

Now, 𝒃𝑇 𝚺𝑋𝑋 𝒃 = 𝒃𝑇 𝑪𝑪𝑇 𝒃 = (𝑪𝑇 𝒃)𝑻 (𝑪𝑇 𝒃) and

𝑻
𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌) = 𝛔𝑇(𝑌) (𝑪𝑪𝑇 )−1 𝝈(𝑌) = (𝑪−1 𝝈(𝑌) ) (𝑪−1 𝝈(𝑌) ) [∵ (𝑪−1 )𝑇 = (𝑪𝑇 )−1 ]

𝑇 𝑇
By Cauchy-Schwarz Inequality, we have, if 𝒂𝑝×1 = (𝑎1 , 𝑎2 , ⋯ , 𝑎𝑝 ) and 𝒅𝑝×1 = (𝑑1 , 𝑑2 , ⋯ , 𝑑𝑝 ) be two 𝑝 × 1 vectors
then

𝑝 2 𝑝 𝑝

(𝒂𝑇 2
𝒅) ≤ (𝒂𝑇 𝑇
𝒂)(𝒅 𝒅) or, (∑ 𝑎𝑖 𝑑𝑖 ) ≤ (∑ 𝑎𝑖2 ) (∑ 𝑑𝑖2 )
𝑖=1 𝑖=1 𝑖=1

with equality holds if and only if 𝒂 = 𝑘𝒅 for some constant 𝑘.

Let 𝒂 = (𝑪𝑇 𝒃) and 𝒅 = (𝑪−1 𝝈(𝑌) ).

∴ 𝒂𝑇 𝒅 = (𝑪𝑇 𝒃)𝑻 (𝑪−1 𝝈(𝑌) ) = 𝒃𝑇 𝝈(𝑌)

By Cauchy-Schwarz Inequality, we obtain

2 𝑻
{𝒃𝑇 𝝈(𝑌) } ≤ {(𝑪𝑇 𝒃)𝑻 (𝑪𝑇 𝒃)} {(𝑪−1 𝝈(𝑌) ) (𝑪−1 𝝈(𝑌) )}

2
or, {𝒃𝑇 𝝈(𝑌) } ≤ (𝒃𝑇 𝚺𝑋𝑋 𝒃)(𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌) )

2
{𝒃𝑇 𝝈(𝑌) } 1 1
or, × ≤ (𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌) ) ×
(𝒃 𝚺𝑋𝑋 𝒃) 𝜎𝑌𝑌
𝑇 𝜎𝑌𝑌

2 𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌)
or, {𝐶𝑜𝑟𝑟(𝑌, (𝑏0 + 𝒃𝑇 𝑿))} ≤
𝜎𝑌𝑌
−1
with equality holds for 𝒃 = 𝚺𝑋𝑋 𝝈(𝑌) = 𝜷.

Hence, 𝑌𝑌.1 2 ⋯𝑝 = 𝐸(𝑌|𝑥1 , 𝑥, ⋯ , 𝑥𝑝 ) = 𝛽0 + 𝜷𝑇 𝑿 = 𝜇𝑌 + 𝛔𝑇(𝑌) 𝚺𝑋𝑋


−1
(𝒙 − 𝝁(𝑋) ) is the linear predictor having maximum
correlation with 𝑌; i.e.,

𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌)
𝜌𝑌.(𝑋) = 𝜌𝑌.1 2 ⋯𝑝 = max{𝐶𝑜𝑟𝑟(𝑌, (𝑏0 + 𝒃𝑇 𝑿))} = 𝐶𝑜𝑟𝑟(𝑌, (𝛽0 + 𝜷𝑇 𝑿)) = +√ .
𝑏0 ,𝒃 𝜎𝑌𝑌

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 34


MULTIVARIATE NORMAL DISTRIBUTION

The alternative expression for the maximum correlation between 𝑌 and its best linear predictor or population multiple
−1
correlation coefficient 𝜌𝑌.1 2 ⋯𝑝 with 𝒃 = 𝚺𝑋𝑋 𝝈(𝑌) = 𝜷 is

𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌) 𝛔𝑇(𝑌) 𝜷 𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝚺𝑋𝑋 𝜷 𝜷𝑇 𝚺𝑋𝑋 𝜷
𝜌𝑌.1 2 ⋯𝑝 = 𝐶𝑜𝑟𝑟(𝑌, (𝛽0 + 𝜷𝑇 𝑿)) = +√ = +√ = +√ = +√
𝜎𝑌𝑌 𝜎𝑌𝑌 𝜎𝑌𝑌 𝜎𝑌𝑌

NOTE:

1. When the population is not normal, the regression function 𝐸(𝑌|𝑥1 , 𝑥, ⋯ , 𝑥𝑝 ) need not be of the form 𝛽0 + 𝜷𝑇 𝑿.
It can be shown that 𝐸(𝑌|𝑥1 , 𝑥, ⋯ , 𝑥𝑝 ), whatever its form, predicts 𝑌 with the smallest mean square error. This
wider optimality among all estimators is possessed by the linear predictor when the population is normal.

2. If 𝜌𝑌.(𝑋) = 𝜌𝑌.1 2 ⋯𝑝 = 0, there is no predictive power in 𝑿. At the other extreme, 𝜌𝑌.(𝑋) = 𝜌𝑌.1 2 ⋯𝑝 = 1 implies
that 𝑌 can be predicted by 𝑿 with no error.

2
3. To test 𝐻0 ∶ 𝜌𝑌.1 2 ⋯𝑝 = 0 or equivalently 𝐻0 ∶ 𝜷 = 𝟎, we use the likelihood ratio test based on 𝐹 and we reject

𝐻0 if

𝑅2 (𝑁 − 𝑝 − 1)
(Observed ) 𝐹 = × ≥ 𝐹𝑝,(𝑁−𝑝−1) (𝛼)
1 − 𝑅2 𝑝

where 𝐹𝑝,(𝑁−𝑝−1) (𝛼) is the upper significance point of 𝐹𝑝,(𝑁−𝑝−1) distribution corresponding to the level of

significance 𝛼 i.e., 𝑃𝐻0 (𝐹 ≥ 𝐹𝑝,(𝑁−𝑝−1) (𝛼)) = 𝛼 and 𝑅 is the sample multiple correlation coefficient.

Multivariate Analysis || Lecture Notes || Anup Kumar Giri 35

Common questions

Powered by AI

Transformations involving the multivariate normal distribution include linear transformations that standardize the variables. Given a symmetric positive-definite matrix A, one can use a transformation Y=CT(X-b) where C is such that CC^T=A, facilitating variable interpretation in a standard normal space. These transformations simplify multivariate normal distribution analysis by converting it into independent normal distributions, isolating variable effects, and clarifying the geometric structure of ellipsoidal distributions in the data .

The joint moments about origin of order (h1 + h2 + ⋯+ hp) for non-negative integers (h1, h2, ⋯, hp) is defined as E(X1^h1 X2^h2 ⋯ Xp^hp). If X ~ is discrete, this is given by ∑∑⋯∑ x1^h1 x2^h2 ⋯ xp^hp∙p_X~(x1, x2, ⋯, xp). If X ~ is continuous, it is defined by the integral form. The significance lies in its ability to capture the higher order dependencies among variables, thus facilitating the understanding of complex relationships in multivariate distributions .

In the context of correlation matrices in multivariate distributions, the cofactor of a matrix element is crucial for computing the adjugate, which is key in finding the inverse of a correlation matrix when it is invertible. This inverse is used for variance-covariance matrix computations, allowing the determination of conditional distributions and other multivariate statistics. The cofactor here specifically refers to the minor of an element, which is used to capture the dependencies between the different variables .

In a multivariate normal distribution, random variables are independent if the variance-covariance matrix is diagonal with non-zero elements only along the diagonal. This occurs when Σ is structured as σ2I_p, meaning all off-diagonal elements (covariances) are zero, allowing each variable to vary independently of the others. Independence is inferred from the lack of any linear relationship amongst variables .

The moment generating function (MGF) of a multivariate distribution serves to uniquely define the distribution by encapsulating all of its moments. In particular, the MGF M_X(t) = E{exp(t^T X)} allows us to derive any order of moments by differentiation, reflecting the dependency structures among variables. Its form as a function of t reveals the distribution's parameters and the interrelationships among its components, assisting in hypothesis testing and predictions .

A multivariate distribution becomes singular when its covariance matrix has a determinant equal to zero (|Σk×k| = 0), indicating that there is a linear dependency among the variables. This implies that not all variables can vary independently, reducing their effective dimensionality by one. Consequently, it complicates computation and interpretation, necessitating analysis of a (k-1) dimensional subset, as full-rank operations like finding inverses are not possible in singular cases .

The singularity of a variance-covariance matrix signifies a linear dependence among variables, leading to complications in performing statistical inference because operations such as inversion, crucial for determining conditional distributions and testing hypotheses, are not feasible. This impacts estimations, revealing the need for dimensionality reduction techniques or regularization methods to obtain meaningful statistical insights .

Central moments are calculated as E{(X - μ)^k}, focusing on deviations from the mean, thus describing the shape of the data relative to its mean. In contrast, joint moments about the origin, E(X1^h1 X2^h2 ⋯ Xp^hp), do not consider the mean but rather the raw moments of the joint distribution. This distinction is crucial because central moments provide insights into variability and shape beyond mere magnitudes, such as variance and covariance, which are vital for understanding dependency and relationships in multivariate data .

The correlation matrix of a multivariate random vector, such as R(k-1)×(k-1), shows the linear relationships between each pair of variables in terms of correlation coefficients. The diagonal elements are all 1, representing the perfect correlation of a variable with itself. Off-diagonal elements like -√(pi pj)/((1−pi)(1−pj)) capture the degree and direction of linear dependence between variables under the constraint that each probability pi is less than 1 .

The multinomial coefficient in a multivariate probability distribution is derived from the expression of the probability mass function as n!/(x1! x2!⋯xs!) (n−∑ xi)! ⋅ p1^x1 p2^x2 ⋯ ps^xs, where xi are the counts of events of differing types, n is the total count, and pi are the probabilities of each type of event. This coefficient represents the number of distinct ways the outcomes can occur, given the constraints of the model .

You might also like