Multivariate Probability Distribution
Multivariate Probability Distribution
RANDOM VECTOR
𝑇
𝑋(𝜔) = (𝑋1 (𝜔), 𝑋2 (𝜔), ⋯ , 𝑋𝑝 (𝜔)) ; 𝜔 ∈ Ω,
~
is called a p-dimensional random vector if the inverse image of each p-dimensional interval
𝑇
𝐼(𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) = {𝑥 = (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) : 𝑥 ∈ ℛ𝑝 or 𝑥𝑖 ∈ ℛ, 𝑖 = 1(1)𝑝}
~ ~
𝑇
Suppose 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 are 𝑝 one dimensional random variable. The vector 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) is a p-
~
𝑇
The random vector 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) may be discrete, continuous or mixed.
~
DISTRIBUTION FUNCTION:
𝑇
The cumulative distribution function (c.d.f.) or d.f. of a random vector 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) is defined as
~
𝐹𝑋 (𝑥 ) = 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) = 𝑃(𝑋1 ≤ 𝑥1 , 𝑋2 ≤ 𝑥2 , ⋯ , 𝑋𝑝 ≤ 𝑥𝑝 ).
~ ~ ~
(𝑖𝑖) continuous from the right with respect to all the realisations or arguments 𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ,
(𝑖𝑖𝑖) 0 ≤ 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ≤ 1,
~
𝑇
(𝑣𝑖) For all 𝑥 = (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ∈ ℛ𝑝 and for all ℎ𝑖 > 0 ( 𝑖 = 1(1)𝑝), the following inequality holds.
~
+ ⋯ + (−1)𝑝 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ≥ 0
~
𝑝𝑋 (𝑥 ) = 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) = 𝑃(𝑋1 = 𝑥1 , 𝑋2 = 𝑥2 , ⋯ , 𝑋𝑝 = 𝑥𝑝 )
~ ~ ~
𝑥𝑝 𝑥𝑝−1 𝑥1
𝑇
The function 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) is called the probability density function p.d.f. of 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) , which
~ ~
satisfies
∞ ∞ ∞
If the distribution function of a p-dimensional random vector 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) is absolutely continuous and
~
𝜕 𝑝 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 )
~
= 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) (almost everywhere), and
𝜕𝑥1 𝜕𝑥2 ⋯ 𝜕𝑥𝑝 ~
𝑥𝑝 𝑥𝑝−1 𝑥1
𝑇
The cumulative distribution function (c.d.f.) or d.f. of a random vector 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) is given by
~
𝐹𝑋 (𝑥 ) = 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) = 𝑃(𝑋1 ≤ 𝑥1 , 𝑋2 ≤ 𝑥2 , ⋯ , 𝑋𝑝 ≤ 𝑥𝑝 )
~ ~ ~
∑ ∑ ⋯ ∑ 𝑝𝑋 (𝑦1 , 𝑦2 , ⋯ , 𝑦𝑝 ) ; if 𝑋 is discrete
~ ~
𝑦1 ≤𝑥1 𝑦2 ≤𝑥2 𝑦𝑝 ≤𝑥𝑝
= 𝑥𝑝 𝑥𝑝−1 𝑥1
where 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) and 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) are the joint probability mass function (p.m.f.) and probability
~ ~
MARGINAL DISTRIBUTION:
𝑇
A. The marginal cumulative distribution function (c.d.f.) of a random vector 𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑞 ) , (𝑞 <
~
𝑝) is given by
= 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑞 , ∞, ⋯ , ∞)
~
where 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) and 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) are the joint probability mass function (p.m.f.) and probability
~ ~
𝑇
In general, the marginal c.d.f. of random vector 𝑋 (𝑘) = (𝑋𝑖1 , 𝑋𝑖2 , ⋯ , 𝑋𝑖𝑘 ) , (1 ≤ 𝑖1 < 𝑖2 < ⋯ < 𝑖𝑘 ≤ 𝑝, 𝑘 =
~
1(1)𝑝) is given by
𝐹𝑋𝑖1 ,𝑋𝑖2 ,⋯,𝑋𝑖 (𝑥𝑖1 , 𝑥𝑖2 , ⋯ , 𝑥𝑖𝑘 ) = 𝐹𝑋 (𝑘) (𝑥𝑖1 , 𝑥𝑖2 , ⋯ , 𝑥𝑖𝑘 )
𝑘 ~
B. The marginal probability mass function (p.m.f.) 𝑝𝑋1 ,𝑋2 ,⋯,𝑋𝑞 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑞 ) of 𝑋 (1) =
~
𝑇
(𝑋1 , 𝑋2 , ⋯ , 𝑋𝑞 ) , (𝑞 < 𝑝) is given by
𝑇
the marginal probability density function (p.d.f.) 𝑓𝑋1 ,𝑋2 ,⋯,𝑋𝑞 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑞 ) of 𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑞 ) , (𝑞 <
~
𝑝) is given by
∞ ∞
where 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) and 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) are the joint probability mass function (p.m.f.) and probability
~ ~
CONDITIONAL DISTRIBUTION:
𝑇
A. The conditional cumulative distribution function (c.d.f.) of a random vector 𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑞 ) , (𝑞 < 𝑝)
~
where 𝑝𝑿 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) and 𝑓𝑿 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) are the joint probability mass function (p.m.f.) and probability density
function of 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 , respectively.
𝑇 𝑇
In general, the conditional c.d.f. of random vector 𝑋 (𝑘) = (𝑋𝑖1 , 𝑋𝑖2 , ⋯ , 𝑋𝑖𝑘 ) given that 𝑋 (𝑙) = (𝑋𝑗1 , 𝑋𝑗2 , ⋯ , 𝑋𝑗𝑙 ) = 𝑥 (𝑙)
~ ~ ~
or (𝑋𝑗1 = 𝑥𝑗1 , 𝑋𝑗2 = 𝑥𝑗2 , ⋯ , 𝑋𝑗𝑙 = 𝑥𝑗𝑙 ); (1 ≤ 𝑖1 < ⋯ < 𝑖𝑘 ≤ 𝑝, 1 ≤ 𝑗1 < ⋯ < 𝑗𝑙 ≤ 𝑝; 𝑖𝑡 ≠ 𝑗𝑠 ) is given by
𝐹𝑋 (𝑘)|𝑋 (𝑙) (𝑥 (𝑘) |𝑥 (𝑙) ) = 𝐹𝑋 (𝑥𝑖1 , 𝑥𝑖2 , ⋯ , 𝑥𝑖𝑘 |𝑥𝑗1 , 𝑥𝑗2 , ⋯ , 𝑥𝑗𝑙 )
~ ~ ~ ~ 𝑖1 ,⋯,𝑋𝑖𝑘 |𝑋𝑗1 ,⋯,𝑋𝑗𝑙
lim
𝑥𝑡 → +∞
𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 )
~
𝑃(𝑋𝑖1 ≤ 𝑥𝑖1 , ⋯ , 𝑋𝑖𝑘 ≤ 𝑥𝑖𝑘 , 𝑋𝑗1 ≤ 𝑥𝑗1 , ⋯ , 𝑋𝑗𝑙 ≤ 𝑥𝑗𝑙 ) 𝑡≠𝑖1 ,𝑖2 ,⋯,𝑖𝑘 ,𝑗1 ,𝑗2 ,⋯,𝑗𝑙
= =
𝑃(𝑋𝑗1 ≤ 𝑥𝑗1 , ⋯ , 𝑋𝑗𝑙 ≤ 𝑥𝑗𝑙 ) 𝑥𝑡 → +∞
lim 𝐹𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 )
~
𝑡≠𝑗1 ,𝑗2 ,⋯,𝑗𝑙
B. The conditional probability mass function (p.m.f.) 𝑝(𝑋 , ⋯ , 𝑋 |𝑋 , ⋯ , 𝑋 ) (𝑥1 , ⋯ , 𝑥𝑞 |𝑥𝑞+1 , ⋯ , 𝑥𝑝 ) of a random
1 𝑞 𝑞+1 𝑝
𝑇
vector 𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑞 ) , (𝑞 < 𝑝) given that 𝑋 (2) = 𝑥 (2) or 𝑋𝑞+1 = 𝑥𝑞+1 , 𝑋𝑞+2 = 𝑥𝑞+2 , ⋯ , 𝑋𝑝 = 𝑥𝑝 is given by
~ ~ ~
𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑞 , 𝑥𝑞+1 , ⋯ , 𝑥𝑝 )
𝑝 (1) (2) (𝑥 (1) |𝑥 (2) ) = 𝑝(𝑋 , ⋯ , 𝑋 |𝑋 , ⋯ , 𝑋 ) (𝑥1 , ⋯ , 𝑥𝑞 |𝑥𝑞+1 , ⋯ , 𝑥𝑝 ) =
(𝑋 |𝑋 ) ~ ~ 1 𝑞 𝑞+1 𝑝 𝑝𝑋𝑞+1,𝑋𝑞+2,⋯,𝑋𝑝 (𝑥𝑞+1 , ⋯ , 𝑥𝑝 )
~ ~
𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑞 , 𝑥𝑞+1 , ⋯ , 𝑥𝑝 )
~
= ; if 𝑿 is discrete, and
∑𝑦1∈ ℛ ∑𝑦2 ∈ ℛ ⋯ ∑𝑦𝑞 ∈ ℛ 𝑝𝑋 (𝑦1 , 𝑦2 , ⋯ , 𝑦𝑞 , 𝑥𝑞+1 , ⋯ , 𝑥𝑝 )
~
the marginal probability density function (p.d.f.) 𝑓(𝑋 , ⋯ , 𝑋 |𝑋 , ⋯ , 𝑋 ) (𝑥1 , ⋯ , 𝑥𝑞 |𝑥𝑞+1 , ⋯ , 𝑥𝑝 ) of a random vector
1 𝑞 𝑞+1 𝑝
𝑇
𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑞 ) , (𝑞 < 𝑝) given that 𝑋 (2) = 𝑥 (2) or 𝑋𝑞+1 = 𝑥𝑞+1 , 𝑋𝑞+2 = 𝑥𝑞+2 , ⋯ , 𝑋𝑝 = 𝑥𝑝 is given by
~ ~ ~
𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑞 , 𝑥𝑞+1 , ⋯ , 𝑥𝑝 )
𝑓 (1) (2) (𝑥 (1) |𝑥 (2) ) = 𝑓(𝑋 , ⋯ , 𝑋 |𝑋 , ⋯ , 𝑋 ) (𝑥1 , ⋯ , 𝑥𝑞 |𝑥𝑞+1 , ⋯ , 𝑥𝑝 ) = ~
(𝑋 |𝑋 ) ~ ~ 1 𝑞 𝑞+1 𝑝 𝑓𝑋𝑞+1,𝑋𝑞+2,⋯,𝑋𝑝 (𝑥𝑞+1 , ⋯ , 𝑥𝑝 )
~ ~
𝑓𝑋 (𝑥1 , ⋯ , 𝑥𝑞 , 𝑥𝑞+1 , ⋯ , 𝑥𝑝 )
~
= ∞ ∞ ; if 𝑋 is continuous,
∫−∞ ⋯ ∫−∞ 𝑓𝑋 (𝑦1 , ⋯ , 𝑦𝑞 , 𝑥𝑞+1 , ⋯ , 𝑥𝑝 )𝑑𝑦1 ⋯ 𝑑𝑦𝑞 ~
~
where 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) and 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) are the joint probability mass function (p.m.f.) and probability density
~ ~
STATISTICAL INDEPENDENCE
𝑇 𝑇
The random vector 𝑋 (𝑘) = (𝑋𝑖1 , 𝑋𝑖2 , ⋯ , 𝑋𝑖𝑘 ) is independent of 𝑋 (𝑙) = (𝑋𝑗1 , 𝑋𝑗2 , ⋯ , 𝑋𝑗𝑙 ) ; (1 ≤ 𝑖1 < ⋯ < 𝑖𝑘 ≤ 𝑝, 1 ≤
~ ~
𝑇
𝑗1 < ⋯ < 𝑗𝑙 ≤ 𝑝; 𝑖𝑡 ≠ 𝑗𝑠 ) iff the conditional d.f. of random vector 𝑋 (𝑘) = (𝑋𝑖1 , 𝑋𝑖2 , ⋯ , 𝑋𝑖𝑘 ) given that 𝑋 (𝑙) =
~ ~
𝑇
(𝑋𝑗1 , 𝑋𝑗2 , ⋯ , 𝑋𝑗𝑙 ) = 𝑥 (𝑙) or (𝑋𝑗1 = 𝑥𝑗1 , 𝑋𝑗2 = 𝑥𝑗2 , ⋯ , 𝑋𝑗𝑙 = 𝑥𝑗𝑙 ) is the marginal d.f. of random vector 𝑋 (𝑘) =
~ ~
𝑇
(𝑋𝑖1 , 𝑋𝑖2 , ⋯ , 𝑋𝑖𝑘 ) i.e.,
𝐹𝑋 (𝑘)|𝑋 (𝑙) (𝑥 (𝑘) |𝑥 (𝑙) ) = 𝐹𝑋 (𝑘) (𝑥𝑖1 , 𝑥𝑖2 , ⋯ , 𝑥𝑖𝑘 ) = 𝐹𝑋 (𝑘) (𝑥 (𝑘) )
~ ~ ~ ~ ~ ~ ~
𝑇 𝑇 𝑇
Any subset 𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑞 ) ; (𝑞 < 𝑝) of 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) is independent of 𝑋 (2) = (𝑋𝑞+1 , 𝑋𝑞+2 , ⋯ , 𝑋𝑝 ) iff
~ ~ ~
MUTUALLY INDEPENDENT:
stochastically independent if and only if the joint cumulative distribution function (c.d.f.) 𝐹𝑋 (𝑥 ) of the random vector
~ ~
𝑇
𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) can be expressed as
~
𝑝
𝑇
𝐹𝑋 (𝑥 ) = 𝐹(𝑋1,𝑋2,⋯,𝑋𝑝) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) = ∏ 𝐹𝑋𝑖 (𝑥𝑖 ) ; ∀ 𝑥 = (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ∈ ℛ𝑝 ,
~ ~ ~
𝑖=1
where 𝐹𝑋𝑖 (𝑥𝑖 ) be the marginal cumulative distribution function (c.d.f.) of 𝑋𝑖 ; 𝑖 = 1(1)𝑝.
completely or stochastically independent if and only if the joint probability mass function (p.m.f.) 𝑝𝑋 (𝑥 ) of the random
~ ~
𝑇
vector 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) can be expressed as
~
𝑝
𝑇
𝑝𝑋 (𝑥 ) = 𝑝(𝑋1,𝑋2,⋯,𝑋𝑝) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) = ∏ 𝑝𝑋𝑖 (𝑥𝑖 ) ; ∀ 𝑥 = (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ∈ ℛ𝑝 ,
~ ~ ~
𝑖=1
where 𝑝𝑋𝑖 (𝑥𝑖 ) be the marginal probability mass function (p.m.f.) of 𝑋𝑖 ; 𝑖 = 1(1)𝑝.
A collection of jointly distributed continuous random variables (RVs) 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 is said to be mutually or completely
or stochastically independent if and only if the joint probability density function (p.d.f.) 𝑓𝑋 (𝑥 ) of the random vector 𝑋 =
~ ~ ~
𝑇
(𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) can be expressed as
𝑝
𝑇
𝑓𝑋 (𝑥 ) = 𝑓(𝑋1,𝑋2,⋯,𝑋𝑝) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) = ∏ 𝑓𝑋𝑖 (𝑥𝑖 ) ; ∀ 𝑥 = (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ∈ ℛ𝑝 ,
~ ~ ~
𝑖=1
where 𝑓𝑋𝑖 (𝑥𝑖 ) be the marginal probability density function (p.d.f.) of 𝑋𝑖 ; 𝑖 = 1(1)𝑝.
PAIRWISE INDEPENDENT:
A. A collection of jointly distributed random variables (RVs) 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 is said to be pairwise independent if and
only if every pair of them are independent i.e.,
where 𝐹(𝑋𝑖,𝑋𝑗 ) (𝑥𝑖 , 𝑥𝑗 ) be the joint cumulative distribution function (c.d.f.) of the random vector (𝑋𝑖 , 𝑋𝑗 ) and 𝐹𝑋𝑖 (𝑥𝑖 ) be
B. A collection of jointly distributed discrete random variables (RVs) 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 is said to be pairwise independent
if and only if every pair of them are independent i.e.,
where 𝑝(𝑋𝑖,𝑋𝑗 ) (𝑥𝑖 , 𝑥𝑗 ) be the joint probability mass function (p.m.f.) of the random vector (𝑋𝑖 , 𝑋𝑗 ) and 𝑝𝑋𝑖 (𝑥𝑖 ) be the
A collection of jointly distributed continuous random variables (RVs) 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 is said to be pairwise independent if
and only if every pair of them are independent i.e.,
where 𝑓(𝑋𝑖 ,𝑋𝑗 ) (𝑥𝑖 , 𝑥𝑗 ) be the joint probability density function (p.d.f.) of the random vector (𝑋𝑖 , 𝑋𝑗 ) and 𝑓𝑋𝑖 (𝑥𝑖 ) be the
NOTE:
1. A collection of jointly distributed discrete random variables (RVs) 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 is said to be pairwise independent
if and only if every pair of them are independent i.e., 𝐹(𝑋 |𝑋 ) (𝑥𝑖 |𝑥𝑗 ) = 𝐹𝑋𝑖 (𝑥𝑖 ) ; ∀ 𝑖 ≠ 𝑗, 𝑥𝑖 , 𝑥𝑗 ∈ ℛ, 𝑖, 𝑗 = 1(1)𝑝,
𝑖 𝑗
where 𝐹(𝑋 |𝑋 ) (𝑥𝑖 |𝑥𝑗 ) be the conditional cumulative distribution function (c.d.f.) of (𝑋𝑖 |𝑋𝑗 ) and 𝐹𝑋𝑖 (𝑥𝑖 ) be the marginal
𝑖 𝑗
2. If 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 are (complete or mutually) independent, every subset {(𝑋𝑖1 , 𝑋𝑖2 , ⋯ , 𝑋𝑖𝑘 ); 1 ≤ 𝑖1 < ⋯ < 𝑖𝑘 ≤ 𝑝}
of (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) is also independent.
3
; 𝑖𝑓 (𝑥1 , 𝑥2 , 𝑥3 ) ∈ {(0, 0, 0), (0, 1, 1), (1, 0, 1), (1, 1, 0)}
𝑃(𝑋1 = 𝑥1 , 𝑋2 = 𝑥2 , 𝑋3 = 𝑥3 ) = { 16
1
; 𝑖𝑓 (𝑥1 , 𝑥2 , 𝑥3 ) ∈ {(0, 0, 1), (0, 1, 0), (1, 0, 0), (1, 1, 1)}
16
Now, we have
1
𝑃(𝑋𝑖 = 𝑥𝑖 , 𝑋𝑗 = 𝑥𝑗 ) = ; 𝑖𝑓 (𝑥𝑖 , 𝑥𝑗 ) ∈ {(0, 0), (0, 1), (1, 0), (1, 1)}; 𝑖 ≠ 𝑗, 𝑖, 𝑗 = 1, 2, 3
4
1
𝑃(𝑋𝑖 = 𝑥𝑖 ) = ; 𝑥𝑖 = 0, 1; 𝑖 = 1, 2, 3
2
The random vectors (𝑋1 , 𝑋2 ), (𝑋1 , 𝑋3 ), and (𝑋2 , 𝑋3 ) are pairwise independent.
Since 𝑃(𝑋1 = 𝑥1 , 𝑋2 = 𝑥2 , 𝑋3 = 𝑥3 ) ≠ 𝑃(𝑋1 = 𝑥1 )𝑃(𝑋2 = 𝑥2 )𝑃(𝑋3 = 𝑥3 ), the random variables 𝑋1 , 𝑋2 , 𝑋3 are not
mutually independent.
THEOREM: Let 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 be independent random variables and 𝑔𝑖 (𝑋𝑖 ) be function of 𝑋𝑖 such that 𝑔𝑖 (𝑋𝑖 ) be a
random variable, 𝑖 = 1(1)𝑝. Then 𝑔1 (𝑋1 ), 𝑔2 (𝑋2 ), ⋯ , 𝑔𝑝 (𝑋𝑝 ) are independent random variables.
MOMENTS:
𝑇 𝑇
Let us define, Mean vector of 𝑋 = 𝐸 (𝑋) = (𝐸(𝑋1 ), 𝐸(𝑋2 ), ⋯ , 𝐸(𝑋𝑝 )) = (𝜇1 , 𝜇2 , ⋯ , 𝜇𝑝 ) = 𝜇 (say)
~ ~ ~
∑ 𝑥 ∙ 𝑝𝑋 (𝑥 ) = ∑ ∑ ⋯ ∑ 𝑥 ∙ 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ; if 𝑋 is discrete
~ ~ ~ ~ ~ ~
𝑥∈ℛ𝑝 𝑥1 ∈ℛ 𝑥2 ∈ℛ 𝑥𝑝 ∈ℛ
~
= ∞ ∞ ∞
𝑇
Variance-Covariance matrix of 𝑋 = Dispersion matrix of 𝑋 = 𝐷 (𝑋) = 𝐸 {(𝑋 − 𝜇 ) (𝑋 − 𝜇 ) } = Σ (say)
~ ~ ~ ~ ~ ~ ~
𝑇
∑ {(𝑥 − 𝜇 ) (𝑥 − 𝜇 ) } ∙ 𝑝𝑋 (𝑥 ) ; if 𝑋 is discrete
~ ~ ~ ~ ~ ~ ~
𝑥∈ℛ𝑝
~
=
𝑇
∫ {(𝑥 − 𝜇 ) (𝑥 − 𝜇 ) } ∙ 𝑓𝑋 (𝑥 ) 𝑑𝑥 ; if 𝑋 is continuous.
~ ~ ~ ~ ~ ~ ~ ~
{𝑥∈ℛ𝑝
~
𝐶𝑜𝑣(𝑋𝑖 , 𝑋𝑗 )
𝐶𝑜𝑟𝑟(𝑋𝑖 , 𝑋𝑗 ) = if 𝑖 ≠ 𝑗
𝜌𝑖𝑗 = √𝑉𝑎𝑟(𝑋𝑖 ) ∙ √𝑉𝑎𝑟(𝑋𝑗 ) and 𝜎𝑖𝑗 = 𝜌𝑖𝑗 𝜎𝑖 𝜎𝑗 ; 𝑖, 𝑗 = 1(1)𝑝.
{𝐶𝑜𝑟𝑟(𝑋𝑖 , 𝑋𝑗 ) = 1 if 𝑖 = 𝑗,
The variance-covariance matrix Σ and the correlation matrix 𝑅 of 𝑋 are always symmetric.
~
𝑇
Let 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) be random vector defined on a probability space (Ω, ℱ, 𝑃), and let 𝑔(∙) be a Borel-measurable
~
𝐸 {𝑔 (𝑋)} = 𝐸{𝑔(𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 )}
~
∑ 𝑔 (𝑥 ) ∙ 𝑝𝑋 (𝑥 ) = ∑ ∑ ⋯ ∑ 𝑔 (𝑥 ) ∙ 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ; if 𝑋 is discrete
~ ~ ~ ~ ~ ~
𝑥∈ℛ𝑝 𝑥1 ∈ℛ 𝑥2 ∈ℛ 𝑥𝑝 ∈ℛ
~
= ∞ ∞ ∞
∑ 𝑔𝑘 (𝑥 ) ∙ 𝑝𝑋 (𝑥 ) ; if 𝑋 is discrete
~ ~ ~ ~
𝑥∈ℛ𝑝
~
𝐸 {𝑔𝑘 (𝑋)} =
~
∫ 𝑔𝑘 (𝑥 ) ∙ 𝑓𝑋 (𝑥 ) 𝑑𝑥 ; if 𝑋 is continuous, provided 𝐸 |𝑔𝑘 (𝑋 )| < ∞. .
~ ~ ~ ~ ~ ~
{~𝑥∈ℛ𝑝
A. The joint moments about origin of order (ℎ1 + ℎ2 + ⋯ + ℎ𝑝 ) for non-negative integers (ℎ1 , ℎ2 , ⋯ , ℎ𝑝 ) is defined
as
′ ℎ ℎ ℎ
𝜇(ℎ1 ,ℎ2 ,⋯,ℎ𝑝 )
= 𝐸 (𝑋1 1 𝑋2 2 ⋯ 𝑋𝑝 𝑝 )
ℎ ℎ ℎ
∑ ∑ ⋯ ∑ 𝑥1 1 𝑥2 2 ⋯ 𝑥𝑝 𝑝 ∙ 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ; if 𝑋 is discrete
~ ~
𝑥1 ∈ℛ 𝑥2 ∈ℛ 𝑥𝑝 ∈ℛ
= ∞ ∞ ∞
ℎ ℎ ℎ
∫ ∫ ⋯ ∫ 𝑥1 1 𝑥2 2 ⋯ 𝑥𝑝 𝑝 ∙ 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) 𝑑𝑥1 𝑑𝑥2 ⋯ 𝑑𝑥𝑝 ; if 𝑋 is continuous.
~ ~
{ −∞ −∞ −∞
ℎ ℎ ℎ
provided 𝐸 |𝑋1 1 𝑋2 2 ⋯ 𝑋𝑝 𝑝 | < ∞.
′ ′
𝜇(1,1,0,⋯,0) = 𝐸(𝑋1 𝑋2 ) = 𝐶𝑜𝑣(𝑋1 , 𝑋2 ) + 𝐸(𝑋1 )𝐸(𝑋2 ) = 𝜎12 + 𝜇1 𝜇2 ; 𝜇(0,⋯,1,⋯,1,⋯,0) = 𝐸(𝑋𝑖 𝑋𝑗 ) = 𝜎𝑖𝑗 + 𝜇𝑖 𝜇𝑗 ; 𝑖 ≠ 𝑗
′ ′
𝜇(2,0,⋯,0) = 𝐸(𝑋12 ) = 𝑉𝑎𝑟(𝑋1 ) + 𝐸 2 (𝑋1 ) = 𝜎11 + 𝜇12 ; 𝜇(0,0,⋯,2,⋯,0) = 𝐸(𝑋𝑖2 ) = 𝑉𝑎𝑟(𝑋𝑖 ) + 𝐸 2 (𝑋𝑖 ) = 𝜎𝑖𝑖 + 𝜇𝑖2
CENTRAL MOMENTS:
B. MARGINAL MOMENTS:
𝑇 𝑇
The joint moments of 𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑞 ) ; (𝑞 < 𝑝), a subset of random vector 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) , of order
~ ~
(ℎ1 + ℎ2 + ⋯ + ℎ𝑞 ) for non-negative integers (ℎ1 , ℎ2 , ⋯ , ℎ𝑞 ) is defined as (computed from the marginal distribution)
′ ℎ ℎ ℎ ℎ ℎ ℎ ℎ ℎ ℎ
𝜇(ℎ1 ,ℎ2 ,⋯,ℎ𝑞 )
= 𝐸 (𝑋1 1 𝑋2 2 ⋯ 𝑋𝑞 𝑞 ) = 𝐸 (𝑋1 1 𝑋2 2 ⋯ 𝑋𝑞 𝑞 𝑋𝑞+1
0 0
𝑋𝑞+2 ⋯ 𝑋𝑝0 ) ; provided 𝐸 |𝑋1 1 𝑋2 2 ⋯ 𝑋𝑞 𝑞 | < ∞.
ℎ ℎ ℎ
∑ ∑ ⋯ ∑ 𝑥1 1 𝑥2 2 ⋯ 𝑥𝑞 𝑞 ∙ 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ; if 𝑋 is discrete
~ ~
𝑥1 ∈ℛ 𝑥2 ∈ℛ 𝑥𝑝 ∈ℛ
= ∞ ∞ ∞
ℎ ℎ ℎ
∫ ∫ ⋯ ∫ 𝑥1 1 𝑥2 2 ⋯ 𝑥𝑞 𝑞 ∙ 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) 𝑑𝑥1 𝑑𝑥2 ⋯ 𝑑𝑥𝑝 ; if 𝑋 is continuous.
~ ~
{ −∞ −∞ −∞
ℎ ℎ ℎ
∑ ∑ ⋯ ∑ 𝑥1 1 𝑥2 2 ⋯ 𝑥𝑞 𝑞 ∙ 𝑝𝑋 (1) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑞 ) ; if 𝑋 (1) is discrete
~ ~
𝑥1 ∈ℛ 𝑥2 ∈ℛ 𝑥𝑞 ∈ℛ
= ∞ ∞ ∞
ℎ ℎ ℎ
∫ ∫ ⋯ ∫ 𝑥1 1 𝑥2 2 ⋯ 𝑥𝑞 𝑞 ∙ 𝑓𝑋 (1) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑞 ) 𝑑𝑥1 𝑑𝑥2 ⋯ 𝑑𝑥𝑞 ; if 𝑋 (1) is continuous.
~ ~
{−∞ −∞ −∞
C. CONDITIONAL MOMENTS:
𝑇
Let 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) be random vector defined on a probability space (Ω, ℱ, 𝑃), and let 𝑔(∙) be a Borel-measurable
~
function. Then the conditional expectation of 𝑔 (𝑋 (1) ) = 𝑔(𝑋1 , 𝑋2 , ⋯ , 𝑋𝑞 ) given 𝑋 (2) = 𝑥 (2) ((𝑋𝑞+1 , 𝑋𝑞+2 , ⋯ , 𝑋𝑝 ) =
~ ~ ~
∑ ∑ ⋯ ∑ 𝑔(𝑥1 , 𝑥2 , ⋯ , 𝑥𝑞 ) ∙ 𝑝 (1) (2) (𝑥 (1) |𝑥 (2) ) ; if 𝑋 is discrete and 𝑝𝑋 (2) (𝑥 (2) ) > 0
(𝑋 |𝑋 ) ~ ~ ~ ~ ~
𝑥1 ∈ℛ 𝑥2 ∈ℛ 𝑥𝑞 ∈ℛ ~ ~
∞ ∞ ∞
∫ ∫ ⋯ ∫ 𝑔(𝑥1 , 𝑥2 , ⋯ , 𝑥𝑞 ) ∙ 𝑓 (1) (2) (𝑥 (1) |𝑥 (2) ) 𝑑𝑥1 𝑑𝑥2 ⋯ 𝑑𝑥𝑞 ; if 𝑋 is continuous and 𝑓𝑋 (2) (𝑥 (2) ) > 0.
(𝑋 |𝑋 ) ~ ~ ~ ~ ~
{−∞ −∞ −∞ ~ ~
𝑇
If the random vector 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) are mutually independent, the joint moments about origin of order
~
′ ℎ ℎ ℎ
𝜇(ℎ1 ,ℎ2 ,⋯,ℎ𝑝 )
= 𝐸 (𝑋1 1 𝑋2 2 ⋯ 𝑋𝑝 𝑝 )
𝑝
ℎ ℎ ℎ ℎ
∑ ∑⋯ ∑ 𝑥1 1 𝑥2 2 ⋯ 𝑥𝑝 𝑝 ∙ 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) = ∏ { ∑ 𝑥𝑖 𝑖 ∙ 𝑝𝑋𝑖 (𝑥𝑖 )} ; if 𝑋 is discrete
~ ~
𝑥1 ∈ℛ 𝑥2 ∈ℛ 𝑥𝑝 ∈ℛ 𝑖=1 𝑥𝑖 ∈ℛ
= ∞ ∞ ∞ 𝑝 ∞
ℎ ℎ ℎ ℎ
∫ ∫⋯ ∫ 𝑥1 1 𝑥2 2 ⋯ 𝑥𝑝 𝑝 ∙ 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) 𝑑𝑥1 𝑑𝑥2 ⋯ 𝑑𝑥𝑝 = ∏ { ∫ 𝑥𝑖 𝑖 ∙ 𝑓𝑋𝑖 (𝑥𝑖 ) 𝑑𝑥𝑖 } ; if 𝑋 is continuous.
~ ~
{−∞ −∞ −∞ 𝑖=1 −∞
𝑝
ℎ ℎ
= ∏ 𝐸(𝑋𝑖 𝑖 ) ; provided 𝐸|𝑋𝑖 𝑖 | < ∞, ∀ 𝑖 = 1(1)𝑝.
𝑖=1
CHARACTERISTIC FUNCTION:
𝑇 𝑇
Let 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) be random vector defined on a probability space (Ω, ℱ, 𝑃) and 𝑡 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 ) ∈ ℛ𝑝 . Then
~ ~
characteristic function of 𝑋 or the joint c.f. of the random variables 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 is denoted by 𝜙𝑋 (𝑡 ) and defined as
~ ~ ~
𝑝
𝑖(𝑡 𝑇 𝑋) 𝑖𝑡1 𝑋1 +𝑖𝑡2 𝑋2 + ⋯+𝑖𝑡𝑝 𝑋𝑝
𝜙𝑋 (𝑡 ) = 𝜙𝑋 (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 ) = 𝐸 {𝑒 ~ ~ } = 𝐸{𝑒 } = 𝐸 {exp (𝑖 ∑ 𝑡𝑗 𝑋𝑗 )} ; 𝑖 = √−1
~ ~ ~
𝑗=1
𝑝 𝑝
= 𝐸 {cos (∑ 𝑡𝑗 𝑋𝑗 )} + 𝑖𝐸 {sin (∑ 𝑡𝑗 𝑋𝑗 )}
𝑗=1 𝑗=1
𝑝
𝑖(𝑡 𝑇 𝑋 ) 𝑖 ∑𝑗=1 𝑡𝑗𝑋𝑗
∑ 𝑒 ~ ~ ∙ 𝑝𝑋 (𝑥 ) = ∑ ∑ ⋯ ∑ 𝑒 ∙ 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ; if 𝑋 is discrete
~ ~ ~ ~
𝑥∈ℛ𝑝 𝑥1 ∈ℛ 𝑥2 ∈ℛ 𝑥𝑝 ∈ℛ
~
= ∞ ∞ ∞
𝑝
𝑖(𝑡 𝑇 𝑋 ) 𝑖 ∑𝑗=1 𝑡𝑗𝑋𝑗
∫ 𝑒 ~ ~ ∙ 𝑓𝑋 (𝑥 ) 𝑑𝑥 = ∫ ∫ ⋯ ∫ 𝑒 ∙ 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) 𝑑𝑥1 𝑑𝑥2 ⋯ 𝑑𝑥𝑝 ; if 𝑋 is continuous.
~ ~ ~ ~ ~
{~𝑥∈ℛ𝑝 −∞ −∞ −∞
𝑇
The characteristic function of 𝑋 i.e., 𝜙𝑋 (𝑡 ) always exists for all 𝑡 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 ) ∈ ℛ𝑝 .
~ ~ ~ ~
𝑝 𝑝
𝑖 ∑𝑗=1 𝑡𝑗 𝑏𝑗 𝑖 ∑𝑗=1 𝑡𝑗𝑎𝑗 𝑋𝑗 𝑖(𝑏 𝑇 𝑡 )
(𝑣𝑖) 𝜙𝑎𝑇 𝑋+𝑏 (𝑡 ) = 𝜙𝑎1 𝑋1 +𝑏1,⋯,𝑎𝑝 𝑋𝑝+𝑏𝑝 (𝑡 ) = 𝑒 ∙ 𝐸 {𝑒 }=𝑒 ~ ~ ∙ 𝜙𝑋 (𝑎1 𝑡1 , 𝑎2 𝑡2 , ⋯ , 𝑎𝑝 𝑡𝑝 )
~ ~ ~ ~ ~ ~
(𝑣𝑖𝑖) The random variables 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 are (mutually) independent iff 𝜙𝑋 (𝑡 ) = 𝜙𝑋1 (𝑡1 ) ∙ 𝜙𝑋2 (𝑡2 ) ⋯ 𝜙𝑋𝑝 (𝑡𝑝 ).
~ ~
ℎ
(𝑣𝑖𝑖𝑖) If 𝐸 |𝑋1ℎ1 𝑋2ℎ2 ⋯ 𝑋𝑝 𝑝 | < ∞, the joint moments about origin 𝜇(ℎ
′
of order (ℎ1 + ℎ2 + ⋯ + ℎ𝑝 ) for non-
1 ,ℎ2 ,⋯,ℎ𝑝 )
𝜕ℎ
In particular, 𝜙𝑋 (𝑡 )| = 𝑖 ℎ 𝐸(𝑋𝑗ℎ ); 𝑗 = 1, 2, ⋯ , 𝑝.
𝜕𝑡𝑗ℎ ~ ~
𝑡 =0
~ ~
Let 𝑔(∙) be a Borel-measurable function. Then characteristic function of 𝑔 (𝑋) = 𝑔(𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) is defined as
~
𝑖𝑡𝑔(𝑋)
𝜙𝑔(𝑋) (𝑡) = 𝐸 {𝑒 ~ } = 𝐸 {exp (𝑖𝑡𝑔 (𝑋))}.
~ ~
𝑇 𝑇
Let 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) be random vector defined on a probability space (Ω, ℱ, 𝑃) and 𝑡 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 ) ∈ ℛ𝑝 . Then
~ ~
the moment generating function (𝑚. 𝑔. 𝑓. ) of 𝑋 or the 𝑚. 𝑔. 𝑓. of the joint distribution of (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) is defined as
~
𝑝
𝑡 𝑇𝑋 𝑡1 𝑋1 +𝑡2 𝑋2 + ⋯+𝑡𝑝 𝑋𝑝 }
𝑀𝑋 (𝑡 ) = 𝑀𝑋 (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 ) = 𝐸 (𝑒 ~ ~ ) = 𝐸{𝑒 = 𝐸 {exp (∑ 𝑡𝑗 𝑋𝑗 )}
~ ~ ~
𝑗=1
𝑝
𝑡 𝑇𝑋 ∑𝑗=1 𝑡𝑗 𝑋𝑗
∑ 𝑒~ ~ ∙ 𝑝𝑋 (𝑥 ) = ∑ ∑ ⋯ ∑ 𝑒 ∙ 𝑝𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ; if 𝑋 is discrete
~ ~ ~ ~
𝑥∈ℛ𝑝 𝑥1 ∈ℛ 𝑥2 ∈ℛ 𝑥𝑝 ∈ℛ
~
= ∞ ∞ ∞
𝑝
𝑡 𝑇𝑋 ∑𝑗=1 𝑡𝑗 𝑋𝑗
∫ 𝑒~ ~ ∙ 𝑓𝑋 (𝑥 ) 𝑑𝑥 = ∫ ∫ ⋯ ∫ 𝑒 ∙ 𝑓𝑋 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) 𝑑𝑥1 𝑑𝑥2 ⋯ 𝑑𝑥𝑝 ; if 𝑋 is continuous.
~ ~ ~ ~ ~
{~𝑥∈ℛ𝑝 −∞ −∞ −∞
𝑝
𝑡 𝑇𝑋 ∑𝑗=1 𝑡𝑗𝑋𝑗
provided 𝐸 |𝑒 ~ ~ | = 𝐸 |𝑒 | < ∞, for |𝑡𝑗 | ≤ ℎ𝑗 ; 𝑗 = 1, 2, ⋯ , 𝑝, for some ℎ𝑗 > 0 ; 𝑗 = 1, 2, ⋯ , 𝑝.
(𝑖𝑖) If 𝑀𝑋 (𝑡 ) = 𝑀𝑌 (𝑡 ) for all 𝑡 ∈ ℛ𝑝 , then 𝑋 and 𝑌 have the same distribution. The m.g.f. 𝑀𝑋 (𝑡 ) uniquely
~ ~ ~ ~ ~ ~ ~ ~ ~
𝑇
determines the joint distribution of 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) , and conversely, if the m.g.f. exists, it is unique.
~
𝑡 𝑇 (𝑎 𝑇 𝑋+𝑏 )
(𝑖𝑖𝑖) 𝑀𝑎𝑇𝑋+𝑏 (𝑡 ) = 𝑀𝑎1 𝑋1+𝑏1,⋯,𝑎𝑝𝑋𝑝 +𝑏𝑝 (𝑡 ) = 𝐸 {𝑒 ~ ~ ~ ~ } = 𝐸{𝑒 𝑡1(𝑎1𝑋1+𝑏1)+ ⋯+𝑡𝑝(𝑎𝑝𝑋𝑝 +𝑏𝑝) }
~ ~ ~ ~ ~
𝑝
∑𝑗=1 𝑡𝑗𝑏𝑗 𝑏𝑇 𝑡
=𝑒 ∙ 𝐸{𝑒 𝑡1𝑎1𝑋1+ ⋯+𝑡𝑝𝑎𝑝 𝑋𝑝 } = 𝑒 ~ ~ ∙ 𝑀𝑋 (𝑎1 𝑡1 , 𝑎2 𝑡2 , ⋯ , 𝑎𝑝 𝑡𝑝 )
~
ℎ ℎ ℎ ′
(𝑣) If 𝐸 |𝑋1 1 𝑋2 2 ⋯ 𝑋𝑝 𝑝 | < ∞, the joint moments about origin 𝜇(ℎ1 ,ℎ2 ,⋯,ℎ𝑝 )
of order (ℎ1 + ℎ2 + ⋯ + ℎ𝑝 ) for non-
′ ℎ ℎ ℎ 𝜕 (ℎ1+ℎ2 +⋯+ℎ𝑝 )
𝜇(ℎ1 ,ℎ2 ,⋯,ℎ𝑝 )
= 𝐸 (𝑋1 1 𝑋2 2 ⋯ 𝑋𝑝 𝑝 ) = ℎ ℎ ℎ
𝑀𝑋 (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 )|
𝜕𝑡1 1 𝜕𝑡2 2 ⋯ 𝜕𝑡𝑝 𝑝 ~
𝑡1 =𝑡2 =⋯=𝑡𝑝 =0
𝜕ℎ
In particular, 𝑀𝑋 (𝑡 )| = 𝐸(𝑋𝑗ℎ ); 𝑗 = 1, 2, ⋯ , 𝑝.
𝜕𝑡𝑗ℎ ~ ~
𝑡 =0
~ ~
𝑇
(𝑣𝑖) The moment generating function of 𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑞 ) ; (𝑞 < 𝑝), a subset of random vector 𝑋 =
~ ~
𝑇 𝑇
(𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) is given by 𝑀𝑋 (1) (𝑡 (1) ) = 𝑀𝑋 (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑞 , 0, 0, ⋯ , 0 ) ; 𝑡 (1) = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑞 ) ∈ ℛ𝑞 .
~ ~ ~ ~
Let 𝑔(∙) be Borel-measurable functions. Then moment generating function of 𝑔 (𝑋) = 𝑔(𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) is defined as
~ ~ ~ ~
𝑡 𝑇 𝑔(𝑋 )
𝑀𝑔(𝑋) (𝑡 ) = 𝐸 {𝑒 ~ ~ ~ } = 𝐸 {exp (𝑡 𝑇 𝑔 (𝑋))}.
~ ~ ~ ~
~ ~
Using the uniqueness property of moment generating function we may find the distribution of 𝑌 = 𝑔(𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ). We
~ ~
compute the m.g.f. of 𝑌 = 𝑔(𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ). If this m.g.f. is one of the known kind, 𝑌 must have this kind of distribution.
~ ~ ~
MULTINOMIAL DISTRIBUTION
INTRODUCTION
Recall the Binomial Distribution. It models the number of "successes" in a fixed number 𝑛 of independent
Bernoulli trials, where each trial has only two possible outcomes (e.g., Success/Failure, Head/Tail, Yes/No)
with constant probabilities 𝑝 and (1 − 𝑝).
But what if each trial has more than two possible outcomes?
o Rolling a die: 6 outcomes {1, 2, 3, 4, 5, 6}
o Survey question with responses: {Strongly Agree, Agree, Neutral, Disagree, Strongly Disagree}
o Drawing balls from an urn with replacement: {Red, Green, Blue}
o Classifying a product from a production line: {Perfect, Minor Flaw, Major Flaw, Unusable}
The Multinomial Distribution is designed for precisely this scenario. It describes the probabilities of the counts
for each category across a fixed number of independent trials.
Let the random variables 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘 denote the number of times each outcome 𝐴1 , 𝐴2 , ⋯ , 𝐴𝑘 occurs in the
𝑛 trials.
Note that these counts must sum to the total number of trials: ∑𝑘𝑖=1 𝑋𝑖 = 𝑛.
The vector of counts 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘 )𝑇 follows a Multinomial Distribution with parameters 𝑛 and 𝑝 =
~ ~
)𝑇
(𝑝1 , 𝑝2 , ⋯ , 𝑝𝑘 . We denote this as: 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘 )𝑇 ~ 𝑀𝑢𝑙𝑡𝑖𝑛𝑜𝑚𝑖𝑎𝑙(𝑛, 𝑝1 , 𝑝2 , ⋯ , 𝑝𝑘 ).
~
The probability of observing exactly 𝑥1 outcomes of type 𝐴1 , 𝑥2 outcomes of type 𝐴2 , ⋯, and 𝑥𝑘 outcomes of
type 𝐴𝑘 , where ∑𝑘𝑖=1 𝑥𝑖 = 𝑛 and each 𝑥𝑖 ≥ 0 (= 0(1)𝑛) is an integer, is given by the 𝑝. 𝑚. 𝑓.:
𝑘
𝑛! 𝑥 𝑥 𝑥
∙ 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑘 𝑘 ; 𝑥𝑖 = 0(1)𝑛, ∑ 𝑥𝑖 = 𝑛,
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑘 !
𝑖=1
𝑝𝑋 (𝑥 ) = 𝑝(𝑋1 ,𝑋2 ,⋯,𝑋𝑘 ) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑘 ) = 𝑘
~ ~
0 < 𝑝𝑖 < 1, ∑ 𝑝𝑖 = 1, 𝑖 = 1(1)𝑘
𝑖=1
{ 0; otherwise
Consider one specific sequence of 𝑛 trial outcomes that results in 𝑥1 𝐴1′ 𝑠, 𝑥2 𝐴′2 𝑠, ⋯ , 𝑥𝑘 𝐴𝑘 ′𝑠.
For example, if 𝑛 = 5, 𝑘 = 3, 𝑥1 = 2, 𝑥2 = 1, 𝑥3 = 2, one such sequence could be (𝐴₁, 𝐴₃, 𝐴₁, 𝐴₂, 𝐴₃). Since
the trials are independent, the probability of this specific sequence is the product of the individual trial
probabilities: 𝑝1 ∙ 𝑝3 ∙ 𝑝1 ∙ 𝑝2 ∙ 𝑝3 = 𝑝12 ∙ 𝑝21 ∙ 𝑝32 .
𝑥 𝑥 𝑥
In general, the probability of any one specific sequence with the desired counts 𝑥1 , 𝑥2 , ⋯ , 𝑥𝑘 is: 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑘 𝑘 .
How many different sequences of trial outcomes yield exactly 𝑥1 𝐴1′ 𝑠, 𝑥2 𝐴′2 𝑠, ⋯ , 𝑥𝑘 𝐴𝑘 ′𝑠? This is a
combinatorial problem of arranging 𝑛 items where there are 𝑥1 identical items of type 1, 𝑥2 identical items of
type 2, ... , 𝑥𝑘 identical items of type 𝑘. The number of distinct permutations is given by the multinomial
𝑛!
coefficient:
𝑥1 !𝑥2 !⋯𝑥𝑘 !
The total probability 𝑃(𝑋1 = 𝑥1 , 𝑋2 = 𝑥2 , ⋯ , 𝑋𝑘 = 𝑥𝑘 ) is the sum of the probabilities of all sequences that
𝑥 𝑥 𝑥
result in these counts. Since each such sequence has the same probability (𝑝1 1 𝑝2 2 ⋯ 𝑝𝑘 𝑘 ) and there are
𝑛!
𝑥1 !𝑥2 !⋯𝑥𝑘 !
such sequences, the total probability is:
𝑛! 𝑥 𝑥 𝑥
(Number of sequences) × (Probability of one sequence) = ∙ 𝑝 1 𝑝 2 ⋯ 𝑝𝑘 𝑘
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑘 ! 1 2
DEFINITION
Suppose a random experiment is repeated 𝑛 times. Each replication of the experiment terminates in one of the
𝑘 mutually exclusive and exhaustive events 𝐴1 , 𝐴2 , ⋯ , 𝐴𝑘 . Let 𝑝𝑖 be the probability that the experiment
terminates in 𝐴𝑖 ; 𝑖 = 1(1)𝑘 and 𝑝𝑖 ; 𝑖 = 1(1)𝑘 remains constant for 𝑛 repetitions. We are assume that the 𝑛
replications are independent. Let 𝑋𝑖 ; 𝑖 = 1(1)𝑘 be the number of times 𝐴𝑖 occurs in 𝑛 replications.
𝑘
𝑛! 𝑥 𝑥 𝑥
∙ 𝑝 1 𝑝 2 ⋯ 𝑝𝑘 𝑘 ; 𝑥𝑖 = 0(1)𝑛, ∑ 𝑥𝑖 = 𝑛,
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑘 ! 1 2
𝑖=1
𝑝𝑋 (𝑥 ) = 𝑝(𝑋1 ,𝑋2 ,⋯,𝑋𝑘 ) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑘 ) = 𝑘
~ ~
0 < 𝑝𝑖 < 1, ∑ 𝑝𝑖 = 1, 𝑖 = 1(1)𝑘
𝑖=1
{ 0; otherwise
Then the random vector 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘 )𝑇 is said to follow the multinomial distribution (singular) with
~
parameters (𝑛, 𝑝1 , 𝑝2 , ⋯ , 𝑝𝑘 ).
NOTATION:
𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘 )𝑇 ~ 𝑀𝑢𝑙𝑡𝑖𝑛𝑜𝑚𝑖𝑎𝑙(𝑛, 𝑝1 , 𝑝2 , ⋯ , 𝑝𝑘 ),
~
𝑘 𝑘
EXAMPLE:
If a fair die is thrown n times and 𝑋𝑖 ; 𝑖 = 1(1)6 is the number of times the 𝑖 th face appears. Then
1 1 1
(𝑋1 , 𝑋2 , ⋯ , 𝑋6 )𝑇 ~ 𝑀𝑢𝑙𝑡𝑖𝑛𝑜𝑚𝑖𝑎𝑙 (𝑛, , , ⋯ , ).
6 6 6
VERIFICATION OF 𝒑. 𝒎. 𝒇.
𝑝𝑋 (𝑥 ) ≥ 0; ∀ 𝑥 ∈ 𝑹𝑘 and
~ ~ ~
𝑛 𝑛 𝑛 𝑛 𝑛 𝑛
𝑛! 𝑥 𝑥 𝑥
∑ ∑ ⋯ ∑ 𝑝𝑋 (𝑥 ) = ∑ ∑ ⋯∑ ∙ 𝑝 1 𝑝 2 ⋯ 𝑝𝑘 𝑘
~ ~ 𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑘 ! 1 2
𝑥1 =0 𝑥2 =0 𝑥𝑘 =0 𝑥1 =0 𝑥2 =0 𝑥𝑘 =0
∑𝑘
𝑖=1 𝑥𝑖 =𝑛 ∑𝑘
𝑖=1 𝑥𝑖 =𝑛
𝑛
𝑘 𝑘
= (𝑝1 + 𝑝2 + ⋯ + 𝑝𝑘 )𝑛 = (∑ 𝑝𝑖 ) = 1 [∵ ∑ 𝑝𝑖 = 1]
𝑖=1 𝑖=1
The moment generating function (𝑚. 𝑔. 𝑓. ) of 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘 )𝑇 or the 𝑚. 𝑔. 𝑓. of the joint distribution of
~
𝑡 𝑇𝑋
𝑀𝑋 (𝑡 ) = 𝑀(𝑋1 ,𝑋2 ,⋯,𝑋𝑘) (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑘 ) = 𝐸 (𝑒 ~ ~ ) = 𝐸{𝑒 𝑡1 𝑋1 +𝑡2 𝑋2 +⋯+𝑡𝑘 𝑋𝑘 }
~ ~
𝑛 𝑛 𝑛
𝑛! 𝑥 𝑥 𝑥
= ∑ ∑ ⋯ ∑ 𝑒 𝑡1 𝑥1 +𝑡2 𝑥2 +⋯+𝑡𝑘 𝑥𝑘 ∙ 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑘 𝑘
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑘 !
𝑥1 =0 𝑥2 =0 𝑥𝑘 =0
∑𝑘
𝑖=1 𝑥𝑖 =𝑛
𝑛 𝑛 𝑛
𝑛!
= ∑ ∑ ⋯∑ ∙ (𝑝1 𝑒 𝑡1 )𝑥1 (𝑝2 𝑒 𝑡2 )𝑥2 ⋯ (𝑝𝑘 𝑒 𝑡𝑘 )𝑥𝑘
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑘 !
𝑥1 =0 𝑥2 =0 𝑥𝑘 =0
∑𝑘
𝑖=1 𝑥𝑖 =𝑛
𝑛
𝑘
= (𝑝1 𝑒 𝑡1 + 𝑝2 𝑒 𝑡2 + ⋯ + 𝑝𝑘 𝑒 𝑡𝑘 )𝑛 = (∑ 𝑝𝑖 𝑒 𝑡𝑖 )
𝑖=1
MOMENTS:
Let 𝐼𝛼𝑖 be an indicator random variable for trial 𝛼; ( 𝛼 = 1(1)𝑛) and category 𝑖; (𝑖 = 1(1)𝑘), i.e.
The total count for category 𝑖, 𝑋𝑖 , is the sum of the indicators for that category across all 𝑛 trials:
𝑋𝑖 = ∑ 𝐼𝛼𝑖 ; 𝑖 = 1(1)𝑘.
𝛼=1
𝑛 𝑛 𝑛
Since 𝐼𝛼𝑖 is a Bernoulli random variable (taking values 1 or 0), its variance is given by 𝑝𝑖 (1 − 𝑝𝑖 ) where 𝑝𝑖 is
the probability of success (being 1). 𝑉𝑎𝑟[𝐼𝛼𝑖 ] = 𝑝𝑖 (1 − 𝑝𝑖 ). Alternatively,
2
𝑉𝑎𝑟(𝐼𝛼𝑖 ) = 𝐸(𝐼𝛼𝑖 ) − 𝐸 2 (𝐼𝛼𝑖 ) = {12 ∙ 𝑃(𝐼𝛼𝑖 = 1) + 02 ∙ 𝑃(𝐼𝛼𝑖 = 0)} − 𝑝𝑖2 = 𝑝𝑖 − 𝑝𝑖2 = 𝑝𝑖 (1 − 𝑝𝑖 ); 𝑖 = 1(1)𝑘
Since the trials are independent, the indicator variables 𝐼𝛼𝑖 (outcome 𝑖 on trial 𝛼) and 𝐼𝛼′ 𝑖 (outcome 𝑖 on trial
𝛼 ′ ) are independent for different trials (𝛼 ≠ 𝛼 ′ ).
𝑛 𝑛 𝑛
𝑛 𝑛 𝑛 𝑛
Now, consider the covariance term 𝐶𝑜𝑣(𝐼𝛼𝑖 , 𝐼𝛼′ 𝑗 ) under two cases based on the trial indices 𝛼 and 𝛼 ′ :
Since trials are independent, the outcome of trial 𝛼 (𝐼𝛼𝑖 ) is independent of the outcome of trial 𝛼 ′ (𝐼𝛼′ 𝑗 ).
Independent random variables have zero covariance. Therefore, 𝐶𝑜𝑣(𝐼𝛼𝑖 , 𝐼𝛼′ 𝑗 ) = 0 when 𝛼 ≠ 𝛼 ′ .
We are considering the covariance between the indicator for category 𝑖 and the indicator for category 𝑗 within
the same trial 𝛼. Since 𝑖 ≠ 𝑗, these represent two different, mutually exclusive outcomes for the same trial, it's
impossible for a single trial to result in both outcomes simultaneously.
∴ 𝐸(𝐼𝛼𝑖 , 𝐼𝛼′𝑗 ) = 0; 𝛼 = 𝛼 ′ and 𝐶𝑜𝑣(𝐼𝛼𝑖 , 𝐼𝛼′ 𝑗 ) = 𝐸(𝐼𝛼𝑖 , 𝐼𝛼′ 𝑗 ) − 𝐸(𝐼𝛼𝑖 )𝐸(𝐼𝛼′ 𝑗 ) = 0 − 𝑝𝑖 𝑝𝑗 = −𝑝𝑖 𝑝𝑗 ; 𝑖 ≠ 𝑗
𝑛 𝑛 𝑛 𝑛
𝑛 𝑛
𝑛−1
𝑘
𝜕 𝜕
𝑀𝑋 (𝑡 ) = 𝑀 (𝑡 , 𝑡 , ⋯ , 𝑡𝑘 ) = 𝑛𝑝𝑖 𝑒 𝑡𝑖 (∑ 𝑝𝑖 𝑒 𝑡𝑖 ) ; 𝑖 = 1(1)𝑘
𝜕𝑡𝑖 ~ ~ 𝜕𝑡𝑖 (𝑋1 ,𝑋2 ,⋯,𝑋𝑘 ) 1 2
𝑖=1
𝜕 𝜕
𝐸(𝑋𝑖 ) = 𝑀𝑋 (𝑡 )| = 𝑀 (𝑡 , 𝑡 , ⋯ , 𝑡𝑘 )| ; 𝑖 = 1(1)𝑘
𝜕𝑡𝑖 ~ ~ 𝑡 =0 𝜕𝑡𝑖 (𝑋1 ,𝑋2 ,⋯,𝑋𝑘 ) 1 2 𝑡 1 =𝑡2 =⋯=𝑡𝑘 =0
~ ~
𝑛−1 𝑛−1
𝑘 𝑘
𝜕2 𝜕2
𝑀𝑋 (𝑡 ) = 𝑀(𝑋1 ,𝑋2 ,⋯,𝑋𝑘 ) (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑘 )
𝜕𝑡𝑖2 ~ ~ 𝜕𝑡𝑖2
𝑛−2 𝑛−1
𝑘 𝑘
𝜕2 𝜕2
𝐸(𝑋𝑖2 ) = 𝑀𝑋 (𝑡 )| = 𝑀(𝑋1 ,𝑋2 ,⋯,𝑋𝑘 ) (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑘 )| ; 𝑖 = 1(1)𝑘
𝜕𝑡𝑖2 ~ ~ 𝑡 =0 𝜕𝑡𝑖2 𝑡 1 =𝑡2 =⋯=𝑡𝑘 =0
~ ~
𝑛−2 𝑛−1
𝑘 𝑘
𝜕2 𝜕2
𝑀𝑋 (𝑡 ) = 𝑀 (𝑡 , 𝑡 , ⋯ , 𝑡𝑘 )
𝜕𝑡𝑖 𝜕𝑡𝑗 ~ ~ 𝜕𝑡𝑖 𝜕𝑡𝑗 (𝑋1 ,𝑋2 ,⋯,𝑋𝑘 ) 1 2
𝑛−2
𝑘
𝜕2 𝜕2
𝐸(𝑋𝑖 𝑋𝑗 ) = 𝑀 (𝑡 )| = 𝑀 (𝑡 , 𝑡 , ⋯ , 𝑡𝑘 )| ; 𝑖, 𝑗 = 1(1)𝑘, 𝑖 ≠ 𝑗
𝜕𝑡𝑖 𝜕𝑡𝑗 𝑋~ ~
𝑡 =0
𝜕𝑡𝑖 𝜕𝑡𝑗 (𝑋1 ,𝑋2 ,⋯,𝑋𝑘) 1 2 𝑡 1 =𝑡2 =⋯=𝑡𝑘 =0
~ ~
𝑛−2
𝑘
𝐶𝑜𝑣(𝑋𝑖 , 𝑋𝑗 ) −𝑛𝑝𝑖 𝑝𝑗
𝜌(𝑋𝑖 , 𝑋𝑗 ) = 𝐶𝑜𝑟𝑟(𝑋𝑖 , 𝑋𝑗 ) = =
√𝑉𝑎𝑟(𝑋𝑖 )√𝑉𝑎𝑟( 𝑋𝑗 ) √𝑛𝑝𝑖 (1 − 𝑝𝑖 )√𝑛𝑝𝑗 (1 − 𝑝𝑗 )
𝑝𝑖 𝑝𝑗
= −√ ; 𝑖, 𝑗 = 1(1)𝑘, 𝑖 ≠ 𝑗
(1 − 𝑝𝑖 )(1 − 𝑝𝑗 )
𝑝1 𝑝2 𝑝1 𝑝𝑘
1 −√ ⋯ −√
(1 − 𝑝1 )(1 − 𝑝2 ) (1 − 𝑝1 )(1 − 𝑝𝑘 )
𝑝2 𝑝1 𝑝2 𝑝𝑘
𝑹𝑘×𝑘 = 𝐶𝑜𝑟𝑟 (𝑋) = −√ 1 ⋯ −√
~
(1 − 𝑝2 )(1 − 𝑝1 ) (1 − 𝑝2 )(1 − 𝑝𝑘 )
⋮ ⋮ ⋱ ⋮
𝑝𝑘−1 𝑝1 𝑝𝑘 𝑝2
−√ −√ ⋯ 1
(1 − 𝑝𝑘 )(1 − 𝑝1 ) (1 − 𝑝𝑘 )(1 − 𝑝2 )
( )𝑘×𝑘
The covariance is negative because the total number of trials 𝑛 is fixed (∑𝑘𝑖=1 𝑋𝑖 = 𝑛). If the count for one
category, 𝑋𝑖 , is higher than expected, the counts for the other categories (𝑋𝑗 ; for 𝑗 ≠ 𝑖) must tend to be lower
to compensate, and vice versa. This inverse relationship is captured by the negative covariance.
MARGINAL DISTRIBUTION
𝑛
𝑘
𝑛
𝑠 𝑘
= (∑ 𝑝𝑖 𝑒 𝑡𝑖 + ∑ 𝑝𝑖 )
𝑖=1 𝑖=𝑠+1
𝑠 𝑠 𝑛
𝑡𝑖
= {∑ 𝑝𝑖 𝑒 + (1 − ∑ 𝑝𝑖 )}
𝑖=1 𝑖=1
𝑠 𝑠
← MGF of 𝐵𝑖𝑛(𝑛, 𝑝𝑖 )
𝑋𝑖 ~ 𝐵𝑖𝑛(𝑛, 𝑝𝑖 ) ; 𝑖 = 1(1)𝑘.
ALTERNATIVE
𝑛! 𝑥 𝑥 𝑥
= ∑ ∑ ⋯ ∑ ∙ 𝑝 1 𝑝 2 ⋯ 𝑝𝑘 𝑘
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑘 ! 1 2
𝑥𝑗 =0(1)𝑛, 𝑗=𝑠+1(1)𝑘, ∑𝑘
𝑗=1 𝑥𝑗 =𝑛
𝑛! 𝑥 𝑥 𝑥 (𝑛 − ∑𝑠𝑖=1 𝑥𝑖 )! 𝑥 𝑥 𝑥
= ∙ 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑠 𝑠 ∙ ∑ ∑ ⋯ ∑ ∙ 𝑝 𝑠+1 𝑝 𝑠+2 ⋯ 𝑝𝑘 𝑘
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! (𝑛 − ∑𝑠𝑖=1 𝑥𝑖 )! 𝑥𝑠+1 ! 𝑥𝑠+2 ! ⋯ 𝑥𝑘 ! 𝑠+1 𝑠+2
𝑥𝑗 =0(1)𝑛, 𝑗=𝑠+1(1)𝑘, ∑𝑘
𝑗=1 𝑥𝑗 =𝑛
𝑛! 𝑥 𝑥 𝑥 𝑠
= ∙ 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑠 𝑠 ∙ (𝑝𝑠+1 + 𝑝𝑠+2 + ⋯ + 𝑝𝑘 )(𝑛−∑𝑖=1 𝑥𝑖)
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! (𝑛 − ∑𝑠𝑖=1 𝑥𝑖 )!
𝑠 (𝑛−∑𝑠𝑖=1 𝑥𝑖 )
𝑛! 𝑥 𝑥 𝑥
= ∙ 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑠 𝑠 ∙ (1 − ∑ 𝑝𝑖 )
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! (𝑛 − ∑𝑠𝑖=1 𝑥𝑖 )!
𝑖=1
𝑠 𝑠
𝑝𝑋𝑖 (𝑥𝑖 ) = ∑ ∑ ⋯ ∑ 𝑝𝑋 (𝑥 )
~ ~
𝑥𝑗 =0(1)𝑛, 𝑥𝑗 ≠𝑥𝑖 , ∑𝑘
𝑗=1 𝑥𝑗 =𝑛
𝑛! 𝑥 𝑥 𝑥
= ∑ ∑ ⋯ ∑ ∙ 𝑝 1 𝑝 2 ⋯ 𝑝𝑘 𝑘 ; 𝑖 = 1(1)𝑘
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑘 ! 1 2
𝑥𝑗 =0(1)𝑛, 𝑥𝑗 ≠𝑥𝑖 , ∑𝑘
𝑗=1 𝑥𝑗 =𝑛
𝑛! 𝑥 (𝑛 − 𝑥𝑖 )! 𝑥 𝑥𝑖−1 𝑥𝑖+1 𝑥
= ∙ 𝑝𝑖 𝑖 ∙ ∑ ∑ ⋯ ∑ ∙ 𝑝1 1 ⋯ 𝑝𝑖−1 𝑝𝑖+1 ⋯ 𝑝𝑘 𝑘
𝑥𝑖 ! (𝑛 − 𝑥𝑖 )! 𝑥1 ! ⋯ 𝑥𝑖−1 ! 𝑥𝑖+1 ! ⋯ 𝑥𝑘 !
𝑥𝑗 =0(1)(𝑛−𝑥𝑖 ), 𝑥𝑗 ≠𝑥𝑖 , ∑𝑘
𝑗=1,𝑗≠𝑖 𝑥𝑗 =𝑛−𝑥𝑖
𝑛! 𝑥
= ∙ 𝑝 𝑖 ∙ (𝑝1 + ⋯ + 𝑝𝑖−1 + 𝑝𝑖+1 + ⋯ + 𝑝𝑘 )(𝑛−𝑥𝑖)
𝑥𝑖 ! (𝑛 − 𝑥𝑖 )! 𝑖
𝑘
𝑛! 𝑥
= ∙ 𝑝 𝑖 ∙ (1 − 𝑝𝑖 )(𝑛−𝑥𝑖) [∵ ∑ 𝑝𝑖 = 1]
𝑥𝑖 ! (𝑛 − 𝑥𝑖 )! 𝑖
𝑖=1
← 𝑝. 𝑚. 𝑓. of 𝐵𝑖𝑛(𝑛, 𝑝𝑖 ) ; 𝑖 = 1(1)𝑘
CONDITIONAL DISTRIBUTION
The conditional distribution of 𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑠 )𝑇 given that 𝑋 (2) = 𝑥 (2) , where 𝑋 (2) =
~ ~ ~ ~
𝑛! 𝑥1 𝑥2 𝑥𝑘
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑘 ! ∙ 𝑝1 𝑝2 ⋯ 𝑝𝑘
=
𝑛! 𝑥𝑠+1 𝑥𝑠+2 𝑥 (𝑛−∑𝑘
𝑖=𝑠+1 𝑥𝑖 )
𝑘 ∙ 𝑝𝑠+1 𝑝𝑠+2 ⋯ 𝑝𝑘 𝑘 (1 − ∑𝑘𝑖=𝑠+1 𝑝𝑖 )
𝑥𝑠+1 ! 𝑥𝑠+2 ! ⋯ 𝑥𝑘 ! (𝑛 − ∑𝑖=𝑠+1 𝑥𝑖 )!
𝑥 𝑥 𝑥 𝑘 𝑠 𝑘
(𝑛 − ∑𝑘𝑖=𝑠+1 𝑥𝑖 )! 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑠 𝑠
= ∙ ∑𝑠 𝑥
[∵ ∑ 𝑥𝑖 = 𝑛, ∑ 𝑥𝑖 = 𝑛 − ∑ 𝑥𝑖 ]
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! (1 − ∑𝑘 𝑝 ) 𝑖=1 𝑖
𝑖=𝑠+1 𝑖 𝑖=1 𝑖=1 𝑖=𝑠+1
𝑥1 𝑥2 𝑥𝑠
(𝑛 − ∑𝑘𝑖=𝑠+1 𝑥𝑖 )! 𝑝1 𝑝2 𝑝𝑠
= ∙{ } { } ⋯{ }
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! (1 − ∑𝑘𝑖=𝑠+1 𝑝𝑖 ) (1 − ∑𝑘𝑖=𝑠+1 𝑝𝑖 ) (1 − ∑𝑘𝑖=𝑠+1 𝑝𝑖 )
𝑥1 𝑥2 𝑥𝑠
(𝑛 − ∑𝑘𝑖=𝑠+1 𝑥𝑖 )! 𝑝1 𝑝2 𝑝𝑠
= ∙{ 𝑠 } { } ⋯{ }
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! ∑𝑖=1 𝑝𝑖 ∑𝑠𝑖=1 𝑝𝑖 ∑𝑠𝑖=1 𝑝𝑖
(𝑛 − ∑𝑘𝑖=𝑠+1 𝑥𝑖 )! 𝑥1 𝑥2 𝑥
= ∙ 𝑄1 𝑄2 ⋯ 𝑄𝑠 𝑠
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 !
𝑘 𝑘 𝑠
𝑝𝑖
← 𝑝. 𝑚. 𝑓. of 𝑀𝑢𝑙𝑡𝑖𝑛𝑜𝑚𝑖𝑎𝑙 [{𝑛 − ∑ 𝑥𝑖 } , 𝑄1 , 𝑄2 , ⋯ , 𝑄𝑠 ] , where ∑ 𝑥𝑖 = 𝑛, 𝑄𝑖 = and ∑ 𝑄𝑖 = 1
∑𝑠𝑖=1 𝑝𝑖
𝑖=𝑠+1 𝑖=1 𝑖=1
𝑛! 𝑥1 𝑥2 𝑥 3
𝑝(𝑋1 ,𝑋2 ,𝑋3 ) (𝑥1 , 𝑥2 , 𝑥3 ) 𝑥1 ! 𝑥2 ! 𝑥3 ! ∙ 𝑝1 𝑝2 𝑝3
𝑝(𝑋1 ,𝑋2 )| 𝑋3 =𝑥3 (𝑥1 , 𝑥2 ) = =
𝑝𝑋3 (𝑥3 ) 𝑛! 𝑥
∙ 𝑝 3 (1 − 𝑝3 )(𝑛−𝑥3 )
𝑥3 ! (𝑛 − 𝑥3 )! 3
𝑥 𝑥 3
(𝑛 − 𝑥3 )! 𝑝1 1 𝑝2 2
= ∙ [∵ ∑ 𝑥𝑖 = 𝑛, 𝑥1 + 𝑥2 = 𝑛 − 𝑥3 ]
𝑥1 ! 𝑥2 ! (1 − 𝑝3 )𝑥1 +𝑥2
𝑖=1
(𝑛 − 𝑥3 )! 𝑝1 𝑥1 𝑝2 𝑥2
= ∙( ) ( )
𝑥1 ! 𝑥2 ! 1 − 𝑝3 1 − 𝑝3
The conditional distribution of 𝑋1 given that 𝑋 (2) = 𝑥 (2) , where 𝑋 (2) = (𝑋2 , 𝑋3 , ⋯ , 𝑋𝑘 )𝑇
~ ~ ~
𝑛! 𝑥1 𝑥2 𝑥𝑘
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑘 ! ∙ 𝑝1 𝑝2 ⋯ 𝑝𝑘
=
𝑛! 𝑥2 𝑥3 𝑥𝑘 𝑘 (𝑛−∑𝑘
𝑖=2 𝑥𝑖 )
𝑘 ∙ 𝑝2 𝑝3 ⋯ 𝑝𝑘 (1 − ∑ 𝑖=2 𝑝𝑖 )
𝑥2 ! 𝑥3 ! ⋯ 𝑥𝑘 ! (𝑛 − ∑𝑖=2 𝑥𝑖 )!
𝑥 𝑘 𝑘
(𝑛 − ∑𝑘𝑖=2 𝑥𝑖 )! 𝑝1 1
= ∙ 𝑥1 [∵ ∑ 𝑥𝑖 = 𝑛, 𝑥1 = 𝑛 − ∑ 𝑥𝑖 ]
𝑥1 ! (1 − ∑𝑘𝑖=2 𝑝𝑖 ) 𝑖=1 𝑖=2
𝑘
(𝑛 − ∑𝑘𝑖=2 𝑥𝑖 )!
= [∵ ∑ 𝑝𝑖 = 1]
𝑥1 !
𝑖=1
0 −𝑛𝑝1 𝑝2 ⋯ −𝑛𝑝1 𝑝𝑘
0 𝑛𝑝2 (1 − 𝑝2 ) ⋯ −𝑛𝑝2 𝑝𝑘
=| |=0
⋮ ⋮ ⋱ ⋮
0 −𝑛𝑝2 𝑝𝑘 ⋯ 𝑛𝑝𝑘 (1 − 𝑝𝑘 )
Since |𝚺𝑘×𝑘 | = 0, the probability distribution of (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘 ) is singular and 𝑅𝑎𝑛𝑘(𝚺𝑘×𝑘 ) = (𝑘 − 1).
To avoid the difficulties associated with singularity, we consider the joint probability distribution of (𝑘 − 1)
random variables.
Let 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘−1 )𝑇 ~ 𝑀𝑢𝑙𝑡𝑖𝑛𝑜𝑚𝑖𝑎𝑙(𝑛, 𝑝1 , 𝑝2 , ⋯ , 𝑝𝑘−1 );
~
𝑘−1 𝑘−1
(𝑛−∑𝑘−1
𝑖=1 𝑥𝑖 )
𝑘−1 𝑘−1
𝑛! 𝑥 𝑥 𝑥𝑘−1
∙ 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑘−1 (1 − ∑ 𝑝𝑖 ) ; 𝑥𝑖 = 0(1)𝑛, ∑ 𝑥𝑖 ≤ 𝑛,
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑘−1 ! (𝑛 − ∑𝑘−1
𝑖=1 𝑥𝑖 )! 𝑖=1 𝑖=1
= 𝑘−1
(𝑛−∑𝑘−1
𝑖=1 𝑥𝑖 )
𝑘−1 𝑘−1 𝑘−1
𝑛! 𝑥
∙ ∏(𝑝𝑖 𝑖 ) (1 − ∑ 𝑝𝑖 ) ; 𝑥𝑖 = 0(1)𝑛, ∑ 𝑥𝑖 ≤ 𝑛,
∏𝑘−1 𝑘−1
𝑖=1 (𝑥𝑖 !) (𝑛 − ∑𝑖=1 𝑥𝑖 )! 𝑖=1 𝑖=1 𝑖=1
= 𝑘−1
∑ ∑ ⋯ ∑ 𝑝𝑋 (𝑥 )
~ ~
𝑥1 =0 𝑥2 =0 𝑥𝑘−1 =0
∑𝑘−1
𝑖=1 𝑥𝑖 ≤𝑛
(𝑛−∑𝑘−1
𝑖=1 𝑥𝑖 )
𝑛 𝑛 𝑛 𝑘−1
𝑛! 𝑥 𝑥 𝑥
= ∑ ∑ ⋯ ∑ ∙ 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑘−1
𝑘−1
(1 − ∑ 𝑝𝑖 )
𝑥 ! 𝑥 ! ⋯ 𝑥𝑘−1 ! (𝑛 − ∑𝑘−1
𝑖=1 𝑥𝑖 )!
𝑥1 =0 𝑥2 =0 𝑥𝑘−1 =0 1 2 𝑖=1
∑𝑘−1
𝑖=1 𝑥𝑖 ≤𝑛
𝑛
𝑘−1
The moment generating function (𝑚. 𝑔. 𝑓. ) of 𝑋 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑘−1 )𝑇 or the 𝑚. 𝑔. 𝑓. of the joint distribution
~
𝑡 𝑇𝑋 𝑘−1
𝑀𝑋 (𝑡 ) = 𝑀(𝑋1 ,𝑋2 ,⋯,𝑋𝑘−1 ) (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑘−1 ) = 𝐸 (𝑒 ~ ~ ) = 𝐸(𝑒 𝑡1 𝑋1 +𝑡2 𝑋2 +⋯+𝑡𝑘−1 𝑋𝑘−1 ) = 𝐸 (𝑒 ∑𝑖=1 𝑡𝑖𝑥𝑖 )
~ ~
𝑛 𝑛 𝑛 𝑘−1 𝑠 (𝑛−∑𝑘−1
𝑖=1 𝑥𝑖 )
∑𝑘−1
𝑛! 𝑥
= ∑ ∑ ⋯ ∑ (𝑒 𝑖=1 𝑡𝑖 𝑥𝑖 ) ∙ ∏(𝑝𝑖 𝑖 ) (1 − ∑ 𝑝𝑖 )
𝑥2 =0
∏𝑘−1 𝑘−1
𝑖=1 (𝑥𝑖 !) (𝑛 − ∑𝑖=1 𝑥𝑖 )!
𝑥1 =0 𝑥𝑘−1 =0 𝑖=1 𝑖=1
∑𝑘−1
𝑖=1 𝑥𝑖 ≤𝑛
𝑛 𝑛 𝑛 𝑘−1 𝑠 (𝑛−∑𝑘−1
𝑖=1 𝑥𝑖 )
𝑛!
= ∑ ∑ ⋯ ∑ ∙ ∏(𝑝𝑖 𝑒 𝑡𝑖 )𝑥𝑖 (1 − ∑ 𝑝𝑖 )
𝑥2 =0
∏𝑘−1 𝑘−1
𝑖=1 (𝑥𝑖 !) (𝑛 − ∑𝑖=1 𝑥𝑖 )!
𝑥1 =0 𝑥𝑘−1 =0 𝑖=1 𝑖=1
∑𝑘−1
𝑖=1 𝑥𝑖 ≤𝑛
𝑛
𝑘−1
𝑛
𝑘−1 𝑘−1
= {∑ 𝑝𝑖 𝑒 𝑡𝑖 + (1 − ∑ 𝑝𝑖 )}
𝑖=1 𝑖=1
ALTERNATIVE:
𝑀(𝑋1 ,𝑋2 ,⋯,𝑋𝑘−1 ) (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑘−1 ) = 𝑀(𝑋1 ,𝑋2 ,⋯,𝑋𝑘−1 ) (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑘−1 , 0)
𝑛 𝑛 𝑛
𝑘 𝑘−1 𝑘−1 𝑘−1 𝑘
𝑡𝑖 𝑡𝑖 𝑡𝑖
= [(∑ 𝑝𝑖 𝑒 ) ] = (∑ 𝑝𝑖 𝑒 + 𝑝𝑘 ) = {∑ 𝑝𝑖 𝑒 + (1 − ∑ 𝑝𝑖 )} [∵ ∑ 𝑝𝑖 = 1]
𝑖=1 𝑖=1 𝑖=1 𝑖=1 𝑖=1
𝑡𝑘−1 =0
MOMENTS:
𝑛−1
𝑘−1 𝑘−1
𝜕 𝜕
𝑀𝑋 (𝑡 ) = 𝑀 (𝑡 , 𝑡 , ⋯ , 𝑡𝑘−1 ) = 𝑛𝑝𝑖 𝑒 𝑡𝑖 {∑ 𝑝𝑖 𝑒 𝑡𝑖 + (1 − ∑ 𝑝𝑖 )} ; 𝑖 = 1(1)(𝑘 − 1)
𝜕𝑡𝑖 ~ ~ 𝜕𝑡𝑖 (𝑋1 ,𝑋2 ,⋯,𝑋𝑘−1 ) 1 2
𝑖=1 𝑖=1
𝜕 𝜕
𝐸(𝑋𝑖 ) = 𝑀𝑋 (𝑡 )| = 𝑀(𝑋1 ,𝑋2 ,⋯,𝑋𝑘−1 ) (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑘−1 )| ; 𝑖 = 1(1)(𝑘 − 1)
𝜕𝑡𝑖 ~ ~ 𝑡 =0 𝜕𝑡𝑖 𝑡
~ ~ 1 =𝑡2 =⋯=𝑡𝑘−1 =0
𝑛−1
𝑘−1 𝑘−1
𝜕2 𝜕2
𝑀 (𝑡 ) = 𝑀(𝑋1 ,𝑋2 ,⋯,𝑋𝑘−1 ) (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑘−1 )
𝜕𝑡𝑖2 𝑋~ ~ 𝜕𝑡𝑖2
𝑛−2 𝑛−1
𝑘−1 𝑘−1 𝑘−1 𝑘−1
𝜕2 𝜕2
𝐸(𝑋𝑖2 ) = 𝑀 (𝑡 )| = 𝑀(𝑋1 ,𝑋2 ,⋯,𝑋𝑘−1 ) (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑘−1 )|
𝜕𝑡𝑖2 𝑋~ ~ 𝑡 =0 𝜕𝑡𝑖2 𝑡 1 =𝑡2 =⋯=𝑡𝑘−1 =0
~ ~
𝑛−2 𝑛−1
𝑘−1 𝑘−1 𝑘−1 𝑘−1
𝜕2 𝜕2
𝑀𝑋 (𝑡 ) = 𝑀 (𝑡 , 𝑡 , ⋯ , 𝑡𝑘−1 )
𝜕𝑡𝑖 𝜕𝑡𝑗 ~ ~ 𝜕𝑡𝑖 𝜕𝑡𝑗 (𝑋1 ,𝑋2 ,⋯,𝑋𝑘−1 ) 1 2
𝑛−2
𝑘−1 𝑘−1
𝜕2 𝜕2
𝐸(𝑋𝑖 𝑋𝑗 ) = 𝑀 (𝑡 )| = 𝑀(𝑋1 ,𝑋2 ,⋯,𝑋𝑘−1 ) (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑘−1 )|
𝜕𝑡𝑖 𝜕𝑡𝑗 𝑋~ ~ 𝜕𝑡𝑖 𝜕𝑡𝑗
𝑡 =0 𝑡 1 =𝑡2 =⋯=𝑡𝑘 =0
~ ~
𝑛−2
𝑘−1 𝑘−1
= [𝑛(𝑛 − 1)𝑝𝑖 𝑝𝑗 𝑒 𝑡𝑖 𝑒 𝑡𝑗 {∑ 𝑝𝑖 𝑒 𝑡𝑖 + (1 − ∑ 𝑝𝑖 )} ]
𝑖=1 𝑖=1
𝑡1 =𝑡2=⋯=𝑡𝑘−1 =0
𝐶𝑜𝑣(𝑋𝑖 , 𝑋𝑗 ) −𝑛𝑝𝑖 𝑝𝑗
𝜌(𝑋𝑖 , 𝑋𝑗 ) = 𝐶𝑜𝑟𝑟(𝑋𝑖 , 𝑋𝑗 ) = =
√𝑉𝑎𝑟(𝑋𝑖 )√𝑉𝑎𝑟( 𝑋𝑗 ) √𝑛𝑝𝑖 (1 − 𝑝𝑖 )√𝑛𝑝𝑗 (1 − 𝑝𝑗 )
𝑝𝑖 𝑝𝑗
= −√ ; 𝑖, 𝑗 = 1(1)(𝑘 − 1), 𝑖 ≠ 𝑗
(1 − 𝑝𝑖 )(1 − 𝑝𝑗 )
𝑝1 𝑝2 𝑝1 𝑝𝑘−1
1 −√ ⋯ −√
(1 − 𝑝1 )(1 − 𝑝2 ) (1 − 𝑝1 )(1 − 𝑝𝑘−1 )
𝑝2 𝑝1 𝑝2 𝑝𝑘−1
= −√ 1 ⋯ −√
(1 − 𝑝2 )(1 − 𝑝1 ) (1 − 𝑝2 )(1 − 𝑝𝑘−1 )
⋮ ⋮ ⋱ ⋮
𝑝𝑘−1 𝑝1 𝑝𝑘−1 𝑝2
−√ −√ ⋯ 1
(1 − 𝑝𝑘−1 )(1 − 𝑝1 ) (1 − 𝑝𝑘−1 )(1 − 𝑝2 )
( )(𝑘−1)×(𝑘−1)
𝑛
𝑘−1 𝑘−1
= [{∑ 𝑝𝑖 𝑒 𝑡𝑖 + (1 − ∑ 𝑝𝑖 )} ]
𝑖=1 𝑖=1
𝑡𝑗 =0; 𝑗=(𝑠+1)(1)(𝑘−1)
𝑛 𝑛
𝑠 𝑘−1 𝑘−1 𝑠 𝑠
𝑡𝑖 𝑡𝑖
= {∑ 𝑝𝑖 𝑒 + ∑ 𝑝𝑖 + (1 − ∑ 𝑝𝑖 )} = {∑ 𝑝𝑖 𝑒 + (1 − ∑ 𝑝𝑖 )}
𝑖=1 𝑖=𝑠+1 𝑖=1 𝑖=1 𝑖=1
𝑠 𝑠
𝑛
𝑘−1 𝑘−1
𝑡𝑖
= 𝑝𝑖 𝑒 + ∑ 𝑝𝑗 + (1 − ∑ 𝑝𝑖 )
𝑗=1 𝑖=1
{ 𝑗≠𝑖 }
𝑛
𝑘−1
𝑛
= (𝑝𝑖 𝑒 𝑡𝑖 + (1 − 𝑝𝑖 )) ; 𝑖 = 1(1)(𝑘 − 1) ← MGF of 𝐵𝑖𝑛(𝑛, 𝑝𝑖 )
ALTERNATIVE
(𝑛−∑𝑘−1
𝑖=1 𝑥𝑖 )
𝑘−1 𝑘−1
𝑛! 𝑥
= ∑ ∑ ⋯ ∑ ∙ ∏(𝑝𝑖 𝑖 ) (1 − ∑ 𝑝𝑖 )
∏𝑘−1
𝑖=1 (𝑥𝑖 !) (𝑛 − ∑𝑘−1
𝑖=1 𝑥𝑖 )!
𝑥𝑗 =0(1)𝑛, 𝑗=(𝑠+1)(1)(𝑘−1), ∑𝑘−1 𝑖=1 𝑖=1
𝑗=1 𝑥𝑗 ≤𝑛
𝑛! 𝑥 𝑥 𝑥
= ∙ 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑠 𝑠
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! (𝑛 − ∑𝑠𝑖=1 𝑥𝑖 )!
(𝑛−∑𝑘−1
𝑖=1 𝑥𝑖 )
𝑘−1 𝑘−1
(𝑛 − ∑𝑠𝑖=1 𝑥𝑖 )! 𝑥𝑖
∙ ∑ ∑ ⋯ ∑ ∙ ∏ (𝑝𝑖 ) (1 − ∑ 𝑝𝑖 )
∏𝑘−1
𝑖=𝑠+1(𝑥𝑖 !) (𝑛 − ∑𝑘−1
𝑖=1 𝑥𝑖 )!
𝑥𝑗 =0(1)𝑛, 𝑗=(𝑠+1)(1)(𝑘−1), ∑𝑘−1
𝑥𝑗 ≤𝑛 𝑖=𝑠+1 𝑖=1
𝑗=1
(𝑛−∑𝑠𝑖=1 𝑥𝑖 )
𝑘−1
𝑛! 𝑥 𝑥 𝑥
= ∙ 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑠 𝑠 ∙ {(1 − ∑ 𝑝𝑖 ) + 𝑝𝑠+1 + 𝑝𝑠+2 + ⋯ + 𝑝𝑘−1 }
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! (𝑛 − ∑𝑠𝑖=1 𝑥𝑖 )!
𝑖=1
𝑠 (𝑛−∑𝑠𝑖=1 𝑥𝑖 )
𝑛! 𝑥 𝑥 𝑥
= ∙ 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑠 𝑠 ∙ (1 − ∑ 𝑝𝑖 )
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! (𝑛 − ∑𝑠𝑖=1 𝑥𝑖 )!
𝑖=1
𝑠 𝑠
(𝑛−∑𝑘−1
𝑖=1 𝑥𝑖 )
𝑘−1 𝑘−1
𝑛! 𝑥
= ∑ ∑ ⋯ ∑ ∙ ∏(𝑝𝑖 𝑖 ) (1 − ∑ 𝑝𝑖 )
∏𝑘−1
𝑖=1 (𝑥𝑖 !) (𝑛 − ∑𝑘−1
𝑖=1 𝑥𝑖 )!
𝑥𝑗 =0(1)𝑛, 𝑥𝑗 ≠𝑥𝑖 , ∑𝑘−1 𝑖=1 𝑖=1
𝑗=1 𝑥𝑗 ≤𝑛
𝑛! 𝑥 (𝑛 − 𝑥𝑖 )!
= ∙𝑝 𝑖 ∙ ∑ ∑ ⋯ ∑
𝑥𝑖 ! (𝑛 − 𝑥𝑖 )! 𝑖 𝑥1 ! ⋯ 𝑥𝑖−1 ! 𝑥𝑖+1 ! ⋯ 𝑥𝑘−1 ! (𝑛 − ∑𝑘−1
𝑖=1 𝑥𝑖 )!
𝑥𝑗 =0(1)(𝑛−𝑥𝑖 ), 𝑥𝑗 ≠𝑥𝑖 , ∑𝑘−1
𝑗=1,𝑗≠𝑖 𝑥𝑗 ≤𝑛−𝑥𝑖
(𝑛−∑𝑘−1
𝑖=1 𝑥𝑖 )
𝑘−1
𝑥 𝑥 𝑥 𝑥
∙ 𝑝1 1 ⋯ 𝑝𝑖−1
𝑖−1
𝑝𝑖+1
𝑖+1
⋯ 𝑝𝑘−1
𝑘−1
(1 − ∑ 𝑝𝑖 )
𝑖=1
(𝑛−𝑥𝑖 )
𝑘−1
𝑛! 𝑥
= ∙ 𝑝 𝑖 ∙ {(1 − ∑ 𝑝𝑖 ) 𝑝1 + ⋯ + 𝑝𝑖−1 + 𝑝𝑖+1 + ⋯ + 𝑝𝑘−1 }
𝑥𝑖 ! (𝑛 − 𝑥𝑖 )! 𝑖
𝑖=1
𝑘
𝑛! 𝑥
= ∙ 𝑝 𝑖 ∙ (1 − 𝑝𝑖 )(𝑛−𝑥𝑖) [∵ ∑ 𝑝𝑖 = 1]
𝑥𝑖 ! (𝑛 − 𝑥𝑖 )! 𝑖
𝑖=1
← 𝑝. 𝑚. 𝑓. of 𝐵𝑖𝑛(𝑛, 𝑝𝑖 ) ; 𝑖 = 1(1)(𝑘 − 1)
CONDITIONAL DISTRIBUTION
The conditional distribution of 𝑋 (1) = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑠 )𝑇 given that 𝑋 (2) = 𝑥 (2) , where 𝑋 (2) =
~ ~ ~ ~
(𝑋𝑠+1 , 𝑋𝑠+2 , ⋯ , 𝑋𝑘−1 )𝑇 [ i. e. given that (𝑋𝑠+1 = 𝑥𝑠+1 , 𝑋𝑠+2 = 𝑥𝑠+2 , ⋯ , 𝑋𝑘−1 = 𝑥𝑘−1 )] is given by
𝑥 𝑥 𝑥 (𝑛−∑𝑘−1 𝑥𝑖 )
(𝑛 − ∑𝑘−1𝑖=𝑠+1 𝑥𝑖 )! 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑠 𝑠 (1 − ∑𝑘−1
𝑖=1 𝑝𝑖 )
𝑖=1
= ∙
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! (𝑛 − ∑𝑘−1
𝑖=1 𝑥𝑖 )! (1 − ∑𝑘−1
{(𝑛−∑𝑘−1 𝑠
𝑖=1 𝑥𝑖 )+∑𝑖=1 𝑥𝑖 }
𝑖=𝑠+1 𝑝𝑖 )
𝑥 𝑥 𝑥 (𝑛−∑𝑘−1 𝑥𝑖 )
(𝑛 − ∑𝑘−1𝑖=𝑠+1 𝑥𝑖 )! 𝑝1 1 𝑝2 2 ⋯ 𝑝𝑠 𝑠 {(1 − ∑𝑘−1 𝑠
𝑖=𝑠+1 𝑝𝑖 ) − ∑𝑖=1 𝑝𝑖 }
𝑖=1
= ∙ ∙
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! (𝑛 − ∑𝑘−1 (∑𝑠 𝑥𝑖 ) (𝑛−∑𝑘−1
𝑖=1 𝑥𝑖 )! (1 − ∑𝑘−1 𝑝𝑖 ) 𝑖=1 (1 − ∑𝑘−1 𝑖=1 𝑥𝑖 )
𝑖=𝑠+1 𝑖=𝑠+1 𝑝𝑖 )
(𝑛 − ∑𝑘−1
𝑖=𝑠+1 𝑥𝑖 )!
=
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! ((𝑛 − ∑𝑘−1 𝑠
𝑖=𝑠+1 𝑥𝑖 ) − ∑𝑖=1 𝑥𝑖 ) !
𝑥1 𝑥2 𝑥𝑠
𝑝1 𝑝2 𝑝𝑠
∙{ } { } ⋯{ }
(1 − ∑𝑘−1
𝑖=𝑠+1 𝑝𝑖 ) (1 − ∑𝑘−1
𝑖=𝑠+1 𝑝𝑖 ) (1 − ∑𝑘−1
𝑖=𝑠+1 𝑝𝑖 )
𝑠 ((𝑛−∑𝑘−1 𝑠
𝑖=𝑠+1 𝑥𝑖 )−∑𝑖=1 𝑥𝑖 )
𝑝𝑖
∙ {1 − ∑ ( )}
𝑖=1
(1 − ∑𝑘−1
𝑖=𝑠+1 𝑝𝑖 )
𝑠 ((𝑛−∑𝑘−1 𝑠
𝑖=𝑠+1 𝑥𝑖 )−∑𝑖=1 𝑥𝑖 )
(𝑛 − ∑𝑘−1
𝑖=𝑠+1 𝑥𝑖 )! 𝑥 𝑥 𝑥
= ∙ 𝑄1 1 𝑄2 2 ⋯ 𝑄𝑠 𝑠 ∙ {1 − ∑ 𝑄𝑖 }
𝑥1 ! 𝑥2 ! ⋯ 𝑥𝑠 ! ((𝑛 − ∑𝑘−1
𝑖=𝑠+1 𝑥𝑖 ) − ∑𝑠𝑖=1 𝑥𝑖 ) ! 𝑖=1
𝑘−1
← 𝑝. 𝑚. 𝑓. of 𝑀𝑢𝑙𝑡𝑖𝑛𝑜𝑚𝑖𝑎𝑙 [{𝑛 − ∑ 𝑥𝑖 } , 𝑄1 , 𝑄2 , ⋯ , 𝑄𝑠 ] ;
𝑖=𝑠+1
𝑘−1 𝑠
𝑝𝑖
where ∑ 𝑥𝑖 ≤ 𝑛, 𝑄𝑖 = ( ) and ∑ 𝑄𝑖 < 1; 𝑖 = 1(1)𝑠
(1 − ∑𝑘−1
𝑖=𝑠+1 𝑝𝑖 )
𝑖=1 𝑖=1
𝑛! 𝑥 𝑥 𝑥
∙ 𝑝 1 𝑝 2 𝑝 3 (1 − 𝑝1 − 𝑝2 − 𝑝3 )(𝑛−𝑥1 −𝑥2 −𝑥3 )
𝑥1 ! 𝑥2 ! 𝑥3 ! (𝑛 − 𝑥1 − 𝑥2 − 𝑥3 )! 1 2 3
=
𝑛! 𝑥
∙ 𝑝 3 (1 − 𝑝3 )(𝑛−𝑥3 )
𝑥3 ! (𝑛 − 𝑥3 )! 3
(𝑛 − 𝑥3 )! 𝑥 𝑥
= ∙ 𝑄1 1 𝑄2 2 ∙ (1 − 𝑄1 − 𝑄2 )((𝑛−𝑥3 )−𝑥1 −𝑥2 )
𝑥1 ! 𝑥2 ! ((𝑛 − 𝑥3 ) − 𝑥1 − 𝑥2 )!
3
𝑝𝑖
← 𝑝. 𝑚. 𝑓. of 𝑀𝑢𝑙𝑡𝑖𝑛𝑜𝑚𝑖𝑎𝑙((𝑛 − 𝑥3 ), 𝑄1 , 𝑄2 ), where ∑ 𝑥𝑖 ≤ 𝑛, 𝑄𝑖 = ( ) and 𝑄1 + 𝑄2 < 1; 𝑖 = 1, 2
1 − 𝑝3
𝑖=1
The conditional distribution of 𝑋1 given that 𝑋 (2) = 𝑥 (2) , where 𝑋 (2) = (𝑋2 , 𝑋3 , ⋯ , 𝑋𝑘−1 )𝑇
~ ~ ~
𝑥 ((𝑛−∑𝑘−1
𝑖=2 𝑥𝑖 )−𝑥1 )
(𝑛 − ∑𝑘−1𝑖=2 𝑥𝑖 )!
𝑝1 1 ((1 − ∑𝑘−1
𝑖=2 𝑝𝑖 ) − 𝑝1 )
= ∙
𝑥1 ! (𝑛 − ∑𝑘−1
𝑖=1 𝑥𝑖 )!
𝑥1 ((𝑛−∑𝑘−1
𝑖=2 𝑥𝑖 )−𝑥1 )
(1 − ∑𝑘−1 𝑘−1
𝑖=2 𝑝𝑖 ) (1 − ∑𝑖=2 𝑝𝑖 )
𝑥1 ((𝑛−∑𝑘−1
𝑖=2 𝑥𝑖 )−𝑥1 )
(𝑛 − ∑𝑘−1
𝑖=2 𝑥𝑖 )! 𝑝1 𝑝1
= ∙{ } ∙ {1 − }
𝑥1 ! ((𝑛 − ∑𝑘−1
𝑖=2 𝑥𝑖 ) − 𝑥1 ) !
(1 − ∑𝑘−1
𝑖=2 𝑝𝑖 ) (1 − ∑𝑘−1
𝑖=2 𝑝𝑖 )
𝑘−1
𝑝1
← 𝑝. 𝑚. 𝑓. of 𝐵𝑖𝑛𝑜𝑚𝑖𝑎𝑙 ({𝑛 − ∑ 𝑥𝑖 } , { })
𝑖=2
(1 − ∑𝑘−1
𝑖=2 𝑝𝑖 )
𝑘−1
𝑝1
← 𝑝. 𝑚. 𝑓. of 𝐵𝑖𝑛𝑜𝑚𝑖𝑎𝑙 ({𝑛 − ∑ 𝑥𝑖 } , 𝑄1 ) , where 𝑄1 = { }
𝑖=2
(1 − ∑𝑘−1
𝑖=2 𝑝𝑖 )
𝑘−1
(2) (2) 𝑝1
Now, 𝐸 (𝑋1 |𝑋 =𝑥 ) = 𝐸(𝑋1 |(𝑥2 , 𝑥3 , ⋯ , 𝑥𝑘−1 )) = {𝑛 − ∑ 𝑥𝑖 } ∙ { }
~ ~
𝑖=2
(1 − ∑𝑘−1
𝑖=2 𝑝𝑖 )
𝑝1 𝑝1 𝑝1 𝑝1
= 𝑛{ }−{ } 𝑥2 − { } 𝑥3 − ⋯ − { } 𝑥𝑘−1
(1 − ∑𝑘−1
𝑖=2 𝑝𝑖 ) (1 − ∑𝑘−1
𝑖=2 𝑝𝑖 ) (1 − ∑𝑘−1
𝑖=2 𝑝𝑖 ) (1 − ∑𝑘−1
𝑖=2 𝑝𝑖 )
𝑝1
𝛽12.34⋯𝑘−1 = − { }
(1 − 𝑝2 − ∑𝑘−1
𝑖=3 𝑝𝑖 )
Similarly, the regression coefficient of 𝑋2 on 𝑋1 , eliminating the effects of (𝑋3 , ⋯ , 𝑋𝑘−1 ), is given by
𝑝2
𝛽21.34⋯𝑘−1 = − { }
(1 − 𝑝1 − ∑𝑘−1
𝑖=3 𝑝𝑖 )
The partial correlation coefficient of 𝑋1 and 𝑋2 , eliminating the effects of (𝑋3 , ⋯ , 𝑋𝑘−1 ), is given by
2 𝑝1 𝑝2
𝜌12.34⋯𝑘−1 =[ ]
(1 − 𝑝2 − ∑𝑖=3 𝑝𝑖 )(1 − 𝑝1 − ∑𝑘−1
𝑘−1
𝑖=3 𝑝𝑖 )
1
𝑝1 𝑝2 2
𝜌12.34⋯𝑘−1 = − [ ] ; [∵ 𝛽12.34⋯𝑘−1 , 𝛽21.34⋯𝑘−1 < 0]
(1 − 𝑝2 − ∑𝑘−1 𝑘−1
𝑖=3 𝑝𝑖 )(1 − 𝑝1 − ∑𝑖=3 𝑝𝑖 )
1
|𝚺| 2
𝜌1.23⋯(𝑘−1) = 𝐶𝑜𝑟𝑟(𝑥1 , 𝑋1.23⋯(𝑘−1) ) = (1 − )
𝜎11 |𝚺2 |
𝑘−1
𝑘−1
𝑘−1
|𝚺(𝑘−1)×(𝑘−1) | = 𝑛 𝑝1 𝑝2 𝑝3 ⋯ 𝑝𝑘−1 {1 − ∑ 𝑝𝑖 } > 0 ;
𝑖=1
𝑘−1
since 0 < 𝑝𝑖 < 1, 𝑖 = 1(1)(𝑘 − 1), ∑ 𝑝𝑖 < 1 and 𝑛 (> 0) is a positive integer.
𝑖=1
|𝚺2 (𝑘−2)×(𝑘−2) | = Determinant obtained by deleting the1st row and 1st column of 𝚺(𝑘−1)×(𝑘−1)
1
|𝚺| 2
∴ 𝜌1.23⋯(𝑘−1) = (1 − )
𝜎11 |𝚺2 |
1
𝑛𝑘−1 𝑝1 𝑝2 𝑝3 ⋯ 𝑝𝑘−1 {1 − ∑𝑘−1
𝑖=1 𝑝𝑖 }
2
= [1 − ]
𝑛𝑝1 (1 − 𝑝1 ) ∙ 𝑛𝑘−2 𝑝2 𝑝3 𝑝4 ⋯ 𝑝𝑘−1 {1 − ∑𝑘−1
𝑖=2 𝑝𝑖 }
1
{1 − ∑𝑘−1
𝑖=1 𝑝𝑖 }
2
= [1 − ]
(1 − 𝑝1 ){1 − ∑𝑘−1
𝑖=2 𝑝𝑖 }
1
1 − ∑𝑘−1 𝑘−1 𝑘−1
𝑖=2 𝑝𝑖 − 𝑝1 + 𝑝1 ∑𝑖=2 𝑝𝑖 − 1 + ∑𝑖=1 𝑝𝑖
2
=[ ]
(1 − 𝑝1 ){1 − ∑𝑘−1
𝑖=2 𝑝𝑖 }
1
𝑝1 ∑𝑘−1
𝑖=2 𝑝𝑖
2
=[ ]
(1 − 𝑝1 ){1 − ∑𝑘−1
𝑖=2 𝑝𝑖 }
1
𝑝1 ∑𝑘−1
𝑖=2 𝑝𝑖
2
𝜌1.23⋯(𝑘−1) =[ ]
(1 − 𝑝1 ){1 − ∑𝑘−1
𝑖=2 𝑝𝑖 }
The partial correlation coefficient between 𝑋1 and 𝑋2 , eliminating the effect of 𝑋 (3) = (𝑋3 , 𝑋4 , ⋯ , 𝑋𝑘−1 )𝑇 is
~
given by
𝚺12
𝜌12.34⋯(𝑘−1) = − ,
√𝚺11 ∙ 𝚺22
= 𝑛𝑘−2 𝑝1 𝑝2 𝑝3 𝑝4 ⋯ 𝑝𝑘−1
− 𝑝2 − ∑ 𝑝𝑖 } (1 − 𝑝3 ) −𝑝4 ⋯ −𝑝𝑘−1
|{1 |
𝑖=3
𝑘−2
=𝑛 𝑝2 𝑝3 𝑝4 ⋯ 𝑝𝑘−1 𝑘−1
since 0 < 𝑝𝑖 < 1, 𝑖 = 1(1)(𝑘 − 1), ∑ 𝑝𝑖 < 1 and 𝑛 (> 0) is a positive integer.
𝑖=1
𝑘−1
𝑘−2
=𝑛 𝑝1 𝑝3 𝑝4 ⋯ 𝑝𝑘−1 {1 − 𝑝1 − ∑ 𝑝𝑖 } > 0 ;
𝑖=3
𝑘−1
since 0 < 𝑝𝑖 < 1, 𝑖 = 1(1)(𝑘 − 1), ∑ 𝑝𝑖 < 1 and 𝑛 (> 0) is a positive integer.
𝑖=1
𝚺12
∴ 𝜌12.34⋯(𝑘−1) = −
√𝚺11 ∙ 𝚺22
𝑛𝑘−2 𝑝1 𝑝2 𝑝3 𝑝4 ⋯ 𝑝𝑘−1
=−
√𝑛𝑘−2 𝑝2 𝑝3 𝑝4 ⋯ 𝑝𝑘−1 {1 − 𝑝2 − ∑𝑘−1
𝑖=3 𝑝𝑖 } ∙ √𝑛
𝑘−2 𝑝 𝑝 𝑝 ⋯ 𝑝
1 3 4
𝑘−1
𝑘−1 {1 − 𝑝1 − ∑𝑖=3 𝑝𝑖 }
1
𝑝1 𝑝2 2
= −[ ]
(1 − 𝑝2 − ∑𝑘−1 𝑘−1
𝑖=3 𝑝𝑖 )(1 − 𝑝1 − ∑𝑖=3 𝑝𝑖 )
The partial correlation coefficient between 𝑋1 and 𝑋2 , eliminating the effect of 𝑋 (3) = (𝑋3 , 𝑋4 , ⋯ , 𝑋𝑘−1 )𝑇 is
~
given by
1
𝑝1 𝑝2 2
𝜌12.34⋯𝑘−1 = − [ ]
(1 − 𝑝2 − ∑𝑘−1 𝑘−1
𝑖=3 𝑝𝑖 )(1 − 𝑝1 − ∑𝑖=3 𝑝𝑖 )
2 𝑝1 𝑝2
𝜌12.34⋯𝑘−1 =[ ]
(1 − 𝑝2 − ∑𝑖=3 𝑝𝑖 )(1 − 𝑝1 − ∑𝑘−1
𝑘−1
𝑖=3 𝑝𝑖 )
ALTERNATIVE
𝑘−1
𝑝1
∴ 𝐸(𝑋1 |(𝑥2 , 𝑥3 , ⋯ , 𝑥𝑘−1 )) = {𝑛 − ∑ 𝑥𝑖 } ∙ { }
𝑖=2
(1 − ∑𝑘−1
𝑖=2 𝑝𝑖 )
𝑝1 𝑝1 𝑝1 𝑝1
= 𝑛{ }−{ } 𝑥2 − { } 𝑥3 − ⋯ − { } 𝑥𝑘−1
(1 − ∑𝑘−1
𝑖=2 𝑝𝑖 ) (1 − ∑𝑘−1
𝑖=2 𝑝𝑖 ) (1 − ∑𝑘−1
𝑖=2 𝑝𝑖 ) (1 − ∑𝑘−1
𝑖=2 𝑝𝑖 )
𝑘−1
𝑝1
𝑋2 |(𝑥1 , 𝑥3 , ⋯ , 𝑥𝑘−1 ) ~ 𝐵𝑖𝑛𝑜𝑚𝑖𝑎𝑙 ({𝑛 − 𝑥1 − ∑ 𝑥𝑖 } , { })
𝑖=3
(1 − 𝑝1 − ∑𝑘−1
𝑖=3 𝑝𝑖 )
𝑘−1
𝑝1
∴ 𝐸(𝑋2 |(𝑥1 , 𝑥3 , ⋯ , 𝑥𝑘−1 )) = {𝑛 − 𝑥1 − ∑ 𝑥𝑖 } ∙ { }
(1 − 𝑝1 − ∑𝑘−1
𝑖=3 𝑝𝑖 )
𝑖=3
𝑝1 𝑝1 𝑝1
= 𝑛{ 𝑘−1
}−{ 𝑘−1
} 𝑥2 − { } 𝑥3 − ⋯
(1 − 𝑝1 − ∑𝑖=3 𝑝𝑖 ) (1 − 𝑝1 − ∑𝑖=3 𝑝𝑖 ) (1 − 𝑝1 − ∑𝑘−1
𝑖=3 𝑝𝑖 )
𝑝1
−{ } 𝑥𝑘−1
(1 − 𝑝1 − ∑𝑘−1
𝑖=3 𝑝𝑖 )
𝑝1
𝛽12.34⋯𝑘−1 = − { }
(1 − 𝑝2 − ∑𝑘−1
𝑖=3 𝑝𝑖 )
Similarly, the regression coefficient of 𝑋2 on 𝑋1 , eliminating the effects of (𝑋3 , ⋯ , 𝑋𝑘−1 ), is given by
𝑝2
𝛽21.34⋯𝑘−1 = − { }
(1 − 𝑝1 − ∑𝑘−1
𝑖=3 𝑝𝑖 )
The partial correlation coefficient between 𝑋1 and 𝑋2 , eliminating the effects of (𝑋3 , ⋯ , 𝑋𝑘−1 ), is given by
2 𝑝1 𝑝2
𝜌12.34⋯𝑘−1 =[ ]
(1 − 𝑝2 − ∑𝑘−1 𝑘−1
𝑖=3 𝑝𝑖 )(1 − 𝑝1 − ∑𝑖=3 𝑝𝑖 )
1
𝑝1 𝑝2 2
𝜌12.34⋯𝑘−1 = − [ ] ; [∵ 𝛽12.34⋯𝑘−1 , 𝛽21.34⋯𝑘−1 < 0]
(1 − 𝑝2 − ∑𝑘−1 𝑘−1
𝑖=3 𝑝𝑖 )(1 − 𝑝1 − ∑𝑖=3 𝑝𝑖 )
𝑇
Suppose 𝑿 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) be a p-dimensional random vector of p-random variables.
𝑇
The p-dimensional random vector 𝑿 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) have multivariate normal distribution if their joint probability
density function can be written in the form
1
𝑓𝑿 (𝒙) = 𝐾 × exp {− (𝒙 − 𝒃)𝑇 𝑨(𝒙 − 𝒃)} ; 𝒙 ∈ ℝ𝑝 , 𝒃 ∈ ℝ𝑝 , 𝑨 is 𝑝. 𝑑.
2
𝑝 𝑝
1
= 𝐾 × exp [− ∑ ∑{𝑎𝑖𝑗 (𝑥𝑖 − 𝑏𝑖 )(𝑥𝑗 − 𝑏𝑗 )}]
2
𝑖=1 𝑗=1
𝑇
where 𝒙 = (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) ∈ ℝ𝑝 and 𝐾(> 0) is a suitable constant.
∞ ∞ ∞
∞ ∞ ∞
1
⟹ 𝐾 × ∫ ∫ ⋯ ∫ exp {− (𝒙 − 𝒃)𝑇 𝑨(𝒙 − 𝒃)} 𝑑𝑥1 𝑑𝑥2 ⋯ 𝑑𝑥𝑝 = 1
2
−∞ −∞ −∞
Since 𝑨 is a positive definite, there exist a non-singular matrix 𝑪 such that 𝑪𝑪𝑇 = 𝑨.
𝑿 1 1 −1 1 −1
|𝐽| = |𝐽 ( )| = 𝑇 = = (√|𝑪𝑪𝑇 |) = = (√|𝑨|) .
𝒀 |𝑪 | √|𝑪𝑪𝑇 | √|𝑨|
𝑇
Now, (𝑿 − 𝒃)𝑇 𝑨(𝑿 − 𝒃) = (𝑿 − 𝒃)𝑇 𝑪𝑪𝑇 (𝑿 − 𝒃) = (𝑪𝑇 (𝑿 − 𝒃)) (𝑪𝑇 (𝑿 − 𝒃)) = 𝒀𝑇 𝒀.
1
∴ 𝐾 × ∫ exp (− (𝒙 − 𝒃)𝑇 𝑨(𝒙 − 𝒃)) 𝑑𝒙 = 1
2
ℝ𝑝
−1 1 𝑇
⟹ 𝐾 ⋅ (√|𝑨|) × ∫ exp (− 𝒚 𝒚) 𝑑𝒚 = 1
2
ℝ𝑝
∞ ∞ ∞ 𝑝
−1 1
⟹ 𝐾 ⋅ (√|𝑨|) × ∫ ∫ ⋯ ∫ exp (− ∑ 𝑦𝑖2 ) 𝑑𝑦1 𝑑𝑦2 ⋯ 𝑑𝑦𝑝 = 1
2
−∞ −∞ −∞ 𝑖=1
∞ ∞ ∞ 𝑝
−1 1 2
⟹ 𝐾 ⋅ (√|𝑨|) × ∫ ∫ ⋯ ∫ {∏ exp (− 𝑦 )} 𝑑𝑦1 𝑑𝑦2 ⋯ 𝑑𝑦𝑝 = 1
2 𝑖
−∞ −∞ −∞ 𝑖=1
𝑝 ∞
−1 1 2
⟹ 𝐾 ⋅ (√|𝑨|) × ∏ { ∫ exp (− 𝑦 ) 𝑑𝑦𝑖 } = 1
2 𝑖
𝑖=1 −∞
𝑝
−1
⟹ 𝐾 ⋅ (√|𝑨|) × ∏(√2𝜋) = 1 [Since 𝑌𝑖 ~𝑖𝑖𝑑 𝑁1 (0, 1) ; 𝑖 = 1(1)𝑝]
𝑖=1
√|𝑨|
⟹ 𝐾= 𝑝 .
(2𝜋)2
√|𝑨| 1
𝑓𝑿 (𝒙) = 𝑝 × exp {− (𝒙 − 𝒃)𝑇 𝑨(𝒙 − 𝒃)} ; 𝒙 ∈ ℝ𝑝 , 𝒃 ∈ ℝ𝑝 , 𝑨 is 𝑝. 𝑑.
(2𝜋)2 2
IN TERMS OF MOMENTS:
Let us define,
𝑇 𝑇
Mean vector of 𝑿 = 𝐸(𝑿) = (𝐸(𝑋1 ), 𝐸(𝑋2 ), ⋯ , 𝐸(𝑋𝑝 )) = (𝜇1 , 𝜇2 , ⋯ , 𝜇𝑝 ) = 𝝁 (𝑠𝑎𝑦) (= ∫ 𝒙 ∙ 𝑓𝑿 (𝒙)𝑑𝒙),
ℝ𝑝
(𝑋1 − 𝜇1 )
(𝑋2 − 𝜇2 )
=𝐸 ( ) ((𝑋1 − 𝜇1 ) (𝑋2 − 𝜇2 ) ⋯ (𝑋𝑝 − 𝜇𝑝 ))
⋮
{ (𝑋𝑝 − 𝜇𝑝 ) }
𝐶𝑜𝑣(𝑋𝑖 , 𝑋𝑗 )
𝐶𝑜𝑟𝑟(𝑋𝑖 , 𝑋𝑗 ) = if 𝑖 ≠ 𝑗
𝜌𝑖𝑗 = √𝑉𝑎𝑟(𝑋𝑖 ) ∙ √𝑉𝑎𝑟(𝑋𝑗 ) and 𝜎𝑖𝑗 = 𝜌𝑖𝑗 𝜎𝑖 𝜎𝑗 ; 𝑖, 𝑗 = 1(1)𝑝.
{𝐶𝑜𝑟𝑟(𝑋𝑖 , 𝑋𝑗 ) = 1 if 𝑖 = 𝑗,
𝑇 1
𝝁 = 𝐸(𝑿) = (𝜇1 , 𝜇2 , ⋯ , 𝜇𝑝 ) = ∫ 𝒙 ∙ 𝑓𝑿 (𝒙)𝑑𝒙 = 𝐾 × ∫ 𝒙 ∙ exp (− (𝒙 − 𝒃)𝑇 𝑨(𝒙 − 𝒃)) 𝑑𝒙 , and
2
ℝ𝑝 ℝ𝑝
Since 𝑨 is a positive definite, there exist a non-singular matrix 𝑪 such that 𝑪𝑪𝑇 = 𝑨.
𝑿 1 1 −1 1 −1
|𝐽| = |𝐽 ( )| = = = (√|𝑪𝑪𝑇 |) = = (√|𝑨|) .
𝒀 ||𝑪𝑇 || √|𝑪𝑪𝑇 | √|𝑨|
𝑇
Now, (𝑿 − 𝒃)𝑇 𝑨(𝑿 − 𝒃) = (𝑿 − 𝒃)𝑇 𝑪𝑪𝑇 (𝑿 − 𝒃) = (𝑪𝑇 (𝑿 − 𝒃)) (𝑪𝑇 (𝑿 − 𝒃)) = 𝒀𝑇 𝒀.
1
𝝁 = 𝐸(𝑿) = 𝐾 × ∫ 𝑥 ∙ exp (− (𝒙 − 𝒃)𝑇 𝑨(𝒙 − 𝒃)) 𝑑𝒙
2
ℝ𝑝
1 𝑇 −1
= 𝐾 × ∫ {(𝑪𝑇 )−1 𝒚 + 𝒃} ∙ exp (− 𝒚 𝒚) ∙ (√|𝑨|) 𝑑𝒚
2
ℝ𝑝
= (𝑪𝑇 )−1 ∙ 𝐸(𝒀) + 𝒃 = 𝒃 [Since 𝒀 ~ 𝑵𝒑 (𝟎, 𝑰𝒑 ) 𝑖. 𝑒. , 𝑌𝑖′ s are 𝑖𝑖𝑑 𝑁(0,1); 𝑖 = 1(1)𝑝 and 𝐸(𝒀) = 𝟎]
∴ 𝒀 = 𝑪𝑇 (𝑿 − 𝝁)
∴ 𝚺 = 𝑨−1 𝑖. 𝑒. , 𝑨 = 𝚺 −1 .
1
∴ 𝐾= 𝑝 , 𝒃=𝝁 and 𝑨 = 𝚺 −1 .
(2𝜋) 2 √|𝚺|
1 1
𝑓𝑿 (𝒙) = 𝑝 × exp {− (𝒙 − 𝝁)𝑇 𝚺 −1 (𝒙 − 𝝁)} ; 𝒙 ∈ ℝ𝑝 , 𝝁 ∈ ℝ𝑝 , 𝚺 is 𝑝. 𝑑.
(2𝜋)2 √|𝚺| 2
NOTE:
𝑥−𝜇 2
The term ( ) = (𝑥 − 𝜇)(𝜎 2 )−1 (𝑥 − 𝜇) in the exponent of the univariate normal density function measures the square
𝜎
of the distance from 𝑥 to 𝜇 in standard deviation units. This can be generalized for a 𝑝 × 1 vector 𝒙𝑝×1 of observations
on several variables as (𝒙 − 𝝁)𝑇 𝚺 −1 (𝒙 − 𝝁).
The 𝑝 × 1 vector 𝝁 represents the expected value of the random vector 𝑿, and the 𝑝 × 𝑝 matrix 𝚺 is the variance-
covariance matrix of 𝑿. We shall assume that the symmetric matrix 𝚺 is positive definite, so the expression
(𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) is the square of the generalized distance from 𝒙 to 𝝁, or the Mahalanobis distance.
𝑝
The constant (2𝜋) 2 √|𝚺| makes the volume under the surface of the multivariate density function unity for any 𝑝.
MODE:
The 𝑝-variate normal density has a maximum value when the squared distance (𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) = 0, i.e.,
when 𝑿 = 𝝁. Thus, 𝝁 is the point of maximum density, or mode, as well as the expected value or mean of 𝑿.
RESULT 1:
The random vector 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) iff 𝒀 = 𝑪−1 (𝑿 − 𝝁) ~ 𝑵𝒑 (𝟎, 𝑰𝒑 ) i. e. , 𝑌𝑖 ′s are iid 𝑁(0, 1), where 𝑪𝑝×𝑝 is
non-singular matrix such that 𝑪𝑪𝑇 = 𝚺.
PROOF:
𝑇
Suppose 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺).
1 1
𝑓𝑿 (𝒙) = 𝑝 × exp {− (𝒙 − 𝝁)𝑇 𝚺 −1 (𝒙 − 𝝁)} ; 𝒙 ∈ ℝ𝑝 , 𝝁 ∈ ℝ𝑝 , 𝚺 is 𝑝. 𝑑.
(2𝜋)2 √|𝚺| 2
Since 𝚺 is a positive definite, there exist a non-singular matrix 𝑪 such that 𝑪𝑪𝑇 = 𝚺.
𝑿
|𝐽| = |𝐽 ( )| = ||𝑪|| = √|𝑪𝑪𝑇 | = √|𝚺|.
𝒀
𝑇
∴ (𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) = (𝑿 − 𝝁)𝑇 (𝑪𝑪𝑇 )−1 (𝑿 − 𝝁) = (𝑪−1 (𝑿 − 𝝁)) (𝑪−1 (𝑿 − 𝝁)) = 𝒀𝑇 𝒀 [∵ (𝑪−1 )𝑇 = (𝑪𝑇 )−1 ]
1 1
𝑔𝒀 (𝒚) = 𝑝 × exp (− 𝒚𝑇 𝒚) × √|𝚺| ; 𝒚 ∈ ℝ𝑝 , 𝚺 is 𝑝. 𝑑.
(2𝜋)2 √|𝚺| 2
𝑝
1 1 2
= 𝑝 ∙ exp (− ∑ 𝑦𝑖 )
(2𝜋)2 2
𝑖=1
𝑝
1 1
= ∏{ ∙ exp (− 𝑦𝑖2 )}
√2𝜋 2
𝑖=1
Hence, 𝒀 = 𝑪−1 (𝑿 − 𝝁) ~ 𝑵𝒑 (𝟎, 𝑰𝒑 ) i.e.,𝑌𝑖′ s are independent and identically distributed with 𝑌𝑖 ~ 𝑁1 (0, 1); 𝑖 = 1(1)𝑝.
Suppose 𝒀𝑝×1 ~ 𝑵𝒑 (𝟎, 𝑰𝒑 ) i. e. , 𝑌𝑖 ′s are iid 𝑁(0, 1), where 𝑪𝑝×𝑝 is non-singular matrix such that 𝑪𝑪𝑇 = 𝚺.
𝑇
The density function (p.d.f.) of 𝒀𝑝×1 = (𝑌1 , 𝑌2 , ⋯ , 𝑌𝑝 ) ~ 𝑵𝒑 (𝟎, 𝑰𝒑 ) is given by
𝑝 𝑝
1 1 𝑇 1 1 2
1 1
𝑓𝒀 (𝒚) = 𝑝 × exp {− 𝒚 𝑦} = 𝑝 ∙ exp (− ∑ 𝑦𝑖 ) = ∏ { ∙ exp (− 𝑦𝑖2 )} ; 𝒚 ∈ ℝ𝑝
(2𝜋) 2 2 (2𝜋) 2 2 √2𝜋 2
𝑖=1 𝑖=1
Since 𝚺 is a positive definite, there exist a non-singular matrix 𝑪 such that 𝑪𝑪𝑇 = 𝚺.
Let us consider the transformation 𝑿 = 𝑪𝒀 + 𝝁, where 𝑪 is a non-singular matrix such that 𝑪𝑪𝑇 = 𝚺.
𝒀 1 1 1
|𝐽| = |𝐽 ( )| = ||𝑪−1 || = = = .
𝑿 ||𝑪|| √|𝑪𝑪 | √|𝚺|
𝑇
𝑇
∴ 𝒀𝑇 𝒀 = (𝑪−1 (𝑿 − 𝝁)) (𝑪−1 (𝑿 − 𝝁)) = (𝑿 − 𝝁)𝑇 (𝑪𝑪𝑇 )−1 (𝑿 − 𝝁) = (𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) [∵ (𝑪−1 )𝑇 = (𝑪𝑇 )−1 ].
1 1 1
𝑓𝑿 (𝒙) = 𝑝 × exp {− (𝒙 − 𝝁)𝑇 𝚺 −1 (𝒙 − 𝝁)} × ; 𝒙 ∈ ℝ𝑝 , 𝝁 ∈ ℝ𝑝 , 𝚺 is 𝑝. 𝑑.
(2𝜋)2 2 √|𝚺|
1 1
= 𝑝 × exp {− (𝒙 − 𝝁)𝑇 𝚺 −1 (𝒙 − 𝝁)}
(2𝜋)2 √|𝚺| 2
Hence, 𝑿𝑝×1 = 𝑪𝒀 + 𝝁 ~ 𝑵𝒑 (𝝁, 𝚺), where 𝒀 ~ 𝑵𝒑 (𝟎, 𝑰𝒑 ) i. e. , 𝑌𝑖 ′s are iid 𝑁(0, 1) and 𝑪𝑝×𝑝 is non-singular matrix such
that 𝑪𝑪𝑇 = 𝚺.
OBSERVATION 1:
Suppose 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺). If 𝚺 = 𝐃𝐢𝐚𝐠(𝜎𝑖𝑖 ; 𝑖 = 1(1)𝑝) = 𝐃𝐢𝐚𝐠(𝜎11 , 𝜎22 , ⋯ , 𝜎𝑝𝑝 ), then 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 are all
mutually independent and 𝑋𝑖 ~ 𝑁1 (𝜇𝑖 , 𝜎𝑖𝑖 ); 𝑖 = 1(1)𝑝.
PROOF:
𝑇
Suppose 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺). The density function (p.d.f.) of 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) is given by
1 1
𝑓𝑿 (𝒙) = 𝑝 × exp {− (𝒙 − 𝝁)𝑇 𝚺 −1 (𝒙 − 𝝁)} ; 𝒙 ∈ ℝ𝑝 , 𝝁 ∈ ℝ𝑝 , 𝚺 is 𝑝. 𝑑.
(2𝜋)2 √|𝚺| 2
𝜎11 0 ⋯ 0
0 𝜎22 ⋯ 0 0 if 𝑖 ≠ 𝑗
Since 𝚺 = 𝐃𝐢𝐚𝐠(𝜎11 , 𝜎22 , ⋯ , 𝜎𝑝𝑝 ) = ( ⋱ ⋮ ), where 𝜎𝑖𝑗 = 𝐶𝑜𝑣(𝑋𝑖 , 𝑋𝑗 ) = { 𝜎𝑖𝑖 if 𝑖 = 𝑗, 𝑖, 𝑗 = 1(1)𝑝.
⋮ ⋮
0 0 ⋯ 𝜎𝑝𝑝
𝑝
−1
1 1 1
∴ 𝚺 = 𝐃𝐢𝐚𝐠 ( , ,⋯, ) and |𝚺| = ∏ 𝜎𝑖𝑖 .
𝜎11 𝜎22 𝜎𝑝𝑝
𝑖=1
𝑝
𝑇 −1 (𝑿
1
Now, (𝑿 − 𝝁) 𝚺 − 𝝁) = ∑ (𝑋 − 𝜇𝑖 )2 .
𝜎𝑖𝑖 𝑖
𝑖=1
𝑇
The density function (p.d.f.) of 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺) is given by
𝑝
1 1 1
𝑓𝑿 (𝒙) = × exp {− ∑ (𝑋𝑖 − 𝜇𝑖 )2 }
𝑝 𝑝 2 𝜎𝑖𝑖
(√2𝜋) √∏𝑖=1 𝜎𝑖𝑖 𝑖=1
𝑝 𝑝
1 1
= ∏[ × exp {− (𝑋𝑖 − 𝜇𝑖 )2 }] = ∏{𝑓𝑋𝑖 (𝑥𝑖 )} ;
√2𝜋√𝜎𝑖𝑖 2 𝜎 𝑖𝑖
𝑖=1 𝑖=1
OBSERVATION 2:
If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝜎 2 𝑰𝒑 ), then 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 are all mutually independent and 𝑋𝑖 ~ 𝑁1 (𝜇𝑖 , 𝜎 2 ); 𝑖 = 1(1)𝑝.
PROOF:
0 if 𝑖 ≠ 𝑗
Same as OBSERVATION 1, where 𝜎𝑖𝑗 = 𝐶𝑜𝑣(𝑋𝑖 , 𝑋𝑗 ) = { 𝑖, 𝑗 = 1(1)𝑝.
𝜎𝑖𝑖 = 𝜎 2 if 𝑖 = 𝑗,
𝑇
The M.G.F. of 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺), for a set of arbitrary real vector 𝒕𝑝×1 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 ) is given by
1 1
= 𝑝 × ∫ exp(𝒕𝑇 𝒙) ∙ exp {− (𝒙 − 𝝁)𝑇 𝚺 −1 (𝒙 − 𝝁)} 𝑑𝒙
(2𝜋)2 √|𝚺| 2
ℝ𝑝
1 1
= 𝑝 × ∫ exp {− (𝒙𝑇 𝚺 −1 𝒙 − 2𝒙𝑇 𝚺 −1 𝝁 + 𝝁𝑇 𝚺 −1 𝝁 − 2𝒕𝑇 𝒙)} 𝑑𝒙 [∵ 𝒙𝑇 𝚺 −1 𝝁 = 𝝁𝑇 𝚺 −1 𝒙]
(2𝜋)2 √|𝚺| 2
ℝ𝑝
1 1
= 𝑝 × ∫ exp [− {𝒙𝑇 𝚺 −1 𝒙 − 2𝒙𝑇 𝚺 −1 (𝝁 + 𝚺𝒕) + 𝝁𝑇 𝚺 −1 𝝁}] 𝑑𝒙 [∵ 𝒕𝑇 𝒙 = 𝒙𝑇 𝒕]
(2𝜋)2 √|𝚺| 2
ℝ𝑝
1 1
= 𝑝 × ∫ exp [− {𝒙𝑇 𝚺 −1 𝒙 − 2𝒙𝑇 𝚺 −1 (𝝁 + 𝚺𝒕) + (𝝁 + 𝚺𝒕)𝑇 𝚺 −1 (𝝁 + 𝚺𝒕)}]
(2𝜋)2 √|𝚺| 2
ℝ𝑝
1
× exp [− {𝝁𝑇 𝚺 −1 𝝁 − (𝝁 + 𝚺𝒕)𝑇 𝚺 −1 (𝝁 + 𝚺𝒕)}] 𝑑𝒙
2
1
= exp [− (𝝁𝑇 𝚺 −1 𝝁 − 𝝁𝑇 𝚺 −1 𝝁 − 𝝁𝑇 𝚺 −1 𝚺𝒕 − 𝒕𝑇 𝚺 𝑇 𝚺 −1 𝝁 − 𝒕𝑇 𝚺 𝑇 𝚺 −1 𝚺𝒕)]
2
1 1 𝑇
× 𝑝 ∫ exp [− (𝒙 − (𝝁 + 𝚺𝒕)) 𝚺 −1 (𝒙 − (𝝁 + 𝚺𝒕))] 𝑑𝒙
(2𝜋)2 √|𝚺| ℝ𝑝 2
1
= exp (𝒕𝑇 𝝁 + 𝒕𝑇 𝚺𝒕) × ∫ 𝑓𝑿 {𝒙|𝑿 ~ 𝑵𝒑 ((𝝁 + 𝚺𝒕), 𝚺)}𝑑𝒙
2
ℝ𝑝
1
= exp (𝒕𝑇 𝝁 + 𝒕𝑇 𝚺𝒕)
2
𝑇 𝝁+1𝒕𝑇 𝚺𝒕
= 𝑒𝒕 2 .
NOTE:
1 1 𝑇
1) The moment generating function for (𝑿 − 𝝁) is 𝑴𝑿−𝝁 (𝒕) = exp ( 𝒕𝑇 𝚺𝒕) = 𝑒 2𝒕 𝚺𝒕
.
2
(i) If two random vectors have the same moment generating function, they have the same density.
(ii) Random vectors are independent if and only if their joint moment generating function factors into the
𝑇
product of their separate moment generating functions; that is, if 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) and any real
𝑇
𝒕𝑝×1 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 ) , then 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 are independent if and only if
𝑇
The C.F. of 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺), for a set of arbitrary real number 𝒕𝑝×1 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 ) is given by
1
𝝓𝑿 (𝒕) = 𝐸{exp(𝑖𝒕𝑇 𝑿)} = exp (𝑖𝒕𝑇 𝝁 − 𝒕𝑇 𝚺𝒕) .
2
RESULT 2:
If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) and 𝒀𝑝×1 = 𝑪𝑿 + 𝒃 where 𝑪𝑝×𝑝 is non-singular matrix and 𝒃 is 𝑝 × 1 vector, then
𝒀 ~ 𝑵𝑝 (𝑪𝝁 + 𝒃, 𝑪𝚺𝑪𝑇 ).
PROOF:
𝑇
Suppose 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺). The density function (p.d.f.) of 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) is given by
1 1
𝑓𝑿 (𝒙) = 𝑝 × exp {− (𝒙 − 𝝁)𝑇 𝚺 −1 (𝒙 − 𝝁)} ; 𝒙 ∈ ℝ𝑝 , 𝝁 ∈ ℝ𝑝 , 𝚺 is 𝑝. 𝑑.
(2𝜋)2 √|𝚺| 2
Let us consider the transformation 𝒀𝑝×1 = 𝑪𝑝×𝑝 𝑿𝑝×1 + 𝒃𝑝×1 i.e., 𝑿 = 𝑪−1 (𝒀 − 𝒃).
Since 𝚺 is a positive definite and 𝑪𝑝×𝑝 is non-singular matrix, 𝑪𝚺𝑪𝑇 is also positive definite.
𝑿 1 1 |𝚺| √|𝚺|
|𝐽| = |𝐽 ( )| = ||𝑪−1 || = = =√ = .
𝒀 ||𝑪|| √|𝑪||𝑪𝑇 | |𝑪| ∙ |𝚺| ∙ |𝑪𝑇 | √|𝑪𝚺𝑪𝑇 |
𝑇
= (𝑪−1 (𝒀 − (𝑪𝝁 + 𝒃))) 𝚺 −1 (𝑪−1 (𝒀 − (𝑪𝝁 + 𝒃)))
𝑇
= (𝒀 − (𝑪𝝁 + 𝒃)) (𝑪𝚺𝑪𝑇 )−1 (𝒀 − (𝑪𝝁 + 𝒃)) [∵ (𝑪−1 )𝑇 = (𝑪𝑇 )−1 ]
1 1 𝑇 √|𝚺|
𝑔𝒀 (𝒚) = 𝑝 ∙ exp {− (𝒚 − (𝑪𝝁 + 𝒃)) (𝑪𝚺𝑪𝑇 )−1 (𝒚 − (𝑪𝝁 + 𝒃))} ∙ ; 𝒚 ∈ ℝ𝑝 , 𝑪𝚺𝑪𝑇 is 𝑝. 𝑑.
(2𝜋)2 √|𝚺| 2 √|𝑪𝚺𝑪𝑇 |
1 1 𝑇
= 𝑝 ∙ exp {− (𝒚 − (𝑪𝝁 + 𝒃)) (𝑪𝚺𝑪𝑇 )−1 (𝒚 − (𝑪𝝁 + 𝒃))}
(2𝜋)2 √|𝑪𝚺𝑪𝑇 | 2
ALTERNATIVE
Let 𝒀𝑝×1 = 𝑪𝑿 + 𝒃.
𝑇
For any real 𝒕𝑝×1 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 ) , the moment generating function (M.G.F.) of 𝒀 is given by
1 1
= exp(𝒕𝑇 𝒃) ∙ exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺) ; [∵ If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺), 𝑴𝑿 (𝑺) = exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺)]
2 2
1
= exp (𝒕𝑇 (𝑪𝝁 + 𝒃) + 𝒕𝑇 (𝑪𝚺𝑪𝑇 )𝒕)
2
Since 𝚺 is a positive definite and 𝑪𝑝×𝑝 is non-singular matrix, 𝑪𝚺𝑪𝑇 is also positive definite.
By uniqueness property of moment generating function (M.G.F.), 𝑴𝒀 (𝒕) is the M.G.F. of 𝑵𝑝 (𝑪𝝁 + 𝒃, 𝑪𝚺𝑪𝑇 ) i.e., the
distribution of 𝒀𝑝×1 = 𝑪𝑿 + 𝒃 is 𝑵𝑝 (𝑪𝝁 + 𝒃, 𝑪𝚺𝑪𝑇 ).
OBSERVATION 3:
If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) and 𝒀𝑝×1 = 𝑪𝑿 where 𝑪𝑝×𝑝 is non-singular matrix, then 𝒀 ~ 𝑵𝑝 (𝑪𝝁, 𝑪𝚺𝑪𝑇 ).
PROOF:
OBSERVATION 4:
If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝜎 2 𝑰𝒑 ) and 𝒀𝑝×1 = 𝑳𝑿 where 𝑳𝑝×𝑝 is an orthogonal matrix, then 𝒀 ~ 𝑵𝑝 (𝑳𝝁, 𝜎 2 𝑰𝒑 ).
PROOF:
𝑇
For any real 𝒕𝑝×1 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 ) , the M.G.F. of 𝒀 is given by
= 𝐸{exp(𝒕𝑇 𝑳𝑿)}
1 1
= exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺) ; [∵ If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺), 𝑴𝑿 (𝑺) = exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺)]
2 2
1
= exp (𝒕𝑇 𝑳𝝁 + 𝒕𝑇 (𝑳𝚺𝑳𝑇 )𝒕)
2
By uniqueness property of moment generating function (M.G.F.), 𝑴𝒀 (𝒕) is the M.G.F. of 𝑵𝑝 (𝑳𝝁, 𝑳𝚺𝑳𝑇 ) i.e., the
distribution of 𝒀𝑝×1 = 𝑳𝑿 is 𝑵𝑝 (𝑳𝝁, 𝜎 2 𝑰𝒑 ).
NOTE:
1. From OBSERVATION 2, we know that, if 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝜎 2 𝑰𝒑 ), then 𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 are all mutually independent and
𝑋𝑖 ~ 𝑁1 (𝜇𝑖 , 𝜎 2 ); 𝑖 = 1(1)𝑝. Since 𝒀𝑝×1 = 𝑳𝑿 ~ 𝑵𝑝 (𝑳𝝁, 𝜎 2 𝑰𝒑 ), 𝑌1 , 𝑌2 , ⋯ , 𝑌𝑝 are all mutually independent with
common variance 𝜎 2 .
The OBSERVATION 4 states that mutually independent normal variables with the same variance remain mutually
independent normal variables with the same variance under orthogonal transformation. This invariance property
is a characterisation of normal distribution.
2. If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝜎 2 𝑰𝒑 ) and 𝒀𝑝×1 = 𝑳𝑇 𝑿 where 𝑳𝑝×𝑝 is an orthogonal matrix, then 𝒀 ~ 𝑵𝑝 (𝑳𝑇 𝝁, 𝜎 2 𝑰𝒑 ).
RESULT 3:
(1)
𝑇 𝑿𝑚×1
Suppose 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) = ( (2) ) ~ 𝑵𝒑 (𝝁, 𝚺),
𝑿(𝑝−𝑚)×1
(1)
The marginal distribution of any subset 𝑿𝑚×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑚 )𝑇 ; (𝑚 < 𝑝) of 𝑿 is 𝑚-dimensional multivariate
(2) 𝑇
normal and the marginal distribution of subset 𝑿(𝑝−𝑚)×1 = (𝑋𝑚+1 , 𝑋𝑚+2 , ⋯ , 𝑋𝑝 ) of 𝑿 is (𝑝 − 𝑚)-dimensional
multivariate normal. The marginal distributions of 𝑿(1) and 𝑿(2) , respectively, are given by
The conditional distribution of 𝑿(1) given for fixed 𝑿(2) = 𝒙(2) is given by
The conditional distribution of 𝑿(2) given for fixed 𝑿(1) = 𝒙(1) is given by
PROOF:
𝑇
Suppose 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺). The density function (p.d.f.) of 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) is given by
1 1
𝑓𝑿 (𝒙) = 𝑝 × exp {− (𝒙 − 𝝁)𝑇 𝚺 −1 (𝒙 − 𝝁)} ; 𝒙 ∈ ℝ𝑝 , 𝝁 ∈ ℝ𝑝 , 𝚺 is 𝑝. 𝑑.
(2𝜋)2 √|𝚺| 2
Let us define
𝑇
𝐸 {(𝑿(1) − 𝝁(1) )(𝑿(1) − 𝝁(1) ) } = 𝚺11 ,
𝑇
𝐸 {(𝑿 (1) − 𝝁(1) )(𝑿(2) − 𝝁(2) ) } = 𝚺12 , and
𝑇
𝐸 {(𝑿(2) − 𝝁(2) )(𝑿(2) − 𝝁(2) ) } = 𝚺22 .
(2) (2)
𝒀(𝑝−𝑚)×1 = 𝑿(𝑝−𝑚)×1 , where 𝑩𝑚×(𝑝−𝑚) is such that 𝐶𝑜𝑣(𝒀(1) , 𝒀(2) ) = 𝟎.
𝑇
Now, 𝐶𝑜𝑣(𝒀(1) , 𝒀(2) ) = 0 ⟹ 𝐸 [(𝒀(1) − 𝐸(𝒀(1) )) (𝒀(2) − 𝐸(𝒀(2) )) ] = 𝟎
𝑇
⟹ 𝐸 [{(𝑿(1) − 𝝁(1) ) + 𝑩(𝑿(2) − 𝝁(2) )}(𝑿(2) − 𝝁(2) ) ] = 𝟎
𝑇 𝑇
⟹ 𝐸 {(𝑿(1) − 𝝁(1) )(𝑿(2) − 𝝁(2) ) } + 𝑩 𝐸 {(𝑿(2) − 𝝁(2) )(𝑿(2) − 𝝁(2) ) } = 𝟎
−1
⟹ 𝚺12 + 𝑩 𝚺22 = 𝟎 ⟹ 𝑩 = −𝚺12 𝚺22 .
(1)
𝒀𝑚×1 𝑰 −1
−𝚺12 𝚺22
−1
𝒀𝑝×1 = 𝑷𝑝×𝑝 𝑿𝑝×1 i. e., 𝑿 = 𝑷 𝒀, where 𝒀𝑝×1 = ( (2) ) and 𝑷𝑝×𝑝 = ( 𝑚 ).
𝒀(𝑝−𝑚)×1 𝟎 (𝑝−𝑚)×𝑚 𝑰(𝑝−𝑚)
𝑰𝑚 −1
−𝚺12 𝚺22 𝝁(1)
Now, 𝑷𝝁 = ( ) ( (2) )
𝟎(𝑝−𝑚)×𝑚 𝑰(𝑝−𝑚) 𝝁
−1 −1 𝑇
𝑰 −𝚺12 𝚺22 𝚺 𝚺12 𝑰𝑚 −𝚺12 𝚺22
and 𝑷𝚺𝑷𝑇 = ( 𝑚 ) ( 11 )( )
𝟎(𝑝−𝑚)×𝑚 𝑰(𝑝−𝑚) 𝚺21 𝚺22 𝟎(𝑝−𝑚)×𝑚 𝑰(𝑝−𝑚)
−1
𝚺 − 𝚺12 𝚺22 𝚺21 𝟎𝑚×(𝑝−𝑚) 𝑰𝑚 𝟎𝑚×(𝑝−𝑚)
= ( 11 ) ( −1 )
𝚺21 𝚺22 −𝚺22 𝚺21 𝑰(𝑝−𝑚)
−1
𝚺11 − 𝚺12 𝚺22 𝚺21 𝟎𝑚×(𝑝−𝑚)
=( )
𝟎(𝑝−𝑚)×𝑚 𝚺22
𝚺11.2 𝟎𝑚×(𝑝−𝑚) −1
=( ) ; (say), where 𝚺11.2 = 𝚺11 − 𝚺12 𝚺22 𝚺21 .
𝟎(𝑝−𝑚)×𝑚 𝚺22
𝑿 1
|𝐽| = |𝐽 ( )| = ||𝑷−1 || = =1.
𝒀 ||𝑷||
𝚺 𝟎 −1 −1
𝚺11.2 𝟎
∴ (𝑷𝚺𝑷𝑇 )−1 = ( 11.2 ) ⟺ (𝑷𝑇 )−1 𝚺 −1 𝑷−1 = ( −1 ).
𝟎 𝚺22 𝟎 𝚺22
𝚺 𝟎
and |𝑷𝚺𝑷𝑇 | = | 11.2 | ⟹ |𝑷||𝚺||𝑷𝑇 | = |𝚺11.2 ||𝚺22 |
𝟎 𝚺22
𝑇
= (𝑷−1 (𝒀 − 𝑷𝝁)) 𝚺 −1 (𝑷−1 (𝒀 − 𝑷𝝁))
−1𝑇
𝒀(1) − 𝜽(1) 𝚺11.2 𝟎 𝒀(1) − 𝜽(1)
=( ) ( −1 ) ( (2) )
𝒀(2) − 𝝁(2) 𝟎 𝚺22 𝒀 − 𝝁(2)
𝑇 𝑇
= (𝒀(1) − 𝜽(1) ) 𝚺11.2
−1
(𝒀(1) − 𝜽(1) ) + (𝒀(2) − 𝝁(2) ) 𝚺22
−1
(𝒀(2) − 𝝁(2) ).
𝑇
The density function (p.d.f.) of 𝒀𝑝×1 = (𝒀(1) , 𝒀(2) ) is given by
1 1 (1) 𝑇 𝑇
𝑔𝒀 (𝒚) = 𝑝 exp [− {(𝒚 − 𝜽(1) ) 𝚺11.2
−1
(𝒚(1) − 𝜽(1) ) + (𝒚(2) − 𝝁(2) ) 𝚺22
−1
(𝒚(2) − 𝝁(2) )}] ;
(2𝜋)2 √|𝚺11.2 ||𝚺22 | 2
1 1 (1) 𝑇
=( 𝑚 exp {− (𝒚 − 𝜽(1) ) 𝚺11.2
−1
(𝒚(1) − 𝜽(1) )})
(2𝜋) 2 √|𝚺11.2 | 2
1 1 (2) 𝑇
×( 𝑝−𝑚 exp {− (𝒚 − 𝝁(2) ) 𝚺22
−1
(𝒚(2) − 𝝁(2) )})
(2𝜋) 2 √|𝚺22 | 2
ALTERNATIVE
𝑇
For any real 𝒕𝑝×1 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 ) , the M.G.F. of 𝒀 is given by
1 1
= exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺) ; [∵ If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺), 𝑴𝑿 (𝑺) = exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺)]
2 2
1
= exp (𝒕𝑇 𝑷𝝁 + 𝒕𝑇 (𝑷𝚺𝑷𝑇 )𝒕) = 𝑀. 𝐺. 𝐹. of 𝑵𝑝 (𝑷𝝁, 𝑷𝚺𝑷𝑇 )
2
By uniqueness property of moment generating function (M.G.F.), 𝑴𝒀 (𝒕) is the M.G.F. of 𝑵𝑝 (𝑷𝝁, 𝑷𝚺𝑷𝑇 ) i.e., the
distribution of 𝒀𝑝×1 is 𝑵𝑝 (𝑷𝝁, 𝑷𝚺𝑷𝑇 ). Hence, 𝒀𝑝×1 ~ 𝑵𝑝 (𝑷𝝁, 𝑷𝚺𝑷𝑇 ).
1
= exp {𝒕𝑇 (𝝁(1) − 𝚺12 𝚺22
−1 (2)
𝝁 ) + 𝒕𝑇 (𝚺11 − 𝚺12 𝚺22
−1
𝚺21 )𝒕}
2
−1 (2)
1
× exp {𝒕𝑇 𝚺12 𝚺22 𝝁 + 𝒕𝑇 𝚺12 𝚺22
−1 −1
𝚺22 𝚺22 𝚺21 𝒕}
2
1 𝑇
= exp (𝒕𝑇 𝝁(1) + 𝒕 𝚺11 𝒕) = 𝑀. 𝐺. 𝐹. of 𝑵𝒎 (𝝁(1) , 𝚺11 )
2
By uniqueness property of moment generating function (M.G.F.), 𝑴𝑿(1) (𝒕) is the M.G.F. of 𝑵𝒎 (𝝁(1) , 𝚺11 ) i.e., the
distribution of 𝑿(1) is 𝑵𝒎 (𝝁(1) , 𝚺11 ). Hence, 𝑿(1) ~ 𝑵𝒎 (𝝁(1) , 𝚺11 ).
The marginal density function (p.d.f.) of 𝑿(2) ~ 𝑵(𝒑−𝒎) (𝝁(2) , 𝚺22 ) is given by
1 1 𝑇 −1
𝑓𝑿(2) (𝒙(2) ) = 𝑝−𝑚 × exp {− (𝒙(2) − 𝝁(2) ) 𝚺22 (𝒙(2) − 𝝁(2) )}.
(2𝜋) 2 √|𝚺22 | 2
1 1 𝑇 −1
𝑓𝑿 (𝒙(1) , 𝒙(2) ) = 𝑚 𝑝−𝑚 × exp {− (𝒙(2) − 𝝁(2) ) 𝚺22 (𝒙(2) − 𝝁(2) )}
(2𝜋) 2 (2𝜋) 2 √|𝚺11.2 |√|𝚺22 | 2
1 𝑇 −1
× exp {− (𝒙(1) − {𝝁(1) + 𝚺12 𝚺22
−1
(𝒙(2) − 𝝁(2) )}) 𝚺11.2 (𝒙(1) − {𝝁(1) + 𝚺12 𝚺22
−1
(𝒙(2) − 𝝁(2) )})}.
2
−1
where 𝚺11.2 = 𝚺11 − 𝚺12 𝚺22 𝚺21 .
The conditional density function (p.d.f.) of 𝑿(1) given for fixed 𝑿(2) = 𝒙(2) is given by
(1)
1 1 𝑇 −1
𝑓 (1) (2) (𝒙 )= 𝑚 × exp {− (𝒙(1) − 𝝀(1) ) 𝚺11.2 (𝒙(1) − 𝝀(1) )},
( 𝑿 |𝒙 ) 2
(2𝜋) 2 √|𝚺11.2 |
The conditional distribution of 𝑿(1) given for fixed 𝑿(2) = 𝒙(2) is given by
1 1 𝑇 −1
𝑓𝑿(1) (𝒙(1) ) = 𝑚 × exp {− (𝒙(1) − 𝝁(1) ) 𝚺11 (𝒙(1) − 𝝁(1) )}.
(2𝜋) 2 √|𝚺11 | 2
Similarly, the joint density function (p.d.f.) of (𝑿(1) , 𝑿(2) ) can be written as
1 1 𝑇 −1
𝑓𝑿 (𝒙(1) , 𝒙(2) ) = 𝑚 𝑝−𝑚 × exp {− (𝒙(1) − 𝝁(1) ) 𝚺11 (𝒙(1) − 𝝁(1) )}
(2𝜋) 2 (2𝜋) 2 √|𝚺11 |√|𝚺22.1 | 2
1 𝑇 −1
× exp {− (𝒙(1) − {𝝁(2) + 𝚺21 𝚺11
−1
(𝒙(1) − 𝝁(1) )}) 𝚺22.1 (𝒙(1) − {𝝁(2) + 𝚺21 𝚺11
−1
(𝒙(1) − 𝝁(1) )})}.
2
−1
where 𝚺22.1 = 𝚺22 − 𝚺21 𝚺11 𝚺12 .
The conditional density function (p.d.f.) of 𝑿(2) given for fixed 𝑿(1) = 𝒙(1) is given by
(2)
1 1 𝑇 −1
𝑓 (2) (1) (𝒙 )= 𝑝−𝑚 × exp {− (𝒙(2) − 𝝀(2) ) 𝚺22.1 (𝒙(2) − 𝝀(2) )},
( 𝑿 |𝒙 ) 2
(2𝜋) 2 √|𝚺22.1 |
The conditional distribution of 𝑿(2) given for fixed 𝑿(1) = 𝒙(1) is given by
NOTE:
The marginal distribution of each component of 𝑿 is univariate normal. Converse is not true in general. The
converse is true if the components of 𝑿 are all independent and normal.
EXAMPLE:
𝐹(𝑋1,𝑋2) (𝑥1 , 𝑥2 ) = Φ(𝑥1 )Φ(𝑥2 ){1 + 𝛼(1 − Φ(𝑥1 ))(1 − Φ(𝑥2 ))}
where |𝛼| < 1, Φ(𝑥) be the distribution function of standard normal distribution. It can be shown that the marginal
distribution of 𝑋1 and 𝑋2 are standard normal.
RESULT 4:
If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) and 𝑩𝑚×𝑝 is a 𝑚 × 𝑝 (𝑚 ≤ 𝑝) matrix with 𝑅𝑎𝑛𝑘(𝑩) = 𝑚 and 𝒅 is 𝑚 × 1 vector, then the 𝑚
linear combinations 𝑩𝑿 ~ 𝑵𝒎 (𝑩𝝁 + 𝒅, 𝑩𝚺𝑩𝑇 ).
PROOF:
(1)
𝒀𝑚×1 𝑩𝑚×𝑝 𝒅𝑚×1
Let us consider the transformation 𝒀𝑝×1 = ( ) =( )𝑿 + (𝟎 ) or, 𝒀 = 𝑷𝑿 + 𝒃,
(2)
𝒀(𝑝−𝑚)×1 𝑪(𝑝−𝑚)×𝑝 𝑝×1 (𝑝−𝑚)×1
From RESULT 2, we know that, if 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) and 𝒀𝑝×1 = 𝑷𝑿 + 𝒃 where 𝑷𝑝×𝑝 is non-singular matrix and 𝒃 is 𝑝 ×
1 vector, then 𝒀 ~ 𝑵𝑝 (𝑷𝝁 + 𝒃, 𝑷𝚺𝑷𝑇 ).
(1)
𝒀𝑚×1 𝑩𝝁 + 𝒅 𝑇
𝑩𝚺𝑪𝑇 ))
Hence, 𝒀𝑝×1 = ( ) ~ 𝑵𝑝 (( ) , (𝑩𝚺𝑩𝑇
(2)
𝒀(𝑝−𝑚)×1 𝑪𝝁 𝑪𝚺𝑩 𝑪𝚺𝑪𝑇
(1)
From RESULT 3, we know that, the marginal distribution of any subset 𝒀𝑚×1 = (𝑌1 , 𝑌2 , ⋯ , 𝑌𝑚 )𝑇 ; (𝑚 < 𝑝) of 𝒀 is 𝑚-
(2) 𝑇
dimensional multivariate normal and the marginal distribution of subset 𝒀(𝑝−𝑚)×1 = (𝑌𝑚+1 , 𝑌𝑚+2 , ⋯ , 𝑌𝑝 ) of 𝒀 is
(𝑝 − 𝑚)-dimensional multivariate normal.
ALTERNATIVE
Let 𝒀𝑚×1 = 𝑩𝑚×𝑝 𝑿𝑝×1 + 𝒅𝑚×1 . For any real 𝒕𝑚×1 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑚 )𝑇 , the moment generating function (M.G.F.) of 𝒀 is
given by
1 1
= exp(𝒕𝑇 𝒅) ∙ exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺) ; [∵ If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺), 𝑴𝑿 (𝑺) = exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺)]
2 2
1
= exp (𝒕𝑇 (𝑩𝝁 + 𝒅) + 𝒕𝑇 (𝑩𝚺𝑩𝑇 )𝒕)
2
Since 𝚺 is a positive definite and 𝑩𝑚×𝑝 is a 𝑚 × 𝑝 (𝑚 ≤ 𝑝) matrix with 𝑎𝑛𝑘(𝑩) = 𝑚 , 𝑩𝚺𝑩𝑇 is also positive definite.
By uniqueness property of moment generating function (M.G.F.), 𝑴𝒀 (𝒕) is the M.G.F. of 𝑵𝑚 (𝑩𝝁 + 𝒅, 𝑩𝚺𝑩𝑇 ) i.e., the
distribution of 𝒀𝑚×1 = 𝑩𝑿 + 𝒅 is 𝑵𝑚 (𝑩𝝁 + 𝒅, 𝑩𝚺𝑩𝑇 ). Hence, 𝒀𝑚×1 = 𝑩𝑿 + 𝒅 ~ 𝑵𝑚 (𝑩𝝁 + 𝒅, 𝑩𝚺𝑩𝑇 ).
OBSERVATION 5:
If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) and 𝑩𝑚×𝑝 is a 𝑚 × 𝑝 (𝑚 ≤ 𝑝) matrix with 𝑅𝑎𝑛𝑘(𝑩) = 𝑚, then 𝑩𝑿 ~ 𝑵𝒎 (𝑩𝝁, 𝑩𝚺𝑩𝑇 ).
PROOF:
ALTERNATIVE
Let 𝒀𝑚×1 = 𝑩𝑚×𝑝 𝑿𝑝×1 . For any real 𝒕𝑚×1 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑚 )𝑇 , the moment generating function (M.G.F.) of 𝒀 is given by
1 1
= exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺) ; [∵ If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺), 𝑴𝑿 (𝑺) = exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺)]
2 2
1
= exp (𝒕𝑇 𝑩𝝁 + 𝒕𝑇 (𝑩𝚺𝑩𝑇 )𝒕)
2
Since 𝚺 is a positive definite and 𝑩𝑚×𝑝 is a 𝑚 × 𝑝 (𝑚 ≤ 𝑝) matrix with 𝑎𝑛𝑘(𝑩) = 𝑚 , 𝑩𝚺𝑩𝑇 is also positive definite.
By uniqueness property of moment generating function (M.G.F.), 𝑴𝒀 (𝒕) is the M.G.F. of 𝑵𝑚 (𝑩𝝁, 𝑩𝚺𝑩𝑇 ) i.e., the
distribution of 𝒀𝑚×1 = 𝑩𝑿 is 𝑵𝑚 (𝑩𝝁, 𝑩𝚺𝑩𝑇 ). Hence, 𝒀𝑚×1 = 𝑩𝑿 ~ 𝑵𝑚 (𝑩𝝁, 𝑩𝚺𝑩𝑇 ).
OBSERVATION 6:
If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) and 𝑪𝑝×𝑚 is a 𝑚 × 𝑝 (𝑚 ≤ 𝑝) matrix with 𝑅𝑎𝑛𝑘(𝑩) = 𝑚, then 𝑪𝑇 𝑿 ~ 𝑵𝒎 (𝑪𝑇 𝝁, 𝑪𝑇 𝚺𝑪).
PROOF:
RESULT 5:
𝑇
The random vector 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) follows 𝑝-dimensional multivariate normal distribution if and only
𝑝 𝑇
if any linear combination 𝒍𝑇 𝑿 = ∑𝑖=1 𝑙𝑖 𝑋𝑖 follows univariate normal distribution, where 𝒍𝑝×1 = (𝑙1 , 𝑙2 , ⋯ , 𝑙𝑝 ) , i.e., the
𝑝
random vector 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) if and only if 𝒍𝑇 𝑿 = ∑𝑖=1 𝑙𝑖 𝑋𝑖 ~ 𝑁1 (𝒍𝑇 𝝁, 𝒍𝑇 𝚺𝒍).
PROOF:
𝑝 𝑇
Suppose 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) and 𝑌 = 𝒍𝑇 𝑿 = ∑𝑖=1 𝑙𝑖 𝑋𝑖 , where 𝒍𝑝×1 = (𝑙1 , 𝑙2 , ⋯ , 𝑙𝑝 ) .
For any real scalar 𝑡 ∈ ℝ, the moment generating function (M.G.F.) of 𝑌 is given by
1 1
= exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺) ; [∵ If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺), 𝑴𝑿 (𝑺) = exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺)]
2 2
1
= exp (𝑡 𝒍𝑇 𝝁 + 𝑡 2 (𝒍𝑇 𝚺𝒍))
2
𝑝 𝑝 𝑝
1
= exp (𝑡 ∑ 𝑙𝑖 𝜇𝑖 + 𝑡 2 ∑ ∑ 𝑙𝑖 𝑙𝑗 𝜎𝑖𝑗 ) ⟶ 𝑀. 𝐺. 𝐹. of 𝑵1 (𝒍𝑇 𝝁, 𝒍𝑇 𝚺𝒍)
2
𝑖=1 𝑖=1 𝑗=1
By uniqueness property of moment generating function (M.G.F.), 𝑀𝑌 (𝑡) is the M.G.F. of 𝑁1 (𝒍𝑇 𝝁, 𝒍𝑇 𝚺𝒍) i.e., the
𝑝
distribution of 𝑌 = 𝒍𝑇 𝑿 = ∑𝑖=1 𝑙𝑖 𝑋𝑖 is 𝑵1 (𝒍𝑇 𝝁, 𝒍𝑇 𝚺𝒍).
𝑝 𝑝 𝑝 𝑝 𝑝
𝑇 (𝒍𝑇 𝑇 𝑇
Hence, 𝑌 = 𝒍 𝑿 = ∑ 𝑙𝑖 𝑋𝑖 ~ 𝑁1 𝝁, 𝒍 𝚺𝒍) i. e. , 𝑌 = 𝒍 𝑿 = ∑ 𝑙𝑖 𝑋𝑖 ~ 𝑁1 (∑ 𝑙𝑖 𝜇𝑖 , ∑ ∑ 𝑙𝑖 𝑙𝑗 𝜎𝑖𝑗 )
𝑖=1 𝑖=1 𝑖=1 𝑖=1 𝑗=1
𝑇
Suppose 𝑌 = 𝒍𝑇 𝑿 follows univariate normal distribution for any real 𝒍𝑝×1 = (𝑙1 , 𝑙2 , ⋯ , 𝑙𝑝 ) .
For any real scalar 𝑡 ∈ ℝ, the moment generating function (M.G.F.) of 𝑌 is given by
1
𝑀𝑌 (𝑡) = 𝐸{exp(𝑡 𝑌)} = exp (𝑡𝑏 + 𝑡 2 𝜎 2 ),
2
1
∴ 𝑀𝑌 (𝑡) = 𝐸{exp(𝑡 𝑌)} = 𝐸{exp(𝑡 𝒍𝑇 𝑿)} = exp (𝑡 𝒍𝑇 𝝁 + 𝑡 2 (𝒍𝑇 𝚺𝒍)).
2
1 𝑇
For 𝑡 = 1, 𝑀𝑌 (1) = 𝐸{exp(𝑌)} = exp (𝑡 𝒍𝑇 𝝁 + 𝑡 2 (𝒍𝑇 𝚺𝒍)) = 𝐸{exp(𝒍𝑇 𝑿)} = 𝑴𝑿 (𝒍), for real 𝒍 = (𝑙1 , 𝑙2 , ⋯ , 𝑙𝑝 ) ,
2
𝑇
which is the M.G.F. of 𝑵𝒑 (𝝁, 𝚺) for real 𝒍𝑝×1 = (𝑙1 , 𝑙2 , ⋯ , 𝑙𝑝 ) .
NOTE:
Suppose 𝑿 = (𝑋1 , 𝑋2 )𝑇 and the marginal distribution of 𝑋1 and 𝑋2 are univariate normal. The random vector 𝑿
follows a bivariate normal distribution iff 𝑙1 𝑋1 + 𝑙2 𝑋2 follows univariate normal for every real 𝑙1 and 𝑙2 .
OBSERVATION 7:
PROOF:
𝑇 1 if 𝑘 = 𝑖
Same as RESULT 5, where 𝒍𝑝×1 = (𝑙1 , 𝑙2 , ⋯ , 𝑙𝑝 ) = (0, 0, ⋯ 0, 1,0, ⋯ , 0)𝑇 i. e., 𝑙𝑘 = { ; 𝑘 = 1(1)𝑝.
0 if 𝑘 ≠ 𝑖
OBSERVATION 8:
𝑇 𝜇𝑖 𝜎𝑖𝑖 𝜎𝑖𝑗
If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺), then (𝑋𝑖 , 𝑋𝑗 ) ~ 𝑵𝟐 (𝜇 , (𝜎 𝜎𝑗𝑗 )) ; 𝑖, 𝑗 = 1(1)𝑝 (𝑖 ≠ 𝑗).
𝑗 𝑗𝑖
PROOF:
𝑇 1 if 𝑘 = 𝑖, 𝑗
Same as RESULT 5, where 𝒍𝑝×1 = (𝑙1 , 𝑙2 , ⋯ , 𝑙𝑝 ) = (0, ⋯ , 1, 0, ⋯ ,0, 1,0, ⋯ , 0)𝑇 i. e., 𝑙𝑘 = { ; 𝑘 = 1(1)𝑝.
0 if 𝑘 ≠ 𝑖, 𝑗
RESULT 6:
Then 𝑿(1) and 𝑿(2) are independently distributed iff 𝚺12 = 𝚺21
𝑇
= 𝟎𝑚×(𝑝−𝑚) [i. e., 𝐶𝑜𝑣(𝑿(1) , 𝑿(2) ) = 𝟎].
PROOF:
𝑇
Suppose 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺).
1 1
𝑓𝑿 (𝒙) = 𝑝 × exp {− (𝒙 − 𝝁)𝑇 𝚺 −1 (𝒙 − 𝝁)} ; 𝒙 ∈ ℝ𝑝 , 𝝁 ∈ ℝ𝑝 , 𝚺 is 𝑝. 𝑑.
(2𝜋)2 √|𝚺| 2
𝚺11 𝚺12 𝚺 𝟎
Now, 𝚺 = ( ) = ( 11 ).
𝚺21 𝚺22 𝟎 𝚺22
−1
𝚺11 𝟎
Hence, |𝚺| = |𝚺11 | ∙ |𝚺22 | and 𝚺 −1 = ( −1 ).
𝟎 𝚺22
𝑇
𝑿(1) − 𝝁(1) 𝚺 −1 𝟎 𝑿(1) − 𝝁(1)
∴ (𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) = ( (2) (2)
) ( 11 −1 ) ( (2) )
𝑿 −𝝁 𝟎 𝚺22 𝑿 − 𝝁(2)
𝑇 𝑇
= (𝑿(1) − 𝝁(1) ) 𝚺11
−1
(𝑿(1) − 𝝁(1) ) + (𝑿(2) − 𝝁(2) ) 𝚺22
−1
(𝑿(2) − 𝝁(2) ).
𝑇 𝑇
The density function (p.d.f.) of 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) = (𝑿(1) , 𝑿(2) ) becomes
1 1 𝑇 −1
𝑓𝑿 (𝒙) = [ 𝑚 × exp {− (𝑿(1) − 𝝁(1) ) 𝚺11 (𝑿(1) − 𝝁(1) )}]
(2𝜋) 2 √|𝚺11 | 2
1 1 𝑇 −1
×[ 𝑝−𝑚 × exp {− (𝒙(2) − 𝝁(2) ) 𝚺22 (𝒙(2) − 𝝁(2) )}] ;
(2𝜋) 2 √|𝚺22 | 2
(1) (2)
Since 𝑿(1) and 𝑿(2) are independently distributed, 𝑋𝑖 and 𝑋𝑗 are independent; 𝑖 = 1(1)𝑚, 𝑗 = 𝑚 + 1(1)𝑝.
∴ 𝑓𝑿 (𝒙) = 𝑓𝑿(1) (𝒙(1) ) ∙ 𝑓𝑿(2) (𝒙(2) ) = 𝑓𝑿(1) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑚 ) ∙ 𝑓𝑿(2) (𝑥𝑚+1 , 𝑥𝑚+2 , ⋯ , 𝑥𝑝 ).
∞ ∞ ∞
(1) (1) (2) (2)
= ∫ ∫ ⋯ ∫ (𝑥𝑖 − 𝜇𝑖 )(𝑥𝑗 − 𝜇𝑗 )𝑓𝑿 (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑝 ) 𝑑𝑥1 𝑑𝑥2 ⋯ 𝑑𝑥𝑝
−∞ −∞ −∞
(1) (1)
= { ∫(𝑥𝑖 − 𝜇𝑖 ) 𝑓𝑿(1) (𝑥1 , 𝑥2 , ⋯ , 𝑥𝑚 ) 𝑑𝑥1 𝑑𝑥2 ⋯ 𝑑𝑥𝑚 }
ℝ𝑚
(2) (2)
×{ ∫ (𝑥𝑗 − 𝜇𝑗 ) 𝑓𝑿(2) (𝑥𝑚+1 , 𝑥𝑚+2 , ⋯ , 𝑥𝑝 ) 𝑑𝑥𝑚+1 𝑑𝑥𝑚+2 ⋯ 𝑑𝑥𝑝 }
ℝ(𝑝−𝑚)
=0
𝑇
⟺ 𝚺12 = 𝚺21 = 𝟎𝑚×(𝑝−𝑚) .
OBSERVATION 9:
𝑇 𝑇
If 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) = (𝑿(1) , 𝑿(2) , ⋯ , 𝑿(𝑟) ) ~ 𝑵𝒑 (𝝁, 𝚺) and 𝐶𝑜𝑣(𝑿(𝑖) , 𝑿(𝑗) ) = 𝟎; 𝑖 ≠ 𝑗 (𝑖, 𝑗 = 1(1)𝑝),
′
then 𝑿(𝑖) s ( 𝑖 = 1(1)𝑝) are mutually independent and not just pairwise independent.
OBSERVATION 10:
𝑇
If 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝜎 2 𝑰𝒑 ), then 𝒂𝑇 𝑿 and 𝒃𝑇 𝑿 are independent iff 𝒂𝑇 𝒃 = 0.
PROOF:
𝑇 𝑇 𝑝
Let 𝒂𝑝×1 = (𝑎1 , 𝑎2 , ⋯ , 𝑎𝑝 ) and 𝒃𝑝×1 = (𝑏1 , 𝑏2 , ⋯ , 𝑏𝑝 ) be two 𝑝 × 1 vectors with 𝒂𝑇 𝒃 = ∑𝑖=1 𝑎𝑖 𝑏𝑖 = 0.
𝑌 𝑇 𝑇
Let us consider the transformation 𝑌1 = 𝒂𝑇 𝑿, 𝑌2 = 𝒃𝑇 𝑿 and 𝒀2×1 = ( 1 ) = (𝒂𝑇 𝑿) = (𝒂𝑇 ) 𝑿 = 𝑷2×𝑝 𝑿.
𝑌2 𝒃 𝑿 𝒃
For any real 𝒕2×1 = (𝑡1 , 𝑡2 )𝑇 , the moment generating function (M.G.F.) of 𝒀 is given by
1 1
= exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺) ; [∵ If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺), 𝑴𝑿 (𝑺) = exp (𝑺𝑇 𝝁 + 𝑺𝑇 𝚺 𝑺)]
2 2
1
= exp (𝒕𝑇 𝑷𝝁 + 𝒕𝑇 (𝑷𝚺𝑷𝑇 )𝒕) ⟶ M.G.F. of 𝑵2 (𝑷𝝁, 𝑷𝚺𝑷𝑇 )
2
Since 𝚺 is a positive definite and 𝑷2×𝑝 is a 2 × 𝑝 (2 ≤ 𝑝) matrix with 𝑎𝑛𝑘(𝑷) = 2 , 𝑷𝚺𝑷𝑇 is also positive definite.
By uniqueness property of moment generating function (M.G.F.), 𝑴𝒀 (𝒕) is the M.G.F. of 𝑵2 (𝑷𝝁, 𝑷𝚺𝑷𝑇 ).
𝒂𝑇 𝝁 𝑇
2 𝒂 𝒂 0 )).
Hence, 𝒀2×1 = 𝑷𝑿 ~ 𝑵2 (( 𝑇 ),𝜎 (
𝒃 𝝁 0 𝒃𝑇 𝒃
𝑇 𝒂𝑇 𝝁 ∑𝑝 𝑎𝑖 𝜇𝑖
Now, 𝑷𝝁 = (𝒂𝑇 ) 𝝁 = ( 𝑇 ) = ( 𝑖=1 )
𝒃 𝒃 𝝁 ∑𝑝𝑖=1 𝑏𝑖 𝜇𝑖
𝑝
𝑇
and 𝑷𝚺𝑷 = 𝜎 𝑷𝑷 = 𝜎 (𝒂𝑇 𝒂
𝑇 2 𝑇 2 𝒂 𝑇 𝒃 ) = (𝜎 2 𝒂 𝑇 𝒂 0 ) 𝑇
[∵ 𝒂 𝒃 = ∑ 𝑎𝑖 𝑏𝑖 = 0]
𝒃 𝒂 𝒃𝑇 𝒃 0 𝜎 2 𝒃𝑇 𝒃
𝑖=1
−1
2 𝑇 (𝜎 2 𝒂𝑇 𝒂)−1 0
∴ (𝑷𝚺𝑷𝑇 )−1 = (𝜎 𝒂 𝒂 0 ) ⟺ (𝑷𝑇 )−1 𝚺 −1 𝑷−1 = ( ).
0 2 𝑇
𝜎 𝒃 𝒃 0 (𝜎 2 𝒃𝑇 𝒃)−1
2 𝑇
and |𝑷𝚺𝑷𝑇 | = |𝜎 𝒂 𝒂 0 | ⟹ |𝑷||𝚺||𝑷𝑇 | = (𝜎 2 𝒂𝑇 𝒂) × (𝜎 2 𝒃𝑇 𝒃)
0 𝜎 2 𝒃𝑇 𝒃
𝑇
𝑌 − 𝒂𝑇 𝝁 (𝜎 2 𝒂𝑇 𝒂)−1 0 𝑌1 − 𝒂𝑇 𝝁
∴ (𝒀 − 𝑷𝝁)𝑇 (𝑷𝚺𝑷𝑇 )−1 (𝒀 − 𝑷𝝁) = ( 1 𝑇 ) ( −1 ) ( )
𝑌2 − 𝒃 𝝁 0 (𝜎 2 𝑇
𝒃 𝒃) 𝑌2 − 𝒃𝑇 𝝁
1 1
𝑓𝒀 (𝑦1 , 𝑦2 ) = 2 × exp {− (𝒀 − 𝑷𝝁)𝑇 (𝑷𝚺𝑷𝑇 )−1 (𝒀 − 𝑷𝝁)}
(2𝜋)2 √|𝑷𝚺𝑷𝑇 | 2
1 1
=[ × exp {− (𝑌1 − 𝒂𝑇 𝝁)𝑇 (𝜎 2 𝒂𝑇 𝒂)−1 (𝑌1 − 𝒂𝑇 𝝁)}]
√2𝜋√𝜎 2 𝒂𝑇 𝒂 2
1 1
×[ × exp {− (𝑌2 − 𝒃𝑇 𝝁)𝑇 (𝜎 2 𝒃𝑇 𝒃)−1 (𝑌2 − 𝒃𝑇 𝝁)}]
√2𝜋√𝜎 2 𝒃𝑇 𝒃 2
∴ 𝒂𝑇 𝒃 = 0 [∵ 𝜎 2 > 0]
RESULT 7:
PROOF:
Suppose 𝑿𝑖 ~ 𝑵𝒑 (𝝁𝒊 , 𝚺𝐢 ); 𝑖 = 1(1)𝑘 be 𝑘 independent 𝑝-dimensional multivariate normal vectors and 𝒀𝑝×1 = ∑𝑘𝑖=1 𝑿𝑖 .
𝑇
For any real 𝒕𝑝×1 = (𝑡1 , 𝑡2 , ⋯ , 𝑡𝑝 ) , the moment generating function (M.G.F.) of 𝒀 is given by
𝑘 𝑘
𝑇 𝑇
𝑴𝒀 (𝒕) = 𝐸{exp(𝒕 𝒀)} = 𝐸 [exp {𝒕 (∑ 𝑿𝑖 )}] = 𝐸 {exp (∑ 𝒕𝑇 𝑿𝑖 )}
𝑖=1 𝑖=1
𝑘
1 1
= ∏ [exp (𝒕𝑇 𝝁𝒊 + 𝒕𝑇 𝚺𝒊 𝒕)] ; [∵ If 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁𝑿 , 𝚺𝐗 ), 𝑴𝑿 (𝒕) = exp (𝒕𝑇 𝝁𝑿 + 𝒕𝑇 𝚺𝐗 𝒕)]
2 2
𝑖=1
𝑘 𝑘
1
𝑇
= exp (𝒕 ∑ 𝝁𝒊 + 𝒕𝑇 (∑ 𝚺𝒊 ) 𝒕)
2
𝑖=1 𝑖=1
1
= exp (𝒕𝑇 𝝁 + 𝒕𝑇 𝚺 𝒕) ⟶ M.G.F. of 𝑵𝑝 (𝝁, 𝚺)
2
Since 𝚺𝒊 ; 𝑖 = 1(1)𝑘 are positive definite matrix, 𝚺 = ∑𝑘𝑖=1 𝚺𝑖 is also positive definite.
By uniqueness property of moment generating function (M.G.F.), 𝑴𝒀 (𝒕) is the M.G.F. of 𝑵𝑝 (𝝁, 𝚺) i.e., the distribution
of 𝒀𝑝×1 = ∑𝑘𝑖=1 𝑿𝑖 is 𝑵𝒑 (𝝁, 𝚺).
RESULT 8:
𝑇
Suppose 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺). Then 𝑸 = (𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) ~ 𝝌𝟐𝒑 .
PROOF:
𝑇
Suppose 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺). The density function (p.d.f.) of 𝑿𝑝×1 ~ 𝑵𝒑 (𝝁, 𝚺) is given by
1 1
𝑓𝑿 (𝒙) = 𝑝 × exp {− (𝒙 − 𝝁)𝑇 𝚺 −1 (𝒙 − 𝝁)} ; 𝒙 ∈ ℝ𝑝 , 𝝁 ∈ ℝ𝑝 , 𝚺 is 𝑝. 𝑑.
(2𝜋)2 √|𝚺| 2
Since 𝚺 is a positive definite, there exist a non-singular matrix 𝑪 such that 𝑪𝑪𝑇 = 𝚺.
𝑿
|𝐽| = |𝐽 ( )| = ||𝑪|| = √|𝑪𝑪𝑇 | = √|𝚺|.
𝒀
𝑇
Now, (𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) = (𝑿 − 𝝁)𝑇 (𝑪𝑪𝑇 )−1 (𝑿 − 𝝁) = (𝑪−1 (𝑿 − 𝝁)) (𝑪−1 (𝑿 − 𝝁)) = 𝒀𝑇 𝒀.
1 1
𝑔𝒀 (𝒚) = 𝑝 × exp (− 𝒚𝑇 𝒚) × √|𝚺| ; 𝒚 ∈ ℝ𝑝 , 𝚺 is 𝑝. 𝑑.
(2𝜋)2 √|𝚺| 2
𝑝
1 1 2
= 𝑝 ∙ exp (− ∑ 𝑦𝑖 )
(2𝜋)2 2
𝑖=1
𝑝
1 1
= ∏{ ∙ exp (− 𝑦𝑖2 )}
√2𝜋 2
𝑖=1
Hence, 𝒀 = 𝑪−1 (𝑿 − 𝝁) ~ 𝑵𝒑 (𝟎, 𝑰𝒑 ) i.e.,𝑌𝑖′ s are independent and identically distributed with 𝑌𝑖 ~ 𝑁1 (0, 1); 𝑖 = 1(1)𝑝.
∴ 𝑸 = (𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) = 𝒀𝑇 𝒀 ~ 𝝌𝟐𝒑 .
RESULT 9:
Suppose 𝚺 is positive definite, so that 𝚺 −1 exists and (𝜆, 𝒆) is an eigenvalue-eigenvector pair for 𝚺. Then, 𝚺𝒆 = 𝜆𝒆
1 1
implies 𝚺 −1 𝒆 = 𝒆, where ( , 𝒆) is the corresponding eigenvalue-eigenvector pair for 𝚺 −1 . Also, 𝚺 −1 is positive definite.
𝜆 𝜆
PROOF:
Suppose 𝚺 is positive definite, so that 𝚺 −1 exists and (𝜆, 𝒆) is an eigenvalue-eigenvector pair for 𝚺 corresponding to the
1
pair ( , 𝒆) for 𝚺 −1 .
𝜆
1
Now, 𝒆 = 𝚺 −1 (𝚺𝒆) = 𝚺 −1 (𝜆𝒆) = 𝜆𝚺 −1 𝒆 ⟹ 𝒆 = 𝚺 −1 𝒆.
𝜆
1
Hence, ( , 𝒆) is an eigenvalue-eigenvector pair for 𝚺 −1 .
𝜆
1
Suppose ( , 𝒆𝒊 ) ; 𝑖 = 1(1)𝑝 are eigenvalue-eigenvector pair for 𝚺 −1 .
𝜆𝑖
𝑝 𝑝
1 1 1
For any vector 𝒙𝑝×1 , 𝒙𝑇 𝚺 −1 𝒙 = 𝒙𝑇 (∑ 𝒆𝒊 𝒆𝑇𝒊 ) 𝒙 = ∑ ( ) (𝒙𝑇 𝒆𝒊 )𝟐 ≥ 0 ; [∵ ( ) (𝒙𝑇 𝒆𝒊 )𝟐 ≥ 0; 𝑖 = 1(1)𝑝]
𝜆𝑖 𝜆𝑖 𝜆𝑖
𝑖=1 𝑖=1
𝑝
𝑇 −1
1
∴ 𝒙≠𝟎 ⟹ 𝒙 𝚺 𝒙 = ∑ ( ) (𝒙𝑇 𝒆𝒊 )𝟐 > 0.
𝜆𝑖
𝑖=1
DIAGONALIZATION:
𝑇
Suppose 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺) and 𝚺 is positive definite with eigen-vectors 𝒆𝟏 , 𝒆𝟐 , ⋯ , 𝒆𝒑
corresponding to the eigen-values 𝜆1 , 𝜆2 , ⋯ , 𝜆𝑝 ; (𝜆𝑖 > 0).
0 if 𝑖 ≠ 𝑗
Now, 𝚺𝒆𝒊 = 𝜆𝑖 𝒆𝒊 ; 𝑖 = 1(1)𝑝 and 𝒆𝑇𝒊 𝒆𝒋 = { .
1 if 𝑖 = 𝑗
𝑇
Let us define 𝑼𝑝×𝑝 = (𝒆𝟏 , 𝒆𝟐 , ⋯ , 𝒆𝒑 ) . Then, 𝑼𝑇 𝑿 ~ 𝑵𝒑 (𝑼𝑇 𝝁, 𝑼𝑇 𝚺𝑼).
𝜆 if 𝑖 = 𝑗
But, 𝒆𝑇𝒋 𝚺𝒆𝒊 = 𝜆𝑖 𝒆𝑇𝒋 𝒆𝒊 = { 𝑖
0 if 𝑖 ≠ 𝑗.
Hence, 𝑼𝑇 𝚺𝑼 = 𝑫𝒊𝒂𝒈(𝜆1 , 𝜆2 , ⋯ , 𝜆𝑝 ).
Thus, given 𝚺, we can always construct an orthogonal matrix 𝑼𝑝×𝑝 such that if 𝒀 = 𝑼𝑇 𝑿, then 𝑌1 , 𝑌2 , ⋯ , 𝑌𝑝 are
independent normal variables with variances 𝜆1 , 𝜆2 , ⋯ , 𝜆𝑝 respectively.
CONTOUR:
For the density of a 𝑝-dimensional normal variable, the paths of 𝒙 values yielding a constant height for the density
are ellipsoids. The multivariate normal density is constant on surfaces where the square of the distance
(𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) is constant. These paths are called contours:
The family of ellipsoids obtained by varying 𝑐 (𝑐 > 0) has the same centre 𝝁, their shapes and orientation are determined
by 𝚺 and their sizes for a given 𝚺 are determined by 𝑐.
The axes of each ellipsoid of constant density are in the direction of the eigenvectors of 𝚺 −1 and their lengths are
proportional to the reciprocals of the square roots of the eigenvalues of 𝚺 −1.
Contours of constant density for the 𝑝-dimensional normal distribution are ellipsoids defined by 𝒙 such that
(𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) = 𝑐 2 , where 𝑐 is a constant.
𝑝
These ellipsoids are centered at 𝝁 and have axes ±𝑐√𝜆𝑖 𝒆𝒊 , where ∑𝑖=1 𝒆𝒊 = 𝜆𝑖 𝒆𝒊 ; 𝑖 = 1(1)𝑝.
FIGURE: The 50% and 90% contours for the bivariate normal distributions with 𝜎11 = 𝜎22 (Johnson, Wichern).
OBSERVATION 11:
The solid ellipsoid of 𝒙 values satisfying (𝑿 − 𝝁)𝑇 𝚺 −1 (𝑿 − 𝝁) ≤ 𝝌𝟐𝒑 (𝛼), has probability (1 − 𝛼) i.e.,
OBSERVATION 12:
𝑝
𝚪 ( + 1)
[ 2 ]; if (𝒙 − 𝝁)𝑇 𝚺 −1 (𝒙 − 𝝁) ≤ 𝑝 + 2 , 𝒙 ∈ ℝ𝑝 , 𝝁 ∈ ℝ𝑝 , 𝚺 is 𝑝. 𝑑.
𝑝
𝑓𝑿 (𝒙) = {(𝑝 + 2)𝜋}2 √|𝚺|
{ 0; otherwise,
has the same mean 𝐸(𝑿) = 𝝁 and the same covariance matrix 𝐸{(𝑿 − 𝝁)(𝑿 − 𝝁)𝑇 } = 𝚺 as the p-variate distribution.
This will be called the ellipsoid of concentration for the given distribution.
CASE I:
Let us define
𝑇
𝐸 {(𝑿(1) − 𝝁(1) )(𝑿(1) − 𝝁(1) ) } = 𝚺11 ,
𝑇
𝐸 {(𝑿 (1) − 𝝁(1) )(𝑿(2) − 𝝁(2) ) } = 𝚺12 , and
𝑇
𝐸 {(𝑿(2) − 𝝁(2) )(𝑿(2) − 𝝁(2) ) } = 𝚺22 .
We know that, the conditional distribution of 𝑿(1) given for fixed 𝑿(2) = 𝒙(2) is given by
(1)
The conditional expectation and variance-covariance matrix of 𝑿𝑚×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑚 )𝑇 given the fixed values
(2) 𝑇
𝒙(𝑝−𝑚)×1 = (𝑥𝑚+1 , 𝑥𝑚+2 , ⋯ , 𝑥𝑝 ) of the predictor variables are, respectively,
This conditional expected value, considered as a function of 𝑥𝑚+1 , 𝑥𝑚+2 , ⋯ , 𝑥𝑝 is called the multivariate regression of
the vector 𝑿(1) on 𝑿(2) = 𝒙(2) . It is composed of 𝑚 univariate linear regressions.
The vector
(1)
𝑿1.(2) = 𝝁(1) + 𝚺12 𝚺22
−1
(𝒙(2) − 𝝁(2) )
The 𝑖-th component of the conditional mean vector 𝐸(𝑋𝑖 |𝑿(2) = 𝒙(2) ) is the linear regression of 𝑋𝑖 on 𝑿(2) = 𝒙(2) i.e.,
which minimizes the mean square error for the prediction of 𝑋𝑖 , where 𝛔𝑇(𝑖) = 𝚺𝑋 𝑿(2) ; 𝑖 = 1(1)𝑚 is the 𝑖-th row of 𝚺12 .
𝑖
(𝑝−𝑚)×1 𝑇 𝑇
The vector 𝜷(𝑖) −1
= (𝚺𝑋 𝑿(2) 𝚺22 ) = (𝛔𝑇(𝑖) 𝚺22
−1
) is the vector of regression coefficients, where 𝛔𝑇(𝑖) ; 𝑖 = 1(1)𝑚 is
𝑖
−1
The 𝑚 × (𝑝 − 𝑚) matrix 𝑩 = 𝚺12 𝚺22 = (𝜷𝑚×1 𝑚×1 𝑚×1
(1) , 𝜷(2) , ⋯ , 𝜷(𝑝−𝑚) ) is called the matrix of regression coefficients.
The element of 𝑖-th row and 𝑘-th column of 𝑩𝑚×(𝑝−𝑚) , denoted by,
𝑇
is the regression coefficient of 𝑋𝑖 on 𝑋𝑘 = 𝑥𝑘 , and it is also the 𝑘-th element of (𝛔𝑇(𝑖) 𝚺22
−1
) .
The correlation between 𝑋𝑖 and its best linear predictor 𝑋𝑖.(2) = 𝐸(𝑋𝑖 |𝑿(2) = 𝒙(2) ) ; 𝑖 = 1(1)𝑚 is called the population
multiple correlation coefficient.
𝛔𝑇(𝑖) 𝚺22
−1
𝝈(𝑖)
= +√ ; [∵ 𝑿(2) ~ 𝑵(𝒑−𝒎) (𝝁(2) , 𝚺22 )]
𝜎𝑖𝑖
The square of the population multiple correlation coefficient, 𝑅𝑖.(2) or 𝜌𝑖.(2) , is called the population coefficient of
determination. The multiple correlation coefficient is a positive square root, so 0 ≤ 𝜌𝑖.(2) ≤ 1.
(1) (1)
The error of prediction vector 𝜺1.(2) = 𝑿(1) − 𝝁(1) − 𝚺12 𝚺22
−1
(𝑿(2) − 𝝁(2) ) = (𝑿(1) − 𝑿1.(2) ) has the expected squares
and cross-products (error-variation) matrix
𝑇
𝚺11.2 = 𝐸[𝑿(1) − 𝝁(1) − 𝚺12 𝚺22
−1
(𝑿(2) − 𝝁(2) )][𝑿(1) − 𝝁(1) − 𝚺12 𝚺22
−1
(𝑿(2) − 𝝁(2) )]
−1 −1 −1 −1
= 𝚺11 − 𝚺12 𝚺22 𝚺21 − 𝚺12 𝚺22 𝚺21 + 𝚺12 𝚺22 𝚺22 𝚺22 𝚺21
−1
= 𝚺11 − 𝚺12 𝚺22 𝚺21 .
(1) (1)
It can be shown that 𝜺1.(2) = (𝑿(1) − 𝑿1.(2) ) ~ 𝑵𝒎 (𝟎, 𝚺11.2 ).
obtained from the best linear predictors to predict 𝑋𝑖 and 𝑋𝑘 on 𝑿(2) = 𝒙(2) ; 𝑖, 𝑘 = 1(1)𝑚.
−1
Let 𝜎𝑖𝑘.(2) = 𝜎𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 = 𝐶𝑜𝑣(𝜀𝑖.(2) , 𝜀𝑘.(2) ) be the (𝑖, 𝑘)-th element of 𝚺11.2 = 𝚺11 − 𝚺12 𝚺22 𝚺21 ; 𝑖, 𝑘 = 1(1)𝑚.
The correlation between 𝜀𝑖.(2) and 𝜀𝑘.(2) ; 𝑖, 𝑘 = 1(1)𝑚 is called the population partial correlation coefficient between
𝑋𝑖 and 𝑋𝑘 eliminating the effects of 𝑿(2) . The population partial correlation coefficient between 𝑋𝑖 and 𝑋𝑘 eliminating
the effects of 𝑿(2) , denoted by 𝜌𝑖𝑘.(2) or, 𝜌𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 , is given by
𝜎𝑖𝑘.(2) 𝜎𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝
= =
√𝜎𝑖𝑖.(2) ∙ √𝜎𝑘𝑘.(2) √𝜎𝑖𝑖.(𝑚+1)(𝑚+2)⋯𝑝 ∙ √𝜎𝑘𝑘.(𝑚+1)(𝑚+2)⋯𝑝
CASE II:
Let us define
𝑋1 𝜇1 𝜎11 𝑇
𝛔(1)
𝑿𝑝×1 = ( (2) ), 𝝁𝑝×1 = (𝝁(2) ), 𝚺=( (𝑝−1)×(𝑝−1) ) ,
𝑿(𝑝−1)×1 (𝑝−1)×1 𝝈(1) 𝚺22
𝑇
where 𝐸(𝑋1 ) = 𝜇1 , 𝐸(𝑿(2) ) = 𝝁(2) = (𝜇2 , 𝜇3 , ⋯ , 𝜇𝑝 ) , 𝐸{(𝑋𝑖 − 𝜇𝑖 )(𝑋𝑗 − 𝜇𝑗 )} = 𝜎𝑖𝑗 ; 𝑖, 𝑗 = 1(1)𝑝
𝑇
𝐸 {(𝑋1 − 𝜇1 )(𝑿(2) − 𝝁(2) ) } = 𝛔𝑇(1) = (𝜎12 , 𝜎13 , ⋯ , 𝜎1𝑝 ) , and
𝑇
𝐸 {(𝑿(2) − 𝝁(2) )(𝑿(2) − 𝝁(2) ) } = 𝚺22 .
We know that, the conditional distribution of 𝑋1 given for fixed 𝑿(2) = 𝒙(2) is given by
(2) 𝑇
The conditional expectation and variance of 𝑋1 given the fixed values 𝒙(𝑝−1)×1 = (𝑥2 , 𝑥3 , ⋯ , 𝑥𝑝 ) of the predictor
variables are, respectively,
This conditional expected value, considered as a function of 𝑥2 , 𝑥3 , ⋯ , 𝑥𝑝 is called the multivariate regression of 𝑋1 on
𝑿(2) = 𝒙(2) . Hence, the function of 𝑥2 , 𝑥3 , ⋯ , 𝑥𝑝 ,
is the linear regression function, which minimizes the mean square error for the prediction of 𝑋1 on 𝑿(2) = 𝒙(2) .
𝑇 −1 𝑇
−1
The vector 𝜷(𝑝−1)×1 = (𝛔(1) 𝚺22 ) = 𝚺22 𝝈(1) is the vector of regression coefficients.
The 𝑘-th element of 𝜷(𝑝−1)×1 , denoted by, 𝛽1𝑘.23⋯(𝑘−1)(𝑘+1)⋯𝑝 is the regression coefficient of 𝑋1 on 𝑋𝑘 = 𝑥𝑘 ; 𝑘 = 2(1)𝑝.
The correlation between 𝑋1 and its best linear predictor 𝑋1.23⋯𝑝 = 𝑋1.(2) = 𝐸(𝑋1 |𝑿(2) = 𝒙(2) ) is called the population
multiple correlation coefficient.
𝛔𝑇(1) 𝚺22
−1
𝝈(1)
= +√ ; [∵ 𝑿(2) ~ 𝑵(𝒑−𝟏) (𝝁(2) , 𝚺22 )]
𝜎11
The square of the population multiple correlation coefficient, 𝑅 (𝑅1.(2) or 𝜌1.(2) or 𝑅1.23⋯𝑝 ), is called the population
coefficient of determination. The multiple correlation coefficient is a positive square root, so 0 ≤ 𝑅 ≤ 1.
𝑇
The error of prediction 𝜀1.23⋯𝑝 = 𝜀1.(2) = (𝑋1 − 𝑋1.(2) ) = 𝑋1 − 𝜇1 − 𝛔(1) −1
𝚺22 (𝒙(2) − 𝝁(2) ) has the expected squares and
cross-products (error-variation) matrix
𝑇 2
= 𝐸{𝑋1 − 𝜇1 − 𝛔(1) −1
𝚺22 (𝒙(2) − 𝝁(2) )}
Suppose
𝑇
where 𝐸(𝑋𝑖 ) = 𝜇𝑖 ; 𝑖 = 1, 2, 𝐸(𝑿(3) ) = 𝝁(3) = (𝜇3 , 𝜇4 , ⋯ , 𝜇𝑝 ) , 𝐸{(𝑋𝑖 − 𝜇𝑖 )(𝑋𝑗 − 𝜇𝑗 )} = 𝜎𝑖𝑗 ; 𝑖, 𝑗 = 1(1)𝑝,
𝑇
𝐸 {(𝑋1 − 𝜇1 )(𝑿(3) − 𝝁(3) ) } = 𝛔𝑇(1) = (𝜎13 , 𝜎14 , ⋯ , 𝜎1𝑝 ),
𝑇
𝐸 {(𝑋2 − 𝜇2 )(𝑿(3) − 𝝁(3) ) } = 𝛔𝑇(2) = (𝜎23 , 𝜎24 , ⋯ , 𝜎2𝑝 ), and
𝑇 (𝑝−2)×(𝑝−2)
𝐸 {(𝑿(3) − 𝝁(3) )(𝑿(3) − 𝝁(3) ) } = 𝚺33 .
obtained from the best linear predictors to predict 𝑋1 and 𝑋2 on 𝑿(3) = 𝒙(3) .
𝑇 −1
𝜎11.(3) = 𝜎11.34⋯𝑝 = 𝑉𝑎𝑟(𝜀1.(3) ) = 𝑉𝑎𝑟(𝜀1.34⋯𝑝 ) = 𝜎11 − 𝛔(1) 𝚺33 𝝈(1) ,
𝑇 −1
𝜎22.(3) = 𝜎22.34⋯𝑝 = 𝑉𝑎𝑟(𝜀2.(3) ) = 𝑉𝑎𝑟(𝜀2.34⋯𝑝 ) = 𝜎22 − 𝛔(2) 𝚺33 𝝈(2) , and
The correlation between 𝜀1.(3) and 𝜀2.(3) is called the population partial correlation coefficient between 𝑋1 and 𝑋2
eliminating the effects of 𝑿(3) . The population partial correlation coefficient between 𝑋1 and 𝑋2 eliminating the effects
𝑇
of 𝑿(3) = (𝑋3 , 𝑋4 , ⋯ , 𝑋𝑝 ) , denoted by 𝜌12.(3) or, 𝜌12.34⋯𝑝 , is given by
𝜎12.(3) 𝜎12.34⋯𝑝
= =
√𝜎11.(3) ∙ √𝜎22.(3) √𝜎11.34⋯𝑝 ∙ √𝜎22.34⋯𝑝
CASE I:
2
Suppose we want to test the null hypothesis 𝐻0 ∶ 𝜌𝑖.(2) = 0, where 𝜌𝑖.(2) is the population multiple correlation coefficient
𝑇
between 𝑋𝑖 ; 𝑖 = 1(1)𝑚 and 𝑿(2) = (𝑋𝑚+1 , 𝑋𝑚+2 , ⋯ , 𝑋𝑝 ) .
𝑇
Let 𝑅𝑖.(2) ; 𝑖 = 1(1)𝑚 be the sample multiple correlation coefficient between 𝑋𝑖 and 𝑿(2) = (𝑋𝑚+1 , 𝑋𝑚+2 , ⋯ , 𝑋𝑝 )
𝑇
based on a sample [(𝑋1𝑘 , 𝑋2𝑘 , ⋯ , 𝑋𝑝𝑘 ) ; 𝑘 = 1(1)𝑁] of size 𝑁 from 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺).
𝑇
The sample multiple correlation coefficient 𝑅𝑖.(2) between 𝑋𝑖 and 𝑿(2) = (𝑋𝑚+1 , 𝑋𝑚+2 , ⋯ , 𝑋𝑝 ) is given by
𝑇 −1
𝑇 −1 (2) 𝐬(𝑖) 𝐒22 𝒔(𝑖)
𝑅𝑖.(2) = 𝐶𝑜𝑟𝑟(𝑋𝑖 , 𝑋𝑖.(2) ) = 𝐶𝑜𝑟𝑟 [𝑋𝑖 , {𝑋𝑖 + 𝐬(𝑖) 𝐒22 (𝑿(2) − 𝑿 )}] = +√ ; 𝑖 = 1(1)𝑚
𝑠𝑖𝑖
2 −1
The null hypothesis 𝐻0 ∶ 𝜌𝑖.(2) = 0 is equivalent to 𝐻0 ∶ 𝜷𝑖.(2) = 𝟎, since 𝜷𝑖.(2) = 𝚺22 𝝈(𝑖) and 𝝈(𝑖) = 𝟎.
2
𝑅𝑖.(2) {𝑁 − (𝑝 − 𝑚) − 1}
𝐹= 2 × ~ 𝐹(𝑝−𝑚),{𝑁−(𝑝−𝑚)−1}
1− 𝑅𝑖.(2) (𝑝 − 𝑚)
2
To test 𝐻0 ∶ 𝜌𝑖.(2) = 0 or equivalently 𝐻0 ∶ 𝜷𝑖.(2) = 𝟎, we use the likelihood ratio test based on 𝐹 and we reject 𝐻0 if
2
𝑅𝑖.(2) {𝑁 − (𝑝 − 𝑚) − 1}
(Observed ) 𝐹 = 2 × ≥ 𝐹(𝑝−𝑚),{𝑁−(𝑝−𝑚)−1} (𝛼)
1− 𝑅𝑖.(2) (𝑝 − 𝑚)
where 𝐹(𝑝−𝑚),{𝑁−(𝑝−𝑚)−1} (𝛼) is the upper significance point of 𝐹(𝑝−𝑚),{𝑁−(𝑝−𝑚)−1} distribution corresponding to the level
CASE II:
2
2
Suppose we want to test the null hypothesis 𝐻0 ∶ 𝜌1.23⋯𝑝 = 0 (or 𝑅1.23⋯𝑝 = 0), where 𝜌1.23⋯𝑝 (or 𝑅1.23⋯𝑝 or 𝑅) is the
𝑇
population multiple correlation coefficient between 𝑋1 and 𝑿(2) = (𝑋2 , 𝑋3 , ⋯ , 𝑋𝑝 ) .
𝑇
Let 𝑅1.23⋯𝑝 or 𝑅 be the sample multiple correlation coefficient between 𝑋1 and 𝑿(2) = (𝑋2 , 𝑋3 , ⋯ , 𝑋𝑝 ) based on a
𝑇
sample [(𝑋1𝑘 , 𝑋2𝑘 , ⋯ , 𝑋𝑝𝑘 ) ; 𝑘 = 1(1)𝑁] of size 𝑁 from 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺).
𝑇
The sample multiple correlation coefficient 𝑅1.23⋯𝑝 or 𝑅 between 𝑋1 and 𝑿(2) = (𝑋2 , 𝑋3 , ⋯ , 𝑋𝑝 ) is given by
𝑇 −1
𝑇
(2) 𝐬(1) 𝐒22 𝒔(1)
𝑅 = 𝑅1.23⋯𝑝 = 𝐶𝑜𝑟𝑟(𝑋1 , 𝑋1.23⋯𝑝 ) = 𝐶𝑜𝑟𝑟 [𝑋1 , {𝑋1 + 𝐬(1) −1
𝐒22 (𝑿(2) − 𝑿 )}] = +√ .
𝑠11
2 −1
The null hypothesis 𝐻0 ∶ 𝜌1.23⋯𝑝 = 0 is equivalent to 𝐻0 ∶ 𝜷1.23⋯𝑝 = 𝟎, since 𝜷 = 𝜷1.23⋯𝑝 = 𝚺22 𝝈(1) and 𝝈(1) = 𝟎.
𝑅2 (𝑁 − 𝑝)
𝐹= × ~ 𝐹(𝑝−1),(𝑁−𝑝)
1−𝑅 2 (𝑝 − 1)
2
To test 𝐻0 ∶ 𝜌1.23⋯𝑝 = 0 or equivalently 𝐻0 ∶ 𝜷 = 𝟎, we use the likelihood ratio test based on 𝐹 and we reject 𝐻0 if
𝑅2 (𝑁 − 𝑝)
(Observed ) 𝐹 = × ≥ 𝐹(𝑝−1),(𝑁−𝑝) (𝛼)
1 − 𝑅2 (𝑝 − 1)
where 𝐹(𝑝−1),(𝑁−𝑝) (𝛼) is the upper significance point of 𝐹(𝑝−1),(𝑁−𝑝) distribution corresponding to the level of
NOTE:
We have
𝑇 −1 𝑇 −1 𝑇 −1
2
𝐬(1) 𝐒22 𝒔(1) 2
𝐬(1) 𝐒22 𝒔(1) 𝑠11 − 𝐬(1) 𝐒22 𝒔(1) 𝑠11.(2)
𝑅 = or, 1−𝑅 =1− = =
𝑠11 𝑠11 𝑠11 𝑠11
𝑇 −1
𝑅2 𝐬(1) 𝐒22 𝒔(1)
∴ 2
=
1−𝑅 𝑠11.(2)
When 𝜌1.23⋯𝑝 = 0 or, 𝑅1.23⋯𝑝 = 0, that is, when 𝜷1.23⋯𝑝 = 𝟎 or when 𝛽1𝑘.23⋯(𝑘−1)(𝑘+1)⋯𝑝 = 0 for each 𝑘 = 2(1)𝑝,
𝑇 𝑁−𝑝 𝑇
𝑠11.(2) = 𝑠11 − 𝐬(1) −1
𝐒22 𝒔(1) is distributed as ∑𝑘=1 𝑍𝑘2 and 𝐬(1) −1
𝐒22 𝒔(1) is distributed as ∑𝑁−1 2
𝑘=𝑁−𝑝+1 𝑍𝑘 , where
𝑇 −1 𝑇 −1
𝑠11.(2) 𝑠11..23⋯𝑝 𝐬(1) 𝐒22 𝒔(1) 𝐬(1) 𝐒22 𝒔(1)
Then = and = are distributed independently as 𝜒 2 -variables with (𝑁 − 𝑝) and (𝑝 − 1)
𝜎11.(2) 𝜎11..23⋯𝑝 𝜎11.(2) 𝜎11..23⋯𝑝
CASE I:
Suppose we want to test the null hypothesis 𝐻0 ∶ 𝜌𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 = 𝜌0 (Given) against two-sided alternative hypothesis
𝐻1 ∶ 𝜌𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 ≠ 𝜌0 , where 𝜌𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 ; 𝑖, 𝑘 = 1(1)𝑚 is the population partial correlation coefficient
𝑇
between 𝑋𝑖 and 𝑋𝑘 ; 𝑖, 𝑘 = 1(1)𝑚 eliminating the effects of 𝑿(2) = (𝑋𝑚+1 , 𝑋𝑚+2 , ⋯ , 𝑋𝑝 ) .
Suppose 𝑟𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 (or 𝑟𝑖𝑘.(2) ) is the sample partial correlation coefficient based on a sample
𝑇
[(𝑋1𝑘 , 𝑋2𝑘 , ⋯ , 𝑋𝑝𝑘 ) ; 𝑘 = 1(1)𝑁] of size 𝑁 from 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺) with population partial
correlation coefficient 𝜌𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 (or 𝜌𝑖𝑘.(2) ); 𝑖, 𝑘 = 1(1)𝑚.
−1
Let 𝑠𝑖𝑘.(2) = 𝑠𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 = 𝐶𝑜𝑣(𝑒𝑖.(2) , 𝑒𝑘.(2) ) be the (𝑖, 𝑘)-th element of 𝐒11.2 = 𝐒11 − 𝐒12 𝐒22 𝐒21 ; 𝑖, 𝑘 = 1(1)𝑚.
𝑇
The sample partial correlation coefficient between 𝑋𝑖 and 𝑋𝑘 eliminating the effects of 𝑿(2) = (𝑋𝑚+1 , 𝑋𝑚+2 , ⋯ , 𝑋𝑝 ) ,
denoted by 𝑟𝑖𝑘.(2) or, 𝑟𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 , is given by
𝑠𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝
𝑟𝑖𝑘.(2) = 𝑟𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 = 𝐶𝑜𝑟𝑟(𝑒𝑖.(2) , 𝑒𝑘.(2) ) = ; 𝑖, 𝑘 = 1(1)𝑚.
√𝑠𝑖𝑖.(𝑚+1)(𝑚+2)⋯𝑝 ∙ √𝑠𝑘𝑘.(𝑚+1)(𝑚+2)⋯𝑝
We use Fisher’s 𝑧 statistic for an approximate significance test of 𝐻0 ∶ 𝜌𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 = 𝜌0 against two-sided
alternatives based on a sample [(𝑋1𝑘 , 𝑋2𝑘 , ⋯ , 𝑋𝑝𝑘 ) ; 𝑘 = 1(1)𝑁] of size 𝑁.
Let
1 1 + 𝑟𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 1 1 + 𝜌0
𝑧= 𝑙𝑛 and 𝜉0 = 𝑙𝑛 .
2 1 − 𝑟𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 2 1 − 𝜌0
Then 𝐻0 is rejected if
|√𝑁 − (𝑝 − 𝑚) − 3 × (𝑧 − 𝜉0 )| > 𝜏𝛼
2
where 𝜏𝛼 is the upper significance point of standard normal distribution corresponding to the level of significance 𝛼 i.e.,
2
𝛼
𝑃𝐻0 (𝑍 ≥ 𝜏𝛼 |𝑍 ~ 𝑁(0, 1)) = .
2 2
NOTE:
Since the distribution of a sample partial correlation 𝑟𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 based on a sample of 𝑁 from a distribution
with population correlation 𝜌𝑖𝑘.(𝑚+1)(𝑚+2)⋯𝑝 equal to a certain value, 𝜌, say, is the same as the distribution of a
simple correlation 𝑟 based on a sample of size 𝑁 − (𝑝 − 𝑚) from a distribution with the corresponding population
correlation of 𝜌, all statistical inference procedures for the simple population correlation can be used for the partial
correlation. The procedure for the partial correlation is exactly the same except that 𝑁 is replaced by 𝑁 − (𝑝 − 𝑚).
CASE II:
Suppose we want to test 𝐻0 ∶ 𝜌12.34⋯𝑝 = 𝜌0 (Given) against two-sided alternative 𝐻1 ∶ 𝜌12.34⋯𝑝 ≠ 𝜌0 , where 𝜌12.34⋯𝑝 is
𝑇
the population partial correlation coefficient between 𝑋1 and 𝑋2 eliminating the effects of 𝑿(3) = (𝑋3 , 𝑋4 , ⋯ , 𝑋𝑝 ) .
Suppose 𝑟12.34⋯𝑝 is the sample partial correlation coefficient based on a sample [(𝑋1𝑘 , 𝑋2𝑘 , 𝑋3𝑘 , ⋯ , 𝑋𝑝𝑘 ); 𝑘 = 1(1)𝑁] of
𝑇
size 𝑁 from 𝑿𝑝×1 = (𝑋1 , 𝑋2 , ⋯ , 𝑋𝑝 ) ~ 𝑵𝒑 (𝝁, 𝚺) with population partial correlation coefficient 𝜌12.34⋯𝑝 .
𝑇
The sample partial correlation coefficient between 𝑋1 and 𝑋2 eliminating the effects of 𝑿(3) = (𝑋3 , 𝑋4 , ⋯ , 𝑋𝑝 ) , denoted
by 𝑟12.(3) or, 𝑟12.34⋯𝑝 , is given by
𝑠12.(3) 𝑠12.34⋯𝑝
𝑟12.(3) = 𝑟12.34⋯𝑝 = 𝐶𝑜𝑟𝑟(𝑒1.34⋯𝑝 , 𝑒2.34⋯𝑝 ) = = .
√𝑠11.(3) ∙ √𝑠22.(3) √𝑠11.34⋯𝑝 ∙ √𝑠22.34⋯𝑝
We use Fisher’s 𝑧 statistic for an approximate significance test of 𝐻0 ∶ 𝜌12.34⋯𝑝 = 𝜌0 against two-sided alternatives based
on a sample [(𝑋1𝑘 , 𝑋2𝑘 , ⋯ , 𝑋𝑝𝑘 ) ; 𝑘 = 1(1)𝑁] of size 𝑁.
Let
1 1 + 𝑟12.34⋯𝑝 1 1 + 𝜌0
𝑧= 𝑙𝑛 and 𝜉0 = 𝑙𝑛 .
2 1 − 𝑟12.34⋯𝑝 2 1 − 𝜌0
Then 𝐻0 is rejected if
|√𝑁 − (𝑝 − 2) − 3 × (𝑧 − 𝜉0 )| > 𝜏𝛼
2
where 𝜏𝛼 is the upper significance point of standard normal distribution corresponding to the level of significance 𝛼 i.e.,
2
𝛼
𝑃𝐻0 (𝑍 ≥ 𝜏𝛼 |𝑍 ~ 𝑁(0, 1)) = .
2 2
RESULT 10:
Let us define
𝑌 𝜇𝑌 𝜎𝑌𝑌 𝛔𝑇(𝑌)
𝒁(𝑝+1)×1 = (𝑿 ) , 𝝁(𝑝+1)×1 = (𝝁𝑝×1 ) , 𝚺 (𝑝+1)×(𝑝+1) = ( 𝑝×𝑝 ) ,
𝑝×1 (𝑋) 𝝈(𝑌) 𝚺𝑋𝑋
𝑇 𝑇
𝑝×1
where 𝐸(𝑌) = 𝜇𝑌 , 𝐸(𝑿𝑝×1 ) = 𝝁(𝑋) = (𝜇𝑋1 , 𝜇𝑋2 , ⋯ , 𝜇𝑋𝑝 ) = (𝜇1 , 𝜇2 , ⋯ , 𝜇𝑝 ) , 𝑉𝑎𝑟(𝑌) = 𝜎𝑌𝑌 ,
𝑇
𝛔𝑇(𝑌) = (𝜎𝑌𝑋1 , 𝜎𝑌𝑋2 , ⋯ , 𝜎𝑌𝑋𝑝 ) , 𝜎𝑌𝑋𝑖 = 𝐶𝑜𝑣(𝑌, 𝑋𝑖 ); 𝑖 = 1(1)𝑝, and 𝐸 {(𝑿 − 𝝁(𝑿) )(𝑿 − 𝝁(𝑋) ) } = 𝚺𝑋𝑋 .
𝑌
Suppose 𝒁(𝑝+1)×1 = (𝑿 ) ~ 𝑵(𝒑+𝟏) (𝝁, 𝚺). Then, the conditional distribution of 𝑌 given for fixed 𝑿 = 𝒙 is given by
𝑝×1
𝑇
(𝑌|𝑿 = 𝒙) ~ 𝑵𝟏 ({𝜇𝑌 + 𝛔(𝑌) −1
𝚺𝑋𝑋 (𝒙 − 𝝁(𝑋) )}, 𝜎𝑌𝑌.(𝑋) ), where 𝜎𝑌𝑌.(𝑋) = 𝜎𝑌𝑌 − 𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌) .
2
𝐸(𝑌 − 𝛽0 − 𝜷𝑇 𝑿)2 = 𝐸 (𝑌 − 𝜇𝑌 − 𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
(𝑿 − 𝝁(𝑋) )) = 𝜎𝑌𝑌 − 𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌) = 𝜎𝑌𝑌.(𝑋) ,
𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌)
𝜌𝑌.(𝑋) = 𝐶𝑜𝑟𝑟(𝑌, (𝛽0 + 𝜷𝑇 𝑿)) = max{𝐶𝑜𝑟𝑟(𝑌, (𝑏0 + 𝒃𝑇 𝑿))} = +√ .
𝑏0 ,𝒃 𝜎𝑌𝑌
PROOF:
𝑌
Suppose 𝒁(𝑝+1)×1 = (𝑿 ) ~ 𝑵(𝒑+𝟏) (𝝁, 𝚺).
𝑝×1
2
𝐸(𝑌 − 𝑏0 − 𝒃𝑇 𝑿)2 = 𝐸{(𝑌 − 𝜇𝑌 ) − 𝒃𝑇 (𝑿 − 𝝁(𝑋) ) + (𝜇𝑌 − 𝑏0 − 𝒃𝑇 𝝁(𝑋) )}
2 2
= 𝐸(𝑌 − 𝜇𝑌 )2 + 𝐸{𝒃𝑇 (𝑿 − 𝝁(𝑋) )} + 𝐸(𝜇𝑌 − 𝑏0 − 𝒃𝑇 𝝁(𝑋) ) − 2𝐸{𝒃𝑇 (𝑿 − 𝝁(𝑋) )(𝑌 − 𝜇𝑌 )}
2
= 𝜎𝑌𝑌 + 𝒃𝑇 𝚺𝑋𝑋 𝒃 + (𝜇𝑌 − 𝑏0 − 𝒃𝑇 𝝁(𝑋) ) − 2𝒃𝑇 𝝈(𝑌)
2
= 𝜎𝑌𝑌 − 𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌) + (𝜇𝑌 − 𝑏0 − 𝒃𝑇 𝝁(𝑋) ) + 𝒃𝑇 𝚺𝑋𝑋 𝒃 − 2𝒃𝑇 𝝈(𝑌) + 𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌)
2 𝑇
= 𝜎𝑌𝑌 − 𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌) + (𝜇𝑌 − 𝑏0 − 𝒃𝑇 𝝁(𝑋) ) + (𝒃 − 𝚺𝑋𝑋
−1 −1
𝝈(𝑌) ) 𝚺𝑋𝑋 (𝒃 − 𝚺𝑋𝑋 𝝈(𝑌) )
−1
The mean square error is minimized by taking 𝒃 = 𝚺𝑋𝑋 𝝈(𝑌) = 𝜷 and then choosing 𝑏0 = 𝜇𝑌 + 𝒃𝑇 𝝁(𝑋) = 𝛽0 to make the
last and third term zero.
Thus, the minimum mean square error is the variance of the conditional distribution of 𝑌 given for fixed 𝑿 = 𝒙,
2
𝐸(𝑌 − 𝛽0 − 𝜷𝑇 𝑿)2 = 𝐸 (𝑌 − 𝜇𝑌 − 𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
(𝒙 − 𝝁(𝑋) )) = 𝜎𝑌𝑌 − 𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌) = 𝜎𝑌𝑌.(𝑋) .
is the best linear predictor of 𝑌 when the population 𝑵(𝒑+𝟏) (𝝁, 𝚺).
Now, 𝑉𝑎𝑟(𝑏0 + 𝒃𝑇 𝑿) = 𝑉𝑎𝑟(𝒃𝑇 𝑿) = 𝒃𝑇 𝚺𝑋𝑋 𝒃 and 𝐶𝑜𝑣(𝑌, (𝑏0 + 𝒃𝑇 𝑿)) = 𝐶𝑜𝑣(𝑌, 𝒃𝑇 𝑿) = 𝒃𝑇 𝝈(𝑌)
2
2 {𝒃𝑇 𝝈(𝑌) }
∴ {𝐶𝑜𝑟𝑟(𝑌, (𝑏0 + 𝒃𝑇 𝑿))} = ; for all 𝑏0 and 𝒃
𝜎𝑌𝑌 (𝒃𝑇 𝚺𝑋𝑋 𝒃)
Since 𝚺𝑋𝑋 is a positive definite, there exist a non-singular matrix 𝑪 such that 𝑪𝑪𝑇 = 𝚺𝑋𝑋 .
𝑻
𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌) = 𝛔𝑇(𝑌) (𝑪𝑪𝑇 )−1 𝝈(𝑌) = (𝑪−1 𝝈(𝑌) ) (𝑪−1 𝝈(𝑌) ) [∵ (𝑪−1 )𝑇 = (𝑪𝑇 )−1 ]
𝑇 𝑇
By Cauchy-Schwarz Inequality, we have, if 𝒂𝑝×1 = (𝑎1 , 𝑎2 , ⋯ , 𝑎𝑝 ) and 𝒅𝑝×1 = (𝑑1 , 𝑑2 , ⋯ , 𝑑𝑝 ) be two 𝑝 × 1 vectors
then
𝑝 2 𝑝 𝑝
(𝒂𝑇 2
𝒅) ≤ (𝒂𝑇 𝑇
𝒂)(𝒅 𝒅) or, (∑ 𝑎𝑖 𝑑𝑖 ) ≤ (∑ 𝑎𝑖2 ) (∑ 𝑑𝑖2 )
𝑖=1 𝑖=1 𝑖=1
2 𝑻
{𝒃𝑇 𝝈(𝑌) } ≤ {(𝑪𝑇 𝒃)𝑻 (𝑪𝑇 𝒃)} {(𝑪−1 𝝈(𝑌) ) (𝑪−1 𝝈(𝑌) )}
2
or, {𝒃𝑇 𝝈(𝑌) } ≤ (𝒃𝑇 𝚺𝑋𝑋 𝒃)(𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌) )
2
{𝒃𝑇 𝝈(𝑌) } 1 1
or, × ≤ (𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌) ) ×
(𝒃 𝚺𝑋𝑋 𝒃) 𝜎𝑌𝑌
𝑇 𝜎𝑌𝑌
2 𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌)
or, {𝐶𝑜𝑟𝑟(𝑌, (𝑏0 + 𝒃𝑇 𝑿))} ≤
𝜎𝑌𝑌
−1
with equality holds for 𝒃 = 𝚺𝑋𝑋 𝝈(𝑌) = 𝜷.
𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌)
𝜌𝑌.(𝑋) = 𝜌𝑌.1 2 ⋯𝑝 = max{𝐶𝑜𝑟𝑟(𝑌, (𝑏0 + 𝒃𝑇 𝑿))} = 𝐶𝑜𝑟𝑟(𝑌, (𝛽0 + 𝜷𝑇 𝑿)) = +√ .
𝑏0 ,𝒃 𝜎𝑌𝑌
The alternative expression for the maximum correlation between 𝑌 and its best linear predictor or population multiple
−1
correlation coefficient 𝜌𝑌.1 2 ⋯𝑝 with 𝒃 = 𝚺𝑋𝑋 𝝈(𝑌) = 𝜷 is
𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝝈(𝑌) 𝛔𝑇(𝑌) 𝜷 𝛔𝑇(𝑌) 𝚺𝑋𝑋
−1
𝚺𝑋𝑋 𝜷 𝜷𝑇 𝚺𝑋𝑋 𝜷
𝜌𝑌.1 2 ⋯𝑝 = 𝐶𝑜𝑟𝑟(𝑌, (𝛽0 + 𝜷𝑇 𝑿)) = +√ = +√ = +√ = +√
𝜎𝑌𝑌 𝜎𝑌𝑌 𝜎𝑌𝑌 𝜎𝑌𝑌
NOTE:
1. When the population is not normal, the regression function 𝐸(𝑌|𝑥1 , 𝑥, ⋯ , 𝑥𝑝 ) need not be of the form 𝛽0 + 𝜷𝑇 𝑿.
It can be shown that 𝐸(𝑌|𝑥1 , 𝑥, ⋯ , 𝑥𝑝 ), whatever its form, predicts 𝑌 with the smallest mean square error. This
wider optimality among all estimators is possessed by the linear predictor when the population is normal.
2. If 𝜌𝑌.(𝑋) = 𝜌𝑌.1 2 ⋯𝑝 = 0, there is no predictive power in 𝑿. At the other extreme, 𝜌𝑌.(𝑋) = 𝜌𝑌.1 2 ⋯𝑝 = 1 implies
that 𝑌 can be predicted by 𝑿 with no error.
2
3. To test 𝐻0 ∶ 𝜌𝑌.1 2 ⋯𝑝 = 0 or equivalently 𝐻0 ∶ 𝜷 = 𝟎, we use the likelihood ratio test based on 𝐹 and we reject
𝐻0 if
𝑅2 (𝑁 − 𝑝 − 1)
(Observed ) 𝐹 = × ≥ 𝐹𝑝,(𝑁−𝑝−1) (𝛼)
1 − 𝑅2 𝑝
where 𝐹𝑝,(𝑁−𝑝−1) (𝛼) is the upper significance point of 𝐹𝑝,(𝑁−𝑝−1) distribution corresponding to the level of
significance 𝛼 i.e., 𝑃𝐻0 (𝐹 ≥ 𝐹𝑝,(𝑁−𝑝−1) (𝛼)) = 𝛼 and 𝑅 is the sample multiple correlation coefficient.
Transformations involving the multivariate normal distribution include linear transformations that standardize the variables. Given a symmetric positive-definite matrix A, one can use a transformation Y=CT(X-b) where C is such that CC^T=A, facilitating variable interpretation in a standard normal space. These transformations simplify multivariate normal distribution analysis by converting it into independent normal distributions, isolating variable effects, and clarifying the geometric structure of ellipsoidal distributions in the data .
The joint moments about origin of order (h1 + h2 + ⋯+ hp) for non-negative integers (h1, h2, ⋯, hp) is defined as E(X1^h1 X2^h2 ⋯ Xp^hp). If X ~ is discrete, this is given by ∑∑⋯∑ x1^h1 x2^h2 ⋯ xp^hp∙p_X~(x1, x2, ⋯, xp). If X ~ is continuous, it is defined by the integral form. The significance lies in its ability to capture the higher order dependencies among variables, thus facilitating the understanding of complex relationships in multivariate distributions .
In the context of correlation matrices in multivariate distributions, the cofactor of a matrix element is crucial for computing the adjugate, which is key in finding the inverse of a correlation matrix when it is invertible. This inverse is used for variance-covariance matrix computations, allowing the determination of conditional distributions and other multivariate statistics. The cofactor here specifically refers to the minor of an element, which is used to capture the dependencies between the different variables .
In a multivariate normal distribution, random variables are independent if the variance-covariance matrix is diagonal with non-zero elements only along the diagonal. This occurs when Σ is structured as σ2I_p, meaning all off-diagonal elements (covariances) are zero, allowing each variable to vary independently of the others. Independence is inferred from the lack of any linear relationship amongst variables .
The moment generating function (MGF) of a multivariate distribution serves to uniquely define the distribution by encapsulating all of its moments. In particular, the MGF M_X(t) = E{exp(t^T X)} allows us to derive any order of moments by differentiation, reflecting the dependency structures among variables. Its form as a function of t reveals the distribution's parameters and the interrelationships among its components, assisting in hypothesis testing and predictions .
A multivariate distribution becomes singular when its covariance matrix has a determinant equal to zero (|Σk×k| = 0), indicating that there is a linear dependency among the variables. This implies that not all variables can vary independently, reducing their effective dimensionality by one. Consequently, it complicates computation and interpretation, necessitating analysis of a (k-1) dimensional subset, as full-rank operations like finding inverses are not possible in singular cases .
The singularity of a variance-covariance matrix signifies a linear dependence among variables, leading to complications in performing statistical inference because operations such as inversion, crucial for determining conditional distributions and testing hypotheses, are not feasible. This impacts estimations, revealing the need for dimensionality reduction techniques or regularization methods to obtain meaningful statistical insights .
Central moments are calculated as E{(X - μ)^k}, focusing on deviations from the mean, thus describing the shape of the data relative to its mean. In contrast, joint moments about the origin, E(X1^h1 X2^h2 ⋯ Xp^hp), do not consider the mean but rather the raw moments of the joint distribution. This distinction is crucial because central moments provide insights into variability and shape beyond mere magnitudes, such as variance and covariance, which are vital for understanding dependency and relationships in multivariate data .
The correlation matrix of a multivariate random vector, such as R(k-1)×(k-1), shows the linear relationships between each pair of variables in terms of correlation coefficients. The diagonal elements are all 1, representing the perfect correlation of a variable with itself. Off-diagonal elements like -√(pi pj)/((1−pi)(1−pj)) capture the degree and direction of linear dependence between variables under the constraint that each probability pi is less than 1 .
The multinomial coefficient in a multivariate probability distribution is derived from the expression of the probability mass function as n!/(x1! x2!⋯xs!) (n−∑ xi)! ⋅ p1^x1 p2^x2 ⋯ ps^xs, where xi are the counts of events of differing types, n is the total count, and pi are the probabilities of each type of event. This coefficient represents the number of distinct ways the outcomes can occur, given the constraints of the model .