BLG 454E Learning From
Data
FALL 2022-2023
Multivariate Methods
(Slides are Prepared by Assoc. Prof. Yusuf Yaslan
& Assist. Prof. Ayşe Tosun)
Lecture Notes from Alpaydın 2010 Introduction to Machine Learning 2e © The MIT Press (V1.0) AND
O. Velksler’s Lecture Notes from CS 434s/541a Pattern Recognition, Uni of Western Ontario
Univariate Normal Density
• So far, we have dealt with univariate x (dimension of 1)
1 1 𝑥−𝜇 2
• 𝑃 𝑥 = exp(− )
2𝜋𝜎 2 𝜎
1 𝑁
• 𝜇𝑀𝐿𝐸 = 𝑚 = σ𝑖=1 𝑥𝑖
𝑛
1 𝑁
• 𝜎 𝑀𝐿𝐸 = 𝑠 = σ𝑖=1(𝑥𝑖 − 𝑚)2
2 2
𝑛
2
Multivariate Normal Density
• What if we have several features 𝑥1 , 𝑥2 , 𝑥3 , … , 𝑥𝑑 𝑥11
X= ⋮
⋯ 𝑥𝑑1
⋱ ⋮
• Each normally distributed 𝑥1𝑁 ⋯ 𝑥𝑑𝑁
• Different variances
• Different means
• May be dependent or independent of each other
1 1
• 𝑃 𝑥 = 𝑑/2 1/2 exp(− (𝑥 − 𝜇)𝑇 Σ −1 (𝑥 − 𝜇))
2𝜋 Σ 2
Mahalonobis distance between x and mean
(elliptical curve between x’s)
Multivariate parameters
• 𝐸 𝑋 = 𝜇 = 𝜇1 , 𝜇2 , … , 𝜇𝑑 𝑇
• Covariance = 𝜎𝑖𝑗 = 𝐶𝑜𝑣 𝑋𝑖 , 𝑋𝑗
𝜎𝑖𝑗
• Correlation =Corr 𝑋𝑖 , 𝑋𝑗 = 𝜌𝑖𝑗 =
𝜎𝑖 𝜎𝑗
• Covariance matrix Σ = 𝐸 (𝑋 − 𝜇)(𝑋 − 𝜇)𝑇
𝜎12 𝜎12 ⋯ 𝜎1𝑑
= ⋮ ⋱ ⋮
𝜎𝑑1 𝜎𝑑2 ⋯ 𝜎𝑑2
Multivariate parameter estimation
σ𝑁 𝑡
𝑥
• Sample mean m: 𝑚𝑖 = 𝑡=1 𝑖
𝑁
, 𝑖 = 1, … , d
σ𝑁 𝑡 𝑡
𝑡=1(𝑥𝑖 −𝑚𝑖 )(𝑥𝑗 −𝑚𝑗 )
• Covariance matrix S: 𝑠𝑖𝑗 = 𝑁
𝑠𝑖𝑗
• Correlation matrix R: 𝑟𝑖𝑗 = 𝑠 𝑠
𝑖 𝑗
• If features 𝑥𝑖 , 𝑥𝑗 are
𝜎12 0 ⋯ 0
• INDEPENDENT, then 𝜎𝑖𝑗 =0 diagonals are non-zero. ⋮ ⋱ ⋮
0 0 ⋯ 𝜎𝑑2
• POSITIVE correlation, 𝜎𝑖𝑗 > 0
• NEGATIVE correlation, 𝜎𝑖𝑗 < 0
Σ in Bivariate Normal
If Σ is diagonal
• Features are independent and
𝑑 1 1
• 𝑃 𝑥 = ς𝑖=1 exp(− 2 𝑥𝑖 − 𝜇𝑖 2 )
2𝜋𝜎𝑖 2𝜎𝑖
• Euclidean distance (circular view between x’s)
• If variances are also equal
1 1
• 𝑃 𝑥 = ς𝑑𝑖=1 exp(− 2 𝑥𝑖 − 𝜇𝑖 2 )
2𝜋𝜎 2𝜎
Σ and µ relations on topological maps of Gaussian surface
Figures from CS 434s/541a Pattern Recognition, Uni of Western Ontario
Discriminant functions for classification
• Classifier can be viewed as m discriminant functions and the
classification is based on selecting the largest discriminant:
• 𝑔𝑖 𝑥 = 𝑃 𝑐𝑖 𝑋 = 𝑃 𝑋 𝑐𝑖 P(𝑐𝑖 )/P(x)
• For normal density, it is more convenient to work on logarithms
• 𝑔𝑖 𝑥 = 𝑙𝑜𝑔𝑃 𝑋 𝑐𝑖 + logP(𝑐𝑖 )
1 𝑑 1
• 𝑔𝑖 𝑥 = − (𝑥 − 𝜇𝑖 )𝑇 Σ𝑖−1 𝑥 − 𝜇𝑖 − 𝑙𝑜𝑔2𝜋 − 𝑙𝑜𝑔 Σ𝑖 + logP(𝑐𝑖 )
2 2 2
2
Case Σ𝑖 = 𝜎 Ι
• Features are independent with different means and equal variances
• 𝜎 2Ι = 𝜎 2 1 0
0 1
• 1 𝑑
𝑔𝑖 𝑥 = − (𝑥 − 𝜇𝑖 )𝑇 Σ−1 𝑥 − 𝜇𝑖 − 𝑙𝑜𝑔2𝜋 − 𝑙𝑜𝑔 Σ
2 2
1
2
+ logP(𝑐𝑖 )
1 1
= − (𝑥 − 𝜇𝑖 )𝑇 ( 2 Ι) 𝑥 − 𝜇𝑖 +logP(𝑐𝑖 )
2 𝜎
1
= − 2 (𝑥 − 𝜇𝑖 )𝑇 𝑥 − 𝜇𝑖 + logP(𝑐𝑖 )
2𝜎
1 𝑇 𝑥 − 𝑥 𝑇 𝜇 − 𝜇 𝑇 𝑥 + 𝜇𝑇 𝜇 )
=− (𝑥 𝑖 𝑖 𝑖 𝑖
2𝜎 2
1
= − 2 −2𝜇𝑖 𝑇 𝑥 + 𝜇𝑖𝑇 𝜇𝑖 + logP(𝑐𝑖 )
2𝜎
• Discriminant function is linear wrt x
• 𝑔𝑖 𝑥 = 𝑤𝑖𝑇 𝑥 + 𝑤𝑖0
2
Case Σ𝑖 = 𝜎 Ι
• Decision boundaries 𝑔𝑖 𝑥 = 𝑔𝑗 𝑥 are linear
• when x has a dimension of 2, lines
• When x has a dimension of 3, plane
• Larger than 3, hyperplanes
Example
1 4 −2 30
• 𝜇1 = , 𝜇2 = , 𝜇3 = , Σ1 = Σ2 = Σ3 =
2 6 4 03
1 1
• 𝑝 𝑐1 = 𝑝 𝑐2 = , 𝑝 𝑐3 =
4 2
𝜇𝑖𝑇 𝜇𝑖𝑇 𝜇𝑖
• 𝑔𝑖 𝑥 = 𝑥 + (− + logP(𝑐𝑖 ))
𝜎2 𝜎2
• First form the discriminants for 𝑔1 𝑥 , 𝑔2 𝑥 , 𝑔3 𝑥
• Then solve 𝑔𝑖 𝑥 = 𝑔𝑗 𝑥
Figure from CS 434s/541a Pattern Recognition, Uni of Western Ontario
Case Σ𝑖 = Σ
• Features are not necessarily independent
• Covariance matrices are equal but arbitrary
1
• 𝑔𝑖 𝑥 = − 2 (𝑥 − 𝜇𝑖 )𝑇 Σ−1 𝑥 − 𝜇𝑖 + logP(𝑐𝑖 )
1
= − 𝑥 𝑇 Σ −1 𝑥 − 2𝜇𝑖𝑇 Σ −1 𝑥 + 𝜇𝑖𝑇 Σ −1 𝜇𝑖 + logP(𝑐𝑖 )
2
𝜇𝑖𝑇 Σ−1 𝜇𝑖
= 𝜇𝑖𝑇 Σ −1 𝑥 + (logP 𝑐𝑖 − )
2
• This is also linear
• 𝑤𝑖𝑇 𝑥 + 𝑤𝑖0
General case
1 1
• 𝑔𝑖 𝑥 = − (𝑥 − 𝜇𝑖 )𝑇 Σ𝑖−1 𝑥 − 𝜇𝑖 − 𝑙𝑜𝑔 Σ𝑖 + logP(𝑐𝑖 )
2 2
1 1
= − (𝑥 𝑇 Σ𝑖−1 𝑥 − 2𝜇𝑖𝑇 Σ𝑖−1 𝑥 + 𝜇𝑖𝑇 Σ𝑖−1 𝜇𝑖 ) − 𝑙𝑜𝑔 Σ𝑖 + logP(𝑐𝑖 )
2 2
= 𝑥 𝑇 Wx + 𝑤 𝑇 𝑥 + 𝑤𝑖0
• Discriminant is quadratic.
• Decision boundaries are ellipses and parabolloids
Model complexity - Bias - Variance
• As we increase complexity, bias decreases and variance increases
• Assume simple models to control variance (regularization)