0% found this document useful (0 votes)
17 views16 pages

Multivariate Methods in Machine Learning

The document covers multivariate methods in machine learning, focusing on normal density functions, parameter estimation, and discriminant functions for classification. It explains the relationships between features, covariance, and correlation, and discusses the implications of different covariance structures on decision boundaries. Additionally, it addresses model complexity, bias, and variance in the context of regularization strategies.

Uploaded by

esadyazar66
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
17 views16 pages

Multivariate Methods in Machine Learning

The document covers multivariate methods in machine learning, focusing on normal density functions, parameter estimation, and discriminant functions for classification. It explains the relationships between features, covariance, and correlation, and discusses the implications of different covariance structures on decision boundaries. Additionally, it addresses model complexity, bias, and variance in the context of regularization strategies.

Uploaded by

esadyazar66
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

BLG 454E Learning From

Data
FALL 2022-2023
Multivariate Methods
(Slides are Prepared by Assoc. Prof. Yusuf Yaslan
& Assist. Prof. Ayşe Tosun)

Lecture Notes from Alpaydın 2010 Introduction to Machine Learning 2e © The MIT Press (V1.0) AND
O. Velksler’s Lecture Notes from CS 434s/541a Pattern Recognition, Uni of Western Ontario
Univariate Normal Density

• So far, we have dealt with univariate x (dimension of 1)


1 1 𝑥−𝜇 2
• 𝑃 𝑥 = exp(− )
2𝜋𝜎 2 𝜎
1 𝑁
• 𝜇𝑀𝐿𝐸 = 𝑚 = σ𝑖=1 𝑥𝑖
𝑛
1 𝑁
• 𝜎 𝑀𝐿𝐸 = 𝑠 = σ𝑖=1(𝑥𝑖 − 𝑚)2
2 2
𝑛

2
Multivariate Normal Density
• What if we have several features 𝑥1 , 𝑥2 , 𝑥3 , … , 𝑥𝑑 𝑥11
X= ⋮
⋯ 𝑥𝑑1
⋱ ⋮
• Each normally distributed 𝑥1𝑁 ⋯ 𝑥𝑑𝑁

• Different variances
• Different means
• May be dependent or independent of each other
1 1
• 𝑃 𝑥 = 𝑑/2 1/2 exp(− (𝑥 − 𝜇)𝑇 Σ −1 (𝑥 − 𝜇))
2𝜋 Σ 2

Mahalonobis distance between x and mean


(elliptical curve between x’s)
Multivariate parameters
• 𝐸 𝑋 = 𝜇 = 𝜇1 , 𝜇2 , … , 𝜇𝑑 𝑇

• Covariance = 𝜎𝑖𝑗 = 𝐶𝑜𝑣 𝑋𝑖 , 𝑋𝑗


𝜎𝑖𝑗
• Correlation =Corr 𝑋𝑖 , 𝑋𝑗 = 𝜌𝑖𝑗 =
𝜎𝑖 𝜎𝑗
• Covariance matrix Σ = 𝐸 (𝑋 − 𝜇)(𝑋 − 𝜇)𝑇
𝜎12 𝜎12 ⋯ 𝜎1𝑑
= ⋮ ⋱ ⋮
𝜎𝑑1 𝜎𝑑2 ⋯ 𝜎𝑑2
Multivariate parameter estimation
σ𝑁 𝑡
𝑥
• Sample mean m: 𝑚𝑖 = 𝑡=1 𝑖
𝑁
, 𝑖 = 1, … , d
σ𝑁 𝑡 𝑡
𝑡=1(𝑥𝑖 −𝑚𝑖 )(𝑥𝑗 −𝑚𝑗 )
• Covariance matrix S: 𝑠𝑖𝑗 = 𝑁
𝑠𝑖𝑗
• Correlation matrix R: 𝑟𝑖𝑗 = 𝑠 𝑠
𝑖 𝑗
• If features 𝑥𝑖 , 𝑥𝑗 are
𝜎12 0 ⋯ 0
• INDEPENDENT, then 𝜎𝑖𝑗 =0 diagonals are non-zero. ⋮ ⋱ ⋮
0 0 ⋯ 𝜎𝑑2
• POSITIVE correlation, 𝜎𝑖𝑗 > 0
• NEGATIVE correlation, 𝜎𝑖𝑗 < 0
Σ in Bivariate Normal
If Σ is diagonal
• Features are independent and
𝑑 1 1
• 𝑃 𝑥 = ς𝑖=1 exp(− 2 𝑥𝑖 − 𝜇𝑖 2 )
2𝜋𝜎𝑖 2𝜎𝑖
• Euclidean distance (circular view between x’s)

• If variances are also equal


1 1
• 𝑃 𝑥 = ς𝑑𝑖=1 exp(− 2 𝑥𝑖 − 𝜇𝑖 2 )
2𝜋𝜎 2𝜎
Σ and µ relations on topological maps of Gaussian surface

Figures from CS 434s/541a Pattern Recognition, Uni of Western Ontario


Discriminant functions for classification
• Classifier can be viewed as m discriminant functions and the
classification is based on selecting the largest discriminant:
• 𝑔𝑖 𝑥 = 𝑃 𝑐𝑖 𝑋 = 𝑃 𝑋 𝑐𝑖 P(𝑐𝑖 )/P(x)
• For normal density, it is more convenient to work on logarithms
• 𝑔𝑖 𝑥 = 𝑙𝑜𝑔𝑃 𝑋 𝑐𝑖 + logP(𝑐𝑖 )

1 𝑑 1
• 𝑔𝑖 𝑥 = − (𝑥 − 𝜇𝑖 )𝑇 Σ𝑖−1 𝑥 − 𝜇𝑖 − 𝑙𝑜𝑔2𝜋 − 𝑙𝑜𝑔 Σ𝑖 + logP(𝑐𝑖 )
2 2 2
2
Case Σ𝑖 = 𝜎 Ι
• Features are independent with different means and equal variances
• 𝜎 2Ι = 𝜎 2 1 0
0 1
• 1 𝑑
𝑔𝑖 𝑥 = − (𝑥 − 𝜇𝑖 )𝑇 Σ−1 𝑥 − 𝜇𝑖 − 𝑙𝑜𝑔2𝜋 − 𝑙𝑜𝑔 Σ
2 2
1
2
+ logP(𝑐𝑖 )
1 1
= − (𝑥 − 𝜇𝑖 )𝑇 ( 2 Ι) 𝑥 − 𝜇𝑖 +logP(𝑐𝑖 )
2 𝜎
1
= − 2 (𝑥 − 𝜇𝑖 )𝑇 𝑥 − 𝜇𝑖 + logP(𝑐𝑖 )
2𝜎
1 𝑇 𝑥 − 𝑥 𝑇 𝜇 − 𝜇 𝑇 𝑥 + 𝜇𝑇 𝜇 )
=− (𝑥 𝑖 𝑖 𝑖 𝑖
2𝜎 2
1
= − 2 −2𝜇𝑖 𝑇 𝑥 + 𝜇𝑖𝑇 𝜇𝑖 + logP(𝑐𝑖 )
2𝜎
• Discriminant function is linear wrt x
• 𝑔𝑖 𝑥 = 𝑤𝑖𝑇 𝑥 + 𝑤𝑖0
2
Case Σ𝑖 = 𝜎 Ι
• Decision boundaries 𝑔𝑖 𝑥 = 𝑔𝑗 𝑥 are linear
• when x has a dimension of 2, lines
• When x has a dimension of 3, plane
• Larger than 3, hyperplanes
Example
1 4 −2 30
• 𝜇1 = , 𝜇2 = , 𝜇3 = , Σ1 = Σ2 = Σ3 =
2 6 4 03
1 1
• 𝑝 𝑐1 = 𝑝 𝑐2 = , 𝑝 𝑐3 =
4 2
𝜇𝑖𝑇 𝜇𝑖𝑇 𝜇𝑖
• 𝑔𝑖 𝑥 = 𝑥 + (− + logP(𝑐𝑖 ))
𝜎2 𝜎2
• First form the discriminants for 𝑔1 𝑥 , 𝑔2 𝑥 , 𝑔3 𝑥
• Then solve 𝑔𝑖 𝑥 = 𝑔𝑗 𝑥

Figure from CS 434s/541a Pattern Recognition, Uni of Western Ontario


Case Σ𝑖 = Σ
• Features are not necessarily independent
• Covariance matrices are equal but arbitrary
1
• 𝑔𝑖 𝑥 = − 2 (𝑥 − 𝜇𝑖 )𝑇 Σ−1 𝑥 − 𝜇𝑖 + logP(𝑐𝑖 )
1
= − 𝑥 𝑇 Σ −1 𝑥 − 2𝜇𝑖𝑇 Σ −1 𝑥 + 𝜇𝑖𝑇 Σ −1 𝜇𝑖 + logP(𝑐𝑖 )
2
𝜇𝑖𝑇 Σ−1 𝜇𝑖
= 𝜇𝑖𝑇 Σ −1 𝑥 + (logP 𝑐𝑖 − )
2
• This is also linear
• 𝑤𝑖𝑇 𝑥 + 𝑤𝑖0
General case
1 1
• 𝑔𝑖 𝑥 = − (𝑥 − 𝜇𝑖 )𝑇 Σ𝑖−1 𝑥 − 𝜇𝑖 − 𝑙𝑜𝑔 Σ𝑖 + logP(𝑐𝑖 )
2 2
1 1
= − (𝑥 𝑇 Σ𝑖−1 𝑥 − 2𝜇𝑖𝑇 Σ𝑖−1 𝑥 + 𝜇𝑖𝑇 Σ𝑖−1 𝜇𝑖 ) − 𝑙𝑜𝑔 Σ𝑖 + logP(𝑐𝑖 )
2 2
= 𝑥 𝑇 Wx + 𝑤 𝑇 𝑥 + 𝑤𝑖0
• Discriminant is quadratic.
• Decision boundaries are ellipses and parabolloids
Model complexity - Bias - Variance
• As we increase complexity, bias decreases and variance increases
• Assume simple models to control variance (regularization)

You might also like