0% found this document useful (0 votes)
3 views14 pages

Normalization v4

This document provides an introduction to Batch Normalization, explaining its role in optimizing deep learning models by normalizing features to improve convergence rates. It discusses the mathematical foundations of batch normalization, including the calculation of mean and standard deviation, and highlights its importance in addressing internal covariate shift. Additionally, it references various normalization techniques and their respective research papers for further reading.

Uploaded by

Chiang Evan
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views14 pages

Normalization v4

This document provides an introduction to Batch Normalization, explaining its role in optimizing deep learning models by normalizing features to improve convergence rates. It discusses the mathematical foundations of batch normalization, including the calculation of mean and standard deviation, and highlights its importance in addressing internal covariate shift. Additionally, it references various normalization techniques and their respective research papers for further reading.

Uploaded by

Chiang Evan
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Quick Introduction of

Batch Normalization
Hung-yi Lee 李宏毅

1
Changing Landscape
w2 Loss L

smooth
w1
+∆y small 𝑤1 +∆𝑤1
𝑦ො 𝑒 𝑦 + 𝑥1 1, 2 ……
+∆e small
small 𝑏
𝑤2
𝐿 = ෍𝑒
1 𝑥2
+∆𝐿
small 2
Changing Landscape
w2 Loss L w2 Loss L
steep

smooth
w1 w1
+∆y large 𝑤1
𝑦ො 𝑒 𝑦 + 𝑥1 1, 2 ……
+∆e small same
large 𝑏
𝑤2 range
𝐿 = ෍𝑒 +∆𝑤2
1 𝑥2 100, 200 ……
+∆𝐿
large large 3
Feature Normalization
𝒙𝟏 𝒙𝟐 𝒙𝟑 𝒙𝒓 𝒙𝑹
𝒙1𝟏 𝒙1𝟐 For each
𝒙𝟏2 𝒙𝟐2 dimension 𝑖:
…… …… mean: 𝑚𝑖
……

……

……

……

……
standard
deviation: 𝜎𝑖
𝒓
𝒓 𝒙𝑖 − 𝑚𝑖 The means of all dims are 0,
෥𝑖 ←
𝒙
𝜎𝑖 and the variances are all 1
In general, feature normalization makes gradient descent
converge faster. 4
Considering Deep Learning
Different dims have different ranges.

Sigmoid
෥𝟏
𝒙 𝑊1 𝒛𝟏 𝑎1 𝑊2 ……

Sigmoid
෥𝟐
𝒙 𝑊1 𝒛𝟐 𝒂𝟐 𝑊2 ……
Sigmoid

෥𝟑
𝒙 𝑊1 𝒛𝟑 𝒂𝟑 𝑊2 ……

Also difficult to optimize


Feature Also need
Normalization normalization 5
Considering Deep Learning

3
1
𝝁 = ෍ 𝒛𝒊
෥𝟏
𝒙 𝑊1 𝒛𝟏 3
𝑖=1

3
1
෥𝟐
𝒙 𝑊1 𝒛𝟐 𝝈= ෍ 𝒛𝒊 − 𝝁 2
3
𝑖=1

෥𝟑
𝒙 𝑊1 𝒛𝟑

𝝁 𝝈
6
Considering Deep Learning 𝒛𝒊
−𝝁
𝒊
𝒛෤ =
This is a large network! 𝝈

Sigmoid
෥𝟏
𝒙 𝑊1 𝒛𝟏 𝒛෤ 𝟏 𝒂𝟏
∆ ∆ ∆

Sigmoid
෥𝟐
𝒙 𝑊1 𝒛𝟐 𝒛෤ 𝟐 𝒂𝟐
∆ ∆

Sigmoid
෥𝟑
𝒙 𝑊1 𝒛𝟑 𝒛෤ 𝟑 𝒂𝟑
∆ ∆
∆ ∆ Consider a batch
𝝁 and 𝝈
𝝁 𝝈 Batch Normalization
depends on 𝒛𝒊
7
𝒛𝒊−𝝁
𝒛෤ 𝒊 =
𝝈
Batch normalization
𝒛ො 𝒊 = 𝜸⨀෤𝒛𝒊 + 𝜷

෥𝟏
𝒙 𝑊1 𝒛𝟏 𝒛෤ 𝟏 𝒛ො 𝟏

෥𝟐
𝒙 𝑊1 𝒛𝟐 𝒛෤ 𝟐 𝒛ො 𝟐

෥𝟑
𝒙 𝑊1 𝒛𝟑 𝒛෤ 𝟑 𝒛ො 𝟑

𝝁 and 𝝈
𝝁 𝝈 𝜷 𝜸
depends on 𝒛𝒊
8
Batch normalization – Testing
𝒛−𝝁 𝝁ഥ
𝒛෤ =

𝝈 𝝈

𝒙 𝑊1 𝒛 𝒛෤ ……
𝝁, 𝝈 are from batch?

We do not always have batch at testing stage.

Computing the moving average of 𝝁 and 𝝈 of the batches


during training.
𝝁𝟏 𝝁𝟐 𝝁𝟑 …… 𝝁𝒕

𝝁 + 1 − 𝑝 𝝁𝒕
ഥ ← 𝑝ഥ
𝝁
Batch normalization
Original paper: [Link]

10
Internal Covariate Shift?
How Does Batch Normalization Help Optimization?
[Link]

𝒙 𝐴 𝒂 𝐵 𝒃 ……

update Good for 𝒂,


But not for 𝒂’

𝒙 𝐴′ 𝒂′ 𝐵′ 𝒃′ ……

Batch normalization make 𝒂 and 𝒂’ have similar statistics.


Experimental results do not support the above idea. 11
Internal Covariate Shift?
How Does Batch Normalization Help Optimization?
[Link]

Experimental results (and theoretically analysis) support batch


normalization change the landscape of error surface.

serendipitous (偶然的)

penicillin
12
To learn more ……
• Batch Renormalization
• [Link]
• Layer Normalization
• [Link]
• Instance Normalization
• [Link]
• Group Normalization
• [Link]
• Weight Normalization
• [Link]
• Spectrum Normalization
• [Link]
13
14

You might also like