0% found this document useful (0 votes)
3 views61 pages

3 BackProp+Regularization

The document outlines key concepts in deep learning, focusing on backpropagation, overfitting, and regularization. It covers gradient descent, forward propagation, and the bias/variance tradeoff, providing mathematical notations and formulas relevant to logistic regression and neural network inference. The instructor for the course is Dr. David C. Anastasiu from Santa Clara University.

Uploaded by

Super Jeromy
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views61 pages

3 BackProp+Regularization

The document outlines key concepts in deep learning, focusing on backpropagation, overfitting, and regularization. It covers gradient descent, forward propagation, and the bias/variance tradeoff, providing mathematical notations and formulas relevant to logistic regression and neural network inference. The instructor for the course is Dr. David C. Anastasiu from Santa Clara University.

Uploaded by

Super Jeromy
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

CSEN 342

Deep Learning

BackProp, Overfitting, Regularization

Dr. David C. Anastasiu


Santa Clara University
Outline
Gradient Descent Recap

Forward Propagation

Computing Derivatives

Backward Propagation

Bias/Variance Tradeoff

Neural Network Regularization

Instructor: David C. Anastasiu CSEN 342: Deep Learning 2


Notation
Data: 𝑥𝑥, 𝑦𝑦 𝑥𝑥 ∈ ℝ𝑛𝑛𝑥𝑥 , 𝑥𝑥 has dimensionality 𝑛𝑛𝑥𝑥 , 𝑦𝑦 ∈ {0,1}
𝑚𝑚 training examples: 𝑥𝑥 1 , 𝑦𝑦 1 , 𝑥𝑥 2 , 𝑦𝑦 2 , … , 𝑥𝑥 𝑚𝑚 , 𝑦𝑦 𝑚𝑚

1
𝑎𝑎1
𝑥𝑥1 = 𝑎𝑎1
0
1
𝑎𝑎2
2
𝑦𝑦� 𝑋𝑋 = 𝑥𝑥 (1) 𝑥𝑥 (2) … 𝑥𝑥 (𝑚𝑚)
𝑥𝑥2 = 𝑎𝑎2
0
𝑎𝑎3
1
𝑎𝑎1

𝑥𝑥3 = 𝑎𝑎3
0
1
𝑎𝑎4

𝑌𝑌 = 𝑦𝑦 1 , 𝑦𝑦 2 , … , 𝑦𝑦 𝑚𝑚

Instructor: David C. Anastasiu CSEN 342: Deep Learning 3


Gradient Descent [Logistic Regression]
𝑇𝑇 1
• Hypothesis: 𝑦𝑦� = 𝑎𝑎 = 𝜎𝜎 𝑧𝑧 , 𝑧𝑧 = 𝑤𝑤 𝑥𝑥 + 𝑏𝑏, 𝜎𝜎 𝑧𝑧 = 1+𝑒𝑒 −𝑧𝑧
• Cost Function:
ℒ 𝑦𝑦,
� 𝑦𝑦 = −(𝑦𝑦 log(𝑦𝑦)
� + (1 − 𝑦𝑦) log(1 − 𝑦𝑦))

𝑚𝑚
1
𝐽𝐽 𝑤𝑤, 𝑏𝑏 = 𝑚𝑚
� ℒ(𝑦𝑦� 𝑖𝑖 , 𝑦𝑦 (𝑖𝑖) )
𝑖𝑖=1
𝑚𝑚
1
= − �
𝑚𝑚
𝑖𝑖=1
𝑦𝑦 (𝑖𝑖) log 𝑦𝑦� 𝑖𝑖 + (1 − 𝑦𝑦 (𝑖𝑖) ) log(1 − 𝑦𝑦� 𝑖𝑖 )

• Want to find 𝑤𝑤, 𝑏𝑏 such that 𝐽𝐽(𝑤𝑤, 𝑏𝑏) is minimized

Instructor: David C. Anastasiu CSEN 342: Deep Learning 4


Gradient Descent [Logistic Regression]
• Initialize 𝑤𝑤
• Repeat: 𝑱𝑱(𝒘𝒘)

𝑦𝑦� (𝑖𝑖) ≔ 𝜎𝜎 𝑤𝑤 𝑇𝑇 𝑥𝑥 (𝑖𝑖) + 𝑏𝑏 , ∀𝑖𝑖 = 1, … , 𝑚𝑚


𝜕𝜕𝜕𝜕(𝑤𝑤, 𝑏𝑏)
𝑤𝑤 ≔ 𝑤𝑤 − 𝛼𝛼
𝜕𝜕𝜕𝜕
𝜕𝜕𝜕𝜕(𝑤𝑤, 𝑏𝑏) 𝒘𝒘
𝑏𝑏 ≔ 𝑏𝑏 − 𝛼𝛼
𝜕𝜕𝜕𝜕

…until 𝑤𝑤, 𝑏𝑏 do not change much.


Instructor: David C. Anastasiu CSEN 342: Deep Learning 5
Forward Propagation

Instructor: David C. Anastasiu CSEN 342: Deep Learning 6


Neural Network Inference
• Each non-input node in the neural network is a separate prediction problem
1
𝑎𝑎1
𝑥𝑥1 = 0
𝑎𝑎1 1
𝑎𝑎2 1 1 𝑇𝑇 [1] [1] 1 1
𝑥𝑥2 = 𝑎𝑎2
0 𝑎𝑎1
2
𝑦𝑦� 𝑧𝑧1 = 𝑤𝑤1 𝑎𝑎[0] + 𝑏𝑏1 , 𝑎𝑎 1 = 𝑔𝑔[1] (𝑧𝑧1 ) = 𝜎𝜎(𝑧𝑧1 )
1
𝑎𝑎3
𝑥𝑥3 = 𝑎𝑎3
0
1
𝑎𝑎4


1
𝑎𝑎1
𝑥𝑥1 = 0
𝑎𝑎1 1
𝑎𝑎2 1 1 𝑇𝑇 [1] [1] 1 1
𝑥𝑥2 = 𝑎𝑎2
0 𝑎𝑎1
2
𝑦𝑦� 𝑧𝑧4 = 𝑤𝑤4 𝑎𝑎[0] + 𝑏𝑏4 , 𝑎𝑎 4 = 𝑔𝑔[1] (𝑧𝑧1 ) = 𝜎𝜎(𝑧𝑧4 )
1
𝑎𝑎3
𝑥𝑥3 = 𝑎𝑎3
0
1
𝑎𝑎4

Instructor: David C. Anastasiu CSEN 342: Deep Learning 7


Neural Network Inference
• Each non-input node in the neural network is a separate prediction problem
1
𝑎𝑎1
𝑥𝑥1 = 0
𝑎𝑎1 1
𝑎𝑎2 2 2 𝑇𝑇 [2] [2] 2 2
𝑥𝑥2 = 𝑎𝑎2
0 𝑎𝑎1
2
𝑦𝑦� 𝑧𝑧1 = 𝑤𝑤1 𝑎𝑎[1] + 𝑏𝑏1 , 𝑎𝑎 1 = 𝑔𝑔[2] (𝑧𝑧1 ) = 𝜎𝜎(𝑧𝑧1 )
1
𝑎𝑎3
𝑥𝑥3 = 𝑎𝑎3
0
1
𝑎𝑎4

Instructor: David C. Anastasiu CSEN 342: Deep Learning 8


Neural Network Inference
• Vectorize operations layer-wise
1 1 𝑇𝑇 [1] [1] 1
𝑎𝑎1
1 𝑧𝑧1 = 𝑤𝑤1 𝑎𝑎[0] + 𝑏𝑏1 , 𝑎𝑎 1 = 𝜎𝜎(𝑧𝑧1 )
1 1 𝑇𝑇 [1] [1] 1
𝑥𝑥1 = 𝑎𝑎1
0
1
𝑎𝑎2
𝑧𝑧2 = 𝑤𝑤2 𝑎𝑎[0] + 𝑏𝑏2 , 𝑎𝑎 2 = 𝜎𝜎(𝑧𝑧2 )

𝑥𝑥2 = 𝑎𝑎2
0
1
2
𝑎𝑎1 𝑦𝑦� 1
𝑧𝑧3 = 𝑤𝑤3
1 𝑇𝑇 [1]
𝑎𝑎[0] + 𝑏𝑏3 ,
[1] 1
𝑎𝑎 3 = 𝜎𝜎(𝑧𝑧3 )
𝑎𝑎3 1 1 𝑇𝑇 [1] [1] 1
𝑥𝑥3 = 𝑎𝑎3
0 𝑧𝑧4 = 𝑤𝑤4 𝑎𝑎[0] + 𝑏𝑏4 , 𝑎𝑎 4 = 𝜎𝜎(𝑧𝑧4 )
1
𝑎𝑎4

Instructor: David C. Anastasiu CSEN 342: Deep Learning 9


Forward Propagation
• Two layer fully connected network
Given input 𝑥𝑥:
1
𝑎𝑎1 𝑎𝑎 0 = 𝑥𝑥
𝑥𝑥1 = 𝑎𝑎1
0
1
𝑎𝑎2 𝑧𝑧 1 = 𝑊𝑊 1 𝑥𝑥 + 𝑏𝑏 1
𝑥𝑥2 = 𝑎𝑎2
0
2
𝑎𝑎1 𝑦𝑦�
𝑎𝑎 1 = 𝜎𝜎(𝑧𝑧 1
1
𝑎𝑎3 )
𝑥𝑥3 = 𝑎𝑎3
0
2 2 𝑎𝑎 1 + 𝑏𝑏 2
1
𝑎𝑎4 𝑧𝑧 = 𝑊𝑊
𝑎𝑎 2 = 𝜎𝜎(𝑧𝑧 2 )
𝑦𝑦 = 𝑎𝑎 2

Instructor: David C. Anastasiu CSEN 342: Deep Learning 10


Vectorization across multiple samples
1
𝑎𝑎1
for i = 1 to m
𝑥𝑥1 = 𝑎𝑎1
0
1
𝑎𝑎2
𝑥𝑥2 = 𝑎𝑎2
0
1
𝑎𝑎1
2
𝑦𝑦� 𝑧𝑧 1 (𝑖𝑖) = 𝑊𝑊 1 𝑥𝑥 (𝑖𝑖) + 𝑏𝑏 1
𝑎𝑎3
𝑥𝑥3 = 𝑎𝑎3
0
1
𝑎𝑎 1 (𝑖𝑖) = 𝜎𝜎(𝑧𝑧 1 𝑖𝑖 )
𝑎𝑎4
𝑧𝑧 2 (𝑖𝑖) = 𝑊𝑊 2 𝑎𝑎 1 (𝑖𝑖) + 𝑏𝑏 2
𝑎𝑎 2 (𝑖𝑖) = 𝜎𝜎(𝑧𝑧 2 𝑖𝑖
)
𝑋𝑋 = 𝑥𝑥 (1)
𝑥𝑥 (2) … 𝑥𝑥 (𝑚𝑚)

𝑍𝑍 1 = 𝑊𝑊 1 𝑋𝑋 + 𝑏𝑏 1
𝐴𝐴 1 = 𝜎𝜎(𝑍𝑍 1 )
A[1] = 𝑎𝑎[1](1) 𝑎𝑎[1](2) … 𝑎𝑎[1](𝑚𝑚) 𝑍𝑍 2 = 𝑊𝑊 2
𝐴𝐴 1 + 𝑏𝑏 2
𝐴𝐴 2 = 𝜎𝜎(𝑍𝑍 2 )
Instructor: David C. Anastasiu CSEN 342: Deep Learning 11
Computing Derivatives

Instructor: David C. Anastasiu CSEN 342: Deep Learning 12


Backpropagation summary
Parameter update:

𝜕𝜕𝜕 𝜕𝜕𝜕 𝜕𝜕𝑎𝑎[𝑘𝑘]


=
𝑤𝑤 [𝑘𝑘] 𝜕𝜕𝑤𝑤 [𝑘𝑘] 𝜕𝜕𝑎𝑎[𝑘𝑘] 𝜕𝜕𝑤𝑤 [𝑘𝑘]
Upstream gradient:
𝜕𝜕ℒ
𝜕𝜕𝑎𝑎[𝑘𝑘]
𝜕𝜕𝑎𝑎[𝑘𝑘]
Local gradient
𝜕𝜕𝑤𝑤 [𝑘𝑘]

𝑔𝑔 𝑘𝑘 𝑧𝑧 𝑘𝑘
𝑎𝑎[𝑘𝑘]
𝜕𝜕𝑎𝑎[𝑘𝑘]
Local gradient
𝜕𝜕𝑎𝑎[𝑘𝑘−1]

𝑎𝑎[𝑘𝑘−1] Downstream gradient:


𝜕𝜕𝜕 𝜕𝜕𝜕 𝜕𝜕𝑎𝑎[𝑘𝑘]
= Forward pass
𝜕𝜕𝑎𝑎[𝑘𝑘−1] 𝜕𝜕𝑎𝑎[𝑘𝑘] 𝜕𝜕𝑎𝑎[𝑘𝑘−1] Backward pass
A detailed example 1
𝑓𝑓 𝑥𝑥, 𝑤𝑤 =
1 + exp − 𝑤𝑤0 𝑥𝑥0 + 𝑤𝑤1 𝑥𝑥1 + 𝑤𝑤2

Source: Stanford 231n


A detailed example 1
𝑓𝑓 𝑥𝑥, 𝑤𝑤 =
1 + exp − 𝑤𝑤0 𝑥𝑥0 + 𝑤𝑤1 𝑥𝑥1 + 𝑤𝑤2

𝑑𝑑𝑑𝑑 1 𝑑𝑑𝑑𝑑 1
𝑓𝑓 𝑥𝑥 = 𝑒𝑒 𝑥𝑥 ⇒ = 𝑒𝑒 𝑥𝑥 𝑓𝑓 𝑥𝑥 = ⇒ =− 2
𝑑𝑑𝑑𝑑 𝑥𝑥 𝑑𝑑𝑑𝑑 𝑥𝑥

𝑑𝑑𝑑𝑑 𝑑𝑑𝑑𝑑
𝑓𝑓 𝑥𝑥 = 𝑎𝑎𝑎𝑎 ⇒ = 𝑎𝑎 𝑓𝑓 𝑥𝑥 = 𝑐𝑐 + 𝑥𝑥 ⇒ =1
𝑑𝑑𝑑𝑑 𝑑𝑑𝑑𝑑

Local Upstream
gradient * gradient

1/𝑥𝑥 ′ = −1/𝑥𝑥 2
1
− ∗ 1 = −0.53
1.372
Source: Stanford 231n
A detailed example 1
𝑓𝑓 𝑥𝑥, 𝑤𝑤 =
1 + exp − 𝑤𝑤0 𝑥𝑥0 + 𝑤𝑤1 𝑥𝑥1 + 𝑤𝑤2

𝑑𝑑𝑑𝑑 1 𝑑𝑑𝑑𝑑 1
𝑓𝑓 𝑥𝑥 = 𝑒𝑒 𝑥𝑥 ⇒ = 𝑒𝑒 𝑥𝑥 𝑓𝑓 𝑥𝑥 = ⇒ =− 2
𝑑𝑑𝑑𝑑 𝑥𝑥 𝑑𝑑𝑑𝑑 𝑥𝑥

𝑑𝑑𝑑𝑑 𝑑𝑑𝑑𝑑
𝑓𝑓 𝑥𝑥 = 𝑎𝑎𝑎𝑎 ⇒ = 𝑎𝑎 𝑓𝑓 𝑥𝑥 = 𝑐𝑐 + 𝑥𝑥 ⇒ =1
𝑑𝑑𝑑𝑑 𝑑𝑑𝑑𝑑

Local Upstream
gradient * gradient

Source: Stanford 231n


A detailed example 1
𝑓𝑓 𝑥𝑥, 𝑤𝑤 =
1 + exp − 𝑤𝑤0 𝑥𝑥0 + 𝑤𝑤1 𝑥𝑥1 + 𝑤𝑤2

𝑑𝑑𝑑𝑑 1 𝑑𝑑𝑑𝑑 1
𝑓𝑓 𝑥𝑥 = 𝑒𝑒 𝑥𝑥 ⇒ = 𝑒𝑒 𝑥𝑥 𝑓𝑓 𝑥𝑥 = ⇒ =− 2
𝑑𝑑𝑑𝑑 𝑥𝑥 𝑑𝑑𝑑𝑑 𝑥𝑥

𝑑𝑑𝑑𝑑 𝑑𝑑𝑑𝑑
𝑓𝑓 𝑥𝑥 = 𝑎𝑎𝑎𝑎 ⇒ = 𝑎𝑎 𝑓𝑓 𝑥𝑥 = 𝑐𝑐 + 𝑥𝑥 ⇒ =1
𝑑𝑑𝑑𝑑 𝑑𝑑𝑑𝑑

Local Upstream
gradient * gradient

exp −1 ∗ −0.53 = −0.20

Source: Stanford 231n


A detailed example 1
𝑓𝑓 𝑥𝑥, 𝑤𝑤 =
1 + exp − 𝑤𝑤0 𝑥𝑥0 + 𝑤𝑤1 𝑥𝑥1 + 𝑤𝑤2

𝑑𝑑𝑑𝑑 1 𝑑𝑑𝑑𝑑 1
𝑓𝑓 𝑥𝑥 = 𝑒𝑒 𝑥𝑥 ⇒ = 𝑒𝑒 𝑥𝑥 𝑓𝑓 𝑥𝑥 = ⇒ =− 2
𝑑𝑑𝑑𝑑 𝑥𝑥 𝑑𝑑𝑑𝑑 𝑥𝑥

𝑑𝑑𝑑𝑑 𝑑𝑑𝑑𝑑
𝑓𝑓 𝑥𝑥 = 𝑎𝑎𝑎𝑎 ⇒ = 𝑎𝑎 𝑓𝑓 𝑥𝑥 = 𝑐𝑐 + 𝑥𝑥 ⇒ =1
𝑑𝑑𝑑𝑑 𝑑𝑑𝑑𝑑

Local Upstream
gradient * gradient

Source: Stanford 231n


A detailed example 1
𝑓𝑓 𝑥𝑥, 𝑤𝑤 =
1 + exp − 𝑤𝑤0 𝑥𝑥0 + 𝑤𝑤1 𝑥𝑥1 + 𝑤𝑤2

𝑑𝑑𝑑𝑑 1 𝑑𝑑𝑑𝑑 1
𝑓𝑓 𝑥𝑥 = 𝑒𝑒 𝑥𝑥 ⇒ = 𝑒𝑒 𝑥𝑥 𝑓𝑓 𝑥𝑥 = ⇒ =− 2
𝑑𝑑𝑑𝑑 𝑥𝑥 𝑑𝑑𝑑𝑑 𝑥𝑥

𝑑𝑑𝑑𝑑 𝑑𝑑𝑑𝑑
𝑓𝑓 𝑥𝑥 = 𝑎𝑎𝑎𝑎 ⇒ = 𝑎𝑎 𝑓𝑓 𝑥𝑥 = 𝑐𝑐 + 𝑥𝑥 ⇒ =1
𝑑𝑑𝑑𝑑 𝑑𝑑𝑑𝑑

Local Upstream
gradient * gradient

Source: Stanford 231n


A detailed example 1
𝑓𝑓 𝑥𝑥, 𝑤𝑤 =
1 + exp − 𝑤𝑤0 𝑥𝑥0 + 𝑤𝑤1 𝑥𝑥1 + 𝑤𝑤2

𝑑𝑑𝑑𝑑 1 𝑑𝑑𝑑𝑑 1
𝑓𝑓 𝑥𝑥 = 𝑒𝑒 𝑥𝑥 ⇒ = 𝑒𝑒 𝑥𝑥 𝑓𝑓 𝑥𝑥 = ⇒ =− 2
𝑑𝑑𝑑𝑑 𝑥𝑥 𝑑𝑑𝑑𝑑 𝑥𝑥

𝑑𝑑𝑑𝑑 𝑑𝑑𝑑𝑑
𝑓𝑓 𝑥𝑥 = 𝑎𝑎𝑎𝑎 ⇒ = 𝑎𝑎 𝑓𝑓 𝑥𝑥 = 𝑐𝑐 + 𝑥𝑥 ⇒ =1
𝑑𝑑𝑑𝑑 𝑑𝑑𝑑𝑑

Local Upstream
gradient * gradient

Source: Stanford 231n


A detailed example 1
𝑓𝑓 𝑥𝑥, 𝑤𝑤 =
1 + exp − 𝑤𝑤0 𝑥𝑥0 + 𝑤𝑤1 𝑥𝑥1 + 𝑤𝑤2

𝑑𝑑𝑑𝑑 1 𝑑𝑑𝑑𝑑 1
𝑓𝑓 𝑥𝑥 = 𝑒𝑒 𝑥𝑥 ⇒ = 𝑒𝑒 𝑥𝑥 𝑓𝑓 𝑥𝑥 = ⇒ =− 2
𝑑𝑑𝑑𝑑 𝑥𝑥 𝑑𝑑𝑑𝑑 𝑥𝑥

𝑑𝑑𝑑𝑑 𝑑𝑑𝑑𝑑
𝑓𝑓 𝑥𝑥 = 𝑎𝑎𝑎𝑎 ⇒ = 𝑎𝑎 𝑓𝑓 𝑥𝑥 = 𝑐𝑐 + 𝑥𝑥 ⇒ =1
𝑑𝑑𝑑𝑑 𝑑𝑑𝑑𝑑

Source: Stanford 231n


A detailed example 1
𝑓𝑓 𝑥𝑥, 𝑤𝑤 =
1 + exp − 𝑤𝑤0 𝑥𝑥0 + 𝑤𝑤1 𝑥𝑥1 + 𝑤𝑤2

𝑑𝑑𝑑𝑑 1 𝑑𝑑𝑑𝑑 1
𝑓𝑓 𝑥𝑥 = 𝑒𝑒 𝑥𝑥 ⇒ = 𝑒𝑒 𝑥𝑥 𝑓𝑓 𝑥𝑥 = ⇒ =− 2
𝑑𝑑𝑑𝑑 𝑥𝑥 𝑑𝑑𝑑𝑑 𝑥𝑥

𝑑𝑑𝑑𝑑 𝑑𝑑𝑑𝑑
𝑓𝑓 𝑥𝑥 = 𝑎𝑎𝑎𝑎 ⇒ = 𝑎𝑎 𝑓𝑓 𝑥𝑥 = 𝑐𝑐 + 𝑥𝑥 ⇒ =1
𝑑𝑑𝑑𝑑 𝑑𝑑𝑑𝑑

Source: Stanford 231n


A detailed example 1
𝑓𝑓 𝑥𝑥, 𝑤𝑤 =
1 + exp − 𝑤𝑤0 𝑥𝑥0 + 𝑤𝑤1 𝑥𝑥1 + 𝑤𝑤2

𝑑𝑑𝑑𝑑 1 𝑑𝑑𝑑𝑑 1
𝑓𝑓 𝑥𝑥 = 𝑒𝑒 𝑥𝑥 ⇒ = 𝑒𝑒 𝑥𝑥 𝑓𝑓 𝑥𝑥 = ⇒ =− 2
𝑑𝑑𝑑𝑑 𝑥𝑥 𝑑𝑑𝑑𝑑 𝑥𝑥

𝑑𝑑𝑑𝑑 𝑑𝑑𝑑𝑑
𝑓𝑓 𝑥𝑥 = 𝑎𝑎𝑎𝑎 ⇒ = 𝑎𝑎 𝑓𝑓 𝑥𝑥 = 𝑐𝑐 + 𝑥𝑥 ⇒ =1
𝑑𝑑𝑑𝑑 𝑑𝑑𝑑𝑑

-0.60

Source: Stanford 231n


A detailed example 1
𝑓𝑓 𝑥𝑥, 𝑤𝑤 =
1 + exp − 𝑤𝑤0 𝑥𝑥0 + 𝑤𝑤1 𝑥𝑥1 + 𝑤𝑤2

𝑑𝑑𝑑𝑑 1 𝑑𝑑𝑑𝑑 1
𝑓𝑓 𝑥𝑥 = 𝑒𝑒 𝑥𝑥 ⇒ = 𝑒𝑒 𝑥𝑥 𝑓𝑓 𝑥𝑥 = ⇒ =− 2
𝑑𝑑𝑑𝑑 𝑥𝑥 𝑑𝑑𝑑𝑑 𝑥𝑥

𝑑𝑑𝑑𝑑 𝑑𝑑𝑑𝑑
𝑓𝑓 𝑥𝑥 = 𝑎𝑎𝑎𝑎 ⇒ = 𝑎𝑎 𝑓𝑓 𝑥𝑥 = 𝑐𝑐 + 𝑥𝑥 ⇒ =1
𝑑𝑑𝑑𝑑 𝑑𝑑𝑑𝑑

-0.40

-0.60

Source: Stanford 231n


A detailed example 1
𝑓𝑓 𝑥𝑥, 𝑤𝑤 =
1 + exp − 𝑤𝑤0 𝑥𝑥0 + 𝑤𝑤1 𝑥𝑥1 + 𝑤𝑤2

𝑑𝑑𝑑𝑑 1 𝑑𝑑𝑑𝑑 1
𝑓𝑓 𝑥𝑥 = 𝑒𝑒 𝑥𝑥 ⇒ = 𝑒𝑒 𝑥𝑥 𝑓𝑓 𝑥𝑥 = ⇒ =− 2
𝑑𝑑𝑑𝑑 𝑥𝑥 𝑑𝑑𝑑𝑑 𝑥𝑥

𝑑𝑑𝑑𝑑 𝑑𝑑𝑑𝑑
0.40 𝑓𝑓 𝑥𝑥 = 𝑎𝑎𝑎𝑎 ⇒ = 𝑎𝑎 𝑓𝑓 𝑥𝑥 = 𝑐𝑐 + 𝑥𝑥 ⇒ =1
𝑑𝑑𝑑𝑑 𝑑𝑑𝑑𝑑

-0.40

-0.60

Source: Stanford 231n


A detailed example 1
𝑓𝑓 𝑥𝑥, 𝑤𝑤 =
1 + exp − 𝑤𝑤0 𝑥𝑥0 + 𝑤𝑤1 𝑥𝑥1 + 𝑤𝑤2

𝑑𝑑𝑑𝑑 1 𝑑𝑑𝑑𝑑 1
𝑓𝑓 𝑥𝑥 = 𝑒𝑒 𝑥𝑥 ⇒ = 𝑒𝑒 𝑥𝑥 𝑓𝑓 𝑥𝑥 = ⇒ =− 2
𝑑𝑑𝑑𝑑 𝑥𝑥 𝑑𝑑𝑑𝑑 𝑥𝑥

𝑑𝑑𝑑𝑑 𝑑𝑑𝑑𝑑
0.40 𝑓𝑓 𝑥𝑥 = 𝑎𝑎𝑎𝑎 ⇒ = 𝑎𝑎 𝑓𝑓 𝑥𝑥 = 𝑐𝑐 + 𝑥𝑥 ⇒ =1
𝑑𝑑𝑑𝑑 𝑑𝑑𝑑𝑑

-0.40

-0.60

Source: Stanford 231n


A detailed example 1
𝑓𝑓 𝑥𝑥, 𝑤𝑤 =
1 + exp − 𝑤𝑤0 𝑥𝑥0 + 𝑤𝑤1 𝑥𝑥1 + 𝑤𝑤2

0.40 Can simplify computation graph

-0.40

-0.60

Sigmoid gate 𝜎𝜎 𝑥𝑥
𝜎𝜎 ′ 𝑥𝑥 = 𝜎𝜎 𝑥𝑥 1 − 𝜎𝜎 𝑥𝑥
𝜎𝜎 1 1 − 𝜎𝜎 1 = 0.73 ∗ 1 − 0.73 = 0.20

Source: Stanford 231n


Patterns in gradient flow

Source: Stanford 231n


Patterns in gradient flow

Add gate: “gradient distributor”

Source: Stanford 231n


Patterns in gradient flow

Add gate: “gradient distributor”

Multiply gate: “gradient switcher”

Source: Stanford 231n


Patterns in gradient flow

Add gate: “gradient distributor”

Multiply gate: “gradient switcher or scaler”

Max gate: “gradient router”

Source: Stanford 231n


Backward Propagation

Instructor: David C. Anastasiu CSEN 342: Deep Learning 32


Network Gradients
Forward Pass
𝑥𝑥 = 𝑎𝑎[0] 𝑧𝑧 1
= 𝑊𝑊 1
𝑎𝑎[0] + 𝑏𝑏 1 𝑎𝑎 1 = 𝜎𝜎 𝑧𝑧 1
𝑧𝑧 2
= 𝑊𝑊 2
𝑎𝑎[1] + 𝑏𝑏 2 𝑎𝑎 2 = 𝜎𝜎 𝑧𝑧 2 ℒ(𝑎𝑎[2] , 𝑦𝑦)
1
𝑊𝑊 𝑊𝑊 2

𝑏𝑏 1 𝑏𝑏 2

Backward Pass
1 2 𝑇𝑇 2
𝑑𝑑𝑑𝑑 1 = 𝑑𝑑𝑑𝑑 1 𝑎𝑎 0 𝑇𝑇 𝑑𝑑𝑑𝑑 = 𝑊𝑊 𝑑𝑑𝑑𝑑 ∗ 𝑔𝑔 1 ′ (𝑧𝑧 1
) 𝑑𝑑𝑑𝑑 2 = 𝑎𝑎 2 − 𝑦𝑦
𝑑𝑑𝑑𝑑 2 = 𝑑𝑑𝑑𝑑 2 𝑎𝑎 1 𝑇𝑇
𝑑𝑑𝑑𝑑 1 = 𝑑𝑑𝑑𝑑 1

𝑑𝑑𝑑𝑑 2 = 𝑑𝑑𝑑𝑑 2

Update

Instructor: David C. Anastasiu CSEN 342: Deep Learning 33


Dealing with vectors
𝜕𝜕𝑧𝑧1 𝜕𝜕𝑧𝑧1

𝜕𝜕𝑥𝑥1 𝜕𝜕𝑥𝑥𝑀𝑀
𝜕𝜕𝜕𝜕 = ⋮ ⋱ ⋮
𝜕𝜕𝜕𝜕 𝜕𝜕𝑧𝑧𝑁𝑁 𝜕𝜕𝑧𝑧𝑁𝑁

𝜕𝜕𝑥𝑥1 𝜕𝜕𝑥𝑥𝑀𝑀
𝑁𝑁 × 𝑀𝑀
Jacobian

𝑥𝑥 𝑓𝑓(𝑥𝑥) 𝑧𝑧
1 × 𝑀𝑀 𝜕𝜕𝜕𝜕 𝜕𝜕𝜕𝜕 𝜕𝜕𝜕𝜕 𝜕𝜕𝜕𝜕 1 × 𝑁𝑁
=
𝜕𝜕𝜕𝜕 𝜕𝜕𝜕𝜕 𝜕𝜕𝜕𝜕
𝜕𝜕𝜕𝜕
1 × 𝑀𝑀 1 × 𝑁𝑁𝑁𝑁 × 𝑀𝑀 1 × 𝑁𝑁
Vectorized Gradient Computation
𝑑𝑑𝑧𝑧 [2] = 𝑎𝑎[2] − 𝑦𝑦 𝑑𝑑𝑍𝑍 [2] = 𝐴𝐴[2] − 𝑌𝑌

1
𝑑𝑑𝑊𝑊 [2] = 𝑑𝑑𝑧𝑧 [2] 𝑎𝑎 1 𝑇𝑇 𝑑𝑑𝑊𝑊 [2] = 𝑑𝑑𝑍𝑍 [2] 𝐴𝐴 1 𝑇𝑇
𝑚𝑚
1
𝑑𝑑𝑏𝑏 [2] = 𝑑𝑑𝑧𝑧 [2] 𝑑𝑑𝑏𝑏 [2] = 𝑛𝑛𝑛𝑛. 𝑠𝑠𝑠𝑠𝑠𝑠(𝑑𝑑𝑍𝑍 2 , 𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎 = 1, 𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘 = 𝑇𝑇𝑇𝑇𝑇𝑇𝑇𝑇)
𝑚𝑚

𝑑𝑑𝑧𝑧 [1] = 𝑊𝑊 2 𝑇𝑇 𝑑𝑑𝑧𝑧 [2] ∗ 𝑔𝑔[1] ′(z 1 ) 𝑑𝑑𝑍𝑍 [1] = 𝑊𝑊 2 𝑇𝑇 𝑑𝑑𝑍𝑍 [2] ∗ 𝑔𝑔[1] ′(Z 1 )
1
𝑑𝑑𝑊𝑊 [1] = 𝑑𝑑𝑧𝑧 [1] 𝑥𝑥 𝑇𝑇 𝑑𝑑𝑊𝑊 [1] = 𝑑𝑑𝑍𝑍 [1] 𝑋𝑋 𝑇𝑇
𝑚𝑚
1
𝑑𝑑𝑏𝑏 [1] = 𝑑𝑑𝑧𝑧 [1] 𝑑𝑑𝑏𝑏 [1] = 𝑛𝑛𝑛𝑛. 𝑠𝑠𝑠𝑠𝑠𝑠(𝑑𝑑𝑍𝑍 1 , 𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎 = 1, 𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘 = 𝑇𝑇𝑇𝑇𝑇𝑇𝑇𝑇)
𝑚𝑚
Andrew Ng
Generalization to L Layers
1
𝑎𝑎1
𝐿𝐿−1
𝑥𝑥1 = 𝑎𝑎1
0
1
𝑎𝑎2
𝑎𝑎1
𝐿𝐿
𝑎𝑎1 𝑦𝑦�
[… ] 𝐿𝐿−1
𝑥𝑥2 = 0
𝑎𝑎2 1
𝑎𝑎3
𝑎𝑎2

𝑥𝑥3 = 𝑎𝑎3
0
1
𝐿𝐿−1
𝑎𝑎3
𝑎𝑎4
𝑑𝑑𝑑𝑑 [𝐿𝐿] = 𝐴𝐴[𝐿𝐿] − 𝑌𝑌
1 𝑇𝑇
𝑑𝑑𝑊𝑊 [𝐿𝐿] = 𝑑𝑑𝑍𝑍 𝐿𝐿 𝐴𝐴 𝐿𝐿
𝑚𝑚
𝑍𝑍 [1] = 𝑊𝑊 [1] 𝑋𝑋 + 𝑏𝑏 [1] 1
𝑑𝑑𝑏𝑏 [𝐿𝐿] = 𝑛𝑛𝑛𝑛. sum(d𝑍𝑍 𝐿𝐿 , 𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎 = 1, 𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘 = 𝑇𝑇𝑇𝑇𝑇𝑇𝑇𝑇)
𝐴𝐴[1] = 𝑔𝑔 1 (𝑍𝑍 1 ) 𝑚𝑚 𝑇𝑇 𝐿𝐿
𝑑𝑑𝑑𝑑 [𝐿𝐿−1] = 𝑑𝑑𝑊𝑊 𝐿𝐿 𝑑𝑑𝑍𝑍 𝐿𝐿 𝑔𝑔′ (𝑍𝑍 𝐿𝐿−1 )
𝑍𝑍 [2] = 𝑊𝑊 [2] 𝐴𝐴[1] + 𝑏𝑏 [2]


𝑇𝑇 1
𝐴𝐴[2] = 𝑔𝑔 2 (𝑍𝑍 2 ) 𝑑𝑑𝑑𝑑 [1] = 𝑑𝑑𝑊𝑊 𝐿𝐿 𝑑𝑑𝑍𝑍 2 𝑔𝑔′ (𝑍𝑍 1 )
1 𝑇𝑇

𝑑𝑑𝑊𝑊 = 𝑑𝑑𝑍𝑍 1 𝐴𝐴 1
[1]
𝐴𝐴[𝐿𝐿] = 𝑔𝑔 𝐿𝐿 𝑍𝑍 𝐿𝐿
= 𝑌𝑌� 𝑚𝑚
1
𝑑𝑑𝑏𝑏 = 𝑛𝑛𝑛𝑛. sum(d𝑍𝑍 1 , 𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎 = 1, 𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘𝑘 = 𝑇𝑇𝑇𝑇𝑇𝑇𝑇𝑇)
[1]
𝑚𝑚
Bias/Variance Tradeoff

Instructor: David C. Anastasiu CSEN 342: Deep Learning 37


Overfitting: Polynomial Regression Example
𝑦𝑦 𝑦𝑦 𝑦𝑦

𝑥𝑥 𝑥𝑥 𝑥𝑥

Overfitting: A model that is too complex (too many features) may learn a
function that fits our training data so well that it fails to generalize to data
outside out training set (predict 𝑦𝑦 on new samples).
Instructor: David C. Anastasiu CSEN 342: Deep Learning 38
Overfitting: Logistic Regression Example
𝑥𝑥2 𝑥𝑥2 𝑥𝑥2

𝑥𝑥1 𝑥𝑥1 𝑥𝑥1

Overfitting: A model that is too complex (too many features) may learn a
function that fits our training data so well that it fails to generalize to data
outside out training set (predict 𝑦𝑦 on new samples).
Instructor: David C. Anastasiu CSEN 342: Deep Learning 39
Regularization Intuition
𝑦𝑦

𝑥𝑥

What happens if we reduce the values of 𝜃𝜃3 , 𝜃𝜃4 , etc.?

Instructor: David C. Anastasiu CSEN 342: Deep Learning 40


Regularization
• Promote small values for parameters 𝜃𝜃0 , 𝜃𝜃1 , … , 𝜃𝜃𝑚𝑚
• Creates a “simpler” hypothesis (some features virtually ignored)
• Less prone to overfitting

• 𝜆𝜆 is called the shrinkage parameter


• It controls the size of the coefficients; the amount of regularization
• As 𝜆𝜆 ↓ 0, 𝐽𝐽(𝜃𝜃) is the least squares solution
• As 𝜆𝜆 ↑ ∞, 𝜃𝜃1 , … , 𝜃𝜃𝑚𝑚 → 0, i.e. we have an intercept-only model (only 𝜃𝜃0 ≠ 0)

Instructor: David C. Anastasiu CSEN 342: Deep Learning 41


Dataset Split

Instructor: David C. Anastasiu CSEN 342: Deep Learning 42


Bias and Variance in Prediction
• Overfitting = High Variance
• Fix overfitting by reducing the Variance:
• Decrease number of features
• Increase regularization
• Add data to training set
• Use simpler/smaller model
• Detect by checking gap between
𝐽𝐽𝑡𝑡𝑡𝑡 (𝜃𝜃) and 𝐽𝐽𝑣𝑣𝑣𝑣𝑣𝑣 (𝜃𝜃)

• Underfitting = High Bias


• Fix underfitting by reducing Bias:
• Use more features
• Decrease regularization
• Use a larger/more complex model
• Detect by checking error at convergence between
𝐽𝐽𝑡𝑡𝑡𝑡 (𝜃𝜃) and 𝐽𝐽𝑣𝑣𝑣𝑣𝑣𝑣 (𝜃𝜃)
Instructor: David C. Anastasiu CSEN 342: Deep Learning 43
Neural Network Regularization

Instructor: David C. Anastasiu CSEN 342: Deep Learning 44


Neural Network Regularization
• Regularization acts as a “weight decay” factor

• Prevents dead neurons

Instructor: David C. Anastasiu CSEN 342: Deep Learning 45


Types of Regularization

• Data augmentation

• Early stopping

• Standard (L-1, L-2, etc.)

• Dropout

Instructor: David C. Anastasiu CSEN 342: Deep Learning 46


Data augmentation
• Introduce transformations not adequately sampled in the training data
• Geometric: flipping, rotation, shearing, multiple crops

Image source Image source

Svetlana Lazebnik
Data augmentation
• Introduce transformations not adequately sampled in the training data
• Geometric: flipping, rotation, shearing, multiple crops
• Photometric: color transformations

Image source
Svetlana Lazebnik
Data augmentation
• Introduce transformations not adequately sampled in the training data
• Geometric: flipping, rotation, shearing, multiple crops
• Photometric: color transformations
• Other: add noise, compression artifacts, lens distortions, etc.

Image source

Svetlana Lazebnik
Data augmentation
• Introduce transformations not adequately sampled in the training data
• Limited only by your imagination and time/memory constraints!
• Avoid introducing artifacts

Image source
Svetlana Lazebnik
Data augmentation
• Introduce transformations not adequately sampled in the training data
• Limited only by your imagination and time/memory constraints!
• Avoid introducing artifacts
• Automatic augmentation strategies: AutoAugment, RandAugment

Svetlana Lazebnik
Beyond Training Error

Better optimization algorithms But we really care about error on


help reduce training loss new data - how to reduce the gap?

Fei-Fei Li & Ranjay Krishna & Danfei Xu


Early Stopping: Always do this
Train
Accuracy Val
Loss

Stop training here

Iteration Iteration

Stop training the model when accuracy on the validation set decreases
Or train for a long time, but always keep track of the model snapshot
that worked best on val

Fei-Fei Li & Ranjay Krishna & Danfei Xu


Regularization: Add term to loss

In common use:
L2 regularization (Weight decay)

L1 regularization
Elastic net (L1 + L2)

Fei-Fei Li & Ranjay Krishna & Danfei Xu


Regularization: A commonpattern
Training: Add random noise
Testing: Marginalize over the noise
Examples:
Dropout
Batch Normalization
Data Augmentation

Fei-Fei Li & Ranjay Krishna & Danfei Xu


Regularization: Fractional Pooling
Training: Use randomized pooling regions
Testing: Average predictions from several regions
Examples:
Dropout
Batch Normalization
Data Augmentation
DropConnect
Fractional Max Pooling

Graham, “Fractional Max Pooling”, arXiv 2014

Fei-Fei Li & Ranjay Krishna & Danfei Xu


Regularization: Stochastic Depth
Training: Skip some layers in the network
Testing: Use all the layer
Examples:
Dropout
Batch Normalization
Data Augmentation
DropConnect
Fractional Max Pooling
Stochastic Depth

Huang et al, “Deep Networks with Stochastic Depth”, ECCV 2016

Fei-Fei Li & Ranjay Krishna & Danfei Xu


Regularization: Cutout
Training: Set random image regions to zero
Testing: Use full image
Examples:
Dropout
Batch Normalization
Data Augmentation
DropConnect
Fractional Max Pooling
Stochastic Depth
Cutout / Random Crop
Works very well for small datasets like CIFAR,
DeVries and Taylor, “Improved Regularization of
Convolutional Neural Networks with Cutout”, arXiv 2017
less common for large datasets like ImageNet

Fei-Fei Li & Ranjay Krishna & Danfei Xu


Regularization: Mixup
Training: Train on random blends of images
Testing: Use original images
Examples:
Dropout Target label:
Batch Normalization CNN cat: 0.4
Data Augmentation dog: 0.6
DropConnect
Fractional Max Pooling
Stochastic Depth Randomly blend the pixels
Cutout / Random Crop of pairs of training images,
e.g. 40% cat, 60% dog
Mixup
Zhang et al, “mixup: Beyond Empirical Risk Minimization”, ICLR 2018

Fei-Fei Li & Ranjay Krishna & Danfei Xu


Regularization - In practice
Training: Add random noise
Testing: Marginalize over the noise
Examples:
Dropout - Consider dropout for large
Batch Normalization fully-connected layers
Data Augmentation - Batch normalization and data
DropConnect augmentation almost always a
Fractional Max Pooling good idea
Stochastic Depth - Try cutout and mixup especially
Cutout / Random Crop for small classification datasets
Mixup

Fei-Fei Li & Ranjay Krishna & Danfei Xu


References
• Andrew Ng – Introduction to Deep Learning
• Andrej Karpathy – Yes, you should understand backprop
• Andrej Karpathy – Stanford CS231n Lecture

Instructor: David C. Anastasiu CSEN 342: Deep Learning 61

You might also like