Convolutional Neural Networks Overview
Convolutional Neural Networks Overview
Mitesh M. Khapra
1/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Module 11.1 : The convolution operation
2/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Suppose we are tracking the position
of an aeroplane using a laser sensor at
discrete time intervals
Now suppose our sensor is noisy
x0 x1 x2
To obtain a less noisy estimate we
would like to average several measure-
Σ∞ ments
st = x t−a w−a = t
a=0 (x∗w) More recent measurements are more
important so we would like to take a
input filter weighted average
convolutio
n
3/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
In practice, we would only sum over a
Σ6 small window
st = x t−a w−a
The weight array (w) is known as the filter
a=0
We just slide the filter over the input and
compute the value of st based on a win-
dow around xt
w w w w w w w
−6 −5 −4 −3 −2 −1 0
W 0.01 0.01 0.02 0.02 0.04 0.4 0.5
X
1.00 1.10 1.20 1.40 1.70 1.80 1.90 2.10 2.20 2.40 2.50 2.70
S
1.80
s =x w +x w +x w +x w +x w +x w +x w
6 6 0 5 −1 4 −2 3 −3 2 −4 1 −5 0 −6
4/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
In practice, we would only sum over a
Σ6 small window
st = x t−a w−a
The weight array (w) is known as the filter
a=0
We just slide the filter over the input and
compute the value of st based on a win-
dow around xt
w w w w w w w
−6 −5 −4 −3 −2 −1 0
W 0.01 0.01 0.02 0.02 0.04 0.4 0.5
X
1.00 1.10 1.20 1.40 1.70 1.80 1.90 2.10 2.20 2.40 2.50 2.70
S
1.80 1.96
s =x w +x w +x w +x w +x w +x w +x w
6 6 0 5 −1 4 −2 3 −3 2 −4 1 −5 0 −6
4/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
In practice, we would only sum over a
Σ6 small window
st = x t−a w−a
The weight array (w) is known as the filter
a=0
We just slide the filter over the input and
compute the value of st based on a win-
dow around xt
w w w w w w w
−6 −5 −4 −3 −2 −1 0
W 0.01 0.01 0.02 0.02 0.04 0.4 0.5
X
1.00 1.10 1.20 1.40 1.70 1.80 1.90 2.10 2.20 2.40 2.50 2.70
S
s =x w +x w +x w +x w +x w +x w +x w
6 6 0 5 −1 4 −2 3 −3 2 −4 1 −5 0 −6
4/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
In practice, we would only sum over a
Σ6 small window
st = x t−a w−a
The weight array (w) is known as the filter
a=0
We just slide the filter over the input and
compute the value of st based on a win-
dow around xt
w w w w w w w
−6 −5 −4 −3 −2 −1 0
W 0.01 0.01 0.02 0.02 0.04 0.4 0.5
X
1.00 1.10 1.20 1.40 1.70 1.80 1.90 2.10 2.20 2.40 2.50 2.70
S
s =x w +x w +x w +x w +x w +x w +x w
6 6 0 5 −1 4 −2 3 −3 2 −4 1 −5 0 −6
4/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
In practice, we would only sum over a
Σ6 small window
st = x t−a w−a
The weight array (w) is known as the filter
a=0
We just slide the filter over the input and
compute the value of st based on a win-
dow around xt
w w w w w w w
−6 −5 −4 −3 −2 −1 0
W 0.01 0.01 0.02 0.02 0.04 0.4 0.5
X
1.00 1.10 1.20 1.40 1.70 1.80 1.90 2.10 2.20 2.40 2.50 2.70
S
s =x w +x w +x w +x w +x w +x w +x w
6 6 0 5 −1 4 −2 3 −3 2 −4 1 −5 0 −6
4/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
In practice, we would only sum over a
Σ6 small window
st = x t−a w−a
The weight array (w) is known as the filter
a=0
We just slide the filter over the input and
compute the value of st based on a win-
dow around xt
w w w w w w w
−6 −5 −4 −3 −2 −1 0
W 0.01 0.01 0.02 0.02 0.04 0.4 0.5
X
1.00 1.10 1.20 1.40 1.70 1.80 1.90 2.10 2.20 2.40 2.50 2.70
S
s =x w +x w +x w +x w +x w +x w +x w
6 6 0 5 −1 4 −2 3 −3 2 −4 1 −5 0 −6
4/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
In practice, we would only sum over a
Σ6 small window
st = x t−a w−a
The weight array (w) is known as the filter
a=0
We just slide the filter over the input and
compute the value of st based on a win-
dow around xt
w w w w w w w Here the input (and the kernel) is one
−6 −5 −4 −3 −2 −1 0
W 0.01 0.01 0.02 0.02 0.04 0.4 0.5 dimensional
X 1.00 1.10 1.20 1.40 1.70 1.80 1.90 2.10 2.20 2.40 2.50 2.70
s =x w +x w +x w +x w +x w +x w +x w
6 6 0 5 −1 4 −2 3 −3 2 −4 1 −5 0 −6
4/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
In practice, we would only sum over a
Σ6 small window
st = x t−a w−a
The weight array (w) is known as the filter
a=0
We just slide the filter over the input and
compute the value of st based on a win-
dow around xt
w w w w w w w Here the input (and the kernel) is one
−6 −5 −4 −3 −2 −1 0
W 0.01 0.01 0.02 0.02 0.04 0.4 0.5 dimensional
Can we use a convolutional operation on a
X 1.00 1.10 1.20 1.40 1.70 1.80 1.90 2.10 2.20 2.40 2.50 2.70 2D input also?
s =x w +x w +x w +x w +x w +x w +x w
6 6 0 5 −1 4 −2 3 −3 2 −4 1 −5 0 −6
4/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
We can think of images as 2D inputs
We would now like to use a 2D filter
(m × n)
First let us see what the 2D formula
looks like
This formula looks at all the preced-
ing neighbours (i − a, j − b)
In practice, we use the following for-
mula which looks at the succeeding
m−1
Σ Σn−1 neighbours
S ij = (I ∗ K) = ij I i+a,j+b K
a,b
a=0 b=0
5/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Inpu Let us apply this idea to a toy ex-
t Kernel ample and see the results
a b c d
w x
e f g h
y z
i j k l
Output
aw+bx+ey+fz
6/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Inpu Let us apply this idea to a toy ex-
t Kernel ample and see the results
a b c d
w x
e f g h
y z
i j k l
Output
aw+bx+ey+fz bw+cx+fy+gz
6/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Inpu Let us apply this idea to a toy ex-
t Kernel ample and see the results
a b c d
w x
e f g h
y z
i j k l
Output
6/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Inpu Let us apply this idea to a toy ex-
t Kernel ample and see the results
a b c d
w x
e f g h
y z
i j k l
Output
ew+fx+iy+jz
6/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Inpu Let us apply this idea to a toy ex-
t Kernel ample and see the results
a b c d
w x
e f g h
y z
i j k l
Output
ew+fx+iy+jz fw+gx+jy+kz
6/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Inpu Let us apply this idea to a toy ex-
t Kernel ample and see the results
a b c d
w x
e f g h
y z
i j k l
Output
6/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Let us see some examples of 2D convolutions applied to images
8/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 1 1
∗ 1 1
1 =
1 1 1
9/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 1 1
∗ 1 1
1 =
1 1 1
9/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
0 -1 0
∗ -1 5
-1 =
0 -1 0
10/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
0 -1 0
∗ -1 5
-1 =
0 -1 0
10/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 1 1
∗ 1 -8
1 =
1 1 1
11/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 1 1
∗ 1 -8
1 =
1 1 1
11/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
We will now see a working example of 2D convolution.
12/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
We just slide the kernel over the input
image
13/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
We just slide the kernel over the input
image
Each time we slide the kernel we get
one value in the output
13/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
We just slide the kernel over the input
image
Each time we slide the kernel we get
one value in the output
13/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
We just slide the kernel over the input
image
Each time we slide the kernel we get
one value in the output
13/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
We just slide the kernel over the input
image
Each time we slide the kernel we get
one value in the output
The resulting output is called a fea-
ture map.
We can use multiple filters to get mul-
tiple feature maps.
13/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Question
In the 1D case, we slide a one dimensional
filter over a one dimensional input
14/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Question A B C B A B C
14/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Question A B C B A B C
14/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Question A B C B A B C
14/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Question A B C B A B C
14/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Question A B C B A B C
14/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Question
In the 1D case, we slide a one dimensional
filter over a one dimensional input
In the 2D case, we slide a two dimen-
stional filter over a two dimensional out-
put
14/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
a b c d
Question
e f g h
In the 1D case, we slide a one dimensional
filter over a one dimensional input
i j k l
In the 2D case, we slide a two dimen-
stional filter over a two dimensional out-
put
14/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
a b c d
Question
e f g h
In the 1D case, we slide a one dimensional
filter over a one dimensional input
i j k l
In the 2D case, we slide a two dimen-
stional filter over a two dimensional out-
put
14/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
a b c d
Question
e f g h
In the 1D case, we slide a one dimensional
filter over a one dimensional input
i j k l
In the 2D case, we slide a two dimen-
stional filter over a two dimensional out-
put
14/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Question
In the 1D case, we slide a one dimensional
filter over a one dimensional input
In the 2D case, we slide a two dimen-
stional filter over a two dimensional out-
put
What would happen in the 3D case?
14/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
RGB
INPUT
15/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
RGB
OUTPUT
INPUT
15/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
RGB
OUTPUT
INPUT
15/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
RGB
OUTPUT
INPUT
15/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
RGB
OUTPUT
INPUT
15/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
RGB
OUTPUT
INPUT
15/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
RGB
OUTPUT
INPUT
15/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
RGB
OUTPUT
INPUT
15/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
RGB
OUTPUT
INPUT
15/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
RGB
OUTPUT
INPUT
15/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
RGB
OUTPUT
INPUT
15/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
RGB
OUTPUT
INPUT
15/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
RGB
OUTPUT
INPUT
15/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
RGB
OUTPUT
INPUT
15/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
RGB
OUTPUT
INPUT
15/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
RGB
INPUT
16/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
So far we have not said anything explicit about the dimensions of the
1 inputs
2 filters
3 outputs
and the relations between them
We will see how they are related but before that we will define a few quantities
17/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
We first define the following quantit-
ies
Width (W1), Height (H1) and Depth (D1)
of the original input
The Stride S (We will come back to this
later)
H
1
W
1
D
1
18/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
We first define the following quantit-
ies
Width (W1), Height (H1) and Depth (D1)
of the original input
The Stride S (We will come back to this
later)
The number of filters K
H
1
W
1
D
1
18/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
We first define the following quantit-
ies
F Width (W1), Height (H1) and Depth (D1)
of the original input
F The Stride S (We will come back to this
D
1 later)
The number of filters K
H
1
The spatial extent (F ) of each filter
(the depth of each filter is same as
the depth of each input)
W
1
D
1
18/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
We first define the following quantit-
ies
F Width (W1), Height (H1) and Depth (D1)
of the original input
F The Stride S (We will come back to this
H
D
1
2
later)
The number of filters K
H
1
The spatial extent (F ) of each filter
(the depth of each filter is same as
the depth of each input)
The output is W2 × H2 × D2 (we will
W
2 soon see a formula for computing W2,
W
1
D
2
H2 and D2)
D
1
18/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Let us compute the dimension (W , H ) of
2 2 the
output
Notice that we can’t place the kernel at the
= corners as it will cross the input boundary
pixel of
interest
19/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Let us compute the dimension (W , H ) of
2 2 the
output
Notice that we can’t place the kernel at the
= corners as it will cross the input boundary
19/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Let us compute the dimension (W , H ) of
2 2 the
output
Notice that we can’t place the kernel at the
= corners as it will cross the input boundary
This is true for all the shaded points (the kernel
crosses the input boundary)
This results in an output which is of smaller
dimensions than the input
As the size of the kernel increases, this be-
comes true for even more pixels
For example, let’s consider a 5 × 5 kernel
20/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Let us compute the dimension (W , H ) of
2 2 the
output
Notice that we can’t place the kernel at the
= corners as it will cross the input boundary
This is true for all the shaded points (the kernel
crosses the input boundary)
This results in an output which is of smaller
pixel of dimensions than the input
interest
As the size of the kernel increases, this be-
comes true for even more pixels
For example, let’s consider a 5 × 5 kernel We
have an even smaller output now
20/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Let us compute the dimension (W , H ) of
2 2 the
output
Notice that we can’t place the kernel at the
= corners as it will cross the input boundary
This is true for all the shaded points (the kernel
crosses the input boundary)
This results in an output which is of smaller
dimensions than the input
pixel of
interest As the size of the kernel increases, this be-
comes true for even more pixels
For example, let’s consider a 5 × 5 kernel We
have an even smaller output now
20/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Let us compute the dimension (W , H ) of
2 2 the
output
Notice that we can’t place the kernel at the
= corners as it will cross the input boundary
This is true for all the shaded points (the kernel
crosses the input boundary)
This results in an output which is of smaller
dimensions than the input
As the size of the kernel increases, this be-
In general, W2 = W1 − F + 1 comes true for even more pixels
For example, let’s consider a 5 × 5 kernel We
H2 = H1 − F + 1
have an even smaller output now
We will refine this formula further
20/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
What if we want the output to be of same size
as the input?
21/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
What if we want the output to be of same size
as the input?
We can use something known as padding
21/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
What if we want the output to be of same size
as the input?
We can use something known as padding
Pad the inputs with appropriate number of 0
inputs so that you can now apply the kernel at
the corners
21/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
What if we want the output to be of same size
0 0 0 0 0 0 0 0 0 as the input?
0 0
We can use something known as padding
0 0
0 0 Pad the inputs with appropriate number of 0
0 0 = inputs so that you can now apply the kernel at
0 0 the corners
0 0 Let us use pad P = 1 with a 3 × 3 kernel
0 0
0 0 0 0 0 0 0 0 0
21/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
What if we want the output to be of same size
0 0 0 0 0 0 0 0 0 as the input?
0 0
We can use something known as padding
0 0
0 0 Pad the inputs with appropriate number of 0
0 0 = inputs so that you can now apply the kernel at
0 0 the corners
0 0 Let us use pad P = 1 with a 3 × 3 kernel
0 0
This means we will add one row and one
0 0 0 0 0 0 0 0 0
column of 0 inputs at the top, bottom, left and
right
21/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
What if we want the output to be of same size
0 0 0 0 0 0 0 0 0 as the input?
0 0
We can use something known as padding
0 0
0 0 Pad the inputs with appropriate number of 0
0 0 = inputs so that you can now apply the kernel at
0 0 the corners
0 0 Let us use pad P = 1 with a 3 × 3 kernel
0 0
This means we will add one row and one
0 0 0 0 0 0 0 0 0
column of 0 inputs at the top, bottom, left and
right
21/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
What if we want the output to be of same size
0 0 0 0 0 0 0 0 0 as the input?
0 0
We can use something known as padding
0 0
0 0 Pad the inputs with appropriate number of 0
0 0 = inputs so that you can now apply the kernel at
0 0 the corners
0 0 Let us use pad P = 1 with a 3 × 3 kernel
0 0
This means we will add one row and one
0 0 0 0 0 0 0 0 0
column of 0 inputs at the top, bottom, left and
right
21/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
What if we want the output to be of same size
0 0 0 0 0 0 0 0 0 as the input?
0 0
We can use something known as padding
0 0
0 0 Pad the inputs with appropriate number of 0
0 0 = inputs so that you can now apply the kernel at
0 0 the corners
0 0 Let us use pad P = 1 with a 3 × 3 kernel
0 0
This means we will add one row and one
0 0 0 0 0 0 0 0 0
column of 0 inputs at the top, bottom, left and
right
21/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
What if we want the output to be of same size
0 0 0 0 0 0 0 0 0 as the input?
0 0
We can use something known as padding
0 0
0 0 Pad the inputs with appropriate number of 0
inputs so that you can now apply the kernel at
0 0 =
0 0
the corners
0 0 Let us use pad P = 1 with a 3 × 3 kernel
0 0
This means we will add one row and one
0 0 0 0 0 0 0 0 0
column of 0 inputs at the top, bottom, left and
right
We now have,
W2 = W1 − F + 2P + 1
H2 = H1 − F + 2P + 1
We will refine this formula further
21/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
0 0 0 0 0 0 0 0 0
What does the stride S do?
0 0
0 0 It defines the intervals at which the
0 0 filter is applied (here S = 2)
0 0 =
0 0
Here, we are essentially skipping
0 0
every 2nd pixel which will again
0 0 result in an output which is of
0 0 0 0 0 0 0 0 0 smaller dimensions
22/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
0 0 0 0 0 0 0 0 0
What does the stride S do?
0 0
0 0 It defines the intervals at which the
0 0 filter is applied (here S = 2)
0 0 =
0 0
Here, we are essentially skipping
0 0
every 2nd pixel which will again
0 0 result in an output which is of
0 0 0 0 0 0 0 0 0 smaller dimensions
22/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
0 0 0 0 0 0 0 0 0
What does the stride S do?
0 0
0 0 It defines the intervals at which the
0 0 filter is applied (here S = 2)
0 0 =
0 0
Here, we are essentially skipping
0 0
every 2nd pixel which will again
0 0 result in an output which is of
0 0 0 0 0 0 0 0 0 smaller dimensions
22/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
0 0 0 0 0 0 0 0 0
What does the stride S do?
0 0
0 0 It defines the intervals at which the
0 0 filter is applied (here S = 2)
0 0 =
0 0
Here, we are essentially skipping
0 0
every 2nd pixel which will again
0 0 result in an output which is of
0 0 0 0 0 0 0 0 0 smaller dimensions
22/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
0 0 0 0 0 0 0 0 0 What does the stride S do?
0 0
0 0 It defines the intervals at which the
0 0 filter is applied (here S = 2)
0 0 = Here, we are essentially skipping
0 0
0 0 every 2nd pixel which will again
0 0 result in an output which is of
0 0 0 0 0 0 0 0 0 smaller dimensions
22/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Finally, coming to the depth of the
output.
Each filter gives us one 2D output.
K filters will give us K such 2D out- puts
H
2 We can think of the resulting output as
K × W2 × H2 volume
H
Thus D2 = K
1
W
2
W
1
D =K
2
D
1
W −F +2P
1
W2 = S +1
H −F +2P
1
H2 = S +1
D =K
2
23/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Let us do a few exercises
H =?
2
11
∗ =
11
227
3
W =?
96 filters 2
227 Stride = 4
Padding = 0
W −F +2P
1
D =?
S 2
W2 = +1
3 H −F +2P
1 S
H2 = +1
24/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Let us do a few exercises
H =?
2
11
∗ =
11
227
3
W =?
96 filters 2
227 Stride = 4
Padding = 0
W −F +2P
1
96
W2 = S +1
3 H −F +2P
1 S
H2 = +1
24/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Let us do a few exercises
227−11
55 = 4
+ 1
11
∗ =
11
227
3
W =?
96 filters 2
227 Stride = 4
Padding = 0
W −F +2P
1
96
W2 = S +1
3 H −F +2P
1 S
H2 = +1
24/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Let us do a few exercises
227−11
55 = 4
+ 1
11
∗ =
11
227
3
227−114
96 filters 55 = + 1
227 Stride = 4
Padding = 0
W −F +2P
1
96
W2 = S +1
3 H −F +2P
1 S
H2 = +1
24/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Let us do a few exercises
H =?
2
5
∗ =
5
32
1
W =?
6 filters 2
32 Stride = 1
Padding = 0
W −F +2P
1
D =?
S 2
W2 = +1
1 H −F +2P
1 S
H2 = +1
25/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Let us do a few exercises
H =?
2
5
∗ =
5
32
1
W =?
6 filters 2
32 Stride = 1
Padding = 0
W −F +2P
1
6
W2 = S +1
1 H −F +2P
1 S
H2 = +1
25/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Let us do a few exercises
28 = 32−5 1+ 1
5
∗ =
5
32
1
W =?
6 filters 2
32 Stride = 1
Padding = 0
W −F +2P
1
6
W2 = S +1
1 H −F +2P
1 S
H2 = +1
25/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Let us do a few exercises
28 = 32−5 1+ 1
5
∗ =
5
32
1
28 = 32−5 1+ 1
6 filters
32 Stride = 1
Padding = 0
W −F +2P
1
6
W2 = S +1
1 H −F +2P
1 S
H2 = +1
25/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Module 11.3 : Convolutional Neural Networks
26/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Putting things into perspective
What is the connection between this operation (convolution) and neural net-
works?
27/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Putting things into perspective
What is the connection between this operation (convolution) and neural net-
works?
We will try to understand this by considering the task of “image classification”
27/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
28/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Features
Raw pixels
28/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Features
Raw pixels
car, bus, monument, flower
28/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Features
Raw pixels
car, bus, monument, flower
28/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Features
Raw pixels
car, bus, monument, flower
Edge Detector
28/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Features
Raw pixels
car, bus, monument, flower
Edge Detector
car, bus, monument, flower
28/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Features
Raw pixels
car, bus, monument, flower
Edge Detector
car, bus, monument, flower
28/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Features
Raw pixels
car, bus, monument, flower
Edge Detector
car, bus, monument, flower
SIFT/HOG
28/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Features
Raw pixels
car, bus, monument, flower
Edge Detector
car, bus, monument, flower
SIFT/HOG
car, bus, monument, flower
28/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Features
Raw pixels
car, bus, monument, flower
Edge Detector
car, bus, monument, flower
SIFT/HOG
car, bus, monument, flower
28/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Input Features Classifier
0 0 0 0 0
0 1 1 1 0
0 1 -8 1 0
0 1 1 1 0
0 0 0 0 0
Instead of using handcrafted kernels such as edge detectors can we learn meaningful ker-
nels/filters in addition to learning the weights of the classifier? 29/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Input Features Classifier
0 0 0 0 0
0 1 1 1 0
0 1 -8 1 0
0 1 1 1 0
0 0 0 0 0
-1.21358689e-03
3.23652686e-03 ··· ···
-2.06615720e-02
... ..
. .
-1.52757822e-03
... ···..
2.36130832e-03 .. . ···
-1.19824838e-02
-8.25322699e-04 . .
. ···
-5.14897937e-03 ···
-9.90395527e-03
Instead of using handcrafted kernels such as edge detectors can we learn meaningful ker-
nels/filters in addition to learning the weights of the classifier? 29/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Input Features Classifier
0 0 0 0 0
0 1 1 1 0
0 1 -8 1 0
0 1 1 1 0
0 0 0 0 0
-1.21358689e-03
3.23652686e-03 ··· ···
-2.06615720e-02
...
-1.52757822e-03
..
. . Learn these weights
... ···..
2.36130832e-03 .. . ···
-1.19824838e-02
-8.25322699e-04 . .
. ···
-5.14897937e-03 ···
-9.90395527e-03
Instead of using handcrafted kernels such as edge detectors can we learn meaningful ker- nels/filters
in addition to learning the weights of the classifier?
29/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Input Features Classifier
0 0 0 0 0
0 1 1 1 0
0 1 -8 1 0
0 1 1 1 0
0 0 0 0 0
-1.21358689e-03
3.23652686e-03 ··· ···
-2.06615720e-02
... ..
. .
-1.52757822e-03
... ··· ..
2.36130832e-03 .. . ···
-1.19824838e-02 -5.14897937e-03
-8.25322699e-04 .. ··· ··· -9.90395527e-03
.
Even better: Instead of using handcrafted kernels (such as edge detectors)can we learn multiple
meaningful kernels/filters in addition to learning the weights of the clas-
sifier? 30/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Input Features Classifier
0 0 0 0 0
0 1 1 1 0
0 1 -8 1 0
0 1 1 1 0
0 0 0 0 0
-0.02337041 -0.03243878
··· ··· -0.04728875
-0.05375158
. . -0.05350766.
···. ···. -0.04323674 ..
. . .
. . ..
.
-0.00792501 -0.00503319 ··· ··· 0.00174674
Even better: Instead of. using handcrafted kernels (such as edge detectors)can we learn multiple
meaningful kernels/filters in addition to learning the weights of the clas- sifier?
30/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Input Features Classifier
0 0 0 0 0
0 1 1 1 0
0 1 -8 1 0
0 1 1 1 0
0 0 0 0 0
-0.01871333 -0.01075948
··· ··· 0.04684572
0.00104325 0.01935937
.. .. .
···. ··· 0.01016542.
.. ..
. .. .
..
0.03008777 0.00335217 ··· ··· -0.02791128
.
Even better: Instead of using handcrafted kernels (such as edge detectors)can we learn multiple
meaningful kernels/filters in addition to learning the weights of the clas-
30/68
sifier?
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Can we learn multiple layers of meaningful kernels/filters in addition to learning the
weights of the classifier?
31/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Input Classifier
31/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Input Classifier
31/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Input Classifier
32/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Okay, I get it that the idea is to learn the kernel/filters by just treating them as
parameters of the classification model
But how is this different from a regular feedforward neural network
32/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Okay, I get it that the idea is to learn the kernel/filters by just treating them as
parameters of the classification model
But how is this different from a regular feedforward neural network Let
us see
32/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
2
33/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
16
2
33/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
10 classes(digits)
16
2
33/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
10 classes(digits)
16
2
33/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
10 classes(digits)
16
2
33/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
10 classes(digits)
..
. 16
2
33/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
10 classes(digits)
..
. 16
2
33/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
10 classes(digits)
..
. 16
2
33/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
10 classes(digits)
..
. 16
2
33/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
10 classes(digits)
..
neural network will look like
. 16
2
33/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
10 classes(digits)
..
neural network will look like
.
There are many dense connections
here
16
2
33/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
10 classes(digits)
..
neural network will look like
.
There are many dense connections
here
2
33/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
10 classes(digits)
..
neural network will look like
.
There are many dense connections
here
2 * =
34/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Only a few local neurons participate in
the computation of h11
..
16 .
2 * =
34/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
h Only a few local neurons participate in
11
the computation of h11
16 .
2
h
11
* =
34/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
h h Only a few local neurons participate in
11 12
the computation of h11
16 .
2
h
12
* =
34/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
h h Only a few local neurons participate in
11 12
the computation of h11
16 .
2
h
13
* =
34/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
h h Only a few local neurons participate in
11 12
the computation of h11
16 .
2
h
14
* =
34/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
h h Only a few local neurons participate in
11 12
the computation of h11
2
h
14
* =
34/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
h h Only a few local neurons participate in
11 12
the computation of h11
2
h between neighboring pixels are more
14
* = interesting)
34/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
h h Only a few local neurons participate in
11 12
the computation of h11
2
h between neighboring pixels are more
14
* = interesting)
34/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
But is sparse connectivity really good
thing ?
∗
Goodfellow-et-al-2016
35/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
But is sparse connectivity really good
thing ?
∗
Goodfellow-et-al-2016
35/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
But is sparse connectivity really good
thing ?
∗
Goodfellow-et-al-2016
35/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
But is sparse connectivity really good
thing ?
∗
Goodfellow-et-al-2016
35/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
But is sparse connectivity really good
thing ?
36/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Another characteristic
of CNNs is weight
sharing
36/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Another characteristic
of CNNs is weight
sharing
Consider the following net-
work
16
Kernel 1
Kernel 2
4x4 Image
36/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Another characteristic
of CNNs is weight
sharing
Consider the following net-
work
Kernel 1
Kernel 2
4x4 Image
36/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Another characteristic
of CNNs is weight
sharing
Consider the following net-
work
36/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Another characteristic
of CNNs is weight
sharing
Consider the following net-
work
16
37/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
In other words shouldn’t the orange
and pink kernels be the same
Yes, indeed
16
37/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
In other words shouldn’t the orange
and pink kernels be the same
Yes, indeed
16
37/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
In other words shouldn’t the orange
and pink kernels be the same
Yes, indeed
16
37/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
In other words shouldn’t the orange
and pink kernels be the same
Yes, indeed
37/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
In other words shouldn’t the orange
and pink kernels be the same
Yes, indeed
37/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
In other words shouldn’t the orange
and pink kernels be the same
Yes, indeed
37/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
In other words shouldn’t the orange
and pink kernels be the same
Yes, indeed
37/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
In other words shouldn’t the orange
and pink kernels be the same
Yes, indeed
38/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
So far, we have focused only on the convolution operation Let us
see what a full convolutional neural network looks like
38/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Convolution Layer 2
Input Convolution Layer 1 Pooling Layer 2
A
Pooling Layer 1 FC 1(120)FC 2(84)
Output(10)
32
28
14
32 Param
10 5 Param Param = 850
28 14 = 10164
= 48120
10 5
S = 1,F = 5, S = 1,F = 2,
K = 6,P = 0, K = 6,P = 0, S = 1,F = 2,
S = 1,F = 5,
Param = 150 Param = 0 K = 16,P = 0,
K = 16,P = 0,
Param = 0
Param = 2400
39/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Convolution Layer 2
Input Convolution Layer 1 Pooling Layer 2
A
Pooling Layer 1 FC 1(120)FC 2(84)
Output(10)
32
28
14
32 Param
10 5 Param Param = 850
28 14 = 10164
= 48120
10 5
S = 1,F = 5, S = 1,F = 2,
K = 6,P = 0, K = 6,P = 0, S = 1,F = 2,
S = 1,F = 5,
Param = 150 Param = 0 K = 16,P = 0,
K = 16,P = 0,
Param = 0
Param = 2400
39/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Convolution Layer 2
Input Convolution Layer 1 Pooling Layer 2
A
Pooling Layer 1 FC 1(120)FC 2(84)
Output(10)
32
28
14
32 Param
10 5 Param Param = 850
28 14 = 10164
= 48120
10 5
S = 1,F = 5, S = 1,F = 2,
K = 6,P = 0, K = 6,P = 0, S = 1,F = 2,
S = 1,F = 5,
Param = 150 Param = 0 K = 16,P = 0,
K = 16,P = 0,
Param = 0
Param = 2400
39/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Convolution Layer 2
Input Convolution Layer 1 Pooling Layer 2
A
Pooling Layer 1 FC 1(120)FC 2(84)
Output(10)
32
28
14
32 Param
10 5 Param Param = 850
28 14 = 10164
= 48120
10 5
S = 1,F = 5, S = 1,F = 2,
K = 6,P = 0, K = 6,P = 0, S = 1,F = 2,
S = 1,F = 5,
Param = 150 Param = 0 K = 16,P = 0,
K = 16,P = 0,
Param = 0
Param = 2400
40/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
*
Input 1 filter
40/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
* =
Input 1 filter
40/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 4 2 1
5 8 3 4
* =
7 6 4 5
1 3 1 2
Input 1 filter
40/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 4 2 1
5 8 3 4 maxpool
* =
7 6 4 5 2x2 filters (stride 2)
1 3 1 2
Input 1 filter
40/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 4 2 1
5 8 3 4 maxpool
* =
7 6 4 5 2x2 filters (stride 2)
1 3 1 2
Input 1 filter
40/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 4 2 1
5 8 3 4 maxpool 8
* =
7 6 4 5 2x2 filters (stride 2)
1 3 1 2
Input 1 filter
40/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 4 2 1
5 8 3 4 maxpool 8 4
* =
7 6 4 5 2x2 filters (stride 2)
1 3 1 2
Input 1 filter
40/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 4 2 1
5 8 3 4 maxpool 8 4
* =
7 6 4 5 2x2 filters (stride 2) 7
1 3 1 2
Input 1 filter
40/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 4 2 1
5 8 3 4 maxpool 8 4
* =
7 6 4 5 2x2 filters (stride 2) 7 5
1 3 1 2
Input 1 filter
40/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 4 2 1
5 8 3 4 maxpool 8 4
* =
7 6 4 5 2x2 filters (stride 2) 7 5
1 3 1 2
Input 1 filter
1 4 2 1
5 8 3 4
7 6 4 5
1 3 1 2
40/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 4 2 1
5 8 3 4 maxpool 8 4
* =
7 6 4 5 2x2 filters (stride 2) 7 5
1 3 1 2
Input 1 filter
1 4 2 1
5 8 3 4 maxpool
1 3 1 2
40/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 4 2 1
5 8 3 4 maxpool 8 4
* =
7 6 4 5 2x2 filters (stride 2) 7 5
1 3 1 2
Input 1 filter
1 4 2 1
5 8 3 4 maxpool
1 3 1 2
40/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 4 2 1
5 8 3 4 maxpool 8 4
* =
7 6 4 5 2x2 filters (stride 2) 7 5
1 3 1 2
Input 1 filter
1 4 2 1
5 8 3 4 maxpool 8
1 3 1 2
40/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 4 2 1
5 8 3 4 maxpool 8 4
* =
7 6 4 5 2x2 filters (stride 2) 7 5
1 3 1 2
Input 1 filter
1 4 2 1
5 8 3 4 maxpool 8 8
1 3 1 2
40/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 4 2 1
5 8 3 4 maxpool 8 4
* =
7 6 4 5 2x2 filters (stride 2) 7 5
1 3 1 2
Input 1 filter
1 4 2 1
5 8 3 4 maxpool 8 8 4
1 3 1 2
40/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 4 2 1
5 8 3 4 maxpool 8 4
* =
7 6 4 5 2x2 filters (stride 2) 7 5
1 3 1 2
Input 1 filter
1 4 2 1
5 8 3 4 maxpool 8 8 4
1 3 1 2
40/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 4 2 1
5 8 3 4 maxpool 8 4
* =
7 6 4 5 2x2 filters (stride 2) 7 5
1 3 1 2
Input 1 filter
1 4 2 1
5 8 3 4 maxpool 8 8 4
1 3 1 2
40/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 4 2 1
5 8 3 4 maxpool 8 4
* =
7 6 4 5 2x2 filters (stride 2) 7 5
1 3 1 2
Input 1 filter
1 4 2 1
5 8 3 4 maxpool 8 8 4
1 3 1 2
40/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 4 2 1
5 8 3 4 maxpool 8 4
* =
7 6 4 5 2x2 filters (stride 2) 7 5
1 3 1 2
Input 1 filter
1 4 2 1
5 8 3 4 maxpool 8 8 4
1 3 1 2 7
40/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 4 2 1
5 8 3 4 maxpool 8 4
* =
7 6 4 5 2x2 filters (stride 2) 7 5
1 3 1 2
Input 1 filter
1 4 2 1
5 8 3 4 maxpool 8 8 4
1 3 1 2 7 6
40/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 4 2 1
5 8 3 4 maxpool 8 4
* =
7 6 4 5 2x2 filters (stride 2) 7 5
1 3 1 2
Input 1 filter
1 4 2 1
5 8 3 4 maxpool 8 8 4
1 3 1 2 7 6 5
40/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 4 2 1
5 8 3 4 maxpool 8 4
* =
7 6 4 5 2x2 filters (stride 2) 7 5
1 3 1 2
Input 1 filter
1 4 2 1
5 8 3 4 maxpool 8 8 4
1 3 1 2 7 6 5
40/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
We will now see some case studies where convolution neural networks have been
successful
41/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
LeNet-5 for handwritten character recognition
Input
A 32
32
42/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
LeNet-5 for handwritten character recognition
A 32
28
32
28
S = 1,F = 5,
K = 6,P = 0,
Param =?
42/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
LeNet-5 for handwritten character recognition
A 32
28
32
28
S = 1,F = 5,
K = 6,P = 0,
Param = 150
42/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
LeNet-5 for handwritten character recognition
A
Pooling Layer 1
32
28
14
32
28 14
S = 1,F = 5, S = 1,F = 2,
K = 6,P = 0, K = 6,P = 0,
Param = 150 Param =?
42/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
LeNet-5 for handwritten character recognition
A
Pooling Layer 1
32
28
14
32
28 14
S = 1,F = 5, S = 1,F = 2,
K = 6,P = 0, K = 6,P = 0,
Param = 150 Param = 0
42/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
LeNet-5 for handwritten character recognition
Convolution Layer 2
Input Convolution Layer 1
A
Pooling Layer 1
32
28
14
32
10
28 14
10
S = 1,F = 5, S = 1,F = 2,
K = 6,P = 0, K = 6,P = 0, S = 1,F = 5,
Param = 150 Param = 0 K = 16,P = 0,
Param =?
42/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
LeNet-5 for handwritten character recognition
Convolution Layer 2
Input Convolution Layer 1
A
Pooling Layer 1
32
28
14
32
10
28 14
10
S = 1,F = 5, S = 1,F = 2,
K = 6,P = 0, K = 6,P = 0, S = 1,F = 5,
Param = 150 Param = 0 K = 16,P = 0,
Param = 2400
42/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
LeNet-5 for handwritten character recognition
Convolution Layer 2
Input Convolution Layer 1 Pooling Layer 2
A
Pooling Layer 1
32
28
14
32
10 5
28 14
10 5
S = 1,F = 5, S = 1,F = 2,
K = 6,P = 0, K = 6,P = 0, S = 1,F = 2,
S = 1,F = 5,
Param = 150 Param = 0 K = 16,P = 0,
K = 16,P = 0,
Param =?
Param = 2400
42/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
LeNet-5 for handwritten character recognition
Convolution Layer 2
Input Convolution Layer 1 Pooling Layer 2
A
Pooling Layer 1
32
28
14
32
10 5
28 14
10 5
S = 1,F = 5, S = 1,F = 2,
K = 6,P = 0, K = 6,P = 0, S = 1,F = 2,
S = 1,F = 5,
Param = 150 Param = 0 K = 16,P = 0,
K = 16,P = 0,
Param = 0
Param = 2400
42/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
LeNet-5 for handwritten character recognition
Convolution Layer 2
Input Convolution Layer 1 Pooling Layer 2
A
Pooling Layer FC 1(120)
1
32
28
14
32 Param
10 5
28 14 =?
10 5
S = 1,F = 5, S = 1,F = 2,
K = 6,P = 0, K = 6,P = 0, S = 1,F = 2,
S = 1,F = 5,
Param = 150 Param = 0 K = 16,P = 0,
K = 16,P = 0,
Param = 0
Param = 2400
42/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
LeNet-5 for handwritten character recognition
Convolution Layer 2
Input Convolution Layer 1 Pooling Layer 2
A
Pooling Layer FC 1(120)
1
32
28
14
32 Param
10 5
28 14 = 48120
10 5
S = 1,F = 5, S = 1,F = 2,
K = 6,P = 0, K = 6,P = 0, S = 1,F = 2,
S = 1,F = 5,
Param = 150 Param = 0 K = 16,P = 0,
K = 16,P = 0,
Param = 0
Param = 2400
42/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
LeNet-5 for handwritten character recognition
Convolution Layer 2
Input Convolution Layer 1 Pooling Layer 2
A
Pooling Layer 1 FC 1(120)
FC 2(84)
32
28
14
32 Param
10
28 14 5 =Param
48120 =?
10 5
S = 1,F = 5, S = 1,F = 2,
K = 6,P = 0, K = 6,P = 0, S = 1,F = 2,
S = 1,F = 5,
Param = 150 Param = 0 K = 16,P = 0,
K = 16,P = 0,
Param = 0
Param = 2400
42/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
LeNet-5 for handwritten character recognition
Convolution Layer 2
Input Convolution Layer 1 Pooling Layer 2
A
Pooling Layer 1 FC 1(120)
FC 2(84)
32
28
14
32 Param
10 5 Param
28 14 = 10164
= 48120
10 5
S = 1,F = 5, S = 1,F = 2,
K = 6,P = 0, K = 6,P = 0, S = 1,F = 2,
S = 1,F = 5,
Param = 150 Param = 0 K = 16,P = 0,
K = 16,P = 0,
Param = 0
Param = 2400
42/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
LeNet-5 for handwritten character recognition
Convolution Layer 2
Input Convolution Layer 1 Pooling Layer 2
A
Pooling Layer 1 FC 1(120)FC 2(84)
Output(10)
32
28
14
32 Param
10 5 Param Param =?
28 14 = 10164
= 48120
10 5
S = 1,F = 5, S = 1,F = 2,
K = 6,P = 0, K = 6,P = 0, S = 1,F = 2,
S = 1,F = 5,
Param = 150 Param = 0 K = 16,P = 0,
K = 16,P = 0,
Param = 0
Param = 2400
42/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
LeNet-5 for handwritten character recognition
Convolution Layer 2
Input Convolution Layer 1 Pooling Layer 2
A
Pooling Layer 1 FC 1(120)FC 2(84)
Output(10)
32
28
14
32 Param
10 5 Param Param = 850
28 14 = 10164
= 48120
10 5
S = 1,F = 5, S = 1,F = 2,
K = 6,P = 0, K = 6,P = 0, S = 1,F = 2,
S = 1,F = 5,
Param = 150 Param = 0 K = 16,P = 0,
K = 16,P = 0,
Param = 0
Param = 2400
42/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
How do we train a convolutional neural network ?
43/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Inpu Kernel
t
b c d l m n o
w x
e f g
y z
c e g i j
h i j b d f h
44/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Inpu Kernel
t
b c d l m n o
w x
e f g
y z
c e g
h i j b d f h i j
44/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Inpu Kernel
t
b c d l m n o
w x
e f g
y z
c e g
h i j b d f h i j
44/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Inpu Kernel
t
b c d l m n o
w x
e f g
y z
c e g i j
h i j b d f h
44/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Inpu Kernel
t
b c d l m n o
w x
e f g
y z
c e g i j
h i j b d f h
44/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Inpu Kernel
t
b c d l m n o
w x
e f g
y z
c e g i j
h i j b d f h
44/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Inpu Kernel
t
b c d l m n o
w x
e f g
y z
c e g i j
h i j b d f h
45/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
ImageNet Success Stories(roadmap for rest of the talk)
AlexNet
46/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
ImageNet Success Stories(roadmap for rest of the talk)
AlexNet
ZFNet
46/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
ImageNet Success Stories(roadmap for rest of the talk)
AlexNet
ZFNet
VGGNet
46/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
28.2
ILSVRC’10
47/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
28.2
25.8
ILSVRC’10 ILSVRC’11
47/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
28.2
25.8
16.4
47/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
28.2
25.8
16.4
11.7
47/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
28.2
25.8
16.4
11.7
7.3
47/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
28.2
25.8
16.4
11.7
7.3 6.7
47/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
28.2
25.8
16.4
11.7
7.3 6.7
3.57
47/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
28.2
25.8
152 layers
16.4
11.7
19 layers 22 layers
7.3 6.7
47/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
28.2
25.8
152 layers
16.4
11.7
19 layers 22 layers
7.3 6.7
47/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
ImageNet Success Stories(roadmap for rest of the talk)
AlexNet
ZFNet
VGGNet
48/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
227
227
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Input: 227 × 227 × 3
Conv1: K = 96,F = 11 S =
4,P = 0
Output:W =?, H =?
2 2
Parameters: ?
227
11
11
227
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Input: 227 × 227 × 3
Conv1: K = 96,F = 11 S =
4,P = 0
Output:W = 55, H = 55
2 2
Parameters: ?
227
55
11
11
55
227
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Input: 227 × 227 × 3
Conv1: K = 96,F = 11 S =
4,P = 0
Output:W = 55, H = 55
2 2
Parameters: (11 × 11 × 3) × 96 = 34K
227
55
11
11
55
227
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Max Pool Input: 55 × 55 × 96
F = 3,S = 2
Output:W =?, H =?
2 2
Parameters: ?
227
55
11
3
3
11
55
227
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Max Pool Input: 55 × 55 × 96
F = 3,S = 2
Output:W = 27, H = 27
2 2
Parameters: ?
227
55
27
11
3
3
11
27
55
227
96
MaxPooling
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Max Pool Input: 55 × 55 × 96
F = 3,S = 2
Output:W = 27, H = 27
2 2
Parameters: 0
227
55
27
11
3
3
11
27
55
227
96
MaxPooling
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Input: 27 × 27 × 96
Conv1: K = 256,F = 5 S =
1,P = 0
Output:W =?, H =?
2 2
Parameters: ?
227
55
27
11
3 5
3
5
11
27
55
227
96
MaxPooling
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Input: 27 × 27 × 96
Conv1: K = 256,F = 5 S =
1,P = 0
Output:W = 23, H = 23
2 2
Parameters: ?
227
55
27 23
11
3 5
3
5
11
27 23
55
227
96 256
Convolution
MaxPooling
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Input: 27 × 27 × 96
Conv1: K = 256,F = 5 S =
1,P = 0
Output:W = 23, H = 23
2 2
Parameters: (5 × 5 × 96) × 256 = 0.6M
227
55
27 23
11
3 5
3
5
11
27 23
55
227
96 256
Convolution
MaxPooling
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Max Pool Input: 23 × 23 × 256
F = 3,S = 2
Output:W =?, H =?
2 2
Parameters: ?
227
55
27 23
11
3 5
3
3 3
5
11
27 23
55
227
96 256
Convolution
MaxPooling
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Max Pool Input: 23 × 23 × 256
F = 3,S = 2
Output:W = 11, H = 11
2 2
Parameters: ?
227
55
27 23
11
11
3 5
3
3 3
5
11
11
27 23
55
227 256
MaxPooling
96 256
Convolution
MaxPooling
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Max Pool Input: 23 × 23 × 256
F = 3,S = 2
Output:W = 11, H = 11
2 2
Parameters: 0
227
55
27 23
11
11
3 5
3
3 3
5
11
11
27 23
55
227 256
MaxPooling
96 256
Convolution
MaxPooling
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Input: 11 × 11 × 256
Conv1: K = 384,F = 3 S =
1,P = 0
Output:W =?, H =?
2 2
Parameters: ?
227
55
27 23
11
11
3 5
3 3
3 3 3
5
11
11
27 23
55
227 256
MaxPooling
96 256
Convolution
MaxPooling
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Input: 11 × 11 × 256
Conv1: K = 384,F = 3 S =
1,P = 0
Output:W = 9, H = 9
2 2
Parameters: ?
227
55
27 23
11
11 9
3 5
3 3
3 3 3
5
11
11 9
27 23
55
384
227 256
Convolution
MaxPooling
96 256
Convolution
MaxPooling
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Input: 11 × 11 × 256
Conv1: K = 384,F = 3 S =
1,P = 0
Output:W = 9, H = 9
2 2
Parameters: (3 × 3 × 256) × 384 = 0.8M
227
55
27 23
11
11 9
3 5
3 3
3 3 3
5
11
11 9
27 23
55
384
227 256
Convolution
MaxPooling
96 256
Convolution
MaxPooling
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Input: 9 × 9 × 384
Conv1: K = 384,F = 3 S =
1,P = 0
Output:W =?, H =?
2 2
Parameters: ?
227
55
27 23
11
11 9
3 5
3 3 3
3 3 3
5 3
11
11 9
27 23
55
384
227 256
Convolution
MaxPooling
96 256
Convolution
MaxPooling
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Input: 9 × 9 × 384
Conv1: K = 384,F = 3 S =
1,P = 0
Output:W = 7, H = 7
2 2
Parameters: ?
227
55
27 23
11
11 9
3 5 7
3 3 3
3 3 3
5 3
11
7
11 9
27 23 384
55 Convolution
384
227 256
Convolution
MaxPooling
96 256
Convolution
MaxPooling
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Input: 9 × 9 × 384
Conv1: K = 384,F = 3 S =
1,P = 0
Output:W = 7, H = 7
2 2
Parameters: (3 × 3 × 384) × 384 = 1.327M
227
55
27 23
11
11 9
3 5 7
3 3 3
3 3 3
5 3
11
7
11 9
27 23 384
55 Convolution
384
227 256
Convolution
MaxPooling
96 256
Convolution
MaxPooling
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Input: 7 × 7 × 384
Conv1: K = 256,F = 3 S =
1,P = 0
Output:W =?, H =?
2 2
Parameters: ?
227
55
27 23
11
11 9
3 5 7
3 3 3 3
3 3 3
5 3 3
11
7
11 9
27 23 384
55 Convolution
384
227 256
Convolution
MaxPooling
96 256
Convolution
MaxPooling
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Input: 7 × 7 × 384
Conv1: K = 256,F = 3 S =
1,P = 0
Output:W = 5, H = 5
2 2
Parameters: ?
227
55
27 23
11
11 9
3 5 7
3 3 3 3 5
3 3 3
5 3 3
11 5
7
11 9 256
27 23 384 Convolution
55 Convolution
384
227 256
Convolution
MaxPooling
96 256
Convolution
MaxPooling
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Input: 7 × 7 × 384
Conv1: K = 256,F = 3 S =
1,P = 0
Output:W = 5, H = 5
2 2
Parameters: (3 × 3 × 384) × 256 = 0.8M
227
55
27 23
11
11 9
3 5 7
3 3 3 3 5
3 3 3
5 3 3
11 5
7
11 9 256
27 23 384 Convolution
55 Convolution
384
227 256
Convolution
MaxPooling
96 256
Convolution
MaxPooling
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Max Pool Input: 5 × 5 × 256
F = 3,S = 2
Output:W =?, H =?
2 2
Parameters: ?
227
55
27 23
11
11 9
3 5 7
3 3 3 3 5
3 3
3 3
5 3 3 3
11 5
7
11 9 256
27 23 384 Convolution
55 Convolution
384
227 256
Convolution
MaxPooling
96 256
Convolution
MaxPooling
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Max Pool Input: 5 × 5 × 256
F = 3,S = 2
Output:W = 2, H = 2
2 2
Parameters: ?
227
55
27 23
11
11 9
3 5 7
3 3 3 3 5
3 3
3 3 2
5 3 3 3
11 5 2
7 256
11 9 256 MaxPooling
27 23 384 Convolution
55 Convolution
384
227 256
Convolution
MaxPooling
96 256
Convolution
MaxPooling
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Max Pool Input: 5 × 5 × 256
F = 3,S = 2
Output:W = 2, H = 2
2 2
Parameters: 0
227
55
27 23
11
11 9
3 5 7
3 3 3 3 5
3 3
3 3 2
5 3 3 3
11 5 2
7 256
11 9 256 MaxPooling
27 23 384 Convolution
55 Convolution
384
227 256
Convolution
MaxPooling
96 256
Convolution
MaxPooling
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
FC1
Parameters: (2 × 2 × 256) × 4096 = 4M
227
55
27 23
11
11 9
3 5 7 dense
3 3 3 3 5
3 3 2
3 3
5 3 3 3
11 5 2
7 256
11 9 256 MaxPooling
27 23 384 Convolution
55 Convolution
384
227 256
Convolution
MaxPooling
96 256
Convolution
MaxPooling 4096
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
FC1
Parameters: 4096 × 4096 = 16M
227
55
27 23
11
11 9
3 5 7 dense
3 3 3 3 5
3 3 2 dense
5 3 3 3 3
11 3
5 2
7 256
11 9 256 MaxPooling
27 23 384 Convolution
55 Convolution
384
227 256
Convolution
MaxPooling
96 256
Convolution
MaxPooling 4096 4096
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
FC1
Parameters: 4096 × 1000 = 4M
227
55
27 23
11
11 9
3 5 7 dense
3 3 3 3 5
3 3 2 dense dense
5 3 3 3 3
11 3
5 2
7 256
11 9 256 MaxPooling
27 23 384 Convolution
55 Convolution
384
227 256
Convolution
MaxPooling 1000
96 256
Convolution 4096 4096
MaxPooling
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Total Parameters: 27.55M
227
55
27 23
11
11 9
3 5 7 dense
3 3 3 3 5
3 3 2 dense dense
5 3 3 3 3
11 3
5 2
7 256
11 9 256 MaxPooling
27 23 384 Convolution
55 Convolution
384
227 256
Convolution
MaxPooling 1000
96 256
Convolution 4096 4096
MaxPooling
96
Convolution
3
Input
49/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Let us look at the
connections in the
fully connected lay-
ers in more detail
2
256
MaxPooling
50/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Let us look at the
connections in the
fully connected lay-
ers in more detail
2
make linear
make it a 1d vector
2 × 2 × 256 = 1024
50/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Let us look at the
connections in the
fully connected lay-
ers in more detail
2
make linear
dense
We will first stretch 2
make it a 1d vector
This 1d vector is 2 × 2 × 256 = 1024
then densely connec-
ted to other lay- 4096
ers just as in a
regular feedforward
neural network
50/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Let us look at the
connections in the
fully connected lay-
ers in more detail
2
make linear
dense dense
We will first stretch 2
make it a 1d vector
This 1d vector is 2 × 2 × 256 = 1024
then densely connec-
ted to other lay- 4096 4096
ers just as in a
regular feedforward
neural network
50/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Let us look at the
connections in the
fully connected lay-
ers in more detail
2
make linear
dense dense dense
We will first stretch 2
make it a 1d vector
This 1d vector is 2 × 2 × 256 = 1024 1000
then densely connec-
ted to other lay- 4096 4096
ers just as in a
regular feedforward
neural network
50/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
ImageNet Success Stories(roadmap for rest of the talk)
AlexNet
ZFNet
VGGNet
51/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
227
227
3
Input
227
227
3
Input 52/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
227
11
11
227
Layer1: F = 11 → 7
3
Input
Difference in Parameters ((112
2
− 7 ) × 3) × 96 = 20.7K
227
227
3
Input 52/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
227
55
11
11
55
227
96
Convolution Layer1: F = 11 → 7
3
Input
Difference in Parameters ((112
2
− 7 ) × 3) × 96 = 20.7K
227
55
55
227
96
Convolution
3
Input 52/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
227
55
11
3
3
11
55
227
96
Convolution
3
Input
Layer2: No difference
227
55
7
3
3
7
55
227
96
Convolution
3
Input 52/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
227
55
27
11
3
3
11
27
55
227
96
MaxPooling
96
Convolution
3
Input
Layer2: No difference
227
55
27
7
3
3
7
27
55
227
96
MaxPooling
96
Convolution
3
Input 52/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
227
55
27
11
3 5
3
11 5
27
55
227
96
MaxPooling
96
Convolution
3
Input
Layer3: No difference
227
55
27
7
3 5
3
7 5
27
55
227
96
MaxPooling
96
Convolution
3
Input 52/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
227
55
27 23
11
3 5
3
11 5
27 23
55
227
96 256
MaxPooling Convolution
96
Convolution
3
Input
Layer3: No difference
227
55
27 23
7
3 5
3
7 5
27 23
55
227
96 256
MaxPooling Convolution
96
Convolution
3
Input 52/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
227
55
27 23
11
3 5
3
3 3
11 5
27 23
55
227
96 256
MaxPooling Convolution
96
Convolution
3
Input
Layer4: No difference
227
55
27 23
7
3 5
3
3 3
7 5
27 23
55
227
96 256
MaxPooling Convolution
96
Convolution
3
Input 52/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
227
55
27 23
11 11
3 5
3
3 3
11 5
11
27 23
55
227 256
256 MaxPooling
96
MaxPooling Convolution
96
Convolution
3
Input
Layer4: No difference
227
55
27 23
11
7
3 5
3
3 3
7 5
11
27 23
55
227 256
256 MaxPooling
96
MaxPooling Convolution
96
Convolution
3
Input 52/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
227
55
27 23
11 11
3 5
3 3
3 3 3
11 5
11
27 23
55
227 256
256 MaxPooling
96
MaxPooling Convolution
96
Convolution Layer5: K = 384 → 512
3
Input
Difference in Parameters
(3 × 3 × 256) × (512 − 384) = 0.29M
227
55
27 23
11
7
3 5
3 3
3 3 3
7 5
11
27 23
55
227 256
256 MaxPooling
96
MaxPooling Convolution
96
Convolution
3
Input 52/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
227
55
27 23
11 11 9
3 5
3 3
3 3 3
11 5
11 9
27 23
55
227 384
256
Convolution
256 MaxPooling
96
MaxPooling Convolution
96
Convolution Layer5: K = 384 → 512
3
Input
Difference in Parameters
(3 × 3 × 256) × (512 − 384) = 0.29M
227
55
27 23
11 9
7
3 5
3 3
3 3 3
7 5
11 9
27 23
55
227 512
256
MaxPooling Convolution
96 256
MaxPooling Convolution
96
Convolution
3
Input 52/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
227
55
27 23
11 11 9
3 5
3 3 3
3 3 3 3
11 5
11 9
27 23
55
227 384
256
Convolution
256 MaxPooling
96
MaxPooling Convolution
96
Convolution Layer6: K = 384 → 1024
3
Input
Difference in Parameters
(3 × 3 × ((384 × 384) − (512 × 1024)) = 0.8M
227
55
27 23
11 9
7
3 5
3 3 3
3 3 3 3
7 5
11 9
27 23
55
227 512
256
MaxPooling Convolution
96 256
MaxPooling Convolution
96
Convolution
3
Input 52/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
227
55
27 23
11 11 9
3 5 7
3 3 3
3 3 3 3
11 5
7
11 9
27 23
384
55
384 Convolution
227 256
Convolution
256 MaxPooling
96
MaxPooling Convolution
96
Convolution Layer6: K = 384 → 1024
3
Input
Difference in Parameters
(3 × 3 × ((384 × 384) − (512 × 1024)) = 0.8M
227
55
27 23
11 9
7
3 5 7
3 3 3
3 3 3 3
7 5
7
11 9
27 23
1024
55
512 Convolution
227 256
MaxPooling Convolution
96 256
MaxPooling Convolution
96
Convolution
3
Input 52/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
227
55
27 23
11 11 9
3 5 7
3 3 3 3
3 3 3 3
11 5 3
7
11 9
27 23
384
55
384 Convolution
227 256
Convolution
256 MaxPooling
96
MaxPooling Convolution
96
Convolution Layer7: K = 256 → 512
3
Input
Difference in Parameters
(3 × 3 × ((384 × 256) − (1024 × 512)) = 0.36M
227
55
27 23
11 9
7
3 5 7
3 3 3 3
3 3 3 3
7 5 3
7
11 9
27 23
1024
55
512 Convolution
227 256
MaxPooling Convolution
96 256
MaxPooling Convolution
96
Convolution
3
Input 52/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
227
55
27 23
11 11 9
3 5 7
3 3 3 3 5
3 3 3 3
11 5 3
5
7
11 9
256
27 23
384 Convolution
55
384 Convolution
227 256
Convolution
256 MaxPooling
96
MaxPooling Convolution
96
Convolution Layer7: K = 256 → 512
3
Input
Difference in Parameters
(3 × 3 × ((384 × 256) − (1024 × 512)) = 0.36M
227
55
27 23
11 9
7
3 5 7
3 3 3 3 5
3 3 3 3
7 5 3
5
7
11 9
23 512
27 1024
55 Convolution
512 Convolution
227 256
MaxPooling Convolution
96 256
MaxPooling Convolution
96
Convolution
3
Input 52/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
227
55
27 23
11 11 9
3 5 7
3 3 3 3 3 5
3 3 3 3
11 5 3
3
7 5
11 9
256
27 23
384 Convolution
55
384 Convolution
227 256
Convolution
256 MaxPooling
96
MaxPooling Convolution
96
Convolution
3
Input
Layer8: No difference
227
55
27 23
11 9
7
3 5 7
3 3 3 3 3 5
3 3 3
7 5 3 3 3
5
7
11 9
23 512
27 1024
55 Convolution
512 Convolution
227 256
MaxPooling Convolution
96 256
MaxPooling Convolution
96
Convolution
3
Input 52/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
227
55
27 23
11 11 9
3 5 7
3 3 3 3 3 5
3 3 2
5 3 3 3
11 3 2
7 5
11 9 256
23 256
27 384 MaxPooling
55 Convolution
384 Convolution
227 256
Convolution
256 MaxPooling
96
MaxPooling Convolution
96
Convolution
3
Input
Layer8: No difference
227
55
27 23
11 9
7
3 5 7
3 3 3 3 3 5
3 3 2
7 5 3 3 3
3 2
7 5
11 9 256
23 512
27 1024 MaxPooling
55 Convolution
512 Convolution
227 256
MaxPooling Convolution
96 256
MaxPooling Convolution
96
Convolution
3
Input 52/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
227
55
27 23
11 11 9
3 5 7
3 3 3 3 3 5
dense
3 3 3 3 2
11 5 3
3 2
7 5
11 9 256
23 256
27 384 MaxPooling
55 Convolution
384 Convolution
227 256
Convolution
256 MaxPooling
96
MaxPooling Convolution 4096
96
Convolution
3
Input
Layer9: No difference
227
55
27 23
11 9
7
3 5 7
3 3 3 3 3 5
3 3 dense
7 5 3 3 3 3 2
5 2
7
11 9 256
23 512
27 1024 MaxPooling
55 Convolution
512 Convolution
227 256
MaxPooling Convolution
96 256
MaxPooling Convolution 4096
96
Convolution
3
Input 52/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
227
55
27 23
11 11 9
3 5 7
3 3 3 3 3 5
3 3 2
dense dense
5 3 3 3
11 3
7 5 2
11 9 256
23 256
27 384 MaxPooling
55 Convolution
384 Convolution
227 256
Convolution
256 MaxPooling
96
MaxPooling Convolution 4096 4096
96
Convolution
3
Input
Layer10: No difference
227
55
27 23
11 9
7
3 5 7
3 3 3 3 3 5
3 3 dense dense
7 5 3 3 3 3 2
5 2
7
11 9 256
23 512
27 1024 MaxPooling
55 Convolution
512 Convolution
227 256
MaxPooling Convolution
96 256
MaxPooling Convolution 4096 4096
96
Convolution
3
Input 52/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
227
55
27 23
11 11 9
3 5 7
3 3 3 3 3 5
3 3 2
dense dense dens e
5 3 3 3
11 3
7 5 2
11 9 256
23 256
27 384 MaxPooling
55 Convolution
384 Convolution
227 256
Convolution 1000
256 MaxPooling
96
MaxPooling Convolution 4096 4096
96
Convolution
3
Input
Layer10: No difference
227
55
27 23
11 9
7
3 5 7
3 3 3 3 3 5
3 3 2
dense dense dens e
7 5 3 3 3
3 2
7 5
11 9 256
23 512
27 1024 MaxPooling
55 Convolution
512 Convolution
227 256
Convolution 1000
256 MaxPooling
96
MaxPooling Convolution 4096 4096
96
Convolution
3
Input 52/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
227
55
27 23
11 11 9
3 5 7
3 3 3 3 3 5
3 3 2
dense dense dens e
5 3 3 3
11 3
7 5 2
11 9 256
23 256
27 384 MaxPooling
55 Convolution
384 Convolution
227 256
Convolution 1000
256 MaxPooling
96
MaxPooling Convolution 4096 4096
96
Convolution Difference in Total No. of Parameters
3
Input 1.45M
227
55
27 23
11 9
7
3 5 7
3 3 3 3 3 5
3 3 2
dense dense dens e
7 5 3 3 3
3 2
7 5
11 9 256
23 512
27 1024 MaxPooling
55 Convolution
512 Convolution
227 256
Convolution 1000
256 MaxPooling
96
MaxPooling Convolution 4096 4096
96
Convolution
3
Input 52/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
ImageNet Success Stories(roadmap for rest of the talk)
AlexNet
ZFNet
VGGNet
53/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
4
22
224
Input
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
4
22
224
Input
Conv
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
4
4
22
22
224
224
64
Input
Conv
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
2
4
11
22
22
112
224
224
64
64 maxpool
Input
Conv
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
2
4
11
22
22
112
224
224
64
64 maxpool
Conv
Input
Conv
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
2
2
4
11
11
22
22
112
112
224
224
64 128
64 maxpool
Conv
Input
Conv
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
2
2
4
11
11
22
22
56
56
112
112
224
224
128
64 128 maxpool
64 maxpool
Conv
Input
Conv
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
2
2
4
11
11
22
22
56
56
112
112
224
224
128
64 128 maxpool Conv
64 maxpool
Conv
Input
Conv
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
2
2
4
11
11
22
22
56
56
112
112
224
224
128
64 128 maxpool Conv
64 maxpool
Conv
Input
Conv
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
2
2
4
11
11
22
22
56
56
56
56
112
112
224
224
128 256
64 128 maxpool Conv
64 maxpool
Conv
Input
Conv
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
2
2
4
11
11
22
22
56
56
28
28
56
56
112
112
224
224
256
128 256
maxpool
64 128 maxpool Conv
64 maxpool
Conv
Input
Conv
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
2
2
4
11
11
22
22
56
56
28
28
56
56
112
112
224
224
256
128 256
maxpool Conv
64 128 maxpool Conv
64 maxpool
Conv
Input
Conv
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
2
2
4
11
11
22
22
56
56
28
28
56
56
112
112
224
224
256
128 256
maxpool Conv
64 128 maxpool Conv
64 maxpool
Conv
Input
Conv
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
2
2
4
11
11
22
22
56
56
28
28
28
28
56
56
112
112
224
224
256 512
128 256
maxpool Conv
64 128 maxpool Conv
64 maxpool
Conv
Input
Conv
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
2
2
4
11
11
22
22
56
56
28
28
14
28
14
28
56
56
112
112
224
224
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
2
2
4
11
11
22
22
56
56
28
28
14
28
14
28
56
56
112
112
224
224
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
2
2
4
11
11
22
22
56
56
28
28
14
28
14
28
56
56
112
112
224
224
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
2
2
4
11
11
22
22
56
56
28
28
14
14
28
14
14
28
56
56
112
112
224
224
512 512
256 512
256
128 maxpool Conv maxpool Conv
64 128 maxpool Conv
64 maxpool
Conv
Input
Conv
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
2
2
4
11
11
22
22
56
56
28
28
14
14
7
7
28
14
14
28
56
56
112
112
512
224
224
512 512
256 512
256
maxpool maxpool maxpool
128 Conv Conv
64 128 maxpool Conv
64 maxpool
Conv
Input
Conv
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
2
2
4
11
11
22
22
56
56
28
28
14
14
7
7
28
14
14
28
56
56
112
112
512
224
224
512 512
256 512
256
maxpool maxpool maxpool
128 Conv Conv
64 128 maxpool Conv
64 maxpool
Conv
Input fc
Conv
4096
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
2
2
4
11
11
22
22
56
56
28
28
14
14
7
7
28
14
14
28
56
56
112
112
512
224
224
512 512
256 512
256
maxpool maxpool maxpool
128 Conv Conv
64 128 maxpool Conv
64 maxpool
Conv
Input fc fc
Conv
4096 4096
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
2
2
4
11
11
22
22
56
56
28
28
14
14
7
7
28
14
14
28
56
56
112
112
512
224
224
512 512
256 512
256
maxpool maxpool maxpool
128 Conv Conv
64 128 maxpool Conv
64 maxpool
Conv
Input fc fc
Conv
4096 4096
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
softmax
2
4
11
11
22
22
56
56
28
28
14
14
7
7
28
14
14
28
56
56
112
112
512
224
224
512 512
256 512
256
maxpool maxpool maxpool
128 Conv Conv
64 128 maxpool Conv
64 maxpool
Conv 1000
Input fc fc
Conv
4096 4096
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
softmax
2
4
11
11
22
22
56
56
28
28
14
14
7
7
28
14
14
28
56
56
112
112
512
224
224
512 512
256 512
256
maxpool maxpool maxpool
128 Conv Conv
64 128 maxpool Conv
64 maxpool
Conv 1000
Input fc fc
Conv
4096 4096
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
softmax
2
4
11
11
22
22
56
56
28
28
14
14
7
7
28
14
14
28
56
56
112
112
512
224
224
512 512
256 512
256
maxpool maxpool maxpool
128 Conv Conv
64 128 maxpool Conv
64 maxpool
Conv 1000
Input fc fc
Conv
4096 4096
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
softmax
2
4
11
11
22
22
56
56
28
28
14
14
7
7
28
14
14
28
56
56
112
112
512
224
224
512 512
256 512
256
maxpool maxpool maxpool
128 Conv Conv
64 128 maxpool Conv
64 maxpool
Conv 1000
Input fc fc
Conv
4096 4096
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
softmax
2
4
11
11
22
22
56
56
28
28
14
14
7
7
28
14
14
28
56
56
112
112
512
224
224
512 512
256 512
256
maxpool maxpool maxpool
128 Conv Conv
64 128 maxpool Conv
64 maxpool
Conv 1000
Input fc fc
Conv
4096 4096
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
softmax
2
4
11
11
22
22
56
56
28
28
14
14
7
7
28
14
14
28
56
56
112
112
512
224
224
512 512
256 512
256
maxpool maxpool maxpool
128 Conv Conv
64 128 maxpool Conv
64 maxpool
Conv 1000
Input fc fc
Conv
4096 4096
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
softmax
2
4
11
11
22
22
56
56
28
28
14
14
7
7
28
14
14
28
56
56
112
112
512
224
224
512 512
256 512
256
maxpool maxpool maxpool
128 Conv Conv
64 128 maxpool Conv
64 maxpool
Conv 1000
Input fc fc
Conv
4096 4096
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
softmax
2
4
11
11
22
22
56
56
28
28
14
14
7
7
28
14
14
28
56
56
112
112
512
224
224
512 512
256 512
256
maxpool maxpool maxpool
128 Conv Conv
64 128 maxpool Conv
64 maxpool
Conv 1000
Input fc fc
Conv
4096 4096
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
softmax
2
4
11
11
22
22
56
56
28
28
14
14
7
7
28
14
14
28
56
56
112
112
512
224
224
512 512
256 512
128 256
maxpool maxpool maxpool
Conv Conv
64 128 maxpool Conv
64 maxpool
Conv 1000
Input fc fc
Conv
4096 4096
54/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Module 11.5 : Image Classification continued
(GoogLeNet and ResNet)
55/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Consider the output at a certain layer
of a convolutional neural network
56/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
H
Consider the output at a certain layer
1
f
of a convolutional neural network
Max Pooling
D
f
After this layer we could apply a max-
H
W
1 pooling layer
1
56/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
H
Consider the output at a certain layer
1
f H
of a convolutional neural network
Max Pooling
D
f
After this layer we could apply a max-
1
H
W
1 convolution pooling layer
1
1
Or a 1 × 1 convolution
W
56/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Consider the output at a certain layer
of a convolutional neural network
H
1
f H
1
H
W
1 convolution H
2
pooling layer
1
3
1
Or a 1 × 1 convolution
convolution
D
3
W
Or a 3 × 3 convolution
W
2
W
56/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Consider the output at a certain layer
of a convolutional neural network
H
1
f H
1
H
W
1 convolution H
2
pooling layer
1
3
1
Or a 1 × 1 convolution
convolution
D
3
W
Or a 3 × 3 convolution
W
2
H
3 Or a 5 × 5 convolution
W
5
convolution
1
5
1
D
W
D 3
56/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Consider the output at a certain layer
of a convolutional neural network
H
1
f H
1
H
W
1 convolution H
2
pooling layer
1
3
1
Or a 1 × 1 convolution
convolution
D
3
W
Or a 3 × 3 convolution
W
2
H
3 Or a 5 × 5 convolution
W
5
1
convolution Question: Why choose between these
5
D
1 options (convolution, maxpool- ing,
D
W
3
filter sizes)?
1
56/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Consider the output at a certain layer
of a convolutional neural network
H
1
f H
1
H
W
1 convolution H
2
pooling layer
1
3
1
Or a 1 × 1 convolution
convolution
D
3
W
Or a 3 × 3 convolution
W
2
H
3 Or a 5 × 5 convolution
W
5
1
convolution Question: Why choose between these
5
D
1 options (convolution, maxpool- ing,
D
W
3
filter sizes)?
1 Idea: Why not apply all of them at the
same time and then concatenate the
feature maps?
56/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
H
Well this naive idea could result in a
1
f H
large number of computations
Max Pooling
f
D
H
1 W
1 convolution H
2
1
1
3
convolution
3
D W
H
3
W
2
W
5
convolution
1
5
1
D
W
D 3
57/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
H
Well this naive idea could result in a
1
f H
large number of computations
Max Pooling
D
f
If P = 0 & S = 1 then convolving a
W × H × D input with a F × F × D
H
1 W
1 convolution H
2
1
1
filter results in a (W − F + 1)(H −
3
3
convolution
F + 1) sized output
D W
H
3
W
2
W
5
convolution
1
5
1
D
W
D 3
57/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
H
Well this naive idea could result in a
1
f H
large number of computations
Max Pooling
D
f
If P = 0 & S = 1 then convolving a
W × H × D input with a F × F × D
H
1 W
1 convolution H
2
1
1
filter results in a (W − F + 1)(H −
3
3
convolution
F + 1) sized output
D W
H
3
Each element of the output requires
O(F × F × D) computations
W
2
W
5
convolution
1
5
1
D
W
D 3
57/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
H
Well this naive idea could result in a
1
f H
large number of computations
Max Pooling
D
f
If P = 0 & S = 1 then convolving a
W × H × D input with a F × F × D
H
1 W
1 convolution H
2
1
1
filter results in a (W − F + 1)(H −
3
3
convolution
F + 1) sized output
D W
H
3
Each element of the output requires
O(F × F × D) computations
W
2
W
5
convolution
1
5
1
Can we reduce the number of compu-
D
D
W
3
tations?
57/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Yes, by using 1 × 1 convolutions
58/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Yes, by using 1 × 1 convolutions Huh??
What does a 1 × 1 convolution
H do ?
58/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Yes, by using 1 × 1 convolutions
Huh?? What does a 1 × 1 convolution
H H do ?
It aggregates along the depth
1
W W
D 1
58/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Yes, by using 1 × 1 convolutions
Huh?? What does a 1 × 1 convolution
H H do ?
It aggregates along the depth
1 So convolving a D×W ×H input with D1
D
1
1×1 (D1 < D) filters will result in a D1 ×
W × H output (S = 1, P = 0)
W W
D D
1
58/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Yes, by using 1 × 1 convolutions
Huh?? What does a 1 × 1 convolution
H H do ?
It aggregates along the depth
1 So convolving a D×W ×H input with D1
D
1
1×1 (D1 < D) filters will result in a D1 ×
W × H output (S = 1, P = 0)
If D1 < D then this effectively re- duces
W W
the dimension of the input and hence
the computations
D D
1
58/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Yes, by using 1 × 1 convolutions
Huh?? What does a 1 × 1 convolution
H H do ?
It aggregates along the depth
1 So convolving a D×W ×H input with D1
D
1
1×1 (D1 < D) filters will result in a D1 ×
W × H output (S = 1, P = 0)
If D1 < D then this effectively re- duces
W W
the dimension of the input and hence
the computations
Specifically instead of O(F × F × D)
D D
1 we will need O(F × F × D1) compu-
tations
58/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Yes, by using 1 × 1 convolutions
Huh?? What does a 1 × 1 convolution
H H do ?
It aggregates along the depth
1 So convolving a D×W ×H input with D1
D
1
1×1 (D1 < D) filters will result in a D1 ×
W × H output (S = 1, P = 0)
If D1 < D then this effectively re- duces
W W
the dimension of the input and hence
the computations
Specifically instead of O(F × F × D)
D D
1 we will need O(F × F × D1) compu-
tations
We could then apply subsequent 3×3,
5 × 5 filter on this reduced output
58/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
But we might want to use different
28
1 × 1 convolutions
dimensionality reductions before the 3
3 × 3 convolutions
(dimensionality re-
duction)
(on reduced input)
× 3 and 5 × 5 filters
5 × 5 convolutions
(on reduced input)
28
256
59/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
But we might want to use different
28
1 × 1 convolutions
dimensionality reductions before the 3
3 × 3 convolutions
(dimensionality re-
duction)
(on reduced input)
× 3 and 5 × 5 filters
So we can use D and1 D 1 × 21 fil-
1 × 1 convolutions 5 × 5 convolutions
(dimensionality re-
(on reduced input)
duction)
59/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
But we might want to use different
28
1 × 1 convolutions
dimensionality reductions before the 3
3 × 3 convolutions
(dimensionality re-
duction)
(on reduced input)
× 3 and 5 × 5 filters
So we can use D and1 D 1 × 21 fil-
1 × 1 convolutions 5 × 5 convolutions
(dimensionality re-
(on reduced input)
duction)
28
3 × 3 Maxpooling
ters before the 3 × 3 and 5 × 5 filters
(dimensionality re-
duction)
1 × 1 convolutions
respectively
256
We can then add the maxpooling layer
followed by dimensionality re- duction
59/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 × 1 convolutions But we might want to use different
28
1 × 1 convolutions
dimensionality reductions before the 3
3 × 3 convolutions
(dimensionality re-
duction)
3 × 3 convolutions
(on reduced input)
(on reduced input) × 3 and 5 × 5 filters
So we can use D and1 D 1 × 21 fil-
1 ×× 1
1 1 convolutions
convolutions 5 × 5 convolutions
(dimensionality
(dimensionality re-
re- 5 × 5 convolutions
(on reduced input)
duction)
duction) (on reduced input)
28
3 × 3 Maxpooling
ters before the 3 × 3 and 5 × 5 filters
(dimensionality re-
duction)
11 ×× 11 convolutions
convolutions respectively
256
We can then add the maxpooling layer
followed by dimensionality re- duction
And a new set of 1 × 1 convolutions
59/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 × 1 convolutions But we might want to use different
28
1 × 1 convolutions
dimensionality reductions before the 3
3 × 3 convolutions
(dimensionality re-
duction)
3 × 3 convolutions
(on reduced input)
(on reduced input)
Filter
× 3 and 5 × 5 filters
concatenation
So we can use D and1 D 1 × 21 fil-
1 ×× 1
1 1 convolutions
convolutions 5 × 5 convolutions
(dimensionality
(dimensionality re-
re- 5 × 5 convolutions
(on reduced input)
duction)
duction) (on reduced input)
28
3 × 3 Maxpooling
ters before the 3 × 3 and 5 × 5 filters
(dimensionality re-
duction)
11 ×× 11 convolutions
convolutions respectively
256
We can then add the maxpooling layer
followed by dimensionality re- duction
And a new set of 1 × 1 convolutions
And finally we concatenate all these
layers
59/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 × 1 convolutions But we might want to use different
28
1 × 1 convolutions
dimensionality reductions before the 3
3 × 3 convolutions
(dimensionality re-
duction)
3 × 3 convolutions
(on reduced input)
(on reduced input)
Filter
× 3 and 5 × 5 filters
concatenation
So we can use D and1 D 1 × 21 fil-
1 ×× 1
1 1 convolutions
convolutions 5 × 5 convolutions
(dimensionality
(dimensionality re-
re- 5 × 5 convolutions
(on reduced input)
duction)
duction) (on reduced input)
28
3 × 3 Maxpooling
ters before the 3 × 3 and 5 × 5 filters
(dimensionality re-
duction)
11 ×× 11 convolutions
convolutions respectively
256
We can then add the maxpooling layer
followed by dimensionality re- duction
And a new set of 1 × 1 convolutions
And finally we concatenate all these
layers
This is called the Inception module
59/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1 × 1 convolutions But we might want to use different
28
1 × 1 convolutions
dimensionality reductions before the 3
3 × 3 convolutions
(dimensionality re-
duction)
3 × 3 convolutions
(on reduced input)
(on reduced input)
Filter
× 3 and 5 × 5 filters
concatenation
So we can use D and1 D 1 × 21 fil-
1 ×× 1
1 1 convolutions
convolutions 5 × 5 convolutions
(dimensionality
(dimensionality re-
re- 5 × 5 convolutions
(on reduced input)
duction)
duction) (on reduced input)
28
3 × 3 Maxpooling
ters before the 3 × 3 and 5 × 5 filters
(dimensionality re-
duction)
11 ×× 11 convolutions
convolutions respectively
256
We can then add the maxpooling layer
followed by dimensionality re- duction
And a new set of 1 × 1 convolutions
And finally we concatenate all these
layers
This is called the Inception module
We will now see GoogLeNet which
contains many such inception mod-
ules 59/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
9
22
229
Input
60/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
9
2
22
11
229
112
64
Conv
Input
60/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
9
56
22
11
56
229
112
64
64 maxpool
Conv
Input
60/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
9
56
56
22
11
56
56
229
112
64 192
64 maxpool
Conv
Conv
Input
60/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
9
56
56
22
11
28
28
56
56
229
112
192
64 192
maxpool
64 maxpool
Conv
Conv
Input
60/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
9
56
56
22
11
28
28
3a
28
28
56
56
229
112
192 256
64 192
maxpoolInception
64 maxpool
Conv
Conv
Input
64 1 × 1 convolutions
28
96 1 × 1 convolu- 128 3 × 3 convolu-
tions (dimensionality
tions (on reduced
reduction)
input)
Filter
concatenation
16 1 × 1 convolu- 32 5 × 5 convolutions
tions (dimensionality
(on reduced input)
reduction)
28
3 × 3 Maxpooling
(dimensionality re-
32 1 × 1 convolutions
duction)
192
60/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
9
56
56
22
11
28
28
28
3a 3b
28
28
28
56
56
229
112
192 256
64 480
192
64 maxpool maxpoolInception Inception
Conv
Conv
Input
128 1 × 1
convolutions
28
128 1 × 1 convolu- 192 3 × 3 convolu-
tions (dimensionality
tions (on reduced
reduction)
input)
Filter
concatenation
32 1 × 1 convolu- 96 5 × 5 convolutions
tions (dimensionality
(on reduced input)
reduction)
28
3 × 3 Maxpooling
(dimensionality re-
64 1 × 1 convolutions
duction)
256
60/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
9
56
56
22
11
28
28
28
14
3a 3b
28
28
28
56
56
14
229
112
60/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
9
56
56
22
11
28
28
28
14
3a 3b
28
28
28
4a
56
56
14
229
112
14
96 1 × 1 convolu- 208 3 × 3 convolu-
tions (dimensionality
tions (on reduced
reduction)
input)
Filter
concatenation
16 1 × 1 convolu- 48 5 × 5 convolutions
tions (dimensionality
(on reduced input)
reduction)
14
3 × 3 Maxpooling
(dimensionality re-
64 1 × 1 convolutions
duction)
480
60/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
9
56
56
22
11
28
28
28
14
3a 3b 4a 4b
28
28
28
56
56
14
229
112
14
112 1 × 1 convolu- 224 3 × 3 convolu-
tions (dimensionality
tions (on reduced
reduction)
input)
Filter
concatenation
24 1 × 1 convolu- 64 5 × 5 convolutions
tions (dimensionality
(on reduced input)
reduction)
14
3 × 3 Maxpooling
(dimensionality re-
64 1 × 1 convolutions
duction)
512
60/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
9
56
56
22
11
28
28
28
14
14
3a 3b 4a 4c
28
28
28
4b
56
56
14
14
229
112
480 512
192 256
64 192 maxpool
480 Inception
maxpoolInception Inception
64 maxpool
Conv
Conv
Input
128 1 × 1
convolutions
14
128 1 × 1 convolu- 256 3 × 3 convolu-
tions (dimensionality
tions (on reduced
reduction)
input)
Filter
concatenation
24 1 × 1 convolu- 64 5 × 5 convolutions
tions (dimensionality
(on reduced input)
reduction)
14
3 × 3 Maxpooling
(dimensionality re-
64 1 × 1 convolutions
duction)
512
60/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
9
56
56
22
11
28
28
28
14
14
14
3a 3b 4a 4b 4c 4d
28
28
28
56
56
14
14
14
229
112
14
144 1 × 1 convolu- 288 3 × 3 convolu-
tions (dimensionality
tions (on reduced
reduction)
input)
Filter
concatenation
32 1 × 1 convolu- 64 5 × 5 convolutions
tions (dimensionality
(on reduced input)
reduction)
14
3 × 3 Maxpooling
(dimensionality re-
64 1 × 1 convolutions
duction)
512
60/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
9
56
56
22
11
28
28
28
14
14
14
14
3a 3b 4a 4b 4c 4d 4e
28
28
28
56
56
14
14
14
14
229
112
14
160 1 × 1 convolu- 320 3 × 3 convolu-
tions (dimensionality
tions (on reduced
reduction)
input)
Filter
concatenation
32 1 × 1 convolu- 128 5 × 5 convolu-
tions (dimensionality
tions (on reduced
reduction)
input)
14
3 × 3 Maxpooling
(dimensionality re- 128 1 × 1
duction) convolutions
528
60/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
9
56
56
22
11
28
28
28
14
14
14
14
7
3a 3b 4a 4c 4e
28
28
28
4b 4d
56
56
14
14
14
14
7
229
112
832
480 512 528 832
192 256 maxpool
64 192 maxpool Inception Inception Inception
maxpoolInception480Inception
64 maxpool
Conv
Conv
Input
60/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
9
56
56
22
11
28
28
28
14
14
14
14
7
3a 3b 4a 4c 4e 5a
28
28
28
4b 4d
56
56
14
14
14
14
7
229
112
832 832
480 512 528 832
192 256 maxpool Inception
64 192 maxpool Inception Inception Inception
maxpoolInception480Inception
64 maxpool
Conv
Conv
Input
256 1 × 1
convolutions
7
160 1 × 1 convolu- 320 3 × 3 convolu-
tions (dimensionality
tions (on reduced
reduction)
input)
Filter
concatenation
32 1 × 1 convolu- 128 5 × 5 convolu-
tions (dimensionality
tions (on reduced
reduction)
input)
7
3 × 3 Maxpooling
(dimensionality re- 128 1 × 1
duction) convolutions
832
60/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
9
56
56
22
11
28
28
28
14
14
14
7
14
7
3a 3b 4a 4c 4d 4e 5a 5b
28
28
28
4b
56
56
14
14
14
14
7
229
112
7
192 1 × 1 convolu- 384 3 × 3 convolu-
tions (dimensionality
tions (on reduced
reduction)
input)
Filter
concatenation
48 1 × 1 convolu- 128 5 × 5 convolu-
tions (dimensionality
tions (on reduced
reduction)
input)
7
3 × 3 Maxpooling
(dimensionality re- 128 1 × 1
duction) convolutions
832
60/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
9
56
56
22
11
28
28
28
14
14
14
7
14
1
7
1
3a 3b 4a 4c 4e 5a 5b 1024
28
28
28
4b 4d
56
56
14
14
14
14
7
avgpool
229
112
60/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
9
56
56
22
11
28
28
28
14
14
14
7
14
1 1
7
1 1
3a 3b 4a 4c 4e 5a 5b 1024
28
28
28
4b 4d 1024
56
56
14
14
14
14
7
avgpool
229
112
60/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
9
56
56
22
11
28
28
28
14
14
14
7
14
1 1
7
1 1
3a 3b 4a 4c 4e 5a 5b 1024
28
28
28
4b 4d 1024
56
56
14
14
14
14
7
avgpool dropout(40%)
229
112
60/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
9
56
56
22
11
28
28
28
14
14
14
7
14
1 1
7
1 1
3a 3b 4a 4c 4e 5a 5b 1024
28
28
28
4b 4d 1024
56
56
14
14
14
14
7
avgpool dropout(40%)
229
112
60/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Important Trick: Got rid of the fully
connected layer
61/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Important Trick: Got rid of the fully
connected layer
Notice that output of the last layer is 7
7 × 7 × 1024
× 7 × 1024 dimensional
flatten
1024
61/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1000 Important Trick: Got rid of the fully
connected layer
Notice that output of the last layer is 7
7 × 7 × 1024
× 7 × 1024 dimensional
flatten What if we were to add a fully connec-
ted layer with 1000 nodes (for 1000
7 classes) on top of this
1024
61/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1000 Important Trick: Got rid of the fully
connected layer
50176×1000
W∈R
Notice that output of the last layer is 7
7 × 7 × 1024
× 7 × 1024 dimensional
flatten What if we were to add a fully connec-
ted layer with 1000 nodes (for 1000
7 classes) on top of this
We would have 7 × 7 × 1024 × 1000 =
1024
49M parameters
61/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1000 Important Trick: Got rid of the fully
connected layer
50176×1000
W∈R
Notice that output of the last layer is 7
7 × 7 × 1024
× 7 × 1024 dimensional
flatten What if we were to add a fully connec-
ted layer with 1000 nodes (for 1000
7 classes) on top of this
pick average We would have 7 × 7 × 1024 × 1000 =
1024
49M parameters
Instead they use an average pooling of
7 size 7 × 7 on each of the 1024 feature
maps
61/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1000 Important Trick: Got rid of the fully
connected layer
50176×1000
W∈R
Notice that output of the last layer is 7
7 × 7 × 1024
× 7 × 1024 dimensional
flatten What if we were to add a fully connec-
ted layer with 1000 nodes (for 1000
7 classes) on top of this
pick average We would have 7 × 7 × 1024 × 1000 =
1024
1024
49M parameters
Instead they use an average pooling of
7 size 7 × 7 on each of the 1024 feature
maps
This results in a 1024 dimensional
output
61/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1000 Important Trick: Got rid of the fully
connected layer
1024×1000
W∈R
Notice that output of the last layer is 7
1024
× 7 × 1024 dimensional
flatten What if we were to add a fully connec-
ted layer with 1000 nodes (for 1000
7 classes) on top of this
pick average We would have 7 × 7 × 1024 × 1000 =
1024
1024
49M parameters
Instead they use an average pooling of
7 size 7 × 7 on each of the 1024 feature
maps
This results in a 1024 dimensional
output
Significantly reduces the number of
parameters 61/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1000 12× less parameters than AlexNet
1024×1000
W∈R
1024
flatten
7
pick average
1024
1024
61/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1000
12× less parameters than AlexNet
2× more computations
1024×1000
W∈R
1024
flatten
7
pick average
1024
1024
61/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
GoogLeNet
ResNet
62/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Suppose we have been able to train a
shallow neural network well
63/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Suppose we have been able to train a
shallow neural network well
Now suppose we construct a deeper
network which has few more layers (in
orange)
63/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Suppose we have been able to train a
shallow neural network well
Now suppose we construct a deeper
network which has few more layers (in
orange)
Intuitively, if the shallow network
works well then the deep network
should also work well by simply learn-
ing to compute identity functions in
the new layers
63/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Suppose we have been able to train a
shallow neural network well
Now suppose we construct a deeper
network which has few more layers (in
orange)
Intuitively, if the shallow network
works well then the deep network
should also work well by simply learn-
ing to compute identity functions in
the new layers
Essentially, the solution space of a
shallow neural network is a subset of
the solution space of a deep neural
network
63/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
But in practice it is observed that this
doesn’t happen
64/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
But in practice it is observed that this
doesn’t happen
Notice that the deep layers have a
higher error rate on the test set
64/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
x
Consider any two stacked layers in a
CNN
relu
relu
H(x)
65/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
x
Consider any two stacked layers in a
CNN
relu
The two layers are essentially
learning some function of the input
relu
H(x)
65/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
x
Consider any two stacked layers in a
CNN
relu
The two layers are essentially
learning some function of the input
relu What if we enable it to learn only a
H(x) residual function of the input?
relu Identity
F (x)
relu
H(x) = F (x) + x
65/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
x
Why would this help?
relu
relu
H(x)
relu Identity
F (x)
relu
H(x) = F (x) + x
66/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
x
Why would this help?
Remember our argument that a
relu
deeper version of a shallow network
would do just fine by learning identity
transformations in the new layers
relu
H(x)
relu Identity
F (x)
relu
H(x) = F (x) + x
66/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
x
Why would this help?
Remember our argument that a
relu
deeper version of a shallow network
would do just fine by learning identity
transformations in the new layers
relu
relu Identity
F (x)
relu
H(x) = F (x) + x
66/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
x
Why would this help?
Remember our argument that a
relu
deeper version of a shallow network
would do just fine by learning identity
transformations in the new layers
relu
F (x)
relu
H(x) = F (x) + x
66/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1st place in all five main tracks
ImageNet Classification: “Ultra- deep”
152-layer nets
67/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1st place in all five main tracks
ImageNet Classification: “Ultra- deep”
152-layer nets
ImageNet Detection: 16% better than
the 2nd best system
67/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1st place in all five main tracks
ImageNet Classification: “Ultra- deep”
152-layer nets
ImageNet Detection: 16% better than
the 2nd best system
ImageNet Localization: 27% bet- ter
than the 2nd best system
67/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1st place in all five main tracks
ImageNet Classification: “Ultra- deep”
152-layer nets
ImageNet Detection: 16% better than
the 2nd best system
ImageNet Localization: 27% bet- ter
than the 2nd best system
COCO Detection: 11% better than the
2nd best system
67/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
1st place in all five main tracks ImageNet
Classification: “Ultra- deep” 152-layer
nets
ImageNet Detection: 16% better than
the 2nd best system
ImageNet Localization: 27% bet- ter
than the 2nd best system
COCO Detection: 11% better than the
2nd best system
COCO Segmentation: 12% better than
the 2nd best system
67/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Bag of tricks
Batch Normalizaton after every
CONV layer
68/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Bag of tricks
Batch Normalizaton after every
CONV layer
Xavier/2 initialization from [He et al]
68/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Bag of tricks
Batch Normalizaton after every
CONV layer
Xavier/2 initialization from [He et al]
SGD + Momentum(0.9)
68/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Bag of tricks
Batch Normalizaton after every
CONV layer
Xavier/2 initialization from [He et al]
SGD + Momentum(0.9)
Learning rate:0.1, divided by 10 when
validation error plateaus
68/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Bag of tricks
Batch Normalizaton after every
CONV layer
Xavier/2 initialization from [He et al]
SGD + Momentum(0.9)
Learning rate:0.1, divided by 10 when
validation error plateaus
Mini-batch size 256
68/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Bag of tricks
Batch Normalizaton after every
CONV layer
Xavier/2 initialization from [He et al]
SGD + Momentum(0.9)
Learning rate:0.1, divided by 10 when
validation error plateaus
Mini-batch size 256
Weight decay of 1e-5
68/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11
Bag of tricks
Batch Normalizaton after every
CONV layer
Xavier/2 initialization from [He et al]
SGD + Momentum(0.9)
Learning rate:0.1, divided by 10 when
validation error plateaus
Mini-batch size 256
Weight decay of 1e-5
No dropout used
68/68
Mitesh M. Khapra CS7015 (Deep Learning) : Lecture 11