Deep Learning in Computer Vision Techniques
Deep Learning in Computer Vision Techniques
Muhammad Faisal
Research Associate @ VisPro Lab & Intelligent Machines Lab
Information Technology University, Lahore
Outline
2
Image Classification
3
Object Detection
4
Segmentation
5
Image Style Transfer
6
Image Colorization
7
CNN Architectures
Case Studies:
AlexNet VGG ResNet DenseNet
8
Review: LeNet-5
[LeCun et al., 1998]
Lecture 9 - 9 9
Case Study: AlexNet
[Krizhevsky et al. 2012]
Architecture:
CONV1
MAX POOL1
NORM1
CONV2
MAX POOL2
NORM2
CONV3
CONV4
CONV5
Max POOL3
FC6
FC7
FC8
Lecture 9 - 1 10
Case Study: AlexNet
[Krizhevsky et al. 2012]
19 layers 22 layers
16
ImageNet Large Scale Visual Recognition Challenge (ILSVRC)
winners
19 layers 22 layers
17
ImageNet Large Scale Visual Recognition Challenge (ILSVRC)
winners
ZFNet: Improved
152 layers 152 layers 152 layers
hyperparameters over
AlexNet
19 layers 22 layers
18
ZFNet [Zeiler and Fergus, 2013]
AlexNet but:
CONV1: change from (11x11 stride 4) to (7x7 stride 2)
CONV3,4,5: instead of 384, 384, 256 filters use 512, 1024,
512
ImageNet top 5 error: 16.4% -> 11.7%
19
ImageNet Large Scale Visual Recognition Challenge (ILSVRC)
winners
19 layers 22 layers
20
Case Study: VGGNet
[Simonyan and Zisserman, 2014]
8 layers (AlexNet)
16 - 19 layers (VGG16Net)
21
Case Study: VGGNet
[Simonyan and Zisserman, 2014]
22
Case Study: VGGNet
[Simonyan and Zisserman, 2014]
23
Receptive Field
25
INPUT: [224x224x3] memory: 224*224*3=150K params: 0 (not counting biases)
CONV3-64: [224x224x64] memory: 224*224*64=3.2M params: (3*3*3)*64 = 1,728
CONV3-64: [224x224x64] memory: 224*224*64=3.2M params: (3*3*64)*64 = 36,864
POOL2: [112x112x64] memory: 112*112*64=800K params: 0
CONV3-128: [112x112x128] memory: 112*112*128=1.6M params: (3*3*64)*128 = 73,728
CONV3-128: [112x112x128] memory: 112*112*128=1.6M params: (3*3*128)*128 = 147,456
POOL2: [56x56x128] memory: 56*56*128=400K params: 0
CONV3-256: [56x56x256] memory: 56*56*256=800K params: (3*3*128)*256 = 294,912 CONV3-
256: [56x56x256] memory: 56*56*256=800K params: (3*3*256)*256 = 589,824 CONV3-256:
[56x56x256] memory: 56*56*256=800K params: (3*3*256)*256 = 589,824
POOL2: [28x28x256] memory: 28*28*256=200K params: 0
CONV3-512: [28x28x512] memory: 28*28*512=400K params: (3*3*256)*512 = 1,179,648 CONV3-
512: [28x28x512] memory: 28*28*512=400K params: (3*3*512)*512 = 2,359,296 CONV3-512:
[28x28x512] memory: 28*28*512=400K params: (3*3*512)*512 = 2,359,296
POOL2: [14x14x512] memory: 14*14*512=100K params: 0
CONV3-512: [14x14x512] memory: 14*14*512=100K params: (3*3*512)*512 = 2,359,296 CONV3-
512: [14x14x512] memory: 14*14*512=100K params: (3*3*512)*512 = 2,359,296 CONV3-512:
[14x14x512] memory: 14*14*512=100K params: (3*3*512)*512 = 2,359,296
POOL2: [7x7x512] memory: 7*7*512=25K params: 0
FC: [1x1x4096] memory: 4096 params: 7*7*512*4096 = 102,760,448 FC: VGG16
[1x1x4096] memory: 4096 params: 4096*4096 = 16,777,216
FC: [1x1x1000] memory: 1000 params: 4096*1000 = 4,096,000
TOTAL memory: 24M * 4 bytes ~= 96MB / image (only forward! ~*2 for bwd)
Details:
- ILSVRC’14 2nd in classification, 1st in
localization
- Similar training procedure as
Krizhevsky 2012
- No Local Response Normalisation (LRN)
- Use VGG16 or VGG19 (VGG19 only
slightly better, more memory)
- Use ensembles for best results
- FC7 features generalize well to other
tasks
AlexNet VGG16 VGG19
29
ImageNet Large Scale Visual Recognition Challenge (ILSVRC)
winners
19 layers 22 layers
30
Case Study: GoogLeNet
[Szegedy et al., 2014]
- 22 layers
- Efficient “Inception” module
- No FC layers
- Only 5 million parameters!
12x less than AlexNet Inception module
- ILSVRC’14 classification
winner (6.7% top 5 error)
31
Case Study: GoogLeNet
[Szegedy et al., 2014]
Inception module
32
Case Study: GoogLeNet
[Szegedy et al., 2014]
33
Case Study: GoogLeNet
[Szegedy et al., 2014]
34
Case Study: GoogLeNet Q: What is the problem with this?
[Szegedy et al., 2014] [Hint: Computational complexity]
Example:
Module input:
28x28x256
35
Case Study: GoogLeNet Q: What is the problem with this?
[Szegedy et al., 2014] [Hint: Computational complexity]
Module input:
28x28x256
36
Case Study: GoogLeNet Q: What is the problem with this?
[Szegedy et al., 2014] [Hint: Computational complexity]
28x28x128
Module input:
28x28x256
37
Case Study: GoogLeNet Q: What is the problem with this?
[Szegedy et al., 2014] [Hint: Computational complexity]
28x28x128
Module input:
28x28x256
38
Case Study: GoogLeNet Q: What is the problem with this?
[Szegedy et al., 2014] [Hint: Computational complexity]
Module input:
28x28x256
39
Case Study: GoogLeNet Q: What is the problem with this?
[Szegedy et al., 2014] [Hint: Computational complexity]
Module input:
28x28x256
40
Case Study: GoogLeNet Q: What is the problem with this?
[Szegedy et al., 2014] [Hint: Computational complexity]
28x28x(128+192+96+256) = 28x28x672
Module input:
28x28x256
41
Case Study: GoogLeNet Q: What is the problem with this?
[Szegedy et al., 2014] [Hint: Computational complexity]
Module input:
28x28x256
42
Case Study: GoogLeNet Q: What is the problem with this?
[Szegedy et al., 2014] [Hint: Computational complexity]
43
Case Study: GoogLeNet Q: What is the problem with this?
[Szegedy et al., 2014] [Hint: Computational complexity]
Module input:
28x28x256
44
Reminder: 1x1 convolutions
1x1 CONV
56 with 32 filters
56
(each filter has size
1x1x64, and performs a
64-dimensional dot
56 product)
56
64 32
45
Reminder: 1x1 convolutions
1x1 CONV
56 with 32 filters
56
preserves spatial
dimensions, reduces depth!
46
Case Study: GoogLeNet
[Szegedy et al., 2014]
47
Case Study: GoogLeNet
[Szegedy et al., 2014]
1x1 conv “bottleneck”
layers
48
Case Study: GoogLeNet Using same parallel layers as
naive example, and adding “1x1
[Szegedy et al., 2014]
conv, 64 filter” bottlenecks:
28x28x480
Conv Ops:
[1x1 conv, 64] 28x28x64x1x1x256
[1x1 conv, 64] 28x28x64x1x1x256
28x28x128 28x28x192 28x28x96 28x28x64
[1x1 conv, 128] 28x28x128x1x1x256
[3x3 conv, 192] 28x28x192x3x3x64
[5x5 conv, 96] 28x28x96x5x5x64
28x28x64 28x28x64 28x28x256 [1x1 conv, 64] 28x28x64x1x1x256
Total: 358M ops
Module input:
Compared to 854M ops for naive version
Bottleneck can also reduce depth after
28x28x256
pooling layer
Inception module with dimension reduction
49
Case Study: GoogLeNet
[Szegedy et al., 2014]
Inception module
50
Case Study: GoogLeNet
[Szegedy et al., 2014]
Full GoogLeNet
architecture
- 22 layers
- Efficient “Inception” module
- No FC layers
- 12x less params than AlexNet
- ILSVRC’14 classification
Deeper networks, with computational efficiency winner (6.7% top 5 error) 51
ImageNet Large Scale Visual Recognition Challenge (ILSVRC) winners
“Revolution of Depth”
19 layers 22 layers
52
Case Study: ResNet
[He et al., 2015]
relu
Very deep networks using residual F(x) + x
connections
..
.
- 152-layer model for ImageNet X
F(x)
- ILSVRC’15 classification relu
identity
winner (3.57% top 5 error)
- Swept all classification and
detection competitions in X
ILSVRC’15 and COCO’15! Residual block
53
Case Study: ResNet
[He et al., 2015]
54
Case Study: ResNet
[He et al., 2015]
55
Case Study: ResNet
[He et al., 2015]
56
Case Study: ResNet
[He et al., 2015]
57
Case Study: ResNet
[He et al., 2015]
Solution: Use network layers to fit a residual mapping instead of directly trying to fit
a desired underlying mapping
relu
H(x) F(x) + x
F(x) X
relu relu
identity
X X
“Plain” layers Residual block
58
Case Study: ResNet
[He et al., 2015]
Solution: Use network layers to fit a residual mapping instead of directly trying to fit
a desired underlying mapping
H(x) = F(x) + x relu
H(x) F(x) + x
Use layers to
fit residual
F(x) relu
X F(x) = H(x) - x
relu identity instead of
H(x) directly
X X
“Plain” layers Residual block
59
Case Study: ResNet
[He et al., 2015]
F(x) X
relu
identity
X
Residual block
60
Case Study: ResNet
[He et al., 2015]
X
Residual block
61
Case Study: ResNet
[He et al., 2015]
- Periodically, double # of
filters and downsample F(x) X
relu
spatially using stride 2 identity
(/2 in each dimension)
- Additional conv layer at
the beginning X
Residual block
Beginning
conv layer
62
Case Study: ResNet layers
No FC
besides FC
[He et al., 2015] 1000 to
output
classes
Full ResNet architecture: Global
relu average
- Stack residual blocks pooling layer
F(x) + x
- Every residual block has ..
after last
conv layer
two 3x3 conv layers .
- Periodically, double # of
filters and downsample F(x) X
relu
spatially using stride 2 identity
(/2 in each dimension)
- Additional conv layer at
the beginning X
- No FC layers at the end Residual block
(only FC 1000 to output
classes)
63
Case Study: ResNet
[He et al., 2015]
64
Case Study: ResNet
[He et al., 2015]
65
Case Study: ResNet
[He et al., 2015]
66
Case Study: ResNet
[He et al., 2015]
Experimental Results
- Able to train very deep
networks without degrading
(152 layers on ImageNet, 1202
on Cifar)
- Deeper networks now achieve
lowing training error as
expected
- Swept 1st place in all ILSVRC
and COCO 2015 competitions
ILSVRC 2015 classification winner
(3.6% top 5 error) -- better than “human
performance”! (Russakovsky 2014)
67
ImageNet Large Scale Visual Recognition Challenge (ILSVRC) winners
19 layers 22 layers
68
Comparing complexity...
Figures copyright Alfredo Canziani, Adam Paszke, Eugenio Culurciello, 2017. Reproduced with permission.
69
Comparing complexity... Inception-v4: Resnet + Inception!
Figures copyright Alfredo Canziani, Adam Paszke, Eugenio Culurciello, 2017. Reproduced with permission.
70
VGG: Highest
Comparing complexity... memory, most
operations
Figures copyright Alfredo Canziani, Adam Paszke, Eugenio Culurciello, 2017. Reproduced with permission.
71
GoogLeNet:
Comparing complexity... most efficient
Figures copyright Alfredo Canziani, Adam Paszke, Eugenio Culurciello, 2017. Reproduced with permission.
72
AlexNet:
Comparing complexity... Smaller compute, still memory
heavy, lower accuracy
Figures copyright Alfredo Canziani, Adam Paszke, Eugenio Culurciello, 2017. Reproduced with permission.
73
ResNet:
Comparing complexity... Moderate efficiency depending on
model, highest accuracy
Figures copyright Alfredo Canziani, Adam Paszke, Eugenio Culurciello, 2017. Reproduced with permission.
Fei-Fei Li & Justin Johnson & Serena Yeung Lecture 9 - May 1, 2018 74
Beyond ResNets...
Densely Connected Convolutional Networks
[Huang et al. 2017]
Lecture 9 - 75
Summary CNN-Architectures
76