0% found this document useful (0 votes)
7 views62 pages

R-CNNs and Deep Learning for Object Detection

The document presents a guest lecture by Ross Girshick on object detection, deep learning, and Region-based Convolutional Networks (R-CNNs). It covers the evolution of object detection techniques, the role of Convolutional Neural Networks (CNNs), and the evaluation metrics used in detection tasks. The lecture highlights the advancements in detection accuracy over the years and introduces R-CNNs as a significant development in the field.

Uploaded by

carlosjulioph
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views62 pages

R-CNNs and Deep Learning for Object Detection

The document presents a guest lecture by Ross Girshick on object detection, deep learning, and Region-based Convolutional Networks (R-CNNs). It covers the evolution of object detection techniques, the role of Convolutional Neural Networks (CNNs), and the evaluation metrics used in detection tasks. The lecture highlights the advancements in detection accuracy over the years and introduces R-CNNs as a significant development in the field.

Uploaded by

carlosjulioph
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd

Object detection,

deep learning, and


R-CNNs
Ross Girshick
Microsoft Research

Guest lecture for UW CSE 455


Nov. 24, 2014
Outline
• Object detection
• the task, evaluation, datasets

• Convolutional Neural Networks (CNNs)


• overview and history

• Region-based Convolutional Networks (R-CNNs)


Image classification
• classes
• Task: assign correct class label to the whole image

Digit classification (MNIST) Object recognition (Caltech-101)


Classification vs. Detection
 Dog

Dog
Dog
Problem formulation
{ airplane, bird, motorbike, person, sofa }
person

motorbik
e

Input Desired output


Evaluating a detector

Test image (previously unseen)


First detection ...

0.9

‘person’ detector predictions


Second detection ...

0.9

0.6

‘person’ detector predictions


Third detection ...
0.2

0.9

0.6

‘person’ detector predictions


Compare to ground truth
0.2

0.9

0.6

‘person’ detector predictions


ground truth ‘person’ boxes
Sort by confidence
0.9 0.8 0.6 0.5 0.2 0.1

... ... ... ... ...

✓ X ✓ ✓ X X
true false
positive positive
(high overlap) (no overlap,
low overlap, or
duplicate)
Evaluation metric
0.9 0.8 0.6 0.5 0.2 0.1

... ... ... ... ...

✓ X ✓ ✓ X X
𝑡

¿ 𝑡𝑟𝑢𝑒 𝑝𝑜𝑠𝑖𝑡𝑖𝑣𝑒𝑠@ 𝑡 ✓
𝑝𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛@ 𝑡=
¿ 𝑡𝑟𝑢𝑒𝑝𝑜𝑠𝑖𝑡𝑖𝑣𝑒𝑠@ 𝑡+¿ 𝑓𝑎𝑙𝑠𝑒 𝑝𝑜𝑠𝑖𝑡𝑖𝑣𝑒𝑠@𝑡 ✓ +X
¿ 𝑡𝑟𝑢𝑒 𝑝𝑜𝑠𝑖𝑡𝑖𝑣𝑒𝑠@ 𝑡
𝑟𝑒𝑐𝑎𝑙𝑙 @𝑡=
¿ 𝑔𝑟𝑜𝑢𝑛𝑑 𝑡𝑟𝑢𝑡h 𝑜𝑏𝑗𝑒𝑐𝑡𝑠
Evaluation metric
0.9 0.8 0.6 0.5 0.2 0.1

... ... ... ... ...

✓ X ✓ ✓ X X

Average Precision (AP)


0% is worst
100% is best

mean AP over classes


(mAP)
Histograms of Oriented Gradients for Human Detection,
Dalal and Triggs, CVPR 2005

Pedestrians
AP ~77%
More sophisticated methods: AP ~90%

(a) average gradient image over training examples


(b) each “pixel” shows max positive SVM weight in the block centered on that pixel
(c) same as (b) for negative SVM weights
(d) test image
(e) its R-HOG descriptor
(f) R-HOG descriptor weighted by positive SVM weights
(g) R-HOG descriptor weighted by negative SVM weights
Why did it work?

Average gradient image


Generic categories

Can we detect people, chairs, horses, cars, dogs, buses, bottles, sheep …?
PASCAL Visual Object Categories (VOC) dataset
Generic categories
Why doesn’t this work (as well)?

Can we detect people, chairs, horses, cars, dogs, buses, bottles, sheep …?
PASCAL Visual Object Categories (VOC) dataset
Quiz time
Warm up

This is an average image of which object class?


Warm up

pedestrian
A little harder

?
A little harder

?
Hint: airplane, bicycle, bus, car, cat, chair, cow, dog, dining table
A little harder

bicycle (PASCAL)
A little harder, yet

?
A little harder, yet

?
Hint: white blob on a green background
A little harder, yet

sheep (PASCAL)
Impossible?

?
Impossible?

dog (PASCAL)
Impossible?

dog (PASCAL)
Why does the mean look like this?
There’s no alignment between the examples!
How do we combat this?
PASCAL VOC detection
history
70%

60%
mean Average Precision (mAP)

50%
41% 41%
40% 37%
DPM++,Selective
30% 28% DPM++
MKL, Search,
23%
DPM, Selective
DPM++,
20% 17% DPM, MKL Search MKL
HOG+
10% DPM
BOW
0%
2006 2007 2008 2009 2010 2011 2012 2013 2014 2015

year
Part-based models & multiple features (MKL)

70%

60%
mean Average Precision (mAP)

50% e nts
m
p rove 41% 41%
i m
40% ance 37%
r m
p erfo DPM++,Selective
id 28% DPM++
30% rap 23%
MKL, Search,
DPM, Selective
DPM++,
20% 17% DPM, MKL Search MKL
HOG+
10% DPM
BOW
0%
2006 2007 2008 2009 2010 2011 2012 2013 2014 2015

year
Kitchen-sink approaches
70%

60%
mean Average Precision (mAP)

increasing complexity & plateau


50%
41% 41%
40% 37%
DPM++,Selective
30% 28% DPM++
MKL, Search,
23%
DPM, Selective
DPM++,
20% 17% DPM, MKL Search MKL
HOG+
10% DPM
BOW
0%
2006 2007 2008 2009 2010 2011 2012 2013 2014 2015

year
Region-based Convolutional Networks (R-CNNs)

70%
62%
60%
mean Average Precision (mAP)

53%R-CNN v2
50%
R-CNN v1
41% 41%
40% 37%
DPM++,Selective
30% 28% DPM++
MKL, Search,
23%
DPM, Selective
DPM++,
20% 17% DPM, MKL Search MKL
HOG+
10% DPM
BOW
0%
2006 2007 2008 2009 2010 2011 2012 2013 2014 2015

year

[R-CNN. Girshick et al. CVPR 2014]


Region-based Convolutional Networks (R-CNNs)

70%

60%
mean Average Precision (mAP)

50% ~1 year

40%

30% ~5 years

20%

10%

0%
2006 2007 2008 2009 2010 2011 2012 2013 2014 2015

year

[R-CNN. Girshick et al. CVPR 2014]


Convolutional Neural
Networks
• Overview
Standard Neural Networks
“Fully connected”

1
𝒙 =( 𝑥 1 , … , 𝑥 784 )𝑇 𝑧 𝑗=𝑔(𝒘 ¿ ¿ 𝑗 𝑇 𝒙)¿ 𝑔 ( 𝑡 )= −𝑡
1+𝑒
From NNs to Convolutional
NNs
• Local connectivity
• Shared (“tied”) weights
• Multiple feature maps
• Pooling
Convolutional NNs
• Local connectivity
compare

• Each green unit is only connected to (3)


neighboring blue units
Convolutional NNs
• Shared (“tied”) weights

𝑤1
𝑤2
𝑤3
• All green units share the same parameters

• Each green unit computes the same function,


𝑤1
but with a different input window
𝑤2
𝑤3
Convolutional NNs
• Convolution with 1-D filter:

𝑤1
𝑤2
𝑤3
• All green units share the same parameters

• Each green unit computes the same function,


but with a different input window
Convolutional NNs
• Convolution with 1-D filter:

𝑤1
𝑤2 • All green units share the same parameters
𝑤3
• Each green unit computes the same function,
but with a different input window
Convolutional NNs
• Convolution with 1-D filter:

𝑤1 • All green units share the same parameters


𝑤2
𝑤3 • Each green unit computes the same function,
but with a different input window
Convolutional NNs
• Convolution with 1-D filter:

• All green units share the same parameters


𝑤1
• Each green unit computes the same function,
𝑤2
but with a different input window
𝑤3
Convolutional NNs
• Convolution with 1-D filter:

• All green units share the same parameters

• Each green unit computes the same function,


𝑤1 but with a different input window
𝑤2
𝑤3
Convolutional NNs
• Multiple feature maps
𝑤 ′1
𝑤 ′2
𝑤 ′3 • All orange units compute the same function
but with a different input windows

• Orange and green units compute


different functions
𝑤1
Feature map 2
𝑤2
(array of orange
𝑤3 units)
Feature map 1
(array of green
units)
Convolutional NNs
• Pooling (max, average)

1
4 • Pooling area: 2 units
4
• Pooling stride: 2 units
0
3
3 • Subsamples feature maps
2D input

Pooling

Convolution

Image
1989

Backpropagation applied to handwritten zip code recognition,


Lecun et al., 1989
Historical perspective –
1980
Historical perspective –
1980
Hubel and Wiesel
1962

Included basic ingredients of ConvNets, but no supervised learning algorithm


Supervised learning – 1986
Gradient descent training with error backpropagation

Early demonstration that error backpropagation can be used


for supervised training of neural nets (including ConvNets)
Supervised learning – 1986

“T” vs. “C” problem Simple ConvNet


Practical ConvNets

Gradient-Based Learning Applied to Document Recognition,


Lecun et al., 1998
Demo
• http://
[Link]/people/karpathy/convnetjs/demo/
[Link]
• ConvNetJS by Andrej Karpathy (Ph.D. student at
Stanford)

Software libraries
• Caffe (C++, python, matlab)
• Torch7 (C++, lua)
• Theano (python)
The fall of ConvNets
• The rise of Support Vector Machines (SVMs)
• Mathematical advantages (theory, convex
optimization)
• Competitive performance on tasks such as digit
classification
• Neural nets became unpopular in the mid 1990s
The key to SVMs
• It’s all about the features
HOG features SVM weights
(+) (-)

Histograms of Oriented Gradients for Human Detection,


Dalal and Triggs, CVPR 2005
Core idea of “deep
learning”
• Input: the “raw” signal (image, waveform, …)

• Features: hierarchy of features is learned from the


raw input
• If SVMs killed neural nets, how did they come back
(in computer vision)?
What’s new since the
1980s?
• More layers
• LeNet-3 and LeNet-5 had 3 and 5 learnable layers
• Current models have 8 – 20+
𝑔 ( 𝑥)
• “ReLU” non-linearities (Rectified Linear Unit)

• Gradient doesn’t vanish 𝑥

• “Dropout” regularization
• Fast GPU implementations
• More data
Ross’s Own System: Region CNNs
Competitive
Results
Top Regions for Six Object Classes

You might also like