Overview
• Introduction to local features
• Harris interest points + SSD, ZNCC, SIFT
• Scale & affine invariant interest point detectors
• Evaluation and comparison of different detectors
• Region descriptors and their performance
Scale invariance - motivation
• Description regions have to be adapted to scale changes
• Interest points have to be repeatable for scale changes
Harris detector + scale changes
Repeatability rate
| {(a i , b i ) | dist ( H (a i ), b i ) } |
R ( )
max(| a i |, | b i |)
Scale adaptation
Scale change between two images
x1 x2 sx1
I1 I 2 I 2
y1 y2 sy1
Scale adapted derivative calculation
Scale adaptation
Scale change between two images
x1 x2 sx1
I1 I 2 I 2
y1 y2 sy1
Scale adapted derivative calculation
x1 x2
I1 Gi i ( ) s I 2 Gi i ( s )
nn
y1 y2
1 n 1 n
Scale adaptation
x ( ) Lx Ly ( )
~ L2
G ( )
L L
x y ( ) L 2
y ( )
where Li ( ) are the derivatives with Gaussian convolution
Scale adaptation
x ( ) Lx Ly ( )
~ L2
G ( )
L L
x y ( ) L 2
y ( )
where Li ( ) are the derivatives with Gaussian convolution
Scale adapted auto-correlation matrix
x ( s ) Lx Ly ( s )
2
~ L
s G ( s )
2
L L
x y ( s ) L 2
y ( s )
Harris detector – adaptation to scale
R ( ) {(a i , b i ) | dist ( H (a i ), b i ) }
Multi-scale matching algorithm
s 1
s3
s5
Multi-scale matching algorithm
s 1
8 matches
Multi-scale matching algorithm
Robust estimation of a global
s 1
affine transformation
3 matches
Multi-scale matching algorithm
s 1
3 matches
s3
4 matches
Multi-scale matching algorithm
s 1
3 matches
s3
4 matches
highest number of matches
correct scale s5
16 matches
Matching results
Scale change of 5.7
Matching results
100% correct matches (13 matches)
Scale selection
• For a point compute a value (gradient, Laplacian etc.) at
several scales
• Normalization of the values with the scale factor
e.g. Laplacian | s 2 ( Lxx Lyy ) |
• Select scale s at the maximum → characteristic scale
| s 2 ( Lxx Lyy ) |
scale
• Exp. results show that the Laplacian gives best results
Scale selection
• Scale invariance of the characteristic scale
s
norm. Lap.
scale
Scale selection
• Scale invariance of the characteristic scale
s
norm. Lap.
norm. Lap.
scale scale
• Relation between characteristic scales s s1 s2
Scale-invariant detectors
• Harris-Laplace (Mikolajczyk & Schmid’01)
• Laplacian detector (Lindeberg’98)
• Difference of Gaussian (Lowe’99)
Harris-Laplace Laplacian
Harris-Laplace
multi-scale Harris points
selection of points at
maximum of Laplacian
invariant points + associated regions [Mikolajczyk & Schmid’01]
Matching results
213 / 190 detected interest points
Matching results
58 points are initially matched
Matching results
32 points are matched after verification – all correct
Matching results
all matches are correct (33)
Laplacian of Gaussian (LOG)
LOG Gxx ( ) Gyy ( )
LOG detector
Detection of maxima and minima of Laplacian in scale space
Difference of Gaussian (DOG)
• Difference of Gaussian approximates the Laplacian
DOG G (k ) G ( )
DOG detector
• Fast computation, scale space processed one octave at a
time
Local features - overview
• Scale invariant interest points
• Affine invariant interest points
• Evaluation of interest points
• Descriptors and their evaluation
Affine invariant regions - Motivation
• Scale invariance is not sufficient for large baseline changes
detected scale invariant region
projected regions, viewpoint changes can locally
be approximated by an affine transformation A
Affine invariant regions - Motivation
Affine invariant regions - Example
Harris/Hessian/Laplacian-Affine
• Initialize with scale-invariant Harris/Hessian/Laplacian
points
• Estimation of the affine neighbourhood with the second
moment matrix [Lindeberg’94]
• Apply affine neighbourhood estimation to the scale-
invariant interest points [Mikolajczyk & Schmid’02,
Schaffalitzky & Zisserman’02]
• Excellent results in a recent comparison
Affine invariant regions
• Based on the second moment matrix (Lindeberg’94)
Lx (x, D ) Lx Ly (x, D )
2
M (x, I , D ) D G( I )
2
Lx Ly (x, D ) Ly (x, D )
2
• Normalization with eigenvalues/eigenvectors
x M x 2
Affine invariant regions
x R Ax L
1
1
xL M xL
2
L xR M xR 2
R
x R Rx L
Isotropic neighborhoods related by image rotation
Affine invariant regions - Estimation
• Iterative estimation – initial points
Affine invariant regions - Estimation
• Iterative estimation – iteration #1
Affine invariant regions - Estimation
• Iterative estimation – iteration #2
Affine invariant regions - Estimation
• Iterative estimation – iteration #3, #4
Harris-Affine versus Harris-Laplace
Harris-Affine Harris-Laplace
Harris/Hessian-Affine
Harris-Affine
Hessian-Affine
Harris-Affine
Hessian-Affine
Matches
22 correct matches
Matches
33 correct matches
Maximally stable extremal regions (MSER) [Matas’02]
• Extremal regions: connected components in a thresholded
image (all pixels above/below a threshold)
• Maximally stable: minimal change of the component
(area) for a change of the threshold, i.e. region remains
stable for a change of threshold
• Excellent results in a recent comparison
Maximally stable extremal regions (MSER)
Examples of thresholded images
high threshold
low threshold
MSER
Overview
• Introduction to local features
• Harris interest points + SSD, ZNCC, SIFT
• Scale & affine invariant interest point detectors
• Evaluation and comparison of different detectors
• Region descriptors and their performance
Evaluation of interest points
• Quantitative evaluation of interest point/region detectors
– points / regions at the same relative location and area
• Repeatability rate : percentage of corresponding points
• Two points/regions are corresponding if
– location error small
– area intersection large
• [K. Mikolajczyk, T. Tuytelaars, C. Schmid, A. Zisserman, J. Matas, F.
Schaffalitzky, T. Kadir & L. Van Gool ’05]
Evaluation criterion
# corresponding regions
repeatability 100%
# detected regions
Evaluation criterion
# corresponding regions
repeatability 100%
# detected regions
intersection
overlap error (1 ) 100%
union
2% 10% 20% 30% 40% 50% 60%
Dataset
• Different types of transformation
– Viewpoint change
– Scale change
– Image blur
– JPEG compression
– Light change
• Two scene types
– Structured
– Textured
• Transformations within the sequence (homographies)
– Independent estimation
Viewpoint change (0-60 degrees )
structured scene
textured scene
Zoom + rotation (zoom of 1-4)
structured scene
textured scene
Blur, compression, illumination
blur - structured scene blur - textured scene
light change - structured scene jpeg compression - structured scene
Comparison of affine invariant detectors
Viewpoint change - structured scene
repeatability % # correspondences
reference image 20 40 60
Comparison of affine invariant detectors
Scale change
repeatability % repeatability %
reference image 2.8 reference image 4
Conclusion - detectors
• Good performance for large viewpoint and scale changes
• Results depend on transformation and scene type, no one best
detector
• Detectors are complementary
– MSER adapted to structured scenes
– Harris and Hessian adapted to textured scenes
• Performance of the different scale invariant detectors is very similar
(Harris-Laplace, Hessian-Laplace, LoG and DOG)
• Scale-invariant detector sufficient up to 40 degrees of viewpoint
change
Overview
• Introduction to local features
• Harris interest points + SSD, ZNCC, SIFT
• Scale & affine invariant interest point detectors
• Evaluation and comparison of different detectors
• Region descriptors and their performance
Region descriptors
• Normalized regions are
– invariant to geometric transformations except rotation
– not invariant to photometric transformations
Descriptors
• Regions invariant to geometric transformations except
rotation
– rotation invariant descriptors
– normalization with dominant gradient direction
• Regions not invariant to photometric transformations
– invariance to affine photometric transformations
– normalization with mean and standard deviation of the image patch
Descriptors
• Sampled image patch
– descriptor dimension is 81
• Gaussian derivative-based descriptors
– Differential invariants (Koenderink and van Doorn’87) (dim. 8)
*
Descriptors
• Gaussian derivative-based descriptors
– Steerable filters (Freeman and Adelson’91)
– “Steering the derivatives in the direction of an angle “
f ' ( ) I x cos I y sin
f ' ' ( ) I xx cos 2 2 IxIy sin cos I yy sin 2
n, i g i /(n 1) i 0...n
• Dominant gradient direction is rotation invariant
Descriptors
• SIFT [Lowe’99]
– 8 orientations of the gradient (dim. 128)
– 4x4 spatial grid
– normalization of the descriptor to norm one
3D histogram
image patch gradient x
y
Descriptors
• Moment invariants [Van Gool et al.’96]
• Shape context [Belongie et al.’02]
• SIFT with PCA dimensionality reduction
• Gradient PCA [Ke and Sukthankar’04]
Comparison criterion
• Descriptors should be
– Distinctive
– Robust to changes on viewing conditions as well as to errors of
the detector
• Detection rate (recall) 1
– #correct matches / #correspondences
• False positive rate
– #false matches / #all matches
• Variation of the distance threshold
– distance (d1, d2) < threshold
1
[K. Mikolajczyk & C. Schmid, PAMI’05]
Viewpoint change (60 degrees)
* *
Scale change (factor 2.8)
* *
Conclusion - descriptors
• SIFT based descriptors perform best
• Significant difference between SIFT and low dimension
descriptors as well as cross-correlation
• Robust region descriptors better than point-wise
descriptors
• Performance of the descriptor is relatively independent of
the detector
Available on the internet
[Link]
• Binaries for detectors and descriptors
– Building blocks for recognition systems
• Carefully designed test setup
– Dataset with transformations
– Evaluation code in matlab
– Benchmark for new detectors and descriptors