Real-Time Object Detection with ConvNets
Real-Time Object Detection with ConvNets
net/publication/335095635
CITATIONS READS
6 306
3 authors, including:
Chih-Hung Li
National Taipei University of Technology
20 PUBLICATIONS 112 CITATIONS
SEE PROFILE
All content following this page was uploaded by Chih-Hung Li on 10 August 2019.
*This work was supported by the National Taipei University of The authors are with the Graduate Institute of Manufacturing Technology,
Technology – King Mongkut’s University of Technology Thonburi Joint National Taipei University of Technology, Taipei, 10608 Taiwan ROC (e-
Research Program: NTUT-KMUTT-107-02. mail: cL4e@[Link]).
photo on the target object is needed and the rest of the data level of abstraction, and makes it possible to identify
preparation process is automated. The average detection time homogeneous but highly variable signals [14], [15].
of our system is less than 0.032 s, making it possible to run in Investigation results also show that it outperforms SIFT on
real time. Low demand on computation and memory is also the descriptor matching [16] and other methods in visual
merit. The rest of the paper is organized as follows. Section II recognition, classification, and detection [17]. The robustness
describes related prior work. Details of the proposed method of ConvNet in image recognition is also one good reason to
and the static experiments are detailed in Section III. adopt it as the feature descriptor [18]. For example, effects of
Manipulator experiments are reported in Section IV. The illumination and view-point difference on place recognition
conclusions are in Section V. were studied; images with high variations due to season and
illumination differences were successfully recognized using a
II. RELATED WORK ConvNet [11]. In object pose estimation, techniques similar to
template matching are still widely adopted in ConvNet-related
Automatic visual positioning has been applied to mobile approaches. For example, R-CNN [8] uses region proposal
robots for fetch-and-delivery, equipment operation [1], SLAM methods to generate potential bounding boxes and then
[2], et al. It is also in great demand in automated production execute a classifier on these proposed boxes. However, sliding
lines such as workpiece placement, product assembly, AOI, windows and running a ConvNet on each window takes a
machinery operation, and more [3]. Traditional visual significant amount of processing time. Redmon et al. [12]
positioning techniques often involve feature detection [4] or reframed detection as a regression problem and achieved a
template matching [5] - [8]. Take pressing an elevator button very high detection speed in YOLO, which sees the entire
as an example, Wang et al. [5] used binary images of the image during training and test time, so it implicitly encodes
characters processed and captured from the whole image to contextual information about classes as well as their
perform pattern matching in order to detect the location of the appearance. As it learns object positions from manually
button. Troniak et al. [6] implemented a template matching annotated training sets, preprocessing the images can still be a
algorithm based on the FastMatchTemplate technique, which tedious task. To be relieved from heavily relying on manually
has been shown to be robust to small changes in lighting and annotated data, Oquab et al. [19] introduced multi-scale
viewing angle. Kang et al. [7] also used template matching for sliding-window training to train from images rescaled to
recognition of the elevator status; a threshold value was multiple different sizes and to treat fully connected adaptation
defined to reject non-objects. However, the feature templates layers as convolutions. They found that ConvNet learns to
often do not react to observations that are subject to localize objects despite having no object location annotation at
environmental or view-point changes [9], [10]. In addition, training. However, it is not clear whether their scheme can
template matching requires sliding the template across the provide the level of precision in coordinate estimation that
image and comparing the template to the local pixels; it often satisfies industrial applications.
consumes considerable computation time [10].
To devise a suitable method for embedded object detection
Recently, Deep ConvNet, which exhibits outstanding and precision coordinate estimation, we propose a method to
image recognition capabilities has been applied in many object adopt ConvNet for visual feature learning. However, the
or place recognition tasks [8], [11] - [13]. It provides higher scheme is not based on template matching nor bounding
Rescale …
(0, 0, 0) (0, 0, 0)
Bound and Annotate
(0, 0, 0) (0, 0, 0)
(1, 0, 0) (1, 0, 0) (1, 0, 0)
(1, 0, 0)
(2, 0, 0) (2, 0, 0) (2, 0, 0)
.. .. (2, 0, 0) ..
. . ..
. . . .
.
(xn, 0, 0) .
(xn, 0, 0) (xn, 0, 0) (xn, 0, 0)
(0, 1, 0) (0, 1, 0) (0, 1, 0)
(0, 1, 0)
(1, 1, 0) (1, 1, 0) (1, 1, 0)
.. .. (1, 1, 0) ..
. . ..
. . . ..
(xn, yn, 0) .
(xn, yn, 0) (xn, yn, 0) (xn, yn, 0)
Camera
ConvNet
For Z
Detection
Figure 2. Illustration of the automatic training data generation and organization frameworks for the proposed 3D coordinate detection scheme. By
altering the size of the cropping window, images of the target object with different scales are generated and annotated automatically. By further sliding
a bounding window, the target object appears at different locations of strictly-defined plane coordinates. Here, each column contains images of the target
object of the same scale but different plane locations. By labeling each column of images with a different depth distance, a ConvNet for depth detection
is trained.
regarded as belonging to the same class of depth distance. By
…… Y=0
combining various classes of depth distances, the data are
organized for training the depth coordinate detection (DCD).
For example, in Fig. 2, images in the left column are labeled
………
………… Y=1 ConvNet
as the deepest and the right column the nearest; a total of 𝑛𝑧 +
For Y
1 levels of depth are constructed.
………
Detection
60%
50%
detection rate
(c)
40%
30%
20%
10%
0% (d)
0 1 2 3 4 5 6 7 8
detected depth
ground truth
0 1 2 3 4 5 6 7
Figure 6. The DCD success rates of the non-rotation cases. For each
depth ground truth, the detection rate of each depth level is presented.
For example, when the ground truth is 4, 55% detects 4, 27% detects 5, Figure 7. Testing images of various viewing angles: (a) yaw, (b) pitch, and
16% detects 3, 2% detects 6, and 0% detects the others. (c) roll. (d) Sample images of out-of-range detections.
60
40
time (s)
20
0
0 degree yaw 15 yaw 30 pitch +15 pitch +30 pitch -15 pitch -30 roll 15 roll 30
degrees degrees degrees degrees degrees degrees degrees degrees
maximum completion time average completion time minimum completion time
Figure 8. Test results of the elevator button-pressing experiment. The average completion time is hardly affected by the attack angle within the orientation
range of ±30 degrees.
cm and the “far” range between 30 cm and 45 cm. Results of where ∆𝜎 denotes the commands on 3D movement. Γ𝜎
the out-of-depth-range PCDs are also shown in Fig. 5 (a). denotes the ConvNet detection, and 𝜋̅𝜎 denotes the policy on
About 80% of the partial inclusion cases still indicated the the manipulator action given the ConvNet detection result
ball-park position of the target object, although the accurate ω𝜎,𝑡 . The overall success rate of the task performance
coordinate estimates might not be provided. evaluated by (5) depends on 𝜋̅𝜎 and Γ𝜎 .
Figure 9. A sample imaging sequence of the Uarm approaching and pressing the elevator button (from left to right) - a successful case of pitch +30
degrees.
[7] J. -G. Kang, S. -Y. An, and S. -Y. Oh, “Navigation strategy for the
service robot in the elevator environment,” in Proc. Int. Conf. Control,
Enlarged Autom. Syst., Seoul, Korea, 2007.
target object
Robot arm [8] R. Girshick, “Fast R-CNN,” Int. Conf. Comput. Vision (ICCV), 2015.
[9] S. Oron, T. Dekel, T. Xue, W. T. Freeman, and S. Avidan, “Best-
Buddies Similarity—robust template matching using mutual nearest
neighbors,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 40, no. 8, pp.
Camera 1799–1813, 2018.
[10] W. Ouyang, F. Tombari, S. Mattoccia, L. Di Stefano, and W. -K. Cham,
“Performance evaluation of full search equivalent pattern matching
algorithms,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 34, no. 1, pp.
PC
127–143, 2012.
[11] N. Sunderhauf, S. Shirazi, F. Dayoub, B. Upcroft, and M. Milford, “On
the performance of ConvNet features for place recognition,” in Proc.
IEEE/RSJ Int. Conf. Intell. Robot. Syst. (IROS), 2015.
Target object [12] J. Redmon, S. Divvala, R. Girshick, A. Farhadi, “You Only Look Once:
unified, real-time object detection,” arXiv:1506.02640v5, 2016.
Figure 10. Setup of the Yaskawa robot arm experiment. The target object [13] Y. Xiang, T. Schmidt, V. Narayanan, and D. Fox, “PoseCNN: a
is movable, and the center of it was randomly placed within a square convolutional neural network for 6D object pose estimation in cluttered
range of 2 × 2 cm2. Only one camera and one target object were used in scenes,” arXiv:1711.00199, 2018.
this test.
[14] Y. Chen, Y. Shen, X. Liu, and B. Zhong, “3D object tracking via image
sets and depth-based occlusion detection,” Signal Processing, vol. ll2,
have a minimum detection increment of 0.5 mm. The center pp. 146–153, 2015.
of the target object was randomly placed within a square range [15] Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature, vol. 521,
of 2 × 2 cm2. For the 100 trials, the robot arm successfully pp. 436–444, 2015.
located the target object within a maximum deviation of ±0.3 [16] P. Fischer, A. Dosovitskiy, and T. Brox, “Descriptor matching with
mm. convolutional neural networks: a comparison to SIFT,” arXiv:
1405.5769, 2014.
V. CONCLUSIONS [17] R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature
hierarchies for accurate object detection and semantic segmentation,”
We proposed an object coordinate detection scheme using in Proc. IEEE Conf. Comput. Vision Pattern Recog. (CVPR), 2014.
rigidly trained ConvNets; we also demonstrated how to build [18] A. S. Razavian, H. Azizpour, J. Sullivan, and S. Carlsson, “CNN
the coordinate-articulate training data based on a single basis features off-the-shelf: an astounding baseline for recognition,” in Proc.
photo. Our experimental evidence shows that the PCD Comput. Vision Pattern Recog. Workshops (CVPRW), 2014.
success rate of perpendicular observation is 100%, and is [19] M. Oquab, L. Bottou, I. Laptev, J. Sivic, “Is object localization for free?
averagely higher than 80% for oblique observations. Using - weakly-supervised learning with Convolutional Neural Networks,”
the proposed real-time detection scheme, a low-precision IEEE Conf. Computer Vision Pattern Recog. (CVPR), 2015.
manipulator was controlled to localize and press a button, [20] S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of
achieving an overall success rate of 98% and an average deep visuomotor policies,” J. Machine Learning Res., vol. 17, pp. 1–40,
completion time of 30 s. When tested on an industrial robot 2016.
arm, a localization precision of ±0.3 mm was achieved with
a 640 × 480 pixels CMOS camera.
REFERENCES
[1] N. Bellotto and H. Hu, “Multisensor-based human detection and
tracking for mobile service robots,” IEEE Trans. Systems, Man,
Cybernetics, Part B: Cybernetics, vol. 39, no. 1, pp. 167–181, February
2009.
[2] M. Veloso, J. Biswas, B. Coltin, and S. Rosenthal, “CoBots: robust
symbiotic autonomous mobile service robots,” in Proc. Twenty-Fourth
Int. Joint Conf. Artificial Intell. (IJCAI), 2015.
[3] Q. Wei, C. Yang, W. Fan, and Y. Zhao, “Design of demonstration-
driven assembling manipulator,” Appl. Sci., vol. 8, no. 5, pp. 797–807,
2018.
[4] J. L. Lowe, “Mobile robot localization and mapping with uncertainty
using scale-invariant visual landmarks,” Int. J. Robot. Res., vol. 21, pp.
735–758, 2002.
[5] W. -J. Wang, C. -H. Huang, I. -H. Lai, and H. -C. Chen, “A robot arm
for pushing elevator buttons,” in Proc. SICE Ann. Conf., August 18-21,
2010.
[6] D. Troniak, J. Sattar, A. Gupta, J. J. Little, W. Chan, E. Calisgan, E.
Croft, and M. Van der Loos, “Charlie rides the elevator–integrating
vision, navigation and manipulation towards multi-floor robot
locomotion,” in Int. Conf. Comput. Robot Vision, 2013.