EPNet: 3D Object Detection Fusion
EPNet: 3D Object Detection Fusion
1 Introduction
The last decade has witnessed significant progress in the 3D object detection
task via different types of sensors, such as monocular images [1,36], stereo cam-
eras [2], and LiDAR point clouds [43,22,39]. Camera images usually contain
plenty of semantic features (e.g., color, texture) while suffering from the lack of
depth information. LiDAR points provide depth and geometric structure infor-
mation, which are quite helpful for understanding 3D scenes. However, LiDAR
points are usually sparse, unordered, and unevenly distributed. Fig. 1(a) illus-
trates a typical example of leveraging the camera image to improve the 3D
detection task. It is challenging to distinguish between the closely packed white
and yellow chairs by only the LiDAR point cloud due to their similar geometric
structure, resulting in chaotically distributed bounding boxes. In this case, uti-
lizing the color information is crucial to locate them precisely. This motivates
us to design an effective module to fuse different sensors for a more accurate 3D
object detector.
However, fusing the representations of LiDAR and camera image is a non-
trivial task for two reasons. On the one hand, they possess highly different data
?
equal contribution
??
corresponding author
2 T. Huang, Z. Liu, et al.
Fig. 1. Illustration of (a) the benefit and (b) potential interference information of the
camera image. (c) demonstrates the inconsistency of the classification confidence and
localization confidence. Green box denotes the ground truth. Blue and yellow boxes
are predicted bounding boxes.
2 Related Work
3D object detection based on camera images. Recent 3D object detection
methods pay much attention to camera images, such as monocular [23,29,12,15,20]
and stereo images [16,35]. Chen et al. [1] obtain 2D bounding boxes with a
CNN-based object detector and infer their corresponding 3D bounding boxes
with semantic, context, and shape information. Mousavian et al. [25] estimate
localization and orientation from 2D bounding boxes of objects by exploiting
the constraint of projective geometry. However, methods based on the camera
image have difficulty in generating accurate 3D bounding boxes due to the lack
of depth information.
3D object detection based on LiDAR. Many LiDAR-based methods [39,24,40]
are proposed in recent years. VoxelNet [43] divides a point cloud into voxels
and employs stacked voxel feature encoding layers to extract voxel features.
4 T. Huang, Z. Liu, et al.
3 Method
𝑆 𝑆
EPNet: Enhancing
C Point Features with Image Semantics
C 5
𝑆
Geometric Stream
𝑵×𝟑
Point
Segmentation
Set Set Set Set Feature Feature Feature Feature 𝑵 × 𝟏𝟐𝟖
Abstraction Abstraction Abstraction Abstraction Propagation Propagation Propagation Propagation
𝑺𝟏 𝑺𝟐 𝑺𝟑 𝑺𝟒 𝑷𝟏 𝑷𝟐 𝑷𝟑 𝑷𝟒 3D Proposal
Generation
Image Stream
𝑭𝟏
𝑭𝟒
C
𝑭𝟐
𝑭𝑼
𝑭𝟏 𝑭𝟐 𝑭𝟑 𝑭𝟒 𝑭𝟑
𝑾×𝑯×𝟑
N×3
FC tanh FC 𝜎 C
Image Sampler
Point-wise Image Feature
Area(D ∩ G)
Lce = −log(c × ) (5)
Area(D ∪ G)
where D and G represents the predicted bounding box and the ground truth.
c denotes the classification confidence for D. Towards optimizing this loss func-
tion, the classification confidence and localization confidence (i.e., the IoU) are
8 T. Huang, Z. Liu, et al.
where Lrpn and Lrcnn denote the training objective for the two-stream RPN and
the refinement network, both of which adopt a similar optimizing goal, including
a classification loss, a regression loss and a CE loss. We adopt the focal loss [19]
as our classification loss to balance the positive and negative samples with the
setting of α = 0.25 and γ = 2.0. For a bounding box, the network needs to
regress its center point (x, y, z), size (l, h, w), and orientation θ.
Since the range of the Y-axis (the vertical axis) is relatively small, we directly
calculate its offset to the ground truth with a smooth L1 loss [7]. Similarly, the
size of the bounding box (h, w, l) is also optimized with a smooth L1 loss. As
for the X-axis, the Z-axis and the orientation θ, we adopt a bin-based regression
loss [31,27]. For each foreground point, we split its neighboring area into several
bins. The bin-based loss first predicts which bin bu the center point falls in, and
then regress the residual offset ru within the bin. We formulate the loss functions
as follows:
Lrpn = Lcls + Lreg + λLcf (7)
Lcls = −α(1 − ct )γ log ct (8)
X X
Lreg = E(bu , bˆu ) + S(ru , rˆu ) (9)
u∈x,z,θ u∈x,y,z,h,w,l,θ
where E and S denote the cross entropy loss and the smooth L1 loss, respectively.
ct is the probability of the point in consideration belong to the ground truth
category. bˆu and rˆu denote the ground truth of the bins and the residual offsets.
4 Experiments
We evaluate our method on two common 3D object detection datasets, including
the KITTI dataset [6] and the SUN-RGBD dataset [33]. KITTI is an outdoor
EPNet: Enhancing Point Features with Image Semantics 9
64 positive candidate boxes which will be refined by the refinement network. For
both datasets, we utilize similar architecture design for the two-stream RPN as
discussed above.
The Training Scheme. Our two-stream RPN and refinement network are end-
to-end trainable. In the training phase, the regression loss Lreg and the CE loss
are only applied to positive proposals, i.e., proposals generated by foreground
points for the RPN stage, and proposals sharing IoU larger than 0.55 with the
ground truth for RCNN stage.
Parameter Optimization. The Adaptive Moment Estimation (Adam) [10] is
adopted to optimize our network. The initial learning rate, weight decay, and
momentum factor are set to 0.002, 0.001, and 0.9, respectively. We train the
model for around 50 epochs on four Titan XP GPUs with a batch size of 12 in
an end-to-end manner. The balancing weights λ in the loss function are set to 5.
Data Augmentation. Three common data augmentation strategies are adopted
to prevent over-fitting, including rotation, flipping, and scale transformations.
First, we randomly rotate the point cloud along the vertical axis within the
range of [−π/18, π/18]. Then, the point cloud is randomly flipped along the
forward axis. Besides, each ground truth box is randomly scaled following the
uniform distribution of [0.95, 1.05]. Many LiDAR-based methods sample ground
truth boxes from the whole dataset and place them into the raw 3D frames to
simulate real scenes with crowded objects following [43,38]. Although effective,
this data augmentation needs the prior information of road plane which is usu-
ally difficult to acquire for kinds of real scenes. Hence, we do not utilize this
augmentation mechanism in our framework for the applicability and generality.
Fig. 4. Visualization of the learned semantic image feature. The image stream mainly
focuses on the foreground objects (cars). The red arrow marks the region under bad
illumination, which show a distinct feature representation to its neighboring region.
<latexit sha1_base64="BGxSzCxroh0Px2WT7inYh/mYkdk=">AAAB/nicbVDLSgMxFL1TX7W+quLKTbAKrspMFXRZcOOyin1AO5RMmmlDM5khyQhlGPBX3LhQxK3f4c6/MTOdhbYeCDmccy85OV7EmdK2/W2VVlbX1jfKm5Wt7Z3dver+QUeFsSS0TUIeyp6HFeVM0LZmmtNeJCkOPE673vQm87uPVCoWigc9i6gb4LFgPiNYG2lYPRoEWE88P8lvgnlyn6aVYbVm1+0caJk4BalBgdaw+jUYhSQOqNCEY6X6jh1pN8FSM8JpWhnEikaYTPGY9g0VOKDKTfL4KTozygj5oTRHaJSrvzcSHCg1CzwzmYVUi14m/uf1Y+1fuwkTUaypIPOH/JgjHaKsCzRikhLNZ4ZgIpnJisgES0y0aSwrwVn88jLpNOrORb1xd1lrnhZ1lOEYTuAcHLiCJtxCC9pAIIFneIU368l6sd6tj/loySp2DuEPrM8fc/eVsw==</latexit>
80
75
0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9
(a) (b)
IoU Loss CE Loss IoU Loss CE Loss
92 90.87 100 97.7 97.0
<latexit sha1_base64="BGxSzCxroh0Px2WT7inYh/mYkdk=">AAAB/nicbVDLSgMxFL1TX7W+quLKTbAKrspMFXRZcOOyin1AO5RMmmlDM5khyQhlGPBX3LhQxK3f4c6/MTOdhbYeCDmccy85OV7EmdK2/W2VVlbX1jfKm5Wt7Z3dver+QUeFsSS0TUIeyp6HFeVM0LZmmtNeJCkOPE673vQm87uPVCoWigc9i6gb4LFgPiNYG2lYPRoEWE88P8lvgnlyn6aVYbVm1+0caJk4BalBgdaw+jUYhSQOqNCEY6X6jh1pN8FSM8JpWhnEikaYTPGY9g0VOKDKTfL4KTozygj5oTRHaJSrvzcSHCg1CzwzmYVUi14m/uf1Y+1fuwkTUaypIPOH/JgjHaKsCzRikhLNZ4ZgIpnJisgES0y0aSwrwVn88jLpNOrORb1xd1lrnhZ1lOEYTuAcHLiCJtxCC9pAIIFneIU368l6sd6tj/loySp2DuEPrM8fc/eVsw==</latexit>
79.59 80
80 78.92 78.73
78 75
Easy Moderate Hard 3D mAP 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9
Fig. 5. Illustration of the ratio of kept positive candidate boxes varying with differ-
ent classification confidence threshold. The CE loss leads to significantly larger ratios
than those of the IoU loss, suggesting its effectiveness in improving the consistency of
localization and classification confidence.
Table 5. Comparisons with state-of-the-art methods on the testing set of the KITTI
dataset (Cars). L and I represent the LiDAR point cloud and the camera image.
3D Detection Bird’s Eye View Orientation
Method Modality
Easy Moderate Hard 3D mAP Easy Moderate Hard BEV mAP Easy Moderate Hard Ori mAP
SECOND [38] L 83.34 72.55 65.82 73.90 89.39 83.77 78.59 83.92 90.93 82.55 73.62 82.37
PointPillars [14] L 82.58 74.31 68.99 75.29 90.07 86.56 82.81 86.48 93.84 90.70 87.47 90.67
TANet [21] L 84.39 75.94 68.82 76.38 91.58 86.54 81.19 86.44 93.52 90.11 84.61 89.41
PointRCNN [31] L 86.96 75.64 70.70 77.77 92.13 87.39 82.72 87.41 95.90 91.77 86.92 91.53
Fast Point R-CNN [4] L 85.29 77.40 70.24 77.64 90.87 87.84 80.52 86.41 - - - -
F-PointNet [27] L+I 82.19 69.79 60.59 70.86 91.17 84.67 74.77 83.54 - - - -
MV3D [3] L+I 74.97 63.63 54.00 64.20 86.62 78.93 69.80 78.45 - - - -
AVOD [11] L+I 76.39 66.47 60.23 67.70 89.75 84.95 78.32 84.34 94.98 89.22 82.14 88.78
AVOD-FPN [11] L+I 83.07 71.76 65.73 73.52 90.99 84.82 79.62 85.14 94.65 88.61 83.71 88.99
ContFuse [18] L+I 83.68 68.78 61.67 71.38 94.07 85.35 75.88 85.10 - - - -
PC-CNN [5] L+I 85.57 73.79 65.65 75.00 91.19 87.40 79.35 85.98 - - - -
MMF [17] L+I 88.40 77.43 70.22 78.68 93.67 88.21 81.99 87.96 - - - -
Ours L+I 89.81 79.28 74.59 81.23 94.22 88.47 83.69 88.79 96.13 94.22 89.68 93.34
where B represents the set of positive candidate boxes. cb denotes the classifi-
cation confidence of the box b. N (·) calculates the number of boxes. It should
be noted that all the boxes in B possess an overlap larger than τ with the cor-
responding ground truth box.
We provide evaluation results on two different settings, i.e., the model trained
with IoU loss and that trained with CE loss. For each frame in the KITTI vali-
dation dataset, the model generates 64 boxes without NMS procedure employed.
Then we get the positive candidate boxes by calculating the overlaps with the
ground truth boxes. We set τ to 0.7 following the evaluation protocol of 3D
detection metric. υ is varied from 0.1 to 0.9 to evaluate the consistency under
different classification confidence thresholds. As is shown in Fig. 5(b), the model
trained with CE loss demonstrates better consistency than that trained with
IoU loss in all the different settings of classification confidence threshold υ.
Table 5 presents quantitative results on the KITTI test set. The proposed
method outperforms multi-sensor based methods F-PointNet [27], MV3D [3],
AVOD-FPN [11], PC-CNN [5], ContFuse [18], and MMF [17] by 10.37%, 17.03%,
7.71%, 6.23%, 9.85% and 2.55% in terms of 3D mAP. It should be noted that
MMF [17] exploits multiple auxiliary tasks (e.g., 2D detection, ground estima-
tion, and depth completion) to boost the 3D detection performance, which re-
quires many extra annotations. These experiments consistently reveal the superi-
ority of our method over the cascading approach [27], as well as fusion approaches
based on RoIs [3,11,5] and voxels [18,17].
We also provide the quantitative results on the KITTI validation split in the
Table 4 for the convenience of comparison with future work. Besides, we present
the qualitative results on the KITTI validation dataset in the supplementary
materials.
14 T. Huang, Z. Liu, et al.
5 Conclusion
We have presented a new 3D object detector named EPNet, which consists of
a two-stream RPN and a refinement network. The two-stream RPN reasons
about different sensors (i.e., LiDAR point cloud and camera image) jointly and
enhances the point features with semantic image features effectively by using
the proposed LI-Fusion module. Besides, we address the issue of inconsistency
between the classification and localization confidence by the proposed CE loss,
which explicitly guarantees the consistency between the localization and classifi-
cation confidence. Extensive experiments have validated the effectiveness of the
LI-Fusion module and the CE loss. In the future, we are going to explore how to
enhance the image feature representation with depth information of the LiDAR
point cloud instead, and its application in 2D detection tasks.
Acknowledgement
This work was supported by National Key R&D Program of China (No.2018YFB
1004600), Xiang Bai was supported by the National Program for Support of Top-
notch Young Professionals and the Program for HUST Academic Frontier Youth
Team 2017QYTD08.
EPNet: Enhancing Point Features with Image Semantics 15
References
1. Chen, X., Kundu, K., Zhang, Z., Ma, H., Fidler, S., Urtasun, R.: Monocular 3d
object detection for autonomous driving. In: Proc. of IEEE Intl. Conf. on Computer
Vision and Pattern Recognition (2016)
2. Chen, X., Kundu, K., Zhu, Y., Ma, H., Fidler, S., Urtasun, R.: 3d object proposals
using stereo imagery for accurate object class detection. IEEE Trans. Pattern Anal.
Mach. Intell. 40(5), 1259–1272 (2017)
3. Chen, X., Ma, H., Wan, J., Li, B., Xia, T.: Multi-view 3d object detection network
for autonomous driving. In: Proc. of IEEE Intl. Conf. on Computer Vision and
Pattern Recognition (2017)
4. Chen, Y., Liu, S., Shen, X., Jia, J.: Fast point r-cnn. In: Porc. of IEEE Intl. Conf.
on Computer Vision (2019)
5. Du, X., Ang, M.H., Karaman, S., Rus, D.: A general pipeline for 3d detection
of vehicles. In: 2018 IEEE International Conference on Robotics and Automation
(ICRA). pp. 3194–3200 (May 2018). [Link]
6. Geiger, A., Lenz, P., Urtasun, R.: Are we ready for autonomous driving. In: Proc.
of IEEE Intl. Conf. on Computer Vision and Pattern Recognition
7. Girshick, R.: Fast r-cnn. In: Proceedings of the IEEE international conference on
computer vision. pp. 1440–1448 (2015)
8. Ioffe, S., Szegedy, C.: Batch normalization: Accelerating deep network training
by reducing internal covariate shift. In: Proc. of Intl. Conf. on Machine Learning
(2015)
9. Jiang, B., Luo, R., Mao, J., Xiao, T., Jiang, Y.: Acquisition of localization confi-
dence for accurate object detection. In: Proc. of European Conference on Computer
Vision (2018)
10. Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization. ICLR (2014)
11. Ku, J., Mozifian, M., Lee, J., Harakeh, A., Waslander, S.L.: Joint 3d proposal
generation and object detection from view aggregation. In: IROS. pp. 1–8. IEEE
(2018)
12. Ku*, J., Pon*, A.D., Waslander, S.L.: Monocular 3d object detection leveraging
accurate proposals and shape reconstruction. In: CVPR (2019)
13. Lahoud, J., Ghanem, B.: 2d-driven 3d object detection in rgb-d images. In: Porc.
of IEEE Intl. Conf. on Computer Vision (2017)
14. Lang, A.H., Vora, S., Caesar, H., Zhou, L., Yang, J., Beijbom, O.: Pointpillars:
Fast encoders for object detection from point clouds. In: Proc. of IEEE Intl. Conf.
on Computer Vision and Pattern Recognition (2019)
15. Li, B., Ouyang, W., Sheng, L., Zeng, X., Wang, X.: Gs3d: An efficient 3d object
detection framework for autonomous driving. In: IEEE Conference on Computer
Vision and Pattern Recognition (CVPR) (2019)
16. Li, P., Chen, X., Shen, S.: Stereo r-cnn based 3d object detection for autonomous
driving. In: CVPR (2019)
17. Liang, M., Yang, B., Chen, Y., Hu, R., Urtasun, R.: Multi-task multi-sensor fusion
for 3d object detection. In: Proc. of IEEE Intl. Conf. on Computer Vision and
Pattern Recognition (2019)
18. Liang, M., Yang, B., Wang, S., Urtasun, R.: Deep continuous fusion for multi-
sensor 3d object detection. In: Proc. of European Conference on Computer Vision
(2018)
19. Lin, T.Y., Goyal, P., Girshick, R., He, K., Dollár, P.: Focal loss for dense object
detection. In: Porc. of IEEE Intl. Conf. on Computer Vision (2017)
16 T. Huang, Z. Liu, et al.
20. Liu, L., Lu, J., Xu, C., Tian, Q., Zhou, J.: Deep fitting degree scoring network
for monocular 3d object detection. In: Proceedings of the IEEE Conference on
Computer Vision and Pattern Recognition. pp. 1057–1066 (2019)
21. Liu, Z., Zhao, X., Huang, T., Hu, R., Zhou, Y., Bai, X.: Tanet: Robust 3d ob-
ject detection from point clouds with triple attention. In: AAAI. pp. 11677–11684
(2020)
22. Luo, W., Yang, B., Urtasun, R.: Fast and furious: Real time end-to-end 3d detec-
tion, tracking and motion forecasting with a single convolutional net. In: Proc. of
IEEE Intl. Conf. on Computer Vision and Pattern Recognition (2018)
23. Ma, X., Wang, Z., Li, H., Zhang, P., Ouyang, W., Fan, X.: Accurate monocular
object detection via color- embedded 3d reconstruction for autonomous driving.
In: Proceedings of the IEEE international Conference on Computer Vision (ICCV)
(2019)
24. Meyer, G.P., Laddha, A., Kee, E., Vallespi-Gonzalez, C., Wellington, C.K.: Laser-
Net: An efficient probabilistic 3D object detector for autonomous driving. In: Pro-
ceedings of the IEEE Conference on Computer Vision and Pattern Recognition
(CVPR) (2019)
25. Mousavian, A., Anguelov, D., Flynn, J., Kosecka, J.: 3d bounding box estimation
using deep learning and geometry. In: Proc. of IEEE Intl. Conf. on Computer
Vision and Pattern Recognition (2017)
26. Qi, C.R., Litany, O., He, K., Guibas, L.J.: Deep hough voting for 3d object detec-
tion in point clouds. Porc. of IEEE Intl. Conf. on Computer Vision (2019)
27. Qi, C.R., Liu, W., Wu, C., Su, H., Guibas, L.J.: Frustum pointnets for 3d object
detection from rgb-d data. In: Proc. of IEEE Intl. Conf. on Computer Vision and
Pattern Recognition (2018)
28. Qi, C.R., Yi, L., Su, H., Guibas, L.J.: Pointnet++: Deep hierarchical feature learn-
ing on point sets in a metric space. In: Advances in neural information processing
systems. pp. 5099–5108 (2017)
29. Qin, Z., Wang, J., Lu, Y.: Monogrnet: A geometric reasoning network for 3d object
localization. The Thirty-Third AAAI Conference on Artificial Intelligence (AAAI-
19) (2019)
30. Ren, Z., Sudderth, E.B.: Three-dimensional object detection and layout prediction
using clouds of oriented gradients. In: Proc. of IEEE Intl. Conf. on Computer
Vision and Pattern Recognition (2016)
31. Shi, S., Wang, X., Li, H.: Pointrcnn: 3d object proposal generation and detection
from point cloud. In: Proc. of IEEE Intl. Conf. on Computer Vision and Pattern
Recognition (2019)
32. Simonelli, A., Bulò, S.R.R., Porzi, L., López-Antequera, M., Kontschieder, P.: Dis-
entangling monocular 3d object detection. arXiv preprint arXiv:1905.12365 (2019)
33. Song, S., Lichtenberg, S.P., Xiao, J.: Sun rgb-d: A rgb-d scene understanding bench-
mark suite. In: Proc. of IEEE Intl. Conf. on Computer Vision and Pattern Recog-
nition (2015)
34. Song, S., Xiao, J.: Deep sliding shapes for amodal 3d object detection in rgb-d
images. In: Proc. of IEEE Intl. Conf. on Computer Vision and Pattern Recognition
(2016)
35. Wang, Y., Chao, W.L., Garg, D., Hariharan, B., Campbell, M., Weinberger, K.:
Pseudo-lidar from visual depth estimation: Bridging the gap in 3d object detection
for autonomous driving. In: CVPR (2019)
36. Xu, B., Chen, Z.: Multi-level fusion based 3d object detection from monocular
images. In: Proc. of IEEE Intl. Conf. on Computer Vision and Pattern Recognition
(2018)
EPNet: Enhancing Point Features with Image Semantics 17
37. Xu, D., Anguelov, D., Jain, A.: Pointfusion: Deep sensor fusion for 3d bounding
box estimation. In: Proc. of IEEE Intl. Conf. on Computer Vision and Pattern
Recognition (2018)
38. Yan, Y., Mao, Y., Li, B.: Second: Sparsely embedded convolutional detection.
Sensors 18(10), 3337 (2018)
39. Yang, B., Luo, W., Urtasun, R.: Pixor: Real-time 3d object detection from point
clouds. In: Proc. of IEEE Intl. Conf. on Computer Vision and Pattern Recognition
(2018)
40. Yang, Z., Sun, Y., Liu, S., Shen, X., Jia, J.: STD: sparse-to-dense 3d object detector
for point cloud. ICCV (2019), [Link]
41. Yu, J., Jiang, Y., Wang, Z., Cao, Z., Huang, T.: Unitbox: An advanced object
detection network. In: Proceedings of the 24th ACM international conference on
Multimedia (2016)
42. Zhao, X., Liu, Z., Hu, R., Huang, K.: 3d object detection using scale invariant and
feature reweighting networks. In: Proceedings of the AAAI Conference on Artificial
Intelligence. vol. 33, pp. 9267–9274 (2019)
43. Zhou, Y., Tuzel, O.: Voxelnet: End-to-end learning for point cloud based 3d object
detection. In: Proc. of IEEE Intl. Conf. on Computer Vision and Pattern Recog-
nition (2018)
18 T. Huang, Z. Liu, et al.
Fig. 6. Qualitative results of our approach on the KITTI validation set. For each pair,
the first and the second row show the camera image and the representative view of
LiDAR point cloud. The ground truth and detected boxes are highlighted with green
and blue boxes, respectively.
EPNet: Enhancing Point Features with Image Semantics 19
We present the qualitative results on the SUN-RGBD test set in Fig. 7. Differ-
ent from the KITTI dataset, SUN-RGBD is an indoor dataset which contains
objects of many categories and various scales. As is shown in Fig. 7, our method
and can accurately detect multiple kinds of objects with significant scale vari-
ations, including large objects (e.g., bed, sofa) and small objects (e.g., dresser,
chair). Predicting the bounding boxes of objects in a crowded area is especially
challenging. For example, Fig. 7 (a) are crowded with lots of chairs and increase
the difficulty for detection significantly. Even under this challenging case, our
method still outputs precise bounding boxes, demonstrating the robustness of
our method for crowded objects.
Fig. 7. Qualitative results of our approach on the SUN-RGBD test set. For each pair,
the first and the second row show the camera image and the representative view of
LiDAR point cloud. The ground truth and detected boxes are highlighted with green
and blue boxes, respectively.
20 T. Huang, Z. Liu, et al.
(b)
Fig. 8. Qualitative analysis of the effect of our LI-Fusion module on the SUN-RGBD
test set. The LI-Fusion module effectively exploits the semantic image information,
which is important for generating more bounding boxes .
Fig. 9. Visualization of the darkened images and lightened images generated by the
illumination transformation to simulate the underexposure and overexposure cases in
the real scenes.
enhance the point features and lead to improved detection performance as shown
in the main manuscript. It demonstrates the superiority of the LI-Fusion mod-
ule in adaptively selecting the beneficial features and suppressing the harmful
features.
The consistency enforcing loss (CE loss) in EPNet is designed to reconcile the inconsistency between localization and classification confidence. This results in a significant performance improvement of 3.93% in 3D mAP without the LI-Fusion module and 5.06% when combined with it, highlighting its effectiveness in enhancing detection reliability .
The LI-Fusion module enhances the 3D detection process by establishing a fine point-wise correspondence between LiDAR and camera image features, leading to more discriminative feature representations across multiple scales . In comparison to simple concatenation, the LI-Fusion yields a 3D mAP improvement of 1.73%, as concatenating raw camera images and LiDAR data at the input level fails to provide sufficient guidance . The performance is further validated against a single scale (SS) fusion approach, where the LI-Fusion modules applied in multiple scales outperform SS by 1.31% in 3D mAP .
The effectiveness of the LI-Fusion module in EPNet is supported by ablation studies conducted on the KITTI validation dataset, where including LI-Fusion resulted in a 1.73% increase in 3D mAP . Additional comparisons with alternative fusion methods further validated its advantage, showing better performance than both simple concatenation and single-scale fusion methods .
EPNet uses the LI-Fusion module to enhance LiDAR point features with corresponding semantic image features across multiple scales, leading to more discriminative feature representations. This multi-scale approach allows for richer feature extraction and contributes significantly to performance improvements in 3D object detection, as evidenced by its superiority in 3D mAP over single-scale fusion alternatives .
The end-to-end training of the two-stream RPN in EPNet allows for joint optimization of LiDAR and image streams, improving feature integration and ensuring consistency between proposals and final detections . This addresses challenges related to sensor data misalignment and the need for multiple annotations, facilitating higher accuracy in 3D object detection as both data types are processed cohesively for better spatial and semantic feature extraction .
MV3D and AVOD refine detection boxes by fusing BEV and camera feature maps for each region of interest, integrating data from both to enhance accuracy . In contrast, EPNet introduces a novel LI-Fusion module that directly operates on LiDAR data, establishing a finer point-wise correspondence between LiDAR and camera image features, leading to a more comprehensive fusion process without relying solely on pre-existing feature maps .
EPNet demonstrates significant qualitative advantages on the KITTI dataset by accurately detecting objects in complex scenarios, such as crowded environments with multiple vehicles and distant objects that are challenging due to sparsity and poor visibility in camera images . This effectiveness is evident from the precise bounding boxes EPNet manages to generate under these challenging conditions .
Data augmentation is crucial for LiDAR-based methods to prevent overfitting by simulating diverse training conditions. In EPNet, three common data augmentation strategies are employed: random rotations of point clouds along the vertical axis, random flipping along the forward axis, and random scaling of ground truth boxes . These augmentations improve generality without relying on hard-to-acquire prior information like road planes .
Illumination changes can introduce harmful interference in camera images, causing issues of underexposure and overexposure . The weight map in the LI-Fusion layer mitigates these challenges by adaptively estimating the importance of semantic image features, which helps to alleviate the interference and improve detection reliability under varying illumination conditions .
The image stream in EPNet learns semantic features without explicit supervision by optimizing together with the geometric stream using the supervision of 3D bounding boxes from the end-to-end RPN setup . The importance of this lies in its ability to accurately differentiate foreground objects from the background and extract rich semantic features despite the absence of 2D detection box annotations, which enhances the overall detection accuracy .