0% found this document useful (0 votes)
23 views11 pages

Enhanced Multi-Sensor Data Fusion Techniques

Uploaded by

report Reort
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
23 views11 pages

Enhanced Multi-Sensor Data Fusion Techniques

Uploaded by

report Reort
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

International Journal of Computing and Digital Systems

2025, VOL. 17, NO. 1, 1–11


[Link]

Multi-Sensor Data Fusion using FPN-ResNET

Vinodh S1 and Ramakanth P2


1,2
Department of Computer Science, R V College of Engineering, Bengaluru, India

Received 12 March 2024, Revised 15 June 2024, Accepted 3 October 2024

Abstract: Multi-sensor data fusion is ubiquitous; therefore, the associated research is significant. There are several instances in the
day-to-day activities where data fusion can be observed. The present generation autonomous driving system requires a thorough
understanding followed by a voluminous dataset for training the model. The earlier efforts have utilized the point-level fusion technique,
thereby supplementing the LiDAR point cloud with the camera features. In experimental data, imagery and proximity sensors are
paramount for the model’s performance. The sole purpose of preserving the semantic density of the imagery is compromised in
point-level technique, rendering the technique ineffective. The present work attempts to enhance the conventional point-level fusion
techniques by allocating prime importance to semantic density without increasing the computational time. This is facilitated by
performance optimization, which identifies the hindrances and enhances the transformation of the view through bird’s-eye-view pooling.
ResNET-FPN is introduced to down-sample the images without affecting the semantic density; the latency is shortened by ≈ 68%. On the
other hand, EKF is used to fuse the sensor data and evaluate the noise-covariance by compensating for the quadratic effects of the data.
The proposed model is compared with the existing models based on their performance in each background class. The IoU of the existing
models is compared with the proposed model, and it is observed to outperform the BEVFusion model by ≈ 3.1%. The detection precision
is found to be 0.9684, and the detection recall is 0.9436, while the mAP is evaluated to be 74.3%, which is ≈ 5.6% better than BEVFusion.

Keywords: Feature Pyramid Network(FPN), LiDAR, Multi-sensor data, Residual-Network(ResNET), Sensor fusion

1. Introduction However, MSDF has associated challenges due to the dif-


A wide variety of sensors are used in autonomous ference in modalities generated by the data of each sensor.
driving systems, rendering them complex. Various sensors To achieve a multi-modal and multi-task fusion, there is a
offer supplementary signals to enhance the overall data need for a unified representation of the data from differ-
collection. The inevitable usage of sensors with different ent sensors. In earlier efforts, two-dimensional perception
modalities demands the usage of the Multi-Sensor Data has been achieved by projecting spatial LiDAR data onto
Fusion(MSDF) for accurate object detection. Camera data semantic camera data. The method is less successful in
is rich in semantic information in the perspective space; detecting objects in three-dimensional space because of the
Light Detection and Ranging(LiDAR) based sensors supply geometric distortion that occurs when the LiDAR data is
spatial information in a three-dimensional space, and radar projected onto the camera(Figure.1a).
estimates the velocity at any particular instant. The present Some of the recent efforts in sensor fusion aim at enhanc-
work considers the fusion of three-dimensional LiDAR ing the LiDAR point cloud data with CNN features[2],
sensor data with two-dimensional camera data. Mapping semantic labels[3], [4] and two-dimensional image-based
the semantic information from the camera with the spatial virtual points[5]. Though there is a commendable detection
information from the LiDAR for accurate object detection performance on large-scale benchmarks, the point-level-
forms a crucial aspect for autonomous driving[1]. based fusion is less impressive on tasks of a semantic
There have been several efforts to develop reliable three- nature such as BEV-Segmentation[6], [7], [8], [9], which
dimensional object detection systems for autonomous driv- can be attributed to the semantically-lossy behavior of the
ing. Regarding information depth, laser-based sensors excel, projection of camera to LiDAR(Figure.1b). Further, the
but cameras can capture semantic data down to a deeper differences in density are more pronounced for sparser
level. Therefore, a fusion of camera and LiDAR-based LiDAR data.
sensors complement each other, permitting the development The present work proposes a fusion of multi-modal features
of a formidable three-dimensional detection system for a by maintaining geometric structure and semantic density,
safe and exceptional autonomous driving experience. which is expected to enhance 3D perception. The Fea-

E-mail address: vinodh2302@[Link], ramakanthkp@[Link]


2 Vinodh S, et al.

(a) Geometric-lossy (b) Semantic-lossy

Figure 1. Projection losses[1]

ture Pyramid Network(FPN) encoder, with Residual Net- Region-based Convolutional Neural Network(R-CNN)
work(ResNET) as the backbone, is used for image classifi- architecture with the existing one-stage-based object
cation and object detection due to the inherent characteristic detection model[21], [22], [23], [24], [25], [26]. The
of deep convolutional layers permitting faster visual data most crucial task for the offline construction of High-
processing. ResNET is more accurate in real-time applica- Definition(HD) maps is the three-dimensional segmentation
tions, which justifies its usage in the present study. Since the of the semantic features. The models[16], [27], [28], [29],
view transformation consumes nearly 80% of the model’s [30] developed to address the seminal task, analogous to
runtime, an Extended Kalman Filter(EKF) kernel is used U-Net, are note-worthy.
for pre-computation and interval reduction, ensuring process To replace the expensive LiDAR sensors for 3D object
speed-up by ≈ 65%. Furthermore, due to the non-linearity in detection, commendable efforts are made to achieve
the measured data, EKF is used to temporally estimate the three-dimensional perception based on Camera-only
state’s mean and covariance on-the-run, iteratively. Lastly, data. The FCOS3D[31] model utilizes three-dimensional
the encoder measurements are prone to uncertainties in ego- regression branches suitably coupled with the image
vehicle localization due to the accumulation of errors, which detectors[32], which is later enhanced to achieve greater
is significant when the measurements are made from larger depth in detection[33], [34]. Irrespective of perspective
distances[10]. Therefore, EKF is used to fuse the sensor data view-based object detection, models that learn from the
by statistically minimizing the error during the estimation object queries in the three-dimensional space coupled with
of the ego vehicle state vector. the Deformable Transformer(DETR)[35] detection model,
The present approach also disproves the conventional wis- viz. DETR3D[36], PETR[37] and Graph-DETR3D[38], are
dom that point-level fusion provides the best multi-sensor also developed. The view transformer-based camera-only
fusion solution. The present model is simple in construc- three-dimensional perception models explicitly transform
tion and operation, rendering it more robust and reliable. camera data to perspective bird’s eye view[6], [39],
Therefore, the work paves the way for future sensor-fusion [40], [7]. The state-of-art models such as BEVDet[41]
developments by building upon the platform reported in the and M2 BEV[42], utilize Lift-Splat- Shoot(LSS)[7] and
document. Orthographic Feature Transform(OFT)[40] for three-
dimensional object detection. Also, the three-dimensional
2. Related Work object detection models through time-dependent cues using
Over a decade, immense efforts have been put forth multiple cameras viz. BEVDet4D[43], BEVFormer[44] and
to develop a reliable and robust method for the fusion PETRv2[37] are some salient developments in single-frame
of sensors with different modalities. However, there is a methods. However, the models such as BEVFormer[44],
great scope for developing more sophisticated and accurate CVT[9], and EGO3RT[45] also perform exceptionally well
models, as the existing models are identified with few through multi-head attention for view transformation.
challenges in overcoming the projection accuracy through Lastly, efforts are put forth to study the models for
reduction in geometric and semantic losses. The earlier multi-task learning. Simultaneous detection of objects and
works to achieve three-dimensional perception based on instant segmentation form the key aspects of multi-task
LiDAR-only data include the single-stage 3D Object learning[46], [47]. Further, the simultaneous detection and
detectors[11], [8], [12], [13], [14], which provided the segmentation is extended to human-object interaction[43],
platform for the evolution of many robust and sophisticated [48], [49], [50]. The models for detection of object and
models. The model is enhanced using PointNets[15] and instance segmentation, simultaneously, viz M2 BEV[42],
SparseConvNet[16] for extracting the flattened point-cloud BEVFormer[44] and BEVerse[51], are not developed by
features. Nevertheless, the restriction offered through considering multi-sensor data fusion. Also, the activities
the bounding box in the earlier models is overcome are performed simultaneously, which demands longer
by introducing the anchorless models[17], [18], [19], computational time and higher hardware requirements,
[20]. Further investigations have led to the development thereby significantly increasing the computational cost.
of two-stage models through the amalgamation of the
International Journal of Computing and Digital Systems 3

On the other hand, the MMF model[52] though performs Network(ResNET) based Feature Pyramid Net-
detection and segmentation simultaneously, it is object- work(FPN) model.
centric, which cannot be extended to BEV Segmentation. 3) EKF is incorporated to handle the non-linearity of
The most recent attempts aim to significantly improve the camera and LiDAR data more effectively.
the detection performance by fusing sensors of different 4) EKF fuses the camera and LiDAR data from
modalities. The methods can be categorized into proposal- ResNET-FPN by minimizing the accumulation error
level and point-level. The proposal-level methods are object- generated due to the usage of ResNET-FPN.
centric and therefore do not support map segmentation 5) The framework is more generic with multi-sensor(3
effectively, whereas point-level techniques are both object- Cameras(C) and 3 LiDAR(L)) perception and mul-
centric and geometric-centric. Some of the exceptional titasking.
contributions towards proposal-level techniques include IS-
FUSION[?], SparseFusion[?], ObjectFusion[?] MV3D[53], 3. Method
F-PointNet[54], F-ConvNet[49], CenterFusion[55]. The three crucial activities that have direct implications
FUTR3D[56] and TransFusion[57], while point-level on the model performance are listed in the section 3-1,
techniques include PointPainting[2], PointAugmenting[3], section 3-2, and section 3-3.
MVP[5], FusionPainting[58], AutoAlign[59],
DeepContinuousFusion[52], Deep Fusion[4], and 1) Unified Representation
FocalSparseCNN[60]. Not all techniques can be Distinct qualities may be present in various viewpoints.
incorporated to process the camera and LiDAR data. LiDAR and radar features, for example, are usually in the
LiDAR data processing can be carried out very effectively three-dimensional bird’s-eye view, whereas camera features
through input-level decoration models viz. PointPainting[2], are in the perspective view. Every camera function, such
PointAugmenting[3], MVP[5], FusionPainting[58], as front, back, left, and right, has a unique viewing angle.
AutoAlign[59], and FocalSparseCNN[60], while Due to this perspective mismatch, feature fusion becomes
camera images require feature-level decoration viz. challenging because the same element may correspond to
DeepContinuousFusion[52], Deep Fusion[4]. entirely different spatial locations in distinct feature tensors
Contrary to the aforementioned models, the proposed (naı̈ve element-wise feature fusion will not operate in this
model has following points, which render it unique: scenario). Thus, it is imperative to identify a shared repre-
label=() sentation that is easily convertible to it without sacrificing
information and appropriate for various purposes[1].
1) It is a point-level fusion approach that performs
2) To Camera
multi-sensor fusion in a shared space by providing
weightage to both semantic and geometric informa- One option is to project the LiDAR point cloud onto the
tion equally, both in the foreground and background camera plane and display the 2.5D sparse depth driven by
2) Faster computation is ensured alongside ob- RGB-D data. This conversion is geometrically lossy. In the
ject detection, facilitated through a Residual 3D space, two neighbors on the depth map may be very far
apart. For activities like 3D object detection that rely on the

Figure 2. Framework
4 Vinodh S, et al.

geometry of the item or scene, this reduces the effectiveness pendency between the outputs. This enables the reduction
of the camera view. in latency by ≈ 44%
3) To LiDAR A. Multi-tasking
The majority of cutting-edge sensor fusion techniques Practically, most 3D perception activity is carried out
[2], [5], [4] embellish LiDAR points with the matching cam- under detection and segmentation. The object center is
era features (e.g., virtual points, CNN features, or semantic evaluated based on the size, velocity, and rotation of the
labels). But this projection from the camera to LiDAR earlier 3D detection articles[57], [17], [5]. On the other
is semantically lossy. Because of the stark differences in hand, the segmentation is carried out by classifying the
densities between LiDAR and camera features (for a 32- features and associating the binary segments with each. The
channel LiDAR scanner), < 5% of camera features match training of the segmentation head is carried out through
a LiDAR point. On semantic-oriented tasks (such as BEV CVT[9], with the focal loss being treated using Lin [Link]
map segmentation), the model’s performance is significantly model[62].
affected by giving up the semantic density of camera
features. More modern fusion techniques in the latent space,
including object query, have comparable demerits[57], [19].
4) To BEV
The lossy identified and explained through Figure.1a
and Figure.1b are considered during the transformation. The
projection of LiDAR data to BEV evens out the sparse fea-
tures in the height dimension, thereby eliminating the aspect
of geometric lossy. On the contrary, the transformation of
camera images to BEV is non-trivial due to its inherent
depth.
The depth distribution of the pixels of the camera images is
predicted using LSS[7] and BEVDet[41], [61]. The features
are re-scaled upon scattering each feature’s pixels to D
discrete points along the ray of the camera. A cloud of
the feature points is generated with a size of NHWD,
where N is the number of the cameras(in the present study,
it is 3 numbers), while, H and W are the height and the
width of the image, respectively. The grid size considered
in the cartesian coordinate system is 0.35m × 0.35m, which
is evened out in the z-direction.
The transformation of camera-to-BEV consumed a compu-
tational time of ≈ 452ms with a Quadro P6000 Graphics
processing. This can be attributed to the large number of
grid points generated per frame of the camera feature. The
Figure 3. ResNET-FPN Architecture
LiDAR features are, therefore, less dense and computation-
ally inexpensive. Nevertheless, curtailing the computational
time for the camera features demands a pre-computation 4. Experiments
and reduction in the interval considered earlier. A. Model
The Pre-computation involves associating the camera fea- ResNET50(Figure.3) is considered the backbone of the
tures to BEV grid points. From the calibration of the FPN due to its superior performance in extracting features
camera, the intrinsic and extrinsic stay the same, permitting with fewer parameters. The bottom-up feature extraction
locating coordinates of the feature cloud of the camera. Pre- is handled by the CNN layers {C1 , C2 , C3 , C4 }, with the
computation is performed by segregating the grid points dimensionality of the output of each CNN layer being 64,
based on indices, and the ranks of the points are recorded. 256, 512, and 1024, respectively. The intermediate CNN
This permits the reordering of the feature points based layers {C1 ′ , C2 ′ , C3 ′ , C4 ′ , C5 ′ }, obtained by 1×1 convolution
on the pre-computed ranks. This task alone reduces grid- and 2× downsampling, eliminate the effects of aliasing
association latency by ≈ 24%, with the remaining latency between convolutional layers and transfer 3×3 convolutional
reduction achieved through interval reduction. kernel. Lastly, the FPN layers {F1 , F2 , F3 , F4 , F5 }, which are
The Interval Reduction aggregates the grid-points generated obtained by top-down operation and 1 × 1 convolution, are
during the pre-computation through symmetric functions responsible for multi-scale information fusion, emanating
viz. mean, maximum, and summation, within the BEV from different convolutional layers. It generates the fea-
grid. The assigned Graphics Processing Unit(GPU) thread ture map by fusing the multi-scale camera images, down-
accelerates the feature aggregation and eliminates the de- sampled to 256 × 704. The settings for FPN are made as
International Journal of Computing and Digital Systems 5

(a) Detected Signal (b) Tracking Signal

Figure 4. Treatment of detection and tracking signals

per[62], with the dimensionality of each FPN layer being state based on the error generated by the MD calculation.
set to 256. Similarly, the LiDAR data, handled by ResNET-
FPN, is down-sampled to 0.075 and 0.1 for detection and D2 (x) = (x − Xi ) × Mi (x − Xi ) (2)
segmentation, respectively. Equation.2 is used to evaluate the MD, where x is the
B. Framework observations to be made, Xi is the calibration data set for
the corresponding ith sensor, while Xi is mean and Mi is
The framework for the present study is depicted in the RMSE of the ith sensor calibration data. Figure.4a and
Figure.2, where 3 cameras and 3 LiDAR sensor data are Figure.4b depict the error minimization, where corrected
input to the respective encoders. The convolutional encoders data is the output from the EKF.
follow the ResNET-FPN architecture depicted in Figure.3.
The features of the camera images are extracted and trans- C. Dataset
formed, which forms the key aspect to achieve higher The present model is trained using Waymo Dataset[67],
accuracy. LSS[7] and BEVDet[61] models are followed to which has 798 sequences for training and 202 sequences for
achieve the transformation of camera images. The trans- validation of vehicles and pedestrians. 64 lanes of LiDAR,
formed images are filtered through an Extended Kalman or 180, 000 points per 0.1 seconds, make up the point
Filter(EKF) and fused. EKF is a non-linear time-invariant clouds. The dataset for the present study comprises 20, 156
state model represented by the Equation.1[63]. annotated samples of three monocular Camera RGB images
ψ(i + 1) = ϕ(ψ(I)) + χ(i)ζ(i) = γ(ψ(i)) + ν(i) (1) capturing a 180◦ field-of-view and three 32-beam LiDAR
data. The camera images are well-nourished with semantic
where χ(i) and ν(i) are non-correlated processes with zero- information, while the LiDAR data precisely provide the
mean, while ϕ and γ are the operators. The state ψ(i + 1) is spatial information.
predicted based on the ζ(i) measurement. EKF is effective
in handling error back-propagation[64], with considerably D. Training
shorter time for training in comparison with second-order The model’s training is carried out end-to-end to avoid
gradient models such as the Gauss-Newton[10] and Least camera-encoder freezing, as observed in earlier models[2],
Mean Squares(LMS)[65] algorithms. Therefore, EKF trans- [3], [57]. The weight decay is ≈ 0.001, and optimization is
forms inherently non-linear LiDAR and Camera data, which achieved through AdamW[68] model.
generates a system matrix and evaluates noise-covariance by
compensating for the quadratic effects of the data. E. Metrics
The filtered data is mapped by estimating the error through The evaluation of the model is made based on the
the update-and-predict of the EKF input state. The root following parameters discussed in section 4-E1 and sec-
mean squared error(RMSE) estimated for a single target is tion 4-E2.
≈ 0.32 during the present study. EKF predicts the model’s
state (ψ(i + 1)), while Mahalanobis distance(MD) matches 1) Intersection over Union(IoU)
the states of multiple sensors[66], and the EKF updates the The accuracy with which the data is predicted can
be obtained through Intersection over Union(IoU), which
6 Vinodh S, et al.

(a) IoU = 0.954(Excellent) (b) IoU = 0.739(Good) (c) IoU = 0.453(Poor)

Figure 5. Intersection over Union

(a) Intersection Over Union (b) X-Error in Signal

(c) Y-Error in Signal (d) Z-Error in Signal

Figure 6. Evaluation of Object Detection Performance

is defined as the percentage of overlap between the ac- mathematically represented by Equation.3.
tual value(ground-truth) and the predicted value(Figure.5), |A ∩ B|
IoU = (3)
|A| ∪ |B|
International Journal of Computing and Digital Systems 7

(a) Accuracy (b) Loss

Figure 7. Performance of FPN-ResNET based MSDF model

where A is the ground truth and B is the predicted value. FPN model’s performance is promising for 3 Cameras and 3
Evaluation and comparison of the present model is facili- LiDAR sensor data. However, it can be further enhanced by
tated by reporting the IoU on the defined background classes considering a larger dataset with 6 Cameras and 6 LiDAR
viz. Drivable space, Pedestrian crossing, walkway, stop-line, sensors.
car parking, and lane divider. The mean IoU is calculated,
forming the basis for comparing other models. Table.I lists 2) Mean Average Precision(mAP)
the performance of the existing models, which are used The mAP is calculated based on the Average Preci-
for assessing the ResNET-FPN model(present study) based sion(AP) obtained from the area under the precision-recall
on the identified background classes. It can be observed curve. The AP is averaged for N samples as indicated by
that the present model outperforms with a mean IoU of
≈ 3.1% greater than the BEVFusion model. The ResNET-

(a) Camera view

(b) Bird’s Eye View

Figure 8. Qualitative results of Camera and LiDAR data indicating object recognition
8 Vinodh S, et al.

Models Modality Drive Crossing Walkway Stop Line Car Parking Divider Mean
OFT[40] C 74 35.3 45.9 27.5 35.9 33.9 42.1
LSS[7] C 75.4 38.8 46.3 30.3 39.1 36.5 44.4
CVT[9] C 74.3 36.8 39.9 25.8 35 29.4 40.2
BEVFusion[1] C 81.7 54.8 58.4 47.4 50.7 46.4 56.6
PointPillars[8] L 72 43.1 53.1 29.7 27.7 37.5 43.8
CenterPoint[17] L 75.6 48.4 57.5 36.5 31.7 41.9 48.6
PointPainting[2] C+L 75.9 48.5 57.1 36.9 34.5 41.9 49.1
MVP[5] C+L 76.1 48.7 57 36.9 33 42.2 49
BEVFusion[1] C+L 85.5 60.5 67.6 52 57 53.7 62.7
ResNET-FPN C+L 88.4 65.2 67.1 51.7 61.1 54.2 64.6
TABLE I. Comparison with the existing models based on IoU. Camera(C); LiDAR(L)

Equation.4 to obtain mAP. in the model’s performance for the given dataset. The
N
detection precision and recall are ≈ 0.9546 and ≈ 0.9344,
1 X respectively, and mAP is 71.2.
mAP = APi (4)
N i=1 The results are compared with the existing models, as
demonstrated in [Link]. It can be observed that the model’s
performance is better than the BEVFusion[1] by ≈ 1.4%.
5. Results and Discussion Usage of ResNET-FPN reduces the latency by ≈ 68%,
Based on the methodology defined in the earlier section, with a corresponding increase in uncertainty due to error
LiDAR signals are processed for detection and tracking. The accumulation[10]. The effect of error accumulation on the
predicted signal generated through the EKF is compared model accuracy is mitigated by using EKF for the data
with the measured signal, which is corrected based on the fusion, which reduces the standard deviation by correcting
root-mean-squared error(RMSE) represented in percentage. the predicted signal based on the measured data[63]. Fur-
The standard deviation between measured and predicted ther, unlike BEVFusion, the non-linearity of camera and
data for the detected signals is ≈ 6%, which is corrected LiDAR data is handled by the introduction of EKF. A
to achieve a standard deviation between measured and better comparison is possible between the present model
corrected signal as ≈ 5%(Figure.4a). Also, for tracking and BEVFusion[1] by considering 6 cameras and 1 LiDAR
signal, the RMSE for predicted values ≈ 7%, which is data, which is left for future work. The results are compared
corrected to achieve an error of ≈ 5.1% (Figure.4b). with the existing models, as demonstrated in [Link]. It can
Figure.6a represents the IoU for the model, which demon- be observed that the model’s performance is better than the
strates a good performance with a minimum score of BEVFusion, which is ≈ 5%. However, for the present study,
0.7, while most of the distribution is within the range of 3 cameras and 3 LiDAR data are used, unlike 6 Cameras
84 − 99%. The error plots Figure.6b, and Figure.6c show and 1 LiDAR data in the case of BEVFusion model[1].
symmetricity about zero with the maximum distribution
close to zero, whereas in the case of error plot in the Z- 6. Conclusion
direction, the range is between 0.5 to 1, with maximum It is evident from the earlier discussion that many MSDF
peaks between to 0.8 − 1. The mean position errors in X, models have demonstrated greater accuracy in the recent
Y, and Z directions are calculated to be 0.0041, 0.0363, past. However, challenges persist that can be attributed
and 0.7243, respectively. The detection performance can be to the environmental or operating conditions that induce
evaluated through accuracy and loss data plots as depicted in errors in the data as discussed in section 1. A formidable
Figure.7a and Figure.7b, respectively. The model’s accuracy correction has to be incorporated, which otherwise can
is ≈ 72% while the loss is calculated to be ≈ 55%. The affect the model’s accuracy. The model presented in this
detection precision and recall are ≈ 0.9684 and ≈ 0.9436, paper attempts to fuse the multi-modal data from different
respectively, and mAP is 74.3. sensors to enhance object detection for future autonomous
A qualitative result of the object detection is demonstrated driving purposes.
in Figure.8b, while the quantitative evaluation is performed The model demonstrates performance fairly well placed
through accuracy and loss data plots as depicted in Fig- against the existing models, particularly BEVFusion. The
ure.7a and Figure.7b, respectively. It is observed from BEVFusion model is developed by considering 6 cameras
Figure.7a that the training accuracy of the model reaches and 1 LiDAR data, while the present model considers
≈ 72% at the end of training without significant variation 3 cameras and 3 LiDAR data for fusion. Hence, there
thereafter. Similarly, the training loss curve depicted in are differences in the modalities handled in the course of
Figure.7b demonstrates a value ≈ 51% and indicates an development of the model. However, the present model
insignificant change at the end of the model’s training. is observed to have an accuracy of 72% with detection
The training accuracy and loss curve also demonstrate that precision and recall of ≈ 0.9684 and ≈ 0.9436, respectively.
further training of the model fetches a meager improvement The mean Average Precision is 74.3%, which is better than
International Journal of Computing and Digital Systems 9

Models Modality mAP


BEVDet[41] C 42.2
M2 BEV[42] C 42.9
BEVFormer[44] C 44.5
BEVDet4D[61] C 45.1
PointPillars[8] L -
SECOND[69] L 52.8
CenterPoint[17] L 60.3
PointPainting[2] C+L −
PointAugmenting[3] C+L 66.8
MVP[5] C+L 66.4
FusionPainting[58] C+L 68.1
AutoAlign[59] C+L −
FUTR3D[56] C+L −
TransFusion[57] C+L 68.9
BEVFusion[1] C+L 70.2
FPN-ResNET(present work) C+L 74.3
TABLE II. Comparison with the existing models. Camera(C); LiDAR(L)

BEVFusion by ≈ 5.6%. [6] B. Pan, J. Sun, H. Y. T. Leung, A. Andonian, and B. Zhou,


Though the present model outperforms BEVFusion and “Cross-view semantic segmentation for sensing surroundings,” IEEE
other point-level methods, there is still great scope for Robotics and Automation Letters, vol. 5, no. 3, pp. 4867–4873, 2020.
developing multi-modal 3D object detection models with
[7] J. Philion and S. Fidler, “Lift, splat, shoot: Encoding images from
inherent challenges associated with accurate depth esti- arbitrary camera rigs by implicitly unprojecting to 3d,” in Computer
mation. The model can be improved by utilizing ground Vision–ECCV 2020: 16th European Conference, Glasgow, UK,
truth to supervise the view-transformer[70], [71] that can August 23–28, 2020, Proceedings, Part XIV 16. Springer, 2020,
be considered for future developments. Also, the present pp. 194–210.
model has considered 180◦ FoV, and there is a scope
[8] A. H. Lang, S. Vora, H. Caesar, L. Zhou, J. Yang, and O. Beijbom,
for improvement by introducing 360◦ FoV with data from “Pointpillars: Fast encoders for object detection from point clouds,”
both front and rear cameras and LiDARs data. Further, the in Proceedings of the IEEE/CVF conference on computer vision and
Single Nearest Neighbour(SNN) association is considered pattern recognition, 2019, pp. 12 697–12 705.
for tracking in the present work, which can be improved
by introducing a Global Nearest Neighbour(GNN) or Joint [9] B. Zhou and P. Krähenbühl, “Cross-view transformers for real-time
Probabilistic Data Association(JPDA). map-view semantic segmentation,” in Proceedings of the IEEE/CVF
conference on computer vision and pattern recognition, 2022, pp.
13 760–13 769.
References
[1] Z. Liu, H. Tang, A. Amini, X. Yang, H. Mao, D. L. Rus, and S. Han, [10] G. Rigatos and S. Tzafestas, “Extended kalman filtering for fuzzy
“Bevfusion: Multi-task multi-sensor fusion with unified bird’s-eye modelling and multi-sensor fusion,” Mathematical and computer
view representation,” in 2023 IEEE International Conference on modelling of dynamical systems, vol. 13, no. 3, pp. 251–266, 2007.
Robotics and Automation (ICRA). IEEE, 2023, pp. 2774–2781.
[11] Y. Zhou and O. Tuzel, “Voxelnet: End-to-end learning for point
[2] S. Vora, A. H. Lang, B. Helou, and O. Beijbom, “Pointpainting: cloud based 3d object detection,” in Proceedings of the IEEE
Sequential fusion for 3d object detection,” in Proceedings of the conference on computer vision and pattern recognition, 2018, pp.
IEEE/CVF conference on computer vision and pattern recognition, 4490–4499.
2020, pp. 4604–4612.
[12] Z. Yang, Y. Sun, S. Liu, and J. Jia, “3dssd: Point-based 3d single
[3] C. Wang, C. Ma, M. Zhu, and X. Yang, “Pointaugmenting: Cross- stage object detector,” in Proceedings of the IEEE/CVF conference
modal augmentation for 3d object detection,” in Proceedings of the on computer vision and pattern recognition, 2020, pp. 11 040–
IEEE/CVF Conference on Computer Vision and Pattern Recognition, 11 048.
2021, pp. 11 794–11 803.
[13] B. Zhu, Z. Jiang, X. Zhou, Z. Li, and G. Yu, “Class-balanced
[4] Y. Li, A. W. Yu, T. Meng, B. Caine, J. Ngiam, D. Peng, J. Shen, grouping and sampling for point cloud 3d object detection,” arXiv
Y. Lu, D. Zhou, Q. V. Le et al., “Deepfusion: Lidar-camera deep preprint arXiv:1908.09492, 2019.
fusion for multi-modal 3d object detection,” in Proceedings of the
IEEE/CVF Conference on Computer Vision and Pattern Recognition, [14] Y. Zhou, P. Sun, Y. Zhang, D. Anguelov, J. Gao, T. Ouyang, J. Guo,
2022, pp. 17 182–17 191. J. Ngiam, and V. Vasudevan, “End-to-end multi-view fusion for
3d object detection in lidar point clouds,” in Conference on Robot
[5] T. Yin, X. Zhou, and P. Krähenbühl, “Multimodal virtual point Learning. PMLR, 2020, pp. 923–932.
3d detection,” Advances in Neural Information Processing Systems,
vol. 34, pp. 16 494–16 507, 2021.
10 Vinodh S, et al.

[15] C. R. Qi, L. Yi, H. Su, and L. J. Guibas, “Pointnet++: Deep volution,” in European conference on computer vision. Springer,
hierarchical feature learning on point sets in a metric space,” 2020, pp. 685–702.
Advances in neural information processing systems, vol. 30, 2017.
[29] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and
[16] B. Graham, M. Engelcke, and L. Van Der Maaten, “3d semantic B. Guo, “Swin transformer: Hierarchical vision transformer using
segmentation with submanifold sparse convolutional networks,” in shifted windows,” in Proceedings of the IEEE/CVF international
Proceedings of the IEEE conference on computer vision and pattern conference on computer vision, 2021, pp. 10 012–10 022.
recognition, 2018, pp. 9224–9232.
[30] X. Zhu, H. Zhou, T. Wang, F. Hong, Y. Ma, W. Li, H. Li, and
[17] T. Yin, X. Zhou, and P. Krahenbuhl, “Center-based 3d object D. Lin, “Cylindrical and asymmetrical 3d convolution networks for
detection and tracking,” in Proceedings of the IEEE/CVF conference lidar segmentation,” in Proceedings of the IEEE/CVF conference on
on computer vision and pattern recognition, 2021, pp. 11 784– computer vision and pattern recognition, 2021, pp. 9939–9948.
11 793.
[31] T. Wang, X. Zhu, J. Pang, and D. Lin, “Fcos3d: Fully convolutional
[18] R. Ge, Z. Ding, Y. Hu, W. Shao, L. Huang, K. Li, and Q. Liu, “1st one-stage monocular 3d object detection,” in Proceedings of the
place solutions to the real-time 3d detection and the most efficient IEEE/CVF International Conference on Computer Vision, 2021, pp.
model of the waymo open dataset challenge 2021,” in IEEE/CVF 913–922.
Conference on Computer Vision and Pattern Recognition Workshops
(CVPRW), vol. 1, 2021. [32] Z. Tian, C. Shen, H. Chen, and T. He, “Fcos: Fully convolutional
one-stage object detection,” in Proceedings of the IEEE/CVF inter-
[19] Q. Chen, L. Sun, Z. Wang, K. Jia, and A. Yuille, “Object as national conference on computer vision, 2019, pp. 9627–9636.
hotspots: An anchor-free 3d object detection approach via firing
of hotspots,” in Computer Vision–ECCV 2020: 16th European [33] T. Wang, Z. Xinge, J. Pang, and D. Lin, “Probabilistic and geometric
Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part depth: Detecting objects in perspective,” in Conference on Robot
XXI 16. Springer, 2020, pp. 68–84. Learning. PMLR, 2022, pp. 1475–1485.

[20] C. R. Qi, Y. Zhou, M. Najibi, P. Sun, K. Vo, B. Deng, and [34] H. Chen, P. Wang, F. Wang, W. Tian, L. Xiong, and H. Li, “Epro-
D. Anguelov, “Offboard 3d object detection from point cloud se- pnp: Generalized end-to-end probabilistic perspective-n-points for
quences,” in Proceedings of the IEEE/CVF Conference on Computer monocular object pose estimation,” in Proceedings of the IEEE/CVF
Vision and Pattern Recognition, 2021, pp. 6134–6144. Conference on Computer Vision and Pattern Recognition, 2022, pp.
2781–2790.
[21] S. Shi, X. Wang, and H. Li, “Pointrcnn: 3d object proposal gen-
eration and detection from point cloud,” in Proceedings of the [35] X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, “Deformable detr:
IEEE/CVF conference on computer vision and pattern recognition, Deformable transformers for end-to-end object detection,” arXiv
2019, pp. 770–779. preprint arXiv:2010.04159, 2020.

[22] C. Yilun, L. Shu, S. Xiaoyong, and J. Jiaya, “Fast point r-cnn,” in [36] Y. Wang, V. C. Guizilini, T. Zhang, Y. Wang, H. Zhao, and
ICCV, 2019. J. Solomon, “Detr3d: 3d object detection from multi-view images
via 3d-to-2d queries,” in Conference on Robot Learning. PMLR,
[23] S. Shi, Z. Wang, J. Shi, X. Wang, and H. Li, “From points to 2022, pp. 180–191.
parts: 3d object detection from point cloud with part-aware and
part-aggregation network,” IEEE transactions on pattern analysis [37] Y. Liu, T. Wang, X. Zhang, and J. Sun, “Petr: Position embedding
and machine intelligence, vol. 43, no. 8, pp. 2647–2664, 2020. transformation for multi-view 3d object detection,” in European
Conference on Computer Vision. Springer, 2022, pp. 531–548.
[24] S. Shi, C. Guo, L. Jiang, Z. Wang, J. Shi, X. Wang, and H. Li, “Pv-
rcnn: Point-voxel feature set abstraction for 3d object detection,” in [38] Z. Chen, Z. Li, S. Zhang, L. Fang, Q. Jiang, and F. Zhao, “Graph-
Proceedings of the IEEE/CVF conference on computer vision and detr3d: rethinking overlapping regions for multi-view 3d object de-
pattern recognition, 2020, pp. 10 529–10 538. tection,” in Proceedings of the 30th ACM International Conference
on Multimedia, 2022, pp. 5999–6008.
[25] S. Shi, L. Jiang, J. Deng, Z. Wang, C. Guo, J. Shi, X. Wang, and
H. Li, “Pv-rcnn++: Point-voxel feature set abstraction with local [39] T. Roddick and R. Cipolla, “Predicting semantic map representations
vector representation for 3d object detection. arxiv 2021,” arXiv from images using pyramid occupancy networks,” in Proceedings
preprint arXiv:2102.00463. of the IEEE/CVF Conference on Computer Vision and Pattern
Recognition, 2020, pp. 11 138–11 147.
[26] Z. Li, F. Wang, and N. Wang, “Lidar r-cnn: An efficient and
universal 3d object detector,” in Proceedings of the IEEE/CVF [40] T. Roddick, A. Kendall, and R. Cipolla, “Orthographic feature
Conference on Computer Vision and Pattern Recognition, 2021, pp. transform for monocular 3d object detection,” arXiv preprint
7546–7555. arXiv:1811.08188, 2018.

[27] C. Choy, J. Gwak, and S. Savarese, “4d spatio-temporal convnets: [41] J. Huang, G. Huang, Z. Zhu, Y. Ye, and D. Du, “Bevdet: High-
Minkowski convolutional neural networks,” in Proceedings of the performance multi-camera 3d object detection in bird-eye-view,”
IEEE/CVF conference on computer vision and pattern recognition, arXiv preprint arXiv:2112.11790, 2021.
2019, pp. 3075–3084.
[42] E. Xie, Z. Yu, D. Zhou, J. Philion, A. Anandkumar, S. Fidler,
[28] H. Tang, Z. Liu, S. Zhao, Y. Lin, J. Lin, H. Wang, and S. Han, P. Luo, and J. M. Alvarez, “Mˆ2 bev: Multi-camera joint 3d detection
“Searching efficient 3d architectures with sparse point-voxel con- and segmentation with unified birds-eye view representation,” arXiv
preprint arXiv:2204.05088, 2022.
International Journal of Computing and Digital Systems 11

[43] K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” [57] X. Bai, Z. Hu, X. Zhu, Q. Huang, Y. Chen, H. Fu, and C.-L. Tai,
in Proceedings of the IEEE international conference on computer “Transfusion: Robust lidar-camera fusion for 3d object detection
vision, 2017, pp. 2961–2969. with transformers,” in Proceedings of the IEEE/CVF conference on
computer vision and pattern recognition, 2022, pp. 1090–1099.
[44] Z. Li, W. Wang, H. Li, E. Xie, C. Sima, T. Lu, Y. Qiao, and
J. Dai, “Bevformer: Learning bird’s-eye-view representation from [58] S. Xu, D. Zhou, J. Fang, J. Yin, Z. Bin, and L. Zhang, “Fusion-
multi-camera images via spatiotemporal transformers,” in European painting: Multimodal fusion with adaptive attention for 3d object
conference on computer vision. Springer, 2022, pp. 1–18. detection,” in 2021 IEEE International Intelligent Transportation
Systems Conference (ITSC). IEEE, 2021, pp. 3047–3054.
[45] J. Lu, Z. Zhou, X. Zhu, H. Xu, and L. Zhang, “Learning ego 3d
representation as ray tracing,” in European Conference on Computer
[59] Z. Chen, Z. Li, S. Zhang, L. Fang, Q. Jiang, F. Zhao, B. Zhou, and
Vision. Springer, 2022, pp. 129–144.
H. Zhao, “Autoalign: pixel-instance feature aggregation for multi-
modal 3d object detection,” arXiv preprint arXiv:2201.06493, 2022.
[46] S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-
time object detection with region proposal networks,” Advances in
neural information processing systems, vol. 28, 2015. [60] Q. Chen, S. Vora, and O. Beijbom, “Polarstream: Streaming object
detection and segmentation with polar pillars,” Advances in Neural
[47] Z. Cai and N. Vasconcelos, “Cascade r-cnn: Delving into high Information Processing Systems, vol. 34, pp. 26 871–26 883, 2021.
quality object detection,” in Proceedings of the IEEE conference
on computer vision and pattern recognition, 2018, pp. 6154–6162. [61] J. Huang and G. Huang, “Bevdet4d: Exploit temporal cues in
multi-camera 3d object detection,” arXiv preprint arXiv:2203.17054,
[48] K. Sun, B. Xiao, D. Liu, and J. Wang, “Deep high-resolution repre- 2022.
sentation learning for human pose estimation,” in Proceedings of the
IEEE/CVF conference on computer vision and pattern recognition, [62] T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for
2019, pp. 5693–5703. dense object detection,” in Proceedings of the IEEE international
conference on computer vision, 2017, pp. 2980–2988.
[49] Z. Wang and K. Jia, “Frustum convnet: Sliding frustums to aggre-
gate local point-wise features for amodal 3d object detection,” in
2019 IEEE/RSJ International Conference on Intelligent Robots and [63] E. W. Kamen and J. K. Su, Introduction to optimal estimation.
Systems (IROS). IEEE, 2019, pp. 1742–1749. Springer Science & Business Media, 2012.

[50] G. Gkioxari, R. Girshick, P. Dollár, and K. He, “Detecting and [64] K. Watanabe and S. G. Tzafestas, “Learning algorithms for neural
recognizing human-object interactions,” in Proceedings of the IEEE networks with the kalman filters,” Journal of Intelligent and Robotic
conference on computer vision and pattern recognition, 2018, pp. Systems, vol. 3, pp. 305–319, 1990.
8359–8367.
[65] D. P. Bertsekas, “Nonlinear programming,” Journal of the Opera-
[51] Y. Zhang, Z. Zhu, W. Zheng, J. Huang, G. Huang, J. Zhou, tional Research Society, vol. 48, no. 3, pp. 334–334, 1997.
and J. Lu, “Beverse: Unified perception and prediction in birds-
eye-view for vision-centric autonomous driving,” arXiv preprint
arXiv:2205.09743, 2022. [66] J. L. Crowley and Y. Demazeau, “Principles and techniques for
sensor data fusion,” Signal processing, vol. 32, no. 1-2, pp. 5–27,
1993.
[52] M. Liang, B. Yang, Y. Chen, R. Hu, and R. Urtasun, “Multi-task
multi-sensor fusion for 3d object detection,” in Proceedings of the
IEEE/CVF Conference on Computer Vision and Pattern Recognition, [67] P. Sun, H. Kretzschmar, X. Dotiwalla, A. Chouard, V. Patnaik,
2019, pp. 7345–7353. P. Tsui, J. Guo, Y. Zhou, Y. Chai, B. Caine et al., “Scalability
in perception for autonomous driving: Waymo open dataset,” in
[53] X. Chen, H. Ma, J. Wan, B. Li, and T. Xia, “Multi-view 3d object Proceedings of the IEEE/CVF conference on computer vision and
detection network for autonomous driving,” in Proceedings of the pattern recognition, 2020, pp. 2446–2454.
IEEE conference on Computer Vision and Pattern Recognition,
2017, pp. 1907–1915. [68] I. Loshchilov and F. Hutter, “Decoupled weight decay regulariza-
tion,” arXiv preprint arXiv:1711.05101, 2017.
[54] C. R. Qi, W. Liu, C. Wu, H. Su, and L. J. Guibas, “Frustum pointnets
for 3d object detection from rgb-d data,” in Proceedings of the IEEE
[69] Y. Yan, Y. Mao, and B. Li, “Second: Sparsely embedded convolu-
conference on computer vision and pattern recognition, 2018, pp.
tional detection,” Sensors, vol. 18, no. 10, p. 3337, 2018.
918–927.

[55] R. Nabati and H. Qi, “Centerfusion: Center-based radar and camera [70] C. Reading, A. Harakeh, J. Chae, and S. L. Waslander, “Categorical
fusion for 3d object detection,” in Proceedings of the IEEE/CVF depth distribution network for monocular 3d object detection,” in
Winter Conference on Applications of Computer Vision, 2021, pp. Proceedings of the IEEE/CVF Conference on Computer Vision and
1527–1536. Pattern Recognition, 2021, pp. 8555–8564.

[56] X. Chen, T. Zhang, Y. Wang, Y. Wang, and H. Zhao, “Futr3d: A [71] D. Park, R. Ambrus, V. Guizilini, J. Li, and A. Gaidon, “Is pseudo-
unified sensor fusion framework for 3d detection,” in Proceedings lidar needed for monocular 3d object detection?” in Proceedings of
of the IEEE/CVF Conference on Computer Vision and Pattern the IEEE/CVF International Conference on Computer Vision, 2021,
Recognition, 2023, pp. 172–181. pp. 3142–3152.

You might also like