3D LiDAR to 2D Dense Depth Map Conversion
3D LiDAR to 2D Dense Depth Map Conversion
Abstract—The 3-D LiDAR scanner and the 2-D charge- as those manufactured by Velodyne, and an accompanying
coupled device (CCD) camera are two typical types of sensors for optical camera. This enables a vehicle to dynamically navigate
surrounding-environment perceiving in robotics or autonomous and map its surroundings as well as locate and recognize objects
driving. Commonly, they are jointly used to improve perception
accuracy by simultaneously recording the distances of surround- in the surrounding environment [1], [2]. In most cases, the in-
ing objects, as well as the color and shape information. In this formation derived from these two kinds of sensors is leveraged
paper, we use the correspondence between a 3-D LiDAR scanner separately. For example, the 3D LiDAR scanner is used for
and a CCD camera to rearrange the captured LiDAR point cloud Simultaneous Localization and Mapping (SLAM) [3]–[5] and
into a dense depth map, in which each 3-D point corresponds to the camera for objection detection and classification [6], [7].
a pixel at the same location in the RGB image. In this paper,
we assume that the LiDAR scanner and the CCD camera are However, hardly had they joined together to improve scene
accurately calibrated and synchronized beforehand so that each understanding. 3D LiDAR point cloud enjoys the superiority
3-D LiDAR point cloud is aligned with its corresponding RGB for perceiving 3D structure and object depth information, while
image. Each frame of the LiDAR point cloud is then projected CCD camera outshines in its complementary capability to
onto the RGB image plane to form a sparse depth map. Then, a capture objects color information. It is very potential to deeply
self-adaptive method is proposed to upsample the sparse depth
map into a dense depth map, in which the RGB image and the combine them together by utilizing their individual advantages.
anisotropic diffusion tensor are exploited to guide upsampling by Object localization and recognition in RGB images have
reinforcing the RGB-depth compactness. Finally, convex optimiza- made tremendous progress with deep learning based methods
tion is applied on the dense depth map for global enhancement. [8]–[13]. The availability of a large number of labeled data
Experiments on the KITTI and Middlebury data sets demonstrate sets and more powerful machines enable us to train more
that the proposed method outperforms several other relevant
state-of-the-art methods in terms of visual comparison and root- complex and powerful models that dramatically boost feature
mean-square error measurement. representations. Recently, an error rate of 4.94% was achieved
on the 1,000-class ImageNet 2012 classification data set [10],
Index Terms—Intelligent vehicle, dense depth map, 3D-2D con-
version, upsampling, global enhancement. compared to an error rate of 5.1% for human-identified objects.
This is a remarkable milestone in machine intelligence as it is
the first to surpass human-level performance (at least in the Im-
I. I NTRODUCTION ageNet data set). Instead of merely using an RGB image, more
recent work [8], [14], [15] has included additional 2D object
N OWADAYS, most intelligent vehicles are equipped with
fast speed three-dimensional (3D) LiDAR scanners, such features, such as the depth map, temporal information, object
surface normal map, and etc. These extra object characteriza-
Manuscript received November 23, 2015; revised March 29, 2016; tions have been shown to dramatically improve deep learning
accepted May 2, 2016. Date of publication June 1, 2016; date of cur- based scene analysis. For example, Gupta et al. [8], [15] pro-
rent version December 23, 2016. This work was supported in part by posed a geocentric embedding system to encode the 3D indoor
the National Natural Science Foundation of China under Grants 41401525,
61301277, and 41371431; by Guangdong Provincial Natural Science under environment as a dense depth map, height above the ground,
Grant 2014A030313209; and by the CCF-Tencent Open Fund under Grant and vertical angle map. The incorporation of these three extra
tIAGR20150114. The Associate Editor for this paper was H. Dia. (Correspond- features helped to achieve a 56% relative improvement on
ing author: Yuhang He.)
L. Chen and J. Chen are with the School of Data and Computer Science, the average precision (37.3%) of object segmentation when
Sun Yat-sen University, Guangzhou 510006, China (e-mail: chenl46@mail. compared to using the RGB image alone.
[Link]; chenjd5@[Link]). Inspired by aforementioned discussion, we focus on a novel
Y. He was with the School of Geodesy and Geomatics, Wuhan University,
Wuhan 430072, China. He is now with Dress-Plus, Beijing 100080, China task: given a pair of 3D LiDAR point cloud-RGB image, we
(e-mail: yuhanghe@[Link]). aim to transform 3D LiDAR point cloud into 2D dense depth
Q. Li is with Shenzhen Key Laboratory of Spatial Smart Sensing and map, which is in pixel-wise compactness with RGB image. To
Services, Shenzhen University, Shenzhen 518060, China (e-mail: liqq@szu.
[Link]). be specific, our proposed framework consists of two consec-
Q. Zou is with the School of Computer Science, Wuhan University, Wuhan utive steps: dynamic projecting 3D LiDAR point cloud into
430072, China (e-mail: qzou@[Link]). 2D camera plane to get sparse depth map and upsampling the
Color versions of one or more of the figures in this paper are available online
at [Link] sparse depth map into dense depth map under the supervision
Digital Object Identifier 10.1109/TITS.2016.2564640 of its corresponding RGB image.
1524-9050 © 2016 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission.
See [Link]
Authorized licensed use limited for moreUTC
to: Universidade de Caxias do Sul (UCS). Downloaded on April 11,2025 at 18:43:06 information.
from IEEE Xplore. Restrictions apply.
166 IEEE TRANSACTIONS ON INTELLIGENT TRANSPORTATION SYSTEMS, VOL. 18, NO. 1, JANUARY 2017
Fig. 1. Pipeline of the proposed framework. With the synchronized and calibrated camera and Velodyne LiDAR scanner, we project the point cloud to the RGB
image plane to get a sparse depth map, which is further used to get a pseudo depth map. The RGB image is used to generate an anisotropic diffusion tensor. The
four maps are passed to a self-adaptive upsampling framework to get the raw dense depth map. Finally, the RGB image and anisotropic diffusion tensor are again
used to give the raw dense depth map a global enhancement framework to get the final dense depth map.
Fig. 1 gives the main framework of our method. Specifically, map is computed through stereo matching [17]–[19], but it can
we assume LiDAR scanner and CCD camera are accurately be problematic if ambiguous occlusion or textureless scene
calibrated and synchronized, which allows us to project each exist. The active method indicates the way in which we directly
3D LiDAR point cloud frame to 2D RGB image plane to form project light into target object to compute depth value, such as
sparse depth map. On the one hand, currently KITTI bench mark time-of-flight (ToF) camera [20], [21], Microsoft Kinect [22],
suite used in our experiment provides us with pre-calibrated and [23] and LiDAR scanner [4], [5], [24], [25] as we discussed
pre-synchronized data set. On the other hand, actually, hard- in this paper. In ToF cameras, depth value is computed by
ware related calibration and calibration is a tough task, we find measuring the phase difference between the emitted light and
the raw sparse depth map often suffers from depth inhomogene- the reflected light, the depth map collected by it often contains
ity and depth-color inconsistency. We propose Depth Uncer- noise and is often with low resolution. Microsoft Kinect uses
tainty Elimination to address this issue, or reduce its effects. With infrared light to project dot pattern on objects, another offset
this assumption, we can easily project the LiDAR point cloud infrared camera receives the pattern and further estimate the
into the RGB image plane with a rotation matrix R and trans- depth information. Again, it suffers from large void holes due
lation vector T to form the sparse depth map. Given an image to occlusions. LiDAR scanner, like Velodyne scanner, also
pair consisting of the calibrated RGB image and sparse depth actively projects lights onto object to calculate the depth value.
map, we upsample the sparse depth map into a dense depth map The depth upsampling process is a fundamental part of the
via a self-adaptive parameter upsampling framework: the depth contributions of this paper. In recent decades, depth upsampling
of each depth-unknown pixel is computed by its neighboring under the supervision of an RGB image received much attention
seed pixel points (depth-known pixels). Each seed point’s con- [26]–[32]. In these methods, a sparse depth map is obtained by
tribution is measured in the spatial, color, and tensor domains. downsampling a dense depth map and either the depth value of
The tensor here is an anisotropic diffusion tensor that is directly each pixel is calculated within a local patch or a global energy
computed from RGB image. The initial dense depth map is function is constructed to estimate depth value of all pixels
then passed to a convex optimizer to be globally enhanced. We simultaneously. However, upsampling a sparse depth map pro-
include the RGB image throughout the entire process because jected from 3D LiDAR point cloud is particularly challenging.
it has a strong correlation with the dense depth map: color The main reasons for this are 1) a large range of depths (the
homogeneous regions correspond to depth homogeneous areas, depth range in a local patch can be as large as 100 meters),
while depth discontinuities often occur around texture edges in 2) large areas of missing depth (e.g., a vehicle window cannot
the RGB image. This strong correlation allows us to interpolate reflect LiDAR points and thus often leaves a large blank area in
the data in a more reliable way. In our experiments, we test the sparse depth map), and 3) depth uncertainty, where 3D-2D
our algorithm on the KITTI object tracking data set [16] and degradation often leads to points of different locations to be pro-
Middlebury data set (depth map only) in terms of root mean jected onto the same area in the sparse depth map. To solve
square error (RMSE). For the KITTI data set, we further eval- these problems, our framework combines local patch upsam-
uate our framework’s performance regarding different object pling to keep local depth details and global enhancement.
categories as well as their horizontal distance to the camera. Additionally, the upsampling is self-adaptive so that there is no
Actually, in a more broad perspective, the raw depth map of need to manually select parameters. The affinity between two
the scene can be obtained by either passive method or the active pixels is measured in a depth-color-tensor aware space domain.
method. In passive method, two or multi-view images of the Our previous work [33], [34] has attempted to involve
same scene are captured, then the disparity map or the depth anisotropic diffusion tensor and spatial distance, color metrics
Authorized licensed use limited to: Universidade de Caxias do Sul (UCS). Downloaded on April 11,2025 at 18:43:06 UTC from IEEE Xplore. Restrictions apply.
CHEN et al.: TRANSFORMING A 3-D LiDAR POINT CLOUD INTO A DENSE DEPTH MAP THROUGH A FRAMEWORK 167
to predict the depth value. For example, in [33], we measure are called seed points. For each depth-unknown pixel px , we
each depth-known pixels contribution in a color-spatial-tensor assume its depth value Dx can be estimated from the seed
Gaussian domain, in [34], we relax the metrics to be self- points in a w × w local patch Nx as follows:
adaptive and propose the parameters to be self-selective in
depth-color-spatial domain. These work both received large Dx = axy Dy (1)
improvements on both indoor and outdoor depth upsampling y∈Nx
tasks. In this paper, we integrate our previous work and further where Dy is the seed point in Nx and axy is the affinity
propose a novel framework encompassing both the parameter measurement between Dx and Dy . To construct axy , we adopt
self-adaptive scheme and as many metrics (color, spatial, tensor, a Gaussian kernel that takes depth, color, tensor, and spatial
depth, etc.) as possible. Thus, the main contribution of this distance information into account
paper lies in:
• A novel framework is proposed to transform 3D LiDAR Dp − Dy 22 Ix − Iy 22
point cloud into 2D dense depth map under the guidance axy = exp 2 · exp 2
2σD 2σC
of RGB image, in the context of scene understanding.
The derived dense depth maps are highly compatible with −x − y22 −Tx − Ty 22
RGB images. We specially designed a novel framework · exp · exp
2σS2 2σT2
to transform 3D LiDAR point cloud into 2D dense depth
map for outdoor complex driving scenario. Comparing where σD , σC , σS , and σT are depth, color, spatial distance and
with existing relative methods, we outperform or achieve the anisotropic diffusion tensor kernel bandwidth, respectively.
comparable results in both indoor and outdoor scenarios. Tx and Ty are the anisotropic diffusion tensor value, whose
• Accurate depth can be obtained by our method. We definition will be covered in the next subsection. We include the
propose to combine local details, including color, spatial anisotropic diffusion tensor here because, in addition to main-
distance, depth variation and anisotropic diffusion tensor, taining the depth smoothness in a color homogeneous region,
to calculate two pixels affinity, and achieve global en- we further hope to guarantee depth gradient and surface nor-
hancement via a convex optimizer building on correlation mal similarity in that region. The anisotropic diffusion tensor
of raw depth map and RGB image. weights the first order of the depth gradient and its direction,
• The proposed self-adaptive scheme can greatly improve thus, it ideally fits this requirement. We give the detailed defini-
the discriminative power of our clustering algorithm. tion of anisotropic diffusion tensor in the next subsection.
All metrics are represented in Gaussian metric manner,
with which the kernel bandwidths are self-adaptive: it
B. Anisotropic Diffusion Tensor
is proportional to a logarithmic distance. Unlike existing
methods that require multiple parameters tuning to get a Anisotropic Diffusion Tensor has already been applied for
fine result, our method is nearly parameter free. affinity measurement in [27] and [33]. It is defined as
The rest of this paper is organized as follows: In Section II, T = exp (−β|∇IH |γ ) nnT + n⊥ n⊥T (2)
we describe the self-adaptive sparse depth upsampling frame-
work. Global enhancement is explained in Section III. Experi- ∇IH is the image gradient, n is the normalized direction of im-
mental results are given in Section IV. We conclude this work age gradient n = ∇IH /|∇IH |, n⊥ is the normal vector of im-
in Section V. age gradient. β and γ adjust the magnitude and sharpness of T .
Anisotropic diffusion tensor is directly calculated from RGB
II. S PARSE D EPTH M AP U PSAMPLING image but it has strong indication to the final depth map that is
derived from RGB image, since most texture edges anisotropic
Given the sparse depth map derived above, we assume the diffusion tensor map most likely corresponds to depth discon-
depth value of each depth-unknown pixel derives from its tinuities. We involve anisotropic diffusion tensor to stress the
spatially neighboring seed points (we call depth-known pixels depth value variation area by trying best to involve the seed
as seed points). Each seed point’s affinity is measured in points with the largest confidence to estimate a query pixel’s
depth, color, spatial and anisotropic diffusion tensor domain depth value. The advantage of involving anisotropic diffusion
in Gaussian metrics. Each Gaussian metric kernel bandwidth tensor will be shown in experiment section.
is dynamically selected: it is inverse to the depth discrepancy
between the two pixels that are taken into consideration. To get
C. Pseudo Depth Estimation
the raw depth value estimation for each depth-unknown pixel,
we assign it with the depth value of the seed point which has Metric axy measures not only the spatial, color, and tensor
shortest geodesic distance that defined below to the query pixel. distance of Dx and Dy , it also measures their mutual depth
disparity: we assume a seed points with larger contribution to
px should share a closer depth value to px than those seed points
A. Affinity Measurement
with less contribution. Depth Dp is the pseudo depth value of
We denote by I the RGB image that accompanies sparse px that we “steal” from the seed point that has the shortest
depth map Ds . In this paper, all the depth-known pixels in Ds geodesic distance to px . Geodesic distance dG (x, y) is defined
Authorized licensed use limited to: Universidade de Caxias do Sul (UCS). Downloaded on April 11,2025 at 18:43:06 UTC from IEEE Xplore. Restrictions apply.
168 IEEE TRANSACTIONS ON INTELLIGENT TRANSPORTATION SYSTEMS, VOL. 18, NO. 1, JANUARY 2017
Fig. 2. Geodesic distance estimation with (dark green color) forward propaga-
tion and (yellow color) backward propagation.
Fig. 5. Visual comparison of depth upsampling result under various conditions. (Left to right) RGB image, dense depth map without depth uncertainty elimination,
with empirical kernel bandwidth, with self-adaptive bandwidth, and with the global enhancement after self-adaptive bandwidth. (a) RGB image. (b) No elimination.
(c) Constant bandwidth. (d) Self-adaptive bandwidth. (e) Global enhancement.
TABLE I
Q UANTITATIVE R ESULTS ON KITTI T RACKING D ATA S ET. O CC .0, O CC .1, AND O CC .2 I NDICATE F ULLY V ISIBLE ,
PARTIALLY O CCLUDED , AND L ARGELY O CCLUDED , R ESPECTIVELY
with the seed point is updated by its 4-connected neighboring seed points are maximally involved and dissimilar seed points
pixels (see Fig. 2, pixels with green color) are minimally involved in the prediction of a query point depth
value. Fig. 4 illustrates how a bandwidth that is too large may
dSGk (x) = +min+ dSGk (x+ ) + dG (x, x+ ) (5) involve unnecessary seed points while a bandwidth that is too
x ∈X
small can easily drop out useful seed points. Our self-adaptive
where X + is x’s 4 neighboring pixels. Similarly, the backward scheme includes similar seed points but drops out dissimilar
propagation is conducted in a reverse order (namely, bottom- seed points automatically.
right to top-left) with the left 4-connected neighboring pixels When computing the color distance Ix − Iy 22 , rather than
(see Fig. 2, pixels with yellow color). Accurate computation merely considering pixel color information alone, we extract
of geodesic distance requires multiple iterations of forward two small u × u patches Px and Py centered at px and py ,
propagation and backward propagation. The iteration number respectively. We then derive a structure-aware filter Bx on Px
is directly affected by the sparsity of seed points in one subset to convolve Py
image, because one seed point can propagate its impact to the I − I 2
P x
whole image within one iteration and one pixel’s association Bx = −exp i∈C x 2
(7)
2σI2
with seed point is most likely to be changed when another seed
point intervenes. In most cases, the optimization can be reached where C is color space, σI is the average value of the maximum
within 10 iterations and we can artificially set the iteration and minimum values of IPx − Ix 22 . Further Bx contains the
number to a smaller number to speed up upsampling without Gaussian color distance distribution of Px , we use it to convolve
much accuracy declination. Py to measure the color distance of two patches
Ix − Iy 22 = Bx ◦ Py − Px 22 (8)
D. Bandwidth Selection i∈C
In contrast to other methods that use empirically defined ker- where ◦ indicates element-wise multiplication. This patch-
nel bandwidths σD and σC , we adopt a self-adaptive parameter- based color distance strategy uses local shape information to
setting scheme in which kernel bandwidth between px and py is measure the color similarity of px , py from a high-level shape
self-defined. That is, σD and σC are bandwidth sets, and each information perspective. (see Fig. 3 for details)
point pair < px , py > has its own bandwidth
E. Seed Point Embedding
σDxy = log (Dp − Dy ) σCxy = log (Ix − Iy ) . (6)
To address the large depth-blank areas that come from a large
This novel bandwidth setting scheme frees us from empirically baseline between camera and laser scanner and the incomplete
choosing the bandwidth, which often results in either over- point clouds reflected from the glass material (see Fig. 7), we
smoothing boundaries (too large a bandwidth) or depth value propose a seed point embedding strategy that includes newly
absence (too small a bandwidth). Moreover, the self-adaptive calculated non-seed points as seed points. After estimating a
scheme automatically differentiates seed points’ contribution pixel’s depth value, we add it to the seed point set so that more
to query point by setting relative low weights to seed points seed points are used when we calculate the next pixel’s depth
with low affinity but large weights to seed points with large value. Note that the inclusion of pseudo seed points is able to
affinity (in color-depth domain). This guarantees that similar fill in the entire depth-blank area because the path of the w × w
Authorized licensed use limited to: Universidade de Caxias do Sul (UCS). Downloaded on April 11,2025 at 18:43:06 UTC from IEEE Xplore. Restrictions apply.
170 IEEE TRANSACTIONS ON INTELLIGENT TRANSPORTATION SYSTEMS, VOL. 18, NO. 1, JANUARY 2017
Fig. 6. Visual comparison on the KITTI data set. The relevant method used to generate each depth map is shown on the image’s top left region.
patch generates pseudo seed points over the whole image. We by four small seed point pi from four directions at the same
give these pseudo seed points a low confidence when measuring direction, we eliminate it from the sparse depth map
their contribution to the query point (e.g., 0.1 times its original
weight). Thus, the new upsampling formula becomes
4
p − pi > θ (i = 1, 2, 3, 4) {p − pi } = 4 (10)
Dx = {f (y)}axy Dy (9) i=1
Fig. 7. Visual comparison of depth upsampling on the KITTI data set. In each subfigure from top to bottom: RGB image superimposed with point cloud image
(point cloud is bolded to be seen more clearly), JOINT [29], TGVL [27], FILTER [26], GEO [28], PREV [34], and OurSD. Note that the other three methods’
results produce blurry boundaries. FILTER [27] even leaves void holes, particularly on various vehicles’ windows. Our proposed algorithm has successfully
avoided these affects, producing edge-sharp dense depth map (please zoom in for better visualization).
The data term is used to guarantee the consistency between with a Velodyne HDL-64E scanner, four cameras, global posi-
the initial dense depth map Ds and final globally enhanced tioning system, an IMU, and other basic sensors. It consists of
dense depth map DF 20 street scene sequences, and each view includes a color image
pair and a 360◦ Velodyne LiDAR point cloud frame. Addition-
H(DS , u) = w|u − DS |2 dx . (12) ally, object labels such as “car” “van” “pedestrian” and “cyclist”
ΩF are available for both the 2D images and LiDAR point cloud.
These labels in particular enable us to evaluate our algorithm
The regularization term used here is the second order total with respect to its intended usage. The Middlebury data set
generalized variation term contains six indoor scenes: “moebius” “books” “dolls” “laun-
⎧ ⎫
⎨ ⎬ dry” “art” and “reindeer.” Each scene includes two color images
R(u) = min α1 |T u −v|dx + α0 | v |dx (13) and a disparity image that serves as the depth ground truth.
v ⎩ ⎭ We compared our algorithm with a MRF optimization based
Ω Ω
method (JOINT) [29], local filter based method (FILTER) [26],
where T is the anisotropic diffusion tensor calculated by Eq. (2). global optimization based method (TGVL) [27], and geodesic
The optimization of energy function (11) can be achieved using distance based method (GEO) [28] and our previous work
a Primal-Dual optimization approach [27] for details). Fig. 5(e) (PREV) [34]. We either use publicly released code or imple-
shows the upsampling result with global enhancement. mented the methods by ourselves. The evaluation metric we
adopt here is RMSE (root mean square error). Given a set of pre-
IV. E XPERIMENT AND E VALUATION dicted values and their corresponding ground truth y = {ypi , ygi }
(i = 1, . . . , N ), RMSE is defined as
We evaluated our algorithm on KITTI tracking visual bench-
mark suite [16] for outdoor environment and the Middlebury N y i 2 − y i 2
i=1 p g
data set [35] for indoor environment. The KITTI tracking data RMSE = . (14)
set was collected by an autonomous driving platform equipped N
Authorized licensed use limited to: Universidade de Caxias do Sul (UCS). Downloaded on April 11,2025 at 18:43:06 UTC from IEEE Xplore. Restrictions apply.
172 IEEE TRANSACTIONS ON INTELLIGENT TRANSPORTATION SYSTEMS, VOL. 18, NO. 1, JANUARY 2017
Fig. 8. Depth and height RMSE variation for object categories, which contain (a) “pedestrian,” (b) “car,” (c) “van,” and (d) “cyclist,” with respect to the distance
to the RGB image plane.
TABLE II
Q UANTITATIVE RMSE R ESULTS ON THE M IDDLEBURY D ATA S ET AT D OWNSAMPLING FACTORS OF 2×, 4×, AND 8×
Fig. 9. Visual comparison on the Middlebury data set. (Left to right) RGB image, FILTER [26], GEO [28], JOINT [29], TGVL [27], PREV [34], OurSEGE, and
ground truth.
Fig. 10. Visual comparison of depth upsampling result on “reindeer” (8×), “dolls” (4×), “art” (2×), “moebius” (8×), “laundry” (4×), and “books” (2×),
respectively. Whereas the other five algorithms produce coarse and blurred results, our algorithm’s results are maximally close to the ground truth.
Fig. 11. Visual comparison of the four versions of our proposed algorithm on the Middlebury data set. To best visualize the difference, we use the downsampling
factor 8×. On one hand, we can see that all the four versions produce similar and fine upsampling results. On the other hand, the involvement of seed point
embedding and global enhancement upgrade the results, particularly around object boundaries.
and upsampling simultaneously, which led to its superiority upsampling. On the contrary, our proposed framework keeps
to our method. Under a large downsampling factor that fewer an appropriate balance between these factors and obtains good
seed points are involved, JOINT [29] failed to infer depth results for all downsampling factors. The qualitative results
details from less seed points, however, our proposed method of our methods and five compared methods with some close-
can remedy this by “borrowing” enough seed points (forward ups are listed in Fig. 9. We can clearly observe that OurSEGE
and backward propagation process) from other area. Similarly, achieves state of the art result regarding depth hierarchy, depth-
the global enhancement scheme TGVL [27] we adopt in this color consistency when comparing with the other five relevant
paper achieves good results for large downsampling factors, methods. In addition, the void hole in the ground truth image
but it fails to retain depth details for small downsampling can also be filled up by our proposed framework.
factors because its unilateral emphasis on global consistency Besides, our proposed framework outperforms our previous
inevitably depletes the detail. FILTER [26] generates a much work on both downsampling factors, which shows that the in-
larger RMSE than all other algorithms because it only includes volvement of anisotropic diffusion tensor on both self-adaptive
depth and spatial information to guide the upsampling process, upsampling and global enhancement dramatically improves the
which in turn attests the importance of the RGB image during final upsampling result.
Authorized licensed use limited to: Universidade de Caxias do Sul (UCS). Downloaded on April 11,2025 at 18:43:06 UTC from IEEE Xplore. Restrictions apply.
CHEN et al.: TRANSFORMING A 3-D LiDAR POINT CLOUD INTO A DENSE DEPTH MAP THROUGH A FRAMEWORK 175
TABLE III
RMSE R ESULTS ON THE M IDDLEBURY AND KITTI D ATA S ETS W ITH R ESPECT TO D IFFERENT A FFINITY M EASUREMENTS . T HE D OWNSAMPLING
FACTOR FOR THE M IDDLEBURY D ATA S ET IS 2×, AND THE O CCLUSION L EVEL FOR THE KITTI D ATA S ET IS O CC .0
Visual comparison of depth upsampling results are shown in many other useful features such as height and local plane sur-
Fig. 10. It is clear that at the downsampling factor 8×, FILTER face values are also important object features and can be applied
[26] and GEO [28] fail to keep depth boundary consistency to image-related applications. We hope our work will motivate
and lose many structural details. JOINT [29] damages depth further research on large-scale or holistic image based scene
hierarchies and generates blurred boundaries. TGVL [27] understanding in robotics or autonomous driving, especially
includes unnecessary texture details in the final depth map with respect to the combination of 3D LiDAR point clouds and
and deliberately produces depth discontinuity around small RGB images.
textures that lie on depth homogeneous regions. In contrast,
our proposed framework successfully avoids these issues and
R EFERENCES
keeps depth boundary sharpness and hierarchies. Further, it
retains the correlations between the RGB image and final depth [1] S. Gidel, P. Checchin, C. Blanc, T. Chateau, and L. Trassoudaine, “Pedes-
trian detection and tracking in an urban environment using a multi-
map maximally. OurSDGE, OurSE, and OurSEGE produce layer laser scanner,” IEEE Trans. Intell. Transp. Syst., vol. 11, no. 3,
similar results to OurSD, so we do not show them here due pp. 579–588, Sep. 2010.
to the space limitations. The direct comparison of those four [2] H. Guan, J. Li, Y. Yu, Z. Ji, and C. Wang, “Using mobile LiDAR data
for rapidly updating road markings,” IEEE Trans. Intell. Transp. Syst.,
versions of our algorithm is show in Fig. 11. vol. 16, no. 5, pp. 2457–2466, Oct. 2015.
[3] P. Corcoran, A. Winstanley, P. Mooney, and R. Middleton, “Background
foreground segmentation for SLAM,” IEEE Trans. Intell. Transp. Syst.,
C. Discussion vol. 12, no. 4, pp. 1177–1183, Dec. 2011.
[4] F. Moosmann and C. Stiller, “Velodyne SLAM,” in Proc. IEEE Intell. Veh.
To compute two pixels’ affinity, we take the spatial, color, Symp., Jun. 2011, pp. 393–398.
[5] J. Choi, “Hybrid map-based SLAM using a Velodyne laser scanner,” in
and tensor distances into consideration. To determine whether Proc. IEEE Int. Conf. Intell. Transp. Syst., Oct. 2014, pp. 3082–3087.
these additional distances improve the final performance, we [6] H. Cho, P. E. Rybski, and W. Zhang, “Vision-based 3D bicycle tracking
conducted another experiment by changing the definition of using deformable part model and interacting multiple model filter,” in
Proc. IEEE Int. Conf. Robot. Autom., May 2011, pp. 4391–4398.
ax,y to include only spatial distance (σS ), only spatial and [7] H. Durrant-Whyte, N. Roy, and P. Abbeel, “Tracking-based semi-
color distance (σS + σC ), and spatial, color and tensor together supervised learning,” in Proc. Robot.: Sci. and Syst., 2012, pp. 329–336.
(σS + σC + σT ). The experimental results on the Middlebury [8] S. Gupta, R. Girshick, P. Arbelaez, and J. Malik, “Learning rich features
from RGB-D images for object detection and segmentation,” in Proc. Eur.
and KITTI data set are listed in Table III. It is clear that Conf. Comput. Vis., 2014, pp. 1–16.
the involvement of color and tensor information dramatically [9] N. Silberman, D. Hoiem, P. Kohli, and R. Fergus, “Indoor segmentation
improves performance, especially on the KITTI data set. The and support inference from RGBD images,” in European Conference
on Computer Vision. Berlin, Germany: Springer, 2012, pp. 746–760.
depth variation on the KITTI data set is much larger than the [10] K. He, X. Zhang, S. Ren, and J. Sun, “Delving deep into rectifiers:
depth variation on the Middlebury data set. The anisotropic Surpassing human-level performance on ImageNet classification,” in
diffusion tensor helps to differentiate these depth value in a Proc. IEEE Int. Conf. Comput. Vis., Dec. 2015, pp. 1026–1034.
[11] D. Ciregan, U. Meier, and J. Schmidhuber, “Multi-column deep neural
more reasonable and robust way. networks for image classification,” in Proc. IEEE Conf. Comput. Vis.
Overall, on both the Middlebury and KITTI benchmark Pattern Recognit., Jun. 2012, pp. 3642–3649.
suites, our proposed LiDAR point cloud organization frame- [12] Y. He, S. Chen, Y. Pan, and K. Ni, “Using edit distance and junction
feature to detect and recognize arrow road marking,” in Proc. IEEE Int.
work can successfully recover object hierarchies, boundary Conf. Intell. Transp. Syst., Oct. 2014, pp. 2317–2323.
sharpness, and global integrity, regardless of the point cloud [13] Y. Sun, Y. Chen, X. Wang, and X. Tang, “Deep learning face rep-
sparsity, large losses, and 3D-2D degradation uncertainty. Even resentation by joint identification-verification,” in Advance in Neural
Information Processing Systems, Z. Ghahramani, M. Welling, C. Cortes,
though the method was initially designed to transform a LiDAR N. Lawrence and K. Weinberger, Eds. New York, NY, USA: Curran
point cloud into a depth map, it still works well on non- Associates, 2014, pp. 1988–1996.
LiDAR point cloud applications (e.g., the Middlebury data set [14] S. Ji, W. Xu, M. Yang, and K. Yu, “3D convolutional neural networks
for human action recognition,” IEEE Trans. Pattern Anal. Mach. Intell.,
experiments presented in this paper). vol. 35, no. 1, pp. 221–231, Jan. 2013.
[15] S. Gupta, P. Arbeláez, and J. Malik, “Perceptual organization and recogni-
tion of indoor scenes from RGB-D images,” in Proc. IEEE Conf. Comput.
V. C ONCLUSION Vis. Pattern Recognit., Jun. 2013, pp. 564–571.
[16] A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for autonomous
In this paper, we proposed a novel framework to transform a driving? The KITTI vision benchmark suite,” in Proc. IEEE Conf.
Comput. Vis. Pattern Recognit., Jun. 2012, pp. 3354–3361.
3D LiDAR point cloud into a 2D dense depth map using its [17] R. Szeliski et al., “A comparative study of energy minimization methods
corresponding RGB image as a guide. Transforming the 3D for Markov random fields with smoothness-based priors,” IEEE Trans.
LiDAR points into RGB image compatible features has many Pattern Anal. Mach. Intell., vol. 30, no. 6, pp. 1068–1080, Jun. 2008.
[18] J. Sun, N.-N. Zheng, and H.-Y. Shum, “Stereo matching using belief
practical applications, especially in image based scene analysis propagation,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 25, no. 7,
and environment perception. In fact, not only depth maps, but pp. 787–800, Jul. 2003.
Authorized licensed use limited to: Universidade de Caxias do Sul (UCS). Downloaded on April 11,2025 at 18:43:06 UTC from IEEE Xplore. Restrictions apply.
176 IEEE TRANSACTIONS ON INTELLIGENT TRANSPORTATION SYSTEMS, VOL. 18, NO. 1, JANUARY 2017
[19] Q. Yang, “Stereo matching using tree filtering,” IEEE Trans. Pattern Anal. Yuhang He received the [Link]. degree in pho-
Mach. Intell., vol. 37, no. 4, pp. 834–846, Apr. 2015. togrammetry and remote sensing from Wuhan Uni-
[20] A. Kolb, E. Barth, R. Koch, and R. Larsen, “Time-of-flight cameras in versity, Wuhan, China, in 2013.
computer graphics,” Comput. Graph. Forum, vol. 29, no. 1, pp. 141–159, He was a Researcher in autonomous driving and
Feb. 2010. deep learning with the Institute of Deep Learning,
[21] C. Schaller, “Time-of-Flight—A New Modality for Radiotherapy,” Baidu Inc., Beijing, China. He is currently a Deep
Ph.D. dissertation, Der Technischen Fakultät der, Friedrich-Alexander- Learning Researcher with Dress-Plus, Beijing. His
Universität Erlangen-Nürnberg, Erlangen, Germany, 2011. research interests include machine vision, deep
[22] D. Miao, J. Fu, Y. Lu, S. Li, and C. W. Chen, “Texture-assisted Kinect learning, and self-driving.
depth inpainting,” in Proc. IEEE Int. Symp. Circuits Syst., May 2012,
pp. 604–607.
[23] J. Hu, R. Hu, Z. Wang, Y. Gong, and M. Duan, “Kinect depth map based
enhancement for low light surveillance image,” in Proc. IEEE Int. Conf.
Image Process., Sep. 2013, pp. 1090–1094.
[24] T. Chen, B. Dai, D. Liu, J. Song, and Z. Liu, “Velodyne-based curb
detection up to 50 meters away,” in Proc. IEEE Intell. Veh. Symp.,
Jun. 2015, pp. 241–248. Jianda Chen has been working toward the bache-
[25] T. Chen, B. Dai, D. Liu, and J. Song, “Performance of global descriptors lor’s degree with the School of Data and Computer
Science, Sun Yat-sen University, Guangzhou, China,
for Velodyne-based urban object recognition,” in Proc. IEEE Intell. Veh.
Symp., Jun. 2014, pp. 667–673. since 2012.
[26] C. Premebida, J. Carreira, J. Batista, and U. Nunes, “Pedestrian detection His research interests include computer vision
and machine learning.
combining RGB and dense LIDAR data,” in Proc. IEEE/RSJ Int. Conf.
Intell. Robot. Syst., Sep. 2014, pp. 4112–4117.
[27] D. Ferstl, C. Reinbacher, R. Ranftl, M. Ruether, and H. Bischof, “Image
guided depth upsampling using anisotropic total generalized variation,” in
Proc. IEEE Int. Conf. Comput. Vis., Dec. 2013, pp. 993–1000.
[28] M. Y. Liu, O. Tuzel, and Y. Taguchi, “Joint geodesic upsampling of depth
images,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., Jun. 2013,
pp. 169–176.
[29] W. Huang, X. Gong, and M. Y. Yang, “Joint object segmentation and depth
upsampling,” IEEE Signal Process. Lett., vol. 22, no. 2, pp. 192–196,
Feb. 2015. Qingquan Li received the M.S. degree in engi-
[30] J. Yang, X. Ye, K. Li, C. Hou, and Y. Wang, “Color-guided depth recovery neering and the Ph.D. degree in photogrammetry
from RGB-D data using an adaptive autoregressive model,” IEEE Trans.
and remote sensing from Wuhan University, Wuhan,
Image Process., vol. 23, no. 8, pp. 3443–3458, Aug. 2014. China, in 1988 and 1998, respectively.
[31] J. Yang, X. Ye, K. Li, and C. Hou, “Depth recovery using an adap- From 1988 to 1996, he was an Assistant Professor
tive color-guided auto-regressive model,” in European Conference on with Wuhan University, where he was an Associate
Computer Vision. Berlin, Germany: Springer-Verlag, 2012, pp. 158–171. Professor from 1996 to 1998 and a Professor with
[32] L. Dai, H. Wang, X. Mei, and X. Zhang, “Depth map upsampling via State Key Laboratory of Information Engineering in
compressive sensing,” in Proc. Asian Conf. Pattern Recognit., Nov. 2013, Surveying, Mapping and Remote Sensing in 1998.
pp. 90–94.
He is currently the President of Shenzhen University,
[33] Y. He, L. Chen, and M. Li, “Sparse depth map upsampling with RGB Shenzhen, China, where he is also the Director of
image and anisotropic diffusion tensor,” in Proc. IEEE Intell. Veh. Symp., Shenzhen Key Laboratory of Spatial Smart Sensing and Services. His research
Jun. 2015, pp. 205–210. interests include photogrammetry, remote sensing, and intelligent transporta-
[34] Y. He, L. Chen, J. Chen, and M. Li, “A novel way to organize 3D LiDAR tion systems.
point cloud as 2D depth map height map and surface normal map,” in Prof. Li is an expert in Modern Traffic with the National 863 Plan and an
Proc. IEEE Int. Conf. Robot. Biomimetics, Dec. 2015, pp. 1383–1388. Editorial Board Member of the Surveying and Mapping Journal and Wuhan
[35] H. Hirschmuller and D. Scharstein, “Evaluation of cost functions for
University Journal-Information Science Edition.
stereo matching,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit.,
Jun. 2007, pp. 1–8.
Authorized licensed use limited to: Universidade de Caxias do Sul (UCS). Downloaded on April 11,2025 at 18:43:06 UTC from IEEE Xplore. Restrictions apply.