0% found this document useful (0 votes)
5 views10 pages

PrevPredMap: Temporal Modeling for HD Maps

The document presents PrevPredMap, a novel temporal modeling framework for online vectorized HD map construction that leverages previous predictions to enhance the accuracy and efficiency of map generation. It introduces two key modules: a previous-predictions-based query generator and a dynamic-position-query decoder, which work together to effectively encode and utilize past predictions for current map predictions. Extensive experiments demonstrate that PrevPredMap achieves state-of-the-art performance on benchmark datasets, validating its innovative approach to temporal modeling in HD map construction.

Uploaded by

wnqnxh
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views10 pages

PrevPredMap: Temporal Modeling for HD Maps

The document presents PrevPredMap, a novel temporal modeling framework for online vectorized HD map construction that leverages previous predictions to enhance the accuracy and efficiency of map generation. It introduces two key modules: a previous-predictions-based query generator and a dynamic-position-query decoder, which work together to effectively encode and utilize past predictions for current map predictions. Extensive experiments demonstrate that PrevPredMap achieves state-of-the-art performance on benchmark datasets, validating its innovative approach to temporal modeling in HD map construction.

Uploaded by

wnqnxh
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

PrevPredMap: Exploring Temporal Modeling with Previous Predictions for

Online Vectorized HD Map Construction

Nan Peng Xun Zhou Mingming Wang


Ruqi Mobility Ruqi Mobility GAC R&D Center
pengnan@[Link] zhouxun@[Link] wangmingming@[Link]

Xiaojun Yang Songming Chen Guisong Chen


arXiv:2407.17378v1 [[Link]] 24 Jul 2024

Ruqi Mobility GAC R&D Center Ruqi Mobility


yangxiaojun@[Link] chensongming@[Link] chenguisong@[Link]

Abstract a labor-intensive and time-consuming endeavor. This is


largely due to the necessity of achieving centimeter-level
Temporal information is crucial for detecting occluded precision for various map elements and the ongoing require-
instances. Existing temporal representations have pro- ment to ensure their currency. The substantial financial in-
gressed from BEV or PV features to more compact query vestment required for this process poses a significant barrier
features. Compared to these aforementioned features, pre- to the broader adoption and regular refresh of HD maps.
dictions offer the highest level of abstraction, providing Online HD map construction is emerging as a promis-
explicit information. In the context of online vectorized ing alternative, offering a cost-effective solution to tradi-
HD map construction, this unique characteristic of pre- tional SLAM-based methods. Utilizing onboard sensors,
dictions is potentially advantageous for long-term tempo- map elements are perceived in real-time through map learn-
ral modeling and the integration of map priors. This pa- ing. Early works [4,13,17,28] considered map learning as a
per introduces PrevPredMap, a pioneering temporal mod- semantic segmentation task, which required complex post-
eling framework that leverages previous predictions for processing and often struggled with efficiency. To over-
constructing online vectorized HD maps. We have metic- come these limitations, recent researches [18, 25] have pro-
ulously crafted two essential modules for PrevPredMap: posed directly predicting vectorized map elements, demon-
the previous-predictions-based query generator and the strating a significant improvement in performance.
dynamic-position-query decoder. Specifically, the previous- Temporal modeling is crucial for detecting occluded map
predictions-based query generator is designed to separately elements, thereby further enhancing the quality of con-
encode different types of information from previous predic- structed HD maps. In the realm of Bird’s-Eye-View (BEV)
tions, which are then effectively utilized by the dynamic- perception, existing temporal representations include BEV
position-query decoder to generate current predictions. features [11, 16, 17, 29], perspective features [20, 24], and
Furthermore, we have developed a dual-mode strategy query features [21, 22, 33]. These representations are all in-
to ensure PrevPredMap’s robust performance across both termediate outcomes, as shown in Fig. 1 (a)-(c). Despite the
single-frame and temporal modes. Extensive experiments significantly smaller information capacity of query features,
demonstrate that PrevPredMap achieves state-of-the-art the effectiveness and efficiency of temporal modeling that
performance on the nuScenes and Argoverse2 datasets. utilizes previous query features have been well-established
Code will be available at https:// [Link]/ [21, 22, 33]. This progress prompts us to delve deeper: Is
pnnnnnnn/PrevPredMap. it still effective to rely solely on previous predictions?
Furthermore, there are at least two potential advantages to
temporal modeling with previous predictions. First, predic-
1. Introduction tions provide explicit information. For long-term temporal
modeling, all preceding predictions can be integrated into
High-definition (HD) maps are essential for the au- one concise prior through advanced post-processing tech-
tonomous driving system, providing vital information for niques such as filtering, tracking, merging, and curve fitting.
the vehicle’s localization, prediction, and planning mod- Second, predictions are vectorized representations, facili-
ules. However, creating and maintaining HD maps is tating the seamless integration of the framework to include

1
Category
Confidence
Location

Current Previous Current Previous Current Previous


BEV Features BEV Features Query Features Query Features Query Features Predictions
Current Previous
Perspective Features Perspective Features

Temporal Module
BEV or PV BEV or PV
PV Decoder Decoder Decoder

BEV Decoder

Current Predictions Current Predictions Current Predictions


Current Predictions

(a) BEV Features (b) Perspective Features (c) Query Features (d) Predictions

Figure 1. The simplified pipeline corresponds to various temporal representations, categorized as follows: (a) BEV features, (b) perspective
features, (c) query features, and (d) predictions. Items highlighted in yellow represent the temporal modules.

map priors like standard definition (SD) maps, outdated HD During training, the previous-predictions-based query gen-
maps, and crowd-sourced maps. erator alternates randomly between single-frame and tem-
In this paper, we introduce our exploration of temporal poral modes, as shown in the upper section of Fig. 2(b).
modeling using previous predictions for the construction In single-frame mode, previous predictions are deliberately
of online vectorized HD maps. This approach presents a omitted. As a result, PrevPredMap excels in both single-
dual challenge. On one hand, previous predictions encapsu- frame and temporal modes during inference.
late compact high-level information, as opposed to the rich In summary, the contributions of this work include:
low-level features. It is imperative that the temporal model
is meticulously crafted to ensure that previous predictions • We introduce a novel temporal modeling framework,
are not only accurately encoded but also effectively lever- named PrevPredMap, which pioneers the use of previ-
aged for current predictions. On the other hand, the tem- ous predictions for the construction of online vector-
poral model must excel in both single-frame and temporal ized HD maps.
modes, guaranteeing that its predictions are reliably propa-
• Two essential modules for PrevPredMap are meticu-
gated from the very first frame.
lously crafted: the previous-predictions-based query
Incorporating the aforementioned considerations, we in- generator and the dynamic-position-query decoder.
troduce PrevPredMap, a novel temporal modeling frame- These modules not only enable PrevPredMap to effec-
work designed to harness previous predictions for con- tively encode and utilize previous predictions but also
structing online vectorized HD maps. Two key modules ensure robust performance across both single-frame
of PrevPredMap are meticulously crafted: the previous- and temporal modes.
predictions-based query generator and the dynamic-
position-query decoder. The previous-predictions-based • We present an enhanced single-frame baseline that in-
query generator employs a separate encoding strategy for curs minimal computational overhead, which is funda-
previous predictions, which typically encompass category, mental to the PrevPredMap framework.
confidence, and location information. Specifically, the cate-
gory and confidence information from previous predictions • PrevPredMap achieves new state-of-the-art results on
are encoded into the content queries, while the location in- existing benchmarks of online vectorized HD map
formation is utilized as the initial location predictions and construction, validating the effectiveness of this inno-
further transformed into the initial position queries. Sub- vative temporal modeling approach.
sequently, the dynamic-position-query decoder introduces
a dynamic update mechanism for the position queries. As 2. Related Work
depicted in Fig. 2(c), the position queries are dynamically
2.1. Online Vectorized HD Map Construction
updated based on the location predictions from the preced-
ing decoder layer. Consequently, previous predictions are Online vectorized HD map construction was initially ap-
effectively encoded and utilized to produce current predic- proached as a semantic segmentation task [4, 13, 17, 28]. To
tions. Moreover, we have devised a dual-mode strategy. build vectorized HD map, HDMapNet [13] first generates

2
BEV semantic segmentation maps and then groups these 3. Method
pixel-wise results to vectorized instances through heuris-
tic post-processing. VectorMapNet [25] introduced the first 3.1. Overall Architecture
end-to-end framework, utilizing an auto-regressive trans- PrevPredMap comprises three primary components, as
former to sequentially retrieve vectorized map instances. illustrated in Fig. 2 (a): the BEV feature extractor,
MapTR [18] further advanced this end-to-end paradigm to the previous-predictions-based query generator, and the
a one-stage, parallel framework, significantly enhancing ef- dynamic-position-query decoder. The BEV feature extrac-
ficiency through the introduction of a unified permutation- tor is a 2D backbone followed by a parameterized PV-
equivalent representation and a hierarchical query embed- to-BEV transformation network, to extract BEV features
ding scheme. Concurrent and follow-up works [2, 5, 7, from surrounding multi-view images. Subsequently, the
10, 15, 19, 23, 26, 30, 36, 39, 41, 42] have proposed vari- previous-predictions-based query generator introduces to
ous insightful methods for further enhancement. These separately encode different types of information from pre-
include the development of concise map representations vious predictions, which are then utilized by the dynamic-
[5,30,39,41], effective attention mechanisms [5,10,19,36], position-query decoder to produce current predictions.
and auxiliary supervisions [10, 19, 41]. Additionally, other
studies have explored the construction of vectorized HD 3.2. Previous-Predictions-Based Query Generator
maps using external information, such as auxiliary maps The previous-predictions-based query generator em-
[6, 27, 31, 38] and temporal information [32, 35, 40, 43]. ploys a dual-mode strategy to ensure PrevPredMap’s robust
performance in both temporal and single-frame modes. As
2.2. Temporal Modeling of BEV Perception illustrated in the upper section of Fig. 2(b), during train-
ing, the previous-predictions-based query generator ran-
Temporal information for perception naturally emerges domly alternates between single-frame and temporal modes
at runtime without incurring any acquisition costs. Re- with probabilities p and 1 − p, respectively. In single-
cent advancements in BEV perception have explored the frame mode, the queries are generated without relying
use of temporal information to enhance perception out- on previous predictions, designated as the non-previous-
comes. BEVDet4D [11] and BEVFormer [17] were the pi- predictions-based queries. In contrast, in temporal mode,
oneers in leveraging previous BEV features for BEV per- the queries incorporate both the non-previous-predictions-
ception. Subsequent works [29, 37] have expanded this based and previous-predictions-based queries. During in-
temporal fusion approach from the BEV features of a sin- ference, the previous-predictions-based query generator se-
gle previous frame to the stacked BEV features of multi- lects between single-frame and temporal modes based on
ple previous frames, facilitating long-term temporal model- whether the input is the first frame, as shown in the lower
ing. In addition, sparse-based detectors that do not rely on section of Fig. 2(b).
dense BEV features have also investigated the use of previ- In addition to the dual-mode strategy, we propose sepa-
ous perspective features through deformable cross-attention rately encoding different types of information from previ-
mechanisms [20, 24]. Recently, VideoBEV [8] introduced a ous predictions. Specifically, the category and confidence
method for propagating previous BEV features in a stream- information from previous predictions are encoded into the
ing manner, which significantly reduces storage and compu- content queries, whereas the location information serves as
tation costs associated with long-term temporal modeling, the initial location predictions and is subsequently encoded
as opposed to stacking multiple previous BEV features in into the initial position queries. Further details are elabo-
a single forward pass. Similarly, in a streaming approach, rated in the two following subsections.
StreamPETR [33] and Sparse4D v2 [21] have proposed uti- Non-Previous-Predictions-Based Query. The queries for
lizing previous query features, further minimizing the re- the decoder typically include two types: content queries
source requirements for temporal information processing. and position queries. The content queries are designed to
extract features and generate predictions in every decoder
In the domain of online vectorized HD map construction,
layer, while the position queries provide the location infor-
StreamMapNet [40] combines previous BEV features and
mation to the attention modules within the decoder. In line
query features in a streaming fashion for enhanced tempo-
with MapTR [18], we employ a hierarchical query embed-
ral modeling. Building upon this foundation, SQD-MapNet
ding scheme to encode each map element. Specifically, the
[32] introduces a stream query denoising strategy to learn
hierarchical content and position queries for the j-th point
the temporal consistency of map elements, thereby further
of the i-th map element are formulated as:
improving the performance. Furthermore, NMP [35] and
NeMO [43] propose region-centric methods that harness con
qij = qicon−ins + qjcon−pt , (1)
temporal information, exhibiting significant potential for
pos
practical applications. qij = qipos−ins + qjpos−pt , (2)

3
BEV
Dynamic-Position-Query
Feature
Decoder
Extractor

Current
Surrounding Multi-view Images Predictions
Previous-Predictions-Based
Query Generator

Lane Divider
Previous Pedestrian Crossing
Predictions Road Boundary

(a) Pipeline Overview

Dual-Mode Strategy of Previous-Predictions-based Query Generator The i-th Layer of Dynamic-Position-Query Decoder Location
Predictions @ i -1
1-p
Random Content Queries @ i Position Queries @ i
Training n
Selector
p

n-k k Self Attention

Cross Attention
yes

Inference First-Frame n Feed Forward


Estimator Class
no Predictor

n-k k Content Queries @ i + 1

Location Location
Non-Previous-Predictions-Based Query Predictor Predictions @ i

Previous-Predictions-Based Query Position Queries @ i + 1

(b) Previous-Predictions-Based Query Generator (c) Dynamic-Position-Query Decoder

Figure 2. (a) Overall architecture of the proposed PrevPredMap, consisting of three primary modules. The BEV feature extractor is a
standard part to obtain BEV features from multi-view images. The previous-predictions-based query generator and the dynamic-position-
query decoder are meticulously designed to effectively encode and utilize previous predictions for producing current predictions. (b) The
dual-mode strategy of the previous-predictions-based query generator. (c) The dynamic update mechanism of the dynamic-position-query
decoder. Yellow arrows indicate the generation of dynamic position queries based on location predictions of the preceding decoder layer.

where qicon−ins and qipos−ins are the instance-level content where qjloc−pt represents the non-previous-predictions-
and position queries for the i-th map element, qjcon−pt and based location for the j-th point, which is also learnable pa-
qjpos−pt are the point-level content and position queries for rameters and randomly initialized. P E() denotes a position
the j-th point. Following MapTR [18], qicon−ins , qipos−ins , encoder, which utilizes a sinusoidal encoding function fol-
qjcon−pt , and qjpos−pt are all learnable parameters and ran- lowed by a linear projection.
domly initialized. In PrevPredMap, however, we modify the Previous-Predictions-Based Query. The top k previous
implementation of qjpos−pt to ensure consistency between predictions, rationed based on confidence, are leveraged to
queries that are not based on previous predictions and those generate previous-predictions-based queries after ego trans-
that are. Specifically, for the non-previous-predictions- formation. These previous predictions encompass category,
based queries, the hierarchical content and position queries confidence, and location information. By employing a sep-
for the j-th point of the i-th map element are formulated as: arate encoding strategy, we encode the category and confi-
dence information into the content queries, and the location
con
qij = qicon−ins + qjcon−pt , (3) information into the position queries. Specifically, for the
previous-predictions-based queries, the hierarchical content
pos
qij = qipos−ins + P E(qjloc−pt ), (4) and position queries for the j-th point of the i-th previous

4
map element are formulated as: requirements. Moreover, the auxiliary group-wise one-to-
many branch is more closely aligned with the main one-
con
qij = Linear(qicate ) + Linear(qiconf ) + qjcon−pt , (5) to-one branch, enhancing inference performance. Ulti-
pos loc mately, this group-wise method, when augmented with an
qij = P E(qij ), (6)
increased number of instance queries, forms an advanced
where qicate and qiconf are the predicted category and con- single-frame baseline that is fundamental to PrevPredMap.
fidence, respectively, for the i-th previous map element.
loc
Linear() denotes a linear projection function, and qij is 4. Experiment
the predicted location for the j-th point of the i-th previous
map element. 4.1. Experimental Setup

3.3. Dynamic-Position-Query Decoder Datasets. We evaluate PrevPredMap on two popular and


large-scale datasets: nuScenes [1] and Argoverse2 [34].
Benefiting from the previous-predictions-based query The nuScenes dataset offers 2D vectorized maps alongside
generator, the content and position queries are encoded with 1000 scenes, with 700 designated for training and 150 for
previous predictions and subsequently fed into the decoder. validation. Each scene encompasses 20 seconds of 2Hz
Static Position Query. In existing methods [18, 19], the RGB images captured by 6 cameras. Argoverse2, on the
position queries remain unchanged throughout all decoder other hand, delivers 3D vectorized maps and consists of
layers. Consequently, the location information supplied to 1000 logs, with 700 allocated for training and 150 for val-
the attention modules becomes outdated beyond the first de- idation. Each log comprises 15 seconds of 20Hz RGB im-
coder layer. ages from 7 ring cameras. To align with the Argoverse2
Dynamic Position Query. To address the issue of outdated setup used by existing HD map construction methods, we
location information, we introduce a dynamic update mech- adjust the camera frame rates from 20Hz to 2Hz and 10Hz,
anism that optimizes the utilization of this information. As respectively.
illustrated in Fig. 2 (c), the position queries at the current de- Evaluation Metrics. Consistent with previous methods
coder layer are dynamically updated based on the predicted [13, 18, 25], we select three static map categories for a fair
locations from the preceding decoder layer. Specifically, the evaluation: pedestrian crossings, lane dividers, and road
position query for the j-th point of the i-th map element at boundaries. The perception range is set as 30m front and
the n-th decoder layer is formulated as: rear and 15m left and right of the vehicle. The common av-
 pos erage precision (AP) based on Chamfer Distance is used as
pos qij if n = 0
qij @n = loc , (7) the evaluation metric under 3 threholds of {0.5, 1.0, 1.5}m.
P E(qij @(n − 1)) if n > 0
Implementation Details. We utilize ResNet50 [9] as the
where qijloc
@(n − 1) represents the predicted location for perspective backbone and LSS-based BEVPoolv2 [12] as
the j-th point of the i-th map element at the (n-1)-th decoder the parameterized PV-to-BEV transformation network. The
layer. Consequently, the dynamic-position-query decoder is optimizer is AdamW with a weight decay 0.01, and the ini-
implemented, enabling the position queries to convey up-to- tial learning rate is set to 0.0006, employing a cosine decay
date location information effectively. schedule. The batch size is 16 and all models are trained
with 4 NVIDIA A100 GPUs. We define the size of each
3.4. An Enhanced Single-Frame Baseline BEV grid as 0.3 meters. The default numbers of instance
Our single-frame baseline starts from MapTRv2 [19], queries, point queries and decoder layers are 100, 20 and
which is open-sourced and competitive on existing bench- 6, respectively. For the previous-predictions-based query
marks. As demonstrated by existing works [14,18], a signif- generator, the probability p is set to 0.5. The amount of
icant improvement can be achieved by increasing the num- the previous-predictions-based queries, denoted as k, is set
ber of instance queries with only a minimal consumption of to 10 for nuScenes and 8 for Argoverse2. For auxiliary
computational resources. However, MapTRv2 [19] incor- group-wise one-to-many matching, the number of groups
porates an auxiliary one-to-many matching branch to expe- is 6. During training, the top k predictions of all training
dite convergence, which, while beneficial, results in sub- samples are stored and updated immediately after their re-
stantial memory usage during training. Moreover, as the spective iterations in a predefined dictionary, allowing for
number of instance queries increases, the memory occupa- the retrieval of previous predictions as needed.
tion in self-attention modules grows quadratically.
4.2. Comparisons with State-of-the-art Methods
Inspired by Group DETR [3], we introduce a group-wise
one-to-many branch as an innovative alternative. This ap- Performance on nuScenes. As depicted in Tab. 1, Pre-
proach enables the self-attention modules to operate in par- vPredMap achieves 66.3 mAP and 71.3 mAP with 24-epoch
allel on a group-wise basis, substantially lowering memory and 110-epoch training schedules, respectively. These re-

5
Table 1. Comparison with SOTA methods on nuScenes. All backbones utilized are ResNet50. ”Temp.” signifies the utilization of temporal
information. The * indicates that MapTRv2 has been re-implemented with the number of instance queries set to 100. The ⋆ denotes that
FPS measurements were conducted on the same machine with NVIDIA RTX A6000 for fair comparison.

Method Epoch Temp. APdiv APped APbou mAP FPS


MapTR ICLR23 [18] 24 no 51.5 46.3 53.1 50.3 15.1
MapTRv2* arxiv23 [19] 24 no 62.8 62.0 65.4 63.4 16.8⋆
StreamMapNet WACV24 [40] 30 yes 66.3 61.7 62.1 63.4 14.9⋆
SQD-MapNet arxiv24 [32] 24 yes 66.6 63.6 64.8 65.0 14.9⋆
PrevPredMap (Ours) 24 yes 66.9 64.5 67.6 66.3 15.7⋆
MGMap CVPR24 [23] 24 no 65.0 61.8 67.5 64.8 11.6
HIMap CVPR24 [42] 30 no 68.4 62.6 69.1 66.7 11.4
MapQR ECCV24 [26] 24 no 68.0 63.4 67.7 66.4 14.7⋆
PrevPredMap++ (Ours) 24 yes 68.7 66.0 68.3 67.6 13.9⋆
MapTRICLR23 [18] 110 no 59.8 56.2 60.1 58.7 15.1
MapTRv2* arxiv23 [19] 110 no 68.8 68.0 71.0 69.2 16.8⋆
PrevPredMap (Ours) 110 yes 70.0 71.2 72.8 71.3 15.7⋆

Table 2. Comparison with SOTA methods on Argoverse2. All backbones utilized are ResNet50. ”Dim.” refers to the dimension of the
predicted coordinates for map elements. ”Freq.” denotes the sampling frequency applied during training. ”Temp.” signifies the utilization
of temporal information. The * indicates that MapTRv2 has been re-implemented with the number of instance queries set to 100.

Method Dim. Epoch Freq. Temp. APdiv APped APbou mAP


MapTRv2* arxiv23 [19] 2d 30 2Hz no 70.9 61.7 65.8 66.1
StreamMapNet WACV24 [40] 2d 30 2Hz yes 62.0 59.5 63.0 61.5
SQD-MapNet arxiv24 [32] 2d 30 2Hz yes 64.9 60.2 64.9 63.3
PrevPredMap(Ours) 2d 30 2Hz yes 72.0 64.2 68.3 68.2
MapTRv2* arxiv23 [19] 2d 6 10Hz no 73.3 65.3 69.1 69.2
HIMap CVPR24 [42] 2d 6 10Hz no 69.5 69.0 70.3 69.4
ADMapv2 ECCV24 [10] 2d 6 10Hz no 72.4 64.5 68.9 68.7
PrevPredMap(Ours) 2d 6 10Hz yes 74.8 65.7 70.9 70.5
MapTRv2* arxiv23 [19] 3d 6 10Hz no 71.7 62.2 66.7 66.9
PrevPredMap(Ours) 3d 6 10Hz yes 71.4 64.1 67.4 67.6

sults outperform all SOTA methods in terms of training con- tion on Argoverse2.
vergence, validation accuracy and inference speed. When
enhanced with GKT-h, a more powerful BEV encoder intro- 4.3. Ablation Study
duced in MapQR [26], PrevPredMap++ attains 67.5 mAP
Components Ablation. As demonstrated in Tab. 3, a
with a 24-epoch training schedule.
straightforward increase in the query amount from 50
Performance on Argoverse 2. Argoverse2 offers a 3D vec- to 100, complemented by an auxiliary group-wise one-
torized map, which includes additional height information to-many branch, achieves a total improvement of 3.3
not present in the nuScenes dataset. Since most existing mAP. Additionally, we independently assess the impact
works do not utilize this height information, with the excep- of the previous-prediction-based query generator and the
tion of MapTRv2 [19], we train PrevPredMap under both dynamic-position-query decoder. The individual enhance-
2D and 3D configurations. For fair comparison, the sam- ments from these modules are modest. However, when in-
pling frequency is set to 2Hz and 10Hz, respectively. It tegrated, these modules effectively encode and utilize previ-
should be noted that the amount of training data at a 2Hz ous predictions, leading to a final enhancement of 4.8 mAP,
sampling frequency is 20% of that at a 10Hz sampling fre- accompanied by only a 6.5% reduction in FPS.
quency. As observed in Tab. 2, PrevPredMap surpasses all Performance Comparison of the Single-Frame and Tem-
SOTA methods in both 2D and 3D vectorized map construc- poral Modes. Benefiting from the proposed dual-mode

6
Table 3. Ablation of the proposed PrevPredMap. The experiments were performed on nuScenes, following a 24-epoch training schedule.
FPS measurements are conducted on the same machine with NVIDIA RTX A6000.

Query Group-Wise Previous-Predictions-based Dynamic-Position-Query mAP FPS


Amount One-to-Many Query Qenerator Decoder
50 61.5 16.8
100 63.4 16.8
100 ✓ 64.8 16.8
100 ✓ ✓ 65.4 16.1
100 ✓ ✓ 64.9 16.0
100 ✓ ✓ ✓ 66.3 15.7

Table 4. Performance comparison of the single-frame and tem- Table 5. Probability of the previous-predictions-based query gen-
poral modes for PrevPredMap on nuScenes, following a 24-epoch erator. All experiments were conducted on nuScenes, following a
training schedule. FPS measurements are conducted on the same 24-epoch training schedule.
machine with NVIDIA RTX A6000.
Probability mAP APdiv APped APbou
Mode mAP APdiv APped APbou FPS
0.4 66.2 65.8 65.4 67.3
Single-Frame 65.7 66.0 63.9 67.3 16.1 0.5 66.3 66.9 64.5 67.6
Temporal 66.3 66.9 64.5 67.6 15.7 0.6 65.7 65.8 63.8 67.5

Table 6. The number of previous-predictions-based queries. All


strategy implemented in the previous-predictions-based experiments were conducted on nuScenes, following a 24-epoch
query generator, PrevPredMap excels in both single-frame training schedule.
and temporal modes. As depicted in Tab. 4, the single-frame
mode of PrevPredMap is 65.7 mAP, which is still compara- k mAP APdiv APped APbou
ble to the existing state-of-the-art methods listed in Tab. 1.
This characteristic is advantageous because online infer- 5 66.1 65.5 65.0 67.5
ence can be interrupted due to various unexpected emer- 8 66.3 65.8 64.9 68.2
gencies. It is possible that the single-frame mode may oc- 10 66.3 66.9 64.5 67.6
casionally be called upon during runtime. Additionally, the 12 66.5 66.9 64.5 68.0
discrepancy between the single-frame and temporal modes 15 65.7 65.3 64.8 66.9
is not particularly noticeable. We offer two possible expla-
nations for this phenomenon. First, map elements are more
likely to be severely occluded in situations such as when a Generator. As depicted in the upper part of Fig. 2 (b),
large truck is nearby or during a traffic jam; however, these the probabilities of p and 1 − p represent the propensity for
scenarios are rare in existing datasets, thereby limiting the temporal and single-frame learning, respectively. Accord-
advantage of the temporal mode. As shown in Table 3 of ing to Tab. 5, PrevPredMap attains optimal performance at
SQD-MapNet [32], the mAP even decreases by 0.7 mAP a probability of 0.5, likely owing to the equilibrium estab-
when the temporal stream is added to the single-frame base- lished between temporal and single-frame learning modes
line. Second, PrevPredMap currently only utilizes predic- in its training regimen.
tions from the previous frame. As demonstrated in Table 5 The Number of Previous-Predictions-Based Queries.
of StreamPETR [33], there is a substantial increase in mAP As illustrated in Tab. 6, PrevPredMap’s performance in-
if the number of frames considered is increased from one creases with an elevated count of previous-predictions-
to seven. We reserve this avenue for future exploration, as based queries, reaching saturation at a number approxi-
the post-processing of predictions from multiple preceding mately 12. It is important to note that the average number
frames is complex and would require significant effort in of map elements per frame is less than 10 for nuScenes. An
refining detailed settings. Finally, the dual-mode setting for excessively high value of k might introduce a multitude of
temporal modeling, as introduced by PrevPredMap, is un- previous predictions characterized by high uncertainty, po-
precedented. There are no existing papers for direct com- tentially leading to adverse effects.
parison. Qualitative Analysis. Fig. 3 displays three common sce-
Probability of the Previous-Predictions-Based Query narios in autonomous driving, where the camera’s field of

7
PrevPredMap PrevPredMap
Surrounding Multi-View Images Ground Truth MapTRv2*
Single-Frame Temporal

Figure 3. Comparison of PrevPredMap with single-frame SOTA methods on qualitative visualization under various occlusion scenarios.
Each sub-part displays four qualitative results: Ground Truth, MapTRv2, PrevPredMap Single-Frame, and PrevPredMap Temporal. The
* indicates that MapTRv2 has been re-implemented with the number of instance queries set to 100. Green, orange and blue lines denote
road boundaries, lane dividers and pedestrian crossings, respectively.

view is partially obstructed by a large truck, plastic water- 5. Conclusion


filled barriers, and roadside railings, respectively. As shown
in the highlighted gray areas of Fig. 3, single-frame models
struggle to accurately perceive map elements due to the var- This paper focuses on temporal modeling using previ-
ious obstacles. In contrast, PrevPredMap effectively lever- ous predictions for the construction of online vectorized
ages previous predictions for constructing vectorized HD HD maps. We introduce two meticulously designed mod-
maps, underscoring the effectiveness of this innovative tem- ules: the previous-predictions-based query generator and
poral modeling approach. the dynamic-position-query decoder. With these, we pro-
pose a novel temporal modeling framework, PrevPredMap,
4.4. Limitations and Future Work
which leverages previous predictions to enhance online vec-
Based on our current understanding, the limitations and torized HD map construction. These modules not only en-
future work of PrevPredMap are discussed in two main as- able PrevPredMap to effectively encode and utilize previ-
pects. Firstly, PrevPredMap currently leverages predictions ous predictions but also ensure robust performance across
solely from the preceding frame. It is anticipated that the both single-frame and temporal modes. Besides, we present
performance of the temporal mode could be significantly an enhanced single-frame baseline that incurs minimal
enhanced by adeptly post-processing and utilizing predic- computational overhead, which is fundamental to the Pre-
tions from multiple preceding frames. Secondly, map priors vPredMap framework. PrevPredMap is simple yet ef-
constitute essential complementary information for online fective, achieving new state-of-the-art results on existing
HD map construction. There is a strong expectation that benchmarks. We hope that PrevPredMap will provide new
incorporating map priors into the PrevPredMap framework insights into the temporal modeling of BEV perception for
will further enhance the reliability of the constructed map. the community.

8
References [14] Toyota Li. Mapnext: Revisiting training and scaling prac-
tices for online vectorized hd map construction. arXiv
[1] Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, preprint arXiv:2401.07323, 2024. 5
Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gi-
[15] Tianyu Li, Peijin Jia, Bangjun Wang, Li Chen, Kun Jiang,
ancarlo Baldan, and Oscar Beijbom. nuscenes: A multi-
Junchi Yan, and Hongyang Li. Lanesegnet: Map learning
modal dataset for autonomous driving. In Proceedings of
with lane segment perception for autonomous driving. arXiv
the IEEE/CVF conference on computer vision and pattern
preprint arXiv:2312.16108, 2023. 3
recognition, pages 11621–11631, 2020. 5
[16] Yinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang, Zengran
[2] Jiacheng Chen, Yuefan Wu, Jiaqi Tan, Hang Ma, and Ya-
Wang, Yukang Shi, Jianjian Sun, and Zeming Li. Bevdepth:
sutaka Furukawa. Maptracker: Tracking with strided mem-
Acquisition of reliable depth for multi-view 3d object detec-
ory fusion for consistent vector hd mapping. arXiv preprint
tion. In Proceedings of the AAAI Conference on Artificial
arXiv:2403.15951, 2024. 3
Intelligence, volume 37, pages 1477–1485, 2023. 1
[3] Qiang Chen, Xiaokang Chen, Jian Wang, Shan Zhang, Kun
[17] Zhiqi Li, Wenhai Wang, Hongyang Li, Enze Xie, Chong-
Yao, Haocheng Feng, Junyu Han, Errui Ding, Gang Zeng,
hao Sima, Tong Lu, Yu Qiao, and Jifeng Dai. Bevformer:
and Jingdong Wang. Group detr: Fast detr training with
Learning bird’s-eye-view representation from multi-camera
group-wise one-to-many assignment. In Proceedings of the
images via spatiotemporal transformers. In European con-
IEEE/CVF International Conference on Computer Vision,
ference on computer vision, pages 1–18. Springer, 2022. 1,
pages 6633–6642, 2023. 5
2, 3
[4] Shaoyu Chen, Tianheng Cheng, Xinggang Wang, Wenming
Meng, Qian Zhang, and Wenyu Liu. Efficient and robust [18] Bencheng Liao, Shaoyu Chen, Xinggang Wang, Tianheng
2d-to-bev representation learning via geometry-guided ker- Cheng, Qian Zhang, Wenyu Liu, and Chang Huang. Maptr:
nel transformer. arXiv preprint arXiv:2206.04584, 2022. 1, Structured modeling and learning for online vectorized hd
2 map construction. In International Conference on Learning
Representations, 2023. 1, 3, 4, 5, 6
[5] Wenjie Ding, Limeng Qiao, Xi Qiu, and Chi Zhang. Pivot-
net: Vectorized pivot learning for end-to-end hd map con- [19] Bencheng Liao, Shaoyu Chen, Yunchi Zhang, Bo Jiang, Qian
struction. In Proceedings of the IEEE/CVF International Zhang, Wenyu Liu, Chang Huang, and Xinggang Wang.
Conference on Computer Vision, pages 3672–3682, 2023. 3 Maptrv2: An end-to-end framework for online vectorized hd
[6] Wenjie Gao, Jiawei Fu, Haodong Jing, and Nanning Zheng. map construction. arXiv preprint arXiv:2308.05736, 2023.
Complementing onboard sensors with satellite map: A 3, 5, 6
new perspective for hd map construction. arXiv preprint [20] Xuewu Lin, Tianwei Lin, Zixiang Pei, Lichao Huang, and
arXiv:2308.15427, 2023. 3 Zhizhong Su. Sparse4d: Multi-view 3d object detec-
[7] Xunjiang Gu, Guanyu Song, Igor Gilitschenski, Marco tion with sparse spatial-temporal fusion. arXiv preprint
Pavone, and Boris Ivanovic. Accelerating online mapping arXiv:2211.10581, 2022. 1, 3
and behavior prediction via direct bev feature attention. [21] Xuewu Lin, Tianwei Lin, Zixiang Pei, Lichao Huang, and
arXiv preprint arXiv:2407.06683, 2024. 3 Zhizhong Su. Sparse4d v2: Recurrent temporal fusion with
[8] Chunrui Han, Jianjian Sun, Zheng Ge, Jinrong Yang, Run- sparse model. arXiv preprint arXiv:2305.14018, 2023. 1, 3
pei Dong, Hongyu Zhou, Weixin Mao, Yuang Peng, and [22] Xuewu Lin, Zixiang Pei, Tianwei Lin, Lichao Huang, and
Xiangyu Zhang. Exploring recurrent long-term tempo- Zhizhong Su. Sparse4d v3: Advancing end-to-end 3d detec-
ral fusion for multi-view 3d perception. arXiv preprint tion and tracking. arXiv preprint arXiv:2311.11722, 2023.
arXiv:2303.05970, 2023. 3 1
[9] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. [23] Xiaolu Liu, Song Wang, Wentong Li, Ruizi Yang, Junbo
Deep residual learning for image recognition. In Proceed- Chen, and Jianke Zhu. Mgmap: Mask-guided learning for
ings of the IEEE conference on computer vision and pattern online vectorized hd map construction. In Proceedings of
recognition, pages 770–778, 2016. 5 the IEEE/CVF Conference on Computer Vision and Pattern
[10] Haotian Hu, Fanyi Wang, Yaonong Wang, Laifeng Hu, Jing- Recognition, pages 14812–14821, 2024. 3, 6
wei Xu, and Zhiwang Zhang. Admap: Anti-disturbance [24] Yingfei Liu, Junjie Yan, Fan Jia, Shuailin Li, Aqi Gao, Tian-
framework for reconstructing online vectorized hd map. cai Wang, and Xiangyu Zhang. Petrv2: A unified framework
arXiv preprint arXiv:2401.13172, 2024. 3, 6 for 3d perception from multi-camera images. In Proceedings
[11] Junjie Huang and Guan Huang. Bevdet4d: Exploit tempo- of the IEEE/CVF International Conference on Computer Vi-
ral cues in multi-camera 3d object detection. arXiv preprint sion, pages 3262–3272, 2023. 1, 3
arXiv:2203.17054, 2022. 1, 3 [25] Yicheng Liu, Tianyuan Yuan, Yue Wang, Yilun Wang, and
[12] Junjie Huang and Guan Huang. Bevpoolv2: A cutting-edge Hang Zhao. Vectormapnet: End-to-end vectorized hd map
implementation of bevdet toward deployment. arXiv preprint learning. In International Conference on Machine Learning,
arXiv:2211.17111, 2022. 5 pages 22352–22369. PMLR, 2023. 1, 3, 5
[13] Qi Li, Yue Wang, Yilun Wang, and Hang Zhao. Hdmapnet: [26] Zihao Liu, Xiaoyu Zhang, Guangwei Liu, Ji Zhao,
An online hd map construction and evaluation framework. In and Ningyi Xu. Leveraging enhanced queries of point
2022 International Conference on Robotics and Automation sets for vectorized map construction. arXiv preprint
(ICRA), pages 4628–4634. IEEE, 2022. 1, 2, 5 arXiv:2402.17430, 2024. 3, 6

9
[27] Katie Z Luo, Xinshuo Weng, Yan Wang, Shuang Wu, Jie range vectorized hd map construction. arXiv preprint
Li, Kilian Q Weinberger, Yue Wang, and Marco Pavone. arXiv:2310.13378, 2023. 3
Augmenting lane perception and topology understanding [40] Tianyuan Yuan, Yicheng Liu, Yue Wang, Yilun Wang, and
with standard definition navigation maps. arXiv preprint Hang Zhao. Streammapnet: Streaming mapping network for
arXiv:2311.04079, 2023. 3 vectorized online hd map construction. In Proceedings of the
[28] Bowen Pan, Jiankai Sun, Ho Yin Tiga Leung, Alex Ando- IEEE/CVF Winter Conference on Applications of Computer
nian, and Bolei Zhou. Cross-view semantic segmentation Vision, pages 7356–7365, 2024. 3, 6
for sensing surroundings. IEEE Robotics and Automation [41] Zhixin Zhang, Yiyuan Zhang, Xiaohan Ding, Fusheng Jin,
Letters, 5(3):4867–4873, 2020. 1, 2 and Xiangyu Yue. Online vectorized hd map construction
[29] Jinhyung Park, Chenfeng Xu, Shijia Yang, Kurt Keutzer, using geometry. arXiv preprint arXiv:2312.03341, 2023. 3
Kris Kitani, Masayoshi Tomizuka, and Wei Zhan. Time will [42] Yi Zhou, Hui Zhang, Jiaqian Yu, Yifan Yang, Sangil Jung,
tell: New outlooks and a baseline for temporal multi-view 3d Seung-In Park, and ByungIn Yoo. Himap: Hybrid represen-
object detection. arXiv preprint arXiv:2210.02443, 2022. 1, tation learning for end-to-end vectorized hd map construc-
3 tion. In Proceedings of the IEEE/CVF Conference on Com-
[30] Limeng Qiao, Wenjie Ding, Xi Qiu, and Chi Zhang. End- puter Vision and Pattern Recognition, pages 15396–15406,
to-end vectorized hd-map construction with piecewise bezier 2024. 3, 6
curve. In Proceedings of the IEEE/CVF Conference on Com- [43] Xi Zhu, Xiya Cao, Zhiwei Dong, Caifa Zhou, Qiangbo Liu,
puter Vision and Pattern Recognition, pages 13218–13228, Wei Li, and Yongliang Wang. Nemo: Neural map growing
2023. 3 system for spatiotemporal fusion in bird’s-eye-view and bdd-
[31] Rémy Sun, Li Yang, Diane Lingrand, and Frédéric Precioso. map benchmark. arXiv preprint arXiv:2306.04540, 2023. 3
Mind the map! accounting for existing map information
when estimating online hdmaps from sensor data. arXiv
preprint arXiv:2311.10517, 2023. 3
[32] Shuo Wang, Fan Jia, Yingfei Liu, Yucheng Zhao, Zehui
Chen, Tiancai Wang, Chi Zhang, Xiangyu Zhang, and Feng
Zhao. Stream query denoising for vectorized hd map con-
struction. arXiv preprint arXiv:2401.09112, 2024. 3, 6, 7
[33] Shihao Wang, Yingfei Liu, Tiancai Wang, Ying Li, and Xi-
angyu Zhang. Exploring object-centric temporal modeling
for efficient multi-view 3d object detection. arXiv preprint
arXiv:2303.11926, 2023. 1, 3, 7
[34] Benjamin Wilson, William Qi, Tanmay Agarwal, John
Lambert, Jagjeet Singh, Siddhesh Khandelwal, Bowen
Pan, Ratnesh Kumar, Andrew Hartnett, Jhony Kaesemodel
Pontes, et al. Argoverse 2: Next generation datasets for
self-driving perception and forecasting. arXiv preprint
arXiv:2301.00493, 2023. 5
[35] Xuan Xiong, Yicheng Liu, Tianyuan Yuan, Yue Wang, Yilun
Wang, and Hang Zhao. Neural map prior for autonomous
driving. In Proceedings of the IEEE/CVF Conference on
Computer Vision and Pattern Recognition, pages 17535–
17544, 2023. 3
[36] Zhenhua Xu, Kenneth KY Wong, and Hengshuang Zhao.
Insightmapper: A closer look at inner-instance informa-
tion for vectorized high-definition mapping. arXiv preprint
arXiv:2308.08543, 2023. 3
[37] Chenyu Yang, Yuntao Chen, Hao Tian, Chenxin Tao, Xizhou
Zhu, Zhaoxiang Zhang, Gao Huang, Hongyang Li, Yu Qiao,
Lewei Lu, et al. Bevformer v2: Adapting modern image
backbones to bird’s-eye-view recognition via perspective su-
pervision. In Proceedings of the IEEE/CVF Conference on
Computer Vision and Pattern Recognition, pages 17830–
17839, 2023. 3
[38] Jiawei Yao, Xiaochao Pan, Tong Wu, and Xiaofeng Zhang.
Building lane-level maps from aerial images. arXiv preprint
arXiv:2312.13449, 2023. 3
[39] Jingyi Yu, Zizhao Zhang, Shengfu Xia, and Jizhang Sang.
Scalablemap: Scalable map learning for online long-

10

You might also like