PrevPredMap: Temporal Modeling for HD Maps
PrevPredMap: Temporal Modeling for HD Maps
1
Category
Confidence
Location
Temporal Module
BEV or PV BEV or PV
PV Decoder Decoder Decoder
BEV Decoder
(a) BEV Features (b) Perspective Features (c) Query Features (d) Predictions
Figure 1. The simplified pipeline corresponds to various temporal representations, categorized as follows: (a) BEV features, (b) perspective
features, (c) query features, and (d) predictions. Items highlighted in yellow represent the temporal modules.
map priors like standard definition (SD) maps, outdated HD During training, the previous-predictions-based query gen-
maps, and crowd-sourced maps. erator alternates randomly between single-frame and tem-
In this paper, we introduce our exploration of temporal poral modes, as shown in the upper section of Fig. 2(b).
modeling using previous predictions for the construction In single-frame mode, previous predictions are deliberately
of online vectorized HD maps. This approach presents a omitted. As a result, PrevPredMap excels in both single-
dual challenge. On one hand, previous predictions encapsu- frame and temporal modes during inference.
late compact high-level information, as opposed to the rich In summary, the contributions of this work include:
low-level features. It is imperative that the temporal model
is meticulously crafted to ensure that previous predictions • We introduce a novel temporal modeling framework,
are not only accurately encoded but also effectively lever- named PrevPredMap, which pioneers the use of previ-
aged for current predictions. On the other hand, the tem- ous predictions for the construction of online vector-
poral model must excel in both single-frame and temporal ized HD maps.
modes, guaranteeing that its predictions are reliably propa-
• Two essential modules for PrevPredMap are meticu-
gated from the very first frame.
lously crafted: the previous-predictions-based query
Incorporating the aforementioned considerations, we in- generator and the dynamic-position-query decoder.
troduce PrevPredMap, a novel temporal modeling frame- These modules not only enable PrevPredMap to effec-
work designed to harness previous predictions for con- tively encode and utilize previous predictions but also
structing online vectorized HD maps. Two key modules ensure robust performance across both single-frame
of PrevPredMap are meticulously crafted: the previous- and temporal modes.
predictions-based query generator and the dynamic-
position-query decoder. The previous-predictions-based • We present an enhanced single-frame baseline that in-
query generator employs a separate encoding strategy for curs minimal computational overhead, which is funda-
previous predictions, which typically encompass category, mental to the PrevPredMap framework.
confidence, and location information. Specifically, the cate-
gory and confidence information from previous predictions • PrevPredMap achieves new state-of-the-art results on
are encoded into the content queries, while the location in- existing benchmarks of online vectorized HD map
formation is utilized as the initial location predictions and construction, validating the effectiveness of this inno-
further transformed into the initial position queries. Sub- vative temporal modeling approach.
sequently, the dynamic-position-query decoder introduces
a dynamic update mechanism for the position queries. As 2. Related Work
depicted in Fig. 2(c), the position queries are dynamically
2.1. Online Vectorized HD Map Construction
updated based on the location predictions from the preced-
ing decoder layer. Consequently, previous predictions are Online vectorized HD map construction was initially ap-
effectively encoded and utilized to produce current predic- proached as a semantic segmentation task [4, 13, 17, 28]. To
tions. Moreover, we have devised a dual-mode strategy. build vectorized HD map, HDMapNet [13] first generates
2
BEV semantic segmentation maps and then groups these 3. Method
pixel-wise results to vectorized instances through heuris-
tic post-processing. VectorMapNet [25] introduced the first 3.1. Overall Architecture
end-to-end framework, utilizing an auto-regressive trans- PrevPredMap comprises three primary components, as
former to sequentially retrieve vectorized map instances. illustrated in Fig. 2 (a): the BEV feature extractor,
MapTR [18] further advanced this end-to-end paradigm to the previous-predictions-based query generator, and the
a one-stage, parallel framework, significantly enhancing ef- dynamic-position-query decoder. The BEV feature extrac-
ficiency through the introduction of a unified permutation- tor is a 2D backbone followed by a parameterized PV-
equivalent representation and a hierarchical query embed- to-BEV transformation network, to extract BEV features
ding scheme. Concurrent and follow-up works [2, 5, 7, from surrounding multi-view images. Subsequently, the
10, 15, 19, 23, 26, 30, 36, 39, 41, 42] have proposed vari- previous-predictions-based query generator introduces to
ous insightful methods for further enhancement. These separately encode different types of information from pre-
include the development of concise map representations vious predictions, which are then utilized by the dynamic-
[5,30,39,41], effective attention mechanisms [5,10,19,36], position-query decoder to produce current predictions.
and auxiliary supervisions [10, 19, 41]. Additionally, other
studies have explored the construction of vectorized HD 3.2. Previous-Predictions-Based Query Generator
maps using external information, such as auxiliary maps The previous-predictions-based query generator em-
[6, 27, 31, 38] and temporal information [32, 35, 40, 43]. ploys a dual-mode strategy to ensure PrevPredMap’s robust
performance in both temporal and single-frame modes. As
2.2. Temporal Modeling of BEV Perception illustrated in the upper section of Fig. 2(b), during train-
ing, the previous-predictions-based query generator ran-
Temporal information for perception naturally emerges domly alternates between single-frame and temporal modes
at runtime without incurring any acquisition costs. Re- with probabilities p and 1 − p, respectively. In single-
cent advancements in BEV perception have explored the frame mode, the queries are generated without relying
use of temporal information to enhance perception out- on previous predictions, designated as the non-previous-
comes. BEVDet4D [11] and BEVFormer [17] were the pi- predictions-based queries. In contrast, in temporal mode,
oneers in leveraging previous BEV features for BEV per- the queries incorporate both the non-previous-predictions-
ception. Subsequent works [29, 37] have expanded this based and previous-predictions-based queries. During in-
temporal fusion approach from the BEV features of a sin- ference, the previous-predictions-based query generator se-
gle previous frame to the stacked BEV features of multi- lects between single-frame and temporal modes based on
ple previous frames, facilitating long-term temporal model- whether the input is the first frame, as shown in the lower
ing. In addition, sparse-based detectors that do not rely on section of Fig. 2(b).
dense BEV features have also investigated the use of previ- In addition to the dual-mode strategy, we propose sepa-
ous perspective features through deformable cross-attention rately encoding different types of information from previ-
mechanisms [20, 24]. Recently, VideoBEV [8] introduced a ous predictions. Specifically, the category and confidence
method for propagating previous BEV features in a stream- information from previous predictions are encoded into the
ing manner, which significantly reduces storage and compu- content queries, whereas the location information serves as
tation costs associated with long-term temporal modeling, the initial location predictions and is subsequently encoded
as opposed to stacking multiple previous BEV features in into the initial position queries. Further details are elabo-
a single forward pass. Similarly, in a streaming approach, rated in the two following subsections.
StreamPETR [33] and Sparse4D v2 [21] have proposed uti- Non-Previous-Predictions-Based Query. The queries for
lizing previous query features, further minimizing the re- the decoder typically include two types: content queries
source requirements for temporal information processing. and position queries. The content queries are designed to
extract features and generate predictions in every decoder
In the domain of online vectorized HD map construction,
layer, while the position queries provide the location infor-
StreamMapNet [40] combines previous BEV features and
mation to the attention modules within the decoder. In line
query features in a streaming fashion for enhanced tempo-
with MapTR [18], we employ a hierarchical query embed-
ral modeling. Building upon this foundation, SQD-MapNet
ding scheme to encode each map element. Specifically, the
[32] introduces a stream query denoising strategy to learn
hierarchical content and position queries for the j-th point
the temporal consistency of map elements, thereby further
of the i-th map element are formulated as:
improving the performance. Furthermore, NMP [35] and
NeMO [43] propose region-centric methods that harness con
qij = qicon−ins + qjcon−pt , (1)
temporal information, exhibiting significant potential for
pos
practical applications. qij = qipos−ins + qjpos−pt , (2)
3
BEV
Dynamic-Position-Query
Feature
Decoder
Extractor
Current
Surrounding Multi-view Images Predictions
Previous-Predictions-Based
Query Generator
Lane Divider
Previous Pedestrian Crossing
Predictions Road Boundary
Dual-Mode Strategy of Previous-Predictions-based Query Generator The i-th Layer of Dynamic-Position-Query Decoder Location
Predictions @ i -1
1-p
Random Content Queries @ i Position Queries @ i
Training n
Selector
p
Cross Attention
yes
Location Location
Non-Previous-Predictions-Based Query Predictor Predictions @ i
Figure 2. (a) Overall architecture of the proposed PrevPredMap, consisting of three primary modules. The BEV feature extractor is a
standard part to obtain BEV features from multi-view images. The previous-predictions-based query generator and the dynamic-position-
query decoder are meticulously designed to effectively encode and utilize previous predictions for producing current predictions. (b) The
dual-mode strategy of the previous-predictions-based query generator. (c) The dynamic update mechanism of the dynamic-position-query
decoder. Yellow arrows indicate the generation of dynamic position queries based on location predictions of the preceding decoder layer.
where qicon−ins and qipos−ins are the instance-level content where qjloc−pt represents the non-previous-predictions-
and position queries for the i-th map element, qjcon−pt and based location for the j-th point, which is also learnable pa-
qjpos−pt are the point-level content and position queries for rameters and randomly initialized. P E() denotes a position
the j-th point. Following MapTR [18], qicon−ins , qipos−ins , encoder, which utilizes a sinusoidal encoding function fol-
qjcon−pt , and qjpos−pt are all learnable parameters and ran- lowed by a linear projection.
domly initialized. In PrevPredMap, however, we modify the Previous-Predictions-Based Query. The top k previous
implementation of qjpos−pt to ensure consistency between predictions, rationed based on confidence, are leveraged to
queries that are not based on previous predictions and those generate previous-predictions-based queries after ego trans-
that are. Specifically, for the non-previous-predictions- formation. These previous predictions encompass category,
based queries, the hierarchical content and position queries confidence, and location information. By employing a sep-
for the j-th point of the i-th map element are formulated as: arate encoding strategy, we encode the category and confi-
dence information into the content queries, and the location
con
qij = qicon−ins + qjcon−pt , (3) information into the position queries. Specifically, for the
previous-predictions-based queries, the hierarchical content
pos
qij = qipos−ins + P E(qjloc−pt ), (4) and position queries for the j-th point of the i-th previous
4
map element are formulated as: requirements. Moreover, the auxiliary group-wise one-to-
many branch is more closely aligned with the main one-
con
qij = Linear(qicate ) + Linear(qiconf ) + qjcon−pt , (5) to-one branch, enhancing inference performance. Ulti-
pos loc mately, this group-wise method, when augmented with an
qij = P E(qij ), (6)
increased number of instance queries, forms an advanced
where qicate and qiconf are the predicted category and con- single-frame baseline that is fundamental to PrevPredMap.
fidence, respectively, for the i-th previous map element.
loc
Linear() denotes a linear projection function, and qij is 4. Experiment
the predicted location for the j-th point of the i-th previous
map element. 4.1. Experimental Setup
5
Table 1. Comparison with SOTA methods on nuScenes. All backbones utilized are ResNet50. ”Temp.” signifies the utilization of temporal
information. The * indicates that MapTRv2 has been re-implemented with the number of instance queries set to 100. The ⋆ denotes that
FPS measurements were conducted on the same machine with NVIDIA RTX A6000 for fair comparison.
Table 2. Comparison with SOTA methods on Argoverse2. All backbones utilized are ResNet50. ”Dim.” refers to the dimension of the
predicted coordinates for map elements. ”Freq.” denotes the sampling frequency applied during training. ”Temp.” signifies the utilization
of temporal information. The * indicates that MapTRv2 has been re-implemented with the number of instance queries set to 100.
sults outperform all SOTA methods in terms of training con- tion on Argoverse2.
vergence, validation accuracy and inference speed. When
enhanced with GKT-h, a more powerful BEV encoder intro- 4.3. Ablation Study
duced in MapQR [26], PrevPredMap++ attains 67.5 mAP
Components Ablation. As demonstrated in Tab. 3, a
with a 24-epoch training schedule.
straightforward increase in the query amount from 50
Performance on Argoverse 2. Argoverse2 offers a 3D vec- to 100, complemented by an auxiliary group-wise one-
torized map, which includes additional height information to-many branch, achieves a total improvement of 3.3
not present in the nuScenes dataset. Since most existing mAP. Additionally, we independently assess the impact
works do not utilize this height information, with the excep- of the previous-prediction-based query generator and the
tion of MapTRv2 [19], we train PrevPredMap under both dynamic-position-query decoder. The individual enhance-
2D and 3D configurations. For fair comparison, the sam- ments from these modules are modest. However, when in-
pling frequency is set to 2Hz and 10Hz, respectively. It tegrated, these modules effectively encode and utilize previ-
should be noted that the amount of training data at a 2Hz ous predictions, leading to a final enhancement of 4.8 mAP,
sampling frequency is 20% of that at a 10Hz sampling fre- accompanied by only a 6.5% reduction in FPS.
quency. As observed in Tab. 2, PrevPredMap surpasses all Performance Comparison of the Single-Frame and Tem-
SOTA methods in both 2D and 3D vectorized map construc- poral Modes. Benefiting from the proposed dual-mode
6
Table 3. Ablation of the proposed PrevPredMap. The experiments were performed on nuScenes, following a 24-epoch training schedule.
FPS measurements are conducted on the same machine with NVIDIA RTX A6000.
Table 4. Performance comparison of the single-frame and tem- Table 5. Probability of the previous-predictions-based query gen-
poral modes for PrevPredMap on nuScenes, following a 24-epoch erator. All experiments were conducted on nuScenes, following a
training schedule. FPS measurements are conducted on the same 24-epoch training schedule.
machine with NVIDIA RTX A6000.
Probability mAP APdiv APped APbou
Mode mAP APdiv APped APbou FPS
0.4 66.2 65.8 65.4 67.3
Single-Frame 65.7 66.0 63.9 67.3 16.1 0.5 66.3 66.9 64.5 67.6
Temporal 66.3 66.9 64.5 67.6 15.7 0.6 65.7 65.8 63.8 67.5
7
PrevPredMap PrevPredMap
Surrounding Multi-View Images Ground Truth MapTRv2*
Single-Frame Temporal
Figure 3. Comparison of PrevPredMap with single-frame SOTA methods on qualitative visualization under various occlusion scenarios.
Each sub-part displays four qualitative results: Ground Truth, MapTRv2, PrevPredMap Single-Frame, and PrevPredMap Temporal. The
* indicates that MapTRv2 has been re-implemented with the number of instance queries set to 100. Green, orange and blue lines denote
road boundaries, lane dividers and pedestrian crossings, respectively.
8
References [14] Toyota Li. Mapnext: Revisiting training and scaling prac-
tices for online vectorized hd map construction. arXiv
[1] Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, preprint arXiv:2401.07323, 2024. 5
Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gi-
[15] Tianyu Li, Peijin Jia, Bangjun Wang, Li Chen, Kun Jiang,
ancarlo Baldan, and Oscar Beijbom. nuscenes: A multi-
Junchi Yan, and Hongyang Li. Lanesegnet: Map learning
modal dataset for autonomous driving. In Proceedings of
with lane segment perception for autonomous driving. arXiv
the IEEE/CVF conference on computer vision and pattern
preprint arXiv:2312.16108, 2023. 3
recognition, pages 11621–11631, 2020. 5
[16] Yinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang, Zengran
[2] Jiacheng Chen, Yuefan Wu, Jiaqi Tan, Hang Ma, and Ya-
Wang, Yukang Shi, Jianjian Sun, and Zeming Li. Bevdepth:
sutaka Furukawa. Maptracker: Tracking with strided mem-
Acquisition of reliable depth for multi-view 3d object detec-
ory fusion for consistent vector hd mapping. arXiv preprint
tion. In Proceedings of the AAAI Conference on Artificial
arXiv:2403.15951, 2024. 3
Intelligence, volume 37, pages 1477–1485, 2023. 1
[3] Qiang Chen, Xiaokang Chen, Jian Wang, Shan Zhang, Kun
[17] Zhiqi Li, Wenhai Wang, Hongyang Li, Enze Xie, Chong-
Yao, Haocheng Feng, Junyu Han, Errui Ding, Gang Zeng,
hao Sima, Tong Lu, Yu Qiao, and Jifeng Dai. Bevformer:
and Jingdong Wang. Group detr: Fast detr training with
Learning bird’s-eye-view representation from multi-camera
group-wise one-to-many assignment. In Proceedings of the
images via spatiotemporal transformers. In European con-
IEEE/CVF International Conference on Computer Vision,
ference on computer vision, pages 1–18. Springer, 2022. 1,
pages 6633–6642, 2023. 5
2, 3
[4] Shaoyu Chen, Tianheng Cheng, Xinggang Wang, Wenming
Meng, Qian Zhang, and Wenyu Liu. Efficient and robust [18] Bencheng Liao, Shaoyu Chen, Xinggang Wang, Tianheng
2d-to-bev representation learning via geometry-guided ker- Cheng, Qian Zhang, Wenyu Liu, and Chang Huang. Maptr:
nel transformer. arXiv preprint arXiv:2206.04584, 2022. 1, Structured modeling and learning for online vectorized hd
2 map construction. In International Conference on Learning
Representations, 2023. 1, 3, 4, 5, 6
[5] Wenjie Ding, Limeng Qiao, Xi Qiu, and Chi Zhang. Pivot-
net: Vectorized pivot learning for end-to-end hd map con- [19] Bencheng Liao, Shaoyu Chen, Yunchi Zhang, Bo Jiang, Qian
struction. In Proceedings of the IEEE/CVF International Zhang, Wenyu Liu, Chang Huang, and Xinggang Wang.
Conference on Computer Vision, pages 3672–3682, 2023. 3 Maptrv2: An end-to-end framework for online vectorized hd
[6] Wenjie Gao, Jiawei Fu, Haodong Jing, and Nanning Zheng. map construction. arXiv preprint arXiv:2308.05736, 2023.
Complementing onboard sensors with satellite map: A 3, 5, 6
new perspective for hd map construction. arXiv preprint [20] Xuewu Lin, Tianwei Lin, Zixiang Pei, Lichao Huang, and
arXiv:2308.15427, 2023. 3 Zhizhong Su. Sparse4d: Multi-view 3d object detec-
[7] Xunjiang Gu, Guanyu Song, Igor Gilitschenski, Marco tion with sparse spatial-temporal fusion. arXiv preprint
Pavone, and Boris Ivanovic. Accelerating online mapping arXiv:2211.10581, 2022. 1, 3
and behavior prediction via direct bev feature attention. [21] Xuewu Lin, Tianwei Lin, Zixiang Pei, Lichao Huang, and
arXiv preprint arXiv:2407.06683, 2024. 3 Zhizhong Su. Sparse4d v2: Recurrent temporal fusion with
[8] Chunrui Han, Jianjian Sun, Zheng Ge, Jinrong Yang, Run- sparse model. arXiv preprint arXiv:2305.14018, 2023. 1, 3
pei Dong, Hongyu Zhou, Weixin Mao, Yuang Peng, and [22] Xuewu Lin, Zixiang Pei, Tianwei Lin, Lichao Huang, and
Xiangyu Zhang. Exploring recurrent long-term tempo- Zhizhong Su. Sparse4d v3: Advancing end-to-end 3d detec-
ral fusion for multi-view 3d perception. arXiv preprint tion and tracking. arXiv preprint arXiv:2311.11722, 2023.
arXiv:2303.05970, 2023. 3 1
[9] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. [23] Xiaolu Liu, Song Wang, Wentong Li, Ruizi Yang, Junbo
Deep residual learning for image recognition. In Proceed- Chen, and Jianke Zhu. Mgmap: Mask-guided learning for
ings of the IEEE conference on computer vision and pattern online vectorized hd map construction. In Proceedings of
recognition, pages 770–778, 2016. 5 the IEEE/CVF Conference on Computer Vision and Pattern
[10] Haotian Hu, Fanyi Wang, Yaonong Wang, Laifeng Hu, Jing- Recognition, pages 14812–14821, 2024. 3, 6
wei Xu, and Zhiwang Zhang. Admap: Anti-disturbance [24] Yingfei Liu, Junjie Yan, Fan Jia, Shuailin Li, Aqi Gao, Tian-
framework for reconstructing online vectorized hd map. cai Wang, and Xiangyu Zhang. Petrv2: A unified framework
arXiv preprint arXiv:2401.13172, 2024. 3, 6 for 3d perception from multi-camera images. In Proceedings
[11] Junjie Huang and Guan Huang. Bevdet4d: Exploit tempo- of the IEEE/CVF International Conference on Computer Vi-
ral cues in multi-camera 3d object detection. arXiv preprint sion, pages 3262–3272, 2023. 1, 3
arXiv:2203.17054, 2022. 1, 3 [25] Yicheng Liu, Tianyuan Yuan, Yue Wang, Yilun Wang, and
[12] Junjie Huang and Guan Huang. Bevpoolv2: A cutting-edge Hang Zhao. Vectormapnet: End-to-end vectorized hd map
implementation of bevdet toward deployment. arXiv preprint learning. In International Conference on Machine Learning,
arXiv:2211.17111, 2022. 5 pages 22352–22369. PMLR, 2023. 1, 3, 5
[13] Qi Li, Yue Wang, Yilun Wang, and Hang Zhao. Hdmapnet: [26] Zihao Liu, Xiaoyu Zhang, Guangwei Liu, Ji Zhao,
An online hd map construction and evaluation framework. In and Ningyi Xu. Leveraging enhanced queries of point
2022 International Conference on Robotics and Automation sets for vectorized map construction. arXiv preprint
(ICRA), pages 4628–4634. IEEE, 2022. 1, 2, 5 arXiv:2402.17430, 2024. 3, 6
9
[27] Katie Z Luo, Xinshuo Weng, Yan Wang, Shuang Wu, Jie range vectorized hd map construction. arXiv preprint
Li, Kilian Q Weinberger, Yue Wang, and Marco Pavone. arXiv:2310.13378, 2023. 3
Augmenting lane perception and topology understanding [40] Tianyuan Yuan, Yicheng Liu, Yue Wang, Yilun Wang, and
with standard definition navigation maps. arXiv preprint Hang Zhao. Streammapnet: Streaming mapping network for
arXiv:2311.04079, 2023. 3 vectorized online hd map construction. In Proceedings of the
[28] Bowen Pan, Jiankai Sun, Ho Yin Tiga Leung, Alex Ando- IEEE/CVF Winter Conference on Applications of Computer
nian, and Bolei Zhou. Cross-view semantic segmentation Vision, pages 7356–7365, 2024. 3, 6
for sensing surroundings. IEEE Robotics and Automation [41] Zhixin Zhang, Yiyuan Zhang, Xiaohan Ding, Fusheng Jin,
Letters, 5(3):4867–4873, 2020. 1, 2 and Xiangyu Yue. Online vectorized hd map construction
[29] Jinhyung Park, Chenfeng Xu, Shijia Yang, Kurt Keutzer, using geometry. arXiv preprint arXiv:2312.03341, 2023. 3
Kris Kitani, Masayoshi Tomizuka, and Wei Zhan. Time will [42] Yi Zhou, Hui Zhang, Jiaqian Yu, Yifan Yang, Sangil Jung,
tell: New outlooks and a baseline for temporal multi-view 3d Seung-In Park, and ByungIn Yoo. Himap: Hybrid represen-
object detection. arXiv preprint arXiv:2210.02443, 2022. 1, tation learning for end-to-end vectorized hd map construc-
3 tion. In Proceedings of the IEEE/CVF Conference on Com-
[30] Limeng Qiao, Wenjie Ding, Xi Qiu, and Chi Zhang. End- puter Vision and Pattern Recognition, pages 15396–15406,
to-end vectorized hd-map construction with piecewise bezier 2024. 3, 6
curve. In Proceedings of the IEEE/CVF Conference on Com- [43] Xi Zhu, Xiya Cao, Zhiwei Dong, Caifa Zhou, Qiangbo Liu,
puter Vision and Pattern Recognition, pages 13218–13228, Wei Li, and Yongliang Wang. Nemo: Neural map growing
2023. 3 system for spatiotemporal fusion in bird’s-eye-view and bdd-
[31] Rémy Sun, Li Yang, Diane Lingrand, and Frédéric Precioso. map benchmark. arXiv preprint arXiv:2306.04540, 2023. 3
Mind the map! accounting for existing map information
when estimating online hdmaps from sensor data. arXiv
preprint arXiv:2311.10517, 2023. 3
[32] Shuo Wang, Fan Jia, Yingfei Liu, Yucheng Zhao, Zehui
Chen, Tiancai Wang, Chi Zhang, Xiangyu Zhang, and Feng
Zhao. Stream query denoising for vectorized hd map con-
struction. arXiv preprint arXiv:2401.09112, 2024. 3, 6, 7
[33] Shihao Wang, Yingfei Liu, Tiancai Wang, Ying Li, and Xi-
angyu Zhang. Exploring object-centric temporal modeling
for efficient multi-view 3d object detection. arXiv preprint
arXiv:2303.11926, 2023. 1, 3, 7
[34] Benjamin Wilson, William Qi, Tanmay Agarwal, John
Lambert, Jagjeet Singh, Siddhesh Khandelwal, Bowen
Pan, Ratnesh Kumar, Andrew Hartnett, Jhony Kaesemodel
Pontes, et al. Argoverse 2: Next generation datasets for
self-driving perception and forecasting. arXiv preprint
arXiv:2301.00493, 2023. 5
[35] Xuan Xiong, Yicheng Liu, Tianyuan Yuan, Yue Wang, Yilun
Wang, and Hang Zhao. Neural map prior for autonomous
driving. In Proceedings of the IEEE/CVF Conference on
Computer Vision and Pattern Recognition, pages 17535–
17544, 2023. 3
[36] Zhenhua Xu, Kenneth KY Wong, and Hengshuang Zhao.
Insightmapper: A closer look at inner-instance informa-
tion for vectorized high-definition mapping. arXiv preprint
arXiv:2308.08543, 2023. 3
[37] Chenyu Yang, Yuntao Chen, Hao Tian, Chenxin Tao, Xizhou
Zhu, Zhaoxiang Zhang, Gao Huang, Hongyang Li, Yu Qiao,
Lewei Lu, et al. Bevformer v2: Adapting modern image
backbones to bird’s-eye-view recognition via perspective su-
pervision. In Proceedings of the IEEE/CVF Conference on
Computer Vision and Pattern Recognition, pages 17830–
17839, 2023. 3
[38] Jiawei Yao, Xiaochao Pan, Tong Wu, and Xiaofeng Zhang.
Building lane-level maps from aerial images. arXiv preprint
arXiv:2312.13449, 2023. 3
[39] Jingyi Yu, Zizhao Zhang, Shengfu Xia, and Jizhang Sang.
Scalablemap: Scalable map learning for online long-
10