0% found this document useful (0 votes)
9 views13 pages

PriorMapNet: Optimizing HD Map Construction

PriorMapNet is a novel framework designed to enhance online vectorized High-Definition (HD) map construction for autonomous driving by integrating position and structure priors into its encoder and decoder. The proposed PPS-Decoder and PF-Encoder improve matching stability and reduce learning difficulty, achieving state-of-the-art performance on nuScenes and Argoverse2 datasets. This method addresses the instability issues found in mainstream approaches by utilizing prior reference points, leading to more accurate and efficient map construction.

Uploaded by

wnqnxh
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
9 views13 pages

PriorMapNet: Optimizing HD Map Construction

PriorMapNet is a novel framework designed to enhance online vectorized High-Definition (HD) map construction for autonomous driving by integrating position and structure priors into its encoder and decoder. The proposed PPS-Decoder and PF-Encoder improve matching stability and reduce learning difficulty, achieving state-of-the-art performance on nuScenes and Argoverse2 datasets. This method addresses the instability issues found in mainstream approaches by utilizing prior reference points, leading to more accurate and efficient map construction.

Uploaded by

wnqnxh
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

PriorMapNet: Enhancing Online Vectorized HD Map Construction with Priors

Rongxuan Wang 1† Xin Lu 2 Xiaoyang Liu2 Xiaoyi Zou2 Tongyi Cao2 Ying Li∗1
1 2
Beijing Institute of Technology [Link]
rongxuan_wang@[Link] [Link]@[Link]
arXiv:2408.08802v2 [[Link]] 20 Aug 2024

Abstract MapTRv2 Baseline PriorMapNet Queris with Priors in PriorMapNet


40% 60%
Online vectorized High-Definition (HD) map construction is
crucial for subsequent prediction and planning tasks in au- 40%
tonomous driving. Following MapTR paradigm, recent works
20%
have made noteworthy achievements. However, reference 20%
points are randomly initialized in mainstream methods, lead-
ing to unstable matching between predictions and ground 0%
0%
truth. To address this issue, we introduce PriorMapNet to en- ut 2nd 3rd 4th 5th 6th 4 8 12 16 20 24
hance online vectorized HD map construction with priors. We Decoder Layer Epoch
propose the PPS-Decoder, which provides reference points (a) on nuScenes val set (b) during training
with position and structure priors. Fitted from the map ele-
ments in the dataset, prior reference points lower the learn-
ing difficulty and achieve stable matching. Furthermore, we Figure 1: Comparison of the unstable matching scores, the
propose the PF-Encoder to enhance the image-to-BEV trans- lower, the better. (a) and (b) denote the unstable matching
formation with BEV feature priors. Besides, we propose the scores during validation and training, respectively. u means
DMD cross-attention, which decouples cross-attention along the percentage of queries whose GT match changed com-
multi-scale and multi-sample respectively to improve effi- pared with the previous decoder layer, and ut means the
ciency. Our proposed PriorMapNet achieves state-of-the-art
performance in the online vectorized HD map construction
percentage of final output queries whose GT match changed
task on nuScenes and Argoverse2 datasets. The code will be compared with the first decoder layer. “Queris with Priors”
released publicly soon. denote the queries corresponding to Prior Reference Points.

1 Introduction MapTR (Liao et al. 2022) and MapTRv2 (Liao et al. 2023)
High-Definition (HD) map is integral to autonomous driv- design an instance-point level hierarchical query embed-
ing, offering detailed information on critical elements such ding scheme and have demonstrated promising outcomes in
as road boundaries, traffic lanes, and pedestrian crossings (Li constructing vectorized HD maps. The mainstream methods
et al. 2022b; Liao et al. 2022). This detailed informa- proposed later follow this pipeline, with improvements fo-
tion is crucial for subsequent tasks like trajectory forecast- cusing on enhancing interactions between queries and inte-
ing (Liang et al. 2020; Zhou et al. 2022) and path plan- grating external features (Xu, Wong, and Zhao 2023; Liu
ning (Hu et al. 2023). Traditionally, HD map has been et al. 2024a; Zhou et al. 2024).
constructed using offline SLAM-based methods, which are In these methods, queries learn the position and structure
time-consuming and do not scale effectively with the rapid of map elements and are matched with the ground truth (GT)
updates of urban environments and road networks. To ad- during training. However, the Hungarian algorithm used
dress these challenges, there is growing interest in online HD for matching is sensitive to small changes in the cost ma-
map construction methods that use vehicle-mounted sensors trix, which leads to unstable matching (Li et al. 2022a). To
to generate maps in real-time. Early approaches (Li et al. quantify the unstability of matching, we define the unstable
2022d; Peng et al. 2023) focus on semantic segmentation in matching score u following Stable-DINO (Liu et al. 2023a),
bird’s eye view (BEV). However, these methods primarily representing the percentage of queries whose GT match
predict rasterized maps, which lack the vectorized map in- changed compared with the previous decoder layer. We also
formation for autonomous driving tasks. measure the total unstable matching score ut , which repre-
Following DETR (Carion et al. 2020) paradigm, recent sents the percentage of final output queries whose GT match
advancements have introduced end-to-end learning frame- changed compared with the first decoder layer. As shown in
works aimed at directly predicting vectorized instances. Fig. 1, MapTRv2 exhibits unstable matching throughout the
† Work done during the internship at [Link]. process of training and validation.
* Corresponding author. Why the matching is unstable? The training process
MapTRv2 Ours
of PriorMapNet. In summary, our contributions are:
• We introduce a novel prior-based framework for online
HD map construction by integrating feature, position,
Abstract and structure priors into encoder and decoder.
“good anchors” Priors

Unstable Stable
• We propose the DMD cross-attention, which decouples
Matching Matching cross-attention along multi-scale and multi-sample re-
spectively to improve efficiency.
Randomly Initialized Reference Points Offline Dataset Map
• We achieve SOTA performance in online vectorized
Ground Truth
Reference Points with Priors Elements Clusters HD map construction on the nuScenes and Argoverse2
datasets, demonstrating both high performance and gen-
Figure 2: Comparison of the matching of MapTRv2 and our eralization capability.
proposed method. Reference points with position and struc-
ture priors achieve stable matching. 2 Related Work
2.1 Online Vectorized HD Map Construction
of DETR-like models has two stages: learning “good an- Unlike traditional offline HD map construction methods, re-
chors” (stage I) and learning relative offsets (stage II) (Li cent studies use vehicle-mounted sensors to construct online
et al. 2022a). In mainstream methods, queries consist of HD map. Early methods (Philion and Fidler 2020; Li et al.
content embeddings and position embeddings. Position em- 2022d; Liu et al. 2022b) tackle map construction as a seg-
beddings generate reference points for sampling (related to mentation task, predicting rasterized maps in BEV space.
stage I), and content embeddings generate sampling offsets HDMapNet (Li et al. 2022b) further converts these raster-
and attention weights (related to stage II). Position embed- ized maps into vectorized maps through post-processing.
dings are learnable and initialized randomly, which leads VectorMapNet (Liu et al. 2023b) introduces the first
to reference points that are distributed without any spe- end-to-end vectorized map model, using a DETR (Car-
cific structure. In contrast, the vectorized HD map consists ion et al. 2020) decoder to detect map elements and opti-
of map elements like polylines or polygons connected in mizing results with an auto-regressive transformer. Subse-
an ordered sequence, with distinct positional distributions quently, MapTR (Liao et al. 2022) and MapTRv2 (Liao et al.
and geometric patterns. As shown in Fig. 2, matching these 2023) design a one-stage map construction paradigm with an
structured map elements with randomly distributed refer- instance-point level hierarchical query embedding scheme.
ence points is challenging and results in unstable matching. The mainstream methods proposed later follow this pipeline,
To solve this problem, we propose the Decoder with Prior with improvements focusing on enhancing interactions of
Position and Structure (PPS-Decoder). By fitting the distri- queries and external features. InsMapper (Xu, Wong, and
bution of map elements in datasets through clustering and Zhao 2023) and HIMap (Zhou et al. 2024) further explore
abstracting these distributions as priors, the reference points the correlation between instances and points and improve the
are enhanced to better match the positional and structural interaction within queries. MapQR (Liu et al. 2024b) implic-
features of map elements. As demonstrated in Tab. 4, prior- itly encodes point-level queries within instance-level queries
aware queries improve both accuracy and matching stability and embeds query positions like Conditional DETR (Meng
by reducing the difficulty of learning “good anchor”. et al. 2021). Despite the above developments, these meth-
In essence, the prior is an effective initialization method, ods randomly initialize reference points, resulting in unsta-
reducing the learning difficulty for the model. To leverage ble matching. To address this issue, our PriorMapNet intro-
this approach, we introduce the Encoder with Prior Fea- duces priors to enhance matching stability.
ture (PF-Encoder). PF-Encoder transforms image features
into initialized BEV features, which are utilized as BEV 2.2 Priors for HD Map Construction
query priors and optimized in the encoder. Discriminative Priors provide effective initialization for map construction
Loss is introduced to better aggregate map elements em- and reduce the difficulty of model learning. We categorize
beddings. Besides, BEV features are downsampled to multi- priors into two types: semantic priors and positional and
scale, bringing computational complexity. To enhance effi- structural priors. For prior semantics, MGMap (Liu et al.
ciency, we propose the Decoupled Multi-Scale Deformable 2024a) proposes Mask-Active Instance (MAI), which learns
Cross-Attention (DMD cross-attention), which decouples map instance segmentation results and provides semantic
cross-attention along multi-scale and multi-sample respec- priors for instance queries. Bi-Mapper (Li et al. 2023a) de-
tively. The combination of the PF-Encoder, PPS-Decoder, signs a two-stream model, using priors from global and local
and DMD cross-attention forms our proposed PriorMapNet. perspectives to enhance semantic map learning. For prior po-
Extensive experiments are conducted to prove our supe- sition and structure, Topo2D (Li et al. 2024a) uses 2D lane
riority. We achieve state-of-the-art (SOTA) performance in detection results as priors to initialize queries. SMERF (Luo
online vectorized HD map construction on nuScenes (Cae- et al. 2023) and P-MapNet (Jiang et al. 2024) introduce Stan-
sar et al. 2020) and Argoverse2 (Wilson et al. 2023) datasets. dard Map (SDMap) as position and structure priors for map
Furthermore, experiments conducted under various settings construction. However, the above methods rely on additional
demonstrate the robustness and generalization capabilities modules, increasing computational complexity. In contrast,
Backbone PF-Encoder PPS-Decoder Prediction Output

Initial
Instance
BEV Feature Query

+
Point
Query … …
Learnable
Multi-view Images Prior Ref.
Prior BEV Queries Hierarchical Ref.
Query
Reference Points
with Priors

Self-Attention

DMD Cross-Attention
Down Sampling
Dis.
×N
Loss Multi-Scale � Road Boundary
Image Backbone BEV Feature Refined Refined � Pedestrian Crossing
+ FPN Hierarchical Query Reference Points � Lane Divider

Figure 3: The overview of our proposed PriorMapNet. Given multi-view images as input, the output is a set of map elements.
PriorMapNet consists of three modules: the backbone, the PF-Encoder and the PPS-Decoder. The backbone extracts image
features by using the ResNet and a FPN neck. The PF-Encoder transforms image features into BEV and downsamples it to
multiple scales. The PPS-Decoder predicts map elements through Transformer, and reference points with priors are used for
stable matching. In the cross-attention layer, the DMD cross-attention is used to achieve efficiency.

PriorMapNet uses offline clustered map elements as position a FPN (Lin et al. 2017) neck. The PF-Encoder transforms
and structure priors, improving performance without addi- images features to BEV features FBEV ∈ RH×W×C and
tional computational consumption. downsamples it to multiple scales, as described in Section
3.3. The PPS-Decoder predicts map elements through trans-
2.3 Image-to-BEV Encoder for Map Construction former, and reference points with priors are used for stable
Map construction usually relies on the BEV feature, which matching, as detailed in Section 3.2. In the cross-attention
is transformed from images by the encoder. There are two layer, we introduce the DMD cross-attention to achieve effi-
types of encoders: bottom-up and top-down. Bottom-up en- ciency, as described in Section 3.4. We start by detailing the
coders (Philion and Fidler 2020; Huang et al. 2021; Li et al. PPS-Decoder, which is the core of our method.
2022c, 2023b) lift images to 3D and use voxel pooling to
generate BEV features. Top-down encoders (Wang et al. 3.2 Decoder with Prior Position and Structure
2022; Li et al. 2022d; Yang et al. 2023; Chen et al. 2022) The pipeline of our PPS-Decoder is shown in Fig. 4c. Com-
generate BEV queries containing 3D information and ex- pared with MapTRv2, which randomly initializes reference
tract image features to BEV queries with the transformer. points, and MGMap, which only provides semantic priors
However, since queries are randomly initialized, the single- without position information, the PPS-Decoder enhances
layer encoder results in low accuracy (Liao et al. 2022), reference points with position and structure priors, providing
and the multi-layer encoder brings more computational com- “good anchor” to improve accuracy and matching stability.
plexity (Liu et al. 2024b; Li et al. 2024b). To overcome these The PPS-Decoder contains several cascaded decoder lay-
limitations, we enhance BEV queries with prior features. ers to refine the hierarchical queries and reference points
iteratively. Hierarchical queries consist of instance-level
3 Method queries qins ∈ RNI ×C and point-level queries qpts ∈
RNP ×C , which are combined through broadcasting:
3.1 Overview
Fig. 3 shows the overall pipeline of our method. Given Nc q = qins + qpts , q ∈ RNI ×NP ×C . (1)
multi-view images {Ii }N c
i=1 as input, the output is a set of
Nm Reference points are initialized with prior position and
Nm map elements {Mi }i=1 . Each map element is defined as structure. To fit the distribution of map elements in the
Np
a class label c and an ordered point sequence {(xi , yi )}i=1 , dataset, we use K-Means to cluster map elements and ab-
where Np is the number of points in each map element. stract the position information of the first Npri elements,
Based on MapTRv2 (Liao et al. 2023), our method con- as shown in Fig. 2. Clustering and abstraction are done of-
sists of three modules: the backbone, the PF-Encoder and fline, ensuring no additional computational burden during
the PPS-Decoder. The backbone extracts multi-scale im- inference. During training and inference, some reference
ages features {Fimgi
}N c
i=1 by ResNet (He et al. 2016) and points obtain the fitted position and structure priors (called
+ + +
Multi-Scale Deformable Decoupled Multi-Scale
Deformable Cross-Attention Cross-Attention Deformable Cross-Attention
V Q Ref. V Q Ref. V Q Ref.

+ + +
BEV
Feature Decoupled Self-Attention Multi-Scale (Decoupled) Self-Attention Multi-Scale Decoupled Self-Attention
BEV Feature BEV Feature
V K Q V K Q V K Q

Position
+ + MAI
+ + + + Embedding

+
Linear

… + … … …
Hierarchical Reference Learnable Instance Point Hierarchical Reference Learnable Hierarchical Learnable
Prior Ref.
Query Points Query Pos. Query Query Query Points Query Pos. Query Ref.

(a) MapTRv2 Decoder (b) MGMap Decoder (c) Our PPS-Decoder

Figure 4: Comparison of the decoder of MapTRv2, MGMap and our proposed PriorMapNet. For simplicity, we only show the
first layer in the transformer decoder. (a) MapTRv2 uses randomly initialized learnable query positions for all layers without any
adaptation, which brings unstable matching results. (b) MGMap adds Mask-Activated Instance to provide semantic priors, but
lacks position information. In contrast, (c) PriorMapNet enhances reference points with priors, which achieves stable matching.

Prior Reference Points, Rpri ∈ RNpri ×NP ×2 ), while the rest 3.3 Encoder with Prior Feature
of the reference points are still from learnable parameters PF-Encoder enhances the image-to-BEV transformation
(called Learnable Reference Points, Rlrn ∈ RNlrn ×NP ×2 ). with BEV feature priors. Built on the foundation of top-
The combined set of reference points is denoted as R = down encoders, such as BEVFormer (Li et al. 2022d) and
{Rpri , Rlrn }, where the total number of instance queries is GKT (Chen et al. 2022), PF-Encoder leverage BEV fea-
NI = Npri + Nlrn . tures as queries to extract relevant image features via cross-
To embed query position, reference points are encoded attentions.
with sinusoidal positions following DAB-DETR (Liu et al.
We first utilize LSS (Philion and Fidler 2020) to transform
2022a). Query position embedding is achieved as follows:
image features into initialized BEV features, which are then
qpos = Linear(PE(R)), RNI ×NP ×2 → RNI ×NP ×C , (2) used as BEV query priors, optimized in a single-layer BEV-
Former (Li et al. 2022d) encoder. Following MGMap (Liu
where PE(·) generates sinusoidal embeddings based on ref- et al. 2024a), BEV features are downsampled to multi-scale
erence points coordinates (Vaswani et al. 2017). The param- with an EML neck.
eters of linear layers are not shared across decoder layers. For queries to better aggregate features from the same
PE(·) is calculated separately on coordinates, and position map element, it is necessary to assimilate the embeddings of
embeddings are concatenated along feature channels: the same instance and distinguish the embeddings of differ-
PE(R) = PE(xr , yr ) = Cat(PE(xr ), PE(yr )). (3) ent instances. Therefore, we introduce Discriminative Loss
of map elements (Neven et al. 2018) to bring the same in-
Reference points and position embeddings are updated stance closer and separate different instances further:
across the PPS-Decoder layers. In each layer, self-attention
and cross-attention mechanisms use the following inputs for 
1
PK 1
Ppn 2
queries, keys, values, and reference points: Lvar = K k=1 pn l=1 [∥µk − el ∥ − δv ]+ ,
(
Self-Attn : Q = q + qpos , K = q + qpos , V = q, Ldist = 1
PK PK
[δd − ∥µi − µj ∥]+ ,
2
K(K−1) i=1 j=1,i̸=j
Cross-Attn : Q = q + qpos , V = FBEV , R = R. (5)
(4) where Lvar pulls the embeddings ei of K map elements to-
The Prior Reference Points fit the position and structure ward their respective means µn , and Ldist pushes away the
distribution of the map elements in the dataset, which helps mean embeddings of different map elements. pn represents
queries concentrate on learning the offsets from reference the number of grids of map elements. ∥·∥ is the L2 distance
points. In addition, we maintain Learnable Reference Points and [x]+ = max(0, x). δv and δd are the borders of the vari-
to capture and represent map elements that deviate from typ- ance and distance loss. The total Discriminative Loss is de-
ical position and structure patterns. The self-attention en- fined as Ldis = λ1 Lvar + λ2 Ldist .
ables interaction between Prior Reference Points and Learn- In the cross-attention layer of the PPS-Decoder, queries
able Reference Points, reducing redundant detections and weighted sample BEV features. PF-Encoder enables the
improving overall detection accuracy. queries to effectively aggregate features associated with the
Scales Samples Scales Samples with 2D vectorized map elements as ground truth. Argov-
erse 2 is designed for perception and prediction studies in
> autonomous driving, containing 1000 scenes of 15 seconds

>

>
>
>
>
> + each. 3D vectorized map elements captured by seven multi-
view cameras are provided as ground truth.
Following the previous studies (Li et al. 2022b; Liao et al.
(a) Vanilla MSDA (b) Our DMD cross-attention 2022), we evaluate performance across three categories of
map elements: lane dividers, pedestrian crossings, and road
Figure 5: Comparison of the vanilla MSDA and our pro- boundaries. The performance of PriorMapNet is assessed
posed DMD cross-attention. DMD cross-attention performs using the Average Precision (AP) metric, where a predic-
cross-attention along multi-scale and multi-sample respec- tion is considered a True Positive if the Chamfer Distance
tively to achieve efficiency. between the prediction and its ground truth is within thresh-
olds of 0.5, 1.0, and 1.5 meters.
same map element while distinguishing between different
4.2 Implementation Details
map instances, improving the accuracy of map construction.
Our model is trained on 8 NVIDIA A100 GPUs with a batch
3.4 Decoupled Multi-Scale Deformable Attention size of 8 × 3. Unless otherwise specified, the number of
To address the computational complexity of multi-scale de- training epochs is 24 on nuScenes and 6 on Argoverse 2.
formable cross-attention (MSDA), we propose the DMD The BEV range is [-30m, 30m] along the longitudinal axis
cross-attention mechanism to decouple cross-attention along and [-15m, 15m] along the lateral axis, with a feature size
multi-scale and multi-sample, as shown in Fig. 5b. H × W of 200×100. The number of instance queries NI ,
In vanilla MSDA (Zhu et al. 2020), each query interacts prior queries Npri and point queries NP are set to 50, 9 and
with M -scale BEV features and N points are sampled at 20, respectively. Both λ1 and λ2 are set to 1. δv is 0.5 and
each scale, whose computation complexity is O(M × N ): δd is 3. Other settings keep in line with MapTRv2. The sup-
plementary material shows more implementation details and
Nh M X
N
X X ′ ablation studies on hyperparameters.
MSDA(Q, V, R) = Wi Aijk ·Wi Vj (R+Oijk ),
i=1 j=1 k=1
4.3 Main Results
(6)
where Nh is the number of attention heads. Aijk ∈ [0, 1] Results on nuScenes. We report quantitative results on
and Oijk ∈ R2 are attention weight and sampling offset re- nuScenes val set in Tab. 1. Under camera modality, Pri-
spectively, which are generated from Q. Aijk ∈ [0, 1] is nor- orMapNet surpasses previous SOTA methods and achieves
PM PN
malized by j=1 k=1 Aijk = 1. Wi ∈ RC×(C/Nh ) and 6.2% mAP improvement compared with our baseline Map-
′ TRv2. On one RTX 4090 GPU, PriorMapNet infers at 13.9
Wi ∈ R(C/Nh )×C are learnable weights. frames per second (FPS). Additionally, under the camera
To improve efficiency, the DMD cross-attention mecha- and lidar fusion modality, PriorMapNet reaches 72.9% mAP
nism decouples the vanilla MSDA process into two stages: and 7.5 FPS, demonstrating strong generalization capabili-
(
Q1 = Linear1 (MSDA1 (Q, V, R)), ties. Qualitative results are shown in Fig. 6, further illustrat-
(7) ing that PriorMapNet achieves improved results. More qual-
Qoutput = Q1 + Linear2 (MSDA2 (Q1 , V1 , R)), itative results are shown in the supplementary material.
where MSDA1 (·) and MSDA2 (·) denote MSDA(·) at N = Results on Argoverse 2. We report quantitative results on
1 and M = 1 respectively. V1 is the largest scale BEV fea- Argoverse 2 val set in Tab. 2. Argoverse 2 provides 3D map
ture. The multi-scale stage performs cross-attention across annotations, allowing predictions for both 2D and 3D map
M scales and samples one point per scale. The multi-sample elements. PriorMapNet surpasses previous SOTA methods
stage uses the output from the multi-scale stage and focuses in both dimensions, achieving 72.0% mAP for 2D map ele-
on the largest scale feature to sample N points. DMD cross- ments and 69.9% mAP for 3D map elements with an infer-
attention reduces the computation complexity to O(M + N ) ence speed of 12.6 FPS. Experimental results demonstrate
and achieves higher performance than vanilla MSDA. the generalizability of our method.
Results on Enlarged BEV Range. We train and evaluate
4 Experiments models on enlarged BEV ranges on nuScenes val set as
shown in Tab. 3. The size of the BEV grid is maintained
4.1 Datasets and Metrics at [0.3m, 0.3m]. To verify the robustness of our method,
To validate the effectiveness of our proposed method Pri- we correspondingly enlarge the prior clustering and position
orMapNet, we evaluate it on the widely used nuScenes range of map elements. Other settings remain in line with the
dataset (Caesar et al. 2020) and Argoverse 2 dataset (Wil- original models. Experimental results demonstrate that Pri-
son et al. 2023) and compare it with the SOTA methods. orMapNet maintains superiority on enlarged BEV ranges.
The nuScenes dataset is a standard benchmark for on- Notably, under the range of 100 × 50m, our method outper-
line vectorized HD map construction, featuring 1000 driv- forms the SOTA method SQD-MapNet (Wang et al. 2024)
ing scenes captured by six multi-view cameras and LiDAR, which integrates stream strategy.
Method Modality Backbone Epoch APdiv APped APbou mAP FPS
MapTRv2 [arxiv23] C R50 24 60.5 60.5 61.8 60.9 16.7
MGMap* [CVPR24] C R50 24 65.0 61.8 67.5 64.8 -
HIMap [CVPR24] C R50 30 68.4 62.6 69.1 66.7 -
InsMapper [ECCV24] C R50 24 65.1 61.6 64.6 63.8 -
MapQR [ECCV24] C R50 24 68.7 63.4 67.7 66.4 16.2
PriorMapNet (Ours) C R50 24 69.0 64.0 68.2 67.1 13.9
MapTRv2 [arxiv23] C R50 110 68.3 68.1 69.7 68.7 16.7
MGMap [CVPR24] C R50 110 64.4 67.6 67.7 66.5 18.0
MapQR [ECCV24] C R50 110 74.4 70.1 73.2 72.6 16.2
PriorMapNet (Ours) C R50 110 73.2 71.5 73.3 72.7 13.9
HDMapNet [ICRA22] C&L EB0 & PP 30 29.6 16.3 46.7 31.0 0.6
MapTRv2 [arxiv23] C&L R50 & Sec 24 66.5 65.6 74.8 69.0 7.8
MGMap [CVPR24] C&L R50 & Sec 24 71.1 67.7 76.2 71.7 7.5
PriorMapNet (Ours) C&L R50 & Sec 24 72.4 70.1 76.2 72.9 7.5

Table 1: Quantitative evaluation on nuScenes val set. “C” and “L” respectively refer to multi-view cameras and LiDAR inputs.
“R50”, “EB0”, “PP” and “Sec” denote ResNet50 (He et al. 2016), EfficientNet-B0 (Tan and Le 2019), PointPillars (Lang et al.
2019) and SECOND (Yan, Mao, and Li 2018) respectively. “MGMap*” means MGMap based on MapTRv2. FPS is tested on
a single RTX 4090 GPU for fair comparison. “-” means the corresponding results are not available.

Dim. Method APdiv APped APbou mAP FPS Range Method APdiv APped APbou mAP
MapTRv2 71.5 63.6 67.4 67.5 15.2 MapTRv2 61.9 56.9 61.4 60.0
90×30m
HIMap 69.5 69.0 70.3 69.6 - PriorMapNet 66.5 62.0 66.1 64.9
2
MapQR 72.3 64.3 68.1 68.2 14.2
MapTRv2 61.5 58.8 61.0 60.5
PriorMapNet 75.4 69.3 71.3 72.0 12.6 60×60m
PriorMapNet 66.6 63.8 66.5 65.6
MapTRv2 68.9 60.7 64.5 64.7 15.0
MapTRv2 62.1 55.3 61.4 59.6
HIMap 68.3 66.7 70.3 68.4 -
3 100×50m SQD-MapNet 65.5 67.0 59.5 64.0
MapQR 71.2 60.1 66.2 65.9 14.1
PriorMapNet 67.4 62.5 65.0 65.0
PriorMapNet 73.4 66.5 69.8 69.9 12.6
Table 3: Quantitative evaluation on nuScenes dataset with
Table 2: Quantitative evaluation on Argoverse 2 val set.
enlarged BEV ranges. PriorMapNet maintains superiority.
PriorMapNet reaches SOTA performance, validating its gen-
eralization ability. Since Argoverse 2 provides 3D map an-
notation, “Dim.” represents the dimension used to model PF PPS DMD APdiv APped APbou mAP ut
map elements. When the map dimension is 2, the height in-
formation of map elements is dropped, and when the map - - - 60.5 60.5 61.8 60.9 0.413
dimension is 3, the 3D map elements are predicted directly. ✓ - - 64.9 61.8 64.9 63.9 -
FPS is tested on a single RTX 4090 GPU for fair compari- - ✓ - 64.6 62.8 65.8 64.4 0.365
son. “-” means the corresponding results are not available. ✓ ✓ - 68.7 63.6 68.0 66.8 -
✓ ✓ ✓ 69.0 64.0 68.2 67.1 0.362

4.4 Ablation Study Table 4: Ablation study of our modules on nuScenes dataset.
We conduct ablation experiments on nuScenes to verify the “PF”, “PPS” and “DMD” denote PF-Encoder, PPS-Decoder
effectiveness of our proposed modules and their designs. All and DMD cross-attention, respectively. ut is the total unsta-
models are trained for 24 epochs, and the BEV range is ble matching score defined in Section 1.
60×30m. As shown in Tab. 4, starting from MapTRv2 as
our baseline, each module improves mAP. In addition, we
also report the total unstable matching score ut on the vali- cause semantic priors lack positional information crucial for
dation set, demonstrating that the prior position and structure vectorized map elements. We then integrate MAI with query
of map elements improve the stability of matching. position, but the mAP decreases as the semantic priors are
Ablation study of decoder query prior. We compare var- not well-suited for query position. In contrast, our PPS pro-
ious query priors under different BEV feature scales in vides prior position and structure, resulting in significant
Tab. 5. MAI in MGMap provides semantic priors for queries, performance improvements. Compared with directly using
but the improvement is limited. This limitation arises be- clustering results as priors, abstracted priors perform better.
Multi-view Images GT123456MapTRv2 Ours Multi-view Images GT123456MapTRv2 Ours

Figure 6: Qualitative results on nuScenes val set. We compare visualization results of PriorMapNet with MapTRv2 and
corresponding GTs. Models are trained for 24 epochs. The green area indicates that our method achieves more accurate results.

FBEV Query Prior APdiv APped APbou mAP FBEV Prior BEV Encoder Dis. Loss mAP
- 64.9 61.8 64.9 63.9 - LSS-d - 60.9
MAI 65.0 62.7 64.8 64.2 - BEVFormer-d - 57.3
Single MAI-pos 58.3 51.0 62.3 57.2
BEVFormer-d BEVFormer - 59.6
Scale PPS-clst 64.6 63.8 66.9 65.1
GKT-d BEVFormer - 61.9
PPS 65.5 62.8 68.5 65.6
LSS-d GKT - 62.4
- 67.5 61.7 65.8 65.0 LSS-d BEVFormer - 63.3
MAI 66.5 62.1 67.1 65.2 LSS-d BEVFormer ✓ 63.9
Multi MAI-pos 58.4 54.7 61.2 58.1
Scales PPS-clst 68.4 63.3 67.5 66.4 Table 6: Ablation study of the PF-Encoder design. “Dis.
PPS 68.7 63.6 68.0 66.8 Loss” denotes Discriminative Loss, and “-d” denotes aux-
iliary depth supervision of image features.
Table 5: Ablation study of decoder query prior. “MAI” de-
notes Mask-Activated Instance in MGMap, and “MAI-pos”
MSDA APdiv APped APbou mAP tc (ms)
denotes integrating MAI with query position. “PPS-clst” de-
notes directly using clustering results as priors. Vanilla 68.7 63.6 68.0 66.8 6.1
Parallel Connection 68.5 65.7 66.5 66.9 4.2
Sample-then-Scale 67.6 64.7 68.0 66.8 4.2
Ablation study of PF-Encoder design. We compare vari- Scale-then-Sample 69.0 64.0 68.2 67.1 4.2
ous BEV feature priors and encoders, as shown in Tab. 6.
All GKT and BEVFormer methods utilize a single-layer ar- Table 7: Ablation study of the DMD cross-attention de-
chitecture. To ensure a fair comparison, we incorporate aux- sign. “Sample-then-Scale” denotes performing MSDA along
iliary depth supervision across all experimental setups. LSS multi-sample then multi-scale, and “Scale-then-Sample” de-
without prior is the encoder of MapTRv2. The third row de- notes performing MSDA along reverse order. tc denotes the
notes a two-layer BEVFormer encoder to exclude the impact inference time of cross-attention.
of the number of encoder layers. The results show that LSS
serves as a good prior generator. Results also demonstrate
the effectiveness of Discriminative Loss. effectively, we propose the PF-Encoder which enhances the
Ablation study of DMD cross-attention design. We com- image-to-BEV transformation with BEV feature priors and
pare different decoupling designs of DMD cross-attention, leverages Discriminative Loss to improve the aggregation
including parallel and serial connection of two MSDA of map element embeddings. To reduce the computation
mechanisms, as shown in Tab. 7. In the serial connection, complexity, we propose the DMD cross-attention, which
we compare different performance orders of MSDA along performs cross-attention respectively along multi-scale and
multi-sample and multi-scale dimensions. Experimental re- multi-sample. Our proposed PriorMapNet achieves state-of-
sults demonstrate that our proposed DMD cross-attention the-art performance on nuScenes and Argoverse2 datasets.
improves efficiency and achieves superior performance. Limitations and Future Work. Despite our development
for online vectorized HD map construction, several limita-
5 Conclusion tions need to be addressed in future work. Firstly, our map
In this paper, we introduce PriorMapNet to enhance online element priors only incorporate positional information and
vectorized HD map construction with priors. To address the lack semantic information, which limits the interaction and
issue of unstable matching, we propose the PPS-Decoder, optimization of queries. Secondly, our method relies solely
which provides reference points with position and structure on single-frame sensor input, constraining the representation
priors clustered from the dataset. To embed BEV features of temporally and spatially continuous map elements.
References Li, Y.; Bao, H.; Ge, Z.; Yang, J.; Sun, J.; and Li, Z. 2022c.
Caesar, H.; Bankiti, V.; Lang, A. H.; Vora, S.; Liong, V. E.; Bevstereo: Enhancing depth estimation in multi-view 3d ob-
Xu, Q.; Krishnan, A.; Pan, Y.; Baldan, G.; and Beijbom, O. ject detection with dynamic temporal stereo. arXiv preprint
2020. nuscenes: A multimodal dataset for autonomous driv- arXiv:2209.10248.
ing. In Proceedings of the IEEE/CVF conference on com- Li, Y.; Ge, Z.; Yu, G.; Yang, J.; Wang, Z.; Shi, Y.; Sun,
puter vision and pattern recognition, 11621–11631. J.; and Li, Z. 2023b. Bevdepth: Acquisition of reliable
depth for multi-view 3d object detection. In Proceedings of
Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, the AAAI Conference on Artificial Intelligence, volume 37,
A.; and Zagoruyko, S. 2020. End-to-end object detection 1477–1485.
with transformers. In ECCV, 213–229. Springer.
Li, Z.; Wang, W.; Li, H.; Xie, E.; Sima, C.; Lu, T.; Qiao,
Chen, S.; Cheng, T.; Wang, X.; Meng, W.; Zhang, Q.; and Y.; and Dai, J. 2022d. Bevformer: Learning bird’s-eye-view
Liu, W. 2022. Efficient and robust 2d-to-bev representa- representation from multi-camera images via spatiotemporal
tion learning via geometry-guided kernel transformer. arXiv transformers. In European conference on computer vision,
preprint arXiv:2206.04584. 1–18. Springer.
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep resid- Liang, M.; Yang, B.; Hu, R.; Chen, Y.; Liao, R.; Feng, S.;
ual learning for image recognition. In Proceedings of the and Urtasun, R. 2020. Learning lane graph representations
IEEE conference on computer vision and pattern recogni- for motion forecasting. In Computer Vision–ECCV 2020:
tion, 770–778. 16th European Conference, Glasgow, UK, August 23–28,
Hu, Y.; Yang, J.; Chen, L.; Li, K.; Sima, C.; Zhu, X.; Chai, 2020, Proceedings, Part II 16, 541–556. Springer.
S.; Du, S.; Lin, T.; Wang, W.; et al. 2023. Planning-oriented Liao, B.; Chen, S.; Wang, X.; Cheng, T.; Zhang, Q.; Liu,
autonomous driving. In Proceedings of the IEEE/CVF W.; and Huang, C. 2022. MapTR: Structured Modeling and
Conference on Computer Vision and Pattern Recognition, Learning for Online Vectorized HD Map Construction. In
17853–17862. ICLR.
Huang, J.; Huang, G.; Zhu, Z.; Ye, Y.; and Du, D. 2021. Liao, B.; Chen, S.; Zhang, Y.; Jiang, B.; Zhang, Q.; Liu, W.;
Bevdet: High-performance multi-camera 3d object detection Huang, C.; and Wang, X. 2023. Maptrv2: An end-to-end
in bird-eye-view. arXiv preprint arXiv:2112.11790. framework for online vectorized hd map construction. arXiv
preprint arXiv:2308.05736.
Jiang, Z.; Zhu, Z.; Li, P.; Gao, H.-a.; Yuan, T.; Shi, Y.; Zhao,
H.; and Zhao, H. 2024. P-MapNet: Far-seeing map gener- Lin, T.-Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.;
ator enhanced by both SDMap and HDMap priors. arXiv and Belongie, S. 2017. Feature pyramid networks for ob-
preprint arXiv:2403.10521. ject detection. In Proceedings of the IEEE conference on
computer vision and pattern recognition, 2117–2125.
Lang, A. H.; Vora, S.; Caesar, H.; Zhou, L.; Yang, J.; and
Beijbom, O. 2019. Pointpillars: Fast encoders for object de- Liu, S.; Li, F.; Zhang, H.; Yang, X.; Qi, X.; Su, H.; Zhu, J.;
tection from point clouds. In Proceedings of the IEEE/CVF and Zhang, L. 2022a. DAB-DETR: Dynamic Anchor Boxes
Conference on Computer Vision and Pattern Recognition, are Better Queries for DETR. In International Conference
12697–12705. on Learning Representations.
Liu, S.; Ren, T.; Chen, J.; Zeng, Z.; Zhang, H.; Li, F.; Li,
Li, F.; Zhang, H.; Liu, S.; Guo, J.; Ni, L. M.; and Zhang,
H.; Huang, J.; Su, H.; Zhu, J.; et al. 2023a. Detection
L. 2022a. Dn-detr: Accelerate detr training by introducing
transformer with stable matching. In Proceedings of the
query denoising. In Proceedings of the IEEE/CVF Confer-
IEEE/CVF International Conference on Computer Vision,
ence on Computer Vision and Pattern Recognition, 13619–
6491–6500.
13627.
Liu, X.; Wang, S.; Li, W.; Yang, R.; Chen, J.; and Zhu, J.
Li, H.; Huang, Z.; Wang, Z.; Rong, W.; Wang, N.; and 2024a. Mgmap: Mask-guided learning for online vector-
Liu, S. 2024a. Enhancing 3D Lane Detection and Topol- ized hd map construction. In Proceedings of the IEEE/CVF
ogy Reasoning with 2D Lane Priors. arXiv preprint Conference on Computer Vision and Pattern Recognition,
arXiv:2406.03105. 14812–14821.
Li, Q.; Wang, Y.; Wang, Y.; and Zhao, H. 2022b. Hdmapnet: Liu, Y.; Yan, J.; Jia, F.; Li, S.; Gao, Q.; Wang, T.; Zhang,
An online hd map construction and evaluation framework. X.; and Sun, J. 2022b. Petrv2: A unified framework for
In ICRA, 4628–4634. IEEE. 3d perception from multi-camera images. arXiv preprint
Li, S.; Yang, K.; Shi, H.; Zhang, J.; Lin, J.; Teng, Z.; and arXiv:2206.01256.
Li, Z. 2023a. Bi-Mapper: Holistic BEV Semantic Mapping Liu, Y.; Yuan, T.; Wang, Y.; Wang, Y.; and Zhao, H. 2023b.
for Autonomous Driving. IEEE Robotics and Automation Vectormapnet: End-to-end vectorized hd map learning. In
Letters. ICML, 22352–22369. PMLR.
Li, T.; Jia, P.; Wang, B.; Chen, L.; JIANG, K.; Yan, J.; and Liu, Z.; Zhang, X.; Liu, G.; Zhao, J.; and Xu, N. 2024b.
Li, H. 2024b. LaneSegNet: Map Learning with Lane Seg- Leveraging Enhanced Queries of Point Sets for Vectorized
ment Perception for Autonomous Driving. In The Twelfth Map Construction. In European Conference on Computer
International Conference on Learning Representations. Vision.
Loshchilov, I.; and Hutter, F. 2016. Sgdr: Stochas- Zhou, Y.; Zhang, H.; Yu, J.; Yang, Y.; Jung, S.; Park, S.-I.;
tic gradient descent with warm restarts. arXiv preprint and Yoo, B. 2024. HIMap: HybrId Representation Learning
arXiv:1608.03983. for End-to-end Vectorized HD Map Construction. In Pro-
Loshchilov, I.; and Hutter, F. 2017. Decoupled weight decay ceedings of the IEEE/CVF Conference on Computer Vision
regularization. arXiv preprint arXiv:1711.05101. and Pattern Recognition, 15396–15406.
Luo, K. Z.; Weng, X.; Wang, Y.; Wu, S.; Li, J.; Weinberger, Zhou, Z.; Ye, L.; Wang, J.; Wu, K.; and Lu, K. 2022. Hivt:
K. Q.; Wang, Y.; and Pavone, M. 2023. Augmenting Lane Hierarchical vector transformer for multi-agent motion pre-
Perception and Topology Understanding with Standard Def- diction. In Proceedings of the IEEE/CVF Conference on
inition Navigation Maps. arXiv preprint arXiv:2311.04079. Computer Vision and Pattern Recognition, 8823–8833.
Meng, D.; Chen, X.; Fan, Z.; Zeng, G.; Li, H.; Yuan, Y.; Zhu, X.; Su, W.; Lu, L.; Li, B.; Wang, X.; and Dai, J. 2020.
Sun, L.; and Wang, J. 2021. Conditional detr for fast training Deformable detr: Deformable transformers for end-to-end
convergence. In Proceedings of the IEEE/CVF international object detection. arXiv preprint arXiv:2010.04159.
conference on computer vision, 3651–3660.
Neven, D.; De Brabandere, B.; Georgoulis, S.; Proesmans,
M.; and Van Gool, L. 2018. Towards end-to-end lane de-
tection: an instance segmentation approach. In 2018 IEEE
intelligent vehicles symposium (IV), 286–291. IEEE.
Peng, L.; Chen, Z.; Fu, Z.; Liang, P.; and Cheng, E. 2023.
BEVSegFormer: Bird’s Eye View Semantic Segmentation
From Arbitrary Camera Rigs. In WACV, 5935–5943.
Philion, J.; and Fidler, S. 2020. Lift, splat, shoot: Encoding
images from arbitrary camera rigs by implicitly unproject-
ing to 3d. In Computer Vision–ECCV 2020: 16th European
Conference, Glasgow, UK, August 23–28, 2020, Proceed-
ings, Part XIV 16, 194–210. Springer.
Tan, M.; and Le, Q. 2019. Efficientnet: Rethinking model
scaling for convolutional neural networks. In International
conference on machine learning, 6105–6114. PMLR.
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones,
L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017. At-
tention is all you need. Advances in neural information pro-
cessing systems, 30.
Wang, S.; Jia, F.; Liu, Y.; Zhao, Y.; Chen, Z.; Wang, T.;
Zhang, C.; Zhang, X.; and Zhao, F. 2024. Stream query de-
noising for vectorized hd map construction. arXiv preprint
arXiv:2401.09112.
Wang, Y.; Guizilini, V. C.; Zhang, T.; Wang, Y.; Zhao, H.;
and Solomon, J. 2022. Detr3d: 3d object detection from
multi-view images via 3d-to-2d queries. In Conference on
Robot Learning, 180–191. PMLR.
Wilson, B.; Qi, W.; Agarwal, T.; Lambert, J.; Singh, J.;
Khandelwal, S.; Pan, B.; Kumar, R.; Hartnett, A.; Pontes,
J. K.; et al. 2023. Argoverse 2: Next generation datasets
for self-driving perception and forecasting. arXiv preprint
arXiv:2301.00493.
Xu, Z.; Wong, K.-Y. K.; and Zhao, H. 2023. InsMapper: Ex-
ploring Inner-instance Information for Vectorized HD Map-
ping. arXiv preprint arXiv:2308.08543.
Yan, Y.; Mao, Y.; and Li, B. 2018. Second: Sparsely embed-
ded convolutional detection. Sensors, 18(10): 3337.
Yang, C.; Chen, Y.; Tian, H.; Tao, C.; Zhu, X.; Zhang, Z.;
Huang, G.; Li, H.; Qiao, Y.; Lu, L.; et al. 2023. BEVFormer
v2: Adapting Modern Image Backbones to Bird’s-Eye-View
Recognition via Perspective Supervision. In Proceedings of
the IEEE/CVF Conference on Computer Vision and Pattern
Recognition, 17830–17839.
Supplementary Material

In the supplementary material, we provide more details, ex- Npri APdiv APped APbou mAP
periments, and analysis of the proposed PriorMapNet, in-
cluding: 0 67.5 61.7 65.8 65.0
5 67.2 62.7 67.8 65.9
• More implementation details of our method; 9 68.7 63.6 68.0 66.8
• Additional ablation studies on hyperparameters; 10 67.3 63.8 68.8 66.7
• Difference with previous SOTA methods; 20 68.1 63.5 68.0 66.5
50 68.4 63.3 67.5 66.4
• More qualitative results and failure cases.

S1 More Implementation Details Table S1: Ablation study of the number of prior queries. Uti-
lizing 9 prior queries achieves the best performance.
Feature extraction and multi-modality fusion. Multi-
modality inputs consist of multi-view cameras and LiDAR
data. For backbones, we utilize ResNet50 (He et al. 2016) λ APdiv APped APbou mAP
for image features extraction and SECOND (Yan, Mao, and 0 64.3 60.9 64.8 63.3
Li 2018) for point cloud features extraction. The input im- 0.5 65.6 61.6 63.7 63.6
age size is 480×800. The voxel size for point cloud is set to 1 64.9 61.8 64.9 63.9
[0.1m, 0.1m, 0.2m], and the LiDAR BEV features are up- 2 65.1 60.7 64.6 63.5
sampled to align with camera BEV features. For BEV fea- 3 64.5 60.1 65.0 63.2
ture fusion, we concatenate camera and LiDAR BEV fea-
tures along feature channels and use a convolution layer to Table S2: Ablation study of the weight of Discriminative
fuse them. Loss. We set λ1 = λ2 = λ and setting λ = 1 achieves
Encoder and decoder. Following MapTRv2 (Liao et al. the best performance.
2023), we add an auxiliary depth loss for LSS to generate
BEV query priors effectively. The PPS-Decoder contains
6 layers to refine the outputs iteratively. In addition to 50 δd APdiv APped APbou mAP
instance-level queries, there are another 300 one-to-many 1 64.8 60.5 64.2 63.2
queries that are used to speed up convergence during train- 3 64.9 61.8 64.9 63.9
ing. Priors are not used for one-to-many queries. 6 66.0 61.2 63.9 63.7
Training. We use the AdamW optimizer (Loshchilov and 12 65.4 61.0 64.5 63.6
Hutter 2017) with an initial learning rate of 6 × 10−4 , and
apply the Cosine Annealing learning rate scheduler with a
linear warm-up phase (Loshchilov and Hutter 2016). Except Table S3: Ablation study of the borders of the variance and
for our Discriminative Loss, other training and loss settings distance loss. We set δd = 6δv and setting δd = 3 achieves
keep in line with our baseline MapTRv2. the best performance.

S2 Additional Ablation Studies


variance and distance loss. Following LaneNet (Neven et al.
Abalation study on the number of prior queries. We clus- 2018), we set δd = 6δv . To embed the BEV features effec-
ter 50 map elements and abstract the position information tively, we compare the results under different borders of the
of the first Npri elements, as shown in Fig. 2. To most ef- distance loss (i.e. δd ), as shown in Tab. S3. Experiments are
fectively utilize the prior position and structure, we compare conducted under λ = 1. Experimental results show that the
the results of different numbers of prior queries (i.e. Npri ), as best performance can be achieved when δd = 3.
shown in Tab. S1. Utilizing 9 prior queries achieves the best
performance. Too few prior queries cannot fully utilize the
prior, while too many prior queries limit the learning ability S3 Difference with Previous SOTA Methods
of learnable queries. We highlight the key distinctions between our PriorMap-
Abalation study on the loss weights. The Discriminative Net and previous SOTA methods, including MapTRv2 (Liao
Loss is defined as Ldis = λ1 Lvar + λ2 Ldist . Following et al. 2023), MGMap (Liu et al. 2024a), HIMap (Zhou
LaneNet (Neven et al. 2018), we set λ1 = λ2 = λ. In order et al. 2024), MapQR (Liu et al. 2024b) and InsMapper (Xu,
to obtain the most effective auxiliary discriminative super- Wong, and Zhao 2023). We compare the decoders in Fig. S1,
vision, we compare the results under different loss weights which are the core of these methods.
(i.e. λ), as shown in Tab. S2. Experimental results show the Motivation and model design. Our motivation and model
effectiveness of Discriminative Loss, and the best perfor- design diverge from those of previous SOTA methods. Pri-
mance can be achieved when λ = 1. orMapNet starts from the issue of unstable matching and
Abalation study on the borders of the variance and dis- utilizes pre-computed clustered priors to initialize reference
tance loss. δv and δd are respectively the borders of the points. Fitted from the map elements in the dataset, prior ref-
Refined

+ + Linear
HIQuery

Multi-Scale Deformable Weighted Sum


+
Deformable Cross-Attention Cross-Attention
Q Q Ref.
Self-Attention
V Ref. V

Point-Element Hybrider
+ +
BEV
+ +
Feature Decoupled Self-Attention Multi-Scale (Decoupled) Self-Attention Element Feat. Extractor Point Feat. Extractor
BEV Feature [Masked Attention] [Deformable Attention]
V K Q V K Q
Q Mask V V Ref. Q

+ + MAI
+ +
+ + BEV
Feature
+
+
Linear

… + … …
Hierarchical Reference Learnable Instance Point Hierarchical Reference Learnable Instance Point Learnable
Query Points Query Pos. Query Query Query Points Query Pos. Query HIQuery Query Ref.

(a) MapTRv2 Decoder (b) MGMap Decoder (c) HIMap Decoder


Gathered
Instance Query + +
Deformable Cross-Attention Decoupled Multi-Scale
MLP
Deformable Cross-Attention

Deformable Cross-Attention V Q Ref.


Inner-Instance
Q Ref. Feature Aggregation
V
[Self-Attention] +
Points
+ Query
BEV BEV
Feature MLP Multi-Scale Decoupled Self-Attention
Feature (Decoupled) Self-Attention
BEV Feature
V K Q
Decoupled Self-Attention
Position
Inner-Instance Query Fusion
V K Q
[Self-Attention]
+ + Embedding

Position
Embedding
+
… … …
Instance Learnable Instance Point Hybrid Hierarchical Learnable
Prior Ref.
Query Ref. Query Query Query Query Ref.

(d) MapQR Decoder (e) InsMapper Decoder (f) Our PPS-Decoder

Figure S1: Comparison of the decoders of previous SOTA methods: (a) MapTRv2, (b) MGMap, (c) HIMap, (d) MapQR, (e)
InsMapper and (f) our proposed PriorMapNet. (a), (b) and (d) are based on their open-source code, and (c) is based on the
implementation details and equations from the paper. For simplicity, we only show the first layer in the transformer decoder.
Since InsMapper currently has no open-source code and the specific implementation process is not detailed in the paper, we
only show the general structure of its decoder. To highlight the key distinctions between our PriorMapNet and previous SOTA
methods, we mark the randomly initialized learnable query position embeddings and reference points in light red, and mark our
prior-aware query position embeddings and reference points in yellow.

erence points lower the learning difficulty and achieve sta- mance under different modalities on nuScenes dataset and
ble matching. In contrast, HIMap, MapQR, and InsMapper different dimensions on Argoverse 2 dataset, demonstrating
ignore query priors and use randomly initialized reference the generalizability of our method.
points. Based on MapTR series methods, they explore the
correlations between instances and points, which introduce S4 Qualitative Results and Failure Cases
additional modules and computational complexity. MGMap We show more qualitative results and failure cases on
proposes Mask-Actived Instance to learn map instance seg- nuScenes val set in Fig. S2 and Fig. S3, respectively. For
mentation results and provides semantic priors for instance each example, the first column represents multi-view cam-
queries. However, semantic priors lack positional informa- era images. The second column shows the prediction GT.
tion, which is essential for vectorized map elements. Our The third and fourth columns respectively show the pre-
PPS-Decoder provides prior position and structure, result- dicted results of MapTRv2 and our PriorMapNet. Visualiza-
ing in significant performance improvements. tion results illustrate that PriorMapNet achieves improved
Performance. Our PriorMapNet surpasses previous SOTA results than MapTRv2. However, in some complex scenes
methods on nuScenes and Argoverse 2 dataset, as illustrated such as intersections, our method fails to predict some map
in Tab. 1 and Tab. 2. PriorMapNet achieves SOTA perfor- elements, showing our limitations and requiring future work.
Multi-view Images GT MapTRv2 Ours
Front

Front Left Front Right

Back Left Back Right

Back

Front

Front Left Front Right

Back Left Back Right

Back

Front

Front Left Front Right

Back Left Back Right

Back

Front

Front Left Front Right

Back Left Back Right

Back

� Road Boundary � Pedestrian Crossing � Lane Divider

Figure S2: Visualization of MapTRv2 and PriorMapNet results and the corresponding GTs. Models are trained for 24 epochs.
Multi-view Images GT MapTRv2 Ours
Front

Front Left Front Right

Back Left Back Right

Back

Front

Front Left Front Right

Back Left Back Right

Back

Front

Front Left Front Right

Back Left Back Right

Back

Front

Front Left Front Right

Back Left Back Right

Back

� Road Boundary � Pedestrian Crossing � Lane Divider

Figure S3: Visualization of failure cases. The green areas indicate locations where our method fails to predict accurate results.

You might also like