FusionMapper: Vectorized Map Construction by Multi-modal and Temporal Fusion based on Transformer

Haoxiang Jie, Yaoyuan Yan, Xinyi Zuo · 2025

Intelligent driving without relying on offline HDmaps has become one of the hot spots in the industry, and many scholars recently focus on building environmental maps online. In this paper, we effectively combine the color, texture information of the image and the position, intensity information in the LiDAR point cloud to realize the multimodal fusion detection of road elements, and propose the FusionMapper algorithm. In the proposed method, we design a dual multi-scale temporal feature fusion (DMTFF) module, which combine feature information of past frames and different scales to improve the continuity and robustness of the network for linear target detection such as lane lines and road curbs. Further, we use the Sparse Map Query based Transformer decoding layer, combined with novel Lane key-point query, Curb key-point query and crosswalk key-point query to achieve inference of three types of target key points, while effectively eliminating the time-consuming postprocessing of existing algorithms. The quantitative results show that the proposed algorithm has strong robustness and excellent accuracy. Specifically, our method reached $70.16 \% \mathrm{mAP}$ on nuScenes map dataset, which outperforms some industry-leading algorithms.

Read the paper · More papers on PaperTik