Asymmetric Explicit Synergy for Multi-Modal 3D Gaussian Pre-Training in Autonomous Driving

Dingwei Zhang, Jie Ji, Chengjun Huang, Bichun Li, Chennian Yu, Chenhui Qu, Zhengyuan Yang, Chen Hua, Biao Yu · World Electric Vehicle Journal · 2026

Generative pre-training via neural rendering has become a cornerstone for scaling 3D perception in autonomous driving. However, prevalent approaches relying on implicit Neural Radiance Fields (NeRFs) face two fundamental limitations: the shape-radiance ambiguity inherent in vision-centric optimization and the prohibitive computational overhead of volumetric ray marching. To address these challenges, we propose AES-Gaussian, a novel multi-modal pre-training framework grounded in the efficient 3D Gaussian Splatting (3DGS) representation. Diverging from symmetric fusion paradigms, our core innovation is an Asymmetric Encoder architecture that couples a deep semantic vision backbone with a lightweight, physics-aware LiDAR branch. In this framework, LiDAR data serve not merely for semantic extraction, but as sparse physical anchors. By employing a novel Explicit Feature Synergy mechanism, we directly inject raw LiDAR intensity and depth priors into the Gaussian decoding process, thereby rigidly constraining scene geometry in open-world environments. Extensive empirical validation on the nuScenes dataset demonstrates the superiority of our approach. AES-Gaussian achieves state-of-the-art transfer performance, yielding a substantial 7.0% improvement in NDS for 3D Object Detection and a 4.8% mIoU gain in 3D semantic occupancy prediction compared to baselines. Notably, our method reduces geometric reconstruction error by over 50% while significantly improving training and inference efficiency, attributed to the streamlined asymmetric design and rapid Gaussian rasterization. Ultimately, by enhancing both perception accuracy and system efficiency, this work contributes to the development of safer and more reliable autonomous driving systems.

Read the paper · More papers on PaperTik