MMSF‐DET : Multi‐Modal Multi‐View Serial Fusion for 3D Object Detection

Xinyu Gao, Anhong Wang, Hao Jing, Shiao Xu, Tammam Tillo · IET Image Processing · 2026

ABSTRACT As autonomous driving progresses toward levels 4/5, researchers increasingly adopt multi‐sensor fusion to break the accuracy bottleneck in 3D object detection. Yet current multi‐modal frameworks are limited in aligning heterogeneous features and enabling joint cross‐modal inference, hindering the full exploitation of inter‐modal complementarity. We propose a novel framework based on multi‐modal, multi‐view serial fusion for 3D detection (MMSF‐DET). In the perspective view (PV), which is closer to the image domain, we introduce a perspective‐view multi‐scale bidirectional fusion module (PV‐MBF) built on a sectorized point‐cloud grid to realize initial cross‐modal fusion. In the birdn the bird. (BEV), we design a BEV‐MBF that leverages bilinear interpolation to perform deep, MBF of cross‐modal features. To overcome the shallow interactions of conventional parallel fusion and strengthen spatial awareness, we devise a serial fusion strategy that operates across the PV and BEV branches. Additionally, we employ a collaborative training scheme with auxiliary tasks—3D semantic segmentation, 2D object detection, and 2D semantic segmentation—and cross‐modal multi‐task joint optimization to further enhance performance. Experiments on KITTI and the Waymo open dataset (WOD) demonstrate superiority over existing multi‐modal detectors, validating the effectiveness and advancement of the proposed fusion strategy and multi‐task learning framework.

Read the paper · More papers on PaperTik