3DifFusionDet: Diffusion Model for 3D Object Detection with Robust Lidar-Camera Fusion

Xinhao Xiang, Simon Dräger, Jiawei Zhang · 2026

LiDAR-camera fusion has proven effective for 3D object detection by leveraging the complementary strengths of each modality, improving robustness and accuracy. Meanwhile, diffusion models—originally designed for image generation—have shown promising properties such as flexibility, dynamic evaluation, and iterative refinement, and have recently been explored in 2D object detection [1]. However, their potential in 3D object detection remains underexplored. In this work, we propose 3DifFusionDet, a novel framework that formulates 3D object detection as a denoising diffusion process. The model learns to iteratively refine a set of randomly initialized 3D bounding boxes into accurate predictions. Built under a LiDAR-camera fusion pipeline, 3DifFusionDet incorporates a two-branch feature alignment module that employs cross-attention to effectively integrate image and point cloud features. Extensive experiments on KITTI and Waymo Open demonstrate that our method achieves competitive performance compared to state-of-the-art approaches, with the added benefit of adjustable inference steps to trade off accuracy and speed. Our findings highlight the feasibility and advantages of applying diffusion models to 3D object detection, especially in multi-modal settings.

Read the paper · More papers on PaperTik