Dense Voxel Reconstruction and Semantic Completion of Vehicle Scenes Based on a Conditional Diffusion Model

Weixuan Lin, Shitong Ye · 2025

For urban traffic environments with multiple obstructions and dense dynamic targets, this paper proposes a method for dense voxel reconstruction and semantic completion based on a conditional diffusion model. The method maps LiDAR-RGB multimodal features to an eight-fold linear scaling latent space, where hierarchical diffusion denoising is performed. Subsequently, a geometric reachability mask and semantic prior KL regularization are introduced as dual guides, significantly reducing ghost voxels while simultaneously improving sparse category mIoU. To ensure cross-frame stability, the framework is designed with implicit optical flow constraints to achieve temporal consistency. Experimental results show that on SemanticKITTI and SSCBench-KITTI-360, compared to the latest baseline CGFormer, the proposed method achieves an average mIoU improvement of 3.2 percentage points and a $24.8 \%$ reduction in phantom voxel rate, yielding favorable experimental results.

Read the paper · More papers on PaperTik