Dual-Modal Dynamic Weight Allocation for Scene Reconstruction
Qunshu Lin, Xinjun Wu, Chengdong Lin, Minghao Liu, Zhi Yu · 2025
Point cloud data annotation is a critical bottleneck in developing deep learning algorithms for 3D perception tasks, with high costs driven by data complexity and labor-intensive labeling processes. Traditional approaches to improve annotation efficiency typically reconstruct a global point cloud from sequential frames to eliminate redundant labeling, but encounter two significant challenges: reconstruction accuracy is constrained by pose estimation errors, and the resulting massive point cloud datasets create severe rendering bottlenecks that paradoxically reduce annotation efficiency. This paper introduces a novel dual-modal adaptive scene reconstruction and dynamic rendering framework that addresses both limitations simultaneously. Our approach innovatively integrates point cloud and image data through a series of carefully designed mechanisms: (1) a dual-modal scene reconstruction architecture that enhances geometric accuracy through complementary modality strengths, (2) a point cloudimage topology association mechanism leveraging multi-level message passing, (3) a modal feature decoupling and recombination strategy that resolves redundancy issues in multi-modal fusion, (4) a bidirectional cross-modal attention mechanism enabling mutual feature enhancement, and (5) a hierarchical point cloud storage structure supporting efficient dynamic rendering of billion-point datasets. Extensive experiments demonstrate that our framework achieves superior reconstruction fidelity while maintaining interactive rendering speeds even with massive point clouds, significantly enhancing the efficiency and accuracy of point cloud annotation processes. These advancements provide higher quality training data for downstream perception tasks, accelerating the development of reliable 3D perception systems for autonomous driving, embodied intelligence, and other applications requiring precise scene understanding.