PP-DETR: Progressive Proposal Detection Transformer for 3D Visual Grounding

Jinyuan Li, Duxin Zhu, Zhuangzhi Liu, Yang Luo, Jinhe Su, Guorong Cai, Yundong Wu · 2024

3D Visual Grounding (3D VG) has become increasingly important in applications such as robotics and human-computer interaction. However, accurately localizing target objects in complex 3D scenes based on natural language descriptions remains a significant challenge. While current methods based on DETR have shown notable improvements in the field of 3D VG, they struggle to achieve precise localization in the presence of multiple objects. To address these challenges, we propose Progressive Proposal Detection Transformer for 3D VG (PP-DETR). We introduce a progressive proposal generation strategy into the DETR-based 3D VG model which is named Progressive Category-Constrained Sampler(PCCS). After thoroughly integrating point clouds features with natural language features, our method generates object queries through a stepwise sampling of point clouds features. These queries are then decoded to produce the final bounding boxes for the target objects. PCCS reduces the interference from irrelevant proposals, enhancing the model's focus on relevant objects and thereby improving overall localization accuracy. Extensive experiments conducted on the ScanRefer dataset demonstrate that our model achieves an overall accuracy of 54.06% on the overall0.25 metric, surpassing the baseline EDA by 0.23%. In more complex multi-object scenarios, our approach yields improvements of 0.45% and 0.49% in the multi0.25 and multi0.5 metrics, respectively.

Read the paper · More papers on PaperTik