Large Language Model-based Semantic Text Decoupling and Cross Attention-based Bidirectional Multimodal Feature Fusion for Generalized Three-dimensional Referring Expression Segmentation

HyeLim Bae, Kyungmin Park, Incheol Kim · Journal of Institute of Control Robotics and Systems · 2025

3D referring expression segmentation (3D-RES) locates target objects within a 3D scene point cloud with precise 3D masks based on natural language descriptions. We propose a new benchmark dataset, OAR3DRef, and a baseline model, TDMF, to overcome the limitations of current research in generalized 3D RES. OAR3DRef provides detailed annotation of referring expressions with rich semantic components, helping a 3D-RES model in understanding the correct meaning of the expressions. TDMF, which is a baseline deep neural network model proposed for effective use of OAR3DRef, utilizes both a linguistic parser and a large language model to decouple the given natural language expression into multiple semantic components. The model also adopts a novel semantic alignment loss function to align decoupled text features with object visual features. It performs cross attention-based bidirectional feature fusion between expression text features and object visual features before their semantic alignment. Furthermore, the model makes use of additional 2D object visual features extracted from multiview RGB-D scene images to enhance exact recognition of object attributes such as color, material, and appearance. We demonstrate the superiority of the proposed baseline model, TDMF, through various quantitative and qualitative experiments using the new OAR3DRef benchmark dataset.

Read the paper · More papers on PaperTik