Listen, Perceive, Grasp: CLIP-Driven Attribute-Aware Network for Language-Conditioned Visual Segmentation and Grasping

Jialong Xie, Jin Liu, Saike Huang, Chaoqun Wang, Fengyu Zhou · IEEE Transactions on Automation Science and Engineering · 2024

Endowing robots with the ability to understand natural language and execute grasping is a challenging task in a human-centric environment. Existing works on language-conditioned grasping achieve end-to-end grasping detection based on language. However, these works lack fine-grained visual grounding, resulting in cognitive deficits for robots. Moreover, they ignore the correlation between visual attributes of objects and grasping, leading to coarse grasp poses. To this end, we propose a CLIP-driven aTtribute-aware network (CTNet) for language-conditioned visual segmentation and grasping, enabling the robots to listen, perceive, and grasp the referred object in real-world applications. Specifically, we first employ Listen stage to understand basic linguistic and visual concepts. Subsequently, we introduce Perceive stage to mine multi-modal features and visual attribute cues (e.g., boundary and spatial location), then yield a language-conditioned segmentation mask. Further, we design Grasp stage to aggregate the perceived attribute information and refine the spatial location and grasping rectangle, generating a high-quality grasp pose. Lastly, we provide an extended large dataset Ref-OCID-Grasp to train and test our method, achieving a grasping accuracy of 97.76% and segmentation OIoU of 91.82%. The real-world robotic applications demonstrate the effectiveness of our proposed approach. The project, video, and dataset can be found athttps://ctnetgrasp.github.io. Note to Practitioners—Most of the existing grasping methods focus on clearing all objects in the workspace. However, as robots integrate into human society, robots should learn to grasp the desired object by understanding human language. Therefore, language-conditioned grasping is a significant skill for human-robot collaboration. The prior works directly complete the grasp detection through the language-grasp paradigm, but they ignore the discussion on whether the robot understands the concept of vision and language expression of the object. Therefore, this paper proposed the Listen-Perceive-Grasp paradigm, in which the Listen-Perceive stage is responsible for the conception alignment of the object in language expression and visual pixels, and the Perceive-Grasp stage achieves the constraining and refining the grasp detection by the perceived visual attributes such as boundary and shape. Experiments show that this method can obtain a refiner grasp pose in cluttered environments and perform language-conditioned grasping well in the real world. In future research, we will work on 6-DoF grasping and multi-object disambiguation conditioned on language.

Read the paper · More papers on PaperTik