Text-driven affordance learning from egocentric vision
Tomoya Yoshida, Shuhei Kurita, Taichi Nishimura, Shinsuke Mori · Advanced Robotics · 2025
Understanding how to interact with objects is crucial for many applications such as robotic manipulation. The aim of this study is to build a robust approach that handles multiple interactions, including tool-object interaction. We introduce text-driven affordance learning, which aims to learn contact points and manipulation trajectories from an egocentric view following textual instruction. To avoid costly manual annotations, we create a large pseudo dataset: TextAFF80K. Then, we apply deconvolutional layers and MLPs to state-of-the-art referring expression comprehension models, training them on TextAFF80K. Experimental results show that our approach robustly handles various affordances and excels in tool-object interaction.