AffordStruct: Weakly Supervised Affordance Grounding Based on Spatial Interaction and Knowledge-Aware

Shiyu Wang, Shanyi Zhang, Fengtao Sun, Wenbai Chen, Peiliang Wu · IEEE Transactions on Automation Science and Engineering · 2025

Affordance grounding refers to identifying the interactive regions of an object that enable it to perform a specific function, which is essential for effective robot interaction with complex environments. Current methods face challenges of incomplete localization and unclear distinction of affordance regions. To address these challenges, we propose a novel framework called AffordStruct, which draws on human experiential knowledge to explore affordance structures from the perspectives of spatial configuration and semantic relationships. Specifically, we first compute multi-order spatial interactions from the spatial configuration perspective to enhance long-range pixel dependencies, thereby reinforcing the network’s sensitivity to affordance-relevant regions and suppressing irrelevant areas. Besides, we design a pyramid hierarchical chain-of-thought prompting to guide the large language model in reasoning about affordance structural attributes. An attention-based fusion strategy is then employed to perform multimodal fusion with egocentric images that contain only the target objects, enabling the model to better perceive affordance structures at the semantic level. Finally, based on the local knowledge transfer mechanism, the affordance knowledge extracted from the exocentric view containing human-object interactions is transferred to the multimodal egocentric view. Extensive experiments on public datasets and robotic grasping tasks demonstrate that our method outperforms existing state-of-the-art affordance grounding models.

Read the paper · More papers on PaperTik