Integrating Pose Features and Cross-Relationship Learning for Human–Object Interaction Detection
Lang Wu, Jie Li, Shuqin Li, Yu Ding, Meng Zhou, Yuntao Shi · AI · 2025
Background: The main challenge in human–object interaction detection (HOI) is how to accurately reason about ambiguous, complex, and difficult to recognize interactions. The model structure of the existing methods is relatively single, and the image input may be occluded and cannot be accurately recognized. Methods: In this paper, we design a Pose-Aware Interaction Network (PAIN) based on transformer architecture and human posture to address these issues through two innovations: A new feature fusion method is proposed, which fuses human pose features and image features early before the encoder to improve the feature expression ability, and the individual motion-related features are additionally strengthened by adding to the human branch; the Cross-Attention Relationship fusion Module (CARM) better fuses the three-branch output and captures the detailed relationship information of HOI. Results: The proposed method achieves 64.51%AProle#1, 66.42%AProle#2 on the public dataset V-COCO and 30.83% AP on HICO-DET, which can recognize HOI instances more accurately.