Learning Human-Object Interactions in Videos by Heterogeneous Graph Neural Networks
Qiyue Li, Jiapeng Yan, Jin Zhang, Yuanqing Li, Zhenxing Zheng · 2024
Video-based human-object interaction (HOI) recognition aims at labeling of human and object sequences with multiple human-object interaction classes. Existing video HOI models typically treat human-object interaction as the homogeneous graph in videos. However, the homogeneous graph only represents one type of node and interaction, which is unable to distinguish whether messages are intra-class or inter-class, or originate from the active or passive side. In this paper, we propose a human-object interaction model based on heterogeneous graph neural networks. The proposed model uses the heterogeneous graph to represent human-object interactions in videos, utilizing heterogeneous attention mechanisms to achieve message passing between nodes. Additionally, the model designs a skip-connected spatial-temporal graph structure for video human-object interaction, focusing on both the motion information of entities in the current frame and the motion states of neighboring frame entities. Our method demonstrates state-of-the-art performance on two pivotal HOI benchmarks, including the CAD-120 dataset and the Something-Else dataset.