Integrating Visual Features and Spatial Relations for Human Interaction Understanding

Nguyen Ngoc Cuc Phuong, Dinh Tuan Tran, Joo‐Ho Lee · 2024

Understanding human interactions is extremely crucial in various applications, including robotics, automated systems, human-computer interaction, and video surveillance. Many studies have leveraged rich visual features and spatial relations of human-human pairs extracted from RGB images to analyze interactions between two people, yielding impressive results. Despite these advancements, there have not been many attempts to effectively integrate both visual features and spatial relations to enhance the performance. To address this shortcoming, we introduce the Visual Refinement Network (VRN), a framework designed to refine visual features with spatial relations to produce robust feature representations for predicting human interactions. Experimental results show that our proposed model achieves superior performance compared to many other state-of-the-art approaches on the public dataset BIT for human interaction understanding.

Read the paper · More papers on PaperTik