Spatio-Temporal Context Reasoning with Heterogeneous Graph Neural Network and Self-Attention for Video Visual Relation Detection
Incheol Kim · Journal of Korea Multimedia Society · 2025
In order to understand a wide range of video scenes in detail, it is necessary not only to detect and track individual entities, but also to find dynamic relationships between them in sequential video scenes. In general, Video Relationship Detection(VRD) has some important issues: (1) how to set the primitive temporal regions to find inter-object relationships, (2) how to infer rich spatio-temporal context to predict each inter-object relationship correctly. In order to address these issues, we propose a novel video relationship detection model, STCN(Spatio-Temporal Context Network). STCN takes an efficient instance-based approach that assigns an unique instance ID to each video entity track found on an entire video, and then makes use of these IDs for both relation detection within each segment and relation association between neighboring segments. Furthermore, STCN adopts a novel spatio-temporal context reasoning module called ST-HSA. ST-HSA performs both spatial context reasoning with heterogeneous graph neural network and temporal context reasoning with Transformer self-attention layers. Through many quantitative and qualitative experiments with the VidOR benchmark dataset, we prove high performance of the proposed model.