Video entity relation detection algorithm based on temporal iterative inference
Zhi Liu, Ziguang He, Haojun Wei, Chuqin Yuan, Qing Li, Hongyun Lu, Yuan Li, Ran Cheng · 2025
Video entity relationship detection aims at recognizing the interaction between entities in a video, and its key task is to classify the relation triplets, including subject, object, and predicate efficiently. Existing methods tend to use static images to detect relationships between entities, ignoring the temporal features of the relationships between entities in the video, making it difficult to effectively capture the dynamic behavior of objects over time. To solve this problem, in this paper, a model based on temporal iterative inference is proposed, which utilizes Temporal Convolutional Network (TCN) to obtain rich temporal features from video sequences to enhance the inference process at each time step, especially in the dynamic change of complex relations. It consists of multiple repeating residual blocks, which help to capture long-term dependencies in the time series, and the dilated causal convolution is included. Dilated convolution is used to obtain a large receptive field without adding complex parameters, to understand the context information of the input data, while causal convolution is used to ensure that the model proceeds chronologically, keeping the directionality of the information flow. Experiments carried out on ImageNet-VidVRD and VidOR datasets show that, the proposed model can significantly improve the accuracy of relationship detection.