Query graph attention for video relation detection
Jian Wang, Haibin Cai · 2023
As a bridge to connect vision and language, visual relations between objects, visual relation provide a more comprehensive visual content understanding beyond objects. Most previous works adopt the track-to-detect framework for video visual relation detection (VidVRD), which cannot capture long-term spatio- temporal contexts in different stages and also suffers from inefficiency. In this work, we propose a query-based method for video visual relation detection. Our model exploits graph structure to autoregressively generate relation graphs with spatio-temporal contexts and uses an attentional graph convolutional network to fuse the contexts. Experiments on benchmark datasets ImageNet-VidVRD demonstrate the accuracy of our method.