Violence Detection In Videos via Motion-Guided Global and Local Views
Ning Su, Lijuan Sun, Yutong Gao, Jingchen Wu, Xu Wu · 2023
Video violence detection aims to locate the time window in which violent behavior occurs. Most methods focus on utilizing RGB features directly or only fusing RGB and audio features, ignoring the effective exploitation of motion information carried in optical flow. This lack of emphasis on motion information may impact the overall accuracy of violence detection. Moreover, we observe that videos contain strong local correlations, so it is insufficient to analyze only from a holistic perspective without capturing finer details. Therefore, in this paper, we design a novel Global-and-Local Cross-Modal Network (GL-CMN) for violence detection, which effectively integrates motion information and multi-granularity features from target videos. Specifically, we first propose a Motion-Guided Attention Module (MGAM) to obtain enhanced visual features by calibrating RGB features through optical flow features. Secondly, The enhanced features are simultaneously fed into two parallel branches of the network. The global branch fuses the visual and audio features into holistic representations. The local branch extracts multi-scale temporal dependencies through dilated convolutions. Experiments demonstrate that our method exhibits significant improvement compared to previous state-of-the-art methods on the XD-Violence dataset.