Multimodal Attention Network for Violence Detection

Rui Zhou Zhenzhen Liu, Xiaoyu Wu · 2022 2nd International Conference on Consumer Electronics and Computer Engineering (ICCECE) · 2022

Violence detection task is an important branch of the field of computer vision, and multimodal method is currently the main method in the violence detection task. Appearance, optical flow, and audio information are commonly used violence feature information. Appearance and optical flow information reflect visual information, and audio information serves as an auxiliary of visual information. At present, appearance and audio information are the commonly used information in violent video detection. However, few studies focused on motion information. Moreover, there are also few effective methods for multimodal violence feature fusion. In this paper, we propose a lightweight feature fusion network based on the multimodal attention model. It can extract the multi-scale motion feature through both appearance feature and optical flow feature, with audio feature serve as auxiliary features. To demonstrate the effect of our model, we performed experiments on the publicly available dataset VSD2015 and the self-built dataset Violence Correspondence Detection (VCD), where VSD2015 achieves an AP of 0.45 and outperforms state-of-the-art performance.

Read the paper · More papers on PaperTik