Multi-Scale Channel Attention Inspiring Multi-Task Network via Self-Supervised Learning for Violence Recognition

Xin Song, Suyuan Li, Zhongcong Zhao, Xiaoqi Wang, Penghui Liu, Zhigang Xie · 2023

Generally, a large amount of training data is essential to train deep learning models for obtaining more accurate detection performance in the computer vision domain. However, collecting and annotating datasets will lead to extensive costs. In this paper, we propose multi-task network based on multi-scale channel attention to learn general video features without adding any human-annotated labels, aiming at improving the performance of violence recognition. Firstly, we propose a violence recognition method based on a convolutional neural network with the self-supervised auxiliary task, which can learn visual features for improving down-stream task (recognizing violence). Secondly, we establish a balance-weighting scheme to solve the crucial problem of balancing the self-supervised auxiliary task and violence recognition task. Thirdly, we develop a multi-scale channel attention module, indicating that the proper use of the channel attention mechanism can effectively reduce the weights of the feature channels and regions representing the background, further improving the semantically meaningful representation of the network. To evaluate the proposed method, two benchmark datasets have been used, and better performance can be shown by the experimental results compared with other state-of-the-art methods.

Read the paper · More papers on PaperTik