Channel-wise Attention in 3D Convolutional Networks for Violence Detection
Bin Jiang, Fangqiang Xu, Wenxuan Tu, Chao Yang · 2019
Aiming at the problems of low computational efficiency and insufficient precision for traditional violent behavior recognition methods, we propose a SELayer-3D Convolutional Neural Network (C3D). Firstly, the C3D model is adopted to extract the spatio-temporal feature information in the video block. Secondly the obtained spatio-temporal features are assigned weights according to the importance degree by SELayer. Finally, the output is predicted by the Softmax classifier. In the test experiment on the CrowdViolence dataset, our method achieves an accuracy of 98.08%, which is 6.08% higher than that of the Deeper 3D Convolutional Neural Network (D3D) model. In the test experiment on the HockeyFight dataset, our method achieves 99.0% accuracy, which is 2.0% higher than that of the FightNet model. The speed can reach 500 FPS or so compared to the artificial feature extraction method of 20-25 FPS. Moreover, experiments show that the proposed method has higher detection accuracy and efficiency.