ESTN: Exacter Spatiotemporal Networks for Violent Action Recognition
Yuqi Chen, Bin Zhang, Ying Liu · 2021 IEEE 6th International Conference on Signal and Image Processing (ICSIP) · 2021
To solve the problems of low efficiency of manual violence recognition and low accuracy of violence recognition method combined with deep learning, this paper proposes an exacter spatiotemporal network for violent action recognition (ESTN). Firstly, the video samples are segmented and pruned by data preprocessing technology and data enhancement technology, and the spatiotemporal features of the preprocessed samples are extracted and fused by a two-dimensional convolutional neural network and three-dimensional convolutional neural network. The two-dimensional network mainly extracts the deep spatial features of the video, and the two branches of the three-dimensional network extract the scene spatial features and temporal features respectively. Finally, the recognition results of the segmented samples are fused to get the final result. Our approach has achieved more than 94% accuracy on HockeyFight and more than 70% accuracy on RWF-2000.