Two-Stream Sparse Feature Non-local Spatiotemporal Residual Convolutional Neural Network for Human Action Recognition
Huimin Qian, Junwei Huang · 2024
Three-dimensional convolutional neural networks and two-stream convolutional neural networks have distinct advantages for recognizing human actions in videos, while both are lack of modeling capability of long-range dependencies in video frames. In view of this, a Two-Stream Pruned Sparse Feature Non-local Spatiotemporal Residual Convolutional Neural Network (TPSFNLST-ResCNN) is proposed in this paper. In which, ST-ResCNN with the presented sparse feature non-local module (SFNL) is applied both in spatial stream and temporal stream. And SFNLST-ResCNN is then compressed based on a channel pruning scheme to reduce the parameters. Among which, the SFNL module can improve the capability of extracting the long-range dependencies information of human actions under less computing load. Experimental results show that the proposed human action recognition model achieves recognition accuracies of 98.36% and 74.70% on the UCF101 and HMDB51 public datasets, respectively. Compared with existing methods, the proposed model offers the advantages of less parameters and higher recognition accuracy.