A lightweight model for action recognition based on depth separable convolution and LSTM
Yuan Liu, Jianing Geng, Zhenli Zu, Mingxuan Ma, Hui Shi · 2024
Addressing the challenges that latest 3D convolutional neural networks face when recognizing actions, such as the excessive number of parameters, disregarding temporal dynamics, and taking too long to process, we propose a lightweight model for action recognition based on depth-separable convolution with LSTM (Long Short-Term Memory). First, Human behaviour and action regions are derived from video sequences through data pre-processing, followed by normalized pre-processing operations. Then, a rich representation of the activity features is obtained through the use of a deep-separable convolutional model that extracts spatial features of the behaviour and action regions. Next, the temporal dynamics of the actions are captured by capturing the extracted feature sequences in the LSTM network. The LSTM network's use of memory cells and gating mechanisms allows for efficient modeling of time series data. Finally, the action classification and recognition is done by utilizing a softmax classifier at the output layer of the network. The model is capable of capturing subtle changes in actions due to its integration of spatial and temporal features. Through the combination of a deep separable convolutional network and LSTM network, the proposed model has better expressive and generalization abilities, and is able to achieve better performance and results in the field of action recognition. The experimental results show that compared with the traditional action recognition algorithms, the algorithm action recognition rate is up to 95.41% on the UCF101 dataset.