Video Human Action Recognition and Classification Based on Channel Attention and LSTM Transformer
Yao He, Yi Yang, Chenxia Li, Jingyue Huang, Pinyao He · 2025
This study introduces a novel model for video-based human action recognition, This model aims to capture spatial features of actions through video frames, while utilizing temporal dependencies between frames for spatiotemporal feature fusion. The proposed model employs a pre-trained ResNet-50 network as an encoder to extract spatial features from video frames. These features are then enhanced by the Convolutional Block Attention Module (CBAM) through channel and spatial attention mechanisms. To capture temporal dependencies between frames, Long Short-Term Memory (LSTM) units are utilized, which model the dynamics of actions over time. Furthermore, to improve long-term dependency modeling, the outputs from the LSTM are processed by a Transformer decoder with a self-attention mechanism, thus enhancing the capture of motion dynamics in the video. Empirical results demonstrate the effectiveness of combining LSTM-Transformer for modeling temporal information, while the CBAM module significantly enhances feature representation. This approach successfully extracts appearance features from individual frames and captures temporal information over extended periods, providing a robust solution for video action recognition. Notably, this method outperforms traditional convolutional network classifiers, especially when applied to datasets with limited data, achieving significant performance improvements in action classification tasks.