Human Action Recognition Research Based on Channel-Temporal Self-Attention Block Network
Xiaowei Han, Yunjing Lu, Qiuyang Guo, Jiwu Liu, Chaolong Fei · 2024
With the rapid development of information technology, video has become an important carrier for people to perceive information in daily life. If you want to analyze the action of people in the video data of various scenes and make use of the results, you must first identify the human action accurately. In this paper, a temporal self-attention module is proposed to help the network dynamically learn the correlation between different data frames, so as to better capture the dependency on the time of action. On this basis, a novel network structure, Channel-Temporal Self-Attention Block (C-TSAB), is further proposed. C-TSAB adaptively emphasizes the key channels of input features and focuses on the time dimension to enhance the model's ability to recognize and represent action features. In the ablation experiment, the Top-1 and Top-5 indicators in HMDB51 dataset were 58.54% and 81.70%, respectively, and the Top-1 and Top-5 indicators in UCF101 dataset were 58.54% and 82.73%, respectively. Comparative experiments were conducted on HMDB51 dataset and UCF101 dataset respectively, our model obtained a better Top-1 indicator, both of which were 58.54%, proving that our model has relatively good performance in video HAR task.