Spatiotemporal Attention Network for Action Recognition Based on Motion Analysis
Mingxiao Yu, Siyu Quan, Jiajia Wang, Yunlong Li · 2024
Action recognition in video content understanding primarily relies on the information extracted by the model, which needs to encompass both spatial and temporal features. Performance enhancement in deep learning models hinges on effectively learning these crucial spatial and temporal features from videos. Building on this analysis, we propose a spatiotemporal attention network for action recognition based on motion analysis(STAN-ARMA), comprising an optical flow integration module, an optical flow keyframe extraction module, and a spatial attention module. Experimental validation on the UCF-101, HMDB51, UTD-MHAD and HAA500 benchmark datasets demonstrates that the proposed network effectively learns spatial motion characteristics at critical moments in videos, thereby improving the model's recognition performance.