Multi-stage motion excitation network for fine-grained action recognition
Yating Liao, Yu Dai, Bohong Liu, Ying Jie Xia · 2025
With the rapid development of sensor devices, video-based human action recognition has emerged as a prevalent research direction. Current approaches have achieved remarkable results by utilizing inter-frame differencing or attention mechanisms to extract short- and long-term motion information. However, as the complexity of actions escalates, some methods struggle to adequately capture rich motion for fine-grained action recognition. To address these issues, we propose a lightweight model named Multi-stage Motion Excitation Network (MMEN), which integrates multi-level motion modeling with attention mechanisms for efficient fine-grained action recognition. MMEN comprises three key components, including the Motion Boundary Perception (MBP) to perceive subtle intra-segment motion changes, the Two-way Motion Selection (TMS) to model inter-segment action evolution, and the Spatio-Temporal Global Attention (STGA) to capture video-level information. Experimental results on the HDMB51, Diving48, and Something-Something V1 datasets demonstrate that MMEN enables the model to learn richer motion information while achieving a better balance between computational cost and recognition accuracy.