Spatial Representation and Motion Attention Fusion Network for Few-shot Action Recognition

Jiabao Wang, Baogui Qi · 2024

Recently advancements in few-shot action recognition technology have been notable. However, certain challenges persist. One issue is that it is difficult to effectively model motion information, which results in subpar performance on datasets that emphasize motion itself. Another issue is the neglect of the integration of global background information and local detail information, leading to a lack of attention to appearance information in the models. To address these challenges, we propose Spatial Representation and Motion Attention (SRMA) fusion network. By constructing a dual-branch structure comprising Motion Attention Enrichment Module (MAEM) and Spatial Attention Enrichment Module (SAEM), the network simultaneously captures and models motion information while also focusing on the global and detailed information of the motion space. Furthermore, we integrate the cross-entropy loss from both branches as the loss function for the entire model. Experimental results demonstrate that our method outperforms most other approaches across multiple datasets.

Read the paper · More papers on PaperTik