HiTAGU-Net: Introducing Hierarchical Temporal-Aware Gated Units for Enhanced Motion Disentanglement in Action Recognition

Arash Asefnejad, Javad Mohammadzadeh, Mitra Mirzarezaee · IEEE Access · 2025

This paper presents HiTAGU-Net, a hierarchical temporal aware network for efficient human action recognition. The proposed architecture integrates a three-state Multi-Gated Unit (MGU) to jointly capture short-, mid-, and long-term temporal dependencies, a dual-level attention mechanism to emphasize key spatial-temporal features, and a difference-aware motion embedding (DME) to encode inter-frame variations without optical flow. We combine cross-entropy, center, and temporal-smoothness losses to enhance stability and intra-class compactness. Evaluations on Kinetics-700 and Something-Something V2 yield Top-1 accuracies of 81.8% and 83.7%, respectively, with only 33.2M parameters and ≈2.1 GFLOPs, demonstrating that HiTAGU-Net achieves high recognition accuracy, robust temporal modeling, and real-time efficiency superior to existing CNN-, RNN-, and Transformer-based approaches.

Read the paper · More papers on PaperTik