Lightweight Multimodal Feature Fusion and Spatiotemporal Learning for Human Action Recognition on Edge Devices
Sougatamoy Biswas, Anup Nandy, Asim Kumar Naskar · IEEE Transactions on Emerging Topics in Computational Intelligence · 2025
Human action recognition (HAR) remains a challenging topic in computer vision, attracting extensive research for applications in surveillance, sports, and human-computer interaction. Existing deep learning-based HAR methods often rely on either RGB-only inputs or global attention mechanisms, which suffer from poor generalization under occlusion, background clutter, and temporal ambiguity. Moreover, conventional methods struggle to capture fine-grained spatiotemporal dependencies due to sudden changes in motion dynamics and the lack of temporal consistency constraints in sequential modeling. To address these limitations, we propose a lightweight multimodal human action recognition framework that combines complementary cues from appearance, motion, and depth modalities. These features are integrated through a mid-level feature fusion strategy to form a unified and discriminative representation of human actions. The architecture employs an Enhanced Long-Term Recurrent Convolutional Network (E-LRCN) to model both spatial and temporal dynamics efficiently. A novel Temporal Causal Self-Attention (TCSA) module is introduced to enforce directional temporal consistency. It emphasizes recent motion context, significantly improving the discrimination of action sequences. Extensive evaluations on the KTH, UCF-101, JHMDB, and HMDB51 datasets show that the proposed framework surpasses state-of-the-art methods, achieving accuracies of 98.10%, 96.28%, 83.47%, and 77.60%, respectively. These results reflect gains of up to 1.27%, 2.08%, 2.81%, and 1.04% over the best-performing benchmark models. The proposed framework improves performance while reducing computational overhead, making it suitable for real-world human action recognition on resource-constrained platforms like Jetson Nano.