Dynamic Temporal Modeling Using Cascaded Deep Networks and Encoder Transformers for Human Action Recognition

Haseeb Amjad, Nija Asif, Umar Shahbaz Khan, Hassan Elahi · 2024

Human action recognition is an important area of research in computer vision with its vast applications from healthcare to industrial environments, solving real-world problems like smart manufacturing, video surveillance, and artificial intelligence. The increased trends of human robot interactions in workplaces and everyday life further signify this area. Gesture-based control that helps a robot recognize motions, gestures, and other physical actions allows them to assist and operate seamlessly in various interactive settings by recognizing and responding to human actions in real-time. The human action recognition models require a very high computational power to process the video frames which usually comes with a lot of cost. This paper proposes a dynamic temporal modelling for effective spatio-temporal human action recognition by simultaneously capturing the short-term, long-term, and global contextual dependencies with a balanced trade-off between accuracy and computational efficiency. Convolutional Neural Networks (CNNs) extract the spatial features, Gated Recurrent Units (GRUs) efficiently handle the short-term patterns, Long Short-Term Memory (LSTMs) support the long-term temporal dependencies and transformers enhance global understanding. The cascaded architecture integrates EfficientNetB0 with LSTMs, GRUs and transformers to achieve a superior performance with computational efficiency. The proposed model outperforms benchmark dataset UCF50 achieving an accuracy of 92% and demonstrating improved inference time and robustness.

Read the paper · More papers on PaperTik