Spatio-Temporal Deep Learning Models for Human Activity Recognition - Performance Evaluation and Optimization
N.V. Rajesh, H Sarojadevi, S. Nitya, B J Abhishek, B. Ananya, Shad Akhtar · 2024
This paper presents an advanced deep learning framework designed to enhance Human Activity Recognition (HAR) by leveraging the combined strengths of Convolutional Neural Networks (CNN), Long Short-Term Memory (LSTM) networks, and Gated Recurrent Units (GRU). HAR plays a crucial role in various applications, including surveillance, autonomous systems, and human-computer interaction. Despite significant advancements in machine learning, accurately recognizing and classifying human activities from video data remains challenging due to the complexity of spatio-temporal dependencies. Our proposed hybrid approach integrates 3D CNNs for efficient spatial feature extraction, ConvLSTMs to capture temporal dynamics, and GRUs to refine sequential learning, creating a robust model capable of handling complex and overlapping human actions. Extensive evaluations were conducted on two benchmark datasets, HMDB51 and UCF50, focused on 12 actions. The hybrid model achieved superior performance, with an accuracy of 96% on the UCF50 dataset and 95.4% on the HMDB51 dataset, surpassing traditional models such as LRCN, ConvLSTM, and 3D CNN in both precision and F1-score. This study highlights the effectiveness of hybrid spatio-temporal models in improving HAR accuracy, offering valuable insights for future research and practical applications in the field.