Spatio-Temporal Fusion for Video-Based Human Activity Recognition: A Novel LRCN TDL Approach
R. Athilakshmi, S Varsha, Urvi Bhanu Hirani, M. Salomi, N. Meenakshi · 2024
Human activity recognition is a crucial task in computer vision, finding applications in diverse fields such as healthcare, sports training, and security. Researchers have explored various approaches to tackle this challenge, some involving sensor data and others capitalizing on video data. This study focuses on two deep learning approaches: LRCN-TDL (Long-term Recurrent Convolutional Networks with Time Distributed Layer) and Convolutional Long Short-Term Memory Networks (ConvLSTM) for the specific task of identifying human actions from videos. The use of CNNs is advantageous because they can automatically extract relevant features, while ConvLSTM excels at handling sequential data, such as video frames. The proposed LRCN-TDL model combines the temporal understanding capabilities of recurrent networks with the spatial feature extraction power of convolutional networks. The Time Distributed layer is employed to efficiently handle sequences of variable lengths. Both models underwent training and evaluation on two datasets: the benchmark UCF50 dataset and the HMDB51 dataset. While both models achieved commendable accuracies, the LRCN-TDL model outperformed the Convolutional LSTM model, achieving an impressive accuracy of 93.6% on the UCF-50 dataset and 90.3% on the HMDB51 dataset.