Unveiling Human Actions: Vision-Based Activity Recognition Using ConvLSTM and LRCN Models

Rahul Kumar Gour, Divyarth Rai · 2024

The crucial problem of vision-based human activity recognition (HAR) has several applications such as human-computer interface, surveillance, and healthcare. The effectiveness of two cutting-edge topologies for HAR, namely Long-term Recurrent Convolutional Networks (LRCN) and Convolutional Long Short-Term Memory (ConvLSTM) networks, is examined in this literature. Spatial and temporal dependencies in video data can be simultaneously modeled by ConvLSTM networks, which integrate LSTM units with convolutional procedures. Convolutional neural networks (CNNs) and long short-term memory (LSTMs) are used in LRCN architectures to automatically extract both spatial and temporal data from video clips. We compare the vision-based HAR performance of ConvLSTM and LRCN architectures through extensive experimentation on UCF 101 and UCF 50 datasets. We investigate how various network designs, training approaches, and fusion techniques affect the accuracy of recognition. Our findings show that LRCN and ConvLSTM architectures are capable of providing competitive performance for vision-based HAR. We draw attention to their potential for real-world use and offer suggestions for new lines of inquiry for this area of study.

Read the paper · More papers on PaperTik