A CNN-LSTM Framework for Real-Time Human Action Recognition in Video Sequences
Ashwin Shenoy M, N. Thillaiarasu, Bharath S, Chaithanya, N V Bhoomika, Chetan Sangalad · 2025
This paper presents an advanced framework for action recognition in video sequences, leveraging the combined strengths of Convolutional Neural Networks (CNNs) and Long Short-Term Memory (LSTM) networks to capture both spatial and temporal dimensions of human activities. The framework processes video input by extracting individual frames and analyzing spatial features through CNNs to determine "what" actions occur in each frame. These extracted features are then passed to an LSTM network, which models the temporal dynamics, allowing the system to understand "how" actions unfold over time. Real-time processing capabilities are enabled through OpenCV, making this approach suitable for applications requiring immediate feedback and dynamic response. The model is trained on a diverse dataset encompassing a range of action classes, achieving high accuracy in recognizing activities from simple gestures to complex behaviors. This CNN-LSTM architecture demonstrates broad applicability across various domains: it can identify abnormal behavior in surveillance settings, enable gesture-based control in human-computer interaction, and analyze athletic performance in sports. By integrating CNNs and LSTMs, this framework enhances automated video analysis capabilities and underscores the potential of deep learning to interpret complex human actions in real-world scenarios. This work represents a significant step forward in video-based action recognition, paving the way for more intelligent, responsive, and robust systems.