Smart Video Monitoring: Advanced Deep Learning for Activity and Object Recognition
Dandinashivara Revanna Shashikumar, N Tejashwini, K N Pushpalatha, Anurag Kumar, Om N Chavan, Atharva Mishra · 2025
This study explores the integration of Convolutional Neural Networks (CNNs) and Long Short-Term Memory (LSTM) networks for the real-time recognition of human activities in video data. By harnessing the advantages of these two approaches, the system achieves high accuracy in detecting complex human actions. Specifically, CNNs address the spatial aspects of the task, while LSTMs handle the temporal sequences. A notable feature of the system is its categorization module, which enables users to select an action and identify similar actions, thereby enhancing productivity and usability. Existing models often face challenges related to real-time inter- action capabilities and resilience to environmental disturbances. This study tackles these shortcomings by refining the CNN-LSTM framework to support real-time functionality and incorporating preprocessing techniques, such as frame extraction and normal- ization, to improve input data quality. The system’s effectiveness is measured using indicators like accuracy, recall, and latency, demonstrating its advantages over traditional rule-based and basic deep learning approaches. The early findings are optimistic, demonstrating significant improvements in performance. Nevertheless, challenges remain, particularly in tracking per- formance under occlusion or in cluttered environments. Future research should explore the integration of multi-modal data and advanced architectures, such as spatio- temporal graph con- volutional networks (STGCN), to further enhance recognition accuracy and system robustness. In conclusion, the proposed CNN-LSTM hybrid architecture for activity recognition demonstrates potential for applications in video surveillance and beyond, including fields like healthcare and sports analytics. The system offers improved automated monitoring capabilities through enhanced accuracy, scalable human action detection, and user- friendly design.