CHRONICLES IN MOTION : Spatiotemporal Activity Detection with Deep Learning
Rakesh Kumar, G. Nagasahithya, Chauhan Khushi Pramod, Apoorva Reddy · 2025
Action recognition and segmentation of video are now needed in many applications including surveillance, health, human-computer interaction, and sports analysis. Classical methods are incapable of dealing with dynamic human action complexity, background variations, and occlusion. This paper describes a sophisticated deep learning model for video action prediction and segmentation in terms of action will be attempted. High recognition accuracies result from applying deep neural networks particularly Convolutional Neural Networks (CNNs) and Recurrent Neural Networks or Transformer models in the system. Attention helps enhance the capability to focus on important features, enhancing robustness in real-world conditions. The approach is based on video frame processing to capture spatiotemporal features and subsequent sequential modelling for understanding action development over time. Pre-trained models like ResNet50 can be utilized in an attempt to improve feature representation. Segmentation is also carried out in an attempt to accurately identify the start and end boundaries of every action so fine-grained knowledge is attained. The model is trained and evaluated using benchmark datasets like UCF101, HMDB51, or ActivityNet to ensure its efficacy. The performance is evaluated in terms of accuracy, precision, recall, and Intersection over Union (IoU) for the segmentation. The value of this work lies in enhancing the accuracy and efficiency of action recognition and segmentation for real-time applications. The results bring the art of developing intelligent video analytics systems to make autonomous cars, monitoring, and assistive technologies more powerful. The action recognition model achieves an overall accuracy of 82%. The future research would involve engaging multimodal data, reinforcement learning, and real-time deployment optimizations to further advance the capabilities of the system. The research is one step closer to explaining video-based AI models in more explainable terms and making them more adaptable to varying environments.