Spatio-temporal transformers and semantic insights: redefining video anomaly detection
Sai Prasad Amara, Pandey Nidhi, Nikita Suresh, Ria R. Kulkarni, Ramamoorthy Srinath · 2025
Traditional video surveillance systems often rely on manual monitoring, resulting in labor-intensive processes that are prone to errors, particularly when detecting complex events that hinder public safety and security. This study addresses these limitations by integrating advanced deep learning models for comprehensive video and text feature extraction, leveraging the power of multi-modal fusion. We employ TimeSformer to capture intricate spatio-temporal patterns within video data and RoBERTa to extract semantic insights from accompanying text captions. Our approach utilizes three multi-modal fusion techniques: Concatenation Fusion, Gated Fusion, and Compact Bilinear Pooling. These methods effectively combine visual and textual representations to enhance anomaly detection capabilities. Experimental results reveal that Concatenation Fusion significantly outperforms the other methods, achieving an accuracy of 85.19% in complex video scenarios. This fusion-based approach not only improves detection accuracy but also reduces the reliance on continuous human oversight, making it a practical solution for applications in public safety and security.