Video Anomaly Detection using Factorized Self-Attention Transformer

R Karthik, S Adithya, P Shalmiya, V. Subramaniyaswamy · 2024

Recently, video anomaly detection has seen significant interest in computer vision for identifying unexpected events in video streams using computational cameras, which has quite important applications in video surveillance, security, safety, and industrial control. Existing methods heavily rely on handcrafted features and rule-based algorithms to detect such anomalies, but these models typically tend to fail to generalize scenarios. Recently, with the success of deep learning-based approaches, particularly deep convolutional and recurrent neural networks, state-of-the-art models have been proposed to exploit the capability of deep learning to learn complex patterns from raw video without handcrafted features. Although deep neural networks have been shown to be the suitable solution to automatically capture complex relationships, there are still issues with addressing a wide spectrum of scenarios and acquiring large, well-labelled training datasets. In this research, a transformer-based architecture is proposed for anomaly detection using labelled training data. A Video Vision Transformer (ViT) is used to extract complex visual patterns from video frames, while audio data is encoded as spectrograms and transformed into audio features using an audio transformer. The outputs of the transformers are fused using a cross-modal relation-aware network and fed to a robust temporal feature magnitude (RTFM) learning module, which provides an anomaly score and classifies the input as anomaly or normal. Furthermore, different evaluation metrics are used to measure how well the model works, highlighting its ability. The proposed model is effective since transformer architectures are capable of grasping rich temporal relationships.

Read the paper · More papers on PaperTik