Video Anomaly Detection with Video Vision Transformer
Biao Guo, Mingrui Liu, Qian He, Ming Yan Jiang · 2023
Video anomaly detection refers to the identification of unusual events or activities and represents irregular behavior. However, it is extremely difficult to localize and extract abnormal events containing potentially valuable information from long video streams. Most of the existing methods rely on autoencoder, which fail to capture the global context information of the video. In this paper, we propose a new video anomaly detection method called VadVVT using Video Vision Transformer (ViViT). This model uses Transformer to automatically learn video representation and global contexts. It utilizes the transform encoder to extract spatio-temporal features from the video sequence and employs a decoder to predict the next frame. Our method models normality in the training phase and identifies an event with unpredictable as anomalies during the test phase. The experiments results on three public datasets demonstrate that our method outperforms existing methods.