MSViT: A video anomaly detection for real time unconstrained environment

Sanjay Roka, Manoj Diwakar · 2023

Most of the previously proposed approaches are either based on CNN or transformer. CNN due to its spatial inductive bias can only learn the local representation whereas due to the absence of this property, the transformer-based approaches can learn only the global representation. Thus, the approaches based on these techniques are generally heavy as they required deeper and wider layers to learn visual representation. In addition, these approaches are mainly designed for the constrained environment. In this paper, we proposed the novel approach MSViT using MobileNet and a transformer that can extract both the local and global features from the incoming frames. Our approach is light, and shallow and can detect the video abnormal activities in both the constrained and unconstrained environment in real-time with high accuracy. The abnormal object recognized by our approach is further detected and tracked using the YOLO and DeepSORT algorithms. For the evaluation, we used our custom unconstrained environment dataset GEU and for the comparison, we used constrained environment dataset Ped1, Ped2, CUHK Avenue, and ShanghaiTech for which we obtain AUC of 97.58%, 97.21%, 94.32%, and 95.56% and EER of 3.25%, 3.49%, 6.8%, and 8.25% respectively. Our approach has an average running time of 0.00785 (127 FPS).

Read the paper · More papers on PaperTik