CNNViT: A robust deep neural network for video anomaly detection

Nikhil Garuda, G. Prasad, P. P. Dev, P. Das, Ebrahim Ghaderpour · IET conference proceedings. · 2024

Detecting anomalies in videos poses a significant challenge due to the unbounded, infrequent, ambiguous, and irregular nature of abnormal events in real-world scenes. Recently, transformers have shown remarkable modeling capabilities for sequential data. As a result, we endeavor to leverage transformers for video anomaly detection. This paper presents a novel prediction-based method for video anomaly detection called CNNViT by integrating the architectural elements of Convolutional Neural Network (CNN) and Vision Transformer (ViT). The purpose of this fusion is to effectively capture enhanced spatial-temporal information and global features. The effectiveness of the proposed method is evaluated on UCSD Ped2 and CUHK Avenue benchmark datasets. Experimental results demonstrate that the proposed method attains considerably superior performance compared to state-of-the-art techniques.

Read the paper · More papers on PaperTik