DCT: A lightweight approach for the real-time video anomaly detection in multi camera
Sanjay Roka, Manoj Diwakar · 2023
CNN and transformer are the popular techniques that have been used in the past to design video anomaly detection approaches. But due to the spatial inductive bias nature of convolution, the CNN-based approaches completely rely upon the local features. Whereas transformer-based approaches are extremely heavy in terms of weight as they required deeper and wider layers to learn the visual representation. In addition, most of the approaches proposed in the past are mainly designed for the single camera-based scenario. In real world scenario, there can be N number of distributed cameras monitoring the different locations of an area. Considering these issues in this paper we design a novel lightweight approach to detect the video anomaly in the multi-camera scenario. We fuse the CNN and transformer in our approach and provide the CT module which has the property like convolution permitting it for global processing. We extracted the local features using CNN and global features using the transformer. We identify the video anomalies using the PSNR and anomaly score. We detect and track an abnormal object using the YOLO and DeepSORT algorithms. To verify the two people in different cameras are the same we use our modified architecture DCT. We evaluated our approach using our custom multicamera dataset GEU, For the comparison with state-of-the-art approaches we use standard datasets Ped1, Ped2, Avenue, and ShanghaiTech in which we achieve AUC of 98.21%, 99.32%, 98.46%, and 98.45% and EER of 2.36%, 1.55%, 1.34%, and 2.15% respectively. The average running time of our approach was 0.0092 (108 FPS).