Video Anomaly Detection Based on 3D-CNN Autoencoder with Bidirectional LSTM

Mohan Sankaran, Nagaraju Jooluri, Balasundaram Subbusundaram, Vijaya Bhaskar Oggu, Alokam Meghana · 2025

Video anomaly detection is a critical component of intelligent surveillance systems aimed at detecting anomalous events in real time. Traditional supervised models require substantial amounts of annotated anomaly data making them costly and time-consuming to acquire. To address these limitations, we propose an unsupervised deep learning-based foundation for anomaly detection, utilizing a hybrid model with 3D Convolutions and a Bidirectional Long Short-Term Memory (Bi-LSTM) model. The primary difficulty is capturing both spatial details and long-range temporal dependencies in surveillance videos. Initially, the input video is pre-processed by resizing each frame using standard resizing methodology with pixel-wise mean and standard deviation normalization to ensure consistency when processing videos. Once the video frames are pre-processed, the next step involves the use of a 3D Convolutional Autoencoder (3D-CNN) with additional processing capabilities through Bi-LSTMs to produce a compact low-dimensional representation of spatiotemporal features in the video frames. The low-dimensional spatiotemporal features are extracted through 3D-CNN which is used to reproduce the Bi-LSTM video-based features. The reconstructed features from the Bi-LSTM were passed into a decoder with a transposed 3D convolutions architecture with skip-connections to once again reconstruct the original frame. Finally, the anomalies are detected by computing the reconstruction error between the original frames and the reconstructed ones. It is clearly shown that the proposed 3D CNN-Bi-LSTM attained better results when compared with 3D CNN in terms of Area Under the Curve (99.8%), True Positive Rate (99.1%) and False Positive Rate (99.7%) respectively.

Read the paper · More papers on PaperTik