X2_PDDVnet: An Explainable AI Based Dual Path Dense Dilated Vision Transformer Network Based Anomaly Detection
Shameem Akthar K., K. Lakshmi Priya · Journal Européen des Systèmes Automatisés · 2025
An essential function of video surveillance systems that are widely utilized for public safety and other purposes is automatic anomaly detection.The interpretability of anomaly detection is essential, though, because the kind and severity of the anomalies in the video dictate the appropriate response.Thus, the proposed study presents an explainable AIbased model, X2_PDDVnet, for video anomaly detection with high accuracy in the detection and interpretation of anomalous events.This model uses a two-path structure combining Vision Transformer (ViT) and AlexNet-inspired Convolutional Neural Network (CNN) hrough ZigZag path learning, which enhances feature extraction based on local and global patterns for such complex and dynamic video scenes.This hybrid approach advances video anomaly detection through combining global and local feature extraction to cover the analysis of high-level and granular details in any video frames.The three datasets include UCSD Anomaly Detection (University if Caifornia, San Diego), Avenue, and Shanghai Tech are utilized.Preprocessing of data is done by applying denoising, contrast enhancement, geometric transformations, and normalization to ensure optimized input.After the pre-processing, the model utilizes a dual-path encoder model.The ViT path captures global relationships between video frames by allowing each frame to be a sequence of tokens for the detection of spatial and temporal anomalies.In the meantime, the Zigzag Alex Net path makes use of dilated convolutions for local feature detection for further multi-scale information capture through an embedded Dilated Multi-Scale Inception Network (DiMS-Inception) in every convolution block.This dual-path structure yields a robust feature representation.The Grad-CAM and LIME offer heatmaps for improved visualization, which emphasize areas in crucial frames to make the mathematical model's decision-making process clearer.In this manner, the model succeeds in producing a transparent and interpretable anomaly detection process, which is of utmost importance for practical surveillance applications.The combination of ViT and CNN with the features of interpretability is promising for accurate and reliable anomaly detection in security systems.