Weakly Supervised Video Anomaly Detection Based on Adaptive Fusion of Multimodal Features
Weichao Dang, Xuzhen Mao, Lihu Pan · International Journal of Pattern Recognition and Artificial Intelligence · 2025
Existing Video Anomaly Detection (VAD) methods predominantly rely on a single visual modality and fail to fully leverage multimodal information due to limited available datasets. However, recent advances in video understanding models have enabled multimodal integration into VAD. In this paper, we propose a novel weakly supervised video anomaly detection network based on adaptive fusion of multimodal features (AFM-WSVAD). This network introduces textual information corresponding to the video through the caption generation model SwinBERT. In VAD, effectively modeling long-term temporal dynamics is crucial. Therefore, we designed a two-stage approach to capture this dynamic information. First, the External Attention mechanism is employed to capture long-range dependencies between samples. Then, the Temporal Context Aggregation (TCA) module and the Multi-scale Temporal Network (MTN) are utilized to model the short-term temporal dynamics of visual and textual features. This design enables the model to handle long-range dependencies and complex dynamics. During multimodal fusion, an adaptive fusion strategy and a Multi-scale Convolutional Attention (MSCA) module are employed to highlight key features and minimize noise, thereby enhancing the model’s detection accuracy. AFM-WSVAD achieves significant performance improvements, with an AUC of 85.4% on UCF-Crime ([Formula: see text]%), 98.1% on ShanghaiTech ([Formula: see text]%), and an AP of 81.8% on XD-Violence ([Formula: see text]%).