Weakly Supervised Video Anomaly Detection Using Dynamic-Weighted Feature Fusion
Hyeonjeong Lim, Donghyeong Kim, Minjung Kim, Chaewon Park, Dongwoo Kang, Sangyoun Lee · 2024
Recent studies have shown that using video and text together works well in weakly supervised video anomaly detection (WSVAD) tasks. However, existing methods often struggle with the presence of both normal and abnormal frames within abnormal videos, leading to suboptimal training and inaccurate anomaly predictions. Additionally, the misalignment of video and textual representations can result in the loss of important contextual information, further hindering performance. To address these issues, we propose a novel approach for weakly supervised video anomaly detection (WSVAD) using dynamic-weighted feature fusion (DWFF). Leveraging vision-language models (VLMs) such as contrastive language-image pre-training (CLIP), we enhance anomaly detection by aligning video and textual representations. Our approach introduces dynamic weighting that emphasizes abnormal frames during training and employs a fusion of video-level and frame-level features to improve anomaly prediction accuracy. We combine frame-level similarities with class embeddings and video-level features, enhancing the model's ability to capture temporal dependencies and refine anomaly detection. We evaluate our method on two benchmark datasets, the UCF-Crime and the XD-Violence, demonstrating its superiority over existing state-of-the-art methods. Our results show significant improvements in detection accuracy, highlighting the effectiveness of dynamic weighting and feature fusion techniques in WSVAD tasks.