A survey on multi-modal and weakly supervised approaches for robust anomaly detection in video data

Rui Z. Barbosa, Hugo Santos Oliveira, João Manuel R.S. Tavares · Information Fusion · 2025

This survey provides a comprehensive overview of Video Anomaly Detection (VAD), focusing on identifying robust and interpretable methods for detecting anomalous events. It emphasizes the limitations of traditional supervised and unsupervised approaches and highlights the advantages of weakly supervised learning, in which models operate with minimal or no explicit anomaly annotations. A key contribution of this survey is its analysis of multi-modal data — integrating visual, audio, and textual modalities — for enhanced anomaly detection. It underscores how audio cues enrich visual features and how textual information during training fosters semantically richer representations. This multi-modal approach demonstrates improved generalization, better detection of subtle anomalies, and interpretable explanations of detected events, marking a paradigm shift in the field, Video Anomaly Understanding (VAU). Synthesizing advancements from benchmarks to methodologies, this work advocates for a centralized platform to enable systematic comparisons across diverse datasets, standardized evaluation metrics, and reproducible ablation studies of novel components. Such a framework would streamline the integration of innovations, address version control and foster transparency, bridging the gap between isolated methodological advances and system-level robustness. By prioritizing contextual understanding, causal reasoning, and real-world interpretability, this initiative aims to elevate weakly supervised VAD beyond detection, ensuring models contextualize and explain anomalies in practical deployments.

Read the paper · More papers on PaperTik