Training-Free VLM-Based Pseudo Label Generation for Video Anomaly Detection

Moshira Abdalla, Sajid Javed · IEEE Access · 2025

In this study, we present a novel framework for Weakly Supervised Video Anomaly Detection (WSVAD) that leverages the vision-language alignment capabilities of the pre-trained CLIP model. Our approach enables pseudo-label generation for fine-grained and coarse-grained classification without any training. The framework employs a triple-branch architecture: the first branch generates pseudo-labels using a similarity-matching mechanism with a threshold-based strategy, while the second and third branches perform coarse-grained binary classification and fine-grained categorical classification, respectively. To enhance video feature representation, we utilize a transformer model to capture short-range dependencies and Graph Convolutional Networks (GCNs) to model long-range temporal relationships. Extensive experiments on the UCF-Crime and XD-Violence datasets validate the effectiveness of our framework, demonstrating competitive performance in both coarse-grained and fine-grained anomaly detection tasks. Additionally, in zero-shot testing on the newly introduced MSAD dataset, our framework achieves top-ranking performance, highlighting its adaptability and robustness.

Read the paper · More papers on PaperTik