Comparative Study of DINOv2, I-JEPA, and ViT Embeddings for Unsupervised Anomaly Detection
Jean-Marc Spat, Houda Chabbi-Drissi · 2025
This paper presents a unified, unsupervised framework for Visual Anomaly Detection (VAD) in dynamic, fixed-view scenes, leveraging modern Vision Transformers (ViTs) without additional fine-tuning. We investigate whether embeddings from backbones like a generic ViT, DINOv2, and I-JEPA, combined with a simple clustering approach, are sufficient for identifying anomalies. Our methodology extracts both global (CLS token) and local (patch-level) embeddings, applies clustering (k-Means, HDBSCAN) to model the distribution of normal scenes, and uses a scalable vector database (ChromaDB) for efficient similarity search. The three backbones delivered comparable performance, with the generic ViT showing a small but consistent advantage in balanced accuracy. Although local embeddings provided a 3 % gain in balanced accuracy, their sequential processing time is considerably higher. This limitation could be mitigated through parallelization, potentially bringing their efficiency closer to that of global embeddings. In contrast, global embeddings deliver comparable performance while being approximately 256 times faster, enabling near real-time batch processing of 15 s video segments. The$\mathbf{k}$-Means clustering and the proposed retrieval strategy proved most effective, achieving a practical operational trade-off with a balanced accuracy of 82 %. These results suggest that ViT embeddings are effective for separating normal and anomalous patterns. Local embeddings appear to capture complementary and pertinent information, and we expect that with modest adjustments in the detection strategy, they could be further leveraged to improve anomaly detection performance.