Low-shot Visual Anomaly Detection with Multimodal Large Language Models

Tobias Schiele, Daria Kern, Anjali DeSilva, Ulrich Klauck · Procedia Computer Science · 2024

This work investigates the potential of Multimodal Large Language Models (MLLM) for visual anomaly detection by framing the problem as a visual question-answering task. We evaluate open-source and closed-source models, as well as zero-shot, one-shot, and few-shot strategies. Additionally, we compare the results with the current state-of-the-art for low-shot anomaly detection. A baseline for solving the MVTec AD dataset on an image level with unmodified MLLM is established: One-shot prompting GPT-4V achieves close-to-state-of-the-art results with an F 1 score of 92%, whereas the open-source model Qwen-VL-Chat behaves close to a random classifier in the one-shot scenario.

Read the paper · More papers on PaperTik