A study on multi-modal LLM reasoning for defect detection
Andrei Alexandru Tulbure, Dermina-Petronela Danciu, Eva Henrietta Dulf, Adrian Alexandru Tulbure · 2024
In the modern landscape of manufacturing, the early and accurate detection of defects is pivotal to maintaining product quality and operational efficiency. This study investigates the application of multi-modal Large Language Models (LLMs) in enhancing defect detection processes. Leveraging the capabilities of LLMs to interpret diverse data types, such as text and visual imagery, this research explores how these models can provide comprehensive defect analysis, and how this analysis can be used for more than just detection. Through our experiments, we demonstrate that multi-modal LLMs have the potential to detect complex visual defects without any training, although with less accuracy and speed than trained convolutional neural network-based models. Moreover, the general context and text-based communication can be useful for deriving further analytics for the production process and for training purposes. On our ceramic defect dataset, the LLMs correctly detected, but not localized, between 10% and 33% of the defects present in the analyzed images, proving that large pre-training is very comprehensive versus a mean average precision (mAP@IoU 0.5) of 0.531 for a YOLO based model. The study also delves into the challenges of integrating different data modalities and the real time constraints of defect detection with are not yet solved with multimodal LLMs, while providing legal context.