Uncovering the limits of visual-language models in engineering knowledge representation

Marco Consoloni, Vito Giordano, Federico A. Galatolo, Mario G. C. A. Cimino, Gualtiero Fantoni · Proceedings of the Design Society · 2025

ABSTRACT: Visual-Language (VL) models offer potential for advancing Engineering Design (ED) by integrating text and visuals from technical documents. We review VL applications across ED phases, highlighting three key challenges: (i) understanding how functional and structural information is complementarily expressed by text and images, (ii) creating large-scale multimodal design datasets and (iii) improving VL models’ ability to represent ED knowledge. A dataset of 1.5 million text-image pairs and an evaluation dataset for cross-modal information retrieval were developed using patents. By Fine-tuning and testing the CLIP base model on these datasets, we identified significant limitations in VL models’ capacity to capture fine-grained technical details required for precision-driven ED tasks. Based on these findings, we propose future research directions to advance VL models for ED applications.

Read the paper · More papers on PaperTik