Evaluating Multi-Modal LLMs for Automatically Recognizing Semantic Elements in UML Use Case Diagram Images
Jameleddine Hassine · 2025
Requirements engineering commonly employs UML Use Case Diagrams (UCD) to visually capture system interactions and functionality, facilitating clear communication between stakeholders. Recognizing and extracting semantic information from UCDs is essential for applications such as automated requirements extraction and system design validation, which improves software analysis accuracy, and streamlines model understanding for both developers and stakeholders. Recent advancements in large language models (LLMs) with visual processing capabilities enable interpreting intricate diagrammatic content. This paper evaluates multi-modal LLMs, specifically GPT-4o and GPT-4o-mini, in accurately identifying semantic elements within UCDs. We conducted experiments on a new dataset of UCDs and other diagrams collected from online sources. Experimental results show that both models struggled to accurately identify and interpret key UCD elements, often misclassifying or overlooking essential ones.