Few-Shot Evaluation of Vision Language Models for Detecting Visual Defects in Autonomous Vehicle Software Requirement Specifications
Nabil Bukhary, M. Omair Ahmad, Karim Rashad, Samarth Rai, Salsabeel Shapsough, Yara Kaddoura, Dana Dghaym, Imran Ahmed Zualkernan · IEEE Access · 2025
Software Requirements Specifications (SRS) are crucial for defining system functionality, constraints, and objectives, particularly in safety-critical applications like autonomous vehicles (AV). Ensuring these requirements are precise, unambiguous, and complete is important for developing reliable self-driving systems. While traditional SRS are predominantly text-based, AV requirements pose unique challenges as they often include visual elements such as images and diagrams. Although Large Language Models (LLMs) have enhanced traditional text-based requirements, Vision-Language Models (VLMs) remain largely unexplored for evaluating defects in visual elements. This paper investigates the few-shot performance of three large-scale VLMs on their ability to detect various types of ambiguities, inconsistencies, and incompleteness in visual traffic scenario images that could be used in AV requirements specifications. Using Soft prompting and Chain-of-Thought (CoT) prompting, we evaluated the effectiveness of GPT-4, GEMINI, and ClaudeAI in detecting six types of visual defects. GPT-4 achieved the highest overall performance among the evaluated models, with the highest F1-score of 0.63 in detecting Spatial Incompleteness. CoT prompting notably enhanced defect detection performance for both GEMINI and ClaudeAI across all defect types. While CoT prompting did not significantly improve GPT-4’s detection capabilities, it significantly enhanced the ability of all models to explain the reasoning behind the identified defects. Despite promising results, none of the models could fully automate the requirements validation process, primarily due to their lack of domain-specific knowledge. Nevertheless, our findings suggest that fine tuning the models and structured prompt design offer the potential to improve VLM precision in visual defect detection.