Test Case Generation for Requirements in Natural Language - An LLM Comparison Study
Brahma Reddy Korraprolu, Pavitra Pinninti, Y. Raghu Reddy · 2025
The rapid evolution of Large Language Models (LLMs) have opened new possibilities in automating tasks across the software developing life cycle, including test case generation This paper presents a comparative analysis of six LLMs in the context of generating test cases for technical requirements written in natural language (in this case English).We compare publicly available general purpose LLMs viz., BARD, ChatGPT3.5,Claude, Gemini, ChatGPT4.o(Omni) and Llama3.The generated test cases are tested against a Simulink model created for the corresponding set of requirements.The coverage metrics thus generated are used for a quantitative comparison of the LLMs. CCS Concepts• Software and its engineering → Software notations and tools.