Comparative Analysis of Large Language Model Tools for Automated Test Data Generation from BDD
Isela Mendoza, Fernando Antonio Grangeiro Filho, Gustavo Medeiros, Aline Paes, Vânia de Oliveira Neves · 2024
Automating processes reduces human workload, particularly in software testing, where automation enhances quality and efficiency. Behavior-driven development (BDD) focuses on software behavior to define and validate required functionalities, using tools to translate functional requirements into automated tests. However, creating BDD scenarios and associated test data inputs is timeconsuming and heavily reliant on a good input data set. Large Language Models (LLMs) such as Microsoft’s Copilot, OpenAI’s ChatGPT-3.5, ChatGPT-4, and Google’s Gemini offer potential solutions by automating test data generation. This study evaluates these LLMs’ ability to understand BDD scenarios and generate corresponding test data across five scenarios ranked by complexity. It assesses the LLMs’ learning, assertiveness, response structuring, quality, representativeness, and coverage of the generated test data. The results indicate that ChatGPT-4 and Gemini stand out as the best tools that met our expectations, showing promise for advancing the automation of test data generation from BDD scenarios.