AI-Powered Unit Test Generation via Multi-LLM Chaining: A Case Study With GPT-4o, Gemini, and Claude-3.5
Chandan Kumar, Usha Sri Ponaka, Pula V Lakshmi Narasimha Naidu, P Bhuvaneswari, Mukesh Prasad, Sunil Kumar Singh · IEEE Access · 2025
Software testing is a critical activity in the software development cycle since it confirms code correctness, reliability, and maintainability. Unit testing, a cornerstone of software testing, involves verifying the correctness of individual components of a program. Manually writing these tests is a resource-intensive and time-consuming activity. Researchers have proposed various automated test generation methods to reduce the possibility of this problem. Large Language Models (LLMs) have demonstrated remarkable efficacy in automatically generating unit tests in recent years. Although a single LLM configuration can produce a satisfactory response, there are possible risks to the effectiveness of the generated tests, such as the repetition of specific test cases, the lack of edge case coverage, and the omission of full assertions. We present a novel method, LLM Chaining, that uses the collaboration of several LLMs, namely Gemini, GPT-4o, and Claude-3.5 Sonnet, to collaborate and iteratively enhance tests. We started by looking into a single LLM writing unit test, but after seeing how well they performed, we looked into LLM chaining as a way to increase the accuracy and comprehensiveness of the tests that were produced. Gemini and GPT-4o yielded the most reliable results of all the configurations we tested. Under this configuration, the LLM system uses Gemini to generate the initial JUnit tests, then sends the generated results to GPT-4o for assimilation of accuracy and depth. The HumanEval dataset was used to create JUnit test cases for this study. JaCoCo was used to collect coverage statistics, and PIT (Pitest) was used for mutation-based testing to gauge an LLM’s capacity to identify test flaws. With 99.05% branch coverage, 90.48% line coverage from 3508 test runs, and 94.32% mutation coverage, we achieved genuinely impressive results. In contrast, under the same circumstances, Randoop, a well-known automated test generation tool, obtained 68.64% branch coverage and 78.84% line coverage. This demonstrates how multi-step LLM refinement works well to advance automated test generation and, in the end, produce software that is more dependable and maintainable.