Leveraging Large Language Models in Software Testing: A Review of Applications and Challenges

Anıl Sezgin, Gürkan Özkan, Esra Coşgun · 2025

The integration of Large Language Models (LLMs), such as GPT and BERT, into software testing has introduced a transformative approach to ensuring software reliability, functionality, and security. This review explores the diverse applications of LLMs in automating and enhancing software testing processes, addressing challenges in traditional methodologies. LLMs excel in automated test case generation by leveraging vast datasets, user stories, and bug reports to produce comprehensive and context-aware test scenarios. These capabilities reduce manual effort, improve coverage, and identify critical edge cases. Moreover, their application in bug detection and code analysis enhances early issue identification, improving software quality and reducing costs. LLMs also streamline the creation of high-quality documentation, enabling better collaboration and scalability in software projects. Despite their potential, LLM adoption in software testing faces challenges, including model interpretability, scalability, and the need for high-quality, diverse training datasets. Security and ethical considerations, such as data privacy and the risk of misuse, also demand attention. The paper evaluates state-of-the-art LLM applications, showcasing advancements in test automation and vulnerability detection while identifying areas for improvement. Future directions include refining hybrid approaches, enhancing domain-specific adaptability, and addressing ethical and governance concerns. This study provides a comprehensive analysis of LLM-driven testing methodologies, emphasizing their transformative potential in modern software development workflows. By fostering interdisciplinary collaboration, the field can harness LLMs to revolutionize software testing, paving the way for efficient, scalable, and secure software solutions.

Read the paper · More papers on PaperTik