Base, Instruct, and Fine-Tune: Evaluating LLMs for Cross-Language Flaky Test Detection
Azeem Ahmad, Xin Sun · IEEE Access · 2025
Flaky tests exhibit non-deterministic outcomes posing a critical challenge to the reliability of continuous integration (CI) pipelines. While prior approaches rely on static token level features and vocabulary-based heuristics, they often fail to generalize across languages and project domains. In this work, we explore the use of Large Language Models (LLMs) for detecting flaky tests across five programming languages: Java, Python, C++, Go, and JavaScript.We evaluate three families of LLMs: LLaMA3, Mistral, and DS-Coder, in their base, instruction-tuned, and fine-tuned forms. Our empirical results demonstrate that fine-tuned LLMs consistently outperform baseline and vocabulary-driven methods, achieving near perfect F1 scores on high resource languages and substantial improvements in low resource settings. Instruction-tuned variants offer notable benefits, especially in early stage adaptation. Moreover, our cross-language evaluation reveals that LLMs trained on one language can moderately generalize to others, with fine-tuning significantly narrowing the performance gap. The findings highlight the practical value of LLMs as language agnostic flaky test detectors and provide guidance on model adaptation strategies for real world CI/CD environments. This work sets the foundation for future research on generalizable and semantically aware testing tools in multilingual software ecosystems.