Comparing LLMs and Proposing an ML-Based Approach for Search String Generation in Systematic Literature Reviews
Diogo Adário Marassi, Juliana Alves Pereira, Katia Romero Felizardo · 2025
The formulation of an effective search string is a critical process in systematic literature reviews (SLRs), as it directly influences both the coverage and precision of the retrieved studies. Traditionally, this process relies on manual keyword selection and expert-driven refinements, making it laborious, susceptible to human bias, and often inaccessible to non-specialists. To address these limitations, this study explores the application of artificial intelligence (AI) to support the generation of search strings.We organized our ongoing investigation into two main phases. In the first phase, we evaluated the performance of search strings generated by different large language models (LLMs), specifically Llama-8B, Gemma-12B, and Mistral-Nemo-12B, using a previously published SLR as a benchmark. Our results suggest that, while LLMs can assist in search string formulation, their effectiveness is inconsistent and sensitive to input conditions. Motivated by these limitations, we propose a semi-automated pipeline based on Machine Learning (ML). Through our preliminary analysis, we proposed a standardized reproducible evaluation framework to assess and compare AI-based search string generation strategies, including our proposed ML-based approach.