Artifacts - Readable but Risky: Evaluating LLM-Generated Evidence Briefings for Software Engineering
Mauro Marcelino, Marcos Antônio Alves, Bianca Trinkenreich, Bruno Cartaxo, Soares, Sérgio, Marcos Kalinowski · arXiv (Cornell University) · 2026
This repository contains all the files necessary for conducting the study designed in the registered report titled "Investigating the Use of LLMs for EvidenceBriefings Generation in Software Engineering". Abstract: Evidence briefings condense the main findings of empirical software engineering studies into a concise and objective format for practitioners. Although widely deemed useful, their manual production is highly labor-intensive, representing a significant barrier to broad adoption. Large Language Models (LLMs) offer a scalable alternative to automate research synthesis, but their reliability in this context remains uncertain. The goal of this study is to evaluate LLM-generated evidence briefings against human-authored baselines regarding content fidelity, ease of understanding, and perceived usefulness. We conducted a controlled crossover experiment, gathering quantitative and qualitative data from 82 industry practitioners (assessing understanding and usefulness) and an expert panel of three researchers (assessing content fidelity). Our results show that AI-generated briefings achieved a statistically significant advantage in Structure, effectively reducing cognitive load, while performing equally well in clarity and conciseness. Furthermore, practitioners rated both treatments as equally useful, revealing that practical applicability is fundamentally dictated by the reader's organizational context. However, the expert evaluation revealed that the LLM significantly degraded precision, generating a \textit{Certainty Illusion} by generalizing context-specific academic findings and fabricating unsupported practical recommendations. We conclude that while LLMs excel at structural design and clarity, their tendency to mask scientific uncertainty prevents autonomous deployment. We highlight the strict necessity for rigorous human-in-the-loop verification steps when employing LLMs for evidence synthesis in software engineering.