Machine Learning Tools To (Semi-)Automate Evidence Synthesis: A Rapid Review and Evidence Map
Gaelen P. Adam, Melinda C. Davies, Jerusha George, Eduardo Lucia Caputo, Ja Mai Htun, Erin L. Coppola, Haley K Holmer, Edi Kuhn, Holly R. Wethington, Ilya Ivlev, Ethan Balk, Thomas A Trikalinos · 2025
Introduction. Tools that leverage machine learning, a subset of artificial intelligence, are becoming increasingly important for conducting evidence synthesis as the volume and complexity of primary literature expands exponentially. In response, we have created a living rapid review and evidence map to understand existing research and identify available tools. Methods. We searched PubMed, Embase, and the ACM Digital Library from January 1, 2021, to April 3, 2024, with update searches on October 3, 2024, and April 3, 2025, for comparative studies, and identified older studies using the reference lists of existing evidence synthesis products (ESPs). We plan update searches every 6 months. We included evaluations of machine learning or artificial intelligence tools to automate or semi-automate any stage of systematic review production. Title and abstract screening was conducted independently by two reviewers, with disagreements resolved through discussion or adjudication by a third reviewer. Full-text screening and data extraction were performed by a single reviewer. We did not assess the quality of individual studies or the strength of evidence across studies. Extracted data included key characteristics of the tools (e.g., type of automation method, systematic review tasks automated), methods used to evaluate the tool performance, and performance results. The protocol was prospectively registered on the AHRQ website. Results. We included 95 studies: 56 in the original report, 21 in the October 2024 interim update, and 18 in the April 2025 update. These studies evaluated the performance of tools primarily relative to standard human processes across various systematic review tasks. Tools for identifying randomized controlled trials performed well, with a median recall of 96% and precision of 79%. Abstract screening tools also showed promising, though variable, results, achieving a median recall of 85% for fully automated screening using zero-shot models; 97% for semi-automated models with a median 51% reduction in screening burden. In contrast, tools for searching had low recall (median 14%) and precision (median 0.09), and data extraction tools varied widely, with a median 66% of data correctly extracted. Risk of bias assessment tools showed moderate agreement with human assessments (Cohen's kappa = 0.20; median agreement: 71%). Discussion. Certain tools, particularly those for automatically identifying randomized controlled trials (RCTs) and prioritizing relevant abstracts in screening, show a high level of recall and precision, suggesting they may soon be appropriate for widespread use with human oversight. However, other tools, such as those for searching and data extraction, show highly variable performance and are not yet reliable enough for to be used for semi-automation of these tasks. These conclusions are largely unchanged from earlier iterations, with the note that commercial large language models (e.g., ChatGPT and Claude) are demonstrating improved performance in fully automating abstract screening and in data extraction. This work revealed the importance of developing standardized evaluation frameworks for assessing the performance of machine learning and artificial intelligence tools in systematic review tasks. We did not assess the risk of bias or methodological quality of the included studies, which may affect the reliability and comparability of the reported performance outcomes. Additionally, the tools were evaluated in a variety of settings, tasks, and review questions, which introduces heterogeneity that makes direct comparisons across tools challenging. Lastly, the rapidly evolving nature of machine learning technologies means that our findings may quickly become outdated. Therefore, we plan ongoing updates every 6 months.