Large-Scale Simulation Study of Active Learning Models for Systematic Reviews
Jelle Jasper Teijema, Jonathan De Bruin, Ayoub Bagheri, Rens van de Schoot · 2023
This study advocates for large-scale simulations as the gold standard for assessing active learning models for the prioritization of screening order in systematic reviews. The use of active learning to prioritize potentially relevant records in the screening phase of systematic reviews has seen considerable progress and innovation. This rapid development, however, has highlighted the disparity between the development of these methodologies and their rigorous evaluation, stemming from constraints in simulation size, lack of infrastructure, and the use of too few datasets.In this study, two large-scale simulations evaluate active learning solutions for systematic review screening, involving over 29 thousand simulations and over 150 million data points. These simulations are designed to provide robust empirical evidence of performance. The first study evaluates 13 combinations of known well-performing classification models and feature extraction techniques such as TF-IDF, SVM, Random Forest, and more across high-quality datasets sourced from the SYNERGY dataset. This baseline is then used to expand the scope of the evaluation further in the second study by incorporating a wider array of classification models and feature extractors such as FastText, pretrained transformer models, and XGBoost, for a total of 92 model combinations.The spectrum of performance varies considerably between datasets, models, and screening progression, from marginally better than random reading to near-flawless results. Still, every single model-feature extraction combination outperforms random screening. Results are publicly available for analysis and replication.