Unveiling the Impact of Sampling on Feature Selection for Performance Prediction in Configurable Systems
João Marcello Bessal, Millena Cavalcanti, Mathieu Acher, Markus Endler, Juliana Alves Pereira · 2025
Modern software systems are highly configurable, offering a vast number of configuration options that can be customized to meet specific functional and non-functional requirements. To support the configuration process, several automated software approaches based on machine learning have been proposed in the literature. These approaches aim to assist developers by predicting non-functional properties based on configuration settings. A recent study demonstrated the potential of leveraging a subset of configuration options (a.k.a. features) to achieve accurate performance predictions in the Linux kernel. The promise of learning over a reduced set of features – instead of all features – is to obtain performance models that are faster to compute, simpler to interpret, and still accurate. Despite the encouraging results of the original study, several questions remain unresolved: Can the findings be generalized to other configurable systems other than Linux? Which learning algorithms deliver the most efficient results when working with a reduced number of features? What are the most effective sampling strategies for building accurate and efficient models? In this work, we extend the original study by conducting an in-depth analysis across eight configurable systems. We evaluate the impact of sampling strategies and learning algorithms on model accuracy and training efficiency. Our goal is to understand whether there is a dominant sampling strategy and learning algorithm for varying systems and performance targets. Our results reveal variability in optimal strategies across systems and advocate for tailored approaches rather than universal solutions.