Data‐driven benchmarking in software development effort estimation: The few define the bulk
Nikolaos Mittas, Lefteris Angelis · Journal of Software Evolution and Process · 2020
Abstract Context The rapid evolvement of software development effort estimation models created the need for empirical evaluation of their quality. The empirical evaluation is based either on hypothesis tests with respect to a single criterion or on aggregating methods for multiple criteria. However, a model can be considered as a multidimensional entity performing differently on alternative datasets and its performance can be divergent when expressed by alternative criteria. Objective In this study, we explore this multidimensional nature of models by considering them as points in two different spaces (domain and criteria spaces). Method Introducing an alternative approach for data‐driven benchmarking, a new framework based on archetypal analysis is proposed for evaluation purposes of multiple models. Results The benefits of the framework are illustrated through a large‐scale experimental setup on a set of 93 effort estimation models, trained and tested on 10 datasets under 8 criteria providing answers to critical research questions. Conclusion The results indicate that a small minority of reference models is enough to define the performance of the bulk of all models. The framework focuses on models that have behavior close to archetypes and especially those that are close to a “best” archetype.