ON PREDICTING SOFTWARE DEVELOPMENT EFFORT USING MACHINE LEARNING TECHNIQUES AND LOCAL DATA
Łukasz Radliński, Władysław Hoffmann · 2010
This paper analyses the accuracy of predictions for software development effort using various machine learning techniques. The main aim is to investigate the stability of these predictions by analyzing if particular techniques achieve a similar level of accuracy for different datasets. Two key assumptions are that (1) predictions are performed using local empirical data and (2) very little expert input is required. The study involves using 23 machine learning techniques with four publicly available datasets: COCOMO, Desharnais, Maxwell and QQDefects. The results show that the accuracy of predictions for each technique varies depending on the dataset used. With feature selection most techniques provide higher predictive accuracy and this accuracy is more stable across different datasets. The highest positive impact of feature selection on the accuracy has been observed for the K * technique, which has generated the most accurate predictions across all datasets.