Predicting Software Effort Estimation Using Machine Learning Techniques

Ahmed BaniMustafa · 2018

In software engineering, estimation plays a vital role in software development. Thus, affecting its cost and required effort and consequently influencing the overall success of software development. The error margin in Expert-Based, Analogy-Based and algorithmic based methods including: COCOMO, Function Point Analysis and Use-Case-Points is quite significant, which exposes software projects to the danger of delays and running over-budget. To obtain better estimation, we propose an alternative method through performing data mining on historical data. This paper suggests performing this prediction using three machine learning techniques that were applied to a preprocessed COCOMO NASA benchmark data which covered 93 projects: Naïve Bayes, Logistic Regression and Random Forests. The generated models were tested using five folds cross-validation and were evaluated using Classification Accuracy, Precision, Recall, and AUC. The estimation results were then compared to COCOMO estimation. All the applied techniques were successful in achieving better results than the compared COCOMO model. However, the best performance was obtained using both Naïve Bayes and Random Forests. Despite the fact that Naïve Bayes outperformed both of the other two techniques in its ROC curve and Recall score, Random Forests has a better Confusion Matrix and scored better in both Classification Accuracy, and Precision measures. The results of this work confirm the validity of data mining in general and the applied technique in particular for software estimation.

Read the paper · More papers on PaperTik