Measuring the Effect on Prediction Accuracy When Data Transformation is Overlooked by Predicting the Development Effort of Software Projects

Cuauhtémoc López-Martín, María Elena Meda-Campaña · 2022

The software effort prediction spent for teams of developers is a main activity of the software planning. When a new prediction model is proposed, a common guideline is to transform data when there exists skewness, heteroscedasticity, or outliers. However, its performance could result awful if actual data is used for a software manager. Thus, in the present study, we empirically measured that impact on the prediction accuracy when nontransformed data value is used in the explanatory variable. Since statistical regression equation has been one of the models mostly used in the mentioned activity, in the present study, eight of them are generated from regression analysis. They were generated involving 1,500 software projects selected from an international public repository (i.e., ISBSG) and applying a hold-out cross validation method. The selection of each the eight data sets observed the guidelines suggested for the ISBSG. The Magnitude of Relative Error (MRE) and Pred$(l)$were used as prediction accuracy measures. We can conclude that, although predictions can seem trustworthy, a software manager should be careful about using nontransformed data in prediction models when these have been generated from transformed data.

Read the paper · More papers on PaperTik