An Agglomerative Clustering Methodology For Data Imputation
Sumanth Yenduri · 2006
The prediction of accurate effort estimates from software project data sets still remains to be a challenging problem. Major amounts of data are frequently found missing in these data sets that are utilized to build effort/cost/time prediction models. Current techniques used in the industry ignore all the missing data and provide estimates based on the remaining complete information. Thus, the very estimates are error prone. In this paper, we investigate the design and application of a hybrid methodology on six real-time software project data sets in order to better the prediction accuracies of the estimates. We perform useful experimental analyses and evaluate the impact of the methodology. Finally, we discuss the findings and elaborate the appropriateness of the methodology