Cross-Project Estimation of Software Development Effort Using In-House Sources and Data Mining Methods - an Experiment

Hrvoje Karna, Ana Masnov, Darija Jurko, Tomislav Peric · 2019

Application of data mining methods is well suited for problems of effort estimation in the field of software engineering. During development process, data mining can provide software engineers with valuable inputs that support decision makings. This way it is possible to overcome the problems caused by erroneous estimates made by the experts. Use of in-house data sources is encouraged because predictive models built on top of them typically provide better estimates compared to the models generated by using data from other sources. Model predictions can either confirm experts statements or propose an alternative solution. This way software development process can become more efficient. In this empirical investigation data from five different software projects originating from the same environment were used to conduct a formal experiment. The experiment uses cross-project estimation in which data sets from other projects are used to build predictive models for the project being estimated. Predictive models implement advanced learners that in general provided more accurate predictions, thus reducing the estimation error. The evaluation of results obtained during data mining process uses established criteria. The organization of the research carried out, distinctive model and structuring of the data together with obtained results encourage the application of similar models in practice.

Read the paper · More papers on PaperTik