Improved mining of software complexity data on evolutionary filtered training sets

Vili Podgorelec · 2009

of software engineering. As the dependency of society on software systems increase, so increases also the importance of efficient software fault prediction. In this paper we present a new approach to improving the classification of faulty software modules. The proposed approach is based on filtering training sets with the introduction of data outliers identification and removal method. The method uses an ensemble of evolutionary induced decision trees to identify the outliers. We argue that a classifier trained by a filtered dataset captures a more general knowledge model and should therefore perform better also on unseen cases. The proposed method is applied on a real-world software reliability analysis dataset and the obtained results are discussed. Key-Words:- data mining, classification, evolutionary decision trees, filtering training sets, software fault prediction, search-based software engineering 1

Read the paper · More papers on PaperTik