A New Attribute Selection Method Based on Maximal Information Coefficient and Automatic Clustering
Haijin Ji, Song Huang, Yaning Wu, Zhanwei Hui, Xuewei Lv · 2017
Software defect prediction (SDP) plays a significant part in identifying the most defect-prone modules before software testing. Software attributes are the key to improve performance indices in the process of SDP. However, there exists redundant information in those attributes, and this may do harm to the accuracy of predictor. To address this issue, a new attribute selection method (NASM) is proposed in this paper. Firstly, we compute the maximal information coefficient matrix between the attributes in the method, and then these attributes are clustered by spectral clustering according to the maximal information coefficient matrix. In order to implement automatic clustering, Calinski-Harabasz criterion is used to find optimal number of clusters in the process of clustering. Consequently, the best set of attributes is selected. The proposed method is evaluated on five public datasets respectively, and compared with other four attribute selection strategies. The experimental results show that the proposed method NASM outperforms the other three attribute selection strategies on the five standard datasets.