AN EFFECTIVE ATTRIBUTE CLUSTERING APPROACH FOR FEATURE SELECTION AND REPLACEMENT
Tzung‐Pei Hong, Po-Cheng Wang, Yeong-Chyi Lee · Cybernetics & Systems · 2009
Feature selection is an important preprocessing step in mining and learning. A good set of features cannot only improve the accuracy of classification, but can also reduce the time to derive rules. It is executed especially when the amount of attributes in a given training data is very large. In this article, an attribute clustering method based on genetic algorithms is proposed for feature selection and feature replacement. It combines both the average accuracy of attribute substitution in clusters and the cluster balance as the fitness function. Experimental comparison with the k-means clustering approach and all combinations of attributes also shows the proposed approach can get a good trade-off between accuracy and time complexity. Besides, after feature selection, the rules derived from only the selected features may usually be hard to use if some values of the selected features cannot be obtained in current environments. This problem can be easily solved in our proposed approach. The attributes with missing values can be replaced by other attributes in the same clusters. The proposed approach is thus more flexible than the previous feature-selection techniques.