Feature selection and replacement by clustering attributes
Tzung‐Pei Hong, Yan-Liang Liou, Shyue-Liang Wang, Bay Vo · Vietnam Journal of Computer Science · 2013
Feature selection is to find useful and relevant features from an original feature space to effectively represent and index a given dataset. It is very important for classification and clustering problems, which may be quite difficult to solve when the amount of attributes in a given training data is very large. They usually need a very time-consuming search to get the features desired. In this paper, we will try to select features based on attribute clustering. A distance measure for a pair of attributes based on the relative dependency is proposed. An attribute clustering algorithm, called Most Neighbors First, is also presented to cluster the attributes into a fixed number of groups. The representative attributes found in the clusters can be used for classification such that the whole feature space can be greatly reduced. Besides, if the values of some representative attributes cannot be obtained from current environments for inference, some other possible attributes in the same clusters can be used to achieve approximate inference results.