Feature selection for clustering problems: a hybrid algorithm that iterates between k-means and a Bayesian filter

Eduardo R. Hruschka, Eduardo R. Hruschka, Thiago Ferreira Covões, Nelson F. F. Ebecken · 2005

There are two fundamentally different approaches for feature selection: wrapper and filter. It is also possible to combine them, obtaining hybrid approaches. This paper describes a hybrid method for selecting relevant features in clustering problems. The proposed approach is based on the combination of the widely known k-means algorithm and a Bayesian filter, which is based on the Markov Blanket concept. Since the number of clusters and the subset of relevant features are usually inter-related, we propose a method that iterates between clustering (assuming that the number of clusters is not known a priori) and filtering. Experiments in a number of datasets show that the proposed approach allows selecting features that provide good partitions.

Read the paper · More papers on PaperTik