Weighing the initial clusters in the ensemble with the help of Heuristic

Mortaza Zolfpour-Arokhlo, Jaafar Partabian · 2015

Abstract—Data clustering means partitioning the samples in similar clusters, in a way that samples in each cluster have the maximum similarity to each other and have the maximum difference with those in other clusters. Due to unsupervised nature of clustering problems, choosing a specific algorithm for clustering an unknown dataset is risky and usually fails its objectives. Because of the complexity of this problem and poor performance of basic clustering methods, today the majority of studies in this field are focused on ensemble clustering methods. Diversity and quality of initial results are two of the most important factors that can affect the quality of final results obtained from ensemble clustering, and both factors have been significantly assessed in the recent studies regarding this field. In this paper, a new framework is proposed to improve the efficiency of ensemble clustering. This framework is based on using a subset of initial clusters and selection of this subset plays a crucial role in the performance of the ensemble, so the process of selection is carried out with the help of two intelligent methods. The main idea of the proposed method for the selection of a subset of clusters is to use the stable clusters with the help of intelligent search algorithms. The stability criterion based on mutual information is used to evaluate the clusters. In the end, the selected clusters are aggregated with the help of multiple final clustering techniques. Experimental results obtained from testing several standard datasets show that the proposed methods can effectively improve the full ensemble method. Index Terms — clustering ensemble, subset of initial results, correlation matrix, cluster evaluation 1.

Read the paper · More papers on PaperTik