Determining the number of clusters and distinguishing overlapping clusters in data analysis (french text)

Haojun Sun · 2005

Data clustering is the process of grouping the data into clusters so that the objects within a cluster are highly similar and the objects in different clusters are highly dissimilar. The main focus of this thesis is to investigate into two important problems in clustering: determining the number of clusters in a given data set and studying the phenomenon of overlapping between clusters. Determining the number of clusters is one of the most important topics in cluster analysis. A common approach for determining the number of clusters is an iterative trial-and-error process based on a cluster validity index. One of the main goal of this thesis is to develop a new validity index for measuring the goodness of trial-clustering and an effective fuzzy algorithm for automatically determining the number of clusters. An application of the new algorithm in subset feature selection is proposed also. The phenomenon of cluster overlap is present in real applications. Many algorithms fail to distinguish overlapping clusters. In this thesis, we establish a theory on the overlap phenomenon in the case of the Gaussian mixture, a fundamental data distribution model for many clustering algorithms. Based on this theory, we develop an algorithm for calculating the overlap rate between two clusters and investigate factors that affect the value of the overlap rate. We show how the theory can be used to generate truthed data sets for evaluating the ability of a validity index for distinguishing overlapping clusters. Another application of the theory to be shown is a hierarchical clustering algorithm for color image segmentation.

Read the paper · More papers on PaperTik