Fuzzy Clustering of Incomplete Data
Ludmila Himmelspach · 2016
Clustering is one of the important and primarily used techniques for the automatic knowledge extraction from large amounts of data. Its task is identifying groups, so-called clusters, of similar objects within a data set. Clustering methods are used in many areas, including database marketing, web analysis, information retrieval, bioinformatics, and many others. However, if clustering methods are applied on real data sets, a problem that often comes up is that missing values occur in the data sets. Since traditional clustering methods were developed to analyze complete data, there is a need for data clustering methods handling incomplete data. Approaches proposed in the literature for adapting the clustering algorithms to incomplete data work well on data sets with equally scattered clusters. In this thesis we present a new approach for adapting the fuzzy c-means clustering algorithm to incomplete data that takes the scatters of clusters into account. In the experiments on artificial and real data sets with differently scattered clusters we show that our approach outperforms the other clustering methods for incomplete data. Since the quality of the partitioning of data produced by the clustering algorithms strongly depends on the assumed number of clusters, in the second part of the thesis we address the problem of finding the optimal nummber of clusters in incomplete data using cluster validity functions. We describe different cluster validity functions and adapt them to incomplete data according to the ``available-case'' approach. We analyze the original and the adapted cluster validity functions using the partitioning results of several artificial and real data sets produced by different fuzzy clustering algorithms for incomplete data. Since both the clustering algorithms and the cluster validity functions are adapted to incomplete data, our aim is finding the factors that are crucial for determining the optimal number of clusters on incomplete data: the adaption of the clustering algorithms, the adaption of the cluster validity functions, or the loss of information in the data itself. Discovering clusters of varying shapes, sizes and densities in a data set is more useful for some applications than just partitioning the complete data set. As a result, density-based clustering methods become more important. Recently presented approaches either require the input parameters involving the information about the structure of the data set, or are restricted to two-dimensional data. In the last part of the thesis, we present a novel density-based clustering algorithm, which uses the fuzzy proximity relations between the data objects for discovering differently dense clusters without any a-priori knowledge of a data set. In experiments, we show that our approach is able to correctly detect the clusters closely located to each other and clusters with wide density variations.