Clustering Characteristics of UCI Dataset

Sun Chang, Yue Shihong, Qi Li · 2020

The UCI (University of California Irvine) machine learning repository currently maintain 488 datasets of various characteristics as a service to the machine learning community. Since the past decades, owing to available cluster labels and data attributes, the UCI datasets have been playing an important role in clustering analysis field. Nevertheless, as the common benchmark to evaluate clustering results, high dimensionality and complexity in UCI datasets lead to that main clustering characteristics fail to be recovered, such as shape, size, density, overlap or separation, subspace or component, noise, key point distributions, and so on. Consequently, the evaluation effect of clustering results from UCI is incomplete, uncertain and inconsistent, and the natural characteristics of the tested datasets cannot be effectively found. In this paper, we apply three most used clustering algorithms to recover the clustering characteristics of UCI datasets, including CM(C-Means), DBSCAN (Density-Based Spatial Clustering of Applications with Noise), and DPC (Density Peak clustering). To deal with high-dimensionality and complexity problem in UCI, MDS (Multidimensional Scaling) is used to map high dimensional data to two dimensional Cartesian coordinates under nearly keeping the distance of any pair of points unchangeable. Experimental results show the clustering characteristics of UCI datasets, and this is very helpful to assess clustering results in real applications.

Read the paper · More papers on PaperTik