A method of incomplete data three-way clustering based on density peaks

Lin Yang, Kaiyan Hou · AIP conference proceedings · 2018

In the study of incomplete data cluster, most scholars consider only the similarity between samples but rarely consider the size distribution of subclasses. When the subclass distribution is unevenly distributed, the edge samples of large distribution clusters are likely to be divided into adjacent small distribution clusters. In order to better cope with edge samples of class clusters, this essay, based on density peak cluster algorithm [1], puts forward a method of incomplete data three-way decision cluster based on density peaks. Firstly, on the basis of sample loss rate, sample sets are divided into three sets: a set of complete samples, a set of samples with low loss rate, and a set of samples with high loss rate. Secondly, samples with low loss rate are filled into complete samples. And cluster algorithm with density peaks is applied on the complete data set to get the original cluster result. Edge samples are observed and rearranged according to membership grade to obtain a three-way decision clustering result. Finally, samples with high loss rate are divided into the boundary region of the closest cluster and a final clustering result is gained. The result of experiment shows that the method put forward in the essay is effective.

Read the paper · More papers on PaperTik