K-means clustering algorithm based on semi-supervised learning
Mingwei Leng, Xiaoyun Chen, Longjie Li, Mingzhe Leng · Lanzhou University Institutional Repository · 2008
In many data mining domains, there is a large supply of unlabeled data but limited labeled data, which can be expensive to generate. Consequently, semi-supervised clustering, which uses a small amount of labeled data to aid unlabeled clustering, has become a topic of significant recent interest. k-means clustering has been one of the popular, simple and faster clustering algorithms, but the right value of k is unknown and selecting effectively initial points is also difficult. In view of this, we present a new algorithm, called k-means clustering algorithm based on semi-supervised learning, which uses the labeled data to aid initial points selecting and clusters merging effectively, and is not restricted by labeled data. The clustering results by using labeled data and influence factor is more meaningful than unsupervised clustering. In order to obtain a faster algorithm, two theorems are proposed and proofed, and using them can improve the algorithm speed greatly. We demonstrate our clustering algorithm with Gaussian dataset and IRIS dataset, and the experimental results confirm that our clustering algorithm significantly improves the accuracy and speed of clustering when given a relatively small amount of supervision. © 2008 Binary Information Press.