Semi-supervised clustering method with constrains
Liu Ying-dong · Computer Engineering and Applications Journal · 2009
In many data mining domains,there is a large supply of unlabeled data but limited labeled data,which can be expensive to generate.Consequently,semi-supervised clustering,which uses a small amount of labeled data to aid unlabeled clustering,has become a topic of significant recent interest.This paper presents a new algorithm,called semi-supervised clustering algorithm based on constrains learning,which obtains the similarity and dissimilarity criterions of data objects,adjusts them in the process of clustering,and uses them to constrain and supervise clustering.Demonstrated the clustering algorithm with Gaussian dataset,and the experimental results confirm that the clustering algorithm significantly improves the accuracy and speed of clustering when given a relatively small amount of supervision.