Semi-Supervised Classification Method Based on Spectral Clustering
Xi Chen · Journal of Networks · 2014
Abstract—With the rapid development of data collection and storage technology, there are plentiful unlabeled data but very few and often expensive labeled data in real-word applications. Thus, semi-supervised learning algorithms have attracted much attention. In this paper, we propose a new semi-supervised classification algorithm benefiting from spectral clustering called SC-SSL. First, we introduce spectral clustering to partition all labeled and unlabeled data into clusters. Second, we build a classifier using all labeled data and predict the probabilities (weights) of classes that each unlabeled instance belongs to for each cluster. Third, for each cluster, we add those unlabeled instances whose labels with the maximum weights as same as the cluster label into the labeled data. Fourth, in terms of the new labeled data set, we reconstruct the classifier. We repeat the above processing of steps 2 and 3 till meeting the stopping condition. Finally, extensive experiments reveal that our SC-SSL algorithm can sufficiently use the information of unlabeled data to get a robust classifier by spectral clustering, and it maintains a higher classification accuracy compared to several well known semi-supervised algorithms.