Label-based semi-supervised fuzzy co-clustering for document categoraization
Yan Tao Yang, Lihui Chen · 2011
Semi-supervised clustering uses a small amount of labeled data to aid and bias the clustering of unlabeled data. In this paper the use of labeled data at the initial state, as well as the use of the constraints generated from the labels during the clustering process is explored. We formulate the clustering process as a constrained optimization problem, and propose a novel semi-supervised fuzzy co-clustering algorithm which incorporated with a few category labels to handle large overlapping text corpus. Simulations on a few large benchmark datasets demonstrate the strength and potentials of this new approach in terms of accuracy, stability and efficiency with limited labels, compared with some existing label-based semi-supervised clustering algorithms.