Text classification with enhanced semi-supervised fuzzy clustering
Garima Keswani, Lawrence Hall · 2003
Given the increasing volume of information available on the Web, it is important to meaningfully organize online documents. Hence, the design of efficient and accurate text classification systems is of interest. In this paper, we explore a framework, in which we improve the performance of a base classifier, by clustering unlabeled data with labeled data using probabilistic and fuzzy approaches. We have used expectation maximization and semi-supervised fuzzy c-means for clustering the unlabeled data with labeled data. The naive Bayes classifier was the base classifier utilizing both the original labeled data and then additional data labeled through clustering. Utilizing unlabeled data from semi-supervised fuzzy clustering results in an improved classifier.