A Closeness-Based Semi-Supervised Text Classification Method

Junyu Niu · Zhongwen xinxi xuebao · 2007

Automatic text categorization has become a very important research area.In most applications,there's only a positive document set with a limited size and a large portion of unlabeled data in the training set while the distribution of the number of the positive set and the negative set is also unbalanced.So,this kind of text categorization task is different from those traditional ones which have not only labeled positive but also labeled negative samples in its training set.Those traditional classification methods can not be directly used in such tasks.This paper proposed a closeness-based method to solve this semi-supervised text categorization problem.It firstly extracts a reliable negative set from the unlabeled set,and then uses the closeness-based algorithm to enlarge initially extracted reliable negative set to a proper size.Based on the labeled positive set and the extracted negative set,the classifier will be constructed.This method will improve the performance of the classifier without any outside resources to help the feature selection,so,it can be used in a lot semi-supervised text categorization tasks in different domains.The experiment on TREC'05 Genomics track data shows that this algorithm performs well in this kind of text categorization tasks.

Read the paper · More papers on PaperTik