Using Clustering and Co-Training to Boost Classification Performance
Antonia Kyriakopoulou · 2007
This paper shows that the performance of a linear SVM classifier can be improved by utilizing meta-information derived from clustering. Clustering aims in discovering extra knowledge concerning the structure of the whole dataset, (both training and testing set). A co-training algorithm is introduced that uses clustering as a complementary step to text classification. At each iteration step of the algorithm the clustering phase augments the feature space with a new meta-feature that for each document reflects cluster membership and the classification phase introduces another meta-feature that indicates class membership. Experimental results obtained using widely used datasets demonstrate the effectiveness of the proposed approaches especially for small training sets.