Multi-view Document Classification with Co-training
Semih Sevım, Ekin Ekıncı, Sevinç İlhan Omurca · 2020
The purpose of the document classification is to assign the most proper label to a document. In real world data, insufficiency of labeled data is a major problem. In this study, a semi-supervised co-training algorithm, which enables classification on a small number of labeled data, is applied on multi-views obtained from 20 Newsgroup datasets. Four different multi-views are obtained using word embedding vectors. Logistic regression and random forest methods are used as the base classification methods. As a result of the experiments, co-training algorithm is found to be more successful than the base classifiers.