Co-training based algorithm for datasets without the natural feature split

Јелена Сливка, Aleksandar Kovačević, Zora Konjović · 2010

The performance of a classification model depends not only on the algorithm by which the model is learned, but also on the training set. Manual annotation of the training data is a tedious and time consuming job. In order to overcome the problem of laborious hand-labeling of a large training set, a set of techniques called semi-supervised learning was designed. Co-training is one of the major semi-supervised learning methods. Its setting applies to datasets that have a natural separation of their features into two disjoint sets. However, in the great majority of practical situations, the natural split of features does not exist. In this paper we propose the new co-training based algorithm which can be applied to such datasets.

Read the paper · More papers on PaperTik