Co-Training for Topic Classification of Scholarly Data

Cornelia Caragea, Florin Bulgarov, Rada F. Mihalcea · 2015

With the exponential growth of scholarly data during the past few years, effective methods for topic classification are greatly needed.Current approaches usually require large amounts of expensive labeled data in order to make accurate predictions.In this paper, we posit that, in addition to a research article's textual content, its citation network also contains valuable information.We describe a co-training approach that uses the text and citation information of a research article as two different views to predict the topic of an article.We show that this method improves significantly over the individual classifiers, while also bringing a substantial reduction in the amount of labeled data required for training accurate classifiers.

Read the paper · More papers on PaperTik