Learning the Semantic Correlation: An Alternative Way to Gain from Unlabeled Text
Yi Zhang, Artur W. Dubrawski, Jeff Schneider · 2008
In this paper, we address the question of what kind of knowledge is gen-erally transferable from unlabeled text. We suggest and analyze the se-mantic correlation of words as a generally transferable structure of the language and propose a new method to learn this structure using an ap-propriately chosen latent variable model. This semantic correlation con-tains structural information of the language space and can be used to control the joint shrinkage of model parameters for any specific task in the same space through regularization. In an empirical study, we con-struct 190 different text classification tasks from a real-world benchmark, and the unlabeled documents are a mixture from all these tasks. We test the ability of various algorithms to use the mixed unlabeled text to en-hance all classification tasks. Empirical results show that the proposed approach is a reliable and scalable method for semi-supervised learn-ing, regardless of the source of unlabeled data, the specific task to be enhanced, and the prediction model used. 1