Learning on the test data: leveraging Unseen features
Ben Taskar, Ming Fai Wong, Daphne Koller · 2003
This paper addresses the problem of classification in situations where the data distribution is not homoge-neous: Data instances might come from different lo-cations or times, and therefore are sampled from re-lated but different distributions. In particular, features may appear in some parts of the data that are rarely or never seen in others. In most situations with non-homogeneous data, the training data is not representa-tive of the distribution under which the classifier must operate. We propose a method, based on probabilistic graphical models, for utilizing unseen features during classification. Our method introduces, for each such unseen feature, a continuous hidden variable describ-ing its influence on the class — whether it tends to be associated with some label. We then use probabilis-tic inference over the test data to infer a distribution over the value of this hidden variable. Intuitively, we “learn ” the role of this unseen feature from the test set, generalizing from those instances whose label we are fairly sure about. Our overall probabilistic model is learned from the training data. In particular, we also learn models for characterizing the role of unseen fea-tures; these models use “meta-features ” of those fea-tures, such as words in the neighborhood of an un-seen feature, to infer its role. We present results for this framework on the task of classifying news arti-cles and web pages, showing significant improvements over models that do not use unseen features. 1.