Semi-Supervised Priors for Microblog Language Identification
S. Carter, E. Tsagkias, Wouter Weerkamp, Boscarino, C., Hofmann, K., Jijkoun, V., Meij, E., de Rijke, M., Wouter Weerkamp · UvA-DARE (University of Amsterdam) · 2011
Offering access to information in microblog posts requires suc-cessful language identification. Language identification on sparse and noisy data can be challenging. In this paper we explore the performance of a state-of-the-art n-gram-based language identifier, and we introduce two semi-supervised priors to enhance perfor-mance at microblog post level: (i) blogger-based prior, using pre-vious posts by the same blogger, and (ii) link-based prior, using the pages linked to from the post. We test our models on five lan-guages (Dutch, English, French, German, and Spanish), and a set of 1,000 tweets per language. Results show that our priors improve accuracy, but that there is still room for improvement.