Accurate Language Identification of Twitter Messages
Marco Lui, Timothy J. Baldwin · 2014
We present an evaluation of "off-theshelf" language identification systems as applied to microblog messages from Twitter.A key challenge is the lack of an adequate corpus of messages annotated for language that reflects the linguistic diversity present on Twitter.We overcome this through a "mostly-automated" approach to gathering language-labeled Twitter messages for evaluating language identification.We present the method to construct this dataset, as well as empirical results over existing datasets and off-theshelf language identifiers.We also test techniques that have been proposed in the literature to boost language identification performance over Twitter messages.We find that simple voting over three specific systems consistently outperforms any specific system, and achieves state-of-the-art accuracy on the task.