langid.py for better language modelling

Paul F. Cook, Marco Lui · 2012

Large corpora are crucial resources for building many statistical language technology systems, and the Web is a readilyavailable source of vast amounts of linguistic data from which to construct such corpora. Nevertheless, little research has considered how to best build corpora from the Web. In this study we consider the importance of language identification in Web corpus construction. Beginning with a Web crawl consisting of documents identified as English using a standard language identification tool, we build corpora of varying sizes both with, and without, further filtering of non-English documents with a stateof-the-art language identifier. We show that the perplexity of a standard English corpus is lower under a language model trained from a Web corpus built with this extra language identification step, demonstrating the importance of state-of-the-art language identification in Web corpus construction. 1 The need for large corpora Corpora are essential resources for building language technology (LT) systems for a variety of applications. For example, frequency estimates for n-grams — which can be used to build a language model, a key component of many contemporary LT systems — are typically derived from corpora. Furthermore, bigger corpora are typically better. Banko and Brill (2001) show that for a classification task central to many LT problems, performance increases as a variety of models are trained on increasingly large corpora. The Web is a source of vast amounts of linguistic data, and the need for large corpora has motivated a wide range of research into techniques for building corpora of various types from the

Read the paper · More papers on PaperTik