Evaluation of a language identification system for mono- and multilingual text documents

Olga Artemenko, Thomas Mandl, Margaryta Shramko, Christa Womser‐Hacker · 2006

Language identification is a classification task between a pre-defined model and a text in an unknown language. This paper presents the implementation of a tool for language identification for mono- and multi-lingual documents. The tool includes four algorithms for language identification. An evaluation for eight languages including Ukrainian and Russian and various text lengths is presented. It could be shown that n-gram-based approaches outperform word-based algorithms for short texts. For longer texts, the performance is comparable. The tool can also identify language changes within one multi-lingual document. Keywords Language identification, n-gram indexing, language model, evaluation 1

Read the paper · More papers on PaperTik