Evaluation of language identification methods using 285 languages
Tommi Jauhiainen, Krister Lindén, Heidi Jauhiainen · DSpace repository (University of Tartu) · 2017
Language identification is the task of giving a language label to a text.It is an important preprocessing step in many automatic systems operating with written text.In this paper, we present the evaluation of seven language identification methods that was done in tests between 285 languages with an out-of-domain test set.The evaluated methods are, furthermore, described using unified notation.We show that a method performing well with a small number of languages does not necessarily scale to a large number of languages.The HeLI method performs best on test lengths of over 25 characters, obtaining an F 1 -score of 99.5 already at 60 characters.