Word-level language identification inThe Chymistry of Isaac Newton
Levi King, Sandra Kübler, Wallace Hooper · Digital Scholarship in the Humanities · 2014
In this article, we introduce the task of word-based language identification in multilingual texts, in which every word needs to be classified with regard to its language. This task is necessary for multilingual texts in which language switches can occur within sentences, often more than once, as is the case in the texts in The Chymistry of Isaac Newton collection. We present a novel method based on character n-grams in combination with a weighting scheme that allows us to model the probability of language switches at different points in sentences. This method reaches the highest accuracy of 89.94% when 5-grams are used.