Language identification: a solved problem suitable for undergraduate instruction
Paul McNamee · Journal of computing sciences in colleges · 2005
Automatic determination of the language of an electronic text is an important problem, which arises when processing natural language. This paper describes the main methods used in attacking this problem and demonstrates how even the most simple of these methods using data obtained from the World Wide Web achieve accuracy approaching 100% on a test suite comprised of ten European languages. The language identification problem and its solution illustrate many fundamental issues in computer science and the processing of human language; accordingly it is well suited for classroom instruction.