Text-based language identification for south african languages

Gerrit Botha, V. Zimu, Etienne Barnard · SAIEE Africa Research Journal · 2007

We investigate the performance of text-based language identification systems on the 11 official languages of South Africa, when n-gram statistics are used as features for classification. In particular, we compare support vector machines, likelihood and frequency difference-based classifiers on different amounts of input, text and for various values of n. With as few as 15 words of input text, reliable language identification is possible. Although the support vector macine is generally more accurate as classifier, the additional computational complexity of training this classifier may not be justified in light of the importance of using a large value for n.

Read the paper · More papers on PaperTik