Text-based language identification for some of the under-resourced languages of South Africa
Tshephisho Joseph Sefara, Madimetja Jonas Manamela, Promise Tshepiso Malatji · 2016
Language identification is the problem of correctly classifying a sample of text/documents based on its language. However, much of the research work focused on the English language corpora and little research work focused on other South African official languages. In a multilingual society like South Africa, the use of automatic language identification in any language-specific system would be a vital step in bridging the digital divide between diverse members of the society. Various machine learning algorithms can be used to solve the problem of identifying the natural language of a document/text. This paper presents a text-based language identification using individual proper names, specifically surnames in a South African context. Three supervised machine learning methods are implemented to perform 3-way multiclass classification using support vector machines, and naïve Bayes language models. These algorithms are applied to the language identification task and evaluated in extensive experiments for three official languages of South Africa: Tshivenda, Xitsonga and S epedi. All three machine learning methods achieved remarkable results in a 10-fold cross validation. The results indicate that a multinomial naïve Bayes method achieved better performance than other algorithms.