A Statistical Approach to Error Correction for isiZulu Spellcheckers
Frida Mjaria, Catharina Maria Keet · IST-Africa Week Conference · 2018
Spellcheckers have become important due to the increase of text-based communication at work and in society on social media. There is, however, very little support for spellchecking in agglutinating Sub-Saharan African (Bantu) languages. While error detection has shown to yield acceptable results for at least isiZulu, error correction has not even been investigated. The aim of this paper is to solve the spelling correction problem by means of a statistical approach such that it can provide candidate corrections to misspelled isiZulu words (non-word errors). Trigrams learned from a corpus, their probabilities, minimum edit distance, and additional optimisations are used in the error corrector. The corrector was evaluated for the four types of non-word errors (substitution, insertions, deletions, and transpositions). It achieved an 89% language recall rate, 84% error recall, 85% language precision, and 88% error precision for error correction. The error corrector was found to have an overall suggestions accuracy rate of 95% and relevance of 61%, performing best for transposition errors. The error corrector has been added to an existing open source isiZulu error detector. This facilitates uptake and, moreover, fills a feature gap that has numerous benefits for society, both for isiZulu speakers and learners, and for bootstrapping spellcheckers for related languages.