Automatic Spelling Corrector for Yoruba Language Using Edit Distance and N-Gram Language Models

Abiola O. B, Erinfolami Oluwaseyi, Lawrence Bunmi Adewole, O. M., Oguntimilehin Abiodun, G.O. Babalola, Bukola Badeji–Ajisafe, Adigun K. A, Sanya O. A · 2024

The lack of tools and resources to support higher-level Natural Language Processing (NLP) tasks for African languages has been a significant obstacle to developing NLP research in Africa. This research proposes an automatic spell correction for the Yoruba language. The procedure comprises four phases: corpus gathering, corpus cleaning normalisation, n-gram model development, and spell correction. Existing publicly accessible Yoruba corpora were combined with corpus sourced using an online corpus elicitation tool for language model development. Various preprocessing steps, such as the removal of foreign words, removal of numbers, and separation of punctuation marks from associated words, among others, were carried out on the corpus. Custom-written Python scripts were used to extract the unigram, bigram, and trigram language models from the corpus gathered. The last phase of the system is spelling error detection and correction, in which a given sentence S is segmented into word tokens. Each word token is checked against a pre-defined lexicon. For each Out-of-Vocabulary Word (OOV), similar words are generated from the lexicon by extracting words of almost equal length using the Levenshtein edit distance. The top k-words, with the minimum Levenshtein edit distance, alongside the word in context, are ranked using the n-gram language model. The original word is returned if no similar word is found within the model’s parameter configuration. The correction accuracy of 71.50% was recorded on a test set of 2000 sentences, with each sentence having at least one misspelt word.

Read the paper · More papers on PaperTik