A Hidden Markov Model-Based Parts-of-Speech Tagger for Yoruba Language

Abiola O. B, Bankole O. H, Adeyemo Oluwaseyi A, Lawrence Bunmi Adewole, Ogundipe A. T, Saka-Balogun O. Y, Obamiyi S. E, Toyin Okebule, Christianah O. Akinduyite · 2024

Parts-of-speech tagging is a linguistics task that assigns the best sequence of tags to a given sequence of input words. The process falls under word sense disambiguation, which determines the function of a given word in a sentence. Many natural language processing tasks have parts of speech tagging in their process pipeline, hence its importance. While resource-rich languages have attained near-human accuracies in parts of speech tagging, most under-resourced languages, where most African languages belong, lack support for automatic tagging. The absence of these tools for resource-scarce languages has prevented these languages from leveraging state-of-the-art methods and algorithms for higher tasks. In this paper, a Hidden Markov model-based parts of speech tagger for the Yoruba language is presented. Issues associated with the Parts of speech tags identified in Yoruba literature and the necessary modifications to enhance the accuracy of the parts-of-speech tagging process were discussed. The process involved the identification of major tags supported by the language in comparison with universal parts of speech Tagset. We trained a tagger on a manually tagged data set of 1000 sentences extracted from different domains while the evaluation process was carried out on 300 sentences. An accuracy of 99.85% was obtained during the testing phase. The model can be used for parts of speech tagging or as a pipeline for other natural language processing tasks.

Read the paper · More papers on PaperTik