Automatic Vowel Elision Resolution in Yoruba Language
Lawrence Bunmi Adewole, Adebayọ Olusọla Adetunmbi, Boniface Kayode Alese, Samuel Adebayo Oluwadare, Oluwatoyin Bunmi Abiola, Olaiya Folorunsho · 2020
Despite advancements in machine translation systems and the development of language-independent frameworks for machine translation, support for African languages is relatively low, while the majority of the supported few are yet to achieve acceptable translation accuracies. Lack of language resources and pre-processing tools have been identified as the major factors limiting the inclusion of most African languages in current translation engines. Yorùbá, a Niger-Congo language largely spoken in the South-Western part of Nigeria with an estimated speaker of about 50 million people, is one of the languages currently supported by machine translation engines such as Google Translate. Unfortunately, translation accuracy involving Yorùbá language is relatively low in comparison with most European languages. One of the reasons for this is the lack of pre-processing tools such as the Elision resolution tool. Elision in Yorùbá is the omission of tone or syllable from a text, often as a result of speech smoothening between adjacent words. This paper presents an Elision resolution framework for the Yorùbá language. The proposed framework uses a hybrid approach to elision resolution. It applies a rule-based word-partitioning algorithm at the syllabication phase, while it uses the n-gram language model for resolution-candidate ranking. Resolution result evaluation was carried out using resolution accuracy. We also established the positive effect of the elision resolution process by translating sample sentences on Google Translate before and after resolution.