Machine Translation for Morphologically Rich Low-Resourced South African Languages

Hlaudi Daniel Masethe, Lawrence M Mothapo, Sunday Olusegun Ojo, Pius Adewale Owolawi, Mosima Anna Masethe, Fausto Giunchigilia · 2024

In the context of machine translation, many of South Africa's languages are categorized as low-resourced, despite the country's vast linguistic diversity. Creating efficient multilingual translation systems that can support languages like Afrikaans, Sesotho sa Leboa (Sepedi), isiNdebele, Setswana, isiXhosa, Xitsonga, siSwati and isiZulu are severely hampered by the lack of high-quality training data. Low translation accuracy and fluency are the result of existing translation models' frequent inability to perform well in these languages. As a result, speakers of these languages encounter difficulties using increasingly accessible online resources for information and services. By utilizing the MarianMT model to create a reliable multilingual machine translation system, this work seeks to address these issues by enhancing the quality of translation for South African languages with limited resources. This study presents a comprehensive approach to multilingual machine translation focused on low-resourced South African languages, utilizing the MarianMT model. Our objective was to enhance translation quality for languages with limited training data, including Afrikaans, Sesotho sa Leboa (Sepedi), isiNdebele, Setswana, IsiXhosa, Xitsonga, siSwati and isiZulu. We conducted extensive experiments, resulting in a BLEU score of 62%, indicating a significant improvement in translation. Our evaluation metrics demonstrated strong performance, with precision at 78%, recall at 81%, F1 score at 79%, and overall accuracy at 82%.

Read the paper · More papers on PaperTik