The Use of Machine Translation to Provide Resources for Under-Resourced Languages - Image Captioning Task

Basem H.A. Ahmed, Motaz Saad · 2021

Image captioning is an NLP task that has many applications such as image search and retrieval. This Task is a challenging task, and it needs a lot of data (image data and their text captions), which might not be available for some languages. In this work, we investigate the use of a machine translation system to provide resources for a low-resourced language (Arabic) for the imaging captioning task. We train a model on captions automatically translated using Google machine translation service. The performance is measured using the BLEU, ROUGE, CIDEr, METEOR metrics. We compare to English model's performance. We also evaluate the generated captions on manually translated captions. The results show that machine translation can be good enough for creating resources for low-resourced languages for the image captioning task and translating training data and building a new model is better than translating the model's output.

Read the paper · More papers on PaperTik