Nepali Image Captioning
Aashish N. Adhikari, Sushil Ghimire · 2019
Neural networks have seen a surge in applications ranging image recognition, time series prediction, and image captioning among others. The task of image captioning combines two fields of machine learning - computer vision and natural language processing. Several works have been done for image captioning in dominant languages like English and Mandarin with impressive results. The challenge, however, increases for Nepali, a language with a complex grammatical structure. With inherently complex grammar, inputting an image and generating the description of the image in Nepali with correct grammar is hard to achieve. Also, a standard data set does not exist for such tasks in Nepali. This work on image captioning in Nepali is the first of its kind that establishes a baseline for tasks of this sort and aims to encourage other researchers to pursue this line of research. For this, it makes public the data set that was generated during this process so as to provide a standard data set for future works. This work builds on top of the model proposed by [24] and generates image descriptions in Nepali. Describing the contents of an image in Nepali has several applications. Additionally, this work serves as a gateway to more challenging problems such as a video descriptor system and image search for search engines in Nepali. This work utilizes two encoder-decoder architectures, one with visual attention and another without visual attention. We empirically show the loss and perplexity of model performances using different optimizers. The captions generated are agreeable and coherent with the images in general and leave room for improvements in the future.