Deep Learning-Based Urdu Image Captioning

Humayun Shakeel Khan, Rimsha Muzaffar, Syed Yasser Arafat, Zain Irshad · 2024

Image captioning is the term used to describe the system that automatically generates the captions/textual descriptions of images and is the fundamental problem in machine learning and computer vision such as human-computer interaction. The development of automated methods to obtain captions from images based on content has gained a lot of attention recently. There are still a few systems or research that implemented the Urdu captioning system for images using the Urdu language. We developed MUST-U rduF8K, the Urdu version of Flickr 8k. Dataset that contains 8092 Flickr8k images and 5 Urdu descriptions per image, totaling 40460 Urdu Descriptions, and, also identified errors in Google translation. We have implemented Convolutional Neural Network (CNN) and Long Short-Term Memory (LSTM-based) deep neural network models that generate Urdu captions and we also performed the line-by-line comparison of Google translation with our native translation and also the self-similarity comparison of the first text description with the other 4 descriptions to verify the correctness of our dataset. The Google translation had an error of around 20%. To improve our results, we developed two models, model 1 is created without optimizing the parameter, and model 2 is parametric optimized. The results of our experiments are 0.5027 and 0.528 BLEU Score corresponding to Model 1 and Model 2.

Read the paper · More papers on PaperTik