Bengali Image Caption Generation using Attention Mechanism
Sayantani De, Ranjita Das, Ashish Singh Patel · 2024
Image captioning has emerged as a rapidly thriving area for the machine learning research community. Generally, image captioning is performed by combining various computer vision features, natural language processing, and machine learning methods with consideration of some additional inputs to get more accurate context-dependent image captions Bengali is a significant language in India, adopted by approximately 100 million people. There exist various state-of-the-art methods for generating captions in the English language; however, for the Bengali language, there are very limited methods, and existing methods in the English language are not particularly helpful. Moreover, translations from English to Bengali may overlook or misinterpret subtle meanings, tones, or cultural nuances. Therefore, this work proposes a machine-learning model for captioning pictures in Bengali using an attention mechanism. The Flickr-8k dataset, which has 8000 images, is used to train the model in this work. The proposed method generates image captions in the Bengali language and attained a BLEU score of 0.66.