A Deep Learning Approach For Bangla Image Captioning System

Toshiba Kamruzzaman, Soomanib Kamruzzaman, Abir Zaman · 2021

Naturalness and generalization are the two challenges while generating automated Image Captioning through a System. There is a lack of research to focus on these challenges for Image Captioning in the Bangla language. Furthermore, the lexical resources for Image Captioning in Bangla are not adequate. An effort has been made in this research to contribute to experimenting and analyzing the results of Image Captioning for Bangla Language using deep learning on a novel dataset “Ovro1.1” which comprises 2000 images along with a caption describing that image. Images on Bangla culture, lifestyle, festivals, etc. are recorded with captions in this dataset which makes the training more immersive for the Bangla language. Two neural networks are used in this model where a Convolutional Neural Network (CNN) is used for extracting the features of the images into a vector representation, whereas the vector representation is trained by a Recurrent Neural Network (RNN) for generating text output as its caption. This is an encoder-decoder model architecture where the CNN acts as an encoder and the RNN acts as a decoder. Inception V4 architecture is used as the CNN model for the encoder. Long Short-Term Memory (LSTM) Cells are used in decoding. The model is trained with the existing “BanglaLekha ImageCaption” dataset along with our developed dataset “Ovro1.1” and an evaluation has been done on the results. The model gives a BLEU score of 0.4224 and gives decent logical captioning for the images. However, the captions are constrained within 50 words and the model does not work well with conceptual arts, cartoons, and in recognizing special persons or famous places.

Read the paper · More papers on PaperTik