Generating Bengali Captions for Images Using Transformer Architecture
Tausif Uddin Ahmed Chowdhury, Abdus Salam · 2024
Image captioning is a challenging task that requires generating meaningful and context-based textual descriptions for given images. It represents a convergence of computer vision and natural language processing, two core areas within the field of artificial intelligence. Despite major advances in this discipline, most studies have concentrated on English languages, overlooking many other native languages, such as the Bengali language spoken by millions of people worldwide. As a result, existing systems often perform poorly on images from different geographical and cultural contexts. Additionally, there has been little study on generating captions using Bengali. In this study, we address this research gap by utilizing a transformer-based approach to produce Bengali captions. Our model leverages a pre-trained convolutional neural network (CNN), specifically Inception V3, for robust image feature extraction. These features are then fed into a transformer architecture, which generates coherent and contextually accurate captions. We conducted extensive experiments using the BanglaLekhaImageCaptions dataset to compare the performance of our model with others. Our findings demonstrate that the proposed model surpasses other models in terms of BLEU scores. This research not only contributes to the field of image captioning but also emphasizes the importance of developing AI technologies that are inclusive of diverse languages and cultural contexts.