Captioning Images with Words: A Transformer-based Image Captioning Model

M Yuvanesh · International Journal for Research in Applied Science and Engineering Technology · 2025

Image captioning represents a complex interdisciplinary task that merges computer vision and natural language processing to produce coherent and contextually meaningful descriptions of visual content. This research focuses on the development of a custom transformer-based model aimed at addressing the limitations of traditional captioning approaches, particularly in terms of semantic accuracy and contextual relevance. The proposed architecture incorporates pre-trained convolutional neural network (CNN) for effective image feature capturing, followed by transformer-based mechanisms for generating natural language descriptions. To assess the effectiveness of the model, a comparative evaluation is conducted against a widely used LSTM-based captioning framework. Experiments are carried out on the Flickr8k dataset, with performance measured using BLEU scores. Results indicate that the transformer- based approach offers notable improvements in the quality and relevance of generated captions, demonstrating its potential for practical applications in areas such as media content analysis, e- commerce, and assistive technologies

Read the paper · More papers on PaperTik