hA Hybrid Model Combining Vision Transformer and Generative Pre-trained Transformer for Image Captioning
Hina Gupta, J Vishwesh, M. Jagadeesh, M. Sucharitha, Vitta Sai Pradyothan · 2025
Creating descriptive captions for images is now becoming a mission-critical application area in the intersection of natural language processing and computer vision. This work provides the hybrid model VisionGPT2, combining Vision Transformer for encoding images and GPT2 for text generation, which enables accurate, coherent, and context-aware image descriptions. The model is tested over a dataset by using BLEU score metrics and has achieved a high BLEU score of 0.4877, which ascertains the strength of cross-modal architectures for image captioning. In this paper, the design and training process of VisionGPT2 has been discussed along with the benefits and challenges faced while proceeding and the potential applications the model has gained in performance for future use.