Vision Transformer-Based Framework for Generating Concise Image Captions: A Vision Language Approach

Raksha Puthran, Thejash, Anusha Prashanth Shetty · 2025

Image Captioning is a great task that comes between the interaction between the computer vision and natural language processing. Deep Learning has a huge process in generating short descriptive captions for any given Images. This paper explores a deep-learning based caption generation model using encoder-decoder structure. Specifically,the model is based on Transformer-based Vision-Language Pre-training (VLP) framework, incorporating the Vision Transformer (ViT) for image feature extraction and a Transformer decoder for generating captions.The Vision Transformer (ViT) replaces traditional convolutional neural networks (CNNs) in the encoder with a patch-based attention mechanism to capture global visual information more effectively. The pre-trained Vision-Language model is fine-tuned on captioning tasks to align image features with corresponding linguistic representations. The encoder captures high-level semantic features from the image, while the decoder generates concise and coherent text based on these features. Our experiments show that this method generates accurate and relevant captions, outperforming conventional CNN-RNN approaches in various image captioning benchmarks

Read the paper · More papers on PaperTik