Context-Aware Image Caption Generation
Ramesh Babu Pittala, Shaik Firoz, Medikonda Asha Kiran, Manyam Thaile, Komma Teja Yadav, Krishna Reddy · 2025
Image captioning has been one of the greatest hustle for research problems in computer vision and natural language processing because of the accurate capturing and presentation of a visual image and caption. This paper seeks to meet the requirement by using VGG16 and EfficientNetB7 Convolutional Neural Networks (CNNN) that extract visual details from pictures. Both models capture critical low-level and high-level features important in the generation of captions that are accurate and meaningful. Data augmentation techniques are used for training. Long Short Term Memory networks (LSTMs) are applied to learn temporal relationships necessary in caption generation. The model is trained on the Flickr8k dataset. BLEU score is used to evaluate relevance and quality. SIGHT-SCRIPT enhances user experience through descriptive, context-relevant captions.