Sight-Script: Automated Captioning for Visual Data

Sukanya Varshini, Sakthitharan Subramanian · 2025

Image captioning has been one of the greatest hustles for research problems in computer vision and natural language processing because of the accurate capturing and presentation of a visual image and caption. This paper seeks to meet the requirement by using VGG16 and EfficientNetB7 Convolutional Neural Networks (CNNs) that extract visual details from pictures. Both models capture critical low-level and high-level features that are important in the generation of captions that are accurate and meaningful. To make the model robust, data augmentation techniques are used which help in creating variety of inputs for training. Long Short-Term Memory networks are applied to accomplish this task and are good for learning temporal relationships between words which are necessary in generating captions. It has been trained and tested on the flickr8k dataset, which comprises of 8000 images with multiple captions for each image to allow for diverse visual-linguistic pairings. BLEU score, computed with respect to human variation, addresses quality and relevance as it estimates the relations of machine-generated captions to those created by people. SIGHT-SCRIPT takes into consideration the accuracy of captions and context of the captions with the intention of enhancing how the user images and improving their experience by interacting with images that are descriptive and context relevant hence explaining the visual aspect to the reader better.

Read the paper · More papers on PaperTik