AI-Driven Image Captioning for Assistive Applications
Y. Ramu Naidu, V.N.V.N. Siva Ram, V. Sateesh Kumar, K. Shri Ramtej · 2026
With millions of visually impaired individuals worldwide, automated image captioning is a crucial assistive tool, bridging the gap between visual content and textual understanding. This paper presents a deep learning-based framework for image caption generation, structured into three key phases: image collection, feature extraction for model training, and performance evaluation. The methodology utilizes the Flickr8K and MSCOCO datasets, with the latter offering a larger and more diverse set of images and captions, providing a robust testbed for model generalization. The proposed framework integrates convolutional neural networks (CNNs) and long short-term memory (LSTM) networks to generate descriptive captions. Specifically, CNN architectures such as VGG16, ResNet, and Inceptionv3 are used to extract image features, while LSTMs process sequential textual data to produce captions. The combined CNN-LSTM model demonstrates superior performance compared to standalone models, achieving higher image captioning accuracy.