Automated Image Caption Generator Using Deep Learning

G. Krishnaveni · International Journal for Research in Applied Science and Engineering Technology · 2025

One of the most important tasks in computer vision and natural language processing is the automatic creation of image captions.. This paper presents an approach to automatically generate descriptive captions for images by combining Convolutional Neural Networks (CNNs) and Inception V3 architecture. The proposed system utilizes a pre-trained Inception V3 model to extract high-level features from input images. These extracted features are then passed to a Recurrent Neural Network (RNN), specifically an LSTM (Long Short-Term Memory) network, to generate coherent and contextually relevant captions. Inception V3, a deep convolutional neural network designed for large-scale image classification, serves as the feature extractor. It helps capture rich spatial hierarchies within the images, making it highly effective for understanding complex visual information. The LSTM network, on the other hand, is used to model the sequence of words in the caption, ensuring grammatical correctness and semantic accuracy. The system is trained on a large dataset of images paired with humangenerated captions, such as the MS-COCO dataset, to ensure robust learning. The proposed method is evaluated based on its performance in generating captions that are semantically and syntactically appropriate. The model’s performance is compared to other existing image captioning methods, demonstrating its effectiveness in generating descriptive and accurate captions for unseen images. This work highlights the synergy between CNNs for visual feature extraction and LSTM networks for sequence generation, offering a promising solution for tasks requiring image-to-text conversion, including image retrieval, content-based indexing, and accessibility applications.

Read the paper · More papers on PaperTik