Image Captioning using CNN and LSTMs

Rajeel Ahmad Ansari, Divyapratap Singh Chauhan, B. Jothi, S Krishnaveni · 2025

Image captioning is a pivotal field in AI, enabling machines to generate descriptive text for images, with applications in accessibility, content creation, and human-computer interaction. This study identifies the shortcomings of the conventional CNN-RNN model by proposing an advanced architecture that incorporates InceptionV3, LSTM, attention mechanism, and beam search. The image features are extracted using Incep-tionV3; captions are generated using LSTMs, where an attention mechanism is used to guide the model to identify important regions of the image for each word. Beam search has been employed to choose the most optimal captions. The suggested model is tested on the Flickr8k data set with the help of the BLEU metric. The outcomes show that the caption relevance, diversity, and interpretability of the model are enhanced, with BLEU-1 and BLEU-2 scores of 0.5365 and 0.3218 respectively, demonstrating its potential in real-world applications such as providing accessibility and machine learning-based content generation.

Read the paper · More papers on PaperTik