Hybrid Deep Learning Approach for Image Captioning: Integrating ResNet50 and GloVe with LSTM
Dhakshana Moorthy, Rajasekar Deepa · 2025
The automatic compilation of textual descriptions for visual content is often referred as image captioning, and it is an important issue in computer vision and natural language processing. In this paper, we provide a hybrid deep learning strategy that combines the advantages of Long Short-Term Memory (LSTM) networks for sequential caption generation, GloVe embeddings for semantic word representation, and the ResNet50 picture feature extraction model. In order to extract 2048- dimensional feature vectors, which contain the visual content of the photos, we first preprocess the images using ResNet50. The GloVe embeddings are then used to encode these aspects into a semantic space, allowing for better comprehension of the visual features tied with textual descriptions. An LSTM network is used to create the sequential captions, that successfully capture the text's temporal dependencies. Our solution improves current techniques by developing captions for a range of images that seem more rational and contextually appropriate. It maximizes the efficiency of the captioning process by improving the quality of the output captions through complete preprocessing, feature extraction, and effective use of semantic embeddings. We provide an evaluation of our model, exhibiting optimistic results on standard image captioning datasets.