Image Captioning Across Object Categories: A CNN ‐ LSTM and BLEU Score‐Based Study
Garima Salgotra, Pawanesh Abrol · Computational Intelligence · 2026
ABSTRACT Image captioning aims to automatically generate semantic descriptions for visual data and is widely applied in assistive systems, content filtering, and recommendation engines. This paper presents a hybrid encoder–decoder architecture integrating Convolutional Neural Networks (CNNs) for visual feature extraction and Long Short‐Term Memory (LSTM) networks for sequence generation. An attention mechanism is incorporated to dynamically weight salient image regions during decoding, improving contextual representation and object interaction modeling. The proposed model is trained and evaluated on the Flickr8k dataset, categorized by object complexity. Performance is assessed using BLEU, METEOR, and ROUGE metrics. Experimental results demonstrate superior performance on low‐complexity images, achieving BLEU scores of 0.61, 0.60, and 0.38 for single‐, double‐, and multi‐object images, respectively. Comparative analysis indicates that the attention‐enhanced model outperforms standard approaches, particularly in complex scenes.