From Pixels to Words: A Comprehensive Study of CNN-RNN Models for Image Captioning
Hayder Jaber Samawi · Journal of Information Systems Engineering & Management · 2025
Image captioning is a challenging task that generates textual descriptions of images using techniques from computer vision and natural language processing. In this paper, we present a large-scale study of CNN-RNN models that have attracted attention as a new approach to the image captioning problem. High-level visual features are extracted from images using Convolutional Neural Networks (CNNs) and the sequential description is generated based on these features using Recurrent Neural Network (RNN), or other RNN variants most popular of which are Long Short-Term Memory (LSTM) and Gated Recurrent Units (GRU). Attention mechanism is one of the main innovations in this method, it enables the model to attend to specific image regions when generating each word in its caption making sure that they are more relevant. The paper describes different types of CNN-RNN architecture, methods and about how the attention enhance their performance. In addition, the paper covers recent datasets, evaluation measures and benchmarks that are used in this area. We examine experimental results from previous works to gain insight into the powers and limitations of CNN-RNN. The paper ends with some challenges and future research directions related to tackling biases in datasets, generalization, and combination of multimodal information. We believe that this systematic review can be a useful reference for further improving image captioning systems.