A Review of Image Captioning Techniques: Types, Deep Learning Advancements, and Limitations

Chaitanya S Bhosale, Pradip Salve, Vishal Shirsath · Cureus Journal of Computer Science. · 2025

Image captioning involves generating text that describes visual content in images. It integrates computer vision and natural language processing domains. The presented study explores recent advancements and ongoing challenges in the image captioning domain. We have explored various image captioning methodologies, including retrieval-based, template-based, and deep learning-based approaches, and highlighted their respective strengths and limitations. We have also studied the critical role of object and action recognition and mapping relationships in images and how they help create clear and meaningful captions. Datasets that provide the foundation for training and evaluating models are also explored, such as Microsoft Common Objects in Context, Flickr8k, Flickr30k, and Visual Genome. Deep learning models such as recurrent neural networks and transformer-based architectures enhance the performance of captioning tasks. Despite these advancements, there are significant challenges that need to be addressed, such as handling complex scenes in images and generating context-aware captions. This study provides a detailed overview of recent research done in image captioning and offers potential directions for future work, focusing on the accuracy and scalability of the system for real-world applications such as assistive technology for the visually impaired, content monitoring on social media, and e-learning.

Read the paper · More papers on PaperTik