Image to Text: Comprehensive Review on Deep Learning Based Unsupervised Image Captioning

Aishwarya D Shetty, Jyothi Shetty · 2023

The task of image captioning has seen considerable success using deep neural networks. This assessment offers a thorough overview of the most cutting-edge approaches for deep learning-based unsupervised image captioning. In order to bridge the gap between computer vision and natural language processing, the effort of creating meaningful textual captions for images is known as image captioning. Without using annotated image-caption pairings for training, unsupervised approaches try to provide captions. The review examines a variety of techniques, including generative adversarial networks (GANs), pre-training, transfer learning, and convolutional neural networks (CNNs) for extracting visual features, as well as recurrent neural networks (RNNs) or transformer models for language modelling. Researchers and practitioners can benefit from this survey's observations, classification of methodologies, and recommendations for the future of deep learning-based unsupervised image captioning.

Read the paper · More papers on PaperTik