A Survey on Image Captioning Using Encoder-Decoder Based Deep Learning Models

Chauhan Harshil Narendrabhai, Chintan Bhupeshbhai Thacker · 2024

The intersection of computer vision and natural language processing has seen remarkable progress, particularly in deep learning techniques. One of the captivating and burgeoning areas of research stemming from this advancement is image caption generation, which aims to automatically extract content from images and describe it in natural language. Image captioning techniques have evolved significantly from classical deep learning approaches to the most advanced methods available today. Presently, Encoder-Decoder based image caption generation techniques have achieved a state-of-the-art level of performance, nearing human-level abilities in describing visual content. This paper provides a comprehensive overview of image caption generation using deep learning, covering various methodologies for content extraction from images, challenges encountered, available image captioning datasets, and the diverse evaluation metrics employed to assess state-of-the-art performance.

Read the paper · More papers on PaperTik