Image to Text Conversion Using Deep Learning Algorithms: Survey
Saranya M D, V. Veera Anusuya, R. Vazhan Arul Santhiya, R. Srimathi · 2024
Image Captioning (ICs) seamlessly combines the realms of Computer Vision (CV) and Natural Language Processing (NLP) task involved in producing textual sentences that summarise the image content in a way which is understandable to humans. The primary objective is to enable machines to convey the content of images in human-readable language. Image understanding and text formation are typically the two primary processes in the image captioning process. A family of Convolutional Neural Network (CNN) is used to retrieve visual information from an input image during the image understanding process. This visual information includes the image's scenes, relationships, and objects. To translate these visual features into coherent and contextually appropriate captions, recurrent neural network, long short-term memory, transformer-based models, etc. are used in the text generation step. In this article, we propose a survey about the previous research paper and the fundamental steps involved in image captioning tasks.