Image Captioning based on Encoder Decoder Architecture

Saranya M D, Veera Anusuya V · 2024

Images are essential for communicating ideas, feelings, and narratives in the era of digital media and content consumption. Computers to produce textual data for an image that replaces humans. Image captioning is a fascinating fusion of computer vision (CV) and Natural Language Processing (NLP) that aims to generate text from an image. By using these technologies, images can be used as an approachable, meaningful, and descriptive form of communication. Because of its wide range of uses, image captioning has attracted a lot of attention. Strong feature representations and context-aware language generation algorithms are necessary to overcome the major challenge of bridging the semantic gap between visuals and text. In this research work, EfficientNetV2B0 is utilized in the encoder part, for extracting objects from an image. Then Long Short-Term Memory (LSTM), a type of recurrent neural network as a decoder for generating descriptions by using encoder output. The suggested approach yields superior results than the earlier research when compared to the current method on the measures of Bilingual Evaluation Understanding and Consensus Based Image Description Evaluation scores.

Read the paper · More papers on PaperTik