En-De-Cap: An Encoder Decoder model for Image Captioning
Nikhil Patwari, Dinesh Naik · 2021
Image captioning is the technique for generating descriptive text of a given image. Modern image captioning techniques are multi-model techniques that uses both Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN) to achieve the task of describing the image . The area of image captioning is considered one of the crucial step in advancement in AI. After achieving high success rate in object detection, the next difficult task is to describe the relation between objects in the image. This paper proposes an image captioning framework which incorporates best in class CNN network with attention based Gated Recurrent Unit (GRU) network to generate better descriptive text for a given image.