Vision to Language: Captioning Images using Deep Learning

Shreyasi Charu, Sourav Mishra, Tapan Kumar Gandhi · 2020

The challenging problem of image captioning and its solution has been approached using an amalgam of Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs). The results obtained using this Encoder-Decoder architecture have proven to be quite satisfactory. This work attempts to study the comparative performance of various encoder architectures for image captioning under CNN-RNN background. We show here the various BLEU scores evaluated using three different experimental settings for the encoder part. Firstly, the encoder used is the VGGNet [1] model which was pre-trained for ILSVRC-2012 [2], followed by ResNet [3] model for ILSVRC-2015 and finally the complex yet less expensive InceptionV3 [7] model.

Read the paper · More papers on PaperTik