Image Captioning using Google's Inception-resnet-v2 and Recurrent Neural Network

Yajurv Bhatia, Aman Bajpayee, Deepanshu Raghuvanshi, Himanshu Mittal · 2019

Given a photograph as input, this paper solves the problem of experiencing a plausible caption of the photograph. The model learns about the correlations between language and images from the provided data-set of labeled images. It proposes a fully automatic approach through a combination of Convolutional Neural Network and a Recurrent Neural Network. The encoder is responsible for understanding the features present in the inputted image that are useful in eventually producing an explanation. The model attempts at producing captions for both the objects and the regions present in the image. Treating language as a big label space, the project generates predictions for the various regions of the image and then stitches them together.

Read the paper · More papers on PaperTik