Automated Image Captioning Using ConvNets and Recurrent Neural Network

Prof. P. Singam · International Journal for Research in Applied Science and Engineering Technology · 2018

We present a model that generates free-form natural language descriptions of image regions. Our model leverages datasets of images and their sentence descriptions to learn about the inter-modal correspondences between text and visual data. Our approach is based on a novel combination of Convolutional Neural Networks over image regions, Recurrent Neural Networks over sentences, and a structured objective that aligns the two modalities through a multimodal embedding. We then describe a Recurrent Neural Network architecture that uses the inferred alignments to learn to generate novel descriptions of image regions. We introduce a system to automatically generate natural language descriptions from images that takes an input image and generates its description in text. It also generates descriptions that are notably more true to the specific image content than previous work.

Read the paper · More papers on PaperTik