Data‐driven image captioning via salient region discovery

Mert Kilickaya, Burak Kerim Akkuş, Ruket Çakıcı, Aykut Erdem, Erkut Erdem, Nazlı İkizler-Cinbiş · IET Computer Vision · 2017

In the past few years, automatically generating descriptions for images has attracted a lot of attention in computer vision and natural language processing research. Among the existing approaches, data‐driven methods have been proven to be highly effective. These methods compare the given image against a large set of training images to determine a set of relevant images, then generate a description using the associated captions. In this study, the authors propose to integrate an object‐based semantic image representation into a deep features‐based retrieval framework to select the relevant images. Moreover, they present a novel phrase selection paradigm and a sentence generation model which depends on a joint analysis of salient regions in the input and retrieved images within a clustering framework. The authors demonstrate the effectiveness of their proposed approach on Flickr8K and Flickr30K benchmark datasets and show that their model gives highly competitive results compared with the state‐of‐the‐art models.

Read the paper · More papers on PaperTik