Colloquial Image Captioning

Xuri Ge, Fuhai Chen, Chen Shen, Rongrong Ji · 2019

Image captioning has recently attracted ever-increasing research attention in multimedia and computer vision. However, most of image captioning models focus on generating the plain description for images, neglecting its colloquialism under a potential topic, e.g., the topic Movie for a poster. In this paper, we consider colloquial image captioning a new task, which, however, is trapped in the visual-textual semantic gap, as well as the semantic and the sentiment diversities. To this end, we propose a novel topic-guided multi-discriminator captioning model, termed TGMD-Cap. In particular, a topic-guided attention module is firstly designed to compose the effective visual semantic features. Secondly, a multi-discriminator generative adversarial module is designed to generate the semantically-diverse and affectively-diverse expression. To facilitate quantitative evaluation, we further construct and release a colloquial image captioning dataset, CIC-dataset, crawled from real-world social network, serving as the first of its kind. Quantitative results show that the proposed TGMD-Cap model outperforms the state-of-the-art approaches under various standard evaluation metrics.

Read the paper · More papers on PaperTik