Image Caption with Synchronous Cross-Attention

Yue Wang, Jinlai Liu, Xiaojie Wang · 2017

The image caption aims to translate images into descriptive sentences, involving both visual and textual resources. The Deep Neural Network (DNN) based models are widely applied to solve this task, due to their impressive performance in the computer vision and natural language processing. Specifically, the attention mechanism is proposed to allow the models to focus on the essential parts of images. However, the previous models ignore both the correlation between the attention at different time, and the supervision of words on attention selection. This paper proposes an Image Caption model with Synchronous Cross-Attention (IC-SCA), which captures a visual sequence of attention with the information of words. Our IC-SCA model has two stages, visual and textual, which jointly model the multimodal information to generate the descriptions. This model is evaluated on one of the largest datasets for image caption, namely the MS-COCO dataset. Experimental results on BLEU-1~4, METEOR and CIDEr metrics demonstrate that our IC-SCA model outperforms the benchmarks. By attention visualization, the effectiveness of our proposed mechanism is also verified.

Read the paper · More papers on PaperTik