A Denoising Framework for Image Caption
Yulong Zhang, Yuxin Ding, Rui Wu, Fuxing Xue · 2019
Image caption is one of the hottest research topics at the moment in the image processing field. However, most image caption models based on the Encoder-Decoder framework cannot accurately find the alignment relationship between objects in the image and objects in the text, resulting in an inaccurate description. In this work, we propose a denoising framework in which the image and the text are processed separately, and we use the caption stem to find the alignment relationship between image and text accurately. Compared with previous work, this framework can more accurately find the alignment of images and texts and generate more accurate captions. Experiments on the MSCOCO dataset show that the caption generated by this model can express the content of an image more accurately. Besides, the model can generate short captions and long captions with rich semantic according to users' needs.