Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation
Ling Cheng, Wei Wei, Xian-Ling Mao, Yong Liu, Chunyan Miao · IEEE Access · 2020
Recently, automatic image caption generation has been an important focus of the work onmultimodaltranslation task. Existing approaches can be roughly categorized into two classes,top-downandbottom-up, the former transfers the image information (called asvisual-level feature) directly into a caption, and the later uses the extracted words (called assemantic-level attribute) to generate a description. However, previous methods either are typically based one-stage decoder or partially utilize part ofvisual-level orsemantic-level information for image caption generation. In this paper, we address the problem and propose an innovative multi-stage architecture (called asStack-VS) for rich fine-grained image caption generation, via combiningbottom-upandtop-downattention models to effectively handle bothvisual-level andsemantic-level information of an input image. Specifically, we also propose a novel well-designed stack decoder model, which is constituted by a sequence of decoder cells, each of which contains two LSTM-layers work interactively to re-optimize attention weights on both visual-level feature vectors and semantic-level attribute embeddings for generating a fine-grained image caption. Extensive experiments on the popular benchmark datasetMSCOCOshow the significant improvements on different evaluation metrics,i.e.,the improvements on BLEU-4 / CIDEr / SPICE scores are 0.372, 1.226 and 0.216, respectively, as compared to the state-of-the-art.