Chinese description generation of dual attention images based on multi-modal fusion

Xiyun Lv · Journal of Physics Conference Series · 2021

Abstract In view of the lack of attention to image features in most of the current research on image Chinese description generation, the low quality of image details, and the low accuracy of generated language description, a multi-modal fusion dual attention image description generation model is proposed. The model uses the pre-trained VGG19 network to extract image features, LSTM extracts key information features of text, and combines image features and text features through attention stitching to obtain the attention weight of each spatial position of the image, and then input it into the LSTM language model Generate descriptive sentences. In the AI Challenger 2017 test set, the BLEU-1 and CIDEr scores of this algorithm can reach 71.2 and 106.1 respectively, which are better than baseline models based on a single attention structure, such as NIC, and perform manual observation and comparison of the effects of several models. The model can generate natural language descriptions that can better represent the details of the picture.

Read the paper · More papers on PaperTik