TB-Transformer: Integrating Mouse Trace with Object Bounding-box for Image Caption
Nan Lin, Shuo Guo, Lipeng Xie · 2022
The traditional image caption algorithms output the image description text based on the image features. However, the predicted results often lack crucial semantic and spatial information. To solve the above problems, this paper proposes an image caption model based on multimodal data integration, which improves the ability of the image caption model to capture critical semantic information by adding mouse trace and object bounding box as additional information sources of the model. Firstly, the mouse traces are segmented, transformed into bounding boxes, and aligned with a text word mark. Secondly, we utilize the object detection network to learn the image features and predict the bounding box used to supplement image features. Finally, the multi-head attention mechanism is used to encode, integrate, and decode the image, caption, and mouse trace, and predict the image caption results. To verify the effectiveness of this method, the proposed method was tested on ADE20k, Flickr30k, and MSCOCO data sets based on localized narrative annotation achieving CIDEr scores of 1.280, 1.983, and 1.355. Compared with the existing methods, the proposed method has obvious advantages.