Modular Attention Network based on Language Model for Referring Expression
Xutao Deng, Liang Xie, Yibo Cui, Meishan Zhang, Huijiong Yan, Ye Yan, Erwei Yin · 2021
Referring expression comprehension is to point out the corresponding object in the image by giving the natural language expression. This requires not only a full understanding of the image content, but also a certain understanding of the relationship between the objects in the text information, such as the understanding of the object appearance, location attributes and the relationship between the objects. In this paper, we propose a module network guided by scene graph and based on pre-training model (pre-training SGAN). The image is modeled as a semantic graph of graph structure representation, and the natural language expression is also analyzed as the representation of language scene diagram. Further, we use attention to extract the feature from different modules. Finally, the image and text are matched by similarity. The experiment shows that our method is better than other algorithms.