A Keyword-Guided GATv2-LSTM Graph Question Answering Network

Lei Wang, Tao Zuo · 2023

Most VQA models are more inclined to object recognition and learning the association between images and text, ignoring image scene understanding and reasoning. Scene graphs, on the other hand, as a structured graphical representation, have shown superiority in image understanding and reasoning. Therefore, in this work, we propose a keyword-guided GATv2-LSTM graph question answering network (called KGL). In particular, scene graphs are used to describe objects and relationships in images, thus narrowing the gap between different modalities. Based on the scene graph, we use a keyword-guided GATv2 network to learn the most important information in the scene graph, and then multiple keyword-guided learned graph features are considered as a sequence, which is fed into an LSTM network to learn the dependency information among the features. Finally, our experiments on the GQA dataset show the superiority of the KGL.

Read the paper · More papers on PaperTik