Semantic-Visual Graph Reasoning for Visual Dialog
Dongze Hao, Qunbo Wang, Jing Liu · 2024
Visual dialog (VisDial) requires models to answer questions based on both the dialog history and image contents. Traditional approaches simply extract relevant visual and textual information for answering questions, ignoring the relationships between the entities in the dialog and the relationships between the objects in the image. These fine-grained information is crucial to help correctly answer the questions in VisDial. In this work, we propose a Semantic-Visual Graph Reasoning framework (SVG) for VisDial. Specifically, we first construct a semantic graph to capture the semantic relationships between different entities in the current question and the dialog history. Secondly, we construct a semantics-aware visual graph to capture high-level visual semantics including key objects of the image and their visual relationships. Extensive experimental results on the VisDial v0.9 and v1.0 show that our method has shown superior performance compared to the state-of-the-art models across most evaluation metrics.