Twice Opportunity Knocks Syntactic Ambiguity: A Visual Question Answering Model with yes/no Feedback

Jianming Wang, Wei Deng, Yukuan Sun, Yuanyuan Li, Kai Wang, Guanghao Jin · 2019

Visual Question Answering (VQA) is a joint task that aims to answer questions based on given images. During dialogs between humans, syntactic ambiguity is a common phenomenon and it also could be found in the questions of VQA systems. Generally, the existing methods for VQA utilize one-shot answering frameworks, which will face a great difficulty if syntactic ambiguity occurs in questions. In human dialogs, people often conquer the problem by feeding back questions for confirmation. Inspired by this observation, we propose a novel method to eliminate the syntactic ambiguity in VQA via the user's feedback. We compared our method with the existing methods on two benchmark datasets, CLEVR and CLEVR-CoGenT. We found that the accuracy of our method is close to 100% on the CLEVR dataset. On the CLEVR-CoGenT dataset, our method is also 21% higher than the state-of-the-art method.

Read the paper · More papers on PaperTik