Non-monotonic Logical Reasoning and Deep Learning for Explainable Visual Question Answering

Heather Riley, Mohan Sridharan · 2018

State of the art visual question answering (VQA) methods rely heavily on deep network architectures. These methods require a large labeled dataset for training, which is not available in many domains. Also, it is difficult to explain the working of deep networks learned from such datasets. Towards addressing these limitations, this paper describes an architecture inspired by research in cognitive systems that integrates commonsense logical reasoning with deep learning algorithms. In the context of answering explanatory questions about scenes and the underlying classification problems, the architecture uses deep networks for processing images and for generating answers to queries. Between these deep networks, it embeds components for non-monotonic logical reasoning with incomplete commonsense domain knowledge and for decision tree induction. Experimental results show that this architecture outperforms an architecture based only on deep networks when the training dataset is small, provides comparable performance on larger datasets, and provides intuitive answers to explanatory questions.

Read the paper · More papers on PaperTik