A Multimodal Interpretable Visual Question Answering Model Introducing Image Caption Processor

He Zhu, Ren Togo, Takahiro Ogawa, Miki Haseyama · 2022 IEEE 11th Global Conference on Consumer Electronics (GCCE) · 2022

This paper presents a multimodal interpretable visual question answering (VQA) model to solve the lack of overall understanding of the previous models, which is caused by focusing too much on the question. Compared with the previous models with only image reference, we newly introduce an image caption processor to improve the prediction process of the VQA model. The experimental results show that our model can generate the answers more accurately and the explanations are more reasonable and closer to the truth than some state-of-art methods.

Read the paper · More papers on PaperTik