Vιsual question answering models Evaluation
S Sarath, J. Amudha · 2020 International Conference for Emerging Technology (INCET) · 2020
Visual question answering (VQA), visual dialogs, visual chat bot are multi-discipline exploration problems, which is a blend of Natural Language Processing (NLP), Image feature extraction and Knowledge Reasoning (KR). Rather than captioning, which is naïve approach of computer vision, VQA problems enhances the perspective by providing interactivity to ask domain specific as well as open ended questions to images and give us the insights based on image features or characteristics. Our research is the evaluate the performance of VQA on counting problems. Given an image, VQA model is expected to answer "how many" question type. We have used few pre-trained models for VQA and visual dialog and tabulated the findings of accuracy of predicted answer with the pre-defined ground truth.