TM-PATHVQA: 90000+ Textless Multilingual Questions for Medical Visual Question Answering
Tonmoy Rajkhowa, Amartya Roy Chowdhury, Sankalp Nagaonkar, Achyut Mani Tripathi, Mahadeva Prasanna · 2024
In healthcare and medical diagnostics, Visual Question Answering (VQA) may emerge as a pivotal tool in scenarios where analysis of intricate medical images becomes critical for accurate diagnoses.Current text-based VQA systems limit their utility in scenarios where hands-free interaction and accessibility are crucial while performing tasks.A speech-based VQA system may provide a better means of interaction where information can be accessed while performing tasks simultaneously.To this end, this work implements a speech-based VQA system by introducing a Textless Multilingual Pathological VQA (TM-PathVQA) dataset, an expansion of the PathVQA dataset, containing spoken questions in English, German & French.This dataset comprises 98,397 multilingual spoken questions and answers based on 5,004 pathological images along with 70 hours of audio.Finally, this work benchmarks and compares TM-PathVQA systems implemented using various combinations of acoustic and visual features.