A Medical Domain Visual Question Generation Model via Large Language Model
He Zhu, Ren Togo, Takahiro Ogawa, Miki Haseyama · 2023
This paper proposes a medical visual question generation model for generating higher-quality questions from medical images. The visual question generation model can guide the diagnostic process and improve the utilization of medical resources by reducing the dependence on physician involvement. Our model uses cross-attention and the large language model to preserve inherent information and addresses the issue of inferior generation performance in the medical domain due to a lack of data. We also control the category of generated questions by setting guidance sentences that include interrogative words. The experimental results demonstrate that our method generates higher-quality questions than previous approaches.