A Finetuned Multimodal Deep Learning Model for Xray Image Diagnosis

Wei Ren Chen, Honghao Wang, Zhihao Feng, Huafeng Mai · 2024

Generative conversational AI has shown significant potential in aiding biomedical professionals, yet existing research predominantly targets single-mode textual communication. While multimodal conversational AI has advanced quickly through utilizing vast numbers of image-text pairs available online, these broad-domain vision-language models are still insufficiently refined in their ability to comprehend and discuss biomedical imagery. This article primarily addresses question-answering (QA) for medical imaging data and cancer diagnosis. We constructed a relevant instruction-based QA dataset and performed directive fine-tuning on the Qwen-VL 7B model. Thanks to the use of QLora technology, our model can be trained quickly on a Nvidia RTX 4090 graphics card and achieves better results compared to other multimodal models. The experiments showed that the Qwen- VL-7B model has best Rouge-1, Rouge-2, Rouge-L score compared to other models, meaning its highly potential for medical question-answering application.

Read the paper · More papers on PaperTik