LLaVA-NeXT-Med: Medical Multimodal Large Language Model
Yunfei Guo, Wu Xin Huang · 2025
This paper introduces the LLaVA-NeXT-Med model, a medical multi-modal system that is pretrained and fine tuned using datasets incorporating LLaVA-Med data to enhance the performance of medical image processing and visual question answering (VQA). The purpose of this study was to evaluate the potential application of LLaVA-NeXT-Med in the medical field and to verify its effectiveness through a series of experiments. First, the LLaVA-NeXT and LLaVA-Med fusion data were used for pretraining, followed by visual fine tuning using the fusion dataset. Through offline experiments on three medical VQA datasets, the results show that LLaVA-NeXT-Med outperforms the original model on several evaluation indicators. The experimental results demonstrate the effectiveness of LLaVA-NeXT-Med in improving medical image analysis and VQA tasks.