Multimodal Medical Image Analysis: Integrating LLM and RAG Deep Learning Strategies

Hanrui Yan, Dan Shao · Journal of Advances in Information Technology · 2025

This study aims to explore a method combining Retrieval-Augmented Generation (RAG), Prompt Learning for Multimodal Large Language Models (MLLM), and Deep Self-Supervised Learning (DSL) to enhance the efficiency and accuracy of medical data management and analysis, particularly in medical image processing and diagnostic tasks.We propose a novel medical MLLM framework that integrates RAG to strengthen knowledge retrieval capabilities and optimizes model generation quality through a carefully designed prompt learning mechanism.Additionally, we incorporate DSL to uncover critical features from unlabeled medical data via self-supervised tasks, thereby improving the model's learning capability.The framework design ensures secure data training and dynamically adjusts retrieval context and prompt formatting to adapt to diverse medical scenarios.Extensive experiments were conducted on various medical datasets, including radiology, ophthalmology, and pathology, covering medical Visual Question Answering (VQA) and report generation tasks.Experimental results demonstrate that the proposed framework significantly outperforms existing methods in factual accuracy, generation quality, and model adaptability.The findings of this study indicate that the integrated approach combining RAG, MLLM prompt learning, and DSL effectively enhances medical data processing performance, verifying its feasibility for secure and efficient data management in medical contexts.This innovative framework provides new ideas and approaches for future medical AI applications, driving the intelligent development of the healthcare industry.

Read the paper · More papers on PaperTik