Multimodal Large Language Models in Medicine and Nursing: A Survey

Jing Liu, Linxiao Gong, Juncen Guo, Jingyi Wu, Liping Sun, Yulai Bi, Kartik Patwari, Boan Chen, L Zhang, Wei Zhou, Liu Yang, Xiaoguang Zhu, Chen‐Nee Chuah, Bala Rajaratnam · 2025

Large Language Models (LLMs) have demonstrated exceptional capabilities in medicine and nursing, garnering significant global attention from researchers worldwide. Building upon textual foundations, Multimodal LLMs (MLLMs) efficiently process diverse medical information, text, images, and audio, while demonstrating considerable advantages in disease diagnosis, prognosis prediction, and patient care. Furthermore, such models offer substantial value in promoting medical resource equity and reducing healthcare professional workloads. Although researchers have developed numerous specialized medical MLLMs with various datasets and training strategies, a comprehensive analysis bridging emerging technology with clinical applications remains essential for advancing commercialization and medical automation. To this end, we present the first systematic review of MLLMs in medicine and nursing, providing detailed guidance for practitioners and researchers. Our survey systematically examines MLLM architectures through taxonomic categorization of multimodal encoders, alignment modules, and LLM backbones. Subsequently, we explore training methodologies and their relationship to benchmark datasets. We analyze existing MLLMs across medical tasks and disease categories, evaluating current developmental landscapes. Additionally, we discuss MLLM roles in healthcare delivery, medical education, and equity promotion. Finally, we present critical challenges and future directions, including hardware limitations, data privacy concerns, and emerging technologies, offering insights for medical AI advancement.

Read the paper · More papers on PaperTik