Multimodal Clinical Prediction with Unified Prompts and Pretrained Large-Language Models
Caleb Winston, Chloe N. Winston, Cailin Winston, Claris Winston, Cleah Winston · 2024
Clinical prediction models (CPMs) increasingly rely on multiple modalities to predict clinical outcomes. The use of free-text data sources (e.g., chief complaint, medical notes) presents new challenges in the heterogeneity of the text across providers and patients and the need to convert the free text to numerical features before combining with other modalities. Prior work has employed multi-head architectures to learn separate heads for different modalities. In this work, we propose a new approach of constructing a unified text prompt that captures information from multiple modalities. We then use existing large-language models (LLMs), optimized with diagnosis-contrastive learning (DCL) to encode the unified prompt and make clinical predictions. We test our approach on the prediction of a variety of outcomes from emergency department visits using a free-text chief complaint and structured numerical data, including demographic information and vital signs. We find that optimized LLMs with unified prompts outperform LLMs that only use the chief complaint by 0.02 weighted F1 score (p < 0.0001) and models that only use the structured data modality as numerical inputs by 0.15 (p < 0.0001) in predicting acuity. We also observe improvements in the prediction of clinical outcomes, including hospital admission, length of stay, and time to revisit.