Robustness tests for biomedical foundation models should tailor to specifications

Rui Patrick Xian, Noah Baker, Tom David, Qiming Cui, A Jay Holmgren, Stefan M. Bauer, Madhumita Sushil, Reza Abbasi-Asl · npj Digital Medicine · 2025

The growing presence of biomedical foundation models (BFMs), including large language models (LLMs), vision-language models (VLMs), and others, trained using biomedical or de-identified healthcare data, suggests they will eventually become integral to healthcare automation. Discussions on the risks of deploying algorithmic decision-making and generative AI in medicine have focused on bias and fairness. Robustness 1 is an equally important topic, which generally refers to the consistency of model prediction to distribution shifts. It is quantified using aggregated performance metrics, stratified comparisons across subsets of data, and worst-case performance. Robustness failures are an origin of the performance gap between model development and deployment, performance degradation over time, and, more alarmingly, the generation of misleading or harmful content by imperfect users or bad actors 2 . The robustness of software also affects the legal responsibilities of providers 3 because the software may cause harm (e.g., misinformation, financial loss, or injury) to users or third parties or require regulatory body authorization for deployment (e.g., medical devices) 4 , 5 .

Read the paper · More papers on PaperTik