“Dr LLM, what do I have?”: The Impact of User Beliefs and Prompt Formulation on Health Diagnoses

Wojciech Kusa, Edoardo Mosca, Aldo Lipani · 2023

The strong capabilities of conversation-based large language models in healthcare applications are attracting an increasingly larger audience.However, the reliability of these models in consistently and accurately providing medical advice based on user-inputted symptoms is a critical concern.This study explores the sensitivity of LLMs to variations in user input, focusing specifically on how different symptom descriptions and prior users' beliefs can potentially lead to different diagnoses.We test two GPT models with five different prompt templates to assess their ability of mentioning the true patient condition based on the symptoms.Our findings reveal a substantial sensitivity to input variations-especially when users have prior assumptions and beliefs-indicating potential inconsistencies in the diagnoses generated by these models.

Read the paper · More papers on PaperTik