Evaluation of the accuracy and readability of large language model responses on menopause and hormone therapy
Jana Karam, CHRISANDRA L. SHUFELT, Nancy Safwan, Ekta Kapoor, Monica Christmas, Stephanie S. Faubion · Menopause The Journal of The North American Menopause Society · 2025
OBJECTIVE: Generative artificial intelligence is rapidly evolving and is now being explored in health care to support patient and clinician education. This study evaluated the accuracy, completeness, and readability of four large language models (LLMs): ChatGPT 3.5, Gemini, ChatGPT 4.0, and OpenEvidence in answering questions about menopause and hormone therapy. METHODS: A total of 35 questions (20 patient-level, 15 clinician-level) were entered into each LLM. OpenEvidence was only used for clinician-level questions. Four blinded expert reviewers rated responses as accurate and complete, accurate but incomplete, or inaccurate. Readability of patient-level responses was assessed using the Flesch Reading Ease Score (FRES) and word count. Analysis used ANOVA for readability, odds ratios for accuracy comparisons. RESULTS: For patient-level questions, ChatGPT 3.5 achieved the highest accuracy (70%), followed by ChatGPT 4.0 (60%) and Gemini (30%); Gemini had significantly lower odds of accuracy compared with ChatGPT 3.5 (OR=0.18, 95% CI=0.05-0.71; P =0.014). FRES scores differed significantly ( P 0.05). CONCLUSION: LLMs demonstrated limited accuracy and frequent incorrect or incomplete responses to menopause-related queries, highlighting the need to improve model performance to ensure accurate and reliable information for both patients and clinicians.