From Chat to Checkup: Can Large Language Models Assist in Diabetes Prediction?
Shadman Sakib, Oishy Fatema Akhand, Ajwad Abrar · 2025
While Machine Learning (ML) and Deep Learning (DL) approaches are commonly used for diabetes prediction, the potential of Large Language Models (LLMs) on structured numerical data remains underexplored. This study investigates the effectiveness of LLMs for diabetes classification using zero-shot, one-shot, and three-shot prompting. We conduct experiments on the Pima Indians Diabetes Dataset (PIDD), evaluating six LLMs-four open-source (Gemma-2-27B, Mistral-7B, LLaMA-3.1-8B, LLaMA-3.2-2B) and two proprietary models (GPT-4o, Gemini Flash 2.0)-and compare them with three traditional ML models: Random Forest, Logistic Regression, and Support Vector Machine (SVM). Performance is assessed using accuracy, precision, recall, and F1-score. Results show that proprietary LLMs outperform their open-source counterparts, with GPT-4o and Gemma-2-27B achieving the highest accuracy in few-shot settings. Gemma-2-27B also surpasses traditional ML models in F1-score. Despite their promise, LLMs exhibit sensitivity to prompt variations and may benefit from domain-specific finetuning. Our findings highlight the potential of LLMs in medical prediction tasks and underscore the importance of prompt design and hybrid modeling strategies in healthcare applications.