Evaluating fine-tuned GPT models on different datasets in the healthcare domain

Yun Hui Lim, Pauline Shan Qing Yeoh, Khin Wee Lai · Innovation and Emerging Technologies · 2025

This study investigates the performance of fine-tuned generative pre-trained transformers (GPT) on different healthcare datasets to enhance public health literacy. The background of this study is rooted in the recognition of the critical role health literacy plays in fostering public awareness and understanding of medical information. Against this backdrop, the objective is to explore domain-specific GPT that enhance accessibility to comprehensive health information. This study fine-tunes the GPT model across different types of datasets, which are PubMed, Medical Information Mart for Intensive Care III (MIMIC-III), MedQA, MedMCQA, and consultation datasets. The models are evaluated using Massive Multitask Language Understanding and Massive Multi-discipline Multimodal Understanding Benchmarks. Results showed that fine-tuning language models on domain-specific datasets, especially in healthcare, substantially improves their performance. In conclusion, this research highlights the potential of domain-specific GPT in the modern healthcare landscape.

Read the paper · More papers on PaperTik