HepatoAudit: A Comprehensive Dataset for Evaluating Consistency of Large Language Models in Hepatobiliary Case Record Diagnosis
HanZhi Zhang, Fangchen Dong, Weibin Li, Yi Ren, Hao Dong · 2025
In recent years, large language models (LLMs) have shown significant potential in the medical field. However, most existing research focuses on general medical test sets, with relatively insufficient exploration in real clinical applications, especially in the field of complex diseases. To fill this gap, we introduce the HepatoAudit dataset – the first comprehensive dataset designed to evaluate the consistency of LLMs in generating diagnostic records for hepatobiliary diseases. This dataset is collected from the Chinese Clinical Cases Outcome Database and contains complete medical records that have undergone peer review, including patient basic information, chief complaints, history of present illness, personal history, family history, past medical history, physical examination, and laboratory tests. We have carefully selected 20 common hepatobiliary diseases, totaling 684 medical records. We systematically evaluated six mainstream large language models using four prompting methods. The experimental results show that although all models perform high in precision (with GPT-4 reaching a maximum of 93.42%), there are still significant deficiencies in recall, with the recall rate generally below 70%. In particular, for diseases with atypical symptoms or caused by multiple factors, such as splenomegaly, fatty liver, hypertension, and type 2 diabetes, the accuracy of the models is even below 40%.The main reasons are the insufficient understanding of the patient's complete medical history and symptoms by the large language models, as well as the lack of effective reflection and error correction mechanisms. Our research not only points out the challenges of LLMs in professional medical diagnostic tasks but also provides important directions for future improvements.