Responses to queries concerning “Performance of large language models on benign prostatic hyperplasia frequently asked questions”

Yuning Zhang, Jianqiao Zhou · The Prostate · 2024

We thank Hinpetch Daungsupawong and Viroj Wiwanitkit for their interest in our work.1 Our study categorized the responses generated by the three different Large language model (LLMs) into four grades based on correctness and comprehensiveness. With this definition, we use the accuracy rate as the main indicator for LLMs' performance. However, as they mentioned, relying solely on the accuracy rate to assess LLMs' performance is limited and incomplete. Other indicators, such as specificity and the depth of responses generated by LLMs, are also important for evaluating their performance. Therefore, in subsequent related studies, we will consider incorporating specificity and depth into the criteria for grading answers or using them as additional indicators for assessing LLMs' performance. We agree with their opinions that any biases or restrictions in the training data that the LLMs were developed using could have an impact on the replies' accuracy and dependability, however, due to the inaccessibility of the data sets that the LLMs were trained on, an assessment in this regard is difficult to do at this time. In addition, their comments on the future direction of LLMs research, such as evaluating the efficacy of LLMs in answering more complex questions, how to improve the reproducibility of LLMs, and developing standards to make LLM-generated content more ethical and transparent, are all very valuable and worth thinking about. Not limited to us, we feel that all researchers interested in the application of LLMs in medicine should fully consider their valuable opinions in future research.

Read the paper · More papers on PaperTik