Leveraging large language models in dermatology
Rubeta Matin, Eleni Linos, Neil Rajan · British Journal of Dermatology · 2023
Dermatology, a specialty where clinical images can play a role in diagnosis, is no stranger to disruptive advances in computational science that have the potential to alter clinical practice.1 A recent expansion in machine learning, a form of artificial intelligence (AI) termed ‘large language models (LLMs)’, is touted as set to change the human experience. Leveraging these for patient benefit in dermatology is key, but how do we best shape the conversation around LLMs to advocate for patient safety and improved care? Also, what research questions do we need to answer before these tools can be incorporated into clinical practice? Firstly, what are LLMs? To the clinician, imagine a resource able to offer a succinct response to any clinical question in real time, drawing on every accessible electronic text resource in existence, from the convenience of a phone. This ability builds on the field of natural language processing, which makes it feasible for LLMs to ‘make sense’ of written language. In effect, LLMs are ‘well read’, having the reading equivalent of thousands of human lifetimes, and can understand and interpret queries, which can help clinicians find relevant information rapidly. In the context of research, the advances are obvious: LLMs are intrinsically able to assess large datasets to extract relevant information, identify patterns and test hypotheses that would otherwise be very time-consuming to perform manually.2 The ability to integrate pictures into LLM queries also aligns well with specialties such as dermatology. For the academic, while this has meant that writer’s block has a new electronic friend, it has spurred a re-think of the role of written assessments for degree-awarding academic institutions. The bar exam for lawyers was used to make the point that one LLM, ChatGPT (Open AI, San Francisco, CA, USA), was an increasingly successful candidate with each iteration, an outcome echoed by its recent US medical board exam success.3 Clinical applications of LLMs have to be considered in the context of regulatory approval for AI/machine learning (ML) in general. Notably, while radiology has > 75% of 500+ US Food and Drug Administration-approved AI/ML enabled devices,4 dermatology has very few in comparison.5 In the UK, the Medicines and Health Regulatory Authority (MHRA) will be tasked to consider the use of AI and LLMs in the context of medical devices. Unlike general purpose LLMs such as ChatGPT and Bard (Google), LLMs intended for use in healthcare settings must be developed and trained using clinical (preferably real-life) data, and each should be designed with specific intended indications. The assessment of the burden of clinical evidence required for approval is key and currently rests with regulatory Notified Bodies and the MHRA (UK) for AI/ML and this is likely to also apply to clinical LLMs. Dermatology has had a head start to consider limitations and concerns of AI in practice. Ethical concerns for LLM use in clinical practice6 include privacy, particularly around the use of personal identifiable information involving human subjects. LLM training and testing needs to address the issue of cognitive biases, and the risk of perpetuating any bias present in training datasets used to develop LLMs. Understanding both the positive and negative implications of LLMs in clinical decision making is essential, as it is possible that LLMs could be used to remove human biases and reduce, instead of perpetuate, health disparities. There is a need to improve transparency of processes used in decision making, particularly if the LLM is being used to inform clinical decisions. This has recently begun to be addressed in skin cancer image AI.7 Clinicians also need to be aware of automation bias; reliance on LLM outputs can result in neglecting other reputable sources of information, critical thinking or multidisciplinary expertise in clinical decision making. Audit and evaluation of LLMs is already recognized as a necessary safeguard, and may help to improve deployed LLMs.8 Perhaps more important than whether LLMs pass the Turing test9 is whether clinicians would trust LLMs to respond to the clinical needs of our friends and family. LLMs do not currently have emotional intelligence, empathy, compassion and understanding that would be expected in clinical scenarios, such as relaying a diagnosis of melanoma. Words matter, and nuances, subtleties and cultural contexts that a human would easily recognize are currently lacking. For patients, dermatologists aided by LLMs that are robustly evaluated with real-world clinical data may improve quality of care, efficiency of clinical trials and triage decisions to improve access. It is crucial that physicians and patients are involved in the development, evaluation and deployment of LLMs. At the BJD we plan to take an active role in moving this field forward by inviting publications of rigorous research on the use of LLMs and AI in dermatology, specifically in order to address some of these key issues. We are grateful for helpful comments on this editorial from John Ingram. This editorial received no specific grant from any funding agency in the public, commercial or not-for-profit sectors. N.R.’s research is supported by the Newcastle NIHR Biomedical Research Centre. No data were generated. Not applicable.