What can you do with a large language model?
Suzanne R. Bakken · Journal of the American Medical Informatics Association · 2024
The Journal of the American Medical Informatics Association (JAMIA) has published papers on large language models (LLM) over the last 5 years. For example, a 2019 paper by Si et al. explored utilizing LLM for clinical concept extraction, including comparing Bidirectional Encoder Representations from Transformers (BERT) to traditional word embedding methods (word2vec, GloVe, fastText). Their experiments demonstrated that contextual embeddings encode valuable semantic information not accounted for in traditional word representations.1 The broad release of ChatGPT 3.5 resulted in a flurry of submissions to JAMIA and motivated a forthcoming focus issue on ChatGPT and LLM in Biomedicine and Health with a particular emphasis on methodological innovation as well as associated ethical, legal, and social implications. In this issue, I highlight 5 papers that explore different areas of application of LLMs including generative AI. Xie et al. evaluated an epilepsy-specific LLM (ClinicalBERT that had been fine-tuned on 700 manually annotated epileptologist notes) for intrinsic bias and used LLM-extracted outcomes to determine if demographic groups varied in freedom from seizures at each office visit.2 In a sample of 84 675 clinic visits from 25 612 unique patients, they found little evidence of bias in prediction accuracy and confidence of outcome classifications across demographic groups (race, ethnicity, sex, income, and health insurance). Females, those with public insurance, and those living in lower-income zip code areas had significantly worse outcomes, that is, freedom from seizures. This study contributes to the body of evidence regarding application of LLMs to examine health disparities.