Large Language Models Lack Sufficient Performance to Provide Definitive Regulatory Guidance
Ryan C. Shean, Jonathan R. Genzen, Nicholas C. Spies · The Journal of Applied Laboratory Medicine · 2025
Clinical laboratory regulatory environments are rapidly evolving and marred with uncertainty. Remaining informed of—and compliant with—these regulations requires laboratory directors to make sense of the oft-unfamiliar language used in regulatory documents and take appropriate action. Large language models (LLMs) have exploded in popularity due to their ability to produce plausible, easily consumable responses with minimal user effort. This makes them a tantalizing option for those hoping to efficiently understand critical and complex regulatory documents. Unfortunately, LLM outputs are not always factually accurate. Response accuracy can, however, be improved through retrieval-augmented generation (RAG), which works by searching a predefined knowledgebase and appending similar text chunks to the beginning of the user-supplied prompt. Given the high-stakes nature of the content in question, it is crucial to objectively evaluate LLM responses focused on regulatory compliance. We formulated 20 questions covering important details in the FDA final rule on laboratory-developed tests (1), 510(k) and premarket approval submissions for medical devices (2), predetermined change control plans (3), clinical decision support (4), and software as a medical device (5).