A Scalable Method for Validated Data Extraction from Electronic Health Records with Large Language Models
Timothy Joseph Stuhlmiller, Adrian Paul J. Rabe, Jeff Rapp, Alaa Awawda, Hiba Kouser, Kristi Lui, Hugh Salamon, Donald Chuyka, William P. Mahoney, John M. Furgason, Madhuri Paul, Frank J. Scarpa, Santosh Kesari, Mika Newton, Kenny K. Wong, Glenn A. Kramer, Mark A. Shapiro · medRxiv · 2025
Abstract Purpose/Background Healthcare organizations increasingly require structured, patient-level clinical variables for treatment decisions, operational workflows, quality measurement, and clinical trial screening. Relevant information is often fragmented across heterogeneous EHR systems, unstructured formats, and scanned documents. Large language models (LLMs) offer an opportunity to enhance medical record data extraction, particularly in oncology where point-of-care structured coding is often insufficient to capture a patient’s longitudinal clinical course. Methods Two complementary LLM-based approaches were developed. In the first, LLMs performed schema-based named entity recognition and relation extraction from unstructured documents; extracted variables were normalized to FHIR and OHDSI vocabularies. In the second, a retrieval-augmented checklist framework queried and integrated embedded document text with extracted structured data using prescriptive prompts to answer specific clinical questions, returning custom outputs with evidence-based justifications and source-document citations. Performance was evaluated through human validation, automated consistency checks, and iterative error analysis. Results Schema-based extractors designed to extract medications, radiation, and surgical procedure fields achieved∼95% accuracy, precision, recall, and F1. Deployed across 3,493 patients, LLM extraction yielded 71% more total medication records and a 207% increase in distinct oncology drug ingredients over structured C-CDA sources. Oncology therapies were captured for 65% of patients versus 40% in structured data, with dramatically improved clinical attribute coverage: indication for prescription (70.5% vs. 4.6%) and reason for discontinuation (9.7% vs. 0%). For radiation and surgical procedures — largely absent from structured records — LLM extraction yielded a 715% and 422% increase, respectively. The checklist extraction framework achieved significant F1 scores when designed to extract cancer diagnosis variables (99.0%) and lines of therapy (97.6%) across 4,802 validated elements. Cancer diagnoses and dates were identified for 93.5% of patients versus 69.8% in structured data; stage and grade were extracted for 64% and 62% of patients versus near-zero structured availability. A lines-of-therapy checklist generated 4,218 treatment lines across 2,320 patients, capturing regimen dates, best response, discontinuation reasons, and progression dates. Among patients with response data, objective response rate declined from 64.7% (95% CI 59.0–70.0%) at first line to 21.5% (95% CI 13.3–33.0%) at third line, and first-line response rates varied across tumor types from 42.9% in colorectal to 85.0% in esophageal cancer. Conclusions Leveraging two complementary LLM-based strategies, we substantially enhanced the completeness and utility of structured clinical data from heterogeneous medical records, supporting scalable generation of interoperable patient-level datasets for clinical analytics, research, and operational workflows.