SEETrials: Leveraging large language models for safety and efficacy extraction in oncology clinical trials
Kyeryoung Lee, Hunki Paek, Liang‐Chin Huang, Cameron Beau Hilton, Surabhi Datta, Josh Higashi, Nneka Ofoegbu, Jingqi Wang, Samuel M. Rubinstein, Andrew J. Cowan, Mary Kwok, Jeremy Lyle Warner, Hua Xu, Xiaoyan Wang · Informatics in Medicine Unlocked · 2024
Background: Initial insights into oncology clinical trial outcomes are often gleaned manually from conference abstracts. We aimed to develop an automated system to extract safety and efficacy information from study abstracts with high precision and fine granularity, transforming them into computable data for timely clinical decision-making. Methods: We collected clinical trial abstracts from key conferences and PubMed (2012-2023). The SEETrials system was developed with three modules: preprocessing, prompt engineering with knowledge ingestion, and postprocessing. We evaluated the system's performance qualitatively and quantitatively and assessed its generalizability across different cancer types- multiple myeloma (MM), breast, lung, lymphoma, and leukemia. Furthermore, the efficacy and safety of innovative therapies, including CAR-T, bispecific antibodies, and antibody-drug conjugates (ADC), in MM were analyzed across a large scale of clinical trial studies. Results: heterogeneity index scores) across several outcome entities analyzed within therapy subgroups. Conclusion: SEETrials demonstrated highly accurate data extraction and versatility across different therapeutics and various cancer domains. Its automated processing of large datasets facilitates nuanced data comparisons, promoting the swift and effective dissemination of clinical insights.