COVID-19 Knowledge Mining Based on Large Language Models and Chain-of-Thought Reasoning
Fushuai Zhang, Yunchao Ling, Jiaxin Yang, P. Zhang, Guoqing Zhang · 2024
This research aims to extract knowledge related to COVID-19 clinical diagnosis and molecular mechanisms from COVID-19 literature by applying large language models, deepening our quantitative understanding of COVID-19 and its causative virus SARS-CoV-2, and providing foundational data resources for subsequent COVID-19 research. The knowledge extraction process is divided into three stages:The first phase: A total of 131 key documents, including COVID-19 research, vaccine clinical trials, and clinical guidelines, were selected for content extraction. Entities, relationships, entity types, and conditions were annotated, and knowledge was described in the form of triplets, using events as the unit. By summarizing the important and frequently occurring events, entities, and relationships, as well as analyzing the clinical diagnosis knowledge of COVID-19, nine COVID-19 clinical diagnosis knowledge patterns were identified. Subsequently, a corpus containing 342,287 articles was constructed using search strategies from PubMed and the CORD-19 database.The Second phase: Chain-of-thought reasoning was applied to enhance the data, simulating human experts in determining the knowledge patterns of COVID-19 sentences and extracting the COVID-19 entities and their types through a step-by-step reasoning process.The Third phase: The focus was on training large language models to achieve superior Knowledge Pattern Discrimination (KPD) and Named Entity Recognition (NER) performance, the Accuracy of the KPD model and the F1-score of the NER model reached 0.869 and 0.887 respectively, surpassing existing methods. Additionally, ablation studies indicated that the data formats used in chain-of-thought reasoning and model fine-tuning affect the performance of the models.