MedIE-Instruct: A Comprehensive Instruction Dataset for Medical Information Extraction
Zhuoyi Xiang, Xinda Wang, Xiaodong Yan, Deng Zhao, Keyan Ding, Qiang Zhang · 2024
Medical information extraction (IE) tasks, including named entity recognition (NER), relation extraction (RE), and event extraction (EE), are crucial for constructing medical knowledge graphs from unstructured text. However, existing IE datasets in the medical domain are often limited, fragmented, and lack uniformity. To address these issues, we present MedIE-Instruct, a comprehensive bilingual (English and Chinese) medical IE instruction corpus comprising 11 datasets with over 100,000 instructions. This corpus was constructed using schema-based instruction generation to create a large-scale, diverse IE dataset to support large language models (LLMs) in medical IE tasks. Experimental results show that fine-tuning state-of-the-art LLMs with MedIE-Instruct significantly enhances model performance, especially in zero-shot scenarios. We hope this dataset and research findings will provide valuable resources and insights for understanding and processing medical content.