Leveraging LLM for Enhancing Document-Level Relation Extraction with Correction and Completion
Huageng Zhong, Xiao Juan Wei, Huiran Zhang · 2025
Document-level Relation Extraction (DocRE) is a crucial task in information extraction, particularly for knowledge graph construction and other downstream applications. Despite its importance, DocRE faces significant challenges due to the complex semantics of documents, long-distance dependencies, and the large number of candidate triplets. While traditional Transformer-based models, such as BERT, have made considerable progress, they rely on large annotated datasets and have reached a performance plateau. In contrast, large language models (LLMs) have demonstrated impressive success in tasks like text generation and understanding. However, they still fall short of traditional methods in terms of performance and encounter limitations in document-level extraction tasks. This paper introduces a novel DocRE approach that integrates BERT's efficient semantic vectorization with the advanced contextual understanding and reasoning capabilities of LLMs. Our method first uses the BERT-base model to retrieve and classify candidate relation triplets, and then employs an LLM to further assess, correct, and refine these relations, addressing issues such as false positives, false negatives, and irrelevant entities. We propose a two-stage framework that ensures both efficiency and accuracy in extracting relations from documents. Experimental results on the DocRED and Re-DocRED datasets demonstrate the effectiveness of our method, highlighting its ability to correct erroneous relations and eliminate unsubstantiated ones. This method provides a scalable solution that can be further fine-tuned and extended, making it a promising direction for future research in DocRE.