Improving Distantly Supervised Document-Level Relation Extraction Through Natural Language Inference

Clara Vania, Grace Lee and Andrea Pierleoni · 2022

The distant supervision (DS) paradigm has been widely used for relation extraction (RE) to alleviate the need for expensive annotations.However, it suffers from noisy labels, which leads to worse performance than models trained on human-annotated data, even when trained using hundreds of times more data.We present a systematic study on the use of natural language inference (NLI) to improve distantly supervised document-level RE.We apply NLI in three scenarios: (i) as a filter for denoising DS labels, (ii) as a filter for model prediction, and (iii) as a standalone RE model.Our results show that NLI filtering consistently improves performance, reducing the performance gap with a model trained on human-annotated data by 2.3 F1. * Work completed at Amazon Alexa.The author now works at Thomson Reuters. 1 According to Yao et al. (2019), at least 40.7% facts in Wikipedia can only be extracted from multiple sentences.

Read the paper · More papers on PaperTik