Efficient Training Method for Phrase Extraction Models using Natural Language Explanations

Ryosuke Saito, Koga Kobayashi, Kei Wakabayashi · 2021

Phrase extraction is an information extraction task that extracts words or phrases in a specific category from text data, which is used in various downstream NLP technologies, including named entity recognition (NER), terminology extraction, question answering, dialogue systems, and information integration. While we need a large amount of annotated corpus for training a phrase extractor using machine learning, building such a corpus requires a lot of manual annotation work by domain experts. This research aims to reduce the annotation cost by developing a method that trains a phrase extractor using natural language explanations from experts. The proposed method transforms the natural language explanations into labeling functions, which allows us to make pseudo annotated corpus from a set of raw sentences. We empirically show the effectiveness of the model through experimental results.

Read the paper · More papers on PaperTik