Data Augmentation for Implicit Discourse Relation Recognition

Yatian Shen, Ning Liu, Xiaojie Hu · 2022

Fine-tuning pre-trained models on downstream tasks tends to have decent performance. Although the pre-training model has learned a lot of knowledge with strong generalization ability in the pre-training stage, the benefits of fine-tuning are not high on the IDRR, a challenging task of NLU. On the one hand, the low amount of data makes the model lose its original generalization ability, and on the other hand, it is difficult for the pre-trained model to obtain enough discourse-level information from the limited data. Therefore, we propose an approach based on data augmentation. The model trains task-related soft prompt on a pre-trained model with frozen parameters. The data is then generated while maintaining the generalization performance of the generative model as much as possible. Furthermore, we use a consistency filtering algorithm to remove the generated noise samples. Experiments on the PDTB benchmark show that proposed method improves the performance of the baseline model.

Read the paper · More papers on PaperTik