Improving Large-scale Paraphrase Acquisition and Generation

Yao Dou, Chao Jiang, Wei Hong Xu · 2022

This paper addresses the quality issues in existing Twitter-based paraphrase datasets, and discusses the necessity of using two separate definitions of paraphrase for identification and generation tasks.We present a new Multi-Topic Paraphrase in Twitter (MULTIPIT) corpus that consists of a total of 130k sentence pairs with crowdsoursing (MULTIPIT CROWD ) and expert (MULTIPIT EXPERT ) annotations using two different paraphrase definitions for paraphrase identification, in addition to a multi-reference test set (MULTIPIT NMR ) and a large automatically constructed training set (MULTIPIT AUTO ) for paraphrase generation.With improved data annotation quality and task-specific paraphrase definition, the best pre-trained language model fine-tuned on our dataset achieves the stateof-the-art performance of 84.2 F 1 for automatic paraphrase identification.Furthermore, our empirical results also demonstrate that the paraphrase generation models trained on MUL-TIPIT AUTO generate more diverse and highquality paraphrases compared to their counterparts fine-tuned on other corpora such as Quora, MSCOCO, and ParaNMT.Topic Domains #Train #Dev #Test Sent/Tweet Len %Paraphrase #Trends/URLs #Uniq Sent %Multi-Ref Our Multi-Topic Paraphrase in Twitter (MULTIPIT CROWD ) Dataset Trends Sports 25,255 3,157 3,157 10.24 / 13.79 40.52% 1,201 34,786 17.89% Entertainment 11,547 1,443 1,444 10.44 / 13.80 62.33% 610 15,784 18.11% Event 8,624 1,078 1,079 10.86 / 15.32 82.83% 359 11,746 17.75% Others 17,751 2,219 2,219 10.41 / 14.56 67.

Read the paper · More papers on PaperTik