Towards Better Characterization of Paraphrases
Timothy Liu, De Wen Soh · Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) · 2022
To effectively characterize the nature of paraphrase pairs without expert human annotation, we propose two new metrics: word position deviation (WPD) and lexical deviation (LD).WPD measures the degree of structural alteration, while LD measures the difference in vocabulary used.We apply these metrics to better understand the commonly-used MRPC dataset and study how it differs from PAWS, another paraphrase identification dataset.We also perform a detailed study on MRPC and propose improvements to the dataset, showing that it improves generalizability of models trained on the dataset.Lastly, we apply our metrics to filter the output of a paraphrase generation model and show how it can be used to generate specific forms of paraphrases for data augmentation or robustness testing of NLP models.