Evaluating Paraphrastic Robustness in Textual Entailment Models

Dhruv Verma, Yash Kumar Lal, Shreyashee Sinha, Benjamin Van Durme, Adam Poliak · 2023

We present P aRT E, a collection of 1,126 pairs of Recognizing Textual Entailment (RTE) examples to evaluate whether models are robust to paraphrasing.We posit that if RTE models understand language, their predictions should be consistent across inputs that share the same meaning.We use the evaluation set to determine if RTE models' predictions change when examples are paraphrased.In our experiments, contemporary models change their predictions on 8-16% of paraphrased examples, indicating that there is still room for improvement. PThe cost of security when world leaders gather near Auchterarder for next year 's G8 summit, is expected to top $150 million.P' The cost of security when world leaders meet for the G8 summit near Auchterarder next year will top $150 million.H More than $150 million will be probably spent for security at next year's G8 summit.H' At the G8 summit next year more than $150 million will likely be spent on security at the event.

Read the paper · More papers on PaperTik