Prompting for explanations improves Adversarial NLI. Is this true? {Yes} it is {true} because {it weakens superficial cues}
Pride Kavumba, Ana Brassard, Benjamin Heinzerling, Kentaro Inui · 2023
Explanation prompts ask language models to not only assign a label to a given input, such as entailment or contradiction in natural language inference (NLI) tasks, but also to generate a free-text explanation that supports this label.While explanation prompts originally introduced aiming to improve model interpretability, here we show that they also improve robustness to superficial cues.Compared to prompting for labels only, explanation prompting shows stronger performance on adversarial NLI benchmarks, outperforming the state of the art on ANLI, Counterfactually-Augmented NLI, and SNLI-Hard datasets.Analysis suggests that the increase in robustness is due to a reduction in the association strength between single tokens and labels, i.e., explanation prompting weakens superficial cues.More specifically, we find that single tokens that are highly predictive of the correct answer in the label-only setting become uninformative when the model also has to generate explanations.