Surface Form Competition: Why the Highest Probability Answer Isn’t Always Right
Ari Holtzman, Peter West, Vered Schwartz, Yejin Choi, Luke Zettlemoyer · Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing · 2021
Large language models have shown promising results in zero-shot settings (Brown et al., 2020;Radford et al., 2019).For example, they can perform multiple choice tasks simply by conditioning on a question and selecting the answer with the highest probability.We introduce Domain Conditional Pointwise Mutual Information, an alternative scoring function that directly compensates for surface form competition by simply reweighing each option according to its a priori likelihood within the context of a specific task.It achieves consistent gains in zero-shot performance over both calibrated (Zhao et al., 2021) and uncalibrated scoring functions on all GPT-2 and GPT-3 models on a variety of multiple choice datasets.