Benchmarking Answer Verification Methods for Question Answering-Based Summarization Evaluation Metrics

Daniel Deutsch, Dan Roth · Findings of the Association for Computational Linguistics: ACL 2022 · 2022

Question answering-based summarization evaluation metrics must automatically determine whether the QA model's prediction is correct or not, a task known as answer verification.In this work, we benchmark the lexical answer verification methods which have been used by current QA-based metrics as well as two more sophisticated text comparison methods, BERTScore and LERC.We find that LERC out-performs the other methods in some settings while remaining statistically indistinguishable from lexical overlap in others.However, our experiments reveal that improved verification performance does not necessarily translate to overall QA-based metric quality: In some scenarios, using a worse verification method -or using none at all -has comparable performance to using the best verification method, a result that we attribute to properties of the datasets.1

Read the paper · More papers on PaperTik