Automatic Text Evaluation through the Lens of Wasserstein Barycenters
Pierre Colombo, Guillaume Staerman, Chloé Clavel, Pablo Piantanida · Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing · 2021
A new metric BaryScore to evaluate text generation based on deep contextualized embeddings (e.g., BERT, Roberta, ELMo) is introduced.This metric is motivated by a new framework relying on optimal transport tools, i.e., Wasserstein distance and barycenter.By modelling the layer output of deep contextualized embeddings as a probability distribution rather than by a vector embedding; this framework provides a natural way to aggregate the different outputs through the Wasserstein space topology.In addition, it provides theoretical grounds to our metric and offers an alternative to available solutions (e.g., Mover-Score and BertScore).Numerical evaluation is performed on four different tasks: machine translation, summarization, data2text generation and image captioning.Our results show that BaryScore outperforms other BERT based metrics and exhibits more consistent behaviour in particular for text summarization.