Tokengram_F, a Fast and Accurate Token-based chrF++ Derivative

Sören Dreano, Derek Molloy, Noel A. Murphy · 2023

The tokengram_F metric presented in this paper is a novel approach to evaluating machine translation that has been submitted as part of the WMT23 challenge.It offers a new perspective on evaluating machine translation that takes advantage of modern tokenization algorithms to provide a more natural representation of the language in comparison to word n-grams.Tokengram_F is an F-score-based evaluation metric for Machine Translation that is heavily inspired by chrF++ and can act as a more accurate replacement.By replacing word n-grams with n-grams obtained from tokenization algorithms, tokengram_F captures similarities between words sharing the same semantic roots.While requiring minimal training based on an open corpus of monolingual datasets, the token-gram_F metric proposed still retains excellent performance that is comparable to more computationally expensive metrics.The tokengram_F metric demonstrates its versatility by showing satisfactory results, even when a tokenizer for a specific language is not available.In such cases, the tokenizer of a related language can be used instead, highlighting the adaptability of the tokengram_F metric to less commonly-used languages.

Read the paper · More papers on PaperTik