Towards a Comprehensive Metric for Evaluating Text Simplification Systems
AlMotasem Bellah Al Ajlouni, Jinlong Li, Mo'ataz A. Ajlouni · 2023
Accurately evaluating text simplification (TS) systems is a major challenge in optimizing existing models and developing new TS approaches. Human evaluation stands as the gold standard for TS quality assessment, it involves expert human annotators who assess simplified texts based on grammaticality (G), meaning preservation (M), and simplicity (S) criteria. Typically, the overall quality of a TS system is calculated by taking the arithmetic mean of these criteria. However, this method can be misleading due to the complex relationships between these criteria. This study investigates the nuanced relationships between the G, M, and S criteria and introduces two new formulas for combining them into a unified metric that more accurately reflects the overall quality of a TS system. The experimental results highlight the effectiveness of BERTscore and BLEU in evaluating TS systems, particularly when high-quality references are available, and their strong correlations with the unified metrics we propose to express overall system performance.