Evaluation of NLG: some analogies and differences with machine translation and reference resolution
Andréi Popescu-Belis · 2007
This short paper first outlines an explanatory model that contrasts the evaluation of systems for which human language appears in their input with systems for which language appears in their output, or in both input and output. The paper then compares metrics for NLG evaluation with those applied to MT systems, and then with the case of reference resolution, which is the reverse task of generating referring expressions. 1 Challenges in NLG Evaluation Defining shared-task evaluation campaigns (STECs) is often the key to making progress in a particular domain, thanks to the convergence of several research teams. However, the definition of STECs requires an acceptable agreement, among a community of researchers, on the relevance of the selected problem to the domain, as well as on common evaluation metrics that indicate progress on this task. In the domain of Natural Language Generation (NLG), recent proposals have started meeting the challenge of STEC definition (Belz and Kilgarriff, 2006), few years after a new metric for Machine Translation (MT) evaluation (Papineni et al., 2001) had revived the interest for common evaluations, thanks to its low application costs, which in turn led to significant improvement of MT systems, and especially statistical ones. So, an important question is: how could NLG benefit from a similarly innovative metric, and how could such a metric be found? This short paper offers an explanation of the difficulty to evaluate NLG systems based on a typology of natural language processing (NLP) systems, and draws from this typology some suggestions for NLG evaluation (Section 2). Then, NLG evaluation is compared to MT evaluation (Section 3). Finally, the focus is set on referring expressions (REs), which have been used in the task proposed at the 2007 UC-NLG+MT workshop, and which might help providing an indirect measure of NLG “quality ” by combining the generation of REs with reference resolution