Best practices for the human evaluation of automatically generated text

Chris van der Lee, Albert Gatt, Emiel van Miltenburg, Sander Wubben, Emiel Krahmer · 2019

Currently, there is little agreement as to how Natural Language Generation (NLG) systems should be evaluated, with a particularly high degree of variation in the way that human evaluation is carried out.This paper provides an overview of how human evaluation is currently conducted, and presents a set of best practices, grounded in the literature.With this paper, we hope to contribute to the quality and consistency of human evaluations in NLG.

Read the paper · More papers on PaperTik