The INLG 2024 Tutorial on Human Evaluation of NLP System Quality: Background, Overall Aims, and Summaries of Taught Units

Anja Belz, João Sedoc, Craig Thomson, Simon Mille, Rudali Huidrom · 2024

Following numerous calls (e.g.van der Lee et al., 2019;Howcroft et al., 2020; Thomson et al., 2024) in the literature for improved practices and standardisation in human evaluation in Natural Language Processing over the past ten years, we held a tutorial on the topic at the 2024 INLG Conference.The tutorial addressed the structure, development, design, implementation, execution and analysis of human evaluations of NLP system quality.Hands-on practical sessions were run, designed to facilitate assimilation of the material presented.Slides, lecture recordings, code and data have been made available on GitHub. 1 In this paper, we provide summaries of the content of the eight units of the tutorial, alongside its research context and aims. Research Context and AimsHuman evaluation is widely considered the most reliable form of evaluation in Natural Language Processing (NLP), but recent research has thrown up a number of concerning issues, including in the design (Belz et al., 2020;Howcroft et al., 2020) and execution (Thomson et al., 2024) of human evaluation experiments, but also obstacles in adopting good practices (Gehrmann et al., 2023).Standardisation and comparability across different experiments is low, as is reproducibility in the sense that repeat runs of the same evaluation often do not support the same main conclusions, quite apart from not producing similar scores.The situation is likely to be in part due to how human evaluation is viewed in NLP: not as something that needs to be studied and learnt before venturing into conducting an evaluation experiment, but something that anyone can throw together without prior knowledge by pulling in a couple of students from the lab next door.1 https://github.com/Human-Evaluation-Tutorial/ INLG-2024-Tutorial

Read the paper · More papers on PaperTik