Cost-based attribute selection for GRE (GRAPH-SC/GRAPH-FP)
Mariët Theune, Pascal Touset, Jette Viethen, Emiel Krahmer · University of Twente Research Information · 2007
In this paper we discuss several approaches to the problem of content determination for the generation of referring expressions (GRE) using the Graphbased framework of Krahmer et al. (2003). This work was carried out in the context of the First NLG Shared Task and Evaluation Challenge on Attribute Selection for Referring Expression Generation. In the shared task proper of the Challenge the output of submitted systems is evaluated against a corpus of human-produced descriptions for the same input. The data used for training, development and testing is a subset of the Aberdeen TUNA Corpus, a human-generated data set designed for the attribute selection task. The data set is divided into a Furniture domain and a People domain, focussing on descriptions of singular entities. Attributes in both domains are type, orientation and x- and y-dimensions. Additional attributes in the Furniture domain are colour and size, and additional attributes in the People domain are age, orientation, hair colour, and the presence or absence of hair, beards, glasses and different clothing items. Each input is a pair consisting of the XML representation of the attributes of all objects contained in a scene and a pointer indicating the intended referent. The output descriptions of the target are lists of attributes, also encoded in XML. Evaluation of the submitted systems is based on the Dice coefficient of similarity between system output and the human-produced data from the TUNA Corpus. There will also be a small-scale task-based evaluation with human participants. 1 1 For details on the TUNA corpus and the GRE Chal-