Evaluating HILDA in the CODA Project: A Case Study in Question Generation Using Automatic Discourse Analysis
Pascal Kuyten, Hugo Hernault, Helmut Prendinger, Mitsuru Ishizuka · 2011
Recent studies on question generation identify the need for automatic discourse analysers. We evaluated the feasibility of integrating an available discourse analyser called HILDA for a specific question generation system called CODA; introduce an approach by extracting a discourse corpus from the CODA parallel corpus; and identified future work towards automatic discourse analysis in the domain of question generation. Question Generation Question generation is an important and challenging component of systems where knowledge extraction and representation in natural language is desired (Rus et al. 2010). Rus and Graesser defined Question Generation as the task of automatically generating of questions from some form of input. The input could vary from information in a database to a deep semantic representation to raw text (Rus and Graesser 2009). Studies on question generation can be classified by considering their scoping. Some generate questions at sentence level and others generate questions at paragraph level (Rus et al. 2010). When considering question generation at paragraph level the discourse relations become important (Heilman 2011). Two studies which are generating questions at paragraph level are the CODA project (Piwek and Stoyanchev 2010a) and the work by Mannem et al. (Mannem, Prasad and Joshi 2010). The CODA project is a two years program that started in 2009 and is dedicated to generation of dialogue. Question generation is identified as an important component in the generation of dialogue. Dialogue is generated from monologue text using the CODA system. A part of Copyright © 2011, Association for the Advancement of Artificial