An Automated Annotation Process for the SciDocAnnot Scientific Document Model
Hélène de Ribaupierre, Gilles Falquet · 2015
Answering precise and complex queries on a corpus of scien- tific documents requires a precise modelling of the document contents. In particular, each document element must be characterised by its dis- course type (hypothesis, definition, result, method, etc.). In this paper we present a scientific document model (SciAnnotDoc) that takes into account the discourse types. Then we show that an automated process can effectively analyse documents to determine the discourse type of each element. The process, based on syntactic rules (patterns), has been evaluated in terms of precision and recall on a representative corpus of more than 1000 articles in Gender studies. It has been used to create a SciDocAnnot representation of the corpus on top of which we built a faceted search interface. Experiments with users show that searching with this interface clearly outperforms standard keyword search for com- plex queries.