A framework for corpus-based analysis of the graphic signalling of discourse structure
Martin Thomas, Judy Delin, Robert Waller · 2010
This paper describes a corpus-based approach to the analysis of graphic text sig- nalling in complex information documents. To make the task of populating the corpus tractable, we have developed software to automate as much of the annotation process as possible. OCR output is first obtained in OpenDocument format. This is post-processed semi-automatically to generate stand-off XML annotations following the GeM model (Henschel, 2003). These gen- erated layers describe the content and layout of the document. This information is augmented with functionally-oriented descriptions and RST analyses (Mann and Thompson, 1987). To- gether these annotations support empirical research into the relationship between the things that are said in documents and the linguistic and graphic resources used to express them. Such research might inform the evaluation of existing documents and the design of new ones.