Document Content Analysis through Inverted Generation
Marc Dymetman · 2003
A paradigm for the deep content analysis of documents in restricted domains is proposed, along with an implementation technique relying on the emergent field of interactive natural language generation. A paradigm for document content analysis The paradigm sees the formal specification of constrained content representations as a prerequisite for making sense of the documents and makes use of these representations for predicting textual aspects of the documents. Access to syntactic information is resorted to only when it is deemed necessary for disambiguating between two well-formed content representations, thus achieving a better division of labor between the highly constrained content space and the much looser syntactic/textual space. The approach relies on a mechanism for producing intermediate structures from the content representations which can be used to perform a fuzzy match with the text of the document to be analysed. The space of well-formed content representations is then heuristically searched based on the fuzzy similarity measure until good enough matches with the text are found. If several candidates remain at this stage, attempts are made to disambiguate between them using shallow syntactic and semantic clues (as may be provided by a shallow parser). If some decisions cannot be made reliably by the system, a human expert may be asked to disambiguate between candidates. The paradigm reverses the traditional picture on content analysis, which tends to view it primarily as a parsing process, where gradually larger syntactic units (at the level of the sentence, then at the level of the discourse) are built and where semantic interpretation is typically done in a compositional manner on the basis of the syntactic struc-tures found (see for example (Allen 1995)). In that picture, document content emerges so to speak as an object derived from well-formed syntactic constructs, and the central tool is a syntactically-oriented grammar. In our view, on the contrary, the central tool should be a formal specification of what counts as a valid semantic object, and a mechanism