Syntactic parsing with NooJ

Max D. Silberztein, Max Silberztein · 2009

ABSTRACT. When parsing a text, NooJ’s parsers store all the annotations that they produce in the Text’s Annotation Structure (TAS). At each level of the various linguistic analyses and the corresponding parser, a given parser may add annotations to, or remove annotations from, the TAS. As annotations are attached to larger and larger sequences of texts, the TAS represents the hierarchical structure of the sentence and its syntactic constituents. We have added a new module that processes this structured information in order to display the structural tree of sentences. We discuss the difference between a structural tree and a derivation tree, and we show how NooJ's structural trees can represent any type of linguistic units, including discontinuous ones. 1. Linguistic Units One characteristic of NooJ is that its parsers process several types of linguistic units in texts: prefixes and suffixes (e.g. dis-, –ization), simple words (e.g. table), multiword units (e.g. as a matter of fact) and discontinuous frozen expressions (e.g. to take … into account). 1 All linguistic units recognized by NooJ’s morphological, lexical, syntactic and semantic parsers are represented as annotations, rather than tags. 2 An annotation might represent either an Atomic Linguistic Unit (ALU), i.e. an element of the vocabulary of a language, or any type of sequences of ALUs that constitutes a meaningful syntactic or semantic unit, such as a noun phrase (e.g. the head of the company), a verbal group (e.g. may not have wanted to read) or an adverbial complement (e.g. Monday February the 11th at 2PM), etc. When parsing a text, NooJ’s parsers store all the annotations that they produce in the

Read the paper · More papers on PaperTik