Parsing Swedish

Atro Voutilainen · DSpace repository (University of Tartu) · 2001

This paper presents two new systems for analysing Swedish texts: a light parser and a functional dependency grammar parser.Their design follows two Helsinki-based frameworks: Constraint Grammar CG (Karlsson &al 1995) and Functional Dependency Grammar FDG (Tapanainen andJärvinen 1997). CG and FDGCG is a reductionistic constraint rule formalism whose input is lexically analysed ambiguous text and whose output is disambiguated text.Disambiguation is carried out by constraints on lemmas and tags that discard alternative analyses on the basis of contextual information, typically coded by a linguist.The ENGCG morphosyntactic tagger was introduced in 1992 (Voutilainen & al.) and compared with a state-of-the-art statistical tagger in 1997 (Samuelsson & Voutilainen).CG was successful in word-class tagging but not adequate for full-scale parsing.A considerable effort on finite-state parsing was made by Koskenniemi, Tapanainen and Voutilainen (see their articles in Roche & Schabes, eds., 1997).A more successful effort was made by Tapanainen and Järvinen, who extended CG into a functional dependency grammar formalism and interpreter/compiler capable of introducing explicit functional dependencies and of applying large grammars efficiently. Earlier work on Swedish tagging and parsingAs discussed in Voutilainen (forthcoming 2001), most efforts at Swedish tagging and parsing have focused on wordclass tagging, mostly in the statistical paradigm.A somewhat more informative analysis is given by Lingsoft's SWECG (morphology + function tags) and shallow finite-state Abney-style parsers (Kokkinakis & Johansson 1999).The Swedish Core Language Engine (Gambäck 1997) produces full syntactic parses, but, as argued by Gambäck, it seems to work only for texts from very restricted domains. Swedish Light SyntaxIn design, SweLite follows Conexor's EngLite (see demo at www.conexor.fi).The first major component is the morphological analyser, based on a recent extension of Koskenniemi's Two-Level formalism.The analyser contains a large lexicon, morphology and guesser for unknown words.The morphological analyser produces analyses, many of them ambiguous.The parser uses mapping statements to introduce light syntactic ambiguity, so before any disambiguation is done, an ambiguous analysis looks like this: " " "tvinga" V INF &MV "tvinga" V PRES &MV "tving#as" N SG/PL NOM &>N &NH "tv#in|gas" N SG NOM &>N &NH Online Proceedings of NODALIDA 2001

Read the paper · More papers on PaperTik