MORPA: A morpheme lexicon based morphological parser

Josée S. Heemskerk, Vincent J. van Heuven · 1993

MORPA is a MORphological PArser developed at Leiden University for use m the text tospeech conversion System for Dutch, SPRAAKMAKER MORPA operates in three successive stages First, it generates all possible segmentations of an mput word mto stnngs of stems and affixes Secondly, it tests each segmentation for morpho syntactic well-formedness while determmmg word class Fmally, all remainmg analyses are ordered, with the most likely analysis in topmost positionIn this paper we shall outline the architecture of MORPA, which compnses a dictionary of 17,087 entnes, a Categonal Grammar based parser, a module for level-ordered attachment of stems and affixes and a module for likehhood determmation The major problem that our System faces is ambiguity, i e, the generation of alternative segmentations and word class assignments for one mput word, many of which are ungrammatical or implausible We shall discuss three kmds of strategy that have been combmed m MORPA to deal with ambiguity Firstly, the System filtere out ungrammatical segmentations by means of linguistic knowledge Secondly, the System Orders the remainmg analyses by means of frequency Information And fmally, morphological Information that is irrelevant for pronunciation determmation is elimmated In conclusion, we shall present illustrative performance data obtamed from an evaluation run, mvolvmg a 3,077 word lest corpus

Read the paper · More papers on PaperTik