Choosing the right lemma when analysing German nouns
Martin Volk · Zurich Open Repository and Archive (University of Zurich) · 1999
Introduction When processing large corpora, it is often necessary to lemmatise the wordforms. This is usually done by a morphological analyser which can, in any case, undo inflection but sometimes even derivation and compounding. The latter is especially useful for German which exhibits very productive compounding. But when using such a system we notice that lemmatisation is a frequent source of ambiguities. Some wordforms genuinely belong to two lemmas of the same part-of-speech such as rasten which can be a form of rasen (`to race') or rasten (`to rest'). Others belong to two lemmas of different word classes such as meinen, which can represent various forms of either the first person possessive pronoun (`my') or the verb `to mean'. This latter ambiguity can easily be resolved by a part-of-speech tagger or a parser. The former ambiguity is much harder to deal with. In the case of verbs a parser might be able to distinguish betw