Towards a corpus-based dictionary of German noun-verb collocations
Ulrich Heid · 1998
We 1 describe our attempts to automatically extract raw material for a dictionary of German noun-verb collocations from large corpora of newspaper text. Such a dictionary should be about collocations and it should include a description of their linguistic properties, rather than listing the mere lexical cooccurrence. Since most statistical collocation finding tools do not provide other than lexical cooccurrence information, we first use symbolic extraction tools, based on a regular grammar over part-of-speech tagged and lemmatized text, and we use statistical filters thereafter. We first list the types of information which should be contained in a collocational dictionary for Natural Language Processing, then sketch our extraction methods and finally discuss and illustrate our initial results. Keywords: Collocations, text corpora, semi-automatic lexical acquisition. 1 Introduction: Motivation and objectives Other than for English (and, to some extent, for French 2 ), th...