Generation of word profiles for large German corpora

Alexander Geyken, Jörg Didakowski, Alexander Siebert · Tokyo University of Foreign Studies · 2009

This paper presents the DWDS word profile system, a software-tool that extracts statistically salient co-occurrences from corpora and clusters them according to their syntactic categories. Due to the difficulties of German, in particular its free word order and long distance dependencies, shallow approaches like phrase chunking are not sufficient for a satisfactory extraction of syntactic relation. Our system uses a syntax parser based entirely on weighted finite state transducers which combines satisfactory extraction of syntactic relations with good performance. Currently, we have built a prototype for two corpora of 160 m tokens (resp. 90 m tokens) from which 68 (resp. 37) million word-pair tokens and 1.26 million (resp. 0.8 million) types have been extracted. Statistical salience has been calculated for all types. For both corpora, a prototype containing all word-pairs with a frequency greater than 10 is accessible on the Internet under http://odo.dwds.de/wortprofil. We will integrate the word profile as an additional information source for the DWDS web-platform, a widely used word information platform for German.

Read the paper · More papers on PaperTik