Characterising a database of spoken German by techniques of data mining

Karl Weilhammer, Susanne Burger · 1998

When designing and characterising large speech corpora one frequently faces the problem of how the growth of vocabulary is related to the number of words uttered or written. In the past there have been several attempts to find a formula that describes this relation efficiently. In this article we will try to derive such a formula for a spontaneous speech corpus. The narrow scenario specifications of the examined dialogues instruct two speakers to negotiate a certain task which is almost the same for all the recording sessions (see Table 1) and therefore has a limited vocabulary. This makes the difference to the classical work in this field, which was done on written text like journal or newspaper articles with a broader range of vocabulary. Introduction In the forties of this century scientists started to work on a formula that describes how the growth of vocabulary is related to the number of written words in a text. J. W. Chotlos (1944) worked on essays written by American pupils, ...

Read the paper · More papers on PaperTik