The Spoken Dutch Corpus and its exploitation environment
Nelleke H. J. Oostdijk, Daan Broeder · 2003
The Spoken Dutch Corpus project (1998-2003) is aimed at the development of a corpus of 1,000 hours of speech. The cor-pus is being annotated with various types of transcriptions and annotations and will be distributed together with the speech recordings. In order for users to access the data efficiently and with relative ease, exploitation software is being developed that can handle both sound files and other types of data files. After a brief introduc-tion in which the goals of the project are outlined, the present paper first describes the Spoken Dutch Corpus as it is pres-ently being constructed and then goes on to describe in some detail the exploitation software. The exploitation environment makes it possible to view the data con-tained in the Spoken Dutch Corpus as one item in a network of other similarly struc-tured corpora. Our outlook on such a cor-pus universe can be found in the ‘Future Work ’ section.