Meta-data in the Spoken Dutch Corpus project
Nelleke H. J. Oostdijk · 2000
The Spoken Dutch Corpus that is currently under construction will constitute a 10-million-word corpus of contemporary Dutch as spoken in Flanders and the Netherlands. A collection of extremely varied data for extremely varied users, the Spoken Dutch Corpus constitutes an ideal case study for evaluating proposals for encoding standards. The paper discusses the nature of the meta-data that are deemed to be relevant for the various user groups of the Spoken Dutch Corpus. It also addresses issues such as how- from a users’ point of view- these meta-data should preferably be structured. In addition, the paper evaluates the extent to which available standard proposals are adequate or need to be adapted to suit these needs. 1.