The JASMIN Speech Corpus: Recordings of Children, Non-natives and Elderly People
Catia Cucchiarini, Hugo Van hamme · Theory and applications of natural language processing · 2012
Large speech corpora (LSC) constitute an indispensable resource for conducting research in speech processing and for developing real-life speech applications. In 2004 the Spoken Dutch Corpus (Corpus Gesproken Nederlands - CGN) became available, a corpus of standard Dutch as spoken by adult natives in the Netherlands and Flanders. CGN does not include speech of children, non-natives, elderly people and recordings of speech produced in human-machine interactions. Since such recordings would be extremely useful for conducting research and for developing HLT applications for these specific groups of speakers of Dutch, the JASMIN-CGN was started with the aim of extending CGN in three dimensions: age, mother tongue and interaction mode. First, by collecting a corpus of contemporary Dutch as spoken by children of different age groups, non-natives with different mother tongues and elderly people in the Netherlands and Flanders (JASMIN-CGN), we aimed at an extension along the age and mother tongue dimensions. In addition, we collected speech material in a communication setting that was not envisaged in CGN: human-machine interaction. One third of the data was collected in Flanders and two thirds in the Netherlands. The corpus has already been used in different ways within the STEVIN programme. In addition, it turned out to be useful for different lines of research. Since 2008 the JASMIN speech corpus has been available through the Dutch-Flemish HLT Agency. We hope that other researchers will make use of the data and the knowledge gathered in this project for further research and development. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.