Jasmin Speech Corpus
Catia Cucchiarini, Hugo Van hamme, F. Smits · OLAC - Open Language Archives Community · 2015
The JASMIN Speech Corpus contains about 115 hours of speech from children, non-natives and senior people. Approx. 50% of the material is read speech and 50% comes from human-machine interaction. The entire corpus was orthographically transcribed and all words in the corpus were automatically provided with a lemma, a POS-tag and a phonetic transcription. Organisations involved in the building of the JASMIN Speech Corpus: CLST, Radboud University; ESAT, KU Leuven and TalkingHome.