Large Corpus of Czech Parliament Plenary Hearings

Jonáš Kratochvíl, Peter Polák, Ondřej Bojar · Americanae (AECID Library) · 2019

We present a large corpus of Czech parliament plenary sessions. The corpus consists of approximately 444 hours of speech data and corresponding text transcriptions. The whole corpus has been segmented to short audio snippets making it suitable for both training and evaluation of automatic speech recognition (ASR) systems. The source language of the corpus is Czech, which makes it a valuable resource for future research as only a few public datasets are available for the Czech language.

Read the paper · More papers on PaperTik