NCHLT Speech II Corpus

Jaco Badenhorst, Febe de Wet, Neil Kleynhans, Thipe Isaiah Modipa · 2016

The speech corpus generated from aligned audio samples from National Parliament using Hansard transcriptions are provided in terms of audio and transcriptions. The XML files provide the following metadata for each session: - audio filename - audio orthography - GOP (goodness of pronunciation) score - start time (seconds) - end time (seconds) The audio files are formatted as 16-bit Signed Integer PCM, single channel, and 16kHz sample rate.

Read the paper · More papers on PaperTik