Development of a large-scale Mandarin Radio Speech Corpus

Yung-hsiang Shawn Chang, Yuan‐Fu Liao, Sheng-Ming Wang, Jenq‐Haur Wang, Sing-Yue Wang, Jhih-wei Chen, You-dian Chen · 2017

The Taiwan Mandarin Radio Speech Corpus consists of roughly 300 (and growing) hours of audio recordings, selected from Taiwan's National Education Radio (NER) archive. The corpus includes speech from hundreds of speakers and various speech styles (spontaneous conversational and read news). This corpus provides a rich resource for research in speech and automatic speech recognition (ASR). In this paper, we briefly introduce the corpus development approach and report two preliminary experimental results using this corpus.

Read the paper · More papers on PaperTik