Formosa Speech Recognition Challenge 2018: Data, Plan and Baselines
Yuan‐Fu Liao, Wu-Hua Hsu, Yu-Chen Lin, Yung-hsiang Shawn Chang, Matúš Pleva, Jozef Juhár, Guang-Feng Deng · 2018
This paper introduces the Formosa speech recognition (FSR) challenge 2018, presents the provided data profile, evaluation plan and reports the experimental results of the baseline systems. This challenge focuses on spontaneous Taiwanese Mandarin speech recognition (TMSR) and it is based on a real-life, multigene broadcast radio speech corpus, NER-Trs-Vol1, selected from the Formosa speech in the wild (FSW) project. To assist participants to establish a good starting system, a set of baseline systems were published based on various deep neural network (DNN) models. NER-Trs-Vol1 is free for participants (noncommercial license), and its corresponding Kaldi recipes for the baselines have been published online. Experimental results show that the combination of NER-Trs-Vol1 and Kaldi recipes is a good resource pack for spontaneous TMSR research and could be used to initialize an advanced semi-supervised training procedure to further improve the recognition performance.