Multilingual data selection for training stacked bottleneck features
Ekapol Chuangsuwanich, Yu Zhang, James Glass · 2016
Deep Neural Networks (DNNs) trained on multilingual data have proven useful for improving speech recognition in languages with limited resources. In this framework, data from rich resource languages are pooled together to train a single system and then adapted to a new language. However, data from a rich language that are similar to the target language are generally more helpful. We explore methods of training bottleneck features by using data that are more similar to the target language. Our experiments on speech recognition and keyword spotting tasks with IARPA-Babel languages show that our proposed methods outperform typical multilingual DNNs.