Context-dependent Phone Mapping for Acoustic Modeling of Under-resourced Languages.

Van Hai, Xiong Xiao, Eng Siong Chng, Haizhou Li · 2015

This paper presents the use of phone mapping for acoustic modeling of a language with limited training data. In this approach, we use well-trained acoustic models of a source language to generate acoustic scores for each feature vector of the target language. These scores are then mapped to the posteriors of context-dependent triphones of the target language using a limited amount of training data. In this paper, English is used as the source language while Malay is used as the target language. Experiments on a Malay large vocabulary continuous speech recognition (LVCSR) task show that with only a few minutes of training data we can achieve a low word error rate which significantly outperforms the best monolingual baseline acoustic model directly trained on the target language data. In addition, our study indicates that a consistent improvement is obtained when source acoustic scores are combined with speech attribute or bi-speech attribute posterior probabilities generated by the source attribute detectors to form the input for phone mapping.

Read the paper · More papers on PaperTik