Effective Training End-to-End ASR systems for Low-resource Lhasa Dialect of Tibetan Language

Lixin Pan, Sheng Li, Longbiao Wang, Jianwu Dang · 2019

The Lhasa dialect is the most important Tibetan dialect and has the largest number of speakers in Tibet and massive written scripts in the long history. Studying how to apply speech recognition techniques to Lhasa dialect has special meaning for preserving Tibet's unique linguistic diversity. Previous research on Tibetan speech recognition focussed on selecting phone-level acoustic modeling units and incorporating tonal information but paid less attention to the problem of limited data. In this paper, we focus on training End-to-End ASR systems for Lhasa dialect using transformer-based models. To solve the low-resource data problem, we investigate effective initialization strategies and introduce highly compressed and reliable sub-character units for acoustic modeling which have never been used before. We jointly training the transformer-based End-to-End acoustic model with two different acoustic unit sets and introduce an error-correction dictionary to further improve the system performance. Experiments show our proposed method can effectively modeling low-resource Lhasa dialect compared to DNN-HMM baseline systems.

Read the paper · More papers on PaperTik