Cross-language Acoustic Model Refinement for the Indonesian Language
T. Martin, S. Sridharan · 2006
Porting ASR capabilities to many languages is hindered by a lack of transcribed acoustic data. Cross-language adaptation techniques seek to address this problem by substituting models trained in resource-rich source languages to recognise speech in resource-poor target languages. The differences in coarticulatory effects between the source and target languages, together with unwanted pronunciation and channel variation, result in recognition rates that are typically much worse then those achieved by well trained monolingual systems. We present a technique which makes more effective use of limited adaptation data by structuring the state distributions to suit the coarticulatory occurrences in the target language. Additionally, the proposed technique provides a more suitable method for synthesising unseen contexts. Evaluation of this technique is presented for a word recognition task using English and Spanish source language acoustic models trained using Switchboard and CallHome databases, respectively. Using 25 minutes of Indonesian speech for target language adaptation data, this technique achieved absolute improvements of 3.69% and 6.31% for English and Spanish sources, respectively, when compared to traditional adaptation techniques. Using 90 minutes of adaptation data, absolute improvements of 3.22% and 3.07% were achieved.