Improving LVCSR system combination using neural network language model cross adaptation

Xiaobing Liu, Mark Gales, Philip C. Woodland · 2011

State-of-the-art large vocabulary continuous speech recognition (LVCSR) systems often combine outputs from multiple sub-systems developed at different sites. Cross system adaptation can be used as an alternative to direct hypothesis level combina-tion schemes such as ROVER. The standard approach involves only cross adapting acoustic models. To fully exploit the com-plimentary features among sub-systems, language model (LM) cross adaptation techniques can be used. Previous research on multi-level n-gram LM cross adaptation is extended to further include the cross adaptation of neural network LMs in this pa-per. Using this improved LM cross adaptation framework, sig-nificant error rate gains of 4.0%-7.1 % relative were obtained over acoustic model only cross adaptation when combining a range of Chinese LVCSR sub-systems used in the 2010 and 2011 DARPA GALE evaluations. 1.

Read the paper · More papers on PaperTik