A hierarchical Bayesian approach for semi-supervised discriminative language modeling

Yik-Cheung Tam, Paul Vozila · 2012

Discriminative language modeling provides a mechanism for differentiating between competing word hypotheses, which are usually ignored in traditional maximum likelihood estimation of N-gram language models. Discriminative language modeling usually requires manual transcription which can be costly and slow to obtain. On the other hand, there are vast amount of untranscribed speech data on which offline adaptation technique can be applied to generate pseudo-truth transcription as an approximation to manual transcription. Viewing manual and pseudo-truth transcriptions as two domains, we perform domain adaptation on the discriminative language models via hierarchical Bayesian, in which the domainspecific models share a common prior model. Domainspecific and prior models are then estimated jointly using training data. On the N-best list rescoring experiment, hierarchical Bayesian has yielded better recognition performance than the model trained only on manual transcription, and is robust against inferior prior.

Read the paper · More papers on PaperTik