Hallucinated n-best lists for discriminative language modeling

Kenji Sagae, Maider Lehr, Emily Tucker Prud’hommeaux, Peng Xu, Nathan Glenn, Damianos Karakos, Sanjeev P. Khudanpur, Brian Roark, Murat Saraçlar, Izhak Shafran, Daniel M. Bikel, Chris Callison-Burch, Yuan Cao, Keith Hall, Eva Hasler, Phillipp Koehn, Antonio Manuel López, Matt Post, Darcey Riley · 2012

This paper investigates semi-supervised methods for discriminative language modeling, whereby n-best lists are “hallucinated” for given reference text and are then used for training n-gram language models using the perceptron algorithm. We perform controlled experiments on a very strong baseline English CTS system, comparing three methods for simulating ASR output, and compare the results with training with “real” n-best list output from the baseline recognizer. We find that methods based on extracting phrasal cohorts - similar to methods from machine translation for extracting phrase tables - yielded the largest gains of our three methods, achieving over half of the WER reduction of the fully supervised methods.

Read the paper · More papers on PaperTik