Unsupervised lattice-based acoustic model adaptation for speaker-dependent conversational telephone speech transcription
K. Thambiratnam, Frank Torsten Bernd Seide · 2009
This paper examines the application of lattice adaptation tech-niques to speaker-dependent models for the purpose of conver-sational telephone speech transcription. Given sufficient train-ing data per speaker, it is feasible to build adapted speaker-dependent models using lattice MLLR and lattice MAP. Experi-ments on iterative and cascaded adaptation are presented. Addi-tionally various strategies for thresholding frame posteriors are investigated, and it is shown that accumulating statistics from the local best-confidence path is sufficient to achieve optimal adaptation. Overall, an iterative cascaded lattice system was able to reduce WER by 7.0 % abs., which was a 0.8 % abs. gain over transcript-based adaptation. Lattice adaptation reduced the unsupervised/supervised adaptation gap from 2.5 % to 1.7%.