Pronunciation modelling for conversational speech recognition: a status report from WS97

Bill Byrne, Michael Finke, Sanjeev P. Khudanpur, J. McDonough, Harriet J. Nock, Michael Riley, Murat Saraçlar, Chuck Wooters, George Zavaliagkos · 2002

Accurately modelling of pronunciation variability in conversational speech is an important component for automatic speech recognition. We describe some of the projects undertaken in this direction at WS97 [the Fifth LVCSR (large-vocabulary conversational speech recognition) Summer Workshop], held at Johns Hopkins University, Baltimore, in July-August 1997. We first illustrate a use of hand-labelled phonetic transcriptions of a portion of the Switchboard corpus, in conjunction with statistical techniques, to learn alternatives to canonical pronunciations of words. We then describe the use of these alternative pronunciations in a recognition experiment as well as in the acoustic training of an automatic speech recognition system. Our results show a reduction of the word error rate in both cases-0.9% without acoustic retraining and 2.2% with acoustic retraining.

Read the paper · More papers on PaperTik