Improving Rapid Unsupervised Speaker Adaptation Based On Hmm Sufficient Statistics

Randy Gómez, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano · 2006

In real-time speech recognition applications, there is a need to implement a fast and reliable adaptation algorithm. We propose a method to reduce adaptation time of the unsupervised speaker adaptation based on HMM-sufficient statistics. We use only a single arbitrary utterance without transcriptions in selecting the N-best speakers' sufficient statistics created offline to provide data for adaptation to a target speaker. Further reduction of N-best implies a reduction in adaptation time. However, it degrades recognition performance due to insufficiency of data needed to robustly adapt the model. Linear interpolation of the global HMM-sufficient statistics offsets this negative effect and achieves a 50% reduction in adaptation time without compromising the recognition performance. We have reduced the adaptation time from 10 sec to 5 sec without degradation of the word accuracy. Furthermore, we compared our method with vocal tract length normalization (VTLN), maximum a posteriori (MAP) and maximum likelihood linear regression (MLLR). Moreover, we tested in office, car, crowd and booth noise environments in 10 dB, 15 dB, 20 dB and 25 dB SNRs

Read the paper · More papers on PaperTik