Rapid speaker adaptation for embedded large vocabulary dictation system with sparse training materials
Wei Huang, Yaxin Zhang, Xin He, Qingfeng Bao · 2008
In this paper, a novel tree-structural maximum a posteriori mapping (SMAP) algorithm is proposed for embedded large vocabulary speech dictation system. A two level tree is created by classing the Gaussian mixtures of a large HMM set into several classes, and an adaptation table is created for each class by MAP observed Gaussian mixture with adaptation data. Based on the adaptation table, the other unobserved Gaussian mixtures are rapidly adapted with negligible computation cost. The experiment results show that the SMAP is better than the conventional MAP estimation even with much less adaptation materials. We have implemented this algorithm on a cellular phone. For 40 short adaptation utterances (about 200 words), the total computation time for adaptation is reduced from about 10 minutes to less than 8 seconds.