Speaker adaptive phoneme recognition based on feature mapping from spectral domain to probabilistic domain

Tetsuo Kobayashi, Y. Uchiyama, Jun Osada, K. Shirai · 1992

A feature parameter space for speech recognition called PRPG (probability ratios between phoneme group pairs) is described, and speaker adaptive phoneme recognition is performed. In the coordinate system proposed, the area with the same information for speech recognition is compressed into one point. The mapping function from spectral coordinate system to the proposed one is realized using a neural network. The code-vectors designed on this coordinate system are guaranteed to be information-theoretically more efficient than that of spectral coordinate system. Moreover, by the definition of the coordinate system, the meaning of axes is equivalent among different speakers, so speaker adaptation can be easily performed without trajectory mapping. Experimental results show that errors are reduced by 40% by coordinate conversion in speaker-dependent tasks. The scores of speaker-adaptive tasks in the proposed feature domain are always superior to those of the speaker-dependent tasks in the spectral domain.>

Read the paper · More papers on PaperTik