Auditory VOCODER: Speech resynthesis from an auditory Mellin representation

Toshio Irino, Roy D. Patterson, Hideki Kawahara · IEEE International Conference on Acoustics Speech and Signal Processing · 2002

We assume that speech rnorphing, noise suppression, and speech segregation would improve if they were more accurately based on human perception. Accordingly, an Auditory VOCODER was developed to resynthesize speech from an auditory Mellin representation used to explain human perception. The Auditory VOCODER has three modules: an Auditory Mellin Image model [9,10], a STRAIGHT VOCODER [2], and a mapping module consisting of warped-frequency cepstral analysis and nonlinear, multivariate regression analysis (MRA). We describe the modules and an evaluation of the system. Informal listening indicates that the sound quality is reasonable.

Read the paper · More papers on PaperTik