New source model for narrow-band vocoders
R. Viswanathan, J. Makhoul, A. W. F. Huggins · The Journal of the Acoustical Society of America · 1977
Present-day narrow-band vocoders employ an idealized source (or excitation) model, which is either a sequence of quasiperiodic pulses for voiced sounds, or white noise for unvoiced sounds. This model seems to be largely responsible for the “buzziness” and lack of naturalness perceived in the resulting synthesized speech. We propose a new source model, which combines both pulse and noise sources in a novel way. Based on the observation that spectra of voiced speech sounds (e.g., voiced fricatives and even certain vowels) exhibit devoiced or incoherent high-frequency bands, the model divides the spectrum into a low-frequency region and a high-frequency region, with the pulse source exciting the low region and the noise source exciting the high region. The cutoff frequency that separates the two regions is adaptively varied in accordance with the changing speech signal. We present the advantages of the proposed model over the pulse/noise model, and describe a method for implementing it. Synthesis experiments conducted using the above model with manually extracted cutoff frequency data indicate the power of the model in almost entirely eliminating the “buzzy” quality. [Work sponsored by ARPA-IPTO.]