Spectral distortion and quality of synthesized speech in cepstral speech analysis‐synthesis system
Tadashi Kitamura, Satoshi Imai · Electronics and Communications in Japan (Part I Communications) · 1982
Abstract The relation between spectral distortion and the quality of synthesized speech in a speech analysis‐synthesis system based on the cepstrum method (cepstral vocoder) is described. In this system, the true logarithmic spectral envelope is estimated by an improved cepstral method in the analysis part and a logarithmic amplitude characteristic approximated filter is used in the synthesis part. The transmission rate for spectral information is reduced using the differential of the cepstrum due to the differential of the spectral envelope, because the spectra do not change very rapidly. The preference score by pair comparison tests is employed as a subjective evaluation and spectral distortion is used as an objective evaluation to establish the relations among the quantization width, word length, frame rate, cepstrum order, spectral distortion and synthesized speech quality. Furthermore, the factors of spectral distortion and its characteristics are clarified and it is shown that spectral distortion can be estimated from the transmission condition. The result is that 2.8‐kbit/s, high‐quality synthesized speech can be obtained by this synthesis system.