Speech synthesis using subband-coded multiband source components and sinusoids

Nobuyuki Nishizawa, Tsuneo Kato · 2013

An improved speech waveform generation method for speech synthesizers using filter banks is proposed where spectral features of synthetic sounds are constructed by amplitude modification and summation of predecomposed source waveforms. In the method, since all operations are performed in the subband-coded domain with a reduced sampling rate, the computational cost can also be reduced. Moreover, to improve the accuracy of spectral reproduction in low frequency domain of voiced sounds, sinusoidal synthesis directly performed on low subbands is also introduced. The result of a subjective test using resynthesized sounds spoken by a male and female narrator indicated that the proposed method was significantly superior to the conventional methods using a mel log spectrum approximation (MLSA) filter and nonmaximally decimated filter bank, which was our previously proposed method.

Read the paper · More papers on PaperTik