Pitch-synchronous analysis/synthesis using models of speech production

S. Gupta, Juergen Schroeter · The Journal of the Acoustical Society of America · 1990

Articulatory analysis/synthesis results, using a fixed frame length of 20 ms, have previously been reported [e.g., Schroeter, Larar, and Sondhi, J. Acoust. Soc. Am. Suppl. 1 82, S54 (1986)]. There, the Ishizaka/Flanagan self-oscillating two-mass model was employed to generate the glottal excitation. It was found, however, that this model causes the analysis-by-synthesis procedure to make a trade-off between pitch accuracy and match of spectral tilt. In this paper, a pitch-synchronous scheme using a parametric model of the time derivative of the glottal area function is presented. As with the two-mass model, the glottal flow is computed taking into account the current tract input impedance (acoustic source-tract interaction). Vocal-tract optimization, too, is done pitch synchronously. For this purpose, an improved articulatory codebook is searched for an optimal time sequence of start-up shapes. In the codebook search, a modified cepstral distance measure assures a good spectral match, and a properly chosen geometric cost function provides tract shapes that vary smoothly over time.

Read the paper · More papers on PaperTik