Parameter estimation from speech signals for tube models
Karl Schnell, Arild Lacroix · The Journal of the Acoustical Society of America · 1999
Discrete time tube models implemented by lattice structures can be used to simulate the propagation of sound waves in the vocal tract. To cover a wide variety of speech sounds, it is necessary to extend the standard all-pole tube model by coupling the nasal tract, and, furthermore, enable the excitation of the tube system at different places in between. The resulting filter structure H(z) now shows poles as well as zeros in the transfer function. The impedance of the lip opening is realized using a model from Laine [U. K. Laine, ‘‘Modeling of lip radiation impedance in the z-domain,’’ Proc. ICASSP-82, 1992–1995 (1982)], which includes one pole and one zero in the z-plane. The parameters of the tube model will then be estimated from speech by minimization of an error, which describes the spectral distance between H(z) and the prefiltered speech signal. The prefilter is used for the separation of signal components due to excitation and radiation. A gradient-based optimization is implemented to find the minimum. Following this procedure, the frequency response is able to approximate the spectral envelope of the speech signal. The resulting areas correspond to those obtained by x-ray or NMR investigations.