Speech modelling and the trigonometric moment problem
P. Delsarte, Yves V. Genin, Yves Kamp, Paul Van Dooren · Digital Access to Libraries (Université catholique de Louvain (UCL), l'Université de Namur (UNamur) and the Université Saint-Louis (USL-B)) · 1982
It is shown that the partial trigonometrie moment problem provides an appropriate unifying framework for some speech modelling techniques like the line spectral pairs and composite sinusoidal wave model recently proposed by Itakura et al., the eigenmodel of Pisarenko used for formant extraction and the classical autoregressive model on which LPC is based. This moment problem is equivalent to an extension problem in the class of impedance functions and hence has a simple circuit theoretical interpretation. The connection with the classical power moment problem is also established. 1. Introduetion Among the many techniques existing today for the modelling of discrete time signals by rational functions, the linear prediction or autoregressive (AR) model is undoubtedly one of the most powerful, especially in the field of speech processing 1). By its very nature, this model provides a simple speech synthesis technique as well as an analysis tool for the estimation of the formants or resonant frequencies of the vocal tract. The popularity of the AR .I method can be justified by many arguments but among these certainly emerges the fact that efficient computational procedures for the derivation of the model are available as well as robust structures for its implementation. In particular, the reflection or partial correlation (PARCOR) coefficients have become a classical characterization ofthe AR model). A few years ago, Itakura and Sugamura 3) showed however that an AR model could equivalently be described in terms of the so-called line spectral pairs leading to a characterization which in some respects is more efficient than the classical PARCOR method. At about the same time, Sagayama and Itakura 4) also proposed a new technique for speech synthesis, called the composite sinusoidal wave model, in which the time signal is represented as a sum of several sinusoidal waves. The amplitudes and frequencies of these sinusoids are adjusted so as to match a limited number of the signal autocorrelation lags. Gueguen 6) for his part introduced still another speech model based on the eigenvector associated