A composite sinusoidal model applied to spectral analysis of speech
Shigeki Sagayama, Fumitada Itakura · Electronics and Communications in Japan (Part I Communications) · 1981
Abstract A new technique of speech analysis based on a composite sinusoidal model is proposed. In this model, parameters (frequency and intensity of each sinusoid) are determined such that the lower‐order autocorrelation functions of the finite number for the reference signal (a sum of several sinusoidal waves) is equal to those for the observed speech signal. The model is classified into four types: (1) the frequencies of the sinusoids are free; (2) one frequency is fixed to 0; (3) one frequency is fixed to π; and (4) two frequencies are fixed to 0 and π. For all of these, the parameters are solved easily by the basic algebraic method. While the problem is essentially related to the theory of orthogonal polynomials, it is proven that the roots exist if only the autocorrelation function is a positive definite matrix. The efficient algorithm is introduced. Itakura's theory of line spectral representation of linear predictive coefficients can be considered as the conversion theory from the LPC to the composite sinusoidal model. Some experiments of speech analysis are demonstrated. Applications to the speech analysis and synthesis, synthesis by rule, speech recognition, speaker identification, etc., are suggested.