Estimation of the parameters of a long-term model for accurate representation of voiced speech

Y. Stettiner, D. Malah, D. Chazan · IEEE International Conference on Acoustics Speech and Signal Processing · 1993

The model is able to describe the slow time-variation (nonstationary) of speech and hence enables the analysis of a whole phoneme in a single frame. This of great importance in the separation of close pitch harmonics common in speech separation problems. It also has potential in speech coding and synthesis applications. The model considered is an extension of the model proposed by L.B. Almeida and J.M. Tribolet (1983). Contrary to their model, which uses a Taylor series approximation and a fixed pitch in the analysis interval, the authors present an efficient iterative algorithm for explicit estimation of the model parameters, including the time-warping function which describes the pitch variation in the analysis frame. Preliminary simulations with voiced speech show that the model has potential in accurately describing whole voiced phonemes, even those several hundred milliseconds in duration, subject to appropriate segmentation.>

Read the paper · More papers on PaperTik