Pitch determination and sinusoidal modeling for time-varying voiced speech
Masashi Ito, Masafumi Yano · The Journal of the Acoustical Society of America · 2006
A voiced speech signal consists of sinusoidal components of which amplitude and frequency are time varying. Usually the signals are analyzed by assuming that they are stationary within a local analysis window, so they include inevitable errors. To solve this problem, we have proposed a method termed local vector transform (LVT), in which amplitude and phase of the sinusoidal components were approximated by quadratic functions uniquely determined from input spectrum. Two types of experiments were carried out to evaluate the validity of this method. First, time-varying pitch frequencies (F0s) of natural utterances were investigated. F0 determined by LVT greatly improved the accuracy compared to the conventional pitch determination algorithms: cepstrum, autocorrelation, and instantaneous frequency estimation. Second, LVT is applied to the sinusoidal modeling and the amplitude and phase were estimated for every component. The results apparently showed that the signal obtained by LVT was very close to the input. This was quantified by a signal to residual power ratio (S/R). For both of synthesized and naturally uttered speech signals, LVT showed the higher S/R compared to the conventional algorithm. These results indicate that the proposed algorithm is highly effective for pitch determination and sinusoidal modeling for time-varying speech signals.