Analysis of non-stationary, multi-component signals with applications to speech
C.S. Ramalingam · 1995
In this thesis we address the problem of decomposing a signal containing multiple non-stationary sinusoids into its components. A non-stationary sinusoid is one whose envelope and frequency are functions of time; we assume that these variations are slow enough so that the components occupy distinct regions in the spectrum at each instant of time. We propose to separate the components by using moving bandpass filters. Each moving bandpass filter is implemented as a cascade of demodulator, low-pass filter, and remodulator. We use the instantaneous frequency of the assigned component for tracking. We point out that while the raw instantaneous frequency may be potentially misleading, its smoothed counterpart is physically very meaningful and can be reliably used for moving the filters. The burden on a filter to isolate its component is eased by subtracting from its input an estimate of other signal components. This is based on the principle of Residual Signal Analysis, originally proposed by Costas. Our procedure results in a cleaner separation of the components when compared with his implementation. We have also proposed a feedforward structure to carry out Residual Signal Analysis. Using these methods for voiced speech analysis we have found that the components need not necessarily be precisely harmonic but only nominally so, and that there can be brief periods of time when some of them can become significantly mistuned; the envelopes of some of the components displayed significant amplitude modulation. Teager has raised a number of objections against the popularly used source-filter model and has claimed that it is completely wrong. Our simulation results indicate that this model reproduces quite accurately the frequency mistuning seen in the examples of natural speech we have looked at. We also demonstrate that the polynomial envelope and phase model for non-stationary signal analysis is sensitive to model mismatch and does not automatically yield the increased accuracy expected of it. The recently proposed DESA-1a algorithm of Maragos, et al. is shown to be equivalent to Prony's method for the noiseless sinusoid case and an ad hoc variant of it in the presence of noise.