Warped Low-Delay CELP for Wideband Audio Coding
Aki Härmä, Unto Kalervo Laine · 1999
In certain bidirectional and multi-directional real-time audio applications there is a need for low-delay audio coding techniques. A typical encoding/decoding delay of current wideband audio codecs is more than 40 ms, whereas the goal for low-delay coding significantly less than that. Both technical and psychoacoustical requirements and limitations of low-delay coding in bidirectional real-time audio applications are discussed in this paper. The coding delay in codecs based on non-parametric spectral estimation, e.g., subband decomposition or MDCT, result from the buffering of signal frames before spectral analysis. Usually there are also several other sources for the algorithmic delay which have been reviewed in the case of MPEG codecs in [1]. In this paper, it is suggested that the algorithmic coding delay should be 2-5 ms. A frame length of 2 ms corresponds to 88 samples of audio at 44.1 kHz sampling rate. It is probable that sufficiently high definition spectral decomposition cannot be obtained using non-parametric techniques in this short frame. The codec introduced in the paper uses parametric spectral estimation, which is a variant of linear predictive spectral modeling. This algorithm is actually a modification of a lowdelay speech coding algorithm, G.728 Low-Delay CELP [2], which is a widely used standard codec in video conferencing applications. Linear predictive analysis and auditory modeling are performed in a backward adaptive manner, which means that the analysis window lies mainly on the already transmitted part of the signal. The coefficients from the LPC analysis are used in a time-varying synthesis filter. The synthesis filter is driven by an excitation signal which consists of a sequence of excitation vectors which have been selected from a vector codebook using a simplified auditory model. A version of the G.728 codec for wideband speech at sampling rate of 32 kHz has already been proposed [3]. A major modification to the conventional algorithm in the current paper is that the linear predictive analysis and synthesis filter are frequencywarped [4, 5, 6, 7]. This means that the frequency resolution of the spectral estimation is matched with the frequency scale of hearing. This technique makes the linear predictive coding scheme applicable to perceptual wideband audio coding. At this time, it seems plausible that high-quality audio coding at approximately 2 bits/sample and with a coding delay of less than 2 ms is a realistic goal.