Combined speech and audio coding using non-linear adaptations
Cheung-Fat Chan · 2002
A combined speech and audio coder is proposed. The coder structure resembles a low-delay CELP coder, however, the excitation gain is adapted non-linearly in a sample-by-sample fashion by using a trained neural network, and the spectral parameters are derived from backward non-linear prediction based on a second-order Volterra filter. A perceptual weighting filter derived from psychoacoustic analysis in the spectral domain is used to shape the coding noise. The proposed non-linear adaptation schemes significantly improve the effectiveness of using an analysis-by-synthesis model for coding audio signal. Simulation results show that transparent coding of wideband (7 kHz) speech and audio at 24 kbps is achieved.