A Hybrid Scheme for Perceptual Audio Coding
Kris Hermus, Werner Verhelst, Patrick Wambacq · 2002
Most parametric audio coders use a traditional sinusoidal model to represent the tonal parts of audio signals, together with dedicated models for the noise and transient-like parts of the audio signal. In this paper we apply Total Least Squares (TLS) algorithms to automatically extract the modeling parameters of a Exponential Sinusoidal Model (ESM) for audio signals. This sum of exponentially damped sinusoids is capable of accurately modeling most transient segments that are readily found in audio signals. In order to turn the SNR optimization criterion of these TLS algorithms into a perceptual modeling strategy we incorporate the psychoacoustic model of MPEG 1- Layer 1 into a subband TLS-ESM scheme. This allows us to model each subband in accordance with its perceptual relevance. Informal listening tests confirm that perceptual ESM achieves the same perceived quality as plain ESM while using substantially less components. The ESM model can serve as the parametric part of a hybrid audio coder in combination with traditional waveform coding techniques for the ESM residual signals. 1.