Hybrid representation for audio eects
David Cournapeau · 2003
Frameworks to decompose signals into different parts for analysis/transformation/synthesis or coding are quite old in the audio DSP field. The idea behind all these frameworks is that some ”basis sets” (here, bases should not necessarilly be taken in the strict mathematical sense) are well suited for some kind of signals, whereas other bases are better for other types of signal. The phase vocoder, introduced by Portnoff in the late seventies, was the first real tool for music DSP, and implicitly assumes that the signal is a sum of sinusoids in each frames. This assumption was the basis for the work of X. Serra in his famous thesis about sine + noise modeling (see [Ser97] for an introduction). This model, also known under the acronym SMS (for Spectral Modeling Synthesis), was the first parametric model for large classes of audio signals 1 It is well known that additive synthesis with sinusoidal basis vectors can approximate the signal as well as we want; the drawback is well known, too: for some kinds of signals, like ”Noise” or ”transients”, we need a lot of basis vectors to approximate them, and it is quite computive intensive. That’s why more recent works, based on Xerra’s ideas, try to decompose the signal in at least 3 differents parts, ie the ”tonal” part, the ”transients” part, and the residual (also called noise, because it is generally coded as a stochastic process). These kinds of models has been successfully used for analysis/synthesis for years in academic areas and in commercial products 2. Laurent Daudet developed [Dau00]a parametric model for audio signals, essentially aiming at compression. This model splits the audio signal into three parts, tonal / transients / residual. But instead of using a sinusoidal model, it uses thresholding on the MDCT to extract the tonal part; in a very similar way, transients are extracted by thresholding on wavelet coefficients of the non-tonal part. Hard thresholding is basically a non-linear approximation3, and is famous since its successful use in image compression and signal denoising[Mal98]. The idea of my internship is to adapt this model to analysis/synthesis, and to use it in some high-level transforms like time-scaling, and more complex schemes like tempo correction (ie changing the time location of some notes to adapt it to a fixed time-grid). At first, we will present the basic framework: the principles, the results on some test samples and the drawbacks for analysis/synthesis purposes, the sinusoidal model was also studied a few years before for speech coding by McAuley and Quatieri [MAR86] See for example the realizer from PPG, figure 1 or more recently the neuron synthetiser project, figure 2 See Annexe 1 Figure 1: The realizer, the first virtual synthetiser? Figure 2: The Neuron synthetiser; may use some resynthesis method mainly on the tonal extraction. The second chapter will present the work I did to try to avoid some of the main problems, like window switching, wavelet filtering and MDCT regularisation. The third chapter is a presentation of the further ideas one can develop to go around the tonal extraction problems, mainly about a ”complex MDCT” which gives phase informations, and so phase can be used for steady state compenent exctraction. Finally, the last chapter will give some details about some implementations I did in matlab/C++ for my work.