Perceptually Weighted Long Term Modeling of Sinusoidal Speech Amplitude Trajectories
Mohammad Firouzmand, Laurent Girin · 2006
In this paper, the problem of modeling the trajectory of the amplitudes of speech signals is addressed within the context of the sinusoidal model of speech. A long-term model of the trajectory of the amplitude of the partials is proposed for each entire voiced section of speech, contrary to standard models, which are defined on a frame-by-frame basis. The complete analysis-modeling-synthesis process is presented. We compare a DCT-based long-term model with classical (frame-by-frame) interpolation schemes, given that the analysis process is identical in both cases. Perceptual constraints are taken into account since the distortion criterion in this approach is the level of modeling noise above the masking threshold. Promising results are given and the interest of the presented models for speech coding and watermarking applications is discussed.