A new technique for joint optimization of excitation and model parameters in parametric speech coders
Khosrow Lashkari, Toshio Miki · The Journal of the Acoustical Society of America · 2001
In a speech coding system, synthesis error (the difference between the original speech at the encoder input and the reproduced speech at the decoder output) is a more relevant measure of signal distortion than linear prediction (LP) error. By minimizing the synthesis error instead of the linear prediction error, the analysis and synthesis stages become more compatible. While LP error is linear in filter parameters, synthesis error is a highly nonlinear function of these parameters, making it computationally intractable for real-time applications. This paper presents a computationally feasible solution for minimizing the synthesis error applicable for joint optimization of the excitation and model parameters in real time. Using a gradient search in the root domain, synthesis error is minimized by re-optimizing the filter parameters for a given excitation. Starting the gradient search from the LPC solution, the resultant synthesis error produced by the proposed technique is guaranteed to be lower than the synthesis error using the LPC filter. By adding an extra minimization step, this technique can be incorporated into any parametric speech coder including LPC, multipulse LPC and CELP speech coders. The new technique and results of application to real speech will be presented.