Fast Implementation of KLT-Based Speech

Yoshifumi Nagata, Kenji Mitsubori, Takahiko Kagi, Toyota Fujioka, Masato Abé · 2006

We propose a new method for implementing Karhunen-Loeve transform (KLT)-based speech enhancement to exploit vector quantization (VQ). The method is suitable for real-time processing. The proposed method consists of a VQ learning stage and a filtering stage. In the VQ learning stage, the autocorrelation vectors comprising the first elements of the autocorrelation function are extracted from learning data. The au- tocorrelation vectors are used as codewords in the VQ codebook. Next, the KLT bases that correspond to all the codeword vectors are estimated through eigendecomposition (ED) of the empirical Toeplitz covariance matrices constructed from the codeword vectors. In the filtering stage, the autocorrelation vectors that are estimated from the input signal are compared to the codewords. The nearest one is chosen in each frame. The precomputed KLT bases corresponding to the chosen codeword are used for filtering instead of performing ED, which is computationally intensive. Speech quality evaluation using objective measures shows that the proposed method is comparable to a conventional KLT-based method that performs ED in the filtering process. Results of sub- jective tests also support this result. In addition, processing time is reduced to about 1/66 that of the conventional method in the case where a frame length of 120 points is used. This complexity reduction is attained after the computational cost in the learning stage and a corresponding increase in the associated memory requirement. Nevertheless, these results demonstrate that the proposed method reduces computational complexity while main- taining the speech quality of the KLT-based speech enhancement. Regarding the colored-noise problem, Ephraim's method is extended and an explicit solution has been given (3). This method utilizes a noise whitening approach, as used in its original method. The present paper includes the assertion that instability in computing the inverse matrix, which is required for whitening, can be avoided through modification of the autocorrelation matrix. However, the computational cost for obtaining the inverse matrix and the filtering to accomplish both whitening and its inverse cannot be disregarded. In contrast, Mittal and Phamdo (4) and Rezayee and Gazor (5) proposed methods that do not require noise-whitening. In those methods, the spectral component power along each KLT axis for the current frame is estimated from a noise signal that has been preserved from a past noise period. This processing corresponds to noise spectrum estimation that is commonly performed in the DFT-based SS, whereas the noise spectrum should be calculated in every speech frame because KLT-based SS uses

Read the paper · More papers on PaperTik