Improved feature decorrelation for HMM-based speech recognition
Kris Demuynck, Jacques Duchateau, Dirk Van Compernolle, Patrick Wambacq · 1998
In most HMM-based recognition systems, a mixture of diagonal covariance gaussians is used to model the observation density functions in the states. The use of diagonal covariance gaussians however assumes that the underlying data vectors have uncorrelated vector components: if each gaussian is replaced with its full covariant counterpart, the off-diagonal elements in the covariance matrices should be small. To that end, most recognition systems have some kind of decorrelation matrix near the end of the preprocessing. Examples are the inverse cosine transform used with cepstral coefficients, and principal component analysis (PCA) or linear discriminant analysis (LDA) of the features. However, none of these transforms is optimal if it comes to reducing the mismatch introduced by setting the off-diagonal elements in the covariance matrices to zero. The algorithm described in this paper reduces the local correlations between feature vector components inside the gaussians with a single global linear transform at the end of the preprocessing stage. The algorithm is optimal in the sense that we calculate the linear transformation that minimises the sum of the square of all off-diagonal elements over all gaussians. The algorithm is compared with principal component analysis, linear discriminant analysis and the recently published maximum likelihood modelling for semi-tied covariance matrices. The decorrelation method is also evaluated on two speech recognition tasks. A significant relative improvement was achieved in both cases.