Speech recognition using regularized minimum variance distortionless response spectrum estimation-based cepstral features

Jahangir Alam, Patrick J Kenny, Douglas D. O’Shaughnessy · 2013

This paper presents regularized minimum variance distortion-less response (MVDR)-based cepstral features for robust continuous speech recognition. The mel-frequency cepstral coefficient (MFCC) features, widely used in speech recognition tasks, are usually computed from a direct spectrum estimate, that is, the squared magnitude of the discrete Fourier transform (DFT) of speech frames. Direct spectrum estimation methods (also known as nonparametric estimators) perform poorly under noisy and adverse conditions. To reduce this performance drop we propose to increase robustness of the speech recognition system by extracting more robust features based on the regularized MVDR technique. The proposed method, when evaluated on the AURORA-4 speech recognition task, provides an average relative improvement in word accuracy of 11.3%, 6.1%, and 5.2% over the conventional MFCC, PLP, MVDR and PMVDR-based MFCC features, respectively.

Read the paper · More papers on PaperTik