Evaluation of different feature extraction methods for speech recognition in car environment
Martin Wolf, Climent Nadeu · 2008
In this paper the performance of robust feature extraction techniques for speech recognition is evaluated in a car noise environment. Starting from the basic log mel-scaled filter-bank energies, both the application of the minimum variance distortionless response (MVDR), and the decorrelating transformation (either DCT or frequency filtering (FF)) are considered. In this way, five different types of feature extraction techniques were compared, using the Spanish version of the SDC-Aurora database, either with or without dropping frames labeled as silence by a voice activity detector. According to the results, which were obtained after extensive parameter tuning, the MVDR method is capable of improving slightly the results for both cepstral coefficients (CC) and FF parameters in most tested conditions. On the other hand, the FF-based techniques show a significantly better performance under the high-mismatched conditions than the CC ones (more than 33% of relative improvement). The best average accuracies are resulting from the new MVDR-FF combination.