A Comparison of Several Approaches to the Feature Extractor Design for ASR Tasks in Telephone Environment

Ascensión Gallardo-Antolín, Javier Macías-Guarasa, Javier Ferreiros, Ricardo de Córdoba, J. M. Montero-Martínez, José Manuel Pardo · 2003

Automatic speech recognition (ASR) systems are usually composed of a parameterization module and a back-end classifier. The performance of the overall system strongly depends on the choice of the feature extraction module. In this paper we investigate two different approaches for designing this module. In the first one (the conventional approach), its main characteristics are chosen based on psychoacoustic knowledge. In the second one, a datadriven technique (“Discriminative Feature Extraction”-DFE-), which performs a simultaneous optimization of the feature extractor and the back-classifier, is used. Both strategies have been applied to a front-end based on the Wavelet Transform (WT). Results show that DFE systematically improves the performance. In fact, applying the DFE strategy to the WT-based acoustic features, a relative error reduction around 23 % (compared to the conventional features based on Short-Time Fourier Transform) is achieved when using the SpeechDat database with a vocabulary of 1000 words. 1.

Read the paper · More papers on PaperTik