A new approach to variable frame rate front-end processing for robust speech recognition

Julien Epps · 2006

Robustness in the presence of various types and severities of environmental noise has been researched extensively over the past several years, however this remains one of the main problems facing automatic speech recognition systems. This paper describes a noise-robust ASR front-end that employs a new variable frame rate analysis, based upon the first-order difference of the log energy for each frame (∆E). Compared with previous variable frame rate methods, this delta energy approach is simpler and achieves similar recognition accuracy improvements but at reduced complexity. Recognition experiments on the Aurora II connected digits database reveal that the proposed front-end achieves an average digit recognition accuracy of 68.69 % for a model set trained from clean data and 85.82 % for a model set trained from data with multiple noise conditions. Compared with the ETSI standard Melcepstral front-end, the proposed front-end obtains a relative error rate reduction of around 20 % for the clean model set, achieved consistently across nearly all signalto-noise ratios and noise conditions tested. 1.

Read the paper · More papers on PaperTik