Comparison of standard ASR front ends and auditory models in neural net-based automatic speech recognition

Mark Terry, Hynek Heřmanský · The Journal of the Acoustical Society of America · 1988

Recent work [Renals and Terry, European Conference on Speech Technology, Edinburgh, United Kingdom (1987)] reported on a small multispeaker isolated digit automatic speech recognition (ASR) experiment using a backward propagation neural net system with a synchrony auditory model as its front end. In spite of fairly crude temporal normalization, the system was capable of better than 80% ASR accuracy. The question remains if the reported performance is due to (a) the paticular front end, (b) the particular neural net-based classification, (c) the particular time-normalization scheme, or (d) the combination of all factors. The goal here is to isolate these factors. In the present contribution, different front ends in the neural net-based ASR are systematically evaluated. Several standard ASR front ends are compared with the synchrony auditory model and the perceptually based linear predictive auditory model front ends in both the speaker-dependent and the speaker-independent ASR. The speed of learning and the ASR accuracy of the compared recognizers are reported and discussed.

Read the paper · More papers on PaperTik