Stop consonant identification using auditory images.
Lawrence L. Feth, Robert A. Fox, Ina Rea Bicknell · The Journal of the Acoustical Society of America · 1992
A number of researchers over the past decade have argued that salient characteristics of auditory processing must be incorporated into speech recognition models before they can accurately represent human speech processing which might, in turn, improve recognition accuracy. Recently, Fox and Feth [Proc. Ninth International Symposium on Hearing, Carcans, France (1991)] processed a series of stop+vowel CVs (using [bdgptk] and the vowels [i schwa æ open aye u]) through Patterson and Holdsworth’s auditory sensation processing (ASP) model producing a set of ‘‘stabilized auditory images.’’ Each image represented a spectral representation of 15 ms of the acoustic signal and the images for each token could be displayed in rapid sequential order producing a ‘‘cartoon’’ of the auditory changes evident over a wide range of auditory analysis channels. A set of viewing experiments conducted using three subjects demonstrated that these stops (and vowels) could be accurately identified on the basis of the dynamic auditory cues contained in the auditory images. Accuracy rates up to 99% for voicing and 92% for place of articulation were obtained. The present study is a continuation of this research using a longer display window that may more accurately reflect short auditory memory in perceptual processing. [Research supported, in part, by an AFOSR grant to L. Feth and a NIA grant to R. Fox.]