Robust automatic speech recognition in reverberation: onset enhancement versus binaural source separation
Hyung‐Min Park, Matthew Maciejewski, Chanwoo Kim, Richard M. Stern · The Journal of the Acoustical Society of America · 2015
The precedence effect describes the auditory system’s ability to suppress later-arriving components of sound in a reverberant environment, maintaining the perceived arrival azimuth of a sound in the direction of the actual source, even though the later reverberant components may arrive from other directions. It is also widely believed that precedence-like processing can also improve speech intelligibility for humans and the accuracy of speech recognition systems in reverberant environments. While the mechanisms underlying the precedence effect have traditionally been assumed to be binaural in nature, it is also possible that the suppression of later-arriving components may take place monaurally, and that the suppression of the corresponding components of the spatial image may be a consequence of this more peripheral processing. This paper compares potential contributions of onset enhancement (and consequent steady-state suppression) of the envelopes of subband components of speech at the monaural and binaural levels. Experimental results indicate that substantial improvement in recognition accuracy can be obtained in reverberant environments if feature extraction includes both onset enhancement and binaural interaction. Recognition accuracy appears to be relatively unaffected by which stage in the binaural processing is the site of the suppression mechanism. [Work supported by the LG Yonam Foundation and Cisco.]