Speech-in-noise perception and recognition

Ann R. Bradlow, Carol Espy-Wilson · The Journal of the Acoustical Society of America · 2007

Recent speech research has established that, for humans, speech-in-noise and speech-in-quiet perception differ along numerous dimensions that span levels of signal encoding and linguistic representation. Listeners place more or less weight on specific acoustic cues, draw more or less on signal-independent, contextual information, and are more or less distracted by lexical neighbors depending on masker type and level. Moreover, noise has different effects at different levels for different listener populations. Machine recognition of noisy speech is well below (at least by an order of magnitude) human recognition of noisy speech. This difference is true for both additive noise and convolutive noise. Researchers have focused on the problem of machine recognition of noisy speech in several ways: Developing robust features, training systems on noisy speech, or developing speech enhancement algorithms to clean up the noisy speech signal before performing recognition. A remaining challenge for understanding how humans do and how computers should handle speech-in-noise is to develop a conceptual framework that goes beyond the division of masking effects into peripheral, energetic versus central, informational.

Read the paper · More papers on PaperTik