Robust speech processing by humans and machines: the role of spectro-temporal modulations

Mounya Elhilali, Sridhar Krishna Nemala · 2012

Most speech processing systems, like automatic speech and speaker recognition systems, suffer from a drastic drop in performance when there is a mismatch between field test data and data used in the system training. The performance degradation is significant even when the levels of mismatch are low. This is quite in contrast to human speech processing which is robust even at relatively high levels of distortion and noise. In this thesis, we propose a number of approaches for robust speech processing and encoding based on the analyses of spectral and temporal modulations which have been shown to be the main carrier of information in speech and critical for perception. We first derive optimal constraints in modulation space for robust speech recognition. We investigate how the boundaries of the spectro-temporal modulation profile are used by humans and machine to provide robust contours along which speech can be successfully interpreted even in presence of strong noise distortions. Extending from the identification of important information regions in the speech modulation space, we show how the information in different subspaces can be effectively used. We propose a multistream feature parameterization that integrates slow and fast dynamics of speech. Three feature streams are carefully designed, in which each stream by itself noise-robust, encodes a full range of either slow and/or fast spectral and temporal modulations. To better characterize spectral information, a multi-resolution representation that focuses on information-rich spectral attributes of speech is explored. The representation allows for identification of message and speaker dominant regions in the speech signal, and we define feature schemes to address two diverse tasks such as speech and speaker recognition in a joint framework. Complimenting the robust feature encoding techniques, we propose a novel approach to speech intelligibility assessment. The model presents a new direction for front-end pruning of acoustic signals (as compared to methods based on voice activity detection), and we show an application of the same for robust speaker recognition. The ability to deal with mismatch train/test acoustic conditions is all the more important and critical as systems are increasingly used to process the speech data input from mobile devices, where the acoustic input is expected to come from a diverse set of acoustic environments and channel conditions. Using the methods proposed in this thesis, we specifically show significant advantages (over state-of-the-art schemes) for speech and speaker recognition tasks, but the methods can be extended to other automatic speech processing tasks as well.

Read the paper · More papers on PaperTik