Non-Stationary Multi-Channel (Multi-Stream) Processing Towards Robust and Adaptive ASR
Hervé A. Bourlard · 1999
In this paper, we discuss the rationale behind multi-channel processing as applied to multi-stream automatic speech recognition (ASR). In this framework, we will develop different mathematical models and discuss some interesting relationships with psycho-acoustic evidence. In the case of multi-channel processing, it is assumed that the speech signal is processed by different "experts", each expert focusing on a different characteristic of the signal, and that the different channels 1 are combined at some (temporal) stage to yield a global recognition output. Although we believe that the discussion below is valid for numerous multi-channel problems (e.g., audio and visual streams, in the case of audio-visual ASR), the present paper will mainly discuss the possible combination strategies (with application to multi-band ASR) and their relationships with different mathematical models. Finally, we will show that the proposed approaches could provide us with a new paradigm for noise robu...