How to make more efficient use of the fact that the speech signal is dynamic and redundant
L.C.W. Pols, Reinier Plomp · 2005
Contrary to human speech perception, most speech-analysis and -processing techniques for front-end automatic speech recognition are based on discrete spectral analysis without implying temporal continuity, dynamic aspects, or context dependency. Speech recognition is presently primarily based on one specific signal aspect at a time, whereas one should make more efficient use of parallel information channels handling multiple cues. The actual global and local conditions should specify which single or combined set of parameters is important at that moment. Dynamic information is worth to receive more attention. The human ear actually seems to be 'overpowered' to process good quality speech; only under critical conditions the full analyzing capability of the ear and all redundant information in the speech signal have to be used. Automatic speech recognition systems not yet have this 'overcapacity' and therefore should rely more on multiple, highly resistant cues.