Exploiting Uncertainties for Binaural Speech Recognition
Soundararajan Srinivasan, Nicoleta Roman, DeLiang Wang · 2007
Recently several algorithms have been proposed to enhance noisy speech by estimating the signal-to-noise ratio (SNR) within a local time-frequency region based on binaural cues of interaural time and intensity differences (ITD and IID). However, the accuracy of the estimated SNR often varies widely across time and frequency, causing uncertainties in the enhanced speech features. We estimate this uncertainty based on statistics of ITD and IID and show that it can be effectively exploited to improve robust speech recognition. Systematic evaluations using the estimated uncertainty show significant improvement in recognition performance compared to the baseline performance.