Separating Speech from Speech Noise Annual Report 2006
Daniel P. W. Ellis · 2007
The main work at Columbia this year has been the development of algorithms for extracting and recognizing speech in nonstationary, noisy environments when only a single microphone channel is available. Our particular approach is based on using trained models to distinguish regions of time-frequency containing speech from nonspeech areas [2], and we have pursued this along several directions: One approach is to use trained models of the speech signal and to find the best set of model parameters that are consistent with the noisy speech observations. An alternative approach is to treat the labeling of each time-frequency cell as a simple classification task, and to train pattern recognition classifiers to perform this task. In that work, the challenge is to find the best classifier architecture and the most effective representation of the context. When more than one microphone channel is available, some different approaches to source separation become possible. The two-channel case is particularly interesting because this is the number of ears possessed by the typical listener. We have been looking at ways to separate sources in recordings made