Speech Source Separation by Combining Localization Cues with Mixture Models of Speech Spectra
Kevin Wilson · 2007
We present a method for simultaneous speech source separation in reverberant environments using both localization cues and a speech model. Previous source separation work has focused primarily on one or the other of these approaches; we use a novel localization cue observation noise model to allow for a natural combination of the approaches. We model speech as a Gaussian mixture model (GMM) of short-time spectral magnitudes and model localization cue noise using a time-varying noise model learned from labeled training data. We show that our technique outperforms competing techniques as measured by segmental signal-to-noise ratio (SNR) and segmental log-spectral distortion (LSD) and also show that our technique is robust to typical levels of audio localization error.