Modelling interaural level and phase cues with Student's t-distribution for robust clustering in MESSL
Zeinab Zohny, Jonathon A. Chambers · 2014
The state of the art model-based expectation maximization source separation and localization (MESSL) algorithm successfully separates multiple sound sources from only two-channel reverberant mixtures. Since MESSL achieves under-determined convolutive blind source separation by essentially clustering spectrogram points based on their interaural spatial cues, the performance of MESSL degrades substantially when the speech sources are in close proximity. In this paper, we therefore enhance its performance by the integration of robust clustering based on the Student's t-distribution. This heavy-tailed distribution, as compared to the Gaussian distribution originally used in MESSL for parametric modelling, can potentially better capture outlier values and thereby lead to more accurate probabilistic masks for source separation. The student's t-distribution is exploited in modelling both the interaural phase difference (IPD) and the interaural level difference (ILD) in order to better represent the uncertainties introduced by noise, reverberations as well as the statistical non-stationarity of speech signals. Simulation studies based on speech mixtures formed from the TIMIT database confirm the advantage of the proposed approach in terms of signal to distortion ratio (SDR) and the perceptual evaluation of speech quality (PESQ).