Bimodal coherence based scale ambiguity cancellation for target speech extraction and enhancement
Qingju Liu, Wenwu Wang, Philip J. B. Jackson · 2010
We present a novel method for extracting target speech from au-ditory mixtures using bimodal coherence, which is statistically characterised by a Gaussian mixture modal (GMM) in the off-line training process, using the robust features obtained from the audio-visual speech. We then adjust the ICA-separated spectral components using the bimodal coherence in the time-frequency domain, to mitigate the scale ambiguities in different frequency bins. We tested our algorithm on the XM2VTS database, and the results show the performance improvement with our pro-posed algorithm in terms of signal to interference ratio (SIR) measurements. Index Terms: speech extraction, bimodal coherence, audio-visual, Gaussian mixture model (GMM), independent compo-