Using the bi-modality of speech for convolutive frequency domain blind source separation

Andrew J. Aubrey, Yulia Hicks, Jonathan Lees, Jonathon A. Chambers · 2006

The problem of blind source separation for the case of convolutive mixtures of speech is considered. A novel algorithm is proposed that exploits the bi-modality of speech. This is achieved by incorporating joint audio-visual features into an existing BSS algorithm for the purpose of improving the convergence rate of the source separation algorithm. The increase in the rate of convergence when using a joint audio-visual model compared to using raw audio data (i.e. no model) is shown with simulations. The difference between using time varying (HMM) and stationary (GMM) statistical models to model the joint audio-visual features is also considered.

Read the paper · More papers on PaperTik