Energy-weighted Mean Shift algorithm for speech source separation
David Ayllón, Roberto Gil‐Pita, P. Jarabo-Amores, Manuel Rosa-Zurera, Cosme Llerena-Aguilar · 2011
Blind Source Separation algorithms have been applied to speech mixtures during many years, taking into account the knowledge and properties of speech signals. A new approach for speech separation based on sparse representations of speech has recently arisen. These methods are commonly known as Time-Frequency Masking methods, being the most famous the DUET algorithm that performs separation of undetermined mixtures from only two microphones. Sparsity property also encourages the idea of applying clustering techniques for source separation. In this work, we introduce an adapted version of the clustering method Mean Shift for the separation of speech sources. Obtained results confirm the validity of the method for speech separation improving the DUET performance and showing better generalization. Furthermore, the use of clustering techniques for separation enables the automatic identification of the number of sources.