Towards online maximum-likelihood-based speech clustering and separation
Mehrez Souden, Keisuke Kinoshita, Tomohiro Nakatani · The Journal of the Acoustical Society of America · 2013
This paper introduces an approach for online speech source clustering and separation, which is based on the utilization of the multichannel location information in a recursive expectation maximization (EM) algorithm. Specifically, the normalized multichannel speech-recording vector is employed as a feature vector and is modeled using Watson mixture model. The model parameters are determined by maximizing the data likelihood at every time-frequency slot in an online processing manner. Consequently, the proposed approach can continuously adjust the speech clusters. Promising results showing the advantage of the proposed approach over the batch EM algorithm in the case of two speakers with speaker movement are obtained.