Single-Channel Source Separation Using Complex Matrix Factorization
Brian J. King, Les Atlas · IEEE Transactions on Audio Speech and Language Processing · 2011
Nonnegative matrix factorization is gaining popularity in speech and audio processing applications. Performing nonnegative matrix factorization on a complex-valued short-time Fourier transform, however, makes assumptions on the signal, such as additivity in the magnitude domain, potentially degrading the results. One application where these assumptions can cause a problem is in single-channel source separation of overlapping speech. In this paper, we present how this problem can be solved by incorporating phase estimation via complex matrix factorization. Another challenge in source separation is how to select reconstruction bases for optimal separation. In this paper, we compare the most common method with a new, simpler method of finding bases that does not share many of the challenges of the current, established method. The paper will conclude by comparing nonnegative with complex matrix factorization as well as the previous and new methods for finding bases on the task of automatic speech recognition of single-channel two-talker overlapping speech.