SOUND SOURCE SEPARATION OF OVERCOMPLETE CONVOLUTIVE MIXTURES USING GENERALIZED SPARSENESS
Masahito Togami, Takashi Sumiyoshi, Akio Amano · 2006
We propose a sound source separation method that works well even if there are more sources than mixtures and signals are recorded in a reverberant room. The proposed method is based on generalized sparseness, where the number of active sources is assumed to vary from 1 to the number of mixtures M at each time-frequency point, and the proposed sparseness estimator estimates the most suitable number of active sources. Approaches using binary masks assume that only one source is active at each time-frequency point. However, when more than two sources are active, separated signals are greatly distorted with musical noise. The separated signals by the shortest-path algorithm are less distorted than those obtained by binary masks. However, when there are fewer than M active sources and noise, the shortestpath algorithm overestimates the source signal’s value. To overcome the overestimation and distortion problems, the proposed method does not fix the number of sources as one or M . Instead, those are estimated at each time-frequency point. Experimental results in a room (reverberation time = 100 ms) indicate that NRR (Noise Reduction Ratio) of signals separated by our proposed method outperform those of binary masks and the shortest-path algorithm by about 3-5db.