Time-Frequency Masking Method Using Wavelet Transform for BSS Problem
Masahiro Yashita, Nozomu Hamada · 2006
In this paper, we propose a novel speech separation method for blind source separation problem using complex wavelet transform. Sound source separation, especially blind source separation (BSS), is necessary for speech-based human-machine interfaces. It is because BSS needs no prior information and estimates source signals only from observed signals with microphones. Time-frequency masking is famous approach for BSS of speech mixtures. It assumes the sparseness property of speech that is called "W-disjoint orthogonality (WDO)". Wavelet transform is known to be useful for analyzing nonstationary signals. We investigate the sound source separation method combining time-frequency masking and complex-valued wavelet transform known as RI-spline wavelet. As a multi-resolution method, we perform the complex multi resolution analysis (CMRA) and subband decomposition (SD). It is shown that the proposed method realizes the high sound source separation ability through simulations