Data-driven mask generation for source separation
Nilesh Madhu · Proceedings of the International Symposium on Auditory and Audiological Research · 2009
Presented is a microphone-array based approach for the extraction of a target signal from a mixture of compteting sources and background noise. The approach builds upon a recent proposal for source localization and tracking in the general M-microphone Q-source case, and extends it to a versatile framework to perform source separation using data-driven soft– or hard– masks. The proposed approach is applicable to any arbitrary array – allowing for its integration into binaural hearing aids. The advantage of the proposed mask generation, in contrast to current algorithms, is the implicit scalability with respect to M, Q, source spread and the amount of reverberation – obviating the need for a heuristic adaptation of the mask generation algorithm in different acoustical scenarios. Further, the individual signals extracted using these soft-masks evince low amounts of musical noise. Additional mask smoothing may be performed to further reduce the musical noise phenomenon, thereby improving the listening experience.