Neural Networks Using Full-Band and Subband Spatial Features for Mask Based Source Separation
Alexander Bohlender, Ann Spriet, Wouter Tirry, Nilesh Madhu · 2021 29th European Signal Processing Conference (EUSIPCO) · 2021
With a microphone array, spatial diversity can be exploited to estimate time-frequency masks that effectively suppress interfering speakers as well as noise. Here, we propose a deep learning approach where the signal components are distinguished based on the associated directions of arrival. To capture the target signal spectrogram more accurately, the estimation can be performed for each subband separately. In order to also take advantage of cross-band dependencies, we additionally consider a combined subband and full-band architecture. Our evaluation indicates that this combination consistently improves the performance in terms of instrumental quality metrics as compared to a pure subband or full-band method. Further, the comparison with two baseline approaches demonstrates the effectiveness of the location based deep learning approach.