Insights into Magnitude and Phase Estimation by Masking and Mapping in DNN-Based Multichannel Speaker Separation
Alexander Bohlender, Ann Spriet, Wouter Tirry, Nilesh Madhu · 2024
Speakers are often separated by time-frequency masking in the short-time Fourier domain to take advantage of the high degree of sparsity of the individual speech spectrograms. Magnitude and phase can be jointly enhanced with complex masks, but prior work suggests that directly mapping the input to the complex spectrogram of the clean signal is a better alternative. For a setup with a compact microphone array, experiments conducted in this paper compare these paradigms with focus on magnitude and phase estimation. Whereas phase is enhanced effectively in general, differences between masking and mapping are minor in this regard. Spectral mapping causes the least target distortion. Complex masking better suppresses interference, but speech quality suffers due to artifacts. Combining magnitude masking with phase mapping presents a compromise, which amounts to the best performance regarding instrumental metrics.