Spatially Selective Speaker Separation Using a DNN With a Location Dependent Feature Extraction

Alexander Bohlender, Ann Spriet, Wouter Tirry, Nilesh Madhu · IEEE/ACM Transactions on Audio Speech and Language Processing · 2023

Deep neural networks (DNNs) have proven themselves as an effective means to separate clean speech from noisy mixtures. When there are multiple concurrent talkers, however, unambiguously defining the target output is not trivial, especially if the mixture is single-channel and the talkers are not known in advance. Although this problem can be addressed with permutation invariant training or deep clustering, the performance still suffers in this case. Approaches for compact arrays of multiple microphones can exploit spatial diversity to resolve the ambiguity: a separate output may be generated for each direction of arrival (DOA), or the speaker assignment can be controlled with a location-based training (LBT). Alternatively, we can narrow down the target definition at the input, to perform a spatially selective speaker separation instead of separating all speakers simultaneously. This is achieved by specifying freely adjustable target DOAs. On the one hand, these can be integrated as location-based input features (LBI). On the other hand, the main contribution of this work is a location dependent feature extraction (LDE): we implicitly introduce a DOA dependence in a small part of the DNN by optimizing its parameters for each DOA separately. Experiments demonstrate that LDE outperforms LBT and LBI in terms of instrumental metrics and speech recognition results. A representative audio example is presented for a qualitative impression. An analysis of the spatial selectivity reveals that target and nontarget directions can be distinguished quite well with LDE, which is also verified by recordings of real moving talkers.

Read the paper · More papers on PaperTik