Meta-Learning For Variable Array Configurations in End-to-End Few-Shot Multichannel Speech Enhancement

Alina Mannanova, Kristina Tesch, Jean-Marie Lemercier, Timo Gerkmann · 2024

Nowadays deep neural networks are a common choice for multichannel speech processing as they may outperform the traditional concatenation of a linear beamformer and a post-filter in challenging scenarios. To obtain strong spatial selectivity, these approaches are typically trained for a specific microphone array configuration. However, it was recently shown that such models are sensitive even to small perturbations in the microphones placements. In this paper we propose a method for handling variable array configurations based on model-agnostic meta-learning. We demonstrate that the proposed approach increases robustness to changes in the array configurations, i.e., mismatched conditions, while maintaining the same performance as the array-specific model on matched conditions.

Read the paper · More papers on PaperTik