Unsupervised Training of Neural Network-Based Virtual Microphone Estimator

Jiachen Wang, Tomoki Toda · 2024

The performance of array signal processing is closely tied to the number of microphones, and technologies designed for scenarios with a limited number of microphones have garnered significant attention due to practical device limitations. In response to this challenge, a solution was proposed in the form of a neural network-based virtual microphone estimator (NN-VME), aimed at generating virtual microphone (VM) signals. However, these methods typically demand data from the VMs' positions for neural network training, thereby increasing the difficulty of applying the method under realistic conditions. This paper presents a practical solution: unsupervised training of NN - VME using a novel loss function, which alleviates the data requirements for model training. The proposed loss function optimizes the model by enhancing the linear correlation between estimated and actual recorded signals at alternative channels, thus eliminating the need for data from specific VMs' positions throughout the training process. Experimental results on the CHiME-4 corpus demonstrate that the introduced loss function significantly improves NN-VME's ability to estimate high-performance VM signals, leading to enhanced beamformer performance. This finding underscores the potential to generate continuous VM signals across the entire array. Building on this insight, we further investigate the efficacy of increasing the number of VMs and explore the potential benefits of incorporating a weighted loss function to capture positional information within the array.

Read the paper · More papers on PaperTik