Neural Virtual Microphone Estimator: Application to Multi-Talker Reverberant Mixtures

Hanako Segawa, Tsubasa Ochiai, Marc Delcroix, Tomohiro Nakatani, Rintaro Ikeshita, Shoko Araki, Takeshi Yamada, Shoji Makino · 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) · 2022

The performance of array processing is limited when the number of available microphones is insufficient. Virtual microphone estimation (VME) aims at tackling this limitation by virtually increasing the number of microphones. Recently, we proposed the neural network-based VME approach (NN-VME), which uses a neural network to predict the signal at a virtual microphone position given observed microphone signals. In our previous study, we only evaluated NN-VME under a single-talker noisy acoustic condition for the supervised beamforming, but its applicability to more difficult acoustic conditions has not been fully explored. In this paper, we apply NN-VME to more challenging acoustic conditions and use different array processing approaches; 1) an underdetermined multi-talker scenario, 2) a far-field reverberant scenario, and 3) blind source separation (BSS). Experimental results demonstrate the applicability of NN-VME to 1) estimating the virtual microphones with high accuracy even when multiple speakers are simultaneously speaking, 2) working in the presence of a certain level of reverberation, although the VME accuracy decreases as the reverberation time becomes larger, and 3) enabling BSS approaches to work under the underdetermined condition where they could not be originally applied.

Read the paper · More papers on PaperTik