A Deep Learning-Based Multi-Feature Ambisonics Speech Enhancement Algorithm

Jiacheng Zhou, Yi Zhou, Hongqing Liu · 2024

Speech enhancement methods for First order Ambisonics (FOA) signals mainly have focused on multi-channel microphone array-based techniques, without exploring the unique characteristics of Ambisonics signals, which limits the speech quality and intelligibility. The subsequent speech recognition task will thus not benefit. This paper proposes a speech enhancement algorithm for FOA speech signals, which combines various features to achieve better speech enhancement effects. The algorithm extends the Dual-Path Convolution Recurrent Network (DPCRN) to multi-channel scenarios. The sound intensity vector is used as an auxiliary feature to provide the energy distribution information of the speech signal. The phase guidance module is used to provide the spatial information, and the amplitude noise reduction of the omnidirectional microphone channel is provided by the W noise reduction module to provide a clearer back-projection reference. Experimental results based on STOI, WER and PESQ show that the proposed algorithm outperforms the baseline model MIMO-UNet on L3DA$S$23 challenge dataset, and ablation studies also demonstrate the effectiveness of the proposed features.

Read the paper · More papers on PaperTik