Dual-Path Federative Magnitude and Phase Estimation Combing Transformer with RNN for Single Channel Speech Enhancement

Manshan Liu · 2025

To enhance the effectiveness and adaptability of existing speech enhancement techniques that rely on masking and spectral mapping, especially in complex environments, we propose a novel single-channel speech enhancement approach that integrates multi-order dual-branch magnitude and phase estimation, called DRP-SENet. This method features an encoder-dual-branch decoder architecture. An intermediate layer with C-DPRNN is designed to improve the capture of local patterns in speech spectrograms and dependencies between consecutive frames. Additionally, to obtain voice channel dimension information, Channel-Aware Multi-Head Attention (CH-MHSA) is used in place of Multi-Head Self-Attention (MHSA). Results on the VoiceBank+DEMAND dataset confirm that our proposed method achieves superior subjective and objective metrics compared to most state-of-the-art related methods.

Read the paper · More papers on PaperTik