Phase-Aware Speech Enhancement with Dual-Stream Architecture and MetricGAN
Hanlong Wang, Xintao Jiao · 2025
Many studies have explored how to utilize phase information. However, the extraction of the relationship between magnitude and phase remains inadequate. To bridge this gap, this paper proposes a phase-aware dual-stream architecture model. The model includes an encoder and two downstream decoders. The two decoders are responsible for modeling magnitude and phase information, respectively. Moreover, in order to capture the implicit association between magnitude and phase, the model establishes a fusion module between the two decoders, allowing reference to predicted magnitude information during the process of predicting phase information. To avoid the compensation effect between magnitude and phase, the predicted magnitude spectrum is applied with a stop-gradient operator before magnitude fusion, thereby blocking phase related gradients from flowing into magnitude arguments. To discard the discrepancy between predicted PESQ and calculated PESQ, additional phase input is introduced to the discriminator, enabling discriminator to accurately simulate the PESQ calculation process. We achieves a PESQ score of 3.53 on the VoiceBank+DEMAND datasets, which indicates a certain improvement over previous work.