Phase-Aware Speech Enhancement Using Multi-Head Attention and Complex Convolutions in SEGAN

Yang He, Qingzhi Du, Lin Duo · 2025

When traditional Generative Adversarial Networks (GANs) are applied to speech enhancement, they typically focus on magnitude spectrum features while paying insufficient attention to the phase spectrum. This oversight can lead to noticeable distortions in enhanced speech, especially in complex noise environments. To address this issue, this paper proposes PASEGAN (Phase-Aware SEGAN), which introduces multi-head self-attention and learnable positional encoding in the generator to enhance long-term dependency modeling. Additionally, a multi-head self-attention-based phase recovery network is employed to jointly optimize the magnitude and phase spectra, mitigating distortion caused by missing phase information. To improve the accuracy of complex-domain speech signal modeling, PASEGAN replaces traditional real-valued convolution with complex convolution and deconvolution and integrates the encoded features with latent variables to enhance the generation performance. Experiments were conducted on the Voice Bank + DEMAND dataset, and results show that under an extreme -10 dB SNR condition, PASEGAN achieves an average PESQ improvement of approximately 4.5% over traditional SEGAN for pink, white, babble, and hfchannel noise types, demonstrating higher enhancement quality and robustness. Both quantitative and qualitative evaluations further validate the superior performance of PASEGAN.

Read the paper · More papers on PaperTik