Improved SEGAN Speech Enhancement by Fusing Residual and Convolutional Attention Mechanisms

Jingtao Huang · 2025

Speech enhancement technology aims to improve speech quality and intelligibility by recovering clear speech from noise-contaminated speech signals, widely used in speech recognition, communication, and hearing assistance. However, the existing speech enhancement models under the Generative Adversarial Network (GAN) architecture, such as SEGAN, are still deficient in deep feature extraction capability and information transfer in the coding and decoding process, even though they can improve speech quality in noisy environments to a certain extent. This is mainly manifested in the fact that SEGAN cannot effectively capture the deep features of speech signals when dealing with complex noise environments. To solve this problem, this paper proposes an improved SEGAN (CA-Res-SEGAN) model that incorporates residual connectivity and a convolutional attention mechanism. The method optimizes the transmission of information flow by introducing residual connections to reduce information loss and improve the expression ability of deep features; meanwhile, CA-Res-SEGAN combines the convolutional attention mechanism to enhance the model's attention to the key features in the speech signal, which improves the noise suppression effect and speech quality recovery. Experimental results show that CA-Res-SEGAN is better than the traditional SEGAN in several evaluation indexes, such as PESQ, STOI, and CSIG, and significantly improves the speech enhancement effect in complex noise environments.

Read the paper · More papers on PaperTik