CCAUNet: Boosting Feature Representation Using Complex Coordinate Attention for Monaural Speech Enhancement

Sixing Liu, Yi Jiang, Qun Yang · 2024

Speech enhancement models based on deep learning, such as the complex U-Net model, have achieved good results. However, traditional methods based on convolutional neural networks often ignore the inherent properties of the speech spectrum, such as long-term temporal dependence, cross-frequency correlation and spatial position information, when processing speech signals. These properties are crucial to helping the model distinguish speech from noise and improve speech quality. In this paper, we propose a new speech enhancement model called CCAUNet. The core of the model is an innovative complex coordinate attention structure that can simultaneously capture and emphasize temporal dependence, frequency dependence and spatial position information in the speech spectrum. In addition, Furthermore, we employ multi-resolution STFT loss and SI -SNR loss for joint optimization of the model, thereby assisting the complex coordinate attention in accurately processing spectral features. The multi-resolution STFT loss can capture detailed information at different frequency scales, while the SI -SNR loss focuses on the quality of the speech signal. Experimental results conducted on the Deep Noise Suppression Challenge dataset show that the proposed CCAUNet outperforms all compared models on WB-PESQ, NB-PESQ, STOI and SI-SNR metrics.

Read the paper · More papers on PaperTik