Improving GCRN via channel-spatial attention and multi-objective learning for monaural speech enhancement
Zhu Xiaojun, Yao Hailong, Zhang Baoxiang, Huang Heming · IET conference proceedings. · 2025
This work presents an improved speech enhancement model based on a classical convolutional recurrent network to alleviate its problems in terms of global context modeling and phase information utilization. In the proposed model, the following improvements have been made: firstly, a multi-scale convolutional block with gated linear units is introduced to obtain speech encoding information from different scales and mitigate the problem of gradient vanishing that may occur during training; secondly, to improve the feature representation and selectively focus on key information regions, a lightweight attention mechanism is integrated after each encoding layer of encoder and long short-term memory layer; thirdly, a shared decoder with spectrum-specific heads is designed for joint estimation the real and imaginary parts of complex spectra while reducing model parameters; finally, a multi-objective training strategy is adopted to solve the problem that learning the real and imaginary spectra cannot directly reduce the time-domain signal distortion. The experimental results on a noisy dataset constructed from the TIMIT corpus show that the proposed model achieves significant improvements in the intelligibility and quality of target speech.