Two-Stage Learning and Fusion Network With Noise Aware for Time-Domain Monaural Speech Enhancement
Xiaoxiao Xiang, Xiaojuan Zhang, Haozhe Chen · IEEE Signal Processing Letters · 2021
The neural network has become a new and powerful paradigm in speech enhancement, triggering the surge of research. However, most existing methods directly predict speech and ignore the noise information. In this paper, we propose a two-stage learning and fusion network with noise awareness for time-domain monaural speech enhancement, which can be regarded as a progressive learning process. More specifically, in the first network, speech and noise are estimated simultaneously. The estimated speech and the signal by subtracting estimated noise from the noisy speech are stacked with noisy speech as input to obtain further refined speech in the second network. Both networks are mainly built on encoders and decoders with skip connections. To better control the information flow in the network, we introduce the gated linear unit in the encoder and the decoder, which also can help model complex interactions. Dilated dense blocks are added after each layer of the encoder and decoder to improve the model efficiency and enlarge the receptive field. Our experiments confirm that the proposed two-stage learning network with noise awareness achieves better performance than several advanced systems under various conditions.