Multiscale Convolutional Fusion Network for Efficient Monaural Speech Separation
Rui Yang, Shanliang Pan · IEEE Access · 2025
Speech separation is crucial for robust speech processing in real-world acoustic environments. To enable deployment in resource-limited scenarios, a speech separation system must balance separation performance, computational efficiency, and inference speed. In time-domain approaches, recurrent convolutional neural network (RCNN)-based models employ multiscale convolutional architectures to model the long sequences of encoded mixture signals. Compared to dual-path models, RCNN-based models offer higher computational efficiency but often struggle to achieve competitive performance. To address this limitation, we propose a multiscale convolutional fusion network (MSCF-Net), an efficient RCNN-based architecture designed to enhance the performance of lightweight speech separation models. The MSCF-Net follows the encoder-mask estimation-decoder pipeline, where the mask estimation process consists of parameter-shared multiscale convolutional fusion (MSCF) modules. MSCF first employs dynamic convolution-based downsampling to enhance the multiscale feature representation. Adaptive gating and multiscale convolutions are then utilized to capture complex acoustic patterns. Finally, cross-scale modulation upsampling efficiently reconstructs the acoustic features. Experiments on three datasets demonstrate that our method achieves state-of-the-art performance among lightweight speech separation models, with low computational complexity and fast inference. Specifically, MSCF-Net achieves 19.1 dB SI-SNRi on the WSJ0-2Mix test set with only 3.4 M parameters and 6.67 G/s multiply-accumulate operations (MACs) per second, and achieves 6 times real-time inference speed on the test CPU.