MixSENet: A Lightweight Model for Speech Enhancement with Multi-Scale Features and Contextual Modeling
Chuike Kong, Guangcun Wei, Shuo Li, Penghao Ma, Changhao Li · 2025
In this paper, we propose a novel lightweight speech enhancement network, MixSENet (Mixed Structure Speech Enhancement Network), designed to address the challenges associated with multi-scale feature processing of speech signals and coping with complex noise environments. The network is based on the U-Net architecture and innovatively introduces a hybrid structural block, which consists of a multi-scale parallel large convolutional kernel module (MSPLCK) and an enhanced parallel attention module (EPAM). The MSPLCK achieves large sensory fields and multi-scale feature extraction through parallel dilated convolution, whereas the EPAM can simultaneously process global shared information and local time-frequency features effectively coping with uneven noise distributions. Extensive experiments demonstrate that MixSENet shows significant performance advantages on the Voice Bank + DEMAND dataset, with a parameter count of only 0.43M, significantly reducing training and deployment costs. The method has substantial practical value in real application scenarios, especially for resource-constrained mobile devices and embedded systems.