Speech Conv-Mamba: Selective Structured State Space Model With Temporal Dilated Convolution for Efficient Speech Separation

Debang Liu, Tianqi Zhang, Ying Wei, Chen Yi, Mads Græsbøll Christensen · IEEE Signal Processing Letters · 2025

As a selective state space model, Mamba exhibits outstanding performance and efficiency in sequence modeling tasks. Therefore, in this paper, we use Mamba as the fundamental network component to construct a novel speech separation model, Speech Conv-Mamba. Specifically, this model embeds Mamba within a U-shaped convolutional network to build the encoder and decoder network for high-dimensional representation and waveform reconstruction of speech signals. Additionally, we stack multiple temporal dilated convolutions and Mamba to create the separation network for separation task. Our comparative experiments on the GRID2Mix and Libri2Mix datasets demonstrate that the proposed model Speech Conv-Mamba, which achieves 98% and 89% of SepFormer's separation accuracy on two datasets using only 9% (2.4 M) of its model size, provides much less computational complexity and training cost.

Read the paper · More papers on PaperTik