CSM2Net: Exploring Effective Fusion of Complex Spectral Masking and Complex Spectral Mapping for Speech Enhancement

Yun He Lu, Tasleem Kausar, Xiaoye Wang · 2024

Speech enhancement is essential to front-end speech signal processing, which can effectively improve speech quality and clarity. Complex spectral masking and complex spectral mapping are the main methods for speech enhancement in the time-frequency domain. This paper proposes a novel fusing complex spectral masking and complex spectral mapping networks (CSM2Net) to combine both advantages. The complex spectral masking and spectral mapping branches share the same encoder and interact in the decoding stage through a dual-branch cross-attention block. In the encoding and decoding stage, a novel soft-threshold attention block extracts speech time-frequency features and makes the model focus on more critical features. Finally, the outputs of the two branches are fused by a weight-generating block, which efficiently selects the better spectrum. On the test set, CSM2Net achieves advanced performance using almost minimal model parameters compared to other state-of-the-art models.

Read the paper · More papers on PaperTik