TF-MCRN: A Lightweight Speech Enhancement Algorithm Based on Mel Spectrogram for Speech Recognition

Ao Li, Tongjia Yan, Mingyang Li, Lin Zhou · 2025

At present, most time-frequency domain speech enhancement algorithms are based on STFT and biased towards human auditory perception. However, the introduction of phase modeling has significantly increased the model complexity, leading to poor compatibility with speech recognition tasks. To address this problem, we propose a lightweight speech enhancement algorithm based on Mel spectrogram(TF-MCRN). TF-MCRN utilizes the Mel filter banks to compress the speech signal spectrum nonlinearly, and employs Depthwise Separable Convolution(DSConv) to extract the local features of the speech signal. Furthermore, it models the global correlation of speech along the frequency and the time dimension using Feedforward Sequential Memory Network (FSMN) and a self-attention equipped Memory module (SAN-M). Additionally, we introduce a linear combination loss function and two post-processing methods to reconstruct speech signal from the Mel spectrogram. Experimental results demonstrate that our proposed algorithm, which requires only 0.52 M parameters and 1.23 GMACs per second, achieves a PESQ score of 3.02 on the VBD dataset and exhibits improved compatibility with speech recognition models.

Read the paper · More papers on PaperTik