LTF-NET: A Low-complexity Three-band Fusion Model For Full-band Speech Enhancement

Yifen Peng, Lingyun Zou, Wenjing Zhao, Zhaohui Chen, Sha Huan, Zhong Cao · 2023

Speech enhancement is a process that improves the quality and intelligibility of speech signals by reducing background noise, reverberation, and other distortions. Early speech enhancement methods neglect phase estimation due to the insensitivity of the human ear to phase and the sensitivity of phase to noise, while recent literature has proved the importance of phase information in speech enhancement. Additionally, most speech enhancement models operate on the wideband (16 kHz) sample rate, while speech sampled at 48 kHz contains more information and features. In this paper, we proposed a low-complexity three-band fusion model for full-band (48 kHz) speech enhancement, which aims to decompose the full-band speech signal into three sub-bands, including low-(0-8 kHz), middle-(8-16 kHz), and high-band (16-24 kHz), and recover each band in steps. Specifically, a lightweight network with two similar branches for estimating magnitude spectrum and phase information in parallel with low-band is proposed and pre-trained. We proposed an adaptive time-frequency linear attention transformer-based module within each branch for temporal sequence modeling. Experiments show that our models outperform previous advanced systems models with relatively low complexity.

Read the paper · More papers on PaperTik