Improved Speech Enhancement by Using Both Clean Speech and ‘Clean’ Noise

Jianqiao Cui, Stefan Bleeck · 2023

Generally, speech enhancement (SE) models based on supervised deep learning technology, use input features from both noisy and clean speech but not from the noise itself. We suggest here that this ‘clean’ background noise, before mixing it with speech, can also help SE and that is to our knowledge not described yet. In our proposed model, not only the speech, but also the noise is enhanced initially and later combined for improved intelligibility and quality. We also present a second innovation to capture better contextual information that traditional networks are often poor in. To leverage both speech and background noise information and long-term context information, this paper describes a sequence-to-sequence (S2S) mapping structure using a novel two-path speech enhancement system, consisting of two parallel paths: a Noise Enhancement Path (NEP) and a Speech Enhancement Path (SEP). In the NEP, the encoder-decoder structure is used for enhancing only the ‘clean’ noise, while the SEP is used to suppress the background noise in the clean speech. In the SEP, a Hierarchical Attention (HA) mechanism is adopted to leverage long-range sequence capture. In the NEP, we us traditional gated controlled mechanism from ConvTasnet but improve it by adding dilated convolution to increase receptive fields. Experiments are conducted on the Librispeech dataset, and results show that the proposed model performs better than recent models in various measures, including ESTOI and PESQ scores. We conclude that the simple speech plus noise paradigm often adopted for training such models is not optimal.

Read the paper · More papers on PaperTik