Dual-Compression Neural Network with Optimized Output Weighting For Improved Single-Channel Speech Enhancement

Stefan Thaleiser, Aleksej Chinaev, Rainer Martin, Gerald Enzner · 2022

Approaches to the problem of single-channel spectral speech enhancement comprise either model-based signal processing or neural networks (NNs) or a combination thereof. Diverse compression functions of input features (noisy spectral magnitudes) and targets (clean speech magnitudes) have been proposed in the NN domain, most commonly the linear and logarithmic compression that will favor speech intelligibility and noise reduction, respectively. An average of a linear and a logarithmic NN prediction (averaged on the same decibel scale) has been proposed recently to benefit from both compression characteristics. In this contribution, we propose an optimized time- and frequency-dependent weighting of linear and logarithmic NN outputs by an additional computationally-light NN stage. This structure is termed dual-compression NN with optimized output (DuCO) weighting. It shows improved performance over NNs with either linear or logarithmic compression, over post-processing by output averaging, and over state-of-the-art model-based processing in terms of objective PESQ, STOI, and segmental SNR.

Read the paper · More papers on PaperTik