Divide and Conquer: A Low-complexity Neural Network for Monophonic Speech Enhancement

Bingxiao Fang, Liang Liu · 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) · 2022

Noise suppression is an essential speech enhancement method to reduce the background noise in communication systems. Although stationary-noise can be successfully suppressed with the conventional signal processing algorithm, the non-stationary noise, especially the instantaneous-noise, remains challenging. Impressive performance has been achieved after neural network introduced in noise suppression task in the past decade. However, good performance and low complexity is still alternative. In this study, we demonstrate a low-complexity neural network for monophonic speech enhancement, in which a divide and conquer strategy has been employed in the design of the neural network to separate the noise suppression task into an envelope enhancement module and a detail enhancement module, denoted as EDNet. This approach achieves significant quality improvements in terms of the perceptual evaluation of speech quality (PESQ), short-time objective intelligibility (STDI), and the Scale-Invariant Signal to Distortion Ratio (SI-SDR) with low computational complexity.

Read the paper · More papers on PaperTik