Light-Weight Causal Speech Enhancement Using Time-Varying Multi-Resolution Filtering
Venkatesh Parvathala, K. Sri Rama Murty · 2024
In the recent past, deep neural network (DNN) approaches have achieved significant improvements in speech enhancement. However, the improved performance is accompanied by a significant increase in computational complexity, which hinders their ability to handle real-time data. Consequently, it is essential to develop networks with reduced computational complexity. In this work, we propose a real-time network named neural time-varying filtering (NTVF) to estimate a time-varying filter in the frequency domain. This network employs dense and convolutional layers to perform context-independent filtering (CIF) and context-dependent filtering (CDF) operations, respectively. We further propose multi-resolution filtering for CDF to adaptively choose the context window using the attention mechanism. Moreover, a combination of spectral, perception, and production related loss functions is proposed to efficiently learn the network parameters. Our experimental evaluations indicate that the proposed causal NTVF network achieves a PESQ of 3.05 on the Valentini dataset with only 0.34 million parameters and 0.01 real-time factor.