A lightweight neural network model for spectral speech enhancement
P. Wang, Chundong Xu, Fengpei Ge · 2023
In order to solve the deep neural network speech enhancement system is currently facing the challenge of high computational resource requirements. In order to address this problem, this paper proposes a lightweight speech enhancement network model, which adopts a bidirectional gated recurrent unit model in the time domain and a Transformer module in the frequency domain. In order to fully utilize the time domain information and the frequency domain information, the BiGRU module and the Transformer module are combined as the intermediate layer of the encoder and the decoder to process the time-frequency domain information alternately. The BiGRU module and the Transformer module are combined as an intermediate layer between the encoder and the decoder to process time-frequency domain information alternately. We stitch the real-virtual encoder, the stacked two-layer BiGRU-Transformer module (BiGTM) and the real-virtual decoder to form our BiGTN (BiGRU-Transformer Network). The BiGTN utilizes the two stacked BiGTM blocks to efficiently extract the local and global information from the encoder output level by level and efficiently combine the combination of speech time-frequency information and hidden features. Finally, we also set up a new joint time-frequency loss function to better train our time-frequency domain model. The effectiveness of the proposed strategy is verified and it is shown experimentally that the proposed model requires only 0.31M parameters, which is smaller than the number of parameters of other baseline models to achieve competitive performance.