ATNSCUNet: Adaptive Transformer Network based Sub-Convolutional U-Net for Single channel Speech Enhancement
Sivaramakrishna Yechuri, Sunny Dayal Vanambathina · 2023
Recent advancements in deep learning-based speech enhancement models have extensively used attention mechanisms. Attention mechanisms-based models achieve state-of-the-art methods by demonstrating their effectiveness. This paper proposes a adaptive transformer network based on sub-convolutional U-Net (ATNSCUNet) for speech enhancement. To overcome the long-dependency problem in sequence modelling networks, we employ a adaptive transformer network (ATN) between the sub-convolutional U-Net encoder and decoder instead of temporal convolutional networks and recurrent neural networks. ATN consists of a adaptive time-frequency attention (ATFA) module made up with Multi-head self attention and Gated recurrent unit. Together, the multiple ATFA modules create an "attention-inattention" structure based on adaptive attention weights. ATFA effectively captures time-frequency dependencies, and aggregates global contextual information. Additionally, to overcome the less receptive area problems in the feature learning process, we utilize the sub-convolutional U-Net (SCUNet) model. SCUNet is a convolutional encoder-decoder model that uses differentsized kernels to produce features at various scales. The proposed ATNSCUNet model outperforms several state-of-the-art methods based on experimental results.