DPATD: Dual-Phase Audio Transformer for Denoising
Junhui Li, Pu Wang, Jialu Li, Xinzhe Wang, Youshan Zhang · 2023
Recent high-performance transformer-based speech enhancement models demonstrate that time-domain methods could achieve similar performance as time-frequency domain methods. However, time-domain speech enhancement applications typically accept audio inputs that are composed of a large number of time steps, making it challenging to model extremely long sequences and train models to perform adequately. In this paper, we utilize smaller audio chunks as input to achieve efficient utilization of audio information to overcome the aforementioned challenges. We present a dual-phase audio transformer for denoising named DPATD, a novel deep model to organize transformer layers to learn clean audio samples for denoising. Using DPATD, audio input sequences are divided into manageable audio chunks, and these short chunks can be effectively and efficiently processed by our memory-compressed explainable attention. It also converges faster compared to the frequently used self-attention module. Extensive experimental results show that our proposed DPATD model is better than state-of-the-art models.