DConvT: Deep Convolution-Transformer Network Utilizing Multi-scale Temporal Attention for Speech Enhancement

Hoang Ngoc Chau, Anh Xuan Tran Thi, Quoc Cuong Nguyen · 2024

Transformers have made significant progress in a wide range of speech processing tasks. However, they are not widely explored in real-time speech enhancement due to the complex distribution of noise over speech signal in challenging scenarios. Pure transformers can hardly capture the structure of speech harmonics and how they change over time in the noisy speech spectrogram. To fill the gap, we propose a deep convolution-transformer network (DConvT), in which the multi-scale temporal attention is designed to effectively model different speech structures from multiple temporal scales in a convolutional encoder-decoder architecture. In DConvT, multi-scale temporal information is extracted by dilated convolutions and projected to the query/key-value sequence of the self-attention operation. In this way, the self-attention can perform global modeling on a variety of temporal structures separately. The multi-scale temporal attention results are subsequently merged by concatenation to fuse the information of multiple temporal patterns. In addition, we adopt a convolutional encoder-decoder (CED) architecture to efficiently model and down-sample the frequency dimension of the speech spectrogram, which produces compact sequences for temporal modeling. The experimental results on 2020 Deep Noise Suppression Challenge (DNS20) dataset show that DConvT achieves state-of-the-art performance among existing speech enhancement baselines.

Read the paper · More papers on PaperTik