Self-Attention for Multi-Channel Speech Separation in Noisy and Reverberant Environments

Conggui Liu, Yoshinao Sato · Asia-Pacific Signal and Information Processing Association Annual Summit and Conference · 2020

Despite recent advances in speech separation technology, there is much to be explored in this field, especially in the presence of noise and reverberation. One of the significant difficulties is that locations where relevant context information is incorporated vary in the time, frequency, and channel directions. To overcome this problem, we investigated the use of self-attention for multi-channel speech separation with timefrequency masking. Our base model is a temporal convolutional network that is the same as Conv-TasNet, except it works in the frequency domain with the short-time Fourier transformation and its inverse. We combined this basis with a self-attention network. We explored nine different types of self-attention network for this purpose. To investigate the effects of the self-attention networks, we evaluated the performance of the proposed model, which we refer to as a confluent self-attention convolutional temporal audio separator network (CACTasNet), on a noisy and reverberant version of the wsjO-2mix dataset. We found that several different self-attention networks substantially improved the performance measured by scale-invariant signal-to-noise ratio and signal-to-distortion ratio. The results indicate that a selfattention mechanism can efficiently locate context information relevant to speech separation.

Read the paper · More papers on PaperTik