Multi-Kernel Attention Encoder For Time-Domain Speech Separation

Zengrun Liu, Diya Shi, Ying Wei · 2024

Recent years, great progress has been made in time-domain single-channel speech separation. In this paper, we propose an innovative method of using multiple convolutional kernels and channel attention in time-domain speech separation codecs called multi-kernel attention encoder. We propose a novel encoder structure that can capture various time-domain features of the input speech signal by using multiple convolutional kernels. Additionally, we introduce a channel attention mechanism that allows the model to adaptively adjust the importance of each channel, thereby further improving speech separation accuracy. Experimental results demonstrate that our method achieves good performance improvements in speech separation tasks, demonstrating its effectiveness and robustness. Our research provides a new and effective solution for time-domain speech separation. The results on the Libri2Mix dataset show that our method has reached the current excellent level, with SDRi of 17.3dB and SISNRi of 16.9dB.

Read the paper · More papers on PaperTik