Convolution-augmented external attention model for time domain speech separation
Yuning Zhang, He Yan, Linshan Du, Mengxue Li · 2023
The ability of the separator to capture the context-detailed features of speech signals and the number of parameters directly affect the accuracy and efficiency of speech separation in time-domain speech separation network (TasNet). This paper combines lightweight external attention with convolution and extends external attention to channel dimension; while satisfying the fine-grained extraction and modeling of spatial-channel correlation, it maintains small parameters and computation. Convolutional position coding is also used to integrate the contextual relationship and relative position information of speech features better. The above module then applies as a separator in the encoder-decoder structure based on TasNet, and a new convolution-augment external attention model for time-domain speech separation is proposed: ExConNet. The comparative experimental results show that ExConNet achieves considerable accuracy of speech separation, while its model parameters and calculation amount are significantly reduced, which can better meet the need for efficiency of speech separation.