SE-Mixer: Towards an Efficient Attention-free Neural Network for Speech Enhancement

Kai Wang, Bengbeng He, Wei‐Ping Zhu · 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) · 2022

In this work, we propose a novel and attention-free architecture based on multi-layer perceptrons (MLPs) for speech enhancement, named SE-Mixer, which consists of an encoder, a decoder and a mixer module in between. The mixer module is designed for efficiently extracting the contextual information of long-range speech sequences. It employs temporal MLP augmented with convolution and frequency MLP to successively extract abundant temporal information from various time scales and capture frequency information within each time step. Our experimental results on a benchmark dataset indicate that the proposed SE-Mixer achieves a competitive performance compared to existing state-of-the-art methods with and without attention mechanism incorporated. Moreover, the proposed model contains fewer trainable parameters (about 710k).

Read the paper · More papers on PaperTik