A Low-cost and Energy-efficient Real-time Speech Isolation Model for AIoT
Kaibo Xu, Wei-Han Yu, Ka-Fai Un, Rui Paulo Martins, Pui‐In Mak · 2024
In the pursuit of efficient speech processing, deep learning-based denoising models have been hindered by a large number of parameters and high computational intensity. In this paper, we present a novel deep learning-based lightweight time-domain speech denoising model that avoids the computational overhead of the Short-Time Fourier Transform (STFT) used in frequency-domain approaches. The design focus of this model is to achieve real-time processing and small model size while also achieving good speech isolation performance. Inspired by the U-Net architecture, the model uses a downsampling and upsampling strategy to compress and then restore data dimensions, with an encoder-separator-decoder structure at its core. The encoder, consisting of five convolutional modules with different stride lengths, captures a wide range of speech features, which are then refined by a transformer-based separator. The transformer, with its unique self-attention mechanism, is capable of capturing long-range dependencies within the speech signal, a crucial aspect for accurately predicting and separating noise components. This enables the model to better understand the context and relationships within the speech waveform, leading to enhanced noise reduction capabilities. The decoder, employing transposed convolutional layers, achieves high fidelity in speech prediction. This model not only reduces the number of parameters to less than 100K, but also maintains denoising effectiveness, making it viable for use on devices with limited computing power. The results indicate a promising direction for time-domain models in speech processing applications.