A Novel Low-Complexity Attention-Driven Composite Model for Speech Enhancement

Mojtaba Hasannezhad, Wei‐Ping Zhu, Benoı̂t Champagne · 2021

Speech exhibits strong dependencies among its samples in both time and frequency domains. In this paper, we propose a low-complexity composite model for speech enhancement (SE) that integrates a convolutional neural network (CNN) and a long short-term memory (LSTM) network. These two modules take full advantage of the spectral and temporal information of input speech and extract in parallel a complementary set of features. The CNN is enabled to capture non-local spectral information via dilated frequency convolutions. It also incorporates an attention mechanism to recalibrate its weights without imposing considerable additional complexity. A grouping strategy is adopted for LSTM implementation to reduce its complexity while keeping performance almost unchanged. Our composite model is carefully designed to address concerns in real-time applications including limited computational resources, low-latency processing, and causal architecture. Through extensive and comparative simulation studies, it is shown that the proposed model significantly outperforms some other DNN-based SE methods in the recent literature.

Read the paper · More papers on PaperTik