DNN-Based Speech Enhancement on Microcontrollers

Zahra Kokhazad, Charalampos Bournas, Dimitrios Gkountelos, Γεώργιος Κεραμίδας · 2024

Deep Neural Networks (DNNs) have made remarkable advances across various domains, capturing the interest of researchers, also in the fields of signal processing and speech enhancement (SE). This growing of attention has led to numerous studies focusing on developing and improving models for SE. One of the persistent challenges in the area is in designing a lightweight model (in terms of memory and computation requirements) that still delivers satisfactory results in terms of sound quality. In this paper, we present a lightweight, DNN-based SE model that offers significantly less computations than the current state-of-the-art SE models. By following a step-by-step approach, we re-evaluate the impact of various parts of typical SE models revealing the trade-offs between speech quality and computational costs. Assuming a single-core, 200MHz microcontroller device, our SE model can achieve real-time SE performance (less than 3M Multiply-Accumulate operations, MACs) with a small degradation in SE capabilities compared to state-of-the-art SE models.

Read the paper · More papers on PaperTik