Deep Learning Based Speech Enhancement on Edge Devices Applied to Assistive Work Equipment
Nika Šljubura, Marko Šimić, Vedran Bilas · 2024
In urban environments, enhancing speech and emergency sounds such as alarms and sirens amongst intense surrounding noise is crucial to ensure clear communication while preserving safety. We aim to develop a deep learning-based speech enhancement system suitable for integration into wearable assistive devices for industrial helmets. While state-of-the-art deep learning techniques show promising results in enhancing both speech and emergency sounds, due to their high memory demand and power consumption, they are not applicable in resource-constrained devices. In this work, we address these limitations by adapting network architecture for low-power microcontrollers. The low memory constraint is satisfied through layer refinement, feature selection, and model quantization, and the proposed model is deployed on a Cortex-M7-based microcontroller. The deployed model achieves a 150 ms inference time with 48 mJ energy consumption per prediction, allowing integration into assistive work equipment. Moreover, our model outperforms state-of-the-art by achieving a 5% increase in classification accuracy, showing improved capability in detecting emergency signals.