Convolutional Neural Network-Based Models for Speech Denoising and Dereverberation: Algorithms and Applications

Chengshi Zheng, Yuxuan Ke, Xiaoxue Luo, Xiaodong Li · River Publishers eBooks · 2023

Recently, speech controlled smart devices play an important role in Internet of Things (IoT) applications. Both reverberation and noise may significantly reduce the efficiency of the human–machine interaction for indoor applications. Therefore, speech enhancement becomes a critical front-end technique to improve the performance, which has attracted increasing attention in recent years. This chapter focuses on deep learning (DL) based monaural speech enhancement algorithms for both denoising and dereverberation, and both single and multiple speakers are considered to be extracted. More specially, convolutional neural network (CNN) based models are presented for this challenging speech enhancement task due to its parameter efficiency and state-of-the-art performance. After describing one-stage and multi-stage CNN-based models, numerous experiments are conducted to show the advantage and disadvantage when applying them to extract one desired speaker and multiple desired speakers. This study reveals that CNN-based models can achieve high performance when there is only one desired speaker to be extracted, while their performance may degrade a lot for multiple desired 66 speakers. Some potential strategies are discussed on improving the performance of extracting multiple desired speakers and future research directions are outlined finally.

Read the paper · More papers on PaperTik