Development of a speech enhancement system using deep neural networks

Daniel Montoro Rodríguez · 2020

In the last years, deep neural networks have become an important tool in speech technologies, yielding notable advances in the fields of speaker and speech recognition and speech synthesis. In this project a design is proposed for a deep neural network for speech enhancement, that is capable of reducing the level of noise in speech recordings taken in real world scenarios such as a public transportation or a cafeteria. The proposed design is intended to reduce the high requirements of computing power of other models that make up the state-of-the-art in audio processing with deep neural networks, as well as the complexity of their architectures. In addition, a loss function is introduced that is based on a measure highly correlated to the perceived quality of speech, and the effect of using it during training is analyzed. The performance of the model is evaluated using objective measures of speech quality.

Read the paper · More papers on PaperTik