Speech Enhancement Using RNN models and Ideal Exponential Mask

Anirban Panda · 2018

This paper aims at Speech Enhancement using Mask- Approximation method. An Ideal Exponential Mask has been proposed to increase the accuracy of the model. The Mel Frequency Cepstral Coefficient is extracted from the speech signal as a feature, which is fed into different models as input. Considering speech as a sequential data, different Recurrent Neural Networks, such as Long Short Term memory and Gated Recurrent Unit, have been used to train the model. Later, considering that the speech signal depends on both the past and future sequence, Bidirectional models have been used. The Signal to Distortion Ratio and Short Time Object Intelligibility has been calculated to measure the accuracy of different models and different masks. A maximum Short Time Object Intelligibility of 0.87 has been achieved using the new mask.

Read the paper · More papers on PaperTik