Speech Enhancement Using Convolutional Denoising Autoencoder
Shaikh Akib Shahriyar, M. A. H. Akhand, Nazmul Haque Siddique, Tetsuya Shimamura · 2019
Speech signals are complex in nature with respect to other forms of communication media such as text or image. Different forms of noises (e.g., additive noise, channel noise, babble noise) interfere with the speech signals and drastically hamper the quality of the speech in the noisy speech signals. Enhancement of speech signals is a daunting task considering multiple forms of noises while denoising speech signals. Certain analog noise eliminator models have been studied over the years for this purpose. Researchers have also delved into some deep learning techniques to enhance speech signals. In this study, a speech enhancement system is investigated using Convolutional Denoising Autoencoder (CDAE). It takes advantages from the 2D structured inputs of the features extracted from speech signals and also considers the local temporal relationship among the features. In the proposed system, CDAE is trained considering features from noisy speech signal as input and clean speech features as desired output. The proposed CDAE based method has been tested on a benchmark dataset, called Speech Command Dataset, and attained 80% similarity between denoised speech and actual clean speech. The proposed system achieved perceptual evaluation of speech quality (PESQ) value of 2.43 which outperformed other related existing methods.