Speech Enhancement Using NMF based on Hierarchical Deep Neural Networks with Joint Learning
Mohammad Mahdi Mirjalili, Sanaz Seyedin · 2020
In this paper, we propose a novel method which includes autoencoder and deep neural networks (DNNs) in a hierarchal structure for speech enhancement. In this method, at first, two parallel autoencoders are employed to obtain the nonnegative matrix factorization (NMF) parameters of speech and noise in a nonlinear mapping. After that, by using the spectrum of noisy speech as the input of the encoder portion of the autoencoders, the outputs of the encoders are calculated and utilized as the input of DNNs in the next hierarchies to further enhance the speech spectrum more efficiently. Also, the last three hierarchies including the decoder portion of autoencoders and the DNN will be trained in a joint learning scenario to improve the results. The proposed method is evaluated on TIMIT corpus with perceptual-evaluation-of-speech-quality (PESQ) and frequency-weighted-segmental-signal-to-noise-ratio (fwSNRseg) criterions. The obtained results show a significant improvement in both seen and unseen noises compared to baselines.