A Deep Neural Framework to Estimate Spectral Phase and Magnitude In Application to Single-Channel Speech Enhancement

Junaid Mushtaq, Jawad Ali, Sher Muhammad, Iqra Batool, Nasir Saleem · 2023

Many speech enhancement (SE) methods that are already in use only estimate the magnitude spectrum and leave the phase spectrum alone. In this paper, we suggest a way for single-channel SE (SCSE) to estimate the spectrum's phase and magnitude. A phase compensation process is applied to noise-contaminated speech and estimates the phase spectrum. Adaptive power law transformation moves the power from the energy-rich voiced segments of the contaminated signal to the weak-energy unvoiced segments, while the entire signal's energy stays same. The energy-redistributed noisy speech signals are used to train the attention-gated long short-term memory (LSTM) recurrent neural network in order to estimate the magnitude of the ideal ratio mask (IRM). Attention-gated LSTM reduces the computational load of conventional LSTM by replacing a forgetting gate without performance loss. During speech synthesis, the speech signals are reconstructed by putting together the estimated phase and magnitude spectra. The test results show that the proposed method did better than the other methods and made speech clearer and better in noisy environments.

Read the paper · More papers on PaperTik