Improvement of Speech Residuals for Speech Enhancement
Samy Elshamy, Tim Fingscheidt · 2019
In this work we present two novel methods to improve speech residuals for speech enhancement. A deep neural network is used to enhance residual signals in the cepstral domain, thereby exceeding a former cepstral excitation manipulation (CEM) approach in different ways: One variant provides higher speech component quality by 0.1 PESQ points in low-SNR conditions, while another one delivers substantially higher noise attenuation by 1.5 dB, without loss of speech component quality or speech intelligibility. Compared to traditional speech enhancement based on the decision-directed (DD) a priori SNR estimation, a gain of even up to 3.5 dB noise attenuation is obtained. A semi-formal comparative category rating (CCR) subjective listening test confirms the superiority of the proposed approach over DD by 0.25 CMOS points (or even by 0.48 if two outlier subjects are not considered).