Phase-sensitive real-time capable speech enhancement under voiced-unvoiced uncertainty
Martin Krawczyk, Robert Rehr, Timo Gerkmann · European Signal Processing Conference · 2013
In many short-time Fourier transform (STFT)-based single channel speech enhancement algorithms, the clean speech spectral amplitude is estimated from a noisy observation to suppress additive noise. For the estimation, only the noisy amplitudes and functions thereof, like the a priori or a posteriori signal-to-noise ratio (SNR), are utilized. Information about the clean speech spectral phase is mostly not employed. In this work we present a comprehensive speech enhancement setup that combines phase-sensitive and phase-insensitive amplitude estimation, improving the perceptual speech quality of the enhanced signal in terms of PESQ compared to phase-insensitive amplitude estimation alone. The proposed algorithm is real-time capable in the sense that it is implemented in a causal block-wise manner and the computational complexity is feasible.