MONAURAL SPEECH SEPARATION USING A PHASE-AWARE DEEP DENOISING AUTO ENCODER
Donald S. Williamson · 2018
Traditional deep denoising autoencoders (DDAE) use magnitude domain features and training targets to separate speech from background noise. Phase enhancement, however, has recently been shown to improve perceptual and objective speech quality. We present an approach that uses a DDAE to estimate phase-aware training targets from phase-aware input features. This network is denoted as a phase-aware deep denoising autoencoder (paDDAE). The short-time Fourier transform (STFT) of noisy speech is the network input, and the network estimates a phase-aware time-frequency mask. The proposed approach is evaluated across multiple conditions, including various signal-to-noise ratios (SNRs), noise types, and speakers. The results show that the paDDAE offers improvements over traditional DDAEs in terms of objective speech quality and intelligibility.