A Novel Training Target of DNN Used for Casa-Based Speech Enhancement

Feng Bao, Waleed Habib Abdulla · 2018

Efficient training target plays a vital role in any training model for speech enhancement. In the Computational Auditory Scene Analysis method based on Deep Neural Networks, the ideal ratio mask or square-root ideal ratio mask is usually considered as the effective training target. In this paper, we propose a novel training target method for speech enhancement. This new training target takes into consideration the inter channel correlations of the power spectra of the noisy speech, clean speech and noise to more efficiently retain speech components and mask the noise components. Additionally, the channel-weight contour based on the equal loudness hearing attribute is introduced to revise the training target in each Gammatone channel to make the resynthesized signal more suitable for hearing characteristics. Moreover, the training target is smoothed to further improve the accuracy. Experiments show that using the proposed training target achieves better performances in terms of speech quality and intelligibility.

Read the paper · More papers on PaperTik