DEEP LEARNING METHODS FOR IMPROVING THE PERCEPTUAL QUALITY OF NOISY AND REVERBERANT SPEECH

Donald S. Williamson · OhioLink ETD Center (Ohio Library and Information Network) · 2016

Speech is a vital form of human communication and it is important for many real-world applications.Voice commands are used to interface with electronic devices and hearing-impaired individuals use hearing aids to understand speech better.In realistic environments, background noise and reverberation are present, resulting in performance degradation.For this reason, it is crucial that speech is separated from interference.Many speech separation approaches have been proposed, but there is a considerable need to produce speech estimates that are both intelligible and high quality, especially at low signal-to-noise ratios (SNRs).Time-frequency (T-F) masking and model-based separation are two common ways to extract speech in a noisy observation.T-F masking involves the estimation of an oracle mask, which can be accomplished using supervised learning.Deep neural networks (DNN) are well suited for T-F mask estimation due to their ability to learn mappings from noisy observations to a desired target.Likewise, model-based separation is suitable due to its ability to represent the spectral structure of speech.This dissertation presents work that develops speech separation systems using combinations of T-F masking, DNNs, and model-based reconstruction.The aim of each system is to improve the perceptual quality of the speech estimates.Ideal binary mask (IBM) estimation has shown success in improving the intelligibility of separated speech, but it often results in poor quality due to estimation errors ix

Read the paper · More papers on PaperTik