Data-Dependent Ensemble of Magnitude Spectrum Predictions for Single Channel Speech Enhancement
Pasi E. Pertila · 2019
Applying a predicted time-frequency mask to the noisy speech spectrogram and directly predicting the clean speech magnitude spectrogram are two common deep learning-based speech enhancement approaches. Ensemble techniques such as averaging and neural network-based fusion of magnitude spectra obtained with these two approaches have been shown to improve the objective perceptual quality of speech using synthetic mixtures of data. This work generalizes the averaging ensemble approach by proposing neural network layers to predict time-frequency varying weights for the combination of the two magnitude spectra obtained by time-frequency masking and by direct prediction. In order to combine the best individual magnitude spectrum estimates, the proposed weight prediction layers are trained after the time-frequency mask and magnitude spectrum networks layers have been separately trained for their corresponding objectives and their weights have been fixed. Using the publicly available CHiME3-challenge data, which consists of both simulated and real speech recordings in everyday environments with noise and interference, the proposed approach leads to significantly higher noise suppression in terms of segmental source-to-distortion ratio over the alternative approaches. In addition, the approach achieves similar improvements in the average objective instrumentally measured intelligibility scores with respect to the best achieved scores.