Improvements of Acoustic Features for Speech Separation
Lujun Li, Chu-Xiong Qin, Dan Qu · 2016
Acoustic features chosen from the input utterances play crucial roles in speech separation.In this paper, we propose a novel complementary feature approach that performs speech separation by combining five promising features, including Gammatone filterbank power spectra (GF) and multi-resolution cochleagram (MRCG) proposed recently especially for speech separation, as a super-vector fed into deep neural network (DNN).Additionally, based on the complementary features, we do experiments with two DNN training strategies, which are restricted Boltzmann machine (RBM) pre-training and dropout combined with Rectified Linear Units (ReLU), to optimize the performance of DNN.The experiment results, obtained in IEEE and TIMIT corpora using four different noises at low SNR levels of 0dB and -5dB, indicate that complementary features and RBM model improve all evaluation metrics.By contrast, dropout combined with ReLU system specializes in noise suppression and objective intelligibility more.