Improvement of mask-based speech source separation using DNN
Ge Zhan, Zhaoqiong Huang, Dongwen Ying, Jielin Pan, Yonghong Yan · 2016
The speech mask is widely used to separate multiple speech sources, wherein the time-frequency bins are classified into clusters that correspond to each source. For each source, the separated signal consists of the components on TF bins that are dominated by this source, whereas the components on the remaining bins are completely masked. Most separation methods ignored the masked components. In fact, the masked components may contain some useful information, and the mask-based speech source separation can be improved by reconstructing the masked components. This paper proposes a post-processing method to reconstruct the masked frequency components through a deep neural network (DNN). We construct a regression from the reliable frequency components to the masked components. After the masked-based separation, the reliable components are kept unchanged, and the masked components are reconstructed by the outputs of DNN. Experimental results confirmed that the proposed method significantly improved the mask-based separation, and that the masked components are still useful to the speech quality.