Distant-talking Speech Recognition Based on Multi-objective Learning using Phase and Magnitude-based Feature
Dongbo Li, Longbiao Wang, Jianwu Dang, Meng Ge, Haotian Guan · 2018
Deep neural network for speech enhancement is an increasingly interesting topic. In this paper, we propose a multi-objective learning method to using the amplitude and phase information for reverberant speech recognition. In previous studies, some researches found that phase information is important for human speech recognition, but phase information is ignored for almost front-end of speech recognition. To address this problem, this paper proposes using a multi-objective neural network method to optimize speech enhancement and feature enhancement simultaneously. For phase information, Modied Group Delay Cepstral Coefcients (MGDCC) and Phase Domain Source-Filter separation based Vocal Tract (PBSFVT) are used. In this paper, we use the data set of Reverb Challenge 2014 to evaluate proposed method on distant-talking speech recognition. The Word Error Rate (WER) of speech recognition was reduced from 26.57% of traditional deep neural work based dereverberation using magnitude feature, to 23.34% of the proposed method and the relative error reduction rate is 12.15%.