On using parameterized multi-channel non-causal Wiener filter-adapted convolutional neural networks for distant speech recognition
Jeehye Lee, Joon‐Hyuk Chang, Jinho Sohn · 2016
Recently, the convolutional neural network (CNN) with multiple microphones was proposed to use the delay-sum (DS) beamformer for distant speech recognition (DSR) and compared to the direct use of multiple acoustic channels as a parallel input to the CNN [1]. We explore the parameterized multi-channel non-causal Wiener filter (PMWF) as the front-end to train the CNN, which is applied to acoustic modeling for DSR. For this, we first present a concise description of the basic PMWF as well as its advantages and then explain how to organize the PMWF into the CNN with a novel architecture. Experimental results on the TIMIT dataset show that the proposed PMWF-based CNN approach outperforms the cross-channel CNN and the DS beamformer when evaluating the word error rate (WER) in various DSR environments.