The TUM system for the REVERB Challenge: Recognition of Reverberated Speech using Multi-Channel Correlation Shaping Dereverberation and BLSTM Recurrent Neural Networks

Jürgen T. Geiger, Erik Marchi, Felix Johannes Weninger, Björn Wolfgang Schuller, Gerhard Rigoll · 2014

This paper presents the TUM contribution to the 2014 REVERB Challenge: we describe a system for robust recognition of reverberated speech. In addition to an HMM-GMM recogniser, we use bidirectional long short-term memory (LSTM) recurrent neural networks. These networks can exploit long-range temporal context by using memory cells in the hidden units, which increases the robustness against reverberation. The LSTM is trained with phonemes as targets, and the predictions are converted into observation likelihoods and used as an acoustic model. Furthermore, we apply a dereverberation method called correlation shaping on the 8-channel recordings. This method applies a reduction of the long-term correlation energy in the received reverberant speech. The linear prediction residual, which generally contains information about reverberation, is processed to suppress the long-term correlation that is mostly due to the speaker-to-receiver impulse response. Using dereverberation as a front-end of the GMM in combination with the LSTM predictions leads to substantial improvements of the word error rate, achieving 11.19 % (relative improvement of about 35 %) and 28.13 % (improvement of about 30 %) with simulated and real data test sets, respectively. In the single-channel case, in which the dereverberation technique can not be applied, improvements of about 20 % (for simulated data) and 7 % (for real data) are obtained with the LSTM technique. Index Terms — Dereverberation, BLSTM recurrent neural networks, multi-channel correlation shaping 1.

Read the paper · More papers on PaperTik