Scattering transform inspired filterbank learning from raw speech for better acoustic modeling

A. Madhavaraj, A. G. Ramakrishnan · 2019

We propose a neural network architecture, which operates on the raw speech signal, where the first layer contains a series of 1D time-domain filters. The output of this layer is fed to the second layer, which is a bank of 2D-convolution filters that capture the spectro-temporal modulations in the speech signal. The outputs of these two layers are concatenated, normalized and then fed to a feed-forward neural network to predict the senone posteriors, which are used for ASR decoding. During the training of the neural network, we have employed different strategies, where the 1D and 2D filters are initialized with (a) Gabor filters and (b) random values and the filter coefficients are either (a) allowed to be updated along with the other affine transform parameters of the network or (b) fixed during training. ASR experiments are conducted on 160 hours of Tamil speech data and the proposed architecture gives an absolute improvement in word error rate (WER) of 1.35% and 1.21% with respect to the neural network models trained on mel-frequency cepstral coefficients and log-filterbank energy features, respectively. We have also compared the performances of various strategies for filter initialization and training and reported the WERs.

Read the paper · More papers on PaperTik