Hierarchical tandem feature extraction
Sunil Sivadas, Hynek Heřmanský · IEEE International Conference on Acoustics Speech and Signal Processing · 2002
We present a hierarchical architecture for tandem acoustic modeling. In the tandem acoustic modeling paradigm a Multi Layer Perceptron (MLP) is discriminatively trained to estimate phoneme posterior probabilities on a labeled database. The outputs of the MLP after nonlinear transformation and whitening are used as features in a Gaussian Mixture Model (GMM) based recognizer. In this paper we replace the large monolithic MLP with hierarchies of MLP experts. We apply this approach on Speech in Noisy Environments (SPINE 1) evaluation conducted by the Naval Research Laboratory (NRL). We observe a reduction in word error rate of 30% with context-independent models and 5% WER with context-dependent models relative to PLP features.