The use of meta-HMM in multistream HMM training for automatic speech recognition
Christian J. Wellekens, Jussi Kangasharju, Cedric Milesi · 1998
Among the different attempts to improve recognition scores and robustness to noise, the recognition of parallel streams of data each one representing partial information on the test signal and the fusion of the decisions have received a great deal of interest. The problem of training such models taking recombination constraints at the level of speech-subunits has not yet been rigorously addressed. This paper shows how equivalence with an extended meta-HMM solves the problem and how reestimation formulas have to be applied to guarantee equivalence between the multistream model and the meta-HMM. 1. INTRODUCTION Since speech is a non-stationary signal, it is analyzed over small time windows where local stationarity is assumed. A typical width for these windows is 10 ms and different kinds of analysis can be conducted on them. The first standard approach is a harmonic analysis that could be obtained using either filter banks or FFT transforms. After elimination of the fundamental period ...