Connectionist Speech Recognition: A Hybrid Approach

Hervé A. Bourlard, Nelson H. Morgan · Infoscience (Ecole Polytechnique Fédérale de Lausanne) · 1993

BACKGROUNDList of Tables 5.1 Comparison of the recognition error rates (I=insertions, S=substitutions, D=deletions) obtained with 1-state phonemic HMMs with discrete emission probabilities, Gaussian emission probabilities, and outputs of a contextual MLP. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .94 6.1 Phonetic classification rates at the frame level obtained by standard approaches."Full Gaussian" refers to the case of one Gaussian with full covariance matrix per phoneme, "MLE" refers to the case of one discrete likelihood density per phoneme estimated by counting, and "MAP" refers to the case of one discrete posterior probability density estimated by counting. . . . . . . . . . . . . . . . . . . . . . .111 6.2 Phonetic classification rates at the frame level obtained from different MLPs, compared with MLE."MLPa b-cd" stands for an MLP with blocs (width of context) of (binary) input units, hidden units and output units.The size of the output layer was kept fixed at 50 units, corresponding to the 50 phonemes to be recognized. . .113 6.3 Phonetic classification rates at the frame level obtained from contextual MLPs, compared with standard likelihoods (MLE) and a posteriori probabilities (MAP).R represents the parametrization ratio, i.e., the number of parameters divided by the number of training patterns. . . . . . . . . .116 6.4 Phonetic classification rates at the frame level on SPI-COS obtained from MLPs with linear and nonlinear outputs. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .117 7.1 Word recognition rate on SPICOS database (speaker m003) for different hybrid HMM/MLP approaches (MLP = no division of output values by priors, MLP/priors = division by priors) compared with standard HMMs trained with MLE criterion. . . . . . . . . . .

Read the paper · More papers on PaperTik