Frame level likelihood transformations for ASR and utterance verification
Konstantin Markov, Satoshi Nakamura · 2000
In most of the current speech recognition systems based on HMM, existing decoding and utterance veri cation methods make use of state output likelihood as a measure of the acoustic match between the input data and the acoustic models. In this paper, we present a new and more generalized approach to the formation of the acoustic match score. The essence of this approach is to transform the likelihood of each acoustic vector with respect to any particular HMM state according to some non-linear function. We have investigated two types of such transformation functions. The rst one, performs likelihood normalization, and the second one transforms likelihoods into exponentially ordered weights. The transformed likelihoods, as new acoustic scores, are used further for decoding, recognition and veri cation instead of the conventional likelihoods. In our evaluation experiments we used TIMIT database for phoneme recognition and veri cation and a database of 710 speakers and a total of 4252 distinct words, for isolated word recognition and veri cation. The results we achieved show that the transformed likelihood scores, in average, increase slightly the recognition accuracy and reduce the veri cation error rates up to 30%.