Discriminative reranking for LVCSR leveraging invariant structure

Masayuki Suzuki, Gakuto Kurata, Masafumi Nishimura, Nobuaki Minematsu · 2012

An invariant structure is one of the long-span acoustic represen-tations, where acoustic variations caused by non-linguistic fac-tors are effectively removed from speech. We present in this pa-per a new method to leverage the invariant structures as features of discriminative reranking for Large Vocabulary Continuous Speech Recognition (LVCSR). First we use a traditional HMM-based LVCSR system to get a list of N-best candidates with phone alignments and construct an invariant structure for each candidate using its phone alignment. Here, the invariant struc-ture is composed of lengths between every two phonemes in the candidate. Then we estimate a score of each phoneme-pair in the invariant structure, and rerank the N-best candidates using a weighted sum of the phoneme-pair scores, where the weights are trained discriminatively by averaged perceptron. Experi-mental results show a relative CER improvement of 6.69 % over the baseline HMM-based LVCSR system. Index Terms: Invariant Structure, LVCSR, Discriminative reranking 1.

Read the paper · More papers on PaperTik