Minimum Bayes risk training of CTC acoustic models in maximum a posteriori based decoding framework
Naoyuki Kanda, Xugang Lu, Hisashi Kawai · 2017
When using connectionist temporal classification (CTC) based acoustic models (AMs) for large vocabulary continuous speech recognition (LVCSR), most previous studies have used a naive interpolation of the CTC-AM score and an additional language model score, although there is no theoretical justification for such an approach. On the other hand, we recently proposed a theoretically more sound decoding framework for CTC-AM called maximum a posteriori (MAP)-based decoding. Although the superiority of the MAP-based decoding framework with CTC-AM has been demonstrated, the effect of additional minimum Bayes risk (MBR) training in the MAP-based decoding framework has not been investigated. In this paper, we report the results of various experiments that examine the effect of MBR training on CTC-AM by comparing two decoding frameworks. Our experiments with English and Japanese LVCSR tasks reveal that the MAP-based decoding framework is superior to the interpolation-based framework, even after the MBR training. In addition, by using about 600 h of training data, we show that the size of the training dataset is a critical factor in achieving good results under CTC-AM.