Discriminative training for improving letter-to-sound conversion performance
Yi‐Ning Chen, Peng Liu, Jiali You, Frank K. Soong · IEEE International Conference on Acoustics Speech and Signal Processing · 2008
In this paper, we propose to use discriminative training (DT) for improving letter-to-sound (LTS) conversion performance. LTS is a critical component in both ASR and TTS for predicting the correct pronunciation of a word not included in the lexicon. For TTS applications, predicting the proper pronunciation of an out-of-vocabulary person/place name, especially a name with foreign origin can be challenging. We utilize discriminative training, which has been successfully used in speech recognition, to sharpen the baseline N-grams of grapheme-phoneme pairs. We address the problem in a unified framework of discriminative training. Two criteria, maximum mutual information (MMI) and minimum phoneme error (MPE), are investigated. Experimental results show that DT yields a small (3.8-4.6% relative) but consistent error reduction across all databases tested. In addition, we observe that by pinpointing the local errors in a finer resolution, we can obtain a better discriminative model.