Online discriminative training for grapheme-to-phoneme conversion

Sittichai Jiampojamarn, Grzegorz Kondrak · 2009

We present an online discriminative training approach to grapheme-to-phoneme (g2p) conversion. We employ a manyto-many alignment between graphemes and phonemes, which overcomes the limitations of widely used one-to-one alignments. The discriminative structure-prediction model incorporates input segmentation, phoneme prediction, and sequence modeling in a unified dynamic programming framework. The learning model is able to capture both local context features in inputs, as well as non-local dependency features in sequence outputs. Experimental results show that our system surpasses the state-of-the-art on several data sets. Index Terms: grapheme-to-phoneme conversion, speech synthesis, discriminative training

Read the paper · More papers on PaperTik