A Posterior-Based Multistream Formulation for G2P Conversion

Marzieh Razavi, Mathew Magimai.-Doss · IEEE Signal Processing Letters · 2017

In the literature, a number of approaches have been proposed for learning grapheme-to-phoneme (G2P) relationship and inferring pronunciations. In this letter, we present a novel multistream framework for G2P conversion, where various machine learning techniques providing different estimates of probability of phonemes given graphemes can be effectively combined during pronunciation inference. More precisely, analogous to multistream automatic speech recognition, the framework involves obtaining different streams of estimates of probability of phonemes given graphemes, combining them based on probability combination rules, and inferring pronunciations by decoding the probabilities resulting after combination. We demonstrate the potential of the proposed approach by combining probabilities estimated by the state-of-the-art conditional random field-based G2P conversion approach and acoustic data-driven G2P conversion approach in the Kullback-Leibler-divergence-based hidden Markov model framework on the PhoneBook 600-word task.

Read the paper · More papers on PaperTik