Posterior-Based Multi-Stream Formulation To Combine Multiple Grapheme-to-Phoneme Conversion Techniques

Marzieh Razavi, Mathew Magimai.-Doss · Infoscience (Ecole Polytechnique Fédérale de Lausanne) · 2015

In the literature, a number of approaches have been proposed for learning grapheme-to-phoneme (G2P) relationship and inferring pronunciations. The paper presents a multi-stream framework where different G2P relationship learning techniques can be effectively combined during pronunciation inference. Specifically, analogous to multi-stream automatic speech recognition in the literature, the framework involves (a) obtaining different streams of estimates of probability of phonemes given graphemes; (b) combining them based on probability combination rules; and (c) inferring pronunciations by decoding the probabilities resulting after combination. We demonstrate the potential of the proposed approach by combining state-of-the-art CRF-based G2P conversion approach and acoustic data-driven G2P conversion approach in the Kullback-Leibler divergence based HMM framework on the PhoneBook 600 words task.

Read the paper · More papers on PaperTik