Towards Multiple Pronunciation Generation in Acoustic G2P Conversion Framework

Marzieh Razavi, Ramya Rasipuram, Mathew Magimai.-Doss · Infoscience (Ecole Polytechnique Fédérale de Lausanne) · 2015

Recently an acoustic data-driven grapheme-to-phoneme (G2P) conversion approach has been proposed in which the G2P relationship is learned through acoustic data, and the learned relationship together with the orthographic transcription of the word are used to infer pronunciations.This paper extends the acoustic G2P conversion approach and proposes two methods for generating multiple pronunciations to better handle pronunciation variations.In the first method, multiple pronunciations are generated by using different cost functions at the learning stage to possibly capture different G2P relationships.The second method generates multiple pronunciations at the inference stage through N-best decoding.Our experimental studies on Phonebook task in English show that (a) the first method yields lower average number of pronunciations per word than the second method; and (b) both methods, without pronunciation selection or pruning, lead to improvements in the performance at the pronunciation level as well as the speech recognition level.

Read the paper · More papers on PaperTik