Grapheme-to-Phoneme Models for (Almost) Any Language

Aliya Deri, Kevin K. Knight · 2016

Grapheme-to-phoneme (g2p) models are rarely available in low-resource languages, as the creation of training and evaluation data is expensive and time-consuming.We use Wiktionary to obtain more than 650k word-pronunciation pairs in more than 500 languages.We then develop phoneme and language distance metrics based on phonological and linguistic knowledge; applying those, we adapt g2p models for highresource languages to create models for related low-resource languages.We provide results for models for 229 adapted languages.

Read the paper · More papers on PaperTik