Phonetic Normalization for Machine Translation of User Generated Content

José Carlos Rosales Núñez, Djamé Seddah, Guillaume Wisniewski · 2019

We present an approach to correct noisy User Generated Content (UGC) in French aiming to produce a pre-processing pipeline to improve Machine Translation for this kind of noncanonical corpora.Our approach leverages the fact that some errors are due to confusion induced by words with similar pronunciation which can be corrected using a phonetic lookup table to produce normalization candidates.We rely on a character-based neural model phonetizer to produce IPA pronunciations of words and a similarity metric based on the IPA representation of words that allow us to identify words with similar pronunciation.These potential corrections are then encoded in a lattice and ranked using a language model to output the most probable corrected phrase.Compared to other phonetizers, our method boosts a Transformer-based machine translation system on UGC.

Read the paper · More papers on PaperTik