On Pronunciations inWiktionary: Extraction and Experiments on Multilingual Syllabification and Stress Prediction
Winston M.C. Wu, David Yarowsky · 2021
We constructed parsers for five non-English editions of Wiktionary, which combined with pronunciations from the English edition, comprises over 5.3 million IPA pronunciations, the largest pronunciation lexicon of its kind.This dataset is a unique comparable corpus of IPA pronunciations annotated from multiple sources.We analyze the dataset, noting the presence of machine-generated pronunciations.We develop a novel visualization method to quantify syllabification.We experiment on the new combined task of multilingual IPA syllabification and stress prediction, finding that training a massively multilingual neural sequence-to-sequence model with copy attention can improve performance on both high-and low-resource languages, and multitask training on stress prediction helps with syllabification.