Automatic phonetic transcription of words based on sparse data

Maria Klara Wolters, Antal P. J. van den Bosch · Research portal (Tilburg University) · 1997

The relation between the orthography and the phonology of a language has traditionally been modelled by hand--crafted rule sets. Machine-learning (ML) approaches offer a means to gather this knowledge automatically. Problems arise when the training material is sparse. Generalising from sparse data is a well-known problem for many ML algorithms. We present experiments in which connectionist, instance--based, and decision--tree learning algorithms are applied to a small corpus of Scottish Gaelic. instance-based learning in the ib1-ig algorithm yields the best generalisation performance, and that most algorithms tested perform tolerably well. Given the availability of a lexicon, even if it is sparse, ML is a valuable and efficient tool for automatic phonetic transcription of written text. 1 The Problem Experienced readers can read text aloud fluently and without pronunciation errors. But can we simulate this performance on a computer? This question is especially relevant for text--to--s...

Read the paper · More papers on PaperTik