Performance comparison of component algorithms for the phonemicization of orthography

Jared Bernstein, Larry Nessly · 1981

A system for converting English text into synthetic speech can be divided into two processes that operate in series: I) a text-to-phoneme converter, and 2) a phonemic-input speech synthesizer.The conversion of orthographic text into a phonemic form may itself comprise several processes in series, for instance, formatting text to expand abbreviations and non-alphabetic expressions, parsing and word class determination, segmental phonemicization of words, word and clause level stress assignment, word internal and word boundary allophonic adjustments, and duration and fundamental frequency settings for phonological units.Comparing the accuracy of different algorithms for text-to-phoneme conversion is often difficult because authors measure and report system performance in incommensurable ways.Furthermore, comparison of the output speech from two complete systems may not always provide a good test of the performance of the corresponding component algorithms in the two systems, because radical performance differences in other components can obscure small differences in the components of interest.The only reported direct comparison of two complete text-to-speech systems (MITALK and TSI's TTS-X) was conducted by Bernsteln and Pisonl (1980).This paper reports one study that compared two algorithms for automatic segmental phonemlcization of words, and a second study that compared two algorithms for automatic assignment of lexical stress.

Read the paper · More papers on PaperTik