DATA - DRIVEN GENERATION OF PRONUNCIATION DICTIONARIES DISCUSSION OF EXPERIMENTAL RESULTS IN THE GERMAN VERBMOBIL PROJECT -

Matthias Eichner, Matthias Wolff · 2000

In the framework of the German Verbmobil project we developed a procedure for the automatic, data - driven generation of pronunciation dictionaries for speech recognition systems. In most recognizers only simple dictionaries containing the canonical pronunciation form are used. They represent the correct pronunciation, but in most cases the canonical pronunciation does not match the actual realization of the word. To solve this problem we chose an approach to derive pronunciation variants automatically from a speech database. The training algorithm bases on a canonical dictionary which is compiled into a graph representation in a first stage. Pronunciation variants are then learned from a training sample consisting of speech signal and its orthographic transcription. In this paper we will focus on the experimental results obtained in the Verbmobil framework and introduce methods to evaluate pronunciation dictionaries generated by the training procedure.

Read the paper · More papers on PaperTik