Low memory decision tree method for text-to-phoneme mapping
J. Suontausta, J. Tian · 2004
Pronunciation models are commonly used in automatic speech recognition (ASR) as well as in text-to-speech (TTS) applications. Decision tree (DT) and neural network (NN) methods have been used for modeling the languages with irregular pronunciation. The DT based methods are usually more accurate and therefore they provide better recognition accuracy than the NN based methods. The main drawback of the DT based methods is their large memory footprint. In the paper, we propose three methods, i.e.: (1) a clipping approach to reduce the pronunciation variability in the aligned dictionary; (2) a revised DT structure; and (3) a Huffman coding scheme for efficient DT parameter storing that can jointly be applied to minimize the memory footprint of the DT models. The results obtained in the simulation experiments indicate that the memory requirements of the DT models can significantly be reduced without degrading the mapping accuracy. The applicability of the approach is also verified in the speech recognition experiments.