Machine learning of word pronunciation: the case against abstraction
Bertjan Busser, Walter Daelemans, Antal van den Bosch · 1999
An adequate approach to speech translation for small to medium sized tasks is the use of subsequential trans-ducers —a finite state model — as language model for a speech recognizer. These transducers can be automati-cally trained from sample corpora. The use of manually defined categories improves the training of the subsequential transducers when the avail-able data are scarce. These categories depend on the source and target languages we want to translate. We introduce an automatic approach to derive cate-gories that can be used in training subsequential transduc-ers. This approach extends monolingual word clustering methods to the bilingual case using alignments obtained from statistical models. Experimental results indicate that the models trained with these categories have lower trans-lation errors. 1