Learning dependency transduction models from unannotated examples
Hiyan Alshawi, Shona Douglas · Philosophical Transactions of the Royal Society A Mathematical Physical and Engineering Sciences · 2000
We present a method for constructing a statistical machine translation system automatically from unannotated examples in a manner consistent with the principles of dependency grammar. The method involves learning a generative statistical model of paired dependency derivations of source and target sentences. Such a dependency transduction model consists of collections of weighted head transducers. Head transducers are finite–state machines with different formal properties from ‘standard’ finite–state transducers. When applied to machine translation, the acquired head transducers are applied ‘middle out’, efficiently converting source head words and dependents directly into their counterparts in the target language. We present experimental results on the accuracy of our models for English–Spanish and English–Japanese translation, the training examples being pairs of transcribed spontaneous utterances and their translations. A hierarchical decomposition of bi–language strings emerges from our training process; this decomposition may or may not correspond to familiar linguistic phrase structure. However, no explicit semantic representations are involved, suggesting an approach to language processing in which natural language itself is the semantic representation.