Recognition performance of a large-scale dependency grammar language model
Adam L. Berger, Harry Printz · 1998
In this paper, we describe a large-scale investigation of dependency grammar language models. Our work includes several significant departures from earlier studies, notably a larger training corpus, improved model structure, different feature types, new feature selection methods, and more coherent training and test data. We report word error rate (wer) results of a speech recognition experiment, in which we used these models to rescore the output of the IBM speech recognition system. 1. INTRODUCTION One promising idea for advancing statistical language modeling is based upon dependency grammars, as reported in [4]. This work has the appeal of integrating, via the maximum entropy / minimum divergence (memd) technique, information from both syntax and ngrams. Moreover, unlike methods based upon context free grammars, grammatical information enters through the words themselves, rather than via abstract constituents like parse-tree node labels. A preliminary investigation of this idea, ju...