Language modeling using efficient best-first bottom-up parsing

Keith Hall, Mark S. Johnson · 2004

In this paper we present a two-stage best-first bottom-up word-lattice parser which we use as a language model for speech recognition. The parser works by using a "figure of merit" that selects lattice paths while simultaneously selecting syntactic category edges for parsing. Additionally, we introduce a modified version of the inside-outside algorithm used as a pruning stage between syntactic context-free parsing and lexicalized context-dependent parsing. We report our results in terms of word error rate on the HUB-1 word-lattices and compare these results to other syntactic language modeling techniques.

Read the paper · More papers on PaperTik