Suffix Trees as Language Models
Casey Kennington, Martin Kay, Annemarie Friedrich · 2012
Suffix trees are data structures that can be used to index a corpus.In this paper, we explore how some properties of suffix trees naturally provide the functionality of an n-gram language model with variable n.We explain how we leverage these properties of suffix trees for our Suffix Tree Language Model (STLM) implementation and explain how a suffix tree implicitly contains the data needed for n-gram language modeling.We also discuss the kinds of smoothing techniques appropriate to such a model.We then show that our STLM implementation is competitive when compared to the state-of-the-art language model SRILM (Stolcke, 2002) in statistical machine translation (SMT) experiments.