A Generalized Framework for Hierarchical Word Sequence Language Model
Xiaoyi Wu, Kevin Duh, Yūji Matsumoto · Institutional Repositories DataBase (IRDB) · 2016
Language modeling is a fundamental research problem that has wide application for many NLP tasks.For estimating probabilities of natural language sentences, most research on language modeling use n-gram based approaches to factor sentence probabilities.However, the assumption under n-gram models is not robust enough to cope with the data sparseness problem, which affects the final performance of language models.At the point, Hierarchical Word Sequence (abbreviated as HWS) language models can be viewed as an effective alternative to normal n-gram method.In this paper, we generalize HWS models into a framework, where different assumptions can be adopted to rearrange word sequences in a totally unsupervised fashion, which greatly increases the expandability of HWS models.For evaluation, we compare our rearranged word sequences to conventional n-gram word sequences.Both intrinsic and extrinsic experiments verify that our framework can achieve better performance, proving that our method can be considered as a better alternative for ngram language models.