Variable-length category-based n-grams for language modelling
TR Niesler, PC Woodland · Cambridge University Engineering Department Publications Database · 1995
This report concerns the theoretical development and subsequent evaluation of n-gram language models based on word categories. In particular, part-of-speech word classifications have been employed as a means of incorporating significant amounts of a-priori grammatical information into the model. The utilisation of categories diminishes the problem of data sparseness which plagues conventional word-based n-gram approaches, and therefore yields a fundamentally more compact model. Furthermore, it allows the use of larger n, and a strategy by means of which successively longer n-grams are selectively added to the model according to a cross-validation likelihood criterion is proposed. This enables the model compactness to be maintained while allowing longer range effects to be modelled where they benefit performance. The language modelling approach was applied to the LOB corpus in order to assess its effectiveness. When compared with models of corresponding complexity constructed according...