The Virtual Corpus Approach to Deriving Ngram Statistics from Large Scale Corpora

Chunyu Kit, Yorick Alexander Wilks · 2002

This paper reports our implementation of the Virtual Corpus approach to deriving ngram statistics for ngrams of any length from large-scale corpora based on the suffix array data structure. In order to enable the VC to accommodate corpora with a vocabulary of different size, we first convert corpus tokens into integer codes. To accelerate the processing, we employ a bucket-radixsort for sorting the VC indices (or pointers, each of which represents a sequence of corpus tokens from its position to the end of the corpus). The time complexity of the sorting algorithm is of O(N log n N) in token code comparisons. 1 Introduction It is known that ngram model is the simplest, most durable and successful statistical model for many natural language processing (NLP) and speech processing applications, e.g., as in [9, 2] and many others. Many high performance part-of-speech tagging systems, e.g., [10, 4, 3] and others, are based on ngram statistics. Multigram model [1, 6, 7] has become a practic...

Read the paper · More papers on PaperTik