A Statistical Approach to Extract Chinese Chunk Candidates from Large Corpora

Tianshun Yao · 2003

The extraction of Chunk candidates from real corpora is one of the fundamental tasks of building example-based machine translation model. This paper presents a statistical approach to extract Chinese chunk candidates from large monolingual corpora. The rst step is to extract large N-grams (up to 20-gram) from raw corpus. Then two newly proposed Fast Statistical Substring Reduction (FSSR) algorithms can be applied to the initial N-gram set to remove some unnecessary N-grams using their frequency information. The two algorithms are ecien t (both have a time complexity of O(n)) and can eectiv ely reduce the size of N-gram set up to 50%. Finally, mutual information is used to obtain chunk candidates from reduced N-gram set. Perhaps the biggest contribution of this paper is that it is the rst time to apply Fast Statistical Substring Reduction algorithm to large corpora and demonstrate the eectiv eness and eciency of this algorithm which, in our hope, will shed new light on large scale corpus oriented research. Experiments on three corpora with dieren t sizes show that this method can extract chunk candidates from corpora of giga bytes ecien tly under current computational power. We get an extraction accuracy of 86.3% from People Daily 2000 news corpus.

Read the paper · More papers on PaperTik