N-gram Chinese Characters Counting for Huge Text Corpora
Yu Yi · 2014
Counting N-gram Chinese characters of huge text corpora is a challenge for Chinese information processing and Cici was developed to count huge Chinese text corpora efficiently.We found that the number of different Chinese strings is maximal when the length of strings is 6,and the number of strings can be estimated by the average length of sentences.Since most Chinese strings appear no more than 10times in the corpora,the N-gram characters are stored in 13separate files according to their frequency,and only highly used strings are sorted.This strategy speeds up the accounting process dramatically.Due to the limited physical memory,huge Chinese text corpora have to be divided into many blocks,whose size is suggested to be 20MB.Every block is counted separately,and then the block statistic results are merged together.We implemented the algorithm of accounting huge corpora efficiently in personal computer.