Compression Methods by Code Mapping and Code Dividing for Chinese Dictionary Stored in a Double-Array Trie

Huidan Liu, Minghua Nuo, Longlong Ma, Jian Wu, Yeping He · International Joint Conference on Natural Language Processing · 2011

There is serious data sparseness problem in Chinese dictionary stored in a doublearray trie. This paper proposes six compression methods by code mapping and code dividing to make it more compact, and a metric called Resource Consumption Ratio is proposed to evaluate these methods. Under the proposed criteria, five of the six methods are better than the baseline. The best method maps the character code into its frequency order, and then divides it into two jump codes. It achieves a space usage reduction of 39.88% and takes only 0.20% time of the baseline on the construction while it takes 13.21% more time on the retrieval. As preprocessing methods, these methods can be used to reduce more space by combining to other compression method which improves the double-array structure itself.

Read the paper · More papers on PaperTik