Modelling of Chinese Texts using Bi-character and Tri-character Chinese Words for Data Compression
Ghim Hwee Ong, Shell Ying Huang, Wing Teck Chong · National University of Singapore · 1994
Text compression operation can be separated into two parts. One part is modelling and the other is coding. Coding algorithms have been well developed. However, due to the different characteristics of the Chinese text files, compression based on the byte oriented models may not yield the best compression ratio. In this paper, a new model of Chinese texts using bi-character words and tri-character words in addition to single characters is presented and evaluated against the byte-based and the character based models. Compression is carried out by adaptive Huffman coding algorithm for five medium sized Chinese texts about different subjects. It is shown that the new model returns the best entropy value and compression ratio.