Effective Modelling Of chinese Text Compression

Mu‐King Tsay, Chia-Hsu Kuo, Rong‐Hauh Ju · International Journal of Modelling and Simulation · 1994

Numerous Chinese characters (about 13,051 words) perform a more uniform frequency distribution in a Chinese text than the distribution of letters in an English text. It is, therefore, difficult and improper to apply the alphanumeric compression methods of English directly to Chinese texts. In this paper, methods of finding an optimal modelling technique for Chinese text from models based on various encoded symbols for high compression efficiency are proposed. The source symbols are divided into four kinds of data symbols: subnibble, nibble, original symbol, and a joint symbol. The stationary model and the Markov model are applied to compress source information using these four different source symbols. The performance of different models, including five hybrid models, is evaluated and compared to the predictive coding method.

Read the paper · More papers on PaperTik