Improving PPM Algorithm Using Dictionaries

Yichuan Hu, Jianzhong Zhang, Farooq Khan, Ying Li · 2011

We propose a method to improve traditional character-based PPM text compression algorithm [1] for natural languages. Consider a text file as a sequence of alternating words and non-words, the basic idea of our algorithm is to encode nonwords and prefixes of words using character-based context models and encode suffixes of words using dictionary models. By using dictionary models, the algorithm can encode multiple characters as a whole, and thus enhance the compression efficiency. The advantages of the proposed algorithm are: 1) it does not require any text preprocessing, 2) it does not need any explicit codeword to identify switch between context and dictionary models, 3) it can be applied to any character-based PPM algorithms without incurring much additional computational cost.

Read the paper · More papers on PaperTik