Procedures of extending the alphabet in combined coding for prediction by partial string matching in text compression

Radu Rădescu, Sever Pasca · 2017

The present work describes the PPM - Prediction by Partial (string) Matching algorithm for lossless compression of text. It studies the procedure of extending the source alphabet for the PPM encoding in order to allow using of symbols not yet present in the source alphabet at the very beginning of the process of text encoding. The work describes the procedure of extending the alphabet and presents the changes needed to be performed in the PPM algorithm so that it can use the extended generated alphabet. There are three ways proposed in extending the source alphabet: manual, running over the text, and adaptive. The arithmetical compression algorithm is used in order to encode the text words with the help of the statistic model derived from the PPM. Experimental results on different formats of text files and useful conclusions for text lossless compression are presented.

Read the paper · More papers on PaperTik