Highly efficient universal coding with classifying to subdictionaries for text compression

Yasuhiko Nakano, H. Yahagi, Yoshiyuki Okada, S. Yoshida · 2002

Describes a practical, locally adaptive data compression algorithm of the LZ78 class. According to the Lempel-Ziv incremental parsing rule, the boundary of a string is not related to the statistical history modeled by finite-state sources. The authors have already reported an algorithm classifying to subdictionaries (CSD), which uses multiple subdictionaries and conditions the current string by using the previous one to obtain a higher compression ratio for image compression. They present a practical implementation of this method for any kind of data, and show that CSD was more efficient than LZC when the UNIX facility for compression. The compression performance of CSD was about 10% better than the LZC with the practical dictionary size, an 8K-entry dictionary when the test data were used form Calgary Compression Corpus. Using hashing, the processing speed of the CSD became as fast as the LZC, though the CSD algorithm was more complicated than the LZC.>

Read the paper · More papers on PaperTik