XAdap: An Adaptive Huffman Coding on Markup Languages

Kalyan Cherukuri, Suneeta Agarwal · 2007

XML documents are used for data exchange and to store large amount of data over the Web. These documents are extremely verbose and require specific compression for efficient transformation. In this paper we are analyzing various existing compressors and propose a new method called XAdap, which uses adaptive Huffman coding. It is based on the principle of extracting data from the document, and grouping it based on semantics. The document is encoded as a sequence of integers, while the data grouping is based on XML tags/attributes/comments. The re-organized data is now compressed by adaptive Huffman coding. We compare the proposed method with other existing (which uses Huffman coding) tools. Performance evaluation shows that XAdap outperforms previously proposed XML specific compression tools.

Read the paper · More papers on PaperTik