Computing preset dictionaries from text corpora for the compression of messages

Marc W. Abel, Soon Myoung Chung · 2014

Rigid length limits of short messages greatly restrict users' ability to express ideas intelligibly. While data compression can help by enabling greater expressivity in short messages, most work in compression has focused on managing large streams of data instead of small ones. We investigated the potential for preset dictionaries to unleash zlib's ability to compress short messages typical of Short Message Service (SMS) texts, microblog updates, and other single-packet transactions. This paper proposes two preset dictionary generation methods and reports strong test results across two dissimilar text corpora: the Enron database of email messages, and the IEEE VAST Challenge 2011 microblog corpus. For exchanges in English using our proposed methods, it is possible to extend "tweets" from 140 to 197 septets on average, and to extend SMS texts from 160 to 227 septets on average. The preset dictionary's role is as important as zlib's, and each requires the other to obtain these gains.

Read the paper · More papers on PaperTik