Text Pre-processing for Lossless Compression

Lu Batista, Luı́s A. Alexandre · DCC · 2008

Textual data holds a number of properties that can be taken into account in order to improve compression. Pre-processing deals with these properties by applying a number of transformations that make the redundancy "more visible" to the compressor. One of the most commonly used concepts in text pre-processing is called capital conversion. Words with capital letters are converted to their lowercase versions while signaling the change with a flag. This way not only context similarities are increased but also dictionaries used for word replacement only need to contain words in their lowercase versions. Word replacement consists of replacing words with shorter codes which are references to their location in a dictionary.

Read the paper · More papers on PaperTik