Better Text Compression from Fewer Lexical n-Grams
Tony C. Smith, Michael G. Lorenz · Research Commons (University of Waikato) · 2001
Word-based context models for text compression have the capacity to outperform more simple character-based models, but are generally unattractive because of inherent problems with exponential model growth and corresponding data sparseness. These ill-effects can be mitigated in an adaptive lossless compression scheme by modelling syntactic and semantic lexical dependencies independently.