Text Compression: Syllables

Jan Lánský, Michal Žemlička · 2005

There are two basic types of text compression by symbols -- in the first case symbols are represented by characters, in the second case by whole words. The first case is useful for very short files, the second case for very long files or large collections. We supposed that there exist yet another way where symbols are represented by units shorter than words -- syllables. This paper is focused to specification of syllables, methods for decomposition of words into syllables, and using syllable-based compression in combination of principles of LZW and Hu#man coding. Above mentioned syllable-based methods are compared with their counterpart variants for characters and whole words.

Read the paper · More papers on PaperTik