Improving the Compression Performance of Turkish Text with PoS Tags.

Ebru Çelikel, Bekir Taner Dinçer · IKE · 2004

Information Retrieval (IR) systems have to deal with large amount of text. This, in turn, introduces a storage problem. We suggest that employing compression reduce that need for vast storage. We designed and developed a part-of-speech (PoS) tagger for Turkish to provide higher retrieval rate of relevant text. To occupy less space, we then compressed the tagged texts via word and tag based Prediction by Partial Matching (PPM) algorithm, yielding high performance on texts, as compared to other common compression tools. In case the tags are kept secret, system also provides security to some degree. We measured the algorithm performance on Turkish and compared the results with that of English. Furthermore, we compared the system’s performance with some other compression tools as Arithmetic Coding, Huffman, LZW and BWT.

Read the paper · More papers on PaperTik