An Efficient Technique for Representation and Compression of Bengali Text

Md Farhad Mokter, Sumya Akter, Md Palash Uddin, Masud Ibn Afjal, Md. Al Mamun, Md. Abu Marjan · 2018

Text representation and compression of natural languages have become one of the challenging research aspects in recent times. It bears more significance for Bengali language as it exposes more complicated structures. Some works have been done on Bengali text compression. However, they may not produce notable compression performance in case of the presence of huge amount of conjugate characters. In this paper, we present an efficient technique for the representation and compression of Bengali document to obtain better compression gain in a computationally inexpensive manner. In the proposed approach, each Bengali single character is represented by a unique 2-digit decimal value whereas a conjugate character is represented by a 4-digit unique decimal value. The decimal value of a word is formed using the decimal values of its constituent characters. Then, indexing and sorting all the word values, a successive subtraction operation is accomplished on the sorted word values to reduce the weight of the numbers. The newly produced decimal values of the words can now be encoded with relatively few bits for the efficient storage or transmission. The experimental result shows that the proposed technique provides a better average improvement on compression ratio using 5 different Bengali datasets than that of the various existing compression schemes such as WinZip (30.74%), Win-RAR (19.75%) and 7-Zip (16.56%).

Read the paper · More papers on PaperTik