Compressing string data in a database by using q-grams

Amir Seyed Danesh, Abdollatif Ahy · 2010

String data is ubiquitous, and its management has a particular importance. Managing string data, especially in databases is really prominent, for we see a large amount of this kind of data in usual databases and this issue persuade us to think about ways for compressing data with a good management. We can do this by using varied kinds of algorithms that nowadays we have. In this paper we present a new algorithm for compressing string data that based on approximate string processing. Commercial databases do not support approximate string queries directly, and it is a challenge to implement this functionality efficiently with user-defined functions. To do this we use small parts of each string that we call them q-gram, and processing them using standard methods available in the DBMS. We can implement this functionality on top of commercial databases by exploiting facilities already available in them.

Read the paper · More papers on PaperTik