A Study on Similar String Matching

Zhenglu Yang, Masaru Kitsuregawa · 2011

Approximate querying on string collections is an important data analysis tool for many applications, and it has been exhaustively studied. However, the scale of the problem has increased dramatically because of the prevalence of the Web. In this paper, we aim to study the efficient top- k similar string matching problem. Several efficient strategies are introduced, such as length aware and adaptive q-gram selection. We present a general q-gram based framework and study the efficient strategies introduced experimentally on real data sets.

Read the paper · More papers on PaperTik