Powerful Database Support for High Performance Data Mining

Christian Böhm · 2001

Larger and larger amounts of data are collected and stored in databases, increasing the need for efficient and effective analysis methods to make use of the information contained implicitly in the data. The extraction of such potentially useful information is called data mining. In the thesis, it is shown that numerous data mining methods such as density based clustering, k-means clustering, outlier detection, or k-nearest neighbor classification can be based on the similarity join as a database primitive. By such a reformulation, the identical result can be achieved at a drastically improved efficiency. The similarity join becomes an important basic operation of advanced database management systems. For a given set of feature vectors, the similarity join determines those object pairs which are similar according to some appropriate similarity measure, in SQL style SELECT * FROM R, S WHERE distance (R.point, S.point) ≤ e. In this thesis, we concentrate on both aspects of the similarity join applications as well as algorithms. For the first aspect we show how typical algorithms of data analysis and data mining can be reformulated such that they are exclusively based on the similarity join. According to several example applications, we demonstrate the enormous performance potential of this database primitive. We also introduce several different kinds of similarity join. The most important variant which corresponds to the SQL statement above is based on the range search as a join

Read the paper · More papers on PaperTik