Using SQL primitives and parallel DB servers to speed up knowledge discovery in large relational databases

Alex Alves Freitas, Simon Lavington · 1996

Efficiency is crucial in KDD (Knowledge Discovery in Databases), due to the huge amount of data stored in commercial databases. We argue that high efficiency in KDD can be achieved by combining two approaches, namely mapping KDD functionality onto standard DBMS operations and executing KDD tasks on a parallel SQL server. We propose generic KDD primitives which underly the candidate-rule evaluation procedures of many KDD algorithms, and we evaluate the speed up achieved by a parallel SQL server when executing a decision-tree learner algorithm implemented via these primitives. 1 Introduction Our approach to Knowledge Discovery in Databases (KDD) is based on Machine Learning (ML) algorithms. However, sequential versions of most ML algorithms are impractical (i.e. take too long to run) on very large data sets. For instance, Catlett [1991] estimated that sequential C4.5 would take several months to learn from 1,000,000 examples by using state-of-the-art hardware at that time. Provost & Aro...

Read the paper · More papers on PaperTik