Speeding up Hypothesis Development

Jörg Schlösser, Peter C. Lockemann, Matthias Gimbel · Kluwer Academic Publishers eBooks · 2005

Knowledge discovery in databases (KDD) is a long-lasting, highly interactive and iterative process, in which the human data analyst searches for the right knowledge by repeatedly querying the database, using so-called derivation streams. Each stream consists of several data exploration steps to select, filter, transform, etc. the input data for the final analysis step (e.g., a data mining algorithm). After executing a stream the data analyst redevelops or adjusts his hypothesis by interpreting the results. This should give him new insights as to how to proceed, or how to backtrack and explore alternatives to earlier derivations. The central premises of this chapter are that current KDD systems do too little in the way of supporting the hypothesis development, that the development — or exploration — phase is closely intertwined with data mining, and that the interactive and spontaneous nature of hypothesis development in the presence of large databases requires much better performance than wha t is common today. The central idea is to exploit the KDD process history to overlay the general optimized query processing capabilities of an underlying relational database system (RDBMS) with a more KDD-process-oriented optimization strategy, which takes the frequent backtracking in a KDD process into account. Our approach is three-fold: First, we propose an information model for documenting the entire derivation history. On its basis we develop subsumption mechanisms for process and result matching, and introduce suitable access and retrieval facilities. Second, based on this model, we develop a strategy and technique for automatically reusing earlier results for the execution of subsequent streams whenever possible in order to avoid expensive recomputations. Since reuse requires prior materialization of intermediate results in persistent store, we introduce a strategy for automatic decisions on materialization and de-materialization. Finally, we translate data preparation steps into corresponding database queries that are transferred to the RDBMS for further optimization and execution.

Read the paper · More papers on PaperTik