A knowledge discovery methodology for the performance evaluation of scientific software

Vassilios S. Verykios, Elias N. Houstis, John R. Rice · 2000

In this paper we define a knowledge discovery in databases (KDD) methodology to automatically generate metadata (i.e., knowledge rules) from software/machine pair performance databases. This metadata can be used to characterize the computational behavior of various classes of software or machines. The core and the most computationally intensive part of the KDD methodology is the data mining phase which identifies "interesting" patterns from the performance data. The discovery patterns are expressed in a high level representation to be used to summarize and predict the computational behavior of the targeted software/machine. This paper presents an implementation and evaluation of the proposed KDD process for a class of scientific software together with three data mining algorithms (ID3, HOODG, and CN2). For this case study we have selected a set of software that implements the "mesh/grid partitioning " phase of the domain decomposition approach used for the parallel processing of partial differential equation (PDEs) computations. The raw performance database is generated from a population of elliptic PDEs and PELLPACK [HRW 98] solvers by varying the PDE domain, mesh, and domain partitioning (DP) algorithm. The goal of the KDD process here is to evaluate the performance of PELLPACK, CHACO [HL95c], METIS [KK95c], and PARTY [PD96] algorithms/software. This case study shows that (a) the three data mining algorithms used are qualitatively and quantitatively equally effective, (b) the knowledge discovered for the DP algorithms by this KDD process is quantitatively similar to that deduced by purely experimental observations [VH97], and (c) the KDD process is not limited by the size of the performance data and its dimensionality.

Read the paper · More papers on PaperTik