A framework for measuring changes in data characteristics

Venkatesh Ganti, Johannes E. Gehrke, Raghu Ramakrishnan · 1999

A data mining algorithm builds a model that captures interesting aspects of the underlying data. We develop a framework for quantifying the difference, called the deviation, between two datasets in terms of the models they induce. Our framework covers a wide variety of models including frequent itemsets, decision tree classifiers, and clusters, and captures standard measures of deviation such as the misclassification rate and the chi-squared metric as special cases. We also show how statistical techniques can be applied to the deviation measure to assess whether the difference between two models is meaningful (i.e., whether the underlying datasets have statistically significant differences in their characteristics), and discuss several practical applications. 1 Introduction The goal of data mining is to discover (predictive) models based on the data maintained in the database [16]. Several algorithms have been proposed for computing novel models [1, 2, 3, 28, 29], for more efficient ...

Read the paper · More papers on PaperTik