Data Mining as Selective Theory Extraction in Probabilistic Logic

Manfred Jaeger, Heikki Mannila, Emil Weydert · 1996

ing from this specific example we are led to the following formulation of data mining problems in general: given a database r, a language L for expressing statements about the data, and a criterion for distinguishing interesting statements about the data from uninteresting ones, we want to find all the interesting statements true for r. Formally, this means finding the set ThI(L; r) = Th(L; r) " I(L; r); where Th(L; r) = f' 2 L j r j= 'g is the set of sentences of L true in r, and I(L; r) ` L is the set of interesting formulas of L about r. This theory extraction formulation is either explicitly or implicitly used in a variety of data mining studies [MR86, DB93, Klo95, AMS + 96]. Its roots are in the use of diagrams of models in model theory. 1 It is based on the view that data mining is principally descriptive: the task is to obtain a collection of statements about the data. In the theory extraction formulation, there are two ways of delimiting the sentences that form the de...

Read the paper · More papers on PaperTik