Good answers from bad data : a data management strategy

Henry B. Kon, Stuart Madnick, Michael D. Siegel · 1995

Data error is an obstacle to effective application, integration, and sharing of data. Error reduction, although desirable, is not always necessary or feasible. Error measurement is a natural alternative. In this paper, we outline an error propagation calculus which models the propagation of an error representation through queries. A closed set of three error types is defined: attribute value inaccuracy (and nulls), object mismembership in a class, and class incompleteness. Error measures are probability distributions over these error types. Given measures of error in query inputs, the calculus both computes and "explains " error in query outputs, so that users and administrators better understand data error. Error propagation is non-trivial as error may be amplified or diminished through query partitions and aggregations. As a theoretical foundation, this work suggests managing error in practice by instituting measurement of persistent tables and extending database output to include a quantitative error term, akin to the confidence interval of a statistical estimate. Two theorems assert the completeness of our error representation.

Read the paper · More papers on PaperTik