Aggregation of imprecise and uncertain information for knowledge discovery in databases

Sally I. McClean, Bryan W. Scotney, Mary Shapcott · 1998

We consider the problem of aggregation for uncertain and imprecise data. For such data, we define aggregation operators and use them to provide information on properties and patterns of data attributes. The aggregates that we define use the Kullback-Leibler information divergence between the aggregated probability distribution and the individual tuple data values. We are thus able to provide a probability distribution for the domain values of an attribute or group of attributes using imperfect data. Information stored in a database is often subject to uncertainty and imprecision. An extended relational data model has previously been proposed for such data which allows us to quantify our uncertainty and imprecision about attribute values by representing them as a probability distribution. Our aggregation operators are defined on such a data model. The provision of such operators is a central requirement in furnishing a database with the capability to perform the operations necessary for Knowledge Discovery in Databases. Background Frequently, real life data are uncertain, i.e. we are not certain about the truth of an attribute value, or imprecise, i.e. we are not certain about the specific value of an attribute. Such imprecision might occur naturally as a result of data being provided at different levels of the concept hierarchy or from the integration of distributed databases. It is therefore important that appropriate functionality is provided for generalised database systems to provide intelligent ways of handling such imperfect information. A recent survey of various approaches to handling imperfect information in Data and Knowledge Bases has been provided by Parsons (1996). A database model which is based on partial values and

Read the paper · More papers on PaperTik