On Bayesian Network and Outlier Detection.

Sakshi Babbar, Sanjay Chawla · Conference on Management of Data · 2010

Existing studies on data mining has largely focused on the design of measures and algorithms to identify outliers in large and high dimensional categorical and numeric databases. However, not much stress has been given on the interestingness of the reported outlier. One way to ascertain interestingness and usefulness of the reported outlier is by making use of a domain knowledge. In this paper, we present a new measure to discover outliers based on background knowledge, represented by a Bayesian network. We define outliers as “unlikely events under the current favored theory of the domain”. We introduce two quantitative rules derived from the Bayesian network to uncover outliers. Furthermore, we use these rules to rank the instances based on joint probability distribution in the Bayesian network. In our approach, we not only identified outliers but also explain why they are likely to be so. A critical analysis on distance based technique is also presented to show why there is a mismatch between outliers as entities “which are far away from their neighbors” and “real” outliers as identified using Bayesian Networks.

Read the paper · More papers on PaperTik