Incorporating domain knowledge into attribute-oriented data mining
Sally I. McClean, Bryan Scotney, Mary Shapcott · International Journal of Intelligent Systems · 2000
It is frequently the case that data mining is carried out in an environment which contains noisy and missing data. This is particularly likely to be true when the data were originally collected for different purposes, as is commonly the case in data warehousing. In this paper we discuss the use of domain knowledge, e.g., integrity constraints or a concept hierarchy, to re-engineer the database and allocate sets to which missing or unacceptable outlying data may belong. Attribute-oriented knowledge discovery has proved to be a powerful approach for mining multi-level data in large databases. Such methods are set-oriented in that attribute values are considered to belong to subsets of the domain. These subsets may be provided directly by the database or derived from a knowledge base using inductive logic programming to re-engineer the database. In this paper we develop an algorithm which allows us to aggregate imprecise data and use it for multi-level rule induction and knowledge discovery. ©2000 John Wiley & Sons, Inc.