Knowledge-rich induction of classification rules
Brent Martin · 1996
Introduction The purpose of this research was to produce a machine learning system that can take advantage of many forms of background knowledge to guide the induction of classification rules. This system will be used for knowledge discovery in databases (also known as "database mining"), as the use of background knowledge can considerably reduce the search space of a database knowledge search. A new system, MARVIN++, is introduced, that attempts to satisfy this aim. The following sections describe some of the problems encountered while database mining and suggest forms of background knowledge that can help to overcome them. Irrelevant data For most machine learning algorithms that produce classification rules based on attribute or attribute-value testing the size of the search space will be proportional to (among other things) the number of attributes available, and the number of examples in the data set. Any irrelevant attributes or examples provided will the