Rule induction based on frequencies of attribute values

Grzegorz Borowik, Karol J. Kowalski · Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE · 2015

Rule induction is one of the most significant issues in data mining. This is due to the fact that decision rules induced from the training data are used to classify new objects. The classification is based on matching the object with the decision rules. Specifically, the generated rules are used to resolve whether or not the object satisfies the conditions specified by the subset of attributes belonging to a given decision class. Most of the rule induction methods are insufficient for large databases and hence do not support today's Big Data issues. This is mainly due to the use of so-called discernibility matrices during calculations. The purpose of this paper is the idea of the implementation of a new efficient rule induction algorithm that is based on statistics of attribute values and that avoids building the discernibility matrix explicitly. Tests have shown that the implementation is much more efficient than currently available solutions for large data sets.

Read the paper · More papers on PaperTik