Reducing redundancy in characteristic rule discovery by using IP-techniques
Tom Brijs, Koen Vanhoof, Geert Wets · 1999
The discovery of characteristic rules is a well-known data mining technique and has lead to several successful applications. Unfortunately, typically a (very) large number of rules is discovered during the mining stage. This makes monitoring and control of these rules extremely costly and difficult. Therefore, a selection of the most promising rules is desirable. In this paper, we propose an integer programming model to solve the problem of selecting the most promising subset of characteristic rules. The proposed technique allows to control a user-defined level of overall quality of the model in combination with a maximum reduction of the redundancy extant in the original ruleset. We use real-world data to evaluate the performance of the proposed technique against the wellknown RuleCover heuristic. 1 Introduction Data mining is the automated search for hidden, previously unknown and potentially useful information from large databases. Moreover, data mining is a crucial pha...