Selection of Significant Rules in Classification Association Rule Mining
Yanbo J. Wang, Qin Xin, Frans Coenen · 2005
Abstract—Classification Rule Mining (CRM) is a Data Mining technique for the extraction of hidden Classification Rules (CRs) from a given database, the objective being to build a classifier to classify “unseen ” data. One recent approach to CRM is to use Association Rule Mining (ARM) techniques to identify the desired CRs, i.e. Classification Association Rule Mining (CARM). Although the advantages of accuracy and efficiency offered by CARM have been established in many papers, one major drawback is the large number of Classification Association Rules (CARs) that may be generated (up to a maximum of 2 n-1, where n represents the number of attributes in a database). However, there are only a limited number (k) of CARs that are required to distinguish between classes. The problem addressed in this paper is how to efficiently select all k such CARs. An algorithm is presented, that addresses the above, that operates in binomial time O(k 2 n 2); as opposed to exponential time O(2 n) – the time required to find all k CARs in a “one-by-one ” manner. Index Terms—classification association rule, database, data mining, selectors.