A Study of Data Reduction Using Multiset Decision Tables
Uday Seelam, Chien-Chung Chan · 2007 IEEE International Conference on Granular Computing (GRC 2007) · 2007
In rough set theory, observations of objects in a domain of interest are stored in a decision table where each row denoting one object. Objects with same description are duplicated. Duplications may be reduced by using information multisystems, which can be further transformed into Multiset Decision Tables (MDT). In this paper, we have demonstrated the efficacy of MDT when dealing with very large data sets. Experimental results based on the well-known Intrusion Detection System (IDS) data set show that the size of MDT is only 1/3 of the original decision table when all features are used. It could be further reduced to 1/7 when a set of 7 features is used. We also showed that the running time of generating an MDT is faster than generating a C4.5-like decision tree based on the MS SQL server 2000.