An Environment for Rule Extraction and Evaluation from Databases

José Augusto Baranauskas, Maria Carolina Monard · 2000

Classi cation for very large databases has many practical applications in Data Mining. Thus, Machine Learning algorithms should be able to operate in massive datasets in order to extract symbolic classi ers. In this context, a symbolic classi er is the one that can be transformed into a set of rules. When a dataset is too big for a particular learning algorithm, there are other ways to make learning feasible such as dataset sampling: for each sample a classi er is extracted for further investigation, like accuracy evaluation. It is also possible to evaluate the performance of combining all extracted classi ers into an ensemble. However, combining symbolic classi ers into a single one by a majority (or other) vote mechanism does not result into a symbolic classi er any more. The approach adopted in this under development work consists in evaluating the induced knowledge as white-box i.e., looking to the rules and trying to combine them using a computational environment. The environment is designed as a test-bed for Data Mining research, as well as a generic knowledge discovery tool for varied database domains. Flexibility is achieved by an open-ended design for extensibility, enabling integration of existing Machine Learning algorithms, support functions for pre-processing as well as new locally developed algorithm and functions.

Read the paper · More papers on PaperTik