Precision And Comprehensibility In The Integration Of Regression Rules
J. B. Pugliesi, Solange Oliveira Rezende · WIT transactions on information and communication technologies · 2003
Data Mining is the process of data analysis and application of algorithms that are able to find a particular pattern relationship from a large amount of data. Regression is an important problem in data analysis and appears in many real world applications. Therefore, there is an increasing interest in the use of the Data Mining to extract patterns from regression problems. The main objective of this paper is to support the users of Data Mining process in the evaluation of integrated knowledge expressed in the form of regression rules. To explore the knowledge evaluation, experiments are executed with Cubist, RT and M5 algorithms. The 10-fold crossvalidation is used to find the mean error. All regression rules are transformed to a standard syntax and put together in a unique set of rules. The worst rules are eliminated from this set to generate a more concise regressor. This final regressor has less comprehensibility and higher precision than a single regressor induced from all data. A method is defined to determine how to predict the new examples with this final regressor. From the data analysis, it is possible to take some conclusions, which show that usually it is possible to eliminate the worst regression rules and have an accurate regressor with a good level of comprehensibility.