Automated selection of rule induction methods based on recursive iteration of resampling methods and multiple statistical testing

Shusaku Tsumoto, Hiroshi Tanaka · 1995

One of the most important problems in rule induction methods is how to estimate which method is the best to use in an applied domain. While some methods are useful in some domains, they aTe not useful in other domains. Therefore it is very dificult to choose one of these methods. FOT this purpose, we introduce mul-tiple testing based on recursive iteration of resampling methods for rule-induction (MULT-RECITE-R). This method consists of four procedures, which includes the inner loop and the outer loop procedures. First, orkg-inal training samples($) are randomly split into new training samples(&) and teat samples(T1) using a Te-sampiing scheme. second, & are again spiii inio training sample(&) and training samples(li) using the same resampling scheme. Rule induction meth-ods ave applied and predefined metrics aTe calculated. This second procedure, as the inner loop, is repeated for 10000 times. Then, third, rule induction methods are applied to 5’1, and the met&s calculated by Tl are cornpaved with those by Tz. If the metrics derived by TZ predicts those by Tl, then we count it as a success. The second and third procedures, as the outeT loop, are iterated foT 10000 times. Finally, fourth, the overall results are interpreted, and the best method is selected if the resampling scheme performs well. In OTdeT to evaluate this system, we apply this MULT-RECITE-R method to three UCI databases. The results show that this method gives the best selection of estimation methods statistically.

Read the paper · More papers on PaperTik