Multiple classifier performance
Ariel Duarte-López · RECERCAT (Consorci de Serveis Universitaris de Catalunya) · 2015
This work aims to discover relationships between characteristics of data sets and machine learning algorithms that lead to selecting the best algorithm for classifying the data set.To this end, we have conducted an extensive experimentation including 8 of the main machine learning methods with different configurations and 215 binary data sets, leading to over 7900 experiments.Several state-of-theart metrics have been used in order to describe the data sets in terms of physical structure, statistics, and information theory.A novelty with respect to prior works is that we use a robust approach based on permutation tests that directly assesses whether a given learning algorithm is able to exploit the attributes in a data set to predict class labels.Using the results of these permutation tests and the F-score, a ranking is calculated in order to propose a set of algorithms that should perform well based on the data sets' characteristics.The results obtained are evaluated using the area under the curve in the histograms of the normalized positions proposed for the ranking.Such ranking should help scientists and practitioners in the selection of learning algorithms when facing a new problem, as well as to shed light into the behaviour of algorithms on different types of data sets.