Matching Data Mining Algorithm Suitability to Data Characteristics Using a Self-Organizing Map

Kate Smith‐Miles, Frederick Woo, Vic Ciesielski, Remzi Ibrahim · 2002

The vast range of data mining algorithms available for learning classification problems has encouraged a trial-and-error approach to finding the best model. This problem is exacerbated by the fact that little is known about which techniques are suited to which types of problems. This paper provides some insights into the data characteristics that suit particular data mining algorithms. Our approach consists of four main stages. First, the performance of six leading data mining algorithms is examined across a collection of 57 well-known classification problems from the machine learning literature. Secondly, a collection of statistics that describe each of the 57 problems in terms of data complexity is collated. Thirdly, a self-organising map (SOM) is used to cluster the 57 problems based on these measures of complexity. Each cluster represents a group of classification problems with similar data characteristics. The performance of each data mining algorithm within each cluster is then examined in the final stage to provide both quantitative and qualitative insights into which techniques perform best on certain problem types. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.

Read the paper · More papers on PaperTik