Applying predictive clustering trees to the inductive logic programming 2005 challenge data

Jan Struyf, Celine Vens, Tom Croonenborghs, Sašo Džeroski, Hendrik Blockeel · Lirias · 2005

This paper presents our submission to the Inductive Logic Programming 2005 Challenge. The goal of this challenge is to build accurate models predicting gene function in the genome of the yeast Saccharomyces cerevisiae based on relational homology and secondary structure data. The approach that we follow consists of two steps. In the first step, the data is propositionalized by constructing features with the relational frequent pattern mining system Warmr (relational features including aggregates for the homology data, subsequences for the secondary structure data). In the second step, the propositional data is used to build the required models. This is non-trivial as each gene is labeled with a set of functional classes instead of just one class and the classes are organized in a class hierarchy (hierarchical multi-classification). A simple approach to solve a multi-classification problem is to build a model for each individual class. In this work, we use a di#erent method. We use Clus, a system for building clustering trees, in combination with a distance metric designed for hierarchical multi-classification. This yields a single clustering tree predicting a reasonable number of the classes. Our model obtains an average precision of 68.9% and covers 51.1% of the examples of an independent test set.

Read the paper · More papers on PaperTik