Redescription Mining Over non-Binary Data Sets Using Decision Trees
Tetiana I. Zinchenko, Pauli Miettinen, Gerhard Weikum · MPG.PuRe (Max Planck Society) · 2014
Scientific data mining is aimed to extract useful information from huge data sets with the help of computational efforts. Recently, scientists encounter an overload of data which describe domain entities from different sides. Many of them provide alternative means to organize information. And every alternative data set offers a different perspective onto the studied problem. Redescription mining is tool with a goal of finding various descriptions of the same objects, i.e. giving information on entity from different perspectives. It is a tool for knowledge discovery which helps uniformly reason across data of diverse origin and integrates numerous forms of characterizing data sets. Redescription mining has important applications. Mainly, redescriptions are useful in biology (e.g. to find bio niches for species), bioinformatics (e.g. dependencies in genes can assist in analysis of diseases) and sociology (e.g. exploration of statistical and political data), etc. We initiate redescription mining with data set consisting of 2 arrays with Boolean and/or real-valued attributes. In redescription mining we are looking for such queries which would describe nearly the same objects from both given arrays. Among all redescription mining algorithms there exist approaches which exploits alternating decision tree induction. Only Boolean variables were involved there so far. In this thesis we extend these approaches to non-Boolean data and adopt two methods which allow redescription mining over non-binary data sets.