On propositionalization for knowledge discovery in relational databases

Mark-André Krogel · Digitalen Hochschulbibliothek Sachsen-Anhalt (Universitäts- und Landesbibliothek Sachsen-Anhalt) · 2005

Propositionalization is a process that leads from relational data and background knowledge to a single-table representation thereof, which serves as the input to widespread systems for knowledge discovery in databases.Systems for propositionalization thus support the analyst during the usually costly phase of data preparation for data mining.Such systems have been applied for more than 15 years, often competitive compared to other approaches to relational learning.However, the broad range of approaches to propositionalization suffered from a number of disadvantages.First, the single approaches were not described in a unified way, which made it difficult for analysts to judge them.Second, the traditional approaches were largely restricted to produce Boolean features as data mining input.This restriction was one of the sources for information loss during propositionalization, which may derogate the quality of learning results.Third, methods for propositionalization often did not scale well.In this thesis, we present a formal framework that allows for a unified description of approaches to propositionalization. Within our framework, we systematically enhance existing approaches with techniques well-known in the area of relational databases.With the application of aggregate functions during propositionalization, we achieve results that preserve more of the information contained in the original representations of learning examples and background knowledge.Further, we suggest special database schema transformations to ensure high efficiency of the whole process.We put special emphasis on empirical investigations into the spectrum of approaches.Here, we use data sets and learning tasks with different characteristics for our experiments.Some of the learning problems are benchmarks from machine learning that have been in use for more than 20 years, others are based on more recent real-life data, which were made available for competitions in the field of knowledge discovery in databases.Data set sizes vary across different orders of magnitude, up to several million data points.Also, the domains are diverse, ranging from biological data sets to financial ones.This way, we demonstrate the broad applicability of propositionalization.Our theoretical and empirical results are promising for other applications as well, in favor of propositionalization for knowledge discovery in relational databases.iii This thesis would not have been possible without all the help that I received from many people.First of all, Stefan Wrobel was a supervisor with superior qualities.His kind and patient advice made me feel able to climb the mountain.He even saw good aspects when I made mistakes, and I repeatedly did so.I will always be very grateful for his support, and I take his positive attitude as a model for myself.Then, there were so many teachers, colleagues and students of influence in my years at Magdeburg University and Edinburgh University, that I cannot name them all.Thank you so much!I am also grateful to the friendly people of Friedrich-Naumann-Stiftung, who generously supported my early steps towards the doctorate with a scholarship and much more.Last not least, my family was a source of constant motivation.So I dedicate this thesis to my children,

Read the paper · More papers on PaperTik