Multirelational Association Rule Mining

Anton Flank · 2004

Data mining is a broad term used to describe various methods for discovering patterns in data. A kind of pattern often considered is association rules, probabilistic rules stating that objects satisfying description A also satisfy description B with certain support and confidence. Association rule mining has previously been restricted to data in a single table, due to the computational complexity of multirelational algorithms. Recently, many methods for mining relational data in multiple tables have been proposed. This thesis explores some of these concepts, aimed at mining multirelational association rules from a standard relational database. Mining is here formulated as a search problem in the space of database queries and L, a subset of tuple calculus that is decidable for satisfiability and syntactic query difference, is used as query language. The search space is ordered based on subsumption by arranging queries in a refinement graph which is built top down. New queries are formed by selecting a node in the graph by a uniform cost strategy and applying a refinement operator, making the query more specific by adding conditions. Support and confidence threshold parameters are used to sort out uninteresting rules. Rules are presented as plain English (or Swedish) sentences, and the user has the option of guiding the search by expressing a positive or negative opinion about discovered rules.

Read the paper · More papers on PaperTik