Discovering interesting information in XML data with association rules
Daniele Braga, Alessandro Siro Campi, Stefano Ceri, Mika Klemettinen, Pier Luca Lanzi · 2003
Data mining algorithms are designed to extract interesting information from large amounts of data. They usually assume that source data are in relational (tabular) from. However, the recent success of XML as a standard to represent semi-structured data and the increasing amount of data available in XML pose new challenges to the data mining community. In this paper we introduce association rules for XML data. To accomplish this, we propose a new operator, based on XPath and inspired by the syntax of XQuery, which allows us to express complex mining tasks, compactly and intuitively. The operator can indifferently (and simultaneously) target both the content and the structure of the data, since the distinction in XML is slight.