Frequent Pattern Mining for XML Query- Answering Support
Alfiya Iqbal, Ahmed Shaikh, Sanchika Bajpai · 2014
Abstract- Extracting information from semi structured documents is difficult task. It is more crucial as there is a huge amount of digital information on the Internet is growing rapidly. Sometimes, documents are often so large that the data set returned as answer to a query may be large to even convey interpretable knowledge. This paper describes an approach which takes RSS feeds as input for which Tree-Based Association Rules (TARs): mined rules are used. It provides more approximate and intentional information on both the structure and the contents of Extensible Markup Language (XML) documents which can then be stored in XML format as well. This generated mined knowledge is later used to provide: 1) The gist of the structure and the content of the XML document and 2) Quick and more approximate answers to queries. This paper focuses on the second feature. In this paper we show a novel approach for finding frequent patterns in XML documents. Keywords:- XML, document, RSS, TARs, approximate. I.