A Hybrid Approximate XML Subtree Matching Method Using Syntactic Features and Word Semantics

Wenxin Liang, Haruo Yokota · 2009

With the exponential increase in the amount and size of XML data on the Internet, XML subtree match- ing has become important for many application areas such as change detection, keyword retrieval and knowledge discoveries over XML documents. In our previous work, we have proposed leaf-clustering based approximate XML subtree matching methods using syntax information of both the clustered leaf nodes and the corresponding paths. In this paper, we propose a hybrid subtree matching method, in which subtree matching is determined by using the word semantics based on WordNet thesaurus in leaf nodes and the syntactic features in the relevant paths. We also propose a one-pass hash join technique to reduce the additional join cost caused by the extra words expanded by the WordNet. We perform experiments to evaluate performance and matching precision and recall comparing the hybrid method with the original syntax-based methods. The experimental results indicate that the proposed hybrid method with one-pass hash join, comparing with the existing path-based SLAX algorithm, can effectively improve the precision and recall with about only 5% increase of the execution time for the leaf-clustering based

Read the paper · More papers on PaperTik