Recognizing and extracting relations and patterns from XML documents

Yangyang Wu · Journal of Tsinghua University(Science and Technology) · 2005

There is a large amount of information that describes interrelation of entities on the web, current search engines lack the ability of manipulating and understanding knowledge and can not discern relations on the web. This paper presents a method for mining relations and patterns from XML documents. The method collected XML documents according to user's requirement, and discerned target XML files which contained relations by calculating similarity between XML documents. The user's searching pattern was established and pattern-matching algorithm was used to extract relation instances from target document. Experimental results show that the similarity calculating method can discern target XML document in a good performance. The pattern-matching algorithm can extract the most target relations from given XML documents accurately.

Read the paper · More papers on PaperTik