An Improved XML Document Clustering Using Path Feature

Jinsha Yuan, Xin-ye Li, Lina Ma · 2008

Extensible markup language (XML) documents clustering is useful to XML application such as XML search engine. The element tags and their position in the document's hierarchy provide valuable information to clustering XML documents. XML path can represent both element tags and their position information. Since common Xpath represents only parts of the XML structural, using common Xpath as XML structural representation is not always efficient to XML clustering, especially when those documents are with dissimilar structure. In this paper, we use all paths less than or equal to length L as feature vectors for XML documents. Since the feature vector matrix is usually sparse, we use bipartite graph to express association relation among XML documents and path features. Based on this idea, we improved the path-based XML clustering algorithm. Experiments are described to demonstrate its efficiency.

Read the paper · More papers on PaperTik