A PROGRESSIVE CLUSTERING ALGORITHM TO GROUP THE XML DATA BY STRUCTURAL AND SEMANTIC SIMILARITY

Richi Nayak, Tien T. T. Tran · International Journal of Pattern Recognition and Artificial Intelligence · 2007

Since the emergence in the popularity of XML for data representation and exchange over the Web, the distribution of XML documents has rapidly increased. It has become a challenge for researchers to turn these documents into a more useful information utility. In this paper, we introduce a novel clustering algorithm PCXSS that keeps the heterogeneous XML documents into various groups according to their similar structural and semantic representations. We develop a global criterion function CPSim that progressively measures the similarity between a XML document and existing clusters, ignoring the need to compute the similarity between two individual documents. The experimental analysis shows the method to be fast and accurate.

Read the paper · More papers on PaperTik