Clustering XML Search Results Based on Content and Structure Similarity
Zhong Minjuan, Wan Chang-Xuan, Liu De-Xi, Xian-Pei Jiao · 2011
Clustering XML search results is an effective way to improve performance. However, the key problem is how to measure similarity between XML documents. In this paper, we propose a semantic similarity measure method combining content with structure, in which a variety of XML document features, including term element frequency, term inverse element frequency, semantic weight of tag label and level information of the term, are analyzed and applied for computing the similarity between XML documents. In addition, two new performance evaluation methodology, namely ClusterRatio_Relevant and DocuRatio_Relevant, for clustering quality are introduced motivated by the observations of relevant documents distribution and the fact that collection has no classification information. Experiment results show that proposed similarity method(CAS measure)outperforms traditional document clustering(CO measure) in ClusterRatio_Relevant and DocuRatio_Relevant and produces better clustering quality.