Measuring similarity of semi-structured documents with context weights

Christopher C. Yang, Nan Liu · 2006

In this work, we study similarity measures for text-centric XML documents based on an extended vector space model, which considers both document content and structure. Experimental results based on a benchmark showed superior performance of the proposed measure over the baseline which ignores structural knowledge of XML documents.

Read the paper · More papers on PaperTik