Classification of XML Documents

Abdelhamid Bouchachia, Marcus Hassler · 2007

With the explosion of XML-based online documents, the task of knowledge discovery from the Web becomes highly significant. As an appropriate machinery, classification allows to categorize documents to facilitate that task. A classification approach is introduced in this paper. It is based on the k-nearest neighborhood algorithm that relies on an edit distance measure. The originality of the work lies in combining both the content and the structure of XML documents to compute the edit distance. The approach is empirically evaluated using real-world XML collections

Read the paper · More papers on PaperTik