XCleaner: A New Method for Clustering XML Documents by Structure
Dariusz Brzeziński, Anna Leśniewska, Tadeusz Morzy, Maciej Piernik · Control and Cybernetics · 2011
With the vastly growing data resources on the In- ternet, XML is one of the most important standards for document management. Not only does it provide enhancements to document exchange and storage, but it is also helpful in a variety of informa- tion retrieval tasks. Document clustering is one of the most inter- esting research areas that utilize XML's semi-structural nature. In this paper, we put forward a new XML clustering algorithm that relies solely on document structure. We propose the use of maximal frequent subtrees and an operator called Satisfy/Violate to divide documents into groups. The algorithm is experimentally evaluated on real and synthetic data sets with promising results.