What can XML do for textual information retrieval on the WEB
Emmanuel Nauer, Rim Al Hulou, Amedeo Napoli · 2000
In this paper, we propose an analysis of the links existing between (1) semistructured data or ssd, i.e. rough and heterogeneous data without xed structure (2) xml, the language of description of textual documents (3) object-based representation systems or obrs. The manipulation of ssd comes from several needs such as data integration, information retrieval on the Web and text mining. Due to its characteristics, xml may be used for describing ssd and can be considered as a bridge between ssd and obrs, obrs being mainly used for their reasoning capabilities. We conclude the paper with a discussion enlightening the bene ts of using xml and obrs for manipulating ssd and solving problems involving ssd.