Semantic Lossy Compression of XML Data.
Mario Cannataro, Gianluca Carelli, Andrea Pugliese, Domenico Saccà · 2001
In the last years a large amount of semistructured data [1, 10] has been managed and exchanged. The largest repository of semistructured data is the World Wide Web, which can be thought of as an enormous database in which data is highly heterogeneous and freely correlated. In this scenario is placed Extensible Markup Language (XML) [14], a language for semistructured data standardised by the World Wide Web Consortium (W3C ), which is candidate to become shortly the de facto standard for web documents. XML allows building machine-readable documents that are naturally convertible in visualisation formats; this is obtained by means of a complete separation among structure, content and style of documents. It is likely that the amount of data available in XML will grow substantially, e.g. in those applications in which the generation of XML documents is performed automatically from data maintained in DBMS. The increasing amount of XML data will lead to the origin of new issues regarding efficiency in the representation of documents. An emerging problem is how to compress the description of an XML document. An interesting solution is XMill [7] that is a lossless (i.e., the original data are eventually restored) compressor/decompressor for XML data, to be used in data exchange and archiving. XMill applies classical entropy-based compression techniques after the execution of ad-hoc compression rules that are driven by the XML structure of the document and by the semantics of data. Its main ideas are (i) separating structure from data, (ii) grouping related data items into homogeneous classes, (iii) applying semantic compressors to those data classes and (iv) applying general-purpose compressors. XMill is a very effective compression tool which overpasses classical generalpurpose compressors such as gzip [3]. The problem with XMill is that, as for classical text compressors, the compression is only used for archiving or exchanging but not for deriving a “synthetic” yet meaningful view of a document as it happens in the compression of images (JPEG) or video sequences (MPEG). Our belief is that lossy compression will become relevant in next applications on Internet. The typical scenario we are envisioning is a multi-channel access to XML documents whereby a document may be required to be displayed e.g. on a small-sized screen using a low bandwidth network. In this case the admissible compression rate