Publishing, locating, and querying networked information sources

Alberto O. Mendelzon, George A. Mihaila · 2000

Most of the current content of the Web has a textual or multimedia nature, and was created for direct human consumption. However, the original motivation behind the creation of the Web was the exchange of data, in a way that can be used by programs for further processing. And indeed, we are witnessing an increasing interest in using this medium for the dissemination of data in various disciplines, such as the environmental sciences, genetic research, astronomy, statistics, etc. However, the highly distributed, autonomous and heterogeneous nature of data sources makes global data sharing difficult, due to a lack of standards for essential activities such as: the discovery of sources containing data relevant to a problem; the assessment of the relative quality of the discovered sources; and the uniform access to data. We propose an infrastructure for the transparent sharing of structured data on the World Wide Web, called WebSemantics (WS). The system provides an easy way to publish metadata about the location, content and quality of sources in WWW documents. This allows the use of a combination of database and information retrieval techniques to locate relevant sources. We define a declarative query language, WSQL, which facilitates source discovery, selection, and querying. We specify a metadata model for the WS architecture and we specify the formal semantics of the WSQL language, by defining a domain calculus over the metadata model. We introduce safety restrictions for calculus queries that guarantee their computability. We also propose an algebra for query evaluation and show it has the same expressive power as the calculus. We describe a prototype implementation of the WS system and we report on our experience using it. We also investigate some theoretical aspects of integrating data from sources with incomplete, overlapping and possibly inconsistent information. Thus, we describe a method to decide the consistency of a source collection, and we study the semantics of data integration in such an environment.

Read the paper · More papers on PaperTik