Query-based selection and integration of semantic web data sources

Jeff Heflin, Abir Qasem · 2009

The Semantic Web extends the present Web through formal semantics, as specified by ontologies. This formal and unambiguous description of web data allows software agents to exploit web data and provide novel query answering solutions that, in an ideal marriage of knowledge representation and web technologies, could help unlock the true potential of the Web as a highly interconnected global knowledge medium. However, traditionally, work on knowledge representation and ontologies assumed well defined, small and static domains. This thesis, while acknowledging that heterogeneity, scalability and dynamicity are issues inherent in any (Semantic) Web system, explores a practical solution that performs reasonably well. The thesis formally defines a source selection problem for the Semantic Web. Given a Semantic Web query, the source selection framework can be used to identify the data sources that are potentially relevant to a given query. Therefore, during query answering, an algorithm need only use the subset of a knowledge base that is necessary to answer a given query. As the size of a knowledge base significantly impacts the query response time, the source selection framework is more efficient than an approach that considers all available data sources, irrespective of its relevance to a given query. Furthermore, the framework supports a dynamic environment because the source is only accessed right before it is actually used to answer a query, and thus the data is as fresh as possible. Two algorithms that use the source selection framework have been developed. The first algorithm is an adaptation of a well known information integration algorithm. The adaptation provides fast source selection and adds the processing of ontological information present in the data. The second algorithm can work with a more expressive ontology language and can handle scenarios in which identical data elements have been expressed using different identifiers. Experiments conducted using synthetic data have demonstrated that a system based on these algorithms performs well over a wide range of queries. For instance, a typical response time for a substantial work load of more than 100 (possibly mapped) ontologies and 500 data sources is less than 2 seconds.

Read the paper · More papers on PaperTik