DISTRIBUTED RDF GRAPH KEYWORD SEARCH
DANILO MORET RODRIGUES · 2013
The goal of this dissertation is to improve RDF keyword search.We propose a scalable approach, based on a tensor representation that allows for distributed storage, and thus the use of parallel techniques to speed up the search over large linked data sets, in particular those published as Linked Data.An unprecedented amount of information is becoming available following the principles of Linked Data, forming what is called the Web of Data.This information, typically codified as RDF subject-predicate-object triples, is commonly abstracted as a graph which subjects and objects are nodes, and predicates are edges connecting them.As a consequence of the widespread adoption of search engines on the World Wide Web, users are familiar with keyword search.For RDF graphs, however, extracting a coherent subset of data graphs to enrich search results is a time consuming and expensive task, and it is expected to be executed on-the-fly at user prompt.The dissertation's goal is to handle this problem.A recent proposal has been made to index RDF graphs as a sparse matrix with the pre-computed information necessary for faster retrieval of sub-graphs, and the use of tensor-based queries over the sparse matrix.The tensor approach can leverage modern distributed computing techniques, e.g., nonrelational database sharding and the MapReduce model.In this dissertation, we propose a design and explore the viability of the tensor-based approach to build a distributed datastore and speed up keyword search with a parallel approach.