Resource planning for SPARQL query execution on data sharing platforms

Stefan Hagedorn, Katja Hose, Kai-Uwe Sattler, Jürgen Umbrich · VBN Forskningsportal (Aalborg Universitet) · 2014

Abstract. To increase performance, data sharing platforms often make use of clusters of nodes where certain tasks can be executed in parallel. Resource plan-ning and especially deciding how many processors should be chosen to exploit parallel processing is complex in such a setup as increasing the number of proces-sors does not always improve runtime due to communication overhead. Instead, there is usually an optimum number of processors for which using more or fewer processors leads to less efficient runtimes. In this paper, we present a cost model based on widely used statistics (VoiD) and show how to compute the optimum number of processors that should be used to evaluate a particular SPARQL query over a particular configuration and RDF dataset. Our first experiments show the general applicability of our approach but also how shortcomings in the used statis-tics limit the potential of optimization. 1

Read the paper · More papers on PaperTik