Scalable in-memory RDFS closure on billions of triples.

Eric L. Goodman, David Mizell · 2010

Abstract. We present an RDFS closure algorithm, specifically designed and implemented on the Cray XMT supercomputer, that obtains infer-ence rates of 13 million inferences per second on the largest system con-figuration we used. The Cray XMT, with its large global memory (4TB for our experiments), permits the construction of a conceptually straight-forward algorithm, fundamentally a series of operations on a shared hash table. Each thread is given a partition of triple data to process, a dedi-cated copy of the ontology to apply to the data, and a reference to the hash table into which it inserts inferred triples. The global nature of the hash table allows the algorithm to avoid a common obstacle for dis-tributed memory machines: the creation of duplicate triples. On LUBM data sets ranging between 1.3 billion and 5.3 billion triples, we obtain nearly linear speedup except for two portions: file I/O, which can be ameliorated with the additional service nodes, and data structure ini-tialization, which requires nearly constant time for runs involving 32 processors or more.

Read the paper · More papers on PaperTik