Topological Analysis of The SPOKE Graph

Geoffrey Sanders, USDOE National Nuclear Security Administration (NNSA), Roger Pearce, Sergio E. Baranzini · 2020

The SPOKE graph [2, 6] is a sparse decorated semantic graph representing a collection of knowledge collected in many scientific databases from the fields of healthcare, biochemistry, chemistry, biology, et cetera. This knowledge graph is stored as a relational dataset decorated with metadata on each constituent vertex and edge. Formally, the graph is G(V, E, D), where V is a set of n vertices V := {1, ..., n} and edges of the form (i, j) ϵ E for i, j ϵ V, and table D that for any item in V υ E stores unstructured data such as vertex/edge type, nature of a relationship, et cetera. D(i) = {data involving vertex i ϵ V}, and D(i, j) = {data involving edge (i, j) ϵ E}. Here, we treat the graph as undirected in the sense that a direct relationship for (i, j) causes a (possibly opposite) reverse direct relationship for (j, i). The SPOKE graph G(V, E, D) is formed by processing a collection of relational datasets from medicine, chemistry, and biology, connecting many entities. Here, we analyze an instance from 2019, Spoke-20190707, where a graph file contains 6.16M edges and associated metadata and a vertex file contains 2.15M vertices and the associated metadata. There are 12 different types of vertex entities; all edge types used are implicit (see §2). There is other metadata in D on edges and vertices, but we just use the topology and the vertex labels in this report. SPOKE is growing as more knowledge is gained and more datasets are added. SPOKE is likely to grow 10x during the next phase of this project, and we therefore would like to consider topoligical analysis techniques that are scalable to several orders of magnitude larger than the current dataset (say >1B edges).

Read the paper · More papers on PaperTik