Review of Graph Databases for Big Data Dynamic Entity Scoring

Montiago LaBute, USDOE, M Dombroski · 2014

Modeling data as a graph enables users to quickly analyze networked phenomenon, such as social network-based marketing data (e.g., linking entities in social media based on their friends and their "likes", their friend-of-a-friend's "likes", etc.) and scientific data, such as in biology to assess a gene's proximity in a graph to particular disease phenotype. Numerous purpose-built graph databases (DBs) called “NoSQL” or “Not Only SQL” DBs are emerging to support this type of graph analysis. These DBs differ from traditional relational DBs in that they are optimized to improve storage and query performance across large, complex graphs. This document reviews the state of the art in graph DBs and identifies potential DBs that show promise for dynamic entity-scoring algorithms. Dynamic entity scoring is a new area of research that allows analysis of data within graphs using undirected methods to weigh the edges of the graph based on the proximity or ambient influence of near-by nodes. The EntityScore (eScore) methodology developed at Lawrence Livermore National Laboratory (LLNL) is one application of these emerging graph-analysis algorithms to help analysts identify and prioritize entities, based on complex, analyst-defined weighting criteria. Our review examines a wide range of different, well-known graph DBs in use since at least 2010 and identifies unique advantages and disadvantages that may impact eScore-specific requirements. Specifically, we review data structures employed, query features employed, load times, query calculation times, memory usage and several other important attributes reported in the literature. We limited our considerations for eScore to Blueprints-compliant graph DBs (https://github.com/tinkerpop/blueprints/) because they provide a standard, java application programming interface (API) based on the property graph model.

Read the paper · More papers on PaperTik